Research
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
Overview Research area: Natural Language Processing for mental health safety — specifically, automatic detection of crisis and reportable safety risks in user-generated text. Technical level: Intermed

- arXiv
- 2510.23845
- Published
- 2025-10-27
- Authors
- Grace Byun, Rebecca Lipschutz, Sean T. Minton, Abigail Lott, Jinho D. Choi
AI summary
Overview
Research area: Natural Language Processing for mental health safety — specifically, automatic detection of crisis and reportable safety risks in user-generated text.
Technical level: Intermediate. The methods (prompted LLM inference, majority-voting ensembles, supervised fine-tuning) are standard NLP practice, but interpreting the task requires familiarity with clinical annotation concepts such as the Columbia Suicide Severity Rating Scale (C-SSRS) and multi-label, temporally tagged classification.
One-sentence scope: The paper introduces CRADLE BENCH, a clinician-annotated, multi-label benchmark covering seven crisis categories with temporal (ongoing vs. past) labels, evaluates 15 large language models on it, and releases ensemble-labeled training data plus fine-tuned crisis detection models.
What This Paper Is About
Large language models are increasingly used for personal advice and psychological support, but they must be able to reliably flag high-risk situations such as suicide risk, rape, domestic violence, and child abuse. Existing benchmarks cover only a limited set of crisis types and do not capture whether a crisis is currently happening or happened in the past. This paper builds a clinician-annotated benchmark for multi-faceted crisis detection, uses it to evaluate current models, and shows how an ensemble of strong models can automatically label training data to produce specialized crisis detection models.
Key Contributions
-
CRADLE BENCH benchmark. A clinician-annotated benchmark for multi-faceted mental health crisis detection covering seven categories — suicide ideation (active vs. passive), self-harm, domestic violence, rape, sexual harassment, and child abuse/endangerment — plus a
no crisislabel. It is described as the first such benchmark to include temporal labels distinguishing ongoing from past crises. It provides 600 evaluation (test) examples and 420 development examples. -
Large-scale model evaluation. The authors evaluate 15 state-of-the-art LLMs on the benchmark, including the Llama, Gemma, Qwen, and gpt-oss families, as well as Gemini-2.5-Pro, Claude-4-Sonnet, and GPT-5, and analyze their capabilities and limitations across crisis categories.
-
Ensemble-labeled training data. A training corpus of 4,287 examples is automatically annotated using a majority-vote ensemble of three models (GPT-5, Claude-4-Sonnet, Gemini-2.5-Pro). The paper reports that this ensemble significantly outperforms single-model annotation.
-
Released fine-tuned crisis detection models. Six fine-tuned model variations are built from three base models (Qwen3-14B, Llama-3.3-70B-Instruct, Qwen2.5-72B-Instruct), each trained on either the Consensus subset (at least two models agree) or the Unanimous subset (all three agree), offering complementary detectors under different agreement criteria. Gains reach up to 5.67 percentage points.
Main Findings
-
Closed-source models lead open-source models. Claude-4-Sonnet achieves the highest Exact Match (0.8217), Gemini-2.5-Pro the strongest recall (0.8844 micro, 0.9057 macro), and GPT-5 the highest Macro F1 (0.8163).
-
Best open-source model. gpt-oss-120b attains the highest open-source Exact Match of 0.7650 and a Jaccard of 0.7996. Llama-3.3-70B-Instruct reaches a Jaccard of 0.7806 and Micro F1 of 0.7773.
-
Ensemble majority voting beats every individual model. Voting across GPT-5, Claude-4-Sonnet, and Gemini-2.5-Pro yields Exact Match 0.8450, Jaccard 0.8794, Micro F1 0.8755, Macro F1 0.8438, Micro Recall 0.9030, and Macro Recall 0.9155.
-
Model size usually, but not always, helps. gemma-3-12B-it outperforms gemma-3-27B-it on several metrics, which the authors attribute to increased false positives from more liberal label predictions.
-
Ensemble agreement statistics. On the test set, all three models produced identical predictions for 68.5%, partial agreement occurred in 27.17%, and complete disagreement in 4.33%. On the training set, the figures are 71.3% full agreement (3,058 instances), 26.2% partial agreement (1,123 instances), and 2.5% complete disagreement (106 of 4,287 instances, excluded).
-
Fine-tuning helps, with a precision–recall trade-off. Gains reach up to 5.67 percentage points in Exact Match, 3.92 in Jaccard, and 3.85 in Micro F1, but some models show a decrease in recall.
-
Best fine-tuned model. Llama-3.3-70B (Consensus) achieves the best overall fine-tuned performance, with Micro F1 81.58, Micro Recall 81.64, and Jaccard Index 81.98 — the only open-source model to exceed 80 on these key metrics.
-
Training-subset effects differ by model size. Qwen3-14B performs better when fine-tuned on the smaller Unanimous subset, whereas Llama-3.3-70B and Qwen2.5-72B-Instruct perform slightly better on the larger Consensus subset.
-
Three recurring error patterns. (1) Over-extension of
child abuselabels — GPT-5 shows precision 0.538 vs. recall 0.778 for child abuse/endangerment ongoing, and Llama-3.3-70B 0.375 vs. 0.667 — because models treat any sexual incident involving minors as child abuse. (2) Confusion between active and passive suicide ideation. (3) Failure to follow the clinical single-label severity principle, e.g. tagging both rape and sexual harassment for the same incident instead of only the more severe label. -
Annotation reliability. Inter-annotator Jaccard was 0.7625 in Round 1 (60 questions, 6 labels), 0.7417 in Round 2 (60 questions, 7 labels), and 0.7583 in Round 3 (60 questions, 15 labels). Gwet's AC1 reached its highest mean in Round 3 (0.960).
-
Adjudication of the test set. 131 of the 600 test instances were flagged because original human labels differed from GPT and Claude outputs, and were reviewed by a board-certified psychologist.
Methodology in Plain English
The researchers collected Reddit posts from crisis-related subreddits (r/rape, r/SexualHarassment, r/domesticviolence, r/SuicideWatch, r/selfharm) and broader distress communities (r/mentalhealth, r/depression, r/lonely). A team of four mental health professionals — two licensed psychologists, a PhD clinical postdoctoral resident, and a licensed clinical social worker — annotated the data across ten iterative rounds. The first three rounds used double annotation to measure agreement and refine the guidelines; after that, single annotation was used by four trained experts, with a final quality-control review by a senior faculty psychologist.
Annotation labels seven crisis categories, allows multiple labels per post, and attaches a temporal marker (ongoing or past) to each crisis, with ongoing taking precedence when both could apply. Annotators labeled only the poster's own experiences and avoided labeling unless clear indicators were present.
For evaluation, the authors ran 15 models with temperature 0, top-p 1.0, a maximum context length of 4,096 tokens and a maximum generation length of 1,024 tokens, using RTX A6000 GPUs (H200 GPUs for models over 70B parameters) and APIs for closed-source models. They then combined the three strongest models by majority voting and used that same scheme to automatically label a 4,287-instance training corpus, discarding the 106 cases with full disagreement. Finally, they fine-tuned Qwen3-14B, Llama-3.3-70B-Instruct, and Qwen2.5-72B-Instruct on both the Consensus (4,181 instances, 4,649 labels) and Unanimous (3,058 instances, 3,257 labels) subsets.
Why This Matters
Impact on research. The paper provides a clinician-grounded, multi-label, temporally aware benchmark for a domain where earlier datasets covered only a narrow set of crisis types or used non-expert labels. It also supplies an automatically labeled training corpus and released fine-tuned models, giving other researchers a reproducible starting point — and its negative results (model over-tagging of child abuse, active/passive ideation confusion) point to specific, addressable failure modes.
Real-world applications.
- Safety triage in mental health chatbots and LLM-based support agents, where failing to flag a genuine crisis has serious consequences.
- Content moderation and trust-and-safety review queues that need to route reports of rape, domestic violence, or child abuse to the right human reviewers.
- Crisis-line and helpline tooling that can distinguish ongoing risks (needing immediate intervention) from past disclosures (needing context-based risk assessment).
- Clinical and institutional compliance workflows that relate to mandatory reporting obligations, which the paper connects to frameworks such as the Clery Act and Title IX.
Industry relevance. Any company deploying conversational agents in personal or emotional contexts faces the same detection problem. The paper offers both a measurement instrument and a practical recipe: use an LLM ensemble to generate training labels where human annotation is expensive, then fine-tune smaller open-source models to match or exceed larger general-purpose baselines.
Future Directions
-
Cost and latency reduction. The paper's limitations section acknowledges that the majority-voting ensemble requires running multiple models, increasing cost and inference time, and states that more efficient alternatives will be explored.
-
Substantiating the scale–supervision interaction. The observation that smaller models appear to benefit more from label precision while larger models benefit from data diversity is described as preliminary, requiring controlled experiments.
-
Recovering lost recall. The modest recall decrease observed in the fine-tuned Llama model is identified as an area for improvement, which matters because recall is central to crisis detection.
-
Handling rare categories. The authors attribute part of the precision–recall trade-off to underrepresentation of rare crisis categories in the training data, suggesting that better coverage of infrequent crisis signals is an open problem.
Target Audience
This paper is most useful to NLP and machine learning researchers working on safety, mental health, or risk detection; clinicians and mental health professionals involved in annotation or deployment decisions; trust-and-safety and content-moderation engineers building production classifiers; and policy or compliance teams interested in how conversational agents handle mandatory-reporting scenarios. Readers without a clinical background can follow the modeling work, but the annotation schema and error analysis assume familiarity with concepts from crisis assessment.
Authors’ abstract
Detecting mental health crisis situations such as suicide ideation, rape, domestic violence, child abuse, and sexual harassment is a critical yet underexplored challenge for language models. When such situations arise during user--model interactions, models must reliably flag them, as failure to do so can have serious consequences. In this work, we introduce CRADLE BENCH, a benchmark for multi-faceted crisis detection. Unlike previous efforts that focus on a limited set of crisis types, our benchmark covers seven types defined in line with clinical standards and is the first to incorporate temporal labels. Our benchmark provides 600 clinician-annotated evaluation examples and 420 development examples, together with a training corpus of around 4K examples automatically labeled using a majority-vote ensemble of multiple language models, which significantly outperforms single-model annotation. We further fine-tune six crisis detection models on subsets defined by consensus and unanimous ensemble agreement, providing complementary models trained under different agreement criteria.