Skip to content
AI.info

The Pulse

OpenAI Releases MentalHealthBench, Built With 80+ Clinicians

OpenAI says its open benchmark evaluates AI responses to realistic mental health conversations, with rubrics developed alongside more than 80 licensed experts.

OpenAI Releases MentalHealthBench, Built With 80+ Clinicians

AI.info Team ·

“Mental health exists on a continuum, from flourishing to everyday stress to acute crisis. AI systems that engage people across that range need to be grounded in both clinical science and lived experience—not only to recognize where someone falls on the continuum, but to know how to respond appropriately at any point.”

Dr. Arthur Evans, chief executive officer of the American Psychological Association

OpenAI tests more than crisis responses

OpenAI released MentalHealthBench on September 23, describing it as an open benchmark for evaluating AI replies to realistic mental health conversations. The company says existing tests often concentrate on emergencies and broad safety rules, leaving less visibility into how systems handle common concerns such as relationships, stress and emotional support. MentalHealthBench aims to assess those ordinary exchanges alongside serious distress and urgent safety situations.

The benchmark contains 1,215 synthetic conversations created with privacy-preserving methods intended to reflect patterns in real AI use. Its scenarios cover adults, teenagers, caregivers and clinicians, and include multiple languages and cultural contexts. OpenAI says the benchmark is available for researchers to inspect, run and build on.

Clinicians write the scoring rules

More than 80 licensed psychologists and psychiatrists from 22 countries helped develop the benchmark. The group speaks 19 languages and represents nearly 20 mental health specialties. For each conversation, experts write weighted criteria for evaluating a model’s response to the final user message; positive weights reward helpful behavior, while negative weights penalize harmful behavior.

At least three experts review each conversation’s criteria. OpenAI says it keeps criteria supported by at least two experts and not contradicted by a third. The resulting rubric can reward specific actions—such as asking a relevant follow-up question—and penalize responses that speculate about a user’s feelings or push them toward a decision.

GPT-6 Astra scores 57.3 on the benchmark

OpenAI reports a task-clipped score of 57.3 for GPT-6 Astra, the highest overall result in the company’s published comparison. GPT-6 Sol scores 53.9, Claude Opus 5.5 scores 52.4, and GPT-4o from March 2025 scores 32.1. Those figures measure performance against the benchmark’s weighted rubrics; they are not clinical outcome rates or a measure of treatment effectiveness.

OpenAI uses GPT-5.6 Sol as an automated grader. The research paper says the company samples four responses for each task and has the grader assess them against the expert criteria. That design lets the benchmark apply consistent rules across many models, but it also means the reported results depend in part on an AI judge deciding whether a response met each criterion.

The dataset spans three levels of urgency: 53.5% of conversations are non-acute, 18.2% high-acuity and 28.3% emergencies. Teenagers account for 21.2% of the examples, while adults account for 68.1%; caregivers and clinicians make up the remainder. OpenAI cautions that the scenario mix was chosen for evaluation and does not represent how frequently those topics arise in ChatGPT.

User preferences and clinical guidance diverge

OpenAI also compared expert evaluations with feedback from 44 adults in 16 countries who had used AI for mental health or emotional support. The user review covered non-acute conversations only. Participants placed more emphasis on practical next steps and tone, while experts more often prioritized gathering context and interpreting ambiguous situations carefully.

The distinction matters because a reply that feels useful to a user may still omit safeguards or make assumptions a clinician would reject. OpenAI’s paper describes MentalHealthBench as an evaluation of responses to the final turn in a synthetic conversation, rather than a test of sustained care or real-world health outcomes. Its figures therefore offer a structured comparison of model behavior under defined scenarios—not evidence that a system can replace therapy or professional care.

Source

Explore

More articles