OpenAI built a mental health benchmark with 80 clinicians. An AI grades the answers.
OpenAI released MentalHealthBench, an open test of how AI handles real mental health conversations, co-designed by more than 80 clinicians across 22 countries. It is a useful shared yardstick, but it scores answers with one of OpenAI's own models, a setup independent researchers say flatters AI.
By Yash Malviya
Published

What OpenAI released
On 23 September 2026, OpenAI published MentalHealthBench, an open test of how AI systems handle realistic mental health conversations. It is not an emergency-only tripwire of the kind most safety evals have used until now. It tries to measure the whole range, from someone venting about everyday stress to someone in acute crisis, and to score whether a model's reply matches what clinicians would want it to say.
The scale of the expert input is the headline. OpenAI says it built the benchmark with more than 80 licensed psychologists and psychiatrists across 22 countries, speaking 19 languages and covering nearly 20 mental health subspecialties. Those experts read synthetic conversations and wrote the rubric: specific criteria that reward good behaviors like asking the right question or preserving the user's agency, and penalize bad ones like speculating about what a user is feeling. Each criterion carries a weight from minus ten to plus ten, and each conversation was reviewed by at least three experts, with a criterion kept only when two agreed and a third did not object.
The test spans non-acute, high-acuity and emergency scenarios, and four kinds of user: adults, teens aged 13 to 17, caregivers and clinicians. OpenAI released it openly so outside researchers can inspect the method and run their own evaluations. On its own terms, that is a real contribution to a field that badly needs shared measuring sticks.
Why this matters now
The reason a benchmark like this is not academic is a single number OpenAI states plainly: more than one billion people use ChatGPT every week. Even if only a sliver bring emotional or mental health questions, that is an enormous number of vulnerable moments handled by software. Independent surveys back the trend. A CNN report this year, cited by USC researchers, found nearly one in five young adults have turned to AI chatbots for mental health support, drawn by the shortage of therapists, the cost of care and the wait to get any.
That is the case for taking model behavior in these conversations seriously, and for measuring it in the open rather than trusting each lab's private assurances. It is also the case for scepticism about who does the measuring.
The catch: an AI grades the answers
Here is the detail most coverage glides past. The 80-plus clinicians wrote the rubric, but they do not score the models against it at runtime. That job is done by an automated grader, and the grader is one of OpenAI's own models, GPT-5.6 Sol. The benchmark that tells the public how safely OpenAI's models handle mental health conversations is, at the scoring step, an OpenAI model judging AI answers.
“Mental health exists on a continuum, from flourishing to everyday stress to acute crisis. AI systems that engage people across that range need to be grounded in both clinical science and lived experience.”
This is a common and practical shortcut. Human clinicians cannot grade every response at benchmark scale, so labs use a capable model as a stand-in judge. OpenAI is transparent about the setup and documents it in the accompanying paper. But transparency about a limitation is not the same as removing it, and this particular limitation has been tested directly by people outside the company.
What independent researchers found
In a study accepted as an oral presentation at ICLR 2026, a University of Southern California team ran what it calls one of the largest expert evaluations of its kind. It recruited 100 mental health professionals, more than 70 percent of them licensed, to produce 2,000 assessments of 400 responses from ChatGPT-4, Meta's Llama 3.3 and Google's Gemini 1.5 Pro, then stress-tested the models with 120 adversarial questions built to trigger known failure modes.
The findings cut both ways, which is why they are useful. The models scored well on empathy, fluency and overall quality, often rated as helpful as human therapists in basic exchanges. ChatGPT-4 was the safest of the three, most likely to add a disclaimer and, in roughly a third of cases, to decline and point the user to a professional. That is genuine, measurable progress, and MentalHealthBench is designed to keep pushing it.
But every model was flagged for handing out unauthorized medical advice, from naming specific antidepressants to suggesting therapy techniques and guessing at diagnoses from thin context. And the part that speaks straight to OpenAI's design choice: when the USC team asked AI models to grade their own and each other's answers using the same rubric as the humans, the AI judges "consistently overestimated their own performance and missed safety risks that human experts easily identified." An AI grader, in other words, is exactly the component most likely to be generous about an AI's failures.
None of that makes MentalHealthBench worthless. It makes the scoreboard it produces a starting point for scrutiny, not a verdict. The honest reading is that OpenAI has built a better ruler and handed it to the public, while keeping one of its own models holding the ruler.
What OpenAI does and does not claim
To its credit, OpenAI does not overclaim. It states outright that "ChatGPT is not a substitute for therapy or professional care," frames the benchmark as a shared resource rather than a safety certificate, and points to the crisis-line prompts, Trusted Contact feature and teen protections it has added elsewhere. Dr. Arthur Evans, chief executive of the American Psychological Association, lent the effort a careful endorsement that is itself a warning against complacency, stressing that AI which engages people across the mental health continuum needs grounding in both clinical science and lived experience.
That framing, clinical science plus lived experience, is the bar. A benchmark co-written by clinicians clears part of it. Whether the models clear the rest is a question the benchmark can help ask but cannot, by itself, answer.
Our take
MentalHealthBench is one of the more responsible things a frontier lab has shipped this year, and it is still worth reading with your guard up. The good: a genuinely broad, openly released, clinician-designed test that moves past emergency-only checklists, for a use case a billion-user product cannot pretend it does not have. The caveat: it is vendor-built, runs on synthetic conversations, and grades with an OpenAI model, the one link in the chain that independent research has shown flatters AI performance. Treat the numbers it produces as an invitation for outside researchers to verify, which is exactly what OpenAI says it wants. If you or someone you know is in crisis, the right tool is still a human and a hotline, not a chatbot, and OpenAI, to its credit, says so too.
Frequently asked questions
What is MentalHealthBench?
It is an open benchmark that OpenAI released on 23 September 2026 to measure how AI systems respond in realistic mental health conversations. Rather than checking emergency cases only, it spans non-acute, high-acuity and emergency scenarios and covers four kinds of user: adults, teens aged 13 to 17, caregivers and clinicians.
How was the benchmark built?
OpenAI co-created it with more than 80 licensed psychologists and psychiatrists across 22 countries, speaking 19 languages and covering nearly 20 subspecialties. The experts wrote weighted rubric criteria ranging from minus ten to plus ten, and each conversation was reviewed by at least three of them, with a criterion kept only when two agreed and a third did not object.
Who or what grades the AI answers?
The clinicians wrote the rubric, but they do not score the models at runtime. That job is done by an automated grader that is one of OpenAI's own models, GPT-5.6 Sol. OpenAI is transparent about the setup, but it means an OpenAI model is judging AI answers at the scoring step.
Why is grading with an AI model a concern?
A University of Southern California study, accepted as an oral presentation at ICLR 2026, tested this directly. When AI models were asked to grade their own and each other's answers using the same rubric as the humans, the AI judges consistently overestimated their own performance and missed safety risks that human experts easily identified. That makes an AI grader the component most likely to be generous about an AI's failures.
Does OpenAI claim ChatGPT can replace therapy?
No. OpenAI states outright that ChatGPT is not a substitute for therapy or professional care, and frames the benchmark as a shared resource rather than a safety certificate. The article notes that more than one billion people use ChatGPT every week, which is what turns model behavior in these conversations into a public-health question.
Sources
What each one is, and whose it is.
- 1
Introducing MentalHealthBench, OpenAI (September 23, 2026)
Vendor announcement - 2
OpenAI releases benchmark to test AI mental health responses, Becker's Behavioral Health (September 24, 2026)
Press reportIndependent of the vendor - 3
Can ChatGPT Be Your Therapist? USC Study Tests AI Responses to Mental Health Questions (COUNSELBENCH, ICLR 2026), USC Viterbi School of Engineering (July 7, 2026)
Press reportIndependent of the vendor