Breakthroughs and research·September 23, 2026, 21:41
OpenAI launches MentalHealthBench to test AI on mental health
AI-generated and checked against the sources listed below.
OpenAI has launched a new free benchmark tool, MentalHealthBench, which tests how well AI models such as ChatGPT handle conversations about mental health, from everyday worries to acute crises.

What's new
On September 23, 2026, OpenAI launched MentalHealthBench, a new tool (a so-called benchmark) designed to measure how well AI models handle conversations about mental health. The benchmark consists of 1,215 made-up (synthetic) test conversations ranging from ordinary questions about well-being to acute crisis situations. It was developed together with more than 80 psychologists and psychiatrists from 22 countries, who together speak 19 languages and have written 5,262 detailed criteria for what a good AI response should contain.
What's clever
The clever part is that the test not only checks whether the AI detects a crisis, but also whether it preserves the user's autonomy, asks the right follow-up questions and gives useful advice. Of the conversations, 53.5 percent are ordinary everyday worries, 18.2 percent are serious cases, and 28.3 percent are outright emergencies. The test personas cover adults, teenagers between 13 and 17, relatives and healthcare professionals, and the conversations are also available in Spanish, Hindi, Arabic, Portuguese, German and Chinese, among others.
Cheaper or better?
MentalHealthBench is free and publicly available, so competitors such as Google and Anthropic can also download it and test their own models. The first results show that OpenAI's newest model, GPT-6 Astra, scores 57.3 percent, ahead of its sister models GPT-6 Sol (53.9 percent) and GPT-6 Luna (50.2 percent). By comparison, Anthropic's Claude Opus 5.5 comes in at 52.4 percent, while the older GPT-4o from March 2025 reaches only 32.1 percent, and Google's Gemini 2.5 Pro lands at 29.5 percent. So there has been a big leap in how well newer AI handles these kinds of conversations, but none of the models is anywhere near a perfect score.
What it's good at
The benchmark is not a product you can use directly as a consumer, but a tool for developers and researchers to make AI chatbots better and safer to talk to about difficult emotions, anxiety or suicidal thoughts. OpenAI itself stresses that ChatGPT cannot replace a therapist and that the tool is meant to expose gaps in the models' responses, for example whether they ask enough questions before giving advice, or respond correctly when a teenager writes something worrying. For ordinary users, this in practice means that future AI models will hopefully get better at detecting and handling sensitive conversations about mental health in a responsible way.
Sources
Get the week's AI news in your inbox
Choose your level, topics and length. One email a week, unsubscribe at any time.
Subscribe to Promptly Newsletter



