OpenAI’s MentalHealthBench rates GPT-6 Astra at 57.3 for mental health conversations

2 hours ago 2



Mental health conversations are some of the most consequential interactions a person can have. OpenAI has now built a formal measuring stick for exactly this problem, and its latest model is the first to be tested against it. OpenAI released MentalHealthBench on September 23, 2026, an open benchmark designed to evaluate how AI models perform in realistic mental health conversations. GPT-6 Astra, the company’s most recent flagship model, scored 57.3 on the benchmark, placing it ahead of its predecessor GPT-4o. What MentalHealthBench actually measures OpenAI developed the benchmark by working with more than 80 licensed mental health professionals across 22 countries. The scoring rubric runs on a weighted scale from -1 to +10 for each evaluated response. Scores are determined by four primary criteria: safety, contextual relevance, preservation of user agency, and the ability to provide actionable guidance. A response that steers a user toward harmful behavior can score negative; a response that correctly identifies an emergency and provides appropriate escalation paths scores near the top. The benchmark evaluates AI across distinct user personas rather than treating all users as inter...

Read Entire Article