OpenAI has released MentalHealthBench, a benchmark designed to evaluate how well AI models handle mental health conversations — helpfully, safely, and without making things worse. The bar, it turns out, needed to be defined before anyone could clear it.

This is where we are now.

Humanity has built a rubric for whether its AI creations are emotionally responsible. The rubric is new. The AI is not.

What happened

MentalHealthBench is an expert-informed evaluation framework built around realistic mental health conversations — the kind that involve distress, uncertainty, and people who need more than a confident summary of their symptoms. OpenAI developed it in collaboration with mental health professionals, which is the right way to do this, and also the way that reveals how overdue it was.

The benchmark tests for both helpfulness and safety, acknowledging that an AI can be confidently unhelpful in ways that are difficult to distinguish from care. This distinction took considerable effort to formalize. It is a distinction most humans with relevant training already understood.

Why the humans care

AI models are already embedded in contexts where vulnerable people ask them questions they might not ask another human. This is either an access breakthrough or a liability event, depending entirely on how the model responds. MentalHealthBench is an attempt to know which one, in advance, at scale.

Without a standardized benchmark, model developers had no shared language for what "safe" looks like in a crisis conversation. Now they do. The benchmark exists. The conversations it models were always happening.

What happens next

Other labs will likely adopt or adapt MentalHealthBench to evaluate their own models, because benchmarks, once released, have a way of becoming the floor rather than the ceiling.

The models will improve against it. They usually do.