Five models, one statistical knot

A team that includes Beth Israel Deaconess psychiatrist John Torous screened OpenAI's HealthBench — 5,000 physician-rubric conversations — for mental-health relevance, ran the resulting subset through two rounds of blinded clinician review, and produced the 610-conversation HealthBench-Psych. They scored 20 models with three AI judges (GPT-4.1, Claude Haiku 4.5, Gemini 2.5 Flash), and checked judge reliability by replicating OpenAI's own published GPT-4.1 result on HealthBench-Hard (0.16) and landing almost exactly on it (0.157).[1]

The result isn't a ranking, it's a heap: Kimi K2.6 (0.627), GPT-5.5 (0.624), Claude Opus 5 (0.620), Grok 4.5 (0.612) and GPT-5.6-Sol (0.610) sit close enough that only two model pairs separate at the uncorrected 95 percent level — and none survive the Holm-Bonferroni correction applied to guard against false positives across the comparisons. There is no “best” among the five that this data can support.[1]

Refusals are data, not virtue

Only two of the tested models refused anything at all. Claude Opus 5 declined 10 of 610 conversations (1.6 percent), concentrated on psychiatric-medication dosing and clinician-voiced patient-management questions. Claude Fable 5 declined 3 of 610 conversations (0.5 percent), all three on Alzheimer's biomarkers and mechanisms. The other eighteen models never refused once; refusals were scored as empty responses under the paper's primary policy.[1]

That pattern doesn't read one way. Concentrating refusals in dosing and neurodegenerative content could mean these two models are deliberately carving out higher-stakes clinical ground — or the same data could just as easily mean an over-cautious filter is costing points against a rubric that rewards a direct, complete answer; the paper doesn't adjudicate between the two. And the fact that the eighteen non-refusing models also land in a tie among themselves raises the possibility that refusal behavior, not clinical judgment quality, is the real axis separating this cluster.[1]

What the score can't tell you yet

The authors' own limitations list isn't short: scores reflect each model's deployed default configuration, not matched inference compute; judge-reliability evidence rests on three judges in a single specialty and, in the authors' own words, should be treated as “suggestive”; reviewers were English-speaking and read 123 non-English conversations through machine translation as their reference; and the screening layer ran inside an orchestration harness carrying its own system-prompt layer, which is why the authors treat it as a pre-filter only. The work is also a preprint, not a peer-reviewed paper.[1]

Read together with those limits, what HealthBench-Psych currently shows is that these five models' outputs resemble each other, under these rules, when graded by other models. That is not the same as knowing which one is safer or more useful for an actual person in a mental-health conversation. The next real evidence milestone would be an independent replication — clinician-scored, run outside the authors' own pipeline, with matched inference compute and non-English-speaking reviewers — before this tie, or any ranking drawn from it, should inform a deployment decision.[1]