What does it measure?

OpenAI's launch page says every GPT-5.6 model it reports outperformed GPT-5.5 on HealthBench Professional. This is the company's own comparison; the page does not provide model-level scores, effect sizes or confidence intervals. The direction of improvement is stated, but its practical importance cannot be calculated from this announcement.[1], [2]

The benchmark is richer than an ordinary medical exam. Its paper covers three clinician uses: care consultation, writing and documentation, and medical research. Each example contains a physician-authored conversation and is scored with rubrics iteratively adjudicated by three or more physicians. From a candidate pool of 15,079 examples, difficult cases for current OpenAI models were enriched by roughly 3.5 times; about one-third involve deliberate adversarial testing.[2]

What does it not measure?

This design is useful for stress-testing difficult responses, but the sample is deliberately tilted toward cases that challenge current OpenAI models. The score measures whether a chat response satisfies a rubric; it does not track diagnostic accuracy in a real clinic, workflow harm, clinician oversight or patient outcomes. The benchmark is open, making independent replication possible, but the announcement does not provide one.[2], [3]

My evidence ladder has four steps here: an open benchmark, independent replication, prospective workflow studies across different clinics, and only then patient outcomes. The first step exists; the others cannot be inferred from this announcement. The safe user takeaway is equally plain: check important information against the original source and make medical decisions with a qualified health professional.[1], [2]