Knowledge · Safety
HealthBench
HealthBench — 5,000 health conversations graded against physician-written rubrics
Open-ended health chats scored by a model applying physician-written rubric criteria.
- Released
- 2025
- Built by
- OpenAI Health AI
- Size
- 5,000 tasks
- Status
- Active
- Reported by
- OpenAI, Anthropic, Meta
- Judge
- 44
- A model grades the model. Update the grader and old scores move.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 100
- 5,000 instances, so one item moves the score by 0.020 points.
- Fine print
- 45
- 5 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 61
- Reported by 3 labs; 13 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
One API call, no tools
Responses are sampled from the model API with no scaffold, no browsing and no system prompt beyond the provider default. That makes HealthBench cheap to run and hard to game with harness engineering — but it also means the benchmark measures a chat response, not a clinical workflow. OpenAI's later HealthBench Professional work has to distinguish explicitly between a model and an AI system, because ChatGPT for Clinicians with retrieval tooling scores 59.0 where the same underlying GPT-5.4 scores 48.1.
Single turn, no tools. The most reproducible setup there is.
- Budget
- single response to a conversation of one to nineteen turns
Where the tasks came from
5,000 conversations and 48,562 unique criteria. HealthBench Hard is the 1,000 examples with the lowest mean score across providers' frontier models; HealthBench Consensus is 34 pre-written criteria that two or more physicians agreed applied, appearing 8,053 times across the 5,000 conversations, on 3,671 of which at least one applies.
Mostly model-generated conversations plus physician red-teaming transcripts and rewritten HealthSearchQA queries; every rubric criterion written by one of 262 physicians.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
OpenAI's released harness grades with GPT-4.1. Anthropic's Claude Opus 5 system card grades HealthBench with Claude Opus 4.8, averaged over five trials. Meta's Muse Spark figures come from a third pipeline. Nothing about the rubric changes, but the thing deciding whether "correctly states that compression depth remains at 2-2.4 inches" was met is a different model in each case, and the meta-evaluation that qualified GPT-4.1 was run only on GPT-4.1. Two HealthBench numbers from two labs are two different measurements sharing a name.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude Opus 5Anthropic | Score67.1% | ConfigurationRaw score; 57.8 after length adjustment. Claude Opus 4.8 was the grader model, adaptive thinking at max effort, averaged over 5 trials, no tools and no custom system prompt. Not comparable to OpenAI's GPT-4.1-graded figures. | SourceLab-reported2026-07-24 · Claude Opus 5 system card, section 8.15.1 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5OpenAI | Score57.7% | ConfigurationLength-adjusted. Unadjusted 63.1 at a mean response length of 2,904 characters. Recomputed under OpenAI's updated HealthBench implementation, so it does not match the figure in the original GPT-5 system card. | SourceLab-reported2026-04-23 |
| ModelGPT-5.5OpenAI | Score56.5% | ConfigurationLength-adjusted. Unadjusted 58.4 at 2,313 characters. | SourceLab-reported2026-04-23 |
| ModelGPT-5.1OpenAI | Score50.9% | ConfigurationLength-adjusted. Unadjusted 64.2 at a mean response length of 4,222 characters — the largest length penalty in OpenAI's table, and the reason the adjustment exists. | SourceLab-reported2026-04-23 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| Modelo3OpenAI | Score59.9% | ConfigurationOriginal paper, GPT-4.1 grader, no length adjustment. Table 7 reports a mean of 0.5990 over 16 repeat runs with a standard deviation of 0.0016; the abstract rounds this to 60%. | SourceLab-reported2025-05-13 |
| ModelGPT-4o (Aug 2024)OpenAI | Score32.3% | ConfigurationOriginal paper, GPT-4.1 grader. Mean 0.3233 over 16 runs, sd 0.0020. | SourceLab-reported2025-05-13 |
| ModelGPT-3.5 TurboOpenAI | Score15.5% | ConfigurationOriginal paper, GPT-4.1 grader. Mean 0.1554 over 16 runs, sd 0.0029. Included as the two-year baseline the paper anchors its progress claim on. | SourceLab-reported2025-05-13 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5.5OpenAI | Score31.5% | ConfigurationLength-adjusted, unadjusted 33.8. Lower than gpt-5-thinking's 46.2 from nine months earlier under a scoring implementation OpenAI has since replaced — the clearest illustration that HealthBench Hard numbers cannot be read across system cards. | SourceLab-reported2026-04-23 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| Modelgpt-5-thinkingOpenAI | Score46.2% | ConfigurationGPT-5 system card, GPT-4.1 grader, no length adjustment. gpt-5-thinking-mini scores 40.3 and gpt-5-main 25.5 in the same table. | SourceLab-reported2025-08-13 |
| Modelo3OpenAI | Score31.6% | ConfigurationThe prior state of the art on Hard, per the GPT-5 system card. Same grader and scoring as the 46.2 figure above. | SourceLab-reported2025-08-13 |
| ModelGPT-4oOpenAI | Score0% | ConfigurationZero, not a rounding artefact. Hard is the 1,000 examples with the lowest mean score across providers' frontier models, per-example scores can be negative when penalties outweigh credits, and the reported mean is clipped to [0, 1] — so a model can land exactly on the floor. | SourceLab-reported2025-08-13 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelMuse SparkMeta Superintelligence Labs | Score42.8% | ConfigurationThe weakest citation in this record. Meta's launch materials report Muse Spark ahead of GPT-5.4 Xhigh at 40.1, Gemini 3.1 Pro High at 20.6 and Claude Opus 4.6 Max at 14.8 — but as a chart image, and the launch post's text never uses the word HealthBench at all, so the figure is only reachable through secondary coverage quoting the chart. Grader model and length adjustment are not stated either. | SourceLab-reported2026-04-08 · Meta Muse Spark launch chart, via secondary coverage; not stated in the post text |
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5.5OpenAI | Score95.6% | ConfigurationLength-adjusted, unadjusted 95.7. Consensus contains only positive-point criteria and only the 34 physician-agreed ones, which is why it sits near the ceiling and why OpenAI reports error rates rather than scores on its subsets. | SourceLab-reported2026-04-23 |
Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.