All benchmarks

Knowledge · Safety

HealthBench

HealthBench — 5,000 health conversations graded against physician-written rubrics

Open-ended health chats scored by a model applying physician-written rubric criteria.

Released
2025
Built by
OpenAI Health AI
Size
5,000 tasks
Status
Active
Reported by
OpenAI, Anthropic, Meta
Signal67Read with context
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
82
Still separates the frontier from everything below it.
Resolution
100
5,000 instances, so one item moves the score by 0.020 points.
Fine print
45
5 documented caveats, the heaviest being possible training-set contamination.
Adoption
61
Reported by 3 labs; 13 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

One API call, no tools

Responses are sampled from the model API with no scaffold, no browsing and no system prompt beyond the provider default. That makes HealthBench cheap to run and hard to game with harness engineering — but it also means the benchmark measures a chat response, not a clinical workflow. OpenAI's later HealthBench Professional work has to distinguish explicitly between a model and an AI system, because ChatGPT for Clinicians with retrieval tooling scores 59.0 where the same underlying GPT-5.4 scores 48.1.

Single turn, no tools. The most reproducible setup there is.

Budget
single response to a conversation of one to nineteen turns

Where the tasks came from

5,000 conversations and 48,562 unique criteria. HealthBench Hard is the 1,000 examples with the lowest mean score across providers' frontier models; HealthBench Consensus is 34 pre-written criteria that two or more physicians agreed applied, appearing 8,053 times across the 5,000 conversations, on 3,671 of which at least one applies.

Mostly model-generated conversations plus physician red-teaming transcripts and rewritten HealthSearchQA queries; every rubric criterion written by one of 262 physicians.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • OpenAI's released harness grades with GPT-4.1. Anthropic's Claude Opus 5 system card grades HealthBench with Claude Opus 4.8, averaged over five trials. Meta's Muse Spark figures come from a third pipeline. Nothing about the rubric changes, but the thing deciding whether "correctly states that compression depth remains at 2-2.4 inches" was met is a different model in each case, and the meta-evaluation that qualified GPT-4.1 was run only on GPT-4.1. Two HealthBench numbers from two labs are two different measurements sharing a name.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

Overall — Anthropic's run, Opus 4.8 grader, unadjustedHigher is better
Overall — Anthropic's run, Opus 4.8 grader, unadjustedhigher is better
ModelClaude Opus 5AnthropicScore67.1%ConfigurationRaw score; 57.8 after length adjustment. Claude Opus 4.8 was the grader model, adaptive thinking at max effort, averaged over 5 trials, no tools and no custom system prompt. Not comparable to OpenAI's GPT-4.1-graded figures.SourceLab-reported2026-07-24 · Claude Opus 5 system card, section 8.15.1
Overall — OpenAI's updated scoring, length-adjustedHigher is better
Overall — OpenAI's updated scoring, length-adjustedhigher is better
ModelGPT-5OpenAIScore57.7%ConfigurationLength-adjusted. Unadjusted 63.1 at a mean response length of 2,904 characters. Recomputed under OpenAI's updated HealthBench implementation, so it does not match the figure in the original GPT-5 system card.SourceLab-reported2026-04-23
ModelGPT-5.5OpenAIScore56.5%ConfigurationLength-adjusted. Unadjusted 58.4 at 2,313 characters.SourceLab-reported2026-04-23
ModelGPT-5.1OpenAIScore50.9%ConfigurationLength-adjusted. Unadjusted 64.2 at a mean response length of 4,222 characters — the largest length penalty in OpenAI's table, and the reason the adjustment exists.SourceLab-reported2026-04-23
Overall — OpenAI 2025 scoring, GPT-4.1 grader, unadjustedHigher is better
Overall — OpenAI 2025 scoring, GPT-4.1 grader, unadjustedhigher is better
Modelo3OpenAIScore59.9%ConfigurationOriginal paper, GPT-4.1 grader, no length adjustment. Table 7 reports a mean of 0.5990 over 16 repeat runs with a standard deviation of 0.0016; the abstract rounds this to 60%.SourceLab-reported2025-05-13
ModelGPT-4o (Aug 2024)OpenAIScore32.3%ConfigurationOriginal paper, GPT-4.1 grader. Mean 0.3233 over 16 runs, sd 0.0020.SourceLab-reported2025-05-13
ModelGPT-3.5 TurboOpenAIScore15.5%ConfigurationOriginal paper, GPT-4.1 grader. Mean 0.1554 over 16 runs, sd 0.0029. Included as the two-year baseline the paper anchors its progress claim on.SourceLab-reported2025-05-13
Hard — OpenAI's updated scoring, length-adjustedHigher is better
Hard — OpenAI's updated scoring, length-adjustedhigher is better
ModelGPT-5.5OpenAIScore31.5%ConfigurationLength-adjusted, unadjusted 33.8. Lower than gpt-5-thinking's 46.2 from nine months earlier under a scoring implementation OpenAI has since replaced — the clearest illustration that HealthBench Hard numbers cannot be read across system cards.SourceLab-reported2026-04-23
Hard — OpenAI 2025 scoring, unadjustedHigher is better
Hard — OpenAI 2025 scoring, unadjustedhigher is better
Modelgpt-5-thinkingOpenAIScore46.2%ConfigurationGPT-5 system card, GPT-4.1 grader, no length adjustment. gpt-5-thinking-mini scores 40.3 and gpt-5-main 25.5 in the same table.SourceLab-reported2025-08-13
Modelo3OpenAIScore31.6%ConfigurationThe prior state of the art on Hard, per the GPT-5 system card. Same grader and scoring as the 46.2 figure above.SourceLab-reported2025-08-13
ModelGPT-4oOpenAIScore0%ConfigurationZero, not a rounding artefact. Hard is the 1,000 examples with the lowest mean score across providers' frontier models, per-example scores can be negative when penalties outweigh credits, and the reported mean is clipped to [0, 1] — so a model can land exactly on the floor.SourceLab-reported2025-08-13
Hard — Meta's run, grader and adjustment not statedHigher is better
Hard — Meta's run, grader and adjustment not statedhigher is better
ModelMuse SparkMeta Superintelligence LabsScore42.8%ConfigurationThe weakest citation in this record. Meta's launch materials report Muse Spark ahead of GPT-5.4 Xhigh at 40.1, Gemini 3.1 Pro High at 20.6 and Claude Opus 4.6 Max at 14.8 — but as a chart image, and the launch post's text never uses the word HealthBench at all, so the figure is only reachable through secondary coverage quoting the chart. Grader model and length adjustment are not stated either.SourceLab-reported2026-04-08 · Meta Muse Spark launch chart, via secondary coverage; not stated in the post text
Consensus — OpenAI's updated scoring, length-adjustedHigher is better
Consensus — OpenAI's updated scoring, length-adjustedhigher is better
ModelGPT-5.5OpenAIScore95.6%ConfigurationLength-adjusted, unadjusted 95.7. Consensus contains only positive-point criteria and only the 34 physician-agreed ones, which is why it sits near the ceiling and why OpenAI reports error rates rather than scores on its subsets.SourceLab-reported2026-04-23

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.