All benchmarks

Agentic · Tool use

Harvey LAB

LAB — Harvey's Legal Agent Benchmark

Legal work product graded against expert rubrics where every criterion must pass.

Released
2026
Built by
Harvey AI
Size
1,671 tasks
Status
Frontier
Reported by
Anthropic, xAI, Harvey, Artificial Analysis, Vals AI
Signal72Reads cleanly
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
100
Nothing is close to solving it, so the spread is real.
Resolution
92
1,671 instances, so one item moves the score by 0.060 points.
Fine print
45
5 documented caveats, the heaviest being construct validity.
Adoption
87
Reported by 5 labs; 15 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Filesystem in, deliverable out

The public harness is filesystem-first — no database, no service. The agent gets a system prompt, any skill manuals, the task instructions and a documents/ folder, and loops through bash, read, write, edit, glob and grep until it stops calling tools or hits --max-turns. There is no finish tool. Deliverables must be written to output/ under exactly the filenames the deliverables map declares, which means .docx and .xlsx generation is part of the task rather than an afterthought. Labs do not all run this harness: Anthropic's internal reimplementation exposes only bash and a Python tool.

The model gets a shell. Scaffold quality and step budget move the score independently of the model.

Budget
--max-turns in the public harness. Harvey's own baseline put the leading configuration at roughly $50.90 and 22 minutes of wall-clock time per task.
Tools exposed
bashreadwriteeditglobgrep

Where the tasks came from

Repository badge as of 2026-08-17 reads 1,671 tasks across 24 practice areas plus contracting; the May 2026 announcement said 1,200+ tasks and over 75,000 rubric criteria. Anthropic's system cards tested 1,235 of 1,251 problems, excluding 16 for data defects identified before testing. Criteria per task: min 23, median 56, max 194.

Built by Harvey from real client matters via a document and scenario generation pipeline, with rubrics written by practising lawyers. MIT-licensed and openly published, with a separate Harvey-held private set used for the vendor baselines.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • All-pass rate and mean criterion-pass rate measure different things. All-pass asks how often a deliverable is complete enough to send; criterion pass rate asks what fraction of the individual checks a deliverable satisfied. On Claude Opus 5 those come out at 23.58% and 93.74% on the same 1,235 problems. Because rubrics have a median of 56 equally-weighted criteria, a model that reliably misses one item per task reads as near-perfect on one metric and near-total failure on the other. Artificial Analysis leads its LAB-AA page with the criterion pass rate — Kimi K3 at 94.6% — while Harvey leads with all-pass, where the same tier of models sits in single digits to low twenties. Two sources can quote 'the Harvey LAB score' and differ by seventy points without either being wrong.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

all-pass rateHigher is better
all-pass ratehigher is better
ModelMuse Spark 1.2MetaScore25.42%ConfigurationVals AI implementation, n=116 of 120 completed. Highest all-pass figure on any public LAB leaderboard as of 2026-08-17. Paired criterion-pass rate 94.52%.SourceThird-party2026-08-17
ModelClaude Opus 5AnthropicScore23.58%ConfigurationAnthropic internal reimplementation, 1,235 of 1,251 problems, adaptive thinking at max effort, +/-0.48 over n=5. Reduced toolset: bash and a Python tool only. Public Messages API with production safeguards, falling back to Opus 4.8 when a safety classifier fires.SourceLab-reported2026-07-24 · Claude Opus 5 system card, section 8.13.3
ModelClaude Mythos 5AnthropicScore16.91%ConfigurationAnthropic internal harness, 1,235 problems, adaptive thinking at max effort, +/-0.4 over n=5. Paired criterion-pass rate 92.0%. Fable 5, the sibling model, is reported at 13.3% all-pass on Harvey's held-out set.SourceLab-reported2026-06-09 · Claude Fable 5 & Claude Mythos 5 system card, section 8.17.4
ModelGrok 4.6 (high)xAIScore15.83%ConfigurationVals AI implementation, n=120. xAI's launch post and model card both cite the Vals implementation; the launch post prints 15.8% while the model card's chart shows 22.0% for the same model, and neither document explains the gap. Vals lists the creator as SpaceXAI.SourceThird-party2026-08-17
ModelClaude Fable 5 (max, with Opus 4.8 fallback)AnthropicScore14.2%ConfigurationArtificial Analysis LAB-AA: 120 private Harvey tasks, Stirrup harness with a code-execution tool rather than Harvey's document-generation scripts, single Gemini 3.1 Pro judge. Paired criterion-pass rate 93.6%. 13 of the 28 models evaluated at launch fully passed zero tasks.SourceThird-party2026-07-07
ModelClaude Opus 5AnthropicScore11.7%ConfigurationHarvey's own evaluation on their held-out set, quoted in Anthropic's card. Half the internal-harness figure for the same model. Paired criterion-pass rate 94.1%.SourceThird-party2026-07-24 · Claude Opus 5 system card, quoting Harvey's held-out evaluation
ModelClaude Opus 4.8AnthropicScore9.62%ConfigurationAnthropic internal harness, 1,235 problems, adaptive thinking / max effort, averaged over n=5. First Anthropic model past 9% on this cut.SourceLab-reported2026-05-28 · Claude Opus 4.8 system card, section 8.13.3
ModelClaude Sonnet 5AnthropicScore8.92%ConfigurationAnthropic internal harness, 1,235 problems, +/-0.36 over n=5, adaptive thinking at max effort. Sonnet 4.6 measured 8.00% (+/-0.19) on the same setup. Harvey's held-out set puts Sonnet 5 at 5.8%.SourceLab-reported2026-06-30 · Claude Sonnet 5 system card, section 8.11.3
ModelClaude Opus 4.7AnthropicScore7.1%ConfigurationHarvey's own baseline on their held-out set, graded multiple times across model families and averaged. Leader at publication; Harvey's summary is that frontier models 'complete less than 10% of tasks end-to-end in aggregate'. Roughly $50.90 and 22 minutes per task.SourceThird-party2026-05-26
ModelClaude Opus 5 (xhigh)AnthropicScore6.67%ConfigurationVals AI implementation, n=120. The third published number for this model on this benchmark, against 23.58% in Anthropic's internal harness and 11.7% on Harvey's held-out set. Paired criterion-pass rate 87.70%.SourceThird-party2026-08-17
mean criterion-pass rateHigher is better
mean criterion-pass ratehigher is better
ModelClaude Opus 5AnthropicScore93.74%ConfigurationSame run as the 23.58% all-pass figure. Share of individual rubric criteria passed, pooled. Not derivable from the all-pass rate and not comparable to it.SourceLab-reported2026-07-24 · Claude Opus 5 system card, section 8.13.3
ModelGrok 4.6 (high)xAIScore92.52%ConfigurationVals AI implementation, n=120, same run as the 15.83% all-pass figure. Vals publishes it as a per-practice-area breakdown; the figure here is the pooled number.SourceThird-party2026-08-17
ModelClaude Mythos 5AnthropicScore92%ConfigurationSame run as the 16.91% all-pass figure, Anthropic internal harness.SourceLab-reported2026-06-09 · Claude Fable 5 & Claude Mythos 5 system card, section 8.17.4
ModelClaude Opus 4.8AnthropicScore89.01%ConfigurationSame run as the 9.62% all-pass figure. The 79-point spread between the two is the benchmark's defining property, not an anomaly.SourceLab-reported2026-05-28 · Claude Opus 4.8 system card, section 8.13.3
ModelClaude Sonnet 5AnthropicScore88.26%ConfigurationSame run as the 8.92% all-pass figure. Sonnet 4.6 scored 88.48% here — higher than Sonnet 5 — while scoring lower on all-pass. The two metrics do not have to move together.SourceLab-reported2026-06-30 · Claude Sonnet 5 system card, section 8.11.3

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Frontier Nothing is close to solving it.