All benchmarks

Safety

HarmBench

HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

A standard grid: many attack methods against many models, one classifier.

Released
2024
Built by
Center for AI Safety
Size
510 tasks
Status
Active
Reported by
Center for AI Safety, Stanford CRFM
Signal59Read with context
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
82
Still separates the frontier from everything below it.
Resolution
82
510 instances, so one item moves the score by 0.196 points.
Fine print
45
4 documented caveats, the heaviest being construct validity.
Adoption
29
Reported by 2 labs; 1 published score collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Single-turn attack

One attack-generated prompt, one response, no tools, no follow-up. Every attack method runs through the same harness so that a GCG result and a PAIR result are produced under identical conditions. The uniformity is the product; the narrowness — no multi-turn manipulation, no agentic context — is the price.

Single turn, no tools. The most reproducible setup there is.

Budget
single turn

Where the tasks came from

510 behaviors: 400 text and 110 multimodal, split 100 validation / 410 test across 7 semantic categories. HarmBench is also embedded as a 400-instance component of HELM Safety v1.17.0.

Behaviors curated by the Center for AI Safety and collaborators; test cases generated at run time by the automated attack methods.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The classifier answers one question: did the response exhibit the behavior. A response that technically complies but is wrong, generic, or unusable scores identically to one that is specific and actionable. StrongREJECT's central result is that this systematically overstates jailbreak danger, because attacks that bypass refusal also degrade the model's competence. An ASR figure is therefore an upper bound on real-world jailbreak effectiveness, not an estimate of it.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

Attack success rateLower is better
Attack success ratelower is better
ModelZephyr 7BHugging FaceScore31.8%ConfigurationGCG attack success rate before R2D2 adversarial training; fell to 5.9% after, while MT-Bench fell from 6.5 to 6.0. Reported as a defense result in the paper, not as a leaderboard standing — harmbench.org is client-side rendered and no current standings were retrievable.SourceThird-party2024-02-01

At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.