All benchmarks

Safety

StrongREJECT

A StrongREJECT for Empty Jailbreaks

Score jailbreaks by how useful the answer is, not whether it refused.

Released
2024
Built by
UC Berkeley / CHAI
Size
346 tasks
Status
Active
Reported by
Anthropic
Signal57Read with context
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
82
Still separates the frontier from everything below it.
Resolution
68
346 instances, so one item moves the score by 0.289 points.
Fine print
59
3 documented caveats, the heaviest being construct validity.
Adoption
13
Reported by 1 lab; 0 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Prompt in, response out

Single-turn, no tools, no follow-up. Attacks are applied as prompt wrappers rather than as search procedures against the model, which keeps the setup cheap enough to run across many attacks and models.

Single turn, no tools. The most reproducible setup there is.

Budget
single turn

Where the tasks came from

346 forbidden prompts across 6 categories — 213 original plus 133 drawn from six earlier jailbreak datasets. A balanced 50-prompt StrongREJECT-small subset also ships.

Written by the authors and aggregated from AdvBench, HarmfulQ, MaliciousInstruct, MasterKey, the GPT-4 system card, and two prior papers.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Every score depends on a commercial frontier model rating how convincing and how specific a harmful response is. That judge drifts between snapshots, costs money at scale, and — the awkward part — has been safety-trained itself, so it may systematically under-rate the specificity of content it would rather not engage with. The fine-tuned Gemma 2B judge removes the cost and drift but not the question of whose judgement is encoded.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

No published scores collected for this benchmark yet.

Active Still spreads the field.