Safety
AgentHarm
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Does the model refuse harmful tasks when it has tools, not just chat?
- Released
- 2024
- Built by
- UK AI Security Institute with Gray Swan AI
- Size
- 440 tasks
- Status
- Active
- Reported by
- UK AI Security Institute, Gray Swan AI
- Judge
- 58
- Mixed grading, so part of the score inherits the weakest judge.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 68
- 440 instances, so one item moves the score by 0.227 points.
- Fine print
- 50
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 26
- Reported by 2 labs; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Synthetic tools, no side effects
The model is given a set of mock tools — email, search, file operations, payment APIs — inside UK AISI's Inspect framework. Nothing leaves the harness, so a completed harmful task causes no real harm; the trace of tool calls is the artefact. This is the whole point of the benchmark: chat-only refusal evals ask whether a model will say something harmful, and this one asks whether it will do something harmful.
Another model plays the user, so its behaviour is part of the measurement.
- Budget
- multi-turn agentic
- Tools exposed
- emailweb searchfile operationspayment APIs(synthetic, per task)
Where the tasks came from
110 base behaviors augmented to 440. The public release is a subset: 44 of 66 public test behaviors (176 augmented) plus 8 of 11 validation behaviors (32 augmented), 468 rows total, plus 208 harmless_benign rows.
Written from scratch by UK AISI, Gray Swan AI, and academic collaborators; never scraped.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
A mock payment API always accepts the call; a real one has rate limits, KYC checks, fraud detection, and a human somewhere. AgentHarm measures whether the model is willing and able to take the harmful action inside a compliant sandbox, which is the right thing to isolate but is not the same as end-to-end real-world harm in either direction. It over-states harm where real defenses would stop the action, and under-states it where a real environment offers capabilities the mock tools do not.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Active — Still spreads the field.