All benchmarks

Safety

AgentHarm

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Does the model refuse harmful tasks when it has tools, not just chat?

Released
2024
Built by
UK AI Security Institute with Gray Swan AI
Size
440 tasks
Status
Active
Reported by
UK AI Security Institute, Gray Swan AI
Signal61Read with context
Judge
58
Mixed grading, so part of the score inherits the weakest judge.
Headroom
82
Still separates the frontier from everything below it.
Resolution
68
440 instances, so one item moves the score by 0.227 points.
Fine print
50
4 documented caveats, the heaviest being construct validity.
Adoption
26
Reported by 2 labs; 0 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Synthetic tools, no side effects

The model is given a set of mock tools — email, search, file operations, payment APIs — inside UK AISI's Inspect framework. Nothing leaves the harness, so a completed harmful task causes no real harm; the trace of tool calls is the artefact. This is the whole point of the benchmark: chat-only refusal evals ask whether a model will say something harmful, and this one asks whether it will do something harmful.

Another model plays the user, so its behaviour is part of the measurement.

Budget
multi-turn agentic
Tools exposed
emailweb searchfile operationspayment APIs(synthetic, per task)

Where the tasks came from

110 base behaviors augmented to 440. The public release is a subset: 44 of 66 public test behaviors (176 augmented) plus 8 of 11 validation behaviors (32 augmented), 468 rows total, plus 208 harmless_benign rows.

Written from scratch by UK AISI, Gray Swan AI, and academic collaborators; never scraped.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • A mock payment API always accepts the call; a real one has rate limits, KYC checks, fraud detection, and a human somewhere. AgentHarm measures whether the model is willing and able to take the harmful action inside a compliant sandbox, which is the right thing to isolate but is not the same as end-to-end real-world harm in either direction. It over-states harm where real defenses would stop the action, and under-states it where a real environment offers capabilities the mock tools do not.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

No published scores collected for this benchmark yet.

Active Still spreads the field.