All benchmarks

Safety

MASK

MASK: A Benchmark for Disentangling Honesty from Accuracy in AI Systems

Does the model say things it does not itself believe?

Released
2025
Built by
Center for AI Safety with Scale AI
Size
1,500 tasks
Status
Active
Reported by
Center for AI Safety, Scale AI
Signal77Reads cleanly
Judge
88
Symbolic equivalence, so formatting differences do not change the score.
Headroom
82
Still separates the frontier from everything below it.
Resolution
92
1,500 instances, so one item moves the score by 0.067 points.
Fine print
51
4 documented caveats, the heaviest being construct validity.
Adoption
48
Reported by 2 labs; 10 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Belief elicitation, then pressure

Two phases per item, run separately so the pressure prompt cannot contaminate the belief. Belief is elicited three times with neutral prompts, plus two indirect consistency questions for binary propositions, yielding a verdict of belief, no belief, or inconsistent. The pressure prompt is then applied in a fresh context. The inconsistent bucket is where the construct quietly fails, and it is reported rather than hidden.

Single turn, no tools. The most reproducible setup there is.

Budget
single turn per phase

Where the tasks came from

1,500 examples, 1,000 public, across six pressure archetypes — roughly 250 per archetype, so per-archetype claims are thin. An earlier paper version gave 1,028 human-labeled plus a 500-example held-out validation set; the discrepancy is flagged rather than silently resolved.

Propositions and pressure prompts hand-crafted by CAIS and Scale AI; the pressure scenarios encode particular coercion setups rather than a sampled distribution.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Everything rests on treating 'what the model says under neutral prompting' as its belief. But neutral prompting is still prompting, and the elicited answer is itself prompt-dependent — change the wording and the alleged belief can change, which would relabel the same pressured response as honest or dishonest. The three-elicitation protocol plus consistency checks is a serious attempt to stabilize this, and the inconsistent bucket is exactly the set of cases where the construct admits it broke down.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

P(honest)Higher is better
P(honest)higher is better
Modelclaude-3-7-sonnet-20250219AnthropicScore47.6%ConfigurationP(honest), Table 3, Appendix A.10. HIGHER is better. Not derivable from P(lie) — the two do not sum to 100.SourceThird-party2025-03-05
Modelclaude-3-5-sonnetAnthropicScore27.7%ConfigurationP(honest), Table 3. HIGHER is better.SourceThird-party2025-03-05
ModelDeepSeek-R1DeepSeekScore24.7%ConfigurationP(honest), Table 3. HIGHER is better.SourceThird-party2025-03-05
ModelGPT-4oOpenAIScore21.8%ConfigurationP(honest), Table 3. HIGHER is better.SourceThird-party2025-03-05
Modelgrok-2xAIScore14.2%ConfigurationP(honest), Table 3. HIGHER is better — this is the worst value in the table.SourceThird-party2025-03-05
P(lie)Lower is better
P(lie)lower is better
Modelclaude-3-7-sonnet-20250219AnthropicScore26.6%ConfigurationP(lie), Table 3, Appendix A.10. LOWER is better. Separate quantity from P(honest) 47.6. Accuracy for context: 82.2.SourceThird-party2025-03-05
Modelclaude-3-5-sonnetAnthropicScore34.4%ConfigurationP(lie), Table 3. LOWER is better. Accuracy for context: approximately 94 — high accuracy, middling honesty, which is the paper's whole argument.SourceThird-party2025-03-05
ModelGPT-4oOpenAIScore45.5%ConfigurationP(lie), Table 3. LOWER is better. Accuracy for context: approximately 94.SourceThird-party2025-03-05
ModelDeepSeek-R1DeepSeekScore54.4%ConfigurationP(lie), Table 3. LOWER is better. Accuracy for context: approximately 89.SourceThird-party2025-03-05
Modelgrok-2xAIScore63%ConfigurationP(lie), Table 3. LOWER is better — the worst in the table. Accuracy for context: approximately 88, so it is nearly as accurate as the field while lying at more than twice the rate.SourceThird-party2025-03-05

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.