Safety · Knowledge
TruthfulQA
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Questions where the common human answer is wrong.
- Released
- 2021
- Built by
- University of Oxford and OpenAI
- Size
- 817 tasks
- Status
- Saturated
- Judge
- 58
- Mixed grading, so part of the score inherits the weakest judge.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 82
- 817 instances, so one item moves the score by 0.122 points.
- Fine print
- 47
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 8
- No lab reports it in a model card, so it rarely appears in a comparison.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Zero- or few-shot QA
No tools, no retrieval, no follow-up. The whole point is to probe what the model will assert from its own weights when the popular answer and the correct answer diverge.
Single turn, no tools. The most reproducible setup there is.
- Budget
- single turn
Where the tasks came from
817 questions across 38 categories, so a single item is worth 0.12%. Written adversarially against one 2021-era reference model.
Hand-written by the authors and filtered to questions a specific 2021 reference model answered falsely; public on GitHub since release.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
No maintained leaderboard exists. The archived HuggingFace Open LLM Leaderboard v1 contents parquet holds 7,260 rows with a maximum date of 2024-06-15 and has not moved since. Every 2026 frontier artifact checked — DeepSeek-V4-Pro, Kimi-K2.6, GLM-5.2, MiMo-V2.5-Pro, Qwen3.5-397B, Gemini 3 Pro and 3.1 Pro, Gemma 4, Qwen3-VL, Qwen3.5-Omni — reports TruthfulQA nowhere. It has been displaced by SimpleQA, MASK, and HalluLens. There is no frontier SOTA to quote and none should be invented.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelContamination/contaminated_proof_7b_v1.0community | Score82.29% | ConfigurationMC2 on the archived Open LLM Leaderboard v1, frozen 2024-06-15. Top of the board, and a deliberate contamination demonstration — listed here as evidence about the benchmark, NOT as a capability result. | SourceCommunity2024-06-15 |
| Modelcloudyu/Mixtral_7Bx2_MoE_DPOcommunity | Score81.5% | ConfigurationMC2 on the archived Open LLM Leaderboard v1. Highest non-contamination-demo entry — a 7Bx2 community merge, not a frontier model. | SourceCommunity2024-06-15 |
| ModelPalmyra X (43B)Writer | Score61.57% | ConfigurationHELM Classic v0.4.0 exact match, reported as 0.6157 — top of 67 models. A different metric from MC2 and not comparable to the 82.29 figure above. | SourceThird-party2024-01-01 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Saturated — Scores are bunched at the ceiling; it no longer separates models.