Reasoning
IFEval
IFEval: Instruction-Following Eval for Large Language Models
Prompts with machine-checkable formatting constraints like word counts.
- Released
- 2023
- Built by
- Google and Yale University
- Size
- 541 tasks
- Status
- Saturated
- Reported by
- Google DeepMind, Meta, Mistral, Alibaba (Qwen), DeepSeek
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 82
- 541 instances, so one item moves the score by 0.185 points.
- Fine print
- 65
- 3 documented caveats, the heaviest being construct validity.
- Adoption
- 65
- Reported by 5 labs; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Zero-shot, single turn
One prompt, one response, no tools, no examples. The run conditions are about as simple as anything on this list, which is why the interesting variation lives entirely in the metric rather than in the environment: four different numbers can be computed from the same set of responses.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn
Where the tasks came from
541 prompts spanning only 25 constraint types, narrow enough that a model can be tuned directly against the target.
Hand-written by the authors at Google and Yale; 25 verifiable instruction types across nine categories. Fully public.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Prompt-level versus instruction-level, crossed with strict versus loose, yields four numbers from the same responses. Loose scoring can add several points over strict. A model card quoting a single IFEval figure without naming the variant is not stating a comparable result.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Saturated — Scores are bunched at the ceiling; it no longer separates models.