Multimodal
GenEval
GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
Does the image contain what the prompt asked for? A detector checks.
- Released
- 2023
- Built by
- University of Washington and Allen Institute for AI
- Size
- 553 tasks
- Status
- Saturated
- Reported by
- Alibaba (Qwen)
- Judge
- 94
- Graded on the final state of the environment, not on prose.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 82
- 553 instances, so one item moves the score by 0.181 points.
- Fine print
- 52
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 21
- Reported by 1 lab; 3 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Offline generation, then detection
Images are generated offline and then scored automatically, with no human in the loop and no interaction. The condition that decides comparability is upstream and invisible in the score: whether a prompt-rewriting or prompt-expansion model was allowed to reformulate the prompt before generation. The benchmark neither forbids nor records this, so the same 553 prompts can measure a bare image model or an image model plus a trained rewriting agent.
Single turn, no tools. The most reproducible setup there is.
- Budget
- single pass, no human in the loop
Where the tasks came from
553 prompts, 4 images each, 2,212 generated images per evaluated model. Task sizes are uneven (80/99/80/94/100/100) but weighted equally in Overall.
Prompts authored against a closed vocabulary of MS COCO object classes, Berlin-Kay color terms, counts of 2-4 and four relative positions.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The best result, PromptRL at 0.97, puts a trainable prompt-refinement agent inside the RL loop, so it is not measuring the image model alone. The best verified score without rewriting is STAR-7B at 0.91. That is a six-point gap, larger than the spread across most of the leaderboard, and neither the benchmark nor most papers record which condition was used. Two GenEval numbers cannot be compared until both conditions are known.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelPromptRLAcademic | Score0.97% | Configurationobtaining scores of 0.97 on GenEval, 0.98 on OCR accuracy, and 24.05 on PickScore — a trainable LM prompt-refinement agent sits inside the RL loop, so this is not the image model alone | SourceLab-reported2026-02 |
| ModelSTAR-7BAcademic | Score0.91% | Configurationno prompt rewriting or prompt expansion (the paper contains no occurrence of 'rewrit'). Per task: single object 0.98, two object 0.94, counting 0.90, colors 0.92, position 0.91, attribute binding 0.80. | SourceLab-reported2025-12 |
| ModelGPT-4oOpenAI | Score0.84% | Configurationtabulated third-party in the STAR comparison table; the rewriting condition is not stated in the source, so this is not safely comparable to either row above | SourceThird-party2025-12 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Saturated — Scores are bunched at the ceiling; it no longer separates models.