All benchmarks

Multimodal

GenEval

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Does the image contain what the prompt asked for? A detector checks.

Released
2023
Built by
University of Washington and Allen Institute for AI
Size
553 tasks
Status
Saturated
Reported by
Alibaba (Qwen)
Signal55Read with context
Judge
94
Graded on the final state of the environment, not on prose.
Headroom
8
Bunched at the ceiling. A win here no longer means much.
Resolution
82
553 instances, so one item moves the score by 0.181 points.
Fine print
52
4 documented caveats, the heaviest being construct validity.
Adoption
21
Reported by 1 lab; 3 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Offline generation, then detection

Images are generated offline and then scored automatically, with no human in the loop and no interaction. The condition that decides comparability is upstream and invisible in the score: whether a prompt-rewriting or prompt-expansion model was allowed to reformulate the prompt before generation. The benchmark neither forbids nor records this, so the same 553 prompts can measure a bare image model or an image model plus a trained rewriting agent.

Single turn, no tools. The most reproducible setup there is.

Budget
single pass, no human in the loop

Where the tasks came from

553 prompts, 4 images each, 2,212 generated images per evaluated model. Task sizes are uneven (80/99/80/94/100/100) but weighted equally in Overall.

Prompts authored against a closed vocabulary of MS COCO object classes, Berlin-Kay color terms, counts of 2-4 and four relative positions.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The best result, PromptRL at 0.97, puts a trainable prompt-refinement agent inside the RL loop, so it is not measuring the image model alone. The best verified score without rewriting is STAR-7B at 0.91. That is a six-point gap, larger than the spread across most of the leaderboard, and neither the benchmark nor most papers record which condition was used. Two GenEval numbers cannot be compared until both conditions are known.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelPromptRLAcademicScore0.97%Configurationobtaining scores of 0.97 on GenEval, 0.98 on OCR accuracy, and 24.05 on PickScore — a trainable LM prompt-refinement agent sits inside the RL loop, so this is not the image model aloneSourceLab-reported2026-02
ModelSTAR-7BAcademicScore0.91%Configurationno prompt rewriting or prompt expansion (the paper contains no occurrence of 'rewrit'). Per task: single object 0.98, two object 0.94, counting 0.90, colors 0.92, position 0.91, attribute binding 0.80.SourceLab-reported2025-12
ModelGPT-4oOpenAIScore0.84%Configurationtabulated third-party in the STAR comparison table; the rewriting condition is not stated in the source, so this is not safely comparable to either row aboveSourceThird-party2025-12

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Saturated Scores are bunched at the ceiling; it no longer separates models.