All benchmarks

Multimodal · Reasoning

CharXiv

CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Real messy charts scraped from arXiv papers, with free-form reasoning questions.

Released
2024
Built by
Princeton Language and Intelligence, with UW-Madison and HKU
Size
11,615 tasks
Status
Near ceiling
Reported by
Google DeepMind
Signal52Read with context
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
40
The top models crowd the ceiling; gaps are shrinking.
Resolution
100
11,615 instances, so one item moves the score by 0.009 points.
Fine print
45
5 documented caveats, the heaviest being possible training-set contamination.
Adoption
24
Reported by 1 lab; 4 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Zero-shot, one chart

Zero-shot, single turn, one chart image, no tools, with a set of natural instructions per question type. Generation and grading are separate steps run from the repo. The uncontrolled variable is image handling: results depend on the resolution the harness feeds the model and on how a 1024px chart is downsampled or tiled, which is a rendering choice rather than a reasoning capability.

Single turn, no tools. The most reproducible setup there is.

Budget
single turn

Where the tasks came from

2,323 charts carrying 9,292 descriptive and 2,323 reasoning questions. Validation is 1,000 charts (5,000 questions) with public answers; the 1,323-chart test split has its answers set to null in the repo, and no test leaderboard file exists, so all public numbers are validation numbers.

Hand-picked figures from arXiv papers dated Jan 2020 to Sep 2023, questions written and verified by human experts. CC-BY-SA-4.0.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Pinning gpt-4o-2024-05-13 is good for reproducibility in principle and bad in practice: the snapshot is subject to deprecation, every evaluation costs OpenAI spend, and anyone substituting a newer judge produces numbers that are not comparable to the leaderboard. Google DeepMind had to compute its own CharXiv numbers for all four models in its comparison because self-reported or official leaderboard numbers were unavailable.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGemini 3 ProGoogle DeepMindScore81.4%ConfigurationCharXiV Reasoning results are on 1000 reasoning questions from the validation split of CharXivSourceLab-reported2025-11
Modelo3 (high)OpenAIScore78.6%ConfigurationReasoning on the validation split; the same row scores 95.00 Descriptive, above the 92.10 human baseline. Top of the official leaderboard CSV, which is frozen around mid-2025.SourceCommunity2026-08-16
ModelGPT-5.1OpenAIScore69.5%ConfigurationReasoning, 1000 validation questions; computed by Google DeepMind, not self-reportedSourceThird-party2025-11
ModelClaude Sonnet 4.5AnthropicScore68.5%ConfigurationReasoning, 1000 validation questions; computed by Google DeepMind, not self-reportedSourceThird-party2025-11

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Near ceiling The top models are close enough to the ceiling to crowd.