All benchmarks

Agentic · Computer use · Multimodal

ScreenSpot-Pro

ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Clicking tiny icons in CAD and video editors on 4K screens.

Released
2025
Built by
National University of Singapore and Hong Kong Baptist University
Size
1,581 tasks
Status
Active
Reported by
Google DeepMind, ByteDance, Alibaba (Qwen), OS-Atlas / UGround
Signal78Reads cleanly
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
82
Still separates the frontier from everything below it.
Resolution
92
1,581 instances, so one item moves the score by 0.063 points.
Fine print
59
3 documented caveats, the heaviest being construct validity.
Adoption
69
Reported by 4 labs; 6 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Static screenshots

There is no environment at all. Nothing is executed, no application state changes, and the model gets exactly one look at one image. That is deliberate: it isolates grounding — knowing where a control is — from planning, recovery and tool use. It is also why a model can do well here and still fail every OSWorld task, and why an OSWorld failure is often really a ScreenSpot-Pro failure in disguise.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn, one coordinate per instance.

Where the tasks came from

1,581 instances across 23 applications, three operating systems and five industry groups plus OS-level tasks.

Captured live from practitioners with 5+ years in each tool, double-reviewed, with bilingual English/Chinese instructions.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The paper's strongest existing baseline, OS-Atlas-7B, managed 18.9%, while the authors' training-free ScreenSeekeR pipeline reached 48.1% — a 254% relative gain achieved largely by zooming and searching rather than by better raw perception. Scores therefore conflate what the model sees with what the cropping strategy hands it, and model cards rarely say which scaffold was used.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGemini 3 ProGoogleScore72.7%ConfigurationClick accuracy on ScreenSpot-Pro; results as of November 2025. Grounding scaffold, if any, not stated in the card.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelScreenSeekeRScreenSpot-Pro authorsScore48.1%ConfigurationThe authors' training-free zoom-and-search pipeline; a 254% relative gain over the strongest existing baseline.SourceLab-reported2025-04
ModelClaude Sonnet 4.5AnthropicScore36.2%ConfigurationMeasured by Google for the Gemini 3 Pro model card comparison table; results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelOS-Atlas-7Bunstated in sourceScore18.9%ConfigurationStrongest pre-existing baseline as measured by the ScreenSpot-Pro authors.SourceThird-party2025-04
ModelGemini 2.5 ProGoogleScore11.4%ConfigurationClick accuracy on ScreenSpot-Pro; results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelGPT-5.1OpenAIScore3.5%ConfigurationMeasured by Google for the Gemini 3 Pro model card comparison table; results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.