Agentic · Computer use · Multimodal
ScreenSpot-Pro
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
Clicking tiny icons in CAD and video editors on 4K screens.
- Released
- 2025
- Built by
- National University of Singapore and Hong Kong Baptist University
- Size
- 1,581 tasks
- Status
- Active
- Reported by
- Google DeepMind, ByteDance, Alibaba (Qwen), OS-Atlas / UGround
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 92
- 1,581 instances, so one item moves the score by 0.063 points.
- Fine print
- 59
- 3 documented caveats, the heaviest being construct validity.
- Adoption
- 69
- Reported by 4 labs; 6 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Static screenshots
There is no environment at all. Nothing is executed, no application state changes, and the model gets exactly one look at one image. That is deliberate: it isolates grounding — knowing where a control is — from planning, recovery and tool use. It is also why a model can do well here and still fail every OSWorld task, and why an OSWorld failure is often really a ScreenSpot-Pro failure in disguise.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn, one coordinate per instance.
Where the tasks came from
1,581 instances across 23 applications, three operating systems and five industry groups plus OS-level tasks.
Captured live from practitioners with 5+ years in each tool, double-reviewed, with bilingual English/Chinese instructions.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The paper's strongest existing baseline, OS-Atlas-7B, managed 18.9%, while the authors' training-free ScreenSeekeR pipeline reached 48.1% — a 254% relative gain achieved largely by zooming and searching rather than by better raw perception. Scores therefore conflate what the model sees with what the cropping strategy hands it, and model cards rarely say which scaffold was used.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGemini 3 ProGoogle | Score72.7% | ConfigurationClick accuracy on ScreenSpot-Pro; results as of November 2025. Grounding scaffold, if any, not stated in the card. | SourceLab-reported2025-11 · Gemini 3 Pro model card |
| ModelScreenSeekeRScreenSpot-Pro authors | Score48.1% | ConfigurationThe authors' training-free zoom-and-search pipeline; a 254% relative gain over the strongest existing baseline. | SourceLab-reported2025-04 |
| ModelClaude Sonnet 4.5Anthropic | Score36.2% | ConfigurationMeasured by Google for the Gemini 3 Pro model card comparison table; results as of November 2025. | SourceLab-reported2025-11 · Gemini 3 Pro model card |
| ModelOS-Atlas-7Bunstated in source | Score18.9% | ConfigurationStrongest pre-existing baseline as measured by the ScreenSpot-Pro authors. | SourceThird-party2025-04 |
| ModelGemini 2.5 ProGoogle | Score11.4% | ConfigurationClick accuracy on ScreenSpot-Pro; results as of November 2025. | SourceLab-reported2025-11 · Gemini 3 Pro model card |
| ModelGPT-5.1OpenAI | Score3.5% | ConfigurationMeasured by Google for the Gemini 3 Pro model card comparison table; results as of November 2025. | SourceLab-reported2025-11 · Gemini 3 Pro model card |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.