Agentic · Computer use · Multimodal
OSWorld-Verified
OSWorld-Verified
OSWorld with the broken tasks and flaky graders fixed.
- Released
- 2025
- Built by
- XLANG Lab (HKU), with contributions from OpenAI, Anthropic, ByteDance Seed TARS, MoonShot AI, Simular and Human Data
- Size
- 369 tasks
- Status
- Active
- Reported by
- XLANG Lab, Anthropic, OpenAI, Google DeepMind, ByteDance, Moonshot AI
- Judge
- 94
- Graded on the final state of the environment, not on prose.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 68
- 369 instances, so one item moves the score by 0.271 points.
- Fine print
- 62
- 3 documented caveats, the heaviest being harness sensitivity.
- Adoption
- 81
- Reported by 6 labs; 1 published score collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Ubuntu VM on AWS
The same Ubuntu desktop VM, but the infrastructure moved off VMware and Docker onto AWS, enabling roughly 50x parallelisation. A full evaluation that used to take more than ten hours now finishes in about an hour. That is not a cosmetic change: wall-clock and cost were a genuine barrier to running OSWorld at all, and cheap runs are what make repeated, lower-variance measurement possible.
Pixel-level control. Slow, flaky, and expensive to run at scale.
- Budget
- Still set by the harness, and still not standardised across reports.
- Tools exposed
- pyautogui.click(x, y)pyautogui.typewritepyautogui.hotkeypyautogui.scrollWAITFAILDONE
Where the tasks came from
The same 369 tasks as OSWorld; only setup scripts and evaluators were repaired.
OSWorld's task set, repaired against 300+ verified community bug reports across six fix categories.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Fixing 300+ issues removed failures that had nothing to do with agent capability, so Verified scores are systematically higher. No before/after per-task table was published, which means the size of the correction is not publicly quantified and historical numbers cannot be rebased — you can only discard them.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelCoACT-1unstated in source | Score60.76% | ConfigurationBest reported system at the OSWorld-Verified launch; XLANG characterised it as 84.4% of human capability. Step budget not stated in the launch post. | SourceThird-party2025-07-28 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.