All benchmarks

Coding · Agentic

SWE-Lancer

SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Real paid freelance tickets, scored in dollars the model could have earned.

Released
2025
Built by
OpenAI
Size
1,400 tasks
Status
Active
Reported by
OpenAI
Signal64Read with context
Judge
58
Mixed grading, so part of the score inherits the weakest judge.
Headroom
82
Still separates the frontier from everything below it.
Resolution
92
1,400 instances, so one item moves the score by 0.071 points.
Fine print
45
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
19
Reported by 1 lab; 2 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Offline app, three hours, 100 calls

Each run is an isolated Docker container on an Azure VM with 64GB of shared memory, holding the repository frozen at the parent commit of the real fix. There is no internet access at all — the actual fix, the PR discussion and Stack Overflow are simply unreachable, which is the primary contamination control. The model gets a terminal, file access, and a user tool that lets it build and exercise the running app across web, desktop and mobile to check its own work. But it is text-only: it cannot see the screenshots that tool produces, and it cannot watch the videos attached to many real issues. Hard budgets are three hours of wall clock and at most 100 tool calls, sampled at temperature 1.0, and it cannot ask a clarifying question the way the human freelancer could.

The model gets a shell. Scaffold quality and step budget move the score independently of the model.

Budget
3 hours wall clock and at most 100 tool calls
Tools exposed
terminalfile read/writeuser tool (run the app interactively)

Where the tasks came from

Over 1,400 tasks worth $1M in real payouts. The public Diamond subset everyone actually reports is 502 tasks worth $500,800 — 237 individual-contributor and 265 manager tasks.

Real resolved Upwork jobs from Expensify's open-source repository, each with its actual payout attached, screened by around 100 professional engineers.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • All 1,400+ tasks are Expensify, so what is measured is Expensify-shaped work. The paper concedes infrastructure and DevOps tasks are underrepresented and that freelance work differs structurally from full-time engineering. Because the repository is frozen at a 2024-era snapshot, it also ages as a capability probe rather than tracking current practice.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGPT-5.1-Codex-MaxOpenAIScore79.9%ConfigurationIC SWE Diamond subset, up from 66.3% for GPT-5.1-Codex. The system card shows the figure only in a chart, so this reaches us via secondary reporting.SourceLab-reported2025-11-01
ModelClaude 3.5 SonnetAnthropicScore26.2%ConfigurationIC SWE Diamond at launch — $58K of the $236,300 available in that lane. The same model scored 44.9% on SWE Manager Diamond, earning $403K of the full $1M across both lanes.SourceThird-party2025-02-01

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.