Human preference · Coding
WebDev Arena
WebDev Arena, now Code Arena (Frontend, Fullstack, HTML, React and Image-to-WebDev boards)
Two models build the same web app; you use both and pick one.
- Released
- 2024
- Built by
- Arena Intelligence Inc. (formerly LMArena / LMSYS)
- Size
- Not published
- Status
- Active
- Reported by
- Anthropic, OpenAI, Google DeepMind, Alibaba (Qwen), Moonshot AI, DeepSeek, xAI, Z.ai
- Judge
- 62
- Human graders. Not reproducible run to run, and expensive to repeat.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 50
- Instance count is not published, so per-item weight is unknown.
- Fine print
- 47
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 97
- Reported by 8 labs; 7 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Sandboxed live preview
Both builds are streamed into a browser sandbox: CodeMirror 6 for the source view and a live, interactive preview the voter can actually click through. This is the distinguishing feature of the board — the artifact under judgement runs, so the vote blends functional correctness with design taste in a way no static coding benchmark reaches. Fullstack Code Arena, whose leaderboard went live 28 July 2026, extends the sandbox to databases, auth and login flows, third-party API keys, and live deployments.
Humans vote. Captures taste, inherits every bias of who showed up to vote.
- Budget
- multi-turn agentic
- Tools exposed
- create_fileedit_fileread_filerun_command
Where the tasks came from
579,848 votes across 115 models on the overall board as of Aug 15 2026. Sub-boards: Frontend 534,069 / 115; Fullstack 42,040 / 52; HTML 142,887 / 114; React 389,696 / 99; Image-to-WebDev 86,653 / 42. Launch in Dec 2024 collected over 80,000 votes.
Build prompts and votes both supplied by self-selected users of the arena; snapshots of every model action retained in Cloudflare R2.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The Code Arena payload carries hasStyleControl: false and serves the leaderboard id 'webdev-overall-raw', while arena.ai/leaderboard/text carries hasStyleControl: true and serves 'text-overall-style_control'. The page still renders a 'Remove Style Control' section header, but the underlying snapshot is the raw one. So the arena's own correction for verbosity and presentation — built because presentation was measurably inflating text scores — is absent from precisely the board where visual polish is hardest to separate from whether the app works.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| Modelclaude-opus-5-maxAnthropic | Score1692 | ConfigurationOverall WebDev board, style control NOT applied. Bounds [1682.1-1701.0] over 6,448 votes. | SourceThird-party2026-08-15 |
| Modelqwen3.8-maxAlibaba (Qwen) | Score1691 | ConfigurationFULLSTACK sub-board (42,040 votes, 52 models), style control NOT applied. Rank 1 here, rank 3 on the overall board. | SourceThird-party2026-08-15 |
| Modelclaude-opus-5-maxAnthropic | Score1678 | ConfigurationFULLSTACK sub-board, style control NOT applied. Rank 2 here, rank 1 on the overall, Frontend, HTML and React boards. | SourceThird-party2026-08-15 |
| Modelkimi-k3-maxMoonshot AI | Score1674 | ConfigurationOverall WebDev board, style control NOT applied. | SourceThird-party2026-08-15 |
| Modelqwen3.8-maxAlibaba (Qwen) | Score1667 | ConfigurationOverall WebDev board, style control NOT applied. The same model ranks FIRST on the Fullstack sub-board at 1691.3. | SourceThird-party2026-08-15 |
| Modelgrok-4.6-highxAI | Score1631 | ConfigurationOverall WebDev board, style control NOT applied. | SourceThird-party2026-08-15 |
| Modelgpt-5.6-sol-xhighOpenAI | Score1622 | ConfigurationOverall WebDev board, style control NOT applied. Run under the codex harness, which is a scaffold choice and not a property of the model. | SourceThird-party2026-08-15 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.