Human preference
LMArena
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference — now operated as Arena by Arena Intelligence Inc.
Anonymous head-to-head chats, voted on by the public, ranked by Elo.
- Released
- 2023
- Built by
- UC Berkeley / LMSYS, now Arena Intelligence Inc.
- Size
- Not published
- Status
- Active
- Reported by
- OpenAI, Anthropic, Google DeepMind, Meta, xAI, Alibaba (Qwen), Moonshot AI, DeepSeek
- Judge
- 62
- Human graders. Not reproducible run to run, and expensive to repeat.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 50
- Instance count is not published, so per-item weight is unknown.
- Fine print
- 50
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 95
- Reported by 8 labs; 6 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Live public chat window
The environment is a web chat box open to anyone, with no task specification, no tool set, no time limit, and no rubric. Conversations can be single-turn or multi-turn. Anonymity is the only real control in the whole apparatus, and it is doing an enormous amount of work: it is what stops the vote from measuring brand recognition instead of output quality.
Humans vote. Captures taste, inherits every bias of who showed up to vote.
- Budget
- unconstrained — single or multi-turn
Where the tasks came from
Live and growing: 7,779,985 votes on the text board as of the Aug 12 2026 update, against 240K in the founding paper. Platform-wide, Arena claims 700M+ conversations and 82M+ votes across all boards.
Prompts written live by self-selected members of the public; votes cast by those same people, with anomalous voters detected via per-vote p-values combined by Fisher's method and Bonferroni-corrected.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Bradley-Terry fits every model's strength jointly against every other model in the fit. There is no absolute anchor. When models are added, deprecated, or sampled at different rates, the whole solution shifts, so an Arena Score quoted in one month is not strictly the same quantity as the one quoted in another. This is also why the arena 'never saturates' — the scale rebases itself continuously rather than approaching a ceiling.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| Modelclaude-fable-5Anthropic | Score1506 | ConfigurationText board, style control applied. +/-5 interval over 21,533 votes — statistically tied with the next four entries. | SourceThird-party2026-08-12 |
| Modelclaude-opus-4-6-highAnthropic | Score1505 | ConfigurationText board, style control applied. +/-4 interval. | SourceThird-party2026-08-12 |
| Modelmuse-spark-1.2 xHighMeta | Score1498 | ConfigurationText board, style control applied. +/-10 interval — the widest in the top ten, so its rank range is correspondingly large. | SourceThird-party2026-08-12 |
| Modelclaude-opus-5-highAnthropic | Score1493 | ConfigurationText board, style control applied. +/-5 interval. | SourceThird-party2026-08-12 |
| Modelqwen3.8-maxAlibaba (Qwen) | Score1491 | ConfigurationText board, style control applied. +/-8 interval. | SourceThird-party2026-08-12 |
| Modelgemini-3.7-flash-highGoogle DeepMind | Score1490 | ConfigurationText board, style control applied. +/-8 interval — 16 points behind the leader, which the intervals do not separate. | SourceThird-party2026-08-12 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.