Evaluation / Methods
verifiedElo Rating
Rank things by who beats whom rather than by scoring each alone. Two models answer the same prompt, a person picks the better one, and each result nudges both ratings — the winner up, the loser down, by an amount that depends on how surprising the result was. It came from chess and it is now how frontier models are compared.
It fits this problem well because pairwise preference is much easier for a person to give than an absolute score, and because it needs no fixed test set to be contaminated. What it inherits is chess's assumptions: that skill is one number, and that it does not depend on the opponent. Neither holds for models — one may be better at code and worse at prose, and a single rating averages that away, so two models with the same Elo can be reliably different on any particular task.
The expected score is a logistic function of the rating difference, so a 400-point gap means the stronger player is expected to score about 91%, and the update moves a rating in proportion to how far the result was from that expectation. The consequence worth knowing is that the confidence interval shrinks with the square root of games played, so early ratings on a new model are noisy in a way the leaderboard's ordering does not show — and a rank based on a few hundred votes is a much weaker claim than the same rank on a hundred thousand.
votes-collected holds 17% of the budget; rest holds the remaining 83%.
Votes behind a rating, against the uncertainty they have not yet removed. Drag the vote count up to watch the unresolved share shrink beside it — it shrinks with the square root, so an early rating is noisy in a way the ordering never shows.
Reviewed by opendroid · 2026-08-18
- arXiv:2403.04132 — Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- arXiv:2306.05685 — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Origin · not linkable
- Elo 1978 — The Rating of Chessplayers, Past and Present · Arco