the.ai

Evaluation / Methods

verified

Elo Rating

Rank things by who beats whom rather than by scoring each alone. Two models answer the same prompt, a person picks the better one, and each result nudges both ratings — the winner up, the loser down, by an amount that depends on how surprising the result was. It came from chess and it is now how frontier models are compared.

Viz primitive · budget-splitvotes-collected = 8

votes-collected holds 17% of the budget; rest holds the remaining 83%.

Votes behind a rating, against the uncertainty they have not yet removed. Drag the vote count up to watch the unresolved share shrink beside it — it shrinks with the square root, so an early rating is noisy in a way the ordering never shows.

8

Reviewed by opendroid · 2026-08-18

  • arXiv:2403.04132 — Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
  • arXiv:2306.05685 — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Origin · not linkable

  • Elo 1978 — The Rating of Chessplayers, Past and Present · Arco