the.ai

Evaluation / Methods

verified

LLM as Judge

Grading open-ended answers needs a human, and humans are slow and expensive. Using a strong model as the grader is fast enough to run on every change, and agrees with human raters about as often as two humans agree with each other. It is also biased in ways a human is not, and the biases are systematic rather than noisy.

Viz primitive · update-spectrumswaps = 1

8 values. The left group decays steeply; the right group is 75% of the way to flat, and reads flatter than the left.

Scores a judge gives the same pair in each presentation order. Drag to average over swapped orders and watch position bias cancel.

1

Reviewed by opendroid · 2026-08-04