the.ai

Alignment / Preference

verified

Reward Model

Nobody can write down a loss function for 'a helpful answer'. But people can reliably say which of two answers is better, and that comparison is cheap. A reward model is trained on those comparisons until it can score a response on its own — turning a judgement nobody can specify into a number an optimiser can use.

Viz primitive · budget-splitagreed-pairs = 40

agreed-pairs holds 59% of the budget; rest holds the remaining 41%.

Pairs the model ranks in the human's order against those it gets backwards. Drag the agreement up to watch the signal every later stage inherits get cleaner.

40

Reviewed by opendroid · 2026-08-04