the.ai

Alignment / Preference

verified

RLHF

Instruction tuning teaches a model to imitate good answers. RLHF goes further: it trains the model against a score for how good an answer is, so it can be pushed toward responses better than anything in the demonstration data. Humans rank outputs, a reward model learns to predict those rankings, and the policy is optimised against it.

Viz primitive · loss-curvesteps = 1000
loss
step 0dashed = held-out1000

Loss over 1000 training steps, starting near 6.0. It falls to about 2.15, with 89% of the total improvement arriving in the first half. A second line shows held-out, ending higher at about 2.69.

Reward-model loss over a policy run, against the same policy measured on held-out preferences. Drag to open the gap — it is the policy fitting the reward model rather than the preference.

0.4

Reviewed by opendroid · 2026-08-04