the.ai

Alignment / Regimes

verified

Instruction Tuning

A pretrained model continues text. Asked a question it may produce more questions, because that is what the training data often did next. Instruction tuning trains on pairs of instruction and good response until the model treats a prompt as something to answer rather than something to extend. The knowledge does not change; the default behaviour does.

Viz primitive · budget-splitsft = 1

sft holds 0% of the budget; rest holds the remaining 100%.

Share of total training compute spent on instruction tuning against pretraining. Drag the size to see how small it stays even as it decides how the model behaves.

1

Reviewed by opendroid · 2026-08-04