the.ai

Alignment / Research

verified

Weak-to-Strong Generalization

Fine-tune a strong model on labels produced by a weak one. Naively the strong model should learn to imitate the weak one's mistakes and end up no better. It does not — it recovers a substantial part of the gap, apparently because the weak labels point at a capability the strong model already has rather than teaching it one.

Viz primitive · budget-splitgap-recovered = 12

gap-recovered holds 38% of the budget; rest holds the remaining 62%.

The performance gap weak supervision recovers, against the part that stays lost, in equal units. Drag the recovered share up to watch most of the gap close — reported values stop well short of the whole, and what remains is the supervisor's errors learned as the task.

12

Reviewed by opendroid · 2026-08-18

  • arXiv:2312.09390 — Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
  • arXiv:2211.03540 — Measuring Progress on Scalable Oversight for Large Language Models