the.ai

Compression / Methods

verified

Knowledge Distillation

Train a small model to copy a large one's outputs rather than the original labels. The large model's full distribution over classes says more than a single correct answer does — that this dog is somewhat wolf-like and not at all a truck — and the small model learns faster from that than from the label alone.

Viz primitive · update-spectrumtemperature = 0.2 · bars = 16

16 values. The left group decays steeply; the right group is 36% of the way to flat, and reads flatter than the left.

The teacher's distribution over classes, as produced and as softened. Drag the temperature up to watch the small probabilities rise into view — those are the signal the student learns from, and at temperature one they are invisible.

0.2

Reviewed by opendroid · 2026-08-18