the.ai

Compression / Methods

verified

Structured Sparsity

Zeros only help if the hardware can skip them. Removing whole channels, heads or blocks produces a smaller dense model that runs faster on anything; removing scattered individual weights produces a model that is smaller on paper and no faster in practice. Structure is what converts compression into latency.

Viz primitive · budget-splitrealised-speedup = 6

realised-speedup holds 25% of the budget; rest holds the remaining 75%.

Compression that turns into latency against compression that stays on paper, in equal units. Drag the structure up to watch the saving become real — unstructured sparsity sits at the far left of this bar however high its zero count.

6

Reviewed by opendroid · 2026-08-18