the.ai

Landscape / Foundations

verified

Sharpness

Two solutions with the same training loss can sit in very different neighbourhoods: one in a narrow valley where moving slightly costs a lot, another on a wide plateau where it barely matters. The wide one usually generalises better, and the reason is intuitive — a test distribution shifts the surface slightly, and a flat solution survives the shift.

Viz primitive · update-spectrumflatness = 0.2 · bars = 16

16 values. The left group decays steeply; the right group is 36% of the way to flat, and reads flatter than the left.

The Hessian's eigenvalues at a solution, as found and as flatness is sought. Drag the flatness up to watch the few enormous directions come down toward the bulk — that spread is what sharpness measures.

0.2

Reviewed by opendroid · 2026-08-18

  • arXiv:2010.01412 — Sharpness-Aware Minimization for Efficiently Improving Generalization
  • arXiv:1609.04836 — On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima