the.ai

Interpretability / Methods

verified

Feature Attribution

Which parts of the input were responsible for this output? Attribution methods answer with one number per input feature, and they are the oldest and most used family in interpretability. They are also the most easily misread: a heatmap over pixels or tokens looks like an explanation whether or not it is one.

Viz primitive · update-spectrumsmoothing-samples = 8 · bars = 16

16 values. The left group decays steeply; the right group is 15% of the way to flat, and reads flatter than the left.

Attribution mass across input tokens, raw and after averaging over noisy copies of the input. Drag the smoothing up to watch the spikes spread into their neighbours — calmer to look at, and no longer pointing at one token.

8

Reviewed by opendroid · 2026-08-18