the.ai

Interpretability / Foundations

verified

Mechanistic Interpretability

Most interpretability asks what a model attended to, or which inputs mattered. Mechanistic interpretability asks a harder question: what algorithm is the network actually running? The bet is that trained networks contain human-legible structure — circuits that do a specific job — and that finding it is difficult rather than hopeless.

Viz primitive · budget-splitexplained-components = 6

explained-components holds 50% of the budget; rest holds the remaining 50%.

Components a circuit account covers against those it leaves unexplained, in components. Drag the coverage up to watch the account grow over the model it is an account of.

6

Reviewed by opendroid · 2026-08-18