the.ai

Interpretability / Methods

verified

Activation Patching

To find out whether a component causes a behaviour, replace its activations with the ones it had on a different input and see whether the answer changes. It is the closest thing interpretability has to an experiment: not "this correlates with the output" but "change this and the output changes".

Viz primitive · budget-splitpatched-components = 4

patched-components holds 33% of the budget; rest holds the remaining 67%.

Components spliced back in from the clean run against those still carrying the corrupted ones, in components. Drag the patch up to watch the behaviour return — restoring it shows the set was enough, not that the model used it.

4

Reviewed by opendroid · 2026-08-18

  • arXiv:2202.05262 — Locating and Editing Factual Associations in GPT
  • arXiv:2211.00593 — Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small