Mechanistic interpretability's landmark result — a two-head attention circuit that copies repeated patterns, forms suddenly during training, and may explain much of what we call in-context learning.
Induction heads arrived at the turn of a decade-long shift in interpretability: from asking what models know to asking how they compute it.
A Transformer is not an opaque blob. It is a wiring diagram of attention heads that read from and write to a shared residual stream — so you can trace a specific computation, component by component. Induction heads were the first circuit traced end-to-end that plausibly implements something as important as in-context learning.
Few-shot prompting is the signature ability of large language models. Yet before 2022, nobody could point to the parts of the network that actually do it.
Think of a trained language model as a city's power grid. Probing is reading the monthly bill — proof that power flows, silence on how. Saliency is watching windows light up from the street — suggestive, but indirect. Circuit analysis opens the panel, follows an actual wire, and pulls it out to see exactly which neighborhood goes dark. That last step — the ablation — is what turns "the model seems to learn in context" into "these two heads cause it."
An induction head is a pattern-completer: when a token repeats in the context, it finds the earlier occurrence, checks what came right after it, and predicts that the same thing comes next again.
The model isn't recalling a memorized sentence, and it isn't using grammar — "Mrs Dursley woke" is only plausible because the pattern is already in the context. The second Dursley triggers a search for the first one, and whatever followed it gets boosted as the prediction. In a real model this happens at the sub-word token level — Durs, ley — but the logic is identical.
No single head can compute "find what followed the earlier A" — building that signal is itself an attention operation. So the circuit spans two layers and uses the residual stream as a message board.
The induction head's query ("cat") must meet a key that says "the token before me is cat." But building that key requires looking one step back — which is itself an attention operation. One head writes the label; a second head, in a later layer, reads it. Two sequential attention steps ⇒ at least two layers. That's why even tiny 2-layer attention-only models can host the full circuit.
In the Transformer Circuits framework, heads communicate through the residual stream. When head A's output lands in a subspace that head B's W_K projects — so A controls what B attends to — that's K-composition. (Q- and V-composition are the sibling cases.) Induction heads are the canonical example: the previous-token head's message is written precisely where the induction head's key computation reads it.
Train a model long enough and something suddenly flips: in-context loss drops sharply, prefix-matching jumps, and induction heads appear. It is the paper's boldest observation.
Midway through training — in small models, around a couple billion training tokens — the loss curve does something dramatic: a small bump as the circuit reorganizes, then a sharp dip as in-context performance lands. On a frozen evaluation of repeated-sequence prediction (the prefix-matching score), the model jumps from barely noticing repetition to exploiting it — and stays there for the rest of training.
Ablate the induction heads in checkpoints from after the transition and in-context loss climbs back to pre-transition levels; ablate them before it and almost nothing changes — the circuit wasn't doing anything yet. Replaying the ablation across training checkpoints shows the same two heads switching in-context learning on, at one moment in time.
The paper argues like a prosecutor: measure, intervene, time it, replicate. Each line is independent; together they make induction heads the leading suspect behind in-context learning.
Measure (prefix-matching predicts ICL across models) → intervene (ablations damage ICL) → time it (the phase change coincides with ICL appearing) → replicate (the circuit is universal). No single result is conclusive on its own; the convergence is the point.
Induction heads showed that a real, trained language model could be reverse-engineered into working parts — and became the field's canonical case study.
Check your understanding of the key concepts from the induction heads paper.
Everything you need to remember about this paper.