History Problem Pattern Circuit Phase Change Evidence Impact Quiz
Interactive Paper Explainer

The Circuit Behind In-Context Learning
Induction Heads

Mechanistic interpretability's landmark result — a two-head attention circuit that copies repeated patterns, forms suddenly during training, and may explain much of what we call in-context learning.

Start Learning Read the Paper ↗
2
Attention Heads in the Circuit
1
Phase Change During Training
1
Mechanism Hypothesis for ICL
2022
Year Published
History

From Probing to Circuits

Induction heads arrived at the turn of a decade-long shift in interpretability: from asking what models know to asking how they compute it.

2015 – 2019
Probing & saliency
Diagnostic classifiers and gradient maps ask "what does the model know?" — correlational evidence, no mechanism.
2020
BERTology peaks
Hundreds of probing papers dissect BERT's layers and neurons — knowledge without wiring diagrams.
2021
Transformer Circuits begins
Elhage et al., "A Mathematical Framework for Transformer Circuits" — decompose models into attention heads, define composition (Q, K, V), study attention-only toys.
2022 · Sep
🚀 Induction heads (Olsson et al.)
A two-head circuit for in-context learning, a phase change that reveals it mid-training, and five lines of evidence tying the two together.
2023 →
Features, circuits, attribution
Sparse autoencoders, feature dictionaries, attribution graphs — mechanistic interpretability goes mainstream.
Key Insight

A Transformer is not an opaque blob. It is a wiring diagram of attention heads that read from and write to a shared residual stream — so you can trace a specific computation, component by component. Induction heads were the first circuit traced end-to-end that plausibly implements something as important as in-context learning.

THE PATTERN THE CIRCUIT COMPLETES
[A] [B] … [A] → [B]
Amber = the repeated token A · Green = the prediction B, copied from context.
Chapter 01

The Problem — In-Context Learning Without a Mechanism

Few-shot prompting is the signature ability of large language models. Yet before 2022, nobody could point to the parts of the network that actually do it.

🕳️
The Mystery Box
  • In-context learning emerges from training — nobody designs it
  • Probing shows a model "knows" something, not how it computes it
  • Saliency highlights inputs; it doesn't identify mechanisms
  • No causal story: remove which components to switch ICL off?
  • Interpretability evidence was correlational, not mechanistic
🔌
The Circuits Approach
  • Decompose the Transformer into heads reading/writing a residual stream
  • Reverse-engineer one circuit end-to-end
  • Causal interventions: ablate heads, measure what breaks
  • Testable predictions — a phase change, universality
  • In-context learning gets a candidate mechanism
Analogy — Bills, Flickers, and Wires

Think of a trained language model as a city's power grid. Probing is reading the monthly bill — proof that power flows, silence on how. Saliency is watching windows light up from the street — suggestive, but indirect. Circuit analysis opens the panel, follows an actual wire, and pulls it out to see exactly which neighborhood goes dark. That last step — the ablation — is what turns "the model seems to learn in context" into "these two heads cause it."

Chapter 02

The Core Idea — the [A][B] … [A] → [B] Pattern

An induction head is a pattern-completer: when a token repeats in the context, it finds the earlier occurrence, checks what came right after it, and predicts that the same thing comes next again.

[ A ] [ B ]  …  [ A ]  →  [ B ]
[A]
Token A
Any token that already appeared earlier in the context.
[B]
What followed A
The token that came right after A's earlier occurrence.
… [A]
A repeats
Later in the same context, A shows up a second time.
→ [B]
Prediction
The head boosts B as the next token — copied, not reasoned.
Worked Example — "Mr Dursley"
… Mr Dursley woke early … Mrs Dursley → woke

The model isn't recalling a memorized sentence, and it isn't using grammar — "Mrs Dursley woke" is only plausible because the pattern is already in the context. The second Dursley triggers a search for the first one, and whatever followed it gets boosted as the prediction. In a real model this happens at the sub-word token level — Durs, ley — but the logic is identical.

Interactive Demo — The Sequence Completer

Pick a sequence, then press Show Attention to watch the two heads cooperate: the previous-token head labels each position with its predecessor, then the induction head uses that label to find — and copy — the answer.

Illustrative — tokens shown as whole words for readability; real models do this at the sub-word level, and individual heads are messier than this clean two-layer story. The mechanism is faithful to the paper's analysis.

Chapter 03

Two Heads, Two Layers — K-Composition

No single head can compute "find what followed the earlier A" — building that signal is itself an attention operation. So the circuit spans two layers and uses the residual stream as a message board.

The Circuit — a Wiring Diagram
Context:  the cat sat . the cat → ?
📥 Token "cat" — 2nd occurrence
Residual stream entering layer n: carries the current token identity, nothing about history yet.
↓
1 · Previous-token head — layer n
Attends one step back at every position; at "sat" it writes prev = "cat" into the key subspace.
↓
💬 Residual stream — the message board
The "sat" position now carries the key prev = "cat", readable by any later head.
↓
2 · Induction head — layer n+1
Query "cat" scans the keys, matches prev = "cat", attends to "sat" and reads its value.
↓
📤 Prediction: "sat"
The head's output boosts "sat" as the next token — [A][B] … [A] → [B] completes.
Why Two Layers?

The induction head's query ("cat") must meet a key that says "the token before me is cat." But building that key requires looking one step back — which is itself an attention operation. One head writes the label; a second head, in a later layer, reads it. Two sequential attention steps ⇒ at least two layers. That's why even tiny 2-layer attention-only models can host the full circuit.

What Is K-Composition?

In the Transformer Circuits framework, heads communicate through the residual stream. When head A's output lands in a subspace that head B's W_K projects — so A controls what B attends to — that's K-composition. (Q- and V-composition are the sibling cases.) Induction heads are the canonical example: the previous-token head's message is written precisely where the induction head's key computation reads it.

Interactive Demo — Circuit Stepper

Press play to light up the circuit one stage at a time: input → the previous-token head writes its label → the induction head matches and copies.

📥 Token "cat" — 2nd occurrence · enters layer n
↓
1 · Previous-token head · layer n — attends one step back, writes prev = "cat"
↓
💬 Residual stream — the "sat" position now carries the key prev = "cat"
↓
2 · Induction head · layer n+1 — query "cat" matches → attends to "sat"
↓
📤 Prediction "sat" — [A][B] … [A] → [B] completes
Chapter 04

The Phase Change — When Models Learn to Learn

Train a model long enough and something suddenly flips: in-context loss drops sharply, prefix-matching jumps, and induction heads appear. It is the paper's boldest observation.

What Happens

Midway through training — in small models, around a couple billion training tokens — the loss curve does something dramatic: a small bump as the circuit reorganizes, then a sharp dip as in-context performance lands. On a frozen evaluation of repeated-sequence prediction (the prefix-matching score), the model jumps from barely noticing repetition to exploiting it — and stays there for the rest of training.

Why It Matters

Ablate the induction heads in checkpoints from after the transition and in-context loss climbs back to pre-transition levels; ablate them before it and almost nothing changes — the circuit wasn't doing anything yet. Replaying the ablation across training checkpoints shows the same two heads switching in-context learning on, at one moment in time.

CIRCUIT SIZE
2
heads, spanning two layers
TRANSITION POINT
≈2.5B
training tokens, small models (illustrative)
EVIDENCE LINES
5
independent threads linking the circuit to ICL
MINIMUM SETUP
2L
attention-only models form them too
Interactive Demo — Loss Curve Explorer

A small model's training run (illustrative, shaped like the paper's curves). Watch the bump-dip at ≈2.5B tokens, then toggle the ablation to see what the post-transition checkpoints would score without their induction heads.

Curves are illustrative of the shapes reported in the paper (small models; the transition lands around a couple billion training tokens — later for larger models). Actual values vary by model and run.

Chapter 05

Five Lines of Evidence

The paper argues like a prosecutor: measure, intervene, time it, replicate. Each line is independent; together they make induction heads the leading suspect behind in-context learning.

📊 Prediction
On a frozen battery of models trained by many different groups, the prefix-matching score predicted in-context performance. Correlation — but across dozens of independently trained models.
⏱️ Timing
The phase change lands exactly when in-context ability appears during training: one moment the model can't use its context, the next it can — and induction heads form at that same step.
✂️ Intervention
Ablating the induction circuit damages tasks believed to need ICL — like in-context translation and letter-string pattern tasks — while other abilities are left comparatively untouched.
🌍 Universality
Similar two-head induction circuits show up across architectures and model families — different labs, different data, same mechanism. The circuit isn't an accident of one codebase.
🧪 Minimal Case
Even 2-layer attention-only transformers — no MLPs, the simplest models that can hold the circuit — form induction heads. Nothing exotic is required.
The Argument in One Line

Measure (prefix-matching predicts ICL across models) → intervene (ablations damage ICL) → time it (the phase change coincides with ICL appearing) → replicate (the circuit is universal). No single result is conclusive on its own; the convergence is the point.

Legacy

Impact — Mechanistic Interpretability Grows Up

Induction heads showed that a real, trained language model could be reverse-engineered into working parts — and became the field's canonical case study.

⚖️ Causality as the standard
Post-2022, interpretability claims increasingly need interventions — ablations, activation patching — not just probes and correlations.
🔬 The circuits agenda
The Transformer Circuits thread continued: attention superposition, toy models, and head-level analysis became a durable research program.
🧩 Features and dictionaries
Sparse autoencoders and feature dictionaries (2023+) scaled the "decompose the model" idea beyond individual heads.
🕸️ Attribution graphs
The 2024+ lineage — "On the Biology of a Large Language Model" — traces circuits through learned features: induction heads' direct descendants.
🤔 The hypothesis stays open
The paper is careful: at frontier scale, ICL may involve more than induction heads. The strongest evidence is for small and mid-size models.
📚 A teaching classic
Induction heads became the "hello world" of circuit analysis — the first worked example in most mechanistic-interpretability courses.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the induction heads paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Induction heads complete the pattern [A][B] … [A] → [B]: find a repeated token, predict what followed its earlier occurrence.
✅ The circuit is two heads across two layers: a previous-token head writes "what came before me"; an induction head reads it.
✅ They compose via K-composition — the early head writes into the subspace the later head's W_K projects.
✅ The circuit forms suddenly during training: a phase change where in-context loss drops and prefix-matching jumps (≈ a couple billion tokens in small models).
✅ Evidence: prefix-matching predicts ICL, ablations damage it, the timing matches, and the circuit appears across model families — even in 2-layer attention-only models.
✅ The ICL link at scale remains a carefully stated hypothesis — and the work launched mechanistic interpretability's circuits agenda.