Subquadratic architectures kept losing to attention because their parameters ignored the input. Mamba makes them input-dependent — selective — and suddenly content-aware reasoning runs at linear time.
Structured state spaces were efficient but content-blind — Mamba's diagnosis and fix.
An SSM compresses history into a running state: h'(t) = A·h(t) + B·x(t). Classical versions keep A, B, C fixed — the model updates its memory identically whether it's reading filler or a name that matters. Mamba's move: let B and C (and Δ, the step size) depend on x. Now the model can choose to write strongly, read strongly, or reset — content-based reasoning, the exact property that made attention win.
The diagnosis that reframed five years of subquadratic architecture research.
A fixed-parameter SSM is a photocopier — every page gets the same exposure; the news and the ad copy blur equally. Selection makes the model a journalist: some sentences go into the notebook verbatim (large Δ·B), some are skimmed, and yesterday's page can be torn out on a whim (gated reset). The notebook stays the same size — the judgment of what enters it became content-aware.
The one equation card that explains selection — and what it buys.
Selection gives an input-dependent gate at every position — recurrent in time, parallel across positions at training.
From the paper: parity-or-better with same-size Transformers, and wins precisely where fixed-parameter models failed.
The selection story is confirmed by ablation: synthetic tasks that require exact recall and copying (induction heads' home turf) are exactly the tasks where selection rescues the SSM — the mechanistic link between content-awareness and the attention-mirroring capability.
The two axes the paper moved simultaneously — the efficiency of recurrence with the selectivity of attention.
Mamba didn't replace attention; it joined it — and changed what hybrid architectures could contain.
Check your understanding of the key concepts from Mamba.
Everything you need to remember about this paper.