History Problem Core Idea Evidence Results Impact Quiz Takeaways
Interactive Paper Explainer

State Spaces,
Selectively
Mamba

Subquadratic architectures kept losing to attention because their parameters ignored the input. Mamba makes them input-dependent — selective — and suddenly content-aware reasoning runs at linear time.

Start Learning Read the Paper ↗
5×
Inference throughput vs Transformer
1M
Sequence length scaled
0 attention
Or even MLP blocks
2023
SSM breakout
History

The Challenger Lineage

Structured state spaces were efficient but content-blind — Mamba's diagnosis and fix.

2021
S4 — structure arrives
Gu et al.'s structured state-space sequence model: continuous-time SSM with structured parameterization, strong long-range results — but parameters fixed across inputs.
2021-22
Linear attention & friends
A wave of subquadratic alternatives (linear attention, gated convolutions, SSM hybrids) — efficient, and consistently behind attention on language.
2023
The diagnosis
Gu & Dao: the weakness is content-blindness — fixed parameters can't selectively remember or forget based on what they're reading.
Dec 2023
🚀 Mamba
Selective SSMs (parameters are functions of the input) + a hardware-aware parallel scan + an attention-free, MLP-free block. 5× inference throughput; million-length scaling; a 3B Mamba matching Transformers 2× its size.
2024+
The hybrid era
Jamba, Zamba, Griffin, and Samba mix Mamba blocks with attention — pure replacements lost, but selective SSMs earned a permanent seat in the architecture toolbox.
Selection: the Missing Verb

An SSM compresses history into a running state: h'(t) = A·h(t) + B·x(t). Classical versions keep A, B, C fixed — the model updates its memory identically whether it's reading filler or a name that matters. Mamba's move: let B and C (and Δ, the step size) depend on x. Now the model can choose to write strongly, read strongly, or reset — content-based reasoning, the exact property that made attention win.

Chapter 01

Efficient but Content-Blind

The diagnosis that reframed five years of subquadratic architecture research.

🥁
The Fixed-Parameter Trap
  • SSMs, linear attention, gated convolutions: efficient recurrence, but parameters don't depend on the input
  • The state updates the same way for crucial tokens and for filler — no selective remembering or forgetting
  • Content-based reasoning (retrieval, induction, copying) is exactly where they lose to attention
  • And making parameters input-dependent breaks the convolution trick that made them fast to train
🎚
The Mamba Answer
  • Selection: Δ, B, C become functions of the input token
  • The convolution path becomes unusable — so a hardware-aware parallel scan (kernel-fusion, activation recomputation) recovers training speed
  • The block design simplifies: no attention, no MLP — the selective SSM IS the block
  • Result: content-aware memory at linear time — 5× inference throughput, scaling to million-token sequences
Analogy — The Photocopier vs the Journalist

A fixed-parameter SSM is a photocopier — every page gets the same exposure; the news and the ad copy blur equally. Selection makes the model a journalist: some sentences go into the notebook verbatim (large Δ·B), some are skimmed, and yesterday's page can be torn out on a whim (gated reset). The notebook stays the same size — the judgment of what enters it became content-aware.

Chapter 02

The Discretized View

The one equation card that explains selection — and what it buys.

The SSM, discretized
  • Continuous: h'(t) = A·h(t) + B·x(t);   y(t) = C·h(t)
  • Step size Δ discretizes: h_t = Ā·h_{t−1} + B̄·x_t, with Ā, B̄ derived from (Δ, A, B)
  • Mamba: Δ, B, C = f(x_t) — the gates now read the content
  • Large Δ → the current token dominates the state; small Δ → history persists

Selection gives an input-dependent gate at every position — recurrent in time, parallel across positions at training.

The hardware-aware scan
  • Input-dependent parameters kill the convolution formulation — so training uses a parallel scan over the recurrence
  • Kernel-fused + activation recomputation: the expanded state never materializes in HBM
  • Inference is a pure recurrence: O(1) per token, no KV cache at all
  • Throughput: 5× a Transformer's at batch-1-style generation
Interactive Demo — Fixed vs Selective Gates

Tab through the same sentence under fixed and input-dependent parameters — watch what each memory model retains.

Chapter 03

What the Benchmarks Said

From the paper: parity-or-better with same-size Transformers, and wins precisely where fixed-parameter models failed.

The Evidence Pattern

The selection story is confirmed by ablation: synthetic tasks that require exact recall and copying (induction heads' home turf) are exactly the tasks where selection rescues the SSM — the mechanistic link between content-awareness and the attention-mirroring capability.

Interactive Demo — Training a Recurrence in Parallel

How Mamba trains fast despite sequential dependencies — the scan that replaced the dead convolution path.

Chapter 05

Linear Time, Content-Aware

The two axes the paper moved simultaneously — the efficiency of recurrence with the selectivity of attention.

INFERENCE
5×
throughput vs Transformers — pure recurrence, no KV cache
QUALITY
2×-size parity
Mamba-3B matches Transformers twice its size
LENGTH
1M tokens
performance improving on real data at scale
BLOCK DESIGN
no attn/MLP
the selective SSM as the entire block
Interactive Demo — Generation Cost per Token

Press run for the per-token generation cost pattern — attention's cache grows with history; Mamba's state stays constant.

Legacy

Legacy — The Third Architecture Family

Mamba didn't replace attention; it joined it — and changed what hybrid architectures could contain.

🧬 The hybrid era
Jamba (Mamba+attention+MoE), Griffin/RecurrentGemma (linear recurrences + attention), Samba — every serious 2024+ architecture search includes selective SSM blocks in the candidate pool.
💾 Serving economics
O(1) per-token generation cost with no KV cache attacked the serving bill from a direction attention-side work (GQA, PagedAttention) could not — constant-memory generation became a real option.
🔬 The selectivity insight
'Input-dependent parameters fix content-blindness' generalized beyond SSMs — gating design became a first-class research topic across sequence architectures.
🧪 Long-sequence science
Million-token scaling on genomics and audio kept Mamba in domains where quadratic attention is truly infeasible — the paper's modality breadth mattered.
⚠️ What it did NOT solve
Exact recall over long contexts still favors attention (associative memory vs compressive state); pure-Mamba models lag on in-context copying; hybrid mixes cost complexity; and the hardware-aware scan needs bespoke kernels per vendor — portability remains engineering work.
🛤 Read next
The alternatives in tension: FlashAttention-2 · Induction Heads · DeepSeek-V3
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Mamba.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Selection = input-dependent Δ, B, C — the SSM learns what to write, read, and reset per token.
✅ Hardware-aware parallel scan replaces the dead convolution path — kernel-fused, log-depth training.
✅ Inference is pure recurrence: O(1) per token, no KV cache, 5× Transformer throughput.
Mamba-3B matches Transformers 2× its size; performance improves up to million-length sequences.
✅ The block is radically simple: no attention, no MLP — the selective SSM is the whole block.
Read it as the paper that made 'attention-free' a serious candidate — and hybrids the outcome.