History Problem Core Idea Verification Acceptance Results Impact Deep Dive Quiz
Interactive Paper Explainer

Draft Fast, Verify Exact
Speculative Decoding

A visual, step-by-step guide to the decoding trick that makes big models generate 2-3× faster with provably identical outputs — a small model guesses, the big model checks everything in parallel.

Start Learning Read the Paper ↗
2–3×
T5-XXL Speedup
Identical
Output Distribution
Parallel
Verification
2023
Year Published
History

An Old Architecture Trick, Reborn

Speculative execution built modern CPUs. The same bet — verify later, in bulk — rebuilt LLM decoding.

1990s
Speculative execution in CPUs
Branch predictors guess the control flow; the pipeline executes ahead and rolls back wrong paths. The trick that made clock speed matter.
2018–22
The autoregressive bottleneck
Decoding K tokens takes K serial runs of the model. Bigger models are slower per token — inference costs grow with quality.
2022 · Nov
🚀 Speculative Decoding (Leviathan, Kalman, Matias — Google)
A small draft model generates γ tokens; the large model verifies all of them in ONE forward pass, accepting the good ones — sampling math guarantees the output distribution is unchanged.
2023 →
The acceleration default
DeepMind's speculative sampling (Chen et al.), Medusa's self-drafting heads, EAGLE, and n-gram drafters — speculative decoding ships in every serious inference stack.
The One-Sentence Idea

Verification is parallel and generation is serial — so never generate directly with the big model. Let a cheap drafter guess several tokens, then let the expensive model check them all at once, accepting the prefix it agrees with and resampling where it doesn't. The math guarantees you always sample from exactly the big model's distribution.

🧭 Pairing
Memory-side speedups: PagedAttention. This page is the token-side story.
Chapter 01

The Serial Tax

Autoregressive decoding is a strict dependency chain: token k+1 cannot start until token k exists.

⛓
One Model, One Token, One Pass
  • K tokens = K serial forward passes of the full model
  • Each pass uses the model's full compute to produce ONE token — massive arithmetic under-utilization
  • Quality demands big models; big models are slow; latency budgets die
  • Parallelism inside a pass (batch, heads, sequence) can't break the cross-token dependency
✅
Guess, Then Check in Bulk
  • A small draft model proposes γ tokens (cheap — it runs γ fast passes)
  • The large model scores ALL γ proposals plus the next position in ONE parallel pass
  • Accepted tokens are kept; the first rejection is resampled from the correct conditional
  • Expected tokens per large pass: 1 + (accepted prefix length) — the chain accelerates without changing its distribution
Analogy — The Senior Architect

A senior architect (the large model) writing every word personally is slow. Instead, a junior drafter (the draft model) writes a paragraph; the architect reviews the whole paragraph in one read, keeps the sentences that are right, and rewrites from the first mistake. The signed document is architect-quality — the speed is junior-quality.

Chapter 02

Why Verification Is Free

The asymmetry that makes it work: scoring k positions in one pass costs roughly the same as scoring one.

draft: x₁…x_γ ~ q  ·  verify: one target pass over positions 1…γ+1  ·  accept while p agrees
γ
Draft length
How many tokens the drafter proposes per round (tuned, typically a handful).
1 pass
Verification cost
The target model's forward pass computes distributions at all γ+1 positions simultaneously.
α
Acceptance rate
Effective similarity between q and p — the paper's key parameter; α grows with draft quality.
E[tokens]
The yield
Expected accepted tokens per target pass — 1 + a geometric-ish sum in α and γ.
The Cost Model

Per round: γ cheap draft passes (cost ~γ·c_d) + one target pass (c_t). If the drafter is ~10-100× cheaper, draft cost is noise; the round costs ≈ one normal token. Whatever the round accepts beyond 1 token is pure speedup. Greedy? Simplify: accept token i iff argmax p = draft token — the outputs are then exactly the greedy decode.

The Two Sampling Modes
  • Greedy: deterministic accept/reject against the target argmax — trivially exact
  • Sampling: the modified rejection-sampling scheme — accept with probability min(1, p/q); on rejection, resample from the residual distribution norm(max(0, p−q)). Provable: the combined process samples exactly from p.
Chapter 03

A Round, Token by Token

Watch one speculative round execute: drafting, parallel verification, prefix acceptance, corrective resample.

Interactive Demo — The Speculative Round
Chapter 04

The α Dial

Everything reduces to one quantity: how often the draft model agrees with the target.

Interactive Demo — Acceptance & Yield Explorer

Set draft agreement α and draft length γ; see expected tokens per target pass and net speedup (assuming a 50× cheaper drafter).

📏 γ is tunable, α is earned
Draft length γ is a free parameter with an interior optimum; α comes from how well the drafter tracks the target — same tokenizer, similar training.
📉 Diminishing γ returns
Longer drafts raise acceptance risk: the chance all k tokens survive decays like α^k. Past the optimum, wasted draft tokens cost more than verification saves.
🎯 Hard tasks are easier
The paper's observation: harder language-modeling tasks have EASIER subtasks — drafts agree more where the text is formulaic (arithmetic templates, boilerplate) — so speculation pays most where latency hurts most.
Chapter 05

2–3× With Identical Outputs

The paper's evaluation: T5-XXL 11B, verified against the T5X implementation, outputs bit-consistent with standard decoding.

T5-XXL 11B
2–3×
acceleration vs standard T5X decoding, identical outputs
NO RETRAINING
0
changes to the target model — works on off-the-shelf weights
DRAFT COST
~T5-3B
a ~3B drafter (same tokenizer family) — negligible per-round overhead
EXACTNESS
proved
the accept/resample scheme provably samples from the target distribution p
Interactive Demo — Wall-Clock Race

Standard decoding vs speculative (α=0.8): same 40-token output, very different pass counts. Press play.

Legacy

Impact — The Free Lunch Everyone Ate

Provably-lossless acceleration became a standard layer of every inference stack — and spawned a research lineage.

🔀 DeepMind's sibling paper
Chen et al.'s "Accelerating LLM Decoding with Speculative Sampling" (2023) landed the same idea for Chinchilla — convergent discovery, shared vocabulary.
🦑 Medusa & self-drafting
Why keep a separate drafter? Medusa attaches multiple decoding heads to the target itself; EAGLE refines it further.
📚 N-gram & retrieval drafters
Drafts from n-gram caches or retrieval — no model at all — for hot-prefix serving workloads.
⚙️ Stack integration
vLLM, TensorRT-LLM and friends ship speculative modes next to paged KV and continuous batching — the three compose.
🧠 The α research agenda
Distilling drafters to maximize agreement became a named optimization target (draft-target alignment).
⚠️ What it did NOT solve
Latency of the FIRST token; batch-heavy throughput regimes (verification gains shrink); draft models for multimodal and MoE targets need care.
Deep Dive

Lossless vs Everything Else

Every other acceleration pays in quality. The deep reason this one doesn't: it never changes what is computed — only when.

⚖️
The Acceleration Marketplace
  • Quantization: cheaper math, some quality risk
  • Distillation: smaller model, permanent quality ceiling
  • Pruning/sparsity: fewer weights, distribution shift
  • Each buys speed by approximating the target — a price paid in outputs
🧬
The Speculative Exception
  • The TARGET model still computes every accepted token's distribution
  • Rejection sampling math: the final token stream is exactly p-distributed
  • Speedup comes from arithmetic idle-time (verify-parallel), not approximation
  • The drafter only proposes; it never decides — agreement or rollback, nothing in between
Interactive Demo — Where the Speedup Lives

One target pass produces distributions for ALL positions. See what's "wasted" in normal decoding vs speculation.

Verdict

Speculative decoding is best understood as latency arbitrage: the GPU could always verify k positions for the price of one — normal decoding simply never asked. Once the ask is made correctly (with exact rejection-sampling bookkeeping), 2-3× arrives without touching the model, the outputs, or the risk profile. The deeper lesson for system designers: audit your pipeline for capabilities the hardware already has but the protocol never uses.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the speculative decoding paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Autoregressive decoding is serial; verification is parallel — speculation converts one into the other.
✅ A cheap draft model proposes γ tokens; the target verifies all of them in ONE forward pass.
✅ Rejection-sampling acceptance provably preserves the target's output distribution — lossless.
✅ T5-XXL: 2-3× acceleration vs standard T5X decoding with identical outputs, no retraining.
✅ Yield is governed by α (draft agreement) and γ (draft length) — with an interior optimum in γ.
✅ Descendant lineage: speculative sampling, Medusa, EAGLE, n-gram drafters — now standard in serving stacks.