A visual, step-by-step guide to the decoding trick that makes big models generate 2-3× faster with provably identical outputs — a small model guesses, the big model checks everything in parallel.
Speculative execution built modern CPUs. The same bet — verify later, in bulk — rebuilt LLM decoding.
Verification is parallel and generation is serial — so never generate directly with the big model. Let a cheap drafter guess several tokens, then let the expensive model check them all at once, accepting the prefix it agrees with and resampling where it doesn't. The math guarantees you always sample from exactly the big model's distribution.
Autoregressive decoding is a strict dependency chain: token k+1 cannot start until token k exists.
A senior architect (the large model) writing every word personally is slow. Instead, a junior drafter (the draft model) writes a paragraph; the architect reviews the whole paragraph in one read, keeps the sentences that are right, and rewrites from the first mistake. The signed document is architect-quality — the speed is junior-quality.
The asymmetry that makes it work: scoring k positions in one pass costs roughly the same as scoring one.
Per round: γ cheap draft passes (cost ~γ·c_d) + one target pass (c_t). If the drafter is ~10-100× cheaper, draft cost is noise; the round costs ≈ one normal token. Whatever the round accepts beyond 1 token is pure speedup. Greedy? Simplify: accept token i iff argmax p = draft token — the outputs are then exactly the greedy decode.
Watch one speculative round execute: drafting, parallel verification, prefix acceptance, corrective resample.
Everything reduces to one quantity: how often the draft model agrees with the target.
The paper's evaluation: T5-XXL 11B, verified against the T5X implementation, outputs bit-consistent with standard decoding.
Provably-lossless acceleration became a standard layer of every inference stack — and spawned a research lineage.
Every other acceleration pays in quality. The deep reason this one doesn't: it never changes what is computed — only when.
Speculative decoding is best understood as latency arbitrage: the GPU could always verify k positions for the price of one — normal decoding simply never asked. Once the ask is made correctly (with exact rejection-sampling bookkeeping), 2-3× arrives without touching the model, the outputs, or the risk profile. The deeper lesson for system designers: audit your pipeline for capabilities the hardware already has but the protocol never uses.
Check your understanding of the key concepts from the speculative decoding paper.
Everything you need to remember about this paper.