A visual, step-by-step guide to the serving breakthrough that borrowed 60 years of operating-system wisdom — paging the KV cache like RAM — and built vLLM, the inference engine that made GPU serving cheap.
The ideas are older than deep learning: virtual memory, paging, and copy-on-write — applied to the KV cache.
Treat the KV cache like RAM: fixed-size physical blocks + a per-request page table + logical-to-physical mapping. Requests never need contiguous memory, fragmentation drops from catastrophic to <4%, and identical prefixes (a system prompt seen by a thousand users) become shared physical blocks.
Before PagedAttention, serving systems reserved contiguous KV memory per request — and paid for it three ways.
Contiguous reservation books every guest a 12-room suite in case their family shows up — most rooms stay dark. Paged allocation gives each guest bunk beds as needed, wherever they exist: a guest's party may sleep on three floors, tracked by a ledger. Occupancy soars; nobody cares about floors.
How one request's KV cache becomes a list of pointers — and why that's enough.
Follow one request generating tokens: new blocks appear anywhere, the table extends, memory never fragments.
PagedAttention is the kernel; vLLM is the serving system built around it.
Same GPUs, same latency budgets, more requests served — measured against the state-of-the-art systems of the day.
PagedAttention pulled LLM inference into the operating-systems tradition — and never left it.
The paper's technique is 60 years old. The lesson is why nobody applied it sooner — and what else it unlocks.
PagedAttention's meta-lesson: when a new field hits an old constraint, inventory the old solutions before inventing new ones. The KV cache's unpredictable growth, scarcity, and sharing needs are a re-run of 1961's virtual-memory problem — and the transfer cost was nearly zero once someone looked. The 2-4x that followed wasn't a cleverer model; it was a correct classification of the problem.
Check your understanding of the key concepts from the PagedAttention paper.
Everything you need to remember about this paper.