History Problem Core Idea Evidence Results Impact Quiz Takeaways
Interactive Paper Explainer

The Attention Compromise
Grouped-Query Attention

Multi-head attention is quality but slow inference; multi-query is fast but degrades. GQA groups key-value heads between the two — and gets there from an existing checkpoint with 5% of the original pre-training compute.

Start Learning Read the Paper ↗
5%
Uptraining compute
MHA → GQA
From checkpoints
~1.7×
Typical decoding speedup
2023
Adopted widely
History

A Memory Problem in Disguise

Decoder inference bottlenecked on the KV cache — and the fixes clustered at two extremes.

2017
MHA — the default
Multi-head attention: h query heads, h key-value heads. Maximum expressivity; every head pays KV-cache rent at decode time.
2019
MQA — the radical cut
Shazeer's multi-query attention keeps a single shared key-value head: inference memory plummets, quality wobbles.
2020-22
Autoregressive scale-up
Chinchilla-class decoding + long contexts made KV-cache bandwidth the visible wall; MQA regained attention.
May 2023
🚀 GQA
Ainslie et al. generalize the spectrum: an intermediate number of KV head groups — MHA quality at near-MQA speed, uptrained in 5% of pre-compute.
2023+
The standard
Llama 2 70B, Mistral, Llama 3, DeepSeek, Gemma … GQA becomes the default decoder attention of open LLMs.
The KV-Cache Arithmetic

At decode time, every new token needs the keys and values of all previous tokens — cached per head. With 32 query heads and 32 KV heads, that is 32× the memory traffic of a 1-KV-head model; memory bandwidth, not FLOPs, sets tokens/sec. Cutting KV heads to 8 (4 query heads per group) shrinks the cache 4× while each query head still reads a distinct-enough key/value projection.

Chapter 01

Quality vs Speed

Two endpoints, no middle — until GQA filled the spectrum.

🐢
The MHA / MQA Dilemma
  • MHA: 32 query heads × 32 KV heads — best quality, but a KV cache that scales with head count
  • MQA: 1 shared KV head — near-1× cache, but measurable quality degradation on hard tasks
  • Training a fresh MQA model wastes the original pre-training investment
  • Decode-time memory bandwidth is the real bottleneck — a wall that grows with context length
🎯
The GQA Answer
  • Group query heads: each group of 4-8 query heads shares one key-value head
  • 8 KV heads for 32 query heads → 4× smaller cache, quality statistically tied with MHA
  • Uptraining recipe: convert an existing MHA checkpoint, then retrain with ~5% of original compute
  • The middle of the spectrum is the sweet spot — adopted as the default by the field
Analogy — The Shared Kitchen

MHA is every chef with a private pantry — stocked redundantly, expensive, fast per chef. MQA is one communal pantry for 32 chefs — cheap, congested at the shelf. GQA assigns 4-8 chefs per pantry: stock cost drops 4-8×, and no queue gets long enough to matter.

Chapter 02

The Spectrum, Parameterized

One number — group size — moves you from MHA quality to MQA speed.

MHA → GQA → MQA
  • MHA: 32 Q heads, 32 KV heads — 1:1 mapping, maximum diversity
  • GQA-8: 32 Q heads, 8 KV heads — 4:1 within-group sharing
  • GQA-4: 32 Q heads, 4 KV heads — 8:1 sharing, more speed
  • MQA: 32 Q heads, 1 KV head — total sharing, quality risk
The uptraining recipe
  • Take an existing MHA checkpoint; mean-pool its KV projections per group
  • Convert to GQA topology, then continue pre-training briefly
  • ~5% of original pre-training compute recovers full quality
  • Far cheaper than training a dedicated MQA model from scratch
Interactive Demo — Walk the Attention Spectrum

Tab from full MHA to MQA and watch the KV cache, quality, and speed trade off. GQA is the middle that won.

Chapter 03

Why the Paper Ends the Debate

Statistically careful evaluation plus the conversion recipe made GQA the low-risk default.

The Evidence

The paper runs paired significance tests across decoding-heavy benchmarks, comparing uptrained GQA, uptrained MQA, and the MHA originals. The finding: uptrained GQA matches MHA quality on statistically significant comparisons, while uptrained MQA degrades on a meaningful fraction. Speed-wise, both approach MQA's decoding throughput. The conversion of Llama-1 65B to GQA demonstrated the recipe on the most-watched open model of the day — which is precisely why Llama 2 70B shipped with GQA native.

Interactive Demo — The 5% Uptraining Recipe

Convert a trained MHA checkpoint into a GQA model without paying for pre-training again — the recipe that made adoption cheap.

Chapter 05

Quality ≈ MHA, Speed ≈ MQA

The compromise chart the field voted on with its architecture choices.

KV CACHE (32→8 heads)
4× smaller
memory bandwidth relief at decode time
QUALITY vs MHA
match
statistically tied on paired benchmarks
UPTRAINING COST
5%
of original pre-training compute
ADOPTION
default
Llama 2 70B, Mistral, Llama 3, Gemma, DeepSeek…
Interactive Demo — The KV-Cache Wall

Why this paper exists: press run and watch cache size scale with KV head count for a 32-query-head model at long context. Bandwidth, not FLOPs, is the wall.

Legacy

Legacy — The Default Head

A quiet infrastructure paper whose mechanism now ships inside most open LLMs.

🦙 Llama 2 70B and beyond
The lineage from Llama 2 70B through Mistral, Llama 3, Gemma, and DeepSeek ships GQA as standard decoder attention — the paper's table of contents became an industry convention.
💰 The conversion economics
Uptraining at 5% of pre-training compute converted an architecture upgrade into a cheap maintenance step — infrastructure habits (vLLM kernels, quantization formats) standardized on it.
🔗 Serving-stack synergy
Smaller KV caches multiply with PagedAttention (entry #25) and quantized caches: the three compose into modern serving throughput.
🧭 The spectrum mindset
GQA reframed MHA/MQA from a binary into a dial — later work (MLA in DeepSeek) continues the dial-turning philosophy.
⚠️ What it did NOT solve
GQA caps the cache, it does not eliminate it; very-long-context serving still needs paging/compression; and grouped sharing is a heuristic, not a learned structure.
🛤 Read next
The serving stack: PagedAttention · Llama 2 · FlashAttention-2
Test Yourself

Quick Quiz

Check your understanding of the key concepts from GQA.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ GQA interpolates between MHA (32 KV heads, best quality) and MQA (1 KV head, best speed).
✅ Fewer KV heads shrink the KV cache — decoding is memory-bandwidth-bound, so cache size is tokens/sec.
✅ Uptraining converts existing MHA checkpoints with ~5% of original pre-training compute.
✅ Statistically tied with MHA on paired benchmarks; near-MQA decode throughput.
✅ Shipped in Llama 2 70B, Mistral, Llama 3, Gemma, DeepSeek — the open-LLM default head.
✅ The deep lesson: at inference time, memory traffic — not FLOPs — sets the ceiling.