Multi-head attention is quality but slow inference; multi-query is fast but degrades. GQA groups key-value heads between the two — and gets there from an existing checkpoint with 5% of the original pre-training compute.
Decoder inference bottlenecked on the KV cache — and the fixes clustered at two extremes.
At decode time, every new token needs the keys and values of all previous tokens — cached per head. With 32 query heads and 32 KV heads, that is 32× the memory traffic of a 1-KV-head model; memory bandwidth, not FLOPs, sets tokens/sec. Cutting KV heads to 8 (4 query heads per group) shrinks the cache 4× while each query head still reads a distinct-enough key/value projection.
Two endpoints, no middle — until GQA filled the spectrum.
MHA is every chef with a private pantry — stocked redundantly, expensive, fast per chef. MQA is one communal pantry for 32 chefs — cheap, congested at the shelf. GQA assigns 4-8 chefs per pantry: stock cost drops 4-8×, and no queue gets long enough to matter.
One number — group size — moves you from MHA quality to MQA speed.
Statistically careful evaluation plus the conversion recipe made GQA the low-risk default.
The paper runs paired significance tests across decoding-heavy benchmarks, comparing uptrained GQA, uptrained MQA, and the MHA originals. The finding: uptrained GQA matches MHA quality on statistically significant comparisons, while uptrained MQA degrades on a meaningful fraction. Speed-wise, both approach MQA's decoding throughput. The conversion of Llama-1 65B to GQA demonstrated the recipe on the most-watched open model of the day — which is precisely why Llama 2 70B shipped with GQA native.
The compromise chart the field voted on with its architecture choices.
A quiet infrastructure paper whose mechanism now ships inside most open LLMs.
Check your understanding of the key concepts from GQA.
Everything you need to remember about this paper.