FlashAttention made attention memory-linear; FlashAttention-2 makes it fast — better work partitioning across warps, parallelism along the sequence dimension, and 2× the speed at 50-73% of peak.
FlashAttention (entry #22) proved exact attention could be IO-aware; the sequel closed the gap to the hardware's ceiling.
The sequel's diagnosis was pure micro-architecture: thread blocks were parallelized only over batch×heads (not the sequence), leaving long sequences under-occupying the GPU; warps within a block redundantly loaded K and V through shared memory for different output tiles. FA2's fixes: (1) reduce non-matmul FLOPs (matmuls are what the tensor cores do well), (2) parallelize over sequence length, (3) split each block's work so warps split K/V — communication shrinks from shared-memory round-trips to register exchanges.
FA1's 25-40%-of-peak number was not a rounding error — it was a scheduling diagnosis.
FA1 was a brilliant cook with a bad floor plan — one station per table (block), staff (warps) all walking to the same pantry (shared memory). FA2 is the same recipes re-stationed: tables split along their length (sequence parallelism), each cook owns one pantry shelf (K/V slice), and dishes pass hand-to-hand (registers) instead of via the counter. Same food — twice the covers per night.
Each one is a scheduling decision; together they roughly double the kernel.
A GPU is a fixed fleet of streaming multiprocessors; a kernel 'occupying' 25% of them wastes 75% of every clock. FA1's grid mapped one block per (batch, head) pair — a 32k-token, 2-head forward pass used a handful of blocks on a device with 108 SMs. FA2's grid maps blocks along the sequence too: the same forward pass spreads across the whole machine. The math is unchanged; the schedule is not.
What changed downstream once attention stopped being the bottleneck.
Once attention ran at 50-73% of peak, the bottleneck moved: long-context training became memory-communication-bound (the PagedAttention era, entry #25) and the quality of long context became the research frontier (Lost in the Middle, entry #33). FA2 is also the kernel everyone else now measures against — a whole family of variants (sliding-window, causal-prefix, GQA-aware) fork from its scheduling skeleton. The paper's honesty is instructive: it openly frames FA1's shortfall as a scheduling failure — engineering candor that made the fix credible.
The efficiency story in one chart — from FA1's 25-40% to FA2's 50-73% of theoretical maximum.
FA2's scheduling skeleton is what 'fast attention' means across the ecosystem.
Check your understanding of the key concepts from FlashAttention-2.
Everything you need to remember about this paper.