History Problem Core Idea Context Results Impact Quiz Takeaways
Interactive Paper Explainer

The Kernel, Retuned
FlashAttention-2

FlashAttention made attention memory-linear; FlashAttention-2 makes it fast — better work partitioning across warps, parallelism along the sequence dimension, and 2× the speed at 50-73% of peak.

Start Learning Read the Paper ↗
~2×
vs FlashAttention
50-73%
of peak A100
230
TFLOPs/s reached
4×
FA2 in GPT-2 era
History

From Exact to Optimal

FlashAttention (entry #22) proved exact attention could be IO-aware; the sequel closed the gap to the hardware's ceiling.

pre-2022
The quadratic tax
Attention materializes an N×N score matrix in HBM — runtime and memory grow quadratically with sequence length; long context priced accordingly.
May 2022
FlashAttention
Dao et al.: tiling + online softmax keeps attention exact while never writing the N×N matrix — linear memory, 2-4× runtime wins (entry #22).
the catch
Only 25-40% of peak
Despite the wins, FA1 left the GPU half-idle: suboptimal work partitioning between thread blocks and warps — low occupancy, wasteful shared-memory traffic.
Jul 2023
🚀 FlashAttention-2
Dao retunes the algorithm: fewer non-matmul FLOPs, parallelism over the sequence dimension even for a single head, warp specialization — ~2× speedup, 50-73% of peak, up to 230 TFLOPs/s.
2023+
The new floor
FA2 becomes the reference attention kernel (PyTorch SDPA, every serving stack); long-context training and 100k-context serving became budget line-items instead of moonshots.
Where the Idle Cycles Were

The sequel's diagnosis was pure micro-architecture: thread blocks were parallelized only over batch×heads (not the sequence), leaving long sequences under-occupying the GPU; warps within a block redundantly loaded K and V through shared memory for different output tiles. FA2's fixes: (1) reduce non-matmul FLOPs (matmuls are what the tensor cores do well), (2) parallelize over sequence length, (3) split each block's work so warps split K/V — communication shrinks from shared-memory round-trips to register exchanges.

Chapter 01

Exact but Underutilized

FA1's 25-40%-of-peak number was not a rounding error — it was a scheduling diagnosis.

📉
FA1's Three Inefficiencies
  • Parallelism only over batch and heads — long sequences, few heads leave the GPU under-occupied
  • All warps in a block redundantly read K/V from shared memory for their output tiles
  • Non-matmul operations (softmax rescaling, masking) consume cycles that tensor cores could be filling
📈
The FA2 Retune
  • Parallelize over the sequence dimension — even a single long-sequence head fills the machine
  • Warp specialization: each warp owns a K/V slice; outputs exchange through registers, not shared memory
  • Fewer non-matmul FLOPs per tile — the matmul:everything ratio climbs toward GEMM's
Analogy — The Kitchen Gets a Head Chef

FA1 was a brilliant cook with a bad floor plan — one station per table (block), staff (warps) all walking to the same pantry (shared memory). FA2 is the same recipes re-stationed: tables split along their length (sequence parallelism), each cook owns one pantry shelf (K/V slice), and dishes pass hand-to-hand (registers) instead of via the counter. Same food — twice the covers per night.

Chapter 02

The Three Changes

Each one is a scheduling decision; together they roughly double the kernel.

1️⃣ Fewer non-matmul FLOPs
Algorithm tweaked so a larger share of work rides the tensor cores — GEMM-class throughput is the target metric.
2️⃣ Sequence parallelism
Thread blocks parallelize over the sequence dimension, even for a single head — long contexts now fill every SM instead of a few.
3️⃣ Warp partitioning
K/V split across warps inside each block: partial results exchanged in registers — shared-memory reads/writes collapse.
📊 The result
~2× FlashAttention's speed; 50-73% of peak on A100s; up to 230 TFLOPs/s — and up to 4× end-to-end speedup training GPT-2-class models (seq 2k-16k) plus 3× on long-seq inference.
Why occupancy was the villain

A GPU is a fixed fleet of streaming multiprocessors; a kernel 'occupying' 25% of them wastes 75% of every clock. FA1's grid mapped one block per (batch, head) pair — a 32k-token, 2-head forward pass used a handful of blocks on a device with 108 SMs. FA2's grid maps blocks along the sequence too: the same forward pass spreads across the whole machine. The math is unchanged; the schedule is not.

What stays true from FA1
  • Exact attention — no approximation, identical outputs
  • Tiled online softmax — the N×N matrix is never materialized in HBM
  • IO-awareness: the memory-hierarchy insight remains the foundation
  • Same asymptotics: linear memory, quadratic compute — but compute at GEMM efficiency
Interactive Demo — FA1 vs FA2 Grids

Tab through the scheduling regimes — watch the GPU fill up when parallelism extends to the sequence dimension.

Chapter 03

The Field Guide Context

What changed downstream once attention stopped being the bottleneck.

The Downstream Story

Once attention ran at 50-73% of peak, the bottleneck moved: long-context training became memory-communication-bound (the PagedAttention era, entry #25) and the quality of long context became the research frontier (Lost in the Middle, entry #33). FA2 is also the kernel everyone else now measures against — a whole family of variants (sliding-window, causal-prefix, GQA-aware) fork from its scheduling skeleton. The paper's honesty is instructive: it openly frames FA1's shortfall as a scheduling failure — engineering candor that made the fix credible.

Interactive Demo — One Tile, FA2 Style

Follow one output tile through the retuned inner loop — matmul-heavy, register-local, shared-memory-light.

Chapter 05

The Peak-FLOPs Ladder

The efficiency story in one chart — from FA1's 25-40% to FA2's 50-73% of theoretical maximum.

SPEEDUP vs FA1
~2×
same math, re-scheduled
PEAK EFFICIENCY
50-73%
of theoretical max FLOPs/s on A100
TOP THROUGHPUT
230 TF/s
reached on attention alone
GPT-2-CLASS TRAINING
up to 4×
end-to-end training speedup (2k-16k ctx)
Interactive Demo — The Efficiency Ladder

Press run to see attention kernels' share of the A100's theoretical maximum — the ladder FA2 climbed from FA1's 25-40%.

Legacy

Legacy — The Reference Kernel

FA2's scheduling skeleton is what 'fast attention' means across the ecosystem.

⚡ The new floor
PyTorch SDPA, vLLM, TensorRT-LLM, JAX stacks — FA2-class scheduling became the default implementation of attention everywhere.
📏 Long context goes mainstream
At 50-73% of peak, 128K contexts train and serve on commodity budgets — the technical precondition for the long-context era (Llama 3.1, etc.).
🔬 Bottleneck migration
With attention fast, the frontier moved to KV-cache serving (entry #25), test-time compute (entry #58), and long-context quality (entry #33) — FA2 quietly redirected the research agenda.
🧬 The variant family
Sliding-window, block-sparse, GQA-aware, and paged variants all fork FA2's partitioning skeleton — a kernel design pattern, not just a kernel.
⚠️ What it did NOT solve
Attention remains quadratic in compute (FA2 fixes IO and scheduling, not asymptotics); exotic attention variants still need bespoke kernels; and gains shrink for short sequences where occupancy was never the issue.
🛤 Read next
Test Yourself

Quick Quiz

Check your understanding of the key concepts from FlashAttention-2.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ FA2 = FA1's exact math, re-scheduled: ~2× speed, 50-73% of peak A100, up to 230 TFLOPs/s.
✅ Three fixes: fewer non-matmul FLOPs, sequence-dimension parallelism, warp-level K/V partitioning.
✅ Occupancy was the villain — batch×heads-only grids starved long sequences.
✅ Register exchanges replace shared-memory traffic inside blocks.
Up to 4× end-to-end speedup on GPT-2-class training runs (2k-16k context).
Read it as the scheduling half of the long-context revolution — the enabler, not the headline.