How to train models too big for any single GPU: split attention and feed-forward matrices across devices with a handful of communication operators — pure PyTorch, no compiler, no framework surgery.
The 2019 memory wall, and the two escape routes the field was weighing.
A Transformer layer's memory lives in its big matrices — attention projections and the FFN. Megatron splits those matrices across GPUs (column for the first of each pair, row for the second) so each device stores a slice. The cost: exactly two communication points per layer — an all-reduce in the forward pass and one in the backward. No compiler, no library changes: 'a few communication operations in native PyTorch.'
The 2019 problem set: memory that wouldn't fit, parallelism that wouldn't compose, and frameworks that wouldn't cooperate.
Pipeline parallelism splits a factory into stations (layers) — complex scheduling, idle bubbles. Megatron instead puts multiple workers inside each station: one worker welds the left half of every panel, another the right, and the finished panel passes one handshake (all-reduce) before the next station. Same stations, same schedule — every station is just wider.
Column-parallel → row-parallel — the pairing that makes each matrix product need only one all-reduce.
What NVIDIA demonstrated in the paper — the numbers that made this the default.
The bonus finding — layer normalization placement — quietly influenced every later Transformer: the difference between stable and unstable multi-billion training was where you put the norm.
Efficiency without exotic tooling — that combination is what every lab copied.
Tensor parallelism is now a fixed ingredient; Megatron is its reference implementation.
Check your understanding of the key concepts from Megatron-LM.
Everything you need to remember about this paper.