History Problem Core Idea Results Results Impact Quiz Takeaways
Interactive Paper Explainer

Slicing the Transformer
Megatron-LM

How to train models too big for any single GPU: split attention and feed-forward matrices across devices with a handful of communication operators — pure PyTorch, no compiler, no framework surgery.

Start Learning Read the Paper ↗
8.3B
Parameters trained
512
GPUs
15.1
PetaFLOPs sustained
76%
Scaling efficiency
History

When Models Outgrew Devices

The 2019 memory wall, and the two escape routes the field was weighing.

2018
The 1GB model era
BERT-Base (110M) trains on one GPU; 'large' means 340M parameters and a 16GB card.
Feb 2019
GPT-2 1.5B
The model that didn't fit anywhere consumer-grade — trained on private clusters, the first mainstream 'too big to ship' moment.
Jun 2019
Pipeline ideas
GPipe formalized pipeline parallelism: micro-batches through layer stages — powerful but complex to schedule and balance.
Sep 2019
🚀 Megatron-LM
NVIDIA's answer: intra-layer parallelism — split each layer's big matrices across GPUs, insert forward/backward all-reduces, keep everything else untouched.
2020+
The standard stack
Megatron-style tensor parallelism composes with ZeRO (entry #21) and pipelines into every serious large-scale trainer — Megatron-DeepSpeed, NeMo, internals of every frontier lab.
One Idea, Two Insertions

A Transformer layer's memory lives in its big matrices — attention projections and the FFN. Megatron splits those matrices across GPUs (column for the first of each pair, row for the second) so each device stores a slice. The cost: exactly two communication points per layer — an all-reduce in the forward pass and one in the backward. No compiler, no library changes: 'a few communication operations in native PyTorch.'

Chapter 01

One Model, Many Devices

The 2019 problem set: memory that wouldn't fit, parallelism that wouldn't compose, and frameworks that wouldn't cooperate.

🧱
The Memory Wall
  • Model states (weights, gradients, optimizer) exceeded any single device's memory past ~1B parameters
  • Data parallelism alone replicated full copies per GPU — no memory relief at all
  • Pipeline parallelism fragmented the model into stages needing careful scheduling and balance
  • Custom compilers/frameworks (Mesh-TensorFlow etc.) demanded code rewrites and rare expertise
⚔️
The Megatron Answer
  • Intra-layer model parallelism: slice attention and FFN matrices across devices
  • Column-parallel then row-parallel placement keeps each pair's math local until a single all-reduce
  • Orthogonal and complementary to pipeline parallelism and data parallelism — they stack
  • Native PyTorch implementation: a few comm operators, reproducible by any team
Analogy — The Factory Line, Inside the Station

Pipeline parallelism splits a factory into stations (layers) — complex scheduling, idle bubbles. Megatron instead puts multiple workers inside each station: one worker welds the left half of every panel, another the right, and the finished panel passes one handshake (all-reduce) before the next station. Same stations, same schedule — every station is just wider.

Chapter 02

The Slicing Recipe

Column-parallel → row-parallel — the pairing that makes each matrix product need only one all-reduce.

The attention block
  • Q, K, V projections: column-parallel — each GPU owns a slice of output heads
  • Attention math stays local per head slice — no communication
  • Output projection: row-parallel — each GPU holds a row slice
  • One all-reduce merges the partial outputs — exactly two per layer forward+backward
The FFN block
  • First FFN matrix (hidden → 4×hidden): column-parallel
  • GeLU/activation applied locally on each slice — the reason column comes first
  • Second FFN matrix (4×hidden → hidden): row-parallel
  • Same contract: local math, one all-reduce at the boundary
Interactive Demo — One Layer Through 8 GPUs

Follow a single Transformer layer's matrices as they are sliced, computed locally, and stitched — the entire Megatron mechanism in four steps.

Chapter 03

The Results Ledger

What NVIDIA demonstrated in the paper — the numbers that made this the default.

From the abstract

The bonus finding — layer normalization placement — quietly influenced every later Transformer: the difference between stable and unstable multi-billion training was where you put the norm.

Interactive Demo — Three Kinds of Parallelism

Tab through the 2019 options — and see why tensor parallelism became the composable middle layer of every training stack.

Chapter 05

15 PetaFLOPs, Plain PyTorch

Efficiency without exotic tooling — that combination is what every lab copied.

PARAMETERS
8.3B
GPT-2-style LM, converged on 512 GPUs
THROUGHPUT
15.1 PF
sustained across the whole application
SCALING
76%
efficiency vs strong single-GPU baseline
BERT SIDE
3.9B
BERT-style SOTA + the layer-norm placement lesson
Interactive Demo — The Scaling Curve

Press run for the parallel efficiency profile the paper reported — near-linear at 8 GPUs, still 76% at 512, against a single-GPU baseline sustaining 39 TeraFLOPs.

Legacy

Legacy — The Layer Beneath Every Trainer

Tensor parallelism is now a fixed ingredient; Megatron is its reference implementation.

🏗 The standard stack
Megatron tensor parallelism + ZeRO sharding (entry #21) + pipelines = the training architecture of essentially every large model since 2020 — often literally as Megatron-DeepSpeed.
🧩 Composability as design
'Orthogonal and complementary' is the paper's most-quoted property: parallelism layers you can stack independently became the engineering norm.
🔬 Accessible scale
Pure PyTorch with a few comm ops meant ordinary labs could train multi-billion models — the democratization moment of 2019-20.
📐 The pre-LN lesson
The layer-norm placement finding (critical as BERT-style models grow) reshaped architecture hygiene well beyond distributed training.
⚠️ What it did NOT solve
All-reduce cost grows with device count; the 76% number is against a 30%-of-peak baseline; and tensor parallelism alone still replicates optimizer states — the gaps ZeRO filled.
🛤 Read next
The companion pieces: ZeRO · FlashAttention · GPT-2
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Megatron-LM.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Intra-layer tensor parallelism: slice each layer's big matrices across GPUs; local math + one all-reduce per pair.
✅ Pure PyTorch: a few comm operators — no compilers, no framework forks.
✅ Demonstrated at 8.3B parameters on 512 GPUs: 15.1 PetaFLOPs sustained, 76% scaling efficiency.
✅ Orthogonal & complementary: stacks with data parallelism and pipelines — the composable middle layer.
✅ Column-then-row split order exists so elementwise activations stay device-local.
✅ The layer-norm placement finding is a durable architecture lesson hidden inside a systems paper.