History Problem Core Idea Results Results Impact Quiz Takeaways
Interactive Paper Explainer

Data Parallelism
Without the Copies
ZeRO

Data parallelism replicates everything; model parallelism fragments compute. ZeRO keeps data-parallel training's simplicity while sharding the memory — optimizer states, gradients, then parameters — across devices.

Start Learning Read the Paper ↗
100B+
Parameters trained
400
GPUs
15
Petaflops
10×
vs prior SOTA
History

The Redundancy Audit

The paper starts with an accounting question: in data-parallel training, what is actually stored per GPU — and how much of it is a copy?

2012-18
Data parallelism's reign
Different batch per GPU, all-reduce the gradients — simple, fast, and every GPU stores a complete model + optimizer replica.
2019
The mixed-precision era
Adam + FP16 training means optimizer states (momenta, variances, master weights) balloon to ~3× model size — redundancy now dominates memory.
Sep 2019
Megatron's lane
Tensor parallelism slices the model — effective, but adds comm complexity per layer and demands code surgery (entry #20).
Oct 2019
🚀 ZeRO
Rajbhandari et al. (DeepSpeed): keep data parallelism, remove its redundancy — partition optimizer state, gradients, and parameters across ranks with partition-and-all-gather.
2020+
DeepSpeed everywhere
ZeRO powers GPT-3-scale training, the ZeRO-Infinity offload lineage, and becomes the default 'just make it fit' trainer for the open ecosystem.
The 16 Bytes per Parameter Audit

With mixed-precision Adam, every parameter costs 16 bytes: 2 for the FP16 weight, 4 for the FP32 master copy, 8 for the two FP32 optimizer moments. Add 2+4 for the gradient and its FP32 copy — a 7B model's training state is ~112GB before activations. Data parallelism multiplies that by the GPU count. ZeRO's insight: none of it needs to be replicated — partition it, and communicate exactly what each rank needs, when it needs it.

Chapter 01

Trained Everywhere, Stored Everywhere

The redundancy problem hiding inside the field's favorite parallelism strategy.

🖨
The Replication Trap
  • Data parallelism replicates parameters, gradients, AND optimizer states on every GPU
  • A 100B model under mixed-precision Adam needs ~1.6TB of replicated state — physically impossible on 2019 clusters
  • Model parallelism fixes memory but complicates code and adds per-layer communication
  • Result: an artificial ceiling — big models were a systems problem, not a research problem
✂️
The ZeRO Answer
  • ZeRO-1: partition optimizer states across ranks (4× memory cut, near-zero comm overhead)
  • ZeRO-2: also partition gradients (8× cut, still data-parallel throughput)
  • ZeRO-3: also partition parameters (memory scales with #devices, Nd× cut)
  • 100B+ parameters trained with super-linear speedup — and up to 13B models need NO model parallelism at all
Analogy — The Shared Notebook

Classic data parallelism: every student in the class has a complete copy of the lab notebook — identical notes, identical backups, 40× the paper. ZeRO: the class keeps one notebook, torn into sections — each student holds a section, passes pages when someone needs them, and everyone still does their own experiments. Nothing is copied; everything is reachable.

Chapter 02

The Three Stages

Each stage shards one more memory class — with a careful eye on the communication bill.

Memory per parameter (mixed-precision Adam)
  • FP16 weight — 2 bytes · kept (ZeRO-3 shards it)
  • FP32 master weight — 4 bytes · optimizer state
  • Adam momentum + variance — 8 bytes · optimizer state
  • FP16 gradient + FP32 copy — 6 bytes · gradient state

Total ≈ 16-20 bytes per parameter — replicated per GPU under naive data parallelism.

What each stage shards
  • ZeRO-1 (Pos): optimizer states partitioned — ~4× memory reduction, communication volume essentially unchanged
  • ZeRO-2 (Pos+g): + gradients — ~8× reduction, still within data-parallel comm budget
  • ZeRO-3 (Pos+g+p): + parameters — memory ∝ 1/Nd GPUs, +50% comm: all-gather before forward, reduce-scatter after backward
Interactive Demo — The Memory Stacks

Tab through baseline data parallelism and the three ZeRO stages — watch the per-GPU memory bill collapse while the communication bill barely moves.

Chapter 03

The Claims, Verified

From the abstract: the 100B run, the throughput, and the usability bonus.

The Results Ledger

The usability clause is the sleeper hit: researchers who could never restructure code for pipeline/tensor parallelism could now train 13B models with an import — DeepSpeed became the open ecosystem's default trainer.

Interactive Demo — One Parameter's Footprint

Slide through the memory line-items one byte-class at a time — the audit that motivated the whole paper.

Byte accounting from the paper's memory analysis (mixed-precision Adam). ZeRO's whole premise: every replicated byte here is a sharding opportunity.
Chapter 05

Super-Linear Scaling

Sharding should cost communication; ZeRO's accounting showed the trade was overwhelmingly profitable.

MODEL SIZE
100B+
trained on 400 GPUs, super-linear speedup
THROUGHPUT
15 PF
sustained across the system
vs SOTA 2019
10×
performance; 8× model size
USABILITY
13B
with zero model parallelism — plain DP code
Interactive Demo — One Training Step Under ZeRO-3

Follow a single optimizer step through partitioned forward, backward, and weight update — the machinery behind the 100B run.

StageShardedMemory reductionExtra communicationPractical ceiling
Baseline DPnothing1×grad all-reduce~1-4B per node
ZeRO-1 (P_os)optimizer states~4×≈ none11B-class single node
ZeRO-2 (P_os+g)+ gradients~8×≈ nonelarger, still DP-simple
ZeRO-3 (P_os+g+p)+ parameters∝ 1/Nd+50% (all-gather / reduce-scatter)100B+ / trillion-class

Reduction factors from the paper's memory analysis for mixed-precision Adam training. The stage ladder is the API surface users actually see in DeepSpeed today.

Legacy

Legacy — The Default Trainer

ZeRO became DeepSpeed, and DeepSpeed became how the open world trains big models.

🚀 DeepSpeed as infrastructure
ZeRO's implementation became Microsoft's DeepSpeed — the training backend for GPT-3-class and beyond, and the default answer to 'my model doesn't fit'.
🧩 Composability with Megatron
ZeRO-1/2 stack under Megatron tensor parallelism (entry #20) — the combined Megatron-DeepSpeed recipe trained the open era's largest models.
📈 The trillion-parameter claim
The analysis that ~1T parameters was reachable on 2019 hardware reframed what counted as a systems limit — later ZeRO-Infinity extended the ladder to CPU/NVMe offload.
🎓 The 13B accessibility clause
Model-parallelism-free training up to 13B meant ordinary research code could suddenly host serious models — a quiet democratization win.
⚠️ What it did NOT solve
Activation memory still grows with batch and sequence (checkpointing needed); ZeRO-3's all-gather latency hurts small-batch regimes; and sharding multiplies failure domains — checkpointing/resilience became its own research area.
🛤 Read next
The system stack: Megatron-LM · FlashAttention · Scaling Laws
Test Yourself

Quick Quiz

Check your understanding of the key concepts from ZeRO.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Data parallelism replicated everything; ZeRO shards optimizer states (4×), gradients (8×), then parameters (∝1/Nd).
✅ Mixed-precision Adam costs ~16-20 bytes per parameter — the audit that started the paper.
✅ Demonstrated: 100B+ parameters, super-linear speedup, 400 GPUs, 15 Petaflops sustained.
✅ Up to 13B parameters with no model parallelism — accessibility as a feature, not a footnote.
✅ DeepSpeed carried this recipe into the open ecosystem's default training stack.
✅ Read it as the systems half of the scaling-laws story: the science said 'grow'; ZeRO said 'here's how'.