Data parallelism replicates everything; model parallelism fragments compute. ZeRO keeps data-parallel training's simplicity while sharding the memory — optimizer states, gradients, then parameters — across devices.
The paper starts with an accounting question: in data-parallel training, what is actually stored per GPU — and how much of it is a copy?
With mixed-precision Adam, every parameter costs 16 bytes: 2 for the FP16 weight, 4 for the FP32 master copy, 8 for the two FP32 optimizer moments. Add 2+4 for the gradient and its FP32 copy — a 7B model's training state is ~112GB before activations. Data parallelism multiplies that by the GPU count. ZeRO's insight: none of it needs to be replicated — partition it, and communicate exactly what each rank needs, when it needs it.
The redundancy problem hiding inside the field's favorite parallelism strategy.
Classic data parallelism: every student in the class has a complete copy of the lab notebook — identical notes, identical backups, 40× the paper. ZeRO: the class keeps one notebook, torn into sections — each student holds a section, passes pages when someone needs them, and everyone still does their own experiments. Nothing is copied; everything is reachable.
Each stage shards one more memory class — with a careful eye on the communication bill.
Total ≈ 16-20 bytes per parameter — replicated per GPU under naive data parallelism.
From the abstract: the 100B run, the throughput, and the usability bonus.
The usability clause is the sleeper hit: researchers who could never restructure code for pipeline/tensor parallelism could now train 13B models with an import — DeepSpeed became the open ecosystem's default trainer.
Sharding should cost communication; ZeRO's accounting showed the trade was overwhelmingly profitable.
| Stage | Sharded | Memory reduction | Extra communication | Practical ceiling |
|---|---|---|---|---|
| Baseline DP | nothing | 1× | grad all-reduce | ~1-4B per node |
| ZeRO-1 (P_os) | optimizer states | ~4× | ≈ none | 11B-class single node |
| ZeRO-2 (P_os+g) | + gradients | ~8× | ≈ none | larger, still DP-simple |
| ZeRO-3 (P_os+g+p) | + parameters | ∝ 1/Nd | +50% (all-gather / reduce-scatter) | 100B+ / trillion-class |
Reduction factors from the paper's memory analysis for mixed-precision Adam training. The stage ladder is the API surface users actually see in DeepSpeed today.
ZeRO became DeepSpeed, and DeepSpeed became how the open world trains big models.
Check your understanding of the key concepts from ZeRO.
Everything you need to remember about this paper.