Conventional MoE activates big experts and hopes they specialize. DeepSeekMoE makes experts finer and reserves a few shared ones for common knowledge — comparable quality with a fraction of the compute.
From 'route to big experts' to 'compose many small ones' — the architecture decision that carried DeepSeek to V3.
MoE efficiency is only real if experts are actually specialized. DeepSeekMoE's diagnosis: coarse experts force each one to carry both common and niche knowledge, so routing wastes capacity. The fix is structural — finer granularity (mN small experts, mK activated: more flexible compositions) plus shared experts (K_s always active: common knowledge gets a permanent home and routed experts stop duplicating it).
Why conventional top-K MoE under-delivers on the specialization it promises.
GShard MoE is three generalist chefs — each can cook anything, so they all keep stock simmering and step on each other. DeepSeekMoE is a brigade: a garde-manger station that is always staffed (shared experts handle the basics), plus many narrow specialists (sauces, grill, pastry) combined per dish. Same total kitchen, far more menu per shift.
Fine-grained segmentation and shared-expert isolation — each independently validated in the paper's 2B ablations.
Every claim is validated at 2B first, then re-validated at 16B and 145B — the paper's methodical center of gravity.
The ladder pattern — validate architecture at 2B, confirm at 16B, extrapolate at 145B — is the methodological export: cheap ablations first, expensive confirmations second.
Not just benchmark wins — evidence that experts actually divide the labor.
| Configuration | Total params | Activated | Comparable to | Compute used |
|---|---|---|---|---|
| DeepSeekMoE 2B | 2B | ~0.2-0.4B | GShard 2.9B | ~2/3 of GShard |
| DeepSeekMoE 16B | 16B | 2.8B | LLaMA2 7B | ~40% |
| DeepSeekMoE 145B | 145B | ~22B | DeepSeek 67B | 28.5% (even 18.2%) |
| DeepSeek-V3 (successor) | 671B | 37B | closed frontier class | MoE efficiency at scale |
Compute figures from the paper's abstract. The V3 row shows where the architecture landed: the same two strategies scaled 46× in total parameters.
DeepSeekMoE became the architectural spine of the models that made frontier-class training look cheap.
Check your understanding of the key concepts from DeepSeekMoE.
Everything you need to remember about this paper.