A visual, step-by-step guide to the paper that simplified Mixture-of-Experts down to a single expert per token — and used that simplicity to scale language models past a trillion parameters while keeping the compute per token flat.
Mixture-of-Experts is thirty years older than the LLM era. Switch Transformers earned its name by throwing away the committee and keeping the switch.
Before Switch, MoE was powerful but fragile: k=2 routing doubled expert bookkeeping, soft weights tangled gradients, and instabilities forced lower precision to be avoided. Switch Transformers showed that removing machinery — routing to exactly one expert — makes sparsity simultaneously cheaper, faster, and more stable. That inversion is why nearly every large sparse model since inherits its design.
A dense model makes every token pay for every parameter. That symmetry is what makes scaling expensive.
A dense model is a kitchen where every order walks through every station — grill, pastry, sushi, whatever you asked for. A sparse model is a kitchen with a maître d' (the router) who sends each order to exactly one station (the expert). The restaurant gets more stations — more total equipment — without any single order taking longer.
The paper's central simplification: route each token to its single best expert — no committee, no weighted blend.
Each expert gets a buffer of expert_capacity = (tokens per batch / experts) × capacity_factor. If one expert attracts too many tokens (a popular specialist), the overflow is dropped — that token's layer output becomes simply its residual connection. Dropped tokens are a visible training signal: too many means routing is collapsing onto a few experts. The paper keeps ~99%+ of tokens routed in healthy runs and logs the drop rate as a health metric.
Watch tokens flow through a layer of experts — including what happens when one expert overflows its capacity.
Sparse models used to fall over in training. Switch's fixes — a router z-loss, selective precision, and careful initialization — made bfloat16 MoE possible for the first time.
Round-off error in the router's softmax blows up in bfloat16: tiny logit perturbations flip argmax decisions, which sends tokens to wrong experts, which perturbs gradients — a feedback loop. The z-loss penalizes large logits directly: L_z = (1/B) Σ (logits²), keeping the router's arithmetic in a numerically comfortable range.
Run the whole model in memory-cheap bfloat16, but keep the router — the most error-sensitive part — in float32. This hybrid is why the paper could claim the first stable large-scale sparse training with low precision formats.
Same compute per token, up to 7× faster pre-training — and a 1.6 trillion parameter model that actually trains.
| Model | Params | Speed vs. dense twin | Setting |
|---|---|---|---|
| T5-Base (dense) | 0.2B | 1× | Same compute per token |
| Switch-Base | 7B | up to 7× | Same compute per token |
| T5-Large (dense) | 0.7B | 1× | Same compute per token |
| Switch-Large | 26B | up to 7× | Same compute per token |
| T5-XXL (dense) | 11B | 1× | Strong baseline |
| Switch-C (1T) | 1.6T | 4× | Trillion-scale, C4 |
Speedups are wall-clock time to reach the same pre-training quality — the sparse model gets there first because more of its capacity is trained in parallel per step.
| Model | GLUE | SuperGLUE | SQuAD |
|---|---|---|---|
| T5-Base | 82.9 | 72.4 | 83.5 |
| Switch-Base | 84.7 | 73.0 | 83.7 |
With identical per-token compute, the sparse twin wins on downstream benchmarks too.
The trillion-parameter model is a research artifact, but its knowledge compresses: distilling Switch-Base into small dense students preserves roughly 30% of the sparse model's quality gain — meaningfully better small models with none of the serving complexity. The paper even distills fine-tuned sparse checkpoints.
Switch Transformers turned MoE from a fragile niche into the standard way to buy knowledge capacity without buying compute.
A sparse parameter is memory, not compute. Understanding that distinction is the key to reading every MoE claim since.
Read MoE claims as "more bookshelf, not more reading speed". The trillion parameters are shelf space for specialist knowledge the router can fetch in constant time. That is genuinely valuable — it is how modern systems hold enormous knowledge at controllable inference cost — but it is a different resource than dense scale, and the two should never be compared parameter-for-parameter.
Check your understanding of the key concepts from the Switch Transformers paper.
Everything you need to remember about this paper.