History Problem Core Idea Routing Training Results Impact Deep Dive Quiz
Interactive Paper Explainer

Trillion-Parameter Routing
Switch Transformers

A visual, step-by-step guide to the paper that simplified Mixture-of-Experts down to a single expert per token — and used that simplicity to scale language models past a trillion parameters while keeping the compute per token flat.

Start Learning Read the Paper ↗
1.6T
Parameters (Switch-C)
2048
Experts per Layer
7×
Faster Pre-Training
2021
Year Published
History

From Committees to Switches

Mixture-of-Experts is thirty years older than the LLM era. Switch Transformers earned its name by throwing away the committee and keeping the switch.

1991
Mixture of Experts (Jacobs et al.)
Neural nets with learned "gates" that divide labor between small expert networks — competitive learning from the connectionist era.
2017
Sparsely-Gated MoE (Shazeer et al.)
MoE inside LSTM layers: a noisy top-k gate activates 2 of thousands of experts per token — 137B parameters, but training was unstable and communication-heavy.
2020
GShard (Lepikhin et al.)
MoE meets the Transformer and mesh parallelism: a 600B-parameter translation model that swapped every other feed-forward block for 1-of-2048 expert routing.
2021 · Jan
🚀 Switch Transformers (Fedus et al.)
Drop k=2 routing to k=1, add router z-loss, run experts in bfloat16. Simpler, faster, stabler — and it scales to 1.6 trillion parameters.
2022 →
GLaM, NLLB, Mixtral…
The Switch recipe becomes the default for frontier efficiency: one-trillion-class sparse models are now routine, and open MoE releases like Mixtral 8x7B follow directly.
Why This Paper Mattered

Before Switch, MoE was powerful but fragile: k=2 routing doubled expert bookkeeping, soft weights tangled gradients, and instabilities forced lower precision to be avoided. Switch Transformers showed that removing machinery — routing to exactly one expert — makes sparsity simultaneously cheaper, faster, and more stable. That inversion is why nearly every large sparse model since inherits its design.

🧭 Where to read next
The Mixtral Guide shows the k=2 variant of this same idea in an open 2024 model.
Chapter 01

The Dense Dilemma

A dense model makes every token pay for every parameter. That symmetry is what makes scaling expensive.

🧱
Dense: Every Token Pays for Everything
  • In a dense Transformer, each token flows through 100% of the feed-forward parameters on every layer
  • Cost grows with capacity: a trillion-parameter dense model would need impossible compute per token
  • Training FLOPs, serving latency, and memory all scale with the full parameter count
  • You cannot buy more knowledge capacity without also buying more compute per token
🔀
Sparse: Different Parameters per Token
  • Replace each feed-forward block with many experts and a router that picks one per token
  • Total parameters explode — but FLOPs per token stay constant
  • The router learns which specialist deserves each token, end-to-end with the rest of the network
  • Knowledge capacity and compute cost become decoupled dials
Analogy — The Kitchen

A dense model is a kitchen where every order walks through every station — grill, pastry, sushi, whatever you asked for. A sparse model is a kitchen with a maître d' (the router) who sends each order to exactly one station (the expert). The restaurant gets more stations — more total equipment — without any single order taking longer.

Chapter 02

The Switch: Top-1 Routing

The paper's central simplification: route each token to its single best expert — no committee, no weighted blend.

r = softmax(Wr·x)  ·  i = argmax(r)  ·  y = ri · Ei(x)
Wr·x
Router logits
A tiny linear layer scores how well each expert fits this token.
argmax(r)
Hard selection
Pick expert #i — exactly one. No k=2 blending, no noisy gating.
ri
Gating weight
The expert's output is scaled by its routing probability — gradients flow to the router.
Ei(x)
Expert computation
A standard feed-forward network — only the selected one runs.
Why k=1 Beats k=2
  • Half the expert compute: prior MoE ran two experts per token and averaged them — Switch runs one.
  • Simpler communication: one expert-to-device round trip per token instead of two, a major cost at scale.
  • Cleaner gradients: routing decisions are a discrete switch scaled by one probability, which trains stably with the right regularization.
  • Scale instead: the compute you save goes into more experts — more capacity at the same FLOPs.
The Capacity Knob

Each expert gets a buffer of expert_capacity = (tokens per batch / experts) × capacity_factor. If one expert attracts too many tokens (a popular specialist), the overflow is dropped — that token's layer output becomes simply its residual connection. Dropped tokens are a visible training signal: too many means routing is collapsing onto a few experts. The paper keeps ~99%+ of tokens routed in healthy runs and logs the drop rate as a health metric.

Chapter 03

Routing in Motion

Watch tokens flow through a layer of experts — including what happens when one expert overflows its capacity.

Interactive Demo — The Top-1 Router

A batch of 6 tokens routes through 4 experts. Press Route tokens to see assignments; lower the capacity factor to starve the popular expert and watch tokens get dropped.

🧠 What the router learns
Without any supervision, experts specialize — by syntax, token frequency, or task domain. The router is just a linear layer trained end-to-end.
⚖️ Load balance
An auxiliary load-balancing loss keeps expert utilization even, so capacity never starves a hot expert.
🪂 Dropped-token safety
Overflow tokens pass through via the residual path — the layer degrades gracefully instead of crashing.
Chapter 04

Taming Instability

Sparse models used to fall over in training. Switch's fixes — a router z-loss, selective precision, and careful initialization — made bfloat16 MoE possible for the first time.

Fix 1 — Router z-loss

Round-off error in the router's softmax blows up in bfloat16: tiny logit perturbations flip argmax decisions, which sends tokens to wrong experts, which perturbs gradients — a feedback loop. The z-loss penalizes large logits directly: L_z = (1/B) Σ (logits²), keeping the router's arithmetic in a numerically comfortable range.

Fix 2 — Selective Precision

Run the whole model in memory-cheap bfloat16, but keep the router — the most error-sensitive part — in float32. This hybrid is why the paper could claim the first stable large-scale sparse training with low precision formats.

Interactive Demo — Router Stability Gym

Watch 20 steps of router logits under bfloat16 round-off. Without the z-loss, logits drift and routing decisions flip; with it, the distribution stays tight.

🔧 Fix 3 — Initialization
Scaled-down random init for router weights avoids starting near decision boundaries where small noise flips routes.
💧 Fix 4 — Regularization
Increased dropout (d=0.2 variants) helps fine-tuned sparse models that would otherwise overfit small downstream datasets.
📈 Fix 5 — More data / longer training
Sparse models underperform dense ones in the under-trained regime — the quality gap disappears with sufficient training steps.
Chapter 05

Speed and Scale

Same compute per token, up to 7× faster pre-training — and a 1.6 trillion parameter model that actually trains.

Pre-training Speed (time to T5-level quality)
ModelParamsSpeed vs. dense twinSetting
T5-Base (dense)0.2B1×Same compute per token
Switch-Base7Bup to 7×Same compute per token
T5-Large (dense)0.7B1×Same compute per token
Switch-Large26Bup to 7×Same compute per token
T5-XXL (dense)11B1×Strong baseline
Switch-C (1T)1.6T4×Trillion-scale, C4

Speedups are wall-clock time to reach the same pre-training quality — the sparse model gets there first because more of its capacity is trained in parallel per step.

Fine-tuning Quality (Base size)
ModelGLUESuperGLUESQuAD
T5-Base82.972.483.5
Switch-Base84.773.083.7

With identical per-token compute, the sparse twin wins on downstream benchmarks too.

Distillation — Gains You Can Ship

The trillion-parameter model is a research artifact, but its knowledge compresses: distilling Switch-Base into small dense students preserves roughly 30% of the sparse model's quality gain — meaningfully better small models with none of the serving complexity. The paper even distills fine-tuned sparse checkpoints.

Interactive Demo — Pre-training Speedup Race

Time-to-quality: press start and watch the dense T5-Base and compute-matched Switch-Base race to the same loss.

Legacy

Impact — The Default Architecture of Scale

Switch Transformers turned MoE from a fragile niche into the standard way to buy knowledge capacity without buying compute.

🏭 The Switch Layer
The paper's layer design — replace FFN with N experts + top-1 router — is now a reusable building block in open frameworks.
🌍 GLaM (2022)
1.2T parameters, activated sparsely, outperforming GPT-3 on many tasks with ~1/3 of the training energy.
🐦 Mixtral 8x7B (2024)
The open-weights MoE that brought this design to everyone — read its guide for the k=2 variant.
📚 NLLB-200 & translation
Meta's 200-language translator is a sparsely gated MoE — the Switch lineage in production.
🔬 Stability toolkit
Router z-loss and selective precision are now standard vocabulary whenever anyone trains sparse models.
⚠️ What it did NOT solve
Serving many experts needs expert-parallelism communication; memory footprint stays huge even though FLOPs are flat.
Deep Dive

1.6T Parameters ≠ 1.6T Knowledge

A sparse parameter is memory, not compute. Understanding that distinction is the key to reading every MoE claim since.

🪞
The Headline Illusion
  • "1.6 trillion parameters" sounds like 1.6T of dense capacity — it is not
  • Each token touches only one expert per layer: compute per token stays T5-Base-sized
  • Knowledge-per-parameter is lower than a dense model of the same size
  • Serving costs: all 1.6T weights must sit in memory (across devices), with routing traffic between them
🧮
The Honest Ledger
  • Total parameters = knowledge storage; active parameters = per-token compute
  • Switch-C: 2048 experts per layer store 1.6T parameters; one expert runs per token per layer
  • The win is time-to-quality: more capacity trained in parallel, same FLOPs
  • Distillation is the honesty test: ~30% of the gain survives compression into dense students
Interactive Demo — The Parameter Ledger

Compare what a token "pays" in a dense T5-Base versus the 1.6T Switch-C — flip between the memory view and the compute view.

Verdict

Read MoE claims as "more bookshelf, not more reading speed". The trillion parameters are shelf space for specialist knowledge the router can fetch in constant time. That is genuinely valuable — it is how modern systems hold enormous knowledge at controllable inference cost — but it is a different resource than dense scale, and the two should never be compared parameter-for-parameter.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the Switch Transformers paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Top-1 routing: each token uses exactly one expert per layer — the simplification that made MoE stable and fast.
✅ Constant compute per token, exploding parameter count: capacity and FLOPs decouple.
✅ Router z-loss + selective precision (bfloat16 model, float32 router) = first stable low-precision sparse training.
✅ Up to 7× faster pre-training than compute-matched dense T5 twins; 4× over T5-XXL at trillion scale.
✅ Switch-C: 2048 experts per layer, 1.6 trillion parameters — trained on the Colossal Clean Crawled Corpus.
✅ Sparse parameters are memory, not compute — distillation keeps ~30% of the quality gain in small dense models.