History Problem Architecture Routing Benchmarks Instruct Impact Deep Dive Quiz
Interactive Paper Explainer

Open Weights, Sparse Brains
Mixtral 8x7B

A visual, step-by-step guide to the open MoE that matched dense models 4× its active size — 47B total parameters, 13B active per token, top-2 routing, 32k context, and an Apache 2.0 license that changed the open-model market.

Start Learning Read the Paper ↗
47B
Total Parameters
13B
Active per Token
8
Experts · Top-2
2024
Year Published
History

The Open-Weights Counterattack

After a year of closed frontier releases, a French lab answered with sparsity + openness.

2021–23
MoE inside the labs
Switch, GLaM, and internal frontier MoEs prove the recipe — but stay closed. Switch Guide.
2023 · Feb–Sep
Mistral 7B — proof of efficiency
A dense 7B with sliding-window attention beats 13B-class models; Mistral AI establishes the "small, sharp, open" playbook.
2023 · Dec → 2024 · Jan
🚀 Mixtral 8x7B (Jiang et al.)
Take the Mistral 7B architecture, put 8 experts in every FFN layer, route top-2: 47B parameters with 13B active — outperforming Llama 2 70B, Apache 2.0, weights on the open internet.
2024 →
The open MoE era
Mixtral 8x22B, DeepSeek MoE families, Qwen MoE — sparse open models become a competitive market segment.
The Strategic Move

Mixtral is a business-model experiment as much as a model: prove that sparsity + permissive licensing lets an open release match closed frontier quality at a fraction of serving cost. Every design choice — 8 experts, top-2, 32k context, Apache 2.0 — serves that thesis. It worked, and the market structure of open AI changed.

🧭 Lineage
Switch proved top-1 at lab scale; Mixtral industrialized top-2 in the open.
Chapter 01

The Open-Model Compute Trap

Dense open models face a brutal trade: quality scales with parameters, and parameters scale with serving cost.

🐘
Dense Open Models Pay Everything
  • To match a 70B dense model, an open release needed… a 70B dense model
  • Every token runs through every parameter — quality purchased with linear serving cost
  • Self-hosting economics collapse beyond ~13B for most teams
  • Meanwhile frontier labs kept the (sparse) efficiency tricks internal
🪶
Sparcity: Capacity Without Cost
  • 8 expert FFNs per layer; a router picks two per token
  • 47B parameters of knowledge storage, ~13B of per-token compute
  • Beats Llama 2 70B across benchmarks — at roughly GPT-3.5-class serving economics
  • All of it under Apache 2.0 — commercial use, modification, redistribution
Analogy — The Studio Orchestra

A dense 70B is a full symphony paid per performance regardless of the piece. Mixtral is a studio of 8 session players per section, where the conductor calls exactly 2 per note: the studio's total skill library is huge, but each note costs a duo's fee.

Chapter 02

Mistral 7B, Eightfold

Architecturally, Mixtral is barely an innovation: the same decoder stack with each FFN replaced by 8 experts + a router.

y = Σi∈Top2 softmax(Wrx)i · Ei(x)
8 experts
Per FFN layer
Each expert is a standard feed-forward network; attention layers stay dense and shared.
Top-2
Routing
Softmax router selects the two highest-scoring experts per token; outputs combine weighted.
47B / 13B
Sparse ratio
Each token sees 47B parameters of storage but computes with ~13B.
32k
Context
Trained with 32k-token context — unusual generosity for the era, inherited from Mistral's sliding-window design.
What Stayed Boring (On Purpose)
  • Same tokenizer, training recipe, and inference stack as Mistral 7B — familiarity was a feature
  • No exotic routing tricks: plain softmax top-2 with standard load-balancing during training
  • The paper is short — the release, not the architecture, is the artifact
The Two-Expert Compromise
  • Switch's k=1 maximizes efficiency; Mixtral's k=2 buys gradient richness — two paths per token smooth training
  • Routing varies by timestep: the same token can see different expert pairs in different contexts
  • Robustness: with two experts, one misroute degrades a token instead of defining it
Chapter 03

Routing Live

Watch tokens pick expert pairs layer by layer — and what the router actually learns.

Interactive Demo — The Top-2 Router
🔬 What experts learn
The paper's analysis: specialization is modest and syntax-oriented, not clean semantic "domains" — some experts lean positional/structural roles rather than topics.
🔁 Sequential routing
Routing is sequential across layers: a token's expert pair at layer 3 can differ from its pair at layer 17 — information composes across specialists.
⚖️ Load balance
Training uses a balancing loss so no expert starves — critical for hardware utilization with expert parallelism.
Chapter 04

The Headline Table

The abstract's claim, unpacked: outperform or match Llama 2 70B and GPT-3.5 across all evaluated benchmarks.

Base-Model Comparison (selected benchmarks, paper-reported)
BenchmarkMixtral 8x7BLlama 2 70BGPT-3.5
MMLU (5-shot)~70.0~69.9~70.0
HellaSwag (10-shot)~81.3~81.3~81.1
Winogrande (5-shot)~72.1~80.1~72.3
Arc-C (5-shot)~61.2~64.8~61.6
GSM8K (5-shot, CoT)~40.6~34.2~57.4
MBPP (3-shot)~38.3~33.2~49.4

Values rounded from the paper's table. The pattern: parity-or-better vs Llama 2 70B (5.4× the active compute), and a decisive win on mathematics/code — "vastly outperforms" per the abstract, with GPT-3.5 still ahead on reasoning-heavy suites.

ACTIVE COMPUTE
13B
vs 70B for Llama 2 — same or better quality at ~1/5 the FLOPs
MULTILINGUAL
strong
beats Llama 2 70B on French/German/Spanish/Italian benchmarks
MATH & CODE
+6 pts
GSM8K and MBPP over Llama 2 70B — the "vastly outperforms" areas
CONTEXT
32k
tokens, with long-range retrieval holding up across the window
Interactive Demo — Quality vs Serving Cost
Chapter 05

The Instruct Sibling

Mixtral 8x7B — Instruct: supervised finetuning + DPO on the sparse base — and the paper's chat evaluations.

Training Recipe
  • SFT on curated public instruction data (multi-turn, multilingual)
  • DPO (direct preference optimization) — no RL loop — for preference alignment. DPO Guide
  • Same sparse architecture; adapters are unnecessary — full sparse finetuning
Human & Benchmark Results
  • On the paper's human benchmarks, Instruct Mixtral surpasses GPT-3.5 Turbo, Claude-2.1, Gemini Pro, and Llama 2 70B-chat
  • MT-Bench-class evaluations place it at the front of the open field for its era
  • Math and code remain the visible gaps versus stronger closed models
Legacy

Impact — Open MoE Becomes a Market

Mixtral's release re-anchored expectations: sparse efficiency is not a frontier secret, and open can compete.

🧬 The open-MoE segment
Mixtral 8x22B, DeepSeek-V2/V3, Qwen-MoE: sparse open models became their own competitive category.
💰 Serving economics
13B-active inference at 47B-quality made self-hosting frontier-class quality feasible — vLLM & friends added MoE kernels. PagedAttention Guide.
⚖️ Apache 2.0 signal
The most permissive practical license on a frontier-adjacent release — commercial adoption without legal friction became the baseline ask.
🔬 Honest analysis culture
The paper ships routing-analysis and data-mixture details openly — including what did NOT cleanly specialize.
⚠️ What it did NOT solve
MoE serving needs expert-parallel plumbing; memory footprint is still 47B; knowledge-per-parameter remains below dense.
🧭 Read with
Switch Transformers for the technique's origin story.
Deep Dive

47B Parameters Is a Category, Not a Fact

The most-quoted number in every MoE debate — and the honest ways to use it.

🥊
The Apples-vs-Oranges Fight
  • "Mixtral is just a 13B model" (compute camp) vs "it's 47B and should beat dense 70B" (memory camp)
  • Both camps quote the same paper; neither is wrong about the arithmetic
  • Parameter count stopped being a single number the moment routing existed
  • Leaderboards forced a naming convention (8x7B ≈ "8 experts of 7B-ish") to dodge the fight
📐
The Honest Accounting
  • Quality per active FLOP: the deployment metric — Mixtral wins decisively here
  • Quality per stored parameter: the capacity metric — dense models still hold an edge
  • Memory × throughput: the economics — 47B of VRAM per replica, batch-dependent amortization
  • Report the axis, then the number — the debate dissolves
Interactive Demo — The Three Axes

Same two models, three scoring axes. Switch axes and watch the winner flip.

Verdict

Mixtral's benchmark table is the moment the field had to grow up about model size. A routed model's "parameters" is a distribution over what a token could use, not a count of what it does use — and every honest comparison since names its axis first. The release itself — sparse, open, commercially unencumbered — did more to reshape the open ecosystem than its (modest, well-executed) architecture ever could.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the Mixtral of Experts paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Mixtral 8x7B: 8 experts per FFN layer, top-2 routing — 47B total, ~13B active per token.
✅ Outperforms or matches Llama 2 70B and GPT-3.5 across all evaluated benchmarks; "vastly" on math, code, multilingual.
✅ 32k-token context trained in — long-range retrieval holds across the window.
✅ Instruct variant (SFT + DPO) surpasses GPT-3.5 Turbo, Claude-2.1, Gemini Pro, Llama 2 70B-chat on human benchmarks.
✅ Apache 2.0: commercial-grade openness that reset market expectations for open models.
✅ Compare sparse models per active FLOP or per stored parameter — never a bare parameter count.