A visual, step-by-step guide to the open MoE that matched dense models 4× its active size — 47B total parameters, 13B active per token, top-2 routing, 32k context, and an Apache 2.0 license that changed the open-model market.
After a year of closed frontier releases, a French lab answered with sparsity + openness.
Mixtral is a business-model experiment as much as a model: prove that sparsity + permissive licensing lets an open release match closed frontier quality at a fraction of serving cost. Every design choice — 8 experts, top-2, 32k context, Apache 2.0 — serves that thesis. It worked, and the market structure of open AI changed.
Dense open models face a brutal trade: quality scales with parameters, and parameters scale with serving cost.
A dense 70B is a full symphony paid per performance regardless of the piece. Mixtral is a studio of 8 session players per section, where the conductor calls exactly 2 per note: the studio's total skill library is huge, but each note costs a duo's fee.
Architecturally, Mixtral is barely an innovation: the same decoder stack with each FFN replaced by 8 experts + a router.
Watch tokens pick expert pairs layer by layer — and what the router actually learns.
The abstract's claim, unpacked: outperform or match Llama 2 70B and GPT-3.5 across all evaluated benchmarks.
| Benchmark | Mixtral 8x7B | Llama 2 70B | GPT-3.5 |
|---|---|---|---|
| MMLU (5-shot) | ~70.0 | ~69.9 | ~70.0 |
| HellaSwag (10-shot) | ~81.3 | ~81.3 | ~81.1 |
| Winogrande (5-shot) | ~72.1 | ~80.1 | ~72.3 |
| Arc-C (5-shot) | ~61.2 | ~64.8 | ~61.6 |
| GSM8K (5-shot, CoT) | ~40.6 | ~34.2 | ~57.4 |
| MBPP (3-shot) | ~38.3 | ~33.2 | ~49.4 |
Values rounded from the paper's table. The pattern: parity-or-better vs Llama 2 70B (5.4× the active compute), and a decisive win on mathematics/code — "vastly outperforms" per the abstract, with GPT-3.5 still ahead on reasoning-heavy suites.
Mixtral 8x7B — Instruct: supervised finetuning + DPO on the sparse base — and the paper's chat evaluations.
Mixtral's release re-anchored expectations: sparse efficiency is not a frontier secret, and open can compete.
The most-quoted number in every MoE debate — and the honest ways to use it.
Mixtral's benchmark table is the moment the field had to grow up about model size. A routed model's "parameters" is a distribution over what a token could use, not a count of what it does use — and every honest comparison since names its axis first. The release itself — sparse, open, commercially unencumbered — did more to reshape the open ecosystem than its (modest, well-executed) architecture ever could.
Check your understanding of the key concepts from the Mixtral of Experts paper.
Everything you need to remember about this paper.