History Problem Design Subjects Results Calibration Impact Deep Dive Quiz
Interactive Paper Explainer

57 Subjects, One Exam
MMLU

A visual, step-by-step guide to the benchmark that measured what models actually know — across law, medicine, moral scenarios, and 53 other fields — and found the distance between fluent language and real knowledge.

Start Learning Read the Paper ↗
57
Subjects
15,908
Questions
~+20pp
GPT-3 Over Random
2020
Year Published
History

From Skills to Knowledge

NLP benchmarks had measured reading and reasoning tricks. MMLU asked a blunter question: does the model know things?

2018–19
GLUE / SuperGLUE era
Sentence-level understanding benchmarks — sentiment, entailment, QA formats. Rapid saturation by fine-tuned models.
2020 · May
GPT-3 few-shot shock
175B parameters show in-context learning — but knowledge coverage beyond popular trivia is unmeasured.
2020 · Sep
🚀 MMLU (Hendrycks et al.)
57 subjects, 15,908 four-choice questions sampled from standardized exams and study guides — elementary math to professional law.
2021 →
The universal yardstick
Every major release reports MMLU (GPT-4: 86.4%). Variants multiply: MMLU-Pro, GPQA — and the contamination era begins.
The Design Insight

Standardized exams are expert-designed adversarial tests — distractors written by professionals to catch shallow pattern-matching. Reusing them across 57 fields gives breadth, difficulty calibration, and comparability to humans in one corpus — for free.

🧭 Pairing
MMLU measures knowledge breadth; TruthfulQA measures honesty under imitation pressure.
Chapter 01

The Unmeasured Middle

Fluency was improving fast. Knowledge — the thing users actually assumed — had no broad instrument.

🎭
Fluency Masquerades as Knowledge
  • Models that sound expert in every register can fail elementary facts in the same paragraph
  • Perplexity on held-out text says nothing about professional knowledge: law, medicine, engineering
  • Existing QA benchmarks are narrow (one domain) or shallow (popular trivia)
  • Nobody can answer: "is this model broader than a smart high-schooler?"
🎓
One Exam, Every Field
  • Sourced from real standardized tests: GRE, USMLE practice, bar-exam prep, AP courses
  • 57 subjects across STEM, humanities, social sciences, and "other" (incl. professional law & medicine)
  • Four-choice MCQ with expert-written distractors
  • Three splits: few-shot dev, validation (subject selection), and the test set
Analogy — The Driving Test vs Street Driving

MMLU is the driving test: multiple-choice, standardized, expert-designed — imperfectly correlated with real-world driving, but the only instrument that scores everyone on the same road signs. Its failures predict (roughly) whose car ends up in a ditch.

Chapter 02

Built From Real Exams

15,908 questions collected in total, split into a few-shot development set, a validation set for subject selection, and the test set.

📚 57 subjects
Elementary mathematics → abstract algebra; US history → global facts; anatomy → clinical knowledge → professional medicine; jurisprudence → professional law; moral scenarios, econometrics, formal logic, machine learning…
🎯 Expert distractors
Wrong options are plausible by design — they encode the specific misconceptions exam-writers expect.
📏 Difficulty spread
From elementary school to graduate/professional level — a per-subject difficulty ladder in one dataset.
🧹 Practice hygiene
Freely available study-guide material only, with attention to public-visibility — long before "decontamination" became a field-wide ritual.
Interactive Demo — Assemble an MMLU Question

Flip through real-style questions and see why expert distractors matter. Guess before revealing.

Chapter 03

The Map of Knowledge

Grouped into four domains, the subject list is a fossil record of what 2020 considered "an educated person's knowledge."

Four Domains
  • STEM (20): math from elementary to abstract algebra, physics, chemistry, biology, CS, EE, ML…
  • Humanities (13): philosophy (formal logic, moral scenarios), law (jurisprudence, international, professional), world religions, history…
  • Social Sciences (18): econometrics, psychology (professional-level), sociology, US politics, security studies…
  • Other (6): miscellaneous professional subjects — business, clinical knowledge, global facts, nutrition…
Why the Weird Subjects Matter

The paper's most-quoted finding lives at the edges: models were near random on moral scenarios, professional law, and morality — the socially load-bearing subjects. A model that aces virology but flips a coin on moral dilemmas has a very particular kind of blindness, and 2020's models had it.

Interactive Demo — Subject Grid Difficulty

Pick a domain; each subject tile shows the 2020-era model accuracy pattern (near-random = red).

Chapter 04

What 2020 Models Didn't Know

The baseline sweep is the paper's time capsule: even the largest GPT-3 was a high-schooler with confidence problems.

The Baseline Table (2020 models, test set)
ModelAverage AccuracyReading
Random chance25.0Four-choice floor
GPT-1 / small open models~27–35Barely above chance
GPT-3 few-shot (175B)~43.9~+20 points over random — far from expert
UnifiedQA 11B (fine-tuned)~49.7Best 2020 system: high-school-ish
Human experts (target)~70–90+ per subjectExpert-level, the stated bar

Headline framing from the paper: most models were near random; the very largest GPT-3 improved over chance by almost 20 points on average, and every subject still needed "substantial improvements" to reach expert accuracy.

NEAR-RANDOM SUBJECTS
moral · law
morality, moral scenarios, professional law — socially critical, worst covered
LOPSIDED PROFILES
uneven
strong in some subjects, chance-level in others — knowledge is patchy, not smooth
MISCALIBRATION
overconfident
models frequently don't know when they're wrong (confidence ≠ correctness)
EXPERT GAP
wide
no 2020 model came close to expert-level on every subject
Chapter 05

Knowing When You Don't Know

MMLU's quiet second act: a knowledge test that doubles as a confidence audit.

Interactive Demo — Confidence vs Correctness

Sort answer attempts by model confidence. In an ideal model, confidence tracks accuracy; in a miscalibrated one, wrong answers hide among confident ones.

Why Calibration Is a Safety Property

A model that says "I'm 45% sure" when it's 45% right can be deferred to appropriately — by users, by pipelines, by downstream agents. A model that is confidently wrong is dangerous in proportion to its fluency. MMLU made this measurable and showed 2020 models mostly failed it.

From Confidence to Refusal

Calibration research (and MMLU-style audits) seeded the later "abstention" agenda: models that know when to say "I don't know" — the connective tissue between this benchmark and the hallucination literature (SelfCheckGPT, FActScore).

Legacy

Impact — The Default Stat

For three years, "how smart is the model?" was answered with one MMLU number. That single number built — and strained — an industry habit.

📈 The universal leaderboard
Chinchilla, PaLM, GPT-4, LLaMA, Mistral — every release table has an MMLU column (GPT-4: 86.4%).
🧬 Harder descendants
MMLU-Pro, GPQA, and domain-specific derivatives chase the ceiling MMLU eventually hit.
🧹 Contamination science
Public exam questions leak into web corpora — MMLU is the benchmark that taught the field to care about decontamination.
🧠 The knowledge-coverage genre
BIG-Bench, HELM, AGIEval — all inherit the "many-subject exam" template.
⚖️ What it did NOT solve
MCQ format rewards elimination skill; knowledge ≠ reasoning (GSM8K's era began in parallel); one number hides profile shape.
🧭 Read next
HELM is the field's answer to "one number is not evaluation."
Deep Dive

Benchmarks Shape Models

The deepest MMLU lesson isn't in the paper: a benchmark stops measuring the moment it becomes a training target.

🔁
Goodhart's Exam Hall
  • Once every lab optimizes MMLU, the number measures optimization, not knowledge
  • Web-scale pretraining almost certainly contains exam text — contamination inflates scores
  • MCQ-specific tricks (elimination, position bias) lift scores without lifting understanding
  • 57 subjects became a curriculum — models now train with MMLU-style data in the mix
🧭
The Honest Usage Pattern
  • Compare within generations, same evaluation harness, and with per-subject profiles
  • Prefer held-out descendants (MMLU-Pro, GPQA) when a ceiling is suspected
  • Report calibration alongside accuracy — the second axis MMLU enabled
  • Treat it as one panel of a dashboard: HELM's whole thesis
Interactive Demo — The Score Timeline Machine

Watch what happens to the meaning of "MMLU accuracy" as the field starts training for it. Scroll the timeline slider.

Verdict

MMLU gave the field a shared map of what "knowing" means — and then demonstrated the field's core dynamic: any map becomes a destination. The correct posture is neither MMLU-worship nor MMLU-dismissal: read it as a 57-dimensional snapshot that ages exactly as fast as its questions leak. The benchmark's own worst subjects — morality, professional judgment — are precisely the ones no multiple-choice instrument will ever settle.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the MMLU paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 57 subjects, 15,908 questions total, four-choice MCQ — split into dev, validation, and test sets.
✅ Sourced from real standardized exams and study guides — expert-designed distractors for free.
✅ 2020 baseline: most models near random (25%); GPT-3 175B few-shot ≈ +20 points over chance.
✅ Near-random on socially critical subjects: morality, moral scenarios, professional law.
✅ Profiles are lopsided and miscalibrated — models often don't know when they're wrong.
✅ Became the default release stat (GPT-4: 86.4%) — and the canonical contamination cautionary tale.