A visual, step-by-step guide to the benchmark that measured what models actually know — across law, medicine, moral scenarios, and 53 other fields — and found the distance between fluent language and real knowledge.
NLP benchmarks had measured reading and reasoning tricks. MMLU asked a blunter question: does the model know things?
Standardized exams are expert-designed adversarial tests — distractors written by professionals to catch shallow pattern-matching. Reusing them across 57 fields gives breadth, difficulty calibration, and comparability to humans in one corpus — for free.
Fluency was improving fast. Knowledge — the thing users actually assumed — had no broad instrument.
MMLU is the driving test: multiple-choice, standardized, expert-designed — imperfectly correlated with real-world driving, but the only instrument that scores everyone on the same road signs. Its failures predict (roughly) whose car ends up in a ditch.
15,908 questions collected in total, split into a few-shot development set, a validation set for subject selection, and the test set.
Grouped into four domains, the subject list is a fossil record of what 2020 considered "an educated person's knowledge."
The paper's most-quoted finding lives at the edges: models were near random on moral scenarios, professional law, and morality — the socially load-bearing subjects. A model that aces virology but flips a coin on moral dilemmas has a very particular kind of blindness, and 2020's models had it.
The baseline sweep is the paper's time capsule: even the largest GPT-3 was a high-schooler with confidence problems.
| Model | Average Accuracy | Reading |
|---|---|---|
| Random chance | 25.0 | Four-choice floor |
| GPT-1 / small open models | ~27–35 | Barely above chance |
| GPT-3 few-shot (175B) | ~43.9 | ~+20 points over random — far from expert |
| UnifiedQA 11B (fine-tuned) | ~49.7 | Best 2020 system: high-school-ish |
| Human experts (target) | ~70–90+ per subject | Expert-level, the stated bar |
Headline framing from the paper: most models were near random; the very largest GPT-3 improved over chance by almost 20 points on average, and every subject still needed "substantial improvements" to reach expert accuracy.
MMLU's quiet second act: a knowledge test that doubles as a confidence audit.
A model that says "I'm 45% sure" when it's 45% right can be deferred to appropriately — by users, by pipelines, by downstream agents. A model that is confidently wrong is dangerous in proportion to its fluency. MMLU made this measurable and showed 2020 models mostly failed it.
Calibration research (and MMLU-style audits) seeded the later "abstention" agenda: models that know when to say "I don't know" — the connective tissue between this benchmark and the hallucination literature (SelfCheckGPT, FActScore).
For three years, "how smart is the model?" was answered with one MMLU number. That single number built — and strained — an industry habit.
The deepest MMLU lesson isn't in the paper: a benchmark stops measuring the moment it becomes a training target.
MMLU gave the field a shared map of what "knowing" means — and then demonstrated the field's core dynamic: any map becomes a destination. The correct posture is neither MMLU-worship nor MMLU-dismissal: read it as a 57-dimensional snapshot that ages exactly as fast as its questions leak. The benchmark's own worst subjects — morality, professional judgment — are precisely the ones no multiple-choice instrument will ever settle.
Check your understanding of the key concepts from the MMLU paper.
Everything you need to remember about this paper.