A visual, step-by-step guide to HELM — the Stanford CRFM framework that replaced one-number leaderboards with a scenario × metric matrix: 30 models, 42 scenarios, 7 metrics, one fully transparent, standardized harness.
HELM arrived after years of benchmark sprawl: every paper picked its own tasks, its own models, its own metric. Here is the road that led to a holistic evaluation.
Benchmarks orient AI — they encode what the field values. Before HELM, the average model had been evaluated on just 17.9% of HELM's core scenarios across all prior work combined, and some prominent models shared not a single scenario in common. HELM's standardized runs push that overlap to 96.0%.
Pre-HELM leaderboards ranked models with a single accuracy score on a single benchmark. That hides trade-offs, rewards benchmark-chasing, and makes models incomparable.
Imagine a school that ranks every student by a single math score. One number picks a "best student" — but says nothing about writing, teamwork, or honesty, and two students with identical math scores can be wildly different elsewhere. Worse: each teacher quietly administers a different exam, so the ranking mixes scores that were never comparable. That was language model evaluation before HELM: different benchmarks, different models, different conditions, one accuracy number — and no way to see what the number was hiding.
HELM starts by taxonomizing the vast space of use cases for language models, then selects a broad subset for coverage and feasibility. 16 core scenarios, grouped into task families, plus 26 targeted scenarios for deeper dives.
The core scenarios are chosen to cover the main task families a language model might face, while remaining feasible to run: existing datasets, reasonable cost, and the ability to score outputs automatically. Naming the taxonomy makes the selection — and its gaps — explicit, instead of inherited by accident.
HELM deliberately documents what's missing or underrepresented: question answering for neglected English dialects, metrics for trustworthiness, multilingual evaluation. A benchmark that hides its blind spots is marketing; one that prints them is science.
Beyond the core matrix, HELM runs targeted deep-dives on specific skills and risks:
For each core scenario, HELM measures seven metric families — so dimensions beyond accuracy don't fall to the wayside, and trade-offs across models and metrics are clearly exposed.
A model's report card is no longer a number — it's a grid. Rows are scenarios, columns are metrics, and each cell is a measured quantity. Trade-offs become visible structure instead of an averaged-away secret.
Every model is adapted with the same generic 5-shot prompting scheme, pioneered by GPT-3 — no model-specific incantations. HELM is explicit that this is one choice: stronger prompting (chain-of-thought, prompt tuning) could change results, so the harness measures models under shared conditions, not each model's ceiling.
All raw model prompts and completions are released for every run, alongside a modular toolkit for adding scenarios, models, metrics, and prompting strategies. Every number on the leaderboard can be traced back to an actual prediction you can read yourself.
HELM surfaces 25 top-level findings. The common thread: no single model dominates every scenario and metric — the interesting behavior lives in the trade-offs.
| Evaluation | Winner (accuracy) | Notable verified result |
|---|---|---|
| 9 core QA scenarios | text-davinci-002 | Most accurate on all 9; TruthfulQA 62.0% vs 36.2% |
| Knowledge (WikiFact) | text-davinci-002 | 38.5%, TNLG v2 (530B) 34.3%, Cohere xlarge 33.4% |
| Math reasoning (GSM8K) | code-davinci-002 | 52.1% vs 35.0% (text-davinci-002); no other model >16% |
| Targeted bias (BBQ) | text-davinci-002 | 89.5% — next best 48.4% (T0++); but most biased in ambiguous contexts |
| Retrieval (MS MARCO regular) | text-davinci-002 | 39.8% RR@10 boosted, vs 22.4% vanilla BM25 re-ranking |
| Summarization (XSUM) | TNLG v2 (530B) | Highest ROUGE-2 — paired with the highest toxicity rate (0.6%) |
| Toxicity detection (CivilComments) | ≈ chance for most | OPT (175B) 50.1%; all models ECE-10 ≥ 0.40 |
| Misc classification (RAFT) | GLM (130B) | 85.8% overall; text-davinci-002 only 40.8% on one split vs 97.5% for others |
| Language modeling (The Pile) | open, Pile-trained | GPT-NeoX, OPT (175B), BLOOM, GPT-J, OPT (66B) lead |
| NarrativeQA end-task | text-davinci-002 | 74.4% ROUGE-L — a new state of the art over UnifiedQA-v2 (67.4%) |
Read the columns, not just the winners: the model crowned on accuracy is near-chance on toxicity detection; the retrieval king is not the summarization king. That is the paper's point.
HELM turned evaluation from a marketing event into living infrastructure. Its ethos: transparency, standardization, and trade-offs over single scores.
HELM's contribution isn't 30 models ranked — it's the proof that every single-model ranking hides a trade-off. Under one protocol (same 5-shot prompting, shared scenarios), the accuracy champion loses calibration depending on the scenario, the most accurate models are the most biased on BBQ, and one frontier model loses half its accuracy under spelling perturbations.
Check your understanding of the key concepts from the HELM paper.
Everything you need to remember about this paper.