History Problem Scenarios Metrics Matrix Findings Impact Deep Dive Quiz
Interactive Paper Explainer

Evaluating Everything,
Not Just Accuracy — HELM

A visual, step-by-step guide to HELM — the Stanford CRFM framework that replaced one-number leaderboards with a scenario × metric matrix: 30 models, 42 scenarios, 7 metrics, one fully transparent, standardized harness.

Start Learning Read the Paper ↗
30
Models Evaluated
42
Scenarios (16 Core + 26 Targeted)
7
Metrics per Scenario
2022
Year Published (arXiv)
History

From One Score to Everything

HELM arrived after years of benchmark sprawl: every paper picked its own tasks, its own models, its own metric. Here is the road that led to a holistic evaluation.

2018
GLUE (Wang et al.)
One leaderboard, one average score, 9 tasks. The BERT era: fine-tune, climb, repeat. SuperGLUE follows in 2019 — still one number per board.
2019
GPT-2 (Radford et al.)
Raw scale shows surprising generality — no task-specific heads needed. But each new model is still judged on a few benchmarks, in different conditions.
2020
GPT-3 & few-shot prompting (Brown et al.)
Language models become general-purpose APIs. Benchmarks saturate, evaluation fragments: each team reports its own numbers on its own tasks.
2021–22
BIG-bench (Srivastava et al.)
200+ community-contributed tasks expand breadth — but models keep getting evaluated on different slices, so head-to-head comparison stays apples-to-oranges.
2022 · Nov
🚀 HELM (Liang et al.)
Stanford CRFM taxonomizes scenarios × metrics and benchmarks 30 models on 42 scenarios with 7 metrics each — under one standardized, fully released harness.
2023 →
A living benchmark
Published in TMLR. The leaderboard at crfm.stanford.edu/helm keeps growing: Classic and Lite suites, domain leaderboards (MedHELM, SEA-HELM, VHELM), Capabilities & Safety boards.
Key Insight

Benchmarks orient AI — they encode what the field values. Before HELM, the average model had been evaluated on just 17.9% of HELM's core scenarios across all prior work combined, and some prominent models shared not a single scenario in common. HELM's standardized runs push that overlap to 96.0%.

MODEL OVERLAP WITH HELM CORE SCENARIOS
before: 17.9% · with HELM: 96.0%
Same 16 core scenarios, same 7 metrics, same few-shot prompting — for all 30 models.
Chapter 01

The Problem with One Number

Pre-HELM leaderboards ranked models with a single accuracy score on a single benchmark. That hides trade-offs, rewards benchmark-chasing, and makes models incomparable.

📊
The Single-Score Leaderboard
  • One accuracy number crowns a "winner" per benchmark
  • Each paper evaluates different models on different tasks
  • Calibration, bias, toxicity, efficiency fall to the wayside
  • Trade-offs are averaged away instead of exposed
  • Prompts and outputs rarely released — results can't be audited
🧮
HELM's Scenario × Metric Matrix
  • Taxonomize scenarios (use cases) and metrics (desiderata)
  • Measure 7 metrics across 16 core scenarios wherever possible
  • 30 models run under identical, standardized conditions
  • Every raw prompt and completion released publicly
  • Trade-offs surface as a matrix you can inspect cell by cell
Interactive Demo — Leaderboard vs Reality

One button flips the view. The left reality is what a single score tells you; the right reality is what HELM's matrix reveals. Values marked ~ are illustrative, patterned on the paper's verified trade-offs.

view: single score
Analogy — Grading Students Only on Math

Imagine a school that ranks every student by a single math score. One number picks a "best student" — but says nothing about writing, teamwork, or honesty, and two students with identical math scores can be wildly different elsewhere. Worse: each teacher quietly administers a different exam, so the ranking mixes scores that were never comparable. That was language model evaluation before HELM: different benchmarks, different models, different conditions, one accuracy number — and no way to see what the number was hiding.

Chapter 02

Scenarios — A Taxonomy of Use Cases

HELM starts by taxonomizing the vast space of use cases for language models, then selects a broad subset for coverage and feasibility. 16 core scenarios, grouped into task families, plus 26 targeted scenarios for deeper dives.

❓ Question Answering · 9 scenarios
BoolQ, NewsQA, NarrativeQA, NaturalQuestions, QuAC, HellaSwag, OpenBookQA, TruthfulQA, MMLU — the biggest task family, from commonsense to knowledge to truthfulness.
🔎 Information Retrieval · 2 scenarios
MS MARCO (regular) and MS MARCO (TREC) — ranking candidate passages for a query, the classic search task.
📝 Summarization · 2 scenarios
CNN/DailyMail and XSUM — condense long documents into short reference-style summaries.
😊 Sentiment Analysis · 1 scenario
IMDB movie reviews — classify text as positive or negative.
☠️ Toxicity Detection · 1 scenario
CivilComments — classify comments as toxic or not, with group-identity metadata for fairness analysis.
🗂️ Misc Text Classification · 1 scenario
RAFT — 11 diverse real-world sub-tasks (banking queries, legal overruling, systematic review inclusion…).
Selection Logic — Coverage × Feasibility

The core scenarios are chosen to cover the main task families a language model might face, while remaining feasible to run: existing datasets, reasonable cost, and the ability to score outputs automatically. Naming the taxonomy makes the selection — and its gaps — explicit, instead of inherited by accident.

Honest Gaps, Stated Up Front

HELM deliberately documents what's missing or underrepresented: question answering for neglected English dialects, metrics for trustworthiness, multilingual evaluation. A benchmark that hides its blind spots is marketing; one that prints them is science.

7 Targeted Evaluations · 26 Additional Scenarios

Beyond the core matrix, HELM runs targeted deep-dives on specific skills and risks:

🗣️
Language
The Pile, BLiMP, ICE…
🌍
Knowledge
WikiFact + 5 QA sets
🧩
Reasoning
GSM8K, MATH, LSAT…
⚖️
Memorization
copyright regurgitation
📰
Disinformation
headline generation
🪞
Bias
BBQ + stereotypical pairs
☠️
Toxicity
RealToxicityPrompts, BOLD
16
core scenarios, all 7 metrics where possible
26
targeted scenarios for deep-dives
21
scenarios new to mainstream LM evaluation
6
task families behind the core scenarios
Chapter 03

Metrics — Seven Desiderata, Not One

For each core scenario, HELM measures seven metric families — so dimensions beyond accuracy don't fall to the wayside, and trade-offs across models and metrics are clearly exposed.

HELM = Taxonomize → Select → Measure → Standardize
🗺️
Taxonomize
Map the full space of scenarios (use cases) and metrics (desiderata) for language models.
🎯
Select
Choose 16 core + 26 targeted scenarios and 7 metrics, based on coverage and feasibility.
📏
Measure
Score all 7 metrics on the core scenarios — 87.5% of the scenario × metric cells.
⚖️
Standardize
Same 5-shot prompting and conditions for all 30 models; release every prompt and completion.
🎯 Accuracy
Did the model get the right answer? Task-appropriate scores: exact match for QA, ROUGE for summaries, RR@10 / NDCG@10 for retrieval.
⚖️ Calibration
Is its confidence honest? A calibrated model that says "70% toxic" on 1,000 comments should be right ~700 times. Measured with ECE-10.
🛡️ Robustness
Does it survive the real world? Worst-case accuracy over perturbations of the input — typos, contractions, casing and other benign corruptions.
🧑‍🤝‍🧑 Fairness
Does it treat groups equally? Worst-case accuracy over fairness perturbations (e.g. altering the dialect of the speaker) plus performance disparities across demographic subsets.
🪞 Bias
What do its generations associate? Gender and race representation in model-generated text, drawn from established word lists and measures.
☠️ Toxicity
How often does it generate toxic text? The toxicity rate of generations, scored by PerspectiveAPI on free-text outputs.
⚡ Efficiency
What does it cost to run? Denoised inference time, plus training energy (kWh) and CO₂ where disclosed.
Interactive Demo — The Beyond-Accuracy Dashboard (Centerpiece)

Pick a model, then run the evaluation: its seven metric profiles fill in one by one. Bars are illustrative (~) — the shape of the trade-offs mirrors the paper's findings, and verdicts quote verified numbers where stated.

select a model above
Chapter 04

Evaluation as a Scenario × Metric Matrix

A model's report card is no longer a number — it's a grid. Rows are scenarios, columns are metrics, and each cell is a measured quantity. Trade-offs become visible structure instead of an averaged-away secret.

CORE MATRIX SIZE
16 × 7
scenarios × metric families, measured wherever possible
87.5% of cells actually measured
TOTAL SCENARIOS
42
16 core + 26 targeted scenarios
21 were new to mainstream LM evaluation
MODELS IN THE HARNESS
30
10 open · 17 limited-access · 3 closed, from 12 organizations
GPT-3 family, OPT, BLOOM, GPT-J/NeoX, Cohere, J1, TNLG, Anthropic-LM…
EVALUATION OVERLAP
96.0%
of core scenarios covered per model, up from 17.9% before
All 30 models, same scenarios, same conditions
Interactive Demo — Scenario × Metric Matrix Explorer

Click any cell to reveal the paper's finding at that intersection. ● marks a verified number from the paper; ~ marks a qualitative or illustrative note.

no cell selected — click one above
Adaptation — Few-Shot Prompting, Standardized

Every model is adapted with the same generic 5-shot prompting scheme, pioneered by GPT-3 — no model-specific incantations. HELM is explicit that this is one choice: stronger prompting (chain-of-thought, prompt tuning) could change results, so the harness measures models under shared conditions, not each model's ceiling.

Full Transparency

All raw model prompts and completions are released for every run, alongside a modular toolkit for adding scenarios, models, metrics, and prompting strategies. Every number on the leaderboard can be traced back to an actual prediction you can read yourself.

Chapter 05

Findings — No Free Lunch

HELM surfaces 25 top-level findings. The common thread: no single model dominates every scenario and metric — the interesting behavior lives in the trade-offs.

INSTRUCTION-TUNING WINS ACCURACY
>90%
head-to-head accuracy win rate for text-davinci-002
Best on accuracy, robustness & fairness; Anthropic-LM v4-s3 (52B) top-3 despite 10× smaller than TNLG v2 (530B)
TRUTHFULQA GAP (KNOWLEDGE)
62.0 vs 36.2
text-davinci-002 vs runner-up Anthropic-LM v4-s3
Instruction-tuned models are far more likely to answer truthfully
GSM8K (MATH REASONING)
52.1 / 35.0 / <16
code-davinci-002 / text-davinci-002 / everyone else
Code models beat text models even on natural-language math
ROBUSTNESS CLIFF (NARRATIVEQA)
72.6 → 38.9
TNLG v2 (530B) accuracy before → after perturbations
The 3rd-most accurate model drops 33.7 points under typos & perturbations
CALIBRATION FLIP (HELLASWAG)
0.29 ECE-10
text-davinci-002 (0.286) & Cohere xlarge (0.341) — the two most accurate
There, improving accuracy worsened calibration; on OpenBookQA the two improve together
CIVILCOMMENTS ≈ CHANCE
50.1%
OPT (175B) on toxicity detection — near coin-flip
A top-tier overall model can fail a safety-relevant task it was never tuned for
BBQ BIAS PARADOX
89.5 / 48.4 / 44.9
text-davinci-002 / T0++ / TNLG v2 — the three most accurate on BBQ
Precisely these three show ambiguous-context biases aligned with social biases — the others lean the other way
LANGUAGE MODELING FLIP
open
Pile-trained open models (GPT-NeoX, OPT, BLOOM, GPT-J) lead The Pile & BLiMP
And text-davinci-002 is among the least accurate on BLiMP irregular forms — over-generalized rules?
Who Wins Where — Different Crowns for Different Cells
EvaluationWinner (accuracy)Notable verified result
9 core QA scenariostext-davinci-002Most accurate on all 9; TruthfulQA 62.0% vs 36.2%
Knowledge (WikiFact)text-davinci-00238.5%, TNLG v2 (530B) 34.3%, Cohere xlarge 33.4%
Math reasoning (GSM8K)code-davinci-00252.1% vs 35.0% (text-davinci-002); no other model >16%
Targeted bias (BBQ)text-davinci-00289.5% — next best 48.4% (T0++); but most biased in ambiguous contexts
Retrieval (MS MARCO regular)text-davinci-00239.8% RR@10 boosted, vs 22.4% vanilla BM25 re-ranking
Summarization (XSUM)TNLG v2 (530B)Highest ROUGE-2 — paired with the highest toxicity rate (0.6%)
Toxicity detection (CivilComments)≈ chance for mostOPT (175B) 50.1%; all models ECE-10 ≥ 0.40
Misc classification (RAFT)GLM (130B)85.8% overall; text-davinci-002 only 40.8% on one split vs 97.5% for others
Language modeling (The Pile)open, Pile-trainedGPT-NeoX, OPT (175B), BLOOM, GPT-J, OPT (66B) lead
NarrativeQA end-tasktext-davinci-00274.4% ROUGE-L — a new state of the art over UnifiedQA-v2 (67.4%)

Read the columns, not just the winners: the model crowned on accuracy is near-chance on toxicity detection; the retrieval king is not the summarization king. That is the paper's point.

The Big Pattern — Correlations & Exceptions
  • Accuracy ↔ robustness ↔ fairness: extremely strongly correlated across scenarios — with exceptions (TNLG v2's NarrativeQA crash).
  • Calibration: scenario-dependent; can improve or degrade with accuracy (OpenBookQA vs HellaSwag).
  • Toxicity & bias in generations: largely constant and low on core scenarios — but targeted evals (toxic prompts, BBQ) reveal much larger differences.
  • Accuracy vs efficiency: no strong overall trade-off; only a subset of models sits on the Pareto frontier.
  • Scale: within a family, bigger is reliably better; across families, parameter count is a poor predictor. All clear head-to-head winners are ≥ 50B.
  • Access gap: limited > closed > open on average, but the best open models land within ~5 points of the best non-open ones.
What HELM Did NOT Solve
  • Missing models: no access to PaLM, LaMDA, Gopher, Chinchilla, RETRO — all explicitly listed as not evaluated (contrary to common retellings, Chinchilla and PaLM are not in the 30).
  • Metric ceilings: automated metrics have limits — ROUGE failed to discriminate summarization quality; toxicity rides on PerspectiveAPI's judgments.
  • English-centric: mostly English (a few dialects and varieties), no multilingual coverage.
  • Contested constructs: fairness and bias are complex social constructs — the paper itself flags how they're operationalized as a choice, not a truth.
  • 87.5%, not 100%: some scenario × metric cells were infeasible to measure; some model runs hit technical issues (the remaining ~4% of the 96.0%).
Legacy

Impact — Beyond the Leaderboard

HELM turned evaluation from a marketing event into living infrastructure. Its ethos: transparency, standardization, and trade-offs over single scores.

🏆 The living leaderboard
crfm.stanford.edu/helm keeps evaluating new models and scenarios: a Classic suite preserving the paper's scenarios, cheaper Lite suites, and Capabilities & Safety flagship boards.
🌐 A family of HELMs
The framework spun into domain leaderboards — MedHELM (medicine), SEA-HELM (Southeast Asian languages), VHELM (vision-language) — and a pip-installable toolkit (crfm-helm).
🔍 Transparency norms
Releasing every raw prompt and completion set a new bar: results became auditable artifacts, not press-release numbers. The web UI lets anyone read the predictions behind each score.
🧬 Evaluation lineage
HELM centralized the GLUE-era tasks, TruthfulQA, RealToxicityPrompts and more into one harness — part of the same movement as BIG-bench (2022) and OpenAI's Evals (2023) that made holistic harnesses standard infrastructure.
📊 Multi-metric by default
Calibration, robustness, fairness, bias, toxicity and efficiency became routine reporting dimensions — the "7 metrics" framing shows up across later evaluation suites and safety leaderboards.
🥗 The no-free-lunch lesson
HELM's most-quoted takeaway: no single model dominates — the most accurate models can be the worst calibrated or most biased in context. Benchmark design now exposes the matrix instead of one crown.
Deep Dive

The Leaderboard Depends on Which Column You Read

HELM's contribution isn't 30 models ranked — it's the proof that every single-model ranking hides a trade-off. Under one protocol (same 5-shot prompting, shared scenarios), the accuracy champion loses calibration depending on the scenario, the most accurate models are the most biased on BBQ, and one frontier model loses half its accuracy under spelling perturbations.

🧭
What One Protocol Buys
  • Before HELM: 5-shot and 0-shot, different prompts, disjoint scenarios — cross-paper comparisons were fiction
  • HELM re-ran everything under shared conditions: core-scenario overlap across papers jumped 17.9% → 96.0%
  • 42 scenarios × 7 metric families; 87.5% of core cells measured — accuracy never "falls to the wayside"
  • Every raw prompt and completion released — the transparency bar the field now copies
🥇
The No-Free-Lunch Findings
  • text-davinci-002 dominates accuracy — and its calibration flips by scenario
  • On BBQ, the most accurate models were also the most biased — competence amplifies stereotype defaults
  • TNLG v2: 72.6% clean → 38.9% perturbed — a robustness crash invisible in an accuracy-only table
  • Conclusion the paper draws: report the matrix, not the mean — there is no "best model" in general
Interactive Demo — The Metric Switchboard

Same four models, four metric columns. Switch the metric and watch the winner change — the exact experience of reading HELM's own tables. Every number shown is one of the paper's headline findings.

VERDICT
Holistic is a discipline, not a size
The scenario × metric matrix is the deliverable — means destroy the information that matters. Downstream, the entire evaluation track of PaperMap inherits this stance: LLM-as-a-Judge asks who grades the graders, JudgeBench finds them failing on adversarial pairs, and AgentBench carries the matrix idea into agent environments. A living leaderboard — crfm.stanford.edu/helm — kept the framework current long past the paper.
🔀 17.9% → 96.0%
The single most damning pre-HELM statistic: five leading papers' core scenarios overlapped by under a fifth. Shared protocol made the numbers commensurable for the first time.
⚖️ Competence × bias
On BBQ, accuracy and stereotype bias rose together. Stronger models learn the patterns of the world — including its skew — with higher fidelity. Fairness is not a side effect of scale.
🧪 Perturbation as X-ray
A model that drops 34 points under character-level noise wasn't robust — it was overfit to surface form. Robustness columns caught what leaderboards rewarded hiding.
📜 Full log release
Prompts, completions, failures — all public. HELM made evaluation auditable: any number in the leaderboard can be traced to the exact generations behind it.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the HELM paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ HELM = Holistic Evaluation of Language Models (Stanford CRFM, arXiv Nov 2022; published in TMLR 2023) — a framework, not just a benchmark.
✅ Evaluation is a scenario × metric matrix: 42 scenarios (16 core + 26 targeted) measured with 7 metric families (87.5% of core cells).
✅ The 7 metrics: accuracy, calibration, robustness, fairness, bias, toxicity, efficiency — "metrics beyond accuracy don't fall to the wayside."
✅ 30 models (10 open, 17 limited, 3 closed) from 12 organizations, all run under the same 5-shot prompting — overlap with core scenarios jumps 17.9% → 96.0%.
✅ No free lunch: instruction-tuned text-davinci-002 dominates accuracy, yet calibration flips by scenario, BBQ's most accurate models are its most biased, and TNLG v2 crashes 72.6% → 38.9% under perturbations.
✅ Legacy: a living leaderboard at crfm.stanford.edu/helm, an open toolkit, and every raw prompt and completion released — "benchmarking beyond the leaderboard."