History Problem Core Idea Gaps Results Impact Quiz Takeaways
Interactive Paper Explainer

Run the Code,
Don't Read It
HumanEval

The Codex paper redefined code evaluation: judge programs by executing their unit tests, not by surface similarity — and discovered that sampling many solutions beats sampling one, 28.8% → 70.2%.

Start Learning Read the Paper ↗
164
Hand-written problems
28.8%
Codex pass@1
70.2%
pass@100 (sampling)
2021
Chen et al. (OpenAI)
History

From Text Metrics to Running Programs

How code generation stopped being graded like translation.

2020-21
Surface-metric era
Code models graded by BLEU/edit-distance against reference code — fluency confused with correctness.
2021 · pre-Codex
GPT-3 on code
General LLMs write plausible-looking Python that often fails at runtime — the evaluation can't tell the difference.
Jul 2021
🚀 Codex + HumanEval
Chen et al.: fine-tune GPT on GitHub code; release 164 hand-written problems (docstring → program) graded by executed unit tests and the pass@k metric.
2021-22
The Copilot era
A production Codex descendant powers GitHub Copilot; pass@k becomes the code-eval standard; every code model reports HumanEval.
2024+
Contamination-era audits
HumanEval saturates and leaks into training sets — LiveCodeBench (entry #71) redesigns for freshness, keeping the functional-correctness gospel.
Correctness Is Execution

The paper's twin contributions fit together: (1) The benchmark — 164 original Python problems, each a function signature + docstring + example test cases, written BY the authors so they cannot be memorized from GitHub; scoring = fraction of generated programs passing all tests. (2) The pass@k metric — sample k programs per problem; pass@k = probability at least one passes; unbiased estimator via n choose k combinatorics. Together they replaced string-matching with semantics: a solution that reorders lines, renames variables, or restructures entirely still scores 1 if it works.

Chapter 01

Programs That Look Right

The evaluation gap the paper closed.

📝
The Text-Metric Trap
  • BLEU/edit-distance reward similarity to one reference — penalize correct solutions that differ stylistically
  • Buggy code and correct code can be near-identical in surface form (an off-by-one is one character)
  • Code benchmarks built from GitHub invite memorization — the model reproduces the reference, not reasoning
  • No principled way to credit 'the model could solve it sometimes' — single-sample scoring is luck-sensitive
▶
The Execution Answer
  • Functional correctness: run the program against test cases — semantics, not surface
  • 164 hand-written problems (not scraped) — original signatures, docstrings, tests
  • pass@k with an unbiased estimator — samples n ≥ k, computes the true probability at least one of k passes
  • Codex 12B: 28.8% pass@1 → 70.2% with 100 samples — sampling as capability amplifier
Analogy — The Driving Test

Grading code by text similarity is judging drivers by how closely their steering motions mimic the instructor's — elegant mimicry, crashes ignored. HumanEval is the actual driving test: here's a route (docstring), drive it (execute), pass or fail. Your hands can be anywhere — the examiner only watches the road.

Chapter 02

The pass@k Metric

The estimator every code benchmark now uses.

Definition
  • Sample n ≥ k programs per problem; count correct ones (c)
  • pass@k = E[ 1 − C(n−c, k) / C(n, k) ] — the unbiased probability that at least one of k randomly-chosen samples passes
  • Avoids the bias of "sample k once and score" — luck averaged out combinatorially
The sampling discovery
  • Codex 12B single-sample: 28.8% of problems solved
  • With 100 samples + reranking by a log-prob + execution filter: 70.2%
  • Repeated sampling is "surprisingly effective" for hard prompts — the pass@k curve's slow decay
  • Early evidence for best-of-N/test-time scaling (entries #49, #58)
The Model Side: Codex

Codex fine-tunes GPT on publicly available GitHub code (Python focus), with careful deduplication against evaluation sets. The paper's engineering chapters cover the knobs that mattered: stop sequences, docstring formats, temperature-vs-pass@k tradeoffs (higher temperature helps high-k sampling, hurts pass@1), and the docstring→function task format that made GitHub Copilot a product. A distinct production version of Codex powered Copilot at launch — the benchmark-to-industry pipeline in one sentence.

Interactive Demo — Surface Metrics vs Execution

Tab through the same generated program under both grading regimes — the case that broke BLEU.

Chapter 03

What the Benchmark Couldn't See

The limitations later papers attacked — stated plainly.

Known Gaps (acknowledged and later exploited)
Interactive Demo — Compute pass@k, Correctly

The unbiased estimator in action — why 'sample n, compute combinatorially' beats 'sample k and score'.

Chapter 05

28.8% → 70.2%: Sampling Is Capability

The paper's second discovery proved as durable as its first.

PASS@1 · CODEX 12B
28.8%
vs GPT-3's 0%, GPT-J's 11.4%
PASS@100
70.2%
100 samples + reranking — repeated sampling works
METRIC
pass@k
unbiased estimator, now universal
PRODUCT
Copilot
a production Codex powers GitHub Copilot at launch
Interactive Demo — The 2021 Leaderboard

Press run for the original HumanEval standings — the spread that launched a product category.

Modelpass@1pass@8pass@100
GPT-3 (175B, no code FT)0%~1%—
GPT-J (6B)11.4%~18%—
Codex (12B)28.8%46.8%70.2%
Codex (12B, reranked)~35%——

Figures from the paper's tables (reranking + careful sample selection push aggregate results further; 70.2% is with 100 samples per problem).

Legacy

Legacy — Execution Forever

Every code benchmark since is HumanEval's child.

⚙ Functional correctness doctrine
Execution-based grading became the unquestioned standard — MBPP, DS-1000, LiveCodeBench (entry #71), and SWE-bench (entry #76) all inherit the test-passing verdict.
📊 pass@k as universal metric
The unbiased estimator moved from this paper into every sampling-based eval — including non-code domains (best-of-N, test-time compute — entry #58).
🚀 Copilot and the copilot economy
The paper shipped a product: a production Codex powered GitHub Copilot, launching the AI-code-assistant industry.
🎲 The sampling insight
'Repeated sampling is surprisingly effective' foreshadowed self-consistency (entry #49) and inference scaling (entry #58) — one of the field's most-replicated observations.
⚠️ What it did NOT solve
164 short single-file problems saturated quickly; visible test examples flatter the model; contamination once public; and real software engineering (multi-file, dependency, issue-driven) needed SWE-bench to enter the conversation.
🛤 Read next
The code-eval lineage: LiveCodeBench · SWE-bench · LLM-as-a-Judge
Test Yourself

Quick Quiz

Check your understanding of the key concepts from HumanEval.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ HumanEval = 164 hand-written docstring→program problems, graded by executed unit tests.
✅ Functional correctness replaced BLEU: semantics over surface — the field's permanent standard.
✅ pass@k with the unbiased combinatorial estimator — the metric of all sampling-based evaluation.
✅ Codex 12B: 28.8% pass@1, 70.2% with 100 samples + selection — sampling as capability amplifier.
✅ A production Codex powered GitHub Copilot at launch.
✅ Read it as the paper that made code evaluation executable and code models deployable.