The Codex paper redefined code evaluation: judge programs by executing their unit tests, not by surface similarity — and discovered that sampling many solutions beats sampling one, 28.8% → 70.2%.
How code generation stopped being graded like translation.
The paper's twin contributions fit together: (1) The benchmark — 164 original Python problems, each a function signature + docstring + example test cases, written BY the authors so they cannot be memorized from GitHub; scoring = fraction of generated programs passing all tests. (2) The pass@k metric — sample k programs per problem; pass@k = probability at least one passes; unbiased estimator via n choose k combinatorics. Together they replaced string-matching with semantics: a solution that reorders lines, renames variables, or restructures entirely still scores 1 if it works.
The evaluation gap the paper closed.
Grading code by text similarity is judging drivers by how closely their steering motions mimic the instructor's — elegant mimicry, crashes ignored. HumanEval is the actual driving test: here's a route (docstring), drive it (execute), pass or fail. Your hands can be anywhere — the examiner only watches the road.
The estimator every code benchmark now uses.
Codex fine-tunes GPT on publicly available GitHub code (Python focus), with careful deduplication against evaluation sets. The paper's engineering chapters cover the knobs that mattered: stop sequences, docstring formats, temperature-vs-pass@k tradeoffs (higher temperature helps high-k sampling, hurts pass@1), and the docstring→function task format that made GitHub Copilot a product. A distinct production version of Codex powered Copilot at launch — the benchmark-to-industry pipeline in one sentence.
The limitations later papers attacked — stated plainly.
The paper's second discovery proved as durable as its first.
| Model | pass@1 | pass@8 | pass@100 |
|---|---|---|---|
| GPT-3 (175B, no code FT) | 0% | ~1% | — |
| GPT-J (6B) | 11.4% | ~18% | — |
| Codex (12B) | 28.8% | 46.8% | 70.2% |
| Codex (12B, reranked) | ~35% | — | — |
Figures from the paper's tables (reranking + careful sample selection push aggregate results further; 70.2% is with 100 samples per problem).
Every code benchmark since is HumanEval's child.
Check your understanding of the key concepts from HumanEval.
Everything you need to remember about this paper.