Static code benchmarks rot as they enter training data. LiveCodeBench continuously collects new contest problems — and scores every model only on problems published AFTER its training cutoff.
The contamination story, told through code evaluation.
The core mechanism is a join on dates: every collected problem carries its contest publication timestamp; every model carries its (known or estimated) training cutoff. Scoring a model means computing its accuracy on problems published after that cutoff — guaranteed unseen. The second mechanism, prompt-style decoupling: the same problem is presented under multiple wrappers (conversation/instruction/code-completion) and formats — revealing that a model's "capability" varies embarrassingly with presentation, and averaging it out of the measurement.
The two diseases: contamination and prompt-format luck.
HumanEval is a crossword book from 2021 — champions have memorized it, and a perfect score means nothing. LiveCodeBench is the daily newspaper crossword: today's puzzle didn't exist when anyone studied, every solver gets the same fresh date, and the champion is whoever solves TODAY'S — with the paper's date stamp as the referee.
Time windows and prompt decoupling — the evaluation machinery.
Beyond function synthesis: self-repair (fix a failing solution given the error), code execution (predict outputs), and test generation (write tests for a spec) — the four capabilities a real coding assistant needs, each under the same fresh-problem regime. The windowed leaderboard then supports honest comparisons across model generations: performance measured only where the test could not have been seen.
What the windowed view immediately revealed.
The benchmark's own findings justified its design.
| Mechanism | What it controls | Disease it cures |
|---|---|---|
| Continuous collection | problem supply never freezes | benchmark aging |
| Time-windowed scoring | models see only post-cutoff problems | contamination |
| Prompt decoupling | wrapper × format grids averaged | presentation luck |
| 4 capability families | synthesis, repair, execution, testing | narrow 'can write a function' tunnel vision |
The paper's contribution is evaluation infrastructure — each mechanism maps to one documented failure of static code benchmarks.
The release-window join became the anti-rot standard across benchmark design.
Check your understanding of the key concepts from LiveCodeBench.
Everything you need to remember about this paper.