History Problem Core Idea Findings Results Impact Quiz Takeaways
Interactive Paper Explainer

Fresh Problems,Timestamped
LiveCodeBench

Static code benchmarks rot as they enter training data. LiveCodeBench continuously collects new contest problems — and scores every model only on problems published AFTER its training cutoff.

Start Learning Read the Paper ↗
3
Contest platforms
Time-windowed
Scoring
4
Code capabilities
2024
Jain et al.
History

Benchmarks Have Expiration Dates

The contamination story, told through code evaluation.

2021
HumanEval sets the standard
Execution-based code eval arrives (entry #67) — 164 problems, instantly canonical, instantly leaking into corpora.
2022-23
Saturation + suspicion
Models score 80-90%+ on HumanEval; nobody can say how much is capability versus memorization — the leaderboard loses meaning.
Mar 2024
🚀 LiveCodeBench
Jain et al.: continuous collection from LeetCode, AtCoder, Codeforces — problems time-stamped by contest date; models evaluated only within their release-window; prompt-style decoupling (wrapper + format modules).
2024+
Contamination forensics
The benchmark's date windows make contamination VISIBLE — models showing suspiciously strong performance on pre-cutoff problems are exposed; training-set audits follow its methodology.
2025+
Live evals everywhere
LiveBench (entry #72), live QA, live agents — the "dated release windows" pattern becomes the anti-rot standard.
Evaluation in Release Windows

The core mechanism is a join on dates: every collected problem carries its contest publication timestamp; every model carries its (known or estimated) training cutoff. Scoring a model means computing its accuracy on problems published after that cutoff — guaranteed unseen. The second mechanism, prompt-style decoupling: the same problem is presented under multiple wrappers (conversation/instruction/code-completion) and formats — revealing that a model's "capability" varies embarrassingly with presentation, and averaging it out of the measurement.

Chapter 01

Leaderboards Rot

The two diseases: contamination and prompt-format luck.

🗓
Static-Benchmark Decay
  • Published problems enter training corpora — scores inflate silently
  • Contamination is invisible in aggregate scores: you cannot tell knowledge from memorization
  • Saturation: 2023+ models ace HumanEval — no headroom, no signal
  • Every 'new benchmark' restarts the same decay clock from zero
🔄
The Live Answer
  • Continuous collection from LeetCode, AtCoder, Codeforces — fresh problems forever
  • Time-stamped problems → score models only on post-cutoff items
  • Contamination becomes measurable: pre-vs-post-cutoff performance gaps expose it
  • Prompt-style decoupling: wrapper × format grids separate capability from presentation luck
Analogy — The Daily Newspaper Crossword

HumanEval is a crossword book from 2021 — champions have memorized it, and a perfect score means nothing. LiveCodeBench is the daily newspaper crossword: today's puzzle didn't exist when anyone studied, every solver gets the same fresh date, and the champion is whoever solves TODAY'S — with the paper's date stamp as the referee.

Chapter 02

The Two Mechanisms

Time windows and prompt decoupling — the evaluation machinery.

Time-windowed scoring
  • Problems continuously collected from three contest platforms
  • Each carries its contest date; models carry their training cutoffs
  • Scoring = accuracy on post-cutoff problems — unseen by construction
  • Drift measurement: the same model scored across date windows tracks true vs. memorized skill
Prompt-style decoupling
  • Wrappers: conversation, instruction, completion — presentation variants of the same problem
  • Format modules control test-visible details and output parsing
  • Revealed: identical models swing notably across wrappers — a hidden variable in every benchmark
  • Reporting the grid (not one lucky cell) makes code eval presentation-robust
Holistic coverage

Beyond function synthesis: self-repair (fix a failing solution given the error), code execution (predict outputs), and test generation (write tests for a spec) — the four capabilities a real coding assistant needs, each under the same fresh-problem regime. The windowed leaderboard then supports honest comparisons across model generations: performance measured only where the test could not have been seen.

Interactive Demo — Same Model, Different Question Ages

Tab through scoring regimes — static benchmarks vs windowed — and watch contamination appear and disappear.

Chapter 03

The Findings

What the windowed view immediately revealed.

Observations from the Paper
Interactive Demo — One Problem's Journey into the Benchmark

Follow a contest problem from platform to windowed, decoupled evaluation item.

Chapter 05

Contamination, Finally Visible

The benchmark's own findings justified its design.

DESIGN
time windows
post-cutoff-only scoring by construction
EXPOSED
pre/post gaps
contamination made measurable across models
REVEALED
prompt swings
wrapper choice materially moves scores
CAPABILITIES
4 families
synthesis · repair · execution · testing
Interactive Demo — The Four Capabilities

Function synthesis is only one of four. Press reveal for what a real coding model must do.

MechanismWhat it controlsDisease it cures
Continuous collectionproblem supply never freezesbenchmark aging
Time-windowed scoringmodels see only post-cutoff problemscontamination
Prompt decouplingwrapper × format grids averagedpresentation luck
4 capability familiessynthesis, repair, execution, testingnarrow 'can write a function' tunnel vision

The paper's contribution is evaluation infrastructure — each mechanism maps to one documented failure of static code benchmarks.

Legacy

Legacy — Dated Evaluation as Doctrine

The release-window join became the anti-rot standard across benchmark design.

🔄 The live-eval pattern
Continuous collection + timestamp joins spread beyond code — LiveBench (entry #72), live QA, agent evals adopt the same decay-proofing contract.
🔍 Contamination forensics
Making pre/post gaps measurable turned contamination from suspicion to evidence — methodology used in training-set audits across the field.
🎭 Presentation-robust reporting
Wrapper × format grids exposed prompt-luck in ALL prior code leaderboards; multi-condition reporting became the credibility norm.
🗺 The frontier map
Scoring repair/execution/testing revealed where code models actually lag — steering the next generation of code-model training.
⚠️ What it did NOT solve
Contest problems skew algorithmic — less real-repo engineering (SWE-bench's domain, entry #76); contest distribution drifts over time; and English/Chinese contest coverage dominates.
🛤 Read next
The eval stack: HumanEval · LiveBench · SWE-bench
Test Yourself

Quick Quiz

Check your understanding of the key concepts from LiveCodeBench.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Continuous collection from LeetCode/AtCoder/Codeforces — the problem supply never freezes.
✅ Time-windowed scoring: models evaluated only on post-cutoff problems — unseen by construction.
✅ Contamination became measurable: pre/post gaps expose training-set leakage.
✅ Prompt-style decoupling: wrapper × format grids end presentation luck in code leaderboards.
✅ Four capabilities: synthesis, self-repair, execution prediction, test generation — repair/testing lag.
✅ Read it as the benchmark that made evaluation renew itself — the anti-rot covenant.