History Problem Core Idea Operations Results Impact Quiz Takeaways
Interactive Paper Explainer

A Benchmark with a Shelf Life
LiveBench

LiveBench refreshes its questions monthly from sources newer than every model — and scores them against objective ground truth: no LLM judges, no crowdsourcing, contamination-limited by design.

Start Learning Read the Paper ↗
Monthly
New questions
0
LLM judges
Versioned
Releases
2024
White et al.
History

Judges and Crowds Have Biases

The evaluation-freshness crisis had two patches; LiveBench rejected both.

2023
The contamination crisis
Benchmark leakage documented across the field; leaderboard gains unattributable between capability and memorization.
2023-24
Patch 1: LLM judges
Crowdsource fresh prompts, grade with strong LLMs — introduces judge bias, position bias, and breakdown on hard questions.
2023-24
Patch 2: human crowdsourcing
Fresh human evals — expensive, noisy, not reproducible, and slow to update.
Jun 2024
🚀 LiveBench
White et al.: monthly question releases from recent sources (competitions, papers, news), objective scoring with ground truth, versioned releases — the first benchmark designed to resist contamination AND judging pitfalls together.
2024-25
The live standard
LiveBench leaderboards track every major release; its versioning pattern (entry #71's sibling) becomes the credibility baseline for general benchmarks.
Three Constraints, One Benchmark

LiveBench's design contract: (1) frequently updated — new questions released monthly (later monthly-ish cadences), sourced from recent material: math competition problems, recent arXiv papers, current news, datasets published after known cutoffs; (2) objective scoring — every task has a ground-truth answer checked by rule (math verification, string/logic checks, exact match) — no LLM judge, no human rater, no preference model in the loop; (3) versioned releases — the whole leaderboard re-runs per release, making cross-time comparisons explicit instead of hand-waved.

Chapter 01

Fresh but Ungradeable

The two patched solutions each failed a different way.

⚖
The Freshness-Grading Dilemma
  • Static benchmarks rot; fresh crowdsourced prompts arrive ungradeable without judges
  • LLM judges bring position, verbosity, and self-preference biases — and break down exactly on hard questions
  • Human grading: expensive, noisy, non-reproducible — and slowest to update
  • Result: the field chose between contaminated-but-objective and fresh-but-biased
🎯
The LiveBench Answer
  • Monthly question releases from sources newer than every model (competitions, papers, news)
  • Task design for ground truth: math with checkable answers, logic, parsing, structured extraction
  • Objective scoring only — verification functions, not judges
  • Versioned releases with full re-scoring — the leaderboard as a time series, not a snapshot
Analogy — The Monthly Audit

A static benchmark is a year-end exam everyone has the answers to. LLM-judged fresh benchmarks are grading essays with a opinionated intern. LiveBench is a monthly financial audit: new transactions each month (recent sources), arithmetic that a calculator checks (ground truth), and every prior month re-verified — an auditor would call the design 'controls'.

Chapter 02

The Design Contract

Three constraints, each attacking a documented failure.

1️⃣ Frequently updated
New questions released monthly (contamination-limited: sourced from recent math competitions, papers, and news — newer than any model's cutoff).
2️⃣ Objective scoring
Ground-truth answers checked programmatically — math verification, exact/logic checks, structured matching. No LLM judge's taste, no crowd's noise.
3️⃣ Versioned releases
Each release re-scores the model set; comparisons carry version labels — the leaderboard behaves like a time series with documented vintages.
🔀 Difficulty preserved
Task families span math, reasoning, language understanding, data analysis, coding — fresh but deliberately non-trivial, resisting the "easy fresh questions" failure mode.
The anti-gaming layer

Objective ground truth is also an anti-gaming choice: with a verifier function as the grader, the only way to score is to actually solve — there is no judge to flatter, no style to optimize, no verbosity premium. Combined with release windows (the LiveCodeBench mechanism, entry #71, applied to general domains), the benchmark removes both failure families at once: memorization (freshness) and eval-hacking (objectivity). The paper frames itself as "contamination-limited" — honest wording: no live benchmark is contamination-proof, but dated windows bound the exposure.

Interactive Demo — The Grading Spectrum

Tab through grading mechanisms — and their specific failure modes. LiveBench picks the rightmost.

Chapter 03

The Leaderboard as a Service

What running LiveBench continuously actually looks like.

Operations
Interactive Demo — One Release Cycle

Follow a monthly release from sourcing to re-scored leaderboard — the operational loop.

Chapter 05

Objective and Alive

A leaderboard that can neither be memorized nor charmed.

RELEASE CADENCE
monthly
questions newer than every model
JUDGES
zero
programmatic ground-truth scoring only
VERSIONS
tracked
leaderboard as a time series
MODELS
open + closed
ranked per release, reproducibly
Interactive Demo — Contamination-Proof vs Contamination-Limited

The paper's careful wording — 'limited', not 'proof'. Press reveal for why the honesty matters.

Evaluation approachFreshnessGrading objectivityUpdate cadence
Static benchmarks (MMLU-era)low — rotshigh (fixed answers)never
LLM-judged fresh setshighlow — judge biasesmoderate
Human-crowdsourcedhighmedium — noisyslow
LiveBenchhigh (dated sources)objective — rule-basedmonthly

The two-axis fix: freshness without sacrificing objectivity — the quadrant the field was missing.

Legacy

Legacy — The Credible Leaderboard

LiveBench became the default 'is this release actually better?' check.

📅 The live-eval standard
Monthly versioned releases with objective scoring set the operational template — the same contract as LiveCodeBench (entry #71), proven in the general domain.
⚖ Judge-free credibility
By removing LLM judges and crowds, LiveBench dodged the bias literature (entry #69 documents the diseases) — objective grading as an anti-eval-hacking measure.
🔁 Release-time reality checks
Frontier model launches now cite LiveBench current-vintage scores — the industry's sanity check against contamination-inflated claims.
🧭 Task-design discipline
'Design questions WITH checkable answers' spread as methodology — constraint as quality control across fresh benchmarks.
⚠️ What it did NOT solve
Checkable-answer tasks constrain coverage (creativity, open-ended reasoning stay hard); closed models' training data remains unauditable; and cadence-limited freshness still trails truly continuous collection.
🛤 Read next
The eval house: LLM-as-a-Judge · LiveCodeBench · GPQA
Test Yourself

Quick Quiz

Check your understanding of the key concepts from LiveBench.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ LiveBench: monthly question releases, objective ground-truth scoring, versioned leaderboards.
✅ No LLM judges, no crowds — rule-based verification removes judge bias and eval-hacking surface.
✅ Questions sourced newer than every model cutoff — contamination-limited by construction.
✅ Task design constraint: if it can't be verified by rule, it doesn't ship.
✅ Versioning turns leaderboards into time series — trends, not snapshots.
✅ Read it with LiveCodeBench as the two halves of the live-eval standard: code and general.