LiveBench refreshes its questions monthly from sources newer than every model — and scores them against objective ground truth: no LLM judges, no crowdsourcing, contamination-limited by design.
The evaluation-freshness crisis had two patches; LiveBench rejected both.
LiveBench's design contract: (1) frequently updated — new questions released monthly (later monthly-ish cadences), sourced from recent material: math competition problems, recent arXiv papers, current news, datasets published after known cutoffs; (2) objective scoring — every task has a ground-truth answer checked by rule (math verification, string/logic checks, exact match) — no LLM judge, no human rater, no preference model in the loop; (3) versioned releases — the whole leaderboard re-runs per release, making cross-time comparisons explicit instead of hand-waved.
The two patched solutions each failed a different way.
A static benchmark is a year-end exam everyone has the answers to. LLM-judged fresh benchmarks are grading essays with a opinionated intern. LiveBench is a monthly financial audit: new transactions each month (recent sources), arithmetic that a calculator checks (ground truth), and every prior month re-verified — an auditor would call the design 'controls'.
Three constraints, each attacking a documented failure.
Objective ground truth is also an anti-gaming choice: with a verifier function as the grader, the only way to score is to actually solve — there is no judge to flatter, no style to optimize, no verbosity premium. Combined with release windows (the LiveCodeBench mechanism, entry #71, applied to general domains), the benchmark removes both failure families at once: memorization (freshness) and eval-hacking (objectivity). The paper frames itself as "contamination-limited" — honest wording: no live benchmark is contamination-proof, but dated windows bound the exposure.
What running LiveBench continuously actually looks like.
A leaderboard that can neither be memorized nor charmed.
| Evaluation approach | Freshness | Grading objectivity | Update cadence |
|---|---|---|---|
| Static benchmarks (MMLU-era) | low — rots | high (fixed answers) | never |
| LLM-judged fresh sets | high | low — judge biases | moderate |
| Human-crowdsourced | high | medium — noisy | slow |
| LiveBench | high (dated sources) | objective — rule-based | monthly |
The two-axis fix: freshness without sacrificing objectivity — the quadrant the field was missing.
LiveBench became the default 'is this release actually better?' check.
Check your understanding of the key concepts from LiveBench.
Everything you need to remember about this paper.