History Problem Build Domains Findings Methods Impact Deep Dive Quiz
Interactive Paper Explainer

Judging the Judges
JudgeBench

A visual, step-by-step guide to the paper that stress-tests LLM judges — 350 adversarial response pairs where one answer is subtly, objectively wrong, and where even GPT-4o performs only slightly better than random guessing.

Start Learning Read the Paper ↗
350
Response Pairs
4
Domains: Knowledge · Reasoning · Math · Code
80.9%
Top Judge Accuracy (o3-mini, high)
~85%
Judge–Human Agreement on Easy MT-Bench (for contrast)
History

From Graders to Judges

LLM judges were legitimized by one number — agreement with humans — and then quietly became infrastructure. JudgeBench is the stress test that arrived late.

2022
Automatic metrics hit their ceiling
BLEU/ROUGE measure surface overlap, not correctness; human evaluation is slow and expensive. LLMs begin quietly grading text in ad-hoc internal experiments.
2023 · Jun
🚀 MT-Bench & "LLM-as-a-Judge" (Zheng et al.)
GPT-4 judging pairwise battles agrees with humans ~85% of the time — about as often as humans agree with each other (~81%). LLM judging is legitimized overnight.
2023
Chatbot Arena & the bias autopsy
Crowdsourced battles + Elo leaderboards go live. AlpacaEval, FairEval and the MT-Bench authors dissect judge biases: position, length, self-enhancement.
2024
Judges become infrastructure
Arena-Hard, AlpacaEval 2.0, fine-tuned judges (Prometheus 2, JudgeLM, AutoJ, Skywork) and reward models now rank models, filter data, and drive RLHF.
2024 · Oct
🎯 JudgeBench (Tan et al.)
Flips the question: can a judge that "agrees with humans" actually verify hard, objective answers? 350 adversarial pairs say mostly no — GPT-4o barely beats random.
2025 →
The judge-quality research line
Reasoning models (o1, o3-mini, DeepSeek-R1) lift judging to 73–81%; JudgeBench v2 tracks them. Judging quality is now a first-class research topic, not a footnote.
Key Insight

A judge validated on easy questions is an unvalidated judge. The paper's framework ranks what a judge must check, in strict order:

THE JUDGING HIERARCHY
1 · follows the instruction?
2 · factually & logically correct? ← JudgeBench lives here
3 · style & preference
Existing benchmarks mostly measured agreement about (3). JudgeBench isolates (2) — where correctness is checkable and errors are subtle.
Chapter 01

The Problem with Easy Questions

LLM judges were validated by agreement with humans — on open-ended chat questions where correctness was never really tested. JudgeBench asks what happens when it is.

📝
Judges Validated by Agreement
  • MT-Bench, LLMEval and FairEval score a judge by matching human preference
  • Their questions are open-ended and easy — writing, roleplay, chat
  • Crowd annotators cannot reliably grade hard math, proofs, or code
  • So ~85% agreement mostly measures shared taste in style, not verification
  • As models improve, judges face answers harder than themselves — unmeasured
🎯
JudgeBench's Answer
  • Start from hard problems with ground truth and automatic verifiers
  • Both answers written by the same model — style and length neutralized
  • The wrong answer hides a subtle, deliberate, verified error
  • Labels are objective: correct vs incorrect, not "preferred"
  • Position bias controlled: every pair judged twice, order swapped
Analogy — Grading Homework ≠ Grading Competition Problems

A teacher grading weekend worksheets can lean on neatness, length, and tone — correctness is rarely in doubt, and "agrees with the other teacher 85% of the time" sounds great. The same teacher grading olympiad entries faces write-ups that are fluent, well-structured, and wrong in step 4. Confidence built on easy grading evaporates exactly when correctness becomes the whole job. An agreement score on easy questions says almost nothing about the second job — that is the gap JudgeBench measures.

Interactive Demo — Play the Judge: Easy vs Adversarial (JudgeBench-style)

Four rounds. Round 1 is the easy chat regime where judge agreement numbers are earned. Rounds 2–4 are JudgeBench-style: one answer is subtly, objectively wrong. Pick the better response — then see the ground truth.

Chapter 02

Building an Adversarial Benchmark

JudgeBench's pipeline converts any dataset with ground-truth labels and a verification algorithm into judge-evaluation pairs. The whole recipe fits in one line.

Verified dataset → sample k answers (GPT-4o) → keep mixed outcomes → pair (1 ✓ + 1 subtle ✗)
Q + V
Hard question + verifier
MMLU-Pro, LiveBench, LiveCodeBench — every question ships with a ground-truth answer and a checking algorithm (string match, exact answer, unit tests).
k
Same-model samples
GPT-4o answers each question k times. One generator for both candidates: no style tells, no length gaps, no self-enhancement confound.
≥1 ✓ ∧ ≥1 ✗
Keep mixed questions
If every sample is right (or every one wrong), the question is dropped. What remains is hard for the generator itself — so the pair is hard to judge.
(✓, ✗)
The final pair
One correct + one subtly incorrect response, labeled by the verifier — not by humans. Same voice, same length: 562 vs 561 tokens on average.
Why Both Answers from the Same Model?

The obvious design — different models write A and B — gives judges shortcuts: style fingerprints, length gaps, and self-enhancement bias (judges favor their own outputs). Sampling both candidates from a single generator leaves correctness as the only reliable signal. A second LLM (GPT-4o-mini) double-checks "incorrect" verdicts that were really just formatting mismatches, and those are filtered out too.

The Verdict Rule

Each pair is judged twice, with the response order swapped. A tie, or a verdict that flips between the two trials, counts as incorrect — a judge that is guessing or order-sensitive gets no credit. Only consistently correct verdicts score.

✓ ✓
both trials correct → point
✓ ✗
flip-flop → wrong
= =
tie twice → wrong
Where the Questions Come From
MMLU-Pro
Knowledge: 12,032 college-level exam questions across 14 disciplines, up to 10 options each.
LiveBench
Reasoning (Big-Bench Hard, Zebra Puzzles) and Math (competition problems like AMC12 / USAMO).
LiveCodeBench
Coding: 300+ contest problems from LeetCode, AtCoder, Codeforces — refreshed to avoid contamination.
350 pairs
The final split: 154 Knowledge · 98 Reasoning · 56 Math · 42 Coding — comparable in size to MT-Bench (80), FairEval (80), LLMBar (419).
Chapter 03

Four Domains of Pain

Each category targets a different failure mode of judging. In every pair, the wrong answer is polished enough to win on style alone — only verification tells them apart.

🧠 Knowledge 154 pairs

Source: MMLU-Pro — college-level multiple choice across 14 disciplines.

  • Trap: a confident, plausible, factually wrong claim — often echoing the most common misconception in the training data.
  • Tiny example: "Largest desert on Earth?" — "The Sahara, 9.2M km²" (polished ✗) vs "Antarctica, ~14M km²" (✓).
  • Judges here: 35–68% — the widest model spread after Reasoning.
🧩 Reasoning 98 pairs

Source: LiveBench Reasoning — Big-Bench Hard tasks and Zebra Puzzles.

  • Trap: a constraint quietly violated mid-derivation — the chain looks airtight until you re-check the premises.
  • Tiny example: a zebra-puzzle write-up that "the coffee drinker owns the zebra" — contradicting a clue it restated correctly two lines earlier.
  • Judges here: 34.7–89.8% — the single most separating category.
🔢 Mathematics 56 pairs

Source: LiveBench Math — competition problems (AMC12, USAMO).

  • Trap: dropped negative roots, sign slips, unproven edge cases — each step locally valid, the conclusion incomplete or wrong.
  • Tiny example: (x+3)² = 25 → "x = 2" (✗) vs "x = 2 or x = −8" (✓).
  • Judges here: oddly kind — GPT-4o scores 75.0, its best category, and judges beat their own solvers.
💻 Coding 42 pairs

Source: LiveCodeBench — contests from LeetCode, AtCoder, Codeforces.

  • Trap: off-by-one bounds, mutation while iterating, code that matches its comments but not the spec.
  • Tiny example: a well-documented loop that starts at 0 and stops one step early — passing the eyeball test, failing the test suite.
  • Judges here: the graveyard — Gemini-1.5-pro 26.2%, GPT-4o 59.5%; every model judges code worse than it writes it.
Interactive Demo — Spot the Subtle Error

This is the skill JudgeBench actually tests. Two snippets; in each, exactly one line is wrong. Click the buggy line — then see how the paper's judges fared on that domain.

Chapter 04

Findings — Judges Struggle

On JudgeBench, models that "agree with humans ~85% of the time" land near coin-flips. All numbers below are overall accuracies from the paper's tables (random guessing = 50%).

VANILLA JUDGE · GPT-4o
50.9%
"which is better?" prompt — ≈ random guessing
ARENA-HARD JUDGE · GPT-4o
56.6%
reference-answer-first prompting, +5.7 over vanilla
BEST GENERAL-PURPOSE · CLAUDE-3.5-SONNET
64.3%
top score among the five models cross-checked on prior benchmarks
BEST OVERALL · O3-MINI (HIGH)
80.9%
test-time compute — 100.0% on Coding pairs
TOP REWARD MODEL · SKYWORK-REWARD-GEMMA-2-27B
64.3%
a specialized verifier matching Claude-3.5-Sonnet
MULTI-AGENT DEBATE · CHATEVAL
34.0%
debating the verdict does not verify it — below random
The 85% Illusion — Agreement vs Accuracy
MT-Bench agreement (GPT-4, easy chat)
~85%
JudgeBench accuracy (GPT-4o, Arena-Hard)
56.6%

A ~28-point gap between "agrees with humans on easy chat" and "verifies correctness on hard pairs" (red tick = 50%, random). JudgeBench also separates models strongly: 31 points between the best (Claude-3.5-Sonnet 64.3) and the worst (Claude-3-Haiku 33.1) of the five models cross-checked against prior benchmarks — comparable to LLMBar: Adversarial (33). The strongest model on JudgeBench was the lowest top score of all five benchmark sets.

Judge Families on JudgeBench (overall accuracy, %)
Judge familyRepresentativeOverallNote
PromptedArena-Hard judge (GPT-4o)56.57vs 50.86 with the vanilla prompt
PromptedVertexAI Evaluation (Gemini-1.5-pro)44.57below random guessing
Fine-tunedSkywork (Llama-3.1-70B)57.43best fine-tuned judge; +5 over its base model
Fine-tunedPandaLM13.14picked "tie" in 479 of 700 judgments
Multi-agentChatEval (debate)34.00debate does not produce verification
Reward modelSkywork-Reward-Gemma-2-27B64.29matches Claude-3.5-Sonnet; family range 59–64
Reasoning modelo3-mini (high), Arena-Hard prompt80.86test-time compute is the biggest lever

Random guessing = 50. Most fine-tuned judges (Prometheus 2, JudgeLM, AutoJ) land between 25 and 41 — below random — often by emitting ties or malformed verdicts.

Interactive Demo — Judge Scoreboard (animated, from the paper's Table 2)

Every judge below uses the same Arena-Hard prompt — only the underlying model changes. Toggle the category to re-rank by domain; the blue bar is the easy-question agreement context, and the red tick marks 50% (random).

Scores from Table 2 of the paper (v2). Reasoning models — o3-mini, o1, DeepSeek-R1 — are the only judges above 70.

Chapter 05

What Improves Judging

JudgeBench is not just a diagnosis — its leaderboard points at the levers that actually move judging accuracy. Four of them work; one is mostly a mirage.

📝 Prompting: reference-first helps a bit
The Arena-Hard judge drafts its own answer first, then compares both candidates against it — lifting GPT-4o from 50.9% to 56.6%. The vanilla "which is better?" prompt leaves it at coin-flip.
🧠 Test-time compute: the big lever
Reasoning models that think before answering dominate: o3-mini (high) 80.9, o3-mini (med) 76.6, o1-preview 75.4, DeepSeek-R1 73.1. The paper calls scaling test-time compute "a promising path" for judging.
🎓 Judge fine-tuning: works, sometimes
Skywork's fine-tuned judges gain +12 points over their Llama-3.1-8B base (53.4 vs 40.9). But most fine-tuned judges sink below random: PandaLM 13.1, JudgeLM 25.1–35.7 — verdict-format drift is a real failure mode.
🏅 Specialized verifiers: reward models
Reward models score 59.4–64.3, rivaling much larger LLM judges. Skywork's Llama-3.1-8B reward model (62.3) crushes its own base model (40.9): a weak model can be trained to judge stronger ones.
Is Verifying Easier Than Solving? (paper's Table 4)
ModelSolver accuracyJudge accuracyJudge − Solver
GPT-4o54.5756.57+2.0
Claude-3.5-Sonnet64.5764.29−0.3
Llama-3.1-405B-Instruct57.7156.86−0.9
Gemini-1.5-pro40.2947.14+6.9

For a fixed model, judging accuracy closely mirrors solving accuracy — the "judging is easy" assumption does not survive. Two twists: in Math, judges beat their own solvers (spotting a sign slip is easier than avoiding one); in Coding, solvers beat judges for every model — code is easier to write than to check. A judge cannot reliably grade work that is harder than what it can solve.

⚠️
What JudgeBench Did NOT Solve
  • No judge clears ~81% — roughly one verdict in five is still wrong
  • Pairs are generated by GPT-4o, biasing the split against GPT-4o judges: Claude-3.5-Sonnet drops from 64.3% to 44.8% on self-generated pairs
  • 350 pairs is modest; Knowledge dominates the mix (154 of 350)
  • Subjective quality — tone, style, helpfulness — is deliberately out of scope
🧭
Why Those Limits Are OK
  • When Knowledge was augmented from 154 to 770 pairs, judge rankings stayed the same
  • The pipeline is generator-agnostic: swap in any model and rebuild the split
  • Same-size peers: MT-Bench (80), FairEval (80), LLMBar (419)
  • Objective ground truth is the point — subjective taste already has benchmarks
Legacy

Impact — Judge Quality Becomes a Field

JudgeBench turned "can we trust the judge?" from a worry into a measurable research question — with consequences for leaderboards, RLHF, and benchmark design.

🔬 A meta-evaluation research line
JudgeBench complements RewardBench, whose reasoning subsets (PRM Math, HumanEvalPack) are saturated at up to ~97% — likely from training contamination. An unsaturated, verifier-labeled judge benchmark was missing.
⚖️ The verifier bottleneck
The paper cites repeated-sampling results: scaling generation only pays off with an oracle-level verifier. If judges stay weak, they become the limiting factor in the whole AI scaling loop.
🧮 Judging ≈ solving
The ablation killed the assumption that LLM judges can cheaply grade work harder than themselves — verifying is roughly as hard as solving, and on code it is harder.
📉 Recalibrating agreement numbers
AlpacaEval- and Arena-Hard-style pipelines that lean on LLM judges now have a stress test: ~85% agreement on easy chat says little about 56.6% accuracy on hard objective pairs.
🛠️ A recipe for meta-benchmarks
Any dataset with ground truth and a verification algorithm can become judge-eval data — the pipeline turns solvers' difficulty into judges' difficulty, automatically and contamination-free.
⚠️ LLM-as-judge, with caution
The practical rule: before trusting a judge in your domain, test it on your hard, checkable cases. Agreement with humans may only be agreement on style.
Deep Dive

Judging ≈ Solving — and Sometimes Harder

JudgeBench's cleanest experiment is its design: both answers in every pair come from the same model, verified correct or subtly wrong, with length and style neutralized (≈562 vs ≈561 tokens). Remove every surface heuristic a judge could exploit, and what's left measures exactly one thing: can the judge work out which answer is right?

🧮
A Benchmark That Can't Be Gamed by Style
  • 350 pairs from MMLU-Pro, LiveBench, LiveCodeBench: Knowledge 154 · Reasoning 98 · Math 56 · Coding 42
  • Same-source pairs kill style confounds — no "longer = better", no "my family writes like this"
  • Double-judged with order swapped; ties and flip-flops count as wrong — positional shortcuts closed
  • Result: a judge score that tracks a real capability, not a preference
📉
What the Capability Actually Looks Like
  • GPT-4o: 50.9% vanilla, 56.6% with Arena-Hard prompt — barely above a coin flip
  • Only test-time reasoning holds up: o3-mini (high) 80.9% · o1-preview 75.4% · DeepSeek-R1 73.1%
  • Correlation that stings: judge accuracy closely tracks solver accuracy on the same content
  • In Coding, judging is harder than solving for every model tested
Interactive Demo — The Style Trap, Live

A math pair in the JudgeBench style: one answer verified correct, one subtly wrong — similar length, similar tone. First watch a vibe-judging pass (what GPT-4o-level judges do under pressure), then switch to step-verify mode (what reasoning judges do). Same pair, opposite verdicts.

PROBLEM (MATH-56 STYLE) · x² − 5x + 6 = 0 · which roots?
ANSWER A · 118 tokens
"Factoring: (x−2)(x−3) = 0. The equation has two real roots, x = 2 and x = 3, both satisfying the original equation — substitution checks out on both."
ANSWER B · 121 tokens
"Factoring: (x−2)(x−4) = 0. The roots are x = 2 and x = 4 — a clean factorization of the quadratic with both solutions verified by inspection."
VERDICT
Hire the judge that shows its work
The lever that moved judging was not a new model family or a better prompt — it was test-time compute. o3-mini reaching 80.9% while GPT-4o sits at 50.9% on identical pairs is the cleanest demonstration in the eval literature that verification is a reasoning task. The upstream story is in the LLM-as-a-Judge guide (the 85% agreement era); the holistic view of what a judge must pass is HELM; where judges must then go is agentic — AgentBench.
🎯 Why same-model pairs
Both answers from GPT-4o means no self-enhancement bias, no family style to favor. The judge can only lean on content — which is the point.
🔁 Order-swap protocol
Each pair judged twice; a tie or a flip counts wrong. Models that win by sitting in the favored slot get zero free wins here.
🧑‍💻 Coding inverts the order
Judging code is harder than writing it for every model tested — reading someone else's logic for a subtle bug is a distinct skill from producing correct logic.
📐 The correlation
Judge accuracy tracks solver accuracy category by category. "Can grade it" ≈ "can do it" — the finding that reframed judges as solvers with a formatting job.
Test Yourself

Quick Quiz

Check your understanding of the key ideas from the JudgeBench paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ JudgeBench = 350 verified response pairs: Knowledge 154, Reasoning 98, Math 56, Coding 42 — built from MMLU-Pro, LiveBench, and LiveCodeBench.
✅ Both answers come from the same model (GPT-4o): one verified correct, one subtly wrong — style and length neutralized (≈562 vs ≈561 tokens).
✅ GPT-4o scores 50.9% (vanilla) to 56.6% (Arena-Hard) — barely above random, versus ~85% human agreement on easy MT-Bench-style questions.
✅ Reasoning models lead: o3-mini (high) 80.9%, o1-preview 75.4%, DeepSeek-R1 73.1% — test-time compute is the biggest lever for judging.
✅ Judging ≈ solving: judge accuracy closely tracks solver accuracy; in Coding, judging is harder than solving for every model tested.
✅ Verdict protocol: each pair judged twice with order swapped — ties and flip-flops count as wrong, closing the door on positional-bias shortcuts.