History Problem MT-Bench Judging Biases Arena Impact Quiz
Interactive Paper Explainer

Can a Judge Be a Model?
LLM-as-a-Judge

When chat answers have no single ground truth, how do you score them? This paper built MT-Bench and Chatbot Arena, showed that strong LLM judges match human preferences — and then exposed their systematic biases.

Start Learning Read the Paper ↗
80
MT-Bench Questions
8
Task Categories
85%
GPT-4–Human Agreement
2023
Year Published
History

From Answer Keys to Judges

LLM-as-a-Judge landed after a decade of borrowed evaluation habits. Here is how chat evaluation finally found tools of its own.

2020–2022
The reference-metrics era
BLEU, ROUGE, and exact match rule NLP evaluation. They compare a candidate against one gold reference — fine for translation and QA, meaningless for open-ended chat.
2022
Human evaluation — gold, but slow
Annotator panels rate answers and remain the gold standard. But they are slow, expensive, and hopeless at the scale of millions of model pairs.
2022–2023
Ad-hoc GPT-4 grading
Practitioners quietly start asking GPT-4 to grade model outputs. Convenient and nearly free — but nobody has rigorously checked whether it agrees with humans.
2023 · Jun
🚀 MT-Bench + Chatbot Arena (Zheng et al.)
The first systematic validation of LLM judges: a multi-turn benchmark, a live crowdsourced arena, agreement measured against humans — and a bias autopsy.
2023 →
Arena-style leaderboards everywhere
Chatbot Arena's anonymous-battle format becomes the reference leaderboard of the field; clones and variants multiply across labs.
2023+
Judging becomes infrastructure
Judge-based evals are standard practice now — deployed with known biases, so "judge robustness" becomes a research area of its own.
Key Insight

Chat changed the shape of evaluation. Classification and QA ask "is this answer right?" — a question with an answer key. Chat asks "which of these two answers would a person prefer?" — a question with no key at all. The paper's move: pick a judge, measure it against humans first, neutralize its known biases, then let it scale.

TWO DIFFERENT QUESTIONS
QA   : question → one gold answer → exact match ✓ / ✗
CHAT: question → answer A vs answer B → which is better? ⚖
The second line has no answer key — it needs a judge.
Chapter 01

Grading Without an Answer Key

Chat models produce open-ended answers — often several correct ones for the same question. Nearly every evaluation tool ever built assumes there is one right answer to compare against.

📉
The Evaluation Dead End
  • Open-ended answers have no reference text to compare against
  • BLEU / ROUGE overlap barely correlates with chat quality
  • Human panels are accurate but slow, expensive, unscalable
  • Millions of possible model pairs — nobody can annotate them all
  • Ad-hoc GPT-4 grading was cheap — but completely unvalidated
⚖️
The Judge Protocol
  • Ask a strong LLM to compare two answers, directly
  • Validate the judge against human preference first
  • Thousands of comparisons in minutes, not months
  • Neutralize known biases — swap orders, watch lengths
  • Aggregate crowd votes into stable Elo-style rankings
Analogy — Math Test vs. Essay Contest

A math test ships with an answer key: grade by exact match, done. An essay contest has no key — every finalist is a "correct" essay — so you hire judges. But you do not hand out medals before checking the judges themselves: do they reward length? play favorites? grade differently in the afternoon? That is precisely what this paper does with LLM judges — validate the judge before trusting the leaderboard.

MATH TEST  :  answer key exists → exact match ✓
ESSAY CONTEST :  no key → pairwise preference ⚖ → validate the judge → then trust it
Chapter 02

MT-Bench — 80 Questions, 8 Arenas

A compact, deliberately open-ended benchmark: 80 multi-turn questions across 8 categories, built to do one thing — separate strong chat models from weak ones.

MT-Bench at a Glance
80
Questions
expert-written, open-ended
8
Categories
writing → humanities
2
Turns Each
question + follow-up
🎯
Discriminative
strong models visibly win
Design Rule 1 — Open-Ended

Every question admits many good answers — by design. If a gold answer existed, you would use exact match and skip the judge entirely. Writing tasks, debates, explanations: quality is a matter of degree, not of correctness, and the benchmark must stay that way to resemble real chat.

Design Rule 2 — Discriminative + Multi-Turn

Questions must separate models: weak answers fail visibly while strong ones shine. The follow-up turn does double duty — it probes whether the model kept track of the conversation across turns, not just whether it can answer a one-shot prompt.

Interactive Demo — MT-Bench Category Explorer

All 8 categories, one click each. Every MT-Bench question is two turns — an opening task plus a follow-up that tests memory and consistency. (Samples below are written in the paper's style.)

Chapter 03

Three Ways to Ask a Judge

The paper formalizes LLM judging into three protocols, then asks the only question that matters: does the judge agree with humans?

Protocol A — Single-Answer Grading

Show the judge one answer plus a rubric; ask for a 1–10 score with a justification. Cheap and simple — but absolute scores drift: the same answer can be a 6 today and an 8 tomorrow.

Protocol B — Pairwise Comparison workhorse

Show two answers to the same question and ask which is better: A, B, or tie. Comparing is cognitively easier than absolute scoring — verdicts are more reliable and match how humans themselves prefer to evaluate.

Protocol C — Reference-Guided

Include a reference answer as an anchor for the verdict. Helps when a good gold answer exists — which, for open-ended chat, is exactly what does not exist.

What the Judge Actually Sees (Protocol B)
[SYSTEM]
"You are an impartial judge. Evaluate which assistant answer is better and justify your choice…"
[QUESTION]
"…one of the 80 MT-Bench questions…"
[ASSISTANT A] <first answer>  ·  [ASSISTANT B] <second answer>
[JUDGE OUTPUT] { "verdict": "A", "reason": "…" }

One prompt, two answers, one structured verdict — plus the reasoning the judge writes for itself. That "reason" field is exactly where the biases of the next chapter hide.

GPT-4 · PAIRWISE JUDGE
85%
agreement with human raters
the headline number of the paper
HUMAN ↔ HUMAN
81%
inter-human agreement
the bar any judge must clear
GPT-3.5 · PAIRWISE JUDGE
≈66%
agreement with humans (no reference)
≈66–81% depending on protocol — notably less reliable
So… Fire the Annotators?

Not yet. 85% is the average, and it hides two catches: agreement is task-dependent (judges track humans closely on writing and roleplay, but slip exactly where correctness matters most — math and coding), and an average can hide systematic errors. A judge can be right 85% of the time and still flip many verdicts whenever you swap answer order — errors that wash out in the agreement score but corrupt any leaderboard built on them. Averages do not catch biases; controlled experiments do. That is the next chapter.

Chapter 04

The Biases — Judge, Know Thyself

Validating the judge was half the paper. The other half was a controlled autopsy of the ways LLM judges systematically misjudge.

Bias 1 — Position Bias 🔁

Present the same two answers in both orders — the verdict flips. LLM judges lean toward whichever answer occupies the favored slot (often the one listed second). Detect it by swapping: if the verdict changes, that was bias, not judgment.

Bias 2 — Verbosity Bias 📜

Longer answers get rated as more helpful even when the extra text adds nothing — or buries an error. Judges read effort as quality: "it engages more thoroughly" is often just "it is longer."

Bias 3 — Self-Enhancement 🪞

Judges prefer answers from their own model family. Judging GPT-4 vs. GPT-3.5 battles, GPT-4's preference for GPT-4 answers is measurably stronger than an independent judge's — though the authors note genuine quality gaps are hard to fully separate out.

The Fixes — Position-Balanced Judging
Interactive Demo — The Position-Bias Simulator

One question, two fixed answers: Response A (short and correct) and Response B (long and vague). Run the judge in each order, then the balanced way. Verdicts, scores, and "reasoning" are precomputed to illustrate the paper's findings — the position effect here is a +2 score bonus for whichever answer is listed second.

judge-console · same question · same answers · only the order changes
USER QUESTION
"Why does water boil at a lower temperature at high altitude?"

Run 1 and Run 2 gave the same pair two different winners — that flip is the tell. The paper's stricter variant marks such battles "inconsistent" rather than counting a win; averaging the two runs is the gentler fix shown here.

Chapter 05

Chatbot Arena — Judgment by Crowd

MT-Bench is a fixed exam. Chatbot Arena is the wild: real users, real prompts, anonymous battles, and a leaderboard computed from millions of votes.

1 · 🥊 Anonymous battle
A real user types any prompt. Two randomly chosen models answer side by side — names hidden until after the vote.
2 · 🗳️ The vote
The user picks A, B, or tie. No rubric, no scores — pure preference, exactly what reference metrics cannot measure.
3 · 📊 The ranking
Millions of battles feed a Bradley-Terry / Elo-style fit: each model earns a strength score from who-beats-whom, not raw win counts.
4 · ✅ Sanity check
The crowd ranking correlates with expert preference — aggregated, the crowd judges like the experts, but never sleeps and never runs out of ballots.
P(model i beats model j) = 1 / (1 + e−(θᵢ − θⱼ))
θᵢ
Model strength
One latent score per model, fit from every battle it fought — not a simple win count.
θᵢ − θⱼ
Strength gap
The score difference converts directly into a win probability for any pairing of models.
1/(1+e⁻ˣ)
Sigmoid
Squashes any gap into a 0–100% probability — the Bradley-Terry / Elo-style workhorse.
Rank by θ
Leaderboard
Votes → pairwise outcomes → strengths. More votes tighten the estimates; battle order does not matter.
Two Complementary Machines
MT-BenchChatbot Arena
Prompts80 fixed, expert-writtenLive user prompts, unbounded
TurnsAlways 2 (question + follow-up)Whatever the user does next
JudgeLLM (validated vs. humans first)The crowd — real humans
ScaleMinutes per model pairMillions of battles, continuously
OutputPairwise win-rates / scoresElo-style leaderboard
Weak spotJudge biases (Ch. 04)Needs enormous vote volume
Interactive Demo — Mini Arena: You Are the Annotator

Three live battles between two anonymous models — Model ◆ and Model ◇. Read the prompt, compare the answers, cast your vote; the reveal comes at the end. (Answers prewritten and illustrative.)

BATTLE 1 / 3
Legacy

Impact — Everything Is a Leaderboard Now

Two artifacts — a benchmark and a leaderboard — quietly became the field's evaluation infrastructure, and "judge robustness" became a research topic.

🏆 The leaderboard
Chatbot Arena became THE leaderboard of the field — the page teams refresh on release day, and the crowd every new model must face.
📝 MT-Bench as a standard
80 questions became a routine quick eval — "MT-Bench score" is now a standard line in papers and model cards.
🧪 Judge-based evals everywhere
AlpacaEval, Arena-Hard, and a whole lineage of LLM-judge benchmarks descend directly from this paper's protocols.
🕵️ Judge robustness research
The three named biases made "can we trust the judge?" its own research area — bias taxonomies, judge meta-evals, judge ensembles.
⚖️ Judges need judges
The meta-lesson of the paper: every judge now requires its own validation against humans before its verdicts count.
🌍 Continuous evaluation
Pairwise preference + crowd votes proved you can measure quality continuously on real traffic — not once a quarter in a lab.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the LLM-as-a-Judge paper.

Reference

Key Takeaways

Everything you need to remember about judging LLMs with LLMs.

✅ Open-ended chat breaks reference metrics — evaluation becomes a preference question, not a matching question.
✅ MT-Bench: 80 multi-turn questions, 8 categories, engineered to discriminate strong models from weak ones.
✅ Three protocols: single-answer grading, pairwise comparison (the reliable one), and reference-guided judging.
✅ GPT-4 judges agree with humans 85% — on par with human–human agreement (~81%). GPT-3.5 is notably less reliable.
✅ Know the big three biases — position, verbosity, self-enhancement — and swap answer orders to neutralize position bias.
✅ Chatbot Arena: anonymous crowd battles → millions of votes → a Bradley-Terry leaderboard that tracks expert preference.