When chat answers have no single ground truth, how do you score them? This paper built MT-Bench and Chatbot Arena, showed that strong LLM judges match human preferences — and then exposed their systematic biases.
LLM-as-a-Judge landed after a decade of borrowed evaluation habits. Here is how chat evaluation finally found tools of its own.
Chat changed the shape of evaluation. Classification and QA ask "is this answer right?" — a question with an answer key. Chat asks "which of these two answers would a person prefer?" — a question with no key at all. The paper's move: pick a judge, measure it against humans first, neutralize its known biases, then let it scale.
Chat models produce open-ended answers — often several correct ones for the same question. Nearly every evaluation tool ever built assumes there is one right answer to compare against.
A math test ships with an answer key: grade by exact match, done. An essay contest has no key — every finalist is a "correct" essay — so you hire judges. But you do not hand out medals before checking the judges themselves: do they reward length? play favorites? grade differently in the afternoon? That is precisely what this paper does with LLM judges — validate the judge before trusting the leaderboard.
A compact, deliberately open-ended benchmark: 80 multi-turn questions across 8 categories, built to do one thing — separate strong chat models from weak ones.
Every question admits many good answers — by design. If a gold answer existed, you would use exact match and skip the judge entirely. Writing tasks, debates, explanations: quality is a matter of degree, not of correctness, and the benchmark must stay that way to resemble real chat.
Questions must separate models: weak answers fail visibly while strong ones shine. The follow-up turn does double duty — it probes whether the model kept track of the conversation across turns, not just whether it can answer a one-shot prompt.
The paper formalizes LLM judging into three protocols, then asks the only question that matters: does the judge agree with humans?
Show the judge one answer plus a rubric; ask for a 1–10 score with a justification. Cheap and simple — but absolute scores drift: the same answer can be a 6 today and an 8 tomorrow.
Show two answers to the same question and ask which is better: A, B, or tie. Comparing is cognitively easier than absolute scoring — verdicts are more reliable and match how humans themselves prefer to evaluate.
Include a reference answer as an anchor for the verdict. Helps when a good gold answer exists — which, for open-ended chat, is exactly what does not exist.
Not yet. 85% is the average, and it hides two catches: agreement is task-dependent (judges track humans closely on writing and roleplay, but slip exactly where correctness matters most — math and coding), and an average can hide systematic errors. A judge can be right 85% of the time and still flip many verdicts whenever you swap answer order — errors that wash out in the agreement score but corrupt any leaderboard built on them. Averages do not catch biases; controlled experiments do. That is the next chapter.
Validating the judge was half the paper. The other half was a controlled autopsy of the ways LLM judges systematically misjudge.
Present the same two answers in both orders — the verdict flips. LLM judges lean toward whichever answer occupies the favored slot (often the one listed second). Detect it by swapping: if the verdict changes, that was bias, not judgment.
Longer answers get rated as more helpful even when the extra text adds nothing — or buries an error. Judges read effort as quality: "it engages more thoroughly" is often just "it is longer."
Judges prefer answers from their own model family. Judging GPT-4 vs. GPT-3.5 battles, GPT-4's preference for GPT-4 answers is measurably stronger than an independent judge's — though the authors note genuine quality gaps are hard to fully separate out.
MT-Bench is a fixed exam. Chatbot Arena is the wild: real users, real prompts, anonymous battles, and a leaderboard computed from millions of votes.
| MT-Bench | Chatbot Arena | |
|---|---|---|
| Prompts | 80 fixed, expert-written | Live user prompts, unbounded |
| Turns | Always 2 (question + follow-up) | Whatever the user does next |
| Judge | LLM (validated vs. humans first) | The crowd — real humans |
| Scale | Minutes per model pair | Millions of battles, continuously |
| Output | Pairwise win-rates / scores | Elo-style leaderboard |
| Weak spot | Judge biases (Ch. 04) | Needs enormous vote volume |
Two artifacts — a benchmark and a leaderboard — quietly became the field's evaluation infrastructure, and "judge robustness" became a research topic.
Check your understanding of the key concepts from the LLM-as-a-Judge paper.
Everything you need to remember about judging LLMs with LLMs.