History Problem Core Idea Scaling Benchmarks Training Impact Deep Dive Quiz
Interactive Paper Explainer

The Predictable Frontier
GPT-4

A visual, step-by-step guide to the technical report behind GPT-4 — a multimodal model whose performance could be predicted from small-scale runs, passing professional exams at human level while revealing the infrastructure science of frontier training.

Start Learning Read the Paper ↗
86.4%
MMLU (5-shot)
Top 10%
Simulated Bar Exam
1/1000
Compute for Prediction
2023
Year Published
History

From Scale to Infrastructure Science

GPT-4 is the point where frontier training became an engineering discipline with predictive laws.

2020
GPT-3 & scaling laws (Kaplan et al.)
175B parameters prove in-context learning — and power-law curves suggest scale predicts performance.
2022 · Mar
InstructGPT / RLHF matures
Human preference tuning turns raw LMs into assistants — the post-training recipe GPT-4 would inherit.
2022 · Nov
ChatGPT
Instruction + RLHF on a GPT-3.5-class model goes viral — the world now knows what an assistant model is.
2023 · Mar
🚀 GPT-4 (OpenAI)
Multimodal (image + text in, text out), 8K/32K context variants, human-level professional exams — and performance predicted from 1/1000th-compute runs.
2023 →
The frontier era
Multimodal assistants, agentic tool use, and eval-driven development — every lineage traces a design decision back to this report.
What Makes This Report Unusual

It is a system report, not a methods paper: capabilities and benchmarks are documented in detail while architecture, parameter count, and dataset construction are deliberately withheld — competitive and safety reasons given. The transferable science is the predictability infrastructure: loss curves, capability metrics, and safety behaviors that extrapolate across three orders of magnitude of compute.

🧭 Study pairing
Read after Scaling Laws — GPT-4 is those laws, industrialized.
Chapter 01

Why Frontier Runs Gamble

Training a model for months on massive clusters is the most expensive bet in software. GPT-4's central problem: make the bet predictable before you place it.

🎲
The Frontier Gamble
  • A frontier run costs months of cluster time — a failed run is catastrophic
  • Models at scale fail in ways small models don't: loss spikes, instabilities, divergence
  • Raw capability is useless if behavior is unsafe or misaligned with users
  • You cannot A/B test a training run — you have to predict it
📐
The Predictability Stack
  • Train families of small models on the same data at 1/1000th the compute
  • Fit loss and capability curves that extrapolate across 4-5 orders of magnitude
  • Predict final loss — and specific benchmark behavior — before committing
  • Pair the prediction with an alignment stage (RLHF) so capability arrives controllable
Analogy — The Flight Simulator

You don't test a new aircraft by building the full plane and hoping. You fly thousands of scaled model flights, learn the physics, and predict the big machine's behavior — then you build it. GPT-4's small-model families are the wind tunnel; the final run is the first flight.

Chapter 02

One Model, Two Senses

GPT-4 accepts interleaved image and text input and produces text output — a Transformer-based model pretrained to predict the next token across both modalities.

🖼 Image + text in
Screenshots, documents, diagrams, photos — the model reasons over visual inputs, not just captions them.
✍️ Text out
Generation stays textual — the interface is a chat, but the input is a page.
📏 8K & 32K contexts
Two variants: 8,192 and 32,768-token windows — the 32K model could hold ~50 pages of text at once, then unheard-of for a chat model.
🎯 Next-token pretraining
Still the core objective — scale, data mixture, and post-training do the rest.
Interactive Demo — Two Senses, One Prompt

Text-only vs multimodal exam questions: toggle the input and watch what becomes answerable.

Chapter 03

Predictable Scaling

The report's most-cited science: final-loss and capability predictions fitted on runs using up to 1/1,000th of GPT-4's compute.

L(C) fits across ~104–105× compute ranges  ·  capability metrics extrapolate too
L(C)
Loss vs compute
Smooth curves fitted on small runs predicted GPT-4's final loss to high accuracy.
1/1,000
Cheap predictions
The computing budget of prediction models is up to three orders of magnitude smaller.
ICL curve
Capability extrapolation
Not just loss: benchmark-style performance (from a suite of internal problems) also extrapolates across scale.
Infra rules
Stability at scale
Loss spikes at scale were predicted and preempted — large runs need engineered interventions, not luck.
Beyond Loss — Training Infrastructure
  • Optimized kernels: custom implementations for throughput on the hardware of the era
  • Loss spike playbook: predicted spike rates, and bail-out strategies (data batch skipping, learning-rate surgery)
  • Data mixture control: carefully staged domain mixtures — a lever Kaplan-era scaling underweighted
What Was NOT Predictable
  • Qualitative jumps: downstream usefulness and emergent behaviors are still not smooth functions
  • Alignment outcomes: RLHF results were tuned by iteration, not by law
  • Capability risk: predictable loss ≠ predictable safety profile — the motivation for red-teaming and the system card
Interactive Demo — The Prediction Ladder

Fit the curve on small runs (1/1000 compute), then extrapolate to the full GPT-4 run. Press Run.

Chapter 04

Professional Exams

The report's signature move: evaluate on human exams — the uniform, adversarially-designed test suite society already maintains.

Human-Level Exam Performance (simulated, no tools)
ExamGPT-4 (overall)GPT-3.5Human bottom–top
Uniform Bar Exam~90th percentile~10th percentile5th–95th
GRE Quantitative163–166 range (~80th percentile)~25th percentile—
AP Biology85–100th percentile62–75th—
AP English Lit & Composition~60th percentile~10–20th—
USABO Semifinal (bio olympiad)~87th percentile~31st—
AMC 12 (math competition)~30th–90th percentile range~1st–10th—

Weak areas remain: English literature, complex math competitions, and any task needing sustained original derivation. "Human-level on exams" is not general human expertise.

MMLU (5-shot)
86.4%
vs 70.0% GPT-3.5 — across 57 subjects
FACTUALITY (internal)
+40%
reduction in hallucination vs GPT-3.5 on internal adversarial factuality eval
SENSITIVITY
↓
calibration and instruction-following improved by RLHF post-training
CONTEXT
32K
tokens in the long-context variant (~50 pages)
Interactive Demo — Percentile Machine

Place GPT-4, GPT-3.5, and a human median on the bar-exam percentile ladder. Click each candidate.

Chapter 05

Post-Training & Alignment

Pretraining buys capability; post-training buys behavior. The report documents an RLHF pipeline inherited from InstructGPT, plus safety-driven evaluation culture.

RLHF Post-Training
  • SFT-style behavior seeding from demonstrations (per the InstructGPT lineage)
  • Preference models trained from human comparisons; RL optimization against them
  • Result: improved factuality, instruction-following, and steerability versus the raw pretrained model
  • Refusal behavior, calibration, and policy compliance are tuned — and also benchmarked
Safety Evaluation Culture
  • Internal and external red-teaming before release; a separate system card documents risks
  • Domain-qualified testers for dangerous-capability areas (e.g., bio, cyber, persuasion)
  • Truthfulness and hallucination measured on adversarial internal suites — the +40% improvement claim
  • Steering and "jailbreak" behavior: post-training narrows, but does not eliminate, misuse surface
Legacy

Impact — The Template for Frontier Reports

Every frontier model report since borrows a piece of this document's structure.

📊 Exam-based evals
Human exams as a public, comparable, adversarially-designed yardstick — now standard in every release.
🧪 Predictable-scaling discipline
Small-model extrapolation is now table stakes for frontier labs (and the seed of the Chinchilla-vs-Kaplan compute debate).
🛡 System cards & red-teaming
Separating capability documentation from risk documentation became an industry norm.
🖼 Multimodal assistants
Image+text chat set the interface expectations for the entire assistant market.
🔒 The secrecy debate
Withholding architecture/data sparked a lasting open-vs-closed argument — and motivated open reproductions.
⚠️ What it did NOT solve
Hallucination persists, long-context ≠ long-memory, and benchmarks saturate — the story continued in evaluation research.
Deep Dive

Predictable Loss, Unpredictable Meaning

The report's deepest tension: the more precisely you can predict training, the less you can infer about what the model will mean to people.

📉
What Curves Can't Tell You
  • Loss predicts loss — not hallucination rates, refusals, or misuse potential
  • Exam percentiles measure test-taking, not judgment or originality
  • Emergent behaviors (good and bad) arrive without a curve announcing them
  • Societal impact depends on deployment context, invisible to any training metric
🔭
Why Prediction Still Wins
  • Predictable training de-risks the biggest capital bets in AI
  • It enables safety work before scale-up: risky-capability evals run on small proxies
  • Scaling transparency (even without architecture disclosure) lets outsiders reason about trends
  • The predictable part funds the careful part — you can't red-team what you can't afford to train
Interactive Demo — Predicted vs Surprise

Sort report outcomes into "fitted in advance" vs "discovered after training" — the honest split of what scaling science covers.

Verdict

Read the GPT-4 report as two documents stapled together: an engineering triumph (capability you can forecast like weather) and a governance experiment (powerful capability released with new norms of documentation and secrecy). The engineering scaled; the governance questions are still open — which is exactly why evaluation papers like MMLU and safety frameworks matter as much as the models they measure.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the GPT-4 Technical Report.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Multimodal: image + text in, text out — still next-token pretrained underneath.
✅ 86.4% MMLU (5-shot) vs 70.0% for GPT-3.5, across all 57 subjects.
✅ Simulated bar exam: ~top 10% of test takers (GPT-3.5: ~bottom 10%).
✅ Performance predicted from models trained with ≤1/1,000th of the compute — including loss spikes.
✅ RLHF post-training improved factuality (internal adversarial suite, ~40% fewer hallucinations) and steerability.
✅ Architecture, size, and data deliberately withheld — the report documents behavior, not the recipe.