A visual, step-by-step guide to the technical report behind GPT-4 — a multimodal model whose performance could be predicted from small-scale runs, passing professional exams at human level while revealing the infrastructure science of frontier training.
GPT-4 is the point where frontier training became an engineering discipline with predictive laws.
It is a system report, not a methods paper: capabilities and benchmarks are documented in detail while architecture, parameter count, and dataset construction are deliberately withheld — competitive and safety reasons given. The transferable science is the predictability infrastructure: loss curves, capability metrics, and safety behaviors that extrapolate across three orders of magnitude of compute.
Training a model for months on massive clusters is the most expensive bet in software. GPT-4's central problem: make the bet predictable before you place it.
You don't test a new aircraft by building the full plane and hoping. You fly thousands of scaled model flights, learn the physics, and predict the big machine's behavior — then you build it. GPT-4's small-model families are the wind tunnel; the final run is the first flight.
GPT-4 accepts interleaved image and text input and produces text output — a Transformer-based model pretrained to predict the next token across both modalities.
The report's most-cited science: final-loss and capability predictions fitted on runs using up to 1/1,000th of GPT-4's compute.
The report's signature move: evaluate on human exams — the uniform, adversarially-designed test suite society already maintains.
| Exam | GPT-4 (overall) | GPT-3.5 | Human bottom–top |
|---|---|---|---|
| Uniform Bar Exam | ~90th percentile | ~10th percentile | 5th–95th |
| GRE Quantitative | 163–166 range (~80th percentile) | ~25th percentile | — |
| AP Biology | 85–100th percentile | 62–75th | — |
| AP English Lit & Composition | ~60th percentile | ~10–20th | — |
| USABO Semifinal (bio olympiad) | ~87th percentile | ~31st | — |
| AMC 12 (math competition) | ~30th–90th percentile range | ~1st–10th | — |
Weak areas remain: English literature, complex math competitions, and any task needing sustained original derivation. "Human-level on exams" is not general human expertise.
Pretraining buys capability; post-training buys behavior. The report documents an RLHF pipeline inherited from InstructGPT, plus safety-driven evaluation culture.
Every frontier model report since borrows a piece of this document's structure.
The report's deepest tension: the more precisely you can predict training, the less you can infer about what the model will mean to people.
Read the GPT-4 report as two documents stapled together: an engineering triumph (capability you can forecast like weather) and a governance experiment (powerful capability released with new norms of documentation and secrecy). The engineering scaled; the governance questions are still open — which is exactly why evaluation papers like MMLU and safety frameworks matter as much as the models they measure.
Check your understanding of the key concepts from the GPT-4 Technical Report.
Everything you need to remember about this paper.