History Problem Core Idea Inputs Results Impact Quiz Takeaways
Interactive Paper Explainer

The First GPT
Generative Pre-Training

Before BERT, before the scaling race, a quiet 2018 result set the template for everything after: pre-train a decoder on unlabeled text, then fine-tune with a task-agnostic input format.

Start Learning Read the Paper ↗
117M
Parameters
12
Decoder layers
2
Training stages
2018
OpenAI report
History

Pre-Training Becomes the Default

Where NLU stood in 2018, and how one recipe changed the training economics of the whole field.

2013
Word2vec
Mikolov's static embeddings: one vector per word, learned without labels — semantics, but no context.
2016-17
Task-specific deep NLP
Every benchmark shipped its own architecture and its own labeled data; transfer meant pretrained word vectors.
Jun 2017
The Transformer
Attention Is All You Need delivered a parallelizable, long-range-friendly backbone — trained as a translator.
Jun 2018
🚀 GPT-1
OpenAI couples a 12-layer decoder-only Transformer with generative pre-training on BooksCorpus and a task-agnostic input format.
Oct 2018+
The recipe spreads
BERT generalizes pre-training further; GPT-2 and GPT-3 push the same decoder line toward scale and zero-shot behavior.
Why 'Generative' Mattered

Supervised data was the bottleneck of 2018 NLP: label thousands of examples per task, retrain per task. GPT-1's bet was that predicting the next token on unlabeled books forces syntax, discourse, and long-range dependency into the weights — and that a small slice of supervised data could then steer those weights. The decoder choice also proved prophetic: the same objective, scaled up, becomes few-shot and zero-shot learning.

Chapter 01

One Model per Task

The 2017-18 status quo that GPT-1 attacked: architectures, data, and objectives rebuilt for every benchmark.

🎯
The Task-First World
  • Each NLU task (entailment, QA, similarity, classification) carried its own bespoke architecture
  • Labeled datasets were small and expensive; unlabeled text was nearly infinite and unused
  • Transfer learning meant pretrained word embeddings — static vectors, no context sensitivity
  • RNN/LSTM encoders processed text slowly and forgot long-range structure
🧠
The GPT-1 Answer
  • One 12-layer decoder-only Transformer trained once on unlabeled books
  • A generative objective (next-token prediction) as the universal pre-training signal
  • Task-agnostic input format: every task serialized into one delimited sequence
  • Fine-tuning the whole network adapts the same weights to any NLU benchmark
Analogy — The Apprenticed Reader

Task-trained 2018 systems are students who only ever did past-exam papers — brilliant inside one exam format, lost outside it. GPT-1 is the apprentice who read the entire library first: the exam prep afterwards is short, and skills spill across subjects.

Chapter 02

The Two-Stage Recipe

Pre-train generatively, fine-tune discriminatively — the loop every later GPT repeats at larger scale.

📚 Stage 1 — generative pre-training
Next-token prediction on BooksCorpus (thousands of unique unpublished books). No labels, no task heads — just learning the shape of language.
🔧 Stage 2 — task fine-tuning
Serialize the task into a delimited sequence, append a classification head, and fine-tune the entire network end-to-end on the labeled set.
🧩 Task-agnostic inputs
Tokens + learned delimiters (start, delimiter, extract) let ONE architecture consume classification, entailment, QA, and similarity — no per-task model surgery.
🔁 Auxiliary LM loss
The fine-tuning objective mixes the classification loss with the language-modeling loss — a small trick that stabilized adaptation across tasks.
Architecture (2018 scale)
  • 12 decoder layers, 12 heads, 768-dim — the same depth family BERT-Base would adopt months later
  • 117M parameters trained with byte-pair encoding (BPE) tokenization
  • Causal masking — each token attends only to its left context, which is what makes the model generative
  • Delimiters ($, #) inserted between task fields are the ancestors of today's prompt templates
What the 2018 paper measured
  • Evaluated across natural language inference, question answering, semantic similarity, and classification suites
  • Reported substantial gains from generative pre-training, with state-of-the-art results on a majority of the studied datasets
  • Ablations: zero-shot behavior emerged with task-agnostic formats — a hint of what scale would later unlock
  • Analysis chapters traced the learned representations' value through layer depth
Interactive Demo — The GPT-1 Pipeline

Walk the full recipe: unlabeled books, generative pre-training, task serialization, fine-tuning — the ancestor of every modern training run.

Chapter 03

The Input Format Trick

How one sequence template absorbed four task families — the conceptual ancestor of prompting.

One Backbone, Many Shapes

The key engineering insight was serialization with learned delimiters: a text-pair becomes one delimited stream, and a linear head reads the final transformer output. No new architecture per task — the task lives in the input, not the model. Swap the delimiter layout and the same weights classify sentiment, judge entailment, or pick the answer span. It is a short conceptual step from here to "prompt in, answer out".

Interactive Demo — Four Tasks, One Sequence

See how each NLU task family was folded into the same delimited-input template — the trick that removed per-task architectures.

Chapter 05

Small Model, Large Signal

The exact numbers era (2018: 117M was 'large'); the durable result is the recipe, not the scorecard.

TASK COVERAGE
NLI · QA
similarity · classification — one backbone, four families
PRE-TRAIN CORPUS
Books
thousands of unique unpublished books, unlabeled
IMPACT
Recipe
pre-train → fine-tune became the default for NLU
ECOSYSTEM
GPT line
GPT-2 zero-shot, GPT-3 few-shot, InstructGPT RLHF — all descendants
Interactive Demo — 2018 vs Today

The same recipe, two decades of scale. Press reveal to see what changed — and what did not.

Legacy

Legacy — The Recipe That Won

GPT-1's specific scores faded within months; its training loop conquered everything.

🏆 The pre-train → fine-tune default
Within a year, every serious NLU system started from a pretrained Transformer. The question changed from 'what architecture?' to 'what pre-training?'
🧬 The GPT lineage
GPT-2 dropped task heads for zero-shot prompting; GPT-3 made few-shot prompting a capability; InstructGPT added RLHF. All are the 2018 recipe with the volume raised.
🤝 The BERT sibling
Two months later, BERT showed the encoder-side variant (masked LM) could dominate NLU benchmarks — the field spent five years triangulating between the two.
🧭 Proto-prompting
Learned delimiters as task containers foreshadowed prompt templates, special tokens, and chat formats — GPT-1 is where the input stopped being fixed and started being designed.
⚠️ What it did NOT solve
117M parameters were never going to produce robust world knowledge; the paper's zero-shot results were a curiosity, not a capability; alignment and safety were a decade away.
🛤 Read next
The family tree: GPT-2 · BERT · GPT-3
Test Yourself

Quick Quiz

Check your understanding of the key concepts from GPT-1.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ GPT-1 = generative pre-training (unlabeled books) + discriminative fine-tuning, on a 117M decoder-only Transformer.
✅ Task-agnostic input formats (learned delimiters) let one backbone serve NLI, QA, similarity, and classification.
✅ Fine-tuning updates the entire network — the task head is tiny, the pre-trained body does the work.
✅ Unlabeled text is the cheap teacher: supervision becomes a light steering layer, not the knowledge source.
✅ The causal decoder choice made 'generation' the native mode — the property scale later turned into prompting.
✅ Read it as the origin point: GPT-2, GPT-3, and every aligned assistant re-run this recipe at larger scale.