Before BERT, before the scaling race, a quiet 2018 result set the template for everything after: pre-train a decoder on unlabeled text, then fine-tune with a task-agnostic input format.
Where NLU stood in 2018, and how one recipe changed the training economics of the whole field.
Supervised data was the bottleneck of 2018 NLP: label thousands of examples per task, retrain per task. GPT-1's bet was that predicting the next token on unlabeled books forces syntax, discourse, and long-range dependency into the weights — and that a small slice of supervised data could then steer those weights. The decoder choice also proved prophetic: the same objective, scaled up, becomes few-shot and zero-shot learning.
The 2017-18 status quo that GPT-1 attacked: architectures, data, and objectives rebuilt for every benchmark.
Task-trained 2018 systems are students who only ever did past-exam papers — brilliant inside one exam format, lost outside it. GPT-1 is the apprentice who read the entire library first: the exam prep afterwards is short, and skills spill across subjects.
Pre-train generatively, fine-tune discriminatively — the loop every later GPT repeats at larger scale.
How one sequence template absorbed four task families — the conceptual ancestor of prompting.
The key engineering insight was serialization with learned delimiters: a text-pair becomes one delimited stream, and a linear head reads the final transformer output. No new architecture per task — the task lives in the input, not the model. Swap the delimiter layout and the same weights classify sentiment, judge entailment, or pick the answer span. It is a short conceptual step from here to "prompt in, answer out".
The exact numbers era (2018: 117M was 'large'); the durable result is the recipe, not the scorecard.
GPT-1's specific scores faded within months; its training loop conquered everything.
Check your understanding of the key concepts from GPT-1.
Everything you need to remember about this paper.