A visual, step-by-step guide to the paper that reframed every NLP task as plain text in, plain text out — one model, one objective, and a systematic study of the entire transfer-learning design space.
T5 arrived after years of transfer-learning one-upmanship. Every new paper changed the architecture, the objective, and the recipe — T5 stepped back and compared them all.
Every NLP task — translation, classification, regression, QA, summarization — can be expressed as text in, text out. Once every task speaks the same language, one model can learn them all, and every design choice can be compared on the same scale.
By 2019, transfer learning clearly worked — but nobody could say which ingredient was doing the work. Every paper changed several at once.
T5 casts every task as a sequence-to-sequence problem: a text instruction goes in, text comes out. Labels, answer spans, and even numbers become generable tokens.
How does one model know which job it's being asked to do? A short prefix prepended to the input: translate English to German:, summarize:, question:. During multi-task training the model learns to condition on them — the task signal lives in the text itself, not in extra parameters or a separate head.
BERT bolted a classifier on top for each task; GPT added per-task output layers and alignment hacks. T5's decoder emits text directly, so class labels ("positive"), answer spans ("Paris"), and numeric scores ("3.8") all share one softmax over the vocabulary — the model "writes" its answers.
Transfer learning runs on unlabeled text. T5 built its own: April 2019's Common Crawl snapshot, scrubbed down to roughly 750GB of clean English text.
C4 dwarfs the classic corpora — and its diversity (every topic on the web, not just encyclopedias) meant the model could train far longer without repeating itself.
Wikipedia and books are clean but limited. When the paper's ablations compared pre-training corpora, bigger and more diverse kept winning: Wikipedia-only pre-training lagged C4, and aggressively filtered "high-quality" subsets gave up too much diversity. For T5's goals — long training at scale — only the web was big enough.
T5's deepest contribution isn't a leaderboard number — it's a controlled study of the entire transfer-learning design space, with every variant scored on the same benchmarks.
| Design Choice | What Was Compared | What Won |
|---|---|---|
| Architecture | encoder-only · decoder-only · encoder-decoder | Encoder-decoder — bidirectional understanding plus autoregressive generation |
| Objective | causal LM · BERT-style MLM · prefix LM · span corruption | Masked / span-style objectives — causal LM lags on understanding tasks |
| Unlabeled data | C4 vs Wikipedia-scale and filtered variants | More, more-diverse data — bigger corpora consistently won |
| Training length | shorter vs longer pre-training | Longer training kept helping — no saturation in sight |
| Scale | 60M → 11B parameters | Bigger models won consistently across every benchmark |
| Transfer method | task prefixes · multi-task training · alternative conditioning schemes | Task prefixes — the task signal stays in the text |
Plug the winning recipe into an 11B-parameter encoder-decoder, and it sets a new record on SuperGLUE — the hardest language-understanding benchmark of the era.
| System | Size | SuperGLUE | What Changed |
|---|---|---|---|
| BERT-Large (2018) | 340M | 69.0 | Encoder-only + masked LM + task-specific heads |
| RoBERTa (2019) | 355M | ~84–85 | Same architecture, tuned recipe — more data, longer training |
| T5-11B (2019) | 11B | 88.9 | Text-to-text + the ablation-selected recipe, at scale |
88.9 was the best SuperGLUE score reported at the time — within sight of the human baseline. T5 also set a new best on GLUE (~89) and scored strong results across summarization, question answering, and translation — all from one model with no task-specific machinery.
Five checkpoints from laptop-scale to research-cluster scale. Toggle the axis to compare them fairly (log scale) or feel the true gap (linear).
At the time of release, T5-11B was the largest, most accurate language model of its kind — and it was shared with researchers through a request process rather than a public download. The smaller family sizes shipped alongside the code and the C4 recipe, making the whole pipeline reproducible.
Because every task is just text in → text out, a multi-task T5 can attempt tasks it was never fine-tuned on — simply describe the task with a prefix and let the decoder write. The paper explores how far this "zero-shot" transfer reaches within the unified framework.
T5's text-to-text framing became the default way to think about task conditioning — and its ablation playbook became the standard way to build models.
Check your understanding of the key concepts from the T5 paper.
Everything you need to remember about this paper.