History Problem Core Idea Templates Results Ablations Impact Deep Dive Quiz
Interactive Paper Explainer

Teaching Models to Follow Instructions
FLAN

A visual, step-by-step guide to the paper that introduced instruction tuning — finetuning a 137B model on 60+ tasks phrased as natural-language instructions so it could solve task types it had never seen, zero-shot.

Start Learning Read the Paper ↗
137B
Parameters (LaMDA-PT)
60+
Instruction-Tuned Tasks
20/25
Beats Zero-Shot GPT-3
2021
Year Published
History

From Prompts to Instructions

FLAN sits at the hinge between the few-shot prompting era and the instruction-following era we live in now.

2019
Task-specific finetuning era
BERT-style workflows: a new task meant new labels, new heads, and a new finetune. Zero-shot transfer was a curiosity.
2020
GPT-3 few-shot prompting (Brown et al.)
175B parameters + a few examples in the prompt. Works — but every task needs hand-crafted prompt formats, and zero-shot lags badly.
2021 · Sep
🚀 FLAN (Wei et al.)
Instruction tuning: finetune LaMDA-PT 137B on 60+ tasks verbalized as instructions. Zero-shot performance jumps on unseen task types.
2021 · Oct
T0 & Natural Instructions (Sanh et al.)
The same idea, open-sourced at scale: T5-based T0 trained on a massive crowdsourced instruction dataset.
2022 →
InstructGPT, Flan-T5, ChatGPT…
Instruction tuning merges with human-feedback alignment and becomes the default final training stage of every assistant model.
The One-Sentence Idea

A model that has been trained on instructions generalizes to instructions it has never seen. FLAN turned "prompt engineering against a raw language model" into "training the model to be promptable" — the ancestor of every instruction-tuned model since.

🧭 Lineage note
FLAN tunes the base model on instructions; InstructGPT adds human preference data and RLHF on top of a similar recipe.
Chapter 01

The Zero-Shot Gap

GPT-3 proved scale could learn tasks in-context — but only if you found the right magic prompt format. Zero-shot was the weakest mode of a powerful model.

🎭
Raw Models Are Prompt-Sensitive
  • Same task, different phrasing: wildly different accuracy — raw models don't reliably "understand" a task description
  • Prompt engineering is hand-crafted, task-by-task, and does not transfer
  • Zero-shot (no examples) trails few-shot by large margins in the 175B GPT-3 evaluations
  • The model's objective — predict the next token — never asked it to be a helpful instruction-follower
📋
Instruction Tuning Closes the Gap
  • Phrase existing supervised datasets as natural-language instructions ("Translate this sentence to German: …")
  • Finetune the 137B model on that mixture — no new labels, just re-phrasing
  • Zero-shot performance on held-out task types jumps by large margins
  • One training stage makes the model usable by anyone who can write an instruction
Analogy — The Operating Manual

A raw pretrained model is a brilliant employee who has read the entire internet but was never told what the company does. Instruction tuning is the onboarding manual: it doesn't add capability, it adds mutual intelligibility — the employee finally knows what you mean when you ask.

Chapter 02

Instruction Tuning, Step by Step

Three moves: verbalize datasets into instructions, group tasks into clusters, and always evaluate on clusters that were held out of tuning.

📝 60+ datasets verbalized
Existing supervised NLP datasets re-written as instruction/response pairs with natural templates.
🗂 12 task clusters
Datasets grouped by task type — NLI, reading comprehension, closed-book QA, translation, sentiment…
🔒 Cluster hold-out
Tune on all clusters except the one being evaluated — the held-out evaluation is genuinely zero-shot at the task-type level.
🤖 LaMDA-PT 137B
The pretrained-only base of Google's LaMDA — a decoder-only Transformer, before dialogue tuning.
Interactive Demo — Instruction Tuning Machine

Take a raw supervised example and verbalize it into an instruction. Flip between raw few-shot prompting and the instruction-tuned format.

Chapter 03

Templates Are the Interface

How a task gets phrased is a first-class design decision — FLAN systematically studies the choices that later papers inherited.

Three Template Decisions
  • Instruction phrasing: a task template can lead with a general instruction or a concrete example; FLAN's default is instruction-first, "task name last"
  • Response format: class labels are mapped into natural answers ("positive"/"negative", or full sentences)
  • Number of templates: multiple phrasings per dataset reduce overfitting to a single format
What the Ablations Showed
  • More tuning datasets: zero-shot transfer improves as the task mixture grows — instruction diversity itself is the signal
  • Scale matters: instruction tuning helps at 137B; smaller models benefit far less — an early "emergent ability" observation
  • Task-cluster overlap: gains concentrate when the held-out task is related to tuned clusters — generalization is graded, not magic
Interactive Demo — Template Flipper

The same sentiment example under three verbalizations. Flip through them and watch what the model must infer in each case.

Chapter 04

Zero-Shot, Measured Honestly

Held-out task clusters, compared against a 175B model with 28% more parameters that never saw instructions.

HELD-OUT TASKS
25
unseen benchmark datasets across 10 evaluated task clusters
FLAN vs GPT-3 ZERO-SHOT
20/25
datasets where 137B FLAN beats zero-shot 175B GPT-3
BEATS FEW-SHOT GPT-3
6
datasets where zero-shot FLAN beats few-shot GPT-3 (ANLI, RTE, BoolQ, ARC, OpenbookQA, StoryCloze)
TUNING COST
1 stage
a finetune over existing datasets — no new human labels
Interactive Demo — Cluster Hold-out Explorer

Pick a held-out cluster and see what FLAN was (and wasn't) tuned on before facing it zero-shot.

Chapter 05

What Actually Drove the Gains

FLAN's ablations previewed the scaling questions that the next five years of instruction-tuning research would ask.

Scale Threshold

Instruction tuning's zero-shot gain is not linear in model size: it shines at the largest scale tested (137B), and the paper's analysis of smaller models foreshadowed the "emergent abilities" debate — a capability that switches on only past a size threshold. Later work (Flan-PaLM, 2022) confirmed the effect grows with scale.

Cluster Relatedness

Hold out a cluster and gains shrink for task types unlike anything in tuning; hold out a related cluster and transfer is strong. FLAN maps generalization as graded — an honest frame that later instruction papers kept.

Legacy

Impact — The Default Final Stage

Instruction tuning went from a 2021 experiment to the finishing move of every serious LLM training pipeline.

💬 The Assistant Era
InstructGPT/ChatGPT-style systems pair instruction tuning with preference alignment — FLAN proved the instruction part alone moves zero-shot scores dramatically.
🔥 Flan-T5 & Flan-PaLM (2022)
Google scaled the recipe to 1,800+ tasks with chain-of-thought data mixed in — Flan-T5 became a default open workhorse.
🌍 T0 & Open Instruction Data
The open-source world answered with T0, Natural Instructions, and eventually community instruction datasets that power open assistants.
🧪 Evaluation Culture
Held-out cluster evaluation — FLAN's careful methodology — became the standard way to claim "zero-shot generalization" honestly.
⚠️ What it did NOT solve
Truthfulness, helpfulness, and calibration: instruction tuning makes models compliant, not correct — the alignment problem needed its own papers.
🧭 Study next
See how instruction tuning meets human preferences in the InstructGPT Guide.
Deep Dive

Capabilities vs Formatting

The sharpest reading of FLAN: much of the "zero-shot gap" was never a knowledge gap — it was an interface gap.

🌫
The Interface Tax
  • A raw LM may "know" how to classify sentiment but fail because it expects a particular format
  • Zero-shot scores conflate task knowledge with format inference
  • Prompt engineering is a patch: human-side adaptation to a fixed model
  • Benchmarks under-measure models that know more than they can show
🔌
The Tuning Dividend
  • Instruction tuning moves the adaptation cost into training, once, for all users
  • The model's latent skills become accessible through plain language
  • Zero-shot FLAN beating few-shot GPT-3 on 6 datasets is the tell: format was the bottleneck
  • Model-side beats human-side adaptation — the economics of the whole field shifted
Interactive Demo — Prompt Format Stress Test

Same knowledge, four formats. Watch the format tax shrink the raw model's score — then vanish after instruction tuning.

Verdict

FLAN is best read as an economics paper disguised as an NLP paper. It showed that a one-time training investment could permanently lower the interaction cost for every user and every task — and that "prompt sensitivity" is a property of untrained interfaces, not of scale itself. Every chat model since has paid that investment forward.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the FLAN paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Instruction tuning: finetune on tasks verbalized as natural instructions — no new labels needed.
✅ LaMDA-PT 137B, 60+ datasets, organized into 12 task clusters; evaluation always holds out the target cluster.
✅ Zero-shot FLAN beats zero-shot 175B GPT-3 on 20 of 25 held-out datasets — with 28% fewer parameters.
✅ It even beats few-shot GPT-3 on ANLI, RTE, BoolQ, AI2-ARC, OpenbookQA, and StoryCloze.
✅ Gains strengthen with model scale and dataset diversity — the two levers later recipes pulled harder.
✅ FLAN fixed the interface, not the truthfulness — alignment arrived with InstructGPT and successors.