Start with 175 human-written seed tasks. Let the model generate new instructions, inputs, and outputs — filter the weak and the near-duplicate — fine-tune on the survivors. Vanilla GPT-3 gains 33% absolutely, matching InstructGPT-001 without its private data.
Instruction tuning worked — and the data to do it belonged to whoever could collect humans at scale.
The pipeline is a four-step loop the model runs on itself: (1) generate new instruction ideas seeded by existing ones (few-shot prompted), (2) expand each into (instruction, input, output), (3) filter — language check, dedup by embedding similarity against existing instructions, drop identical or degenerate outputs, (4) fine-tune the base model on the survivors. The model's own distribution generates the curriculum; cheap heuristics enforce quality and diversity; the fine-tuned model becomes a better generator for the next round.
The 2022 economics that Self-Instruct attacked.
Instruction tuning was a chef cooking from recipes commissioned from specialists — expensive, slow, coverage-limited. Self-Instruct hands the chef 175 sample dishes and says: invent new recipes inspired by these, throw out the flops and the repeats, and practice on the keepers. The cookbook grows by the thousands — and the practice itself makes the invention better.
Generate → expand → filter → finetune — the loop that became the synthetic-data canon.
Where the idea went next — the $600 assistant.
Stanford's Alpaca (March 2023) re-ran the pattern with a stronger generator: GPT-3.5 (text-davinci-003) produced 52K instruction-following demonstrations from the Self-Instruct seed set for about $600 of API calls, then LLaMA-7B was fine-tuned on them — producing a single-GPU assistant that famously closed most of the gap to text-davinci-003 in blind evaluations. The loop's economics: synthetic data + a capable generator + a cheap open student = the open-weights ecosystem's assembly line. Every later recipe — Evol-Instruct (deepen complexity iteratively), UltraFeedback, distillation cascades — is a Self-Instruct variant with different plumbing.
The result that made synthetic instruction data the default ingredient.
Self-Instruct legitimized the loop that now trains half the field.
Check your understanding of the key concepts from Self-Instruct.
Everything you need to remember about this paper.