History Problem Signatures Modules Teleprompters Results Impact Deep Dive Quiz
Interactive Paper Explainer

Stop Writing Prompts. Program Them.
DSPy

A visual, step-by-step guide to the framework that treats prompts as compiled artifacts — declare what a pipeline stage does, and let an optimizer find the demonstrations, instructions, and strategies that make it work.

Start Learning Read the Paper ↗
+25–65%
Over Few-Shot Baselines
3
Core Abstractions
Minutes
To Compile
2023
Year Published
History

From Prompt String to Program

DSPy reframed the central artifact of the LLM era: the prompt stopped being a hand-crafted asset and became compiler output.

2020–22
The prompt-crafting era
GPT-3 few-shot magic; a folklore of templates, magic phrases, and "say you're an expert" — found by trial and error, versioned in notebooks.
2022–23
Stacked pipelines appear
Multi-hop QA, retrieval, agents: programs of chained LM calls — each link a fragile hand-written prompt. CoT · ReAct.
2023 · Oct
🚀 DSPy (Khattab et al., Stanford)
A programming model: declarative Signatures + parameterized Modules + a compiler (teleprompter) that optimizes any pipeline for a metric — prompts become build artifacts.
2024 →
The optimizer ecosystem
BootstrapFewShot → MIPRO/MIPROv2, GEPA, and a generation of "prompt optimizers" that treat prompting as search over programs.
The Analogy That Explains Everything

prompts are assembly. Writing them by hand was fine for one function; for pipelines it's unmaintainable folklore. Declare intent; compile for your model, your metric, your data.

Chapter 01

Prompt Strings Rot

The paper's diagnosis of 2023's standard practice: pipelines of hard-coded templates, each discovered by trial and error.

📜
Hand-Written Prompt Symptoms
  • Long template strings discovered via trial and error — no systematic path to improvement
  • Every model change (or model version) silently breaks the magic phrases
  • Each pipeline stage tuned in isolation; no notion of end-task optimization
  • Demos, instructions, and control flow tangled inside one text artifact
🧩
The Programmatic Answer
  • Separate what a stage does (Signature) from how it's prompted (compiled artifact)
  • Compose stages as Modules with learnable parameters
  • Optimize the whole pipeline against YOUR metric with a compiler
  • Re-compile for a new model in minutes — the pipeline is the asset, the prompt is build output
The Cost of Folklore (Real Numbers)

The paper's two case studies: multi-hop QA (HotpotQA) and math word problems (GSM8K-class). Succinct DSPy programs — a few lines — compile to prompts that beat standard few-shot by over 25% (GPT-3.5) and 65% (llama2-13b-chat), and beat pipelines with expert-written demonstrations by up to 5-46% and 16-40% respectively. Expert hours replaced by compile-minutes.

Chapter 02

Signatures — Declarative Intent

The atomic unit: a typed declaration of a transformation, stripped of any prompting strategy.

"context, question → answer"  ·  "document → summary"  ·  "question, choices → rationale, selection"
inputs
Left side
Typed fields the stage consumes — the contract's dependencies.
outputs
Right side
Typed fields the stage produces — including intermediate fields like rationale.
zero strategy
The discipline
No demonstrations, no instruction wording, no output formatting tricks — those are compiler decisions.
portable
The payoff
The same Signature compiles to a 5-shot prompt, a finetuned T5-770M, or a chain-of-thought instruction.
Interactive Demo — Signature Splitter
Chapter 03

Modules — Parameterized Strategy

If Signatures say WHAT, Modules say HOW the call is structured — and carry the learnable parameters a compiler can tune.

Predict
The simplest module: one LM call fulfilling a Signature — the substrate everything else decorates.
ChainOfThought
Adds a rationale field and reasoning instructions — the CoT strategy as a reusable wrapper. CoT paper.
ReAct / Retrieve
Agent loops and retrieval as first-class modules — the paper's compiler composes them into multi-hop pipelines.
learnable params
Demonstrations, instructions, hyper-strategy choices: the knobs teleprompters search over — created and collected automatically, not hand-written.
The Pipeline as Text-Transformation Graph

The paper formalizes programs as imperative computational graphs where LMs are invoked through declarative modules — a graph you can execute, inspect, and optimize end-to-end. A multi-hop QA system in DSPy: a few composed modules and one metric function. The same graph is the optimizer's search space.

Control Flow Is Code

Loops, retries, tool dispatch — plain Python in the graph, not prompt-embedded pseudo-logic. The division of labor: Python decides control; the compiler decides prompting; the LM decides content. Each concern finally lives where it belongs.

Chapter 04

The Teleprompter Compiles

Optimization over pipeline space: search demonstrations, instructions, and strategies to maximize the metric.

Interactive Demo — Bootstrap Compilation
BootstrapFewShot (the paper's workhorse)
  • Run the pipeline on training examples; keep the ones where the final answer scores well
  • Backpropagate credit through the graph: trace WHICH intermediate traces led to success
  • Those traces become in-context demonstrations for each module — self-generated few-shot data
  • Result: prompts whose examples are drawn from the model's own successful behavior
What "Compiling" Emits
  • Per-module instruction wording (optionally proposed/selected by an LM or search)
  • Per-module demonstration sets (bootstrapped or retrieved)
  • Strategy choices (e.g., CoT on/off per stage)
  • Or: finetuned weights — the same program compiles to a small open model (T5-770M) competitive with expert-prompted GPT-3.5
Chapter 05

Numbers That Hurt Folklore

Two case studies, four models of improvement, one consistent message.

Improvement Over Standard Few-Shot Prompting
SettingGain over few-shotOver expert demonstrations
GPT-3.5 pipelines> +25%+5–46%
llama2-13b-chat pipelines> +65%+16–40%
Small open models (T5-770M, compiled)competitive with expert-written prompt chains for GPT-3.5

Within minutes of compiling — no prompt engineering hours, per case studies on multi-hop QA and math word problems.

Interactive Demo — Quality Stack
Legacy

Impact — Prompting Becomes a Build Step

The framework's descendants and the research field it named.

🧰 The DSPy ecosystem
MIPRO/MIPROv2, GEPA, SIMBA — optimizer research exploded; dspy is a standard layer of the LLM stack.
🧬 Offspring & cousins
TextGrad, OPRO, prompt-space search: an entire genre of "prompts as optimization targets".
✅ DSPy Assertions
The direct sequel: computational constraints inside compiled programs — read it next.
🎓 Pedagogy shift
"Learn to program LMs" replaced "learn to prompt" in courses and docs across the industry.
⚠️ What it did NOT solve
Metrics are the new bottleneck (optimizers optimize what you measure); compile cost & data needs; not all tasks decompose cleanly.
🧭 Study path
Strategy modules' roots: CoT → this page → constraints: DSPy Assertions.
Deep Dive

The Compiler Bet, Five Years Later

DSPy's deepest claim is an analogy — and analogies in systems design are testable bets.

🎲
Where the Analogy Strains
  • Classical compilers have semantics; LLM "compilation" is stochastic search — no correctness guarantee
  • The optimization target (metric) becomes the spec — garbage metric, eloquently compiled garbage
  • Bootstrapped demos can entrench model biases: self-training loops amplify quirks
  • Debugging a compiled artifact is its own skill — the folklore moved, it didn't vanish
🏗
Where It Paid Off
  • Model churn: re-compile beats re-prompt — portability is now demonstrably cheap
  • Credit assignment through multi-stage graphs was genuinely unsolvable by hand-tuning
  • Optimizers discovered non-obvious wins (fewer/different demos than experts choose)
  • The ecosystem effect: measurable, comparable prompt-improvement research replaced anecdotes
Interactive Demo — Who Writes the Prompt?

The same QA stage under three regimes. Toggle and watch who decides the prompt's parts.

Verdict

DSPy's bet — abstraction beats craftsmanship at scale — is the same bet every systems discipline took, and the LLM stack took it the moment pipelines got long enough. The honest residue: the compiler moved expertise from writing strings to writing metrics and signatures. That's not the end of craft; it's craft promoted a level — exactly what compilers have always done.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the DSPy paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Prompts are build artifacts, not source code — declare intent, let a compiler produce the prompt.
✅ Signatures (typed what) + Modules (parameterized how) = pipelines as optimizable text-transformation graphs.
✅ Teleprompters search demonstrations/instructions/strategies against your metric — minutes, not expert-hours.
✅ Compiled pipelines beat standard few-shot by >25% (GPT-3.5) and >65% (llama2-13b-chat); up to 5-46% over expert demos.
✅ The same program compiles to small open models: T5-770M competitive with expert-prompted GPT-3.5 chains.
✅ The sequel adds constraints: DSPy Assertions — reliability as a programming construct.