A visual, step-by-step guide to the framework that treats prompts as compiled artifacts — declare what a pipeline stage does, and let an optimizer find the demonstrations, instructions, and strategies that make it work.
DSPy reframed the central artifact of the LLM era: the prompt stopped being a hand-crafted asset and became compiler output.
prompts are assembly. Writing them by hand was fine for one function; for pipelines it's unmaintainable folklore. Declare intent; compile for your model, your metric, your data.
The paper's diagnosis of 2023's standard practice: pipelines of hard-coded templates, each discovered by trial and error.
The paper's two case studies: multi-hop QA (HotpotQA) and math word problems (GSM8K-class). Succinct DSPy programs — a few lines — compile to prompts that beat standard few-shot by over 25% (GPT-3.5) and 65% (llama2-13b-chat), and beat pipelines with expert-written demonstrations by up to 5-46% and 16-40% respectively. Expert hours replaced by compile-minutes.
The atomic unit: a typed declaration of a transformation, stripped of any prompting strategy.
If Signatures say WHAT, Modules say HOW the call is structured — and carry the learnable parameters a compiler can tune.
The paper formalizes programs as imperative computational graphs where LMs are invoked through declarative modules — a graph you can execute, inspect, and optimize end-to-end. A multi-hop QA system in DSPy: a few composed modules and one metric function. The same graph is the optimizer's search space.
Loops, retries, tool dispatch — plain Python in the graph, not prompt-embedded pseudo-logic. The division of labor: Python decides control; the compiler decides prompting; the LM decides content. Each concern finally lives where it belongs.
Optimization over pipeline space: search demonstrations, instructions, and strategies to maximize the metric.
Two case studies, four models of improvement, one consistent message.
| Setting | Gain over few-shot | Over expert demonstrations |
|---|---|---|
| GPT-3.5 pipelines | > +25% | +5–46% |
| llama2-13b-chat pipelines | > +65% | +16–40% |
| Small open models (T5-770M, compiled) | competitive with expert-written prompt chains for GPT-3.5 | |
Within minutes of compiling — no prompt engineering hours, per case studies on multi-hop QA and math word problems.
The framework's descendants and the research field it named.
DSPy's deepest claim is an analogy — and analogies in systems design are testable bets.
DSPy's bet — abstraction beats craftsmanship at scale — is the same bet every systems discipline took, and the LLM stack took it the moment pipelines got long enough. The honest residue: the compiler moved expertise from writing strings to writing metrics and signatures. That's not the end of craft; it's craft promoted a level — exactly what compilers have always done.
Check your understanding of the key concepts from the DSPy paper.
Everything you need to remember about this paper.