A 6.7B-parameter model that inserts its own API calls into text, decides for itself when a tool helps, and reads the results back in — beating GPT-3 (175B) on zero-shot knowledge tasks while being 26× smaller.
Models had touched tools before — retrieval welded into the weights, reasoning chains inside one context window, actions prompted by hand. Toolformer's move was to make tool use a learned behavior, trained on data the model generated and filtered itself.
The model is its own annotator. Nobody labels tool calls for Toolformer — a large LM proposes the calls itself, executes them for real, and keeps only the ones whose results make the following text more predictable. Whatever survives becomes fine-tuning data.
A plain language model takes every test closed-book: it can't touch a calculator, look up a fact, or check today's date. Teaching it to use tools used to mean massive human annotation or brittle pipelines — and the tool always sat outside the text, in some special harness.
A language model is a student forced to take every exam closed-book. Humans don't work that way: the moment a problem exceeds memory, we reach for a calculator, a dictionary, a search bar. Handing the model a pencil case sounds easy — but teaching when to reach in used to require millions of hand-labeled examples. Toolformer flips the roles: let the student draft its own practice problems, grade them with a perplexity test, and study only the ones it aced.
Toolformer generates its own tool-use training data in four moves. No humans in the loop: the only hand-written input is a handful of demonstrations per tool, used to prompt the sampler.
Left alone, a sampler proposes plenty of calls — but many are noise: wrong tools, useless positions, unhelpful results. Human intuition about what "should" help doesn't scale, and may not even match what the model itself needs. The perplexity test is ground truth from the model's own perspective: if a result lowers the loss on the next tokens, it genuinely helped this model, in this sentence. The filter turns millions of noisy proposals into exactly the data this model needs — and nothing else.
Toolformer ships with five simple APIs. The only requirements: input and output are plain text, and someone writes a handful of demonstrations for the sampling prompt. Every call uses the same bracketed grammar — which tool earns its keep, and when, is what the model learns.
Fine-tuned on its own filtered data, the 6.7B Toolformer decides at inference time when to pause, call a tool, and read the result back — and it climbs past models many times its size on knowledge and math tasks.
| Model | Size | Tools | LAMA SQuAD | LAMA T-REx | SVAMP (math) |
|---|---|---|---|---|---|
| GPT-J | 6.7B | none | 17.8 | 31.9 | 5.2 |
| OPT | 66B | none | 21.6 | 30.1 | 6.0 |
| GPT-3 | 175B | none | 26.8 | 39.8 | 10.0 |
| Toolformer | 6.7B | 5 self-taught | 33.8 | 53.5 | 29.4 |
LAMA evaluation disables Toolformer's Wikipedia search for fairness (LAMA facts come from Wikipedia) — it still wins, choosing to ask its QA tool in 98.1% of cases. On SVAMP it calls the calculator for 97.9% of examples.
On WebQS, Natural Questions and TriviaQA — where the QA tool is disabled and Wikipedia search is the only option — Toolformer beats every 6.7B baseline but still trails GPT-3 175B (e.g. TriviaQA 48.8 vs 65.9). The paper blames the simple one-shot search: the model can't reformulate a failed query or browse multiple hits. Interaction is flagged as future work.
Fine-tuning on tool-call data doesn't break the base model: perplexity on held-out language modeling (WikiText, CCNet) stays essentially unchanged. The tool habit is additive — the model learns when to reach out without forgetting how to just write. And even with tools disabled at test time, math scores improve, apparently from reading so many worked calculations.
Toolformer's tool use is real but narrow — a first proof that self-taught tool use works, not a finished agent. The paper is candid about what the model still can't do.
Within months of Toolformer, tool use went from research demo to a line in the API docs. Most of what a modern model does by default descends from this line of work.
Check your understanding of the key concepts from the Toolformer paper.
Everything you need to remember about this paper.