History Problem Core Idea Tools Results Limits Impact Quiz
Interactive Paper Explainer

Models That Teach Themselves Tools
Toolformer

A 6.7B-parameter model that inserts its own API calls into text, decides for itself when a tool helps, and reads the results back in — beating GPT-3 (175B) on zero-shot knowledge tasks while being 26× smaller.

Start Learning Read the Paper ↗
6.7B
Parameters (GPT-J)
5
Self-Taught Tools
0
Human-Labeled Calls
2023
Year Published
History

The Road to Self-Taught Tool Use

Models had touched tools before — retrieval welded into the weights, reasoning chains inside one context window, actions prompted by hand. Toolformer's move was to make tool use a learned behavior, trained on data the model generated and filtered itself.

2020
Augmented LMs (REALM & friends)
Retrieval moves inside the forward pass — knowledge is baked into special architectures, not callable on demand mid-generation.
2022 · Jan
Chain-of-Thought prompting
Models reason step by step in text — big gains on math and logic, but every step stays inside the model's own head. CoT Guide →
2022 · Oct
ReAct
Thoughts and actions interleaved by prompting — a model can act in an environment, but every tool step is hand-prompted, example by example. ReAct Guide →
2023 · Feb
🚀 Toolformer
Self-supervised tool use: the model samples API calls inside raw text, filters them with a perplexity test, and fine-tunes on the survivors — calls live in the text it generates.
2023 →
Tool use goes mainstream
Function-calling ships in production LLM APIs; Gorilla and ToolLLM train models on thousands of APIs; tool use becomes a standard part of agent stacks.
Key Insight

The model is its own annotator. Nobody labels tool calls for Toolformer — a large LM proposes the calls itself, executes them for real, and keeps only the ones whose results make the following text more predictable. Whatever survives becomes fine-tuning data.

A CALL THAT SURVIVED THE FILTER
The population of Marseille is [QA("What is the population of Marseille?", "Marseille") → 861,635] as of 2017.
The call sits mid-sentence; the result comes back inline. Fine-tune on millions of passages like this and the model learns to write the bracket itself.
Chapter 01

Trapped Inside Their Own Weights

A plain language model takes every test closed-book: it can't touch a calculator, look up a fact, or check today's date. Teaching it to use tools used to mean massive human annotation or brittle pipelines — and the tool always sat outside the text, in some special harness.

🧱
Teaching Tool Use, the Old Way
  • Human-annotated call sites — labeling where, what, and how to call at scale is far too expensive
  • Hand-crafted pipelines bolt one tool onto one task at a time
  • Brittle special-casing: hard-coded rules for when a tool "should" fire
  • Task-specific fine-tunes that don't survive the next task
  • Tool use walled off from normal, free-form text generation
🛠️
Toolformer's Way
  • The model itself samples where a call belongs — mid-sentence, in natural text
  • It picks the tool and writes the arguments
  • Calls execute for real; results are read back into the generation
  • A perplexity filter keeps only calls that provably help
  • Zero human-labeled tool calls in the fine-tuning data
The Analogy — The Closed-Book Student

A language model is a student forced to take every exam closed-book. Humans don't work that way: the moment a problem exceeds memory, we reach for a calculator, a dictionary, a search bar. Handing the model a pencil case sounds easy — but teaching when to reach in used to require millions of hand-labeled examples. Toolformer flips the roles: let the student draft its own practice problems, grade them with a perplexity test, and study only the ones it aced.

3482 × 12
Human: reaches for a calculator → 41,784
Plain LM: guesses from its weights → "about 40,000?"
Chapter 02

The Core Idea — A Self-Annotating Loop

Toolformer generates its own tool-use training data in four moves. No humans in the loop: the only hand-written input is a handful of demonstrations per tool, used to prompt the sampler.

The Self-Supervised Loop
🎲
1 · Sample
Prompted with a few demos, GPT-J inserts candidate calls — position, tool, arguments — into a huge text corpus.
⚙️
2 · Execute
Each candidate is executed for real; the result is pasted back between the call and the continuation.
🎯
3 · Filter
Score the model's loss on the continuation with vs. without the result — keep only calls whose results help.
🔁
4 · Fine-Tune
Train GPT-J on the surviving annotated text. It now writes calls itself — and pauses generation to use them.
sampled:  e₁  [ Tool(args) → r ]  e₂
keep the call only if  L(e₂ | e₁, r) < L(e₂ | e₁)
e₁
Text before
Everything up to the position where the model decided a tool might help.
Tool(args) → r
The sampled call
Tool and arguments chosen by the model; r is the real, executed result.
e₂
The continuation
The text the model still has to predict after the call.
L(·)
Loss / perplexity
Keep the call only if seeing r makes e₂ cheaper to predict.
Why the Filter Is the Whole Trick

Left alone, a sampler proposes plenty of calls — but many are noise: wrong tools, useless positions, unhelpful results. Human intuition about what "should" help doesn't scale, and may not even match what the model itself needs. The perplexity test is ground truth from the model's own perspective: if a result lowers the loss on the next tokens, it genuinely helped this model, in this sentence. The filter turns millions of noisy proposals into exactly the data this model needs — and nothing else.

Interactive Demo — You Are the Perplexity Filter

For each sampled call, two perplexities are computed on the text that follows it — one without the result, one with it (illustrative values). Read the bars, then decide like the filter: keep this call or discard it.

Chapter 03

Five Tools, One Grammar

Toolformer ships with five simple APIs. The only requirements: input and output are plain text, and someone writes a handful of demonstrations for the sampling prompt. Every call uses the same bracketed grammar — which tool earns its keep, and when, is what the model learns.

🧮 Calculator
Four basic arithmetic operations, results rounded to two decimals. A 6.7B LM can't multiply multi-digit numbers; the calculator is exact.
[Calculator(27 + 4 * 2) → 35]
❓ Question Answering
A factoid QA system — Atlas, a retrieval-augmented model fine-tuned on Natural Questions. Ask a question, get a short answer.
[QA("Where was the Knights of Columbus founded?", "Knights of Columbus") → New Haven, Connecticut]
📚 Wikipedia Search
A search engine over Wikipedia: give a term, get short snippets back. More comprehensive than QA — but the model has to read and use what returns.
[WikiSearch("Spin fishing") → "Spin fishing is distinguished from fly fishing…"]
📅 Calendar
Takes no arguments; returns the current date. Grounds "today", "tomorrow", and "next Friday" in actual time — something frozen weights can never know.
[Calendar() → 2023-02-04, Saturday]
🌐 Machine Translation
A 600M-parameter NLLB model covering 200 languages, translating any phrase into English (source language auto-detected).
[MT("sûreté nucléaire", "French", "English") → nuclear safety]
📐 The one constraint
Text in, text out. That's it. Any API whose inputs and outputs can be written as text sequences — and which has a few demonstrations available — can join the bracket grammar the same way.
Interactive Demo — Tool Trace Builder

A model is writing a travel guide and pauses at the blank. Insert one of three tool calls and watch the completed sentence — and what each result buys the continuation. This is exactly the choice Toolformer learned to make on its own.

Chapter 04

Small Model, Long Reach

Fine-tuned on its own filtered data, the 6.7B Toolformer decides at inference time when to pause, call a tool, and read the result back — and it climbs past models many times its size on knowledge and math tasks.

Interactive Demo — Size vs Knowledge (real paper scores)

Toggle the benchmark and watch three models race. All scores are zero-shot, straight from the paper's Tables 3 and 4.

Zero-Shot Results (paper Tables 3 & 4)
ModelSizeToolsLAMA SQuADLAMA T-RExSVAMP (math)
GPT-J6.7Bnone17.831.95.2
OPT66Bnone21.630.16.0
GPT-3175Bnone26.839.810.0
Toolformer6.7B5 self-taught33.853.529.4

LAMA evaluation disables Toolformer's Wikipedia search for fairness (LAMA facts come from Wikipedia) — it still wins, choosing to ask its QA tool in 98.1% of cases. On SVAMP it calls the calculator for 97.9% of examples.

LAMA · T-REx
53.5
Toolformer (6.7B) — vs 39.8 for the 26× larger GPT-3
LAMA · SQUAD
33.8
vs 26.8 for GPT-3 (175B) — +11.7 pts over the best baseline
ASDIV · MATH
40.4
vs 14.0 for GPT-3 (175B) — calculator in 97.9% of cases
AUTONOMY
98.1%
of LAMA answers where the model chose to call the QA tool itself
Where It Doesn't Win

On WebQS, Natural Questions and TriviaQA — where the QA tool is disabled and Wikipedia search is the only option — Toolformer beats every 6.7B baseline but still trails GPT-3 175B (e.g. TriviaQA 48.8 vs 65.9). The paper blames the simple one-shot search: the model can't reformulate a failed query or browse multiple hits. Interaction is flagged as future work.

What It Keeps

Fine-tuning on tool-call data doesn't break the base model: perplexity on held-out language modeling (WikiText, CCNet) stays essentially unchanged. The tool habit is additive — the model learns when to reach out without forgetting how to just write. And even with tools disabled at test time, math scores improve, apparently from reading so many worked calculations.

Chapter 05

The Catch

Toolformer's tool use is real but narrow — a first proof that self-taught tool use works, not a finished agent. The paper is candid about what the model still can't do.

📏 Locked syntax
The model is bound to the API call format it was fine-tuned on. It can't invent new tools, rename arguments, or call an API it never saw in the bracket grammar.
🔗 No chaining
One call, one result. A call's output can't feed the next call, so there's no multi-step tool use — no "search, then read, then compute". The paper flags interaction as future work.
🖊️ Humans design the tools
Each API needs a handful of hand-written demonstrations to seed the sampler — plus a real executor behind every call. Self-supervised tool use, not tool creation.
💸 Sampling is compute-heavy
Annotating a corpus means generating many candidate calls per position, executing them all, and scoring two losses for each. The bootstrap isn't free.
📈 Scale gates the trick
The paper's scaling study shows tool learning improves with model size — smaller backbones (its GPT-2 tests) barely benefit. GPT-J at 6.7B is where self-annotation starts to pay off.
🧭 Still, the direction is clear
Bigger models are better tool learners — which means this recipe compounds with scale rather than fighting it. The 2023 function-calling wave proved the point.
Legacy

Impact — Function Calling Becomes a Feature

Within months of Toolformer, tool use went from research demo to a line in the API docs. Most of what a modern model does by default descends from this line of work.

⚡ Function calling (2023)
Production LLM APIs ship structured tool calling: the model decides when to emit a call, the platform executes it. Model-initiated, mid-generation tool use is Toolformer's pattern at scale.
🦍 Gorilla (2023)
An LLM fine-tuned to call thousands of real APIs accurately — tool use scaled from Toolformer's five hand-seeded tools to entire API corpora.
🧰 ToolLLM & API-Bank (2023)
Open tool-use instruction data and the first benchmarks for evaluating tool-using agents — the ecosystem that grew around self-taught tool use.
🪞 Models writing their own data
The deeper idea is the bootstrap: a model generating and filtering its own supervision. That lineage runs through self-generated instruction data into today's self-improvement loops.
🤖 Agent stacks (2023–24)
Planner + tools + memory became the standard agent blueprint, usually wired through ReAct-style thought-action loops — tool use as a default module, not a special case.
📉 The scale lesson
A 6.7B model with the right tool beat a 175B model without one — the strongest early argument that tool augmentation can substitute for raw parameter count.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the Toolformer paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Self-supervised tool use: the model samples its own API calls, executes them, and filters by perplexity — 0 human-labeled calls.
✅ The filter is the trick: a call survives only if its result makes the next tokens more predictable (lower loss).
✅ Five text-in/text-out tools: calculator, QA (Atlas), Wikipedia search, calendar, machine translation (NLLB) — called inline in plain text.
✅ GPT-J 6.7B backbone, fine-tuned on self-annotated data; at inference it pauses generation, calls the tool, and reads the result back.
✅ Zero-shot wins: LAMA T-REx 53.5 and SVAMP 29.4 — beating GPT-3 175B with a model 26× smaller.
✅ Limits: fixed call syntax, no tool chaining, human-designed tools — but the recipe became the function-calling standard.