Language models decode left-to-right and commit instantly. ToT turns reasoning into search: branch over coherent thought-steps, self-evaluate the promising ones, backtrack from the dead ones — GPT-4 on Game of 24 goes from 4% to 74%.
The decoding constraint ToT named and broke: token-level, left-to-right, no looking back.
ToT's four moves: (1) thoughts — coherent intermediate steps, not single tokens; (2) proposal — the LM generates several candidate next-steps from a partial solution; (3) self-evaluation — the LM scores states' promise as a search heuristic; (4) search — BFS or DFS over the tree, expanding high-value nodes and backtracking from dead ends. The general recipe subsumes CoT (a degenerate single-branch tree) and self-consistency (a star, not a tree) — and restores what decoding lost: lookahead, global decisions, and second thoughts.
Why left-to-right generation fails exactly when deliberation matters.
Standard decoding is touch-typing a novel: every letter is final the instant it exists. Self-consistency types several novels and holds a vote. ToT plays chess: consider candidate moves, evaluate the positions they create, look ahead along the promising line, and — crucially — take the piece back when line 3 turns sour. The board state is the partial solution; thinking is navigation, not transcription.
The paper's flagship task, end to end — the anatomy of a solved search.
Creative writing (plan-then-write with re-ranking) and mini crosswords (constrained DFS) — plus the honest bill.
The bill, paid honestly: each ToT solve means many more model calls (proposals + evaluations per node) than a single CoT pass — the trade later quantified by test-time-compute research (entry #58). ToT is a demonstration that structure buys correctness; industry's answer was to fold that structure into training.
The headline and the cost, together.
| Inference paradigm | Paths explored | Mid-course control | Game of 24 (GPT-4) |
|---|---|---|---|
| Standard prompting | 1 | none | ~4%-class failure |
| Chain-of-thought | 1, visible | none — linear | 4% |
| Self-consistency | N independent | vote at the end only | improved, still capped |
| Tree of Thoughts | branching tree | propose · evaluate · backtrack | 74% |
Game of 24 figures from the paper (CoT 4% vs ToT 74% with GPT-4). The structural contrast — not the specific task — is the export.
ToT made reasoning control explicit, visual, and measurable.
Check your understanding of the key concepts from Tree of Thoughts.
Everything you need to remember about this paper.