History Problem Core Idea The Bill Results Impact Quiz Takeaways
Interactive Paper Explainer

Frontier Quality,
Discount Training
DeepSeek-V3

A 671B-parameter MoE (37B activated) trained on 14.8T tokens in 2.788M H800 GPU-hours — Multi-head Latent Attention, DeepSeekMoE, auxiliary-loss-free balancing, and FP8 mixed precision, published as a reproducible report.

Start Learning Read the Paper ↗
671B
Total parameters
37B
Activated / token
14.8T
Training tokens
2.788M
H800 GPU-hours
History

The Cost Revolution

Each DeepSeek generation attacked a different line of the training-cost equation — V3 attacked all of them at once.

2023-24
The cost assumption
Frontier training = tens of millions of dollars, opaque, closed. Efficiency work was open-model territory, capability work was closed-lab territory.
2024 · V2
MLA arrives
Multi-head Latent Attention compresses KV cache into a latent vector; DeepSeekMoE (entry #17) becomes the family spine — the 'cheap inference' generation.
Dec 2024
🚀 DeepSeek-V3
671B/37B activated, 14.8T tokens, FP8 mixed-precision training, auxiliary-loss-free load balancing, multi-token prediction — 2.788M H800 GPU-hours total.
Jan 2025
The R1 shockwave
V3's base + RL produces DeepSeek-R1 reasoning (entry #47); markets re-price the entire AI stack overnight.
2025+
Efficiency as strategy
The industry re-learns that architecture + numerics + data engineering can substitute for raw capital — V3 is the reference case.
The Thesis in One Number

2.788M H800 GPU-hours. At published rental rates that is a few million dollars for a model that outperforms other open-source models and is comparable to leading closed-source models — not by using less compute to do less, but by removing every inefficiency the stack had learned to tolerate: attention (MLA), routing (aux-loss-free MoE), numerics (FP8), and curriculum (multi-token prediction).

Chapter 01

Frontier Training, Priced as Frontier

The assumption under attack: capable models must be expensive models, and only closed labs can afford the attempt.

💸
The 2024 Cost Wall
  • Frontier-class pre-training was believed to cost an order of magnitude more than V3's bill
  • MoE load balancing via auxiliary losses degrades quality — an invisible tax on every sparse model
  • FP16/BF16 training wasted half the bits the hardware could carry
  • KV-cache attention costs scaled with heads and context — serving 671B-class models looked unsustainable
🧮
The V3 Answer
  • MLA: keys/values folded into a compressed latent — tiny KV cache at any context
  • DeepSeekMoE with auxiliary-loss-free balancing: bias-based routing, quality no longer taxed
  • FP8 mixed precision end-to-end on 2,048 H800s with fine-grained scaling strategies
  • Multi-token prediction (MTP): the training signal predicts future tokens, densifying every FLOP
Analogy — The Budget Airline

Closed frontier training was a flagship carrier: first-class FLOPs, everyone paying full precision. V3 is the budget airline that flies the same route — same destination (frontier quality), every cost line engineered: thinner seats (MLA), dynamic crew assignment (aux-loss-free routing), fuel hedged (FP8), and every flight carrying freight too (multi-token prediction). The ticket price is not a discount on quality — it is the absence of waste.

Chapter 02

The Four Efficiency Pillars

Each pillar was validated in V2 and scaled in V3 — the report is their industrial integration.

🧠 MLA — Multi-head Latent Attention
KV pairs compressed into a low-rank latent vector; cache one latent per token instead of per-head keys/values. Long context becomes affordable.
🔀 Aux-loss-free MoE
DeepSeekMoE topology, but load balancing via per-expert bias adjustment instead of an auxiliary loss — balance without fighting the gradient.
🔢 FP8 mixed precision
End-to-end FP8 training with fine-grained block-wise scaling on 2,048 H800s — the numerics gamble that halved memory and bandwidth.
🎯 Multi-token prediction
An auxiliary head predicts the next several tokens simultaneously, densifying the training signal and feeding better inference acceleration.
Scale facts (from the report)
  • 671B total / 37B activated per token
  • 14.8T diverse, high-quality pre-training tokens
  • 2.788M H800 GPU-hours total (pre-train + context ext + post-train)
  • Post-training: SFT + RL stages; distillation from earlier DeepSeek reasoning series
The pipeline
  • Pre-training with MTP; then long-context extension (128K)
  • SFT on instruction/reasoning/code data
  • RL stages — including rule-based rewards for verifiable domains (the R1 lineage, entry #47)
  • Evaluation: outperforms open peers; comparable to leading closed models
Interactive Demo — Four Pillars, Four Savings

Tab through the pillars and see what each one removes from the cost equation.

Chapter 03

What the Bill Actually Means

The GPU-hour number that reset industry assumptions — and what it does and does not imply.

Reading 2.788M GPU-Hours Honestly

The report's cost covers this run's GPU-hours — not the research program, the failed runs, the V2 line's development, or the data pipeline's amortized cost. What it proves is narrower and stronger: a frontier-comparable open model can be trained for a few million dollars of compute, because architecture (MLA+MoE), numerics (FP8), and training design (MTP, aux-free balancing) compound. Every efficiency's existence was known separately; V3's contribution is the disciplined integration that made them safe at 671B scale.

Interactive Demo — One Token Through DeepSeek-V3

Follow a single token through the full forward pass — latent attention, aux-free routing, and the MTP head looking ahead.

Chapter 05

Open Frontier, Printed Receipt

The evaluation claim and the cost claim, side by side — that juxtaposition is the paper.

671B total · 37B activated · 14.8T tokens · 2.788M H800 GPU-hours
MLA
Latent attention
KV cache compressed into latent vectors — per-token, not per-head, storage.
MoE + bias
Aux-loss-free
Load balancing by expert bias updates, outside the loss function — no quality tax.
FP8
Numerics
Fine-grained block-wise scaling keeps 8-bit training stable through 14.8T tokens.
MTP
Dense signal
One forward pass trains predictions for several future positions — free extra supervision.
vs OPEN MODELS
outperforms
comprehensive evals across knowledge, reasoning, coding
vs CLOSED MODELS
comparable
performance on par with leading closed-source models
TOTAL COST
2.788M
H800 GPU-hours — pre-train through post-train
SPARSITY
37B / 671B
5.5% of parameters active per token
Interactive Demo — The Activation Gap

Press run to compare total vs activated parameters — the MoE story that makes a 671B model serve like a 37B one.

Legacy

Legacy — The Repricing Event

V3's bill changed how the industry budgets, benchmarks, and believes.

📉 The cost-belief reset
2.788M H800 GPU-hours for frontier-comparable quality forced every lab and investor to re-derive what frontier training must cost — and what moats compute actually buys.
🧠 MLA and aux-free MoE go canonical
The two architecture pieces became the default spine for the DeepSeek family and were widely studied as the reference 'efficient attention + routing' combination.
🔢 FP8 legitimacy
A flagship open model trained end-to-end in FP8 — 8-bit training moved from research curiosity to production default.
🚀 The R1 on-ramp
V3's base is the substrate DeepSeek-R1 (entry #47) reasoning models were trained on — the efficiency shock and the reasoning shock share a foundation.
⚠️ What it did NOT solve
The bill excludes the amortized research program; expert-parallel serving of 671B parameters still demands serious infrastructure; English/Chinese dominate the training mix; and safety alignment at frontier capability remains the open project it always was.
🛤 Read next
The family: DeepSeekMoE · DeepSeek-R1 · the alternative spine: Mamba
Test Yourself

Quick Quiz

Check your understanding of the key concepts from DeepSeek-V3.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 671B total / 37B activated MoE, trained on 14.8T tokens — outperforms open peers, comparable to leading closed models.
✅ Four pillars: MLA (latent KV cache), aux-loss-free MoE (bias-based balancing), FP8 numerics, multi-token prediction.
✅ The full run cost 2.788M H800 GPU-hours — frontier-class quality without frontier-class waste.
✅ Post-training (SFT + RL) carries the R1 lineage forward; the base is the reasoning substrate.
✅ Every efficiency was known separately; V3's contribution is disciplined integration at scale.
✅ Read it as the proof that architecture and engineering — not just capital — set the frontier.