History Problem Core Idea Alpha Results Impact Quiz Takeaways
Interactive Paper Explainer

Move the Difficulty,
Keep the Math
SmoothQuant

Weights quantize easily; activations don't — a few outlier channels ruin the activation grid. SmoothQuant migrates that difficulty into the weights with an equivalent transform, and W8A8 INT8 serving works.

Start Learning Read the Paper ↗
W8A8
Weights + activations
1.56×
Speedup (with 2× memory cut)
530B
On a single node
2022
Xiao et al.
History

The Activation Outlier Wall

Why INT8 serving worked for weights but not the rest of the matmul.

2022
Weights quantize fine
GPTQ-class methods (entry #112) compress weights to 4 bits — but the matmul still multiplies FP16 activations: memory saves, compute drags.
2022
The outlier discovery
LLM activations develop a few persistent outlier channels — magnitudes tens of times the rest, wrecking the INT8 activation grid.
Nov 2022
🚀 SmoothQuant
Xiao et al.: scale activations down, weights up (per channel, by migration strength α) — a mathematically equivalent transform. Both sides become quantization-friendly; W8A8 works everywhere.
2022-23
Serving-era standard
Fully-integrated INT8 kernels ship in inference stacks; the technique generalizes across model families and later quantization schemes.
2023+
The efficiency stack
SmoothQuant composes with weight-only methods and FP8 (entry #19) — difficulty migration becomes a standard quantization design pattern.
Equivalence Is the Trick

The mathematical core: scale each channel's activations by s and its matching weight columns by 1/s — the product is unchanged, so outputs are identical: Y = (X·diag(s)) · (diag(1/s)·W). The insight behind choosing s: activation difficulties are spiky outliers in a few channels, while weight difficulties are smooth and spread out. Migration strength α tunes how much of each channel's scale comes from activations (X) versus weights (W) — pushing activation difficulty onto weights where the grid can absorb it. INT8 for BOTH operands of every matmul, with an equivalence proof in hand.

Chapter 01

Weights Yield, Activations Don't

The asymmetric quantization problem at the heart of INT8 serving.

📈
The Outlier Regime
  • LLM activations grow a few persistent outlier channels — huge, systematic, model-wide
  • INT8 activation grids clip or lose precision around outliers: quantization error explodes
  • Weight-only quantization (GPTQ) saves memory but leaves activation matmuls in FP16 — compute speedup forgone
  • Mixed-precision W8A16 workarounds are hardware-inefficient — the INT8 units sit idle
⚖
The SmoothQuant Answer
  • Per-channel, mathematically equivalent rescaling: activations ÷ s, weights × 1/s
  • Migration strength α: 0 = all-weight difficulty; 1 = all-activation — tuned per model
  • Full W8A8 INT8 for every matmul — hardware-efficient, fully-integrated kernels
  • Up to 1.56× speedup + 2× memory reduction; 530B serving on a single node
Analogy — The See-Saw

Two kids on a see-saw: the heavy one (spiky activations) pins the light one (smooth weights) — the game (INT8 quantization) can't be played. SmoothQuant slides the fulcrum (per-channel scale s): the heavy side lightens exactly as the light side heavies, the balance (the model's outputs) never changes — and both sides now sit in the weight class the game requires. Equivalent math, fairer fight.

Chapter 02

The Transform

The see-saw, precisely.

The math
  • Y = X·W = (X·diag(s)) · (diag(s)⁻¹·W)
  • s_j chosen per input channel j — typically s_j = max(|X_j|)^α / max(|W_j|)^(1−α)
  • α = migration strength: 0.5 default; higher for spikier models
  • Offline: transform weights once; runtime: activations pass through a cheap diag multiply
Why it works
  • Activation difficulty: concentrated (a few outlier channels)
  • Weight difficulty: diffuse (spread across channels)
  • Migration moves concentrated spikiness to where the grid is roomy — both sides end in INT8's comfort zone
Coverage and results

Validated across the model zoo of its era — OPT, BLOOM, GLM, MT-NLG, Llama-1/2, Falcon, Mistral, Mixtral — with negligible accuracy loss at W8A8. Integrated INT8 kernels deliver up to 1.56× speedup and 2× memory reduction, and the flagship deployment claim: serving a 530B model on a single node (8× GPUs). Training-free, general-purpose, turn-key: the paper's own three adjectives.

Interactive Demo — Turn the Alpha Dial

Slide the migration strength and watch difficulty move between activations and weights — with the failure modes at both ends.

s_j = max|X_j|^α / max|W_j|^(1−α) — the dial is calibrated per model on held-out data; 0.5 is the shipped default.
Chapter 03

The Alpha Dial

One hyperparameter, per model — the tuning story.

Migration Strength
Interactive Demo — One Matmul, Transformed

Follow a single matrix multiplication through the equivalent rescaling — same output, INT8-friendly operands.

Chapter 05

INT8, Everywhere

The fully-integrated serving claim, validated across the model zoo.

Y = X·W = (X·diag(s)) · (diag(1/s)·W)  ·  s_j = max|X_j|^α / max|W_j|^(1−α)
diag(s)
The see-saw
Per-channel scaling — activations divided, weights multiplied, product unchanged.
α
Migration strength
How much outlier difficulty moves from activations to weights — the one dial, calibrated per model.
outliers
The problem
A few activation channels with huge persistent magnitudes — the INT8 activation grid's nemesis.
W8A8
The prize
Both operands in INT8 for every matmul: real compute speedup, not just memory savings.
PRECISION
W8A8
every matmul, both operands, INT8
SPEEDUP
up to 1.56×
with 2× memory reduction
FLAGSHIP
530B / 1 node
single-node serving, 8 GPUs
COVERAGE
9+ families
OPT/BLOOM/GLM/MT-NLG/Llama/Falcon/Mistral/Mixtral…
Interactive Demo — Why Not Just Fix the Outliers in Training?

Outlier channels could be regularized away at pre-training. Press reveal.

Legacy

Legacy — The Migration Pattern

Difficulty migration became a permanent quantization design principle.

⚡ The W8A8 serving era
Fully-integrated INT8 kernels became standard in inference stacks — real compute speedup (not just memory savings) for deployment economics.
🧩 The composability
The pattern composes with weight-only quantization and FP8 training — difficulty migration appears across the modern efficiency stack (entry #19).
📊 The outlier reframing
Treating activation outliers as migratable rather than pathological influenced quantization research broadly — including AWQ's salient-channel view (entry #114).
⚠️ What it did NOT solve
INT8 activation precision still bounds very-low-bit regimes (W4A8 and below need more); α calibration is per-model manual work; and the 1.56× ceiling reflects memory-bandwidth limits — FP8 hardware later leapfrogged INT8 serving entirely.
🛤 Read next
The quantization ladder: GPTQ · AWQ · DeepSeek-V3 (FP8)
Test Yourself

Quick Quiz

Check your understanding of the key concepts from SmoothQuant.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ SmoothQuant: equivalent per-channel rescaling migrates activation-outlier difficulty into weights.
✅ s_j = max|X_j|^α / max|W_j|^(1−α) — one calibration dial per model.
✅ Full W8A8 INT8 for every matmul: up to 1.56× speedup, 2× memory reduction.
✅ 530B-class serving on a single node; validated across 9+ model families.
✅ Training-free, turn-key, equivalence-proven — deployment-era engineering.
✅ Read it as the pattern that made activations quantizable: difficulty can move.