Weights quantize easily; activations don't — a few outlier channels ruin the activation grid. SmoothQuant migrates that difficulty into the weights with an equivalent transform, and W8A8 INT8 serving works.
Why INT8 serving worked for weights but not the rest of the matmul.
The mathematical core: scale each channel's activations by s and its matching weight columns by 1/s — the product is unchanged, so outputs are identical: Y = (X·diag(s)) · (diag(1/s)·W). The insight behind choosing s: activation difficulties are spiky outliers in a few channels, while weight difficulties are smooth and spread out. Migration strength α tunes how much of each channel's scale comes from activations (X) versus weights (W) — pushing activation difficulty onto weights where the grid can absorb it. INT8 for BOTH operands of every matmul, with an equivalence proof in hand.
The asymmetric quantization problem at the heart of INT8 serving.
Two kids on a see-saw: the heavy one (spiky activations) pins the light one (smooth weights) — the game (INT8 quantization) can't be played. SmoothQuant slides the fulcrum (per-channel scale s): the heavy side lightens exactly as the light side heavies, the balance (the model's outputs) never changes — and both sides now sit in the weight class the game requires. Equivalent math, fairer fight.
The see-saw, precisely.
Validated across the model zoo of its era — OPT, BLOOM, GLM, MT-NLG, Llama-1/2, Falcon, Mistral, Mixtral — with negligible accuracy loss at W8A8. Integrated INT8 kernels deliver up to 1.56× speedup and 2× memory reduction, and the flagship deployment claim: serving a 530B model on a single node (8× GPUs). Training-free, general-purpose, turn-key: the paper's own three adjectives.
One hyperparameter, per model — the tuning story.
The fully-integrated serving claim, validated across the model zoo.
Difficulty migration became a permanent quantization design principle.
Check your understanding of the key concepts from SmoothQuant.
Everything you need to remember about this paper.