Post-training quantization with approximate second-order information: OPT-175B and friends compressed to 3-4 bits per weight in about four GPU-hours — no retraining, negligible loss.
Getting small without training — the efficiency frontier before GPTQ.
Round-to-nearest ignores that weights matter unequally: perturbing a weight in a flat region of the loss is free; perturbing one in a sharp region is expensive. GPTQ quantizes layer by layer, column by column, using the local Hessian (second-order approximation of the loss): when a column is quantized, the error it introduces is compensated by adjusting the not-yet-quantized columns — error is pushed onto the directions the loss surface says are cheapest. The result: extreme bit-widths at accuracy round-to-nearest cannot touch, in one pass, without backpropagation through the model.
The deployment wall: inference costs lock research models out of production and hobbyists out entirely.
Round-to-nearest rounds every number to the nearest cent and lets the books drift. GPTQ hires an accountant who knows which accounts matter (the Hessian): round a number here, and immediately adjust the accounts that can absorb the difference — the ledger balances at every step. At the end of the pass, the totals (model outputs) barely moved, though every number is stored in cents (3-4 bits).
Layer-wise, column-wise, Hessian-guided — the machinery.
Verified results — the compression ladder.
The released code became the tooling backbone for the 4-bit open-model ecosystem — quantized checkpoints of every major open LLM shipped as GPTQ artifacts throughout 2023.
The one-shot quantization frontier, moved decisively.
GPTQ made open models privately ownable.
Check your understanding of the key concepts from GPTQ.
Everything you need to remember about this paper.