History Problem Core Idea Results Results Impact Quiz Takeaways
Interactive Paper Explainer

175 Billion Parameters,
One GPU-Hour Each
GPTQ

Post-training quantization with approximate second-order information: OPT-175B and friends compressed to 3-4 bits per weight in about four GPU-hours — no retraining, negligible loss.

Start Learning Read the Paper ↗
3-4
Bits per weight
~4
GPU-hours for 175B
2×+
Compression vs prior PTQ
2022
Frantar et al.
History

Quantization's One-Shot Dream

Getting small without training — the efficiency frontier before GPTQ.

2020-21
PTQ for CNNs, dreaming big
Post-training quantization works on vision models; LLM scale breaks the assumptions (one-shot error accumulates layer by layer).
2021
Round-to-nearest fails at scale
Naive RTN quantization of GPT/OPT-class models at 3-4 bits degrades severely — outliers and depth compound the error.
Oct 2022
🚀 GPTQ
Frantar et al.: one-shot quantization via approximate second-order (Hessian-based) information — accurate at 3-4 bits, fast enough for 175B in hours.
2022-23
The consumer-model wave
4-bit GPTQ models run on single GPUs; llama.cpp-class ecosystems and hobbyist serving adopt the recipe (with AWQ to follow, entry #114).
2023+
The quantization ladder
SmoothQuant (entry #113), AWQ, and FP8 training inherit the problem definition — GPTQ set the PTQ baseline every method beats against.
Use the Curvature

Round-to-nearest ignores that weights matter unequally: perturbing a weight in a flat region of the loss is free; perturbing one in a sharp region is expensive. GPTQ quantizes layer by layer, column by column, using the local Hessian (second-order approximation of the loss): when a column is quantized, the error it introduces is compensated by adjusting the not-yet-quantized columns — error is pushed onto the directions the loss surface says are cheapest. The result: extreme bit-widths at accuracy round-to-nearest cannot touch, in one pass, without backpropagation through the model.

Chapter 01

Big Models, Small Machines

The deployment wall: inference costs lock research models out of production and hobbyists out entirely.

💸
The Inference Wall
  • 175B-class models need multiple high-end GPUs just to serve — usability, not capability, is the bottleneck
  • Fine-tuning-based quantization (QAT) is absurdly expensive at this scale
  • Round-to-nearest PTQ collapses at 3-4 bits on LLMs — error accumulates catastrophically with depth
  • Prior accurate PTQ research didn't scale to models this large
📐
The GPTQ Answer
  • One-shot PTQ with approximate second-order (Hessian) information per layer
  • Quantize column by column; compensate error on remaining columns — always
  • 175B parameters to 3-4 bits in ~4 GPU-hours, negligible degradation vs FP16
  • More than doubles compression vs prior one-shot methods at matched quality
Analogy — The Cost Accountant

Round-to-nearest rounds every number to the nearest cent and lets the books drift. GPTQ hires an accountant who knows which accounts matter (the Hessian): round a number here, and immediately adjust the accounts that can absorb the difference — the ledger balances at every step. At the end of the pass, the totals (model outputs) barely moved, though every number is stored in cents (3-4 bits).

Chapter 02

The Algorithm

Layer-wise, column-wise, Hessian-guided — the machinery.

Step by step
  • For each layer: gather calibration inputs; compute the Hessian approximation (input second-moment matrix)
  • Process the weight matrix column by column: quantize column j to the grid
  • Compute the induced error; propagate compensation onto the remaining un-quantized columns via the Hessian
  • Efficiency trick: a Cholesky-based reformulation of the inverse Hessian makes the whole pass near-matrix-multiply speed
Why it is fast AND accurate
  • One forward pass of calibration data — no backprop, no training loop
  • Approximate second-order ≈ exact-enough local error model, cheap to compute
  • Error never accumulates silently — it is re-assigned at every column
  • 175B parameters in ~4 GPU-hours — quantization at model scale, finally
Interactive Demo — One Layer, Column by Column

Watch the Hessian-guided pass quantize a weight matrix — every column's error gets re-homed, not ignored.

Chapter 03

The Ledger

Verified results — the compression ladder.

From the paper

The released code became the tooling backbone for the 4-bit open-model ecosystem — quantized checkpoints of every major open LLM shipped as GPTQ artifacts throughout 2023.

Interactive Demo — RTN vs GPTQ at Low Bits

Tab through bit-widths — watch round-to-nearest collapse while Hessian guidance holds.

Chapter 05

4 Bits, 4 Hours

The one-shot quantization frontier, moved decisively.

quantize column j  →  update remaining columns by −H⁻¹-based error compensation
column j
The unit of work
Weights quantize one column at a time — never a whole matrix blind.
H
The Hessian (approx.)
Second-moment of layer inputs: how sensitive the loss is to each weight direction — the cost map.
compensation
The balance
The quantization error is re-assigned to un-quantized columns along cheap directions — the ledger stays near-balanced.
Cholesky trick
The speed
A reformulated inverse-Hessian update makes the pass run at near-GEMM speed — 175B in hours.
BIT-WIDTH
3-4 bits
OPT-175B / BLOOM-176B, negligible loss
TIME
~4 GPU-hrs
for the 175B class — no retraining
vs PRIOR PTQ
2×+ compression
at matched accuracy
DEPLOYMENT
single GPU
175B-class 4-bit inference in practice
Interactive Demo — The Compression Ladder

Press run for the compression ratios by bit-width — and where GPTQ sits against prior one-shot PTQ.

Legacy

Legacy — The 4-Bit Commonwealth

GPTQ made open models privately ownable.

🖥 Consumer-scale LLMs
4-bit GPTQ checkpoints let 30B-70B models run on single consumer GPUs — the local-AI movement's enabling technology.
🔧 The PTQ baseline
Every subsequent method (SmoothQuant, AWQ, extensions) benchmarks against GPTQ — it defined the accuracy-at-bits frontier.
📦 Tooling standard
The released implementation became the ecosystem's quantization toolchain — model repos shipped GPTQ artifacts as a default format.
⚠️ What it did NOT solve
Weights only — activations stay FP16 (dequantize-then-compute: speedups lag memory savings); outlier features still bite at extreme bits; and per-model calibration runs are required (vs training-free AWQ's speed).
🛤 Read next
The quantization ladder: SmoothQuant · AWQ · QLoRA
Test Yourself

Quick Quiz

Check your understanding of the key concepts from GPTQ.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ GPTQ: one-shot, Hessian-guided layer-wise PTQ — no retraining, no backprop.
✅ Column-by-column quantization with error compensation on remaining columns.
✅ 175B models to 3-4 bits in ~4 GPU-hours, negligible degradation vs FP16.
✅ More than 2× the usable compression of prior one-shot PTQ at matched quality.
✅ Enabled the 4-bit open-model ecosystem: 30B-70B models on consumer GPUs.
✅ Read it as the deployment-wall removal that made open weights privately ownable.