⚡ Model Quantization
Calculate what low precision saves, follow rounding and clipping by hand, and test whether those savings survive in a real serving workload.
On this page
Before you start
Review Numerical Computing for floating-point formats and memory units, Linear Algebra for linear layers, and Deep Learning Basics for weights and activations.
You will learn to quantize a small vector, calculate storage including scale metadata, distinguish quantization methods from file formats, and choose a quality/performance experiment for a deployment.
The problem: smaller weights, same useful behavior
A 7-billion-parameter model stores 14 GB of raw FP16 weights at two bytes per parameter. Four-bit weights have a raw payload of 3.5 GB. These are decimal gigabytes, 1 GB = 10⁹ bytes. A deployed model also needs scales, possible zero points, padding, any unquantized layers, activations, kernel workspaces and a KV cache.
Quantization replaces a large set of possible values with a smaller codebook. The engineering question is whether those approximations keep the behavior you need and whether the target runtime can use the compressed representation efficiently.
Work one quantizer by hand
For an affine integer quantizer:
q = clamp(round(x/s) + z, q_min, q_max)
x_hat = s × (q − z)
Here x is the original value, q its integer code, s > 0 the scale in original-value units per code step, z an integer zero point, and x_hat the reconstructed approximation. The clamp prevents codes leaving the available range. Rounding rules must match the implementation.
For a symmetric toy example, use z=0, s=0.5, and codes from −7 to 7. This uses 15 symmetric levels in a four-bit container; one of the 16 bit patterns is unused. Round to nearest, breaking ties away from zero.
| x | x/s | q | x_hat | x_hat − x |
|---|---|---|---|---|
| −3.2 | −6.4 | −6 | −3.0 | 0.2 |
| −0.3 | −0.6 | −1 | −0.5 | −0.2 |
| 0.2 | 0.4 | 0 | 0.0 | −0.2 |
| 2.9 | 5.8 | 6 | 3.0 | 0.1 |
Mean squared reconstruction error is (0.04 + 0.04 + 0.04 + 0.01)/4 = 0.0325, in squared original-value units. For x=4, rounding proposes code 8, which clips to 7 and reconstructs as 3.5. Clipping error differs from ordinary rounding error. Making the range wider can avoid clipping while making each quantization step coarser.
This example minimizes neither task loss nor layer-output error. Two equal weight errors can matter very differently if one multiplies a large activation.
Granularity has a storage cost
Per-tensor quantization shares one scale over a tensor. Per-channel quantization uses separate scales along a defined channel axis. Per-group quantization assigns scales to small groups, such as 128 weights. Always specify the grouped axis and group size; the name alone does not define the layout.
For 128 four-bit weights, payload is 128 × 4/8 = 64 bytes. Add one two-byte scale: 66 bytes, or 4.125 bits/weight. Zero points, alignment and padding would add more. Smaller groups can represent local ranges better, but add metadata and may not have equally efficient kernels.
Where the speed can come from
Low-batch autoregressive decoding can spend much of its time moving weights from device memory. Fewer weight bytes can help there. Larger batches reuse weights across examples, and long contexts add cache traffic; compute, communication or KV access may then dominate.
- Weight-only, such as W4A16: weights use four bits; activations use 16 bits. A runtime may unpack/dequantize weights into higher-precision arithmetic inside a fused kernel. Compression is not itself a native four-bit matrix multiply.
- Weights and activations, such as W8A8: compatible hardware and kernels can use lower-precision matrix operations. Activation ranges vary with input, making calibration and outlier handling important.
- KV-cache quantization: compresses cached attention keys and values. This is a separate choice from weight quantization.
Measure time to first token, decode latency, throughput and peak memory on the real prompt lengths and concurrency. A kernel's peak arithmetic rate is not an end-to-end speedup.
PTQ, QAT and the methods inside them
Post-training quantization (PTQ) starts from trained weights and chooses a quantized representation. Some methods use representative calibration inputs; simple round-to-nearest can use weight statistics alone. PTQ need not update weights by task-loss training, although methods may optimize scales or reconstruction.
Quantization-aware training (QAT) includes quantization effects during training or fine-tuning. Fake-quantization operations round/clamp in the forward pass while a surrogate, often a straight-through estimator, supplies backward gradients. QAT may recover useful behavior, but its training setup must match deployment and improvement is not guaranteed.
| Method or artifact | What it does | What to check |
|---|---|---|
| GPTQ | Uses calibration activations and second-order information to compensate weight-rounding errors | Calibration coverage, grouping, damping and runtime support |
| AWQ | Uses activation statistics to choose channel scaling that reduces important weight errors | Scaling/packing support and quality on your task |
| SmoothQuant | Moves scale between activations and weights before quantizing both | Residual outliers and supported W8A8 kernels |
| GGUF | A container for model tensors and metadata, including quantized types | Exact tensor format and backend compatibility |
GGUF is not a quantization algorithm and is not restricted to CPU execution. GPTQ and AWQ describe methods; their exported checkpoints still need a compatible representation and runtime.
Why calibration activations matter
Write a linear layer as Y = XW: calibration samples are rows of X, with shape n × d_in; W has shape d_in × d_out. GPTQ aims to keep the squared output difference ‖XW − XW_hat‖² small. For each output column, the quadratic curvature is 2XᵀX, often with damping added for stability. Correlated input features can let an unquantized weight compensate for error in a quantized one.
SmoothQuant uses an invertible diagonal scale matrix S: XW = (XS⁻¹)(SW) before quantization. Reducing a large activation channel increases the corresponding weight scale. This trades quantization difficulty between the two; it does not make information loss disappear.
Calibration should cover the intended languages, lengths, domains and modalities. Keep the final evaluation set separate. A calibration sample count is a tuning choice, not a universal constant.
Floating-point formats are not integer grids
Integer quantization usually reconstructs values on a uniformly spaced grid within a group. FP8 and FP4 use nonuniform floating-point codebooks, often with additional tensor or block scaling. E4M3 and E5M2 allocate different exponent/fraction bits; exact finite ranges and special values depend on the format specification.
OCP MXFP4 combines four-bit E2M1 elements with a shared scale per 32-element block. Four bits mean 16 bit patterns total, including signed-zero patterns where defined. In the basic MXFP4 layout, 32 elements use 16 payload bytes plus one scale byte: 4.25 bits/element before alignment. Vendor-specific four-bit formats need not use the same scaling scheme.
Lower precision still needs clipping/range decisions and validation. Check the accelerator generation, supported operations, accumulation precision, packing and kernels instead of assuming every GPU or TPU supports the same format.
Check yourself
Your W4 checkpoint is smaller, but peak serving memory barely changed. What could explain it?
Solution: First separate weight payload from total allocation. A large KV cache or activation/workspace budget may dominate. The loader may expand weights, keep duplicate copies, or leave many tensors at higher precision. Inspect actual allocations and runtime kernels before concluding that quantization failed.
For a calculation: if 128 four-bit values have a two-byte scale and a one-byte zero point, storage is 64 + 2 + 1 = 67 bytes, or 4.1875 bits/value, excluding padding.
Where to go next
Model Serving combines weight storage with cache and scheduler budgets. Evaluation & Benchmarking shows how to decide whether a quality change is acceptable.
References
- GPTQ: layer-output reconstruction and error compensation.
- AWQ: activation-aware weight scaling.
- SmoothQuant: moving quantization difficulty between activations and weights.
- OCP Microscaling Formats specification: MX element formats and shared scales.