Memory & Efficiency

Quantization & Precision

Intermediate

Quantization is the single highest-leverage lever for fitting big models on real hardware. By storing each number in fewer bits, you shrink memory and bandwidth - often with near-zero quality loss. This page builds from how a float is laid out in bits, to the methods that do it, to the hardware that makes it fast.

Fewer bits, same model

A model's weights are just billions of numbers. Store each one in 2 bytes instead of 4, and the model halves in size. Drop to 1 byte () or half a byte () and it halves again - and again. The catch is precision: too few bits and the numbers can no longer represent what the model learned. Modern quantization is the science of cutting bits right up to that line, and Blackwell-class hardware now makes the low-precision math faster, not just smaller.

Anatomy of a number - where the bits go

Every format splits its bits into a sign, an exponent (range) and a mantissa (precision). Pick a format, flip its bits, or store a number and compare how every format holds it.

FP8 E4M3 - 8-bit, more precision, 8 bits · click a bit to flip it

neighbor

+1 × 211−7 × 1.25 = 20

sign × 2exponent − bias × 1.mantissa

Store a number:

Decoded value 20

Exponent bits buy range

Smallest to largest positive value, log scale.

1e-401e-301e-201e-1011e101e201e30FP32BF16FP16±65,504FP8 E4M3±448FP8 E5M2±57,344FP4 E2M1±6

Pale = subnormals. The orange line is your number. FP32 and BF16 span the same range; FP8 E5M2 gives up a mantissa bit for about 128× more range than E4M3.

Mantissa bits buy precision

Every FP8 E4M3 value from 16 to 32, linear scale.

1618202224262830328 values, each 2 apart

3 mantissa bits = 23 values per doubling. The next band up has the same count, twice as far apart.

sign / your numberexponent - sets the rangemantissa - sets the precision hollow = bit is 0

FP8 E4M3 - Forward pass weights & activations on Hopper/Blackwell. More mantissa, less range. Precision is relative: every step up in exponent doubles the gap between neighboring values, so a format with few mantissa bits is fine-grained near zero and coarse near its maximum.

The format ladder

Each rung halves the bytes. is the lossless default; FP8 is the current serving sweet spot; /INT4 push the frontier with per-block scaling to stay accurate.

Bytes per value, rung by rung

Each bar is drawn to scale in bytes, with a tick per bit. Tap a rung to see what it means for a 70B-parameter model.

Per value

2 B

16 bits

vs FP32

2× smaller

50% of the bytes

70B weights

140 GB

70B × 2 B

Same bytes is not the same number: FP16 and BF16 are both 2 B but split their bits differently (see the anatomy above), and INT4 is an evenly spaced grid while FP4 is a tiny float.

Size it: precision vs. what fits

The same model swings wildly in footprint depending on precision. A 70B model needs 4 GPUs at FP32 but fits a single H100 at INT4. Pick a model and watch the GPU count - and the quality cost - change.

Precision explorer - how many GPUs does it take?

Each box is one GPU's memory (HBM). The weights pour in first, then the KV cache for the workload you pick. Fewer bytes per parameter, fewer boxes.

Model
GPU
KV cache
Context
Sequences
weights (tinted by precision) KV cache free HBM●●●○ quality (illustrative)

INT4 (~97% quality) - GPTQ/AWQ 4-bit. Small quality hit, huge memory win.

Llama 3.3-70B at INT4: 35 GB of weights + 21 GB of KV cache (8 × 8K tokens in BF16) → fits one H100. The cache is 38% of what has to fit - shrinking the weights alone stops paying off; quantize the cache too.

Poured in order for clarity - real deployments shard weights and cache evenly across GPUs and need headroom for activations. Workload and quality figures are illustrative; real degradation is task- and method-dependent. Below INT4, per-block scaling (NVFP4/MXFP4) becomes essential to hold quality.

What, exactly, gets quantized

Quantization isn't one knob. You can quantize the weights, the activations, or the KV cache - independently - and each targets a different bottleneck.

Three targets in one layer

Each block's width is drawn to scale in bits; the dashed outline is its BF16 size. Pick a scheme to see which tensors shrink.

One layer, one forward step

X activations · W weights · Y output

matmul in BF16 - weights dequantized on the fly

KV cache16-bit · BF16

one slot per cached token - grows with every request

Weight-only

Quantize weights, keep activations in BF16. Dequantize on the fly.

Cuts the dominant memory term; minimal quality loss.

The three are independent knobs - a deployment can combine them. Block shapes are illustrative.

How it's done

The toolbox splits into (fast, no retraining) and (best quality, expensive). These are the names you'll meet in every model card.

PTQ

minutes · no retraining

Post-Training Quantization

Trained BF16 model

frozen weights, already converged

↓

Calibrate on a few batches

measure activation ranges → pick scales

↓

Round weights to low-bit

GPTQ / AWQ / SmoothQuant

↓

Quantized model ships

small quality drop, no GPUs spent

QAT

full training run · best quality

Quantization-Aware Training

Train with fake-quant inserted

forward pass rounds, gradients stay full-precision

↓

Model learns to tolerate rounding

weights adapt around the coarse grid

↓

Bake in the real low-bit weights

no calibration gap to recover

↓

Quantized model ships

holds accuracy even at 2–4 bits

The named recipes, mapped

Columns: when quantization happens. Rows: what gets quantized. Tap any recipe - or PTQ / QAT themselves.

after training · minutes

during training · a full run

Weights only

4-bit GPU serving

-

Weights + activations

W8A8 · INT8 / FP8

-

Fine-tune a quantized base

4-bit frozen weights

-

CPU / edge formats

block-quantized files

-

GPTQ / AWQ

Weight-only PTQ to 4-bit. GPTQ minimizes layer-wise error; AWQ protects the most salient weight channels. The backbone of 4-bit serving.

Try it: round the weights yourself

Quantization is, at heart, snapping each weight onto a coarse grid. Drop the bit-width to see the grid get coarser and the rounding error grow - then toggle per-block scaling to watch that error shrink back. This is the exact trick behind NVFP4/MXFP4.

Quantization playground - round the weights, watch the error

A weight matrix is just numbers. Quantizing snaps each one onto a grid of allowed values (the faint lines). Fewer bits means a coarser grid and bigger rounding error - per-block scaling fights back by giving each block of 4 its own grid.

Precision
+0.920−0.92scale 0.92scale 0.67scale 0.91scale 0.60original 0.920quantized 0.920original -0.130quantized -0.131original 0.450quantized 0.394original -0.780quantized -0.789original 0.210quantized 0.191original 0.670quantized 0.670original -0.340quantized -0.383original 0.080quantized 0.096original -0.550quantized -0.520original 0.390quantized 0.390original 0.070quantized 0.130original -0.910quantized -0.910original 0.140quantized 0.171original -0.270quantized -0.257original 0.600quantized 0.600original -0.480quantized -0.514block 1 RMS 0.028block 2 RMS 0.025block 3 RMS 0.034block 4 RMS 0.024
original weight quantized value allowed levels ± scale (block max) rounding error · bars = RMS error per block

Levels

15

16 codes, 1 unused (symmetric)

Bytes / value

0.50

vs 2 B (BF16)

RMS error

0.0279

lower is better

Max error

0.0600

worst single weight

Per-block scaling gives each block of 4 its own range, so a block of small weights doesn't waste its levels on an outlier elsewhere - see the tighter grid in blocks 2 and 4. Turn it off to see the error jump - that's the whole idea behind NVFP4/MXFP4.

Hardware makes low precision fast

Smaller numbers aren't just cheaper to store - on the right silicon they're cheaper to multiply. Each GPU generation adds native support for a lower format, roughly doubling throughput. Hopper made FP8 a first-class citizen (~2× BF16); Blackwell adds FP4/NVFP4 (~2× FP8 again), with per-block micro-scaling to hold accuracy at 4 bits.

Each generation adds a faster, smaller format

Relative Tensor Core throughput, stacking the ~2× steps described here - not measured TFLOPS. Tap a generation.

BF16

2 B / value

FP8

1 B / value

FP4

0.5 B / value

Tap Hopper to see the ladder before native FP4 arrived.

Quantization composes with everything

Lower precision shrinks weights and the KV cache, and it changes how many GPUs your parallelism plan needs. It's the cheapest lever you have - pull it first, then size the rest.