Quantization is the process of converting an AI model’s internal numbers from a high-precision format, like 32-bit floating point, into a lower-precision format, like 8-bit or 4-bit integers — shrinking the model’s file size and speeding up its calculations, usually with only a small, carefully managed drop in accuracy. It’s the main reason a model with billions of parameters can run on a laptop, a phone, or a single GPU instead of a rack of servers.
How Does Quantization Actually Work?
Every number inside a neural network — its weights and activations — is normally stored as a 32-bit floating-point value (FP32), which can represent roughly 4 billion distinct values. Quantization maps those values onto a much smaller set of numbers, such as the 256 values available in 8-bit integers (INT8), using a scaling factor and, in some methods, a zero-point offset to keep the mapping as accurate as possible. Two common approaches are:
- Absolute max scaling: divides every value by the largest value in the set, then scales the result into the target range.
- Affine (zero-point) quantization: combines a scaling factor with an offset, which tends to preserve accuracy better when the original values aren’t centered around zero.
Because INT8 math requires far simpler circuitry than FP32 math, a quantized model needs less memory bandwidth and fewer processor cycles to run the same calculation — which is what makes it faster and cheaper to serve.
What Do FP32, FP16, INT8, and INT4 Actually Mean?
These labels describe how many bits are used to store each number in the model, and each step down roughly halves the memory footprint:
- FP32 (32-bit float): the default training precision, most accurate, largest and slowest.
- FP16 / BF16 (16-bit float): half the size of FP32, the common precision for running (not training) large models with minimal accuracy loss.
- INT8 (8-bit integer): a quarter the size of FP32, widely used for deployment on standard hardware.
- INT4 (4-bit integer): an eighth the size of FP32, used when memory is extremely limited, at a higher risk of accuracy loss.
What’s the Accuracy Tradeoff?
Squeezing a model’s numbers into fewer bits introduces what’s called quantization error — small rounding differences that accumulate across the millions or billions of calculations a model performs. For most everyday tasks, INT8 quantization is nearly indistinguishable from the full-precision original, which is why it has become a default for deployment. Pushing further to INT4 saves even more memory but increases the chance that the model’s outputs become noticeably less reliable, especially on tasks that require precise reasoning or rare, fine-grained distinctions. Teams that quantize aggressively typically re-test the model on their own evaluation set afterward, rather than assuming the published benchmarks still hold.
When Should You Actually Use a Quantized Model?
Quantization is worth reaching for whenever the deployment target has limited memory, limited power, or a tight latency budget: running a model on a phone, a browser, a single consumer GPU, or any setup where every millisecond and every gigabyte of VRAM is being paid for. It’s a complementary technique to approaches like knowledge distillation, which shrinks a model by training a smaller one to mimic a larger one, and to fine-tuning, which adapts a model’s behavior rather than its size — the three techniques are often combined: distill or fine-tune first, then quantize the result for deployment.
Frequently Asked Questions
Does quantization make a model “dumber”?
Not usually at INT8, where the accuracy drop is typically small enough to be unnoticeable in everyday use. The risk grows at INT4 and below, where more aggressive compression can measurably hurt performance on tasks that need precise reasoning, so it’s worth testing the quantized version against your own use cases rather than assuming it behaves identically.
Is quantization the same thing as compression?
They’re related but not identical. Quantization specifically reduces the numerical precision of a model’s weights and activations. Other compression techniques — like pruning (removing less-important weights entirely) or distillation (training a smaller model to copy a larger one) — shrink a model in different ways, and are often used alongside quantization rather than instead of it.
Can you quantize any AI model?
Most modern neural networks, including large language models, can be quantized, and popular open-weight models are frequently released with ready-made quantized versions. The practical limit is usually how much accuracy loss a given task can tolerate, not whether quantization is technically possible.
Photo: computer processor chips by Edward Wilders, licensed CC BY-SA 4.0, via Wikimedia Commons.
Further reading: IBM’s explainer on quantization and the research paper comparing BF16 and quantized precision go deeper into the math behind these tradeoffs. For related reading on this blog, see What Is Knowledge Distillation? and What Is Fine-Tuning?



