Open Weight Models
Home›Knowledge›Hardware
Hardware · source-first reference

LLM Quantization Explained

Quantization reduces the precision used to represent model weights — and sometimes activations — so a model can consume less memory and often run faster. The hard part is understanding what a label like 4-bit actually means for quality, runtime compatibility and hardware.

Direct answer

Quantization represents model weights or activations with lower-precision numeric formats. Moving from FP16/BF16 toward 8-bit or 4-bit can substantially reduce weight memory, making larger models practical on smaller hardware. The trade-off is that quantization is not free: quality, throughput and compatibility depend on the algorithm, calibration, kernel support, model architecture and runtime. “4-bit” describes a precision target, not one universal file format or quality level.

GoalReduce memory + compute
Common levels8-bit and 4-bit
Trade-offAccuracy + compatibility
Not equivalentINT4 ≠ every 4-bit method
Precision ladder — approximate raw weight storage
FP324 bytes / param
FP16/BF162 bytes / param
FP8~1 byte / param
INT8~1 byte + metadata
4-bit~0.5 byte + metadata

What quantization changes

Neural-network weights are numeric values. A high-precision checkpoint may store them in FP32, FP16 or BF16. Quantization maps those values into lower-precision representations, often using scaling factors and group-wise metadata so the approximation remains useful.

The most obvious benefit is memory. A weight tensor stored at half the bits needs roughly half the raw storage, although real formats add metadata and alignment overhead. Lower precision can also improve inference throughput when the hardware and runtime have optimized kernels for the chosen representation.

Some methods quantize only weights. Others also quantize activations, caches or intermediate computations. That distinction matters when interpreting a format label.

FP16, FP8, INT8 and 4-bit are families, not guarantees

PrecisionTypical roleMain advantageMain caution
FP16 / BF16High-quality inference and trainingBroad support, low quantization lossLarge memory footprint
FP8Modern accelerator inference/trainingHigh throughput on supported GPUsHardware and kernel dependent
INT8Memory-efficient inferenceRoughly halves weight memory vs FP16Outliers and calibration matter
4-bitLocal inference and QLoRAVery large memory reductionQuality varies strongly by method/model

Hugging Face’s Transformers documentation supports multiple quantization approaches, including bitsandbytes, AWQ and GPTQ. These are not interchangeable. A runtime may support one method well and another poorly.

Why quantization can affect quality

Quantization replaces many precise numbers with a smaller set of representable values. If important weight differences are lost, model behavior can change. Good algorithms reduce this error by choosing scales carefully, handling outliers, quantizing in groups or calibrating on representative data.

Quality sensitivity is workload-dependent. A quantization that looks nearly identical for casual chat may have a larger effect on code generation, long-context retrieval, multilingual output or mathematical reasoning. Evaluate on the tasks that matter instead of assuming a bit-width guarantees a fixed loss.

Operational rule: Compare quantizations on your own prompt distribution, not only on file size or a single benchmark.

Quantization method versus file format

GGUF is commonly associated with local quantized models because llama.cpp and related tools use GGUF to store tensors plus metadata. But GGUF is a container format, not a synonym for one quantization algorithm. A GGUF file can use different tensor types and quantization schemes.

Safetensors is also a tensor serialization format rather than a quantization method. A repository can contain Safetensors weights in BF16, FP16 or a quantized representation supported by the surrounding framework.

Keep these layers separate: numeric representation → quantization algorithm → serialization format → runtime kernel.

Quantization for fine-tuning: QLoRA

QLoRA made 4-bit base-model loading especially important for adaptation. The method keeps the pretrained model quantized and frozen while gradients train small LoRA adapters. The original paper introduced NF4, double quantization and paged optimizers to reduce memory during fine-tuning.

This does not mean “train all 4-bit weights.” In common QLoRA workflows, the quantized base weights remain fixed and the adapter parameters are trained.

How to choose a quantization

  1. Start with the model and runtime you actually plan to use.
  2. Find publisher- or ecosystem-supported quantizations for that architecture.
  3. Choose the highest precision that comfortably fits your hardware and latency target.
  4. Measure task quality, context behavior and throughput.
  5. Record the exact quantization method and file, not just “4-bit.”

For local use

4-bit GGUF-style quantizations are often attractive when RAM/VRAM is the limiting factor. For production GPU serving, FP8, AWQ, GPTQ or vendor-specific formats may be more appropriate when kernels and accelerators support them well.

Frequently asked questions

Does 4-bit make a model four times smaller than FP16?

Raw weight storage can approach one quarter of FP16 because 4 bits is one quarter of 16 bits, but real files include scales, metadata and other tensors, so the actual ratio is not exact.

Is 4-bit always worse than 8-bit?

Lower precision generally introduces more approximation, but quality depends on the quantization algorithm, model and workload. A good 4-bit method can be surprisingly strong for many tasks.

Is GGUF a quantization algorithm?

No. GGUF is a tensor and metadata file format used heavily by llama.cpp. GGUF files can contain different quantized tensor types.

Can quantized models be fine-tuned?

Yes, but the method matters. QLoRA commonly trains LoRA adapters while keeping a 4-bit quantized base model frozen.

Primary sources and technical references

OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.

Continue the knowledge path