Quantization represents model weights or activations with lower-precision numeric formats. Moving from FP16/BF16 toward 8-bit or 4-bit can substantially reduce weight memory, making larger models practical on smaller hardware. The trade-off is that quantization is not free: quality, throughput and compatibility depend on the algorithm, calibration, kernel support, model architecture and runtime. “4-bit” describes a precision target, not one universal file format or quality level.
What quantization changes
Neural-network weights are numeric values. A high-precision checkpoint may store them in FP32, FP16 or BF16. Quantization maps those values into lower-precision representations, often using scaling factors and group-wise metadata so the approximation remains useful.
The most obvious benefit is memory. A weight tensor stored at half the bits needs roughly half the raw storage, although real formats add metadata and alignment overhead. Lower precision can also improve inference throughput when the hardware and runtime have optimized kernels for the chosen representation.
Some methods quantize only weights. Others also quantize activations, caches or intermediate computations. That distinction matters when interpreting a format label.
FP16, FP8, INT8 and 4-bit are families, not guarantees
| Precision | Typical role | Main advantage | Main caution |
|---|---|---|---|
| FP16 / BF16 | High-quality inference and training | Broad support, low quantization loss | Large memory footprint |
| FP8 | Modern accelerator inference/training | High throughput on supported GPUs | Hardware and kernel dependent |
| INT8 | Memory-efficient inference | Roughly halves weight memory vs FP16 | Outliers and calibration matter |
| 4-bit | Local inference and QLoRA | Very large memory reduction | Quality varies strongly by method/model |
Hugging Face’s Transformers documentation supports multiple quantization approaches, including bitsandbytes, AWQ and GPTQ. These are not interchangeable. A runtime may support one method well and another poorly.
Why quantization can affect quality
Quantization replaces many precise numbers with a smaller set of representable values. If important weight differences are lost, model behavior can change. Good algorithms reduce this error by choosing scales carefully, handling outliers, quantizing in groups or calibrating on representative data.
Quality sensitivity is workload-dependent. A quantization that looks nearly identical for casual chat may have a larger effect on code generation, long-context retrieval, multilingual output or mathematical reasoning. Evaluate on the tasks that matter instead of assuming a bit-width guarantees a fixed loss.
Quantization method versus file format
GGUF is commonly associated with local quantized models because llama.cpp and related tools use GGUF to store tensors plus metadata. But GGUF is a container format, not a synonym for one quantization algorithm. A GGUF file can use different tensor types and quantization schemes.
Safetensors is also a tensor serialization format rather than a quantization method. A repository can contain Safetensors weights in BF16, FP16 or a quantized representation supported by the surrounding framework.
Keep these layers separate: numeric representation → quantization algorithm → serialization format → runtime kernel.
Quantization for fine-tuning: QLoRA
QLoRA made 4-bit base-model loading especially important for adaptation. The method keeps the pretrained model quantized and frozen while gradients train small LoRA adapters. The original paper introduced NF4, double quantization and paged optimizers to reduce memory during fine-tuning.
This does not mean “train all 4-bit weights.” In common QLoRA workflows, the quantized base weights remain fixed and the adapter parameters are trained.
How to choose a quantization
- Start with the model and runtime you actually plan to use.
- Find publisher- or ecosystem-supported quantizations for that architecture.
- Choose the highest precision that comfortably fits your hardware and latency target.
- Measure task quality, context behavior and throughput.
- Record the exact quantization method and file, not just “4-bit.”
For local use
4-bit GGUF-style quantizations are often attractive when RAM/VRAM is the limiting factor. For production GPU serving, FP8, AWQ, GPTQ or vendor-specific formats may be more appropriate when kernels and accelerators support them well.
Frequently asked questions
Does 4-bit make a model four times smaller than FP16?
Raw weight storage can approach one quarter of FP16 because 4 bits is one quarter of 16 bits, but real files include scales, metadata and other tensors, so the actual ratio is not exact.
Is 4-bit always worse than 8-bit?
Lower precision generally introduces more approximation, but quality depends on the quantization algorithm, model and workload. A good 4-bit method can be surprisingly strong for many tasks.
Is GGUF a quantization algorithm?
No. GGUF is a tensor and metadata file format used heavily by llama.cpp. GGUF files can contain different quantized tensor types.
Can quantized models be fine-tuned?
Yes, but the method matters. QLoRA commonly trains LoRA adapters while keeping a 4-bit quantized base model frozen.
Apply this knowledge
Move from the concept to a concrete deployment shortlist.
Primary sources and technical references
OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.