Open Weight Models
Home›Knowledge›Hardware
Hardware · source-first reference

How Much RAM and VRAM Do Open-Weight Models Need?

The model file is only part of inference memory. A useful estimate starts with weight storage, then adds KV cache, context, runtime buffers, multimodal components and headroom.

Direct answer

A first-pass estimate for weight memory is parameter count multiplied by bytes per parameter: FP16/BF16 ≈ 2 bytes, 8-bit ≈ 1 byte plus metadata, and 4-bit ≈ 0.5 byte plus metadata. Real RAM/VRAM requirements are higher because inference also needs KV cache, runtime buffers, temporary tensors, tokenizer/runtime state and sometimes vision or audio encoders. Longer context and larger batch sizes can increase memory substantially even when the weights themselves fit.

8B FP16 weights~16 GB
8B 4-bit weights~4–6 GB practical
Context costKV cache grows
Plan forHeadroom + runtime
Inference memory budget
Model weightsUsually the largest fixed component.
KV cacheGrows with context and concurrent sequences.
Runtime buffersKernels, temporary tensors, allocator overhead.
HeadroomLeave capacity for peaks, OS and application.

Start with weight memory, but do not stop there

The easiest calculation is raw weights. An 8B dense model contains roughly eight billion learned parameters. At 16-bit precision, each parameter uses about two bytes, giving roughly 16 GB of weight storage. A 32B model at the same precision is roughly 64 GB.

Quantization changes that arithmetic. An 8-bit representation approaches one byte per parameter; a 4-bit representation approaches half a byte per parameter. In practice, scales, metadata, alignment and unquantized tensors add overhead.

Model scaleFP16/BF16 raw weights8-bit rough weights4-bit rough weights
3B~6 GB~3 GB + overhead~1.5 GB + overhead
8B~16 GB~8 GB + overhead~4 GB + overhead
14B~28 GB~14 GB + overhead~7 GB + overhead
32B~64 GB~32 GB + overhead~16 GB + overhead
70B~140 GB~70 GB + overhead~35 GB + overhead

KV cache: why long context changes the answer

Autoregressive transformers cache attention keys and values from previous tokens so they do not need to recompute the entire prefix on every generation step. This KV cache grows with sequence length, number of concurrent requests, architecture and cache precision.

A model that fits at a short context can therefore run out of memory at a long context. Production servers face the same issue when many user sequences are active at once. Techniques such as paged KV caching, lower-precision caches and prefix sharing can improve utilization, but the memory is still a real capacity constraint.

System RAM versus GPU VRAM

GPU inference is fastest when the active model tensors and caches fit in VRAM. If they do not, some runtimes can offload layers or tensors to system RAM. That can make a model runnable, but data movement across PCIe or other interconnects may reduce throughput.

On unified-memory systems such as Apple Silicon, CPU and GPU share a common memory pool. The practical limit is therefore system memory minus operating-system and application needs. On discrete-GPU systems, large system RAM does not automatically compensate for limited VRAM at full GPU speed.

Local deployment should distinguish “can load” from “runs at acceptable speed.”

MoE memory: active parameters do not equal footprint

For sparse Mixture-of-Experts models, the total checkpoint can be much larger than the active parameter count. A 120B/12B-active model still contains roughly 120B total learned parameters, even though only a subset is used for each token.

Use total parameters to estimate weight storage. Use active parameters to understand compute intensity. Multi-GPU expert sharding can distribute memory, but then communication and runtime support become part of the capacity plan.

Training and fine-tuning need much more memory than inference

Full training requires gradients, optimizer states and activations in addition to model weights. Hugging Face’s memory documentation shows why mixed-precision training can consume several times the raw parameter storage. Parameter-efficient methods reduce this burden by freezing most base weights.

QLoRA is especially relevant when memory is tight because it can keep a 4-bit base model frozen while training small low-rank adapters. Even then, sequence length, activations and optimizer state for the trainable components matter.

A practical sizing workflow

  1. Record total model parameters and whether the architecture is dense or MoE.
  2. Choose the exact precision or quantized checkpoint.
  3. Estimate raw weight storage.
  4. Add context/KV-cache requirements for your target sequence length and concurrency.
  5. Add runtime and application overhead.
  6. Leave safety headroom instead of planning to 100% memory utilization.
  7. Benchmark throughput and latency on the actual runtime.
Good sizing question: “What memory do we need for this checkpoint, precision, context, batch size and runtime?” not “How much VRAM does a 14B model need?”

Frequently asked questions

Can an 8B model run on an 8 GB GPU?

Sometimes with a suitable 4-bit quantization and modest context, but the exact answer depends on model architecture, runtime overhead and KV cache. Weight size alone is not enough.

Does more system RAM help if VRAM is too small?

Some runtimes can offload to system RAM, making larger models runnable, but throughput may fall because data moves between memory domains.

Why does context length increase memory?

Attention KV cache stores information for prior tokens. Longer sequences and more concurrent requests increase cache memory.

Should I size hardware from the model's active MoE parameters?

No. Use total parameters for weight-memory planning; active parameters are primarily a compute-routing metric.

Primary sources and technical references

OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.

Continue the knowledge path