Home›Compare›Hardware shortlist
Hardware shortlist · source-first decision page

Open-Weight Models for 16 GB VRAM

A practical 16 GB VRAM shortlist spanning 7B–14B-class open-weight models with quantization, context and runtime caveats.

Updated 1 Oct 20268 referenced modelsNo universal ranking
Direct answer

Sixteen gigabytes of VRAM is a strong local-inference tier for 7B–8B models and a practical quantized tier for many 12B–14B checkpoints. Do not size from parameter count alone: precision, KV cache, context length, multimodal components and runtime overhead determine whether a configuration has usable headroom.

018B-class models usually provide the most comfortable local headroom.
0212B–14B models often benefit from 4-bit or other memory-efficient formats.
03Long-context work can require more conservative model sizing.
04Evaluate quality per watt and latency, not only the largest model that fits.

Models to evaluate

ModelArchitecture / sizeLicenseHardware classProvider
Qwen3 8BQwen / Alibaba8B · local general-purpose modelApache 2.0≤8B · consumer/local🇨🇳 China
Qwen3 14BQwen / Alibaba14B · dense general-purpose modelApache 2.09–16B · workstation/local🇨🇳 China
Phi-4Microsoft14B · denseMIT9–16B · workstation/local🇺🇸 United States
Gemma 3 12B ITGoogle DeepMind12B · multimodalGemma Terms9–16B · workstation/local🇺🇸 United States
Mistral Nemo Instruct 2407Mistral AI / NVIDIA12B-class · multilingualApache 2.09–16B · workstation/local🇫🇷 France
Ministral 3 8B Instruct 2512Mistral AI8B-class · vision-language edge modelApache 2.0≤8B · consumer/local🇫🇷 France
Granite 4.2 8BIBM8B · 128K enterprise reasoningApache 2.0≤8B · consumer/local🇺🇸 United States
Qwen2.5-Coder 7B InstructQwen / Alibaba7B · code-specialized instructApache 2.0≤8B · consumer/local🇨🇳 China
Shortlist, not ranking: these candidates span different capability and hardware classes. Remove incompatible models first, then benchmark the remainder on the exact workload.

Decision criteria

8B-class models usually provide the most comfortable local headroom.

12B–14B models often benefit from 4-bit or other memory-efficient formats.

Long-context work can require more conservative model sizing.

Evaluate quality per watt and latency, not only the largest model that fits.

Deployment reality

Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.

OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.

Primary model sources