Home›Compare›Hardware shortlist
Hardware shortlist · source-first decision page

Open-Weight Models for 8 GB VRAM

A conservative shortlist for 8 GB VRAM, emphasizing compact checkpoints, quantization, context headroom and realistic local inference.

Updated 1 Oct 20268 referenced modelsNo universal ranking
Direct answer

With 8 GB of VRAM, compact models are the practical starting point. Prefer sub-8B checkpoints and leave headroom for KV cache, runtime buffers and the operating environment. Larger models may load with aggressive quantization or partial CPU offload, but that is a different latency and throughput profile. Treat 8 GB as a hard deployment constraint and benchmark the exact quantized artifact.

01Prefer models whose quantized weights leave several gigabytes of working headroom.
02Measure the longest realistic context; KV cache can erase apparent free memory.
03CPU/GPU split can expand what loads, but may materially change latency.
04For multimodal models, account for vision/audio components in addition to language weights.

Models to evaluate

ModelArchitecture / sizeLicenseHardware classProvider
Qwen3 0.6BQwen / Alibaba0.6B · compact reasoning-capable text modelApache 2.0≤8B · consumer/local🇨🇳 China
Qwen3 1.7BQwen / Alibaba1.7B · compact reasoning-capable text modelApache 2.0≤8B · consumer/local🇨🇳 China
Qwen3 4BQwen / Alibaba4B · compact general-purpose modelApache 2.0≤8B · consumer/local🇨🇳 China
SmolLM3 3BHugging Face3B · hybrid reasoningApache 2.0≤8B · consumer/local🇺🇸 United States
Phi-4 Mini InstructMicrosoftcompact · multilingualMIT≤8B · consumer/local🇺🇸 United States
Gemma 3 1B ITGoogle DeepMind1B · compact text instructGemma Terms≤8B · consumer/local🇺🇸 United States
Ministral 3 3B Instruct 2512Mistral AI3B-class · vision-language edge modelApache 2.0≤8B · consumer/local🇫🇷 France
Gemma 3n E2B ITGoogle DeepMindE2B-class · mobile multimodalGemma Terms≤8B · consumer/local🇺🇸 United States
Shortlist, not ranking: these candidates span different capability and hardware classes. Remove incompatible models first, then benchmark the remainder on the exact workload.

Decision criteria

Prefer models whose quantized weights leave several gigabytes of working headroom.

Measure the longest realistic context; KV cache can erase apparent free memory.

CPU/GPU split can expand what loads, but may materially change latency.

For multimodal models, account for vision/audio components in addition to language weights.

Deployment reality

Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.

OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.

Primary model sources