Home›Compare›Hardware shortlist
Hardware shortlist · source-first decision page

Open-Weight Models for 24 GB VRAM

A practical 24 GB VRAM shortlist with quantization headroom, context/KV-cache caveats and source-backed model profiles.

Updated 1 Oct 20267 referenced modelsNo universal ranking
Direct answer

For a 24 GB GPU, start with 3B–14B models if you want comfortable headroom, and treat 20B–32B-class models as quantized configurations that need careful context and runtime testing. Weight memory is only the first budget item: KV cache, runtime overhead, vision components and batching can consume several additional gigabytes. OWM therefore treats 24 GB as a deployment constraint, not a model-quality ranking.

01Weight budget
02Quantization
03KV cache
04Real workload

Shortlist snapshot

ModelArchitecture / sizeLicenseOWM hardware classProvider origin
Qwen3 8BQwen / Alibaba8B · local general-purpose modelApache 2.0≤8B · consumer/localChina
Qwen3 14BQwen / Alibaba14B · dense general-purpose modelApache 2.09–16B · workstation/localChina
Phi-4Microsoft14B · denseMIT9–16B · workstation/localUnited States
Granite 4.2 8BIBM8B · 128K enterprise reasoningApache 2.0≤8B · consumer/localUnited States
SmolLM3 3BHugging Face3B · hybrid reasoningApache 2.0≤8B · consumer/localUnited States
Ministral 3 14B Instruct 2512Mistral AI14B-class · vision-language edge modelApache 2.09–16B · workstation/localEU provider
gpt-oss-20bOpenAI20B · compact reasoningApache 2.017–32B · high-memory workstationUnited States
Important: a shortlist is not a ranking. Eliminate incompatible models first, then benchmark the survivors on the exact workload.

What should drive the decision?

A useful first-pass estimate for weight-only memory is parameters × bytes per parameter. Four-bit storage is roughly 0.5 bytes per parameter before metadata and runtime overhead; eight-bit is roughly 1 byte. Real implementations can differ materially.

Models in the 8B and 14B class generally leave more room for long context and concurrent requests than a 20B–32B model squeezed into 24 GB through aggressive quantization. That headroom can matter more than a larger checkpoint for interactive use.

If a model only barely loads, test the longest realistic prompt and generation length before calling the configuration viable. Context growth and KV cache often expose memory limits that a short smoke test misses.

Models to evaluate

🇨🇳 Qwen / AlibabaApache 2.0

Qwen3 8B

Reasoning, multilingual chat, coding and local RAG

8B · local general-purpose model≤8B · consumer/local

8B-class Apache-licensed candidate with substantial memory headroom for local text work.

🇨🇳 Qwen / AlibabaApache 2.0

Qwen3 14B

Reasoning, multilingual work, coding and self-hosted assistants

14B · dense general-purpose model9–16B · workstation/local

14B dense candidate; quantization can preserve room for context and runtime overhead.

🇺🇸 MicrosoftMIT

Phi-4

Reasoning, mathematics and code

14B · dense9–16B · workstation/local

14B MIT-licensed dense model for reasoning, mathematics and code.

🇺🇸 IBMApache 2.0

Granite 4.2 8B

Enterprise assistants, reasoning, tool use and RAG

8B · 128K enterprise reasoning≤8B · consumer/local

8B enterprise-oriented Apache candidate for assistants, tools and RAG.

🇺🇸 Hugging FaceApache 2.0

SmolLM3 3B

Compact reasoning, agents and local inference

3B · hybrid reasoning≤8B · consumer/local

Compact 3B option where low memory and local iteration matter most.

🇺🇸 OpenAIApache 2.0

gpt-oss-20b

Reasoning, local and edge-class inference

20B · compact reasoning17–32B · high-memory workstation

Sparse reasoning candidate; validate the exact quantized/runtime configuration rather than relying on total parameters alone.

Hardware and runtime reality

Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.

Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.

License, provider origin and Europe

Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.

Evaluation checklist

Task qualityRepresentative prompts and hard cases.
ReliabilityTool errors, malformed output and regressions.
Latency + throughputMeasure target concurrency.
Peak memoryUse realistic context lengths.
License fitExact checkpoint and distribution model.
Data pathInference, retrieval, logs and backups.

Primary sources and related references