Home›Compare›Hardware shortlist
Hardware shortlist · source-first decision page

Open-Weight Models for 48 GB VRAM

A 48 GB VRAM shortlist for 20B–32B-class open-weight models, with precision, context, modality and runtime considerations.

Updated 1 Oct 20267 referenced modelsNo universal ranking
Direct answer

A 48 GB GPU opens a useful 20B–32B model tier, especially with 8-bit or 4-bit weight formats, but it still does not guarantee every 32B-class deployment at every context length. Dense 27B–32B checkpoints, sparse 30B-class models and coding-focused models can all be candidates. Reserve memory for KV cache, runtime workspaces, multimodal towers and concurrency before choosing precision.

01Model class
02Precision
03Context headroom
04Concurrency

Shortlist snapshot

ModelArchitecture / sizeLicenseOWM hardware classProvider origin
Qwen3-32BQwen / Alibaba32B · denseApache 2.017–32B · high-memory workstationChina
Qwen3-30B-A3BQwen / Alibaba30B / 3B active · MoEApache 2.017–32B · high-memory workstationChina
Qwen3-Coder-30B-A3B-InstructQwen / Alibaba30B / ~3B active · coding MoEApache 2.017–32B · high-memory workstationChina
Gemma 3 27B ITGoogle DeepMind27B · multimodalGemma Terms17–32B · high-memory workstationUnited States
Devstral Small 2 24B Instruct 2512Mistral AI24B · coding agent modelApache 2.017–32B · high-memory workstationEU provider
OLMo 3 32BAi232B · fully open research stackApache 2.017–32B · high-memory workstationUnited States
gpt-oss-20bOpenAI20B · compact reasoningApache 2.017–32B · high-memory workstationUnited States
Important: a shortlist is not a ranking. Eliminate incompatible models first, then benchmark the survivors on the exact workload.

What should drive the decision?

At 48 GB, the practical question shifts from “can the weights load?” to “how much service headroom remains?” A configuration that uses nearly all memory for weights may have poor long-context or concurrent-serving behavior.

Dense and MoE models should not be compared using active parameters alone. Sparse activation may reduce compute per token, while total checkpoint size and expert routing still shape memory and serving complexity.

For production, measure time-to-first-token, generation throughput and peak memory at the target context and concurrency. Consumer workstation behavior can differ sharply from datacenter serving frameworks.

Models to evaluate

🇨🇳 Qwen / AlibabaApache 2.0

Qwen3-32B

Reasoning, multilingual, general-purpose

32B · dense17–32B · high-memory workstation

Dense 32B-class Apache model; benchmark quantized and higher-precision paths separately.

🇨🇳 Qwen / AlibabaApache 2.0

Qwen3-30B-A3B

Efficient reasoning and multilingual inference

30B / 3B active · MoE17–32B · high-memory workstation

Sparse 30B / 3B-active architecture; compute and memory trade-offs differ from dense 32B.

🇺🇸 Google DeepMindGemma Terms

Gemma 3 27B IT

Multimodal, multilingual, long context

27B · multimodal17–32B · high-memory workstation

27B multimodal candidate with Gemma-specific terms and a 128K context window.

🇺🇸 Ai2Apache 2.0

OLMo 3 32B

Open research, reproducibility, instruction and reasoning variants

32B · fully open research stack17–32B · high-memory workstation

32B Apache-licensed research/reproducibility-oriented candidate.

🇺🇸 OpenAIApache 2.0

gpt-oss-20b

Reasoning, local and edge-class inference

20B · compact reasoning17–32B · high-memory workstation

Sparse reasoning model with 21B total / 3.6B active according to OpenAI.

Hardware and runtime reality

Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.

Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.

License, provider origin and Europe

Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.

Evaluation checklist

Task qualityRepresentative prompts and hard cases.
ReliabilityTool errors, malformed output and regressions.
Latency + throughputMeasure target concurrency.
Peak memoryUse realistic context lengths.
License fitExact checkpoint and distribution model.
Data pathInference, retrieval, logs and backups.

Primary sources and related references