Shortlist snapshot
| Model | Architecture / size | License | OWM hardware class | Provider origin |
|---|---|---|---|---|
| Qwen3-32BQwen / Alibaba | 32B · dense | Apache 2.0 | 17–32B · high-memory workstation | China |
| Qwen3-30B-A3BQwen / Alibaba | 30B / 3B active · MoE | Apache 2.0 | 17–32B · high-memory workstation | China |
| Qwen3-Coder-30B-A3B-InstructQwen / Alibaba | 30B / ~3B active · coding MoE | Apache 2.0 | 17–32B · high-memory workstation | China |
| Gemma 3 27B ITGoogle DeepMind | 27B · multimodal | Gemma Terms | 17–32B · high-memory workstation | United States |
| Devstral Small 2 24B Instruct 2512Mistral AI | 24B · coding agent model | Apache 2.0 | 17–32B · high-memory workstation | EU provider |
| OLMo 3 32BAi2 | 32B · fully open research stack | Apache 2.0 | 17–32B · high-memory workstation | United States |
| gpt-oss-20bOpenAI | 20B · compact reasoning | Apache 2.0 | 17–32B · high-memory workstation | United States |
What should drive the decision?
At 48 GB, the practical question shifts from “can the weights load?” to “how much service headroom remains?” A configuration that uses nearly all memory for weights may have poor long-context or concurrent-serving behavior.
Dense and MoE models should not be compared using active parameters alone. Sparse activation may reduce compute per token, while total checkpoint size and expert routing still shape memory and serving complexity.
For production, measure time-to-first-token, generation throughput and peak memory at the target context and concurrency. Consumer workstation behavior can differ sharply from datacenter serving frameworks.
Models to evaluate
Qwen3-32B
Reasoning, multilingual, general-purpose
Dense 32B-class Apache model; benchmark quantized and higher-precision paths separately.
Qwen3-30B-A3B
Efficient reasoning and multilingual inference
Sparse 30B / 3B-active architecture; compute and memory trade-offs differ from dense 32B.
Qwen3-Coder-30B-A3B-Instruct
Coding agents, repository work, tool use and software engineering
Coding-agent MoE with 256K native context; long-context tests are essential.
Gemma 3 27B IT
Multimodal, multilingual, long context
27B multimodal candidate with Gemma-specific terms and a 128K context window.
Devstral Small 2 24B Instruct 2512
Software engineering, repository-scale coding and agentic development
24B coding-agent candidate from an EU provider under Apache 2.0.
OLMo 3 32B
Open research, reproducibility, instruction and reasoning variants
32B Apache-licensed research/reproducibility-oriented candidate.
gpt-oss-20b
Reasoning, local and edge-class inference
Sparse reasoning model with 21B total / 3.6B active according to OpenAI.
Hardware and runtime reality
Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.
Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.
License, provider origin and Europe
Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.