Shortlist snapshot
| Model | Architecture / size | License | OWM hardware class | Provider origin |
|---|---|---|---|---|
| Qwen3-Coder-30B-A3B-InstructQwen / Alibaba | 30B / ~3B active · coding MoE | Apache 2.0 | 17–32B · high-memory workstation | China |
| Devstral Small 2 24B Instruct 2512Mistral AI | 24B · coding agent model | Apache 2.0 | 17–32B · high-memory workstation | EU provider |
| Qwen2.5-Coder-32B-InstructQwen / Alibaba | 32B · code | Apache 2.0 | 17–32B · high-memory workstation | China |
| Qwen2.5-Coder 7B InstructQwen / Alibaba | 7B · code-specialized instruct | Apache 2.0 | ≤8B · consumer/local | China |
| Phi-4Microsoft | 14B · dense | MIT | 9–16B · workstation/local | United States |
| DeepSeek-R1DeepSeek | reasoning model · MoE | MIT | Model-specific · large / specialized | China |
| Mistral Small 4 119B A6BMistral AI | 119B / 6.5B active · multimodal MoE | Apache 2.0 | Model-specific · large / specialized | EU provider |
What should drive the decision?
Coding workloads are unusually sensitive to the surrounding agent. File search, shell execution, test feedback, tool-call format and context construction can change results as much as the base checkpoint.
For repository work, evaluate complete issue-resolution loops: identify files, make a minimal patch, run tests, repair failures and stop cleanly. Record malformed tool calls, unnecessary edits and regressions in addition to task success.
Long context can help with codebases, but indiscriminate context stuffing can be slower and noisier than retrieval. Compare full-context, retrieval-assisted and agentic search approaches on the same repository set.
Models to evaluate
Qwen3-Coder-30B-A3B-Instruct
Coding agents, repository work, tool use and software engineering
Purpose-built agentic coding MoE with 256K native context.
Devstral Small 2 24B Instruct 2512
Software engineering, repository-scale coding and agentic development
Mistral software-engineering agent model with 256K context.
Qwen2.5-Coder-32B-Instruct
Code generation, repair and reasoning
Dense 32B code-specialized option in the Qwen2.5 generation.
Qwen2.5-Coder 7B Instruct
Local code generation, repair, repository assistance and coding RAG
Compact local code model for lower-memory deployments.
Phi-4
Reasoning, mathematics and code
General reasoning model with mathematics and code focus rather than a dedicated coding-agent identity.
DeepSeek-R1
Reasoning, mathematics and coding
Very large reasoning system; infrastructure requirements differ radically from workstation models.
Mistral Small 4 119B A6B
Instruction, reasoning, coding, agents and vision
Large multimodal MoE whose broader agent/coding capabilities may fit server deployments.
Hardware and runtime reality
Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.
Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.
License, provider origin and Europe
Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.