Comparison snapshot
| Model | Architecture / size | License | OWM hardware class | Provider origin |
|---|---|---|---|---|
| Qwen3-Coder-30B-A3B-InstructQwen / Alibaba | 30B / ~3B active · coding MoE | Apache 2.0 | 17–32B · high-memory workstation | China |
| Devstral Small 2 24B Instruct 2512Mistral AI | 24B · coding agent model | Apache 2.0 | 17–32B · high-memory workstation | EU provider |
What should drive the decision?
Qwen’s official material emphasizes agentic coding, repository-scale understanding, function calling and a 256K native context window for the 30B-A3B checkpoint. Mistral positions Devstral Small 2 specifically around exploring codebases, editing multiple files and software-engineering agents.
The MoE-versus-dense distinction changes inference economics. Qwen3-Coder activates a small subset of its total parameters per token, but the full checkpoint still has to be available. Devstral’s 24B-class footprint is easier to reason about in conventional dense-model memory estimates, subject to its published checkpoint precision.
Do not decide this comparison from SWE-bench or a single coding benchmark. Build a repository-level evaluation with issue reproduction, multi-file edits, test execution, tool-call correctness, patch size and regression rate.
Models in this comparison
Qwen3-Coder-30B-A3B-Instruct
Coding agents, repository work, tool use and software engineering
Devstral Small 2 24B Instruct 2512
Software engineering, repository-scale coding and agentic development
Hardware and runtime reality
Start from the exact checkpoint and weight format. Estimate weight memory, then add KV cache, runtime workspaces, multimodal components where relevant, and concurrency headroom. A configuration that merely loads is not yet a production configuration.
For local deployments, test the exact quantized artifact and runtime you intend to ship. For server deployments, record parallelism, time to first token, steady-state throughput and peak memory at target context.
License, provider origin and Europe
The checkpoint license does not automatically describe every adapter, quantization, tokenizer, dataset or application component. Provider origin is also separate from data residency. Map inference, retrieval, embeddings, logs, observability, backups and subprocessors before making a residency claim.