Home›Compare›Coding comparison
Coding comparison · source-first decision page

Qwen3-Coder-30B-A3B vs Devstral Small 2

Compare two self-hostable coding-agent models by architecture, context, tool use, hardware class, Apache 2.0 licensing and repository-scale workflows.

Updated 1 Oct 2026No universal rankingPrimary model sources
Direct answer

Both models are credible Apache-2.0 coding-agent candidates in the workstation/server class, but they emphasize different architectures. Qwen3-Coder-30B-A3B is a 30.5B-total MoE model with roughly 3.3B active parameters and native 256K context; Devstral Small 2 is a 24B model positioned by Mistral for software-engineering agents with 256K context. Choose by repository tasks, edit reliability, tool protocol, runtime maturity and memory—not benchmark headlines alone.

01Repo tasks
02Tool protocol
03Memory + context
04Patch evaluation

Comparison snapshot

ModelArchitecture / sizeLicenseOWM hardware classProvider origin
Qwen3-Coder-30B-A3B-InstructQwen / Alibaba30B / ~3B active · coding MoEApache 2.017–32B · high-memory workstationChina
Devstral Small 2 24B Instruct 2512Mistral AI24B · coding agent modelApache 2.017–32B · high-memory workstationEU provider
Important: hardware classes are planning guidance, not guaranteed minimum requirements. Precision, quantization, context, batching and runtime overhead change real memory use.

What should drive the decision?

Qwen’s official material emphasizes agentic coding, repository-scale understanding, function calling and a 256K native context window for the 30B-A3B checkpoint. Mistral positions Devstral Small 2 specifically around exploring codebases, editing multiple files and software-engineering agents.

The MoE-versus-dense distinction changes inference economics. Qwen3-Coder activates a small subset of its total parameters per token, but the full checkpoint still has to be available. Devstral’s 24B-class footprint is easier to reason about in conventional dense-model memory estimates, subject to its published checkpoint precision.

Do not decide this comparison from SWE-bench or a single coding benchmark. Build a repository-level evaluation with issue reproduction, multi-file edits, test execution, tool-call correctness, patch size and regression rate.

Models in this comparison

Hardware and runtime reality

Start from the exact checkpoint and weight format. Estimate weight memory, then add KV cache, runtime workspaces, multimodal components where relevant, and concurrency headroom. A configuration that merely loads is not yet a production configuration.

For local deployments, test the exact quantized artifact and runtime you intend to ship. For server deployments, record parallelism, time to first token, steady-state throughput and peak memory at target context.

License, provider origin and Europe

The checkpoint license does not automatically describe every adapter, quantization, tokenizer, dataset or application component. Provider origin is also separate from data residency. Map inference, retrieval, embeddings, logs, observability, backups and subprocessors before making a residency claim.

Evaluation checklist

Task qualityRepresentative real prompts and edge cases.
ReliabilityMalformed output, tool errors and regressions.
Latency + throughputTTFT and throughput at target concurrency.
Peak memoryLongest realistic context and generation.
License fitExact checkpoint terms and distribution model.
Data pathInference, RAG, logs, backups and external services.

Primary sources and related references