Home›Compare›Workload shortlist
Workload shortlist · source-first decision page

Open-Weight Coding Models — A Deployment Shortlist

A source-first shortlist of coding and software-engineering open-weight models across local, workstation and server deployments.

Updated 1 Oct 20267 referenced modelsNo universal ranking
Direct answer

Do not choose a coding model from a single leaderboard. Separate inline code generation from repository-scale agents, tool use, long-context codebase reading, patch generation and reasoning. The OWM registry contains compact code models, workstation-class coding checkpoints and large reasoning/MoE systems; the right shortlist depends first on repository size, agent loop and hardware.

01Coding task
02Agent tools
03Repository context
04Tests + regressions

Shortlist snapshot

ModelArchitecture / sizeLicenseOWM hardware classProvider origin
Qwen3-Coder-30B-A3B-InstructQwen / Alibaba30B / ~3B active · coding MoEApache 2.017–32B · high-memory workstationChina
Devstral Small 2 24B Instruct 2512Mistral AI24B · coding agent modelApache 2.017–32B · high-memory workstationEU provider
Qwen2.5-Coder-32B-InstructQwen / Alibaba32B · codeApache 2.017–32B · high-memory workstationChina
Qwen2.5-Coder 7B InstructQwen / Alibaba7B · code-specialized instructApache 2.0≤8B · consumer/localChina
Phi-4Microsoft14B · denseMIT9–16B · workstation/localUnited States
DeepSeek-R1DeepSeekreasoning model · MoEMITModel-specific · large / specializedChina
Mistral Small 4 119B A6BMistral AI119B / 6.5B active · multimodal MoEApache 2.0Model-specific · large / specializedEU provider
Important: a shortlist is not a ranking. Eliminate incompatible models first, then benchmark the survivors on the exact workload.

What should drive the decision?

Coding workloads are unusually sensitive to the surrounding agent. File search, shell execution, test feedback, tool-call format and context construction can change results as much as the base checkpoint.

For repository work, evaluate complete issue-resolution loops: identify files, make a minimal patch, run tests, repair failures and stop cleanly. Record malformed tool calls, unnecessary edits and regressions in addition to task success.

Long context can help with codebases, but indiscriminate context stuffing can be slower and noisier than retrieval. Compare full-context, retrieval-assisted and agentic search approaches on the same repository set.

Models to evaluate

🇺🇸 MicrosoftMIT

Phi-4

Reasoning, mathematics and code

14B · dense9–16B · workstation/local

General reasoning model with mathematics and code focus rather than a dedicated coding-agent identity.

🇨🇳 DeepSeekMIT

DeepSeek-R1

Reasoning, mathematics and coding

reasoning model · MoEModel-specific · large / specialized

Very large reasoning system; infrastructure requirements differ radically from workstation models.

🇫🇷 Mistral AIApache 2.0

Mistral Small 4 119B A6B

Instruction, reasoning, coding, agents and vision

119B / 6.5B active · multimodal MoEModel-specific · large / specialized

Large multimodal MoE whose broader agent/coding capabilities may fit server deployments.

Hardware and runtime reality

Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.

Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.

License, provider origin and Europe

Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.

Evaluation checklist

Task qualityRepresentative prompts and hard cases.
ReliabilityTool errors, malformed output and regressions.
Latency + throughputMeasure target concurrency.
Peak memoryUse realistic context lengths.
License fitExact checkpoint and distribution model.
Data pathInference, retrieval, logs and backups.

Primary sources and related references