Home›Compare›Workload shortlist
Workload shortlist · source-first decision page

Open-Weight Reasoning Models

Compare open-weight reasoning models across local, workstation and large-server classes without collapsing them into a universal leaderboard.

Updated 1 Oct 20268 referenced modelsNo universal ranking
Direct answer

Reasoning models differ in much more than benchmark score: some expose explicit reasoning controls, some are dense, some are MoE, and their memory and latency profiles span from local workstations to distributed servers. Define the reasoning workload first—mathematics, code, planning, tools or domain analysis—and compare quality together with token usage, latency and failure modes.

01Measure reasoning quality and total generated tokens together.
02Include tool-use and structured-output reliability where relevant.
03Compare fast/non-thinking modes separately when a model supports them.
04Large MoE checkpoints need infrastructure evaluation independent of active-parameter counts.

Models to evaluate

ModelArchitecture / sizeLicenseHardware classProvider
Phi-4 Reasoning PlusMicrosoft14B · dense reasoning modelMIT9–16B · workstation/local🇺🇸 United States
Qwen3 14BQwen / Alibaba14B · dense general-purpose modelApache 2.09–16B · workstation/local🇨🇳 China
Qwen3-32BQwen / Alibaba32B · denseApache 2.017–32B · high-memory workstation🇨🇳 China
gpt-oss-20bOpenAI20B · compact reasoningApache 2.017–32B · high-memory workstation🇺🇸 United States
Magistral Small 2506Mistral AI24B-class · reasoningApache 2.017–32B · high-memory workstation🇫🇷 France
DeepSeek-R1DeepSeekreasoning model · MoEMITModel-specific · large / specialized🇨🇳 China
Qwen3-235B-A22BQwen / Alibaba235B / 22B active · MoEApache 2.0Model-specific · large / specialized🇨🇳 China
GLM-4.5Z.ai355B / 32B active · MoEMITModel-specific · large / specialized🇨🇳 China
Shortlist, not ranking: these candidates span different capability and hardware classes. Remove incompatible models first, then benchmark the remainder on the exact workload.

Decision criteria

Measure reasoning quality and total generated tokens together.

Include tool-use and structured-output reliability where relevant.

Compare fast/non-thinking modes separately when a model supports them.

Large MoE checkpoints need infrastructure evaluation independent of active-parameter counts.

Deployment reality

Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.

OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.

Primary model sources