Home›Compare›Large reasoning comparison
Large reasoning comparison · source-first decision page

DeepSeek-R1 vs Qwen3-235B-A22B

Compare two large MoE reasoning models by total and active parameters, context, licensing, infrastructure and evaluation strategy.

Updated 1 Oct 2026No universal rankingPrimary model sources
Direct answer

DeepSeek-R1 and Qwen3-235B-A22B are large MoE reasoning systems that require serious infrastructure planning. DeepSeek’s official model card lists 671B total / 37B active parameters and 128K context; Qwen lists 235B total / 22B active with 32K native context and 131K via YaRN. DeepSeek-R1 uses MIT terms and Qwen3 uses Apache 2.0. The smaller total checkpoint of Qwen3 materially changes storage and serving design, but task quality must be measured separately.

01Quality gate
02MoE topology
03Cluster design
04Governance

Comparison snapshot

ModelArchitecture / sizeLicenseOWM hardware classProvider origin
DeepSeek-R1DeepSeekreasoning model · MoEMITModel-specific · large / specializedChina
Qwen3-235B-A22BQwen / Alibaba235B / 22B active · MoEApache 2.0Model-specific · large / specializedChina
Important: hardware classes are planning guidance, not guaranteed minimum requirements. Precision, quantization, context, batching and runtime overhead change real memory use.

What should drive the decision?

Active parameters are useful for understanding sparse compute, but they are not a VRAM number. The full expert set, quantization, tensor/expert parallelism, KV cache and serving framework determine the deployable configuration.

Both publishers provide reasoning-oriented model families, but their recommended prompting, serving and tool-use paths differ. That makes end-to-end evaluation more important than comparing raw architecture fields.

For organizations evaluating EU-hosted inference, both can in principle be deployed on infrastructure you control when the weights and required software are available. Provider origin does not by itself determine data residency; logs, RAG, embeddings, monitoring and backups remain part of the data path.

Models in this comparison

Hardware and runtime reality

Start from the exact checkpoint and weight format. Estimate weight memory, then add KV cache, runtime workspaces, multimodal components where relevant, and concurrency headroom. A configuration that merely loads is not yet a production configuration.

For local deployments, test the exact quantized artifact and runtime you intend to ship. For server deployments, record parallelism, time to first token, steady-state throughput and peak memory at target context.

License, provider origin and Europe

The checkpoint license does not automatically describe every adapter, quantization, tokenizer, dataset or application component. Provider origin is also separate from data residency. Map inference, retrieval, embeddings, logs, observability, backups and subprocessors before making a residency claim.

Evaluation checklist

Task qualityRepresentative real prompts and edge cases.
ReliabilityMalformed output, tool errors and regressions.
Latency + throughputTTFT and throughput at target concurrency.
Peak memoryLongest realistic context and generation.
License fitExact checkpoint terms and distribution model.
Data pathInference, RAG, logs, backups and external services.

Primary sources and related references