DeepSeek-R1: what matters beyond the model card
The model card lists 671B total parameters, 37B active parameters and a 128K context window. That scale shifts the sovereignty question away from laptop ownership and toward data-center choice: can the organization control where the model runs, which serving stack it uses and whether it can change providers without changing the model itself?
DeepSeek-R1 is strategically important because it proved that a very large reasoning model could be distributed as weights under a permissive license. OWM does not treat it as a “local model” in the casual sense: the full 671B checkpoint is infrastructure-heavy. Its significance is that organizations can still operate that infrastructure themselves or choose among independent providers.
Model facts
Runtime paths recorded by OWM: vLLM · SGLang · Transformers · Docker Model Runner. Runtime support is version-sensitive; a named runtime should not be read as a guarantee that every quantization, context size or feature works identically.
Why this model matters
The model card lists 671B total parameters, 37B active parameters and a 128K context window. That scale shifts the sovereignty question away from laptop ownership and toward data-center choice: can the organization control where the model runs, which serving stack it uses and whether it can change providers without changing the model itself?
OWM evaluates a model as infrastructure, not only as a benchmark entry. That means the exact checkpoint, license, runtime ecosystem, memory footprint, ability to move between providers and the quality of the evidence all matter alongside model capability.
Hardware reality
The full model is enormous. Even an idealized 4-bit raw-weight estimate is hundreds of gigabytes before runtime overhead. DeepSeek-R1 is therefore fundamentally a distributed-serving model unless a heavily transformed derivative is used. Distilled R1 variants are separate checkpoints and should not be confused with the full model.
OWM deliberately separates raw-weight arithmetic, publisher guidance and measured runtime evidence. A model that theoretically fits into a memory budget can still fail in practice because of KV cache, runtime buffers, vision components, tensor-parallel overhead or concurrent requests.
License reality
DeepSeek releases the R1 repository under MIT and states support for commercial use. OWM still recommends checking the exact checkpoint and any derivative’s upstream licensing because distilled models can introduce separate license considerations.
This is an informational deployment summary, not legal advice. Production users should review the exact current license, usage policy, derivative-model terms and applicable law before shipping a product.
OWM Sovereignty Lens
OWM does not assign a single sovereignty score. We describe the layers separately because a model can be highly portable technically while remaining conditional legally — or permissively licensed while requiring infrastructure that limits practical choice.
Runtime evidence
SemiAnalysis InferenceX has published an independent DeepSeek-R1 H100 measurement using Dynamo SGLang. OWM records it as external laboratory evidence and does not present it as an OWM reproduction.
“Third-party measured” means the result was measured outside Open Weight Models and is shown with provenance. “OWM runtime tested” is reserved for configurations that OWM physically reproduces with an exact checkpoint, runtime version, hardware configuration, workload and date.
Change history
First OWM verification snapshot created. From this date forward, material changes to the model card, license, checkpoints, runtime support and deployment facts can be appended without reconstructing unobserved history.
Open the global OWM Change History →
Where DeepSeek-R1 fits — and where it does not
Where it fits
- Multi-GPU or multi-node reasoning infrastructure.
- Research labs and enterprises that need control of a large reasoning checkpoint.
- Provider-hosted deployments where the customer still wants model portability between infrastructure vendors.
- Teams comparing the economics of sparse large models against closed reasoning APIs.
Where it does not fit
- Typical single-GPU consumer hardware for the full checkpoint.
- Teams that interpret “37B active” as meaning the full 671B checkpoint has 37B-model storage requirements.
- Small deployments where operational complexity outweighs the benefit of owning the model layer.
Open-weight significance
The strategic value of this model is not simply that its weights can be downloaded. The important question is what those weights let an operator control: infrastructure, data location, runtime, adaptation and the ability to exit a provider relationship without discarding the model layer. Those freedoms remain bounded by the model’s license and the practical hardware required to run it.
Frequently asked questions
How large is DeepSeek-R1?
DeepSeek lists 671B total parameters and 37B activated parameters.
What is the context window of DeepSeek-R1?
The official model card lists a 128K context length.
Can DeepSeek-R1 run on one GPU?
The full checkpoint is generally a distributed-inference model. Distilled variants are much smaller but are different checkpoints.
Is DeepSeek-R1 commercially usable?
The main project is released under MIT and DeepSeek states commercial use is supported; always verify exact derivative terms.
Primary sources and OWM data
Last verified by Open Weight Models: 2026-09-27. Facts can change as model repositories, licenses and runtime support evolve.