Devstral Small 1.1: what matters beyond the model card
Coding agents can see entire repositories, internal tickets and proprietary source code. Keeping the model inside controlled infrastructure can therefore matter more than it does for generic public chat. Devstral combines a 128K context window, function calling, open weights and Apache 2.0 licensing in a size class that does not inherently require a GPU cluster.
Devstral Small 1.1 is one of the most practical examples of open-weight AI moving into software-engineering infrastructure. Mistral explicitly positions it for agentic coding and says the 24B checkpoint is light enough for a single RTX 4090 or a Mac with 32 GB RAM. That makes private code-agent deployment materially more accessible.
Model facts
Runtime paths recorded by OWM: mistral-inference · Transformers · vLLM · LM Studio · llama.cpp. Support is version-sensitive and does not imply identical feature parity across runtimes.
Why this model matters
Coding agents can see entire repositories, internal tickets and proprietary source code. Keeping the model inside controlled infrastructure can therefore matter more than it does for generic public chat. Devstral combines a 128K context window, function calling, open weights and Apache 2.0 licensing in a size class that does not inherently require a GPU cluster.
OWM evaluates a checkpoint as infrastructure: exact weights, license, runtime portability, memory reality, evidence quality and provider exit all matter alongside capability.
Hardware reality
Mistral states that Devstral Small 1.1 can run on a single RTX 4090 or a Mac with 32 GB RAM. This is publisher guidance and depends on the actual runtime and precision, but it is unusually concrete hardware guidance for a 24B coding model.
OWM separates publisher guidance, engineering estimates and measured runtime evidence. Memory arithmetic alone does not include every KV-cache, runtime, multimodal or concurrency cost.
License reality
Apache 2.0 allows commercial and non-commercial use, modification and redistribution under its terms. OWM views this as a strong fit for internal developer platforms and productized coding tools.
This is an informational deployment summary, not legal advice. Always review the exact current license and policies before production use.
OWM Sovereignty Lens
OWM does not collapse sovereignty into one score. Technical portability and legal freedom can differ substantially.
Runtime evidence
The publisher documents Mistral inference, vLLM and local deployment paths. OWM has not yet published a reproduced coding-agent throughput benchmark for the exact 2507 checkpoint.
“OWM runtime tested” remains reserved for configurations physically reproduced by the project with exact hardware, runtime version, workload and date.
Change history
First OWM verification snapshot created. From this date forward, material changes to the model card, license, checkpoints, runtime support and deployment facts can be appended without reconstructing unobserved history.
Open the global OWM Change History →
Where Devstral Small 1.1 fits — and where it does not
Where it fits
- Private coding agents over proprietary repositories.
- Local developer workstations with high-memory GPUs or Apple Silicon.
- OpenHands-style agentic software-engineering workflows.
- Organizations that want Apache 2.0 rather than a closed coding-service dependency.
Where it does not fit
- Vision-based coding tasks; the vision encoder was removed from the underlying Mistral Small 3.1 lineage.
- Small-memory systems below the practical 24B deployment range.
- Teams expecting coding-agent reliability without sandboxing, testing and human review.
Open-weight significance
The strategic value is not simply that weights can be downloaded. The relevant question is what the operator can control: infrastructure, data location, runtime, adaptation and provider exit — all bounded by the license and practical hardware requirements.
Frequently asked questions
Can Devstral Small 1.1 run on an RTX 4090?
Mistral says the 24B model is light enough for a single RTX 4090.
What is the context length?
The model card documents 128K context.
Is Devstral Small 1.1 commercially usable?
It is released under Apache 2.0.
Is Devstral multimodal?
This Devstral Small checkpoint is text-only; Mistral notes that the vision encoder was removed before coding fine-tuning.
Primary sources and OWM data
Last verified by Open Weight Models: 2026-09-27. Facts can change as model repositories, licenses and runtime support evolve.