Shortlist snapshot
| Model | Architecture / size | License | OWM hardware class | Provider origin |
|---|---|---|---|---|
| Ministral 3 3B Instruct 2512Mistral AI | 3B-class · vision-language edge model | Apache 2.0 | ≤8B · consumer/local | EU provider |
| Ministral 3 8B Instruct 2512Mistral AI | 8B-class · vision-language edge model | Apache 2.0 | ≤8B · consumer/local | EU provider |
| Ministral 3 14B Instruct 2512Mistral AI | 14B-class · vision-language edge model | Apache 2.0 | 9–16B · workstation/local | EU provider |
| Devstral Small 2 24B Instruct 2512Mistral AI | 24B · coding agent model | Apache 2.0 | 17–32B · high-memory workstation | EU provider |
| Magistral Small 2506Mistral AI | 24B-class · reasoning | Apache 2.0 | 17–32B · high-memory workstation | EU provider |
| Mistral Small 3.1 24B InstructMistral AI | 24B · multimodal | Apache 2.0 | 17–32B · high-memory workstation | EU provider |
| Mistral Small 4 119B A6BMistral AI | 119B / 6.5B active · multimodal MoE | Apache 2.0 | Model-specific · large / specialized | EU provider |
What should drive the decision?
Provider country answers who publishes the model. Data residency answers where prompts, retrieved documents, embeddings, logs, telemetry and backups are actually processed or stored. Those are different questions.
Self-hosting downloaded weights on EU/EEA infrastructure can give an organization more control over the inference path, but it does not by itself establish GDPR compliance. Legal basis, processor relationships, security, retention, user rights and international transfers still require analysis.
The EU AI Act likewise depends on role and use context. OWM does not assign blanket compliance badges to a model because the same checkpoint can be used in very different systems.
Models to evaluate
Ministral 3 3B Instruct 2512
Edge chat, vision and compact local assistants
Compact Mistral AI vision-language option for edge/local deployments.
Ministral 3 8B Instruct 2512
Balanced local chat, vision and instruction following
8B-class balanced local vision-language model.
Ministral 3 14B Instruct 2512
Higher-capability local chat, vision and instruction following
14B-class higher-capability local vision-language model.
Devstral Small 2 24B Instruct 2512
Software engineering, repository-scale coding and agentic development
24B software-engineering agent model.
Magistral Small 2506
Reasoning and step-by-step problem solving
24B-class reasoning model.
Mistral Small 3.1 24B Instruct
General-purpose multimodal inference
24B multimodal general-purpose model.
Mistral Small 4 119B A6B
Instruction, reasoning, coding, agents and vision
Large multimodal MoE for server-class deployments.
Hardware and runtime reality
Weight-only estimates are a starting point. Add KV cache, runtime workspaces, multimodal components, batching and concurrency headroom. For local inference, validate the exact quantized artifact. For server inference, measure time to first token, throughput and peak memory at target concurrency.
Long context can make an otherwise comfortable model exceed the practical memory budget. Test the longest realistic prompt and generation, not only a short loading test.
License, provider origin and Europe
Review the exact checkpoint license and any separate usage terms. Provider origin is supply-chain metadata, not an inference-location claim. If EU/EEA residency matters, map inference, RAG, embeddings, logs, telemetry, backups and subprocessors.