Models to evaluate
| Model | Architecture / size | License | Hardware class | Provider |
|---|---|---|---|---|
| Gemma 3 4B ITGoogle DeepMind | 4B · multimodal | Gemma Terms | ≤8B · consumer/local | 🇺🇸 United States |
| Gemma 3n E2B ITGoogle DeepMind | E2B-class · mobile multimodal | Gemma Terms | ≤8B · consumer/local | 🇺🇸 United States |
| Gemma 3n E4B ITGoogle DeepMind | E4B · mobile multimodal | Gemma Terms | ≤8B · consumer/local | 🇺🇸 United States |
| Phi-4 Multimodal InstructMicrosoft | text + vision + audio | MIT | ≤8B · consumer/local | 🇺🇸 United States |
| Qwen2.5-VL 7B InstructQwen / Alibaba | 7B · vision-language instruct | Apache 2.0 | ≤8B · consumer/local | 🇨🇳 China |
| Ministral 3 3B Instruct 2512Mistral AI | 3B-class · vision-language edge model | Apache 2.0 | ≤8B · consumer/local | 🇫🇷 France |
| Ministral 3 8B Instruct 2512Mistral AI | 8B-class · vision-language edge model | Apache 2.0 | ≤8B · consumer/local | 🇫🇷 France |
| Ministral 3 14B Instruct 2512Mistral AI | 14B-class · vision-language edge model | Apache 2.0 | 9–16B · workstation/local | 🇫🇷 France |
Decision criteria
Confirm the exact supported modalities; “multimodal” is not a single capability.
Test document/image resolution and preprocessing, not just chat prompts.
Measure peak memory with real image counts and context sizes.
Check runtime support for the multimodal processor as well as the language model.
Deployment reality
Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.
OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.