Home›Compare›Workload shortlist
Workload shortlist · source-first decision page

Local Multimodal Open-Weight Models

Compare locally deployable open-weight vision-language and multimodal models across compact, workstation and larger hardware classes.

Updated 1 Oct 20268 referenced modelsNo universal ranking
Direct answer

Local multimodal deployment adds more variables than text-only inference. Besides language-model weights, budget for vision or audio encoders, image resolution, preprocessing and larger prompt payloads. Start by fixing the modality you actually need—images, documents, audio or video—then choose the smallest checkpoint that passes the task-quality threshold on your hardware.

01Confirm the exact supported modalities; “multimodal” is not a single capability.
02Test document/image resolution and preprocessing, not just chat prompts.
03Measure peak memory with real image counts and context sizes.
04Check runtime support for the multimodal processor as well as the language model.

Models to evaluate

ModelArchitecture / sizeLicenseHardware classProvider
Gemma 3 4B ITGoogle DeepMind4B · multimodalGemma Terms≤8B · consumer/local🇺🇸 United States
Gemma 3n E2B ITGoogle DeepMindE2B-class · mobile multimodalGemma Terms≤8B · consumer/local🇺🇸 United States
Gemma 3n E4B ITGoogle DeepMindE4B · mobile multimodalGemma Terms≤8B · consumer/local🇺🇸 United States
Phi-4 Multimodal InstructMicrosofttext + vision + audioMIT≤8B · consumer/local🇺🇸 United States
Qwen2.5-VL 7B InstructQwen / Alibaba7B · vision-language instructApache 2.0≤8B · consumer/local🇨🇳 China
Ministral 3 3B Instruct 2512Mistral AI3B-class · vision-language edge modelApache 2.0≤8B · consumer/local🇫🇷 France
Ministral 3 8B Instruct 2512Mistral AI8B-class · vision-language edge modelApache 2.0≤8B · consumer/local🇫🇷 France
Ministral 3 14B Instruct 2512Mistral AI14B-class · vision-language edge modelApache 2.09–16B · workstation/local🇫🇷 France
Shortlist, not ranking: these candidates span different capability and hardware classes. Remove incompatible models first, then benchmark the remainder on the exact workload.

Decision criteria

Confirm the exact supported modalities; “multimodal” is not a single capability.

Test document/image resolution and preprocessing, not just chat prompts.

Measure peak memory with real image counts and context sizes.

Check runtime support for the multimodal processor as well as the language model.

Deployment reality

Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.

OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.

Primary model sources