Home›Compare›Workload shortlist
Workload shortlist · source-first decision page

Open-Weight Models for RAG

A source-first RAG shortlist focused on context handling, instruction following, structured output, deployment footprint and private data paths.

Updated 1 Oct 20268 referenced modelsNo universal ranking
Direct answer

For RAG, retrieval quality and prompt construction often matter more than choosing the largest generator. Shortlist models by instruction following, answer grounding, context behavior, structured output and deployment footprint, then evaluate them against the same retriever and document set. Keep embeddings, vector storage, inference and logging inside the intended data boundary.

01Use the same retrieval corpus and retriever when comparing generators.
02Score citation/grounding behavior separately from fluency.
03Test abstention when the answer is not present in retrieved evidence.
04Map embeddings, vector stores, logs and backups as part of the data path.

Models to evaluate

ModelArchitecture / sizeLicenseHardware classProvider
Granite 4.2 8BIBM8B · 128K enterprise reasoningApache 2.0≤8B · consumer/local🇺🇸 United States
Qwen3 8BQwen / Alibaba8B · local general-purpose modelApache 2.0≤8B · consumer/local🇨🇳 China
Qwen3 14BQwen / Alibaba14B · dense general-purpose modelApache 2.09–16B · workstation/local🇨🇳 China
Qwen3-32BQwen / Alibaba32B · denseApache 2.017–32B · high-memory workstation🇨🇳 China
Mistral Nemo Instruct 2407Mistral AI / NVIDIA12B-class · multilingualApache 2.09–16B · workstation/local🇫🇷 France
Ministral 3 8B Instruct 2512Mistral AI8B-class · vision-language edge modelApache 2.0≤8B · consumer/local🇫🇷 France
gpt-oss-20bOpenAI20B · compact reasoningApache 2.017–32B · high-memory workstation🇺🇸 United States
OLMo 3 32BAi232B · fully open research stackApache 2.017–32B · high-memory workstation🇺🇸 United States
Shortlist, not ranking: these candidates span different capability and hardware classes. Remove incompatible models first, then benchmark the remainder on the exact workload.

Decision criteria

Use the same retrieval corpus and retriever when comparing generators.

Score citation/grounding behavior separately from fluency.

Test abstention when the answer is not present in retrieved evidence.

Map embeddings, vector stores, logs and backups as part of the data path.

Deployment reality

Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.

OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.

Primary model sources