Home›Compare›Capability shortlist
Capability shortlist · source-first decision page

Long-Context Open-Weight Models

Compare long-context open-weight models while separating advertised maximum context from usable quality, KV-cache memory and serving latency.

Updated 1 Oct 20268 referenced modelsNo universal ranking
Direct answer

A large advertised context window is not automatically useful context. Long-context deployment must be evaluated for retrieval quality, attention behavior, KV-cache growth, time to first token and accuracy at the positions your workload actually uses. Treat publisher context limits as capability boundaries, then measure the practical operating range yourself.

01Benchmark realistic document lengths instead of only the published maximum.
02Track KV-cache memory and prefill latency as context grows.
03Test information placed at the beginning, middle and end of long prompts.
04For RAG, compare retrieval-assisted prompts with full-context stuffing.

Models to evaluate

ModelArchitecture / sizeLicenseHardware classProvider
Qwen3-Coder-30B-A3B-InstructQwen / Alibaba30B / ~3B active · coding MoEApache 2.017–32B · high-memory workstation🇨🇳 China
Qwen3-VL-30B-A3B-InstructQwen / Alibaba30B-class sparse vision-language modelApache 2.017–32B · high-memory workstation🇨🇳 China
Mistral Small 4 119B A6BMistral AI119B / 6.5B active · multimodal MoEApache 2.0Model-specific · large / specialized🇫🇷 France
Command A+ 05-2026Cohere / Cohere Labs218B / 25B active · multimodal MoEApache 2.0Model-specific · large / specialized🇨🇦 Canada
gpt-oss-20bOpenAI20B · compact reasoningApache 2.017–32B · high-memory workstation🇺🇸 United States
Gemma 3 27B ITGoogle DeepMind27B · multimodalGemma Terms17–32B · high-memory workstation🇺🇸 United States
Granite 4.2 8BIBM8B · 128K enterprise reasoningApache 2.0≤8B · consumer/local🇺🇸 United States
Mistral Nemo Instruct 2407Mistral AI / NVIDIA12B-class · multilingualApache 2.09–16B · workstation/local🇫🇷 France
Shortlist, not ranking: these candidates span different capability and hardware classes. Remove incompatible models first, then benchmark the remainder on the exact workload.

Decision criteria

Benchmark realistic document lengths instead of only the published maximum.

Track KV-cache memory and prefill latency as context grows.

Test information placed at the beginning, middle and end of long prompts.

For RAG, compare retrieval-assisted prompts with full-context stuffing.

Deployment reality

Validate the exact checkpoint, precision or quantization, runtime, context length and concurrency target. Weight memory alone does not capture KV cache, runtime workspaces, multimodal encoders or distributed-serving overhead.

OWM keeps license, provider origin and data residency separate. A provider-country label is provenance metadata; the deployer determines where inference and connected services run.

Primary model sources