Direct weights
Maximum control. The operator owns loading, sharding, batching, scaling, safeguards, observability and lifecycle. Precision and model format are explicit choices.
Compare how 32 open-weight models can be run — not just whether weights exist.
Deployment is recorded as a combination of checkpoint, runtime/provider, precision and hardware. The Explorer separates local inference, self-hosted serving and managed APIs while preserving provider-specific context and feature differences.
The OpenWeightModels Deployment Explorer compares the practical ways downloadable AI models can be served: directly from weights, through inference runtimes, in local applications, or through managed cloud platforms.
The central idea is that deployment is not a binary “self-hosting: yes” field. A 3B quantized model running in llama.cpp on a laptop and a 671B sparse model spread across datacenter GPUs are both self-hosted, but they are radically different operational systems.
Likewise, a base model's headline context window does not guarantee that every provider exposes the same context, output limit, tool support or region. Llama 4 Maverick demonstrates the problem clearly: its base model is documented at 1M tokens, while provider-specific implementations can expose different limits by cloud or GPU class.
The Explorer therefore treats provider and runtime deployments as child records of a model, not as universal capabilities inherited by every implementation.
Choosing a model is also choosing how much infrastructure responsibility the operator wants to own.
Maximum control. The operator owns loading, sharding, batching, scaling, safeguards, observability and lifecycle. Precision and model format are explicit choices.
vLLM and SGLang provide production-oriented servers for supported models; Transformers covers direct model execution; llama.cpp and Ollama emphasize accessible local or portable workflows.
Products such as NVIDIA NIM package validated model/runtime/hardware combinations, which can make the supported operating envelope more specific than a generic self-host setup.
A provider owns much of the serving infrastructure but can impose its own region, context, endpoint, feature, price and lifecycle envelope.
Hardware class is an editorial navigation aid based primarily on total parameter scale and documented paths. Open each Passport for exact evidence.
| Model | Parameters | Context | Editorial hardware class | Documented runtimes | Managed/cloud | Local/self-host |
|---|---|---|---|---|---|---|
| DeepSeek-R1 DeepSeek | See passport | 128K tokens | Model-specific | vLLM, SGLang | DeepSeek API | Self-hosted |
| DeepSeek-V3.1 DeepSeek | See passport | 128K tokens | Model-specific | vLLM, SGLang | — | Self-hosted |
| Devstral Small 2505 Mistral AI | See passport | 128K tokens | Model-specific | vLLM | — | Yes / documented |
| Gemma 3 12B IT Google DeepMind | See passport | 128K tokens | Model-specific | Transformers, llama.cpp, Ollama | Google Cloud | Yes / documented |
| Gemma 3 27B IT Google DeepMind | See passport | 128K tokens | Model-specific | Transformers, llama.cpp, Ollama | Google Cloud | Yes / documented |
| Gemma 3 4B IT Google DeepMind | See passport | 128K tokens | Model-specific | Transformers, llama.cpp, Ollama | Google Cloud | Yes / documented |
| Gemma 3n E4B IT Google DeepMind | See passport | 32K tokens | Model-specific | Transformers | — | Yes / documented |
| GLM-4.5 Z.ai | See passport | 128K tokens | Model-specific | Transformers, vLLM, SGLang | — | Self-hosted |
| GLM-4.5 Air Z.ai | See passport | 128K tokens | Model-specific | Transformers, vLLM, SGLang | — | Self-hosted |
| gpt-oss-120b OpenAI | See passport | 131,072 tokens | Model-specific | Transformers, vLLM, SGLang, Ollama | — | Yes / documented |
| gpt-oss-20b OpenAI | See passport | 131,072 tokens | Model-specific | Transformers, vLLM, SGLang, Ollama | — | Yes / documented |
| Granite 3.3 8B Instruct IBM | See passport | 128K tokens | Model-specific | Transformers, vLLM | — | Yes / documented |
| Kimi K2 Instruct Moonshot AI | See passport | 128K tokens | Model-specific | Transformers, vLLM, SGLang | — | Yes / documented |
| Llama 3.3 70B Instruct Meta | See passport | 128K tokens | Model-specific | Transformers, vLLM, SGLang | — | Yes / documented |
| Llama 4 Maverick Meta | See passport | 1000000 tokens | Model-specific | Transformers, vLLM, SGLang, NVIDIA NIM | AWS / Bedrock, Microsoft Foundry | Self-hosted |
| Llama 4 Scout Meta | See passport | 10,000,000 tokens | Model-specific | Transformers, vLLM, SGLang | — | Self-hosted |
| Magistral Small 2506 Mistral AI | See passport | 128K nominal · 40K recommended for quality | Model-specific | vLLM | — | Yes / documented |
| Mathstral 7B v0.1 Mistral AI | See passport | 32,768 tokens | Model-specific | Transformers, vLLM | — | Self-hosted |
| Mistral Nemo Instruct 2407 Mistral AI / NVIDIA | See passport | 128K tokens | Model-specific | Transformers, vLLM | — | Self-hosted |
| Mistral Small 3.1 24B Instruct Mistral AI | See passport | 128K tokens | Model-specific | Transformers, vLLM | — | Yes / documented |
| OLMo 2 13B Instruct Ai2 | See passport | 4,096 tokens | Model-specific | Transformers, SGLang | — | Self-hosted |
| OLMo 2 32B Instruct Ai2 | See passport | 4,096 tokens | Model-specific | Transformers, SGLang | — | Self-hosted |
| Phi-4 Microsoft | See passport | 16K tokens | Model-specific | Transformers | — | Yes / documented |
| Phi-4 Mini Instruct Microsoft | See passport | 128K tokens | Model-specific | Transformers | — | Yes / documented |
| Phi-4 Multimodal Instruct Microsoft | See passport | 128K tokens | Model-specific | Transformers | — | Self-hosted |
| Qwen2.5-Coder-32B-Instruct Qwen / Alibaba | See passport | 32K native · long-context extension documented by Qwen | Model-specific | Transformers, vLLM, SGLang | — | Yes / documented |
| Qwen2.5-VL-72B-Instruct Qwen / Alibaba | See passport | 32K recommended default · config exposes 128K positions | Model-specific | Transformers, vLLM, SGLang | — | Self-hosted |
| Qwen3-235B-A22B Qwen / Alibaba | See passport | 32,768 native · 131,072 with YaRN | Model-specific | Transformers, vLLM, SGLang | — | Self-hosted |
| Qwen3-30B-A3B Qwen / Alibaba | See passport | 32,768 native · 131,072 with YaRN | Model-specific | Transformers, vLLM, SGLang | — | Yes / documented |
| Qwen3-32B Qwen / Alibaba | See passport | 32,768 native · 131,072 with YaRN | Model-specific | Transformers, vLLM, SGLang, llama.cpp, Ollama | — | Yes / documented |
| Qwen3-Coder-480B-A35B-Instruct Qwen / Alibaba | See passport | 256K tokens | Model-specific | Transformers, vLLM, SGLang | — | Self-hosted |
| SmolLM3 3B Hugging Face | See passport | 64K native · up to 128K extended | Model-specific | Transformers, vLLM, llama.cpp | — | Yes / documented |
OpenWeightModels uses conservative editorial classes to prevent a misleading “self-hostable” badge from flattening very different infrastructure realities.
Generally small models up to roughly 8B total parameters, often practical with quantization on consumer GPUs, Apple Silicon or CPU/GPU hybrid setups. Exact context and precision still matter.
Roughly 8B–32B dense-scale models. Quantization can make many practical on high-memory workstations; full-precision or long-context use may push them into server GPUs.
Roughly 32B–80B total parameters. These often require high-memory accelerators, sharding or substantial system RAM for useful throughput.
Very large dense or MoE checkpoints. Active parameters can reduce per-token compute without shrinking the total weight set that must be stored and orchestrated.
Long context can dominate additional memory. “Fits at 4K” and “supports 128K at production throughput” are not equivalent deployment claims.
Lower precision can materially reduce weight memory, but format, runtime compatibility, quantization recipe and quality effects must be tracked separately.
Counts below reflect the 32 current Passports, not universal compatibility claims for every model in the ecosystem.
Direct Python inference and model integration across text, vision, audio and multimodal tasks. Useful as a reference implementation and application library.
High-throughput model serving with an OpenAI-compatible HTTP server for supported models. Production hardening remains the operator's responsibility.
Serving and structured-generation runtime used by several modern large/MoE model releases, including model-specific deployment guidance.
C/C++ inference across a broad hardware range. GGUF and quantization make many smaller/medium models practical locally and enable CPU+GPU hybrid execution.
Developer-oriented local model management. Where a publisher explicitly documents Ollama, the Passport records it as an official/documented path rather than inferring it from community availability.
Packaged, validated inference profiles can publish precise GPU and context envelopes. Maverick's H100/H200 NIM differences are a key example of deployment-specific data.
A provider catalog changes independently of a model release. Regions, APIs and model lifecycle must therefore be checked against the named platform.
AWS documents supported foundation models, API compatibility, endpoints and region availability separately. For model-specific records, OpenWeightModels uses the current Bedrock documentation rather than assuming every Meta/Mistral model is present.
Google documents both managed model APIs and one-click/self-deploy open-model paths with verified machine and accelerator options for supported entries.
Microsoft distinguishes models sold by Azure, partner/community models, managed compute and serverless deployment. Capability and lifecycle data belong to the specific catalog entry.
OpenWeightModels stores the base model context and the provider/runtime context separately whenever evidence supports both.
This avoids one of the most common errors in model directories: copying a model-card maximum into every cloud-provider row. Provider endpoints may choose a smaller maximum context, a different maximum output length, restricted image counts, tool support or a particular precision checkpoint.
Hardware also matters. NVIDIA's documented Maverick NIM profile supports 1M context on H200 but 430K on H100. Those are valid deployment records without contradicting Meta's 1M base-model figure.
The same principle applies to context-extension techniques. A 32K native context with YaRN extension to 128K is not written as if 128K were the unqualified native window.
A useful deployment estimate is a chain of assumptions. Skipping one of them can turn an apparently precise VRAM number into a poor capacity plan.
Base vs instruct, BF16 vs FP8, official vs community quantization and dense vs MoE determine the stored weight set and runtime compatibility.
Parameter count × bytes per parameter is a first-order estimate only. For MoE, use total stored parameters for weight storage, not only active parameters per token.
KV cache grows with context length, batch/concurrency and architecture. A long-context production endpoint can require far more memory than a short single-user test.
Tensor, pipeline and expert parallelism, CPU offload, runtime buffers and replication determine how the model is split and what throughput is achievable.
Quantization changes the deployment artifact, memory profile and often the runtime path.
An official FP8 checkpoint released by the model developer carries different provenance from a community 4-bit GGUF conversion. Both may be useful, but they should not be described as if they were the same weights with a smaller file size.
OpenWeightModels therefore tries to preserve four facts: publisher, precision/format, quantization status and runtime. Where a publisher validates a specific checkpoint on specific hardware — for example NVIDIA NIM with an official FP8 model — that combination is a stronger deployment claim than a generic “FP8 supported” label.
Quality impact is not assumed to be zero. If a publisher provides evaluation evidence for a quantized checkpoint, it can be recorded; otherwise the Explorer avoids promising equivalence to BF16.
Choosing direct weights also chooses which production responsibilities stay with the operator.
Authentication, network isolation, request validation and rate limiting belong to the deployment stack. A runtime's OpenAI-compatible server is not automatically a complete secure edge.
Replica strategy, model loading, GPU failures, rolling updates and capacity headroom become operational engineering tasks.
Useful production telemetry includes time-to-first-token, inter-token latency, throughput, queueing, KV-cache pressure, GPU utilization and error rates.
The deployed artifact, hash, quantization, system prompt and runtime version should be traceable if reproducibility or regulated change control matters.
Managed providers may add content controls or policy layers. Direct weights generally leave application-level safeguards to the operator.
Owning accelerators can be efficient at sustained load; managed APIs can be attractive at variable load. Model size alone does not determine total cost.
The Explorer is designed to narrow a deployment decision without pretending that one route is best for every team.
Check that the intended commercial, hosted, redistribution and jurisdictional use fits the model terms before optimizing infrastructure.
Define modality, context, output length, concurrency, latency target, tool/function needs and privacy requirements.
Decide whether direct weights, a packaged runtime or a managed API best matches the team's operational capabilities and governance needs.
Use a documented combination where possible, then benchmark the exact checkpoint and quantization under the intended prompt/context mix.
If using cloud inference, verify region, endpoint/API compatibility, context, output limits, tools, model version and lifecycle from the current provider documentation.
Benchmark latency, throughput, quality and memory on representative traffic before committing to a production capacity plan.
SmolLM3 3B, Phi-4 Mini and compact Gemma variants illustrate the local end of the spectrum: quantized checkpoints can be practical on consumer hardware, workstations or edge-oriented stacks. That accessibility is a deployment property, not a different definition of weight access.
At the other end, DeepSeek-R1, Kimi K2, Qwen3-Coder 480B-A35B and Llama 4 Maverick are large sparse systems. Their active parameter counts can make inference compute more efficient than an equally large dense model, but the total expert weights still create datacenter-scale storage, sharding and interconnect requirements.
This is why the Explorer exposes both total parameter scale and documented deployment paths. Neither number alone is sufficient.
No. It means the trained weights are available. Some models are compact enough for local deployment; others contain hundreds of billions or a trillion total parameters and remain datacenter workloads.
There is no universal best runtime. The answer depends on model support, hardware, quantization, throughput/latency targets, API needs and operational constraints. The Explorer reports documented paths rather than assigning a winner.
Because precision, context, KV cache, batching, runtime overhead and sharding can change the requirement materially. A precise number without those conditions often creates false confidence.
Not automatically. Guardrails, function calling wrappers, endpoint compatibility, output limits and regions can be platform features or provider envelopes. They are attached to deployment records.
Weight access and technical feasibility do not override model-specific legal terms.