Operational reference for open weights

Deployment Explorer

Compare how 32 open-weight models can be run — not just whether weights exist.

Deployment is recorded as a combination of checkpoint, runtime/provider, precision and hardware. The Explorer separates local inference, self-hosted serving and managed APIs while preserving provider-specific context and feature differences.

Self-hosted ≠ localA model can expose downloadable weights yet still require a multi-GPU datacenter deployment.
Context is implementation-specificProvider limits can differ from the base model's documented maximum.
Quantization is a deployment choiceBF16, FP8 and GGUF variants change memory, compatibility and provenance.
Hardware classes are editorialThey guide navigation; they are not guaranteed minimum system requirements.
Definition

What is an open-weight model Deployment Explorer?

The OpenWeightModels Deployment Explorer compares the practical ways downloadable AI models can be served: directly from weights, through inference runtimes, in local applications, or through managed cloud platforms.

The central idea is that deployment is not a binary “self-hosting: yes” field. A 3B quantized model running in llama.cpp on a laptop and a 671B sparse model spread across datacenter GPUs are both self-hosted, but they are radically different operational systems.

Likewise, a base model's headline context window does not guarantee that every provider exposes the same context, output limit, tool support or region. Llama 4 Maverick demonstrates the problem clearly: its base model is documented at 1M tokens, while provider-specific implementations can expose different limits by cloud or GPU class.

The Explorer therefore treats provider and runtime deployments as child records of a model, not as universal capabilities inherited by every implementation.

Deployment paths

Four layers between weights and an application.

Choosing a model is also choosing how much infrastructure responsibility the operator wants to own.

01

Direct weights

Maximum control. The operator owns loading, sharding, batching, scaling, safeguards, observability and lifecycle. Precision and model format are explicit choices.

02

Serving runtime

vLLM and SGLang provide production-oriented servers for supported models; Transformers covers direct model execution; llama.cpp and Ollama emphasize accessible local or portable workflows.

03

Packaged inference

Products such as NVIDIA NIM package validated model/runtime/hardware combinations, which can make the supported operating envelope more specific than a generic self-host setup.

04

Managed API / cloud

A provider owns much of the serving infrastructure but can impose its own region, context, endpoint, feature, price and lifecycle envelope.

Interactive matrix

Compare 32 deployment profiles.

Hardware class is an editorial navigation aid based primarily on total parameter scale and documented paths. Open each Passport for exact evidence.

32 models
ModelParametersContextEditorial hardware classDocumented runtimesManaged/cloudLocal/self-host
DeepSeek-R1
DeepSeek
See passport128K tokensModel-specificvLLM, SGLangDeepSeek APISelf-hosted
DeepSeek-V3.1
DeepSeek
See passport128K tokensModel-specificvLLM, SGLang—Self-hosted
Devstral Small 2505
Mistral AI
See passport128K tokensModel-specificvLLM—Yes / documented
Gemma 3 12B IT
Google DeepMind
See passport128K tokensModel-specificTransformers, llama.cpp, OllamaGoogle CloudYes / documented
Gemma 3 27B IT
Google DeepMind
See passport128K tokensModel-specificTransformers, llama.cpp, OllamaGoogle CloudYes / documented
Gemma 3 4B IT
Google DeepMind
See passport128K tokensModel-specificTransformers, llama.cpp, OllamaGoogle CloudYes / documented
Gemma 3n E4B IT
Google DeepMind
See passport32K tokensModel-specificTransformers—Yes / documented
GLM-4.5
Z.ai
See passport128K tokensModel-specificTransformers, vLLM, SGLang—Self-hosted
GLM-4.5 Air
Z.ai
See passport128K tokensModel-specificTransformers, vLLM, SGLang—Self-hosted
gpt-oss-120b
OpenAI
See passport131,072 tokensModel-specificTransformers, vLLM, SGLang, Ollama—Yes / documented
gpt-oss-20b
OpenAI
See passport131,072 tokensModel-specificTransformers, vLLM, SGLang, Ollama—Yes / documented
Granite 3.3 8B Instruct
IBM
See passport128K tokensModel-specificTransformers, vLLM—Yes / documented
Kimi K2 Instruct
Moonshot AI
See passport128K tokensModel-specificTransformers, vLLM, SGLang—Yes / documented
Llama 3.3 70B Instruct
Meta
See passport128K tokensModel-specificTransformers, vLLM, SGLang—Yes / documented
Llama 4 Maverick
Meta
See passport1000000 tokensModel-specificTransformers, vLLM, SGLang, NVIDIA NIMAWS / Bedrock, Microsoft FoundrySelf-hosted
Llama 4 Scout
Meta
See passport10,000,000 tokensModel-specificTransformers, vLLM, SGLang—Self-hosted
Magistral Small 2506
Mistral AI
See passport128K nominal · 40K recommended for qualityModel-specificvLLM—Yes / documented
Mathstral 7B v0.1
Mistral AI
See passport32,768 tokensModel-specificTransformers, vLLM—Self-hosted
Mistral Nemo Instruct 2407
Mistral AI / NVIDIA
See passport128K tokensModel-specificTransformers, vLLM—Self-hosted
Mistral Small 3.1 24B Instruct
Mistral AI
See passport128K tokensModel-specificTransformers, vLLM—Yes / documented
OLMo 2 13B Instruct
Ai2
See passport4,096 tokensModel-specificTransformers, SGLang—Self-hosted
OLMo 2 32B Instruct
Ai2
See passport4,096 tokensModel-specificTransformers, SGLang—Self-hosted
Phi-4
Microsoft
See passport16K tokensModel-specificTransformers—Yes / documented
Phi-4 Mini Instruct
Microsoft
See passport128K tokensModel-specificTransformers—Yes / documented
Phi-4 Multimodal Instruct
Microsoft
See passport128K tokensModel-specificTransformers—Self-hosted
Qwen2.5-Coder-32B-Instruct
Qwen / Alibaba
See passport32K native · long-context extension documented by QwenModel-specificTransformers, vLLM, SGLang—Yes / documented
Qwen2.5-VL-72B-Instruct
Qwen / Alibaba
See passport32K recommended default · config exposes 128K positionsModel-specificTransformers, vLLM, SGLang—Self-hosted
Qwen3-235B-A22B
Qwen / Alibaba
See passport32,768 native · 131,072 with YaRNModel-specificTransformers, vLLM, SGLang—Self-hosted
Qwen3-30B-A3B
Qwen / Alibaba
See passport32,768 native · 131,072 with YaRNModel-specificTransformers, vLLM, SGLang—Yes / documented
Qwen3-32B
Qwen / Alibaba
See passport32,768 native · 131,072 with YaRNModel-specificTransformers, vLLM, SGLang, llama.cpp, Ollama—Yes / documented
Qwen3-Coder-480B-A35B-Instruct
Qwen / Alibaba
See passport256K tokensModel-specificTransformers, vLLM, SGLang—Self-hosted
SmolLM3 3B
Hugging Face
See passport64K native · up to 128K extendedModel-specificTransformers, vLLM, llama.cpp—Yes / documented
Hardware classes

Scale is a map, not a minimum-spec sheet.

OpenWeightModels uses conservative editorial classes to prevent a misleading “self-hostable” badge from flattening very different infrastructure realities.

LOCAL

Local / consumer-friendly

Generally small models up to roughly 8B total parameters, often practical with quantization on consumer GPUs, Apple Silicon or CPU/GPU hybrid setups. Exact context and precision still matter.

WORKSTATION

Workstation / server GPU

Roughly 8B–32B dense-scale models. Quantization can make many practical on high-memory workstations; full-precision or long-context use may push them into server GPUs.

MULTI-GPU

Large-memory / multi-GPU

Roughly 32B–80B total parameters. These often require high-memory accelerators, sharding or substantial system RAM for useful throughput.

DATACENTER

Datacenter / multi-GPU

Very large dense or MoE checkpoints. Active parameters can reduce per-token compute without shrinking the total weight set that must be stored and orchestrated.

CONTEXT

KV cache changes the answer

Long context can dominate additional memory. “Fits at 4K” and “supports 128K at production throughput” are not equivalent deployment claims.

PRECISION

BF16, FP8, INT4, GGUF

Lower precision can materially reduce weight memory, but format, runtime compatibility, quantization recipe and quality effects must be tracked separately.

First-order memory rule: raw weight memory is approximately parameters × bytes per parameter. That estimate excludes KV cache, activations, runtime buffers, fragmentation and replication, so it is never presented as a complete production requirement.
Serving & local runtimes

The runtime changes the operating model.

Counts below reflect the 32 current Passports, not universal compatibility claims for every model in the ecosystem.

Transformers · 28 records

Direct Python inference and model integration across text, vision, audio and multimodal tasks. Useful as a reference implementation and application library.

vLLM · 23 records

High-throughput model serving with an OpenAI-compatible HTTP server for supported models. Production hardening remains the operator's responsibility.

SGLang · 18 records

Serving and structured-generation runtime used by several modern large/MoE model releases, including model-specific deployment guidance.

llama.cpp / GGUF · 5 records

C/C++ inference across a broad hardware range. GGUF and quantization make many smaller/medium models practical locally and enable CPU+GPU hybrid execution.

Ollama · 6 records

Developer-oriented local model management. Where a publisher explicitly documents Ollama, the Passport records it as an official/documented path rather than inferring it from community availability.

NVIDIA NIM · 1 deep record

Packaged, validated inference profiles can publish precise GPU and context envelopes. Maverick's H100/H200 NIM differences are a key example of deployment-specific data.

Managed deployment

Cloud availability is a separate dataset.

A provider catalog changes independently of a model release. Regions, APIs and model lifecycle must therefore be checked against the named platform.

Amazon Bedrock

AWS documents supported foundation models, API compatibility, endpoints and region availability separately. For model-specific records, OpenWeightModels uses the current Bedrock documentation rather than assuming every Meta/Mistral model is present.

Google Model Garden / Agent Platform

Google documents both managed model APIs and one-click/self-deploy open-model paths with verified machine and accelerator options for supported entries.

Microsoft Foundry

Microsoft distinguishes models sold by Azure, partner/community models, managed compute and serverless deployment. Capability and lifecycle data belong to the specific catalog entry.

Context & provider variance

The same weights can have different limits.

OpenWeightModels stores the base model context and the provider/runtime context separately whenever evidence supports both.

This avoids one of the most common errors in model directories: copying a model-card maximum into every cloud-provider row. Provider endpoints may choose a smaller maximum context, a different maximum output length, restricted image counts, tool support or a particular precision checkpoint.

Hardware also matters. NVIDIA's documented Maverick NIM profile supports 1M context on H200 but 430K on H100. Those are valid deployment records without contradicting Meta's 1M base-model figure.

The same principle applies to context-extension techniques. A 32K native context with YaRN extension to 128K is not written as if 128K were the unqualified native window.

Sizing workflow

From model card to hardware plan.

A useful deployment estimate is a chain of assumptions. Skipping one of them can turn an apparently precise VRAM number into a poor capacity plan.

STEP 1

Choose the exact weights

Base vs instruct, BF16 vs FP8, official vs community quantization and dense vs MoE determine the stored weight set and runtime compatibility.

STEP 2

Estimate raw weight memory

Parameter count × bytes per parameter is a first-order estimate only. For MoE, use total stored parameters for weight storage, not only active parameters per token.

STEP 3

Add context & concurrency

KV cache grows with context length, batch/concurrency and architecture. A long-context production endpoint can require far more memory than a short single-user test.

STEP 4

Choose parallelism & runtime

Tensor, pipeline and expert parallelism, CPU offload, runtime buffers and replication determine how the model is split and what throughput is achievable.

Quantization

Lower precision is not one interchangeable switch.

Quantization changes the deployment artifact, memory profile and often the runtime path.

An official FP8 checkpoint released by the model developer carries different provenance from a community 4-bit GGUF conversion. Both may be useful, but they should not be described as if they were the same weights with a smaller file size.

OpenWeightModels therefore tries to preserve four facts: publisher, precision/format, quantization status and runtime. Where a publisher validates a specific checkpoint on specific hardware — for example NVIDIA NIM with an official FP8 model — that combination is a stronger deployment claim than a generic “FP8 supported” label.

Quality impact is not assumed to be zero. If a publisher provides evaluation evidence for a quantized checkpoint, it can be recorded; otherwise the Explorer avoids promising equivalence to BF16.

Operational responsibility

Self-hosting moves more than inference.

Choosing direct weights also chooses which production responsibilities stay with the operator.

Security

Endpoint hardening

Authentication, network isolation, request validation and rate limiting belong to the deployment stack. A runtime's OpenAI-compatible server is not automatically a complete secure edge.

Reliability

Autoscaling & failure domains

Replica strategy, model loading, GPU failures, rolling updates and capacity headroom become operational engineering tasks.

Observability

Latency, tokens & memory

Useful production telemetry includes time-to-first-token, inter-token latency, throughput, queueing, KV-cache pressure, GPU utilization and error rates.

Governance

Model/version control

The deployed artifact, hash, quantization, system prompt and runtime version should be traceable if reproducibility or regulated change control matters.

Safety

Guardrails are not inherent

Managed providers may add content controls or policy layers. Direct weights generally leave application-level safeguards to the operator.

Cost

Utilization decides economics

Owning accelerators can be efficient at sustained load; managed APIs can be attractive at variable load. Model size alone does not determine total cost.

Decision workflow

How to use the Explorer.

The Explorer is designed to narrow a deployment decision without pretending that one route is best for every team.

01Confirm the license

Check that the intended commercial, hosted, redistribution and jurisdictional use fits the model terms before optimizing infrastructure.

02Set the workload envelope

Define modality, context, output length, concurrency, latency target, tool/function needs and privacy requirements.

03Choose control level

Decide whether direct weights, a packaged runtime or a managed API best matches the team's operational capabilities and governance needs.

04Select precision & runtime

Use a documented combination where possible, then benchmark the exact checkpoint and quantization under the intended prompt/context mix.

05Validate provider variance

If using cloud inference, verify region, endpoint/API compatibility, context, output limits, tools, model version and lifecycle from the current provider documentation.

06Measure real workload

Benchmark latency, throughput, quality and memory on representative traffic before committing to a production capacity plan.

Local vs datacenter

Two models can both be “open weight” and live in different worlds.

SmolLM3 3B, Phi-4 Mini and compact Gemma variants illustrate the local end of the spectrum: quantized checkpoints can be practical on consumer hardware, workstations or edge-oriented stacks. That accessibility is a deployment property, not a different definition of weight access.

At the other end, DeepSeek-R1, Kimi K2, Qwen3-Coder 480B-A35B and Llama 4 Maverick are large sparse systems. Their active parameter counts can make inference compute more efficient than an equally large dense model, but the total expert weights still create datacenter-scale storage, sharding and interconnect requirements.

This is why the Explorer exposes both total parameter scale and documented deployment paths. Neither number alone is sufficient.

FAQ

Deployment questions, separated.

Does “open weight” mean I can run the model on my laptop?

No. It means the trained weights are available. Some models are compact enough for local deployment; others contain hundreds of billions or a trillion total parameters and remain datacenter workloads.

Which runtime is best?

There is no universal best runtime. The answer depends on model support, hardware, quantization, throughput/latency targets, API needs and operational constraints. The Explorer reports documented paths rather than assigning a winner.

Why don't you list one minimum VRAM number?

Because precision, context, KV cache, batching, runtime overhead and sharding can change the requirement materially. A precise number without those conditions often creates false confidence.

Are provider features part of the model?

Not automatically. Guardrails, function calling wrappers, endpoint compatibility, output limits and regions can be platform features or provider envelopes. They are attached to deployment records.

Deployment starts with rights

Check the license before choosing the runtime.

Weight access and technical feasibility do not override model-specific legal terms.

Open License Explorer