Weights available
The trained parameters must be obtainable from an official or clearly attributable distribution channel. “Open API” without weight access is not enough.
How OpenWeightModels verifies models, licenses and deployment information.
The methodology is designed to make every model passport auditable. We separate what a model is, what its license permits, and how a specific runtime or cloud provider deploys it. When those facts are uncertain, the uncertainty stays visible.
The OpenWeightModels methodology is a source-first process for documenting open-weight AI models as three related but separate records: the base model, its legal terms, and its deployment implementations.
This separation exists because a single row such as “Llama 4 Maverick · 1M context · commercial use · tool calling” can silently combine claims that come from different authorities. Meta can document the base architecture and base context; a cloud provider can expose a smaller context window; a license can permit broad use while an incorporated policy adds restrictions; a serving runtime can add an OpenAI-compatible API without changing the model itself.
OpenWeightModels therefore tries to preserve provenance. A model-card statement remains a model-card statement. A provider feature remains attached to that provider. A calculated memory figure is labeled as an estimate. An editorial classification is explicitly separated from a publisher claim.
The methodology is optimized for technical selection, reproducibility, search and agentic retrieval — not for producing a single leaderboard or a one-number “openness” score.
The registry is curated rather than exhaustive. Inclusion requires that trained weights are available to obtain and that the release has enough first-party evidence to document identity, legal terms and a plausible deployment path.
The trained parameters must be obtainable from an official or clearly attributable distribution channel. “Open API” without weight access is not enough.
The checkpoint, family, developer and tuning state must be sufficiently specific to avoid combining materially different artifacts under one name.
A license, terms document, model card or official repository must make it possible to assess usage and redistribution conditions.
The model should add meaningful architectural, licensing, modality, scale or deployment coverage to the curated set.
Status labels describe the evidence behind a field. They are not quality scores for the model.
A current publisher, official repository, exact license text or named deployment provider directly supports the claim.
The fact appears in official documentation, but may describe a family, environment or implementation rather than the exact checkpoint in every condition.
No sufficiently reliable current source was found. The field remains unknown rather than being inferred from neighboring models.
OpenWeightModels derives a comparison label from documented facts — for example “datacenter / multi-GPU” — and marks it as editorial rather than publisher language.
Source priority is contextual. A cloud provider is authoritative for its own endpoint, while the model developer is normally authoritative for the underlying checkpoint.
Exact model card, repository, configuration file, license text, release note or technical report for the checkpoint being documented.
Official family documentation, developer blog, FAQ or platform documentation when the exact checkpoint page does not contain the field.
AWS, Google Cloud, Microsoft, NVIDIA or another named provider is primary for its own context limits, regions, APIs, hardware profiles and lifecycle.
Official vLLM, SGLang, llama.cpp, Transformers or Ollama documentation for runtime-specific compatibility and behavior.
Third-party quantizations, benchmarks or deployment notes may be useful, but they are labeled as community evidence and never silently promoted to official status.
Before comparing capability, OpenWeightModels fixes the artifact identity: exact checkpoint, base vs instruct, precision, modality, architecture and release.
For MoE models, total stored parameters and active parameters per token are separate fields. Active parameters are not used as a proxy for storage footprint.
Developer context limits are stored separately from provider- or hardware-specific limits. YaRN or other extension methods are identified as extensions rather than native context.
Base, Instruct, Thinking, FP8, GGUF and distilled variants can have different behavior, sizes and terms. Family names do not replace checkpoint identity.
“Multimodal” is not enough. Text, image, audio and video inputs are listed separately from generated output modalities.
Training tokens, datasets, GPU hours, post-training methods and cutoffs are included when publishers disclose them and omitted when they do not.
A single minimum-VRAM number is avoided unless the checkpoint/runtime combination is documented. KV cache, batch size and precision materially change real memory use.
For every model, the license record asks separate questions about commercial use, modification, redistribution, attribution, hosted services, derivative models, acceptable-use policies, scale thresholds and geography.
Apache 2.0 and MIT are tracked as OSI-approved permissive software licenses. Custom model agreements such as Llama, Gemma and Qwen are described by their actual conditions rather than mapped into a misleading permissive/non-permissive binary.
A separate usage policy is not automatically treated as part of a software license. When a license incorporates another policy by reference, that relationship is recorded explicitly. When a model is Apache-2.0 licensed but its publisher also publishes a separate usage policy, both facts can coexist without rewriting the Apache license.
OpenWeightModels does not provide legal advice and does not determine enforceability. It summarizes the text and highlights clauses that materially affect technical deployment decisions.
Deployment records belong to a model + checkpoint + runtime/provider + precision + hardware context, not to the family name alone.
Self-hosted weights, vLLM, SGLang, llama.cpp, Ollama, NVIDIA NIM, AWS Bedrock, Google Cloud and Microsoft Foundry can expose different context windows, tools, output limits, regions and safety layers. Those differences stay attached to the named implementation.
Hardware classifications such as “local / consumer-friendly” or “datacenter / multi-GPU” are editorial navigation aids. They are based primarily on total parameter scale and documented deployment paths, and are never presented as exact minimum requirements.
Quantization is tracked as a deployment variable. An official FP8 checkpoint is distinct from a community GGUF quantization; the publisher, recipe and format matter for provenance.
OpenWeightModels records selected evaluation evidence but does not collapse heterogeneous benchmark tables into an overall ranking.
Scores are attributed to the developer or evaluator. Precision, checkpoint and shot setting are retained where available.
Accuracy, pass@1, exact match and other metrics are not treated as interchangeable. Benchmark version and date can matter.
Calculated weight-memory values use transparent assumptions. They never replace documented hardware requirements and exclude runtime/KV-cache overhead unless explicitly modeled.
Every Gold Passport carries a verification date and change history so corrections and provider changes do not disappear into silent edits.
License changes, model revisions, provider lifecycle events, context-limit changes and major new official checkpoints trigger a review of affected records.
If a source was misread or a field was attached to the wrong layer, the field is corrected and the change history records the material correction.
Human pages and JSON records are updated together where practical. `verified_at`, source URLs and explicit classifications support agentic retrieval and auditing.
The methodology is published as `/data/methodology.json`. License classifications live in `/data/licenses.json`, deployment records in `/data/deployments.json`, and model records in `/data/models.json` plus per-model `model.json` files.
This makes the site's definitions and status vocabulary usable by tools without forcing them to reverse-engineer prose or tables.
Uncertainty is treated as data. The site prefers a visible unknown to a convenient but unsupported value.
A missing field on one checkpoint is not automatically copied from another size, tuning variant or family member. Closely related models can differ in license, context, modality, precision and supported runtimes.
If two official sources disagree, we first ask whether they govern different things. A base model card, an FP8 repository and a managed endpoint can all publish different limits without any source being “wrong”.
Provider catalogs and runtime support change faster than architecture facts. Deployment records therefore need a current verification date and can become stale independently from the underlying model passport.
The methodology asks which source is actually competent to support a specific claim rather than using one favorite website for everything.
| Field | Preferred authority | Fallback | What we avoid |
|---|---|---|---|
| Parameters / architecture | Exact publisher model card or config | Official technical report | Third-party size summaries when the publisher provides exact data |
| Context window | Exact model config/card | Official family docs with checkpoint match | Applying a provider maximum to the base model |
| Commercial rights | Exact license/terms | Official legal FAQ that quotes/interprets the terms | Model-hub metadata as the sole legal source |
| Cloud regions/features | Named cloud provider | Developer docs linking to that provider | Assuming all providers expose the same capabilities |
| Runtime compatibility | Publisher or runtime project docs | Reproducible community evidence, clearly labeled | Inferring support because a related architecture works |
| Hardware | Validated publisher/provider profile | Transparent editorial estimate | One “minimum VRAM” number without precision/context conditions |
The human page and the structured record intentionally mirror the same conceptual layers.
A model record stores identity, architecture, training disclosures, weights and high-level capability evidence. A license record stores the exact legal instrument and decomposed rights. A deployment record stores runtime/provider-specific facts such as model ID, precision, context, hardware and feature envelope.
This structure avoids destructive normalization. For example, a deployment provider's context window is not copied into `architecture.context`, and a separate usage policy is not rewritten into the `license.name` field. Agents can therefore retrieve a value together with the layer that owns it.
Where a field is editorial — for example a hardware class — the structured data says so. This allows downstream systems to decide whether they want only source-backed values or also curated interpretation.
A correction should improve both the fact and its provenance, not merely replace visible text.
Determine whether the error belongs to the base model, license, runtime or provider record.
Use a closer or newer authority and retain the current URL in the relevant data record.
Correct human text, JSON, tables and any explorer that aggregates the affected field.
Substantive corrections belong in the Passport change history with a new verification date.
No. VERIFIED means OpenWeightModels reviewed a current primary source that supports the field. It is a provenance status, not a third-party certification of the model, provider or license.
A single score would combine legally and technically different dimensions — weight access, software license, training transparency, redistribution rights and deployment freedom. The registry keeps those dimensions separate so readers can apply their own requirements.
Useful navigation sometimes requires synthesis. Hardware class is a good example: parameter scale and documented runtimes can support a practical category even when no developer uses that exact phrase. The label remains explicitly editorial.
Yes. If official weights disappear, terms become unverifiable or a release is superseded in a way that makes the current record misleading, the entry can be revised, archived or removed. Historical facts should not be presented as current deployment guidance.
The two explorers apply these rules across all 32 curated model passports.