Source-first research standard

Methodology

How OpenWeightModels verifies models, licenses and deployment information.

The methodology is designed to make every model passport auditable. We separate what a model is, what its license permits, and how a specific runtime or cloud provider deploys it. When those facts are uncertain, the uncertainty stays visible.

Primary sources firstPublisher model cards, repositories, licenses and provider documentation.
No hidden interpolationMissing facts stay missing instead of being converted into confident estimates.
Deployment is contextualContext, tools, regions and hardware can differ by runtime or provider.
Machine-readableThe same classifications are published in JSON alongside human-readable pages.
Definition

What is the OpenWeightModels methodology?

The OpenWeightModels methodology is a source-first process for documenting open-weight AI models as three related but separate records: the base model, its legal terms, and its deployment implementations.

This separation exists because a single row such as “Llama 4 Maverick · 1M context · commercial use · tool calling” can silently combine claims that come from different authorities. Meta can document the base architecture and base context; a cloud provider can expose a smaller context window; a license can permit broad use while an incorporated policy adds restrictions; a serving runtime can add an OpenAI-compatible API without changing the model itself.

OpenWeightModels therefore tries to preserve provenance. A model-card statement remains a model-card statement. A provider feature remains attached to that provider. A calculated memory figure is labeled as an estimate. An editorial classification is explicitly separated from a publisher claim.

The methodology is optimized for technical selection, reproducibility, search and agentic retrieval — not for producing a single leaderboard or a one-number “openness” score.

Scope

What qualifies for the registry.

The registry is curated rather than exhaustive. Inclusion requires that trained weights are available to obtain and that the release has enough first-party evidence to document identity, legal terms and a plausible deployment path.

01

Weights available

The trained parameters must be obtainable from an official or clearly attributable distribution channel. “Open API” without weight access is not enough.

02

Identifiable release

The checkpoint, family, developer and tuning state must be sufficiently specific to avoid combining materially different artifacts under one name.

03

Terms discoverable

A license, terms document, model card or official repository must make it possible to assess usage and redistribution conditions.

04

Operational relevance

The model should add meaningful architectural, licensing, modality, scale or deployment coverage to the curated set.

Open weight is a scope term, not a legal verdict. Weight availability does not automatically mean the full AI system meets every definition of open source. The Open Source Initiative's Open Source AI Definition considers broader freedoms and access to preferred forms for modification, not only the presence of downloadable weights.
Verification status

Four labels, four meanings.

Status labels describe the evidence behind a field. They are not quality scores for the model.

VERIFIED

Primary-source confirmed

A current publisher, official repository, exact license text or named deployment provider directly supports the claim.

DOCUMENTED

Officially stated

The fact appears in official documentation, but may describe a family, environment or implementation rather than the exact checkpoint in every condition.

NOT CONFIRMED

Evidence insufficient

No sufficiently reliable current source was found. The field remains unknown rather than being inferred from neighboring models.

EDITORIAL

Explicit classification

OpenWeightModels derives a comparison label from documented facts — for example “datacenter / multi-GPU” — and marks it as editorial rather than publisher language.

Source hierarchy

Closer to the artifact wins.

Source priority is contextual. A cloud provider is authoritative for its own endpoint, while the model developer is normally authoritative for the underlying checkpoint.

LEVEL 1Exact primary artifact

Exact model card, repository, configuration file, license text, release note or technical report for the checkpoint being documented.

LEVEL 2Developer documentation

Official family documentation, developer blog, FAQ or platform documentation when the exact checkpoint page does not contain the field.

LEVEL 3Deployment authority

AWS, Google Cloud, Microsoft, NVIDIA or another named provider is primary for its own context limits, regions, APIs, hardware profiles and lifecycle.

LEVEL 4Runtime authority

Official vLLM, SGLang, llama.cpp, Transformers or Ollama documentation for runtime-specific compatibility and behavior.

LEVEL 5Community evidence

Third-party quantizations, benchmarks or deployment notes may be useful, but they are labeled as community evidence and never silently promoted to official status.

Conflict rule: when two sources disagree, the site does not average them. It identifies what each source governs. A provider's 524K context limit can coexist with a developer's 1M base-model context because they describe different layers of the system.
Model facts

Identity before comparison.

Before comparing capability, OpenWeightModels fixes the artifact identity: exact checkpoint, base vs instruct, precision, modality, architecture and release.

Parameters

Total ≠ active

For MoE models, total stored parameters and active parameters per token are separate fields. Active parameters are not used as a proxy for storage footprint.

Context

Base ≠ provider

Developer context limits are stored separately from provider- or hardware-specific limits. YaRN or other extension methods are identified as extensions rather than native context.

Variants

Checkpoint matters

Base, Instruct, Thinking, FP8, GGUF and distilled variants can have different behavior, sizes and terms. Family names do not replace checkpoint identity.

Modalities

Input and output split

“Multimodal” is not enough. Text, image, audio and video inputs are listed separately from generated output modalities.

Training

Only what is disclosed

Training tokens, datasets, GPU hours, post-training methods and cutoffs are included when publishers disclose them and omitted when they do not.

Hardware

No universal GPU claim

A single minimum-VRAM number is avoided unless the checkpoint/runtime combination is documented. KV cache, batch size and precision materially change real memory use.

License methodology

Rights are decomposed, not reduced to “open”.

For every model, the license record asks separate questions about commercial use, modification, redistribution, attribution, hosted services, derivative models, acceptable-use policies, scale thresholds and geography.

Apache 2.0 and MIT are tracked as OSI-approved permissive software licenses. Custom model agreements such as Llama, Gemma and Qwen are described by their actual conditions rather than mapped into a misleading permissive/non-permissive binary.

A separate usage policy is not automatically treated as part of a software license. When a license incorporates another policy by reference, that relationship is recorded explicitly. When a model is Apache-2.0 licensed but its publisher also publishes a separate usage policy, both facts can coexist without rewriting the Apache license.

OpenWeightModels does not provide legal advice and does not determine enforceability. It summarizes the text and highlights clauses that materially affect technical deployment decisions.

Deployment methodology

A model does not have one deployment.

Deployment records belong to a model + checkpoint + runtime/provider + precision + hardware context, not to the family name alone.

Self-hosted weights, vLLM, SGLang, llama.cpp, Ollama, NVIDIA NIM, AWS Bedrock, Google Cloud and Microsoft Foundry can expose different context windows, tools, output limits, regions and safety layers. Those differences stay attached to the named implementation.

Hardware classifications such as “local / consumer-friendly” or “datacenter / multi-GPU” are editorial navigation aids. They are based primarily on total parameter scale and documented deployment paths, and are never presented as exact minimum requirements.

Quantization is tracked as a deployment variable. An official FP8 checkpoint is distinct from a community GGUF quantization; the publisher, recipe and format matter for provenance.

Benchmarks & estimates

Evidence without a synthetic score.

OpenWeightModels records selected evaluation evidence but does not collapse heterogeneous benchmark tables into an overall ranking.

Keep the evaluator

Scores are attributed to the developer or evaluator. Precision, checkpoint and shot setting are retained where available.

Keep the metric

Accuracy, pass@1, exact match and other metrics are not treated as interchangeable. Benchmark version and date can matter.

Keep estimates labeled

Calculated weight-memory values use transparent assumptions. They never replace documented hardware requirements and exclude runtime/KV-cache overhead unless explicitly modeled.

Maintenance

Verification is dated and revisable.

Every Gold Passport carries a verification date and change history so corrections and provider changes do not disappear into silent edits.

Material changes

License changes, model revisions, provider lifecycle events, context-limit changes and major new official checkpoints trigger a review of affected records.

Corrections

If a source was misread or a field was attached to the wrong layer, the field is corrected and the change history records the material correction.

Machine-readable parity

Human pages and JSON records are updated together where practical. `verified_at`, source URLs and explicit classifications support agentic retrieval and auditing.

Current review state: 32 curated model records and 32 Gold Passports are marked reviewed on 29 September 2026. “Verified” means the cited evidence was reviewed for this registry; it is not a warranty that a third party has independently audited the model.
Machine-readable methodology

The rules are data too.

The methodology is published as `/data/methodology.json`. License classifications live in `/data/licenses.json`, deployment records in `/data/deployments.json`, and model records in `/data/models.json` plus per-model `model.json` files.

This makes the site's definitions and status vocabulary usable by tools without forcing them to reverse-engineer prose or tables.

Uncertainty & conflicts

What happens when the evidence is incomplete.

Uncertainty is treated as data. The site prefers a visible unknown to a convenient but unsupported value.

MISSING

No silent inheritance

A missing field on one checkpoint is not automatically copied from another size, tuning variant or family member. Closely related models can differ in license, context, modality, precision and supported runtimes.

CONFLICT

Split the governing layer

If two official sources disagree, we first ask whether they govern different things. A base model card, an FP8 repository and a managed endpoint can all publish different limits without any source being “wrong”.

STALE

Date the evidence

Provider catalogs and runtime support change faster than architecture facts. Deployment records therefore need a current verification date and can become stale independently from the underlying model passport.

Example: a family README saying “128K context” does not automatically override an exact checkpoint configuration that exposes 32K native positions with a documented extension method. Both can be recorded, but they must be labeled correctly.
Field-level provenance

Every important field has an authority.

The methodology asks which source is actually competent to support a specific claim rather than using one favorite website for everything.

FieldPreferred authorityFallbackWhat we avoid
Parameters / architectureExact publisher model card or configOfficial technical reportThird-party size summaries when the publisher provides exact data
Context windowExact model config/cardOfficial family docs with checkpoint matchApplying a provider maximum to the base model
Commercial rightsExact license/termsOfficial legal FAQ that quotes/interprets the termsModel-hub metadata as the sole legal source
Cloud regions/featuresNamed cloud providerDeveloper docs linking to that providerAssuming all providers expose the same capabilities
Runtime compatibilityPublisher or runtime project docsReproducible community evidence, clearly labeledInferring support because a related architecture works
HardwareValidated publisher/provider profileTransparent editorial estimateOne “minimum VRAM” number without precision/context conditions
Data model

Why the JSON is separated.

The human page and the structured record intentionally mirror the same conceptual layers.

A model record stores identity, architecture, training disclosures, weights and high-level capability evidence. A license record stores the exact legal instrument and decomposed rights. A deployment record stores runtime/provider-specific facts such as model ID, precision, context, hardware and feature envelope.

This structure avoids destructive normalization. For example, a deployment provider's context window is not copied into `architecture.context`, and a separate usage policy is not rewritten into the `license.name` field. Agents can therefore retrieve a value together with the layer that owns it.

Where a field is editorial — for example a hardware class — the structured data says so. This allows downstream systems to decide whether they want only source-backed values or also curated interpretation.

Corrections & change control

How a Passport is corrected.

A correction should improve both the fact and its provenance, not merely replace visible text.

01

Identify the layer

Determine whether the error belongs to the base model, license, runtime or provider record.

02

Replace the source

Use a closer or newer authority and retain the current URL in the relevant data record.

03

Update dependent fields

Correct human text, JSON, tables and any explorer that aggregates the affected field.

04

Record material change

Substantive corrections belong in the Passport change history with a new verification date.

Methodology FAQ

How to read the registry.

Does VERIFIED mean independently audited?

No. VERIFIED means OpenWeightModels reviewed a current primary source that supports the field. It is a provenance status, not a third-party certification of the model, provider or license.

Why not score openness from 1 to 10?

A single score would combine legally and technically different dimensions — weight access, software license, training transparency, redistribution rights and deployment freedom. The registry keeps those dimensions separate so readers can apply their own requirements.

Why are some facts described as editorial?

Useful navigation sometimes requires synthesis. Hardware class is a good example: parameter scale and documented runtimes can support a practical category even when no developer uses that exact phrase. The label remains explicitly editorial.

Can a model move out of the registry?

Yes. If official weights disappear, terms become unverifiable or a release is superseded in a way that makes the current record misleading, the entry can be revised, archived or removed. Historical facts should not be presented as current deployment guidance.

Apply the methodology

Explore licenses and deployment paths.

The two explorers apply these rules across all 32 curated model passports.