Open weight means the trained parameters are obtainable. Open Source AI is a broader claim about the freedoms to use, study, modify and share an AI system, with access to the preferred form for making modifications. Proprietary or API-only models may expose capabilities through a service without distributing the underlying weights. These categories overlap imperfectly, so license and artifact access must be checked separately.
Artifact access
Are weights, code, training information and checkpoints available?
Legal freedoms
What may you use, study, modify and share under the governing terms?
Operational access
Can you run the model yourself or only call a provider service?
Reproducibility
Can an independent party meaningfully reproduce or modify the system?
Why the terminology causes confusion
AI models are made of multiple artifacts: architecture definitions, training code, preprocessing logic, datasets or data descriptions, checkpoints, final weights, tokenizers and inference software. A publisher can open some of these and keep others closed.
That makes a single word such as “open” ambiguous. A downloadable checkpoint may be highly useful for self-hosting even if the training set is not available. Conversely, a project can publish extensive research code but use model terms that restrict some downstream uses.
OpenWeightModels therefore avoids treating “open” as a score. It records concrete facts: are weights available, what is the license, what source material is published, and what deployment options are technically realistic?
What the Open Source AI Definition adds
The Open Source Initiative’s Open Source AI Definition 1.0 describes Open Source AI in terms of freedoms to use, study, modify and share. For machine-learning systems, it also describes the preferred form for making modifications and explicitly distinguishes weights from the broader materials needed to understand and modify an AI system.
This means that open weights alone are not enough to establish Open Source AI under that definition. A model may be extremely useful and widely self-hosted while still failing that broader test.
This distinction is important for procurement and research. “Can we run it ourselves?” and “Can we study and reproduce how it was made?” are different questions.
A practical comparison matrix
| Property | Open-weight model | Open Source AI | API-only / proprietary |
|---|---|---|---|
| Weights obtainable | Usually yes | Expected where weights are part of the system | Usually no |
| Independent inference | Usually possible | Possible where artifacts support it | No, except through provider service |
| Training transparency | Varies | Broader modification materials expected | Usually limited |
| Commercial rights | License-specific | Must align with open-source freedoms | Service terms |
| Fine-tuning | Often technically possible | Permitted within open-source freedoms | Only if provider offers it |
| Regional self-hosting | Potentially | Potentially | Provider-dependent |
Why licenses still matter after weights are downloadable
A file being publicly downloadable does not answer what you may legally do with it. Apache 2.0 and MIT are permissive software licenses with familiar grant structures. Other AI models use community licenses or custom terms with conditions that may relate to attribution, redistribution, acceptable use, scale or competitive use.
For production decisions, record the exact model/checkpoint license. Do not infer rights from the organization name, family name or the fact that a Hugging Face repository is public.
Why the distinction matters operationally
For an infrastructure team, an open-weight model may solve the most important problem: the ability to operate inference on chosen hardware. For a research team, reproducibility and access to training materials may matter more. For a company, licensing and support obligations may dominate the decision.
Using precise language prevents teams from arguing about labels when they actually need different capabilities. A useful architecture review should ask separately about deployment control, modification rights, training transparency, redistribution and service dependencies.
Recommended language for technical documentation
- Say “open-weight” when the claim is specifically about weight availability.
- Say “Open Source AI” only when using a defined standard and the relevant materials satisfy it.
- Say “self-hostable” when discussing deployment capability.
- Say “commercial use permitted under …” when making a license-specific statement.
- Say “source-available” or describe the exact artifacts rather than inventing a vague openness tier.
Frequently asked questions
Can a model be open weight but not Open Source AI?
Yes. Weight availability is narrower than the Open Source AI Definition and does not by itself establish access to all materials or freedoms required by that definition.
Does Apache 2.0 make a model Open Source AI?
A permissive license on a checkpoint is an important signal, but Open Source AI is a claim about the broader system and preferred form for modification, not just the license name on one artifact.
Are proprietary APIs always closed models?
A hosted service can expose a model without distributing weights. Some providers may offer both an API and downloadable weights, so the deployment channel should be checked per release.
Why does OpenWeightModels avoid an openness score?
Because artifact access, legal rights, reproducibility and deployment control are distinct dimensions. A single score hides the trade-offs teams actually need to evaluate.
Apply this knowledge
Move from the concept to a concrete deployment shortlist.
Primary sources and technical references
OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.