Detailed OWM Model Reference · verified 2026-09-27

gpt-oss-120b

OpenAI · Mixture-of-Experts Transformer · Apache 2.0 + gpt-oss usage policy

Direct answer: gpt-oss-120b is the more infrastructure-oriented member of OpenAI’s open-weight family. OWM sees it as strategically important because it brings a much larger reasoning model into a deployment envelope that OpenAI says can fit on a single 80 GB GPU, while still preserving the operator’s ability to choose the serving stack.

117B total · 5.1B active/token 128K native context Apache 2.0 + gpt-oss usage policy Text → text
OWM View

gpt-oss-120b: what matters beyond the model card

For organizations already operating data-center GPUs, gpt-oss-120b makes the open-weight decision more consequential than a small local model. It can sit inside private infrastructure, be exposed through an internal API and remain decoupled from a single external model provider. That is exactly the type of separation between model and service layer that OWM considers important for long-term AI optionality.

Open Weight Models editorial view

gpt-oss-120b is the more infrastructure-oriented member of OpenAI’s open-weight family. OWM sees it as strategically important because it brings a much larger reasoning model into a deployment envelope that OpenAI says can fit on a single 80 GB GPU, while still preserving the operator’s ability to choose the serving stack.

Model facts

DeveloperOpenAI
Exact model IDopenai/gpt-oss-120b
Parameters117B total · 5.1B active/token
Context128K native
ArchitectureMixture-of-Experts Transformer
ModalitiesText → text
LicenseApache 2.0 + gpt-oss usage policy
AccessDirect open-weight download

Runtime paths recorded by OWM: Transformers · vLLM · llama.cpp · Ollama. Runtime support is version-sensitive; a named runtime should not be read as a guarantee that every quantization, context size or feature works identically.

Why this model matters

For organizations already operating data-center GPUs, gpt-oss-120b makes the open-weight decision more consequential than a small local model. It can sit inside private infrastructure, be exposed through an internal API and remain decoupled from a single external model provider. That is exactly the type of separation between model and service layer that OWM considers important for long-term AI optionality.

OWM evaluates a model as infrastructure, not only as a benchmark entry. That means the exact checkpoint, license, runtime ecosystem, memory footprint, ability to move between providers and the quality of the evidence all matter alongside model capability.

Hardware reality

OpenAI states that gpt-oss-120b can run efficiently on a single 80 GB GPU. That statement is especially notable for a 117B-parameter model because the MoE architecture activates only 5.1B parameters per token. Storage, loading, context length, KV cache and runtime implementation still affect real deployment requirements.

OWM deliberately separates raw-weight arithmetic, publisher guidance and measured runtime evidence. A model that theoretically fits into a memory budget can still fail in practice because of KV cache, runtime buffers, vision components, tensor-parallel overhead or concurrent requests.

License reality

The same Apache 2.0 plus usage-policy structure applies as with gpt-oss-20b. OWM considers the permissive software license a major deployment advantage, while still recording the policy layer explicitly so organizations do not confuse “permissive license” with “no conditions anywhere.”

This is an informational deployment summary, not legal advice. Production users should review the exact current license, usage policy, derivative-model terms and applicable law before shipping a product.

OWM Sovereignty Lens

OWM does not assign a single sovereignty score. We describe the layers separately because a model can be highly portable technically while remaining conditional legally — or permissively licensed while requiring infrastructure that limits practical choice.

Weight control
Strong
The checkpoint can be obtained and retained outside an OpenAI-hosted API.
License freedom
Strong with policy layer
Apache 2.0 is permissive; the usage policy remains part of the operational review.
Deployment control
Strong
The model is intended for self-hosted and provider-hosted deployment.
Runtime portability
Strong
Multiple mainstream inference stacks support the family.
Exit capability
Strong
The model layer can survive a change of infrastructure provider.

Runtime evidence

Evidence status

OWM currently records publisher-stated runtime support for the family. Independent or OWM-reproduced 120B deployment measurements should be added only with exact hardware, runtime version and workload.

“Third-party measured” means the result was measured outside Open Weight Models and is shown with provenance. “OWM runtime tested” is reserved for configurations that OWM physically reproduces with an exact checkpoint, runtime version, hardware configuration, workload and date.

Change history

Initial snapshot

First OWM verification snapshot created. From this date forward, material changes to the model card, license, checkpoints, runtime support and deployment facts can be appended without reconstructing unobserved history.

Open the global OWM Change History →

Where gpt-oss-120b fits — and where it does not

Where it fits

  • Private enterprise inference on H100/H200-class or other 80 GB accelerator infrastructure.
  • Organizations that want a larger reasoning model while keeping model weights under their own operational control.
  • Internal agent platforms where data location, provider exit and runtime choice matter.
  • Teams comparing self-hosted inference economics with premium managed APIs.

Where it does not fit

  • Consumer laptops or typical single consumer GPUs without aggressive offload strategies.
  • Teams whose inference volume is too low to justify operating an 80 GB-class accelerator.
  • Multimodal workloads requiring native vision or audio.

Open-weight significance

The strategic value of this model is not simply that its weights can be downloaded. The important question is what those weights let an operator control: infrastructure, data location, runtime, adaptation and the ability to exit a provider relationship without discarding the model layer. Those freedoms remain bounded by the model’s license and the practical hardware required to run it.

Frequently asked questions

How much GPU memory does gpt-oss-120b need?

OpenAI says it can run efficiently on a single 80 GB GPU. Production needs can rise with long contexts, concurrency and runtime overhead.

How many parameters does gpt-oss-120b have?

OpenAI documents 117B total parameters with 5.1B active parameters per token.

Does gpt-oss-120b have a 128K context window?

Yes. OpenAI documents native context support up to 128K tokens.

Can an enterprise self-host gpt-oss-120b?

Yes in principle; the weights are available for self-hosting, subject to license, policy, infrastructure and operational requirements.

Primary sources and OWM data

Last verified by Open Weight Models: 2026-09-27. Facts can change as model repositories, licenses and runtime support evolve.

Related model references