Detailed OWM Model Reference · verified 2026-09-27

Kimi K2 Instruct

Moonshot AI · Mixture-of-Experts · 384 experts · 8 selected · Modified MIT

Direct answer: Kimi K2 Instruct shows that open-weight distribution can extend to trillion-parameter-class MoE models. OWM sees it less as a local model and more as a model-layer portability asset for infrastructure providers and large organizations. Its 32B active parameter count improves compute efficiency, but the 1T total checkpoint keeps storage and distributed serving firmly in data-center territory.

1T total · 32B active128K contextModified MITText → text
OWM View

Kimi K2 Instruct: what matters beyond the model card

Moonshot documents 1T total parameters, 32B active parameters, 384 experts with eight selected per token and 128K context. The model is optimized for agentic capabilities. For OWM, Kimi K2 demonstrates that sovereignty can mean the ability to choose among infrastructure providers even when self-hosting requires substantial clusters.

Open Weight Models editorial view

Kimi K2 Instruct shows that open-weight distribution can extend to trillion-parameter-class MoE models. OWM sees it less as a local model and more as a model-layer portability asset for infrastructure providers and large organizations. Its 32B active parameter count improves compute efficiency, but the 1T total checkpoint keeps storage and distributed serving firmly in data-center territory.

Model facts

DeveloperMoonshot AI
Exact model IDmoonshotai/Kimi-K2-Instruct
Parameters1T total · 32B active
Context128K
ArchitectureMixture-of-Experts · 384 experts · 8 selected
ModalitiesText → text
LicenseModified MIT
Verified2026-09-27

Runtime paths recorded by OWM: Transformers with custom code · provider runtimes · distributed inference stacks. Support is version-sensitive and does not imply identical feature parity across runtimes.

Why this model matters

Moonshot documents 1T total parameters, 32B active parameters, 384 experts with eight selected per token and 128K context. The model is optimized for agentic capabilities. For OWM, Kimi K2 demonstrates that sovereignty can mean the ability to choose among infrastructure providers even when self-hosting requires substantial clusters.

OWM evaluates a checkpoint as infrastructure: exact weights, license, runtime portability, memory reality, evidence quality and provider exit all matter alongside capability.

Hardware reality

The total checkpoint is enormous. Active parameters do not eliminate the need to store and distribute the broader expert set. Practical deployment therefore depends on high-capacity storage, accelerator memory, interconnect and a distributed serving stack. It should not be described as a 32B-memory model.

OWM separates publisher guidance, engineering estimates and measured runtime evidence. Memory arithmetic alone does not include every KV-cache, runtime, multimodal or concurrency cost.

License reality

The repository identifies a modified MIT license. OWM intentionally labels it “Modified MIT” rather than plain MIT because custom provisions should be reviewed directly before commercial deployment.

This is an informational deployment summary, not legal advice. Always review the exact current license and policies before production use.

OWM Sovereignty Lens

OWM does not collapse sovereignty into one score. Technical portability and legal freedom can differ substantially.

Weight control
Strong
The checkpoint is publicly distributed.
License freedom
Conditional
The license is modified MIT, not identical to standard MIT.
Deployment control
Strong in principle, infrastructure-heavy in practice
Multiple hosting choices exist, but all require substantial capacity.
Runtime portability
Moderate
Custom code and large-scale serving requirements narrow practical portability.
Exit capability
Strong at model layer
Large operators can retain the checkpoint while moving between infrastructure vendors.

Runtime evidence

Evidence status

OWM currently records architecture and source-verified deployment information but does not present a single universal throughput figure for a 1T MoE model whose performance is topology-dependent.

“OWM runtime tested” remains reserved for configurations physically reproduced by the project with exact hardware, runtime version, workload and date.

Change history

Initial snapshot

First OWM verification snapshot created. From this date forward, material changes to the model card, license, checkpoints, runtime support and deployment facts can be appended without reconstructing unobserved history.

Open the global OWM Change History →

Where Kimi K2 Instruct fits — and where it does not

Where it fits

  • Large-scale provider or enterprise agentic inference.
  • Organizations that want a portable trillion-parameter checkpoint rather than a single proprietary API.
  • Distributed infrastructure teams experienced with MoE serving.
  • Research into large sparse agentic models.

Where it does not fit

  • Single-node consumer deployment.
  • Teams interpreting 32B active parameters as 32B checkpoint storage.
  • Organizations unwilling to review the modified license carefully.

Open-weight significance

The strategic value is not simply that weights can be downloaded. The relevant question is what the operator can control: infrastructure, data location, runtime, adaptation and provider exit — all bounded by the license and practical hardware requirements.

Frequently asked questions

How large is Kimi K2 Instruct?

Moonshot documents 1T total parameters and 32B activated parameters.

What is the context length?

The model card lists 128K.

Is the license standard MIT?

No. The repository labels it modified MIT, so the exact terms should be reviewed.

Can Kimi K2 run on a single consumer GPU?

No realistic full-checkpoint deployment fits that description; it is a distributed-infrastructure model.

Primary sources and OWM data

Last verified by Open Weight Models: 2026-09-27. Facts can change as model repositories, licenses and runtime support evolve.

Related model references