Note #23 •

K2 Horizon: The Largest Fully Open AI Model Fleet in History

On September 3, 2026, MBZUAI's Institute of Foundation Models (IFM) dropped what might be the most important open-source AI release of the year: K2 Horizon, a fleet of six models ranging from 0.9 billion to 375 billion parameters — every single one fully open under Apache 2.0.

Not open-weights. Fully open. Weights, code, training data, methodology, and recipes. The kind of transparency the AI industry keeps promising and rarely delivers.

The Fleet

Six models, one architecture, one license. Here's the lineup:

Model Parameters Active Best For
K2 Horizon 0.9B 0.9B 0.9B Watches, glasses, edge devices — state-of-the-art at its size
K2 Horizon 3.7B 3.7B 3.7B Phones, on-device — best reasoning under 4B
K2 Horizon 7B 7B 7B Best model under 10B — strong coding & research
K2 Horizon 32B 32B 32B Local hosting on laptops and on-prem servers
K2 Horizon 36B 36B 4B (MoVA) Cost-effective local hosting with Mixture of Value Attention
K2 Horizon 375B 375B 23B (MoE) Enterprise reasoning and agentic workloads

Why This Matters

The headline isn't the parameter count. It's the commitment to verifiability.

Most "open" model releases stop at publishing weights. You can run inference, but you can't see the data, reproduce the training, or understand how the model was built. K2 Horizon ships all of it — every model comes with its training data, recipe, and evaluations. As IFM founder Eric Xing put it: "Science works when others can see the data, follow the method, reproduce the result, and improve on it."

That's a direct challenge to the industry's prevailing open-weights-but-closed-everything-else model. OpenAI, Google, and Anthropic publish nothing. Meta and Mistral publish weights but not data. K2 Horizon is the first fleet-level release to clear the transparency bar entirely.

Key Technical Innovations

Diffusion Distillation

A technique that generates blocks of tokens in parallel instead of one-by-one, roughly 3× speedup without degrading response quality. This is significant for production deployments where latency matters — the smallest models on edge hardware suddenly become practical for real-time use.

Mixture of Value Attention (MoVA)

A new attention architecture that improves reasoning without adding computation. The 36B model activates only 4B parameters per forward pass but outperforms many larger dense models. Efficient sparsity done right.

Dynamic Model Routing

IFM's routing system directs each task to the most cost-effective model in the fleet. Developers prototype on the 0.9B model and scale up to the 375B flagship without changing code or deployment workflows.

The Small End Is the Interesting End

The 0.9B model is worth pausing on. It achieves state-of-the-art results in math, reasoning, and tool use at a size that fits on a smartwatch. Not "good for its size" — genuinely competitive at constrained deployment. For edge AI, IoT, and on-device agents, this is the most practical release in the fleet.

The 3.7B and 7B models follow the same pattern: best-in-class at their scale for coding, agentic tasks, and deep research. Fine-tuning-friendly sizes that actually run on consumer hardware.

Availability

All models are on Hugging Face under Apache 2.0. Supported in vLLM and SGLang. API access through Compass, Cerebras, AWS, and Nebius.

What This Means for Open-Source AI

K2 Horizon raises the bar for what "open" means in AI. It's no longer enough to publish weights and call it a day. The fleet is an existence proof that end-to-end transparency is possible at scale — from a 0.9B watch model to a 375B enterprise flagship.

If you're building on open models, this gives you a single architecture to target from edge to data center. If you're evaluating model transparency, K2 Horizon is the new benchmark everyone else has to beat.

The question now isn't whether open-source AI can compete. It's whether closed models can justify staying closed.


Sources: IFM Press Release, MBZUAI Announcement