← Back to explorer

K2 Horizon: six fully open models from 0.9B to 375B (IFM)

Type
model
Venue
IFM blog / HuggingFace
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

The Institute of Foundation Models (frontier lab of Abu Dhabi's MBZUAI) released K2 Horizon, a fleet of six fully open models (0.9B, 3.7B, 7B, 32B dense; 36B-A4B and 375B-A23B MoE) for reasoning, coding and agentic work, under Apache 2.0 with weights, training code, intermediate checkpoints, data and recipes published. The 36B-A4B introduces Mixture-of-Value Attention (MoVA); Uno freezes autoregressive params and trains diffusion params for parallel block emission. Tool definitions were presented in JSON, XML and Markdown during training.

Keywords

open-models · moe · long-context · agentic · code · reasoning

Topics

open-models, moe, long-context, agentic, code

Research notes

  • Also: https://x.com/IFM_AI/status/2099925030570455369, https://huggingface.co/IFM/K2-Horizon-3.7B
  • Discovery: Posted to #random-papers 2026-09-03 (blog), 2026-09-15 (IFM X post + K2-Horizon-3.7B HF page)
  • Method: Shared core architecture/vocabulary/training methodology across all sizes; 3.7B trained on 22.9T tokens with 4 midtraining stages extending context 32K to 512K, then branched RL (math/code/STEM-code experts merged), then two SFT phases at 512K.
  • Key findings: 3.7B: 68.6 SWE-bench Verified, 65.4 GPQA Diamond, 512K native context; all intermediate checkpoints released; 36B-A4B: 80.8% GPQA Diamond, 58.6% Terminal-Bench 2.1; 375B-A23B: 87.3 GPQA Diamond, 70.2 Terminal-Bench 2.1 (audited to 66.9); Full disclosure package (data, checkpoints, logs) goes well beyond typical open-weight releases
  • Limitations: Benchmark figures are company-reported; independent reproduction at 20T-token scale is out of reach for most researchers despite open artifacts. Technical report not yet released (expected end of September 2026).
  • Merged from 3 Discord posts (blog 9/3, X post 9/15, HF 3.7B page 9/15). Blog page itself was rate-limited during research; substance verified via multiple secondary reports.