MiniCPM5-2B
- Type
- model
- Venue
- Hugging Face (openbmb)
- Year
- 2026
- Source
- huggingface
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
MiniCPM5-2B is a dense 2.5B-parameter causal language model (LlamaForCausalLM, 42 layers, 131K context) built for on-device and resource-constrained deployment. OpenBMB claims 2B-class open-source SOTA with strengths in coding, math, long-context understanding, tool use, and agentic tasks, and released the training datasets (UltraX, UltraData-Code, UltraData-SFT-Agent-2609, UltraData-RL-2609) alongside the model.
Keywords
small language models · on-device · open weights · reinforcement learning · distillation · MiniCPM
Topics
small language models, on-device, open weights, reinforcement learning, distillation
Research notes
- Method: Three-stage training: base training, mid-training, then post-training in SFT (400B tokens deep-thinking SFT) → RL (critic-based algorithm) → OPD (On-Policy Distillation merging 16 RL expert models via full-vocabulary reverse KL as advantage estimate).
- Key findings: Self-reported 2B-class SOTA: 53.9 average vs 33.2/28.0/24.6 for LFM2.5-2.6B, Qwen3.5-2B, Gemma-4-E2B-it; AIME 2025/2026 86.5, SWE-bench Verified 46.4, GAIA 88.7 per the model card's own table; Released in many formats: BF16, GGUF, MLX, GPTQ, LiteRT plus a DSpark speculative-decoding draft model
- Limitations: All benchmark figures are self-reported on the model card with a card-selected comparison set; independent verification not checked.
- Release date 2026-09-07 taken from the OpenBMB/MiniCPM repo changelog.