HRM-Text: Efficient Pretraining Beyond Scaling
- Type
- paper
- Venue
- arXiv / Sapient Intelligence
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:20:00Z
- Verified
- 2026-08-14T19:20:00Z
Summary
Sapient Intelligence 1B Hierarchical Recurrent Model for language: slow H-module and fast L-module (2 H cycles x 3 L updates = 8 steps / 4 effective recursions), MagicNorm, warmup deep credit assignment (K=2 to K=5), sigmoid-gated attention, RoPE. Trains only on instruction-response pairs with response-only NLL and PrefixLM bidirectional prompt attention. 40B unique / 60B training tokens, ~$1500, 1.9 days on 16 H100s. Reports MMLU 60.7, ARC-C 81.9, DROP 82.2, GSM8K 84.5, MATH 56.2; competitive with several 2-7B open models at 96-432x less estimated compute and 100-900x fewer tokens. Ablations: HRM beats FLOP-matched Transformer/looped/RINS; task-completion + PrefixLM each move the needle. Discord post is a third-party ablation read (huskydogewoof) of Table 3/4, not the authors.
Keywords
hrm · hrm-text · prefixlm · recurrent · sapient · pretraining · latent-reasoning
Topics
hierarchical recurrent models, efficient pretraining, latent reasoning
Research notes
- Primary: arxiv abs 2605.20613. Discord/X https://x.com/huskydogewoof/status/2057043774996394409 is a review thread linking https://sapientinc.github.io/HRM-Text/assets/HRM_Text.pdf and quoting https://github.com/sapientinc/HRM-Text. Intro blog https://sapient.inc/introducing-hrm-text/. Contact research@sapient.inc. Uses public instruction corpora (FLAN, Tasksource, SYNTH, OpenMathInstruct2, etc.), not a new hosted corpus, so no datasets_local row. Authors Cai Zhou / Chenyu Wang listed with MIT affiliation on abs.