Towards Looped Models Done Right — Part I: Topology, Input Injection, Recurrent-State Design
- Type
- other
- Venue
- Institute of Foundation Models
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:25:00Z
- Verified
- 2026-08-14T19:25:00Z
Summary
IFM living blog (Huang IFM/CMU corresponding). Controlled Ouro-to-Huginn ablations at matched parameter scale, logical depth, and token budget. Three axes: iteration envelope, input injection, latent-state organization. Sandwich prelude–loop–coda helps instance-conditioned reasoning (MATH500 +12.00, DROP +2.61 at 730M / 336B tokens); input injection helps context/spec tasks (BBH-CoT, DROP, code) but can hurt math; random init and shared H/L states are mixed or negative. Full Huginn beats Ouro on all ten dense benches at 730M. MoE transfer on TxT360: 8.0B resident / 0.8B active Huginn MoE vs Ouro vs a 32B-A3.2B feedforward MoE, all 500B tokens, matched train+infer FLOPs. Huginn MoE GSM8K 83.6% vs feedforward 80.8% with 75% fewer resident parameters; more even expert load, and forcing later loops to reuse iter-1 experts hurts. Code “release soon.”
Keywords
loop-models · huginn · ouro · moe · ifm · mbzuai · blog · x
Topics
looped transformers, MoE, architecture ablations
Research notes
- Primary: IFM Notion living blog. Discord/X https://x.com/huskydogewoof/status/2083242247945126203 via fxtwitter. Correspondence benhaoh@andrew.cmu.edu. Affiliations IFM/USC/CMU. Code not released at check. Distinct from papers_local 513 (LT2) and 525 (DiffusionBlocks mentions Huginn). Blog/ablations, not a hosted corpus, so no datasets_local row.