← Back to explorer

Towards Looped Models Done Right — Part I: Topology, Input Injection, Recurrent-State Design

Type
other
Venue
Institute of Foundation Models
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:25:00Z
Verified
2026-08-14T19:25:00Z

Summary

IFM living blog (Huang IFM/CMU corresponding). Controlled Ouro-to-Huginn ablations at matched parameter scale, logical depth, and token budget. Three axes: iteration envelope, input injection, latent-state organization. Sandwich prelude–loop–coda helps instance-conditioned reasoning (MATH500 +12.00, DROP +2.61 at 730M / 336B tokens); input injection helps context/spec tasks (BBH-CoT, DROP, code) but can hurt math; random init and shared H/L states are mixed or negative. Full Huginn beats Ouro on all ten dense benches at 730M. MoE transfer on TxT360: 8.0B resident / 0.8B active Huginn MoE vs Ouro vs a 32B-A3.2B feedforward MoE, all 500B tokens, matched train+infer FLOPs. Huginn MoE GSM8K 83.6% vs feedforward 80.8% with 75% fewer resident parameters; more even expert load, and forcing later loops to reuse iter-1 experts hurts. Code “release soon.”

Keywords

loop-models · huginn · ouro · moe · ifm · mbzuai · blog · x

Topics

looped transformers, MoE, architecture ablations

Research notes

  • Primary: IFM Notion living blog. Discord/X https://x.com/huskydogewoof/status/2083242247945126203 via fxtwitter. Correspondence benhaoh@andrew.cmu.edu. Affiliations IFM/USC/CMU. Code not released at check. Distinct from papers_local 513 (LT2) and 525 (DiffusionBlocks mentions Huginn). Blog/ablations, not a hosted corpus, so no datasets_local row.