← Back to explorer

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator

Type
paper
Venue
arXiv / MBZUAI
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T17:02:26Z
Verified
2026-08-14T17:02:26Z

Summary

Interprets a TRM step as mapping a pre-reasoning policy to a post-reasoning policy whose log-ratio is an advantage-like score. Deep Improvement Supervision (DIS) trains each of N_sup=6 steps toward a less-corrupted discrete-masking target, dropping ACT/halting (T=1, n=2 vs TRM T=3, n=6, 16 steps; ~18x fewer forwards). DIS-compact 0.8M: ARC-AGI-1 24% vs TRM-compact 12%. DIS 7M: ARC-1 41.3 / ARC-2 6.0 vs reported TRM 40.4 / 3.3. Matched 0.69 on N-Queens. No official code on the abs page.

Keywords

trm · dis · arc-agi · latent-reasoning · policy-improvement · mbzuai

Topics

latent reasoning, tiny recursive models, ARC-AGI

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.LG). MBZUAI; correspondence arip.asadulaev@mbzuai.ac.ae. Discord posted AlphaXiv 2511.16886; canonical abs recorded. HF paper page title is Deep Improvement Supervision (1 upvote); arxiv title used. HF lists 1 citing model and 1 Space; not official paper artifacts, so hf_* left blank. No official code on abs. Uses public ARC-AGI/N-Queens; no new corpus, so no datasets_local row.