Probabilistic Tiny Recursive Model
- Type
- other
- Venue
- arXiv / Mila / ILLS & ETS Montreal
Summary
PTRM injects Gaussian noise into the TRM latent at each deep recursion step, runs K parallel rollouts, and selects the candidate with highest Q-head score (the ACT correctness classifier unused at standard TRM inference). No retraining or task-specific test-time augmentation. PPBench golden 62.6%→91.2% (K=100, D=48, σ=0.2), beating a 7-LLM ensemble with oracle verifier (55.1%) at ~$0.001 vs $38.51 per correct. Sudoku-Extreme 87.4%→98.75%; Maze-Hard 83.8%→86.73%; ARC-AGI-2 pass@1 7.36%→8.47%. Width scaling dominates depth; Q nearly matches pass@K on PPBench/Sudoku but lags on Maze-Hard.
Keywords
ptrm · trm · recursive-reasoning · test-time-scaling · q-head · ppbench · sudoku · arc-agi · mila
Topics
recursive reasoning, test-time scaling, TRM
Research notes
- Primary: arxiv abs (cs.AI). CC BY 4.0 on HTML. Sghaier Mila/ILLS & ETS Montreal, Parviz Mila, Jolicoeur-Martineau Independent. Correspondence amin.sghaier/ali.parviz@mila.quebec, alexia.jolicoeur-martineau@mail.mcgill.ca. No official code on abs. HF paper page 0 upvotes; no linked models/datasets. Discord posted AlphaXiv. Uses public PPBench/Sudoku-Extreme/Maze-Hard/ARC-AGI-2; no new corpus, so no datasets_local row. License field left blank per catalog convention.