← Back to explorer

Continuous Thought Machines

Type
other
Venue
arXiv / Sakana AI

Summary

CTM unfolds an internal tick timeline: a synapse MLP produces pre-activations, each neuron has a private MLP over a rolling history, and pairwise post-activation synchronization is the latent used for attention queries and outputs. Native adaptive compute comes from a min-loss + max-certainty loss over ticks, with no separate halt head. On 39x39 mazes with no positional embeddings it traces routes and generalizes to 99x99 by re-application, beating LSTM/FF baselines. ImageNet-1K with a ResNet-152 backbone and 50 ticks: 72.47% top-1 / 89.89% top-5, with emergent looking-around attention. Also parity, sorting, Q&A MNIST, and RL POMDPs. Goal is the architecture, not SOTA. Code https://github.com/SakanaAI/continuous-thought-machines; demo https://pub.sakana.ai/ctm/.

Keywords

ctm · sakana · neural-synchronization · adaptive-compute · maze · imagenet · neurips

Topics

recurrent reasoning, neural dynamics, biologically inspired models

Research notes

  • Primary: arxiv abs (cs.LG/AI). Sakana AI / Tsukuba / ITU Copenhagen; correspondence luke/ciaran/sebastianrisi/jeffrey/llion@sakana.ai. Code Apache-2.0 https://github.com/SakanaAI/continuous-thought-machines (2019 stars at check); project https://pub.sakana.ai/ctm/. HF paper page 15 upvotes, org SakanaAI. Linked official models SakanaAI/ctm-imagenet (186M, 7 likes) and SakanaAI/ctm-maze-large (32M). Discord posted abs. Uses ImageNet/CIFAR/synthetic mazes/parity; no new standalone public corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.