First Sparks of Thermodynamic Recursive Intelligence
- Type
- article
- Venue
- Extropic blog, Oct 1, 2026
- Year
- 2026
- Source
- blog
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Extropic post-trains Qwen3.6-35B-A3B with GRPO to reproduce classic connectionist experiments (Boltzmann machines, wake-sleep) as the first step toward Thermodynamic Recursive Self-Improvement: held-out reward nearly triples (0.127 to 0.361) after 100 steps, competitive with much larger frontier models. Next steps: TSU-based sandboxes and real-chip feedback from thermodynamic chips coming online in 2027.
Keywords
thermodynamic computing · recursive self-improvement · GRPO · Qwen3.6 · Boltzmann machines · wake-sleep · Extropic · TSU
Topics
thermodynamic computing, recursive self-improvement, post-training, GRPO
Research notes
- Discovery: @extropic X post 2026-10-01 (https://x.com/extropic/status/2105699687739343120)
- Authors: Alexander Neagoe, Guillaume Verdon, Seth Morton (Extropic, San Francisco); post-training platform Prime Intellect
- Post-trained Qwen3.6-35B-A3B (35B total, 3B active) with GRPO on ~50 coding tasks adapted from open-source reproductions of classic Hinton-tradition experiments (Boltzmann machines, wake-sleep); Nemotron 3 Super 120B as LLM judge; reward r = 0.7*exec + 0.3*rubric; held-out reward 0.127 -> 0.361 after 100 steps (~3x base), competitive with much larger frontier models
- Framed as first step toward Thermodynamic Recursive Self-Improvement (Thermo RSI): agents that discover thermo algorithms for Extropic's thermodynamic sampling units (TSUs); next: TSU sandboxes, real-chip feedback when first large-scale chips come online in 2027
- Resources: THRML (docs.thrml.ai), simulator API (extropic.dev), task data on GitHub (link in post), Torx/Thermalizers arXiv papers
- No arXiv/HF release for this post itself; license not stated