← Back to explorer

First Sparks of Thermodynamic Recursive Intelligence

Type
article
Venue
Extropic blog, Oct 1, 2026
Year
2026
Source
blog
Access
free
Language
English
Added
2026-10-01
Verified
2026-10-01

Summary

Extropic post-trains Qwen3.6-35B-A3B with GRPO to reproduce classic connectionist experiments (Boltzmann machines, wake-sleep) as the first step toward Thermodynamic Recursive Self-Improvement: held-out reward nearly triples (0.127 to 0.361) after 100 steps, competitive with much larger frontier models. Next steps: TSU-based sandboxes and real-chip feedback from thermodynamic chips coming online in 2027.

Keywords

thermodynamic computing · recursive self-improvement · GRPO · Qwen3.6 · Boltzmann machines · wake-sleep · Extropic · TSU

Topics

thermodynamic computing, recursive self-improvement, post-training, GRPO

Research notes

  • Discovery: @extropic X post 2026-10-01 (https://x.com/extropic/status/2105699687739343120)
  • Authors: Alexander Neagoe, Guillaume Verdon, Seth Morton (Extropic, San Francisco); post-training platform Prime Intellect
  • Post-trained Qwen3.6-35B-A3B (35B total, 3B active) with GRPO on ~50 coding tasks adapted from open-source reproductions of classic Hinton-tradition experiments (Boltzmann machines, wake-sleep); Nemotron 3 Super 120B as LLM judge; reward r = 0.7*exec + 0.3*rubric; held-out reward 0.127 -> 0.361 after 100 steps (~3x base), competitive with much larger frontier models
  • Framed as first step toward Thermodynamic Recursive Self-Improvement (Thermo RSI): agents that discover thermo algorithms for Extropic's thermodynamic sampling units (TSUs); next: TSU sandboxes, real-chip feedback when first large-scale chips come online in 2027
  • Resources: THRML (docs.thrml.ai), simulator API (extropic.dev), task data on GitHub (link in post), Torx/Thermalizers arXiv papers
  • No arXiv/HF release for this post itself; license not stated