Self-Play Search Distillation for Large Language Model Reasoning
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- x
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
SPSD generates superhuman synthetic reasoning data via self-play of MuZero-like networks trained on board games. Executable environments turn search into structured reasoning problems: at each state the expert records a preferred decision, plausible alternatives, plausible opponent replies, and value estimates; the self-play search records are converted into superhuman chains-of-thought for environment-grounded LLM supervision.
Keywords
synthetic data · MuZero · self-play · reasoning · distillation
Topics
synthetic data, MuZero, self-play, reasoning, distillation
Research notes
- Discovery: Shared in #random-papers as an X post by Lorenzo Molfetta; resolved to arXiv 2609.30936.
- Method: Self-play of MuZero-like networks on board games; per-state expert annotations (preferred move, alternatives, opponent replies, value estimates); search-record-to-chain-of-thought conversion; supervised training of LLMs on the resulting records.
- Key findings: Although trained only on self-play search records, SPSD transfers to unseen mathematics: on Qwen3-4B-Base it raises the mean over six math benchmarks from 24.1 to 36.6 and increases the held-out-game win rate from 15% to 45%.
- Authors span University of Bologna, University of Edinburgh, Huawei R&D UK, and Miniml.AI.