← Back to explorer

Self-Play Search Distillation for Large Language Model Reasoning

Type
paper
Venue
arXiv
Year
2026
Source
x
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

SPSD generates superhuman synthetic reasoning data via self-play of MuZero-like networks trained on board games. Executable environments turn search into structured reasoning problems: at each state the expert records a preferred decision, plausible alternatives, plausible opponent replies, and value estimates; the self-play search records are converted into superhuman chains-of-thought for environment-grounded LLM supervision.

Keywords

synthetic data · MuZero · self-play · reasoning · distillation

Topics

synthetic data, MuZero, self-play, reasoning, distillation

Research notes

  • Discovery: Shared in #random-papers as an X post by Lorenzo Molfetta; resolved to arXiv 2609.30936.
  • Method: Self-play of MuZero-like networks on board games; per-state expert annotations (preferred move, alternatives, opponent replies, value estimates); search-record-to-chain-of-thought conversion; supervised training of LLMs on the resulting records.
  • Key findings: Although trained only on self-play search records, SPSD transfers to unseen mathematics: on Qwen3-4B-Base it raises the mean over six math benchmarks from 24.1 to 36.6 and increases the held-out-game win rate from 15% to 45%.
  • Authors span University of Bologna, University of Edinburgh, Huawei R&D UK, and Miniml.AI.