Allspark: Weak to Strong Transfer via Alternating Chain of Thought
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Asks whether reasoning improvements learned by a small, weak model can benefit a larger, stronger model without using the strong model's rollouts during training. Introduces Allspark, a training and inference framework for weak-to-strong transfer through alternating chains of thought: a weak teacher is trained with RL alongside a frozen copy of the same model, the two alternate reasoning segments, and the frozen model produces the final answer. At inference time a stronger student replaces the frozen training partner while both models remain fixed. Because they communicate through text, the teacher can steer students from different model families and with different tokenizers. Studied at two scales: controlled Qwen3-1.7B/4B experiments across math and reasoning domains, and larger-scale experiments with an Inkling-Small teacher on 96 ARC-AGI-2 development problems. The Inkling experiments show accuracy gains in within-family and cross-family settings, including transfer to Kimi K2.6 and Nemotron 3 Ultra, with benefits varying across inference settings; analyses also characterize the accuracy-token tradeoff. Motivates reusing a trained weak teacher across strong students.
Keywords
weak-to-strong transfer · RL · reasoning · chain-of-thought · cross-family transfer · inference
Topics
weak-to-strong transfer, reinforcement learning, chain-of-thought, reasoning
Research notes
- User-supplied arXiv link on 2026-09-29: https://arxiv.org/html/2609.32913v1
- Authors: Kaizhao Liang (UT Austin), Junxiong Wang (Together AI), Chen Liang (Microsoft), Zhendong Wang (Microsoft), Qiang Liu (UT Austin).
- Submitted 2026-09 (v1), cs.AI/cs.CL.
- Weak-to-strong transfer without strong-model rollouts; complements the collection's RL-for-reasoning and distillation entries.