Learning Adaptive Parallel Reasoning with Language Models
- Type
- other
- Venue
- arXiv / UC Berkeley / UCSF
Summary
APR adds spawn()/join() parent-child threads on SGLang, SFT on hybrid symbolic search traces then end-to-end GRPO. On Countdown vs serialized SoS+: 83.4% vs 60.0% at 4k context, 80.1% vs 66.6% at 20k total tokens, 75.2% vs 57.3% at ~5s latency. RL mainly widens search (6.1→8.2 child threads). 228M Llama-2-arch trained from scratch on 500k traces. Accepted at COLM 2025. Code https://github.com/Parallel-Reasoning/APR.
Keywords
apr · parallel-reasoning · spawn-join · grpo · countdown · colm · berkeley · sglang
Topics
parallel reasoning, inference-time compute, reinforcement learning
Research notes
- Primary: arxiv abs (cs.AI; also cs.CL). License not stated on abs/HTML at check. Equal contrib Pan/Li/Lian. UC Berkeley / UCSF. Code https://github.com/Parallel-Reasoning/APR (145 stars at check). HF paper page 44 upvotes; githubRepo linked; 1 unofficial linked model (Parallel-Reasoning/llama-sosp) not copied into hf_* fields. Discord posted abs. Countdown traces are generated by a symbolic solver, not a substantial hosted corpus, so no datasets_local row. License field left blank per catalog convention.