← Back to explorer

SPICE: Self-Play In Corpus Environments Improves Reasoning

Type
other
Venue
arXiv / Meta FAIR

Summary

One model plays Challenger (mines a raw doc, writes MCQ or typed free-form QA with document-extracted gold) and Reasoner (solves without the doc). Challenger reward is a Gaussian on answer-variance peaked at 50% pass; Reasoner gets binary correctness (DrGRPO, role-specific advantages). 20k docs from Nemotron-CC-Math + NaturalReasoning. Qwen3-4B-Base 35.8→44.9 (+9.1); Qwen3-8B 43.0→48.7; OctoThinker-3B 14.7→25.2; OctoThinker-8B 20.5→32.4, beating R-Zero and Absolute Zero. Math avg +8.9 / general +9.8. Fixed-Reasoner pass 55%→35% as the Challenger hardens.

Keywords

spice · self-play · corpus-grounding · rlvr · fair · meta · challenger-reasoner

Topics

self-play, reasoning, RLVR

Research notes

  • Primary: arxiv abs (cs.CL). License not stated on abs/HTML at check. FAIR at Meta / NUS. Joint second Jin/Kim; joint last Lanchantin/Weston; work done at Meta. Correspondence benjaminliu.eecs@gmail.com, {jacklanchantin, jase}@meta.com. No official code on abs. HF paper page 18 upvotes; no linked models/datasets. Discord posted abs. Uses existing Nemotron-CC-Math and NaturalReasoning slices, not a new public release, so no datasets_local row. License field left blank per catalog convention.