Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
- Type
- other
- Venue
- arXiv / Meta Superintelligence Labs / University of Oxford
Summary
CaT aggregates G=8 policy rollouts into a pseudo-reference (synthesis via a frozen anchor) then derives rewards: answer-match in verifiable domains, self-proposed binary rubrics scored by an LLM judge in non-verifiable ones. On HealthBench, relative gains up to +30% vs the initial policy and self-proposed rubrics match physician rubrics; on MATH-500 up to +33%. Trained policy matches or exceeds inference-time synthesis at 9× less test-time compute. Gemma 3 4B / Qwen 3 4B / Llama 3.1 8B. No official code on abs.
Keywords
cat · compute-as-teacher · grpo · rubrics · healthbench · math-500 · icml · meta · oxford
Topics
RL post-training, reference-free supervision, non-verifiable rewards
Research notes
- Primary: arxiv abs (cs.LG). License not stated on abs/HTML at check. Comment: Published as a conference paper at ICML 2026. 23 pages, 6 figures, 12 tables. Work done at Meta Superintelligence Labs; affiliations Oxford / ELLIS Tübingen / MPI-IS / Anthropic / Meta. Correspondence dulhan@robots.ox.ac.uk, agi@meta.com, alanschelten@meta.com. No official code on abs. HF paper page 6 upvotes; no linked models/datasets. Discord posted abs. Uses public HealthBench/MATH-500; no new corpus, so no datasets_local row. License field left blank per catalog convention.