← Back to explorer

LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Type
other
Venue
arXiv / Fudan University / Shanghai AI Laboratory / Stanford

Summary

SR-MCTS treats a full solution as a state and Self-Refine (critique+rewrite) as the action; PPRM (Gemma2-2B DPO on 7.78M PRM800K/OpenMathInstruct pairs) predicts pairwise prefs; Enhanced Borda Count plus Floyd-Warshall lifts them to global quantiles. Untrained LLaMA-3.1-8B-Instruct at 16 rollouts: GSM8K 96.1 rm, MATH 75.3, AIME24 8/30 (26.7) vs 2/30 greedy, AMC23 54.2, GPQA-Diamond 92.4. Beats ToT/rStar at fewer rollouts. NAACL 2025. No official code on abs.

Keywords

llama-berry · mcts · self-refine · pprm · ebc · aime · olympiad-math · naacl · shanghai-ai-lab

Topics

mathematical reasoning, MCTS, reward models

Research notes

  • Primary: arxiv abs (cs.AI; also cs.CL). License CC BY 4.0 on HTML at check. Equal contrib Zhang/Wu/Lei; corresponding Yuqiang Li and Dongzhan Zhou (zhoudongzhan/liyuqiang @pjlab.org.cn). Fudan / Shanghai AI Lab / UC Merced / HK PolyU / UNSW / SJTU / Stanford (Pavone). No official code URL on abs. Related later repo SimpleBerry/LLaMA-O1 not treated as abs-linked. HF paper page 53 upvotes; unofficial/follow-up linked models (SimpleBerry/LLaMA-O1-Supervised-1129, di-zhang-fdu/PPRM-gemma-2-2b-it, etc.) not copied into hf_* fields. Linked HF sets (pprm_math_preference, OpenLongCoT-*, AIME_1983_2024) are preference/CoT derivatives of public MATH/GSM8K/AIME, not a new hosted pretraining corpus, so no datasets_local row. Discord posted abs.