GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
- Type
- other
- Venue
- arXiv / Apple
Summary
100 GSM8K-test questions become symbolic templates (names, numbers, conditions); 50 instantiations each (5,000 eval items) plus harder P1/P2 clause variants and GSM-NoOp (irrelevant but topical clauses). Across 25 open/closed models, accuracy is a wide distribution; GSM8K often sits >1σ above the GSM-Symbolic mean. Models are more robust to name swaps than number swaps. Difficulty M1→P2 shifts the mean down and variance up. GSM-NoOp drops SOTA models up to 65% (Phi-3-mini) even with 8 shots of the same question. Hypothesis: pattern-matching, not formal reasoning. ICLR 2025. Data https://github.com/apple/ml-gsm-symbolic and HF apple/GSM-Symbolic.
Keywords
gsm-symbolic · gsm8k · mathematical-reasoning · apple · iclr · pattern-matching · noop · evaluation
Topics
mathematical reasoning, LLM evaluation, GSM8K
Research notes
- Primary: arxiv abs (cs.LG; also cs.AI). License: arXiv.org perpetual non-exclusive on HTML at check. Mirzadeh/Alizadeh/Tuzel/Bengio/Farajtabar Apple; Shahrokhi Washington State University (internship). Correspondence {imirzadeh,farajtabar}@apple.com. Code/data on HTML https://github.com/apple/ml-gsm-symbolic (HF githubRepo; 90 stars at check; GitHub license NOASSERTION). Project https://machinelearning.apple.com/research/gsm-symbolic. Official HF apple/GSM-Symbolic (CC BY-NC-ND 4.0; 12,500 rows; 8,590 downloads / 24 likes at check) — see datasets_local. HF paper page 22 upvotes; unofficial linked sets vojewale/Cross-lingualGSMSymbolic and WT-solutions/BG-GSM-Synthetic not copied into hf_* fields. Discord posted PDF. License field left blank per catalog convention (paper arXiv default; dataset license lives on the datasets_local row).