RLVR-MATH
- Type
- dataset
- Venue
- allenai (HuggingFace)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.768351+00:00
- Verified
- 2026-07-17T20:18:03.768351+00:00
Summary
RLVR-MATH is a 7,500-sample dataset consisting of the MATH training set reformatted for reinforcement learning with verifiable rewards (RLVR), with each example containing a math problem (as messages), a ground-truth answer string, and the source dataset label. It was used to train the final Tulu 3 models with RL as part of AI2's open-instruct ecosystem, enabling rule-based reward verification of model-generated math solutions.
Keywords
math reasoning rlvr reinforcement-learning verifiable-rewards tulu-3 allenai
Topics
Math / Reasoning
Research notes
- Part of the Tulu 3 RLVR training data. The companion dataset allenai/RLVR-GSM-MATH-IF-Mixed-Constraints combines this with GSM8k (7,473 samples) and IF prompts (14,973 samples). MATH subset is MIT licensed.