← Back to explorer

RLVR-MATH

Type
dataset
Venue
allenai (HuggingFace)
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.768351+00:00
Verified
2026-07-17T20:18:03.768351+00:00

Summary

RLVR-MATH is a 7,500-sample dataset consisting of the MATH training set reformatted for reinforcement learning with verifiable rewards (RLVR), with each example containing a math problem (as messages), a ground-truth answer string, and the source dataset label. It was used to train the final Tulu 3 models with RL as part of AI2's open-instruct ecosystem, enabling rule-based reward verification of model-generated math solutions.

Keywords

math reasoning rlvr reinforcement-learning verifiable-rewards tulu-3 allenai

Topics

Math / Reasoning

Research notes

  • Part of the Tulu 3 RLVR training data. The companion dataset allenai/RLVR-GSM-MATH-IF-Mixed-Constraints combines this with GSM8k (7,473 samples) and IF prompts (14,973 samples). MATH subset is MIT licensed.