Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards
- Type
- paper
- Venue
- arXiv / UIUC
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:36:00Z
- Verified
- 2026-08-14T16:36:00Z
Summary
Shows prior 100%-noisy RLVR sets were contaminated: a GPT-5 Pro + math-verify + LLM-judge + manual pipeline found 16.4% of DeepScaleR labels marked incorrect were actually correct. After sanitizing to 12,769 truly incorrect items, Qwen2.5-Math-7B GRPO on 100% noise is 8–10% worse than clean data (9% on MATH-500) and no better than format-only rewards; random labels fall below the base model. Dr.GRPO, TIS, PGFC, DAPO, and SAPO fail to beat GRPO under 50% noise. On a manually corrected 600-example BIRD subset (372/600 originally noisy), real annotation errors cost 5–12% vs the cleaned set. Code/data https://github.com/uiuc-kang-lab/rlvr-noisy-data.
Keywords
rlvr · grpo · noisy-labels · deepscaler · bird · text2sql · uiuc
Topics
RLVR, data quality, mathematical reasoning, Text2SQL
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG). UIUC; Kang also Bridgewater AIA Labs. Code https://github.com/uiuc-kang-lab/rlvr-noisy-data. HF uiuc-kang-lab/DeepScaleR-Qwen2.5-Math-7B-Incorrect-Answers (12,769, 10K<n<100K, downloads=27 likes=0, lastModified 2026-04-07) plus Falsely-Incorrect-Answers (2,498). HF cards do not state a license. Also added to datasets_local.csv. Discord posted abs link.