← Back to explorer

Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards

Type
paper
Venue
arXiv / UIUC
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:36:00Z
Verified
2026-08-14T16:36:00Z

Summary

Shows prior 100%-noisy RLVR sets were contaminated: a GPT-5 Pro + math-verify + LLM-judge + manual pipeline found 16.4% of DeepScaleR labels marked incorrect were actually correct. After sanitizing to 12,769 truly incorrect items, Qwen2.5-Math-7B GRPO on 100% noise is 8–10% worse than clean data (9% on MATH-500) and no better than format-only rewards; random labels fall below the base model. Dr.GRPO, TIS, PGFC, DAPO, and SAPO fail to beat GRPO under 50% noise. On a manually corrected 600-example BIRD subset (372/600 originally noisy), real annotation errors cost 5–12% vs the cleaned set. Code/data https://github.com/uiuc-kang-lab/rlvr-noisy-data.

Keywords

rlvr · grpo · noisy-labels · deepscaler · bird · text2sql · uiuc

Topics

RLVR, data quality, mathematical reasoning, Text2SQL

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG). UIUC; Kang also Bridgewater AIA Labs. Code https://github.com/uiuc-kang-lab/rlvr-noisy-data. HF uiuc-kang-lab/DeepScaleR-Qwen2.5-Math-7B-Incorrect-Answers (12,769, 10K<n<100K, downloads=27 likes=0, lastModified 2026-04-07) plus Falsely-Incorrect-Answers (2,498). HF cards do not state a license. Also added to datasets_local.csv. Discord posted abs link.