Webscale-RL
- Type
- dataset
- Venue
- Salesforce AI Research
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- English
- Added
- 2026-07-17T20:18:03.665935+00:00
- Verified
- 2026-07-17T20:18:03.665935+00:00
Summary
Webscale-RL is a large-scale reinforcement learning dataset containing approximately 1.2 million verifiable question-answer pairs across more than 9 domains, created by an automated pipeline that converts large-scale pre-training documents into RL-ready data. It was released by Salesforce AI Research alongside the paper 'Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels' (arXiv:2510.06499). Experiments show that models trained on this dataset achieve performance comparable to continual pre-training with up to 100x fewer tokens.
Keywords
reinforcement-learning qa-pairs llm-training rl salesforce synthetic pretraining
Topics
NLP / Reinforcement Learning
Research notes
- The HuggingFace page returned HTTP 401 (gated access). Information sourced from the arXiv paper (2510.06499) and the GitHub repo (SalesforceAIResearch/PretrainRL-pipeline). The dataset was generated using GPT and the authors note it should not be used to develop models that compete with OpenAI. Stack-v2 derived data was excluded due to license issues.