ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Proposes ReScraper, a unified 0.6B-parameter language model that replaces the entire heuristic pretraining-corpus cleaning stack (HTML scraper + dozens of rule-based filters). ReScraper first extracts the main content from raw pages, then chooses among four operations: keep as extracted, edit out noisy lines/spans, delete entirely, or rewrite when poorly written but informative. Trained on supervised data curated from three teacher models. On the same crawled data pool, pretraining 400M, 1.4B, and 2.8B models on the curated data improves DCLM Core score by a relative 3.8-4.7% over the strongest baseline at each scale, including costly multi-agent curation. Each operation plays a distinct, complementary role; one-model extraction+cleaning beats a cascade of separate models; ReScraper concentrates edits on the pages that need them, raising poor-page quality while keeping corpus diversity.
Keywords
data curation · pretraining · web scraping · HTML extraction · AI4AI · DCLM
Topics
data curation, pretraining, web scraping
Research notes
- Discovery: shared by Zichun Yu (@Zichun_Yu, PhD student @ CMU LTI with Chenyan Xiong) in an X thread on 2026-09-29: https://x.com/Zichun_Yu/status/2105024721716752608
- Blog: https://cxcscmu.github.io/ReScraper
- Model: https://huggingface.co/cx-cmu/ReScraper
- Data: https://huggingface.co/datasets/cx-cmu/ReScraper-Data (logged separately in the Datasets sheet)
- Code: https://github.com/cxcscmu/ReScraper
- Submitted to arXiv 2026-09-28.
- Positioned as AI4AI for pretraining data curation: a small learned model taking over an entire pipeline stage from hand-written heuristics.