Falcon RefinedWeb
- Type
- corpus
- Venue
- Technology Innovation Institute (TII)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.786400+00:00
- Verified
- 2026-07-17T20:18:03.786400+00:00
Summary
Falcon RefinedWeb is a large-scale English web corpus extracted from Common Crawl and rigorously filtered and deduplicated, used to train the Falcon LLM family. The accompanying paper (arXiv:2306.01116) demonstrates that properly filtered web data alone can match or exceed curated sources for pretraining. This HuggingFace release contains 968M documents in parquet format.
Keywords
web-corpus pretraining common-crawl falcon filtered deduplicated
Topics
Web text / General
Research notes
- Associated papers: arXiv:2306.01116 (RefinedWeb), arXiv:2203.15556, arXiv:2107.06499. DOI: 10.57967/hf/0737. One of the most influential open web corpora for LLM pretraining.