FineWeb
- Type
- corpus
- Venue
- Hugging Face (FineData / HuggingFaceFW)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.602572+00:00
- Verified
- 2026-07-17T20:18:03.602572+00:00
Summary
FineWeb is a 15-trillion-token (now 18.5T+) dataset of cleaned and deduplicated English web text derived from 96 CommonCrawl snapshots spanning summer 2013 to April 2024. Created by Hugging Face using the datatrove library, its processing pipeline includes URL filtering, Trafilatura text extraction, language filtering, MassiveText quality filters, C4 filters, custom FineWeb filters, MinHash deduplication, and PII reformatting, producing better-performing LLMs than other open pretraining datasets. Published at NeurIPS 2024, it also includes FineWeb-Edu (1.3T tokens of educational content) as a companion dataset.
Keywords
web-text common-crawl pretraining english deduplication llm-training huggingface neurips-2024 large-scale
Topics
Web Text / General
Research notes
- Published at NeurIPS 2024 (Datasets and Benchmarks Track). arXiv:2406.17557. DOI:10.57967/hf/2493. Includes sample subsets (10BT, 100BT, 350BT) for smaller experiments.