← Back to explorer

FineWeb

Type
corpus
Venue
Hugging Face (FineData / HuggingFaceFW)
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.602572+00:00
Verified
2026-07-17T20:18:03.602572+00:00

Summary

FineWeb is a 15-trillion-token (now 18.5T+) dataset of cleaned and deduplicated English web text derived from 96 CommonCrawl snapshots spanning summer 2013 to April 2024. Created by Hugging Face using the datatrove library, its processing pipeline includes URL filtering, Trafilatura text extraction, language filtering, MassiveText quality filters, C4 filters, custom FineWeb filters, MinHash deduplication, and PII reformatting, producing better-performing LLMs than other open pretraining datasets. Published at NeurIPS 2024, it also includes FineWeb-Edu (1.3T tokens of educational content) as a companion dataset.

Keywords

web-text common-crawl pretraining english deduplication llm-training huggingface neurips-2024 large-scale

Topics

Web Text / General

Research notes

  • Published at NeurIPS 2024 (Datasets and Benchmarks Track). arXiv:2406.17557. DOI:10.57967/hf/2493. Includes sample subsets (10BT, 100BT, 350BT) for smaller experiments.