FineWeb 2
- Type
- corpus
- Venue
- Hugging Face (FineData / HuggingFaceFW)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- 1000+ languages (1870 language subsets including Arabic, Bengali, Chinese, French, German, Spanish, and many low-resource languages)
- Added
- 2026-07-17T20:18:03.606193+00:00
- Verified
- 2026-07-17T20:18:03.606193+00:00
Summary
FineWeb 2 is a multilingual web text dataset extending the FineWeb methodology to over 1000 languages (1870 subsets), derived from CommonCrawl data and processed with the same datatrove-based pipeline as the original FineWeb. It provides cleaned, deduplicated text for languages ranging from major ones (Arabic, Chinese, French, German) to extremely low-resource languages, making it the largest publicly available multilingual clean web pretraining corpus. Published with an associated paper (arXiv:2506.20920).
Keywords
web-text multilingual common-crawl pretraining low-resource llm-training huggingface large-scale fineweb
Topics
Web Text / Multilingual
Research notes
- arXiv:2506.20920. DOI:10.57967/hf/3744. Successor to FineWeb. Many subsets have very few rows (e.g., some languages have <100 rows).