← Back to explorer

FineWeb 2

Type
corpus
Venue
Hugging Face (FineData / HuggingFaceFW)
Year
2026
Source
huggingface
Access
free
Language
1000+ languages (1870 language subsets including Arabic, Bengali, Chinese, French, German, Spanish, and many low-resource languages)
Added
2026-07-17T20:18:03.606193+00:00
Verified
2026-07-17T20:18:03.606193+00:00

Summary

FineWeb 2 is a multilingual web text dataset extending the FineWeb methodology to over 1000 languages (1870 subsets), derived from CommonCrawl data and processed with the same datatrove-based pipeline as the original FineWeb. It provides cleaned, deduplicated text for languages ranging from major ones (Arabic, Chinese, French, German) to extremely low-resource languages, making it the largest publicly available multilingual clean web pretraining corpus. Published with an associated paper (arXiv:2506.20920).

Keywords

web-text multilingual common-crawl pretraining low-resource llm-training huggingface large-scale fineweb

Topics

Web Text / Multilingual

Research notes

  • arXiv:2506.20920. DOI:10.57967/hf/3744. Successor to FineWeb. Many subsets have very few rows (e.g., some languages have <100 rows).