Ultra-FineWeb
- Type
- corpus
- Venue
- HuggingFace (openbmb)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English, Chinese
- Added
- 2026-07-17T20:18:03.625111+00:00
- Verified
- 2026-07-17T20:18:03.625111+00:00
Summary
Ultra-FineWeb is a large-scale web text corpus by OpenBMB containing 1.29 billion rows (1.16B English, 131M Chinese) with quality scores, used for pretraining the MiniCPM4 series of language models. It is part of the UltraData Collection and applies advanced data filtering to FineWeb-sourced content, retaining documents with high quality scores (0.5–1.0 range). The dataset is described in arXiv papers 2505.05427, 2602.09003, and 2412.04315, and is licensed under Apache 2.0.
Keywords
pretraining web-corpus data-filtering high-quality fineweb minicpm llm bilingual
Topics
NLP / Pretraining Data
Research notes
- Part of the UltraData Collection. Each row includes a content field, a float quality score, and a source string. The dataset is a re-filtered subset of FineWeb with quality scoring.