← Back to explorer

Ultra-FineWeb

Type
corpus
Venue
HuggingFace (openbmb)
Year
2026
Source
huggingface
Access
free
Language
English, Chinese
Added
2026-07-17T20:18:03.625111+00:00
Verified
2026-07-17T20:18:03.625111+00:00

Summary

Ultra-FineWeb is a large-scale web text corpus by OpenBMB containing 1.29 billion rows (1.16B English, 131M Chinese) with quality scores, used for pretraining the MiniCPM4 series of language models. It is part of the UltraData Collection and applies advanced data filtering to FineWeb-sourced content, retaining documents with high quality scores (0.5–1.0 range). The dataset is described in arXiv papers 2505.05427, 2602.09003, and 2412.04315, and is licensed under Apache 2.0.

Keywords

pretraining web-corpus data-filtering high-quality fineweb minicpm llm bilingual

Topics

NLP / Pretraining Data

Research notes

  • Part of the UltraData Collection. Each row includes a content field, a float quality score, and a source string. The dataset is a re-filtered subset of FineWeb with quality scoring.