Ultra-FineWeb-L3
- Type
- dataset
- Venue
- OpenBMB (Hugging Face)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English, Chinese
- Added
- 2026-07-17T20:18:03.818066+00:00
- Verified
- 2026-07-17T20:18:03.818066+00:00
Summary
Ultra-FineWeb-L3 is the L3 (late-stage refined) tier of OpenBMB's UltraData L0-L4 framework, transforming high-value web corpora from Ultra-FineWeb into structured, high-learnability training data with clearer reasoning signals. It uses MiniCPM4 and Qwen3 to perform Q&A Pair Generation and Multi-style Rewriting, producing 400B+ English tokens and 200B+ Chinese tokens — the largest open-source Chinese pre-training synthetic corpus to date. It serves as key training data for the decay phase of MiniCPM5-1B.
Keywords
pretraining synthetic data-synthesis qa-generation multi-style-rewriting ultrafineweb ultradata minicpm high-quality chinese
Topics
NLP / Pretraining Data
Research notes
- Part of the UltraData L0-L4 tiered data management framework. Built on Ultra-FineWeb (1T English + 120B Chinese tokens). Technical report: arxiv 2505.05427.