← Back to explorer

Ultra-FineWeb-L3

Type
dataset
Venue
OpenBMB (Hugging Face)
Year
2026
Source
huggingface
Access
free
Language
English, Chinese
Added
2026-07-17T20:18:03.818066+00:00
Verified
2026-07-17T20:18:03.818066+00:00

Summary

Ultra-FineWeb-L3 is the L3 (late-stage refined) tier of OpenBMB's UltraData L0-L4 framework, transforming high-value web corpora from Ultra-FineWeb into structured, high-learnability training data with clearer reasoning signals. It uses MiniCPM4 and Qwen3 to perform Q&A Pair Generation and Multi-style Rewriting, producing 400B+ English tokens and 200B+ Chinese tokens — the largest open-source Chinese pre-training synthetic corpus to date. It serves as key training data for the decay phase of MiniCPM5-1B.

Keywords

pretraining synthetic data-synthesis qa-generation multi-style-rewriting ultrafineweb ultradata minicpm high-quality chinese

Topics

NLP / Pretraining Data

Research notes

  • Part of the UltraData L0-L4 tiered data management framework. Built on Ultra-FineWeb (1T English + 120B Chinese tokens). Technical report: arxiv 2505.05427.