Nemotron Pre-Training Datasets
- Type
- collection
- Venue
- NVIDIA (HuggingFace)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.759066+00:00
- Verified
- 2026-07-17T20:18:03.759066+00:00
Summary
A HuggingFace collection of large-scale pretraining datasets used to train NVIDIA's Nemotron family of LLMs, including Nemotron-CC (a 6.3T-token refined Common Crawl dataset with 4.4T real + 1.9T synthetic tokens), Nemotron-CC-Math (133B tokens of math content), and specialized code, legal, and SFT subsets. The datasets use classifier ensembling, synthetic data rephrasing, and quality tiers to achieve better trade-offs between accuracy and data quantity for long-horizon pretraining, enabling 8B models to outperform Llama 3.1 8B.
Keywords
pretraining common-crawl nemotron nvidia synthetic math code legal llm
Topics
NLP / Pretraining
Research notes
- Collection includes 13 datasets + 2 papers. Key paper: Nemotron-CC (arXiv:2412.02595, ACL 2025). Nemotron-CC-Math paper: arXiv:2508.15096. Also used by GAIR's daVinci-LLM project (~4.28T tokens from Nemotron-CC-v1).