← Back to explorer

Nemotron Pre-Training Datasets

Type
collection
Venue
NVIDIA (HuggingFace)
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.759066+00:00
Verified
2026-07-17T20:18:03.759066+00:00

Summary

A HuggingFace collection of large-scale pretraining datasets used to train NVIDIA's Nemotron family of LLMs, including Nemotron-CC (a 6.3T-token refined Common Crawl dataset with 4.4T real + 1.9T synthetic tokens), Nemotron-CC-Math (133B tokens of math content), and specialized code, legal, and SFT subsets. The datasets use classifier ensembling, synthetic data rephrasing, and quality tiers to achieve better trade-offs between accuracy and data quantity for long-horizon pretraining, enabling 8B models to outperform Llama 3.1 8B.

Keywords

pretraining common-crawl nemotron nvidia synthetic math code legal llm

Topics

NLP / Pretraining

Research notes

  • Collection includes 13 datasets + 2 papers. Key paper: Nemotron-CC (arXiv:2412.02595, ACL 2025). Nemotron-CC-Math paper: arXiv:2508.15096. Also used by GAIR's daVinci-LLM project (~4.28T tokens from Nemotron-CC-v1).