Nemotron-CC-v2 (Nemotron Pre-Training Dataset v1)
- Type
- dataset
- Venue
- NVIDIA
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- English (with translated QA in 15 languages)
- Added
- 2026-07-17T20:18:03.819941+00:00
- Verified
- 2026-07-17T20:18:03.819941+00:00
Summary
Nemotron-CC-v2 is NVIDIA's updated English web crawl dataset based on Nemotron-CC, with eight additional Common Crawl snapshots (2024-2025), synthetic rephrasing using Qwen3-30B-A3B and Mistral-Nemo-12B, English filtering, and global deduplication. It is part of the larger Nemotron Pre-Training Dataset v1 (6.58T total tokens across English CC, synthetic CC, diverse QA, translated QA, math, and code categories) used to train the NVIDIA Nemotron Nano 2 family of LLMs (9B/12B parameters, 128K context).
Keywords
pretraining common-crawl nvidia nemotron synthetic multilingual math code web-crawl deduplication
Topics
NLP / Pretraining Data
Research notes
- Requires agreeing to NVIDIA Data Agreement for Model Training. Intended for model training purposes only. Part of a 4-dataset release including Nemotron-CC-Math-v1, Nemotron-Pretraining-Code-v1, and Nemotron-Pretraining-SFT-v1. Paper: arxiv 2508.14444.