Nemotron Post-Training Dataset v2
- Type
- dataset
- Venue
- NVIDIA / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- English, German, Italian, Spanish, French, Japanese
- Added
- 2026-07-17T20:18:03.697198+00:00
- Verified
- 2026-07-17T20:18:03.697198+00:00
Summary
The Nemotron Post-Training Dataset v2 is NVIDIA's open post-training dataset extending SFT and RL data into five target languages (Spanish, French, German, Italian, Japanese) to support the NVIDIA-Nemotron-Nano-9B-v2 model family. It contains ~6.3M samples across categories including math (239K), code (175K), STEM (355K), chat (628K), and multilingual splits (~1M each for JA, DE, IT, ES, FR). Prompts were sourced from public corpora or synthetically generated, and responses were synthetically generated by DeepSeek-R1-0528, Qwen2.5-14B-Instruct, Qwen2.5-32B-Instruct-AWQ, Qwen3-30B-A3B, and Qwen3-235B-A22B, with reasoning traces in English only. Released August 2025.
Keywords
nvidia post-training multilingual sft rl reasoning math code synthetic
Topics
NLP / Multilingual / Reasoning
Research notes
- Requires agreeing to share contact info. Associated with arxiv: 2508.14444 (NVIDIA Nemotron Nano 2). Some prompts from Wildchat (ODC-BY) and StackOverflow (CC-BY-SA). Qwen and DeepSeek license terms may apply to derived models.