← Back to explorer

Nemotron Post-Training Dataset v2

Type
dataset
Venue
NVIDIA / Hugging Face
Year
2026
Source
huggingface
Access
restricted
Language
English, German, Italian, Spanish, French, Japanese
Added
2026-07-17T20:18:03.697198+00:00
Verified
2026-07-17T20:18:03.697198+00:00

Summary

The Nemotron Post-Training Dataset v2 is NVIDIA's open post-training dataset extending SFT and RL data into five target languages (Spanish, French, German, Italian, Japanese) to support the NVIDIA-Nemotron-Nano-9B-v2 model family. It contains ~6.3M samples across categories including math (239K), code (175K), STEM (355K), chat (628K), and multilingual splits (~1M each for JA, DE, IT, ES, FR). Prompts were sourced from public corpora or synthetically generated, and responses were synthetically generated by DeepSeek-R1-0528, Qwen2.5-14B-Instruct, Qwen2.5-32B-Instruct-AWQ, Qwen3-30B-A3B, and Qwen3-235B-A22B, with reasoning traces in English only. Released August 2025.

Keywords

nvidia post-training multilingual sft rl reasoning math code synthetic

Topics

NLP / Multilingual / Reasoning

Research notes

  • Requires agreeing to share contact info. Associated with arxiv: 2508.14444 (NVIDIA Nemotron Nano 2). Some prompts from Wildchat (ODC-BY) and StackOverflow (CC-BY-SA). Qwen and DeepSeek license terms may apply to derived models.