← Back to explorer

Nemotron-Pretraining-SFT-v1

Type
dataset
Venue
Hugging Face / NVIDIA
Year
2026
Source
huggingface
Access
restricted
Added
2026-07-17T20:18:03.649318+00:00
Verified
2026-07-17T20:18:03.649318+00:00

Summary

Nemotron-Pretraining-SFT-v1 is a diverse synthetically generated and curated SFT-style dataset spanning STEM, multilingual, academic, and reasoning domains, part of the broader 6.5-trillion-token Nemotron pretraining dataset. STEM data was expanded from high-quality math/science seeds using multi-iteration generation with Qwen3 and DeepSeek models (DeepSeek-R1, DeepSeek-R1-0528, Qwen2.5-Math-72B, Qwen2.5-32B-Instruct, Mixtral-8x22B, Qwen3-30B-A3B, and others), producing varied, harder, multiple-choice questions with solutions. It supports the NVIDIA Nemotron Nano 2 model family (9B/12B) with 128K context.

Keywords

nvidia nemotron sft synthetic stem math reasoning multilingual pretraining deepseek qwen

Topics

NLP / Multilingual / STEM / Reasoning

Research notes

  • Gated; requires agreeing to NVIDIA Data Agreement for Model Training. Synthetic data generated with DeepSeek-R1/R1-0528, Qwen2.5/Qwen3 family, Mixtral-8x22B, Nemotron-4-340B-Instruct, and others. Paper: arxiv 2508.14444.