Nemotron-ClimbMix (ClimbMix)
- Type
- corpus
- Venue
- NVIDIA
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.738648+00:00
- Verified
- 2026-07-17T20:18:03.738648+00:00
Summary
Nemotron-ClimbMix is a 400-billion-token pre-training dataset from NVIDIA, created using the CLIMB (CLustering-based Iterative Data Mixture Bootstrapping) framework. Data is grouped into 1,000 topic-based clusters, filtered by advertisement detection and educational value classifiers, then mixed using optimized weights. A 1B model trained on this mixture exceeds Llama-3.2-1B by 2.0% averaged across 12 reasoning tasks. The dataset is tokenized with the GPT-2 tokenizer.
Keywords
pretraining data-mixing clustering nvidia llm filtered educational climb
Topics
NLP / Pre-training
Research notes
- Paper: arxiv 2504.13161. Non-commercial license. Tokenized with GPT-2 tokenizer; detokenization script provided. Companion to ClimbLab (1.2T-token corpus with 20 clusters).