← Back to explorer

Nemotron-ClimbMix (ClimbMix)

Type
corpus
Venue
NVIDIA
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.738648+00:00
Verified
2026-07-17T20:18:03.738648+00:00

Summary

Nemotron-ClimbMix is a 400-billion-token pre-training dataset from NVIDIA, created using the CLIMB (CLustering-based Iterative Data Mixture Bootstrapping) framework. Data is grouped into 1,000 topic-based clusters, filtered by advertisement detection and educational value classifiers, then mixed using optimized weights. A 1B model trained on this mixture exceeds Llama-3.2-1B by 2.0% averaged across 12 reasoning tasks. The dataset is tokenized with the GPT-2 tokenizer.

Keywords

pretraining data-mixing clustering nvidia llm filtered educational climb

Topics

NLP / Pre-training

Research notes

  • Paper: arxiv 2504.13161. Non-commercial license. Tokenized with GPT-2 tokenizer; detokenization script provided. Companion to ClimbLab (1.2T-token corpus with 20 clusters).