Nemotron-Pretraining-Code-v1
- Type
- dataset
- Venue
- Hugging Face / NVIDIA
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- 11 programming languages (for synthetic code QA); source code from GitHub in many languages
- Added
- 2026-07-17T20:18:03.651165+00:00
- Verified
- 2026-07-17T20:18:03.651165+00:00
Summary
Nemotron-Pretraining-Code-v1 is a large-scale curated source-code dataset mined from GitHub, processed through multi-stage filtering including license-based removal (BigCode-inspired, with a stricter license set), exact and fuzzy deduplication, and heuristic quality filters from OpenCoder, with all files annotated with metadata to guide filtering. It additionally includes large-scale synthetic code question-answer data in 11 programming languages generated by prompting LLMs (Mixtral-8x22B) on curated code snippets, solving generated problems, and filtering for correctness, producing diverse natural-language-code pairs for pretraining.
Keywords
nvidia nemotron code github pretraining synthetic deduplication opencoder bigcode
Topics
Code / Software Engineering
Research notes
- Gated; requires NVIDIA Data Agreement acceptance. Part of the Nemotron pretraining dataset (arxiv 2508.14444). Combines real GitHub source code (human) with synthetic code QA (AI-generated via Mixtral-8x22B).