← Back to explorer

Nemotron-Pretraining-Code-v1

Type
dataset
Venue
Hugging Face / NVIDIA
Year
2026
Source
huggingface
Access
restricted
Language
11 programming languages (for synthetic code QA); source code from GitHub in many languages
Added
2026-07-17T20:18:03.651165+00:00
Verified
2026-07-17T20:18:03.651165+00:00

Summary

Nemotron-Pretraining-Code-v1 is a large-scale curated source-code dataset mined from GitHub, processed through multi-stage filtering including license-based removal (BigCode-inspired, with a stricter license set), exact and fuzzy deduplication, and heuristic quality filters from OpenCoder, with all files annotated with metadata to guide filtering. It additionally includes large-scale synthetic code question-answer data in 11 programming languages generated by prompting LLMs (Mixtral-8x22B) on curated code snippets, solving generated problems, and filtering for correctness, producing diverse natural-language-code pairs for pretraining.

Keywords

nvidia nemotron code github pretraining synthetic deduplication opencoder bigcode

Topics

Code / Software Engineering

Research notes

  • Gated; requires NVIDIA Data Agreement acceptance. Part of the Nemotron pretraining dataset (arxiv 2508.14444). Combines real GitHub source code (human) with synthetic code QA (AI-generated via Mixtral-8x22B).