GitHub Code 2025
- Type
- dataset
- Venue
- Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- multiple programming languages
- Added
- 2026-07-17T20:18:03.659632+00:00
- Verified
- 2026-07-17T20:18:03.659632+00:00
Summary
A large-scale code dataset containing 148 million source code files harvested from GitHub's top 1 million repositories with 2+ stars, totaling approximately 1 TB. Each row includes the repository ID, file size, file path, and full file content, making it suitable for training code generation models and performing code analysis tasks. At least 17 models have been trained or fine-tuned on this dataset.
Keywords
code github source-code code-generation programming large-scale
Topics
Code
Research notes
- Published by Nick Saga (nick007x) on HuggingFace. Dataset viewer and parquet conversion available.