← Back to explorer

GitHub Code 2025

Type
dataset
Venue
Hugging Face
Year
2026
Source
huggingface
Access
free
Language
multiple programming languages
Added
2026-07-17T20:18:03.659632+00:00
Verified
2026-07-17T20:18:03.659632+00:00

Summary

A large-scale code dataset containing 148 million source code files harvested from GitHub's top 1 million repositories with 2+ stars, totaling approximately 1 TB. Each row includes the repository ID, file size, file path, and full file content, making it suitable for training code generation models and performing code analysis tasks. At least 17 models have been trained or fine-tuned on this dataset.

Keywords

code github source-code code-generation programming large-scale

Topics

Code

Research notes

  • Published by Nick Saga (nick007x) on HuggingFace. Dataset viewer and parquet conversion available.