The Stack
- Type
- corpus
- Venue
- Hugging Face / BigCode
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- 358 programming languages; comments/docstrings in 40+ natural languages (EN, ZH, FR, PT, ES, RU, DE, KO, JA, etc.)
- Added
- 2026-07-17T20:18:03.654516+00:00
- Verified
- 2026-07-17T20:18:03.654516+00:00
Summary
The Stack is a collection of over 6TB of permissively-licensed source code files covering 358 programming languages, created as part of the BigCode Project (an open scientific collaboration for responsible Code LLM development). It serves as a pretraining dataset for code-generating LLMs and provides per-file provenance metadata (repository name, licenses, stars/forks/issues counts, git hashes) to support attribution and license compliance. The dataset has an opt-out mechanism and is regularly updated to enact validated data-removal requests.
Keywords
code pretraining github permissive-licenses bigcode starcoder corpus opt-out provenance
Topics
Code / Software Engineering
Research notes
- Gated; requires accepting Terms of Use and contact-info sharing. v1.0 had 30 languages/18 licenses (3TB); v1.1 expanded to 358 languages/193 licenses (6TB) and removed copyleft (MPL/EPL/LGPL). Papers: arxiv 2211.15533, 2107.03374, 2207.14157. A newer version, The Stack v2, also exists.