← Back to explorer

The Stack

Type
corpus
Venue
Hugging Face / BigCode
Year
2026
Source
huggingface
Access
restricted
Language
358 programming languages; comments/docstrings in 40+ natural languages (EN, ZH, FR, PT, ES, RU, DE, KO, JA, etc.)
Added
2026-07-17T20:18:03.654516+00:00
Verified
2026-07-17T20:18:03.654516+00:00

Summary

The Stack is a collection of over 6TB of permissively-licensed source code files covering 358 programming languages, created as part of the BigCode Project (an open scientific collaboration for responsible Code LLM development). It serves as a pretraining dataset for code-generating LLMs and provides per-file provenance metadata (repository name, licenses, stars/forks/issues counts, git hashes) to support attribution and license compliance. The dataset has an opt-out mechanism and is regularly updated to enact validated data-removal requests.

Keywords

code pretraining github permissive-licenses bigcode starcoder corpus opt-out provenance

Topics

Code / Software Engineering

Research notes

  • Gated; requires accepting Terms of Use and contact-info sharing. v1.0 had 30 languages/18 licenses (3TB); v1.1 expanded to 358 languages/193 licenses (6TB) and removed copyleft (MPL/EPL/LGPL). Papers: arxiv 2211.15533, 2107.03374, 2207.14157. A newer version, The Stack v2, also exists.