The Stack v2
- Type
- corpus
- Venue
- BigCode Project / Software Heritage / INRIA
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- 658 programming and markup languages
- Added
- 2026-07-17T20:18:03.735027+00:00
- Verified
- 2026-07-17T20:18:03.735027+00:00
Summary
The Stack v2 is a large-scale pre-training dataset of source code containing over 3 billion files in 600+ programming and markup languages, derived from the Software Heritage archive. It serves as a pre-training corpus for Code LLMs (used to train StarCoder2), with the full dataset at 67.5TB and the deduplicated version at 32.1TB (~900B tokens). The dataset provides SWHIDs (Software Heritage IDs) for provenance and compliance, with file contents stored on Software Heritage's S3 bucket.
Keywords
code pretraining source-code software-heritage starcoder2 multilingual bigcode llm
Topics
Code
Research notes
- Bulk download requires agreement with Software Heritage and INRIA. Regularly updated for data removal requests. Paper: arxiv 2402.19173. Multiple variants available: the-stack-v2, the-stack-v2-dedup, the-stack-v2-train-full-ids, the-stack-v2-train-smol-ids.