← Back to explorer

The Stack v2

Type
corpus
Venue
BigCode Project / Software Heritage / INRIA
Year
2026
Source
huggingface
Access
restricted
Language
658 programming and markup languages
Added
2026-07-17T20:18:03.735027+00:00
Verified
2026-07-17T20:18:03.735027+00:00

Summary

The Stack v2 is a large-scale pre-training dataset of source code containing over 3 billion files in 600+ programming and markup languages, derived from the Software Heritage archive. It serves as a pre-training corpus for Code LLMs (used to train StarCoder2), with the full dataset at 67.5TB and the deduplicated version at 32.1TB (~900B tokens). The dataset provides SWHIDs (Software Heritage IDs) for provenance and compliance, with file contents stored on Software Heritage's S3 bucket.

Keywords

code pretraining source-code software-heritage starcoder2 multilingual bigcode llm

Topics

Code

Research notes

  • Bulk download requires agreement with Software Heritage and INRIA. Regularly updated for data removal requests. Paper: arxiv 2402.19173. Multiple variants available: the-stack-v2, the-stack-v2-dedup, the-stack-v2-train-full-ids, the-stack-v2-train-smol-ids.