The Stack GitHub Issues
- Type
- dataset
- Venue
- BigCode / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- restricted
- Language
- English (primarily), code snippets in 300+ programming languages
- Added
- 2026-07-17T20:18:03.682350+00:00
- Verified
- 2026-07-17T20:18:03.682350+00:00
Summary
The Stack GitHub Issues dataset contains ~30.9 million conversations from GitHub issues and pull requests, extracted as part of the larger Stack code dataset. Each conversation includes events such as issue opening, comments, and closures, with author usernames masked for PII protection. The dataset was cleaned from 180GB down to 54GB by removing bot comments, automated email replies, and low-quality conversations, and is used to train code generation models like Stable Code 3B.
Keywords
code github issues conversations software-development llm-training
Topics
Code / NLP
Research notes
- Requires agreeing to Terms of Use and sharing contact info. Part of the BigCode Stack initiative. PII (usernames, IPs, emails) has been redacted via regex masking.