← Back to explorer

The Stack GitHub Issues

Type
dataset
Venue
BigCode / Hugging Face
Year
2026
Source
huggingface
Access
restricted
Language
English (primarily), code snippets in 300+ programming languages
Added
2026-07-17T20:18:03.682350+00:00
Verified
2026-07-17T20:18:03.682350+00:00

Summary

The Stack GitHub Issues dataset contains ~30.9 million conversations from GitHub issues and pull requests, extracted as part of the larger Stack code dataset. Each conversation includes events such as issue opening, comments, and closures, with author usernames masked for PII protection. The dataset was cleaned from 180GB down to 54GB by removing bot comments, automated email replies, and low-quality conversations, and is used to train code generation models like Stable Code 3B.

Keywords

code github issues conversations software-development llm-training

Topics

Code / NLP

Research notes

  • Requires agreeing to Terms of Use and sharing contact info. Part of the BigCode Stack initiative. PII (usernames, IPs, emails) has been redacted via regex masking.