OpenCoder Datasets
- Type
- collection
- Venue
- OpenCoder-LLM / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- 607 programming languages; English and Chinese (natural language)
- Added
- 2026-07-17T20:18:03.608795+00:00
- Verified
- 2026-07-17T20:18:03.608795+00:00
Summary
The OpenCoder Datasets collection comprises the full reproducible training data pipeline for the OpenCoder code LLM family, including RefineCode (960B tokens of high-quality code across 607 programming languages with 130+ language-specific filtering rules), FineWeb-code-corpus (148GB), FineWeb-math-corpus (10GB), an annealing corpus (24GB, 15.6M examples), and two-stage SFT datasets (4.21M + 436K examples). Released alongside the OpenCoder paper (ACL 2025), it provides complete transparency into the data curation process for training top-tier code LLMs.
Keywords
code programming llm-training refinecode sft pretraining open-cookbook reproducible multilingual synthetic
Topics
Code / Programming
Research notes
- Published at ACL 2025 (arXiv:2411.04905). Collection includes 6 datasets: opc-sft-stage1, opc-sft-stage2, opc-fineweb-math-corpus, opc-fineweb-code-corpus, opc-annealing-corpus, RefineCode-code-corpus-meta. RefineCode built on The Stack v2 with additional filtering.