← Back to explorer

OpenCoder Datasets

Type
collection
Venue
OpenCoder-LLM / Hugging Face
Year
2026
Source
huggingface
Access
free
Language
607 programming languages; English and Chinese (natural language)
Added
2026-07-17T20:18:03.608795+00:00
Verified
2026-07-17T20:18:03.608795+00:00

Summary

The OpenCoder Datasets collection comprises the full reproducible training data pipeline for the OpenCoder code LLM family, including RefineCode (960B tokens of high-quality code across 607 programming languages with 130+ language-specific filtering rules), FineWeb-code-corpus (148GB), FineWeb-math-corpus (10GB), an annealing corpus (24GB, 15.6M examples), and two-stage SFT datasets (4.21M + 436K examples). Released alongside the OpenCoder paper (ACL 2025), it provides complete transparency into the data curation process for training top-tier code LLMs.

Keywords

code programming llm-training refinecode sft pretraining open-cookbook reproducible multilingual synthetic

Topics

Code / Programming

Research notes

  • Published at ACL 2025 (arXiv:2411.04905). Collection includes 6 datasets: opc-sft-stage1, opc-sft-stage2, opc-fineweb-math-corpus, opc-fineweb-code-corpus, opc-annealing-corpus, RefineCode-code-corpus-meta. RefineCode built on The Stack v2 with additional filtering.