CodeFIM-Data
- Type
- dataset
- Venue
- HuggingFace (Etherll)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- Rust, TypeScript, Python, Go (programming languages)
- Added
- 2026-07-17T20:18:03.614276+00:00
- Verified
- 2026-07-17T20:18:03.614276+00:00
Summary
CodeFIM-Data is a HuggingFace dataset of 238,000 rows formatted for fill-in-the-middle (FIM) code completion training. Each row contains a file name, a code prefix, a code suffix, the missing middle segment, and a FIM type label. The data is drawn from multiple programming languages (primarily Rust, but also TypeScript, Python, and Go) and is used to fine-tune code completion models such as Etherll's Qwen2.5-CodeFIM series for IDE-style autocomplete via Continue.
Keywords
code fill-in-the-middle fim code-completion rust autocomplete
Topics
Code
Research notes
- No explicit license stated on the dataset card. The creator also published a larger multi-language variant (code-fim-v2, 64k rows) and a Rust-specific variant (CodeFIM-Rust-Mellum).