KStack
- Type
- corpus
- Venue
- JetBrains (HuggingFace)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- Kotlin, Java, English
- Added
- 2026-07-17T20:18:03.766622+00:00
- Verified
- 2026-07-17T20:18:03.766622+00:00
Summary
KStack is the largest collection of permissively licensed Kotlin source code, containing 5.52 million files scraped from GitHub repositories with metadata like owner, repo ID, language distribution, license, and commit SHA. It is part of JetBrains' Kotlin ML Pack (alongside the curated KStack-clean and KExercises datasets) and is intended for pretraining and fine-tuning code generation models for Kotlin, demonstrating that even small high-quality subsets can yield up to a 16-point increase in HumanEval pass rate.
Keywords
kotlin code corpus github code-generation jetbrains permissively-licensed
Topics
Code
Research notes
- Part of JetBrains' Kotlin ML Pack. Associated papers: arXiv:2211.15533 (2022) and arXiv:2402.19173 / arXiv:2405.19250 (Kotlin ML Pack Technical Report, 2024). KStack-clean is a filtered 25k-example subset for higher-quality training.