Jupyter Agent Dataset
- Type
- dataset
- Venue
- jupyter-agent
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- code
- Added
- 2026-07-17T19:57:55.130883+00:00
- Verified
- 2026-07-17T19:57:55.130883+00:00
Summary
The dataset uses real Kaggle notebooks processed through a multi-stage pipeline to de-duplicate, fetch referenced datasets, score educational quality, filter to data-analysis–relevant content, generate dataset-grounded question–answer (QA) pairs, and produce executable reasoning traces by running notebooks. The resulting examples include natural questions about a dataset/notebook, verified answers, and step-by-step execution traces suitable for agent training.
Keywords
hf-dataset qa · -language-modeling · -code machine-generated monolingual parquet text datasets dask mlcroissant polars jupyter kaggle agents code synthetic
Topics
QA, Language Modeling, Code
Research notes
- downloads=1353; likes=170