← Back to explorer

Jupyter Agent Dataset

Type
dataset
Venue
jupyter-agent
Year
2026
Source
huggingface
Access
free
Language
code
Added
2026-07-17T19:57:55.130883+00:00
Verified
2026-07-17T19:57:55.130883+00:00

Summary

The dataset uses real Kaggle notebooks processed through a multi-stage pipeline to de-duplicate, fetch referenced datasets, score educational quality, filter to data-analysis–relevant content, generate dataset-grounded question–answer (QA) pairs, and produce executable reasoning traces by running notebooks. The resulting examples include natural questions about a dataset/notebook, verified answers, and step-by-step execution traces suitable for agent training.

Keywords

hf-dataset qa · -language-modeling · -code machine-generated monolingual parquet text datasets dask mlcroissant polars jupyter kaggle agents code synthetic

Topics

QA, Language Modeling, Code

Research notes

  • downloads=1353; likes=170