ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
- Type
- paper
- Venue
- arXiv:2609.18805
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
ProgramDistill (KAIST/Microsoft) is a benchmark and pipeline that turns working web apps into verifiable coding tasks: it mines 1,975 replay-verified behaviors from 26 live apps and constructs 4,063 tasks with no human labeling, asking coding agents to infer missing behavior from a live reference and reimplement it. GPT-6 Astra leads at 49.2% on full-application reconstruction; the associated dataset is public on Hugging Face.
Keywords
benchmark · coding-agents · swe · evals · web
Topics
benchmark, coding-agents, swe, evals, web
Research notes
- Method: Factorize working web apps into features of different granularities, each with replayable behaviors executable via a gold patch; a mine-craft-patch pipeline (Planner/Collector/Relabeler/Reflector roles + Playwright replay verifier) masks implementations and asks coding agents to restore behavior from a live reference.
- Key findings: mine-craft-patch pipeline: 1,975 replay-verified behaviors mined across 26 interactive web apps → 4,063 tasks, zero human labeling; Three difficulty bands: atomic repair, cumulative repair, full-application reconstruction; Nine frontier coding agents evaluated: GPT-6 Astra 49.2%, Claude Opus 5 28.8% on cumulative full-app reconstruction; success collapses with restoration depth (100%→64%, 96%→32% from depth 1→8); Trajectory analysis: observation effort per required behavior declines as tasks deepen — effort allocation across observation/validation/editing is a key dimension alongside raw capability
- Limitations: Web-app domain only; replay-verified behaviors exclude nondeterministic interactions. Partially a benchmark-marketing artifact of Microsoft's debug-gym effort.
- KAIST, Microsoft Research Montreal, Microsoft AI. Blog: https://microsoft.github.io/debug-gym/blog/2026/09/programdistill/. Dataset (MIT for benchmark artifacts): https://huggingface.co/datasets/microsoft/ProgramDistill. Leaderboard on the blog page.