← Back to explorer

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

Type
paper
Venue
arXiv:2609.18805
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

ProgramDistill (KAIST/Microsoft) is a benchmark and pipeline that turns working web apps into verifiable coding tasks: it mines 1,975 replay-verified behaviors from 26 live apps and constructs 4,063 tasks with no human labeling, asking coding agents to infer missing behavior from a live reference and reimplement it. GPT-6 Astra leads at 49.2% on full-application reconstruction; the associated dataset is public on Hugging Face.

Keywords

benchmark · coding-agents · swe · evals · web

Topics

benchmark, coding-agents, swe, evals, web

Research notes

  • Method: Factorize working web apps into features of different granularities, each with replayable behaviors executable via a gold patch; a mine-craft-patch pipeline (Planner/Collector/Relabeler/Reflector roles + Playwright replay verifier) masks implementations and asks coding agents to restore behavior from a live reference.
  • Key findings: mine-craft-patch pipeline: 1,975 replay-verified behaviors mined across 26 interactive web apps → 4,063 tasks, zero human labeling; Three difficulty bands: atomic repair, cumulative repair, full-application reconstruction; Nine frontier coding agents evaluated: GPT-6 Astra 49.2%, Claude Opus 5 28.8% on cumulative full-app reconstruction; success collapses with restoration depth (100%→64%, 96%→32% from depth 1→8); Trajectory analysis: observation effort per required behavior declines as tasks deepen — effort allocation across observation/validation/editing is a key dimension alongside raw capability
  • Limitations: Web-app domain only; replay-verified behaviors exclude nondeterministic interactions. Partially a benchmark-marketing artifact of Microsoft's debug-gym effort.
  • KAIST, Microsoft Research Montreal, Microsoft AI. Blog: https://microsoft.github.io/debug-gym/blog/2026/09/programdistill/. Dataset (MIT for benchmark artifacts): https://huggingface.co/datasets/microsoft/ProgramDistill. Leaderboard on the blog page.