TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development
- Type
- paper
- Venue
- Carnegie Mellon University; arXiv:2608.26086 (accepted at NeurIPS 2026 E&D track, COLM 2026 Workshop on Agent Behavior); annotations/schemas/dataset code CC BY 4.0; toolkit Apache-2.0
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
TraceML: a trajectory-level analysis tool for auto-research agents and systematic agent-human comparisons. Reconstructs ML development trajectories (human Kaggle notebook histories; agent git commits and search journals) into a unified schema: every code version labeled with ML-pipeline stages (8 coarse, 136 fine tags) and each edit with action, intent, magnitude, and score effect. Released as the traceml CLI so users can analyze their own agent runs (AIDE, MLEvolve journals, any git workspace). Scale: 4,465 human Kaggle trajectories across 134 competitions (151,088 labeled code versions); paired subset of 430 human + 207 agent trajectories on 7 competitions. Findings: agents collapse into narrow loops -- Codex spends 89% of edits tuning the submission (62nd percentile finish), MLEvolve 63% mutating its model (48th percentile), while top-10% humans mix data work, validation, model changes, and ensembling (97th/92nd percentile finishes); humans pivot on 25% of transitions vs 9% (Codex) / 58% (MLEvolve); top humans return to earlier approaches on 9% of eligible versions (78% end higher), Codex 1 of 658, MLEvolve 0 of 344. A ~1,000-token planning prompt distilled from human practice improved scores in 5 of 7 competitions (2 within noise, none regressed) -- "closes only the instructable part of the gap."
Keywords
agents · ML research · trajectory analysis · Kaggle · auto-research
Topics
agents, ML research, trajectory analysis, Kaggle, auto-research
Research notes
- Discovery: Weiwei Sun (@sunweiwei12) 2026-09-30 quote-post of Jiarui Yan (@jiaruiyan123): https://x.com/sunweiwei12/status/2105405100172722496?s=20
- Project site: https://jerryyan123.github.io/TraceML/
- Code/toolkit: https://github.com/JerryYan123/TraceML
- Labeling models: https://huggingface.co/jerryyan/TraceML-Labelers
- Dataset: https://huggingface.co/datasets/jerryyan/TraceML (separate Datasets sheet row).
- Connects to the collection's agent, research-automation, and benchmark entries.