← Back to explorer

who-ate-my-flops: bring your own PyTorch job, let an agent speed it up end to end

Type
tool
Venue
2026-09-30; github.com/OpenPerfAgent
Year
2026
Source
code
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

An open-source Claude Code / Codex plugin (OpenPerfAgent) for iteratively optimizing PyTorch training and inference jobs as a whole, not just individual GPU kernels. Design philosophy: rely on the foundation model's reasoning rather than prescribing how the agent reasons; the harness supplies (1) context -- computation structure and runtime profiling, (2) verification -- tools to check correctness against the baseline as the agent iterates, and (3) user alignment -- guidance to ask the right questions and align on goals/constraints. Claimed results: performance PRs developed with it merged into FunASR, FastVideo, gsplat, Ultralytics, Unsloth, and SGLang, with speedups up to 3.6x on tested workloads. Case study: in FastVideo the agent found CPU-GPU weight transfers consuming ~2/3 of the video-decoding stage, removed redundant transfers and sped up the rest, reducing generation time 15.20s -> 6.34s (2.40x) on 4 B200s. Usage: give the agent a repo, a launch command, and GPU access; run /who-ate-my-flops::init to clarify goals, then /who-ate-my-flops::optimize. Includes experiments, profiles, and lessons learned on the blog.

Keywords

MLSys · performance optimization · agents · PyTorch · profiling

Topics

MLSys, performance optimization, agents, PyTorch, profiling

Research notes

  • Discovery: Hexu Zhao (@zhaohexu2001, verified, NYU Courant MLSys PhD) 2026-09-30 thread: https://x.com/zhaohexu2001/status/2105338131285492120?s=20
  • Blog: https://openperfagent.github.io/who-ate-my-flops
  • Merged FastVideo PR example: https://github.com/hao-ai-lab/FastVideo/pull/1867
  • License not stated in the thread; team credits in the blog.
  • Connects to the collection's agentic performance-tuning and MLSys entries (PyTorch, profiling, whole-job optimization).