cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
- Type
- paper
- Venue
- arXiv:2609.40284 (cs.LG), submitted 30 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Proposes cua-speedrun, a standardized framework for evaluating computer-use agents (CUAs) on speed, cost, and performance under a uniform virtual machine setup, execution pipeline, and common agent interface, addressing a reproducibility crisis where varying machine/container configurations confound speed measurement. Across four CUA benchmarks (OSWorld, OSWorld 2.0, CUA-World, MyPCBench), no single model family is optimal across performance, speed, and cost, and no open-weight models lie on the efficiency frontier. Counterintuitively, for some models increasing reasoning effort speeds up overall task completion (better actions, shorter trajectories), while faster environment input-output (40-1000x) can slow agents down because they act before the UI finishes updating. The benchmark task sets can also be substantially reduced without degrading statistical power, enabling faster, cheaper evaluation. All code, infrastructure, and analysis are public.
Keywords
cua-speedrun · computer-use agents · CUA · OSWorld · benchmarking speed · cost-performance Pareto · reasoning effort · environment latency · Modal
Topics
computer-use agents, benchmarking, evaluation infrastructure, OSWorld, efficiency, agent speed
Research notes
- Discovery: @rsalakhu X post 2026-10-01 (https://x.com/rsalakhu/status/2105715719300112530) pointing to @kohjingyu's detailed thread (https://x.com/kohjingyu/status/2105680456587137295)
- Website/leaderboard: https://cuaspeedrun.com
- Code: https://github.com/cua-speedrun/cua-speedrun
- Standardizes desktop environments, VM hardware (through Modal), task set, timing rules, and agent interfaces so single-agent implementations run across benchmarks; motivation: recent models score 78.7-86.1% vs humans 72.4% on OSWorld-Verified, but success rates hide speed — e.g. GPT-6 Astra and Kimi K3 score the same on OSWorld while Kimi K3 takes 4.4x longer per task
- Findings: agents with similar scores have very different speed/cost Pareto frontiers and the right choice depends on benchmark and use case; on OSWorld-Verified the time-performance frontier is dominated by a mix of model families and reasoning efforts (incl. Opus 5.5, GPT-6 Astra, GPT-5.6 Luna); more reasoning can mean less waiting (more tokens per step but better actions -> shorter, faster trajectories); making VMs 40-1000x faster at environment step times can make agents slower because they act before the UI has finished updating; evaluation task sets can be substantially reduced while preserving statistical power
- Team: Carnegie Mellon SCS (@scsatcmu), co-led by Jing Yu Koh, Pranjal Aggarwal, Lawrence Jang, with Sean Welleck, Daniel Fried, Ruslan Salakhutdinov
- License: CC BY 4.0