AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
- Type
- paper
- Venue
- arXiv:2609.31590 (cs.MA), submitted 25 Sep 2026; published as a COLM 2026 conference paper
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Existing multi-agent benchmarks test competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance — failing to isolate genuine collaboration. AgentWorld is a benchmark of 100 human-annotated tasks (plus 100 augmented variants) for long-horizon multi-agent collaboration: tasks span 50+ interaction rounds in a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness beyond binary task success, it introduces Causal Collaboration Effectiveness (CCE), a graph-based metric tracing causal dependencies between agent actions to measure what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B: even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. Fully open-source.
Keywords
AgentWorld · multi-agent LLM · benchmark · collaboration effectiveness · CCE · coordination · long-horizon tasks · sandbox
Topics
multi-agent LLM, benchmark, collaboration, long-horizon, coordination, sandbox
Research notes
- Discovery: @omarsar0 (dair_ai) X post 2026-10-02 (https://x.com/omarsar0/status/2105855550588428584)
- Chat-with-paper page: https://academy.dair.ai/papers/agentworld-benchmarking-long-horizon-collaboration-of-multi-agent-llms-2609.31590
- Key finding: fewer than a third of a multi-agent team's actions actually help finish the task — more agents do not mean higher performance; coordination is the bottleneck, with coordination tasks the hardest category (12% success) and common failures including communication breakdowns, role confusion, and lost shared plans
- Authors: OpenAgents, Columbia University, University of Pennsylvania, Seoul National University, Penn State University
- License: CC BY 4.0