← Back to explorer

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Type
paper
Venue
arXiv:2609.31590 (cs.MA), submitted 25 Sep 2026; published as a COLM 2026 conference paper
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Existing multi-agent benchmarks test competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance — failing to isolate genuine collaboration. AgentWorld is a benchmark of 100 human-annotated tasks (plus 100 augmented variants) for long-horizon multi-agent collaboration: tasks span 50+ interaction rounds in a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness beyond binary task success, it introduces Causal Collaboration Effectiveness (CCE), a graph-based metric tracing causal dependencies between agent actions to measure what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B: even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. Fully open-source.

Keywords

AgentWorld · multi-agent LLM · benchmark · collaboration effectiveness · CCE · coordination · long-horizon tasks · sandbox

Topics

multi-agent LLM, benchmark, collaboration, long-horizon, coordination, sandbox

Research notes

  • Discovery: @omarsar0 (dair_ai) X post 2026-10-02 (https://x.com/omarsar0/status/2105855550588428584)
  • Chat-with-paper page: https://academy.dair.ai/papers/agentworld-benchmarking-long-horizon-collaboration-of-multi-agent-llms-2609.31590
  • Key finding: fewer than a third of a multi-agent team's actions actually help finish the task — more agents do not mean higher performance; coordination is the bottleneck, with coordination tasks the hardest category (12% success) and common failures including communication breakdowns, role confusion, and lost shared plans
  • Authors: OpenAgents, Columbia University, University of Pennsylvania, Seoul National University, Penn State University
  • License: CC BY 4.0