Multi-Agent LLMs Fail to Explore Each Other
- Type
- paper
- Venue
- arXiv / University of Wisconsin–Madison / UC Santa Barbara
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:40:00Z
- Verified
- 2026-08-14T16:40:00Z
Summary
In a two-armed delegation bandit, GPT-4, GPT-5, and Qwen2.5-7B-Instruct lock onto one peer within the first rounds (polarized 0-or-50 selections) unlike UCB. Formalizes Multi-Agent Exploration as a POSG and reduces it to per-agent LinUCB over relational features (n-gram response diversity, peer distinctiveness, historical reward, round). Across contextual diversity (10 Qwen2.5-7B agents on HotpotQA distractors) and parametric diversity (GPT-5/Qwen/Llama3.1/Mistral on Math500 and GPQA Diamond), in-context exploration can underperform random while MACE cuts regret; frozen MACE parameters transfer to 2WikiMultihopQA. Theory: MACE O(√(T log T)) regret vs greedy Ω(δT); exploration value grows with agent diversity. Code promised at https://github.com/deeplearning-wisc/mace.
Keywords
mace · multi-agent · exploration · linucb · hotpotqa · gpqa · math500 · uw-madison
Topics
multi-agent LLMs, exploration, contextual bandits
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.MA/AI). Corresponding froilanchoi / sharonli @cs.wisc.edu. Choi UW–Madison; Jiatong Li UCSB. Code promised at https://github.com/deeplearning-wisc/mace (not confirmed released at catalog time). Discord posted abs link.