← Back to explorer

Multi-Agent LLMs Fail to Explore Each Other

Type
paper
Venue
arXiv / University of Wisconsin–Madison / UC Santa Barbara
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:40:00Z
Verified
2026-08-14T16:40:00Z

Summary

In a two-armed delegation bandit, GPT-4, GPT-5, and Qwen2.5-7B-Instruct lock onto one peer within the first rounds (polarized 0-or-50 selections) unlike UCB. Formalizes Multi-Agent Exploration as a POSG and reduces it to per-agent LinUCB over relational features (n-gram response diversity, peer distinctiveness, historical reward, round). Across contextual diversity (10 Qwen2.5-7B agents on HotpotQA distractors) and parametric diversity (GPT-5/Qwen/Llama3.1/Mistral on Math500 and GPQA Diamond), in-context exploration can underperform random while MACE cuts regret; frozen MACE parameters transfer to 2WikiMultihopQA. Theory: MACE O(√(T log T)) regret vs greedy Ω(δT); exploration value grows with agent diversity. Code promised at https://github.com/deeplearning-wisc/mace.

Keywords

mace · multi-agent · exploration · linucb · hotpotqa · gpqa · math500 · uw-madison

Topics

multi-agent LLMs, exploration, contextual bandits

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.MA/AI). Corresponding froilanchoi / sharonli @cs.wisc.edu. Choi UW–Madison; Jiatong Li UCSB. Code promised at https://github.com/deeplearning-wisc/mace (not confirmed released at catalog time). Discord posted abs link.