← Back to explorer

Decentralized Multi-Agent Systems with Shared Context

Type
other
Venue
arXiv / Stanford University

Summary

DeLM has parallel agents claim subtasks from a queue, read compact verified gists, and admit updates only after evidence checks, with hierarchical unfold (gist→summary→raw) for long sources. On SWE-bench Verified with Gemini 3 Flash: Avg.@1 65.7 / Pass@4 77.4 at $0.12/task vs AOrchestra-Parallel 56.4 / 71.8 at $0.25. Claude Opus 4.6 gains are smaller but still best. LongBench-v2 Multi-Doc QA: highest average across GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash, DeepSeek-V4-Pro (up to +5.7 pp). Hybrid DeLM+RLM is best on OOLONG and LongBench-v2. Code https://github.com/yuzhenmao/DeLM; project https://yuzhenmao.github.io/DeLM/.

Keywords

delm · multi-agent · shared-context · swe-bench · longbench · rlm · stanford

Topics

multi-agent systems, test-time scaling, long-context agents

Research notes

  • Primary: arxiv abs (cs.MA). Stanford SAIL/HAI; correspondence yuzhenm/azalia@stanford.edu. Code https://github.com/yuzhenmao/DeLM (MIT badge; GitHub license field Other; 110 stars on HF paper page at check); project https://yuzhenmao.github.io/DeLM/. HF paper page 4 upvotes, org StanfordUniversity; no linked models/datasets. Discord posted abs. Uses existing SWE-bench/LongBench-v2/OOLONG; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so license left blank.