Agensh: Scaling Organizational Intelligence to 1,024 Agents
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
A scalable, self-organized multi-agent harness with no central orchestrator: concurrent workers run an async cooperate loop (gather context, claim/self-assign sub-tasks, act and share findings, verify, merge progress) over three shared infrastructure components — a workspace of proposed/ongoing/completed work, a message interface, and shared context of reusable findings and work intentions.
Keywords
multi-agent · agent-harness · scaling · orchestration · ProgramBench
Topics
multi-agent, agent-harness, scaling, orchestration, ProgramBench
Research notes
- Discovery: Shared in #random-papers as a bare arXiv link with no comment.
- Method: Decentralized orchestration: workers self-assign sub-tasks from a shared workspace and coordinate asynchronously via message passing and shared context, removing the central orchestrator bottleneck. Evaluated on the five hardest ProgramBench tasks with GPT-5.6-sol (high), scaling 1 to 1,024 agents.
- Key findings: Scaling 1 to 128 agents raises mean final test-pass rate from 19.31% to 28.78% (~49% relative gain) on the five hardest ProgramBench tasks; larger organizations reach comparable pass rates earlier. On pandoc, scaling 1 to 1,024 agents raises the final test-pass rate from 33.89% to 55.06%. Worker trajectories show self-organized cooperation forms gradually emerging and standardizing as the organization grows.
- Limitations: Evaluation is on a narrow slice (five hardest ProgramBench tasks, one model family); no cost/latency breakdown of coordination overhead at 1,024 agents was captured from the abstract page.
- Presents agent count as a new scaling dimension for multi-agent organizations under hard latency/time budgets.