← Back to explorer

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

A scalable, self-organized multi-agent harness with no central orchestrator: concurrent workers run an async cooperate loop (gather context, claim/self-assign sub-tasks, act and share findings, verify, merge progress) over three shared infrastructure components — a workspace of proposed/ongoing/completed work, a message interface, and shared context of reusable findings and work intentions.

Keywords

multi-agent · agent-harness · scaling · orchestration · ProgramBench

Topics

multi-agent, agent-harness, scaling, orchestration, ProgramBench

Research notes

  • Discovery: Shared in #random-papers as a bare arXiv link with no comment.
  • Method: Decentralized orchestration: workers self-assign sub-tasks from a shared workspace and coordinate asynchronously via message passing and shared context, removing the central orchestrator bottleneck. Evaluated on the five hardest ProgramBench tasks with GPT-5.6-sol (high), scaling 1 to 1,024 agents.
  • Key findings: Scaling 1 to 128 agents raises mean final test-pass rate from 19.31% to 28.78% (~49% relative gain) on the five hardest ProgramBench tasks; larger organizations reach comparable pass rates earlier. On pandoc, scaling 1 to 1,024 agents raises the final test-pass rate from 33.89% to 55.06%. Worker trajectories show self-organized cooperation forms gradually emerging and standardizing as the organization grows.
  • Limitations: Evaluation is on a narrow slice (five hardest ProgramBench tasks, one model family); no cost/latency breakdown of coordination overhead at 1,024 agents was captured from the abstract page.
  • Presents agent count as a new scaling dimension for multi-agent organizations under hard latency/time budgets.