• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology
Report

Agensh coordinates 1,024 coding agents without a central orchestrator

Microsoft researchers report that self-organizing workers improved test-pass rates under a fixed six-hour coding benchmark.

elvisEL
1 Source, 13d ago, first seen 13d ago

TLDR

Microsoft Research’s Agensh replaces a central multi-agent coordinator with a shared workspace, message interface and reusable context. Workers claim subtasks, share findings, verify changes and merge progress asynchronously. On five difficult ProgramBench tasks using GPT-5.6-sol, the paper reports that raising the team from one agent to 128 increased the mean final test-pass rate from 19.31% to 28.78%. A separate 1,024-agent pandoc run reached 55.06%, compared with 33.89% for one agent. The results show useful parallel scaling in a controlled six-hour test, not production-ready autonomous software development.

Combined views

34.9K

1 Source, first seen 13d ago

364 likes61 comments469 saves57 reposts

Combined views

34.9K

1 Source, first seen 13d ago

364 likes61 comments469 saves57 reposts

Microsoft Research has introduced Agensh, a coding-agent harness designed to scale without asking one central model to divide and supervise all the work.

Featured Source

Instead, each worker looks at the project’s current state, claims a subtask, makes changes, shares what it learned, checks results and merges useful progress. The researchers argue that this self-organizing loop avoids a bottleneck common to multi-agent systems: as more workers join, a central orchestrator can struggle to keep every assignment and dependency straight.

A shared workspace replaces the manager

Agensh coordinates workers through three pieces of lightweight infrastructure. A shared workspace records proposed, active and completed work. A message interface lets workers communicate. Shared context retains findings and intentions that another worker can reuse.

That setup does not eliminate coordination. It moves coordination into a common state that every agent can read and update. Workers can divide labor, integrate one another’s changes and settle into repeatable roles without waiting for a single manager to issue each instruction.

The approach is especially relevant to long software tasks, where many small investigations, implementations and tests can proceed at the same time. It also creates familiar distributed-systems risks: workers can duplicate effort, pursue stale assumptions or collide while merging changes. Agensh’s verification and shared-state steps are meant to keep that parallel work useful.

More workers improved the benchmark, but did not solve it

The team evaluated Agensh with GPT-5.6-sol at high reasoning effort on the five hardest ProgramBench tasks: rebuilding FFmpeg, gromacs, pandoc, PHP-src and ctags. Each run had a six-hour budget and no internet access.

Across those five tasks, the paper reports a mean final test-pass rate of 19.31% with one agent and 28.78% with 128 agents — about a 49% relative increase. Larger teams generally reached a given pass rate sooner. On pandoc, for example, 128 agents cleared 30% after 30 minutes, while teams of 32 and eight first reached that mark after 60 and 90 minutes.

The largest demonstration put 1,024 agents on pandoc. Its final test-pass rate reached 55.06%, up from 33.89% for a single agent. Researchers also observed coordination patterns become more structured as the group grew, moving from peer-to-peer help toward standardized workflows and specialized roles.

Those numbers are promising, but they are not a claim that 1,024 agents can autonomously ship dependable production software. The experiments use one model family, a small set of difficult reconstruction tasks and a fixed laboratory setup. Even the best reported pandoc run left many tests failing. The clearest result is narrower: under these conditions, adding self-organizing parallel workers improved both speed and final coverage without requiring a central task dispatcher.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

elvis@omarsar0Banger paper from Microsoft Research. (bookmark it) They run 1K+ coding agents at once to test a scalable self-organized multi-agent harness. This is an interesting test because most multi-agent systems today have some hierarchy or structure. Agensh has no central orchestrator. It coordinates parallel coding agents through a shared state instead of a central orchestrator. Each agent gathers context, claims a sub-task, does the work, shares what it found, verifies the result, and merges it, all asynchronously through a shared workspace and a message channel. On the five hardest ProgramBench tasks with GPT-5.6-sol, increasing agents from 1 to 128 raises the mean final test-pass rate from 19.31% to 28.78%. Larger teams also reach a given pass rate sooner. On pandoc, 1,024 agents take the test-pass rate from 33.89% to 55.06%. The authors also report forms of cooperation that the agents start on their own and that become standard practice as the team grows. Paper: https://arxiv.org/abs/2609.26781 Chat with Paper: https://academy.dair.ai/papers/agensh-scaling-organizational-intelligence-to-1-024-agents-2609.2678113d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AgenshMicrosoft Research

    1 Source

    elvis@omarsar0Banger paper from Microsoft Research. (bookmark it) They run 1K+ coding agents at once to test a scalable self-organized multi-agent harness. This is an interesting test because most multi-agent systems today have some hierarchy or structure. Agensh has no central orchestrator. It coordinates parallel coding agents through a shared state instead of a central orchestrator. Each agent gathers context, claims a sub-task, does the work, shares what it found, verifies the result, and merges it, all asynchronously through a shared workspace and a message channel. On the five hardest ProgramBench tasks with GPT-5.6-sol, increasing agents from 1 to 128 raises the mean final test-pass rate from 19.31% to 28.78%. Larger teams also reach a given pass rate sooner. On pandoc, 1,024 agents take the test-pass rate from 33.89% to 55.06%. The authors also report forms of cooperation that the agents start on their own and that become standard practice as the team grows. Paper: https://arxiv.org/abs/2609.26781 Chat with Paper: https://academy.dair.ai/papers/agensh-scaling-organizational-intelligence-to-1-024-agents-2609.2678113d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet