| Paper | Year | Why read it |
|---|---|---|
| Multi-agent architecture and orchestration | ||
| AutoGenEnabling Next-Gen LLM Applications via Multi-Agent Conversation | 2023 |
Background. By 2023 the single-agent loop of a model plus tools plus a scratchpad was a settled pattern. Building an application out of several such agents meant writing the orchestration by hand every time, because there was no vocabulary for how one agent hands work to another. Problem. Multi-agent applications had no programming model. What an agent is, how it receives a turn, who decides whose turn comes next, and where a human sits in the arrangement were all answered per application, so nothing composed and nothing was reusable. Key idea. Make conversation the programming model. Agents are conversable objects with configurable roles and tools, a human can be one of the participants on equal terms, and control flow is whatever message-passing pattern you wire between them, including a group chat with a manager choosing the next speaker. Findings. On the full MATH test set, an AutoGen conversation between an assistant and a code executor reaches 69.48% accuracy, compared with 55.18% for GPT-4 alone. On 134 unseen ALFWorld tasks, adding a grounding agent improves success by 15% on average over the same conversational system without that agent. Why it matters. It gives you the vocabulary for deciding when a second agent earns its bill. Read against the fan-out arithmetic it also prices the arrangement, since every additional participant carries its own context and pays its own tokens on every turn it takes. |
| MetaGPTMeta Programming for A Multi-Agent Collaborative Framework | 2023 |
Background. Conversational multi-agent frameworks let agents talk freely and let coordination emerge from the conversation. Human software teams do not work that way. They assign roles and pass documents whose shape is agreed in advance. Problem. Free-form conversation between agents drifts. Each agent restates the task in its own words, an error made early propagates as prose that later agents treat as given, and no step has a contract the next step can check before acting on it. Key idea. Assign roles borrowed from a software organization, product manager, architect, engineer, and quality assurance, and require each role to emit a structured artifact that the next role consumes. The hand-off is explicit and the return contract between agents becomes the design rather than a convention. Findings. With GPT-4, MetaGPT reaches 85.9% pass@1 on HumanEval and 87.7% on MBPP. Its executable-feedback loop improves MBPP by 5.4 percentage points over the otherwise identical pipeline without execution feedback, isolating part of the gain from the broader multi-agent arrangement. Why it matters. It is the clearest statement of the position that structure, not more conversation, is what makes a pipeline of agents work. The structured hand-off is also what makes the pipeline debuggable, since you can read the artifacts and see which stage was already wrong. |
| multi-agent research systemHow we built our multi-agent research system (Anthropic) | 2025 |
Background. Some agent builders favor one compacted thread. Anthropic's research feature instead uses a lead agent to delegate parallel searches and synthesize their results. Problem. A multi-agent design must specify each worker's task, its return format, and the lead's stopping rule. Every worker also pays for a fresh context. Key idea. The lead delegates explicit tasks, workers explore in parallel, and each returns a summary rather than its transcript. Prompting concentrates on teaching the lead to scope work. Findings. Agents use about 4× the tokens of a chat interaction, and multi-agent systems use about 15×. On BrowseComp, token use alone explains 80% of performance variance. Why it matters. Compare it with Cognition's single-thread position. Parallelism pays when subtasks are genuinely independent and read-only, and the 15× token cost determines whether the fan-out is affordable. |
| scaling agent systemsTowards a Science of Scaling Agent Systems | 2025 |
Background. The frameworks above show how to wire multiple agents, and Anthropic shows one task where parallel workers pay off. Neither establishes which arrangement to choose for a different task. Problem. Comparisons of agent teams often change the model, prompts, tools, or compute budget along with the coordination design. A higher score then cannot show whether the architecture helped, and adding agents remains a guess. Key idea. Hold prompts, tools, and compute fixed while comparing a single agent with independent, centralized, decentralized, and hybrid multi-agent arrangements. Vary model capability and task structure across 260 configurations on six agentic benchmarks. Findings. Relative to a single agent, multi-agent performance ranged from an 80.8% gain on decomposable financial reasoning to a 70.0% loss on sequential planning. Independent agents amplified trace-level errors 17.2×, compared with 4.4× under centralized coordination. Why it matters. It gives the architecture choices above an empirical test. Parallel work can pay when a task decomposes, while sequential work can lose accuracy to coordination; a central verifier also changes how far errors propagate. Choose the topology from the task and its failure modes rather than the number of agents alone. |
| Coordination and communication | ||
| don’t build multi-agentsPrinciples of Context Engineering (Cognition) | 2025 |
Background. The frameworks above assume decomposition helps: split the task, give each part its own agent, run the parts in parallel, and merge. Cognition builds Devin, a long-running coding agent, and reached the opposite conclusion from running one in production. Problem. Most of what an agent decides rests on context that was never written down. Split the work across parallel agents and each one makes those implicit decisions independently, so the pieces disagree and the disagreement surfaces only at the merge. Decomposition can multiply inconsistency faster than useful work. Key idea. Keep a single thread. Carry the whole trajectory forward and compact it when it grows too long, and where a subagent is unavoidable, give it read-only work and take back a summary instead of letting it act. Context engineering, not orchestration, becomes the primary design activity. Why it matters. It is the counterargument to AutoGen and MetaGPT, written by a team shipping a product rather than a paper. Read it as an engineering position paper and ask which of its claims the evaluation papers still need to test. |