Multi-Agent Systems

A multi-agent system splits work across several LLM-driven agents, each with its own instructions, tools, and context, coordinated by code or by another agent. A research system might have a lead agent that plans and several subagents that search different angles in parallel. A customer service system might route conversations between billing, technical, and account agents via handoffs. A coding agent might spawn subagents to explore a codebase without cluttering its own context.

Multiple agents can deliver real gains: parallelism, context isolation (each agent works in a clean window), and specialization. They also multiply cost, latency, and failure modes. A well-designed single agent often beats a poorly coordinated team, so the key skill is knowing when splitting work helps.

TL;DR

Quick Example

An orchestrator that fans research out to parallel subagents and synthesizes the results:

Each subagent spends its own context on searching and reading; the orchestrator sees only the distilled briefs.

Core Concepts

Architectures

Frameworks support these directly: LangGraph (graphs of agents), the OpenAI Agents SDK (handoffs), the Claude Agent SDK (subagents), CrewAI, and AutoGen. See AI agent frameworks.

Why Multiple Agents Help

Anthropic reported that its multi-agent research system substantially outperformed a single agent on broad research tasks, while using many times more tokens. That trade is worth it for high-value tasks, and wasteful for simple ones.

Communication and State

Failure Modes

Debugging requires tracing every agent's prompts, tool calls, and outputs as one linked trace. See LLMOps.

When to Use Multiple Agents

Good fits:

Poor fits:

Best Practices

Start With One Agent

Build a strong single agent first. Split only when you hit concrete limits (context overflow, slow sequential exploration, conflicting tool needs) that multiple agents would solve.

Design Delegation Like an API

Treat subagent tasks as function contracts: inputs, expected output schema, scope limits, and effort guidance (how many tool calls are reasonable). Vague delegation is the most common failure.

Scale Effort to Task Complexity

Simple questions need one agent and a few tool calls; complex ones may need several subagents. Teach the orchestrator explicit heuristics for how many subagents to spawn, so it doesn't over-invest in easy tasks.

Evaluate End to End and per Agent

Measure final task success, cost, and latency against a single-agent baseline, and inspect individual agents' behavior to find weak links. See LLM evaluation.

Common Mistakes

Role-Play Teams Without a Reason

Creating "CEO", "PM", "engineer", and "QA" agents for a task one agent could do adds conversation overhead without new capability. Agents should exist because of isolation, parallelism, or specialization needs, not org-chart aesthetics.

Passing Full Transcripts Between Agents

Forwarding a subagent's entire conversation to the orchestrator defeats context isolation. Summarize, or store artifacts externally and pass references.

Parallel Agents Writing to the Same Resources

Concurrent subagents editing the same files or records produce conflicts and lost work. Partition ownership, or have subagents propose changes that one agent applies.

FAQ

Are multi-agent systems better than single agents?

For broad, parallelizable tasks, often yes: they explore more in less wall-clock time and keep contexts clean. For tightly coupled or simple tasks, a single agent is usually as good or better, and much cheaper. Decide with evaluations, not by default.

How do agents communicate with each other?

Most commonly through the orchestrating code: a lead agent calls subagents like tools, passing task descriptions and receiving results. Alternatives include shared state (files, databases), message queues, and standardized protocols like A2A for cross-system agent communication.

Why do multi-agent systems cost so much?

Each agent consumes its own tokens, re-reading instructions, calling tools, and reasoning, and orchestration adds more calls. Several agents working in parallel can easily use many times the tokens of one agent. Reserve them for tasks whose value justifies the spend, and use cheaper models for simpler roles.

What's the difference between a subagent and a tool?

Functionally, an orchestrator often invokes a subagent as a tool. The difference is that a subagent is itself an LLM loop with its own instructions, tools, and context, able to take many steps before returning. An ordinary tool executes a single deterministic operation.

Related Topics

References