AI Agent Design Patterns
"Agent" covers a spectrum, from a single LLM call with retrieval, to fixed multi-step workflows, to fully autonomous systems where a model decides its own steps in a loop. Most production value comes from choosing the simplest pattern that solves the problem: predictable workflows where the steps are known, and autonomous agents only where flexibility is genuinely needed and worth the extra cost and unpredictability.
This page catalogs the core patterns, largely following Anthropic's "Building effective agents" taxonomy, plus classic techniques like ReAct, planning, and reflection. They're framework-agnostic: you can implement them in a few dozen lines of code or with any agent framework.
TL;DR
- Building block: the augmented LLM, a model with retrieval, tools, and memory.
- Workflows (predefined code paths): prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer.
- Agents: the model dynamically directs its own process in a loop of reason → act (tool) → observe → repeat.
- ReAct interleaves reasoning and actions; plan-and-execute plans first, then carries out steps; reflection critiques and revises.
- Add human-in-the-loop checkpoints for high-stakes actions.
- Start simple, measure, and add complexity only when evaluation shows it helps.
Quick Example
An evaluator-optimizer workflow: generate, critique against criteria, revise until it passes (or a limit is hit).
And a minimal autonomous agent loop:
Core Concepts
The Augmented LLM
Every pattern builds on one unit: an LLM that can retrieve information (RAG), call tools, and use memory. Often, a well-designed augmented LLM call, with good retrieval and a clear prompt, is all an application needs.
Workflow Patterns
Workflows are predictable, testable, and cheaper. Control flow lives in your code, and the LLM handles the fuzzy parts.
Autonomous Agents
An agent runs a loop where the model chooses the next action based on observations, until it decides the task is done or hits limits. Agents fit open-ended problems where the number and order of steps can't be predicted, such as debugging, research, multi-system operations, and coding. The costs: higher latency and token spend, compounding errors, and harder evaluation. They need good tools, clear stopping criteria, sandboxing, and guardrails.
ReAct
ReAct (Reason + Act) interleaves reasoning traces with tool actions: think about what's needed → call a tool → observe → think again. Modern tool-calling models do this natively, and reasoning models and extended thinking make the reasoning step stronger (see reasoning models). It's the default loop behind most agents.
Planning Patterns
- Plan-and-execute: the model first writes an explicit plan (a to-do list), then executes steps, updating the plan as results come in. It improves coherence on long tasks, and it makes progress visible and resumable.
- Hierarchical planning: a planner decomposes the task, and executors (sometimes subagents) handle the pieces. See multi-agent systems.
- Re-planning on failure: when an observation contradicts the plan, revise it rather than pushing forward.
Reflection and Self-Correction
After producing output, the model (or a separate critic) reviews it against the task, tests, or criteria, then revises. Reflection works best with external feedback (test results, linters, validators, retrieved facts) rather than pure self-critique, which tends to be less reliable.
Human-in-the-Loop
Insert approval checkpoints before irreversible actions (payments, deployments, emails to customers), at plan approval for long tasks, and when confidence is low. Well-placed checkpoints let you grant agents more autonomy elsewhere.
Choosing a Pattern
- Can a single augmented LLM call do it? Stop there.
- Are the steps known in advance? Use a workflow: chaining, routing, or parallelization.
- Are subtasks dynamic but the overall goal clear? Use orchestrator-workers.
- Is there a clear quality criterion and iteration improves results? Add evaluator-optimizer.
- Is the path genuinely open-ended, with tools and feedback available? Use an agent, with limits, sandboxing, and checkpoints.
Best Practices
Invest in the Agent-Computer Interface
Tool design, error messages, and environment feedback matter more than clever orchestration. Agents succeed when tools are clear and results informative. See tool calling.
Keep Control Flow in Code Where Possible
Deterministic code for sequencing, retries, validation, and routing is easier to test and debug than asking the model to manage it. Let the model handle judgment, language, and ambiguity.
Set Budgets and Stopping Conditions
Cap steps, tokens, time, and cost per task, detect loops (repeated identical actions), and define what "done" means (tests pass, criteria met).
Evaluate Each Pattern Change
Measure task success, cost, and latency before and after adding complexity. Multi-step patterns often raise quality for hard tasks and waste money on easy ones, so route accordingly. See LLM evaluation.
Common Mistakes
Starting With a Fully Autonomous Multi-Agent System
Complex architectures are hard to debug and often underperform a simple workflow with good prompts and tools. Build up from the simplest version.
Hiding Prompts Inside Frameworks
Heavy abstractions can obscure the actual prompts and tool definitions sent to the model, which makes failures mysterious. Make sure you can inspect every request, or implement core loops directly.
Unverified Self-Reflection
"Review your answer and fix any errors" without external signals often just rephrases the same mistakes. Give critics concrete criteria, reference data, or executable checks.
FAQ
What's the difference between a workflow and an agent?
In a workflow, your code defines the sequence of LLM calls and tool uses; the model fills in each step. In an agent, the model decides which actions to take and in what order, looping until it judges the task complete. Workflows are more predictable; agents are more flexible.
Is ReAct still relevant with tool-calling models?
The idea is. Modern tool-calling APIs and reasoning models implement the reason-act-observe loop natively, so you rarely need ReAct's original text-parsing prompt format, but the pattern itself underlies nearly every agent loop.
Do I need an agent framework?
Not necessarily. Many patterns are a few dozen lines on top of a model API. Frameworks (LangGraph, the OpenAI Agents SDK, the Claude Agent SDK, CrewAI, and others) help with state management, persistence, tracing, and multi-agent orchestration. Adopt one when those needs are real, and make sure you understand what it sends to the model. See AI agent frameworks.
How do I make agents more reliable?
Better tools and descriptions, clearer instructions and success criteria, external feedback (tests, validators), limited autonomy with checkpoints, strong models or reasoning for hard steps, and systematic evaluation on realistic tasks, with transcript review to find failure modes.
Related Topics
- AI Agents — What agents are and how they work
- Tool Calling — The action mechanism in every pattern
- Multi-Agent Systems — Orchestrating several agents
- Agent Memory — State across steps and sessions
- AI Agent Frameworks — Libraries implementing these patterns
- Reasoning Models — Stronger planning and reflection