Choreography vs Orchestration
A business process like "place an order" spans several services: reserve inventory, charge payment, arrange shipping, send a confirmation. Once these live in separate microservices with their own databases, someone has to coordinate the steps, handle failures, and undo partial work. There are two fundamental styles:
- Choreography: no central coordinator. Each service reacts to events and publishes its own, and the process emerges from these reactions, like dancers who each know their part.
- Orchestration: a central orchestrator tells each service what to do (commands), waits for results, and decides the next step, like a conductor directing an orchestra.
Both implement long-running, distributed processes such as sagas. The choice shapes coupling, visibility, failure handling, and how easy it is to change the process later.
TL;DR
- Choreography: services publish and subscribe to events. It's decoupled and simple for short flows, but the process is implicit, spread across services.
- Orchestration: a coordinator issues commands and tracks state. The process is explicit and easy to see, change, and monitor, but the coordinator is an additional component with more coupling.
- Choreography fits simple, stable flows with few steps. Orchestration fits complex flows with branching, timeouts, compensations, and human steps.
- Workflow engines (durable execution: Temporal, Step Functions, Camunda) make orchestration reliable and code-first.
- Both need idempotency, compensations, and timeouts. Many systems mix them: choreography between domains, orchestration within one.
Quick Example
Choreography: each service reacts to the previous event.
Orchestration: a durable workflow drives the steps (Temporal-style TypeScript).
The workflow engine persists progress, so a crash mid-process resumes exactly where it stopped.
Core Concepts
Choreography
Services subscribe to events they care about and publish events describing their own outcomes. No service knows the whole process.
Strengths
- Loose coupling: producers don't know consumers, and adding a new reaction (a loyalty service awarding points on
OrderPlaced) requires no change to existing services. - No central component to build, scale, or fail.
- Natural for notifications and fan-out reactions.
Weaknesses
- The process is implicit: to understand "how does an order complete?", you read code in five services.
- Cyclic dependencies and event chains grow hard to reason about. A change to the flow can require coordinated changes across teams.
- Failure handling is distributed: each service must know which compensating events to react to.
- Monitoring "where is order 123 stuck?" requires correlating events across services (correlation IDs and tracing help).
Orchestration
A coordinator (a saga orchestrator, process manager, or workflow) owns the process state and sends commands to participants, which reply with results.
Strengths
- The process is explicit in one place: readable, testable, versionable.
- Centralized failure handling: retries, timeouts, compensations, and escalations.
- Visibility: query the workflow's state, see history, and find stuck instances.
- Handles complex logic: branching, parallel steps, waiting days for a human approval or an external callback.
Weaknesses
- The orchestrator knows about participants (more coupling to their APIs).
- It's an extra component that must be highly available, and durable if hand-built.
- Risk of a "god service" that accumulates business logic belonging to domain services.
Comparison
Sagas Both Ways
A saga is a sequence of local transactions with compensating actions for rollback. It can be choreographed (each service listens for failure events and compensates) or orchestrated (the orchestrator calls compensations in reverse order). Orchestrated sagas are easier to reason about as the number of steps grows. See saga pattern.
Workflow Engines and Durable Execution
Hand-building a reliable orchestrator means persisting state, handling retries, timers, and crashes, and versioning in-flight processes. Durable execution engines do this for you:
- Temporal / Restate / DBOS: workflows as code, with persisted history, automatic retries, timers, and signals.
- AWS Step Functions, Azure Durable Functions, Google Workflows: managed state machines.
- Camunda / Zeebe: BPMN-based process orchestration with visual models.
See durable execution.
Reliability Essentials (Both Styles)
- Idempotent handlers and activities, since messages and retries repeat. See idempotency.
- Reliable event publishing, via the outbox pattern.
- Timeouts for every wait: what happens if payment never replies?
- Compensations that are themselves retryable and idempotent. Some actions (sending an email) can't be undone, only followed up.
- Correlation IDs across all messages, for tracing and debugging.
Choosing and Combining
- Use choreography for simple flows (2–4 steps), independent reactions, and cross-domain notifications where the publisher shouldn't care who listens.
- Use orchestration for multi-step business processes with ordering, compensations, deadlines, human tasks, or frequent changes. It's also useful where compliance requires a clear audit of process state.
- Combine: domains communicate via events (choreography between bounded contexts), while each domain uses an orchestrator internally for its own complex processes. An orchestrated workflow can publish events as it progresses, for other domains to react to.
Best Practices
Keep Business Rules in Domain Services
Orchestrators should coordinate the sequence, timeouts, and compensations, not implement pricing rules or eligibility logic that belongs in domain services.
Make the Process Visible
For choreography, build process dashboards from correlated events, or add a lightweight process tracker. For orchestration, expose workflow state and history to support teams.
Design Compensations Up Front
For each step, define how to undo or mitigate it, and what happens if the compensation fails (retry, alert, manual queue).
Version Long-Running Processes
In-flight workflows may run for days. Workflow engines provide versioning APIs, and choreographed flows need backward-compatible events. See event schema evolution.
Common Mistakes
Event Chains Nobody Understands
A 12-step process implemented as choreography becomes a distributed puzzle. When flows grow complex, introduce an orchestrator.
Orchestrator as a Distributed Monolith
Centralizing all logic in a giant orchestrator that micro-manages every service recreates a monolith with network calls. Keep orchestration thin.
No Timeout Handling
Waiting forever for an event that never arrives leaves processes stuck silently. Every step needs a deadline and a fallback.
FAQ
What is the difference between choreography and orchestration?
In choreography, services coordinate by reacting to each other's events, with no central controller. In orchestration, a central coordinator explicitly directs each step by sending commands and tracking the process state. Choreography maximizes decoupling, while orchestration maximizes visibility and control.
Which is better for sagas?
For simple sagas with a few steps, choreography works well. For sagas with many steps, branching, compensations, or timeouts, orchestration is usually easier to build, understand, and operate, especially with a workflow engine.
Is an orchestrator a single point of failure?
It can be, if hand-built naively. Workflow engines and managed services are designed to be highly available and durable: they persist state, so workflows resume after crashes. Treat the orchestrator as critical infrastructure.
Can I use both choreography and orchestration?
Yes, and most mature systems do: event-based choreography between bounded contexts for loose coupling, and orchestration inside a context for complex processes, with orchestrated workflows publishing events that other contexts react to.
Related Topics
- Event-Driven Architecture — Pillar overview
- Saga Pattern — Distributed transactions with compensation
- Durable Execution — Workflow engines for orchestration
- Domain Events — The currency of choreography
- Microservices Communication — Sync vs async interactions
- Outbox Pattern — Reliable event publishing