Circuit Breakers & Resilience Patterns
In a distributed system, something is always failing: a database slows down, a dependency deploys a bug, a network link drops packets. Without protection, one slow dependency can exhaust threads and connections in every caller, which then time out their own callers, a cascading failure that takes down healthy services too. Resilience patterns contain failures: timeouts bound waiting, retries absorb transient errors, circuit breakers stop calling a dependency that's clearly failing, bulkheads isolate resources, and fallbacks degrade gracefully.
These patterns are essential in microservices, and useful anywhere code calls over a network. They're implemented in libraries (Resilience4j, Polly, failsafe-go, and Hystrix before it was retired) or in infrastructure (Envoy, service meshes).
TL;DR
- Timeouts on every remote call, derived from latency budgets. They're the foundation of every other pattern.
- Retries only for transient failures and idempotent operations, with exponential backoff + jitter and caps and budgets.
- Circuit breaker: after a failure threshold it opens (fails fast), later lets trial calls through (half-open), and closes on success.
- Bulkheads limit concurrent calls per dependency, so one bad dependency can't consume all resources.
- Fallbacks return cached, default, or partial results when a dependency is unavailable.
- Load shedding and rate limiting protect services from overload. Test all of it with chaos experiments.
Quick Example
Resilience4j in a Spring Boot service calling a recommendations API:
Core Concepts
Timeouts
Every network call needs a connect timeout and a request (read) timeout. Derive them from the caller's overall latency budget: if a request must finish in 1s and makes two sequential calls, neither can take 1s. Deadline propagation passes the remaining budget downstream (gRPC deadlines do this natively), so downstream services don't keep working on requests the caller has already abandoned.
Retries
Retries turn transient blips into successes, and they amplify load during real outages. Rules:
- Retry only transient errors (connection resets, 503, 429 with
Retry-After, timeouts), and only idempotent operations, or ones with idempotency keys. - Use exponential backoff with jitter: randomized delays spread retries out and prevent synchronized retry storms.
- Cap attempts (2–3), and use retry budgets (for example, retries at most 10% of requests) to avoid multiplying traffic.
- Retry at one layer. Retries at every layer (client, gateway, service, mesh) multiply exponentially: 3 × 3 × 3 = 27 attempts.
Circuit Breaker States
Benefits: callers stop wasting threads and time on a failing dependency, the dependency gets breathing room to recover, and users get fast responses (fallbacks or errors) rather than timeouts.
Bulkheads
Named after ship compartments: isolate resources per dependency, so one failing dependency can't sink the whole service:
- Semaphore bulkheads limit concurrent calls per dependency.
- Thread pool bulkheads give separate pools to separate dependencies.
- Connection pool limits per downstream service.
- At the infrastructure level, separate deployments, node pools, or cells for critical and non-critical workloads.
Fallbacks and Graceful Degradation
When a dependency is unavailable, return something useful: cached data (possibly stale), defaults (generic recommendations), partial responses (the page without the reviews widget), or queuing work for later. Not everything can degrade (a payment can't succeed without the payment provider), but many features can, and users prefer a partially working page to an error page.
Load Shedding and Rate Limiting
When a service is overloaded, accepting more work makes everything slower until nothing succeeds. Load shedding rejects excess requests early (fast 503 or 429 responses), prioritizing critical traffic. Rate limiting caps per-client usage. Adaptive concurrency limits (for example Netflix's concurrency-limits, and Envoy's adaptive concurrency) adjust limits from observed latency.
Library vs Infrastructure
Many teams combine them: mesh-level timeouts, retries, and outlier detection, plus library-level fallbacks where business logic matters.
Best Practices
Set Timeouts Everywhere, Consistently
Audit every HTTP client, database driver, and messaging client for timeouts. Defaults are frequently infinite, or far too long.
Make Breakers Observable
Export circuit breaker state changes, failure rates, rejected calls, and retry counts as metrics, and alert when breakers open. An open breaker is a symptom worth investigating. See monitoring.
Test Failure Modes Deliberately
Inject latency, errors, and outages in staging and, carefully, in production (chaos engineering), and verify that timeouts fire, breakers open, fallbacks work, and alerts trigger.
Tune From Data
Base thresholds on observed latency percentiles and error rates. Breakers that open too eagerly cause self-inflicted outages; ones that never open provide no protection.
Common Mistakes
Retrying Non-Idempotent Operations
Retrying a timed-out "charge card" call without an idempotency key can charge customers twice. Make writes idempotent, or don't retry them automatically.
Retries Without Backoff and Jitter
Immediate retries from thousands of clients hit a recovering service at the same instant, knocking it over again. Always use backoff with jitter.
Circuit Breakers Without Timeouts
A breaker counts failures, but if calls hang for 60 seconds before failing, threads are exhausted long before the breaker opens. Timeouts come first.
FAQ
What does a circuit breaker do?
It monitors calls to a dependency, and when failures or slow responses exceed a threshold, it "opens", immediately failing further calls (or serving fallbacks) instead of waiting on a broken dependency. After a cooldown, it lets trial requests through, and closes again if they succeed. It prevents cascading failures and gives dependencies time to recover.
What's the difference between a retry and a circuit breaker?
Retries handle brief, transient failures by trying again. Circuit breakers handle sustained failures by stopping attempts altogether for a while. They complement each other: retries inside a closed circuit, and no retries once the circuit is open.
Where should resilience logic live: in code or in the service mesh?
Both have roles. Meshes provide consistent, language-agnostic timeouts, retries, and outlier detection. Application libraries add business-aware fallbacks and fine-grained control. Avoid duplicating retries across both layers, which multiplies load.
What is a bulkhead?
A pattern that partitions resources (threads, connections, concurrency slots) per dependency or workload, so a failure or slowdown in one area can't consume resources needed by others, like watertight compartments in a ship.
Related Topics
- Microservices — The architecture overview
- Microservices Communication — Where these patterns apply
- Idempotency — Making retries safe
- Rate Limiting — Protecting services from overload
- Chaos Engineering — Testing resilience
- High Availability — Designing for failure