Microservices Communication

Splitting a system into microservices turns in-process function calls into network calls, and networks are slow, unreliable, and asynchronous. How services communicate (synchronously over HTTP or gRPC, or asynchronously through messages and events) shapes latency, availability, coupling, and how failures spread. Many microservice failures, such as cascading outages, distributed monoliths, and inconsistent data, trace back to communication design.

The core decision is synchronous vs asynchronous. Synchronous calls are simple and give immediate answers, but they couple availability: if B is down, A fails. Asynchronous messaging decouples services in time, and absorbs load spikes, at the cost of eventual consistency and harder debugging. Most real systems use both, deliberately.

TL;DR

Quick Example

Checkout combining sync and async communication:

The user waits only for pricing and payment. Stock reservation, email, analytics, and loyalty happen asynchronously and independently, so a slow email service never delays checkout.

Core Concepts

Synchronous Communication

Synchronous calls couple availability and latency: the caller's response time includes every downstream call, and the caller's availability is roughly the product of its dependencies' availability. Five services at 99.9% each chained together gives about 99.5%.

Asynchronous Communication

Benefits: temporal decoupling (consumers can be down temporarily), load leveling (queues absorb spikes), and easy fan-out to new consumers. Costs: eventual consistency, duplicate and out-of-order delivery, harder tracing, and operating a broker. See message queues, event-driven architecture, and Kafka.

Choosing Sync vs Async

Resilience for Network Calls

Edge Patterns: API Gateway and BFF

Internally, service meshes can add mTLS, retries, and traffic policies without library code. See service mesh.

Contracts and Versioning

Best Practices

Minimize Synchronous Call Depth

Design so a user request touches few services synchronously. Deep chains (A→B→C→D) multiply latency and failure rates. Denormalize data via events, or restructure service boundaries.

Prefer Events for Cross-Domain Side Effects

When an action in one domain triggers work in others, publish an event rather than calling each service. New consumers can then be added without changing the publisher.

Propagate Context

Pass trace context, correlation IDs, and deadlines across calls and messages, so failures can be traced end to end. See distributed tracing.

Design for Duplicates and Reordering

Asynchronous consumers must handle messages delivered more than once and occasionally out of order: idempotent handlers, version checks, and upserts.

Common Mistakes

The Distributed Monolith

Services that must all be deployed together, call each other synchronously for every operation, and share databases have microservice costs without the benefits. Revisit boundaries, and reduce synchronous coupling.

No Timeouts

A downstream service hanging with no caller timeouts ties up threads and connections everywhere, and one slow dependency takes down the whole system.

Chatty Interfaces

Fetching a list, then calling another service once per item (the N+1 problem across the network) destroys performance. Provide batch endpoints, or aggregate data via events.

FAQ

Should microservices communicate with REST or gRPC?

REST is universal, simple to debug, and great for public and edge APIs. gRPC offers strongly typed contracts, better performance, and streaming, which suits internal service-to-service traffic. Many organizations use REST or GraphQL at the edge and gRPC internally.

When should microservices use messaging instead of HTTP?

When the caller doesn't need an immediate result, several services react to the same occurrence, workloads are bursty, or you want services to keep working when others are temporarily unavailable. Use HTTP or gRPC for queries and interactions that need an answer now.

What is an API gateway for?

It's the single entry point for external clients, handling cross-cutting concerns (routing, authentication, rate limiting, TLS, aggregation, logging) so individual services don't each implement them. It also shields clients from internal service topology changes.

How do I avoid cascading failures between services?

Set timeouts on every call, use circuit breakers and bulkheads, limit retries with backoff and jitter, provide fallbacks, prefer asynchronous communication for non-critical work, and load-test failure scenarios. Chaos engineering helps validate these defenses.

Related Topics

References