Service Discovery

In a static world, service A calls service B at a fixed hostname and port. In microservices running on autoscaling infrastructure, instances come and go constantly: deployments replace them, autoscalers add and remove them, failures kill them, and IP addresses change every time. Service discovery answers "where are the healthy instances of B right now?", continuously and automatically.

On Kubernetes, discovery is largely built in through Services and DNS. Outside Kubernetes, or across clusters and hybrid environments, registries like Consul, Eureka, and cloud-native discovery services fill the role. Service meshes build on discovery to add load balancing, mTLS, and traffic policy.

TL;DR

Quick Example

Discovery on Kubernetes: a Service gives a stable name for a changing set of pods.

Registration with Consul outside Kubernetes (agent service definition):

Core Concepts

The Service Registry

A registry is a database of service instances: name, address, port, metadata (version, zone, tags), and health. It must be highly available, since it's on the critical path of every call. That's why registries like Consul, etcd, and ZooKeeper use consensus, and why clients cache results.

Client-Side vs Server-Side Discovery

Service meshes blur the line: a sidecar proxy next to each service does client-side-style balancing, transparently to the application.

Registration Patterns

Health Checking

Registries must stop routing to unhealthy instances:

Deregister promptly on shutdown, and drain connections, to avoid errors during deployments.

DNS-Based Discovery

DNS is the universal discovery interface: A or AAAA records for instance IPs, SRV records for ports, and short TTLs for freshness. Kubernetes (CoreDNS), Consul, and AWS Cloud Map all expose DNS. Limitations: client DNS caching can serve stale addresses, and plain DNS doesn't carry health or metadata richly. Many clients combine DNS with health-aware registries or proxies. See DNS.

Kubernetes Discovery

See Kubernetes networking.

Service Mesh and Multi-Cluster

A service mesh consumes discovery data to provide per-request load balancing (least requests, locality-aware), retries and timeouts, mTLS identity, and traffic splitting for canaries. For multi-cluster or hybrid deployments, discovery must span clusters: mesh multi-cluster modes, Consul federation, Kubernetes multi-cluster Services (MCS API), or global load balancers with health-aware DNS.

Best Practices

Let the Platform Handle Registration

On Kubernetes, rely on Services and readiness probes rather than embedding registry clients in apps. Outside Kubernetes, use agents or orchestrator integrations (Consul with Nomad, ECS with Cloud Map) for third-party registration.

Make Health Checks Meaningful

Readiness should reflect ability to serve (warm-up done, critical dependencies reachable), and liveness only process health. Poor checks route traffic to broken instances, or remove healthy ones during dependency blips.

Cache and Degrade Gracefully

Clients should cache discovery results and keep using last-known-good instances if the registry is briefly unavailable. A registry outage shouldn't become a total outage.

Prefer Locality-Aware Routing

Route to instances in the same zone when possible, which reduces latency and cross-zone data transfer costs, with failover to other zones when local capacity is unhealthy.

Common Mistakes

Hard-Coding IPs or Hostnames per Instance

Static lists of instance addresses break with every deploy and autoscaling event. Use service names resolved by discovery.

Long DNS TTLs and Client Caching

Clients (notably older JVMs with default DNS caching) holding stale addresses keep calling terminated instances. Use short TTLs, and configure client DNS caching appropriately.

Long-Lived Connections Ignoring New Instances

HTTP/2 and gRPC connections established to a few instances don't rebalance when new ones appear. Use client-side load balancing with periodic re-resolution, connection max-age settings, or a mesh.

FAQ

What is service discovery?

The mechanism by which services locate the network addresses of healthy instances of other services in dynamic environments where instances start, stop, and move frequently. It typically involves a registry of instances, health checking, and a lookup method (DNS, API, or a proxy).

Do I need Consul or Eureka on Kubernetes?

Usually not. Kubernetes Services, EndpointSlices, and cluster DNS provide discovery and basic load balancing out of the box. Consul or a mesh becomes useful for multi-cluster, hybrid (VMs plus Kubernetes), or advanced traffic management scenarios.

What's the difference between client-side and server-side discovery?

In client-side discovery, the calling service queries the registry and chooses an instance itself. In server-side discovery, the caller sends requests to a load balancer or proxy that looks up instances and routes the request. Server-side keeps clients simple; client-side avoids an extra hop and enables smarter balancing.

How do health checks relate to discovery?

Discovery should only return instances able to serve traffic. Health checks (active probes, or heartbeats) mark instances healthy or unhealthy, and registries and load balancers exclude unhealthy ones automatically, so failures don't reach callers.

Related Topics

References