Nginx Load Balancing
When one application server isn't enough, or you need redundancy so one failure doesn't take the site down, Nginx can distribute requests across a pool of backends defined in an upstream block. It supports several balancing methods (round robin, least connections, hashing), weights for uneven hardware, passive health checks that stop sending traffic to failing servers, retries on another server, backup servers, and connection reuse.
Load balancing is a core building block of scalability and high availability. Nginx's open-source edition covers most needs; NGINX Plus adds active health checks, dynamic reconfiguration, and cookie-based session persistence.
TL;DR
- Define backends in an
upstreamblock and pointproxy_passat it. - The default method is round robin; use
least_connfor uneven request durations, andhash … consistentorip_hashfor affinity. weightsends proportionally more traffic to bigger servers;backupservers only receive traffic when the others are down.- Passive health checks:
max_failsandfail_timeouttemporarily remove failing servers;proxy_next_upstreamretries safely on another server. keepalivereuses backend connections for lower latency.- Prefer stateless backends over sticky sessions. Use
split_clientsor weights for canary rollouts.
Quick Example
Core Concepts
Balancing Methods
hash $request_uri consistent uses consistent hashing, so adding or removing a server remaps only a fraction of keys, which is important when backends hold caches.
Server Parameters
Passive Health Checks and Retries
Open-source Nginx detects failures passively, from real requests: connection errors, timeouts, and (if configured) specific status codes. After max_fails, the server is skipped for fail_timeout. proxy_next_upstream controls when a failed request is retried on another server:
- By default, only
errorandtimeouttrigger retries, and non-idempotent requests (POST, PATCH) aren't retried once sent, unless you addnon_idempotent. Leave that off unless the backend handles duplicates. See idempotency. - Limit retries with
proxy_next_upstream_triesandproxy_next_upstream_timeoutto avoid amplifying load during incidents.
Active health checks (periodic probes to /health) require NGINX Plus, or external tooling such as orchestrator health checks or the community nginx_upstream_check_module. In Kubernetes, readiness probes remove unhealthy pods from Service endpoints before Nginx ingress ever sees them.
Keepalive to Upstreams
Without keepalive, Nginx opens a new TCP (and possibly TLS) connection to a backend per request. The keepalive N directive keeps up to N idle connections per worker for reuse. It requires proxy_http_version 1.1 and clearing the Connection header.
Session Affinity
Sticky sessions route a user to the same backend, which is sometimes necessary for legacy apps with in-memory sessions. Options: ip_hash (breaks with NAT and mobile networks), hash $cookie_sessionid consistent, or NGINX Plus sticky cookie. Better: make backends stateless, keeping sessions in Redis or the database, so any server can handle any request and failover is seamless.
Canary and Blue-Green Routing
split_clients sends a stable percentage of clients to the canary pool. Blue-green deployments swap which upstream proxy_pass points to, followed by a reload. See deployment strategies.
Best Practices
Choose least_conn for Mixed Workloads
When some requests take milliseconds and others seconds (reports, uploads, AI calls), round robin piles slow requests onto the same servers. least_conn balances by actual load.
Retry Only Idempotent Requests
Retrying POSTs on another server after a timeout can create duplicate orders or payments. Keep the default non-idempotent behavior, or make endpoints idempotent with idempotency keys.
Drain Servers Before Maintenance
Mark a server down (or remove it) and reload, wait for in-flight requests to finish, then deploy. Orchestrators automate this with readiness gates and graceful termination.
Monitor Upstream Health
Log $upstream_addr, $upstream_status, and $upstream_response_time in access logs, and track per-backend error rates and latency. Uneven latency across backends often reveals a bad node. See monitoring.
Common Mistakes
Forgetting keepalive Requirements
Add proxy_http_version 1.1; and proxy_set_header Connection "";.
ip_hash Behind a CDN or Corporate NAT
Many users share a few egress IPs, so ip_hash concentrates traffic on a few backends. Hash on a user or session identifier instead, or eliminate the need for affinity.
Infinite Retry Amplification
Allowing retries across all servers on timeouts during an overload multiplies traffic exactly when backends are struggling. Cap tries and total retry time.
FAQ
What load balancing algorithms does Nginx support?
Round robin (default), least connections, IP hash, generic hash (optionally consistent), and random (with "two choices" variants). All support server weights. NGINX Plus adds least time (lowest latency) and additional session persistence options.
Does open-source Nginx do health checks?
Passive ones: it marks servers unavailable after max_fails failed requests within fail_timeout. Active health checks, periodic probes independent of traffic, are an NGINX Plus feature, or can come from your orchestrator or service discovery.
How do sticky sessions work in Nginx?
Open-source Nginx can approximate stickiness with ip_hash, or hash on a cookie or header value with consistent. NGINX Plus offers cookie-based sticky directives. Stateless backends with shared session storage avoid the need altogether.
Can Nginx load balance TCP and UDP?
Yes. The stream module balances arbitrary TCP and UDP traffic (databases, DNS, MQTT, custom protocols) with the same upstream concepts, outside the http context.
Related Topics
- Nginx — The web server overview
- Nginx Reverse Proxy — Proxying fundamentals
- Load Balancing — Algorithms and architectures in general
- Consistent Hashing — Minimal remapping for cache-friendly affinity
- High Availability — Redundancy and failover
- Deployment Strategies — Canary and blue-green releases