Kubernetes Networking

Kubernetes networking answers three questions: how Pods talk to each other, how clients find a set of Pods that keeps changing, and how traffic gets in from outside the cluster. The answers are the Pod network, Services, and Ingress / Gateway API, with NetworkPolicies on top to decide who may talk to whom.

The model is simple on paper. Every Pod gets its own IP and can reach every other Pod without NAT. Most confusion comes from the layers built on top of that model, so this page walks through them in the order a packet meets them.

TL;DR

Quick Example

Expose a Deployment internally with a Service, then publicly with an Ingress:

Inside the cluster, other Pods call http://api (same namespace) or http://api.prod.svc.cluster.local. Outside, clients call https://api.example.com.

Core Concepts

The Pod Network

Kubernetes requires that all Pods can reach each other by IP without NAT, and that agents on a node can reach all Pods on that node. It doesn't implement this itself; a CNI (Container Network Interface) plugin does. Overlay CNIs such as Flannel's VXLAN encapsulate traffic between nodes. Routed CNIs such as Calico's BGP mode or cloud VPC CNIs give Pods real VPC IPs. eBPF-based CNIs like Cilium replace much of kube-proxy with kernel programs for speed and visibility (see eBPF).

Services

Pods come and go, so you address them through a Service. A Service selects Pods by label, and the control plane maintains an EndpointSlice listing the IPs of matching Pods that are ready. kube-proxy (or the CNI) programs iptables, IPVS, or eBPF rules on every node so traffic to the Service's virtual IP is load-balanced to one of those endpoints.

Cluster DNS

CoreDNS runs in the cluster and serves records for Services (api.prod.svc.cluster.local) and, for headless Services, individual Pods (postgres-0.postgres.prod.svc.cluster.local). Pods get a resolv.conf with search domains so a short name like api resolves within the same namespace. See DNS for the underlying protocol.

Ingress and the Gateway API

An Ingress resource declares HTTP routing rules (host, path, TLS); an Ingress controller such as ingress-nginx, Traefik, HAProxy, or a cloud controller watches those resources and configures a real proxy. The Gateway API is the successor: it splits concerns into GatewayClass (the infrastructure), Gateway (listeners, owned by the platform team), and HTTPRoute, GRPCRoute, and similar (routing, owned by app teams). It also supports traffic splitting, header matching, and cross-namespace routing without controller-specific annotations.

NetworkPolicies

By default all Pods accept traffic from anywhere. A NetworkPolicy selects Pods and whitelists ingress and/or egress by Pod label, namespace, or CIDR. Once any policy selects a Pod for a direction, everything not explicitly allowed in that direction is denied. Policies are enforced by the CNI, so your plugin must support them (Calico and Cilium do; basic Flannel doesn't).

Locking Down Traffic

A common baseline is to default-deny ingress in each namespace, then allow specific paths:

This is the in-cluster half of zero trust; a service mesh adds mutual TLS and identity-based authorization on top.

Best Practices

Use ClusterIP by Default

Most Services should be internal. Expose the edge through one Ingress or Gateway (with TLS, rate limits, and auth) instead of a LoadBalancer Service per app. That saves money on cloud load balancers and gives you one place to enforce policy.

Rely on Readiness, Not Luck

Services only route to ready Pods. A good readiness probe is what makes rolling updates and scale-ups invisible to clients. Without one, traffic hits Pods that are still booting.

Default-Deny, Then Allow

Start every namespace with a deny-all policy and add explicit allows. Include an egress policy that allows DNS (UDP/TCP 53 to CoreDNS), or everything breaks in confusing ways.

Prefer the Gateway API for New Setups

It's portable across implementations and replaces piles of vendor-specific Ingress annotations with typed fields. Existing Ingress setups keep working, so there's no rush to migrate.

Common Mistakes

Selector and Label Mismatch

Check with kubectl get endpointslices -l kubernetes.io/service-name=api. An empty list means the selector matches nothing or no Pods are ready.

Confusing port and targetPort

port is what clients connect to on the Service; targetPort is the container port. Getting them backwards produces "connection refused" even though everything looks healthy.

Long-Lived Connections Not Rebalancing

kube-proxy balances connections, not requests. HTTP/2 and gRPC clients hold one connection open, so all requests from a client stick to one Pod even after you scale up. Use client-side load balancing with a headless Service, or a mesh or L7 proxy that balances per request.

FAQ

What's the difference between Ingress and a LoadBalancer Service?

A LoadBalancer Service provisions one cloud L4 load balancer for one Service. Ingress is an L7 router: one entry point that routes many hostnames and paths to many Services, with TLS termination. Most clusters run one LoadBalancer Service for the ingress controller and route everything else through it.

Should I migrate from Ingress to the Gateway API?

For new platforms, start with the Gateway API; it's the direction upstream is investing in, and ingress-nginx has entered maintenance mode. For existing clusters, migrate when you need something Ingress can't express cleanly, such as weighted canaries, header routing, or delegated ownership.

Do I need a service mesh?

Not to get started. Services, Ingress, and NetworkPolicies cover most needs. A mesh adds mTLS everywhere, per-request load balancing, retries, and deep traffic telemetry, at the cost of operational complexity. Adopt one when those needs are concrete.

How do I debug "can't connect to service"?

Work outward: is the Pod ready? Does the Service have endpoints? Does DNS resolve (nslookup api from a debug Pod)? Can you curl the Pod IP directly? Is a NetworkPolicy blocking it? kubectl debug with a netshoot image is the standard toolbox.

Related Topics

References