Load balancing

Follow traffic from anycast and global routing through L4 and L7 balancers, then experiment with IP hashing, sticky sessions, health failures, and backend selection algorithms.

15 min read

“Put a load balancer in front of it” sounds like one box. In a real system a request is balanced several times on its way in: once to find a region, once to find a machine, and again to find the exact service that handles it. Each of those decisions can only use what that hop can see, and that one constraint explains most of the subject.

The load balancer is usually a chain

Follow one request from a browser. Anycast or DNS picks an edge. A network (L4) balancer picks a machine using only IP addresses and ports. An application (L7) gateway reads the HTTP request and picks a service. A service-mesh sidecar may then pick an instance. Four or more decisions, each with a different slice of the request in hand.

ClientBackends
L3/L4 load balancing
Can see
Source and destination IP, TCP or UDP, and port: the 5-tuple
Best used to route
Connections or packets to a healthy service endpoint
Cannot
Read /checkout, a Host header, a cookie or a GraphQL operation
Usually: NAT, DSR, ECMP, TCP and UDP proxies
Each hop balances on what it can see. Pick a stop to compare them.

A note on “L5”: in the OSI model it is the session layer, but “layer 5 load balancer” is not a consistent product category. Real products call themselves L4, L7, a TLS proxy or a session gateway. Ask what a box can inspect and where it terminates the connection, not which layer number it claims.

One IP, several edges: anycast

The first decision often happens before any load balancer sees the request. With anycast, several locations announce the same IP prefix over BGP, and the internet's own routing carries each client to one of them. There is only one address, so DNS has nothing to choose between.

Connect from
A client in London dials the same address as everyone else. BGP carries it to the London edge.
Three edges announce one address. Routing, not DNS, decides which one you reach.

BGP optimises for network topology (path length and peering policy), not kilometres. Usually that lands you somewhere near. Occasionally a London client is routed to Virginia because that is the cheaper path, and no application config will change it. The alternative, geo-DNS, returns a different address per region by policy, but it is slowed by DNS caching and TTLs. Anycast is also why big providers absorb DDoS floods: the attack is spread across every edge announcing the address.

What makes a request stick?

Sometimes you want related requests on the same backend, usually because it holds a session in memory. That is affinity, and the question is what you key it on. With no affinity, round-robin spreads requests anywhere, which is perfect for stateless apps. Source-IP hashing works at L4 but can pile everyone behind one office NAT onto one server. A cookie is precise but needs an L7 balancer that can read HTTP.

Keep a client on one server by
  • Ada from 10.0.1.8server 1
  • Ben from 10.0.1.9server 2
  • Cy from 10.0.1.8server 1
  • Dee from 10.0.2.4server 1
Click a server to fail it, and watch which clients have to move.
Fail a server under source-IP hashing: clients that were never on it move too, because the hash is taken over the servers that are left.

The failure case is the interesting one. Plain hash(ip) % N changes for most clients when N changes, so losing one backend reshuffles sessions that had nothing to do with it. It is the same failure consistent hashing exists to fix (see the hashing lab). The durable fix is not stickier routing but stateless servers, with sessions in a replicated store. Then affinity becomes an optimisation instead of a correctness requirement.

Put the decision where its signal exists

Every balancing goal depends on a piece of information, and that information only exists at certain hops. Routing by URL path needs a hop that has decrypted and parsed HTTP. Routing a raw TCP database connection needs a hop that does not try to. Match the goal to the layer that can see it.

Send users to a healthy regionGlobalGeo-DNS, GSLB or anycast
Balance TCP, UDP, databases or raw socketsL4Network load balancer
Route api.example.com and shop.example.com apartL7Reverse proxy or ingress
Send /images and /api to different poolsL7HTTP-aware proxy
Keep a client on one backendL4 or L7IP hash or affinity cookie
Canary 5% of checkout trafficL7Gateway or service mesh

Whatever the layer, three things decide whether it works in production. Health: probe what you actually promise, a readiness endpoint rather than just an open port. Capacity: set connection limits, timeouts and autoscaling from measured saturation, not CPU alone. Truth: preserve the client's identity on purpose, with the PROXY protocol or trusted forwarded headers, so the backend still knows who it is talking to.

Once traffic arrives, which server gets it?

The last decision is picking a backend from a healthy pool. Round-robin deals requests out in a strict cycle. Random picks any server. Least connections sends each new request to the server with the fewest open right now.

Pick a server by
server 0
server 1
server 2
server 3
0 served · busiest 0 vs quietest 0 in flight · 58% of capacityClick a server to fail it
4 per tick
4
Most requests are quick and a few are slow. Compare busiest and quietest under round-robin, then under least connections.

Round-robin balances arrivals, not load. Requests take wildly different amounts of time, and round-robin cannot see that a server is still chewing on a slow one, so it keeps dealing it work. Random is surprisingly close to round-robin in practice. Least connections steers on the one signal the balancer really has, how many requests each server has open, so busiest and quietest stay close together.

Fail a server and the balancer routes around it on the next request, which is exactly what a health-check failure looks like in production. Push the traffic past 100% of capacity and no algorithm saves you: requests arrive faster than servers drain them. Then the only levers are less traffic or more servers.

The short version

  • A request is balanced several times, and each hop can only use what it can see.
  • Anycast lets the network pick an edge for one shared address.
  • Affinity keyed on hash % N breaks when a server fails; keep servers stateless instead.
  • Least connections beats round-robin when request times vary, but nothing beats missing capacity.