Design nearby friends

Stream a moving user's location to only the friends within a 5-mile radius. Watch a peer-to-peer mesh collapse, push updates through WebSockets and Redis Pub/Sub, filter by distance, and shard the channels across a ring as load explodes.

12 min read

“Nearby friends” shows which of your friends are within five miles right now, updating as everyone moves. Unlike a proximity service, the data never sits still: phones report their location every 30 seconds or so, and the whole design is about pushing a constant stream of updates to exactly the right people, fast.

Why not let the phones talk to each other?

The appealing first idea is peer to peer: each phone sends its location straight to its nearby friends. No servers, no fan-out. Count the connections, though.

Connections in the group
36
Held by each phone
8
8
A mesh grows with the square of the group, and every phone pays for it in battery. A backend needs one connection per phone.

Every phone would hold a live connection to every nearby friend, a mesh that grows with the square of the group, over flaky mobile links on a battery budget. Routing through a backend instead means each phone holds one connection. The backend receives each update, decides which friends care, and forwards it to just them.

One update in, only the friends within five miles out

The rule the backend applies is simple: your update goes only to friends inside a radius around you. Everyone else is filtered out, so a distant friend's phone never wakes up.

AvaBenColeDiaEliFayGusIvyJaxKaiGeofence radius
Drag the handle to resize; double-click the map to move yourself
Friends who get your update
3
Filtered out, nothing sent
7
5 miles
Each location update goes only to friends inside the circle. A distant friend's phone never wakes up.

The scale is what makes it hard. With 100 million daily users, 10% online and an update every 30 seconds, that is about 334,000 location updates a second, and with around 40 nearby friends each, about 14 million forwards a second. Each update travels a short pipeline:

  1. 1Your phone sends its latitude, longitude and a timestamp over its WebSocket about every 30 seconds.
  2. 2Load balancer routes it to the WebSocket server holding your connection.
  3. 3WebSocket server stateful: keeps your connection open, and updates the location cache.
  4. 4Location history appends the new position to a write-heavy store such as Cassandra.
  5. 5Redis Pub/Sub publishes it to your channel, which your friends' handlers subscribe to.
  6. 6Each friend's handler checks the distance and forwards only if you are within range.

The publisher should not know who is listening

The step that does the fan-out is Redis Pub/Sub. Every active user gets a channel, and each of their friends' WebSocket handlers subscribes to it. Publishing to the channel delivers to every subscriber at once: no loop over a friend list, no database read, and the message lives in memory only for that instant.

In range
One publish, five subscribers, 3 forwarded. The rest are dropped by the distance check.
You publish once and never learn who is listening. The channel fans out; each friend's handler decides whether you are close enough.

Pub/Sub delivers to all subscribers, so the distance check still happens at the edge, in each friend's handler. Subscribing is cheap, which is why it is fine to subscribe every online friend and let the check decide; a friend moving into range is just a subscriber whose check starts passing.

Millions of channels do not fit on one Redis

Fourteen million forwards a second is far more than one Redis server can push, so channels are spread across a cluster. Which server owns which channel has to stay stable as the cluster grows, or every resize would disconnect every subscriber.

S1
8
S2
2
S3
4
S4
2
Each channel belongs to the first server clockwise. Add or remove a server and count what moves.
Add a server and only the channels in one arc move. With hash modulo N, almost all of them would.

That is the consistent hashing ring again, applied to Pub/Sub channels instead of cache keys. WebSocket servers learn the current ring through service discovery and re-subscribe only the channels that moved. The bottleneck of the whole system is this fan-out CPU, not storage, so the bus is what you scale first.

WebSocket serversStateful: they hold every live connection. Scale up on connection count; scaling down means draining.
API serversStateless HTTP for sign-in, friends and profiles. Auto-scale freely.
Redis location cacheEach active user's latest position, with a TTL so stale users expire.
Redis Pub/Sub clusterThe routing layer, bound by CPU rather than memory; shard channels on the ring.
Location historyEvery position over time; write-heavy, sharded by user.
User databaseProfiles and the friend graph; replicated and sharded by user.

The short version

  • Phones connect to a backend, one socket each, rather than to each other.
  • Updates go only to friends inside the radius; everyone else is filtered out.
  • Each user publishes to their own Pub/Sub channel; friends' handlers subscribe and check distance.
  • Channels are sharded across the Pub/Sub cluster on a consistent-hash ring.