Rohit Swami
India Resume ↗

Writing · Real-time systems · 5 min read

Push, don't poll

Getting a notification from a background job to a browser in under 100 milliseconds is easy once. The hard parts arrive when it has to happen for everyone at once.

Plenty of software does its real work in the background. Someone starts an analysis, a worker picks it up, and minutes later it finishes. Between those two moments there's a person with a browser tab open, waiting to find out.

Getting the news to them sounds trivial. I architected a notification microservice on Redis and WebSockets that delivers in under 100 milliseconds, and the interesting parts weren't Redis or WebSockets, which are both well understood. They were three problems that only appear when the thing has to work for a lot of people at once.

1The cost of asking

The obvious design is polling. Every open tab asks the server, every few seconds, whether anything has happened. It's easy to build, easy to reason about, and it works perfectly in a demo.

It has two costs, and they pull in opposite directions. The first is latency: poll every T seconds and news waits on average T/2 before anyone asks for it, and up to T in the worst case. The second is load: a thousand tabs polling every two seconds is five hundred requests a second, and nearly all of them get the answer "nothing yet". Improve one and you worsen the other.

every
users

requests/s –typical wait –empty answers –

Fig. 1 Ten of the users, over one minute. Grey ticks are polls; red lines are updates waiting to be asked for. The readouts are for the whole population, assuming each user gets an update about every five minutes.

Pushing inverts this. Each browser opens one long-lived WebSocket, and the server sends a message when, and only when, there's something to say. Latency becomes one trip across the network. Load becomes proportional to things actually happening, rather than to the number of people waiting for them.

2One connection, many servers

The first problem appears as soon as there's more than one server. Connections are spread across gateway instances behind a load balancer, so each user's socket lives on one particular instance. The worker that finishes a job has no idea which one. Hand the event to the wrong instance and there's nobody there to receive it.

Redis pub/sub solves this neatly. Each gateway subscribes to a channel for every user connected to it, and unsubscribes when they leave. The worker publishes the event once, to that user's channel, and Redis delivers it to whichever gateway is listening. Nobody needs to know where anyone is connected.

delivered 0lost 0

Fig. 2 The worker sends one update at a time, each for a different user. Without a bus, an update lands wherever the load balancer sends it. Through Redis, it goes to the gateway subscribed to that user's channel.

In code, a gateway is little more than a map from users to sockets and one subscriber connection to Redis:

import redis.asyncio as redis

r = redis.Redis()
pubsub = r.pubsub()            # one subscriber connection per gateway
sockets = {}                   # user id -> this gateway's sockets for that user

async def on_connect(user_id, ws):
    sockets.setdefault(user_id, set()).add(ws)
    await pubsub.subscribe(f"user:{user_id}")

async def on_disconnect(user_id, ws):
    sockets[user_id].discard(ws)
    if not sockets[user_id]:
        del sockets[user_id]
        await pubsub.unsubscribe(f"user:{user_id}")

async def relay():             # runs for the life of the gateway
    await pubsub.subscribe("gateway:ping")       # listen() needs a subscription to start
    async for msg in pubsub.listen():
        if msg["type"] != "message":
            continue
        user_id = msg["channel"].decode().split(":", 1)[1]
        for ws in sockets.get(user_id, ()):
            await ws.send(msg["data"])

# and in the worker, when a job finishes:
#     await r.publish(f"user:{user_id}", json.dumps(event))

Inside a data centre, the publish and the hop through Redis take on the order of a millisecond. Nearly all of a 100 ms budget goes on the last stretch to the browser, which is exactly the stretch that polling makes longer. That changes where the effort should go. Shaving the server is rarely worth much; not adding a polling interval on top of the network is worth everything.

3When everyone reconnects at once

The second problem arrives with the first deploy. Restart the gateways and every connected browser drops at the same instant. Each one, quite reasonably, tries to reconnect. If they all retry immediately, the new instances meet a wall of simultaneous handshakes. Many fail, and every failure retries. A routine deploy turns into a self-inflicted outage.

Waiting longer between retries isn't enough by itself. Exponential backoff spaces out each client's attempts, but if every client starts at the same moment and follows the same schedule, they stay in lockstep and arrive in synchronised waves. What breaks the waves is randomness. With full jitter, each client waits a random time between zero and its current backoff, and the herd smears into a steady stream the servers can absorb.

peak –refused –all back by –

Fig. 3 Four thousand browsers lose their connection at once. Each bar is reconnection attempts in one tenth of a second: blue were accepted, red were refused: the gateways take only so many new connections at a time, and fewer when they're swamped. The line is how many are still offline.
function nextDelay(attempt) {
  const base = 500, cap = 30_000;                     // milliseconds
  const backoff = Math.min(cap, base * 2 ** attempt);
  return Math.random() * backoff;                     // full jitter: anywhere in [0, backoff)
}

4Pub/sub forgets

The third problem is quieter. Redis pub/sub is fire and forget: a message published while nobody is subscribed is simply gone. A browser that was mid-reconnect at the moment its job finished will never hear about it, and a notification system that silently drops notifications is worse than one that is merely slow.

So the fast path needs a safety net. Every event is also written somewhere durable with an increasing ID; a Redis stream works, and so does a database table. Each client remembers the last ID it saw and asks for anything newer when it reconnects. Pub/sub gives you speed, the log gives you the guarantee, and neither is enough without the other.

5Boring, on purpose

None of these pieces is novel, and the best compliment a notification system can get is that nobody thinks about it. People see their result the moment it's ready. Engineers deploy without a reconnect storm. And nothing important depends on a message arriving through the one path that's allowed to drop it.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.