Design a notification system
Fan a single event out to push, SMS and email through queues and workers, with retries and rate limits at each provider.
9 min read
“Your order shipped” is one event and three deliveries: a push notification, a text and an email, each through a third party you do not control and cannot make faster. Everything about the design follows from that last part.
You are now as slow and as unreliable as Twilio
The simplest version calls Apple's push service, Twilio and Sendgrid from the request handler that triggered the event. It works on a laptop. In production it hands a third party control of your latency and your availability.
Each event holds a thread for 2.4 seconds while three providers answer in turn, so 500 events a second needs 1,200 threads held open just to wait, against a pool of a couple of hundred. The threads are not working; they are waiting, very expensively, and every unrelated request to the service queues behind them. One slow SMS gateway has taken down your checkout page. The fix is not a faster provider or a shorter timeout. It is to stop having the caller wait at all.
Write it down, answer at once, deliver later
Instead, the service drops one copy of the event onto a queue per channel and returns immediately. Pools of workers drain those queues at whatever pace each provider allows, and retry failures without anyone waiting on them.
A queue turns a failure into a delay. When Twilio fails 30% of attempts, the worker catches the error and puts the message back, so the chance of exhausting three attempts is about 2.7%. Those few go to a dead-letter queue for inspection rather than vanishing. A provider outage becomes a delivery delay instead of a delivery loss, and nothing upstream even notices.
A system that always delivers will happily deliver forty
Once delivery is reliable and cheap, the failure mode flips. A retry loop, a batch job or one busy comment thread can fan out dozens of notifications to one person in a minute, and the user's response is not to read them. It is to turn notifications off for good.
You should fix the bug that fired forty events, of course. But the notification system is the last place that can see “this user has had thirty-nine already”, and the only place that protects against the next bug nobody has written yet. A per-user cap, plus collapsing related notifications into one, stands between an incident and a permanent opt-out, and opt-outs are the one thing in this system you can never retry.
Around that core sit the parts with no clever algorithm and plenty of operational weight: a contact store of device tokens, phone numbers and verified emails; per-user, per-category settings; and delivery, open and opt-out analytics. The queues make the system reliable. The settings and the cap keep it welcome.
The short version
- Calling providers inside the request ties your latency and uptime to theirs.
- Queue one message per channel and let workers deliver at the provider's pace.
- Retries turn provider failures into delays; dead-letter what still fails.
- Cap and collapse per user, because an opt-out cannot be retried.