Ayobami ZenthosGet in touch
All insights

Push notifications that never repeat and never get lost

How Zenthos Gym delivers every alert to every device exactly once, with database claims, retries that expire and a five-minute backstop.

Live check-in at the front deskMember home screenManager overview

For a gym, notifications are not marketing. A receptionist needs to know a member has just paid. A member needs to know their membership has started. A manager needs to know a new member joined.

Two things can go wrong, and both are bad. An alert can be lost, so the desk never hears about a payment. Or it can arrive twice, so a phone buzzes for something that already happened, and people learn to ignore it.

Zenthos Gym is built so that neither happens. This is how.

Every alert starts as a row

Nothing sends a push directly. Every alert is first written to a notifications table, which doubles as the in-app inbox. A notification that never reaches a phone still exists, and the person still sees it when they open the app.

Writing a row also wakes the sender. A database trigger calls the push dispatcher with its own secret, not the service key, so the dispatcher can only do this one job.

Claim before you send

Several alerts can be written in the same second, which means several dispatcher runs can start at once. If they all read “notifications not yet pushed” and send them, the member gets the same alert three times.

So a run never just reads. It claims:

update notifications
   set push_claimed_at = now()
 where id in (
   select id from notifications
    where pushed_at is null
      and (push_claimed_at is null or push_claimed_at < now() - interval '2 minutes')
    order by created_at
    limit p_limit
    for update skip locked
 )
returning *

for update skip locked lets overlapping runs pass each other: whatever one run has locked, the next one skips. Each alert ends up with exactly one run responsible for it.

A claim that expires

If a run claims ten alerts and then crashes halfway, those alerts must not be stuck forever. That is what the two-minute window is for. A claim older than two minutes counts as abandoned, and the next run picks those alerts up again.

And in case a trigger call itself never arrives, a scheduled job wakes the dispatcher every five minutes as a backstop. Between the trigger, the expiring claim and the backstop, there is always a second chance.

Deciding what counts as done

Sending to one person can mean sending to several devices: a phone, a tablet, a shared desk phone. The dispatcher sends to all of them at once and then reads the results carefully:

const status = (outcome.reason as { statusCode?: number })?.statusCode
if (status === 404 || status === 410) expired.add(devices[index].id)
else retryable = true

A 404 or 410 means that device has unsubscribed or the app was removed. It will never work again, so the subscription is deleted. Any other failure, such as a timeout or the push service having a bad moment, is worth retrying.

An alert is marked as pushed if at least one device accepted it, or if there is nothing left worth retrying:

// With no device left to try the alert still lives in the in-app inbox.
if (accepted || !retryable) handled.push(alert.id)

Otherwise it is left unmarked, its claim expires, and a later run tries again.

Urgency where it matters

Not every alert is equal. Payment alerts and renewals that are due are sent as high urgency, so the phone wakes for them straight away. Everything else is normal urgency, which lets phones batch them and save battery.

What I took from it

  • Write first, send second. The table is the source of truth; the push is just delivery.
  • Claim work instead of reading it. Locks with skip locked turn a race into a queue.
  • Let claims expire. Crashes are normal, so give abandoned work back automatically.
  • Read failures properly. “Gone forever” and “try again” are different answers.
  • Always have a backstop. A scheduled sweep catches whatever the fast path misses.
See Zenthos Gym