Nostr publish failures are silently swallowed — inventory counters drift with no retry and no operator signal #35
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Incident
On 2026-09-05 an nsecbunkerd signing outage caused every NIP-52
republish on aio-demo to fail for roughly 25 minutes:
Three tickets sold for event
hyhNjddtsZ7JNbezwX8ERBduring thatwindow (18:11:21, 18:21:28, 18:29:25). All three republishes were
dropped. The relay kept serving a copy of the calendar event published
2026-08-09, with month-old inventory:
Meanwhile the DB read
amount_tickets=9, sold=7. Every client showed astale "tickets left" badge, and nothing anywhere indicated a problem.
The drift was only discovered because a human noticed a number that
looked wrong on an event card.
nsecbunkerd has since been fixed and publishing recovered on its own,
but the swallow-and-forget path remains.
Mechanism
Two nested catch-alls, neither of which propagates.
nostr_publisher.py:202-204— swallows and returnsNone:nostr_hooks.py:39-46— swallows whatever the first one missed, andtreats a
Nonereturn as an ordinary no-op:publish_or_delete_nostr_eventreturnsNoneon both success andfailure, so none of its 11 call sites can tell the difference —
services.py:68(the sale path) plus ten inviews_api.py.Swallowing here is deliberate and partly correct: a Nostr outage must
not fail the payment flow that triggered the publish. The bug is that
the failure then leaves no trace anywhere except a
WARNINGline, andnothing ever tries again.
Impact
— until the next sale that happens to land while the signer is
healthy. Buyers see wrong availability; a sold-out event can keep
advertising tickets.
nostr_event_created_atstays pinned to the last success, so themonotonic-timestamp logic anchors on a stale value.
WARNINGis indistinguishable from routinenoise in the journal, and there is no metric, alert, or persisted
failure state.
Existing partial mitigation
views_api.py:136/republish-all(admin) andviews_api.py:160/republish-mine(owner) force-republish approved events. So recoveryis possible — but only by a human who already knows the drift happened,
which is exactly what the current logging fails to tell them.
Proposed fix
ERRORfor publish failures. These are notroutine.
nostr_publish_failed/
nostr_publish_pendingmarker on the event row so drift isqueryable rather than buried in the journal.
A periodic sweep republishing rows whose marker is set is probably
simpler and more robust than an in-request retry loop, and it reuses
the logic already behind
/republish-all.publish_or_delete_nostr_eventso callers can branch, while keepingthe default non-fatal.
Signing is
NostrSigner-backend agnostic, so the same failure modeapplies to
LocalSigner— a bunker outage is just the most likelytrigger, not the only one.
Related
Split out of #34, which covers the
amount_ticketscapacity-semanticsbug found during the same investigation. The two are independent: #34
makes the published numbers wrong by construction, this one stops
correct numbers from reaching the relay at all.
Second occurrence on cfaun — and a mechanism this issue doesn't cover
Same symptom, different path in. Worth recording because the fix proposed here wouldn't have caught it.
Event
Q6mhQQzQPSXqkQuFndzXnf(URU ECSTATIC DANCE), oyez.ariege.io. DB readamount_tickets=45, sold=5; the relay served a copy published 2026-09-12 withtickets_available=48 tickets_sold=2— 14 days stale. Found the same way as last time: a human noticed a wrong number on a public page.events.nostr_event_created_atis still1789211268andnostr_event_idstill756fed21…, matching the relay exactly — so no publish succeeded after Sep 12.The difference from the Sep 5 incident: that one produced six
Failed to publish to NostrWARNINGs — the swallow-and-forget path this issue describes. This one produced no log line at all, success or failure.soldreached 5 and onlyset_ticket_paidwrites it, so the hook definitely ran three more times; it returned through one of the two paths that log nothing at INFO:nostr_hooks.py:34—if signer is None: return, a bare return with no log at any levelnostr_publisher.py:175—if not nostr_client: logger.debug(...), invisible at INFONeither is an exception, so neither reaches the
exceptblocks quoted above, and raising those to ERROR wouldn't have surfaced this. Filed as #51 for the logging half.Implication for the fix proposed here: a
nostr_publish_pendingmarker set on failure would also have missed this, because from the code's point of view nothing failed — it declined to publish. The marker wants to be set before attempting and cleared on confirmed success, so "never attempted" and "attempt failed" both leave drift queryable. The periodic sweep then covers both without caring why.That strengthens the case for the sweep over in-request retry: it's the only part of the proposal that catches a cause nobody predicted.
(I filed #52 against this before spotting this issue — closed as a duplicate, nothing in it that isn't here or in #51.)