A lost NIP-52 republish is permanent — nothing reconciles the relay copy #52

Closed
opened 2026-09-26 12:42:47 +00:00 by padreug · 1 comment
Owner

Ticket inventory reaches clients only through the NIP-52 republish in set_ticket_paid. That republish is fire-and-forget, and if it's lost the relay keeps serving the stale count forever — nothing retries, and nothing ever compares what the relay holds against the DB.

That's what happened on cfaun (see #51 for the diagnosis): the relay served tickets_available=48 for 14 days while the DB said 45, and the only reason it got noticed was someone reading the public page.

Every layer drops the event quietly:

  • NostrClient.publish_nostr_event just does await self.send_req_queue.put(["EVENT", e.dict()]) — no OK, no confirmation.
  • run_forever (nostr/nostr_client.py:78-89) dequeues, then self.ws.send(...). If that raises, the event is already off the queue and is never retried; the loop logs a warning and sleeps 60s.
  • is_websocket_connected returns self.ws.keep_running, which stays True on a half-dead socket, so sends can also vanish with no exception at all.
  • publish_event_to_nostr logs Published NIP-52 calendar event right after the queue put, so that line means "queued", not "the relay has it".

What would fix it

Roughly in order of value:

  1. Reconciliation. Something that periodically (or on organizer page load) checks whether the relay's copy matches event.sold / event.amount_tickets and republishes when it doesn't. This is the piece that turns a lost publish from permanent drift into a blip, and it covers every cause including ones we haven't thought of.
  2. Don't lose the dequeued event. In run_forever, on a send failure put the request back (or hold it) rather than dropping it on the floor.
  3. Read the relay's OK. publish_nostr_event could await the OK frame and surface a rejection, so "Published" means published.

(1) alone would have kept oyez correct through this incident even with everything else unchanged.

Ticket inventory reaches clients only through the NIP-52 republish in `set_ticket_paid`. That republish is fire-and-forget, and if it's lost the relay keeps serving the stale count forever — nothing retries, and nothing ever compares what the relay holds against the DB. That's what happened on cfaun (see #51 for the diagnosis): the relay served `tickets_available=48` for 14 days while the DB said 45, and the only reason it got noticed was someone reading the public page. Every layer drops the event quietly: - `NostrClient.publish_nostr_event` just does `await self.send_req_queue.put(["EVENT", e.dict()])` — no OK, no confirmation. - `run_forever` (`nostr/nostr_client.py:78-89`) dequeues, then `self.ws.send(...)`. If that raises, the event is **already off the queue and is never retried**; the loop logs a warning and sleeps 60s. - `is_websocket_connected` returns `self.ws.keep_running`, which stays True on a half-dead socket, so sends can also vanish with no exception at all. - `publish_event_to_nostr` logs `Published NIP-52 calendar event` right after the queue put, so that line means "queued", not "the relay has it". ## What would fix it Roughly in order of value: 1. **Reconciliation.** Something that periodically (or on organizer page load) checks whether the relay's copy matches `event.sold` / `event.amount_tickets` and republishes when it doesn't. This is the piece that turns a lost publish from permanent drift into a blip, and it covers every cause including ones we haven't thought of. 2. **Don't lose the dequeued event.** In `run_forever`, on a send failure put the request back (or hold it) rather than dropping it on the floor. 3. **Read the relay's OK.** `publish_nostr_event` could await the OK frame and surface a rejection, so "Published" means published. (1) alone would have kept oyez correct through this incident even with everything else unchanged.
Author
Owner

Duplicate of #35, which covers the same drift with more evidence (the 2026-09-05 aio-demo incident) and a fuller proposal. Filed this before checking the tracker — my mistake.

The one thing here that wasn't in #35 — the cfaun occurrence and the fact that it came in through a silent path rather than a swallowed exception — is now recorded as a comment on #35. The logging half is #51.

Closing.

Duplicate of #35, which covers the same drift with more evidence (the 2026-09-05 aio-demo incident) and a fuller proposal. Filed this before checking the tracker — my mistake. The one thing here that wasn't in #35 — the cfaun occurrence and the fact that it came in through a *silent* path rather than a swallowed exception — is now recorded as a comment on #35. The logging half is #51. Closing.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/events#52
No description provided.