`run_forever` took a req off the queue and then sent it; if the send
raised, the req was already gone and the publish was lost outright,
with the caller long since told it succeeded (the queue put returns
immediately, and the publisher logs "Published" straight after).
Hold the req across the reconnect and retry, bounded at three attempts
so one unsendable message can't wedge every later publish behind it.
This narrows but does not close the gap: a send is still confirmed at
the queue, not by the relay's OK, so a half-dead socket can accept
bytes that never arrive. Closing that needs OK handling in
publish_nostr_event.
Refs #35
A flagged row recovers on its own instead of waiting for the next sale
that happens to land while the signer is healthy — or for an operator
who already knows to run /republish-all, which was the only recovery
path and requires knowing about drift that nothing reported.
Retrying from the DB rather than an in-memory queue means the retry
survives a restart, and it needs no theory about why the publish didn't
land: the sweep covers the signer outage of #35 and the silent skip of
#51 identically, along with causes nobody has hit yet.
Runs every 5 minutes, take-down branch mirroring the publish/delete
split the CRUD endpoints already use. Quiet by design — on a healthy
instance the query returns nothing and it logs nothing.
Refs #35
Inventory reaches clients only through the republished calendar event,
and until now a publish that failed or was skipped left no durable
trace — only a log line, if that. Twice the drift was caught by a human
reading a wrong number on a public page (#35 on aio-demo, #51 on cfaun,
where an event's relay copy sat 14 days behind the DB).
Adds `events.nostr_publish_pending`, set before every attempt and
cleared only on a confirmed success. Ordering it that way is what makes
"the attempt was never made" — no signer resolved, no NostrClient, the
process died mid-flight — as discoverable as "the attempt raised". Both
shapes have now been observed in production; only the second one was
ever visible.
`set_ticket_paid` raises the flag inside its own update so the counters
and "the relay doesn't know about them yet" commit atomically, and the
sale path pays no extra write.
`publish_or_delete_nostr_event` now returns a bool so callers can
branch. The flag, not the return value, is the durable record — the
existing call sites stay correct ignoring it.
Publish failures move from WARNING to ERROR: the published ticket count
has stopped tracking reality, which is not routine journal noise.
Refs #35