feat(nostr): make publish drift queryable and self-healing #55

Merged
padreug merged 3 commits from feat/nostr-publish-reconciliation into main 2026-09-26 21:47:41 +00:00
Owner

Closes #35

Three commits, one concern each.

1. events.nostr_publish_pending (migrations_fork.m003) — set before every publish attempt, cleared only on a confirmed success.

The ordering is the point. #35 proposed a marker set on failure, which would have caught the Sep 5 signer outage but not the cfaun case in #51, where nothing failed — the hook declined to publish and returned. Flagging first makes "never attempted" as discoverable as "attempted and raised", and both shapes have now been seen in production.

set_ticket_paid raises the flag inside its own existing update, so the counters and "the relay doesn't know yet" commit atomically and the sale path pays no extra write. publish_or_delete_nostr_event returns a bool now; the flag is the durable record, so the ten existing call sites stay correct ignoring it. Publish failures move WARNING → ERROR.

2. The sweep — every 5 minutes, republish whatever is still flagged, with the same publish/delete branching the CRUD endpoints use. Retrying from the DB rather than an in-memory queue means it survives a restart, and it needs no theory about why a publish didn't land: signer outage, silent skip, and causes nobody has hit yet all recover identically. Quiet by design — a healthy instance returns zero rows and logs nothing.

Before this, /republish-all was the only recovery, and it requires an operator who already knows about drift that nothing reported.

3. run_forever no longer drops a dequeued req. It took the req off the queue and then sent it; a raised send lost the publish outright while the caller had already been told it succeeded. Now held across the reconnect and retried, bounded at three attempts so one unsendable message can't wedge everything behind it.

Known remaining gap

A send is still confirmed at the queue, not by the relay's OK — publish_nostr_event returns as soon as the req is queued and the publisher logs "Published" immediately after. So a half-dead socket can accept bytes that never arrive, and the flag gets cleared anyway. Closing that needs OK handling in NostrClient; worth its own issue rather than bolting onto this one. Both observed incidents were upstream of that gap, so this PR covers them.

Testing

74 passed (67 existing + 7 new). The new tests pin the flag lifecycle across all four outcomes — success, missing signer, publisher returning None, publisher raising — plus the no-double-write path the sweep takes on an already-flagged row, the delete case not clobbering the coordinate, and set_ticket_paid flagging within its own write.

ruff and black clean on every touched file; mypy reports only the pre-existing crud.py / nostr_sync.py annotations.

Deploy note

The migration defaults existing rows to FALSE, not TRUE — on upgrade there's no evidence they're stale, and flagging the whole table would stampede the signer with a full-table republish on first boot. /republish-all stays the deliberate way to force that, and cfaun will want it once this is deployed, since Q6mhQQzQPSXqkQuFndzXnf is drifted today and won't be flagged by the migration.

No config.json bump here, per the release procedure.

Closes #35 Three commits, one concern each. **1. `events.nostr_publish_pending`** (`migrations_fork.m003`) — set before every publish attempt, cleared only on a confirmed success. The ordering is the point. #35 proposed a marker set *on failure*, which would have caught the Sep 5 signer outage but not the cfaun case in #51, where nothing failed — the hook declined to publish and returned. Flagging first makes "never attempted" as discoverable as "attempted and raised", and both shapes have now been seen in production. `set_ticket_paid` raises the flag inside its own existing update, so the counters and "the relay doesn't know yet" commit atomically and the sale path pays no extra write. `publish_or_delete_nostr_event` returns a bool now; the flag is the durable record, so the ten existing call sites stay correct ignoring it. Publish failures move WARNING → ERROR. **2. The sweep** — every 5 minutes, republish whatever is still flagged, with the same publish/delete branching the CRUD endpoints use. Retrying from the DB rather than an in-memory queue means it survives a restart, and it needs no theory about *why* a publish didn't land: signer outage, silent skip, and causes nobody has hit yet all recover identically. Quiet by design — a healthy instance returns zero rows and logs nothing. Before this, `/republish-all` was the only recovery, and it requires an operator who already knows about drift that nothing reported. **3. `run_forever` no longer drops a dequeued req.** It took the req off the queue and then sent it; a raised send lost the publish outright while the caller had already been told it succeeded. Now held across the reconnect and retried, bounded at three attempts so one unsendable message can't wedge everything behind it. ## Known remaining gap A send is still confirmed at the queue, not by the relay's OK — `publish_nostr_event` returns as soon as the req is queued and the publisher logs "Published" immediately after. So a half-dead socket can accept bytes that never arrive, and the flag gets cleared anyway. Closing that needs OK handling in `NostrClient`; worth its own issue rather than bolting onto this one. Both observed incidents were upstream of that gap, so this PR covers them. ## Testing `74 passed` (67 existing + 7 new). The new tests pin the flag lifecycle across all four outcomes — success, missing signer, publisher returning None, publisher raising — plus the no-double-write path the sweep takes on an already-flagged row, the delete case not clobbering the coordinate, and `set_ticket_paid` flagging within its own write. ruff and black clean on every touched file; mypy reports only the pre-existing `crud.py` / `nostr_sync.py` annotations. ## Deploy note The migration defaults existing rows to FALSE, not TRUE — on upgrade there's no evidence they're stale, and flagging the whole table would stampede the signer with a full-table republish on first boot. `/republish-all` stays the deliberate way to force that, and cfaun will want it once this is deployed, since `Q6mhQQzQPSXqkQuFndzXnf` is drifted today and won't be flagged by the migration. No `config.json` bump here, per the release procedure.
Inventory reaches clients only through the republished calendar event,
and until now a publish that failed or was skipped left no durable
trace — only a log line, if that. Twice the drift was caught by a human
reading a wrong number on a public page (#35 on aio-demo, #51 on cfaun,
where an event's relay copy sat 14 days behind the DB).

Adds `events.nostr_publish_pending`, set before every attempt and
cleared only on a confirmed success. Ordering it that way is what makes
"the attempt was never made" — no signer resolved, no NostrClient, the
process died mid-flight — as discoverable as "the attempt raised". Both
shapes have now been observed in production; only the second one was
ever visible.

`set_ticket_paid` raises the flag inside its own update so the counters
and "the relay doesn't know about them yet" commit atomically, and the
sale path pays no extra write.

`publish_or_delete_nostr_event` now returns a bool so callers can
branch. The flag, not the return value, is the durable record — the
existing call sites stay correct ignoring it.

Publish failures move from WARNING to ERROR: the published ticket count
has stopped tracking reality, which is not routine journal noise.

Refs #35
A flagged row recovers on its own instead of waiting for the next sale
that happens to land while the signer is healthy — or for an operator
who already knows to run /republish-all, which was the only recovery
path and requires knowing about drift that nothing reported.

Retrying from the DB rather than an in-memory queue means the retry
survives a restart, and it needs no theory about why the publish didn't
land: the sweep covers the signer outage of #35 and the silent skip of
#51 identically, along with causes nobody has hit yet.

Runs every 5 minutes, take-down branch mirroring the publish/delete
split the CRUD endpoints already use. Quiet by design — on a healthy
instance the query returns nothing and it logs nothing.

Refs #35
fix(nostr): retry a dequeued req instead of dropping it
Some checks failed
lint.yml / fix(nostr): retry a dequeued req instead of dropping it (pull_request) Failing after 0s
dc2a296bad
`run_forever` took a req off the queue and then sent it; if the send
raised, the req was already gone and the publish was lost outright,
with the caller long since told it succeeded (the queue put returns
immediately, and the publisher logs "Published" straight after).

Hold the req across the reconnect and retry, bounded at three attempts
so one unsendable message can't wedge every later publish behind it.

This narrows but does not close the gap: a send is still confirmed at
the queue, not by the relay's OK, so a half-dead socket can accept
bytes that never arrive. Closing that needs OK handling in
publish_nostr_event.

Refs #35
padreug deleted branch feat/nostr-publish-reconciliation 2026-09-26 21:47:41 +00:00
padreug referenced this pull request from a commit 2026-09-27 07:27:48 +00:00
padreug referenced this pull request from a commit 2026-09-27 21:12:16 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/events!55
No description provided.