feat(nostr): make publish drift queryable and self-healing #55
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "feat/nostr-publish-reconciliation"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Closes #35
Three commits, one concern each.
1.
events.nostr_publish_pending(migrations_fork.m003) — set before every publish attempt, cleared only on a confirmed success.The ordering is the point. #35 proposed a marker set on failure, which would have caught the Sep 5 signer outage but not the cfaun case in #51, where nothing failed — the hook declined to publish and returned. Flagging first makes "never attempted" as discoverable as "attempted and raised", and both shapes have now been seen in production.
set_ticket_paidraises the flag inside its own existing update, so the counters and "the relay doesn't know yet" commit atomically and the sale path pays no extra write.publish_or_delete_nostr_eventreturns a bool now; the flag is the durable record, so the ten existing call sites stay correct ignoring it. Publish failures move WARNING → ERROR.2. The sweep — every 5 minutes, republish whatever is still flagged, with the same publish/delete branching the CRUD endpoints use. Retrying from the DB rather than an in-memory queue means it survives a restart, and it needs no theory about why a publish didn't land: signer outage, silent skip, and causes nobody has hit yet all recover identically. Quiet by design — a healthy instance returns zero rows and logs nothing.
Before this,
/republish-allwas the only recovery, and it requires an operator who already knows about drift that nothing reported.3.
run_foreverno longer drops a dequeued req. It took the req off the queue and then sent it; a raised send lost the publish outright while the caller had already been told it succeeded. Now held across the reconnect and retried, bounded at three attempts so one unsendable message can't wedge everything behind it.Known remaining gap
A send is still confirmed at the queue, not by the relay's OK —
publish_nostr_eventreturns as soon as the req is queued and the publisher logs "Published" immediately after. So a half-dead socket can accept bytes that never arrive, and the flag gets cleared anyway. Closing that needs OK handling inNostrClient; worth its own issue rather than bolting onto this one. Both observed incidents were upstream of that gap, so this PR covers them.Testing
74 passed(67 existing + 7 new). The new tests pin the flag lifecycle across all four outcomes — success, missing signer, publisher returning None, publisher raising — plus the no-double-write path the sweep takes on an already-flagged row, the delete case not clobbering the coordinate, andset_ticket_paidflagging within its own write.ruff and black clean on every touched file; mypy reports only the pre-existing
crud.py/nostr_sync.pyannotations.Deploy note
The migration defaults existing rows to FALSE, not TRUE — on upgrade there's no evidence they're stale, and flagging the whole table would stampede the signer with a full-table republish on first boot.
/republish-allstays the deliberate way to force that, and cfaun will want it once this is deployed, sinceQ6mhQQzQPSXqkQuFndzXnfis drifted today and won't be flagged by the migration.No
config.jsonbump here, per the release procedure.