feat(nostr): confirm publishes against the relay's OK #58
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "feat/confirm-publish-with-relay-ok"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Closes #56
publish_nostr_eventreturned as soon as the EVENT was on the send queue, and the publisher loggedPublishedon the next line. Queueing is not delivery — nostrclient drops an EVENT outright when no relay is connected and answersOK false "error: no relays connected". We discarded that reply.That cost a completed repair on cfaun this morning: the #55 sweep republished a 14-day-stale calendar event 21s before nostrclient finished connecting, got
OK false, loggedPublished, reported1/1 recoveredand cleared the pending flag — leaving the count stale, the row unflagged and the log asserting success. So #55's contract was never true: it claimed to clear only on confirmed success but cleared on confirmed enqueue.publish_nostr_eventnow registers a future per event id, awaits theOK, and returns whether it was accepted.publish_event_to_nostrreturnsNonewhen unconfirmed, which keeps the row flagged for the sweep. Correlation lives inget_event— the one place relay messages cross from the websocket thread into the event loop, so there's no cross-thread future juggling. A disconnect settles in-flight publishes immediately rather than making callers wait out the timeout.Verified on aio-demo, not just in tests
This is a change about the gap between what the code believes and what actually happened, so unit tests can only go so far. Ran it on
aio-demo(which had the releasedv1.6.1-aio.16, i.e. #54 and #55 but not this) by patching the two files over the installed extension, then restored the originals — md5s confirmed identical to the pre-test values.Delivery confirmed:
created_aton the relay matches the DB, so the OK tracked a real delivery rather than merely not timing out.The cfaun failure, reproduced and now caught. Removed the relay from nostrclient to force
OK false:Under
aio.16this exact sequence logsPublished, says1/1 recovered, and clears the flag. The relay's own diagnostic now reaches the journal verbatim —error: no relays connectedis the sentence that would have ended this morning's investigation in seconds.Recovery closes the loop. Restored the relay; the next sweep finished the job with nobody touching the event:
On latency — I had this wrong in the first draft
I initially wrote that a relay outage adds up to 12s per paid ticket on the invoice-listener task. The live run shows a disconnected relay is rejected in 230ms, because nostrclient's router short-circuits without touching a socket. The timeout budget only applies when relays are connected but silent, which nostrclient itself bounds at
PUBLISH_TIMEOUT_SECONDS = 10.That narrow case does still stall
set_ticket_paid. If it ever matters, the remedy is to stop awaiting on the sale path while leaving the flag set — the sweep already guarantees eventual delivery — not to return to reporting unverified success. Commit message corrected before pushing.Known and accepted:
wait_for_nostr_eventssleeps 10s after a dropped socket and nothing drains the receive queue meanwhile, so a publish landing in that window times out and gets flagged. Outcome is correct (the sweep retries), just noisier than necessary.Testing
81 passed(74 existing + 7 new). The new tests cover accept, reject, timeout, an OK for a different event not settling ours,get_eventswallowing OK frames while forwarding everything else, and a disconnect settling in-flight publishes. ruff and black clean on all three files.No
config.jsonbump — version gets cut on main per the release procedure.