From e4742963a3da3a4708662768ab46192b2e4e3a05 Mon Sep 17 00:00:00 2001 From: Padreug Date: Tue, 22 Sep 2026 23:01:25 +0200 Subject: [PATCH 1/2] docs(adr): record the cassette-state synchronization model MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit This protocol spans two repos and decides how much cash a machine will pay out, and its only specification was a closed issue and a chat log. That is how four separate divergence bugs went unnoticed. Records what the transport actually permits — an addressable event is an unconditional overwrite ordered by a second-granularity clock, and a relay acknowledges an event it then discards, so a losing writer is never told — and why that rules out compare-and-swap and leads to the ATM owning the count while the operator publishes operations. Decisions 5 through 8 shipped in #104 and spirekeeper#44; 1 through 4 are the v2 operations wire and are not yet built. Co-Authored-By: Claude Fable 5.1 --- .../adr/004-cassette-state-synchronization.md | 132 ++++++++++++++++++ 1 file changed, 132 insertions(+) create mode 100644 docs/adr/004-cassette-state-synchronization.md diff --git a/docs/adr/004-cassette-state-synchronization.md b/docs/adr/004-cassette-state-synchronization.md new file mode 100644 index 0000000..e82e7f9 --- /dev/null +++ b/docs/adr/004-cassette-state-synchronization.md @@ -0,0 +1,132 @@ +# ADR-004: Cassette-State Synchronization + +**Status:** Accepted +**Date:** 2026-09-22 +**Context:** Cassette counts exist on two machines that both write them, over a transport +that cannot report a losing write. This has been load-bearing since #56 shipped, and until +now its only specification was a closed issue and a chat log — which is how four separate +divergence bugs went unnoticed. + +## The problem + +The ATM holds per-bay rows in `state.db` (`position` → `denomination`, `count`). spirekeeper +holds its own `cassette_configs` view for the operator dashboard. They are kept in step over +Nostr kind-30078, one addressable document per direction: + +| d-tag | Direction | Author | +| --------------------------------------- | --------------------- | -------- | +| `bitspire-cassettes-state:` | ATM reports counts up | ATM | +| `bitspire-cassettes:` | operator pushes down | operator | + +Counts drive cash dispensing and the public availability beacon, so a wrong number either +strands a customer at a machine that will not pay out or advertises cash that is not there. + +The transport shapes everything else. Per NIP-01, an addressable event is identified by +`kind:pubkey:d` and ordered by `created_at` at **second granularity**, ties broken by lowest +event id. Relays MAY discard the loser, and a relay returns `OK` for an event it then +discards — so **acceptance is not persistence, and a losing writer is never told**. That +single fact rules out the obvious design. + +## Decisions + +### 1. The ATM owns `count`. The operator publishes operations, not counts. + +A value with one writer cannot be clobbered. Compare-and-swap was considered and rejected: +CAS works because the writer learns it failed and retries, and every standard implementation +of it — HTTP `412`, Kubernetes `409`, a zero rowcount, `CMPXCHG` returning false — delivers +that signal. A kind-30078 publish cannot. Bolting a version onto the current design would let +the ATM refuse a stale push but leave the operator believing they set a count they did not, +trading a wrong number for a phantom edit. + +So the operator publishes `refill`, `empty`, `recount` and `set_denomination` operations. The +vocabulary mirrors lamassu-server's `cash_unit_operation_type`, which is the same shape the +ancestor of this HAL arrived at. Absolute writes survive only as `recount`, which is what an +operator opening a bay and counting actually does. + +`denomination` stays operator-authoritative: the machine cannot know what was physically +loaded into a bay. + +### 2. Idempotency is explicit, because deltas are not idempotent. + +Addressable events are re-delivered on reconnect, so a naive delta would be applied twice. +Every operation carries an operator-minted `id`; the ATM records applied ids and ignores +duplicates. This is lamassu-server's `pullNewBills` pattern — a client-minted UUID per unit of +work, making resend free and ordering irrelevant — rather than a sequence number. + +### 3. The operator publishes a window of recent operations, not one. + +An event the ATM missed self-heals on the next publish, because the next event still carries +the earlier operations. This is the same trick as Lightning.Pub piggybacking `latest_balance` +on every incremental message so a client that missed events corrects itself. + +### 4. The ATM echoes applied ids back, which is the acknowledgement. + +The state document carries `applied_ops`, so the dashboard can render each published operation +as applied or pending. This supplies the feedback leg a replaceable event cannot, without +needing the transport to report failures. + +### 5. Ordering is decided by `created_at`, never by arrival order, on both sides. + +The ATM forces each stamp strictly above its last published one, so a same-second publish or a +clock stepping backwards cannot silently discard a report. spirekeeper applies an event only +when strictly newer than the **oldest** stamp on file for that machine. + +Oldest, not newest, because LNbits' `Connection.execute` commits per call: a multi-row apply +cannot be made atomic through that data layer, so a crash mid-apply leaves some rows advanced. +Gating on the oldest means a partial apply is re-applied rather than mistaken for a complete +one, and the ATM's heartbeat makes it converge. + +### 6. The machine's bay set is authoritative for layout. + +Bay count is hardware-determined. spirekeeper deletes positions absent from a report rather +than leaving them; the operator cannot add or remove bays. + +### 7. Unverified counts are declared, not guessed. + +When a dispense ends with no per-bay report — a driver throw, or the dispense timeout — bills +may have reached the customer with nothing knowing how many. The ATM flags +`counts_uncertain_since` and carries it in the state document rather than letting a number +known to read high stand as measurement. An operator `recount` clears it. + +### 8. State is published on every change and on a heartbeat. + +A publish is one fire-and-forget event with no retry. The heartbeat is what makes the channel +self-healing after a relay outage, and the only way an out-of-band edit to the table ever +reaches the operator. + +## What this replaces + +The original design published a single hello-event gated on a one-shot flag, deduplicated on +one remembered event id, and never compared `created_at` at all. In practice that produced: + +| Failure | Issue | +| --------------------------------------------------------------- | -------------- | +| Layout changes after first boot never published | bitspire#94 | +| Remediation dispenses debited HAL but not the rows | bitspire#76 | +| Absent positions never deleted; publishes then rejected forever | spirekeeper#43 | +| A stale dashboard publish overwriting a newer report | spirekeeper#43 | +| A drained machine advertising bills it had already dispensed | found in audit | +| A re-delivered A, B, A applied three times | found in audit | + +## Status of implementation + +Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format. +Decisions 1 through 4 are the v2 operations wire and are not yet built. + +Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on +their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an +operator refill does not land until it updates — which fails safe, since the machine +under-counts and will not dispense bills it believes it lacks. + +## Alternatives considered + +- **Compare-and-swap on absolute writes.** Rejected: see decision 1. Viable only with a + feedback leg the transport cannot provide, and decision 4 gets the same benefit without + pretending the transport is something it is not. +- **NIP-77 negentropy for reconciliation.** Rejected: it reconciles sets of event ids and + still requires a separate fetch. For a single mutable document it costs more than + re-reading it. +- **One envelope carrying all operator config.** Rejected earlier and still right: a fee edit + that republished a stale cassette inventory is a real failure mode. One d-tag per lifecycle. +- **Publishing the operation log as kind-78.** Deferred. The operator authors the operations + and the ATM records what it applied, so both sides already hold an audit trail. From 7e3112987ded6990e9a524341ee14f6f70bc79ad Mon Sep 17 00:00:00 2001 From: Padreug Date: Wed, 23 Sep 2026 00:06:29 +0200 Subject: [PATCH 2/2] docs(adr): record the publish warning, and the live verification MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two things learned after the ADR was first written. The dashboard's publish dialog already warns that the publish overwrites the ATM's tracked counts and that decrements since the last baseline will be lost, and says v2 reconciliation will replace it. That changes how the gap should be read: a known risk with a human-factors mitigation, not an oversight, and the product had already reached the same conclusion these decisions formalise. Worth stating that a warning is the weakest control available — it depends on an operator reading a dialog, and cannot help when the stale value is the one already in the form. Also records that decisions 5 to 8 were verified live on sintra rather than only by unit test, and which of the listed failures those decisions do not close. Co-Authored-By: Claude Fable 5.1 --- .../adr/004-cassette-state-synchronization.md | 32 +++++++++++++++++-- 1 file changed, 30 insertions(+), 2 deletions(-) diff --git a/docs/adr/004-cassette-state-synchronization.md b/docs/adr/004-cassette-state-synchronization.md index e82e7f9..ebe35d0 100644 --- a/docs/adr/004-cassette-state-synchronization.md +++ b/docs/adr/004-cassette-state-synchronization.md @@ -65,6 +65,22 @@ The state document carries `applied_ops`, so the dashboard can render each publi as applied or pending. This supplies the feedback leg a replaceable event cannot, without needing the transport to report failures. +### 4a. Until decisions 1 to 4 ship, the overwrite is warned about, not prevented. + +The dashboard's publish dialog already states the failure plainly — that the publish will +overwrite the ATM's tracked counts, that decrements since the last baseline will be lost, and +that it should follow a physical refill rather than a mid-day tweak. It also says v2 +reconciliation will replace it. + +Recording this because it changes how the gap should be read. It is a known, deliberately +accepted risk carrying a human-factors mitigation, not an oversight, and the product had +already reached the same conclusion these decisions formalise. It is worth keeping in mind +that a warning is the weakest control available: it depends on an operator reading a dialog +at the end of a refill round, and it cannot help at all when the stale value is the one +already in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53 +while a form loaded at 54 stayed open, and nothing but that dialog stood between the operator +and discarding the decrement. + ### 5. Ordering is decided by `created_at`, never by arrival order, on both sides. The ATM forces each stamp strictly above its last published one, so a same-second publish or a @@ -108,10 +124,22 @@ one remembered event id, and never compared `created_at` at all. In practice tha | A drained machine advertising bills it had already dispensed | found in audit | | A re-delivered A, B, A applied three times | found in audit | +The last of these is the only one decisions 5 to 8 do not close, because it is not a defect +in the mechanism: the operator is permitted to write the count, so a stale write is +indistinguishable from an intended one. Only decisions 1 to 4 remove it, by removing the +operator's ability to write counts at all. + ## Status of implementation -Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format. -Decisions 1 through 4 are the v2 operations wire and are not yet built. +Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format, +and were verified against the deployed code on sintra on 2026-09-22: a zeroed machine reported +drained rather than freezing its beacon, a re-delivered operator config was dropped as stale on +eight consecutive restarts, three heartbeat republishes carried strictly increasing stamps read +back off the relay, and the machine, the relay and the operator dashboard agreed on the counts +with timestamps correlated to the second. + +Decisions 1 through 4 are the v2 operations wire and are not yet built. Decision 4a describes +what stands in for them meanwhile. Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an