From e4742963a3da3a4708662768ab46192b2e4e3a05 Mon Sep 17 00:00:00 2001 From: Padreug Date: Tue, 22 Sep 2026 23:01:25 +0200 Subject: [PATCH] docs(adr): record the cassette-state synchronization model MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit This protocol spans two repos and decides how much cash a machine will pay out, and its only specification was a closed issue and a chat log. That is how four separate divergence bugs went unnoticed. Records what the transport actually permits — an addressable event is an unconditional overwrite ordered by a second-granularity clock, and a relay acknowledges an event it then discards, so a losing writer is never told — and why that rules out compare-and-swap and leads to the ATM owning the count while the operator publishes operations. Decisions 5 through 8 shipped in #104 and spirekeeper#44; 1 through 4 are the v2 operations wire and are not yet built. Co-Authored-By: Claude Fable 5.1 --- .../adr/004-cassette-state-synchronization.md | 132 ++++++++++++++++++ 1 file changed, 132 insertions(+) create mode 100644 docs/adr/004-cassette-state-synchronization.md diff --git a/docs/adr/004-cassette-state-synchronization.md b/docs/adr/004-cassette-state-synchronization.md new file mode 100644 index 0000000..e82e7f9 --- /dev/null +++ b/docs/adr/004-cassette-state-synchronization.md @@ -0,0 +1,132 @@ +# ADR-004: Cassette-State Synchronization + +**Status:** Accepted +**Date:** 2026-09-22 +**Context:** Cassette counts exist on two machines that both write them, over a transport +that cannot report a losing write. This has been load-bearing since #56 shipped, and until +now its only specification was a closed issue and a chat log — which is how four separate +divergence bugs went unnoticed. + +## The problem + +The ATM holds per-bay rows in `state.db` (`position` → `denomination`, `count`). spirekeeper +holds its own `cassette_configs` view for the operator dashboard. They are kept in step over +Nostr kind-30078, one addressable document per direction: + +| d-tag | Direction | Author | +| --------------------------------------- | --------------------- | -------- | +| `bitspire-cassettes-state:` | ATM reports counts up | ATM | +| `bitspire-cassettes:` | operator pushes down | operator | + +Counts drive cash dispensing and the public availability beacon, so a wrong number either +strands a customer at a machine that will not pay out or advertises cash that is not there. + +The transport shapes everything else. Per NIP-01, an addressable event is identified by +`kind:pubkey:d` and ordered by `created_at` at **second granularity**, ties broken by lowest +event id. Relays MAY discard the loser, and a relay returns `OK` for an event it then +discards — so **acceptance is not persistence, and a losing writer is never told**. That +single fact rules out the obvious design. + +## Decisions + +### 1. The ATM owns `count`. The operator publishes operations, not counts. + +A value with one writer cannot be clobbered. Compare-and-swap was considered and rejected: +CAS works because the writer learns it failed and retries, and every standard implementation +of it — HTTP `412`, Kubernetes `409`, a zero rowcount, `CMPXCHG` returning false — delivers +that signal. A kind-30078 publish cannot. Bolting a version onto the current design would let +the ATM refuse a stale push but leave the operator believing they set a count they did not, +trading a wrong number for a phantom edit. + +So the operator publishes `refill`, `empty`, `recount` and `set_denomination` operations. The +vocabulary mirrors lamassu-server's `cash_unit_operation_type`, which is the same shape the +ancestor of this HAL arrived at. Absolute writes survive only as `recount`, which is what an +operator opening a bay and counting actually does. + +`denomination` stays operator-authoritative: the machine cannot know what was physically +loaded into a bay. + +### 2. Idempotency is explicit, because deltas are not idempotent. + +Addressable events are re-delivered on reconnect, so a naive delta would be applied twice. +Every operation carries an operator-minted `id`; the ATM records applied ids and ignores +duplicates. This is lamassu-server's `pullNewBills` pattern — a client-minted UUID per unit of +work, making resend free and ordering irrelevant — rather than a sequence number. + +### 3. The operator publishes a window of recent operations, not one. + +An event the ATM missed self-heals on the next publish, because the next event still carries +the earlier operations. This is the same trick as Lightning.Pub piggybacking `latest_balance` +on every incremental message so a client that missed events corrects itself. + +### 4. The ATM echoes applied ids back, which is the acknowledgement. + +The state document carries `applied_ops`, so the dashboard can render each published operation +as applied or pending. This supplies the feedback leg a replaceable event cannot, without +needing the transport to report failures. + +### 5. Ordering is decided by `created_at`, never by arrival order, on both sides. + +The ATM forces each stamp strictly above its last published one, so a same-second publish or a +clock stepping backwards cannot silently discard a report. spirekeeper applies an event only +when strictly newer than the **oldest** stamp on file for that machine. + +Oldest, not newest, because LNbits' `Connection.execute` commits per call: a multi-row apply +cannot be made atomic through that data layer, so a crash mid-apply leaves some rows advanced. +Gating on the oldest means a partial apply is re-applied rather than mistaken for a complete +one, and the ATM's heartbeat makes it converge. + +### 6. The machine's bay set is authoritative for layout. + +Bay count is hardware-determined. spirekeeper deletes positions absent from a report rather +than leaving them; the operator cannot add or remove bays. + +### 7. Unverified counts are declared, not guessed. + +When a dispense ends with no per-bay report — a driver throw, or the dispense timeout — bills +may have reached the customer with nothing knowing how many. The ATM flags +`counts_uncertain_since` and carries it in the state document rather than letting a number +known to read high stand as measurement. An operator `recount` clears it. + +### 8. State is published on every change and on a heartbeat. + +A publish is one fire-and-forget event with no retry. The heartbeat is what makes the channel +self-healing after a relay outage, and the only way an out-of-band edit to the table ever +reaches the operator. + +## What this replaces + +The original design published a single hello-event gated on a one-shot flag, deduplicated on +one remembered event id, and never compared `created_at` at all. In practice that produced: + +| Failure | Issue | +| --------------------------------------------------------------- | -------------- | +| Layout changes after first boot never published | bitspire#94 | +| Remediation dispenses debited HAL but not the rows | bitspire#76 | +| Absent positions never deleted; publishes then rejected forever | spirekeeper#43 | +| A stale dashboard publish overwriting a newer report | spirekeeper#43 | +| A drained machine advertising bills it had already dispensed | found in audit | +| A re-delivered A, B, A applied three times | found in audit | + +## Status of implementation + +Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format. +Decisions 1 through 4 are the v2 operations wire and are not yet built. + +Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on +their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an +operator refill does not land until it updates — which fails safe, since the machine +under-counts and will not dispense bills it believes it lacks. + +## Alternatives considered + +- **Compare-and-swap on absolute writes.** Rejected: see decision 1. Viable only with a + feedback leg the transport cannot provide, and decision 4 gets the same benefit without + pretending the transport is something it is not. +- **NIP-77 negentropy for reconciliation.** Rejected: it reconciles sets of event ids and + still requires a separate fetch. For a single mutable document it costs more than + re-reading it. +- **One envelope carrying all operator config.** Rejected earlier and still right: a fee edit + that republished a stale cassette inventory is a real failure mode. One d-tag per lifecycle. +- **Publishing the operation log as kind-78.** Deferred. The operator authors the operations + and the ATM records what it applied, so both sides already hold an audit trail.