docs(adr): record the cassette-state synchronization model

This protocol spans two repos and decides how much cash a machine will
pay out, and its only specification was a closed issue and a chat log.
That is how four separate divergence bugs went unnoticed.

Records what the transport actually permits — an addressable event is an
unconditional overwrite ordered by a second-granularity clock, and a
relay acknowledges an event it then discards, so a losing writer is
never told — and why that rules out compare-and-swap and leads to the
ATM owning the count while the operator publishes operations.

Decisions 5 through 8 shipped in #104 and spirekeeper#44; 1 through 4
are the v2 operations wire and are not yet built.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
Padreug 2026-09-22 23:01:25 +02:00
commit e4742963a3

View file

@ -0,0 +1,132 @@
# ADR-004: Cassette-State Synchronization
**Status:** Accepted
**Date:** 2026-09-22
**Context:** Cassette counts exist on two machines that both write them, over a transport
that cannot report a losing write. This has been load-bearing since #56 shipped, and until
now its only specification was a closed issue and a chat log — which is how four separate
divergence bugs went unnoticed.
## The problem
The ATM holds per-bay rows in `state.db` (`position` → `denomination`, `count`). spirekeeper
holds its own `cassette_configs` view for the operator dashboard. They are kept in step over
Nostr kind-30078, one addressable document per direction:
| d-tag | Direction | Author |
| --------------------------------------- | --------------------- | -------- |
| `bitspire-cassettes-state:<atm_pubkey>` | ATM reports counts up | ATM |
| `bitspire-cassettes:<atm_pubkey>` | operator pushes down | operator |
Counts drive cash dispensing and the public availability beacon, so a wrong number either
strands a customer at a machine that will not pay out or advertises cash that is not there.
The transport shapes everything else. Per NIP-01, an addressable event is identified by
`kind:pubkey:d` and ordered by `created_at` at **second granularity**, ties broken by lowest
event id. Relays MAY discard the loser, and a relay returns `OK` for an event it then
discards — so **acceptance is not persistence, and a losing writer is never told**. That
single fact rules out the obvious design.
## Decisions
### 1. The ATM owns `count`. The operator publishes operations, not counts.
A value with one writer cannot be clobbered. Compare-and-swap was considered and rejected:
CAS works because the writer learns it failed and retries, and every standard implementation
of it — HTTP `412`, Kubernetes `409`, a zero rowcount, `CMPXCHG` returning false — delivers
that signal. A kind-30078 publish cannot. Bolting a version onto the current design would let
the ATM refuse a stale push but leave the operator believing they set a count they did not,
trading a wrong number for a phantom edit.
So the operator publishes `refill`, `empty`, `recount` and `set_denomination` operations. The
vocabulary mirrors lamassu-server's `cash_unit_operation_type`, which is the same shape the
ancestor of this HAL arrived at. Absolute writes survive only as `recount`, which is what an
operator opening a bay and counting actually does.
`denomination` stays operator-authoritative: the machine cannot know what was physically
loaded into a bay.
### 2. Idempotency is explicit, because deltas are not idempotent.
Addressable events are re-delivered on reconnect, so a naive delta would be applied twice.
Every operation carries an operator-minted `id`; the ATM records applied ids and ignores
duplicates. This is lamassu-server's `pullNewBills` pattern — a client-minted UUID per unit of
work, making resend free and ordering irrelevant — rather than a sequence number.
### 3. The operator publishes a window of recent operations, not one.
An event the ATM missed self-heals on the next publish, because the next event still carries
the earlier operations. This is the same trick as Lightning.Pub piggybacking `latest_balance`
on every incremental message so a client that missed events corrects itself.
### 4. The ATM echoes applied ids back, which is the acknowledgement.
The state document carries `applied_ops`, so the dashboard can render each published operation
as applied or pending. This supplies the feedback leg a replaceable event cannot, without
needing the transport to report failures.
### 5. Ordering is decided by `created_at`, never by arrival order, on both sides.
The ATM forces each stamp strictly above its last published one, so a same-second publish or a
clock stepping backwards cannot silently discard a report. spirekeeper applies an event only
when strictly newer than the **oldest** stamp on file for that machine.
Oldest, not newest, because LNbits' `Connection.execute` commits per call: a multi-row apply
cannot be made atomic through that data layer, so a crash mid-apply leaves some rows advanced.
Gating on the oldest means a partial apply is re-applied rather than mistaken for a complete
one, and the ATM's heartbeat makes it converge.
### 6. The machine's bay set is authoritative for layout.
Bay count is hardware-determined. spirekeeper deletes positions absent from a report rather
than leaving them; the operator cannot add or remove bays.
### 7. Unverified counts are declared, not guessed.
When a dispense ends with no per-bay report — a driver throw, or the dispense timeout — bills
may have reached the customer with nothing knowing how many. The ATM flags
`counts_uncertain_since` and carries it in the state document rather than letting a number
known to read high stand as measurement. An operator `recount` clears it.
### 8. State is published on every change and on a heartbeat.
A publish is one fire-and-forget event with no retry. The heartbeat is what makes the channel
self-healing after a relay outage, and the only way an out-of-band edit to the table ever
reaches the operator.
## What this replaces
The original design published a single hello-event gated on a one-shot flag, deduplicated on
one remembered event id, and never compared `created_at` at all. In practice that produced:
| Failure | Issue |
| --------------------------------------------------------------- | -------------- |
| Layout changes after first boot never published | bitspire#94 |
| Remediation dispenses debited HAL but not the rows | bitspire#76 |
| Absent positions never deleted; publishes then rejected forever | spirekeeper#43 |
| A stale dashboard publish overwriting a newer report | spirekeeper#43 |
| A drained machine advertising bills it had already dispensed | found in audit |
| A re-delivered A, B, A applied three times | found in audit |
## Status of implementation
Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format.
Decisions 1 through 4 are the v2 operations wire and are not yet built.
Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on
their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an
operator refill does not land until it updates — which fails safe, since the machine
under-counts and will not dispense bills it believes it lacks.
## Alternatives considered
- **Compare-and-swap on absolute writes.** Rejected: see decision 1. Viable only with a
feedback leg the transport cannot provide, and decision 4 gets the same benefit without
pretending the transport is something it is not.
- **NIP-77 negentropy for reconciliation.** Rejected: it reconciles sets of event ids and
still requires a separate fetch. For a single mutable document it costs more than
re-reading it.
- **One envelope carrying all operator config.** Rejected earlier and still right: a fee edit
that republished a stale cassette inventory is a real failure mode. One d-tag per lifecycle.
- **Publishing the operation log as kind-78.** Deferred. The operator authors the operations
and the ATM records what it applied, so both sides already hold an audit trail.