Merge pull request 'docs(adr): record the cassette-state synchronization model' (#105) from docs/adr-cassette-sync into dev

Reviewed-on: #105
This commit is contained in:
padreug 2026-09-22 22:26:11 +00:00
commit 9077f9c299

View file

@ -0,0 +1,160 @@
# ADR-004: Cassette-State Synchronization
**Status:** Accepted
**Date:** 2026-09-22
**Context:** Cassette counts exist on two machines that both write them, over a transport
that cannot report a losing write. This has been load-bearing since #56 shipped, and until
now its only specification was a closed issue and a chat log — which is how four separate
divergence bugs went unnoticed.
## The problem
The ATM holds per-bay rows in `state.db` (`position` → `denomination`, `count`). spirekeeper
holds its own `cassette_configs` view for the operator dashboard. They are kept in step over
Nostr kind-30078, one addressable document per direction:
| d-tag | Direction | Author |
| --------------------------------------- | --------------------- | -------- |
| `bitspire-cassettes-state:<atm_pubkey>` | ATM reports counts up | ATM |
| `bitspire-cassettes:<atm_pubkey>` | operator pushes down | operator |
Counts drive cash dispensing and the public availability beacon, so a wrong number either
strands a customer at a machine that will not pay out or advertises cash that is not there.
The transport shapes everything else. Per NIP-01, an addressable event is identified by
`kind:pubkey:d` and ordered by `created_at` at **second granularity**, ties broken by lowest
event id. Relays MAY discard the loser, and a relay returns `OK` for an event it then
discards — so **acceptance is not persistence, and a losing writer is never told**. That
single fact rules out the obvious design.
## Decisions
### 1. The ATM owns `count`. The operator publishes operations, not counts.
A value with one writer cannot be clobbered. Compare-and-swap was considered and rejected:
CAS works because the writer learns it failed and retries, and every standard implementation
of it — HTTP `412`, Kubernetes `409`, a zero rowcount, `CMPXCHG` returning false — delivers
that signal. A kind-30078 publish cannot. Bolting a version onto the current design would let
the ATM refuse a stale push but leave the operator believing they set a count they did not,
trading a wrong number for a phantom edit.
So the operator publishes `refill`, `empty`, `recount` and `set_denomination` operations. The
vocabulary mirrors lamassu-server's `cash_unit_operation_type`, which is the same shape the
ancestor of this HAL arrived at. Absolute writes survive only as `recount`, which is what an
operator opening a bay and counting actually does.
`denomination` stays operator-authoritative: the machine cannot know what was physically
loaded into a bay.
### 2. Idempotency is explicit, because deltas are not idempotent.
Addressable events are re-delivered on reconnect, so a naive delta would be applied twice.
Every operation carries an operator-minted `id`; the ATM records applied ids and ignores
duplicates. This is lamassu-server's `pullNewBills` pattern — a client-minted UUID per unit of
work, making resend free and ordering irrelevant — rather than a sequence number.
### 3. The operator publishes a window of recent operations, not one.
An event the ATM missed self-heals on the next publish, because the next event still carries
the earlier operations. This is the same trick as Lightning.Pub piggybacking `latest_balance`
on every incremental message so a client that missed events corrects itself.
### 4. The ATM echoes applied ids back, which is the acknowledgement.
The state document carries `applied_ops`, so the dashboard can render each published operation
as applied or pending. This supplies the feedback leg a replaceable event cannot, without
needing the transport to report failures.
### 4a. Until decisions 1 to 4 ship, the overwrite is warned about, not prevented.
The dashboard's publish dialog already states the failure plainly — that the publish will
overwrite the ATM's tracked counts, that decrements since the last baseline will be lost, and
that it should follow a physical refill rather than a mid-day tweak. It also says v2
reconciliation will replace it.
Recording this because it changes how the gap should be read. It is a known, deliberately
accepted risk carrying a human-factors mitigation, not an oversight, and the product had
already reached the same conclusion these decisions formalise. It is worth keeping in mind
that a warning is the weakest control available: it depends on an operator reading a dialog
at the end of a refill round, and it cannot help at all when the stale value is the one
already in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53
while a form loaded at 54 stayed open, and nothing but that dialog stood between the operator
and discarding the decrement.
### 5. Ordering is decided by `created_at`, never by arrival order, on both sides.
The ATM forces each stamp strictly above its last published one, so a same-second publish or a
clock stepping backwards cannot silently discard a report. spirekeeper applies an event only
when strictly newer than the **oldest** stamp on file for that machine.
Oldest, not newest, because LNbits' `Connection.execute` commits per call: a multi-row apply
cannot be made atomic through that data layer, so a crash mid-apply leaves some rows advanced.
Gating on the oldest means a partial apply is re-applied rather than mistaken for a complete
one, and the ATM's heartbeat makes it converge.
### 6. The machine's bay set is authoritative for layout.
Bay count is hardware-determined. spirekeeper deletes positions absent from a report rather
than leaving them; the operator cannot add or remove bays.
### 7. Unverified counts are declared, not guessed.
When a dispense ends with no per-bay report — a driver throw, or the dispense timeout — bills
may have reached the customer with nothing knowing how many. The ATM flags
`counts_uncertain_since` and carries it in the state document rather than letting a number
known to read high stand as measurement. An operator `recount` clears it.
### 8. State is published on every change and on a heartbeat.
A publish is one fire-and-forget event with no retry. The heartbeat is what makes the channel
self-healing after a relay outage, and the only way an out-of-band edit to the table ever
reaches the operator.
## What this replaces
The original design published a single hello-event gated on a one-shot flag, deduplicated on
one remembered event id, and never compared `created_at` at all. In practice that produced:
| Failure | Issue |
| --------------------------------------------------------------- | -------------- |
| Layout changes after first boot never published | bitspire#94 |
| Remediation dispenses debited HAL but not the rows | bitspire#76 |
| Absent positions never deleted; publishes then rejected forever | spirekeeper#43 |
| A stale dashboard publish overwriting a newer report | spirekeeper#43 |
| A drained machine advertising bills it had already dispensed | found in audit |
| A re-delivered A, B, A applied three times | found in audit |
The last of these is the only one decisions 5 to 8 do not close, because it is not a defect
in the mechanism: the operator is permitted to write the count, so a stale write is
indistinguishable from an intended one. Only decisions 1 to 4 remove it, by removing the
operator's ability to write counts at all.
## Status of implementation
Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format,
and were verified against the deployed code on sintra on 2026-09-22: a zeroed machine reported
drained rather than freezing its beacon, a re-delivered operator config was dropped as stale on
eight consecutive restarts, three heartbeat republishes carried strictly increasing stamps read
back off the relay, and the machine, the relay and the operator dashboard agreed on the counts
with timestamps correlated to the second.
Decisions 1 through 4 are the v2 operations wire and are not yet built. Decision 4a describes
what stands in for them meanwhile.
Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on
their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an
operator refill does not land until it updates — which fails safe, since the machine
under-counts and will not dispense bills it believes it lacks.
## Alternatives considered
- **Compare-and-swap on absolute writes.** Rejected: see decision 1. Viable only with a
feedback leg the transport cannot provide, and decision 4 gets the same benefit without
pretending the transport is something it is not.
- **NIP-77 negentropy for reconciliation.** Rejected: it reconciles sets of event ids and
still requires a separate fetch. For a single mutable document it costs more than
re-reading it.
- **One envelope carrying all operator config.** Rejected earlier and still right: a fee edit
that republished a stale cassette inventory is a real failure mode. One d-tag per lifecycle.
- **Publishing the operation log as kind-78.** Deferred. The operator authors the operations
and the ATM records what it applied, so both sides already hold an audit trail.