docs(adr): record the cassette-state synchronization model #105
1 changed files with 160 additions and 0 deletions
160
docs/adr/004-cassette-state-synchronization.md
Normal file
160
docs/adr/004-cassette-state-synchronization.md
Normal file
|
|
@ -0,0 +1,160 @@
|
||||||
|
# ADR-004: Cassette-State Synchronization
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
**Date:** 2026-09-22
|
||||||
|
**Context:** Cassette counts exist on two machines that both write them, over a transport
|
||||||
|
that cannot report a losing write. This has been load-bearing since #56 shipped, and until
|
||||||
|
now its only specification was a closed issue and a chat log — which is how four separate
|
||||||
|
divergence bugs went unnoticed.
|
||||||
|
|
||||||
|
## The problem
|
||||||
|
|
||||||
|
The ATM holds per-bay rows in `state.db` (`position` → `denomination`, `count`). spirekeeper
|
||||||
|
holds its own `cassette_configs` view for the operator dashboard. They are kept in step over
|
||||||
|
Nostr kind-30078, one addressable document per direction:
|
||||||
|
|
||||||
|
| d-tag | Direction | Author |
|
||||||
|
| --------------------------------------- | --------------------- | -------- |
|
||||||
|
| `bitspire-cassettes-state:<atm_pubkey>` | ATM reports counts up | ATM |
|
||||||
|
| `bitspire-cassettes:<atm_pubkey>` | operator pushes down | operator |
|
||||||
|
|
||||||
|
Counts drive cash dispensing and the public availability beacon, so a wrong number either
|
||||||
|
strands a customer at a machine that will not pay out or advertises cash that is not there.
|
||||||
|
|
||||||
|
The transport shapes everything else. Per NIP-01, an addressable event is identified by
|
||||||
|
`kind:pubkey:d` and ordered by `created_at` at **second granularity**, ties broken by lowest
|
||||||
|
event id. Relays MAY discard the loser, and a relay returns `OK` for an event it then
|
||||||
|
discards — so **acceptance is not persistence, and a losing writer is never told**. That
|
||||||
|
single fact rules out the obvious design.
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
### 1. The ATM owns `count`. The operator publishes operations, not counts.
|
||||||
|
|
||||||
|
A value with one writer cannot be clobbered. Compare-and-swap was considered and rejected:
|
||||||
|
CAS works because the writer learns it failed and retries, and every standard implementation
|
||||||
|
of it — HTTP `412`, Kubernetes `409`, a zero rowcount, `CMPXCHG` returning false — delivers
|
||||||
|
that signal. A kind-30078 publish cannot. Bolting a version onto the current design would let
|
||||||
|
the ATM refuse a stale push but leave the operator believing they set a count they did not,
|
||||||
|
trading a wrong number for a phantom edit.
|
||||||
|
|
||||||
|
So the operator publishes `refill`, `empty`, `recount` and `set_denomination` operations. The
|
||||||
|
vocabulary mirrors lamassu-server's `cash_unit_operation_type`, which is the same shape the
|
||||||
|
ancestor of this HAL arrived at. Absolute writes survive only as `recount`, which is what an
|
||||||
|
operator opening a bay and counting actually does.
|
||||||
|
|
||||||
|
`denomination` stays operator-authoritative: the machine cannot know what was physically
|
||||||
|
loaded into a bay.
|
||||||
|
|
||||||
|
### 2. Idempotency is explicit, because deltas are not idempotent.
|
||||||
|
|
||||||
|
Addressable events are re-delivered on reconnect, so a naive delta would be applied twice.
|
||||||
|
Every operation carries an operator-minted `id`; the ATM records applied ids and ignores
|
||||||
|
duplicates. This is lamassu-server's `pullNewBills` pattern — a client-minted UUID per unit of
|
||||||
|
work, making resend free and ordering irrelevant — rather than a sequence number.
|
||||||
|
|
||||||
|
### 3. The operator publishes a window of recent operations, not one.
|
||||||
|
|
||||||
|
An event the ATM missed self-heals on the next publish, because the next event still carries
|
||||||
|
the earlier operations. This is the same trick as Lightning.Pub piggybacking `latest_balance`
|
||||||
|
on every incremental message so a client that missed events corrects itself.
|
||||||
|
|
||||||
|
### 4. The ATM echoes applied ids back, which is the acknowledgement.
|
||||||
|
|
||||||
|
The state document carries `applied_ops`, so the dashboard can render each published operation
|
||||||
|
as applied or pending. This supplies the feedback leg a replaceable event cannot, without
|
||||||
|
needing the transport to report failures.
|
||||||
|
|
||||||
|
### 4a. Until decisions 1 to 4 ship, the overwrite is warned about, not prevented.
|
||||||
|
|
||||||
|
The dashboard's publish dialog already states the failure plainly — that the publish will
|
||||||
|
overwrite the ATM's tracked counts, that decrements since the last baseline will be lost, and
|
||||||
|
that it should follow a physical refill rather than a mid-day tweak. It also says v2
|
||||||
|
reconciliation will replace it.
|
||||||
|
|
||||||
|
Recording this because it changes how the gap should be read. It is a known, deliberately
|
||||||
|
accepted risk carrying a human-factors mitigation, not an oversight, and the product had
|
||||||
|
already reached the same conclusion these decisions formalise. It is worth keeping in mind
|
||||||
|
that a warning is the weakest control available: it depends on an operator reading a dialog
|
||||||
|
at the end of a refill round, and it cannot help at all when the stale value is the one
|
||||||
|
already in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53
|
||||||
|
while a form loaded at 54 stayed open, and nothing but that dialog stood between the operator
|
||||||
|
and discarding the decrement.
|
||||||
|
|
||||||
|
### 5. Ordering is decided by `created_at`, never by arrival order, on both sides.
|
||||||
|
|
||||||
|
The ATM forces each stamp strictly above its last published one, so a same-second publish or a
|
||||||
|
clock stepping backwards cannot silently discard a report. spirekeeper applies an event only
|
||||||
|
when strictly newer than the **oldest** stamp on file for that machine.
|
||||||
|
|
||||||
|
Oldest, not newest, because LNbits' `Connection.execute` commits per call: a multi-row apply
|
||||||
|
cannot be made atomic through that data layer, so a crash mid-apply leaves some rows advanced.
|
||||||
|
Gating on the oldest means a partial apply is re-applied rather than mistaken for a complete
|
||||||
|
one, and the ATM's heartbeat makes it converge.
|
||||||
|
|
||||||
|
### 6. The machine's bay set is authoritative for layout.
|
||||||
|
|
||||||
|
Bay count is hardware-determined. spirekeeper deletes positions absent from a report rather
|
||||||
|
than leaving them; the operator cannot add or remove bays.
|
||||||
|
|
||||||
|
### 7. Unverified counts are declared, not guessed.
|
||||||
|
|
||||||
|
When a dispense ends with no per-bay report — a driver throw, or the dispense timeout — bills
|
||||||
|
may have reached the customer with nothing knowing how many. The ATM flags
|
||||||
|
`counts_uncertain_since` and carries it in the state document rather than letting a number
|
||||||
|
known to read high stand as measurement. An operator `recount` clears it.
|
||||||
|
|
||||||
|
### 8. State is published on every change and on a heartbeat.
|
||||||
|
|
||||||
|
A publish is one fire-and-forget event with no retry. The heartbeat is what makes the channel
|
||||||
|
self-healing after a relay outage, and the only way an out-of-band edit to the table ever
|
||||||
|
reaches the operator.
|
||||||
|
|
||||||
|
## What this replaces
|
||||||
|
|
||||||
|
The original design published a single hello-event gated on a one-shot flag, deduplicated on
|
||||||
|
one remembered event id, and never compared `created_at` at all. In practice that produced:
|
||||||
|
|
||||||
|
| Failure | Issue |
|
||||||
|
| --------------------------------------------------------------- | -------------- |
|
||||||
|
| Layout changes after first boot never published | bitspire#94 |
|
||||||
|
| Remediation dispenses debited HAL but not the rows | bitspire#76 |
|
||||||
|
| Absent positions never deleted; publishes then rejected forever | spirekeeper#43 |
|
||||||
|
| A stale dashboard publish overwriting a newer report | spirekeeper#43 |
|
||||||
|
| A drained machine advertising bills it had already dispensed | found in audit |
|
||||||
|
| A re-delivered A, B, A applied three times | found in audit |
|
||||||
|
|
||||||
|
The last of these is the only one decisions 5 to 8 do not close, because it is not a defect
|
||||||
|
in the mechanism: the operator is permitted to write the count, so a stale write is
|
||||||
|
indistinguishable from an intended one. Only decisions 1 to 4 remove it, by removing the
|
||||||
|
operator's ability to write counts at all.
|
||||||
|
|
||||||
|
## Status of implementation
|
||||||
|
|
||||||
|
Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format,
|
||||||
|
and were verified against the deployed code on sintra on 2026-09-22: a zeroed machine reported
|
||||||
|
drained rather than freezing its beacon, a re-delivered operator config was dropped as stale on
|
||||||
|
eight consecutive restarts, three heartbeat republishes carried strictly increasing stamps read
|
||||||
|
back off the relay, and the machine, the relay and the operator dashboard agreed on the counts
|
||||||
|
with timestamps correlated to the second.
|
||||||
|
|
||||||
|
Decisions 1 through 4 are the v2 operations wire and are not yet built. Decision 4a describes
|
||||||
|
what stands in for them meanwhile.
|
||||||
|
|
||||||
|
Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on
|
||||||
|
their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an
|
||||||
|
operator refill does not land until it updates — which fails safe, since the machine
|
||||||
|
under-counts and will not dispense bills it believes it lacks.
|
||||||
|
|
||||||
|
## Alternatives considered
|
||||||
|
|
||||||
|
- **Compare-and-swap on absolute writes.** Rejected: see decision 1. Viable only with a
|
||||||
|
feedback leg the transport cannot provide, and decision 4 gets the same benefit without
|
||||||
|
pretending the transport is something it is not.
|
||||||
|
- **NIP-77 negentropy for reconciliation.** Rejected: it reconciles sets of event ids and
|
||||||
|
still requires a separate fetch. For a single mutable document it costs more than
|
||||||
|
re-reading it.
|
||||||
|
- **One envelope carrying all operator config.** Rejected earlier and still right: a fee edit
|
||||||
|
that republished a stale cassette inventory is a real failure mode. One d-tag per lifecycle.
|
||||||
|
- **Publishing the operation log as kind-78.** Deferred. The operator authors the operations
|
||||||
|
and the ATM records what it applied, so both sides already hold an audit trail.
|
||||||
Loading…
Add table
Add a link
Reference in a new issue