docs(adr): record the cassette-state synchronization model #105
1 changed files with 132 additions and 0 deletions
docs(adr): record the cassette-state synchronization model
This protocol spans two repos and decides how much cash a machine will pay out, and its only specification was a closed issue and a chat log. That is how four separate divergence bugs went unnoticed. Records what the transport actually permits — an addressable event is an unconditional overwrite ordered by a second-granularity clock, and a relay acknowledges an event it then discards, so a losing writer is never told — and why that rules out compare-and-swap and leads to the ATM owning the count while the operator publishes operations. Decisions 5 through 8 shipped in #104 and spirekeeper#44; 1 through 4 are the v2 operations wire and are not yet built. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
commit
e4742963a3
132
docs/adr/004-cassette-state-synchronization.md
Normal file
132
docs/adr/004-cassette-state-synchronization.md
Normal file
|
|
@ -0,0 +1,132 @@
|
|||
# ADR-004: Cassette-State Synchronization
|
||||
|
||||
**Status:** Accepted
|
||||
**Date:** 2026-09-22
|
||||
**Context:** Cassette counts exist on two machines that both write them, over a transport
|
||||
that cannot report a losing write. This has been load-bearing since #56 shipped, and until
|
||||
now its only specification was a closed issue and a chat log — which is how four separate
|
||||
divergence bugs went unnoticed.
|
||||
|
||||
## The problem
|
||||
|
||||
The ATM holds per-bay rows in `state.db` (`position` → `denomination`, `count`). spirekeeper
|
||||
holds its own `cassette_configs` view for the operator dashboard. They are kept in step over
|
||||
Nostr kind-30078, one addressable document per direction:
|
||||
|
||||
| d-tag | Direction | Author |
|
||||
| --------------------------------------- | --------------------- | -------- |
|
||||
| `bitspire-cassettes-state:<atm_pubkey>` | ATM reports counts up | ATM |
|
||||
| `bitspire-cassettes:<atm_pubkey>` | operator pushes down | operator |
|
||||
|
||||
Counts drive cash dispensing and the public availability beacon, so a wrong number either
|
||||
strands a customer at a machine that will not pay out or advertises cash that is not there.
|
||||
|
||||
The transport shapes everything else. Per NIP-01, an addressable event is identified by
|
||||
`kind:pubkey:d` and ordered by `created_at` at **second granularity**, ties broken by lowest
|
||||
event id. Relays MAY discard the loser, and a relay returns `OK` for an event it then
|
||||
discards — so **acceptance is not persistence, and a losing writer is never told**. That
|
||||
single fact rules out the obvious design.
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. The ATM owns `count`. The operator publishes operations, not counts.
|
||||
|
||||
A value with one writer cannot be clobbered. Compare-and-swap was considered and rejected:
|
||||
CAS works because the writer learns it failed and retries, and every standard implementation
|
||||
of it — HTTP `412`, Kubernetes `409`, a zero rowcount, `CMPXCHG` returning false — delivers
|
||||
that signal. A kind-30078 publish cannot. Bolting a version onto the current design would let
|
||||
the ATM refuse a stale push but leave the operator believing they set a count they did not,
|
||||
trading a wrong number for a phantom edit.
|
||||
|
||||
So the operator publishes `refill`, `empty`, `recount` and `set_denomination` operations. The
|
||||
vocabulary mirrors lamassu-server's `cash_unit_operation_type`, which is the same shape the
|
||||
ancestor of this HAL arrived at. Absolute writes survive only as `recount`, which is what an
|
||||
operator opening a bay and counting actually does.
|
||||
|
||||
`denomination` stays operator-authoritative: the machine cannot know what was physically
|
||||
loaded into a bay.
|
||||
|
||||
### 2. Idempotency is explicit, because deltas are not idempotent.
|
||||
|
||||
Addressable events are re-delivered on reconnect, so a naive delta would be applied twice.
|
||||
Every operation carries an operator-minted `id`; the ATM records applied ids and ignores
|
||||
duplicates. This is lamassu-server's `pullNewBills` pattern — a client-minted UUID per unit of
|
||||
work, making resend free and ordering irrelevant — rather than a sequence number.
|
||||
|
||||
### 3. The operator publishes a window of recent operations, not one.
|
||||
|
||||
An event the ATM missed self-heals on the next publish, because the next event still carries
|
||||
the earlier operations. This is the same trick as Lightning.Pub piggybacking `latest_balance`
|
||||
on every incremental message so a client that missed events corrects itself.
|
||||
|
||||
### 4. The ATM echoes applied ids back, which is the acknowledgement.
|
||||
|
||||
The state document carries `applied_ops`, so the dashboard can render each published operation
|
||||
as applied or pending. This supplies the feedback leg a replaceable event cannot, without
|
||||
needing the transport to report failures.
|
||||
|
||||
### 5. Ordering is decided by `created_at`, never by arrival order, on both sides.
|
||||
|
||||
The ATM forces each stamp strictly above its last published one, so a same-second publish or a
|
||||
clock stepping backwards cannot silently discard a report. spirekeeper applies an event only
|
||||
when strictly newer than the **oldest** stamp on file for that machine.
|
||||
|
||||
Oldest, not newest, because LNbits' `Connection.execute` commits per call: a multi-row apply
|
||||
cannot be made atomic through that data layer, so a crash mid-apply leaves some rows advanced.
|
||||
Gating on the oldest means a partial apply is re-applied rather than mistaken for a complete
|
||||
one, and the ATM's heartbeat makes it converge.
|
||||
|
||||
### 6. The machine's bay set is authoritative for layout.
|
||||
|
||||
Bay count is hardware-determined. spirekeeper deletes positions absent from a report rather
|
||||
than leaving them; the operator cannot add or remove bays.
|
||||
|
||||
### 7. Unverified counts are declared, not guessed.
|
||||
|
||||
When a dispense ends with no per-bay report — a driver throw, or the dispense timeout — bills
|
||||
may have reached the customer with nothing knowing how many. The ATM flags
|
||||
`counts_uncertain_since` and carries it in the state document rather than letting a number
|
||||
known to read high stand as measurement. An operator `recount` clears it.
|
||||
|
||||
### 8. State is published on every change and on a heartbeat.
|
||||
|
||||
A publish is one fire-and-forget event with no retry. The heartbeat is what makes the channel
|
||||
self-healing after a relay outage, and the only way an out-of-band edit to the table ever
|
||||
reaches the operator.
|
||||
|
||||
## What this replaces
|
||||
|
||||
The original design published a single hello-event gated on a one-shot flag, deduplicated on
|
||||
one remembered event id, and never compared `created_at` at all. In practice that produced:
|
||||
|
||||
| Failure | Issue |
|
||||
| --------------------------------------------------------------- | -------------- |
|
||||
| Layout changes after first boot never published | bitspire#94 |
|
||||
| Remediation dispenses debited HAL but not the rows | bitspire#76 |
|
||||
| Absent positions never deleted; publishes then rejected forever | spirekeeper#43 |
|
||||
| A stale dashboard publish overwriting a newer report | spirekeeper#43 |
|
||||
| A drained machine advertising bills it had already dispensed | found in audit |
|
||||
| A re-delivered A, B, A applied three times | found in audit |
|
||||
|
||||
## Status of implementation
|
||||
|
||||
Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format.
|
||||
Decisions 1 through 4 are the v2 operations wire and are not yet built.
|
||||
|
||||
Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on
|
||||
their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an
|
||||
operator refill does not land until it updates — which fails safe, since the machine
|
||||
under-counts and will not dispense bills it believes it lacks.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Compare-and-swap on absolute writes.** Rejected: see decision 1. Viable only with a
|
||||
feedback leg the transport cannot provide, and decision 4 gets the same benefit without
|
||||
pretending the transport is something it is not.
|
||||
- **NIP-77 negentropy for reconciliation.** Rejected: it reconciles sets of event ids and
|
||||
still requires a separate fetch. For a single mutable document it costs more than
|
||||
re-reading it.
|
||||
- **One envelope carrying all operator config.** Rejected earlier and still right: a fee edit
|
||||
that republished a stale cassette inventory is a real failure mode. One d-tag per lifecycle.
|
||||
- **Publishing the operation log as kind-78.** Deferred. The operator authors the operations
|
||||
and the ATM records what it applied, so both sides already hold an audit trail.
|
||||
Loading…
Add table
Add a link
Reference in a new issue