bitspire/docs/adr/004-cassette-state-synchronization.md
Padreug 65f5f982ec docs(adr): record that the operations wire shipped
Decisions 1 through 4 are built on both sides, so the status section says
what landed rather than what is planned, and decision 4a is marked
superseded.

4a is kept rather than deleted. It is the calibration for what a warning
dialog is worth: it depended on an operator reading it at the end of a
refill round, and it could not help at all when the stale value was
already in the form. The dialog and the endpoint behind it are both gone
now, which is a stronger guarantee than any wording could be.

Also records the one gap left open on purpose. An operation recorded
while the relay is unreachable waits for the operator's next action,
because only an operator action triggers a publish.
2026-09-23 12:56:19 +02:00

178 lines
10 KiB
Markdown

# ADR-004: Cassette-State Synchronization
**Status:** Accepted
**Date:** 2026-09-22
**Context:** Cassette counts exist on two machines that both write them, over a transport
that cannot report a losing write. This has been load-bearing since #56 shipped, and until
now its only specification was a closed issue and a chat log — which is how four separate
divergence bugs went unnoticed.
## The problem
The ATM holds per-bay rows in `state.db` (`position` → `denomination`, `count`). spirekeeper
holds its own `cassette_configs` view for the operator dashboard. They are kept in step over
Nostr kind-30078, one addressable document per direction:
| d-tag | Direction | Author |
| --------------------------------------- | --------------------- | -------- |
| `bitspire-cassettes-state:<atm_pubkey>` | ATM reports counts up | ATM |
| `bitspire-cassettes:<atm_pubkey>` | operator pushes down | operator |
Counts drive cash dispensing and the public availability beacon, so a wrong number either
strands a customer at a machine that will not pay out or advertises cash that is not there.
The transport shapes everything else. Per NIP-01, an addressable event is identified by
`kind:pubkey:d` and ordered by `created_at` at **second granularity**, ties broken by lowest
event id. Relays MAY discard the loser, and a relay returns `OK` for an event it then
discards — so **acceptance is not persistence, and a losing writer is never told**. That
single fact rules out the obvious design.
## Decisions
### 1. The ATM owns `count`. The operator publishes operations, not counts.
A value with one writer cannot be clobbered. Compare-and-swap was considered and rejected:
CAS works because the writer learns it failed and retries, and every standard implementation
of it — HTTP `412`, Kubernetes `409`, a zero rowcount, `CMPXCHG` returning false — delivers
that signal. A kind-30078 publish cannot. Bolting a version onto the current design would let
the ATM refuse a stale push but leave the operator believing they set a count they did not,
trading a wrong number for a phantom edit.
So the operator publishes `refill`, `empty`, `recount` and `set_denomination` operations. The
vocabulary mirrors lamassu-server's `cash_unit_operation_type`, which is the same shape the
ancestor of this HAL arrived at. Absolute writes survive only as `recount`, which is what an
operator opening a bay and counting actually does.
`denomination` stays operator-authoritative: the machine cannot know what was physically
loaded into a bay.
### 2. Idempotency is explicit, because deltas are not idempotent.
Addressable events are re-delivered on reconnect, so a naive delta would be applied twice.
Every operation carries an operator-minted `id`; the ATM records applied ids and ignores
duplicates. This is lamassu-server's `pullNewBills` pattern — a client-minted UUID per unit of
work, making resend free and ordering irrelevant — rather than a sequence number.
### 3. The operator publishes a window of recent operations, not one.
An event the ATM missed self-heals on the next publish, because the next event still carries
the earlier operations. This is the same trick as Lightning.Pub piggybacking `latest_balance`
on every incremental message so a client that missed events corrects itself.
### 4. The ATM echoes applied ids back, which is the acknowledgement.
The state document carries `applied_ops`, so the dashboard can render each published operation
as applied or pending. This supplies the feedback leg a replaceable event cannot, without
needing the transport to report failures.
### 4a. Superseded. Before decisions 1 to 4 shipped, the overwrite was warned about.
The dashboard's publish dialog stated the failure plainly — that the publish would overwrite
the ATM's tracked counts, that decrements since the last baseline would be lost, and that it
should follow a physical refill rather than a mid-day tweak.
Kept here rather than deleted, because it is the calibration for how much a warning is worth.
It was a known, deliberately accepted risk carrying a human-factors mitigation, not an
oversight, and the product had already reached the same conclusion these decisions formalise.
It was also the weakest control available: it depended on an operator reading a dialog at the
end of a refill round, and it could not help at all when the stale value was the one already
in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53 while a form
loaded at 54 stayed open, and nothing but that dialog stood between the operator and
discarding the decrement.
The dialog and the endpoint behind it are both gone. The operator dashboard no longer has a
field that accepts a count, which is a stronger guarantee than any wording could be.
### 5. Ordering is decided by `created_at`, never by arrival order, on both sides.
The ATM forces each stamp strictly above its last published one, so a same-second publish or a
clock stepping backwards cannot silently discard a report. spirekeeper applies an event only
when strictly newer than the **oldest** stamp on file for that machine.
Oldest, not newest, because LNbits' `Connection.execute` commits per call: a multi-row apply
cannot be made atomic through that data layer, so a crash mid-apply leaves some rows advanced.
Gating on the oldest means a partial apply is re-applied rather than mistaken for a complete
one, and the ATM's heartbeat makes it converge.
### 6. The machine's bay set is authoritative for layout.
Bay count is hardware-determined. spirekeeper deletes positions absent from a report rather
than leaving them; the operator cannot add or remove bays.
### 7. Unverified counts are declared, not guessed.
When a dispense ends with no per-bay report — a driver throw, or the dispense timeout — bills
may have reached the customer with nothing knowing how many. The ATM flags
`counts_uncertain_since` and carries it in the state document rather than letting a number
known to read high stand as measurement. An operator `recount` clears it.
### 8. State is published on every change and on a heartbeat.
A publish is one fire-and-forget event with no retry. The heartbeat is what makes the channel
self-healing after a relay outage, and the only way an out-of-band edit to the table ever
reaches the operator.
## What this replaces
The original design published a single hello-event gated on a one-shot flag, deduplicated on
one remembered event id, and never compared `created_at` at all. In practice that produced:
| Failure | Issue |
| --------------------------------------------------------------- | -------------- |
| Layout changes after first boot never published | bitspire#94 |
| Remediation dispenses debited HAL but not the rows | bitspire#76 |
| Absent positions never deleted; publishes then rejected forever | spirekeeper#43 |
| A stale dashboard publish overwriting a newer report | spirekeeper#43 |
| A drained machine advertising bills it had already dispensed | found in audit |
| A re-delivered A, B, A applied three times | found in audit |
The last of these is the only one decisions 5 to 8 do not close, because it is not a defect
in the mechanism: the operator is permitted to write the count, so a stale write is
indistinguishable from an intended one. Only decisions 1 to 4 remove it, by removing the
operator's ability to write counts at all.
## Status of implementation
Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format,
and were verified against the deployed code on sintra on 2026-09-22: a zeroed machine reported
drained rather than freezing its beacon, a re-delivered operator config was dropped as stale on
eight consecutive restarts, three heartbeat republishes carried strictly increasing stamps read
back off the relay, and the machine, the relay and the operator dashboard agreed on the counts
with timestamps correlated to the second.
Decisions 1 through 4 are the v2 operations wire and shipped in spirekeeper#46 and
bitspire#106.
On the operator side there is no longer any endpoint that accepts a count: the absolute
publish, its CRUD write and its request model were removed rather than deprecated. The
dashboard records operations and renders each as applied or pending from the machine's
`applied_ops` echo. On the machine side, schema v13 adds a `cassette_ops` dedup ledger, the
`created_at` watermark on this path is retired in favour of per-op ids, and the state document
carries `schema_version`, `seq` and `applied_ops`.
The wire shapes are those given under decisions 1 and 4 above.
Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on
their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an
operator refill does not land until it updates — which fails safe, since the machine
under-counts and will not dispense bills it believes it lacks. In the other direction an
updated machine drops a v1 absolute-count payload on the missing `ops` array, which is the
same safe direction: the machine keeps the counts it is now the only writer of.
One gap stays open deliberately. An operation recorded while the relay is unreachable waits
for the operator's next action to be published, because only an operator action triggers a
publish. The window makes that self-healing once anything is published, but nothing on the
operator side republishes on its own. An operator-side heartbeat is the fix; it is not built.
## Alternatives considered
- **Compare-and-swap on absolute writes.** Rejected: see decision 1. Viable only with a
feedback leg the transport cannot provide, and decision 4 gets the same benefit without
pretending the transport is something it is not.
- **NIP-77 negentropy for reconciliation.** Rejected: it reconciles sets of event ids and
still requires a separate fetch. For a single mutable document it costs more than
re-reading it.
- **One envelope carrying all operator config.** Rejected earlier and still right: a fee edit
that republished a stale cassette inventory is a real failure mode. One d-tag per lifecycle.
- **Publishing the operation log as kind-78.** Deferred. The operator authors the operations
and the ATM records what it applied, so both sides already hold an audit trail.