Decisions 1 through 4 are built on both sides, so the status section says what landed rather than what is planned, and decision 4a is marked superseded. 4a is kept rather than deleted. It is the calibration for what a warning dialog is worth: it depended on an operator reading it at the end of a refill round, and it could not help at all when the stale value was already in the form. The dialog and the endpoint behind it are both gone now, which is a stronger guarantee than any wording could be. Also records the one gap left open on purpose. An operation recorded while the relay is unreachable waits for the operator's next action, because only an operator action triggers a publish.
178 lines
10 KiB
Markdown
178 lines
10 KiB
Markdown
# ADR-004: Cassette-State Synchronization
|
|
|
|
**Status:** Accepted
|
|
**Date:** 2026-09-22
|
|
**Context:** Cassette counts exist on two machines that both write them, over a transport
|
|
that cannot report a losing write. This has been load-bearing since #56 shipped, and until
|
|
now its only specification was a closed issue and a chat log — which is how four separate
|
|
divergence bugs went unnoticed.
|
|
|
|
## The problem
|
|
|
|
The ATM holds per-bay rows in `state.db` (`position` → `denomination`, `count`). spirekeeper
|
|
holds its own `cassette_configs` view for the operator dashboard. They are kept in step over
|
|
Nostr kind-30078, one addressable document per direction:
|
|
|
|
| d-tag | Direction | Author |
|
|
| --------------------------------------- | --------------------- | -------- |
|
|
| `bitspire-cassettes-state:<atm_pubkey>` | ATM reports counts up | ATM |
|
|
| `bitspire-cassettes:<atm_pubkey>` | operator pushes down | operator |
|
|
|
|
Counts drive cash dispensing and the public availability beacon, so a wrong number either
|
|
strands a customer at a machine that will not pay out or advertises cash that is not there.
|
|
|
|
The transport shapes everything else. Per NIP-01, an addressable event is identified by
|
|
`kind:pubkey:d` and ordered by `created_at` at **second granularity**, ties broken by lowest
|
|
event id. Relays MAY discard the loser, and a relay returns `OK` for an event it then
|
|
discards — so **acceptance is not persistence, and a losing writer is never told**. That
|
|
single fact rules out the obvious design.
|
|
|
|
## Decisions
|
|
|
|
### 1. The ATM owns `count`. The operator publishes operations, not counts.
|
|
|
|
A value with one writer cannot be clobbered. Compare-and-swap was considered and rejected:
|
|
CAS works because the writer learns it failed and retries, and every standard implementation
|
|
of it — HTTP `412`, Kubernetes `409`, a zero rowcount, `CMPXCHG` returning false — delivers
|
|
that signal. A kind-30078 publish cannot. Bolting a version onto the current design would let
|
|
the ATM refuse a stale push but leave the operator believing they set a count they did not,
|
|
trading a wrong number for a phantom edit.
|
|
|
|
So the operator publishes `refill`, `empty`, `recount` and `set_denomination` operations. The
|
|
vocabulary mirrors lamassu-server's `cash_unit_operation_type`, which is the same shape the
|
|
ancestor of this HAL arrived at. Absolute writes survive only as `recount`, which is what an
|
|
operator opening a bay and counting actually does.
|
|
|
|
`denomination` stays operator-authoritative: the machine cannot know what was physically
|
|
loaded into a bay.
|
|
|
|
### 2. Idempotency is explicit, because deltas are not idempotent.
|
|
|
|
Addressable events are re-delivered on reconnect, so a naive delta would be applied twice.
|
|
Every operation carries an operator-minted `id`; the ATM records applied ids and ignores
|
|
duplicates. This is lamassu-server's `pullNewBills` pattern — a client-minted UUID per unit of
|
|
work, making resend free and ordering irrelevant — rather than a sequence number.
|
|
|
|
### 3. The operator publishes a window of recent operations, not one.
|
|
|
|
An event the ATM missed self-heals on the next publish, because the next event still carries
|
|
the earlier operations. This is the same trick as Lightning.Pub piggybacking `latest_balance`
|
|
on every incremental message so a client that missed events corrects itself.
|
|
|
|
### 4. The ATM echoes applied ids back, which is the acknowledgement.
|
|
|
|
The state document carries `applied_ops`, so the dashboard can render each published operation
|
|
as applied or pending. This supplies the feedback leg a replaceable event cannot, without
|
|
needing the transport to report failures.
|
|
|
|
### 4a. Superseded. Before decisions 1 to 4 shipped, the overwrite was warned about.
|
|
|
|
The dashboard's publish dialog stated the failure plainly — that the publish would overwrite
|
|
the ATM's tracked counts, that decrements since the last baseline would be lost, and that it
|
|
should follow a physical refill rather than a mid-day tweak.
|
|
|
|
Kept here rather than deleted, because it is the calibration for how much a warning is worth.
|
|
It was a known, deliberately accepted risk carrying a human-factors mitigation, not an
|
|
oversight, and the product had already reached the same conclusion these decisions formalise.
|
|
It was also the weakest control available: it depended on an operator reading a dialog at the
|
|
end of a refill round, and it could not help at all when the stale value was the one already
|
|
in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53 while a form
|
|
loaded at 54 stayed open, and nothing but that dialog stood between the operator and
|
|
discarding the decrement.
|
|
|
|
The dialog and the endpoint behind it are both gone. The operator dashboard no longer has a
|
|
field that accepts a count, which is a stronger guarantee than any wording could be.
|
|
|
|
### 5. Ordering is decided by `created_at`, never by arrival order, on both sides.
|
|
|
|
The ATM forces each stamp strictly above its last published one, so a same-second publish or a
|
|
clock stepping backwards cannot silently discard a report. spirekeeper applies an event only
|
|
when strictly newer than the **oldest** stamp on file for that machine.
|
|
|
|
Oldest, not newest, because LNbits' `Connection.execute` commits per call: a multi-row apply
|
|
cannot be made atomic through that data layer, so a crash mid-apply leaves some rows advanced.
|
|
Gating on the oldest means a partial apply is re-applied rather than mistaken for a complete
|
|
one, and the ATM's heartbeat makes it converge.
|
|
|
|
### 6. The machine's bay set is authoritative for layout.
|
|
|
|
Bay count is hardware-determined. spirekeeper deletes positions absent from a report rather
|
|
than leaving them; the operator cannot add or remove bays.
|
|
|
|
### 7. Unverified counts are declared, not guessed.
|
|
|
|
When a dispense ends with no per-bay report — a driver throw, or the dispense timeout — bills
|
|
may have reached the customer with nothing knowing how many. The ATM flags
|
|
`counts_uncertain_since` and carries it in the state document rather than letting a number
|
|
known to read high stand as measurement. An operator `recount` clears it.
|
|
|
|
### 8. State is published on every change and on a heartbeat.
|
|
|
|
A publish is one fire-and-forget event with no retry. The heartbeat is what makes the channel
|
|
self-healing after a relay outage, and the only way an out-of-band edit to the table ever
|
|
reaches the operator.
|
|
|
|
## What this replaces
|
|
|
|
The original design published a single hello-event gated on a one-shot flag, deduplicated on
|
|
one remembered event id, and never compared `created_at` at all. In practice that produced:
|
|
|
|
| Failure | Issue |
|
|
| --------------------------------------------------------------- | -------------- |
|
|
| Layout changes after first boot never published | bitspire#94 |
|
|
| Remediation dispenses debited HAL but not the rows | bitspire#76 |
|
|
| Absent positions never deleted; publishes then rejected forever | spirekeeper#43 |
|
|
| A stale dashboard publish overwriting a newer report | spirekeeper#43 |
|
|
| A drained machine advertising bills it had already dispensed | found in audit |
|
|
| A re-delivered A, B, A applied three times | found in audit |
|
|
|
|
The last of these is the only one decisions 5 to 8 do not close, because it is not a defect
|
|
in the mechanism: the operator is permitted to write the count, so a stale write is
|
|
indistinguishable from an intended one. Only decisions 1 to 4 remove it, by removing the
|
|
operator's ability to write counts at all.
|
|
|
|
## Status of implementation
|
|
|
|
Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format,
|
|
and were verified against the deployed code on sintra on 2026-09-22: a zeroed machine reported
|
|
drained rather than freezing its beacon, a re-delivered operator config was dropped as stale on
|
|
eight consecutive restarts, three heartbeat republishes carried strictly increasing stamps read
|
|
back off the relay, and the machine, the relay and the operator dashboard agreed on the counts
|
|
with timestamps correlated to the second.
|
|
|
|
Decisions 1 through 4 are the v2 operations wire and shipped in spirekeeper#46 and
|
|
bitspire#106.
|
|
|
|
On the operator side there is no longer any endpoint that accepts a count: the absolute
|
|
publish, its CRUD write and its request model were removed rather than deprecated. The
|
|
dashboard records operations and renders each as applied or pending from the machine's
|
|
`applied_ops` echo. On the machine side, schema v13 adds a `cassette_ops` dedup ledger, the
|
|
`created_at` watermark on this path is retired in favour of per-op ids, and the state document
|
|
carries `schema_version`, `seq` and `applied_ops`.
|
|
|
|
The wire shapes are those given under decisions 1 and 4 above.
|
|
|
|
Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on
|
|
their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an
|
|
operator refill does not land until it updates — which fails safe, since the machine
|
|
under-counts and will not dispense bills it believes it lacks. In the other direction an
|
|
updated machine drops a v1 absolute-count payload on the missing `ops` array, which is the
|
|
same safe direction: the machine keeps the counts it is now the only writer of.
|
|
|
|
One gap stays open deliberately. An operation recorded while the relay is unreachable waits
|
|
for the operator's next action to be published, because only an operator action triggers a
|
|
publish. The window makes that self-healing once anything is published, but nothing on the
|
|
operator side republishes on its own. An operator-side heartbeat is the fix; it is not built.
|
|
|
|
## Alternatives considered
|
|
|
|
- **Compare-and-swap on absolute writes.** Rejected: see decision 1. Viable only with a
|
|
feedback leg the transport cannot provide, and decision 4 gets the same benefit without
|
|
pretending the transport is something it is not.
|
|
- **NIP-77 negentropy for reconciliation.** Rejected: it reconciles sets of event ids and
|
|
still requires a separate fetch. For a single mutable document it costs more than
|
|
re-reading it.
|
|
- **One envelope carrying all operator config.** Rejected earlier and still right: a fee edit
|
|
that republished a stale cassette inventory is a real failure mode. One d-tag per lifecycle.
|
|
- **Publishing the operation log as kind-78.** Deferred. The operator authors the operations
|
|
and the ATM records what it applied, so both sides already hold an audit trail.
|