Consume operator cassette operations instead of counts #106

Merged
padreug merged 2 commits from feat/cassette-ops-consumer into dev 2026-09-23 21:31:08 +00:00
Showing only changes of commit 65f5f982ec - Show all commits

docs(adr): record that the operations wire shipped

Decisions 1 through 4 are built on both sides, so the status section says
what landed rather than what is planned, and decision 4a is marked
superseded.

4a is kept rather than deleted. It is the calibration for what a warning
dialog is worth: it depended on an operator reading it at the end of a
refill round, and it could not help at all when the stale value was
already in the form. The dialog and the endpoint behind it are both gone
now, which is a stronger guarantee than any wording could be.

Also records the one gap left open on purpose. An operation recorded
while the relay is unreachable waits for the operator's next action,
because only an operator action triggers a publish.
Padreug 2026-09-23 12:56:19 +02:00

View file

@ -65,21 +65,23 @@ The state document carries `applied_ops`, so the dashboard can render each publi
as applied or pending. This supplies the feedback leg a replaceable event cannot, without
needing the transport to report failures.
### 4a. Until decisions 1 to 4 ship, the overwrite is warned about, not prevented.
### 4a. Superseded. Before decisions 1 to 4 shipped, the overwrite was warned about.
The dashboard's publish dialog already states the failure plainly — that the publish will
overwrite the ATM's tracked counts, that decrements since the last baseline will be lost, and
that it should follow a physical refill rather than a mid-day tweak. It also says v2
reconciliation will replace it.
The dashboard's publish dialog stated the failure plainly — that the publish would overwrite
the ATM's tracked counts, that decrements since the last baseline would be lost, and that it
should follow a physical refill rather than a mid-day tweak.
Recording this because it changes how the gap should be read. It is a known, deliberately
accepted risk carrying a human-factors mitigation, not an oversight, and the product had
already reached the same conclusion these decisions formalise. It is worth keeping in mind
that a warning is the weakest control available: it depends on an operator reading a dialog
at the end of a refill round, and it cannot help at all when the stale value is the one
already in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53
while a form loaded at 54 stayed open, and nothing but that dialog stood between the operator
and discarding the decrement.
Kept here rather than deleted, because it is the calibration for how much a warning is worth.
It was a known, deliberately accepted risk carrying a human-factors mitigation, not an
oversight, and the product had already reached the same conclusion these decisions formalise.
It was also the weakest control available: it depended on an operator reading a dialog at the
end of a refill round, and it could not help at all when the stale value was the one already
in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53 while a form
loaded at 54 stayed open, and nothing but that dialog stood between the operator and
discarding the decrement.
The dialog and the endpoint behind it are both gone. The operator dashboard no longer has a
field that accepts a count, which is a stronger guarantee than any wording could be.
### 5. Ordering is decided by `created_at`, never by arrival order, on both sides.
@ -138,13 +140,29 @@ eight consecutive restarts, three heartbeat republishes carried strictly increas
back off the relay, and the machine, the relay and the operator dashboard agreed on the counts
with timestamps correlated to the second.
Decisions 1 through 4 are the v2 operations wire and are not yet built. Decision 4a describes
what stands in for them meanwhile.
Decisions 1 through 4 are the v2 operations wire and shipped in spirekeeper#46 and
bitspire#106.
On the operator side there is no longer any endpoint that accepts a count: the absolute
publish, its CRUD write and its request model were removed rather than deprecated. The
dashboard records operations and renders each as applied or pending from the machine's
`applied_ops` echo. On the machine side, schema v13 adds a `cassette_ops` dedup ledger, the
`created_at` watermark on this path is retired in favour of per-op ids, and the state document
carries `schema_version`, `seq` and `applied_ops`.
The wire shapes are those given under decisions 1 and 4 above.
Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on
their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an
operator refill does not land until it updates — which fails safe, since the machine
under-counts and will not dispense bills it believes it lacks.
under-counts and will not dispense bills it believes it lacks. In the other direction an
updated machine drops a v1 absolute-count payload on the missing `ops` array, which is the
same safe direction: the machine keeps the counts it is now the only writer of.
One gap stays open deliberately. An operation recorded while the relay is unreachable waits
for the operator's next action to be published, because only an operator action triggers a
publish. The window makes that self-healing once anything is published, but nothing on the
operator side republishes on its own. An operator-side heartbeat is the fix; it is not built.
## Alternatives considered