Publish cassette operations instead of counts #46

Merged
padreug merged 8 commits from feat/cassette-ops-publisher into main 2026-09-23 21:16:34 +00:00
Owner

Phase 3 operator side: the dashboard records cassette operations instead of writing counts.

The operator can no longer set a count, because there is no longer an endpoint that accepts one. POST .../cassettes/ops records a refill, an empty, a recount or a set_denomination and publishes the machine's recent window; GET .../cassettes/ops lists them with acked_at so the dashboard shows each as applied or pending. The old publish endpoint, update_cassette_config and UpsertCassetteConfigData are gone.

Why: both sides used to write the same value over a transport that never tells a writer it lost. A dashboard form loaded before a dispense discarded that dispense on publish and nothing could detect it. Confirmed live on sintra on 2026-09-22.

Also here:

  • m013 cassette_ops, the append-only op log. The id is the idempotency key the machine dedups on.
  • m014 stores counts_uncertain_since. The machine has been publishing it since v1 and we were parsing it into a field nobody read, which defeats the point: it exists to tell a human to open the bay and recount.
  • m015 stores the machine's seq and uses it to break a same-second tie in the ordering gate. created_at is one-second granular, and a dispense plus the publish after it share a second routinely.

The op is recorded before the publish and is not rolled back if the publish fails. Notes went into a bay whether or not a relay was reachable. Each publish carries a window rather than just the newest op, so one that missed its own publish rides out with the next.

Deploy this before merging aiolabs/bitspire#106 (the machine consumer). Strict cutover, no compatibility code.

264 tests pass. Needs an LNbits restart on ariege plus a catalog release.

Design: aiolabs/bitspire docs/adr/004-cassette-state-synchronization.md.

Phase 3 operator side: the dashboard records cassette operations instead of writing counts. The operator can no longer set a count, because there is no longer an endpoint that accepts one. POST `.../cassettes/ops` records a refill, an empty, a recount or a set_denomination and publishes the machine's recent window; GET `.../cassettes/ops` lists them with `acked_at` so the dashboard shows each as applied or pending. The old publish endpoint, `update_cassette_config` and `UpsertCassetteConfigData` are gone. Why: both sides used to write the same value over a transport that never tells a writer it lost. A dashboard form loaded before a dispense discarded that dispense on publish and nothing could detect it. Confirmed live on sintra on 2026-09-22. Also here: - m013 `cassette_ops`, the append-only op log. The id is the idempotency key the machine dedups on. - m014 stores `counts_uncertain_since`. The machine has been publishing it since v1 and we were parsing it into a field nobody read, which defeats the point: it exists to tell a human to open the bay and recount. - m015 stores the machine's `seq` and uses it to break a same-second tie in the ordering gate. `created_at` is one-second granular, and a dispense plus the publish after it share a second routinely. The op is recorded before the publish and is not rolled back if the publish fails. Notes went into a bay whether or not a relay was reachable. Each publish carries a window rather than just the newest op, so one that missed its own publish rides out with the next. Deploy this before merging aiolabs/bitspire#106 (the machine consumer). Strict cutover, no compatibility code. 264 tests pass. Needs an LNbits restart on ariege plus a catalog release. Design: aiolabs/bitspire `docs/adr/004-cassette-state-synchronization.md`.
First piece of the v2 wire (bitspire ADR-004). The operator stops
publishing counts and starts publishing what it DID; the machine, which
holds the notes, keeps the running total. A value with one writer cannot
be clobbered, which is the whole point: the absolute-count wire let a
form loaded before a dispense discard that dispense when published, and
nothing in an addressable event can tell the loser it lost.

m013 adds cassette_ops, append-only. The id is minted here and is the
idempotency key the machine dedups on, because a delta applied twice is
wrong and addressable events are re-delivered on reconnect. acked_at is
set when the machine reports that id back, which is the only
acknowledgement this transport can carry.

The models enforce that an op carries exactly the one field its type
means, so an instance is publishable by construction — the same contract
FeeConfigPayload has — and nulls never reach the wire for the machine to
disambiguate. recount is the only absolute, deliberately: it is what an
operator opening a bay and counting actually does, and it stays
auditable as its own act rather than looking like a stale form.

Vocabulary mirrors lamassu-server's cash_unit_operation_type.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Append-only, and deliberately nothing here writes cassette_configs. That
table now holds only what the machine has reported; letting an operation
write it would put back the second writer this whole design exists to
remove.

get_cassette_ops_window returns oldest-first because order is meaning: a
recount followed by a refill is not the same as the reverse. It takes the
most recent N and reverses, so the window slides without the machine ever
seeing them out of sequence.

The window is what makes the channel self-healing, so it has to cover a
plausible outage rather than just the newest change — a machine that
missed one event still sees the operation in the next.

_should_ack_op is extracted pure, the same way the state-event gate is,
because three of its rules are easy to get wrong and none need a database
to test: an unknown id closes out nothing, one machine must never be able
to ack another machine's operation, and the FIRST acknowledgement is the
one worth keeping. That last one matters because the machine echoes a
window, so every id comes back many times over; overwriting would keep
sliding the timestamp forward and lose when the operation actually landed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The v2 operator to ATM wire. Same kind-30078 document and the same d-tag
the counts wire used, because the machine subscribes by that tag and the
document is addressable, so v2 replaces v1 in place.

Sends a WINDOW of recent operations, oldest-first, not just the newest
change. Each publish replaces the last, so a machine that was offline for
one of them would otherwise never see that operation again; carrying the
recent history means the channel heals itself without anyone noticing it
broke. Re-delivery costs nothing because every op carries an id the
machine dedups on.

Tests pin the contract rather than the implementation: the d-tag, that
the payload declares v2 and carries no positions key, that window order
survives the publisher untouched, that an empty window still ships a
well-formed document so a machine can tell "no operations" from "operator
still on v1", and that an npub entered in the UI is normalised to hex —
get that last one wrong and the machine's subscription filter silently
never matches.

Additive. The endpoints still publish counts until the next commit, so
the tree is not left half-switched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The machine echoes the operation ids it has applied in its state
document, and this records them. That echo is the only acknowledgement
this transport can carry: an addressable event gives its publisher no
failure signal at all, since the relay returns OK for an event it then
discards. Without it the dashboard could only ever show an operation as
sent, never as delivered.

Deliberately not gated on whether the state event advanced the counts.
The machine echoes its applied ids on every publish, heartbeats included,
so an event carrying nothing new about the counts can still be the first
one to tell us an operation landed.

The consumer goes in before the producer on purpose. The machine does not
send applied_ops yet, and every new field on the state payload defaults to
a value meaning "this machine does not report that yet" rather than to
one that would be wrong — an absent list reads as nothing acknowledged,
which is exactly right for a machine that has applied nothing.

Also picks up seq and counts_uncertain_since, which the machine already
publishes and this side was dropping on the floor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The operator can no longer write a count. POST .../cassettes/ops records
one operation — refill, empty, recount, set_denomination — and publishes
the machine's recent window; GET .../cassettes/ops lists them newest
first with acked_at, so the dashboard can tell a delivered operation from
one merely sent.

POST .../cassettes/publish is gone, along with update_cassette_config and
UpsertCassetteConfigData. Nothing in the operator can now set a count,
which is the point: a value with one writer cannot be clobbered. Under
the old endpoint a dashboard form loaded before a dispense silently
discarded that dispense on publish, and neither side could detect it —
addressable events order by created_at at second granularity and a relay
returns OK for an event it then drops, so the losing writer is never told.

The op is recorded before the publish and is deliberately not rolled back
when the publish fails. It records something that physically happened;
notes went into a bay whether or not a relay was reachable. The window
carries recent operations rather than just the newest, so an op that
missed its own publish rides out with the next one.

Validation rejects an unpaired machine and a position the machine has not
reported. Bay count stays hardware-determined.
The machine has been publishing counts_uncertain_since since v1 of the
state document and spirekeeper has been parsing it into a field nobody
read. That defeats the point of the marker: it exists so a human opens
the bay and recounts.

A dispenser can throw, or time out, after notes have physically moved.
The machine cannot know how many left, so rather than decrement a number
it would be guessing at, it stamps the moment (bitspire ADR-004,
decision 3). m014 gives that stamp a home on dca_machines, and the
consumer mirrors it on every state event.

Stored on the machine rather than the bay because the uncertainty is
about the dispense as a whole; a multi-bay dispense that fails midway
gives no reliable way to attribute it to one position.

Written through even when the machine reports None. The machine clearing
the marker is as important as setting it — the operator recounted, the
bay is trustworthy again — and a banner that never goes away is a banner
nobody reads.
The cassettes tab no longer has editable count fields, because there is
no longer an endpoint that would accept them. The bays render read-only
from the machine's own report, and a Record-operation dialog captures
what the operator did: a refill in notes added, an empty, a recount, a
denomination change.

This removes the failure the tab used to invite. A form loaded before a
dispense held a count that was already wrong, and publishing it
overwrote the dispense with no error on either side. Recording a delta
instead means a dispense that happened while the dialog was open is kept
rather than discarded, and a recount is now an explicit act — what an
operator opening a bay and counting actually does — rather than being
indistinguishable from a stale form.

A recent-operations list shows each one as Applied or Pending from
acked_at, which is the machine echoing the id back. Pending needs no
retry button: every publish carries the recent window, so an operation
that missed its own publish keeps being re-offered until it lands, and
saying so in the panel is more useful than a button that would do
nothing new.

The counts-uncertain banner tells the operator when the machine cannot
vouch for its own numbers and asks for the recount that clears it. The
machine row is re-read on every cassette refresh, since that flag is set
by the consumer while the dialog is open.
fix(cassettes): break the same-second tie with the machine's counter
Some checks failed
ci.yml / fix(cassettes): break the same-second tie with the machine's counter (pull_request) Failing after 0s
5f60b3fe31
The ordering gate compares created_at, which NIP-01 defines at one-second
granularity. A dispense and the publish that follows it land inside one
second routinely, so the report was dropped and the operator kept the
pre-dispense count until the next heartbeat five minutes later.

The machine bumps a counter on every local change to a bay count and
carries it in its state document. m015 stores it per row, and the gate
consults it only when the stamps are equal, where created_at carries no
information at all.

Only on equality, deliberately. A machine whose state.db was replaced
restarts its counter at zero while its wall clock keeps moving forward;
gating on the counter across different stamps would lock that machine out
for good. Equal stamps with no counter on either side stay closed, which
costs one heartbeat and risks nothing.
padreug deleted branch feat/cassette-ops-publisher 2026-09-23 21:16:34 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/spirekeeper!46
No description provided.