fix(cassettes): close the machine-side divergence paths #104

Merged
padreug merged 5 commits from fix/cassette-sync-machine into dev 2026-09-22 20:30:24 +00:00
Owner

Phase 1 of the cassette synchronization audit. Machine-side only, no wire format change, so it ships independently of the spirekeeper half.

Five failures, each its own commit:

  • A remediation dispense (manual_dispense) debited HAL's memory but not the cassette rows, and HAL re-seeds from those rows at boot, so the machine came back believing it held bills already handed to a customer. Closes #76.
  • getInventory dropped zero-count bays, so a drained machine read as "nothing known" and callers fell back to a stale snapshot. The availability beacon went on advertising bills that were gone.
  • The kind-21003 management dispense and the main-process command poller both moved bays without refreshing the renderer or republishing.
  • The state publish was gated on a one-shot flag, so layout changes after first boot were never announced, and a publish lost to a relay outage was lost for good. Now published on every start and on a heartbeat, with each stamp forced above the last so a same-second publish or a backwards clock cannot silently discard a report. Closes #94.
  • A dispense that ends with no per-bay report now records that the counts are unverified rather than leaving a number known to read high as fact.

Typecheck clean; 125 tests; full Electron build green. The new decrement tests were confirmed to fail against the unfixed code.

Phase 1 of the cassette synchronization audit. Machine-side only, no wire format change, so it ships independently of the spirekeeper half. Five failures, each its own commit: - A remediation dispense (manual_dispense) debited HAL's memory but not the cassette rows, and HAL re-seeds from those rows at boot, so the machine came back believing it held bills already handed to a customer. Closes #76. - getInventory dropped zero-count bays, so a drained machine read as "nothing known" and callers fell back to a stale snapshot. The availability beacon went on advertising bills that were gone. - The kind-21003 management dispense and the main-process command poller both moved bays without refreshing the renderer or republishing. - The state publish was gated on a one-shot flag, so layout changes after first boot were never announced, and a publish lost to a relay outage was lost for good. Now published on every start and on a heartbeat, with each stamp forced above the last so a same-second publish or a backwards clock cannot silently discard a report. Closes #94. - A dispense that ends with no per-bay report now records that the counts are unverified rather than leaving a number known to read high as fact. Typecheck clean; 125 tests; full Electron build green. The new decrement tests were confirmed to fail against the unfixed code.
recordTransaction only debited the cassette rows for type 'cash_out'.
An operator remediation is recorded as 'manual_dispense', so HAL's
in-memory bays went down while the persisted rows did not — and HAL
re-seeds from those rows on the next boot, so the machine came back
believing it still held bills a customer had already been handed.

A remediation against a partly-dispensed original debits again on
purpose: the original only ever debited what physically left, and this
is a second lot of bills leaving the bay.

Closes #76

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
getInventory dropped zero-count bays, so a fully dispensed machine
returned an empty map — identical to a machine with no cassettes
configured. Every caller reads an empty map as "nothing known, ask the
hardware": reloadPersistedInventory skipped the update entirely, so the
last non-empty snapshot stuck and the public availability beacon went on
advertising bills that had already gone out the slot.

Zero-count bays are kept, so an empty map now means exactly one thing:
no cassettes are configured. Consumers already filter for > 0 before
offering a denomination. loadInventoryFromDb returns null when the DB
could not be asked at all (browser dev, failed IPC) so callers can still
tell "no answer" from an answer of "the bays are empty", and only the
former defers to HAL.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two dispense paths bypassed the refresh-and-publish step that cash-out
does. The kind-21003 management command persisted the transaction and
stopped there, and the operator-command poller runs entirely in the main
process, where the renderer cannot see the bays move at all. In both
cases the renderer kept serving a stale inventory and the operator's
cassette view stayed frozen until the next customer cash-out.

The management handler takes an after-hook, and the main process emits
'cassettes:changed' when it mutates the table so the renderer can catch
up. Both land on one helper that reloads the inventory and republishes
the state document.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three linked failures in one mechanism, so one commit.

The state publish was gated on a one-shot 'have we said hello' flag. It
fired once on first boot and then only after a dispense or an applied
operator config, so any change to the layout itself — a reseed, an
atm-tui edit, direct SQL — was never announced. The operator kept
validating against a bay set the machine no longer had, and a publish
from the dashboard could overwrite a fresh seed (#94). State is now
published on every start.

A publish is one fire-and-forget event with no retry. If the relay was
unreachable at the moment of a dispense, that update was gone until the
next customer bought cash. A five-minute heartbeat makes the channel
self-healing and is also the only way an out-of-band edit to the table
ever reaches the operator.

Addressable events are ordered by created_at at second granularity with
ties broken by lowest event id, and a relay acknowledges an event it
then discards. Two publishes inside one second therefore left the winner
decided by a hash, permanently, and a clock stepping backwards would
have made every report from this machine vanish silently. Each publish
now takes a stamp strictly above the last, recorded in the meta row that
used to hold the gate — same key, no migration, honest name.

Closes #94

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
When the dispenser throws or the dispense times out there is no per-bay
report, so nothing is debited — not the cassette rows, not HAL's bays.
Bills may well have reached the customer, and both counters then read
high with nothing to indicate it. The machine went on treating a number
it had reason to doubt as measurement.

A dispense that ends with no report now latches a countsUncertainSince
flag, which rides along in the state document as counts_uncertain_since
so the operator can see the numbers need a recount. The field is
additive: a consumer reading positions ignores it, so this needs no
coordinated release. An operator config apply clears the flag inside the
same transaction, since asserting authoritative counts is precisely what
a recount is.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
padreug deleted branch fix/cassette-sync-machine 2026-09-22 20:30:24 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/bitspire!104
No description provided.