ADR-005 rollout step 1: value-confirmed dispense, fault screens, cash-out latch, dispense-report outbox #123

Merged
padreug merged 6 commits from feat/dispense-outcome into dev 2026-10-10 19:54:56 +00:00
Owner

The bitspire half of ADR-005 (docs/adr/005-cash-out-dispense-outcome.md), the response to #122. Six commits, one layer each.

HAL — every dispenser returns a tagged error: errorCode (family), rawCode (78 42), errorClass (terminal / recoverable / inventory), human decode. F56 decode table seeded from sintra's exit jam and the Tejo's long-bill reject; unknown codes fail safe as terminal. Drops the borrowed statusCode 570. First tests in packages/hal — the bill-length table test would have caught the GTQ window months ago.

State machine — dispensed (driver boolean) → dispenseConfirmed (value equality, computed by the HAL). dispenseError splits into dispenseFault (hardware error; paid, owed; 120 s with evidence) and outOfCash (no hardware error; 30 s). A terminal fault latches cash-out off (cashOutHeld), preserved across resets, released only by CASH_OUT_RELEASED. Payment hash carried into context. 46 tests.

Machine app — value confirmation in both HAL glue paths; cash-out hold persisted in meta, restored on boot, released by a recount or a new resume_cash_out operator op (honoured only if stamped after the hold began); hold mirrored into the cassettes-state doc and the beacon; idle Sell button disabled with the reason. A zero-dispensed report with an error now sets countsUncertainSince instead of being trusted. Fault/out-of-cash screens show "your payment went through", amounts, txid as QR + text, payment hash, time; raw code is not shown.

Outbox — dispense_reports table (schema v14), written in the same SQLite transaction as the transactions row, drained to report_dispense after each persist, on relay reconnect, and every 60 s, acked only on OK, backing off 30 s · 2^attempts (cap 1 h). Sent on success too — the success report is what spirekeeper will capture on.

Not in this PR: the spirekeeper half (ADR rollout step 2 — the report_dispense handler, awaiting_dispense / cash_owed, worklist + notification). Until that lands the RPC rejects as unregistered and the outbox rows simply wait with backoff; you'll see one [ATM] Dispense report … not delivered warning per attempt, then quiet. That is the intended shipping order.

Verified: hal 20/20, state-machine 46/46, lnbits 29/29, nostr-client 43/43, clink 11/11; vue-tsc, electron tsc, preload tsc clean; full pnpm build of apps/machine passes (after the qrcode fix that already went to dev as cbff654).

Behaviour changes on a machine running this: a terminal dispenser fault takes cash-out out of service until an operator recounts or publishes resume_cash_out; customers who pay and don't receive cash see a reference screen instead of a 30 s countdown; the kiosk theme preference resets once (storage key renamed earlier today).

Refs #122, #27.

The bitspire half of ADR-005 (`docs/adr/005-cash-out-dispense-outcome.md`), the response to #122. Six commits, one layer each. **HAL** — every dispenser returns a tagged error: `errorCode` (family), `rawCode` (`78 42`), `errorClass` (`terminal` / `recoverable` / `inventory`), human decode. F56 decode table seeded from sintra's exit jam and the Tejo's long-bill reject; unknown codes fail safe as terminal. Drops the borrowed `statusCode 570`. First tests in `packages/hal` — the bill-length table test would have caught the GTQ window months ago. **State machine** — `dispensed` (driver boolean) → `dispenseConfirmed` (value equality, computed by the HAL). `dispenseError` splits into `dispenseFault` (hardware error; paid, owed; 120 s with evidence) and `outOfCash` (no hardware error; 30 s). A terminal fault latches cash-out off (`cashOutHeld`), preserved across resets, released only by `CASH_OUT_RELEASED`. Payment hash carried into context. 46 tests. **Machine app** — value confirmation in both HAL glue paths; cash-out hold persisted in `meta`, restored on boot, released by a `recount` or a new `resume_cash_out` operator op (honoured only if stamped after the hold began); hold mirrored into the cassettes-state doc and the beacon; idle Sell button disabled with the reason. A zero-dispensed report *with* an error now sets `countsUncertainSince` instead of being trusted. Fault/out-of-cash screens show "your payment went through", amounts, txid as QR + text, payment hash, time; raw code is not shown. **Outbox** — `dispense_reports` table (schema v14), written in the same SQLite transaction as the `transactions` row, drained to `report_dispense` after each persist, on relay reconnect, and every 60 s, acked only on OK, backing off 30 s · 2^attempts (cap 1 h). Sent on success too — the success report is what spirekeeper will capture on. **Not in this PR:** the spirekeeper half (ADR rollout step 2 — the `report_dispense` handler, `awaiting_dispense` / `cash_owed`, worklist + notification). Until that lands the RPC rejects as unregistered and the outbox rows simply wait with backoff; you'll see one `[ATM] Dispense report … not delivered` warning per attempt, then quiet. That is the intended shipping order. **Verified:** hal 20/20, state-machine 46/46, lnbits 29/29, nostr-client 43/43, clink 11/11; vue-tsc, electron tsc, preload tsc clean; full `pnpm build` of apps/machine passes (after the `qrcode` fix that already went to `dev` as cbff654). **Behaviour changes on a machine running this:** a terminal dispenser fault takes cash-out out of service until an operator recounts or publishes `resume_cash_out`; customers who pay and don't receive cash see a reference screen instead of a 30 s countdown; the kiosk theme preference resets once (storage key renamed earlier today). Refs #122, #27.
Every dispenser now returns a tagged DispenseError: errorCode (the
family name, e.g. F56DispenseError), rawCode (driver-native, '78 42'),
errorClass (terminal | recoverable | inventory) and a human decode. The
class is what the state machine routes on: terminal latches cash-out off,
recoverable shows the fault screen but stays in service, inventory means
nothing was asked of the hardware.

The F56 table is built empirically and from the Fujitsu F56-BDU Error
Code List, seeded with sintra's 78 42 (note stopped at the cassette exit,
terminal) and the Tejo's 82 00 (long-bill reject, recoverable), plus the
83/84/86 00 checks and the 85 0n / B5 .. families. An unknown code fails
SAFE — terminal — so an unfamiliar fault latches rather than letting the
next customer pay into it. f56-rs232 surfaces the raw code structurally
instead of only inside the message string.

Drops the borrowed statusCode 570 from the F56 driver: lamassu-server
read 570 as "insufficient funds", so a jam told operators to refill full
cassettes (their 34ba9203 fix). Puloon gets the same contract with every
fault terminal until it has a decode table.

packages/hal had no tests at all (ADR-005 finding 10). Adds the first
two: the decode table, and the bill-length table — every window [hi, lo]
sane, and GTQ/USD(/HNL when added) sharing one window for what is
physically the same 156 mm note. That second test would have caught the
GTQ fault months ago.
DispenseCashResult.dispensed (a driver boolean) is replaced by
dispenseConfirmed — Σ(denomination × dispensed) equals the requested
value, computed by the HAL — plus errorCode / rawCode / errorClass.
dispensingCash.onDone guards on dispenseConfirmed and nothing else.

The single dispenseError state becomes two. dispenseFault: the dispenser
reported an error, the customer has paid and is owed — 120 s screen with
evidence, ACKNOWLEDGE_FAULT to dismiss. outOfCash: a shortfall with no
hardware error or an inventory refusal — 30 s. A hung dispense is a
terminal fault.

A terminal errorClass latches cash-out off: context.cashOutHeld, set by
latchCashOutIfTerminal, preserved across resetContext (it is machine
health, not transaction state), guarding idle's SELECT_CASH_OUT. Cash-in
is unaffected. Only CASH_OUT_RELEASED clears it — the store sends that
when an operator recount or resume_cash_out op lands; re-initialising
the dispenser never does, because re-init does not move a stuck note.
CASH_OUT_HELD lets the store restore a persisted hold on boot.

PAYMENT_RECEIVED now carries the payment hash into context.paymentHash
so the fault screen can show the reference the server indexes.

Tests: the dispense section is rewritten around outcomes — value
confirmation, fault vs out-of-cash routing, terminal latch + release,
recoverable does not latch, partial-with-error is a fault, inventory
refusal is out-of-cash, boot-restored hold gates, 30 s vs 120 s timers,
acknowledge/cancel, timeout latches. 46/46.
HAL glue (electron/hal-service.ts and the renderer-side services/hal.ts):
dispenseConfirmed is Σ(denomination × dispensed) === Σ(denomination ×
requested), computed on value. The driver's tagged error is carried
through as errorCode / rawCode / errorClass / human; pre-dispense
inventory refusals are errorClass 'inventory' so they route to outOfCash
rather than the fault screen. The manual-dispense command result keeps
its wire key `dispensed` (spirekeeper's poller reads it) and gains the
new fields alongside.

Cash-out hold: state-store persists it in meta as one JSON value beside
countsUncertainSince, idempotent on set (the first fault's `since` is
kept); IPC get/set/clear through preload. The store persists the hold the
moment the machine sets it and restores it into the machine on boot. A
recount clears it in the store (same gesture that clears counts-
uncertain); operator-config also honours a new resume_cash_out op — not a
cassette op, split off before applyOperatorCassetteOps, and honoured only
when stamped after the hold began so a re-delivered old resume cannot
clear a fresh fault. Either release calls back into the store, which
sends CASH_OUT_RELEASED. The cassettes-state document carries
cash_out_held_since / _reason / _code (additive, like
counts_uncertain_since); the availability beacon reports cash_out false
while held; the idle Sell button is disabled with the reason.

Store watcher: dispenseFault and outOfCash both record dispense_error /
partial (the customer has paid either way). A report of zero dispensed
WITH a hardware error now sets countsUncertainSince instead of being
trusted as zero — a note stopped in the transport completes neither
counter (sintra 2026-10-09: bay read 66, held 65, one in the transport).

Fault screen: both terminal states show "your payment went through",
amount paid, per-denomination dispensed, the txid as QR and text, the
payment hash (threaded from the settlement watch through PAYMENT_RECEIVED)
and the time, with "keep this reference" and an acknowledge button. The
raw dispenser code is not shown; it travels in the report.
One cash-out's dispense outcome, sent on success as well as failure —
the success report is what captures the settlement server-side. Field
names follow lamassu-server's cash_out_txs / cash_out_actions
(dispense_confirmed, error, error_code) with raw_code and error_class
alongside, per-denomination bills with `requested`, per-bay cassettes
verbatim, the payment hash as the join key, and counts_uncertain.

Idempotent on txid (the server upserts), so the call is wrapped in
idempotent() and safe for the machine's outbox to retry. Until
spirekeeper registers the RPC it rejects with LnbitsRpcError, which the
outbox treats like any other transient failure.
Every cash-out now produces one report_dispense — on success as well as
failure — and the machine does not stop sending it until spirekeeper
acknowledges it.

state.db gains a dispense_reports table (migration v13 → v14): the report
is written INSIDE recordTransaction's SQLite transaction, alongside the
transactions row, so a crash between the two cannot lose it. Rows carry
attempts / last_attempt_at / last_error / acked_at. Three IPC calls
(pending / ack / note-attempt) expose it to the renderer.

The store builds the report when a cash-out reaches complete,
dispenseFault or outOfCash: txid, payment hash, dispense_confirmed,
error / error_code / raw_code / error_class, per-denomination requested
vs dispensed vs rejected, the per-bay cassette record verbatim, and
counts_uncertain. The success report is what lets the server capture
(distribute) the settlement; the failure report is what puts a customer
on the owed-cash worklist instead of leaving the only record on the ATM.

Delivery is at-least-once: a flusher drains pending rows after each
persist, on relay (re)connect, and every 60 s, acking only on an OK reply
and backing off 30 s · 2^attempts (capped 1 h) otherwise. While
spirekeeper has not registered the RPC every send fails the same way; the
backoff keeps that quiet and the rows wait — this half ships first.

The lightning service exposes reportDispense; the function pointer is
set at all three lightning-init sites so the flusher works on every path.
padreug deleted branch feat/dispense-outcome 2026-10-10 19:54:56 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/bitspire!123
No description provided.