From 483e44aaa964c560c5b63c701d03fb5ab3751bdb Mon Sep 17 00:00:00 2001 From: Padreug Date: Fri, 9 Oct 2026 16:34:54 +0200 Subject: [PATCH] =?UTF-8?q?docs(adr):=20ADR-005=20=E2=80=94=20cash-out=20d?= =?UTF-8?q?ispense=20outcome=20and=20settlement=20capture?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A customer paid a 40 EUR cash-out, a note jammed at the cassette exit, and the dashboard showed `processed`. The machine had recorded the failure correctly. Nothing it knew ever left the box. The root is ordering, not display: spirekeeper spawns process_settlement the instant the payment lands, which is before the machine has begun to dispense. The legs are paid sub-second; the dispense fails afterwards; and the one remediation tool refuses once any leg has completed. It is unreachable for the exact case it was built for. ADR-005 makes payment the authorization and dispense confirmation the capture — distribution waits for the machine's report. The report is a report_dispense RPC (not the state doc: ADR-004's losing-writer problem), carried through a durable outbox, sent on success and failure, adopting lamassu's dispense_confirmed / error / error_code taxonomy and its per-bay action log — which the machine already records in cassette_bills and simply never ships. Deviates from lamassu in three places it got wrong or never did: a zero-dispensed report that arrives with an error is treated as unverified, not as zero; a mechanical fault is its own customer screen with evidence and is not "out of cash"; and terminal dispenser faults latch cash-out off until a recount or an explicit operator op, because re-initialising does not move a stuck note. Closes the review loop with ten findings outside the ADR's decisions. Refs #122, #27, #78 --- docs/adr/005-cash-out-dispense-outcome.md | 315 ++++++++++++++++++++++ 1 file changed, 315 insertions(+) create mode 100644 docs/adr/005-cash-out-dispense-outcome.md diff --git a/docs/adr/005-cash-out-dispense-outcome.md b/docs/adr/005-cash-out-dispense-outcome.md new file mode 100644 index 0000000..43378ce --- /dev/null +++ b/docs/adr/005-cash-out-dispense-outcome.md @@ -0,0 +1,315 @@ +# ADR-005: Cash-Out Dispense Outcome and Settlement Capture + +**Status:** Proposed +**Date:** 2026-10-09 +**Context:** On 2026-10-09 a customer paid a 40 EUR cash-out on sintra, a note jammed at the +cassette exit, and the operator dashboard showed the settlement as `processed` with no sign +anything was wrong (aiolabs/bitspire#122). The machine had recorded the failure correctly and +in detail. Nothing it knew ever reached anyone. This ADR specifies how a dispense outcome +becomes a first-class fact on both sides of the wire, and fixes the structural reason the +existing remediation tool could not have helped. + +## The problem + +A cash-out moves value in two steps that today are not connected: + +1. **Payment.** The customer pays the ATM's BOLT11 invoice. LNbits lands it in the machine + wallet, `spirekeeper._handle_payment` verifies attribution, inserts a `dca_settlements` row, + and — in the same breath — spawns `process_settlement` as a background task. +2. **Dispense.** The machine, which learns of the payment through `watchInvoice`, commands the + dispenser. The hardware reports per-bay `dispensed` / `rejected` counts and, on failure, an + error code. + +Step 1 does not wait for step 2. `process_settlement` pays the super fee, the operator's +commission splits, and the DCA legs the moment the payment lands, which on an LNbits-internal +transfer is sub-second. The dispense begins afterwards. So by the time the F56 reported +`78 42` at T+2 s, the settlement's legs were already `completed` and its status was +`processed` — which is what the dashboard faithfully displayed. `processed` means *all +distribution legs paid*. It has never meant *cash reached a hand*, because the server has no +input that could tell it. + +The tool built for this situation, `apply_partial_dispense_and_redistribute`, carries a hard +guard: it refuses once any leg has completed, because a Lightning payment cannot be clawed +back. Under the current ordering that guard is reached on every real failure. The remediation +is structurally unreachable for the exact case it was written for, except by winning a race +against a sub-second transfer. + +Downstream of that, four smaller gaps compound it (all in #122): + +- The machine's `dispenseError` state is terminal and local. No report, no notification. The + only record that a customer is owed money lives in `state.db` on the ATM. +- The dispenser's counters do not see a note that leaves the bay and stops in the transport — + it is neither `dispensed` nor `rejected`. The machine trusts the resulting `dispensed: 0`, + leaves the bay count untouched, and republishes it as fact. The `countsUncertainSince` + safety net fires only when the report is *absent*, not when it is present and wrong. +- Nothing reads dispenser health. The availability beacon derives `cash_out` from + `totalBills > 0` alone and kept advertising a jammed machine as available. (Nothing + consumes that beacon today, so it could not have been the enforcement point in any case.) +- The customer sees a 30-second countdown and a txid QR, then the idle screen. There is no + claim reference, no statement that they have paid, and nothing distinguishes a mechanical + fault from an out-of-cash condition. + +## Prior art + +lamassu-machine / lamassu-server ran this exact hardware in production for a decade. Their +model, which this ADR adopts where it fits and deviates from where it is wrong: + +- **Three fields on the transaction:** `error` (human message), `error_code` (the error's + *name*, machine-readable), `dispense_confirmed` (boolean). `dispense_confirmed` is computed on + **value** — `tx.fiat.eq(Σ denomination × dispensed)` — not taken from the driver. +- **An append-only action log, `cash_out_actions`,** one row per dispense attempt with + per-bay `provisioned_N` / `denomination_N` / `dispensed_N` / `rejected_N`, written by + `logDispense` as `action: 'dispense'` or `'dispenseError'` purely on whether `error` is set. +- **The operator is notified in the same atomic block that logs the dispense** + (`notifyOperator`, cash-out-atomic.js). Push, not a worklist. +- **A mechanical fault is not "out of cash."** Their 2026-09-29 fix routes a + dispenser-reported error to the screen that asks the customer to photograph their receipt, + *"the right prompt when they have paid and are owed money,"* and reserves `outOfCash` for a + shortfall with no error. The same commit removed a borrowed `statusCode 570` from the F56 + driver because 570 meant "insufficient funds" to the server — a jammed BDU was being + reported as a hot-wallet problem. +- **What they did not have:** any notion of disabling a machine on a dispenser fault. + `getMachineStatuses` is inferential — ping age, stuck-screen age — and `cashOut` is an + operator config toggle. A jammed dispenser on a responsive machine reads *Fully + functional*. That is sintra's beacon exactly, so the latch below is new work, not a port. +- **What they got wrong and we will not copy:** `dispenseOccurred(bills)` returns true if the + bill entries merely *have* `dispensed` and `rejected` keys, and `updateCassettes` then + decrements by those numbers. A jam reporting `dispensed: 0` passes and decrements by zero. + Their cassette counts drift the same way sintra's did. + +## Decisions + +### 1. Payment is authorization. Dispense confirmation is capture. Distribution waits for capture. + +A `cash_out` settlement lands as `pending` exactly as now, but `_handle_payment` no longer +spawns `process_settlement` for it. The row moves to a new status, `awaiting_dispense`, and +stays there until the machine reports. + +| Machine reports | Settlement becomes | Then | +| ----------------------------------------- | ------------------ | --------------------------------------- | +| `dispense_confirmed: true` | `pending` | claim + distribute → `processed` | +| partial (some notes out, value short) | `partial_pending` | operator confirms → distribute scaled | +| `dispense_confirmed: false`, nothing out | `cash_owed` | legs never run; funds stay in wallet | +| no report within `DISPENSE_REPORT_TTL` | `dispense_unreported` | worklist; operator investigates | + +This is the card-processing shape — authorize, then capture — and it is the same ordering +lamassu-server enforces between `dispense_confirmed` and `updateCassettes`. The cost is that +operator and DCA legs land seconds later than they do today, which is the dispense time. +The benefit is that `apply_partial_dispense_and_redistribute` is always reachable, because +no leg has run yet, and `cash_owed` is a state the money has not left. + +`cash_in` settlements are unaffected: there is no dispense to wait for, and +`_pay_dca_distributions` already branches on `tx_type` for exactly this kind of asymmetry. + +**Rejected:** keeping immediate distribution and adding a compensating reversal. Internal +legs *are* reversible — they are LNbits-internal invoices, so a compensating internal +payment is mechanically possible and the guard's "Lightning can't be clawed back" is only +true of the `autoforward` leg. But undoing money movement is strictly harder than not +moving it yet, and the autoforward leg stays irreversible either way. Compensation is kept +as a secondary tool for settlements that distributed before this ADR landed. + +### 2. The machine reports every cash-out outcome over a `report_dispense` RPC. + +Not over the kind-30078 state document. ADR-004 established why: an addressable event gives +its publisher no failure signal, and a losing writer is never told. A per-transaction +outcome is an append-only fact that must be acknowledged, which is a request/reply. + +The RPC follows `create_withdraw` and `get_machine_config`: `register_rpc` at +`AUTH_ACCOUNT`, identity taken from the **verified** `sender_pubkey`, never from the body. +It is sent on **success as well as failure** — a success report is what captures (Decision +1). Payload, adopting the lamassu field names: + +```jsonc +{ + "txid": "tx_mv0madw6_wdhtea1v", + "payment_hash": "6f216df32c36…", + "tx_type": "cash_out", + "dispense_confirmed": false, // value equality, Decision 3 + "error": "Dispensing, code: 78 42", // human, null on success + "error_code": "F56DispenseError", // the error's NAME, null on success + "raw_code": "78 42", // driver-native, for the decode table + "error_class": "terminal", // "terminal" | "recoverable" | null, Decision 5 + "fiat_cents": 4000, + "bills": [{ "denomination": 20, "requested": 2, "dispensed": 0, "rejected": 0 }], + "cassettes": [ // the machine's cassette_bills rows, verbatim + { "position": 2, "denomination": 20, "provisioned": 2, "dispensed": 0, "rejected": 0 } + ], + "counts_uncertain": true, // Decision 3 + "at": 1791529353 +} +``` + +**Delivery is at-least-once with a durable outbox.** The machine writes the report to +`state.db` in the same transaction as the `transactions` row (`dispense_reports`: +`txid PRIMARY KEY, payload, created_at, acked_at`), then sends. It resends on boot, on relay +reconnect, and on a timer until an `OK` reply sets `acked_at`. The server upserts on `txid`, +so a resend is a no-op. This is the cassette-ops idempotency pattern applied to the other +direction. + +The server stores every report in an append-only `dispense_reports` table (one row per +attempt, lamassu's `cash_out_actions` shape, keyed to the settlement by `bitspire_txid`, +which is already populated from `extra.txid`), copies the per-bay detail into +`dca_settlements.bills_json` / `cassettes_json` (columns that exist today and are never +written), and sets `dispense_confirmed`, `error`, `error_code` on the settlement. + +### 3. `dispense_confirmed` is computed on value, separately from `error`, and a zero report with an error is unverified. + +On the machine, after the HAL returns: + +``` +confirmed = requestedFiatCents === Σ(denomination × dispensed) × 100 +``` + +`error` is carried independently. The state machine's `dispensingCash.onDone` guard moves +from `output.dispensed === true` to `output.dispenseConfirmed`, and `DispenseCashResult` +gains `dispenseConfirmed`, `errorCode`, `rawCode`, `errorClass`. + +**The deviation from both bitSpire-today and lamassu:** when a report arrives with `error` +set and `dispensed === 0` on every bay, the machine does *not* treat that zero as a count. A +note in the transport path completes neither counter. The machine sets `countsUncertainSince` +exactly as it already does for an absent report, carries `counts_uncertain: true` in the RPC, +and the bay stays flagged until a `recount` op clears it. `atm-reconcile` then shows the gap +instead of a clean ledger. + +### 4. A dispenser fault is its own customer screen, with evidence, and it is not "out of cash." + +Two distinct terminal states replace the single `dispenseError`: + +- **`outOfCash`** — the request could not be met from inventory and the dispenser reported + **no error**. Nothing was charged beyond what was dispensed. +- **`dispenseFault`** — the dispenser reported an error. The customer **has paid** and is + owed the shortfall. + +`dispenseFault` shows: the amount paid, the amount dispensed (per denomination, as now), the +txid as QR (as now) **and as text**, the first 12 characters of the payment hash, the time, +and the sentence *"You have paid. The operator has been notified and holds your transaction +record. Keep this reference."* The raw error code is **not** shown to the customer; it is in +the report. The 30-second auto-return is extended to 120 s and the screen offers "I've saved +this" rather than only "Return to Start." This is lamassu's `fiatTransactionError` prompt +without the receipt camera. + +### 5. Terminal dispenser faults latch cash-out off. A recount or an explicit operator op clears it. + +The HAL classifies each error as `terminal` or `recoverable` (#27's split: jam, motor stop, +diverter and sensor faults, dispense timeout are terminal; pickup error and bill-end are +recoverable). The machine persists `cashOutHeld: { reason, errorCode, since }` in `meta` +when a terminal fault lands, and `SELECT_CASH_OUT` is guarded on it. The idle screen shows +cash-out unavailable with the reason. + +The hold is **not** cleared by re-initialising the dispenser. `dispenseCash` already re-inits +on the next attempt, and re-initialising does not move a note that is stuck. It is cleared +by: + +- a `recount` operator op on any bay — the same "operator opened the machine" gesture that + clears `countsUncertainSince`, so one physical act resolves both; or +- a new `resume_cash_out` operator op, idempotent-id'd like the cassette ops, for the case + where the operator cleared the jam without touching a bay count. + +The machine mirrors the hold into its cassettes-state document (`cash_out_held_since`, +`cash_out_held_reason`) and spirekeeper writes it onto `dca_machines` beside +`counts_uncertain_since`. The availability beacon reports `cash_out: false` while held. + +Separately, the operator gets a manual switch: `cash_out_enabled` on `dca_machines`, +published as an operator op, defaulting true. This is lamassu's `cashOutConfig.active` +shape. Cash-out is offered only when the machine is not held **and** the switch is on. +`dca_machines.is_active` is **not** used for either — it is the roster filter in +`get_machine_by_wallet`, and flipping it makes the machine unknown to the RPC handlers +rather than pausing it. + +### 6. Owed cash is a first-class state on both sides, with an off-machine settle path. + +Server: `cash_owed` and `dispense_unreported` are two new buckets on +`StuckSettlementsResponse`. They are the only buckets whose meaning is *a customer is owed +money*, and they render first. Arrival in either bucket triggers the operator notification +path (whatever `notifyOperator` equivalent spirekeeper grows; at minimum the dashboard banner +— but the push is the point, and it belongs in the same transaction that writes the row). + +Resolution closes **both** ledgers: + +- **On-machine remediation.** `manual_dispense` with `ref_txid` already flips the machine row + to `remediated` via `remediateTransaction`. The machine sends a `report_dispense` for the + remediation with `remediates_txid`, and the server moves the settlement from `cash_owed` to + `pending` and distributes. +- **Off-machine settlement.** The operator paid the customer by hand. A new `settle_cash_owed` + operator action records provenance (free text, author, time) on the settlement, moves it to + `pending`, and publishes a `settle_transaction { txid, note }` operator op; the machine + applies it by setting `remediated_by` to the note and `status = 'remediated'`. Today there is + no way to record this at all, and the machine's ledger asserts the debt forever. + +`PartialDispenseData` is pre-filled from the report's `bills` so the operator confirms a +number the hardware produced rather than typing one. + +### 7. Raw codes get a decode table in the driver, built empirically. + +`packages/hal` owns a `rawCode → { errorCode, errorClass, human }` table per dispenser. It is +seeded with what has been observed — `78 42` on an F56 is a note stopped at the cassette +exit (sintra, 2026-10-09) — and grows as codes occur; an unknown code reports as +`F56DispenseError` / `terminal` / `"unrecognised dispenser error "`, failing safe. No +code is borrowed from another layer's vocabulary (the lamassu 570 lesson). + +## Consequences + +- One cash-out now produces one `report_dispense`; the server's `dispense_reports` table is + the audit trail, and `atm-reconcile`'s natural sibling is a settlement↔transaction + reconciliation that joins on `txid` and flags any settlement without a report. +- Distribution for `cash_out` is delayed by the dispense (seconds). Operators watching the + dashboard will see `awaiting_dispense` briefly on every sale. +- `apply_partial_dispense_and_redistribute`'s hard guard still exists but is reached only for + settlements that pre-date this ADR; its message should say which leg type blocked it. +- The state machine gains `dispenseFault`, `outOfCash` and a `cashOutHeld` guard; tests in + `packages/state-machine` cover all three (cash-out is the critical path). +- A machine on an old build keeps working: the server treats a `cash_out` settlement with no + report after `DISPENSE_REPORT_TTL` as `dispense_unreported`, not as failed, and the operator + can capture manually. That is also the upgrade path. + +## Rollout + +1. **bitspire:** `DispenseCashResult` gains the new fields; HAL computes `dispenseConfirmed`, + classifies, decodes; `dispensingCash` guards on it; `dispenseFault` / `outOfCash` screens; + `dispense_reports` outbox + `report_dispense` client; `cashOutHeld` latch and guard; + `counts_uncertain` on zero-with-error. Ships first — with no server handler the RPC + returns an error and the outbox simply retries, so the machine is never blocked on it. +2. **spirekeeper:** `report_dispense` handler + `dispense_reports` table; `awaiting_dispense` / + `cash_owed` / `partial_pending` / `dispense_unreported` statuses; gate in `_handle_payment`; + worklist buckets + notification; `settle_cash_owed`; `cash_out_enabled` and the two new + operator ops; `dca_machines.cash_out_held_*`. +3. **Both:** `resume_cash_out` / `settle_transaction` ops on the machine; beacon reflects + held; docs: `docs/nostr-patterns` entry for the outbox-RPC pattern, this ADR → Accepted. + +## Review findings outside this ADR + +Found while tracing the money path end to end for this document. Not decided here; each is +a candidate issue. + +1. **The settlement push did not deliver — `[ATM Service] Invoice paid (poll)!`** The + `subscribe_payments` stream missed the payment and the polling fallback caught it. Worth + knowing before relying on push latency anywhere; related to #78. +2. **`waitingForCashTaken` auto-advances to `complete` after 30 s "assume taken."** A + transaction can be recorded complete with notes still in the slot. The HAL already blocks + in `waitForBillsRemoved` inside `dispenseCash`, so the state is doing a second, weaker + version of the same job. +3. **The machine and server keep two ledgers with no reconciliation.** `transactions` and + `dca_settlements` join on `txid` today and nothing ever joins them. Decision 2 gives the + server everything a reconciliation needs. +4. **The availability beacon has no consumers and overlaps the cassettes-state document.** + Two machine-authored state documents with overlapping fields is drift waiting to happen. + Either fold availability into the cassettes-state doc or give the beacon a reader. +5. **`is_active` reads as a service gate and is a roster flag.** Rename to something like + `enrolled`, or document at the column. +6. **`generateInvoice` carries a comment deferring `bills`/`cassettes` onto the invoice + `extra`.** With Decision 2 that would be the wrong place — provisioned is not dispensed — + and the comment invites a future contributor to wire it there. Remove it. +7. **The HAL's `dispensed` boolean is count-based (`totalRequested === totalDispensed`), + not value-based.** Equivalent only while each bay dispenses its own denomination. + Decision 3 replaces it. +8. **`_handle_payment` processes cash-in and cash-out through one path and only `tx_type` + tells them apart.** Decision 1 adds a second direction-specific branch. If a third + arrives, split the handler. +9. **The partial-dispense guard's message is wrong for internal legs.** "Lightning payments + can't be clawed back" is true of `autoforward` and false of the LNbits-internal legs, + which are compensatable. Make the guard leg-aware or correct the message. +10. **Review scope.** This document traced the cash-out path: state machine → HAL → ledger → + transport → settlement → distribution → dashboard. The cash-in path shares the settlement + pipeline and has its own money-at-risk shape in `create_withdraw` (server-side amounts, + `max_cash_in_sats`); it has not been reviewed to the same depth and is the obvious next + slice. The HAL drivers, access layer and deploy module were not in scope.