docs(adr): ADR-005 — cash-out dispense outcome and settlement capture

A customer paid a 40 EUR cash-out, a note jammed at the cassette exit,
and the dashboard showed `processed`. The machine had recorded the
failure correctly. Nothing it knew ever left the box.

The root is ordering, not display: spirekeeper spawns process_settlement
the instant the payment lands, which is before the machine has begun to
dispense. The legs are paid sub-second; the dispense fails afterwards;
and the one remediation tool refuses once any leg has completed. It is
unreachable for the exact case it was built for.

ADR-005 makes payment the authorization and dispense confirmation the
capture — distribution waits for the machine's report. The report is a
report_dispense RPC (not the state doc: ADR-004's losing-writer problem),
carried through a durable outbox, sent on success and failure, adopting
lamassu's dispense_confirmed / error / error_code taxonomy and its
per-bay action log — which the machine already records in cassette_bills
and simply never ships.

Deviates from lamassu in three places it got wrong or never did: a
zero-dispensed report that arrives with an error is treated as
unverified, not as zero; a mechanical fault is its own customer screen
with evidence and is not "out of cash"; and terminal dispenser faults
latch cash-out off until a recount or an explicit operator op, because
re-initialising does not move a stuck note.

Closes the review loop with ten findings outside the ADR's decisions.

Refs #122, #27, #78
This commit is contained in:
Padreug 2026-10-09 16:34:54 +02:00
commit 483e44aaa9

View file

@ -0,0 +1,315 @@
# ADR-005: Cash-Out Dispense Outcome and Settlement Capture
**Status:** Proposed
**Date:** 2026-10-09
**Context:** On 2026-10-09 a customer paid a 40 EUR cash-out on sintra, a note jammed at the
cassette exit, and the operator dashboard showed the settlement as `processed` with no sign
anything was wrong (aiolabs/bitspire#122). The machine had recorded the failure correctly and
in detail. Nothing it knew ever reached anyone. This ADR specifies how a dispense outcome
becomes a first-class fact on both sides of the wire, and fixes the structural reason the
existing remediation tool could not have helped.
## The problem
A cash-out moves value in two steps that today are not connected:
1. **Payment.** The customer pays the ATM's BOLT11 invoice. LNbits lands it in the machine
wallet, `spirekeeper._handle_payment` verifies attribution, inserts a `dca_settlements` row,
and — in the same breath — spawns `process_settlement` as a background task.
2. **Dispense.** The machine, which learns of the payment through `watchInvoice`, commands the
dispenser. The hardware reports per-bay `dispensed` / `rejected` counts and, on failure, an
error code.
Step 1 does not wait for step 2. `process_settlement` pays the super fee, the operator's
commission splits, and the DCA legs the moment the payment lands, which on an LNbits-internal
transfer is sub-second. The dispense begins afterwards. So by the time the F56 reported
`78 42` at T+2 s, the settlement's legs were already `completed` and its status was
`processed` — which is what the dashboard faithfully displayed. `processed` means *all
distribution legs paid*. It has never meant *cash reached a hand*, because the server has no
input that could tell it.
The tool built for this situation, `apply_partial_dispense_and_redistribute`, carries a hard
guard: it refuses once any leg has completed, because a Lightning payment cannot be clawed
back. Under the current ordering that guard is reached on every real failure. The remediation
is structurally unreachable for the exact case it was written for, except by winning a race
against a sub-second transfer.
Downstream of that, four smaller gaps compound it (all in #122):
- The machine's `dispenseError` state is terminal and local. No report, no notification. The
only record that a customer is owed money lives in `state.db` on the ATM.
- The dispenser's counters do not see a note that leaves the bay and stops in the transport —
it is neither `dispensed` nor `rejected`. The machine trusts the resulting `dispensed: 0`,
leaves the bay count untouched, and republishes it as fact. The `countsUncertainSince`
safety net fires only when the report is *absent*, not when it is present and wrong.
- Nothing reads dispenser health. The availability beacon derives `cash_out` from
`totalBills > 0` alone and kept advertising a jammed machine as available. (Nothing
consumes that beacon today, so it could not have been the enforcement point in any case.)
- The customer sees a 30-second countdown and a txid QR, then the idle screen. There is no
claim reference, no statement that they have paid, and nothing distinguishes a mechanical
fault from an out-of-cash condition.
## Prior art
lamassu-machine / lamassu-server ran this exact hardware in production for a decade. Their
model, which this ADR adopts where it fits and deviates from where it is wrong:
- **Three fields on the transaction:** `error` (human message), `error_code` (the error's
*name*, machine-readable), `dispense_confirmed` (boolean). `dispense_confirmed` is computed on
**value** — `tx.fiat.eq(Σ denomination × dispensed)` — not taken from the driver.
- **An append-only action log, `cash_out_actions`,** one row per dispense attempt with
per-bay `provisioned_N` / `denomination_N` / `dispensed_N` / `rejected_N`, written by
`logDispense` as `action: 'dispense'` or `'dispenseError'` purely on whether `error` is set.
- **The operator is notified in the same atomic block that logs the dispense**
(`notifyOperator`, cash-out-atomic.js). Push, not a worklist.
- **A mechanical fault is not "out of cash."** Their 2026-09-29 fix routes a
dispenser-reported error to the screen that asks the customer to photograph their receipt,
*"the right prompt when they have paid and are owed money,"* and reserves `outOfCash` for a
shortfall with no error. The same commit removed a borrowed `statusCode 570` from the F56
driver because 570 meant "insufficient funds" to the server — a jammed BDU was being
reported as a hot-wallet problem.
- **What they did not have:** any notion of disabling a machine on a dispenser fault.
`getMachineStatuses` is inferential — ping age, stuck-screen age — and `cashOut` is an
operator config toggle. A jammed dispenser on a responsive machine reads *Fully
functional*. That is sintra's beacon exactly, so the latch below is new work, not a port.
- **What they got wrong and we will not copy:** `dispenseOccurred(bills)` returns true if the
bill entries merely *have* `dispensed` and `rejected` keys, and `updateCassettes` then
decrements by those numbers. A jam reporting `dispensed: 0` passes and decrements by zero.
Their cassette counts drift the same way sintra's did.
## Decisions
### 1. Payment is authorization. Dispense confirmation is capture. Distribution waits for capture.
A `cash_out` settlement lands as `pending` exactly as now, but `_handle_payment` no longer
spawns `process_settlement` for it. The row moves to a new status, `awaiting_dispense`, and
stays there until the machine reports.
| Machine reports | Settlement becomes | Then |
| ----------------------------------------- | ------------------ | --------------------------------------- |
| `dispense_confirmed: true` | `pending` | claim + distribute → `processed` |
| partial (some notes out, value short) | `partial_pending` | operator confirms → distribute scaled |
| `dispense_confirmed: false`, nothing out | `cash_owed` | legs never run; funds stay in wallet |
| no report within `DISPENSE_REPORT_TTL` | `dispense_unreported` | worklist; operator investigates |
This is the card-processing shape — authorize, then capture — and it is the same ordering
lamassu-server enforces between `dispense_confirmed` and `updateCassettes`. The cost is that
operator and DCA legs land seconds later than they do today, which is the dispense time.
The benefit is that `apply_partial_dispense_and_redistribute` is always reachable, because
no leg has run yet, and `cash_owed` is a state the money has not left.
`cash_in` settlements are unaffected: there is no dispense to wait for, and
`_pay_dca_distributions` already branches on `tx_type` for exactly this kind of asymmetry.
**Rejected:** keeping immediate distribution and adding a compensating reversal. Internal
legs *are* reversible — they are LNbits-internal invoices, so a compensating internal
payment is mechanically possible and the guard's "Lightning can't be clawed back" is only
true of the `autoforward` leg. But undoing money movement is strictly harder than not
moving it yet, and the autoforward leg stays irreversible either way. Compensation is kept
as a secondary tool for settlements that distributed before this ADR landed.
### 2. The machine reports every cash-out outcome over a `report_dispense` RPC.
Not over the kind-30078 state document. ADR-004 established why: an addressable event gives
its publisher no failure signal, and a losing writer is never told. A per-transaction
outcome is an append-only fact that must be acknowledged, which is a request/reply.
The RPC follows `create_withdraw` and `get_machine_config`: `register_rpc` at
`AUTH_ACCOUNT`, identity taken from the **verified** `sender_pubkey`, never from the body.
It is sent on **success as well as failure** — a success report is what captures (Decision
1). Payload, adopting the lamassu field names:
```jsonc
{
"txid": "tx_mv0madw6_wdhtea1v",
"payment_hash": "6f216df32c36…",
"tx_type": "cash_out",
"dispense_confirmed": false, // value equality, Decision 3
"error": "Dispensing, code: 78 42", // human, null on success
"error_code": "F56DispenseError", // the error's NAME, null on success
"raw_code": "78 42", // driver-native, for the decode table
"error_class": "terminal", // "terminal" | "recoverable" | null, Decision 5
"fiat_cents": 4000,
"bills": [{ "denomination": 20, "requested": 2, "dispensed": 0, "rejected": 0 }],
"cassettes": [ // the machine's cassette_bills rows, verbatim
{ "position": 2, "denomination": 20, "provisioned": 2, "dispensed": 0, "rejected": 0 }
],
"counts_uncertain": true, // Decision 3
"at": 1791529353
}
```
**Delivery is at-least-once with a durable outbox.** The machine writes the report to
`state.db` in the same transaction as the `transactions` row (`dispense_reports`:
`txid PRIMARY KEY, payload, created_at, acked_at`), then sends. It resends on boot, on relay
reconnect, and on a timer until an `OK` reply sets `acked_at`. The server upserts on `txid`,
so a resend is a no-op. This is the cassette-ops idempotency pattern applied to the other
direction.
The server stores every report in an append-only `dispense_reports` table (one row per
attempt, lamassu's `cash_out_actions` shape, keyed to the settlement by `bitspire_txid`,
which is already populated from `extra.txid`), copies the per-bay detail into
`dca_settlements.bills_json` / `cassettes_json` (columns that exist today and are never
written), and sets `dispense_confirmed`, `error`, `error_code` on the settlement.
### 3. `dispense_confirmed` is computed on value, separately from `error`, and a zero report with an error is unverified.
On the machine, after the HAL returns:
```
confirmed = requestedFiatCents === Σ(denomination × dispensed) × 100
```
`error` is carried independently. The state machine's `dispensingCash.onDone` guard moves
from `output.dispensed === true` to `output.dispenseConfirmed`, and `DispenseCashResult`
gains `dispenseConfirmed`, `errorCode`, `rawCode`, `errorClass`.
**The deviation from both bitSpire-today and lamassu:** when a report arrives with `error`
set and `dispensed === 0` on every bay, the machine does *not* treat that zero as a count. A
note in the transport path completes neither counter. The machine sets `countsUncertainSince`
exactly as it already does for an absent report, carries `counts_uncertain: true` in the RPC,
and the bay stays flagged until a `recount` op clears it. `atm-reconcile` then shows the gap
instead of a clean ledger.
### 4. A dispenser fault is its own customer screen, with evidence, and it is not "out of cash."
Two distinct terminal states replace the single `dispenseError`:
- **`outOfCash`** — the request could not be met from inventory and the dispenser reported
**no error**. Nothing was charged beyond what was dispensed.
- **`dispenseFault`** — the dispenser reported an error. The customer **has paid** and is
owed the shortfall.
`dispenseFault` shows: the amount paid, the amount dispensed (per denomination, as now), the
txid as QR (as now) **and as text**, the first 12 characters of the payment hash, the time,
and the sentence *"You have paid. The operator has been notified and holds your transaction
record. Keep this reference."* The raw error code is **not** shown to the customer; it is in
the report. The 30-second auto-return is extended to 120 s and the screen offers "I've saved
this" rather than only "Return to Start." This is lamassu's `fiatTransactionError` prompt
without the receipt camera.
### 5. Terminal dispenser faults latch cash-out off. A recount or an explicit operator op clears it.
The HAL classifies each error as `terminal` or `recoverable` (#27's split: jam, motor stop,
diverter and sensor faults, dispense timeout are terminal; pickup error and bill-end are
recoverable). The machine persists `cashOutHeld: { reason, errorCode, since }` in `meta`
when a terminal fault lands, and `SELECT_CASH_OUT` is guarded on it. The idle screen shows
cash-out unavailable with the reason.
The hold is **not** cleared by re-initialising the dispenser. `dispenseCash` already re-inits
on the next attempt, and re-initialising does not move a note that is stuck. It is cleared
by:
- a `recount` operator op on any bay — the same "operator opened the machine" gesture that
clears `countsUncertainSince`, so one physical act resolves both; or
- a new `resume_cash_out` operator op, idempotent-id'd like the cassette ops, for the case
where the operator cleared the jam without touching a bay count.
The machine mirrors the hold into its cassettes-state document (`cash_out_held_since`,
`cash_out_held_reason`) and spirekeeper writes it onto `dca_machines` beside
`counts_uncertain_since`. The availability beacon reports `cash_out: false` while held.
Separately, the operator gets a manual switch: `cash_out_enabled` on `dca_machines`,
published as an operator op, defaulting true. This is lamassu's `cashOutConfig.active`
shape. Cash-out is offered only when the machine is not held **and** the switch is on.
`dca_machines.is_active` is **not** used for either — it is the roster filter in
`get_machine_by_wallet`, and flipping it makes the machine unknown to the RPC handlers
rather than pausing it.
### 6. Owed cash is a first-class state on both sides, with an off-machine settle path.
Server: `cash_owed` and `dispense_unreported` are two new buckets on
`StuckSettlementsResponse`. They are the only buckets whose meaning is *a customer is owed
money*, and they render first. Arrival in either bucket triggers the operator notification
path (whatever `notifyOperator` equivalent spirekeeper grows; at minimum the dashboard banner
— but the push is the point, and it belongs in the same transaction that writes the row).
Resolution closes **both** ledgers:
- **On-machine remediation.** `manual_dispense` with `ref_txid` already flips the machine row
to `remediated` via `remediateTransaction`. The machine sends a `report_dispense` for the
remediation with `remediates_txid`, and the server moves the settlement from `cash_owed` to
`pending` and distributes.
- **Off-machine settlement.** The operator paid the customer by hand. A new `settle_cash_owed`
operator action records provenance (free text, author, time) on the settlement, moves it to
`pending`, and publishes a `settle_transaction { txid, note }` operator op; the machine
applies it by setting `remediated_by` to the note and `status = 'remediated'`. Today there is
no way to record this at all, and the machine's ledger asserts the debt forever.
`PartialDispenseData` is pre-filled from the report's `bills` so the operator confirms a
number the hardware produced rather than typing one.
### 7. Raw codes get a decode table in the driver, built empirically.
`packages/hal` owns a `rawCode → { errorCode, errorClass, human }` table per dispenser. It is
seeded with what has been observed — `78 42` on an F56 is a note stopped at the cassette
exit (sintra, 2026-10-09) — and grows as codes occur; an unknown code reports as
`F56DispenseError` / `terminal` / `"unrecognised dispenser error <code>"`, failing safe. No
code is borrowed from another layer's vocabulary (the lamassu 570 lesson).
## Consequences
- One cash-out now produces one `report_dispense`; the server's `dispense_reports` table is
the audit trail, and `atm-reconcile`'s natural sibling is a settlement↔transaction
reconciliation that joins on `txid` and flags any settlement without a report.
- Distribution for `cash_out` is delayed by the dispense (seconds). Operators watching the
dashboard will see `awaiting_dispense` briefly on every sale.
- `apply_partial_dispense_and_redistribute`'s hard guard still exists but is reached only for
settlements that pre-date this ADR; its message should say which leg type blocked it.
- The state machine gains `dispenseFault`, `outOfCash` and a `cashOutHeld` guard; tests in
`packages/state-machine` cover all three (cash-out is the critical path).
- A machine on an old build keeps working: the server treats a `cash_out` settlement with no
report after `DISPENSE_REPORT_TTL` as `dispense_unreported`, not as failed, and the operator
can capture manually. That is also the upgrade path.
## Rollout
1. **bitspire:** `DispenseCashResult` gains the new fields; HAL computes `dispenseConfirmed`,
classifies, decodes; `dispensingCash` guards on it; `dispenseFault` / `outOfCash` screens;
`dispense_reports` outbox + `report_dispense` client; `cashOutHeld` latch and guard;
`counts_uncertain` on zero-with-error. Ships first — with no server handler the RPC
returns an error and the outbox simply retries, so the machine is never blocked on it.
2. **spirekeeper:** `report_dispense` handler + `dispense_reports` table; `awaiting_dispense` /
`cash_owed` / `partial_pending` / `dispense_unreported` statuses; gate in `_handle_payment`;
worklist buckets + notification; `settle_cash_owed`; `cash_out_enabled` and the two new
operator ops; `dca_machines.cash_out_held_*`.
3. **Both:** `resume_cash_out` / `settle_transaction` ops on the machine; beacon reflects
held; docs: `docs/nostr-patterns` entry for the outbox-RPC pattern, this ADR → Accepted.
## Review findings outside this ADR
Found while tracing the money path end to end for this document. Not decided here; each is
a candidate issue.
1. **The settlement push did not deliver — `[ATM Service] Invoice paid (poll)!`** The
`subscribe_payments` stream missed the payment and the polling fallback caught it. Worth
knowing before relying on push latency anywhere; related to #78.
2. **`waitingForCashTaken` auto-advances to `complete` after 30 s "assume taken."** A
transaction can be recorded complete with notes still in the slot. The HAL already blocks
in `waitForBillsRemoved` inside `dispenseCash`, so the state is doing a second, weaker
version of the same job.
3. **The machine and server keep two ledgers with no reconciliation.** `transactions` and
`dca_settlements` join on `txid` today and nothing ever joins them. Decision 2 gives the
server everything a reconciliation needs.
4. **The availability beacon has no consumers and overlaps the cassettes-state document.**
Two machine-authored state documents with overlapping fields is drift waiting to happen.
Either fold availability into the cassettes-state doc or give the beacon a reader.
5. **`is_active` reads as a service gate and is a roster flag.** Rename to something like
`enrolled`, or document at the column.
6. **`generateInvoice` carries a comment deferring `bills`/`cassettes` onto the invoice
`extra`.** With Decision 2 that would be the wrong place — provisioned is not dispensed —
and the comment invites a future contributor to wire it there. Remove it.
7. **The HAL's `dispensed` boolean is count-based (`totalRequested === totalDispensed`),
not value-based.** Equivalent only while each bay dispenses its own denomination.
Decision 3 replaces it.
8. **`_handle_payment` processes cash-in and cash-out through one path and only `tx_type`
tells them apart.** Decision 1 adds a second direction-specific branch. If a third
arrives, split the handler.
9. **The partial-dispense guard's message is wrong for internal legs.** "Lightning payments
can't be clawed back" is true of `autoforward` and false of the LNbits-internal legs,
which are compensatable. Make the guard leg-aware or correct the message.
10. **Review scope.** This document traced the cash-out path: state machine → HAL → ledger →
transport → settlement → distribution → dashboard. The cash-in path shares the settlement
pipeline and has its own money-at-risk shape in `create_withdraw` (server-side amounts,
`max_cash_in_sats`); it has not been reviewed to the same depth and is the obvious next
slice. The HAL drivers, access layer and deploy module were not in scope.