521 lines
34 KiB
Markdown
521 lines
34 KiB
Markdown
# ADR-005: Cash-Out Dispense Outcome and Settlement Capture
|
||
|
||
**Status:** Accepted (2026-10-10)
|
||
**Date:** 2026-10-09
|
||
**Context:** On 2026-10-09 a customer paid a 40 EUR cash-out on sintra, a note jammed at the
|
||
cassette exit, and the operator dashboard showed the settlement as `processed` with no sign
|
||
anything was wrong (aiolabs/bitspire#122). The machine had recorded the failure correctly and
|
||
in detail. Nothing it knew ever reached anyone. This ADR specifies how a dispense outcome
|
||
becomes a first-class fact on both sides of the wire, and fixes the structural reason the
|
||
existing remediation tool could not have helped.
|
||
|
||
## The problem
|
||
|
||
A cash-out moves value in two steps that today are not connected:
|
||
|
||
1. **Payment.** The customer pays the ATM's BOLT11 invoice. LNbits lands it in the machine
|
||
wallet, `spirekeeper._handle_payment` verifies attribution, inserts a `dca_settlements` row,
|
||
and — in the same breath — spawns `process_settlement` as a background task.
|
||
2. **Dispense.** The machine, which learns of the payment through `watchInvoice`, commands the
|
||
dispenser. The hardware reports per-bay `dispensed` / `rejected` counts and, on failure, an
|
||
error code.
|
||
|
||
Step 1 does not wait for step 2. `process_settlement` pays the super fee, the operator's
|
||
commission splits, and the DCA legs the moment the payment lands, which on an LNbits-internal
|
||
transfer is sub-second. The dispense begins afterwards. So by the time the F56 reported
|
||
`78 42` at T+2 s, the settlement's legs were already `completed` and its status was
|
||
`processed` — which is what the dashboard faithfully displayed. `processed` means *all
|
||
distribution legs paid*. It has never meant *cash reached a hand*, because the server has no
|
||
input that could tell it.
|
||
|
||
The tool built for this situation, `apply_partial_dispense_and_redistribute`, carries a hard
|
||
guard: it refuses once any leg has completed, because a Lightning payment cannot be clawed
|
||
back. Under the current ordering that guard is reached on every real failure. The remediation
|
||
is structurally unreachable for the exact case it was written for, except by winning a race
|
||
against a sub-second transfer.
|
||
|
||
Downstream of that, four smaller gaps compound it (all in #122):
|
||
|
||
- The machine's `dispenseError` state is terminal and local. No report, no notification. The
|
||
only record that a customer is owed money lives in `state.db` on the ATM.
|
||
- The dispenser's counters do not see a note that leaves the bay and stops in the transport —
|
||
it is neither `dispensed` nor `rejected`. The machine trusts the resulting `dispensed: 0`,
|
||
leaves the bay count untouched, and republishes it as fact. The `countsUncertainSince`
|
||
safety net fires only when the report is *absent*, not when it is present and wrong.
|
||
- Nothing reads dispenser health. The availability beacon derives `cash_out` from
|
||
`totalBills > 0` alone and kept advertising a jammed machine as available. (Nothing
|
||
consumes that beacon today, so it could not have been the enforcement point in any case.)
|
||
- The customer sees a 30-second countdown and a txid QR, then the idle screen. There is no
|
||
claim reference, no statement that they have paid, and nothing distinguishes a mechanical
|
||
fault from an out-of-cash condition.
|
||
|
||
## Prior art
|
||
|
||
lamassu-machine / lamassu-server ran this exact hardware in production for a decade. Their
|
||
model, which this ADR adopts where it fits and deviates from where it is wrong:
|
||
|
||
- **Three fields on the transaction:** `error` (human message), `error_code` (the error's
|
||
*name*, machine-readable), `dispense_confirmed` (boolean). `dispense_confirmed` is computed on
|
||
**value** — `tx.fiat.eq(Σ denomination × dispensed)` — not taken from the driver.
|
||
- **An append-only action log, `cash_out_actions`,** one row per dispense attempt with
|
||
per-bay `provisioned_N` / `denomination_N` / `dispensed_N` / `rejected_N`, written by
|
||
`logDispense` as `action: 'dispense'` or `'dispenseError'` purely on whether `error` is set.
|
||
- **The operator is notified in the same atomic block that logs the dispense**
|
||
(`notifyOperator`, cash-out-atomic.js). Push, not a worklist.
|
||
- **A mechanical fault is not "out of cash."** Their 2026-09-29 fix routes a
|
||
dispenser-reported error to the screen that asks the customer to photograph their receipt,
|
||
*"the right prompt when they have paid and are owed money,"* and reserves `outOfCash` for a
|
||
shortfall with no error. The same commit removed a borrowed `statusCode 570` from the F56
|
||
driver because 570 meant "insufficient funds" to the server — a jammed BDU was being
|
||
reported as a hot-wallet problem.
|
||
- **What they did not have:** any notion of disabling a machine on a dispenser fault.
|
||
`getMachineStatuses` is inferential — ping age, stuck-screen age — and `cashOut` is an
|
||
operator config toggle. A jammed dispenser on a responsive machine reads *Fully
|
||
functional*. That is sintra's beacon exactly, so the latch below is new work, not a port.
|
||
- **What they got wrong and we will not copy:** `dispenseOccurred(bills)` returns true if the
|
||
bill entries merely *have* `dispensed` and `rejected` keys, and `updateCassettes` then
|
||
decrements by those numbers. A jam reporting `dispensed: 0` passes and decrements by zero.
|
||
Their cassette counts drift the same way sintra's did.
|
||
|
||
## Decisions
|
||
|
||
### 1. Payment is authorization. Dispense confirmation is capture. Distribution waits for capture.
|
||
|
||
A `cash_out` settlement lands as `pending` exactly as now, but `_handle_payment` no longer
|
||
spawns `process_settlement` for it. The row moves to a new status, `awaiting_dispense`, and
|
||
stays there until the machine reports.
|
||
|
||
| Machine reports | Settlement becomes | Then |
|
||
| ----------------------------------------- | ------------------ | --------------------------------------- |
|
||
| `dispense_confirmed: true` | `pending` | claim + distribute → `processed` |
|
||
| partial (some notes out, value short) | `partial_pending` | nothing moves until the shortfall is resolved (below) |
|
||
| `dispense_confirmed: false`, nothing out | `cash_owed` | legs never run; funds stay in wallet |
|
||
| no report within `DISPENSE_REPORT_TTL` | `dispense_unreported` | worklist; operator investigates |
|
||
|
||
**A partial dispense distributes once, when the outcome is final.** Some notes reached the
|
||
customer and some did not, so the sale's true amount is not yet known: it is the full amount
|
||
if the shortfall is remediated (an on-machine `manual_dispense` against the `txid`, or an
|
||
off-machine payout recorded with `settle_cash_owed`), and the scaled amount if the shortfall
|
||
is vouchered or written off. `partial_pending` therefore holds *everything* — including the
|
||
operator's and LPs' share of the notes that did dispense — until the operator records which
|
||
of those it was. Then one distribution runs, at that amount, using the existing
|
||
`apply_partial_dispense_and_redistribute` arithmetic for the scaled case (linear scale; the
|
||
fee split by the ratio locked at landing; operator absorbs rounding). A vouchered remainder
|
||
stays undistributed until redemption or expiry, per the voucher rules under *Future
|
||
directions*.
|
||
|
||
The alternative — distribute the scaled part immediately and the remainder on resolution —
|
||
pays the LPs the same afternoon but needs a second, additive distribution pass keyed to the
|
||
same settlement, which the repo does not have and which the completed-legs guard in the
|
||
current tool would fight. It is deferred, not rejected: a partial is almost always a
|
||
terminal-class fault that has also latched cash-out off, so the operator is coming to the
|
||
machine anyway and resolution is hours, not weeks. Revisit if prompt LP payout ever matters
|
||
more than one-distribution-per-settlement. (Decided 2026-10-10.)
|
||
|
||
This is the card-processing shape — authorize, then capture — and it is the same ordering
|
||
lamassu-server enforces between `dispense_confirmed` and `updateCassettes`. The cost is that
|
||
operator and DCA legs land seconds later than they do today, which is the dispense time.
|
||
The benefit is that `apply_partial_dispense_and_redistribute` is always reachable, because
|
||
no leg has run yet, and `cash_owed` is a state the money has not left.
|
||
|
||
`cash_in` settlements are unaffected: there is no dispense to wait for, and
|
||
`_pay_dca_distributions` already branches on `tx_type` for exactly this kind of asymmetry.
|
||
|
||
**Rejected:** keeping immediate distribution and adding a compensating reversal. Internal
|
||
legs *are* reversible — they are LNbits-internal invoices, so a compensating internal
|
||
payment is mechanically possible and the guard's "Lightning can't be clawed back" is only
|
||
true of the `autoforward` leg. But undoing money movement is strictly harder than not
|
||
moving it yet, and the autoforward leg stays irreversible either way. Compensation is kept
|
||
as a secondary tool for settlements that distributed before this ADR landed.
|
||
|
||
### 2. The machine reports every cash-out outcome over a `report_dispense` RPC.
|
||
|
||
Not over the kind-30078 state document. ADR-004 established why: an addressable event gives
|
||
its publisher no failure signal, and a losing writer is never told. A per-transaction
|
||
outcome is an append-only fact that must be acknowledged, which is a request/reply.
|
||
|
||
The RPC follows `create_withdraw` and `get_machine_config`: `register_rpc` at
|
||
`AUTH_ACCOUNT`, identity taken from the **verified** `sender_pubkey`, never from the body.
|
||
It is sent on **success as well as failure** — a success report is what captures (Decision
|
||
1). Payload, adopting the lamassu field names:
|
||
|
||
```jsonc
|
||
{
|
||
"txid": "tx_mv0madw6_wdhtea1v",
|
||
"payment_hash": "6f216df32c36…",
|
||
"tx_type": "cash_out",
|
||
"dispense_confirmed": false, // value equality, Decision 3
|
||
"error": "Dispensing, code: 78 42", // human, null on success
|
||
"error_code": "F56DispenseError", // the error's NAME, null on success
|
||
"raw_code": "78 42", // driver-native, for the decode table
|
||
"error_class": "terminal", // "terminal" | "recoverable" | null, Decision 5
|
||
"fiat_cents": 4000,
|
||
"bills": [{ "denomination": 20, "requested": 2, "dispensed": 0, "rejected": 0 }],
|
||
"cassettes": [ // the machine's cassette_bills rows, verbatim
|
||
{ "position": 2, "denomination": 20, "provisioned": 2, "dispensed": 0, "rejected": 0 }
|
||
],
|
||
"counts_uncertain": true, // Decision 3
|
||
"at": 1791529353
|
||
}
|
||
```
|
||
|
||
**Delivery is at-least-once with a durable outbox.** The machine writes the report to
|
||
`state.db` in the same transaction as the `transactions` row (`dispense_reports`:
|
||
`txid PRIMARY KEY, payload, created_at, acked_at`), then sends. It resends on boot, on relay
|
||
reconnect, and on a timer until an `OK` reply sets `acked_at`. The server upserts on `txid`,
|
||
so a resend is a no-op. This is the cassette-ops idempotency pattern applied to the other
|
||
direction.
|
||
|
||
The server stores every report in an append-only `dispense_reports` table (one row per
|
||
attempt, lamassu's `cash_out_actions` shape, keyed to the settlement by `bitspire_txid`,
|
||
which is already populated from `extra.txid`), copies the per-bay detail into
|
||
`dca_settlements.bills_json` / `cassettes_json` (columns that exist today and are never
|
||
written), and sets `dispense_confirmed`, `error`, `error_code` on the settlement.
|
||
|
||
### 3. `dispense_confirmed` is computed on value, separately from `error`, and a zero report with an error is unverified.
|
||
|
||
On the machine, after the HAL returns:
|
||
|
||
```
|
||
confirmed = requestedFiatCents === Σ(denomination × dispensed) × 100
|
||
```
|
||
|
||
`error` is carried independently. The state machine's `dispensingCash.onDone` guard moves
|
||
from `output.dispensed === true` to `output.dispenseConfirmed`, and `DispenseCashResult`
|
||
gains `dispenseConfirmed`, `errorCode`, `rawCode`, `errorClass`.
|
||
|
||
**The deviation from both bitSpire-today and lamassu:** when a report arrives with `error`
|
||
set and `dispensed === 0` on every bay, the machine does *not* treat that zero as a count. A
|
||
note in the transport path completes neither counter. The machine sets `countsUncertainSince`
|
||
exactly as it already does for an absent report, carries `counts_uncertain: true` in the RPC,
|
||
and the bay stays flagged until a `recount` op clears it. `atm-reconcile` then shows the gap
|
||
instead of a clean ledger.
|
||
|
||
### 4. A dispenser fault is its own customer screen, with evidence, and it is not "out of cash."
|
||
|
||
Two distinct terminal states replace the single `dispenseError`:
|
||
|
||
- **`outOfCash`** — the request could not be met and the dispenser reported **no error**
|
||
(an inventory refusal, or simply short). In the cash-out flow this state is reached *after*
|
||
payment, so the customer **has paid** and is owed the shortfall exactly as below; the
|
||
difference is the cause — no hardware fault, so the machine stays in service and nothing
|
||
latches. (An earlier draft said "nothing was charged beyond what was dispensed"; that is
|
||
only true of the inventory check *before* payment, which already prevents the sale.)
|
||
- **`dispenseFault`** — the dispenser reported an error. The customer **has paid** and is
|
||
owed the shortfall, and a `terminal` class also latches cash-out off (Decision 5).
|
||
|
||
Both screens therefore show the same evidence; the heading and the latch differ.
|
||
|
||
`dispenseFault` shows: the amount paid, the amount dispensed (per denomination, as now), the
|
||
txid as QR (as now) **and as text**, the first 12 characters of the payment hash, the time,
|
||
and the sentence *"You have paid. The operator has been notified and holds your transaction
|
||
record. Keep this reference."* The raw error code is **not** shown to the customer; it is in
|
||
the report. The 30-second auto-return is extended to 120 s and the screen offers "I've saved
|
||
this" rather than only "Return to Start." This is lamassu's `fiatTransactionError` prompt
|
||
without the receipt camera.
|
||
|
||
### 5. Terminal dispenser faults latch cash-out off. A recount or an explicit operator op clears it.
|
||
|
||
The HAL classifies each error as `terminal` or `recoverable` (#27's split: jam, motor stop,
|
||
diverter and sensor faults, dispense timeout are terminal; pickup error and bill-end are
|
||
recoverable). The machine persists `cashOutHeld: { reason, errorCode, since }` in `meta`
|
||
when a terminal fault lands, and `SELECT_CASH_OUT` is guarded on it. The idle screen shows
|
||
cash-out unavailable with the reason.
|
||
|
||
The hold is **not** cleared by re-initialising the dispenser. `dispenseCash` already re-inits
|
||
on the next attempt, and re-initialising does not move a note that is stuck. It is cleared
|
||
by:
|
||
|
||
- a `recount` operator op on any bay — the same "operator opened the machine" gesture that
|
||
clears `countsUncertainSince`, so one physical act resolves both; or
|
||
- a new `resume_cash_out` operator op, idempotent-id'd like the cassette ops, for the case
|
||
where the operator cleared the jam without touching a bay count.
|
||
|
||
The machine mirrors the hold into its cassettes-state document (`cash_out_held_since`,
|
||
`cash_out_held_reason`) and spirekeeper writes it onto `dca_machines` beside
|
||
`counts_uncertain_since`. The availability beacon reports `cash_out: false` while held.
|
||
|
||
Separately, the operator gets a manual switch: `cash_out_enabled` on `dca_machines`,
|
||
published as an operator op, defaulting true. This is lamassu's `cashOutConfig.active`
|
||
shape. Cash-out is offered only when the machine is not held **and** the switch is on.
|
||
`dca_machines.is_active` is **not** used for either — it is the roster filter in
|
||
`get_machine_by_wallet`, and flipping it makes the machine unknown to the RPC handlers
|
||
rather than pausing it.
|
||
|
||
### 6. Owed cash is a first-class state on both sides, with an off-machine settle path.
|
||
|
||
Server: `cash_owed` and `dispense_unreported` are two new buckets on
|
||
`StuckSettlementsResponse`. They are the only buckets whose meaning is *a customer is owed
|
||
money*, and they render first. Arrival in `cash_owed` or `partial_pending` sends the operator
|
||
a **NIP-17 gift-wrapped DM** (kind 14 → 13 → 1059) to their own LNbits-account pubkey, or to
|
||
`super_config.alerts_pubkey` when set — a note to self any NIP-46 client renders. It is signed
|
||
through the operator's signer (no key at rest), is best-effort (a failed publish is logged and
|
||
the report is still acked — the worklist is the durable record), and sets
|
||
`operator_notified_at` so a report resend never re-alerts. Not email, not NIP-04.
|
||
*(Implemented: spirekeeper `notify.py`, slice 2.)*
|
||
|
||
Resolution closes **both** ledgers:
|
||
|
||
- **On-machine remediation.** `manual_dispense` with `ref_txid` already flips the machine row
|
||
to `remediated` via `remediateTransaction`. The machine sends a `report_dispense` for the
|
||
remediation with `remediates_txid`, and the server moves the settlement from `cash_owed` to
|
||
`pending` and distributes.
|
||
- **Off-machine settlement.** The operator paid the customer by hand.
|
||
`POST /settlements/{id}/settle-cash-owed` records provenance (free text, author, time) on the
|
||
settlement, moves it to `pending` and distributes **at the full amount** (the customer is
|
||
whole), and publishes a machine-wide `settle_transaction { id, at, txid, note }` operator op;
|
||
the machine applies it through `remediateTransaction(txid, "settled-off-machine:<op id>:<note>")`,
|
||
which only touches rows still in an error state, so re-delivery is harmless. Before slice 2
|
||
there was no way to record this at all, and the machine's ledger asserted the debt forever.
|
||
|
||
`PartialDispenseData` is pre-filled from the report's `bills` so the operator confirms a
|
||
number the hardware produced rather than typing one.
|
||
|
||
### 7. Raw codes get a decode table in the driver, built empirically.
|
||
|
||
`packages/hal` owns a `rawCode → { errorCode, errorClass, human }` table per dispenser. It is
|
||
seeded with what has been observed — `78 42` on an F56 is a note stopped at the cassette
|
||
exit (sintra, 2026-10-09) — and grows as codes occur; an unknown code reports as
|
||
`F56DispenseError` / `terminal` / `"unrecognised dispenser error <code>"`, failing safe. No
|
||
code is borrowed from another layer's vocabulary (the lamassu 570 lesson).
|
||
|
||
## Consequences
|
||
|
||
- One cash-out now produces one `report_dispense`; the server's `dispense_reports` table is
|
||
the audit trail, and `atm-reconcile`'s natural sibling is a settlement↔transaction
|
||
reconciliation that joins on `txid` and flags any settlement without a report.
|
||
- Distribution for `cash_out` is delayed by the dispense (seconds). Operators watching the
|
||
dashboard will see `awaiting_dispense` briefly on every sale.
|
||
- `apply_partial_dispense_and_redistribute`'s hard guard still exists but is reached only for
|
||
settlements that pre-date this ADR; its message should say which leg type blocked it.
|
||
- The state machine gains `dispenseFault`, `outOfCash` and a `cashOutHeld` guard; tests in
|
||
`packages/state-machine` cover all three (cash-out is the critical path).
|
||
- A machine on an old build keeps working: the server treats a `cash_out` settlement with no
|
||
report after `DISPENSE_REPORT_TTL` as `dispense_unreported`, not as failed, and the operator
|
||
can capture manually. That is also the upgrade path.
|
||
|
||
## Rollout
|
||
|
||
1. **bitspire:** `DispenseCashResult` gains the new fields; HAL computes `dispenseConfirmed`,
|
||
classifies, decodes; `dispensingCash` guards on it; `dispenseFault` / `outOfCash` screens;
|
||
`dispense_reports` outbox + `report_dispense` client; `cashOutHeld` latch and guard;
|
||
`counts_uncertain` on zero-with-error. Ships first — with no server handler the RPC
|
||
returns an error and the outbox simply retries, so the machine is never blocked on it.
|
||
2. **spirekeeper:** `report_dispense` handler + `dispense_reports` table; `awaiting_dispense` /
|
||
`cash_owed` / `partial_pending` / `dispense_unreported` statuses; gate in `_handle_payment`;
|
||
worklist buckets + notification; `settle_cash_owed`; `cash_out_enabled` and the two new
|
||
operator ops; `dca_machines.cash_out_held_*`.
|
||
3. **Both:** `resume_cash_out` / `settle_transaction` ops on the machine; beacon reflects
|
||
held; docs: `docs/nostr-patterns` entry for the outbox-RPC pattern, this ADR → Accepted.
|
||
|
||
## Future directions
|
||
|
||
Recorded 2026-10-09 so the decisions above are made with the destination in view. None of
|
||
these are decided; several would change what "owed cash" even means.
|
||
|
||
### Hold invoices — capture at the Lightning layer instead of the application layer
|
||
|
||
Decision 1 implements authorize/capture in spirekeeper because a plain BOLT11 payment is
|
||
final the moment it lands. A **hold (HODL) invoice** moves that boundary into the protocol:
|
||
the payer's HTLC is accepted but not settled until the receiver reveals the preimage, and
|
||
can be cancelled instead, returning the funds with no second payment. The ATM would mint the
|
||
preimage, create the hold invoice, dispense, and then `settle` on `dispense_confirmed` or
|
||
`cancel` on a fault. A cancelled hold means nobody is owed anything — the customer's funds
|
||
were never taken.
|
||
|
||
What is already there: LNbits core has `create_hold_invoice`, `settle_hold_invoice(preimage)`
|
||
and `cancel_hold_invoice` (`lnbits/core/services/payments.py`), tagging `extra.hold_invoice`.
|
||
It is implemented for the **lndrest and lndgrpc** funding sources only; other backends raise
|
||
*"Hold invoices are not supported by the funding source."* None of the three is exposed over
|
||
the nostr-transport yet, so three RPCs are needed before the machine can use them. Spark's
|
||
`createLightningHodlInvoice({ amountSats, paymentHash, … })` and RoboSats' escrow bonds are
|
||
the reference shapes — the payer-visible behaviour (a pending payment that later settles or
|
||
cancels) is identical.
|
||
|
||
Two properties bound what this buys:
|
||
|
||
- **Settlement is all-or-nothing per HTLC.** A hold invoice cannot be partially settled, so
|
||
the machine's rule is **`dispensed > 0 → settle; dispensed == 0 → cancel`**. A fault with
|
||
nothing presented cancels cleanly — and that includes an exit jam like sintra's 2026-10-09,
|
||
where the note stopped in the transport and the counter read zero: the customer is charged
|
||
nothing and the note is the operator's to recover, with `countsUncertainSince` covering the
|
||
inventory side. A **partial** — notes actually in the customer's hand — cannot cancel without
|
||
refunding someone holding cash, so it settles the full amount and from that instant is
|
||
identical to a plain-BOLT11 partial: `partial_pending`, one distribution when the shortfall
|
||
is resolved (Decision 1). Hold invoices eliminate owed-cash for full faults and exit jams,
|
||
not for true partials.
|
||
- **The hold window locks the payer's funds and route liquidity**, and some wallets surface
|
||
a long-pending payment as a failure. The window should equal the dispense window — seconds,
|
||
capped at a minute or two — with an automatic `cancel` on timeout, never an open-ended hold.
|
||
|
||
Where it lands in this ADR: a settlement created from a held payment reaches
|
||
`_handle_payment` only on **settle** (`payment.success` is false while held), so for hold-paid
|
||
transactions the `awaiting_dispense` state in Decision 1 is unnecessary — the Lightning layer
|
||
is already waiting. Decision 1 stays for plain BOLT11 and for any CLINK path that resolves to
|
||
an ordinary invoice. The two coexist; `report_dispense` (Decision 2) is what triggers the
|
||
settle/cancel either way.
|
||
|
||
### Vouchers — a fiat-denominated claim instead of a debt
|
||
|
||
For the shortfall a hold invoice cannot refund, and for any dispense failure on a plain
|
||
invoice, the customer could be issued a **voucher**: a claim on the operator for *X fiat value
|
||
of notes*, redeemable at a machine at a later date **at the original exchange rate**,
|
||
independent of the BTC price at redemption and requiring no further Lightning payment. The
|
||
machine prints or displays it; redemption is a cash-out whose "payment" is the voucher.
|
||
|
||
Vouchers could also be a general product — buy X fiat value now, collect later — but the
|
||
first use is fault recovery. Accounting constraints that must hold whichever form ships:
|
||
|
||
- A voucher is a **liability** row on the server, linked to its origin `txid` / settlement
|
||
when it has one (nullable for the general case), with the fiat amount, the locked rate, the
|
||
issuing machine and an expiry.
|
||
- The sats the origin settlement received for the undispensed portion must **not** be
|
||
distributed while the voucher is open: redemption or expiry is what releases them, so the
|
||
operator never pays out commission on cash that has not left a machine. This is the same
|
||
rule as Decision 1 applied over a longer window.
|
||
- Redemption records a `cash_out` with `tx_type = 'voucher_redeem'`, `wire_sats = 0`, and a
|
||
reference to the voucher; the machine's own ledger treats it as a dispense like any other
|
||
(bays decrement, `cassette_bills` written, `report_dispense` sent).
|
||
- A voucher is a bearer instrument unless bound to a pubkey. Both are possible; the choice
|
||
decides whether a lost voucher is lost money.
|
||
|
||
Taken together, hold invoices plus vouchers could remove the *owed cash* state entirely:
|
||
full fault → cancel, nobody pays; partial fault → settle, voucher for the difference; and
|
||
`cash_owed` in Decision 6 becomes the fallback for a plain-invoice path, not the norm.
|
||
|
||
### CLINK — the protocol this machine is moving toward
|
||
|
||
bitSpire is a Nostr machine and the long-term direction is to implement, and eventually
|
||
favour, Shocknet's **CLINK** (Common Lightning Interface for Nostr Keys) for customer-facing
|
||
payment flows: kind 21001 Offers (`noffer`), 21002 Debits (`ndebit`), 21003 Manage, 21004
|
||
Enroll, and the **CLINK Beacon** — a kind-30078 service heartbeat carrying liveness, persona
|
||
and fee disclosure. Reference tree: `~/dev/refs/repos/shocknet/shocknet/{CLINK,ClinkSDK,
|
||
clink-demo,Lightning.Pub}`. The `@bitSpire/clink` package is in tree and dormant.
|
||
|
||
Two consequences for this ADR. First, BOLT11 stays alongside CLINK rather than being
|
||
replaced, so Decisions 1–2 remain load-bearing for the invoice path. Second, review finding 4
|
||
below — the availability beacon has no readers and overlaps the cassettes-state document —
|
||
should be resolved by aligning the beacon with the **CLINK Beacon** spec rather than by
|
||
inventing a third shape: a machine advertising itself to Nostr clients should do so in the
|
||
form those clients will read.
|
||
|
||
### Operator notification is Nostr-native
|
||
|
||
Decision 6 calls for the operator to be notified when a settlement reaches `cash_owed`.
|
||
That notification is **a Nostr event to the operator's pubkey**, not email or SMS (the
|
||
lamassu `notifyOperator` channel). The pieces exist:
|
||
|
||
- The operator's pubkey is already on their LNbits account — `get_machine_config` refuses to
|
||
run without it ("operator has no Nostr pubkey on file").
|
||
- The sender signs through the bunker (`sign_as_operator` / `resolve_operator_signer` in
|
||
`nostr_publish.py`), so no key is at rest on the server. A dedicated server identity for
|
||
alerts is preferable to the operator messaging themself.
|
||
- The operator does **not** need their private key to read it. nsecbunkerd supports
|
||
`nip44_encrypt` / `nip44_decrypt` for its users (see its ACL tests), so any NIP-46 client —
|
||
our webapp, Amber, nsec.app — decrypts the message through the bunker. nsecbunkerd does hold
|
||
a `decryptNsec` path, but it is the bunker *admin's*, gated on the passphrase; exposing it to
|
||
operators would reverse the "no nsec outside the bunker" principle this stack is built on.
|
||
- Operators may still want alerts on a phone identity that is not their LNbits pubkey. An
|
||
optional per-operator `alerts_pubkey` covers that without changing the default.
|
||
|
||
NIP-17 (kind 14 in a kind-1059 gift wrap) is the right wire for metadata privacy; NIP-04 is an
|
||
acceptable interim if the receiving clients are ours. Pick one in the notification issue.
|
||
|
||
### An operator-facing error glossary
|
||
|
||
Every `error_code` surfaced by Decisions 2 and 7 should link to a glossary entry: what the
|
||
code means, what the operator will find when they open the machine, what clears it, and
|
||
whether it is `terminal` or `recoverable`. The authoritative source for the F56 is the Fujitsu
|
||
Frontech **F56-BDU Error Code List** (K3KD03234–K3KD03236-0001, edition E02), per the Lamassu
|
||
port backlog; it is not held locally yet and should be obtained. Seed entries, from the
|
||
backlog and from this incident:
|
||
|
||
| raw | meaning (F56-BDU) | class | observed |
|
||
| ------- | --------------------------------- | ----------- | ------------------------- |
|
||
| `78 42` | note stopped at the cassette exit | terminal | sintra, 2026-10-09 |
|
||
| `82 00` | bill length — long | recoverable | Tejo GTQ, 2026-09-26 |
|
||
| `83 00` | bill length — short | recoverable | |
|
||
| `84 00` | bill thickness | recoverable | |
|
||
| `85 0n` | pick from another safe | recoverable | |
|
||
| `86 00` | bill spacing | recoverable | |
|
||
| `B5 ..` | reject box overflow | terminal | |
|
||
|
||
The glossary lives in `docs/` and is served by spirekeeper so the dashboard can deep-link
|
||
from a report. Neither lamassu codebase ever decoded these bytes.
|
||
|
||
### The Lamassu port backlog
|
||
|
||
A provenance-gated analysis of what bitSpire and spirekeeper can take from lamassu-machine
|
||
and lamassu-server exists as *Lamassu Port Backlog* (27 Sep 2026, claude.ai artifact
|
||
`61af38f6`). Items that bear directly on this ADR: **structured fault reporting**
|
||
(`routes/diagnosticsRoutes.js` → `machine-loader.updateDiagnostics` — a fault record carrying
|
||
the driver error, the raw device frame and cassette state at failure, which is `report_dispense`
|
||
by another name); the **interactive hardware test harness** (`lib/hardware-testing/`, an F56
|
||
dispense-and-count case is the obvious first addition); the **denomination solver**
|
||
(`lib/coin-change.js`, portable from its upstream `git.sr.ht/~siiky/coin-change`); and the
|
||
ID003 startup and hang fixes. The backlog also records that bitSpire carried lamassu's narrow
|
||
GTQ note-length window byte for byte — fixed on `dev` alongside this ADR, mirroring
|
||
lamassu-machine `b1cc3622`.
|
||
|
||
### Bay layout, machine identity, and the operator docs that do not exist yet
|
||
|
||
How many bays a machine has is **machine-authoritative**: on first boot the layout comes from
|
||
`VITE_LAMASSU_CASSETTES` in the machine's `.env` (documented in `docs/device-configuration.md`
|
||
— an installer document, with a stale `LAMASSU` prefix), after which `state.db` owns it and
|
||
the operator adjusts counts and denominations by publishing ops. spirekeeper adopts whatever
|
||
the machine reports: `apply_reported_state` treats the payload as the full bay set and
|
||
*deletes* positions the machine no longer reports ("the bay count is hardware-determined").
|
||
There is **no operator-facing document** describing any of this — spirekeeper's README still
|
||
describes the satmachineadmin era and has no setup section — so an operator with a two-bay
|
||
Tejo and one with a four-bay Tejo have no page telling them where the difference is set.
|
||
|
||
That gap is one face of a larger one: every fleet target in `flake.nix` is a **specific
|
||
machine** (fiat, upgrade timer, card reader, address are keyed by hostname), so a second Tejo
|
||
cannot join the fleet by pointing at the same flake. Generalising the install means splitting
|
||
**model** (the hardware preset: `sintra`, `tejo`, `douro`, `batm3`) from **identity** (the
|
||
per-install provisioning: seed, cassettes, fiat, network), with the latter arriving through
|
||
pairing and operator ops rather than through Nix. The operator docs should be written against
|
||
that split, not the current one.
|
||
|
||
## Review findings outside this ADR
|
||
|
||
Found while tracing the money path end to end for this document. Not decided here; each is
|
||
a candidate issue.
|
||
|
||
1. **The settlement push did not deliver — `[ATM Service] Invoice paid (poll)!`** The
|
||
`subscribe_payments` stream missed the payment and the polling fallback caught it. Worth
|
||
knowing before relying on push latency anywhere; related to #78.
|
||
2. **`waitingForCashTaken` auto-advances to `complete` after 30 s "assume taken."** A
|
||
transaction can be recorded complete with notes still in the slot. The HAL already blocks
|
||
in `waitForBillsRemoved` inside `dispenseCash`, so the state is doing a second, weaker
|
||
version of the same job.
|
||
3. **The machine and server keep two ledgers with no reconciliation.** `transactions` and
|
||
`dca_settlements` join on `txid` today and nothing ever joins them. Decision 2 gives the
|
||
server everything a reconciliation needs.
|
||
4. **The availability beacon has no consumers and overlaps the cassettes-state document.**
|
||
Two machine-authored state documents with overlapping fields is drift waiting to happen.
|
||
Either fold availability into the cassettes-state doc or give the beacon a reader.
|
||
5. **`is_active` reads as a service gate and is a roster flag.** Rename to something like
|
||
`enrolled`, or document at the column.
|
||
6. **`generateInvoice` carries a comment deferring `bills`/`cassettes` onto the invoice
|
||
`extra`.** With Decision 2 that would be the wrong place — provisioned is not dispensed —
|
||
and the comment invites a future contributor to wire it there. Remove it.
|
||
7. **The HAL's `dispensed` boolean is count-based (`totalRequested === totalDispensed`),
|
||
not value-based.** Equivalent only while each bay dispenses its own denomination.
|
||
Decision 3 replaces it.
|
||
8. **`_handle_payment` processes cash-in and cash-out through one path and only `tx_type`
|
||
tells them apart.** Decision 1 adds a second direction-specific branch. If a third
|
||
arrives, split the handler.
|
||
9. **The partial-dispense guard's message is wrong for internal legs.** "Lightning payments
|
||
can't be clawed back" is true of `autoforward` and false of the LNbits-internal legs,
|
||
which are compensatable. Make the guard leg-aware or correct the message.
|
||
10. **`packages/hal` has no tests.** `pnpm test` there exits 1 with "No test files found."
|
||
The F56 note-length table that produced a production fault on a GTQ Tejo was carried
|
||
byte for byte from lamassu with nothing over it; the fix landed the same way. A table test
|
||
asserting every currency's window is centred on its note length is a few lines, and
|
||
`dispenseConfirmed` (Decision 3) needs the harness to exist before it can be tested.
|
||
11. **Review scope.** This document traced the cash-out path: state machine → HAL → ledger →
|
||
transport → settlement → distribution → dashboard. The cash-in path shares the settlement
|
||
pipeline and has its own money-at-risk shape in `create_withdraw` (server-side amounts,
|
||
`max_cash_in_sats`); it has not been reviewed to the same depth and is the obvious next
|
||
slice. The HAL drivers, access layer and deploy module were not in scope.
|