docs(adr): ADR-005 — future directions and the operator-docs gap

Records where the cash-out design is heading so Decisions 1–7 are made
with the destination in view:

- Hold invoices move authorize/capture from spirekeeper into the
  Lightning layer. LNbits core already has create/settle/cancel
  (lndrest + lndgrpc only; not yet on the nostr-transport). Settlement is
  all-or-nothing per HTLC, so a full fault cancels cleanly but a partial
  fault still needs a voucher — hold invoices remove owed-cash for the
  common case, not every case.
- Vouchers: a fiat-denominated claim at the original rate, a liability
  row linked to its origin settlement, whose undispensed sats stay
  undistributed until redemption or expiry.
- CLINK as the eventual favoured customer protocol; the availability
  beacon should align with the CLINK Beacon spec rather than grow a third
  shape.
- Operator notification is a Nostr event to the operator's pubkey, not
  email/SMS. The pubkey is already on the LNbits account; nsecbunkerd
  supports nip44 so no operator ever needs their nsec.
- An operator-facing error glossary, seeded from the F56-BDU error list
  and this incident's 78 42.
- Cross-reference to the Lamassu Port Backlog, and the finding that
  bitSpire carried lamassu's narrow GTQ window byte for byte.
- Bay layout is machine-authoritative (VITE_LAMASSU_CASSETTES on first
  boot, then state.db); spirekeeper adopts and deletes absent positions;
  no operator document describes any of it, and fleet targets are keyed
  by hostname so a second Tejo cannot join without a new flake target.

Refs #122
This commit is contained in:
Padreug 2026-10-09 21:42:03 +02:00
commit 6a974054ce

View file

@ -276,6 +276,170 @@ code is borrowed from another layer's vocabulary (the lamassu 570 lesson).
3. **Both:** `resume_cash_out` / `settle_transaction` ops on the machine; beacon reflects
held; docs: `docs/nostr-patterns` entry for the outbox-RPC pattern, this ADR → Accepted.
## Future directions
Recorded 2026-10-09 so the decisions above are made with the destination in view. None of
these are decided; several would change what "owed cash" even means.
### Hold invoices — capture at the Lightning layer instead of the application layer
Decision 1 implements authorize/capture in spirekeeper because a plain BOLT11 payment is
final the moment it lands. A **hold (HODL) invoice** moves that boundary into the protocol:
the payer's HTLC is accepted but not settled until the receiver reveals the preimage, and
can be cancelled instead, returning the funds with no second payment. The ATM would mint the
preimage, create the hold invoice, dispense, and then `settle` on `dispense_confirmed` or
`cancel` on a fault. A cancelled hold means nobody is owed anything — the customer's funds
were never taken.
What is already there: LNbits core has `create_hold_invoice`, `settle_hold_invoice(preimage)`
and `cancel_hold_invoice` (`lnbits/core/services/payments.py`), tagging `extra.hold_invoice`.
It is implemented for the **lndrest and lndgrpc** funding sources only; other backends raise
*"Hold invoices are not supported by the funding source."* None of the three is exposed over
the nostr-transport yet, so three RPCs are needed before the machine can use them. Spark's
`createLightningHodlInvoice({ amountSats, paymentHash, … })` and RoboSats' escrow bonds are
the reference shapes — the payer-visible behaviour (a pending payment that later settles or
cancels) is identical.
Two properties bound what this buys:
- **Settlement is all-or-nothing per HTLC.** A hold invoice cannot be partially settled. A
**full** fault (nothing dispensed) is cleanly cancelled. A **partial** fault — notes out,
value short — cannot be: cancelling would refund a customer who is holding cash, and
settling takes the full amount. Partial therefore still needs the voucher or refund path
below. Hold invoices eliminate owed-cash for the common full-fault case, not for every case.
- **The hold window locks the payer's funds and route liquidity**, and some wallets surface
a long-pending payment as a failure. The window should equal the dispense window — seconds,
capped at a minute or two — with an automatic `cancel` on timeout, never an open-ended hold.
Where it lands in this ADR: a settlement created from a held payment reaches
`_handle_payment` only on **settle** (`payment.success` is false while held), so for hold-paid
transactions the `awaiting_dispense` state in Decision 1 is unnecessary — the Lightning layer
is already waiting. Decision 1 stays for plain BOLT11 and for any CLINK path that resolves to
an ordinary invoice. The two coexist; `report_dispense` (Decision 2) is what triggers the
settle/cancel either way.
### Vouchers — a fiat-denominated claim instead of a debt
For the shortfall a hold invoice cannot refund, and for any dispense failure on a plain
invoice, the customer could be issued a **voucher**: a claim on the operator for *X fiat value
of notes*, redeemable at a machine at a later date **at the original exchange rate**,
independent of the BTC price at redemption and requiring no further Lightning payment. The
machine prints or displays it; redemption is a cash-out whose "payment" is the voucher.
Vouchers could also be a general product — buy X fiat value now, collect later — but the
first use is fault recovery. Accounting constraints that must hold whichever form ships:
- A voucher is a **liability** row on the server, linked to its origin `txid` / settlement
when it has one (nullable for the general case), with the fiat amount, the locked rate, the
issuing machine and an expiry.
- The sats the origin settlement received for the undispensed portion must **not** be
distributed while the voucher is open: redemption or expiry is what releases them, so the
operator never pays out commission on cash that has not left a machine. This is the same
rule as Decision 1 applied over a longer window.
- Redemption records a `cash_out` with `tx_type = 'voucher_redeem'`, `wire_sats = 0`, and a
reference to the voucher; the machine's own ledger treats it as a dispense like any other
(bays decrement, `cassette_bills` written, `report_dispense` sent).
- A voucher is a bearer instrument unless bound to a pubkey. Both are possible; the choice
decides whether a lost voucher is lost money.
Taken together, hold invoices plus vouchers could remove the *owed cash* state entirely:
full fault → cancel, nobody pays; partial fault → settle, voucher for the difference; and
`cash_owed` in Decision 6 becomes the fallback for a plain-invoice path, not the norm.
### CLINK — the protocol this machine is moving toward
bitSpire is a Nostr machine and the long-term direction is to implement, and eventually
favour, Shocknet's **CLINK** (Common Lightning Interface for Nostr Keys) for customer-facing
payment flows: kind 21001 Offers (`noffer`), 21002 Debits (`ndebit`), 21003 Manage, 21004
Enroll, and the **CLINK Beacon** — a kind-30078 service heartbeat carrying liveness, persona
and fee disclosure. Reference tree: `~/dev/refs/repos/shocknet/shocknet/{CLINK,ClinkSDK,
clink-demo,Lightning.Pub}`. The `@bitSpire/clink` package is in tree and dormant.
Two consequences for this ADR. First, BOLT11 stays alongside CLINK rather than being
replaced, so Decisions 1–2 remain load-bearing for the invoice path. Second, review finding 4
below — the availability beacon has no readers and overlaps the cassettes-state document —
should be resolved by aligning the beacon with the **CLINK Beacon** spec rather than by
inventing a third shape: a machine advertising itself to Nostr clients should do so in the
form those clients will read.
### Operator notification is Nostr-native
Decision 6 calls for the operator to be notified when a settlement reaches `cash_owed`.
That notification is **a Nostr event to the operator's pubkey**, not email or SMS (the
lamassu `notifyOperator` channel). The pieces exist:
- The operator's pubkey is already on their LNbits account — `get_machine_config` refuses to
run without it ("operator has no Nostr pubkey on file").
- The sender signs through the bunker (`sign_as_operator` / `resolve_operator_signer` in
`nostr_publish.py`), so no key is at rest on the server. A dedicated server identity for
alerts is preferable to the operator messaging themself.
- The operator does **not** need their private key to read it. nsecbunkerd supports
`nip44_encrypt` / `nip44_decrypt` for its users (see its ACL tests), so any NIP-46 client —
our webapp, Amber, nsec.app — decrypts the message through the bunker. nsecbunkerd does hold
a `decryptNsec` path, but it is the bunker *admin's*, gated on the passphrase; exposing it to
operators would reverse the "no nsec outside the bunker" principle this stack is built on.
- Operators may still want alerts on a phone identity that is not their LNbits pubkey. An
optional per-operator `alerts_pubkey` covers that without changing the default.
NIP-17 (kind 14 in a kind-1059 gift wrap) is the right wire for metadata privacy; NIP-04 is an
acceptable interim if the receiving clients are ours. Pick one in the notification issue.
### An operator-facing error glossary
Every `error_code` surfaced by Decisions 2 and 7 should link to a glossary entry: what the
code means, what the operator will find when they open the machine, what clears it, and
whether it is `terminal` or `recoverable`. The authoritative source for the F56 is the Fujitsu
Frontech **F56-BDU Error Code List** (K3KD03234–K3KD03236-0001, edition E02), per the Lamassu
port backlog; it is not held locally yet and should be obtained. Seed entries, from the
backlog and from this incident:
| raw | meaning (F56-BDU) | class | observed |
| ------- | --------------------------------- | ----------- | ------------------------- |
| `78 42` | note stopped at the cassette exit | terminal | sintra, 2026-10-09 |
| `82 00` | bill length — long | recoverable | Tejo GTQ, 2026-09-26 |
| `83 00` | bill length — short | recoverable | |
| `84 00` | bill thickness | recoverable | |
| `85 0n` | pick from another safe | recoverable | |
| `86 00` | bill spacing | recoverable | |
| `B5 ..` | reject box overflow | terminal | |
The glossary lives in `docs/` and is served by spirekeeper so the dashboard can deep-link
from a report. Neither lamassu codebase ever decoded these bytes.
### The Lamassu port backlog
A provenance-gated analysis of what bitSpire and spirekeeper can take from lamassu-machine
and lamassu-server exists as *Lamassu Port Backlog* (27 Sep 2026, claude.ai artifact
`61af38f6`). Items that bear directly on this ADR: **structured fault reporting**
(`routes/diagnosticsRoutes.js` → `machine-loader.updateDiagnostics` — a fault record carrying
the driver error, the raw device frame and cassette state at failure, which is `report_dispense`
by another name); the **interactive hardware test harness** (`lib/hardware-testing/`, an F56
dispense-and-count case is the obvious first addition); the **denomination solver**
(`lib/coin-change.js`, portable from its upstream `git.sr.ht/~siiky/coin-change`); and the
ID003 startup and hang fixes. The backlog also records that bitSpire carried lamassu's narrow
GTQ note-length window byte for byte — fixed on `dev` alongside this ADR, mirroring
lamassu-machine `b1cc3622`.
### Bay layout, machine identity, and the operator docs that do not exist yet
How many bays a machine has is **machine-authoritative**: on first boot the layout comes from
`VITE_LAMASSU_CASSETTES` in the machine's `.env` (documented in `docs/device-configuration.md`
— an installer document, with a stale `LAMASSU` prefix), after which `state.db` owns it and
the operator adjusts counts and denominations by publishing ops. spirekeeper adopts whatever
the machine reports: `apply_reported_state` treats the payload as the full bay set and
*deletes* positions the machine no longer reports ("the bay count is hardware-determined").
There is **no operator-facing document** describing any of this — spirekeeper's README still
describes the satmachineadmin era and has no setup section — so an operator with a two-bay
Tejo and one with a four-bay Tejo have no page telling them where the difference is set.
That gap is one face of a larger one: every fleet target in `flake.nix` is a **specific
machine** (fiat, upgrade timer, card reader, address are keyed by hostname), so a second Tejo
cannot join the fleet by pointing at the same flake. Generalising the install means splitting
**model** (the hardware preset: `sintra`, `tejo`, `douro`, `batm3`) from **identity** (the
per-install provisioning: seed, cassettes, fiat, network), with the latter arriving through
pairing and operator ops rather than through Nix. The operator docs should be written against
that split, not the current one.
## Review findings outside this ADR
Found while tracing the money path end to end for this document. Not decided here; each is