bug(cash-out): a jammed dispense takes the customer's sats, tells the server nothing, over-counts the cassette, and leaves the machine advertising itself as available #122
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Sintra, 2026-10-09 01:02 CST. A customer paid a 40 EUR cash-out (54 440 sats) and got no cash. A 20 EUR note was found jammed just past the cassette exit. Three separate defects stacked up on one transaction; the money one is first.
Machine-side record is complete and correct:
Timeline from the journal:
1. The failure never leaves the machine (the money defect)
dispenseErroris a terminal local state. The row above is written tostate.db, the screen counts down 30s, and that is the end of it. Nothing publishes the outcome, nothing alerts, nothing refunds —grep -riE "refund|reversal|compensat|notify.*fail"acrossapps/andpackages/returns nothing on the money path. So the operator's view is an LNbits payment markedprocessedwith no hint of trouble, and the only record that a customer is owed 40 EUR lives on the ATM's own disk.The hook for fixing this already exists:
generateInvoiceputsextra.txidon the invoice (apps/machine/src/services/lightning.ts:1004-1011), so the server can join a payment to a machine transaction. What's missing is the machine ever telling it the join came out bad. #78 noted in passing that "only spirekeeper-side reconciliation would ever notice" — this is that gap arriving in production, on a transaction that was watched correctly and recorded correctly.A terminal dispense failure should publish the outcome for the txid the payment already carries, so the operator sees an owed-cash queue instead of a clean
processed.2. A mid-transport jam is recorded as
dispensed: 0, so the ledger over-countsThe F56 reports per-bay dispensed/rejected alongside the error. Decoding the response bytes: error code at
res[3..5]=78 42; dispensed at0x27=30 30→ 0; rejected at0x2f=30 30→ 0. A note that left the bay and stopped in the transport completes neither counter, so the driver faithfully reports nothing moved.Consequences, all visible in the log above:
hal-service.ts:320decrements each bay byresult.value[i].dispensed— zero, so the bay stays at 66.atm.ts:628-634seestotalDispensed === 0, keepsstatus = dispense_error(notpartial) and writesbills = [].markCountsUncertain(atm.ts:645-657) is the safety net for exactly this — "bills may well have reached the customer, and nothing knows how many" — but it only fires whendispenseResultis absent. A structured report claimingdispensed: 0is trusted outright, socountsUncertainSincewas never set (metahas no such key).publishCassettesState()then republished 20x66 as fact, three times since.cassettessays 66 twenties. The cassette holds 65 and the transport holds one. Adispensed: 0accompanied by an error is not evidence that nothing left the bay — it should route to the same uncertainty flag as a missing report.3. The availability beacon has no notion of dispenser health
useAvailabilityBroadcast.computeSnapshot()derivescashOutfromtotalBills > 0and nothing else. So 64 seconds after a jam that will block every subsequent dispense, sintra published{"cash_in":true,"cash_out":true,"cash_level":"full"}, and has kept republishing it every 5 minutes since.This is what turns one bad transaction into many.
dispenseCashre-inits the dispenser on the next attempt (hal-service.ts:255-259), but a note physically stuck in the transport path isn't cleared by re-initialising — so the next customer to tap, pay, and wait would have lost their money the same way. Nothing in the beacon, and nothing in thelockedstate the machine returned to, stops them.#27 already classifies JAM as terminal ("bill stuck in transport path, will block all cassettes") and correctly says don't retry it. The missing half is that terminal means stop advertising cash-out until an operator clears it and confirms a recount.
Scope note
Distinct from #78 (payment with no machine row at all — unwatched window) and #27 (retry/fallback across cassettes on recoverable pickup errors). Here the watcher worked, the record is right, and the hardware error was correctly terminal; what failed is everything after the error.
Suggested split if these want to be separate: (1) is the release-blocker, (2) is a correctness fix in the
dispenseResultinterpretation, (3) is a beacon/health-gate feature. They're filed together because one transaction produced all three and (3) is what makes (1) recur.A fourth gap, surfaced by actually responding to this incident: there is no way to record that a failed dispense was settled off-machine.
The operator cleared the jam and handed the customer 40 EUR by hand, pulling the notes from the cassette directly rather than through the machine. That is the obvious real-world response, and the ledger has no vocabulary for it:
remediateTransaction(txid, remediatedByTxid)requires amanual_dispensetxid to point at. A hand-settled payout produces none, and running a real manual dispense just to mint one would push another 40 EUR out of the machine.tx_mv0madw6_wdhtea1vstaysstatus = dispense_errorindefinitely, and the machine's own ledger keeps asserting the customer is owed money that has in fact been paid.remediated_bycolumn isTEXT, so the storage is already there — what's missing is a path that writes a non-txid provenance (an operator note / settlement reference) and flips the status.This compounds defect 1 rather than being separate from it: once failures do reach the server, the operator needs a way to close one out, and "the only way to clear an owed-cash row is to dispense more cash from the machine that just jammed" is not it.
Worth considering as part of the same design: a
settleoperator op, carried the same way the cassette ops are (operator-authored, id-deduped, applied by the machine as sole writer), so an off-machine payout is as auditable as a refill.Also for the record — how NOT to fix the count
The cassette row read
20x66while the bay physically held fewer. The temptation is aUPDATE cassettes SET count = …againststate.db, and that is wrong:applyOperatorCassetteOpsis deliberately the single writer ("Nobody but this process writes a count any more, so there is no second writer to lose a race to"), and a hand-edit lands in neithercassette_opsnor the published cassette state, so the operator's view and the machine's diverge with no audit trail.The supported repair is an operator-published
recountop for the affected position — which is also the only op that clearscountsUncertainSince, precisely because a recount is someone opening the bay and counting it. Any fix for defect 2 should make a jammed dispense set that flag so the next recount is the thing that resolves it.Spec'd as ADR-005 —
docs/adr/005-cash-out-dispense-outcome.md(ondev, lands with the next push).Tracing the whole cash-out path for it turned up a structural root underneath the four defects above, and it changes what "fix" means here:
Distribution runs before the dispense.
_handle_paymentspawnsprocess_settlementthe instant the payment lands — before the machine has even learned of the payment viawatchInvoice, let alone commanded the dispenser. The legs are LNbits-internal transfers and complete sub-second. So when the F56 reported78 42at T+2 s, the settlement was alreadyprocessed— which the dashboard reported accurately.processedmeans all legs paid; it has never had a dispense input.That also means
apply_partial_dispense_and_redistribute— the tool built for exactly this — is structurally unreachable: its hard guard refuses once any leg has completed, and under the current ordering a leg has always completed by the time a dispense can fail.ADR-005's central decision is therefore authorize/capture: the payment is the authorization, the machine's dispense report is the capture, and distribution waits for it. Everything else (the
report_dispenseRPC with a durable outbox, the lamassudispense_confirmed/error/error_codetaxonomy, the per-bay action log the machine already records incassette_bills, the fault-vs-out-of-cash customer screen, the terminal-fault latch cleared by recount, the off-machine settle path) hangs off that.Three deliberate deviations from the lamassu baseline are called out in the doc, including the one that caused defect 2: their
dispenseOccurredtrusts adispensed: 0that arrives with an error, and their counts drift the same way sintra's did.The ADR ends with ten review findings outside its scope — candidates, not filed.
Status 2026-10-10 — both halves of ADR-005's capture logic are merged.
dev): value-confirmed dispense,dispenseFaultvsoutOfCashscreens with evidence, terminal-fault cash-out latch, counts-uncertain on a zero-with-error report, and thedispense_reportsoutbox →report_dispense. Defects 1–4 above and the fourth gap (off-machine settle) all have their machine side in place.main, released as v0.1.7, in the catalog):cash_outsettlements land asawaiting_dispenseand are captured on the machine's report — confirmed → distributes; partial →partial_pending, held whole (ADR-005 Decision 1); nothing out →cash_owed, first on the worklist. Migration m016 verified on bohm's dev instance;report_dispenseregistered.Deploy state: not yet on hardware. Sintra's nightly upgrade has been failing every night since Oct 6 on the same 60 s build timeout as batm3 (#116), so nothing merged this week has reached it; the fix is pushing the toplevel to the aiolabs cachix (
deploy/push-cache.sh sintra), which in turn surfaced thatnix/mkAtmApp.nix'spnpmDeps.hashwas stale after today's lockfile changes — being corrected ondevnow. ariege's LNbits needs the spirekeeper upgrade to 0.1.7 via the admin UI.Still open for this issue (ADR-005 slice 2): Nostr operator notification on
cash_owed, the off-machinesettle_cash_owedpath +settle_transactionop so a hand-paid customer closes both ledgers, thecash_out_enabledoperator switch, the error glossary, operator docs. Keeping this open until the first real report has round-tripped from sintra and the slice-2 items are either landed or split out.