ADR-005 §6: the operator paid the customer by hand and recorded it in
spirekeeper. The op carries the txid and the note; the machine flips its
own dispense_error/partial row to remediated via remediateTransaction,
which only touches rows still in an error state, so re-delivery is a
no-op. Closes the machine side of the ledger for an owed-cash sale
without dispensing anything.
The fleet table called sintra a working dev unit; its nixos-upgrade had
failed every night 2026-10-06 → 10-09 on the same 60 s atm-app timeout
as batm3, so nothing merged that week reached it. Recorded with the
interim rule — push-cache after every dev push, because mkAtmApp's
src = self makes any commit invalidate the cached toplevel — and the two
traps found while applying it: a lockfile change silently reuses a stale
pnpmDeps store unless the hash is re-derived, and activation scripts
don't have grep/sed on PATH.
NixOS activation scripts run with a minimal PATH that has coreutils but
not gnugrep or gnused. On sintra the snippet printed its success line and
then failed with "sed: command not found" (127) — the keys were never
renamed, the new app booted on fallback config, and switch-to-
configuration exited 2. Both binaries are now referenced by store path.
The atm-app's pnpm store is a fixed-output derivation keyed by this hash.
pnpm-lock.yaml changed twice today (the dead-script dependency prune in
763817b, then qrcode declared where fund-atm.ts uses it in cbff654) and
the hash was not updated, so nix reused the stale store and the
sandboxed `pnpm install --offline` failed on @types/qrcode
(ERR_PNPM_NO_OFFLINE_TARBALL). That is what sintra's nightly upgrade and
the cachix push were dying on.
Rule to carry forward: any commit that touches pnpm-lock.yaml must
re-derive this hash (blank it, build, paste the `got:` value) — the
local `pnpm build` passing says nothing about the nix build.
Every cash-out now produces one report_dispense — on success as well as
failure — and the machine does not stop sending it until spirekeeper
acknowledges it.
state.db gains a dispense_reports table (migration v13 → v14): the report
is written INSIDE recordTransaction's SQLite transaction, alongside the
transactions row, so a crash between the two cannot lose it. Rows carry
attempts / last_attempt_at / last_error / acked_at. Three IPC calls
(pending / ack / note-attempt) expose it to the renderer.
The store builds the report when a cash-out reaches complete,
dispenseFault or outOfCash: txid, payment hash, dispense_confirmed,
error / error_code / raw_code / error_class, per-denomination requested
vs dispensed vs rejected, the per-bay cassette record verbatim, and
counts_uncertain. The success report is what lets the server capture
(distribute) the settlement; the failure report is what puts a customer
on the owed-cash worklist instead of leaving the only record on the ATM.
Delivery is at-least-once: a flusher drains pending rows after each
persist, on relay (re)connect, and every 60 s, acking only on an OK reply
and backing off 30 s · 2^attempts (capped 1 h) otherwise. While
spirekeeper has not registered the RPC every send fails the same way; the
backoff keeps that quiet and the rows wait — this half ships first.
The lightning service exposes reportDispense; the function pointer is
set at all three lightning-init sites so the flusher works on every path.
One cash-out's dispense outcome, sent on success as well as failure —
the success report is what captures the settlement server-side. Field
names follow lamassu-server's cash_out_txs / cash_out_actions
(dispense_confirmed, error, error_code) with raw_code and error_class
alongside, per-denomination bills with `requested`, per-bay cassettes
verbatim, the payment hash as the join key, and counts_uncertain.
Idempotent on txid (the server upserts), so the call is wrapped in
idempotent() and safe for the machine's outbox to retry. Until
spirekeeper registers the RPC it rejects with LnbitsRpcError, which the
outbox treats like any other transient failure.
HAL glue (electron/hal-service.ts and the renderer-side services/hal.ts):
dispenseConfirmed is Σ(denomination × dispensed) === Σ(denomination ×
requested), computed on value. The driver's tagged error is carried
through as errorCode / rawCode / errorClass / human; pre-dispense
inventory refusals are errorClass 'inventory' so they route to outOfCash
rather than the fault screen. The manual-dispense command result keeps
its wire key `dispensed` (spirekeeper's poller reads it) and gains the
new fields alongside.
Cash-out hold: state-store persists it in meta as one JSON value beside
countsUncertainSince, idempotent on set (the first fault's `since` is
kept); IPC get/set/clear through preload. The store persists the hold the
moment the machine sets it and restores it into the machine on boot. A
recount clears it in the store (same gesture that clears counts-
uncertain); operator-config also honours a new resume_cash_out op — not a
cassette op, split off before applyOperatorCassetteOps, and honoured only
when stamped after the hold began so a re-delivered old resume cannot
clear a fresh fault. Either release calls back into the store, which
sends CASH_OUT_RELEASED. The cassettes-state document carries
cash_out_held_since / _reason / _code (additive, like
counts_uncertain_since); the availability beacon reports cash_out false
while held; the idle Sell button is disabled with the reason.
Store watcher: dispenseFault and outOfCash both record dispense_error /
partial (the customer has paid either way). A report of zero dispensed
WITH a hardware error now sets countsUncertainSince instead of being
trusted as zero — a note stopped in the transport completes neither
counter (sintra 2026-10-09: bay read 66, held 65, one in the transport).
Fault screen: both terminal states show "your payment went through",
amount paid, per-denomination dispensed, the txid as QR and text, the
payment hash (threaded from the settlement watch through PAYMENT_RECEIVED)
and the time, with "keep this reference" and an acknowledge button. The
raw dispenser code is not shown; it travels in the report.
DispenseCashResult.dispensed (a driver boolean) is replaced by
dispenseConfirmed — Σ(denomination × dispensed) equals the requested
value, computed by the HAL — plus errorCode / rawCode / errorClass.
dispensingCash.onDone guards on dispenseConfirmed and nothing else.
The single dispenseError state becomes two. dispenseFault: the dispenser
reported an error, the customer has paid and is owed — 120 s screen with
evidence, ACKNOWLEDGE_FAULT to dismiss. outOfCash: a shortfall with no
hardware error or an inventory refusal — 30 s. A hung dispense is a
terminal fault.
A terminal errorClass latches cash-out off: context.cashOutHeld, set by
latchCashOutIfTerminal, preserved across resetContext (it is machine
health, not transaction state), guarding idle's SELECT_CASH_OUT. Cash-in
is unaffected. Only CASH_OUT_RELEASED clears it — the store sends that
when an operator recount or resume_cash_out op lands; re-initialising
the dispenser never does, because re-init does not move a stuck note.
CASH_OUT_HELD lets the store restore a persisted hold on boot.
PAYMENT_RECEIVED now carries the payment hash into context.paymentHash
so the fault screen can show the reference the server indexes.
Tests: the dispense section is rewritten around outcomes — value
confirmation, fault vs out-of-cash routing, terminal latch + release,
recoverable does not latch, partial-with-error is a fault, inventory
refusal is out-of-cash, boot-restored hold gates, 30 s vs 120 s timers,
acknowledge/cancel, timeout latches. 46/46.
Every dispenser now returns a tagged DispenseError: errorCode (the
family name, e.g. F56DispenseError), rawCode (driver-native, '78 42'),
errorClass (terminal | recoverable | inventory) and a human decode. The
class is what the state machine routes on: terminal latches cash-out off,
recoverable shows the fault screen but stays in service, inventory means
nothing was asked of the hardware.
The F56 table is built empirically and from the Fujitsu F56-BDU Error
Code List, seeded with sintra's 78 42 (note stopped at the cassette exit,
terminal) and the Tejo's 82 00 (long-bill reject, recoverable), plus the
83/84/86 00 checks and the 85 0n / B5 .. families. An unknown code fails
SAFE — terminal — so an unfamiliar fault latches rather than letting the
next customer pay into it. f56-rs232 surfaces the raw code structurally
instead of only inside the message string.
Drops the borrowed statusCode 570 from the F56 driver: lamassu-server
read 570 as "insufficient funds", so a jam told operators to refill full
cassettes (their 34ba9203 fix). Puloon gets the same contract with every
fault terminal until it has a decode table.
packages/hal had no tests at all (ADR-005 finding 10). Adds the first
two: the decode table, and the bill-length table — every window [hi, lo]
sane, and GTQ/USD(/HNL when added) sharing one window for what is
physically the same 156 mm note. That second test would have caught the
GTQ fault months ago.
The fund-atm esbuild bundle imports `qrcode`, but apps/machine never
declared it — it resolved only through packages/nostr-client's
devDependency, which 763817b removed along with the dead scripts that
were the only reason it was there. The full `pnpm build` then failed at
its last step ("Could not resolve qrcode"), which is what the nix image
build runs. Declared (with @types/qrcode) in the package that imports it.
Decision 1's partial row said "operator confirms → distribute scaled"
and left the undispensed remainder's fate implicit. Made explicit:
partial_pending holds everything — including the share of the notes
that did dispense — until the operator records how the shortfall was
resolved (remediated → full amount; vouchered or written off → scaled),
then one distribution runs at that amount with the existing scaling
arithmetic. A vouchered remainder waits for redemption or expiry.
The alternative (scaled part now, remainder on resolution) is recorded
as deferred, not rejected: it needs a second additive distribution pass
the repo lacks, and a partial is almost always a terminal fault that
has latched cash-out off, so resolution is hours. Decided 2026-10-10.
Also carries the hold-invoice decision rule that fell out of the same
analysis: dispensed > 0 → settle, dispensed == 0 → cancel. An exit jam
that reports zero cancels cleanly — the customer is charged nothing and
the stuck note is the operator's to recover — so hold invoices remove
owed-cash for full faults and exit jams, not for true partials.
.gitignore still carried rules for docker/**/data/ and docker/.state/ —
paths that no longer exist — under a now-empty "# Docker" header. The
docs skill's example sync report named LIGHTNING_PUB_URL as its sample
env var; swapped for a variable that exists.
devenv regenerates this file on every `devenv shell`; the committed copy
was pinned to a directory that no longer exists
(~/Work/tries/2026-01-22-lamassu-refactor-packages/lamassu-next). Untracked
and gitignored beside .devenv/. The file stays on disk.
docker/ (two compose stacks, dev.sh, regtest.sh, start-with-regtest.sh,
regtest-bootstrap.sh, strfry.conf), packages/nostr-client/dev/ (nine
agent and test scripts), and the seventeen devenv commands that drove
them — infra-*, lncli, btccli, mine-blocks, auto-mine, setup-channel,
alice-*, fund-atm, test-setup, test-payment, node-info — along with the
devenv postgres service, DATABASE_URL, LIGHTNING_PUB_URL, pgcli,
docker-compose and the `just` runner (no justfile exists).
None of it could talk to the app on `dev`. Every piece was built around
Lightning.Pub (a `lightning-pub` service in both compose files, 37
references in dev.sh, LIGHTNING_PUB_PUBKEY and the :1776 API in every
dev script, a NIP-44 v1 implementation the project forbids), and the
last substantive change predates the LNbits cutover that deleted
packages/lightning. No container under either name exists on any
machine. Development runs against LNbits: FakeWallet needs nothing,
bohm's native instance answers on :5001, and the shared regtest stack
lives at ~/dev/local/docker/regtest, outside this repo.
devenv.nix keeps the toolchain, the hardware/serial utilities, the git
hooks and `relay-test`, now pointed at LNbits's bundled nostrrelay. The
Rust toolchain stays for the orphaned crate until that is removed on its
own. nostr-client drops the four dependencies and two devDependencies
only the dead scripts imported (@noble/curves, @scure/base,
@shocknet/clink-sdk, @stablelib/xchacha20, qrcode, ws); lockfile
regenerated, −272 lines.
Verified: devenv.nix parses; nostr-client 43/43 + tsc; clink 11/11 + tsc;
machine app vue-tsc clean. The ndebit-cash-in-flow doc, kept as CLINK
design history, now says the commands it quotes no longer exist here.
- @lamassu/clink import examples → @bitSpire/clink, the package's real name.
- machine-installation.md: the service user is `bitspire`, not `lamassu`
(renamed in configuration.nix long ago; the doc never followed).
- README: clone aiolabs/bitspire, not lamassu-next; the fleet sentence
claiming batm3/douro run `main` against Lightning.Pub was stale.
- nostr-check skill: table headers say bitSpire.
- hal-check skill: the boundary is c0b69d1, not v8.1.5 (CLAUDE.md corrected
this 2026-07-04; the skill kept asserting the wrong tag), and the
"forbidden operations" now reflect the recorded permission —
reference over port, name the source commit — plus a rule born of the
GTQ window: no value table without a test over it.
Deliberately kept: every `aiolabs/lamassu-next#NN` issue citation, the
provenance sections, "Ported from lamassu-machine" driver headers, and
the hardware names "Lamassu Sintra/Tejo/Douro" — those are the machines.
The aliases were added for the brand transition with a note to drop them
once nothing referenced them. Nothing does. The autoUpgrade comment still
said the legacy aiolabs/lamassu-next repo fed batm3 and douro; every live
machine pulls from this repo now (CLAUDE.md → Branch model).
LamassuEventKind → BitSpireEventKind (nostr-client; no consumers outside
the package), the kiosk theme localStorage keys lamassu-theme /
lamassu-color-mode → bitspire-* (a one-time theme reset on existing
kiosks), the ui-shared UMD global LamassuUIShared → BitSpireUIShared, and
the orphaned Rust HAL's Cargo name/description/repository plus its lib.rs
header — now also labelled as the unbuilt leftover it is.
nostr-client: 43/43 tests, tsc clean. Machine app: vue-tsc + electron tsc clean.
Container names (relay, bitcoind, lnd, lnd-alice, lightning-pub, miner,
postgres), the regtest bitcoind rpcuser/rpcpassword, the postgres role and
database (bitspire_dev), the dev admin token, BITSPIRE_HOST_IP, the devenv
project name, and the banners/log prefixes in the scripts. Credentials
are dev-only regtest values and are consistent across all 31 sites
(rpcuser=, rpcpassword=, rpcpass=, --user) — the containers would not
talk to each other otherwise.
Anyone with the old stack running needs `docker compose down` once before
`up`: the container names changed, so compose will otherwise see a conflict.
Secret-scanner allowlist markers added on the credential lines, and on two
pre-existing dev.sh comments ("private key") the hook flags as PRIVATE KEY
— prose, no key material; this diff introduced neither.
MACHINE_MODEL, FIAT_CODE, VALIDATOR_DEVICE, DISPENSER_DEVICE and CASSETTES
carried the old brand in their names. Renamed everywhere they are read
(device.ts, electron/main.ts), written (flake.nix, mkAtmApp.nix, live.nix,
provision-atm.sh, factory-reset-atm.sh) and documented (.env.example,
docs/device-configuration.md). No compatibility fallback in code: the
machine reads VITE_BITSPIRE_* and nothing else.
The deployed .env files are the one place the old names persist — sintra's
/var/lib/bitspire/.env holds all three keys today — and the machine reads
MACHINE_MODEL / FIAT_CODE / CASSETTES from that file on every boot. Renaming
the keys in code alone would boot a live machine on preset defaults (wrong
bays, wrong fiat) at the next nightly pull. So configuration.nix gains an
activation script, beside the existing lamassu→bitspire user migration,
that rewrites VITE_LAMASSU_* → VITE_BITSPIRE_* in that file. Idempotent;
runs before bitspire.service starts.
Every quetzal note is 156 x 67 mm — physically the same note as USD and
HNL, which the F56 table accepts at 146–166 (±10). GTQ was configured
at ±5 (151–161), with Q5 and Q20 further centred on 158 rather than 156.
This table was carried byte for byte from lamassu-machine, narrow window
included.
On a Tejo in GTQ it produced F56 error 82 00 (bill length, long) on
every Q100 pick — 5/5 notes rejected, 0 dispensed, a false "out of
cash" — byte-identical across transactions days apart, so the BDU was
reading >161 mm consistently. The reject tray held single notes, not
pairs, which rules out the offset double-pick the narrow window exists
to catch.
Every denomination now uses 0xa6 0x92 (146–166), identical to USD/HNL.
Mirrors lamassu-machine b1cc3622 (2026-09-29), ported with permission as
prior art.
Verification: tsc clean. packages/hal has no test files (vitest exits 1,
"No test files found") — recorded as ADR-005 review finding 10.
Records where the cash-out design is heading so Decisions 1–7 are made
with the destination in view:
- Hold invoices move authorize/capture from spirekeeper into the
Lightning layer. LNbits core already has create/settle/cancel
(lndrest + lndgrpc only; not yet on the nostr-transport). Settlement is
all-or-nothing per HTLC, so a full fault cancels cleanly but a partial
fault still needs a voucher — hold invoices remove owed-cash for the
common case, not every case.
- Vouchers: a fiat-denominated claim at the original rate, a liability
row linked to its origin settlement, whose undispensed sats stay
undistributed until redemption or expiry.
- CLINK as the eventual favoured customer protocol; the availability
beacon should align with the CLINK Beacon spec rather than grow a third
shape.
- Operator notification is a Nostr event to the operator's pubkey, not
email/SMS. The pubkey is already on the LNbits account; nsecbunkerd
supports nip44 so no operator ever needs their nsec.
- An operator-facing error glossary, seeded from the F56-BDU error list
and this incident's 78 42.
- Cross-reference to the Lamassu Port Backlog, and the finding that
bitSpire carried lamassu's narrow GTQ window byte for byte.
- Bay layout is machine-authoritative (VITE_LAMASSU_CASSETTES on first
boot, then state.db); spirekeeper adopts and deletes absent positions;
no operator document describes any of it, and fleet targets are keyed
by hostname so a second Tejo cannot join without a new flake target.
Refs #122
Three corrections to facts a session reads before touching code.
The maintainer taking over the Lamassu codebase has given permission to
use lamassu-machine and lamassu-server, post-boundary included, as prior
art (relayed by padreug, 2026-10-09). The hard rule that limited us to
c0b69d1 is superseded; the guidance is now reference-over-port, with
verbatim ports naming their source commit.
lamassu-server DID have a public-domain era — last open commit adbc9709,
licence added in d06a8f54, both 2023-09-19 — contrary to what the
squashed ~/lamassu/lamassu-server checkout suggests (its first commit is
already Appendix A). History survives in ~/dev/repos/ and at Software
Heritage. The earlier "not re-verified" note is replaced with the
verified boundary.
The hardware-driver table claimed ccnet, cashflow_sc, bnr_advance,
genmega, hcm2, gsr50 and three printers. The tree has validators id003
and ebds, dispensers f56 and puloon, no printers directory, and orphaned
Rust files from an abandoned HAL. The table now matches the tree.
A customer paid a 40 EUR cash-out, a note jammed at the cassette exit,
and the dashboard showed `processed`. The machine had recorded the
failure correctly. Nothing it knew ever left the box.
The root is ordering, not display: spirekeeper spawns process_settlement
the instant the payment lands, which is before the machine has begun to
dispense. The legs are paid sub-second; the dispense fails afterwards;
and the one remediation tool refuses once any leg has completed. It is
unreachable for the exact case it was built for.
ADR-005 makes payment the authorization and dispense confirmation the
capture — distribution waits for the machine's report. The report is a
report_dispense RPC (not the state doc: ADR-004's losing-writer problem),
carried through a durable outbox, sent on success and failure, adopting
lamassu's dispense_confirmed / error / error_code taxonomy and its
per-bay action log — which the machine already records in cassette_bills
and simply never ships.
Deviates from lamassu in three places it got wrong or never did: a
zero-dispensed report that arrives with an error is treated as
unverified, not as zero; a mechanical fault is its own customer screen
with evidence and is not "out of cash"; and terminal dispenser faults
latch cash-out off until a recount or an explicit operator op, because
re-initialising does not move a stuck note.
Closes the review loop with ten findings outside the ADR's decisions.
Refs #122, #27, #78
The column was renamed to fee_fraction (schema_version 13), so the main
query has been failing outright with "no such column: t.fee_percent" —
the tool only ever worked in --summary and --inventory mode. Caught
while reconciling sintra's cassettes, where listing transactions was
the obvious first step and didn't work.
Refs #40
The cassettes table is a running total, so it can be re-derived: an
absolute truth point (a recount, or an empty) plus the refills and
dispenses since. A derived count that disagrees with the stored one is
evidence of something the ledger never saw.
Reconciliation deliberately refuses to start from a refill. A refill is
a delta, and applying deltas on top of a wrong number just carries the
error forward — which is how sintra's 20-EUR bay ran 10 notes high for
weeks while its 50-EUR bay, zeroed by an `empty` before refilling,
reconciled exactly. A bay with no baseline is reported as
unreconcilable rather than silently assumed good.
Also surfaces the two things that make a count untrustworthy: the
counts-uncertain flag, and any transaction still sitting in
dispense_error / partial.
The SQL uses scalar subqueries rather than joins on purpose — joining
transaction_bills to cassettes fans out across bays, and a LEFT JOIN
whose rows are all excluded by the baseline cutoff collapses to NULL
and poisons the arithmetic downstream (the first draft read "expected:
blank" for exactly that reason).
Exits non-zero on any gap or missing baseline so it can be run as a
check after a test session.
Refs #122
`networking.wireguard.interfaces.wg0.ips` was set in hardware/douro.nix
and hardware/batm3.nix, but hardware/upboard.nix is shared by tejo and
sintra — an address there would be claimed by both machines on the same
/24, so neither got one. tejo therefore evaluated to `wg0.ips = [ ]`:
the interface comes up with no IP and the tunnel is silently dead. On a
machine with no other route in, that is how you lose a box.
Replace the two per-hardware definitions with one `wireguardIpForModel`
table in flake.nix, keyed on model like fiatCodeForModel /
upgradeWindowForModel / nfcReaderForModel, and give tejo 10.0.0.3/24 —
the address it answers on today under its factory Debian.
douro (10.0.0.4/24) and batm3 (10.0.0.5/24) evaluate unchanged; sintra
stays deliberately unlisted, since it is reachable on the LAN and has
never had a tunnel address.
The address is only half of it: the VPS maps peer pubkey to tunnel IP,
so the machine still needs /var/lib/wireguard/wg0.key carried over from
its previous install (or a fresh key added to the VPS peer list). Both
wireguard units are ConditionPathExists-guarded on that key, so a
keyless first boot is clean and the tunnel starts once it is dropped in.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The tejo still runs its factory Debian (ubilinux4, kernel 4.9) on
internal storage and has never had bitspire on it. Rather than flash
that drive, give it the run-from-USB shape douro and batm3 already use:
the stick is the system and the internal install is never touched.
- nixosConfigurations.tejo-usb — tejo-installed + usbBootModule +
usbBusHardening + usbGrubHybridModule. Evaluates identically to
sintra-usb, which shares hardware/upboard.nix.
- packages.disk-image-tejo-usb — hybrid table, GRUB, BIOS + UEFI.
NOT the efi/systemd-boot shape douro uses. The tejo is the same Aaeon
UP Board as sintra, whose firmware was found to USB-boot in Legacy/BIOS
mode; systemd-boot is UEFI-only, so a dd'd systemd-boot stick would not
be recognised as bootable at all. The hybrid image boots either path, so
it is also the safe choice if the firmware turns out to differ.
README documents both bootloader shapes and which models take which.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
disk-image-sintra-usb was a 60-line inline copy of everything
mkUsbDiskImage already does, plus the GRUB/hybrid bits the Aaeon
firmware needs — so the two implementations had already drifted: the
sintra image never picked up the `nofail` /boot that keeps a slow
ESP-USB enumeration out of emergency mode, nor the uas/autosuspend
hardening batm3.nix and douro.nix carry.
- mkUsbDiskImage takes named args with `partitionTableType` ("efi" for
systemd-boot, "hybrid" for GRUB) and `grubBiosDevice`. The ESP relabel
is layout-independent: the hybrid table creates the ESP first and
bios_grub second, so it stays partition 1 either way.
- New `usbGrubHybridModule` + `usbBusHardening` modules. The hardening is
scoped to the -usb configs rather than hardware/upboard.nix, which
sintra's eMMC install also reads — no cmdline change on a production
machine.
- `nixosConfigurations.sintra-usb` is now a named config, so a running
stick can be updated in place (nix copy + switch-to-configuration)
like batm3-usb and douro-usb.
- grub.devices is "nodev" in the config and mkForce'd to the build VM's
disk only for the image: an in-place switch on a live stick has no
/dev/vda, and GRUB's embedded core.img reads grub.cfg off the
partition, so the MBR stage needs no per-generation rewrite. Plain
definition rather than mkForce, since two mkForce lists merge into
[ "/dev/vda" "nodev" ] instead of replacing.
douro-usb and batm3-usb evaluate to byte-identical kernelParams,
blacklistedKernelModules, fileSystems, bootloader and autoUpgrade
config as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
setCassettes updated bays and dispenserInitData before re-initialising the
device, deliberately, so that "subsequent dispense calls see the new layout
even if the dispenser re-init is slow / fails". With the re-init failing
every time on douro, that meant the app kept a layout the hardware had
never taken, the operator-config consumer logged "Applied ops", and a
cassettes-state event went out to the operator advertising it. The douro
spent the afternoon reporting bay1:100x60 bay2:200x0 while the device was
still running the boot-time 100x50/200x50. A dispense in that state picks
bays by a layout the device does not share.
Roll the in-memory layout back when the re-init throws, and let the error
propagate as before. Reversing that earlier choice deliberately: a stale
but honest layout beats a fresh but fictional one when the difference is
which cassette pays out.
Also await dispenser.close() here and in cleanup(), now that close()
reports completion.
Closes#118
The Dispenser interface declared close(): void, so no caller could know when
the port was free — and both drivers returned well before it was.
puloon deferred serial.close() behind a 100 ms setTimeout and returned
immediately. A close-then-reopen caller (setCassettes -> init) therefore
raced a handle that was still open and got EAGAIN "Cannot lock port" every
single time, the overlap being the full 100 ms. Worse, had the reopen ever
won, the pending timer would then have closed the *new* handle and nulled
the field, leaving a silently dead dispenser rather than a loud error.
f56 has no timer but serialport's close() is asynchronous regardless, so it
had the same race with a much narrower window — intermittent rather than
deterministic, on sintra/tejo/gaia/batm3.
Both now resolve on serialport's close callback, claiming the handle up
front so concurrent calls can't double-close. puloon keeps its 100 ms drain
(it lets an in-flight write land) but awaits it instead of firing and
forgetting.
pcscd was enabled in hardware/batm3.nix and hardware/upboard.nix, which
cannot express "is a reader fitted": upboard.nix is shared by sintra (HID
Global OMNIKEY 5022) and tejo (nothing fitted), so tejo inherited pcscd it
has no use for, while the douro — with its own hardware file — got none and
wedged on every boot.
Make it a machine capability instead. services.bitspire.nfc.enable owns
pcscd, the two polkit rules and the wedge-recovery unit, and hands the app
a BITSPIRE_NFC_ENABLED flag so it doesn't initialise nfc-pcsc at all on a
machine with no reader. Per-model truth lives in nfcReaderForModel in
flake.nix next to fiatCodeForModel and upgradeWindowForModel, since a
shared hardware file can't answer the question. batm3 and sintra are true;
douro and tejo flip to true when readers are fitted.
The flag goes through the unit's Environment rather than
/var/lib/bitspire/.env, because .env is only written when absent — a
machine provisioned months ago would never pick up a new value.
nfc-pcsc's pcsclite binding does not fail when pcscd is not running — it
retries SCardEstablishContext in a tight loop on the calling thread, which
here is Electron's main thread. ~12k stat()s a second on
/run/pcscd/pcscd.comm, event loop dead: the window never paints, the
renderer is never reaped, and the watchdog can't fire because it needs the
same event loop. The douro sat like that for 11 hours at 80% CPU (its
CPUQuota ceiling), ignoring SIGTERM, with nothing in the journal after
[StateStore].
Check the socket exists before touching the binding. This is what makes
the "best-effort, every failure swallowed into a status callback" contract
in the module header true, and it also covers pcscd dying at runtime on a
machine that does have a reader.
The douro cutover to bitspire is being done remotely with a USB stick
and the machine's internal drive is not NixOS, so the stick has to be
the system rather than an installer medium. Give douro the same
run-from-USB shape batm3 already has.
flake.nix
- Lift the batm3-usb module and image post-processing into shared
`usbBootModule` / `mkUsbDiskImage` helpers (distinct nixos-usb/ESP-USB
labels, nofail /boot, no growPartition, autoUpgrade off, ESP relabel).
batm3-usb evaluates to the same fileSystems/upgrade config as before.
- Add `nixosConfigurations.douro-usb` and
`packages.disk-image-douro-usb` on top of douro-installed.
douro.nix
- Blacklist uas and set usbcore.autosuspend=-1, the same bus-drop
hardening batm3.nix carries, so a stick is a reliable boot medium on
the Bay Trail box.
README
- Document the -usb outputs and the flash-with-Etcher, no-installer flow.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The kiosk has launched with --disable-gpu AND
--disable-software-rasterizer since the first ISO commit (19d43c2).
Together those turn off GPU compositing and the SwiftShader fallback,
leaving Chromium to rasterise every pixel on the CPU — on Atom-class
hardware, for no reason anyone wrote down. No comment, no issue, no
commit message ever justified the pair, and /etc/bitspire/config.env has
claimed ELECTRON_DISABLE_GPU=false the whole time, contradicting the
actual command line.
Tested on sintra today. With the flags gone the GPU process is stable —
zero crashes, zero service restarts — and genuinely on hardware:
/proc/<gpu-pid>/maps shows libgallium, libGLX_mesa and dri_gbm, with no
swrast and no SwiftShader. It renders through crocus on Braswell.
Confirmed by eye on the panel, which is the part no log could answer.
Worth noting what the first attempt looked like, because it read as a
failure and was not. Restarting the unit logged "GPU process exited
unexpectedly: exit_code=15" and "has crashed 1 time(s)" — but those came
from the OUTGOING process being SIGTERMed by the restart. The incoming
one logged nothing. Checking crash counts and the gpu-process pid across
an interval, rather than grepping the last sixty lines, is what separates
the two.
DOURO KEEPS THE OLD FLAGS. Bay Trail already carries three display
workarounds — a 5.15 kernel pin for an i915 eDP regression,
i915.enable_psr=0, and vt.handoff=7 to preserve the BIOS display init —
which makes it the one machine where the original flags plausibly fixed
something real rather than being bring-up scaffolding. It is also down
pending a reflash, so it cannot be tested. Shipping an untested display
change to the most display-fragile box in the fleet, to be discovered
whenever it comes back, is not a trade worth making for one machine's
frame rate. Drop the exemption once douro is back and accelerates
cleanly.
The env override still works on every machine, douro included, so this
can be flipped either way without a rebuild.
Three claims in this file were stale, and I took all three at face value
today before checking any of them.
`dev` is not a staging branch. Every live machine runs it. batm3's
nixos-upgrade unit pulls ?ref=dev#batm3-installed daily at 04:00, which
makes "push freely to dev" actively dangerous advice — a bad commit
reaches production hardware overnight, unattended. The file said the
production ATMs ran `main` against Lightning.Pub and only sintra was on
dev. batm3 runs bitspire.service out of /var/lib/bitspire with a
VITE_SPIRE_SEED and no Lightning.Pub vars at all.
bitspire.service runs as `bitspire`, not `lamassu`. Leftover from the
rename in 46e52f6.
Added a surveyed fleet table, because two facts in it are load-bearing
for anything touching hardware. Every GPU binds crocus, including
sintra's Braswell which does so despite being Gen8. And batm3's ethernet
is DOWN — its only working network path is an Intel 7260 over WiFi — so
intel/iwlwifi firmware is what keeps that machine reachable at all.
Also recorded that batm3's nightly upgrade is currently failing. It dies
building the ATM app locally against the 60s nix.settings.timeout,
because the app is in neither aiolabs.cachix.org nor cache.nixos.org.
The timeout comment in flake.nix assumes heavy derivations are
upstream-cached; that holds for nixpkgs and not for our own app. The
machine is therefore pinned to its last successful generation and
nothing merged to dev reaches it. Same class as #98, different mechanism.
The fix is publishing atm-app-* to the cachix, not raising the ceiling.
Noted douro as down, pending a reflash and WireGuard reconnection, and
tejo as still Debian (ubilinux4, kernel 4.9) and never installed with
bitspire — a flake target rather than a deployment. Both matter when
reading "all four models build".
nixpkgs builds Mesa with 21 gallium drivers, the full Vulkan stack and the
VDPAU and VA state trackers, so one binary can serve every GPU and
cross-build case. This fleet is four Intel boards and a kiosk that never
asks for Vulkan.
mesa closure 974.5 -> 88.5 MiB
sintra 3940 -> 3201 MB tejo 3940 -> 3201 MB
batm3 3918 -> 3179 MB douro 3894 -> 3156 MB
Keep crocus, i915 and softpipe. softpipe earns its place: it is the
software rasterizer that does NOT use LLVM, so a board whose KMS driver
fails still brings up X slowly rather than dying headless somewhere
nobody can reach it.
IRIS IS OUT, and that is what makes the rest of this possible. Mesa's
meson puts with_gallium_iris in with_driver_using_cl and then
with_llvm.enable_if(with_clc, error_message : 'CLC requires LLVM')
so asking for iris drags in the OpenCL frontend and with it 540MB of
llvm-lib, and -Dllvm=disabled fails at configure. Nothing here needs iris.
The fleet was surveyed rather than assumed: sintra and tejo are Braswell
[8086:22b0], batm3 is Haswell GT2 [8086:0412], douro is Bay Trail. sintra
and batm3 were read off their running X logs and both say crocus.
That survey corrected an assumption an earlier draft of this commit was
built on. It claimed the UP Boards were the iris machines and crocus was
only for douro and batm3. sintra's Braswell is Gen8 and binds crocus
anyway. Had the list been trimmed to iris on that reasoning, which looked
like the tidier option, sintra would have dropped to software rendering or
lost its display outright.
With iris gone, LLVM goes: verified with patchelf, libgallium.so has no
libLLVM in its DT_NEEDED, not merely absent from the closure listing.
Dropping llvmpipe alone never achieved that.
THE COST IS FUTURE HARDWARE. A newer x86 board — a modern NUC, the "build
it from these parts" kiosk — will need iris, and re-adding it re-adds the
540MB. Until then such a board falls back to softpipe and renders in
software: it boots, it displays, it looks fine, and it is very slow. The
driver list carries that warning. Check `DRI driver:` in /var/log/X.0.log
on any new hardware rather than trusting the list still covers it.
Five secondary failures on the way here, each now a comment where it bites.
Two are nixpkgs' meson hook forcing auto_features=enabled, which turns
Mesa's soft driver guards into hard errors, so gallium-vdpau and gallium-va
must be disabled explicitly once the AMD and NVIDIA drivers are gone. One
is that outputs lists spirv2dxil and cross_tools unconditionally while only
d3d12, asahi and panfrost populate them; nix fails a build that leaves a
declared output unproduced, so they are created empty — and mesa sets
__structuredAttrs, so $outputs is a bash array and the obvious
`for o in $outputs` loop silently does nothing. The last two are the
asahi/panfrost cross tools and install-mesa-clc, which reference
prog_mesa_clc and so must go with LLVM.
Tested on sintra at the previous revision (with iris, 3744MB): X restarted
onto the pruned Mesa, glamor reported hardware acceleration on crocus, no
errors. This revision removes iris and LLVM and has NOT been on hardware
yet. douro and batm3 want their own nixos-rebuild test regardless; batm3
runs a different kernel and douro is the only Bay Trail.
An upgrade restarts the app and the Fujitsu dispenser runs an audible
init routine when it does. On sintra that was firing at 10:17 in the
morning, in the room, because `dates = "04:00"` is local time and every
machine inherits America/Guatemala from the shared base config.
The obvious fix is to set the system timezone per machine. This does the
narrower thing instead: systemd 252+ accepts a timezone suffix on a
calendar spec, so the timer follows Europe/Paris and its DST while the
system clock stays a fleet default nobody maintains per host. The only
other consumer of machine-local time is a technician reading the journal
at the machine, and for correlating against relay created_at stamps, UTC
is easier anyway.
The operator dashboard never needed this. It renders timestamps in the
reader's own browser locale, which is why the cassettes tab read
correctly while the journal did not.
Verified by resolving the config for all four hosts, and against
systemd-analyze on sintra itself: 04:00 Europe/Paris is 02:00 UTC in
summer, 04:00 America/Guatemala is 10:00 UTC.
A warning is in the comment because it nearly caught me: the same check
under `nix-shell -p systemd` silently computes EVERY named zone as UTC
while echoing the zone back in its normalized form. It looks accepted and
is wrong. Test on a real system.