feat: Raspberry Pi 4 (aarch64) target #110

Open
padreug wants to merge 12 commits from feat/rpi4-target into feat/rpi5-apex-hal

12 commits

Author SHA1 Message Date
76f3c2ff9d fix(hal): enable Apex escrow, use the real return bit, checksum the whole frame
Three protocol corrections from Pyramid's spec (RS_232 Rev G), all of which
this driver had guessed at because it was written clean-room without it.

**Escrow was never enabled.** BYTE 1 bit 4 is an enable, and the driver left
it clear on every poll. With it clear the acceptor does not stop at escrow,
so the host is never offered the stack-or-return decision and notes are
banked before anything has validated them. The entire FSM here is built
around that decision point, so this was not a missing nicety — the driver's
central flow could not have happened. It is now asserted on every poll.

**Return used the wrong mechanism.** The driver expressed "give the note
back" by zeroing the enable mask mid-escrow, under an in-code assumption
that "Apex has no distinct return opcode". It has one: BYTE 1 bit 6. The old
approach was flagged in a comment as needing hardware verification; the spec
settles it instead. Note the spec also distinguishes Returning (host refused
a valid note) from Rejected (acceptor judged it invalid), which is the
distinction this bit exists to express.

**The checksum range was hardcoded to the host frame.** computeChecksum
always XORed bytes 1..5, which is right for the 8-byte poll and wrong for
the 11-byte reply, where it should span 1..8. So every reply failed
validation. That was masked by the check being non-fatal "pending hardware
verification", which logged a warning and parsed anyway.

The range is now derived from the frame length, and verified against the two
reset frames the spec spells out literally with their checksums — the only
ground truth available without hardware. Those same frames are now a test.

With the range confirmed, a mismatch becomes a hard drop rather than a
warning. A corrupt frame carries a denomination field, and crediting a note
from a frame known to be damaged is the one outcome worth refusing. The raw
bytes are logged so a systematic framing error stays diagnosable.

Also records two operational facts from the spec that were not written down:
the interface is Mars/MEI GL5-compatible (hence the resemblance to the EBDS
driver), and polls must not fall more than 5s apart or the acceptor may dump
an escrowed note and stop accepting until the host resumes. Our 100ms cadence
is comfortably inside that.
2026-09-30 21:59:56 +02:00
f7b1942f38 fix(hal): correct the Apex USD channel map — it credited notes at the wrong value
The credit channel is a fixed protocol constant, not an index into whichever
notes a given unit has enabled. Pyramid's RS-232 spec (Rev G, BYTE 2 bits
3-5) fixes it:

    001=$1  010=$2  011=$5  100=$10  101=$20  110=$50  111=$100

This table omitted $2, with a comment calling it "rarely enabled". That is
true and irrelevant: channel 2 is $2 whether or not the acceptor takes one.
Dropping it shifted every larger note down a slot, so the machine would have
credited:

    $5   as $10
    $10  as $20
    $20  as $50
    $50  as $100
    $100 as nothing at all (channel 7 ran off the end and read as unmapped,
         which the driver treats as an invalid note and returns)

Every error is in the customer's favour and none of them is visible — the
value never appears on the wire, only the channel, so there is nothing to
reconcile against. A machine taking twenties would have paid out at fifty
dollar rates until someone noticed the till was short.

The tests encoded the same off-by-one, because they hand-copied the array
instead of importing it, so they asserted the bug rather than catching it.
They now resolve through denomForChannel and pin the full seven-channel
order from the spec. Reverting the table alone fails three of them.

Non-USD is unverified. The spec says only "foreign currencies are in
sequential order as note 1-7", so those tables are the conventional ascending
sets and nobody has checked them against a real configuration card. Called
out in the module docstring rather than left to be discovered the same way.
2026-09-30 21:59:56 +02:00
dd542eda88 fix(hal): seed the validator's fiat code from config, or it rejects every note
The Apex returned every bill inserted. Cause is not wiring or DIP
switches: the driver had no fiat code, so no credit channel could resolve
to a value.

ApexValidator initialises `fiatCode` to null and only assigns it in
setFiatCode(). NOTHING in this repo calls setFiatCode() — not
hal-service, not the store, nothing. It is declared on the BillValidator
interface and implemented three times, and it is dead code.

So the resolver installed in run():

    this.rs232.setDenomResolver((ch) => denomForChannel(this.fiatCode, ch))

is always called with null, and denomForChannel bails on its first line
(`if (!fiatCode || channel < 1) return null`). Every note is read,
resolves to null denomination, hits the `!bill.denomination` branch, and
is handed straight back.

id003 is unaffected because it never relies on the field: run() threads
`config.fiatCode` into the rs232 config, so the value reaches the layer
that needs it regardless. Apex and EBDS both read `this.fiatCode`
instead, so both are broken the same way. EBDS is fixed here too — it has
the identical dead-field dependency in _denominations() and would fail
identically the first time it met hardware.

Both constructors now seed from `config.fiatCode ?? config.rs232.fiatCode`,
which is what the callers have been passing all along. setFiatCode() stays
as a later override rather than the only path in.

Also makes the failure loud, because the old log line is what sent us
looking at the acceptor instead of the driver. "Bill rejected:
unsupported/unmapped channel" reads as a dataset/hardware mismatch and
gives no hint that the driver simply has no currency. It now names which
of the two causes it is, and run() logs an error up front when there is no
fiat code at all, since in that state every note is guaranteed to be
returned.
2026-09-29 21:08:24 +02:00
ea4c1f406b fix(machine): make the dispenser optional, as the validator already was
initializeHal created and initialised the dispenser unconditionally, so a
missing dispenser device threw and aborted the WHOLE of HAL init — taking
the validator down with it, even when the validator was present and
working.

The Pi bring-up hit exactly that. With a Pyramid Apex correctly wired and
enumerated on /dev/ttyValidator0:

  [ATM] Validator device: /dev/ttyValidator0
  [ATM] Dispenser device: /dev/ttyDispenser-not-fitted
  [Electron] HAL init failed: cannot open /dev/ttyDispenser-not-fitted
  [Recovery] Reloading renderer to re-attempt initialization

and round again, forever, with a perfectly good acceptor attached.

The validator has been optional since it was written — it checks the
device exists, catches init failures, and logs "running dispenser-only".
The dispenser had no equivalent. That asymmetry was the bug, not the
placeholder device path that exposed it: a cash-in-only machine is a
legitimate configuration, and the Raspberry Pi reference build is one.

Mirrors the validator's handling exactly: existence check, try/catch,
null on failure, and a log line saying what the machine will do instead
("running cash-in only"). Three call sites then need guarding —
dispenseCash returns a clear "No dispenser fitted on this machine —
cash-out unavailable" rather than dereferencing null, setCassettes still
records the layout but skips the re-init, and cleanup uses an optional
call.

This also removes the sharp edge from the rpi4/rpi5 presets added in the
previous commit. Their dispenser block points at a path that does not
exist because DispenseType has no 'none' variant and DeviceConfig
requires the field. That is still worth fixing properly with a real
'none' variant, but the machine no longer has to care.
2026-09-29 18:16:41 +02:00
64a19da924 feat(machine): add rpi4/rpi5 device presets, wire 'apex' into the app types
The Pi 4 bring-up got as far as a rendering kiosk and then failed every
init cycle with

  [App] Initialization failed: TypeError: Cannot read properties of
        undefined (reading 'validator')

MACHINE_PRESETS had entries for sintra, tejo, douro, gaia and batm3 but
none for rpi4, so getDeviceConfig() dereferenced undefined. Pairing never
happened either: the throw lands before the signer runs, so a correctly
provisioned VITE_SPIRE_SEED sat in the process environment while
bunker_binding stayed at zero. rpi5 had the same hole and would have hit
it the moment anyone booted that target.

This is the third instance today of the same shape: the Pi targets reuse
the shared runtime, and the shared runtime carries per-model tables that
nobody added the Pi to. pcscd was the first (a busy-loop, no window), the
Electron cassette presets the second (silently seeded nothing).

Also widens the validator union from 'id003' | 'ebds' to include 'apex'
in both DeviceConfig and HalConfig. packages/hal has had
ValidatorType = 'id003' | 'ebds' | 'apex' since the Pyramid Apex driver
landed with the Pi 5 work, but these app-side unions were never widened,
so no machine could be configured to use the driver at all. The presets
below are the first thing that needed it, which is presumably why nobody
noticed.

The dispenser block in both presets is a placeholder, not a claim.
DispenserType has no 'none' variant and DeviceConfig requires the field,
so it points at a path that does not exist and carries no cassettes.
These boards are cash-in only until real hardware lands. A 'none'
dispenser variant would be the honest fix and is worth doing separately.

Validator device defaults to /dev/ttyValidator0, the FTDI udev symlink
raspberry-pi-4.nix creates. Swap to ttyValidator1 (CP210x) or
ttyValidator2 (CH340) to match the adapter fitted; `ls -l /dev/ttyValidator*`
after plugging it in says which appeared.
2026-09-25 22:56:57 +02:00
2572a14c61 fix(rpi4): enable pcscd — without it the kiosk hangs before drawing
This is the white screen. Not graphics, not the bundle, not pairing.

The app constructs @pokusew/pcsclite at startup. That calls
SCardEstablishContext(), which calls SCardCheckDaemonAvailability(), which
— finding no pcscd — BUSY-LOOPS in fstatat64 at ~92% CPU rather than
returning an error. It runs on Electron's main thread before the
BrowserWindow is created, so no window is ever made and the panel stays
white.

It is a hang, not a crash, which is why it presented so badly. Nothing
throws. Nothing is logged after "[StateStore] Initialized database". The
process looks healthy: systemd reports the service active, Electron is
running, and there is even a gpu-process. But a setInterval registered
before startup never fires once in 32 seconds, the main thread sits in
state R, and the remote debugger reports zero page targets.

V8's own tooling cannot see it either, because the thread never yields to
the inspector: Debugger.pause returns nothing and Profiler.stop times out.
A native backtrace was the only thing that worked:

  #0 fstatat64                     libc
  #1 SCardCheckDaemonAvailability  libpcsclite
  #2 SCardEstablishContext         libpcsclite
  #3 PCSCLite::PCSCLite()          pcsclite.node
  #4 PCSCLite::New(...)

Ruled out along the way, each by direct test on the machine: graphics (it
fails identically with the GPU fully disabled), Electron on aarch64 (a
minimal app renders fine), the Vue bundle (loading the real index.html
from a minimal main process mounts the app and reaches "[ATM] State
machine initialized"), the preload script, the CSP, kiosk and fullscreen
window options, better-sqlite3, /dev/shm, memory, page size, X
authorisation, and isDev.

upboard.nix and batm3.nix both enable pcscd for their real readers, which
is why no x86 machine has ever hit this. This module did not, and that was
the entire difference. pcscd with no reader attached simply idles, so
enabling it costs nothing.

Worth noting for the wider fleet: any future board that omits pcscd
inherits this, and it presents as a blank screen with a healthy-looking
service. The robust fix is for the app to not block its main thread on a
card-reader handshake at all — the NFC path is already documented as
best-effort — but that is an app change and this unblocks the hardware.
2026-09-25 22:09:36 +02:00
8917b8b967 fix(rpi4): raise CMA to 256M, document the firmware-partition step
Two findings from the first Pi 4 (actually a CM4) bring-up, chased from a
white screen to hardware-accelerated X.

CMA. The vc4 display pipeline allocates its framebuffer from the contiguous
memory area, and the default reservation here is 32MiB with ~11MiB free. The
attached panel is 3840x1080, whose framebuffer is ~16.6MB before double
buffering, so X picked the mode and then died:

  Output HDMI-1 using initial mode 3840x1080 +0+0
  (EE) AddScreen/ScreenInit failed for driver 0

nixos-hardware's fkms-3d injected a CMA overlay alongside its display one;
cma=256M replaces that half. This part is declarative, since kernelParams
reach the extlinux APPEND line.

The display itself is NOT declarative, and that is the uncomfortable part.
This board boots the FIRMWARE's vendor DTB, not the DTBs NixOS builds: the
live device tree carries __symbols__ and mainline's do not, and U-Boot found
no FDTDIR match for compatible "raspberrypi,4-compute-module" so it passed
the firmware's DTB straight through. hardware.deviceTree.overlays therefore
cannot reach the running tree at all, which is also why the earlier fkms-3d
removal fixed a build error without fixing the display.

In the vendor DTB every display node ships disabled, so config.txt needs
`dtoverlay=vc4-kms-v3d,noaudio` and /boot/firmware/overlays/ has to be
populated from raspberrypifw. Three traps in that one line, each of which
failed silently:

  - the NixOS sd-image writes the DTBs to the firmware partition but NOT the
    overlays, so the directory ships empty and dtoverlay= does nothing
  - copying only vc4-kms-v3d.dtbo is insufficient; the firmware remaps that
    to vc4-kms-v3d-pi4.dtbo on this board, so the whole directory goes
  - without noaudio, vc4_hdmi cannot register its PCM component, returns
    -517 (EPROBE_DEFER) forever, and the DRM device never registers, so X
    finds no card. We deleted the audio stack anyway.

Result on the machine: card1 is vc4-drm with HDMI-A-1 connected, card2 is
v3d, and X reports

  glamor X acceleration enabled on V3D 4.2.14.0

against swrast before. The manual steps are written into the module so the
next person does not rediscover them from a blank screen, but they are lost
on a reflash and belong in the image builder. Follow-up.
2026-09-25 18:02:31 +02:00
11cc1b88d6 fix(rpi4): drop fkms-3d, mainline does full KMS without an overlay
With the mainline kernel from the previous commit, the device-tree overlay
step fails outright:

  Applying overlay rpi4-cma-overlay
  Applying overlay rpi4-vc4-fkms-v3d-overlay
  libfdt.FdtException: pylibfdt error -1: FDT_ERR_NOTFOUND

nixos-hardware's fkms-3d applies two overlays that patch nodes present in
the Raspberry Pi VENDOR kernel's DTBs and absent from mainline's. Predicted
when the kernel changed; this is it arriving.

It is also the wrong thing to want. "fkms" is FIRMWARE KMS, the older
arrangement where the VideoCore firmware owns the display and Linux drives
it at arm's length. Mainline does full KMS, and mainline's own
bcm2711-rpi-4-b.dtb already describes the hardware: it carries
brcm,bcm2711-vc5 and brcm,2711-v3d nodes, confirmed by decompiling the DTB
with dtc. The vc4 and v3d drivers bind to those directly, no overlay
involved.

So the overlay was not providing capability, it was translating for a
kernel we no longer use.

hardware.deviceTree.overlays is now empty, so there is nothing left for the
overlay builder to fail on. Verified by evaluating the config.

No replacement needed for videoDrivers either. fkms-3d used to set it as a
side effect, but the shared configuration.nix already declares modesetting,
which is correct for full KMS and is what the x86 machines use. Setting it
again here just produced ["modesetting" "modesetting"].

Still unproven on hardware: whether X comes up on vc4 rather than falling
back to a framebuffer. That is the next thing to read out of
/var/log/X.0.log once the machine boots, and it is now a minutes-long
iteration rather than a kernel compile per attempt.
2026-09-25 15:45:56 +02:00
b2bf2ffe4b fix(rpi4): use the mainline kernel, not the uncached Pi vendor one
nixos-hardware's raspberry-pi/4 module mkDefaults boot.kernelPackages to
the Raspberry Pi vendor kernel (linux-rpi, via common/kernel.nix). That
kernel is in no binary cache: Hydra does not build nixos-hardware's
overlay kernels, and it is not in aiolabs.cachix.org either. So every Pi
compiles a kernel from source, on an SD card, and recompiles on every
bump.

Found the hard way during the first Pi 4 bring-up, which spent hours on

  building linux-rpi-6.18.39-stable_20260724 (buildPhase):
    CC [M]  fs/overlayfs/inode.o

before anyone looked closely enough to notice it was not the app. I had
told the operator only the app would build, having checked that Electron
and the Pi firmware were cached and never checked the kernel.

Mainline aarch64 kernels are cached, and mainline demonstrably boots a
Pi 4 — it is what the stock NixOS aarch64 SD image runs, which is how this
machine was bootstrapped in the first place. The vendor kernel's
Pi-specific patches buy nothing this kiosk needs: display, USB serial and
WiFi are all mainline, and vc4/v3d KMS has been mainline for years.

  before: linux-rpi-6.18.39-stable_20260724   not cached, hours to build
  after:  linux-6.12.90                       cached, downloads

Verified the config still evaluates, which also clears nixos-hardware's
assertion that the kernel be at least 6.1.

Watch the graphics path. fkms-3d is a nixos-hardware overlay built around
the vendor kernel's firmware-KMS route; the mainline equivalent is full KMS
(vc4-kms-v3d). If X lands on the framebuffer with Electron rendering in
software, that overlay is where to look, not this kernel choice.

rpi5 is left alone deliberately and still carries the vendor kernel, so it
will pay the same cost whenever someone first builds it. Mainline Pi 5
support is younger than Pi 4's and the RP1 southbridge needed vendor
patches for longer, so that one wants its own check rather than the same
change applied on faith.
2026-09-25 15:27:12 +02:00
6568618811 fix(nix): keep only the serialport prebuild this system can load
The first ever aarch64 build of the app died here, on a Raspberry Pi 4:

  auto-patchelf could not satisfy dependency liblog.so wanted by
    node_modules/@serialport/bindings-cpp/prebuilds/android-arm64/node.napi.armv8.node
  auto-patchelf could not satisfy dependency libc++_shared.so wanted by
    (the same file)

@serialport/bindings-cpp ships prebuilds for every platform it supports:
android-arm, android-arm64, darwin, linux-arm, linux-arm64, linux-x64 in
both glibc and musl, win32-ia32 and win32-x64. installPhase copied the
whole directory.

On x86_64 that was harmless because autoPatchelf skips ELF files whose
architecture does not match the host, so the Android and ARM prebuilds
were never touched. On aarch64 the android-arm64 prebuild IS the host
architecture, so autoPatchelf picks it up and goes looking for Android's
liblog.so and libc++_shared.so, which NixOS does not have. The failure was
invisible until someone built for a second architecture.

Keep only the prebuild the target can load, selected from
stdenv.hostPlatform: linux-arm64 on aarch64, linux-x64 elsewhere, glibc
rather than musl.

Pruning rather than adding those two libraries to
autoPatchelfIgnoreMissingDeps, which would have been the one-line fix.
Teaching autoPatchelf to tolerate a binary we never load, for a platform
we do not target, leaves the foreign prebuilds in the closure and leaves
the same trap set for the next architecture. The existing
"libc.musl-x86_64.so.1" entry in that list is this same problem solved the
other way; it is now redundant, and is left in place only because this
commit is unblocking a machine mid-build and is not the moment to find out
whether something else depended on it.

Verified on x86_64: the app still builds, ships linux-x64/node.napi.glibc.node
alone where it previously carried nine platforms, and comes to 25M.
2026-09-25 12:30:25 +02:00
91c6994dd4 feat(deploy): add aarch64 Raspberry Pi 4 target
The Pi 4 twin of the Pi 5 build: `rpi4-installed` (in-place rebuild target),
`rpi4-image`, and `packages.aarch64-linux.{sd-image-rpi4,atm-app-rpi4}`, all
through the board-keyed machinery of the previous commit. The shared runtime
is untouched; only the board pair is new.

deploy/nixos/hardware/raspberry-pi-4.nix mirrors raspberry-pi-5.nix line for
line except where the boards differ:
- KMS for the kiosk display is an opt-in on the Pi 4
  (`hardware.raspberry-pi."4".fkms-3d`), which also injects the CMA + vc4
  device-tree overlays; the Pi 5 gets it by default. Without it X falls back
  to the framebuffer and Electron renders in software.
- fkms-3d sets videoDrivers itself, so the module doesn't.
Everything else — extlinux boot, console pinned to tty0 so the GPIO UART is
free for a validator, no-suspend, the ttyValidator{0,1,2} udev symlinks — is
identical by design.

Evaluation-verified only: rpi4-installed/rpi4-image instantiate, and against
rpi5 they differ solely in the expected places (bcm2711 device tree, the two
fkms overlays, the rpiVersion=4 kernel, no clk-rp1 in initrd, machine model
in the env seed). Not yet booted on hardware; the doc says so.

docs/raspberry-pi-setup.md covers both boards — build, flash, first boot +
provisioning via the spire seed, in-place updates, peripherals — since #87
shipped the Pi 5 without one. It replaces a never-committed Pi 4 sketch
(parked on wip/rpi4-sketch) whose flake wiring didn't evaluate and whose
provisioning section predated the pairing seed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013A6683cCHnQxFUosx1krY4
2026-09-24 21:44:25 +02:00
029bb09745 refactor(deploy): key the Pi build machinery by board
piBaseModules hardcoded the nixos-hardware raspberry-pi-5 module and our
raspberry-pi-5.nix glue, so mkPiInstalled/mkPiImage could only ever produce
a Pi 5. Everything else in the Pi runtime is board-agnostic.

Introduce `piBoards`, keyed by machine model — the parameter already threaded
through both builders — pairing each board's nixos-hardware module with its
glue file, and have piBaseModules look the pair up. Same modules in the same
order for rpi5, so its evaluated configuration is unchanged (compared on 16
app-independent facets: kernel, params, initrd modules, loader, device tree,
video drivers, udev, filesystems, swap, nix settings, sleep targets, service
exec/memory, env seed, sshd). No new board yet — that's the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013A6683cCHnQxFUosx1krY4
2026-09-24 21:44:25 +02:00