perf: cut the image 38% and turn GPU acceleration back on #112

Merged
padreug merged 9 commits from perf/gpu-acceleration into dev 2026-09-25 05:48:15 +00:00

9 commits

Author SHA1 Message Date
0347530511 perf(deploy): enable GPU acceleration by default, except on douro
The kiosk has launched with --disable-gpu AND
--disable-software-rasterizer since the first ISO commit (19d43c2).
Together those turn off GPU compositing and the SwiftShader fallback,
leaving Chromium to rasterise every pixel on the CPU — on Atom-class
hardware, for no reason anyone wrote down. No comment, no issue, no
commit message ever justified the pair, and /etc/bitspire/config.env has
claimed ELECTRON_DISABLE_GPU=false the whole time, contradicting the
actual command line.

Tested on sintra today. With the flags gone the GPU process is stable —
zero crashes, zero service restarts — and genuinely on hardware:
/proc/<gpu-pid>/maps shows libgallium, libGLX_mesa and dri_gbm, with no
swrast and no SwiftShader. It renders through crocus on Braswell.
Confirmed by eye on the panel, which is the part no log could answer.

Worth noting what the first attempt looked like, because it read as a
failure and was not. Restarting the unit logged "GPU process exited
unexpectedly: exit_code=15" and "has crashed 1 time(s)" — but those came
from the OUTGOING process being SIGTERMed by the restart. The incoming
one logged nothing. Checking crash counts and the gpu-process pid across
an interval, rather than grepping the last sixty lines, is what separates
the two.

DOURO KEEPS THE OLD FLAGS. Bay Trail already carries three display
workarounds — a 5.15 kernel pin for an i915 eDP regression,
i915.enable_psr=0, and vt.handoff=7 to preserve the BIOS display init —
which makes it the one machine where the original flags plausibly fixed
something real rather than being bring-up scaffolding. It is also down
pending a reflash, so it cannot be tested. Shipping an untested display
change to the most display-fragile box in the fleet, to be discovered
whenever it comes back, is not a trade worth making for one machine's
frame rate. Drop the exemption once douro is back and accelerates
cleanly.

The env override still works on every machine, douro included, so this
can be flipped either way without a rebuild.
2026-09-24 23:53:53 +02:00
db5f433706 Merge branch 'perf/mesa-no-llvm' into perf/gpu-acceleration 2026-09-24 23:39:39 +02:00
f565e004e5 docs: correct the fleet and branch model against the actual machines
Three claims in this file were stale, and I took all three at face value
today before checking any of them.

`dev` is not a staging branch. Every live machine runs it. batm3's
nixos-upgrade unit pulls ?ref=dev#batm3-installed daily at 04:00, which
makes "push freely to dev" actively dangerous advice — a bad commit
reaches production hardware overnight, unattended. The file said the
production ATMs ran `main` against Lightning.Pub and only sintra was on
dev. batm3 runs bitspire.service out of /var/lib/bitspire with a
VITE_SPIRE_SEED and no Lightning.Pub vars at all.

bitspire.service runs as `bitspire`, not `lamassu`. Leftover from the
rename in 46e52f6.

Added a surveyed fleet table, because two facts in it are load-bearing
for anything touching hardware. Every GPU binds crocus, including
sintra's Braswell which does so despite being Gen8. And batm3's ethernet
is DOWN — its only working network path is an Intel 7260 over WiFi — so
intel/iwlwifi firmware is what keeps that machine reachable at all.

Also recorded that batm3's nightly upgrade is currently failing. It dies
building the ATM app locally against the 60s nix.settings.timeout,
because the app is in neither aiolabs.cachix.org nor cache.nixos.org.
The timeout comment in flake.nix assumes heavy derivations are
upstream-cached; that holds for nixpkgs and not for our own app. The
machine is therefore pinned to its last successful generation and
nothing merged to dev reaches it. Same class as #98, different mechanism.
The fix is publishing atm-app-* to the cachix, not raising the ceiling.

Noted douro as down, pending a reflash and WireGuard reconnection, and
tejo as still Debian (ubilinux4, kernel 4.9) and never installed with
bitspire — a flake target rather than a deployment. Both matter when
reading "all four models build".
2026-09-24 23:29:07 +02:00
1691511aea perf(deploy): drop the gallium drivers this fleet cannot use
nixpkgs builds Mesa with 21 gallium drivers, the full Vulkan stack and the
VDPAU and VA state trackers, so one binary can serve every GPU and
cross-build case. This fleet is four Intel boards and a kiosk that never
asks for Vulkan.

  mesa closure  974.5 -> 88.5 MiB
  sintra 3940 -> 3201 MB    tejo  3940 -> 3201 MB
  batm3  3918 -> 3179 MB    douro 3894 -> 3156 MB

Keep crocus, i915 and softpipe. softpipe earns its place: it is the
software rasterizer that does NOT use LLVM, so a board whose KMS driver
fails still brings up X slowly rather than dying headless somewhere
nobody can reach it.

IRIS IS OUT, and that is what makes the rest of this possible. Mesa's
meson puts with_gallium_iris in with_driver_using_cl and then

  with_llvm.enable_if(with_clc, error_message : 'CLC requires LLVM')

so asking for iris drags in the OpenCL frontend and with it 540MB of
llvm-lib, and -Dllvm=disabled fails at configure. Nothing here needs iris.
The fleet was surveyed rather than assumed: sintra and tejo are Braswell
[8086:22b0], batm3 is Haswell GT2 [8086:0412], douro is Bay Trail. sintra
and batm3 were read off their running X logs and both say crocus.

That survey corrected an assumption an earlier draft of this commit was
built on. It claimed the UP Boards were the iris machines and crocus was
only for douro and batm3. sintra's Braswell is Gen8 and binds crocus
anyway. Had the list been trimmed to iris on that reasoning, which looked
like the tidier option, sintra would have dropped to software rendering or
lost its display outright.

With iris gone, LLVM goes: verified with patchelf, libgallium.so has no
libLLVM in its DT_NEEDED, not merely absent from the closure listing.
Dropping llvmpipe alone never achieved that.

THE COST IS FUTURE HARDWARE. A newer x86 board — a modern NUC, the "build
it from these parts" kiosk — will need iris, and re-adding it re-adds the
540MB. Until then such a board falls back to softpipe and renders in
software: it boots, it displays, it looks fine, and it is very slow. The
driver list carries that warning. Check `DRI driver:` in /var/log/X.0.log
on any new hardware rather than trusting the list still covers it.

Five secondary failures on the way here, each now a comment where it bites.
Two are nixpkgs' meson hook forcing auto_features=enabled, which turns
Mesa's soft driver guards into hard errors, so gallium-vdpau and gallium-va
must be disabled explicitly once the AMD and NVIDIA drivers are gone. One
is that outputs lists spirv2dxil and cross_tools unconditionally while only
d3d12, asahi and panfrost populate them; nix fails a build that leaves a
declared output unproduced, so they are created empty — and mesa sets
__structuredAttrs, so $outputs is a bash array and the obvious
`for o in $outputs` loop silently does nothing. The last two are the
asahi/panfrost cross tools and install-mesa-clc, which reference
prog_mesa_clc and so must go with LLVM.

Tested on sintra at the previous revision (with iris, 3744MB): X restarted
onto the pruned Mesa, glamor reported hardware acceleration on crocus, no
errors. This revision removes iris and LLVM and has NOT been on hardware
yet. douro and batm3 want their own nixos-rebuild test regardless; batm3
runs a different kernel and douro is the only Bay Trail.
2026-09-24 23:17:28 +02:00
645fd57e5b perf(deploy): prune linux-firmware to the hardware bitSpire runs on
hardware.enableRedistributableFirmware installed the entire linux-firmware
tree: 752MB compressed, 16% of the image and its single largest component.
The fleet is four fixed Intel boards. The rest is firmware for Qualcomm,
Mellanox, NVIDIA, Marvell, AMD and MediaTek parts that will never be in
one of these machines.

Keep i915 for the GPU, intel/iwlwifi, rtl_nic, rtw88, rtw89 and brcm for
whatever NIC a given box turns out to have. Turning the option off also
drops the extras it bundles (sof-firmware, libreelec-dvb, alsa-firmware,
intel2200BG, zd1211fw), none of which applies to a soundless kiosk on a
wired Intel board. The regulatory database is normally implied by that
same option so it is now requested explicitly; without it WiFi is pinned
to the most restrictive channel set.

Intel WiFi is 89MB and most of what survives. That is the deliberately
conservative half of the trade: losing the network on a fielded ATM is
not recoverable remotely, and 89MB is cheap next to a site visit.

  sintra 4604 -> 3940 MB    tejo  4604 -> 3940 MB
  batm3  4582 -> 3918 MB    douro 4544 -> 3894 MB

VERIFIED ON HARDWARE. sintra was switched to this and rebooted. It came
back with ethernet up (r8169, RTL8168g), the kiosk running, and no
firmware load failures. Before the reboot, for every module these boards
use, the firmware the kernel declares was confirmed present: i915 44 of
44, r8169 23 of 23, r8152 7 of 7. iwlwifi declares 67 and 28 are absent,
but all 28 are absent from the full upstream tree too, so the module
simply names more files than linux-firmware ships.

The reboot is what earned the intel/fw_sst_* entries. The first boot
after pruning logged

  intel_sst_acpi: Direct firmware load for intel/fw_sst_22a8.bin failed
  with error -2

the Intel Smart Sound DSP that Cherry Trail boards probe at startup. The
audio stack is already gone so nothing was functionally broken, but a
recurring error in a payment terminal's boot log is worth 420KB to
remove: an error people learn to ignore is one they will ignore when it
matters. No static check would have found this — the firmware a driver
requests at probe time is not what modinfo reports.

Two traps found while building it, both carrying comments where they bite:

The symlink loop originally ended in `[ -e ... ] && ln ...`, which makes
the loop's exit status depend on whether the LAST candidate matched. A
non-match returns 1 and set -e fails the build, so whether it worked was
a function of readdir order. It passed standalone and failed once spliced
in.

Kept directories contain symlinks pointing outside themselves: brcm's
blobs are links into cypress/. Left dangling they fail nixpkgs'
compression step, and deleting them would silently drop firmware a device
needs, so the targets get pulled in instead and anything still dangling
is a hard error.

The tree is left uncompressed because NixOS compresses each
hardware.firmware entry itself, zstd or xz depending on the kernel.
Confirmed: sintra gets -zstd, douro's 5.15 gets -xz.

system.forbiddenDependenciesRegexes rejects the upstream package by its
versioned name, so a nixpkgs bump or a stray module re-enabling the
option fails the build instead of quietly putting 750MB back.
2026-09-24 22:52:59 +02:00
425f00a71d perf(deploy): make Electron's GPU flags tunable without a rebuild
The kiosk has launched with --disable-gpu AND
--disable-software-rasterizer since the first ISO commit (19d43c2).
Together those turn off GPU compositing and the SwiftShader fallback,
which leaves Chromium rasterizing every pixel on the CPU. On a Bay Trail
Atom that is expensive, and it is very likely the largest single
contributor to a sluggish UI.

Nothing in git ever justified the pair. There is no comment, no issue and
no commit message about it; the flags arrived with the original hardware
bring-up and were carried through every refactor since. The descriptive
config at /etc/bitspire/config.env has even claimed
ELECTRON_DISABLE_GPU=false this whole time, contradicting the actual
command line. So this looks like bring-up scaffolding rather than a
diagnosed workaround, and it is worth re-testing now that the Mesa work
gives known-good crocus and iris drivers for all three GPU generations in
the fleet.

Testing it by rebuilding is the wrong loop. These are remote machines
with no one at the screen, a wrong flag is a black display, and each
attempt is a large closure copy over WireGuard. So the GPU flags move out
of ExecStart into a shell variable read from /var/lib/bitspire/.env: set
BITSPIRE_ELECTRON_GPU_FLAGS, restart the unit, look at the panel. A bad
value is one edit and a restart away from being undone.

Behaviour is unchanged by default. The variable uses ${VAR-default}, not
${VAR:-default}, so an absent line means today's flags while an
explicitly empty value means no GPU flags at all, i.e. full acceleration.
That distinction is the whole point and is why the .env template ships
the line commented out rather than set: a present-but-empty value would
silently enable the GPU on every machine that regenerates its .env.

The live ISO takes the same launcher via specialArgs, so the ISO and the
installed image cannot drift apart on this.

Closure is unchanged at 4604MB.
2026-09-24 19:13:06 +02:00
e516ab449a perf(deploy): trim systemPackages to kiosk essentials
Every entry here ships to each ATM and eats the eMMC headroom the
nightly nixos-rebuild needs, which is tight enough already that GC runs
at 03:30 purely to clear room for the 04:00 upgrade.

Out: git, at 70MB, since nixos-rebuild fetches the flake with its own
git-minimal that unit-nixos-upgrade.service keeps in the closure, so
auto-upgrade is unaffected. nodejs_22, at 94MB, which nothing runs: the
app is Electron and embeds its own node, and fund-atm references
pkgs-unstable.nodejs by store path. wget, which curl covers. And vim,
replaced by nano.

Keeping an editor at all is deliberate. Field edits to
/var/lib/bitspire/.env happen over ssh, and nano costs a few MB where
vim costs 43. minicom and screen stay for the same reason: the validator
and dispenser sit on ttyJ5 and ttyJ7, those two are how a serial fault
gets diagnosed, and they cost about 2MB between them.

With the two preceding commits the sintra-installed closure goes from
5144MB to 4604MB across 131 fewer store paths.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-24 15:27:17 +02:00
cda3f3f17e perf(deploy): force the unused audio stack off
The block enabling pipewire was commented "for transaction sounds", but
no such sounds exist: nothing under apps/machine or packages/ constructs
an Audio element or ships an audio file. It has been dead weight for as
long as it has been there.

Removing our own `enable = true` is not enough. services.xserver pulls
in NixOS's graphical-desktop module, which mkDefault-enables pipewire
exactly as it does speechd, so the stack survived the first attempt at
this. That is why the line sits beside the speechd mkForce rather than
where the old block was.

Most of PipeWire's dependency chain is shared with the GStreamer that
Electron drags in, and that stays in the closure either way, so this
frees 21MB rather than the whole stack. The remainder comes out with
gtk4/gst, which wants a launch test on the sintra first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-24 15:27:06 +02:00
73a77c82b3 perf(deploy): stop the app closure retaining its build toolchain
node-gyp leaves its scaffolding beside the addons it compiles, and
several of those files carry absolute store paths to the tools that did
the compiling: build/node_gyp_bins/python3 is an ELF copy of python3
with an RPATH into it, build/config.gypi names python3, nodejs and npm,
the .o.d files under build/Release/.deps name pcsclite's dev output, and
pnpm rewrote a few CLI helpers' shebangs to the full nodejs.

Nix scans $out for store hashes, so each of those became a runtime
reference. Every ATM was carrying python311, nodejs, npm and
pcsclite.dev -- 212MB of closure -- for files nothing reads after the
build. Only build/Release/*.node is ever loaded, through bindings and
node-gyp-build.

Drop the scaffolding, and point the stray shebangs at PATH rather than
deleting files a package might still require. All three addons survive
with their RPATHs intact: better_sqlite3.node, pcsclite.node and the
serialport prebuilds. The derivation's references are now down to bash,
pcsclite.lib and the two gcc runtime libs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-24 15:26:58 +02:00