perf: cut the image 38% and turn GPU acceleration back on #112

Merged
padreug merged 9 commits from perf/gpu-acceleration into dev 2026-09-25 05:48:15 +00:00
Owner

Takes sintra-installed from 5144 MB to 3201 MB, a 38% cut. Every figure was measured by building, and the whole stack has been through a cold boot on sintra.

Time-sensitive. Sintra is running this now, deployed by hand. Its nixos-upgrade fires Fri 04:00 CST, which is 12:00 CEST, and until this is on dev that run will rebuild from the old dev and revert all of it.

Merge #111 first. It fixes the upgrade-window timezone and applies cleanly on top of this. An earlier revision of this PR carried a competing system-timezone change; that has been dropped, see the comment below.

Sizes

Machine Before After
sintra 5144 MB 3201 MB
tejo 5144 MB 3201 MB
batm3 4582 MB 3179 MB
douro 4544 MB 3156 MB

What's in it

Stop the app closure retaining its build toolchain (212 MB). node-gyp scaffolding kept python311, nodejs, npm and pcsclite.dev alive as runtime references for files nothing reads after the build.

Force the audio stack off (21 MB). services.xserver re-enables PipeWire through NixOS's graphical-desktop module, so removing our own enable = true did nothing; it needs mkForce. The app has never played a sound.

Trim systemPackages (117 MB). git, vim, nodejs_22, wget out; nano in, so there's still an editor for field edits.

Prune linux-firmware (664 MB), 752 MB down to 113 MB. Verified per module: i915 44 of 44 present, r8169 23 of 23, r8152 7 of 7. A reboot then caught a missing Intel Smart Sound blob, since firmware requested at probe time is not what modinfo reports — 420 KB, added back.

Drop the gallium drivers this fleet can't use (739 MB), Mesa 975 MB down to 88.5 MB, LLVM gone entirely. Only possible because no machine here uses iris: sintra and tejo are Braswell, batm3 is Haswell, douro is Bay Trail, and all bind crocus. sintra's and batm3's X logs say so directly.

GPU acceleration on by default. The kiosk has run with --disable-gpu --disable-software-rasterizer since the first ISO commit with nothing in git justifying it. Removing them gives a stable GPU process on real hardware. Douro is exempt — it already carries three display workarounds and is down pending a reflash.

CLAUDE.md corrections. It claimed production ran main against Lightning.Pub and only sintra was on dev. Every live machine runs dev. Also records that batm3 networks over WiFi with ethernet down, so intel/iwlwifi firmware is what keeps it reachable.

Verified on sintra

Cold boot on the final generation. Network up, crocus with glamor acceleration, no firmware errors, no X errors, GPU process stable with zero crashes and hardware GL mapped. Screen confirmed by eye.

Not verified

Only sintra. Douro is down pending a reflash and is the only Bay Trail; batm3 runs a different kernel and networks over WiFi. Both want their own nixos-rebuild test.

Separate problem this surfaced

batm3's nightly upgrade is failing and has been. It times out building the ATM app locally against the 60 s nix.settings.timeout, because the app is in neither aiolabs.cachix.org nor cache.nixos.org. Nothing merged to dev reaches that machine until atm-app-* is published to the cachix. Raising the timeout would be the wrong fix.

Takes `sintra-installed` from **5144 MB to 3201 MB**, a 38% cut. Every figure was measured by building, and the whole stack has been through a cold boot on sintra. **Time-sensitive.** Sintra is running this now, deployed by hand. Its `nixos-upgrade` fires Fri 04:00 CST, which is 12:00 CEST, and until this is on `dev` that run will rebuild from the old `dev` and revert all of it. **Merge #111 first.** It fixes the upgrade-window timezone and applies cleanly on top of this. An earlier revision of this PR carried a competing system-timezone change; that has been dropped, see the comment below. ### Sizes | Machine | Before | After | |---|---|---| | sintra | 5144 MB | 3201 MB | | tejo | 5144 MB | 3201 MB | | batm3 | 4582 MB | 3179 MB | | douro | 4544 MB | 3156 MB | ### What's in it **Stop the app closure retaining its build toolchain** (212 MB). node-gyp scaffolding kept python311, nodejs, npm and pcsclite.dev alive as runtime references for files nothing reads after the build. **Force the audio stack off** (21 MB). `services.xserver` re-enables PipeWire through NixOS's `graphical-desktop` module, so removing our own `enable = true` did nothing; it needs `mkForce`. The app has never played a sound. **Trim systemPackages** (117 MB). git, vim, nodejs_22, wget out; nano in, so there's still an editor for field edits. **Prune linux-firmware** (664 MB), 752 MB down to 113 MB. Verified per module: i915 44 of 44 present, r8169 23 of 23, r8152 7 of 7. A reboot then caught a missing Intel Smart Sound blob, since firmware requested at probe time is not what modinfo reports — 420 KB, added back. **Drop the gallium drivers this fleet can't use** (739 MB), Mesa 975 MB down to 88.5 MB, LLVM gone entirely. Only possible because no machine here uses iris: sintra and tejo are Braswell, batm3 is Haswell, douro is Bay Trail, and all bind **crocus**. sintra's and batm3's X logs say so directly. **GPU acceleration on by default**. The kiosk has run with `--disable-gpu --disable-software-rasterizer` since the first ISO commit with nothing in git justifying it. Removing them gives a stable GPU process on real hardware. Douro is exempt — it already carries three display workarounds and is down pending a reflash. **CLAUDE.md corrections.** It claimed production ran `main` against Lightning.Pub and only sintra was on `dev`. Every live machine runs `dev`. Also records that batm3 networks over WiFi with ethernet down, so `intel/iwlwifi` firmware is what keeps it reachable. ### Verified on sintra Cold boot on the final generation. Network up, crocus with glamor acceleration, no firmware errors, no X errors, GPU process stable with zero crashes and hardware GL mapped. Screen confirmed by eye. ### Not verified Only sintra. Douro is down pending a reflash and is the only Bay Trail; batm3 runs a different kernel and networks over WiFi. Both want their own `nixos-rebuild test`. ### Separate problem this surfaced **batm3's nightly upgrade is failing** and has been. It times out building the ATM app locally against the 60 s `nix.settings.timeout`, because the app is in neither `aiolabs.cachix.org` nor `cache.nixos.org`. Nothing merged to `dev` reaches that machine until `atm-app-*` is published to the cachix. Raising the timeout would be the wrong fix.
Every machine inherited America/Guatemala from the shared base config,
which is right for douro and tejo and wrong for sintra in France. It read
six hours behind, so its journal timestamps had to be converted by hand
against anything on the relay.

The clock is not cosmetic here. system.autoUpgrade's `dates = "04:00"` is
local time, so the zone decides when the nightly rebuild restarts the app
and the Fujitsu dispenser runs its audible init routine. On sintra that
was firing at 10:17 in the morning, in the room, rather than at 4am.

timeZoneForModel mirrors fiatCodeForModel and lists only the exceptions.
The base config keeps America/Guatemala as the fleet default, now under
mkDefault so a per-model value wins without mkForce. Verified by eval:
sintra resolves to Europe/Paris, douro, tejo and batm3 are unchanged.

batm3 is deliberately left alone. It is USD and in the field, and I do
not know where.
node-gyp leaves its scaffolding beside the addons it compiles, and
several of those files carry absolute store paths to the tools that did
the compiling: build/node_gyp_bins/python3 is an ELF copy of python3
with an RPATH into it, build/config.gypi names python3, nodejs and npm,
the .o.d files under build/Release/.deps name pcsclite's dev output, and
pnpm rewrote a few CLI helpers' shebangs to the full nodejs.

Nix scans $out for store hashes, so each of those became a runtime
reference. Every ATM was carrying python311, nodejs, npm and
pcsclite.dev -- 212MB of closure -- for files nothing reads after the
build. Only build/Release/*.node is ever loaded, through bindings and
node-gyp-build.

Drop the scaffolding, and point the stray shebangs at PATH rather than
deleting files a package might still require. All three addons survive
with their RPATHs intact: better_sqlite3.node, pcsclite.node and the
serialport prebuilds. The derivation's references are now down to bash,
pcsclite.lib and the two gcc runtime libs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The block enabling pipewire was commented "for transaction sounds", but
no such sounds exist: nothing under apps/machine or packages/ constructs
an Audio element or ships an audio file. It has been dead weight for as
long as it has been there.

Removing our own `enable = true` is not enough. services.xserver pulls
in NixOS's graphical-desktop module, which mkDefault-enables pipewire
exactly as it does speechd, so the stack survived the first attempt at
this. That is why the line sits beside the speechd mkForce rather than
where the old block was.

Most of PipeWire's dependency chain is shared with the GStreamer that
Electron drags in, and that stays in the closure either way, so this
frees 21MB rather than the whole stack. The remainder comes out with
gtk4/gst, which wants a launch test on the sintra first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every entry here ships to each ATM and eats the eMMC headroom the
nightly nixos-rebuild needs, which is tight enough already that GC runs
at 03:30 purely to clear room for the 04:00 upgrade.

Out: git, at 70MB, since nixos-rebuild fetches the flake with its own
git-minimal that unit-nixos-upgrade.service keeps in the closure, so
auto-upgrade is unaffected. nodejs_22, at 94MB, which nothing runs: the
app is Electron and embeds its own node, and fund-atm references
pkgs-unstable.nodejs by store path. wget, which curl covers. And vim,
replaced by nano.

Keeping an editor at all is deliberate. Field edits to
/var/lib/bitspire/.env happen over ssh, and nano costs a few MB where
vim costs 43. minicom and screen stay for the same reason: the validator
and dispenser sit on ttyJ5 and ttyJ7, those two are how a serial fault
gets diagnosed, and they cost about 2MB between them.

With the two preceding commits the sintra-installed closure goes from
5144MB to 4604MB across 131 fewer store paths.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The kiosk has launched with --disable-gpu AND
--disable-software-rasterizer since the first ISO commit (19d43c2).
Together those turn off GPU compositing and the SwiftShader fallback,
which leaves Chromium rasterizing every pixel on the CPU. On a Bay Trail
Atom that is expensive, and it is very likely the largest single
contributor to a sluggish UI.

Nothing in git ever justified the pair. There is no comment, no issue and
no commit message about it; the flags arrived with the original hardware
bring-up and were carried through every refactor since. The descriptive
config at /etc/bitspire/config.env has even claimed
ELECTRON_DISABLE_GPU=false this whole time, contradicting the actual
command line. So this looks like bring-up scaffolding rather than a
diagnosed workaround, and it is worth re-testing now that the Mesa work
gives known-good crocus and iris drivers for all three GPU generations in
the fleet.

Testing it by rebuilding is the wrong loop. These are remote machines
with no one at the screen, a wrong flag is a black display, and each
attempt is a large closure copy over WireGuard. So the GPU flags move out
of ExecStart into a shell variable read from /var/lib/bitspire/.env: set
BITSPIRE_ELECTRON_GPU_FLAGS, restart the unit, look at the panel. A bad
value is one edit and a restart away from being undone.

Behaviour is unchanged by default. The variable uses ${VAR-default}, not
${VAR:-default}, so an absent line means today's flags while an
explicitly empty value means no GPU flags at all, i.e. full acceleration.
That distinction is the whole point and is why the .env template ships
the line commented out rather than set: a present-but-empty value would
silently enable the GPU on every machine that regenerates its .env.

The live ISO takes the same launcher via specialArgs, so the ISO and the
installed image cannot drift apart on this.

Closure is unchanged at 4604MB.
hardware.enableRedistributableFirmware installed the entire linux-firmware
tree: 752MB compressed, 16% of the image and its single largest component.
The fleet is four fixed Intel boards. The rest is firmware for Qualcomm,
Mellanox, NVIDIA, Marvell, AMD and MediaTek parts that will never be in
one of these machines.

Keep i915 for the GPU, intel/iwlwifi, rtl_nic, rtw88, rtw89 and brcm for
whatever NIC a given box turns out to have. Turning the option off also
drops the extras it bundles (sof-firmware, libreelec-dvb, alsa-firmware,
intel2200BG, zd1211fw), none of which applies to a soundless kiosk on a
wired Intel board. The regulatory database is normally implied by that
same option so it is now requested explicitly; without it WiFi is pinned
to the most restrictive channel set.

Intel WiFi is 89MB and most of what survives. That is the deliberately
conservative half of the trade: losing the network on a fielded ATM is
not recoverable remotely, and 89MB is cheap next to a site visit.

  sintra 4604 -> 3940 MB    tejo  4604 -> 3940 MB
  batm3  4582 -> 3918 MB    douro 4544 -> 3894 MB

VERIFIED ON HARDWARE. sintra was switched to this and rebooted. It came
back with ethernet up (r8169, RTL8168g), the kiosk running, and no
firmware load failures. Before the reboot, for every module these boards
use, the firmware the kernel declares was confirmed present: i915 44 of
44, r8169 23 of 23, r8152 7 of 7. iwlwifi declares 67 and 28 are absent,
but all 28 are absent from the full upstream tree too, so the module
simply names more files than linux-firmware ships.

The reboot is what earned the intel/fw_sst_* entries. The first boot
after pruning logged

  intel_sst_acpi: Direct firmware load for intel/fw_sst_22a8.bin failed
  with error -2

the Intel Smart Sound DSP that Cherry Trail boards probe at startup. The
audio stack is already gone so nothing was functionally broken, but a
recurring error in a payment terminal's boot log is worth 420KB to
remove: an error people learn to ignore is one they will ignore when it
matters. No static check would have found this — the firmware a driver
requests at probe time is not what modinfo reports.

Two traps found while building it, both carrying comments where they bite:

The symlink loop originally ended in `[ -e ... ] && ln ...`, which makes
the loop's exit status depend on whether the LAST candidate matched. A
non-match returns 1 and set -e fails the build, so whether it worked was
a function of readdir order. It passed standalone and failed once spliced
in.

Kept directories contain symlinks pointing outside themselves: brcm's
blobs are links into cypress/. Left dangling they fail nixpkgs'
compression step, and deleting them would silently drop firmware a device
needs, so the targets get pulled in instead and anything still dangling
is a hard error.

The tree is left uncompressed because NixOS compresses each
hardware.firmware entry itself, zstd or xz depending on the kernel.
Confirmed: sintra gets -zstd, douro's 5.15 gets -xz.

system.forbiddenDependenciesRegexes rejects the upstream package by its
versioned name, so a nixpkgs bump or a stray module re-enabling the
option fails the build instead of quietly putting 750MB back.
nixpkgs builds Mesa with 21 gallium drivers, the full Vulkan stack and the
VDPAU and VA state trackers, so one binary can serve every GPU and
cross-build case. This fleet is four Intel boards and a kiosk that never
asks for Vulkan.

  mesa closure  974.5 -> 88.5 MiB
  sintra 3940 -> 3201 MB    tejo  3940 -> 3201 MB
  batm3  3918 -> 3179 MB    douro 3894 -> 3156 MB

Keep crocus, i915 and softpipe. softpipe earns its place: it is the
software rasterizer that does NOT use LLVM, so a board whose KMS driver
fails still brings up X slowly rather than dying headless somewhere
nobody can reach it.

IRIS IS OUT, and that is what makes the rest of this possible. Mesa's
meson puts with_gallium_iris in with_driver_using_cl and then

  with_llvm.enable_if(with_clc, error_message : 'CLC requires LLVM')

so asking for iris drags in the OpenCL frontend and with it 540MB of
llvm-lib, and -Dllvm=disabled fails at configure. Nothing here needs iris.
The fleet was surveyed rather than assumed: sintra and tejo are Braswell
[8086:22b0], batm3 is Haswell GT2 [8086:0412], douro is Bay Trail. sintra
and batm3 were read off their running X logs and both say crocus.

That survey corrected an assumption an earlier draft of this commit was
built on. It claimed the UP Boards were the iris machines and crocus was
only for douro and batm3. sintra's Braswell is Gen8 and binds crocus
anyway. Had the list been trimmed to iris on that reasoning, which looked
like the tidier option, sintra would have dropped to software rendering or
lost its display outright.

With iris gone, LLVM goes: verified with patchelf, libgallium.so has no
libLLVM in its DT_NEEDED, not merely absent from the closure listing.
Dropping llvmpipe alone never achieved that.

THE COST IS FUTURE HARDWARE. A newer x86 board — a modern NUC, the "build
it from these parts" kiosk — will need iris, and re-adding it re-adds the
540MB. Until then such a board falls back to softpipe and renders in
software: it boots, it displays, it looks fine, and it is very slow. The
driver list carries that warning. Check `DRI driver:` in /var/log/X.0.log
on any new hardware rather than trusting the list still covers it.

Five secondary failures on the way here, each now a comment where it bites.
Two are nixpkgs' meson hook forcing auto_features=enabled, which turns
Mesa's soft driver guards into hard errors, so gallium-vdpau and gallium-va
must be disabled explicitly once the AMD and NVIDIA drivers are gone. One
is that outputs lists spirv2dxil and cross_tools unconditionally while only
d3d12, asahi and panfrost populate them; nix fails a build that leaves a
declared output unproduced, so they are created empty — and mesa sets
__structuredAttrs, so $outputs is a bash array and the obvious
`for o in $outputs` loop silently does nothing. The last two are the
asahi/panfrost cross tools and install-mesa-clc, which reference
prog_mesa_clc and so must go with LLVM.

Tested on sintra at the previous revision (with iris, 3744MB): X restarted
onto the pruned Mesa, glamor reported hardware acceleration on crocus, no
errors. This revision removes iris and LLVM and has NOT been on hardware
yet. douro and batm3 want their own nixos-rebuild test regardless; batm3
runs a different kernel and douro is the only Bay Trail.
Three claims in this file were stale, and I took all three at face value
today before checking any of them.

`dev` is not a staging branch. Every live machine runs it. batm3's
nixos-upgrade unit pulls ?ref=dev#batm3-installed daily at 04:00, which
makes "push freely to dev" actively dangerous advice — a bad commit
reaches production hardware overnight, unattended. The file said the
production ATMs ran `main` against Lightning.Pub and only sintra was on
dev. batm3 runs bitspire.service out of /var/lib/bitspire with a
VITE_SPIRE_SEED and no Lightning.Pub vars at all.

bitspire.service runs as `bitspire`, not `lamassu`. Leftover from the
rename in 46e52f6.

Added a surveyed fleet table, because two facts in it are load-bearing
for anything touching hardware. Every GPU binds crocus, including
sintra's Braswell which does so despite being Gen8. And batm3's ethernet
is DOWN — its only working network path is an Intel 7260 over WiFi — so
intel/iwlwifi firmware is what keeps that machine reachable at all.

Also recorded that batm3's nightly upgrade is currently failing. It dies
building the ATM app locally against the 60s nix.settings.timeout,
because the app is in neither aiolabs.cachix.org nor cache.nixos.org.
The timeout comment in flake.nix assumes heavy derivations are
upstream-cached; that holds for nixpkgs and not for our own app. The
machine is therefore pinned to its last successful generation and
nothing merged to dev reaches it. Same class as #98, different mechanism.
The fix is publishing atm-app-* to the cachix, not raising the ceiling.

Noted douro as down, pending a reflash and WireGuard reconnection, and
tejo as still Debian (ubilinux4, kernel 4.9) and never installed with
bitspire — a flake target rather than a deployment. Both matter when
reading "all four models build".
The kiosk has launched with --disable-gpu AND
--disable-software-rasterizer since the first ISO commit (19d43c2).
Together those turn off GPU compositing and the SwiftShader fallback,
leaving Chromium to rasterise every pixel on the CPU — on Atom-class
hardware, for no reason anyone wrote down. No comment, no issue, no
commit message ever justified the pair, and /etc/bitspire/config.env has
claimed ELECTRON_DISABLE_GPU=false the whole time, contradicting the
actual command line.

Tested on sintra today. With the flags gone the GPU process is stable —
zero crashes, zero service restarts — and genuinely on hardware:
/proc/<gpu-pid>/maps shows libgallium, libGLX_mesa and dri_gbm, with no
swrast and no SwiftShader. It renders through crocus on Braswell.
Confirmed by eye on the panel, which is the part no log could answer.

Worth noting what the first attempt looked like, because it read as a
failure and was not. Restarting the unit logged "GPU process exited
unexpectedly: exit_code=15" and "has crashed 1 time(s)" — but those came
from the OUTGOING process being SIGTERMed by the restart. The incoming
one logged nothing. Checking crash counts and the gpu-process pid across
an interval, rather than grepping the last sixty lines, is what separates
the two.

DOURO KEEPS THE OLD FLAGS. Bay Trail already carries three display
workarounds — a 5.15 kernel pin for an i915 eDP regression,
i915.enable_psr=0, and vt.handoff=7 to preserve the BIOS display init —
which makes it the one machine where the original flags plausibly fixed
something real rather than being bring-up scaffolding. It is also down
pending a reflash, so it cannot be tested. Shipping an untested display
change to the most display-fragile box in the fleet, to be discovered
whenever it comes back, is not a trade worth making for one machine's
frame rate. Drop the exemption once douro is back and accelerates
cleanly.

The env override still works on every machine, douro included, so this
can be flipped either way without a rebuild.
padreug force-pushed perf/gpu-acceleration from 8899f4ef48 to 0347530511 2026-09-25 05:32:14 +00:00 Compare
Author
Owner

Force-pushed to drop the fix/sintra-timezone merge. I had merged it in before noticing #111, which takes the narrower approach to the same bug and explicitly supersedes the system-timezone route. Keeping both would have conflicted in flake.nix.

So this PR is now purely the image-size and GPU work. #111 owns the timezone fix, and it applies cleanly on top of this — verified.

Rebuilt after the drop: sintra 3201 MB, douro 3156 MB, both with tz=America/Guatemala and dates=04:00 untouched.

Merge #111 first, then this one.

One correction to #111's deploy note

It expects the timer's last-run stamp to be newer than the new window's most recent occurrence, so Persistent=true won't cause a catch-up fire. On sintra right now that does not hold:

Persistent=yes
LastTriggerUSec=Thu 2026-09-24 04:00:02 CST

04:00 Europe/Paris is 20:00 CST the previous day. The most recent occurrence of that is Thu 20:00 CST, which is newer than the Thu 04:00 CST last-run stamp — so systemd should see a missed window and fire on switch.

That is harmless provided both PRs are on dev before sintra next switches, since the catch-up would then rebuild what is already deployed. It is only dangerous in the gap where #111 is merged and this one is not: a catch-up run would revert the image work and install the new window at the same time.

Force-pushed to drop the `fix/sintra-timezone` merge. I had merged it in before noticing #111, which takes the narrower approach to the same bug and explicitly supersedes the system-timezone route. Keeping both would have conflicted in `flake.nix`. So this PR is now purely the image-size and GPU work. #111 owns the timezone fix, and it applies cleanly on top of this — verified. Rebuilt after the drop: sintra 3201 MB, douro 3156 MB, both with `tz=America/Guatemala` and `dates=04:00` untouched. Merge #111 first, then this one. ### One correction to #111's deploy note It expects the timer's last-run stamp to be newer than the new window's most recent occurrence, so `Persistent=true` won't cause a catch-up fire. On sintra right now that does not hold: ``` Persistent=yes LastTriggerUSec=Thu 2026-09-24 04:00:02 CST ``` `04:00 Europe/Paris` is `20:00 CST` the previous day. The most recent occurrence of that is Thu 20:00 CST, which is *newer* than the Thu 04:00 CST last-run stamp — so systemd should see a missed window and fire on switch. That is harmless provided both PRs are on `dev` before sintra next switches, since the catch-up would then rebuild what is already deployed. It is only dangerous in the gap where #111 is merged and this one is not: a catch-up run would revert the image work and install the new window at the same time.
padreug changed title from perf: cut the image 38%, turn GPU acceleration back on, fix sintra's timezone to perf: cut the image 38% and turn GPU acceleration back on 2026-09-25 05:46:01 +00:00
padreug deleted branch perf/gpu-acceleration 2026-09-25 05:48:15 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/bitspire!112
No description provided.