From f565e004e5e456e482302c94d2bc504fd33cddc7 Mon Sep 17 00:00:00 2001 From: Padreug Date: Thu, 24 Sep 2026 23:29:07 +0200 Subject: [PATCH] docs: correct the fleet and branch model against the actual machines MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three claims in this file were stale, and I took all three at face value today before checking any of them. `dev` is not a staging branch. Every live machine runs it. batm3's nixos-upgrade unit pulls ?ref=dev#batm3-installed daily at 04:00, which makes "push freely to dev" actively dangerous advice — a bad commit reaches production hardware overnight, unattended. The file said the production ATMs ran `main` against Lightning.Pub and only sintra was on dev. batm3 runs bitspire.service out of /var/lib/bitspire with a VITE_SPIRE_SEED and no Lightning.Pub vars at all. bitspire.service runs as `bitspire`, not `lamassu`. Leftover from the rename in 46e52f6. Added a surveyed fleet table, because two facts in it are load-bearing for anything touching hardware. Every GPU binds crocus, including sintra's Braswell which does so despite being Gen8. And batm3's ethernet is DOWN — its only working network path is an Intel 7260 over WiFi — so intel/iwlwifi firmware is what keeps that machine reachable at all. Also recorded that batm3's nightly upgrade is currently failing. It dies building the ATM app locally against the 60s nix.settings.timeout, because the app is in neither aiolabs.cachix.org nor cache.nixos.org. The timeout comment in flake.nix assumes heavy derivations are upstream-cached; that holds for nixpkgs and not for our own app. The machine is therefore pinned to its last successful generation and nothing merged to dev reaches it. Same class as #98, different mechanism. The fix is publishing atm-app-* to the cachix, not raising the ceiling. Noted douro as down, pending a reflash and WireGuard reconnection, and tejo as still Debian (ubilinux4, kernel 4.9) and never installed with bitspire — a flake target rather than a deployment. Both matter when reading "all four models build". --- CLAUDE.md | 53 +++++++++++++++++++++++++++++++++++++++++++++++++---- 1 file changed, 49 insertions(+), 4 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 878be52..2cc2db7 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -4,7 +4,7 @@ Guidance for Claude Code when working in this repo. Read this before touching co ## Project Overview -**bitSpire** is a Nostr-native Lightning ATM. Production ATMs (`batm3`, `douro`) currently run from `main` against Lightning.Pub; the `dev` branch — which is what this file describes — has been migrated to **LNbits over the nostr-native-transport**. +**bitSpire** is a Nostr-native Lightning ATM running **LNbits over the nostr-native-transport**. The `dev` branch, which this file describes, is what the machines run. Core principles: @@ -25,8 +25,53 @@ bitSpire is an independent project under AGPL-3.0 and is not affiliated with Lam ## Branch model -- `main` — production. Lightning.Pub backend. The two production ATMs auto-pull from here daily at 04:00 (`flake.nix:152-160`). **DO NOT** push to `main` casually — a wrong commit gets baked into prod ATMs the next morning. -- `dev` — staging. LNbits backend. The Sintra dev unit auto-pulls from here (`?ref=dev` pin on this branch's `flake.nix`). Push freely; tag `pre-bitspire-cutover` is the rollback target if the migration ever needs to be reverted on prod. +- `dev` — **what every live machine runs.** Not a staging branch any more. Verified + 2026-09-24 on batm3, whose `nixos-upgrade` unit pulls + `git+ssh://…/bitspire.git?ref=dev#batm3-installed` daily at 04:00. "Push freely + to dev" is no longer safe advice: a bad commit reaches production hardware the + next morning, unattended. +- `main` — Lightning.Pub era, historical. Tag `pre-bitspire-cutover` is the + rollback target if the migration ever has to be reverted. + +> This section previously said the production ATMs ran `main` against +> Lightning.Pub and that only Sintra was on `dev`. That was stale and it was +> repeatedly taken at face value. Check the machine, not this file, before +> relying on which stack a given box runs: `systemctl cat nixos-upgrade` gives +> the branch, `/var/lib/bitspire` vs `/var/lib/lamassu-atm` gives the era. + +### Fleet state (surveyed 2026-09-24) + +| Machine | Reachable | Stack | GPU | Notes | +|---|---|---|---|---| +| `sintra` | LAN `192.168.0.252` | dev / LNbits | Braswell `8086:22b0` → crocus | dev unit; ethernet `r8169` | +| `batm3` | wg `10.0.0.5` | dev / LNbits | Haswell GT2 `8086:0412` → crocus | **networks over WiFi**, `iwlwifi` 7260; ethernet down | +| `douro` | **down** | — | Bay Trail (Gen7) | needs reflashing with the current image and reconnecting to WireGuard | +| `tejo` | wg `10.0.0.3` | **Debian** (`ubilinux4`, kernel 4.9) | Braswell `8086:22b0` | never had bitspire installed; a flake target, not a deployment | + +Two consequences worth holding onto. Every GPU in the fleet binds **crocus**, not +iris — sintra's Braswell does so despite being Gen8. And batm3's only working +network path is Intel WiFi, so `intel/iwlwifi` firmware is load-bearing there; +trimming it would strand the machine with no way back in. + +### batm3's nightly upgrade is currently FAILING + +Confirmed 2026-09-24. The run dies at: + +``` +04:03:26 building '…-bitspire-atm-app-0.1.0.drv'... +04:04:28 error: timed out after 60 seconds +``` + +The ATM app is built in-house and is **not in `aiolabs.cachix.org` or +`cache.nixos.org`**, so batm3 has to build it locally, and `nix.settings.timeout += 60` in `flake.nix` kills it. The comment there assumes heavy derivations are +"effectively cache-only … upstream-cached", which is true of nixpkgs and false of +our own app. + +So the machine is pinned to whatever generation last succeeded, and nothing +merged to `dev` reaches it. This is the same class of silent-updater failure as +#98, in a new form. The fix is pushing `atm-app-*` to the aiolabs cachix as part +of releasing, not raising the timeout — a 60s ceiling on ATM hardware is correct. ## Architecture @@ -220,7 +265,7 @@ UP Board enumerates its eMMC controller via ACPI, not PCI. `upboard.nix` force-l - The renderer logs prefix every line with a tag: `[Lightning]`, `[ATM]`, `[ATM Service]`, `[LNURL Session]`, `[CLINK]`, `[StateStore]`. `journalctl -u bitspire | grep '\['` is your friend. - **Never pass an object as a console argument in the renderer.** Electron's console bridge stringifies each argument, so `console.log('msg:', { a, b })` reaches the journal as `msg: [object Object]` and every field is lost. Interpolate instead. Cost a debugging session on 2026-09-23, when a cassette publish that had worked looked like it had done nothing. -- `bitspire.service` runs as the `lamassu` user; `/var/lib/bitspire` is its `dataDir` (ReadWritePaths). DB lives at `/var/lib/bitspire/state.db` (we previously had `/var/lib/lamassu-atm` — that path is gone on dev, see commit `9c455d6`). +- `bitspire.service` runs as the `bitspire` user (verified on sintra 2026-09-24; this line used to say `lamassu`, left over from the rename in `46e52f6`); `/var/lib/bitspire` is its `dataDir` (ReadWritePaths). DB lives at `/var/lib/bitspire/state.db` (we previously had `/var/lib/lamassu-atm` — that path is gone on dev, see commit `9c455d6`). - The `lightning.lightningPub` field on `LightningServices` is a `LightningBackend` *adapter*, not a `LightningPubClient`. Don't try to call LP-only methods on it. ## Related documentation