Multi-location deployment: extract per-machine config to runtime site files #41

Open
opened 2026-06-13 22:02:58 +00:00 by padreug · 0 comments
Owner

Migrated from aiolabs/lamassu-next#41 — opened by @padreug on 2026-04-07.\n\n## Problem

Each ATM hardware model (douro, batm3, etc.) currently has a single flake target like douro-installed or batm3-installed. Per-machine values are baked into the system derivation at build time, which means:

  1. Two machines of the same model can't use the same closure if they differ in any location-specific way (WireGuard IP, hostname, USB serials, fiat code, etc.).
  2. Cachix pushes scale linearly with the number of machines. With 20 douros across different locations, you'd need 20 separate builds and pushes — each producing a unique top-level system closure.
  3. New deployments require flake edits and a CI cycle, even when only the location differs (no code changes).

This is fine for the current scale (1 BATM3, 1 douro) but becomes a major friction point as the fleet grows.

What's currently location-specific (baked into the flake)

Item Where Notes
WireGuard IP deploy/nixos/hardware/batm3.nix:141 (10.0.0.5/24) Two BATM3s would collide
Fiat code flake.nix:60 (always USD for batm3) A Mexico BATM3 needs MXN
USB serial symlinks batm3.nix:117-119 Each F56/MEI/NFC has different serial #s
Touchscreen calibration matrix batm3.nix:136 Each panel calibrates slightly differently
Hostname deploy/nixos/configuration.nix All machines named lamassu-atm
WireGuard private key currently outside flake Per-machine secret

What's already location-aware (good)

These already live outside the flake on each machine (in /var/lib/lamassu-atm/.env or DB):

  • ATM Nostr keypair
  • Lightning.Pub admin token, URL, relay
  • WiFi credentials (wifi.conf)
  • Cassette inventory (SQLite DB)
  • Maintenance mode
  • Operator pubkeys
  • Fee rates (VITE_CASH_IN_FEE, VITE_CASH_OUT_FEE)

Proposed solution: runtime site config

Make the system derivation identical for all machines of the same hardware model. Per-machine values come from a small site file (e.g., /var/lib/lamassu-atm/site.json) that lives outside the repo and is provisioned per machine.

Example layout

// /var/lib/lamassu-atm/site.json (per-machine, not in git)
{
  "hostname": "atm-antigua-1",
  "fiatCode": "USD",
  "wireguard": {
    "ip": "10.0.0.5/24",
    "privateKeyFile": "/var/lib/lamassu-atm/wg-private.key"
  },
  "serials": {
    "f56": "DDDLb103Y23",
    "mei": "A9YW78OC",
    "nfc": "A9ZF8ELY"
  },
  "touchscreen": {
    "calibrationMatrix": "0 -1.268 1.147 -1.224 0 1.118 0 0 1"
  }
}

How it gets applied

NixOS activation script reads site.json and:

  • Sets hostname via hostname command
  • Generates udev rules for USB serial symlinks into /etc/udev/rules.d/ (writable)
  • Generates a NetworkManager/systemd-networkd WireGuard connection from the IP + key
  • Installs a systemd unit that re-runs xinput set-prop with the matrix on graphical-session start
  • Sets VITE_LAMASSU_FIAT_CODE in /var/lib/lamassu-atm/.env if not already set

The flake-level config becomes:

  • One douro-installed and one batm3-installed target — shared by all locations
  • Hardware files (batm3.nix, douro.nix) only contain things that truly differ by hardware: kernel modules, firmware, drivers, partition layout, etc.

Cachix impact

Setup Pushes needed for 20 douros
Current 20 (one per unique closure)
Per-location flake targets + push script 20 (still 20 unique closures, just automated)
Runtime site config 1 (all 20 share one closure)

Provisioning flow (after refactor)

For a new deployment:

# 1. Flash the disk image (same image for all locations)
sudo dd if=disk-image-batm3.img of=/dev/sda

# 2. First boot: provision per-machine secrets and site config
ssh lamassu@new-atm
sudo mkdir -p /var/lib/lamassu-atm
sudo tee /var/lib/lamassu-atm/site.json <<EOF
{ "hostname": "atm-quetzaltenango-1", ... }
EOF
sudo cat > /var/lib/lamassu-atm/wifi.conf <<EOF
SSID=...
PSK=...
EOF
# Run existing provision-atm.sh for Nostr keys, LP token, etc.

# 3. Reboot — activation script applies site.json on first activation
sudo reboot

No flake change. No rebuild. No cachix push.

Migration tasks

  • Define site.json schema and document it (docs/site-config.md)
  • Write activation script that reads site.json and applies values
  • Move WireGuard from declarative NixOS to runtime-configured (NetworkManager or systemd-networkd template unit)
  • Move udev USB serial rules to /etc/udev/rules.d/ generated at activation
  • Move touchscreen calibration matrix to systemd unit reading from site config
  • Move hostname to activation script (instead of networking.hostName)
  • Move fiat code from mkAtmApp { fiatCode = "USD" } to runtime .env (already partially supported)
  • Update provision-atm.sh to write site.json interactively
  • Update CLAUDE.md and deploy docs
  • Test on existing BATM3 and douro to ensure no regressions
  • Once stable, remove per-location flake targets and consolidate to one per model

Open questions

  1. WireGuard private key — should this stay in /var/lib/lamassu-atm/wg-private.key (file) or live inside site.json? File is safer (can be chmod 600), but site.json is one less thing to provision.
  2. Cachix push automation — even after this refactor, we should have a deploy/push-all.sh that builds all current model closures and pushes them. Should this run in CI on every main push?
  3. Site config validation — should we validate site.json schema at activation time and fail early with a clear error if malformed? (Yes, probably with a small Nix-eval-time check or a JSON schema validator.)
  4. Fallback when site.json is missing — should activation fail (forces operator to provision it) or fall back to safe defaults (hostname = lamassu-atm, no WG)? Failing loud is probably safer to prevent silent misconfiguration.

Priority

Medium. Not blocking current operations, but worth tackling before deploying the 3rd machine of any model — at that point the per-target approach starts feeling painful, and refactoring becomes harder once 20 machines are already in production with location-specific flake targets.

> _Migrated from [aiolabs/lamassu-next#41](https://git.atitlan.io/aiolabs/lamassu-next/issues/41) — opened by @padreug on 2026-04-07._\n\n## Problem Each ATM hardware model (`douro`, `batm3`, etc.) currently has a single flake target like `douro-installed` or `batm3-installed`. Per-machine values are **baked into the system derivation at build time**, which means: 1. **Two machines of the same model can't use the same closure** if they differ in any location-specific way (WireGuard IP, hostname, USB serials, fiat code, etc.). 2. **Cachix pushes scale linearly with the number of machines.** With 20 douros across different locations, you'd need 20 separate builds and pushes — each producing a unique top-level system closure. 3. **New deployments require flake edits and a CI cycle**, even when only the location differs (no code changes). This is fine for the current scale (1 BATM3, 1 douro) but becomes a major friction point as the fleet grows. ## What's currently location-specific (baked into the flake) | Item | Where | Notes | |------|-------|-------| | WireGuard IP | `deploy/nixos/hardware/batm3.nix:141` (`10.0.0.5/24`) | Two BATM3s would collide | | Fiat code | `flake.nix:60` (always USD for batm3) | A Mexico BATM3 needs MXN | | USB serial symlinks | `batm3.nix:117-119` | Each F56/MEI/NFC has different serial #s | | Touchscreen calibration matrix | `batm3.nix:136` | Each panel calibrates slightly differently | | Hostname | `deploy/nixos/configuration.nix` | All machines named `lamassu-atm` | | WireGuard private key | currently outside flake | Per-machine secret | ## What's already location-aware (good) These already live outside the flake on each machine (in `/var/lib/lamassu-atm/.env` or DB): - ATM Nostr keypair - Lightning.Pub admin token, URL, relay - WiFi credentials (`wifi.conf`) - Cassette inventory (SQLite DB) - Maintenance mode - Operator pubkeys - Fee rates (`VITE_CASH_IN_FEE`, `VITE_CASH_OUT_FEE`) ## Proposed solution: runtime site config Make the system derivation **identical for all machines of the same hardware model**. Per-machine values come from a small site file (e.g., `/var/lib/lamassu-atm/site.json`) that lives outside the repo and is provisioned per machine. ### Example layout ```json // /var/lib/lamassu-atm/site.json (per-machine, not in git) { "hostname": "atm-antigua-1", "fiatCode": "USD", "wireguard": { "ip": "10.0.0.5/24", "privateKeyFile": "/var/lib/lamassu-atm/wg-private.key" }, "serials": { "f56": "DDDLb103Y23", "mei": "A9YW78OC", "nfc": "A9ZF8ELY" }, "touchscreen": { "calibrationMatrix": "0 -1.268 1.147 -1.224 0 1.118 0 0 1" } } ``` ### How it gets applied NixOS activation script reads `site.json` and: - Sets hostname via `hostname` command - Generates udev rules for USB serial symlinks into `/etc/udev/rules.d/` (writable) - Generates a NetworkManager/systemd-networkd WireGuard connection from the IP + key - Installs a systemd unit that re-runs `xinput set-prop` with the matrix on graphical-session start - Sets `VITE_LAMASSU_FIAT_CODE` in `/var/lib/lamassu-atm/.env` if not already set The flake-level config becomes: - One `douro-installed` and one `batm3-installed` target — **shared by all locations** - Hardware files (`batm3.nix`, `douro.nix`) only contain things that **truly differ by hardware**: kernel modules, firmware, drivers, partition layout, etc. ### Cachix impact | Setup | Pushes needed for 20 douros | |-------|-----------------------------| | Current | 20 (one per unique closure) | | Per-location flake targets + push script | 20 (still 20 unique closures, just automated) | | Runtime site config | **1** (all 20 share one closure) | ## Provisioning flow (after refactor) For a new deployment: ```bash # 1. Flash the disk image (same image for all locations) sudo dd if=disk-image-batm3.img of=/dev/sda # 2. First boot: provision per-machine secrets and site config ssh lamassu@new-atm sudo mkdir -p /var/lib/lamassu-atm sudo tee /var/lib/lamassu-atm/site.json <<EOF { "hostname": "atm-quetzaltenango-1", ... } EOF sudo cat > /var/lib/lamassu-atm/wifi.conf <<EOF SSID=... PSK=... EOF # Run existing provision-atm.sh for Nostr keys, LP token, etc. # 3. Reboot — activation script applies site.json on first activation sudo reboot ``` No flake change. No rebuild. No cachix push. ## Migration tasks - [ ] Define `site.json` schema and document it (`docs/site-config.md`) - [ ] Write activation script that reads `site.json` and applies values - [ ] Move WireGuard from declarative NixOS to runtime-configured (NetworkManager or systemd-networkd template unit) - [ ] Move udev USB serial rules to `/etc/udev/rules.d/` generated at activation - [ ] Move touchscreen calibration matrix to systemd unit reading from site config - [ ] Move hostname to activation script (instead of `networking.hostName`) - [ ] Move fiat code from `mkAtmApp { fiatCode = "USD" }` to runtime `.env` (already partially supported) - [ ] Update `provision-atm.sh` to write `site.json` interactively - [ ] Update CLAUDE.md and deploy docs - [ ] Test on existing BATM3 and douro to ensure no regressions - [ ] Once stable, remove per-location flake targets and consolidate to one per model ## Open questions 1. **WireGuard private key** — should this stay in `/var/lib/lamassu-atm/wg-private.key` (file) or live inside `site.json`? File is safer (can be `chmod 600`), but `site.json` is one less thing to provision. 2. **Cachix push automation** — even after this refactor, we should have a `deploy/push-all.sh` that builds all current model closures and pushes them. Should this run in CI on every `main` push? 3. **Site config validation** — should we validate `site.json` schema at activation time and fail early with a clear error if malformed? (Yes, probably with a small Nix-eval-time check or a JSON schema validator.) 4. **Fallback when `site.json` is missing** — should activation fail (forces operator to provision it) or fall back to safe defaults (hostname = `lamassu-atm`, no WG)? Failing loud is probably safer to prevent silent misconfiguration. ## Priority **Medium.** Not blocking current operations, but worth tackling before deploying the 3rd machine of any model — at that point the per-target approach starts feeling painful, and refactoring becomes harder once 20 machines are already in production with location-specific flake targets. ## Related - [Auto-close issues in commit messages](https://git.atitlan.io/aiolabs/lamassu-next/wiki) (existing convention) - BATM3 hardware deployment notes
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/bitspire#41
No description provided.