feat(machine): connectivity auto-recovery + on-screen Retry #82

Merged
padreug merged 3 commits from feat/connection-recovery into dev 2026-08-05 02:33:01 +00:00

3 commits

Author SHA1 Message Date
Patrick Mulligan
f8f2037100 refactor(deploy): expose batm3-usb as a named nixosConfiguration
Lift the USB-variant config (distinct fs labels, nofail /boot, no
growPartition, autoUpgrade off) out of the inline disk-image-batm3-usb
`let` into `nixosConfigurations.batm3-usb`, and build the disk-image from
that same config. Enables in-place app deploys to a running stick via
`nix copy` + `switch-to-configuration` (build the toplevel, copy the
closure, activate) — no reflash, preserving pairing + /var/lib state.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-04 18:49:59 +02:00
Patrick Mulligan
6a833d357e feat(machine): connectivity auto-recovery + on-screen Retry
A connectivity-type init failure (e.g. "No connected relays" when the box
boots before the network) landed on "ATM Unavailable" permanently: init is
one-shot and the nostr reconnect only helps after a first successful
connect, so a machine never self-healed when internet returned.

Recover by reloading the renderer, which re-runs init from a clean JS
context (no leaked actors/subscriptions) while the main process keeps HAL:
- main.ts: new `app:recover` IPC → reloadRenderer() (resets secretsConsumed).
- hal:init is now idempotent (reuse the existing instance) so the reload —
  and the pre-existing watchdog crash-reload — can't double-open serial ports.
- App.vue: when initError is a connectivity type (not the operator/
  self-clearing states unpaired/awaiting-fees/maintenance), watch for the
  `online` event (recover immediately) plus a 45s backoff safety net, and
  render a kiosk-sized Retry button for a person at the machine.

Preserves pairing + /var/lib state (renderer reload, not a process restart).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-04 18:11:40 +02:00
Patrick Mulligan
44f5c0dbcf docs(adr): amend ADR-002 — app recovery is first-line, SSH/NetBird last-resort
The access/recovery plane (SSH/NetBird) stands and may carry recovery
procedures, but it is explicitly NOT the only or first-line recovery. Add
a layered, cheapest-first recovery model: (1) app auto-recovery of its own
relay/Lightning connectivity, (2) an on-screen Retry for an operator at the
kiosk, (3) SSH/NetBird as the last-resort remote plane for genuine app/OS
failure. A public kiosk must not need remote shell access to recover from a
transient/boot-before-network outage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-04 18:11:40 +02:00