diff --git a/docs/adr/002-remote-access-and-fleet-management.md b/docs/adr/002-remote-access-and-fleet-management.md index 4931daa..a518930 100644 --- a/docs/adr/002-remote-access-and-fleet-management.md +++ b/docs/adr/002-remote-access-and-fleet-management.md @@ -104,6 +104,36 @@ The only honest way for a machine operator to exclude the SaaS operator is to ** - The machine operator's Nostr key can become the single root of trust across all three planes — SSH `authorized_keys` + VPN enrollment at install, `AddOperator`/`RevokeOperator` (#42) for delegation — so granting/revoking any party (including the SaaS operator) is one scoped, revocable capability model. - `sshd` posture should be tightened to key-only for deployed boxes (password auth is currently forced on for installed configs for first-boot provisioning; scope it to the LAN/first-boot window). Tracks with [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51). +## Amendment (2026-08-04): the access/recovery plane is not the *only* recovery + +**Status:** Accepted · **Context:** the ATM app had no way to recover its own +connectivity — a machine that booted with no internet (or whose init otherwise +failed) sat on "ATM Unavailable" until a manual `systemctl restart bitspire`, +even after the network came back. + +This ADR's SSH/NetBird recovery plane stands — it is the operator's +**app-and-OS-independent** path for the unanticipated and the broken, and may +carry recovery *procedures* (restart the service, inspect logs, re-provision). +But it is explicitly **not the first-line and not the only recovery method.** +Recovery is layered, cheapest-first: + +1. **App auto-recovery (first-line, no human).** The ATM app recovers its own + relay/Lightning connectivity when possible: the nostr client already + reconnects with backoff, and the app now re-initializes when connectivity + returns (a fresh renderer reload — HAL is preserved in the main process), + so "internet came back" self-heals without anyone touching the machine. +2. **On-screen manual retry (operator at the machine).** The maintenance + ("ATM Unavailable") screen carries a **Retry** button so a person standing + at the kiosk can force an immediate recovery attempt without shell access. +3. **SSH/NetBird (operator remote, last resort).** This plane — for when the + app *can't* self-heal or the box is genuinely broken. Unchanged by this + amendment beyond the reframing: it is the floor, not the front line. + +Rationale: the common failure (transient network / boot-before-network) must +not require remote shell access to a public kiosk. Reserve the heavyweight +recovery plane for genuine app/OS failure. Implemented on branch +`feat/connection-recovery`. + ## References - [#41](https://git.atitlan.io/aiolabs/bitspire/issues/41) — Multi-location deployment: runtime site config (the access plane's per-machine identity is provisioned here, not baked into the closure).