From 44f5c0dbcf4ac4b9edde895cbcb646ee6c504b9c Mon Sep 17 00:00:00 2001 From: Patrick Mulligan Date: Tue, 4 Aug 2026 18:11:40 +0200 Subject: [PATCH] =?UTF-8?q?docs(adr):=20amend=20ADR-002=20=E2=80=94=20app?= =?UTF-8?q?=20recovery=20is=20first-line,=20SSH/NetBird=20last-resort?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The access/recovery plane (SSH/NetBird) stands and may carry recovery procedures, but it is explicitly NOT the only or first-line recovery. Add a layered, cheapest-first recovery model: (1) app auto-recovery of its own relay/Lightning connectivity, (2) an on-screen Retry for an operator at the kiosk, (3) SSH/NetBird as the last-resort remote plane for genuine app/OS failure. A public kiosk must not need remote shell access to recover from a transient/boot-before-network outage. Co-Authored-By: Claude Opus 4.8 --- .../002-remote-access-and-fleet-management.md | 30 +++++++++++++++++++ 1 file changed, 30 insertions(+) diff --git a/docs/adr/002-remote-access-and-fleet-management.md b/docs/adr/002-remote-access-and-fleet-management.md index 4931daa..a518930 100644 --- a/docs/adr/002-remote-access-and-fleet-management.md +++ b/docs/adr/002-remote-access-and-fleet-management.md @@ -104,6 +104,36 @@ The only honest way for a machine operator to exclude the SaaS operator is to ** - The machine operator's Nostr key can become the single root of trust across all three planes — SSH `authorized_keys` + VPN enrollment at install, `AddOperator`/`RevokeOperator` (#42) for delegation — so granting/revoking any party (including the SaaS operator) is one scoped, revocable capability model. - `sshd` posture should be tightened to key-only for deployed boxes (password auth is currently forced on for installed configs for first-boot provisioning; scope it to the LAN/first-boot window). Tracks with [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51). +## Amendment (2026-08-04): the access/recovery plane is not the *only* recovery + +**Status:** Accepted · **Context:** the ATM app had no way to recover its own +connectivity — a machine that booted with no internet (or whose init otherwise +failed) sat on "ATM Unavailable" until a manual `systemctl restart bitspire`, +even after the network came back. + +This ADR's SSH/NetBird recovery plane stands — it is the operator's +**app-and-OS-independent** path for the unanticipated and the broken, and may +carry recovery *procedures* (restart the service, inspect logs, re-provision). +But it is explicitly **not the first-line and not the only recovery method.** +Recovery is layered, cheapest-first: + +1. **App auto-recovery (first-line, no human).** The ATM app recovers its own + relay/Lightning connectivity when possible: the nostr client already + reconnects with backoff, and the app now re-initializes when connectivity + returns (a fresh renderer reload — HAL is preserved in the main process), + so "internet came back" self-heals without anyone touching the machine. +2. **On-screen manual retry (operator at the machine).** The maintenance + ("ATM Unavailable") screen carries a **Retry** button so a person standing + at the kiosk can force an immediate recovery attempt without shell access. +3. **SSH/NetBird (operator remote, last resort).** This plane — for when the + app *can't* self-heal or the box is genuinely broken. Unchanged by this + amendment beyond the reframing: it is the floor, not the front line. + +Rationale: the common failure (transient network / boot-before-network) must +not require remote shell access to a public kiosk. Reserve the heavyweight +recovery plane for genuine app/OS failure. Implemented on branch +`feat/connection-recovery`. + ## References - [#41](https://git.atitlan.io/aiolabs/bitspire/issues/41) — Multi-location deployment: runtime site config (the access plane's per-machine identity is provisioned here, not baked into the closure).