docs(adr): amend ADR-002 — app recovery is first-line, SSH/NetBird last-resort

The access/recovery plane (SSH/NetBird) stands and may carry recovery
procedures, but it is explicitly NOT the only or first-line recovery. Add
a layered, cheapest-first recovery model: (1) app auto-recovery of its own
relay/Lightning connectivity, (2) an on-screen Retry for an operator at the
kiosk, (3) SSH/NetBird as the last-resort remote plane for genuine app/OS
failure. A public kiosk must not need remote shell access to recover from a
transient/boot-before-network outage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Patrick Mulligan 2026-08-04 18:11:40 +02:00
commit 44f5c0dbcf

View file

@ -104,6 +104,36 @@ The only honest way for a machine operator to exclude the SaaS operator is to **
- The machine operator's Nostr key can become the single root of trust across all three planes — SSH `authorized_keys` + VPN enrollment at install, `AddOperator`/`RevokeOperator` (#42) for delegation — so granting/revoking any party (including the SaaS operator) is one scoped, revocable capability model.
- `sshd` posture should be tightened to key-only for deployed boxes (password auth is currently forced on for installed configs for first-boot provisioning; scope it to the LAN/first-boot window). Tracks with [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51).
## Amendment (2026-08-04): the access/recovery plane is not the *only* recovery
**Status:** Accepted · **Context:** the ATM app had no way to recover its own
connectivity — a machine that booted with no internet (or whose init otherwise
failed) sat on "ATM Unavailable" until a manual `systemctl restart bitspire`,
even after the network came back.
This ADR's SSH/NetBird recovery plane stands — it is the operator's
**app-and-OS-independent** path for the unanticipated and the broken, and may
carry recovery *procedures* (restart the service, inspect logs, re-provision).
But it is explicitly **not the first-line and not the only recovery method.**
Recovery is layered, cheapest-first:
1. **App auto-recovery (first-line, no human).** The ATM app recovers its own
relay/Lightning connectivity when possible: the nostr client already
reconnects with backoff, and the app now re-initializes when connectivity
returns (a fresh renderer reload — HAL is preserved in the main process),
so "internet came back" self-heals without anyone touching the machine.
2. **On-screen manual retry (operator at the machine).** The maintenance
("ATM Unavailable") screen carries a **Retry** button so a person standing
at the kiosk can force an immediate recovery attempt without shell access.
3. **SSH/NetBird (operator remote, last resort).** This plane — for when the
app *can't* self-heal or the box is genuinely broken. Unchanged by this
amendment beyond the reframing: it is the floor, not the front line.
Rationale: the common failure (transient network / boot-before-network) must
not require remote shell access to a public kiosk. Reserve the heavyweight
recovery plane for genuine app/OS failure. Implemented on branch
`feat/connection-recovery`.
## References
- [#41](https://git.atitlan.io/aiolabs/bitspire/issues/41) — Multi-location deployment: runtime site config (the access plane's per-machine identity is provisioned here, not baked into the closure).