bitspire/docs/adr/002-remote-access-and-fleet-management.md
Patrick Mulligan 44f5c0dbcf docs(adr): amend ADR-002 — app recovery is first-line, SSH/NetBird last-resort
The access/recovery plane (SSH/NetBird) stands and may carry recovery
procedures, but it is explicitly NOT the only or first-line recovery. Add
a layered, cheapest-first recovery model: (1) app auto-recovery of its own
relay/Lightning connectivity, (2) an on-screen Retry for an operator at the
kiosk, (3) SSH/NetBird as the last-resort remote plane for genuine app/OS
failure. A public kiosk must not need remote shell access to recover from a
transient/boot-before-network outage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-04 18:11:40 +02:00

144 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# ADR-002: Remote Access & Fleet Management — Three Planes, Operator-Owned Access via NetBird
**Status:** Accepted
**Date:** 2026-06-14
**Context:** Multi-operator bitSpire fleet — separating the payment, control, and recovery planes by who owns them.
## Decision
1. **Separate three planes by trust owner**, and never conflate them:
- **Payment plane** — ATM ↔ LNbits over the nostr-native-transport. Owned by the **SaaS operator**. Implies *no* machine access.
- **Fleet control plane** — routine ops/telemetry/enrollment over Nostr (see [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42)). Authorized by the **machine operator's** key.
- **Access / recovery plane** — SSH for the unanticipated and the broken. Owned by the **machine operator**.
2. **The machine operator's own recovery access is provisioned at install and is app-independent.** Their SSH key and their VPN/NetBird enrollment are established when the machine is set up, so they can always reach a box even when the bitSpire app or OS is broken. Access for *anyone else* is runtime-granted, scoped, and revocable — never the owner's own path.
3. **Adopt NetBird as the standard access/recovery plane**, chosen for fleet scale and a **self-hostable, fully FOSS control plane**. The platform may provide a default NetBird setup as a convenience; a machine operator who does not wish to trust whoever runs that control plane **disables it and provisions their own access plane** (self-hosted NetBird, or their own WireGuard hub).
4. **We will NOT build a "revoke SaaS-operator access" toggle in the operator dashboard.** It is a false promise of security: the SaaS operator runs LNbits (and, in the default deployment, the NetBird control plane), so a toggle they ultimately control cannot protect a machine operator against them. The honest boundary is **exclusion-by-ownership, not exclusion-by-toggle** — an operator who wants to exclude the SaaS operator takes ownership of the access plane.
## Context
### The players
A deployed bitSpire machine sits between two distinct principals:
- **SaaS operator** — runs the LNbits instance and provides the Lightning backend as a service.
- **Machine operator** — owns the physical ATM(s) and is identified by a Nostr key (the operator pubkey in the [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42) allow-list).
These are different parties with different interests. A machine operator will want to SSH to their own machine for support and recovery, and **may or may not want to grant the SaaS operator that same access.**
### Why the SaaS operator needs zero box access by design
The whole nostr-native architecture (no admin tokens on the kiosk, no inbound network surface, payment over Nostr) means the SaaS operator can deliver the full service **without ever touching the machine**. So "the machine operator may refuse the SaaS operator access" is not a constraint to engineer around — it is the **default that costs nothing**. SaaS-operator box access is a *support convenience*, never a service requirement. The natural posture is therefore **default-deny for the SaaS operator**.
### Why SSH can't be replaced by the Nostr control plane
The Nostr control plane (#42) is a fixed menu of structured, capability-scoped commands dispatched by a handler *inside the app*. It is excellent for routine, auditable, fleet-wide ops on **healthy** machines, and strictly better than SSH for those (signed, scoped, logged, fan-out). But:
- It can only do what a handler was written for; incidents are by definition unanticipated.
- The listener lives in the app, so it dies exactly when the app dies — the case you most need recovery for.
SSH (arbitrary, interactive, app-independent) is therefore irreducible as the **recovery plane**. The two are complements, not substitutes.
### Why the recovery path must be app-independent
The whole point of a recovery path is to survive the failure of the thing it recovers. So it must not be gated by the bitSpire app, nor by a Nostr command the app dispatches. The kernel/agent that carries the tunnel and `sshd` must come up at boot independent of the app. (`allowedTCPPorts = []` already means `sshd` is unreachable except across the tunnel — the VPN handshake is the outer lock, the SSH key the inner one.)
### Why the single shared hub had to change
The pre-existing design used one WireGuard hub (`170.75.161.21`) run by platform infra. Whoever runs that hub has a standing network path to every enrolled box — i.e. the SaaS operator having access to machines they don't own. Multi-tenancy requires the access plane to be **per-operator or policy-isolated**, rooted in the machine operator, not the platform.
## Options Considered
### Access-plane mechanism
#### Option A: Always-up minimal WireGuard hub
**Pros:** After boot, zero userspace dependency — the kernel holds the tunnel, nothing can crash it short of a kernel/networking fault; smallest, most battle-tested trusted-code surface; simplest possible recovery floor.
**Cons:** Manual peer management; no policy/ACL/enrollment ergonomics; a single shared hub re-creates the multi-tenant trust problem (must be run per-operator to avoid it); does not scale operationally to many operators × many machines.
#### Option B: NetBird (Selected)
**Pros:** Policy/ACL-based, revocable, per-peer access control; enrollment + audit out of the box; **self-hostable, fully FOSS control plane** — we retain the ability to run and modify every layer; scales to the many-operators × many-machines world #42 anticipates.
**Cons:** The NetBird agent is a userspace daemon, so the recovery path depends on that daemon being up (less bulletproof than kernel-level always-up WG) — mitigated by it being independent of the bitSpire app, mature, and systemd-restarted; running a control plane is operational weight (acceptable: the platform provides a default; sovereignty-seeking operators self-host).
#### Option C: Tailscale
**Pros:** Best-in-class ergonomics and NAT traversal.
**Cons:** **Control plane is closed source with no FOSS alternative** (headscale only reimplements the coordination server, chasing an upstream we don't control). Fails the hard requirement that we can always self-host and modify any software we depend on. Rejected on that basis alone.
#### Option D: On-demand tunnel toggled by the Nostr control plane
**Pros:** No standing reachability; every access window is a signed, audited, time-boxed event.
**Cons:** If the toggle is handled by the app, it fails in the exact recovery scenario (listener died with the app). If handled by a separate daemon, it reintroduces a privileged userspace listener into the recovery path and grows, rather than shrinks, the trusted-code surface. Acceptable only as an **audited convenience layer on top of** an always-available floor (and designed to fail open), never as the load-bearing gate. Not adopted as the primary mechanism.
### Trust model for excluding the SaaS operator
#### Option 1: Dashboard toggle to revoke SaaS-operator access (Rejected)
The SaaS operator controls LNbits (the machine's wallet/account is an LNbits user they can administer) and, in the default deployment, the NetBird control plane. A toggle whose enforcement they ultimately control gives the machine operator no real protection against them — it is security theater. **Rejected as a false promise.**
#### Option 2: Exclusion by ownership (Selected)
The only honest way for a machine operator to exclude the SaaS operator is to **own the access plane**: disable the default (platform-provided) NetBird enrollment and stand up their own — self-hosted NetBird, or their own WireGuard hub. The default deployment trusts whoever runs the control plane *and says so plainly*; operators who won't extend that trust take ownership. Control = ownership; we do not pretend otherwise.
## Consequences
### Positive
- Honest trust boundaries: the SaaS operator has no standing box access by default, and the limits of platform-provided convenience are stated rather than faked.
- Scales to many operators × many machines via NetBird policy/enrollment, while preserving a self-hosting escape hatch for sovereignty.
- The recovery plane survives app and OS failure because it is provisioned at install and independent of the runtime.
- Every dependency remains FOSS and self-hostable — no closed control plane anywhere in the stack.
### Negative
- The NetBird agent is a standing userspace daemon; a box where *both* the app and the agent are down falls to the physical/LAN floor (same floor as any remote scheme — only pure kernel-WG narrows it, at the cost of NetBird's ergonomics). Operators who weight reliability over ergonomics can choose self-hosted plain WireGuard.
- Sovereignty for a distrusting operator costs them operational work (running their own access plane). This is inherent to "control = ownership," not incidental.
- Two enrollment surfaces at provisioning: app/payment identity (#42 seed URL) and system/access identity (this plane). They must be kept conceptually distinct.
### Future Considerations
- An **audited convenience layer** (Nostr `OpenAccess`/`CloseAccess` that opens a time-boxed SSH window and logs it as a signed event) may be added *on top of* the always-available floor, designed to fail open, for the routine "let me in" case. It is explicitly not the recovery gate.
- The machine operator's Nostr key can become the single root of trust across all three planes — SSH `authorized_keys` + VPN enrollment at install, `AddOperator`/`RevokeOperator` (#42) for delegation — so granting/revoking any party (including the SaaS operator) is one scoped, revocable capability model.
- `sshd` posture should be tightened to key-only for deployed boxes (password auth is currently forced on for installed configs for first-boot provisioning; scope it to the LAN/first-boot window). Tracks with [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51).
## Amendment (2026-08-04): the access/recovery plane is not the *only* recovery
**Status:** Accepted · **Context:** the ATM app had no way to recover its own
connectivity — a machine that booted with no internet (or whose init otherwise
failed) sat on "ATM Unavailable" until a manual `systemctl restart bitspire`,
even after the network came back.
This ADR's SSH/NetBird recovery plane stands — it is the operator's
**app-and-OS-independent** path for the unanticipated and the broken, and may
carry recovery *procedures* (restart the service, inspect logs, re-provision).
But it is explicitly **not the first-line and not the only recovery method.**
Recovery is layered, cheapest-first:
1. **App auto-recovery (first-line, no human).** The ATM app recovers its own
relay/Lightning connectivity when possible: the nostr client already
reconnects with backoff, and the app now re-initializes when connectivity
returns (a fresh renderer reload — HAL is preserved in the main process),
so "internet came back" self-heals without anyone touching the machine.
2. **On-screen manual retry (operator at the machine).** The maintenance
("ATM Unavailable") screen carries a **Retry** button so a person standing
at the kiosk can force an immediate recovery attempt without shell access.
3. **SSH/NetBird (operator remote, last resort).** This plane — for when the
app *can't* self-heal or the box is genuinely broken. Unchanged by this
amendment beyond the reframing: it is the floor, not the front line.
Rationale: the common failure (transient network / boot-before-network) must
not require remote shell access to a public kiosk. Reserve the heavyweight
recovery plane for genuine app/OS failure. Implemented on branch
`feat/connection-recovery`.
## References
- [#41](https://git.atitlan.io/aiolabs/bitspire/issues/41) — Multi-location deployment: runtime site config (the access plane's per-machine identity is provisioned here, not baked into the closure).
- [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42) — Fleet management: Nostr-native remote control & telemetry (the control plane this ADR sits beside).
- [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51) — NixOS systemd hardening (sshd posture tightening).
- [#52](https://git.atitlan.io/aiolabs/bitspire/issues/52) — Sidecar bunker for the ATM key (related key-handling direction).
- `deploy/nixos/configuration.nix` — current WireGuard hub + `sshd` config (to be reworked per this decision).
- NetBird — <https://github.com/netbirdio/netbird> (self-hostable, FOSS control plane).