From 627d5e63e5ac1b8400d6618e291becdf0251cd1f Mon Sep 17 00:00:00 2001 From: Padreug Date: Sun, 14 Jun 2026 11:17:02 +0200 Subject: [PATCH] =?UTF-8?q?docs(adr):=20ADR-002=20remote=20access=20&=20fl?= =?UTF-8?q?eet=20management=20=E2=80=94=20three=20planes,=20NetBird?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Separate payment (Nostr↔LNbits, SaaS-operator-owned), fleet control (Nostr #42, machine-operator-owned), and access/recovery (SSH) planes by trust owner. Recovery access is provisioned at install and app-independent. Adopt NetBird for the access plane (scale + fully FOSS self-hostable control plane; rejects Tailscale's closed control plane). Reject a dashboard 'revoke SaaS-operator access' toggle as a false promise — the SaaS operator controls LNbits and the default control plane, so exclusion is by ownership (operator self-hosts), not by toggle. --- .../002-remote-access-and-fleet-management.md | 114 ++++++++++++++++++ 1 file changed, 114 insertions(+) create mode 100644 docs/adr/002-remote-access-and-fleet-management.md diff --git a/docs/adr/002-remote-access-and-fleet-management.md b/docs/adr/002-remote-access-and-fleet-management.md new file mode 100644 index 0000000..4931daa --- /dev/null +++ b/docs/adr/002-remote-access-and-fleet-management.md @@ -0,0 +1,114 @@ +# ADR-002: Remote Access & Fleet Management — Three Planes, Operator-Owned Access via NetBird + +**Status:** Accepted +**Date:** 2026-06-14 +**Context:** Multi-operator bitSpire fleet — separating the payment, control, and recovery planes by who owns them. + +## Decision + +1. **Separate three planes by trust owner**, and never conflate them: + - **Payment plane** — ATM ↔ LNbits over the nostr-native-transport. Owned by the **SaaS operator**. Implies *no* machine access. + - **Fleet control plane** — routine ops/telemetry/enrollment over Nostr (see [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42)). Authorized by the **machine operator's** key. + - **Access / recovery plane** — SSH for the unanticipated and the broken. Owned by the **machine operator**. + +2. **The machine operator's own recovery access is provisioned at install and is app-independent.** Their SSH key and their VPN/NetBird enrollment are established when the machine is set up, so they can always reach a box even when the bitSpire app or OS is broken. Access for *anyone else* is runtime-granted, scoped, and revocable — never the owner's own path. + +3. **Adopt NetBird as the standard access/recovery plane**, chosen for fleet scale and a **self-hostable, fully FOSS control plane**. The platform may provide a default NetBird setup as a convenience; a machine operator who does not wish to trust whoever runs that control plane **disables it and provisions their own access plane** (self-hosted NetBird, or their own WireGuard hub). + +4. **We will NOT build a "revoke SaaS-operator access" toggle in the operator dashboard.** It is a false promise of security: the SaaS operator runs LNbits (and, in the default deployment, the NetBird control plane), so a toggle they ultimately control cannot protect a machine operator against them. The honest boundary is **exclusion-by-ownership, not exclusion-by-toggle** — an operator who wants to exclude the SaaS operator takes ownership of the access plane. + +## Context + +### The players + +A deployed bitSpire machine sits between two distinct principals: + +- **SaaS operator** — runs the LNbits instance and provides the Lightning backend as a service. +- **Machine operator** — owns the physical ATM(s) and is identified by a Nostr key (the operator pubkey in the [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42) allow-list). + +These are different parties with different interests. A machine operator will want to SSH to their own machine for support and recovery, and **may or may not want to grant the SaaS operator that same access.** + +### Why the SaaS operator needs zero box access by design + +The whole nostr-native architecture (no admin tokens on the kiosk, no inbound network surface, payment over Nostr) means the SaaS operator can deliver the full service **without ever touching the machine**. So "the machine operator may refuse the SaaS operator access" is not a constraint to engineer around — it is the **default that costs nothing**. SaaS-operator box access is a *support convenience*, never a service requirement. The natural posture is therefore **default-deny for the SaaS operator**. + +### Why SSH can't be replaced by the Nostr control plane + +The Nostr control plane (#42) is a fixed menu of structured, capability-scoped commands dispatched by a handler *inside the app*. It is excellent for routine, auditable, fleet-wide ops on **healthy** machines, and strictly better than SSH for those (signed, scoped, logged, fan-out). But: + +- It can only do what a handler was written for; incidents are by definition unanticipated. +- The listener lives in the app, so it dies exactly when the app dies — the case you most need recovery for. + +SSH (arbitrary, interactive, app-independent) is therefore irreducible as the **recovery plane**. The two are complements, not substitutes. + +### Why the recovery path must be app-independent + +The whole point of a recovery path is to survive the failure of the thing it recovers. So it must not be gated by the bitSpire app, nor by a Nostr command the app dispatches. The kernel/agent that carries the tunnel and `sshd` must come up at boot independent of the app. (`allowedTCPPorts = []` already means `sshd` is unreachable except across the tunnel — the VPN handshake is the outer lock, the SSH key the inner one.) + +### Why the single shared hub had to change + +The pre-existing design used one WireGuard hub (`170.75.161.21`) run by platform infra. Whoever runs that hub has a standing network path to every enrolled box — i.e. the SaaS operator having access to machines they don't own. Multi-tenancy requires the access plane to be **per-operator or policy-isolated**, rooted in the machine operator, not the platform. + +## Options Considered + +### Access-plane mechanism + +#### Option A: Always-up minimal WireGuard hub + +**Pros:** After boot, zero userspace dependency — the kernel holds the tunnel, nothing can crash it short of a kernel/networking fault; smallest, most battle-tested trusted-code surface; simplest possible recovery floor. +**Cons:** Manual peer management; no policy/ACL/enrollment ergonomics; a single shared hub re-creates the multi-tenant trust problem (must be run per-operator to avoid it); does not scale operationally to many operators × many machines. + +#### Option B: NetBird (Selected) + +**Pros:** Policy/ACL-based, revocable, per-peer access control; enrollment + audit out of the box; **self-hostable, fully FOSS control plane** — we retain the ability to run and modify every layer; scales to the many-operators × many-machines world #42 anticipates. +**Cons:** The NetBird agent is a userspace daemon, so the recovery path depends on that daemon being up (less bulletproof than kernel-level always-up WG) — mitigated by it being independent of the bitSpire app, mature, and systemd-restarted; running a control plane is operational weight (acceptable: the platform provides a default; sovereignty-seeking operators self-host). + +#### Option C: Tailscale + +**Pros:** Best-in-class ergonomics and NAT traversal. +**Cons:** **Control plane is closed source with no FOSS alternative** (headscale only reimplements the coordination server, chasing an upstream we don't control). Fails the hard requirement that we can always self-host and modify any software we depend on. Rejected on that basis alone. + +#### Option D: On-demand tunnel toggled by the Nostr control plane + +**Pros:** No standing reachability; every access window is a signed, audited, time-boxed event. +**Cons:** If the toggle is handled by the app, it fails in the exact recovery scenario (listener died with the app). If handled by a separate daemon, it reintroduces a privileged userspace listener into the recovery path and grows, rather than shrinks, the trusted-code surface. Acceptable only as an **audited convenience layer on top of** an always-available floor (and designed to fail open), never as the load-bearing gate. Not adopted as the primary mechanism. + +### Trust model for excluding the SaaS operator + +#### Option 1: Dashboard toggle to revoke SaaS-operator access (Rejected) + +The SaaS operator controls LNbits (the machine's wallet/account is an LNbits user they can administer) and, in the default deployment, the NetBird control plane. A toggle whose enforcement they ultimately control gives the machine operator no real protection against them — it is security theater. **Rejected as a false promise.** + +#### Option 2: Exclusion by ownership (Selected) + +The only honest way for a machine operator to exclude the SaaS operator is to **own the access plane**: disable the default (platform-provided) NetBird enrollment and stand up their own — self-hosted NetBird, or their own WireGuard hub. The default deployment trusts whoever runs the control plane *and says so plainly*; operators who won't extend that trust take ownership. Control = ownership; we do not pretend otherwise. + +## Consequences + +### Positive + +- Honest trust boundaries: the SaaS operator has no standing box access by default, and the limits of platform-provided convenience are stated rather than faked. +- Scales to many operators × many machines via NetBird policy/enrollment, while preserving a self-hosting escape hatch for sovereignty. +- The recovery plane survives app and OS failure because it is provisioned at install and independent of the runtime. +- Every dependency remains FOSS and self-hostable — no closed control plane anywhere in the stack. + +### Negative + +- The NetBird agent is a standing userspace daemon; a box where *both* the app and the agent are down falls to the physical/LAN floor (same floor as any remote scheme — only pure kernel-WG narrows it, at the cost of NetBird's ergonomics). Operators who weight reliability over ergonomics can choose self-hosted plain WireGuard. +- Sovereignty for a distrusting operator costs them operational work (running their own access plane). This is inherent to "control = ownership," not incidental. +- Two enrollment surfaces at provisioning: app/payment identity (#42 seed URL) and system/access identity (this plane). They must be kept conceptually distinct. + +### Future Considerations + +- An **audited convenience layer** (Nostr `OpenAccess`/`CloseAccess` that opens a time-boxed SSH window and logs it as a signed event) may be added *on top of* the always-available floor, designed to fail open, for the routine "let me in" case. It is explicitly not the recovery gate. +- The machine operator's Nostr key can become the single root of trust across all three planes — SSH `authorized_keys` + VPN enrollment at install, `AddOperator`/`RevokeOperator` (#42) for delegation — so granting/revoking any party (including the SaaS operator) is one scoped, revocable capability model. +- `sshd` posture should be tightened to key-only for deployed boxes (password auth is currently forced on for installed configs for first-boot provisioning; scope it to the LAN/first-boot window). Tracks with [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51). + +## References + +- [#41](https://git.atitlan.io/aiolabs/bitspire/issues/41) — Multi-location deployment: runtime site config (the access plane's per-machine identity is provisioned here, not baked into the closure). +- [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42) — Fleet management: Nostr-native remote control & telemetry (the control plane this ADR sits beside). +- [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51) — NixOS systemd hardening (sshd posture tightening). +- [#52](https://git.atitlan.io/aiolabs/bitspire/issues/52) — Sidecar bunker for the ATM key (related key-handling direction). +- `deploy/nixos/configuration.nix` — current WireGuard hub + `sshd` config (to be reworked per this decision). +- NetBird — (self-hostable, FOSS control plane).