docs(adr): ADR-002 remote access & fleet management — three planes, NetBird

Separate payment (Nostr↔LNbits, SaaS-operator-owned), fleet control
(Nostr #42, machine-operator-owned), and access/recovery (SSH) planes
by trust owner. Recovery access is provisioned at install and
app-independent. Adopt NetBird for the access plane (scale + fully FOSS
self-hostable control plane; rejects Tailscale's closed control plane).

Reject a dashboard 'revoke SaaS-operator access' toggle as a false
promise — the SaaS operator controls LNbits and the default control
plane, so exclusion is by ownership (operator self-hosts), not by
toggle.
This commit is contained in:
Padreug 2026-06-14 11:17:02 +02:00
commit 627d5e63e5

View file

@ -0,0 +1,114 @@
# ADR-002: Remote Access & Fleet Management — Three Planes, Operator-Owned Access via NetBird
**Status:** Accepted
**Date:** 2026-06-14
**Context:** Multi-operator bitSpire fleet — separating the payment, control, and recovery planes by who owns them.
## Decision
1. **Separate three planes by trust owner**, and never conflate them:
- **Payment plane** — ATM ↔ LNbits over the nostr-native-transport. Owned by the **SaaS operator**. Implies *no* machine access.
- **Fleet control plane** — routine ops/telemetry/enrollment over Nostr (see [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42)). Authorized by the **machine operator's** key.
- **Access / recovery plane** — SSH for the unanticipated and the broken. Owned by the **machine operator**.
2. **The machine operator's own recovery access is provisioned at install and is app-independent.** Their SSH key and their VPN/NetBird enrollment are established when the machine is set up, so they can always reach a box even when the bitSpire app or OS is broken. Access for *anyone else* is runtime-granted, scoped, and revocable — never the owner's own path.
3. **Adopt NetBird as the standard access/recovery plane**, chosen for fleet scale and a **self-hostable, fully FOSS control plane**. The platform may provide a default NetBird setup as a convenience; a machine operator who does not wish to trust whoever runs that control plane **disables it and provisions their own access plane** (self-hosted NetBird, or their own WireGuard hub).
4. **We will NOT build a "revoke SaaS-operator access" toggle in the operator dashboard.** It is a false promise of security: the SaaS operator runs LNbits (and, in the default deployment, the NetBird control plane), so a toggle they ultimately control cannot protect a machine operator against them. The honest boundary is **exclusion-by-ownership, not exclusion-by-toggle** — an operator who wants to exclude the SaaS operator takes ownership of the access plane.
## Context
### The players
A deployed bitSpire machine sits between two distinct principals:
- **SaaS operator** — runs the LNbits instance and provides the Lightning backend as a service.
- **Machine operator** — owns the physical ATM(s) and is identified by a Nostr key (the operator pubkey in the [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42) allow-list).
These are different parties with different interests. A machine operator will want to SSH to their own machine for support and recovery, and **may or may not want to grant the SaaS operator that same access.**
### Why the SaaS operator needs zero box access by design
The whole nostr-native architecture (no admin tokens on the kiosk, no inbound network surface, payment over Nostr) means the SaaS operator can deliver the full service **without ever touching the machine**. So "the machine operator may refuse the SaaS operator access" is not a constraint to engineer around — it is the **default that costs nothing**. SaaS-operator box access is a *support convenience*, never a service requirement. The natural posture is therefore **default-deny for the SaaS operator**.
### Why SSH can't be replaced by the Nostr control plane
The Nostr control plane (#42) is a fixed menu of structured, capability-scoped commands dispatched by a handler *inside the app*. It is excellent for routine, auditable, fleet-wide ops on **healthy** machines, and strictly better than SSH for those (signed, scoped, logged, fan-out). But:
- It can only do what a handler was written for; incidents are by definition unanticipated.
- The listener lives in the app, so it dies exactly when the app dies — the case you most need recovery for.
SSH (arbitrary, interactive, app-independent) is therefore irreducible as the **recovery plane**. The two are complements, not substitutes.
### Why the recovery path must be app-independent
The whole point of a recovery path is to survive the failure of the thing it recovers. So it must not be gated by the bitSpire app, nor by a Nostr command the app dispatches. The kernel/agent that carries the tunnel and `sshd` must come up at boot independent of the app. (`allowedTCPPorts = []` already means `sshd` is unreachable except across the tunnel — the VPN handshake is the outer lock, the SSH key the inner one.)
### Why the single shared hub had to change
The pre-existing design used one WireGuard hub (`170.75.161.21`) run by platform infra. Whoever runs that hub has a standing network path to every enrolled box — i.e. the SaaS operator having access to machines they don't own. Multi-tenancy requires the access plane to be **per-operator or policy-isolated**, rooted in the machine operator, not the platform.
## Options Considered
### Access-plane mechanism
#### Option A: Always-up minimal WireGuard hub
**Pros:** After boot, zero userspace dependency — the kernel holds the tunnel, nothing can crash it short of a kernel/networking fault; smallest, most battle-tested trusted-code surface; simplest possible recovery floor.
**Cons:** Manual peer management; no policy/ACL/enrollment ergonomics; a single shared hub re-creates the multi-tenant trust problem (must be run per-operator to avoid it); does not scale operationally to many operators × many machines.
#### Option B: NetBird (Selected)
**Pros:** Policy/ACL-based, revocable, per-peer access control; enrollment + audit out of the box; **self-hostable, fully FOSS control plane** — we retain the ability to run and modify every layer; scales to the many-operators × many-machines world #42 anticipates.
**Cons:** The NetBird agent is a userspace daemon, so the recovery path depends on that daemon being up (less bulletproof than kernel-level always-up WG) — mitigated by it being independent of the bitSpire app, mature, and systemd-restarted; running a control plane is operational weight (acceptable: the platform provides a default; sovereignty-seeking operators self-host).
#### Option C: Tailscale
**Pros:** Best-in-class ergonomics and NAT traversal.
**Cons:** **Control plane is closed source with no FOSS alternative** (headscale only reimplements the coordination server, chasing an upstream we don't control). Fails the hard requirement that we can always self-host and modify any software we depend on. Rejected on that basis alone.
#### Option D: On-demand tunnel toggled by the Nostr control plane
**Pros:** No standing reachability; every access window is a signed, audited, time-boxed event.
**Cons:** If the toggle is handled by the app, it fails in the exact recovery scenario (listener died with the app). If handled by a separate daemon, it reintroduces a privileged userspace listener into the recovery path and grows, rather than shrinks, the trusted-code surface. Acceptable only as an **audited convenience layer on top of** an always-available floor (and designed to fail open), never as the load-bearing gate. Not adopted as the primary mechanism.
### Trust model for excluding the SaaS operator
#### Option 1: Dashboard toggle to revoke SaaS-operator access (Rejected)
The SaaS operator controls LNbits (the machine's wallet/account is an LNbits user they can administer) and, in the default deployment, the NetBird control plane. A toggle whose enforcement they ultimately control gives the machine operator no real protection against them — it is security theater. **Rejected as a false promise.**
#### Option 2: Exclusion by ownership (Selected)
The only honest way for a machine operator to exclude the SaaS operator is to **own the access plane**: disable the default (platform-provided) NetBird enrollment and stand up their own — self-hosted NetBird, or their own WireGuard hub. The default deployment trusts whoever runs the control plane *and says so plainly*; operators who won't extend that trust take ownership. Control = ownership; we do not pretend otherwise.
## Consequences
### Positive
- Honest trust boundaries: the SaaS operator has no standing box access by default, and the limits of platform-provided convenience are stated rather than faked.
- Scales to many operators × many machines via NetBird policy/enrollment, while preserving a self-hosting escape hatch for sovereignty.
- The recovery plane survives app and OS failure because it is provisioned at install and independent of the runtime.
- Every dependency remains FOSS and self-hostable — no closed control plane anywhere in the stack.
### Negative
- The NetBird agent is a standing userspace daemon; a box where *both* the app and the agent are down falls to the physical/LAN floor (same floor as any remote scheme — only pure kernel-WG narrows it, at the cost of NetBird's ergonomics). Operators who weight reliability over ergonomics can choose self-hosted plain WireGuard.
- Sovereignty for a distrusting operator costs them operational work (running their own access plane). This is inherent to "control = ownership," not incidental.
- Two enrollment surfaces at provisioning: app/payment identity (#42 seed URL) and system/access identity (this plane). They must be kept conceptually distinct.
### Future Considerations
- An **audited convenience layer** (Nostr `OpenAccess`/`CloseAccess` that opens a time-boxed SSH window and logs it as a signed event) may be added *on top of* the always-available floor, designed to fail open, for the routine "let me in" case. It is explicitly not the recovery gate.
- The machine operator's Nostr key can become the single root of trust across all three planes — SSH `authorized_keys` + VPN enrollment at install, `AddOperator`/`RevokeOperator` (#42) for delegation — so granting/revoking any party (including the SaaS operator) is one scoped, revocable capability model.
- `sshd` posture should be tightened to key-only for deployed boxes (password auth is currently forced on for installed configs for first-boot provisioning; scope it to the LAN/first-boot window). Tracks with [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51).
## References
- [#41](https://git.atitlan.io/aiolabs/bitspire/issues/41) — Multi-location deployment: runtime site config (the access plane's per-machine identity is provisioned here, not baked into the closure).
- [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42) — Fleet management: Nostr-native remote control & telemetry (the control plane this ADR sits beside).
- [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51) — NixOS systemd hardening (sshd posture tightening).
- [#52](https://git.atitlan.io/aiolabs/bitspire/issues/52) — Sidecar bunker for the ATM key (related key-handling direction).
- `deploy/nixos/configuration.nix` — current WireGuard hub + `sshd` config (to be reworked per this decision).
- NetBird — <https://github.com/netbirdio/netbird> (self-hostable, FOSS control plane).