The access/recovery plane (SSH/NetBird) stands and may carry recovery procedures, but it is explicitly NOT the only or first-line recovery. Add a layered, cheapest-first recovery model: (1) app auto-recovery of its own relay/Lightning connectivity, (2) an on-screen Retry for an operator at the kiosk, (3) SSH/NetBird as the last-resort remote plane for genuine app/OS failure. A public kiosk must not need remote shell access to recover from a transient/boot-before-network outage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
144 lines
12 KiB
Markdown
144 lines
12 KiB
Markdown
# ADR-002: Remote Access & Fleet Management — Three Planes, Operator-Owned Access via NetBird
|
||
|
||
**Status:** Accepted
|
||
**Date:** 2026-06-14
|
||
**Context:** Multi-operator bitSpire fleet — separating the payment, control, and recovery planes by who owns them.
|
||
|
||
## Decision
|
||
|
||
1. **Separate three planes by trust owner**, and never conflate them:
|
||
- **Payment plane** — ATM ↔ LNbits over the nostr-native-transport. Owned by the **SaaS operator**. Implies *no* machine access.
|
||
- **Fleet control plane** — routine ops/telemetry/enrollment over Nostr (see [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42)). Authorized by the **machine operator's** key.
|
||
- **Access / recovery plane** — SSH for the unanticipated and the broken. Owned by the **machine operator**.
|
||
|
||
2. **The machine operator's own recovery access is provisioned at install and is app-independent.** Their SSH key and their VPN/NetBird enrollment are established when the machine is set up, so they can always reach a box even when the bitSpire app or OS is broken. Access for *anyone else* is runtime-granted, scoped, and revocable — never the owner's own path.
|
||
|
||
3. **Adopt NetBird as the standard access/recovery plane**, chosen for fleet scale and a **self-hostable, fully FOSS control plane**. The platform may provide a default NetBird setup as a convenience; a machine operator who does not wish to trust whoever runs that control plane **disables it and provisions their own access plane** (self-hosted NetBird, or their own WireGuard hub).
|
||
|
||
4. **We will NOT build a "revoke SaaS-operator access" toggle in the operator dashboard.** It is a false promise of security: the SaaS operator runs LNbits (and, in the default deployment, the NetBird control plane), so a toggle they ultimately control cannot protect a machine operator against them. The honest boundary is **exclusion-by-ownership, not exclusion-by-toggle** — an operator who wants to exclude the SaaS operator takes ownership of the access plane.
|
||
|
||
## Context
|
||
|
||
### The players
|
||
|
||
A deployed bitSpire machine sits between two distinct principals:
|
||
|
||
- **SaaS operator** — runs the LNbits instance and provides the Lightning backend as a service.
|
||
- **Machine operator** — owns the physical ATM(s) and is identified by a Nostr key (the operator pubkey in the [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42) allow-list).
|
||
|
||
These are different parties with different interests. A machine operator will want to SSH to their own machine for support and recovery, and **may or may not want to grant the SaaS operator that same access.**
|
||
|
||
### Why the SaaS operator needs zero box access by design
|
||
|
||
The whole nostr-native architecture (no admin tokens on the kiosk, no inbound network surface, payment over Nostr) means the SaaS operator can deliver the full service **without ever touching the machine**. So "the machine operator may refuse the SaaS operator access" is not a constraint to engineer around — it is the **default that costs nothing**. SaaS-operator box access is a *support convenience*, never a service requirement. The natural posture is therefore **default-deny for the SaaS operator**.
|
||
|
||
### Why SSH can't be replaced by the Nostr control plane
|
||
|
||
The Nostr control plane (#42) is a fixed menu of structured, capability-scoped commands dispatched by a handler *inside the app*. It is excellent for routine, auditable, fleet-wide ops on **healthy** machines, and strictly better than SSH for those (signed, scoped, logged, fan-out). But:
|
||
|
||
- It can only do what a handler was written for; incidents are by definition unanticipated.
|
||
- The listener lives in the app, so it dies exactly when the app dies — the case you most need recovery for.
|
||
|
||
SSH (arbitrary, interactive, app-independent) is therefore irreducible as the **recovery plane**. The two are complements, not substitutes.
|
||
|
||
### Why the recovery path must be app-independent
|
||
|
||
The whole point of a recovery path is to survive the failure of the thing it recovers. So it must not be gated by the bitSpire app, nor by a Nostr command the app dispatches. The kernel/agent that carries the tunnel and `sshd` must come up at boot independent of the app. (`allowedTCPPorts = []` already means `sshd` is unreachable except across the tunnel — the VPN handshake is the outer lock, the SSH key the inner one.)
|
||
|
||
### Why the single shared hub had to change
|
||
|
||
The pre-existing design used one WireGuard hub (`170.75.161.21`) run by platform infra. Whoever runs that hub has a standing network path to every enrolled box — i.e. the SaaS operator having access to machines they don't own. Multi-tenancy requires the access plane to be **per-operator or policy-isolated**, rooted in the machine operator, not the platform.
|
||
|
||
## Options Considered
|
||
|
||
### Access-plane mechanism
|
||
|
||
#### Option A: Always-up minimal WireGuard hub
|
||
|
||
**Pros:** After boot, zero userspace dependency — the kernel holds the tunnel, nothing can crash it short of a kernel/networking fault; smallest, most battle-tested trusted-code surface; simplest possible recovery floor.
|
||
**Cons:** Manual peer management; no policy/ACL/enrollment ergonomics; a single shared hub re-creates the multi-tenant trust problem (must be run per-operator to avoid it); does not scale operationally to many operators × many machines.
|
||
|
||
#### Option B: NetBird (Selected)
|
||
|
||
**Pros:** Policy/ACL-based, revocable, per-peer access control; enrollment + audit out of the box; **self-hostable, fully FOSS control plane** — we retain the ability to run and modify every layer; scales to the many-operators × many-machines world #42 anticipates.
|
||
**Cons:** The NetBird agent is a userspace daemon, so the recovery path depends on that daemon being up (less bulletproof than kernel-level always-up WG) — mitigated by it being independent of the bitSpire app, mature, and systemd-restarted; running a control plane is operational weight (acceptable: the platform provides a default; sovereignty-seeking operators self-host).
|
||
|
||
#### Option C: Tailscale
|
||
|
||
**Pros:** Best-in-class ergonomics and NAT traversal.
|
||
**Cons:** **Control plane is closed source with no FOSS alternative** (headscale only reimplements the coordination server, chasing an upstream we don't control). Fails the hard requirement that we can always self-host and modify any software we depend on. Rejected on that basis alone.
|
||
|
||
#### Option D: On-demand tunnel toggled by the Nostr control plane
|
||
|
||
**Pros:** No standing reachability; every access window is a signed, audited, time-boxed event.
|
||
**Cons:** If the toggle is handled by the app, it fails in the exact recovery scenario (listener died with the app). If handled by a separate daemon, it reintroduces a privileged userspace listener into the recovery path and grows, rather than shrinks, the trusted-code surface. Acceptable only as an **audited convenience layer on top of** an always-available floor (and designed to fail open), never as the load-bearing gate. Not adopted as the primary mechanism.
|
||
|
||
### Trust model for excluding the SaaS operator
|
||
|
||
#### Option 1: Dashboard toggle to revoke SaaS-operator access (Rejected)
|
||
|
||
The SaaS operator controls LNbits (the machine's wallet/account is an LNbits user they can administer) and, in the default deployment, the NetBird control plane. A toggle whose enforcement they ultimately control gives the machine operator no real protection against them — it is security theater. **Rejected as a false promise.**
|
||
|
||
#### Option 2: Exclusion by ownership (Selected)
|
||
|
||
The only honest way for a machine operator to exclude the SaaS operator is to **own the access plane**: disable the default (platform-provided) NetBird enrollment and stand up their own — self-hosted NetBird, or their own WireGuard hub. The default deployment trusts whoever runs the control plane *and says so plainly*; operators who won't extend that trust take ownership. Control = ownership; we do not pretend otherwise.
|
||
|
||
## Consequences
|
||
|
||
### Positive
|
||
|
||
- Honest trust boundaries: the SaaS operator has no standing box access by default, and the limits of platform-provided convenience are stated rather than faked.
|
||
- Scales to many operators × many machines via NetBird policy/enrollment, while preserving a self-hosting escape hatch for sovereignty.
|
||
- The recovery plane survives app and OS failure because it is provisioned at install and independent of the runtime.
|
||
- Every dependency remains FOSS and self-hostable — no closed control plane anywhere in the stack.
|
||
|
||
### Negative
|
||
|
||
- The NetBird agent is a standing userspace daemon; a box where *both* the app and the agent are down falls to the physical/LAN floor (same floor as any remote scheme — only pure kernel-WG narrows it, at the cost of NetBird's ergonomics). Operators who weight reliability over ergonomics can choose self-hosted plain WireGuard.
|
||
- Sovereignty for a distrusting operator costs them operational work (running their own access plane). This is inherent to "control = ownership," not incidental.
|
||
- Two enrollment surfaces at provisioning: app/payment identity (#42 seed URL) and system/access identity (this plane). They must be kept conceptually distinct.
|
||
|
||
### Future Considerations
|
||
|
||
- An **audited convenience layer** (Nostr `OpenAccess`/`CloseAccess` that opens a time-boxed SSH window and logs it as a signed event) may be added *on top of* the always-available floor, designed to fail open, for the routine "let me in" case. It is explicitly not the recovery gate.
|
||
- The machine operator's Nostr key can become the single root of trust across all three planes — SSH `authorized_keys` + VPN enrollment at install, `AddOperator`/`RevokeOperator` (#42) for delegation — so granting/revoking any party (including the SaaS operator) is one scoped, revocable capability model.
|
||
- `sshd` posture should be tightened to key-only for deployed boxes (password auth is currently forced on for installed configs for first-boot provisioning; scope it to the LAN/first-boot window). Tracks with [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51).
|
||
|
||
## Amendment (2026-08-04): the access/recovery plane is not the *only* recovery
|
||
|
||
**Status:** Accepted · **Context:** the ATM app had no way to recover its own
|
||
connectivity — a machine that booted with no internet (or whose init otherwise
|
||
failed) sat on "ATM Unavailable" until a manual `systemctl restart bitspire`,
|
||
even after the network came back.
|
||
|
||
This ADR's SSH/NetBird recovery plane stands — it is the operator's
|
||
**app-and-OS-independent** path for the unanticipated and the broken, and may
|
||
carry recovery *procedures* (restart the service, inspect logs, re-provision).
|
||
But it is explicitly **not the first-line and not the only recovery method.**
|
||
Recovery is layered, cheapest-first:
|
||
|
||
1. **App auto-recovery (first-line, no human).** The ATM app recovers its own
|
||
relay/Lightning connectivity when possible: the nostr client already
|
||
reconnects with backoff, and the app now re-initializes when connectivity
|
||
returns (a fresh renderer reload — HAL is preserved in the main process),
|
||
so "internet came back" self-heals without anyone touching the machine.
|
||
2. **On-screen manual retry (operator at the machine).** The maintenance
|
||
("ATM Unavailable") screen carries a **Retry** button so a person standing
|
||
at the kiosk can force an immediate recovery attempt without shell access.
|
||
3. **SSH/NetBird (operator remote, last resort).** This plane — for when the
|
||
app *can't* self-heal or the box is genuinely broken. Unchanged by this
|
||
amendment beyond the reframing: it is the floor, not the front line.
|
||
|
||
Rationale: the common failure (transient network / boot-before-network) must
|
||
not require remote shell access to a public kiosk. Reserve the heavyweight
|
||
recovery plane for genuine app/OS failure. Implemented on branch
|
||
`feat/connection-recovery`.
|
||
|
||
## References
|
||
|
||
- [#41](https://git.atitlan.io/aiolabs/bitspire/issues/41) — Multi-location deployment: runtime site config (the access plane's per-machine identity is provisioned here, not baked into the closure).
|
||
- [#42](https://git.atitlan.io/aiolabs/bitspire/issues/42) — Fleet management: Nostr-native remote control & telemetry (the control plane this ADR sits beside).
|
||
- [#51](https://git.atitlan.io/aiolabs/bitspire/issues/51) — NixOS systemd hardening (sshd posture tightening).
|
||
- [#52](https://git.atitlan.io/aiolabs/bitspire/issues/52) — Sidecar bunker for the ATM key (related key-handling direction).
|
||
- `deploy/nixos/configuration.nix` — current WireGuard hub + `sshd` config (to be reworked per this decision).
|
||
- NetBird — <https://github.com/netbirdio/netbird> (self-hostable, FOSS control plane).
|