Remote access & recovery plane: NetBird-based, operator-owned (implements ADR-002) #59

Open
opened 2026-06-14 09:36:51 +00:00 by padreug · 0 comments
Owner

Tracking issue for implementing the remote access / recovery plane decided in ADR-002 (docs/adr/002-remote-access-and-fleet-management.md, on the dev branch). The full architectural reasoning lives in the ADR; this issue tracks the implementation, to be picked up in the future.

Decision recap

Three planes, separated by trust owner:

  • Payment — ATM ↔ LNbits over the nostr-native-transport (SaaS operator; implies no box access).
  • Fleet control — Nostr commands (#42; authorized by the machine operator's key).
  • Access / recovery — SSH, this issue (machine-operator-owned).

Key calls:

  • The machine operator's own recovery access is provisioned at install and app-independent — it survives bitSpire app/OS failure (the case you most need it). Everyone else's access is runtime-granted, scoped, and revocable.
  • NetBird is the standard access plane — chosen for fleet scale and a fully FOSS, self-hostable control plane. Tailscale rejected (closed control plane, no FOSS alternative).
  • No dashboard "revoke SaaS-operator access" toggle. It is a false promise of security: the SaaS operator controls LNbits (the machine's wallet is an LNbits user) and, in the default deployment, the NetBird control plane. Exclusion is therefore by ownership, not by toggle — an operator who won't extend that trust disables the default NetBird and stands up their own access plane (self-hosted NetBird, or their own WireGuard hub).

Implementation tasks

  • Replace the single shared WireGuard hub (170.75.161.21, deploy/nixos/configuration.nix) with NetBird agent enrollment.
  • Provision the machine operator's access authority at install — NetBird enrollment + their SSH key in authorized_keys — independent of the bitSpire app/runtime, so a crashed box is still reachable.
  • Keep this a distinct provisioning step from the LNbits/payment enrollment (#42 seed URL): two enrollment surfaces, kept conceptually separate (relates to #41).
  • Default deployment: platform-provided NetBird control plane (convenience); document plainly that whoever runs it is trusted at that layer.
  • Sovereignty path: documented procedure for an operator to disable the default and self-host (NetBird or plain WG).
  • Tighten sshd to key-only on deployed boxes; scope password auth to the LAN/first-boot window (relates to #51).
  • (Optional, later) audited convenience layer: a Nostr OpenAccess/CloseAccess command that opens a time-boxed SSH window and logs it as a signed event — on top of the always-available floor and fail-open, never the recovery gate.
  • Docs: deploy/nixos/README.md + CLAUDE.md provisioning sections.

Open questions

  1. Who runs the default NetBird control plane (aiolabs-hosted for tenants vs per-operator), and how multi-tenant isolation is expressed in NetBird policy so one operator can't reach another's machines.
  2. Reliability trade: the NetBird agent is a userspace daemon (less bulletproof than kernel-level always-up WireGuard). Offer self-hosted plain WG as the reliability-first option for operators who weight it over ergonomics?
  3. Relationship to #41: the runtime site-config plan needs revising so the WireGuard ip field gives way to NetBird enrollment, while hostname / USB serials / touchscreen calibration stay in site.json.
  • ADR-002 (this issue implements it) — docs/adr/002-remote-access-and-fleet-management.md
  • #41 — Multi-location deployment: runtime site config
  • #42 — Fleet management: Nostr-native remote control & telemetry
  • #51 — NixOS systemd hardening (sshd posture)
  • #52 — Sidecar bunker for the ATM key
Tracking issue for implementing the remote access / recovery plane decided in **ADR-002** (`docs/adr/002-remote-access-and-fleet-management.md`, on the `dev` branch). The full architectural reasoning lives in the ADR; this issue tracks the implementation, to be picked up in the future. ## Decision recap Three planes, separated by **trust owner**: - **Payment** — ATM ↔ LNbits over the nostr-native-transport (SaaS operator; implies *no* box access). - **Fleet control** — Nostr commands (#42; authorized by the machine operator's key). - **Access / recovery** — SSH, *this issue* (machine-operator-owned). Key calls: - The machine operator's own recovery access is **provisioned at install and app-independent** — it survives bitSpire app/OS failure (the case you most need it). Everyone else's access is runtime-granted, scoped, and revocable. - **NetBird** is the standard access plane — chosen for fleet scale and a **fully FOSS, self-hostable control plane**. Tailscale rejected (closed control plane, no FOSS alternative). - **No dashboard "revoke SaaS-operator access" toggle.** It is a false promise of security: the SaaS operator controls LNbits (the machine's wallet is an LNbits user) and, in the default deployment, the NetBird control plane. Exclusion is therefore **by ownership, not by toggle** — an operator who won't extend that trust disables the default NetBird and stands up their own access plane (self-hosted NetBird, or their own WireGuard hub). ## Implementation tasks - [ ] Replace the single shared WireGuard hub (`170.75.161.21`, `deploy/nixos/configuration.nix`) with NetBird agent enrollment. - [ ] Provision the machine operator's access authority at install — NetBird enrollment + their SSH key in `authorized_keys` — independent of the bitSpire app/runtime, so a crashed box is still reachable. - [ ] Keep this a **distinct provisioning step** from the LNbits/payment enrollment (#42 seed URL): two enrollment surfaces, kept conceptually separate (relates to #41). - [ ] Default deployment: platform-provided NetBird control plane (convenience); document **plainly** that whoever runs it is trusted at that layer. - [ ] Sovereignty path: documented procedure for an operator to disable the default and self-host (NetBird or plain WG). - [ ] Tighten `sshd` to key-only on deployed boxes; scope password auth to the LAN/first-boot window (relates to #51). - [ ] *(Optional, later)* audited convenience layer: a Nostr `OpenAccess`/`CloseAccess` command that opens a time-boxed SSH window and logs it as a signed event — **on top of** the always-available floor and **fail-open**, never the recovery gate. - [ ] Docs: `deploy/nixos/README.md` + `CLAUDE.md` provisioning sections. ## Open questions 1. Who runs the default NetBird control plane (aiolabs-hosted for tenants vs per-operator), and how multi-tenant isolation is expressed in NetBird policy so one operator can't reach another's machines. 2. Reliability trade: the NetBird agent is a userspace daemon (less bulletproof than kernel-level always-up WireGuard). Offer self-hosted plain WG as the reliability-first option for operators who weight it over ergonomics? 3. Relationship to #41: the runtime site-config plan needs revising so the WireGuard `ip` field gives way to NetBird enrollment, while hostname / USB serials / touchscreen calibration stay in `site.json`. ## Related - **ADR-002** (this issue implements it) — `docs/adr/002-remote-access-and-fleet-management.md` - #41 — Multi-location deployment: runtime site config - #42 — Fleet management: Nostr-native remote control & telemetry - #51 — NixOS systemd hardening (sshd posture) - #52 — Sidecar bunker for the ATM key
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/bitspire#59
No description provided.