A relay outage at boot is misread as a revoked pairing, and latches the machine on the pairing wizard for good #121

Open
opened 2026-10-08 05:34:59 +00:00 by padreug · 0 comments
Owner

Sintra was power-cycled on 2026-10-07 and came up on the QR pairing wizard. Its pairing was never lost — the bunker_binding row in state.db is intact and the signer resumed fine ([Signer] Resuming bunker session for spire df2003…). What actually happened is that a transport failure got laundered into a pairing-revoked verdict, and that verdict is non-recoverable.

Sequence on sintra

Boot at 22:10, no usable network yet — every outbound call failing ([ExchangeRate] LNbits failed: Failed to fetch), relay in backoff 2s → 4s → 8s → 16s → 32s. At 22:18:54 the watchdog reloads the renderer, the LNbits client initializes against a relay that is connected-but-flapping, and the first bunker call dies:

[Lightning] LNbits client initialized
[ATM] HAL initialization failed: BunkerRejectedError: bunker nip44_encrypt: All promises were rejected

All promises were rejected is nostr-tools' nip46 Promise.any over the relay set — it means the request never reached a relay. Nothing was rejected by the bunker.

The three steps that turn that into a dead machine

  1. packages/nostr-client/src/bunker-signer.ts:105-106 — #call wraps everything that isn't a BunkerTimeoutError as BunkerRejectedError. A publish that never left the box is indistinguishable from the bunker saying no.
  2. apps/machine/src/services/init-error.ts:17 — BunkerRejectedError → unpaired, on the documented assumption that it means revoked / TTL-expired / off-policy.
  3. apps/machine/src/App.vue:191 — unpaired is in NON_RECOVERABLE, so the watchdog stops retrying.

The network came back within minutes (exchange rates were flowing again by 22:25), but step 3 means the machine never re-attempted. It sat on the wizard for ~75 minutes until I restarted the unit; after the restart it initialized fully in 35s — wallet, operator pubkey, fee subscription, HAL.

Why this one deserves a fix rather than a shrug

The wizard is the one screen where an operator is invited to burn a fresh one-shot connect token. A machine that boots a minute before its network is ready — which is the normal shape of a cold start, and the exact shape of a power cut — asks to be re-paired when nothing is wrong with its pairing. Had the operator obliged, we'd have spent a seed and rotated the binding to fix a DHCP race.

Suggested shape

Separate "could not reach the bunker" from "the bunker said no" in #call: a transport-layer failure (Promise.any aggregate, no connected relays) belongs with BunkerTimeoutError → signer-unreachable, which is already modelled as transient and already retries. Reserve BunkerRejectedError for an actual error string coming back from the bunker over the wire.

Worth considering alongside: gate initializeServices on at least one live relay socket before the first bunker call, so a flapping connection doesn't get one shot at the worst possible moment. [Lightning] Connected to relay: … currently logs optimistically and printed at 22:11:08 on the same line as [Nostr] Relay …: reconnecting in 2s (attempt 1) — the log reads as healthy while the socket is down, which is its own small trap when debugging this.

Sintra was power-cycled on 2026-10-07 and came up on the QR pairing wizard. Its pairing was never lost — the `bunker_binding` row in `state.db` is intact and the signer resumed fine (`[Signer] Resuming bunker session for spire df2003…`). What actually happened is that a transport failure got laundered into a pairing-revoked verdict, and that verdict is non-recoverable. ### Sequence on sintra Boot at 22:10, no usable network yet — every outbound call failing (`[ExchangeRate] LNbits failed: Failed to fetch`), relay in backoff `2s → 4s → 8s → 16s → 32s`. At 22:18:54 the watchdog reloads the renderer, the LNbits client initializes against a relay that is connected-but-flapping, and the first bunker call dies: ``` [Lightning] LNbits client initialized [ATM] HAL initialization failed: BunkerRejectedError: bunker nip44_encrypt: All promises were rejected ``` `All promises were rejected` is nostr-tools' nip46 `Promise.any` over the relay set — it means the request never reached a relay. Nothing was rejected by the bunker. ### The three steps that turn that into a dead machine 1. `packages/nostr-client/src/bunker-signer.ts:105-106` — `#call` wraps *everything* that isn't a `BunkerTimeoutError` as `BunkerRejectedError`. A publish that never left the box is indistinguishable from the bunker saying no. 2. `apps/machine/src/services/init-error.ts:17` — `BunkerRejectedError` → `unpaired`, on the documented assumption that it means revoked / TTL-expired / off-policy. 3. `apps/machine/src/App.vue:191` — `unpaired` is in `NON_RECOVERABLE`, so the watchdog stops retrying. The network came back within minutes (exchange rates were flowing again by 22:25), but step 3 means the machine never re-attempted. It sat on the wizard for ~75 minutes until I restarted the unit; after the restart it initialized fully in 35s — wallet, operator pubkey, fee subscription, HAL. ### Why this one deserves a fix rather than a shrug The wizard is the one screen where an operator is invited to burn a fresh one-shot connect token. A machine that boots a minute before its network is ready — which is the normal shape of a cold start, and the exact shape of a power cut — asks to be re-paired when nothing is wrong with its pairing. Had the operator obliged, we'd have spent a seed and rotated the binding to fix a DHCP race. ### Suggested shape Separate "could not reach the bunker" from "the bunker said no" in `#call`: a transport-layer failure (`Promise.any` aggregate, no connected relays) belongs with `BunkerTimeoutError` → `signer-unreachable`, which is already modelled as transient and already retries. Reserve `BunkerRejectedError` for an actual `error` string coming back from the bunker over the wire. Worth considering alongside: gate `initializeServices` on at least one live relay socket before the first bunker call, so a flapping connection doesn't get one shot at the worst possible moment. `[Lightning] Connected to relay: …` currently logs optimistically and printed at 22:11:08 on the same line as `[Nostr] Relay …: reconnecting in 2s (attempt 1)` — the log reads as healthy while the socket is down, which is its own small trap when debugging this.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/bitspire#121
No description provided.