audit id003 _send 100ms setTimeout — possible escrow→stack race on Sintras #46

Open
opened 2026-06-13 22:03:00 +00:00 by padreug · 0 comments
Owner

Migrated from aiolabs/lamassu-next#46 — opened by @padreug on 2026-05-19.\n\n## Context

While investigating the MEI/EBDS bill-tear root cause on the BATM3
(PR #45), the
id003 driver was inspected for an analogous timing concern.

The MEI tear is fixed by polling at 100 ms inside the validator's
escrow grace window. id003 already polls at 100 ms — but it has a
separate concern that warrants a focused audit before we declare the
JCM iVIZION path equally safe.

The concern

packages/hal/src/validators/id003/index.ts:276-289:

private _send(command: string): void {
  if (this.disablePolling) {
    this.rs232?.send(command)
    return
  }

  this._stopPolling()
  // Timeout to prevent interleaved commands
  setTimeout(() => {
    this.rs232?.send(command)
    if (!this.disablePolling) {
      this._startPolling()
    }
  }, POLLING_INTERVAL)
}

Every host command is delayed by POLLING_INTERVAL (100 ms) before the
serial write actually fires. Functionally this is a defensive pause to
"prevent interleaved commands" — but it means a host stack() call
arrives at the validator 100 ms after we decide to stack.

For id003 / JCM iVIZION the bill is mechanically gripped and stationary
during the host-decision phase, so this 100 ms is normally harmless.
But:

  • We have no measurement of how often escrow → stack actually closes
    within 100 ms vs how often it slips longer (no equivalent of the
    EBDS escrow watchdog landing in PR #45).
  • Under high host load (e.g. concurrent state-machine evaluations,
    large Vue re-renders, GC pauses on the Aaeon UP Board — which
    memory/sintra_perf_constraints.md notes is constrained) the
    effective delay could be 100 ms + scheduler latency.
  • If the JCM ever enters its own internal escrow grace timeout (the
    spec describes one, default behavior varies by firmware), we'd be
    in the same family of failure as the MEI.

The "interleaved commands" rationale is also undocumented — it'd be
worth tracing back where that came from. If it's a paranoia hold-over
from the original lamassu-machine and not actually load-bearing on
modern JCM firmware, the setTimeout could be dropped entirely.

Suggested work

  1. Measure first. Port the escrow watchdog landing in PR #45 to the
    id003 FSM (packages/hal/src/validators/id003/id003-fsm.ts).
    Deploy to one Sintra (Atitlan?), collect a week of journals, see
    whether escrow → vendValid is ever over ~250 ms in practice.
  2. Trace the rationale. Find the original commit / upstream
    reference for "Timeout to prevent interleaved commands." If no
    protocol justification surfaces, remove the setTimeout and make
    _send synchronous (with the polling stop/start brackets kept).
  3. Optional: if the setTimeout is load-bearing for some firmware
    versions, gate it behind a per-model flag rather than always-on.

Scope explicitly out

This isn't urgent. We have zero field reports of id003-related
bill tears on the Sintras and Douros — the only torn bills came from
the BATM3, which is the EBDS code path. This is preventive work to
make sure we'd actually catch a similar issue if it surfaced.

References

  • Root-cause fix for BATM3: PR #45
  • Sintra perf context: per-cwd memory sintra_perf_constraints.md
    (Aaeon UP Board scheduler constraints)
  • id003 driver: packages/hal/src/validators/id003/index.ts

🤖 Filed via Claude Code

> _Migrated from [aiolabs/lamassu-next#46](https://git.atitlan.io/aiolabs/lamassu-next/issues/46) — opened by @padreug on 2026-05-19._\n\n## Context While investigating the MEI/EBDS bill-tear root cause on the BATM3 ([PR #45](https://git.atitlan.io/aiolabs/lamassu-next/pulls/45)), the id003 driver was inspected for an analogous timing concern. The MEI tear is fixed by polling at 100 ms inside the validator's escrow grace window. id003 already polls at 100 ms — but it has a separate concern that warrants a focused audit before we declare the JCM iVIZION path equally safe. ## The concern `packages/hal/src/validators/id003/index.ts:276-289`: ```ts private _send(command: string): void { if (this.disablePolling) { this.rs232?.send(command) return } this._stopPolling() // Timeout to prevent interleaved commands setTimeout(() => { this.rs232?.send(command) if (!this.disablePolling) { this._startPolling() } }, POLLING_INTERVAL) } ``` Every host command is delayed by `POLLING_INTERVAL` (100 ms) before the serial write actually fires. Functionally this is a defensive pause to "prevent interleaved commands" — but it means a host `stack()` call arrives at the validator **100 ms after** we decide to stack. For id003 / JCM iVIZION the bill is mechanically gripped and stationary during the host-decision phase, so this 100 ms is normally harmless. But: - We have no measurement of how often escrow → stack actually closes within 100 ms vs how often it slips longer (no equivalent of the EBDS escrow watchdog landing in PR #45). - Under high host load (e.g. concurrent state-machine evaluations, large Vue re-renders, GC pauses on the Aaeon UP Board — which `memory/sintra_perf_constraints.md` notes is constrained) the effective delay could be 100 ms + scheduler latency. - If the JCM ever enters its own internal escrow grace timeout (the spec describes one, default behavior varies by firmware), we'd be in the same family of failure as the MEI. The "interleaved commands" rationale is also undocumented — it'd be worth tracing back where that came from. If it's a paranoia hold-over from the original lamassu-machine and not actually load-bearing on modern JCM firmware, the `setTimeout` could be dropped entirely. ## Suggested work 1. **Measure first.** Port the escrow watchdog landing in PR #45 to the id003 FSM (`packages/hal/src/validators/id003/id003-fsm.ts`). Deploy to one Sintra (Atitlan?), collect a week of journals, see whether escrow → vendValid is ever over ~250 ms in practice. 2. **Trace the rationale.** Find the original commit / upstream reference for "Timeout to prevent interleaved commands." If no protocol justification surfaces, remove the setTimeout and make `_send` synchronous (with the polling stop/start brackets kept). 3. **Optional:** if the setTimeout is load-bearing for some firmware versions, gate it behind a per-model flag rather than always-on. ## Scope explicitly out This isn't urgent. We have **zero field reports** of id003-related bill tears on the Sintras and Douros — the only torn bills came from the BATM3, which is the EBDS code path. This is preventive work to make sure we'd actually catch a similar issue if it surfaced. ## References - Root-cause fix for BATM3: PR #45 - Sintra perf context: per-cwd memory `sintra_perf_constraints.md` (Aaeon UP Board scheduler constraints) - id003 driver: `packages/hal/src/validators/id003/index.ts` 🤖 Filed via Claude Code
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
aiolabs/bitspire#46
No description provided.