diff --git a/docs/adr/004-cassette-state-synchronization.md b/docs/adr/004-cassette-state-synchronization.md index e82e7f9..ebe35d0 100644 --- a/docs/adr/004-cassette-state-synchronization.md +++ b/docs/adr/004-cassette-state-synchronization.md @@ -65,6 +65,22 @@ The state document carries `applied_ops`, so the dashboard can render each publi as applied or pending. This supplies the feedback leg a replaceable event cannot, without needing the transport to report failures. +### 4a. Until decisions 1 to 4 ship, the overwrite is warned about, not prevented. + +The dashboard's publish dialog already states the failure plainly — that the publish will +overwrite the ATM's tracked counts, that decrements since the last baseline will be lost, and +that it should follow a physical refill rather than a mid-day tweak. It also says v2 +reconciliation will replace it. + +Recording this because it changes how the gap should be read. It is a known, deliberately +accepted risk carrying a human-factors mitigation, not an oversight, and the product had +already reached the same conclusion these decisions formalise. It is worth keeping in mind +that a warning is the weakest control available: it depends on an operator reading a dialog +at the end of a refill round, and it cannot help at all when the stale value is the one +already in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53 +while a form loaded at 54 stayed open, and nothing but that dialog stood between the operator +and discarding the decrement. + ### 5. Ordering is decided by `created_at`, never by arrival order, on both sides. The ATM forces each stamp strictly above its last published one, so a same-second publish or a @@ -108,10 +124,22 @@ one remembered event id, and never compared `created_at` at all. In practice tha | A drained machine advertising bills it had already dispensed | found in audit | | A re-delivered A, B, A applied three times | found in audit | +The last of these is the only one decisions 5 to 8 do not close, because it is not a defect +in the mechanism: the operator is permitted to write the count, so a stale write is +indistinguishable from an intended one. Only decisions 1 to 4 remove it, by removing the +operator's ability to write counts at all. + ## Status of implementation -Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format. -Decisions 1 through 4 are the v2 operations wire and are not yet built. +Decisions 5 through 8 shipped in bitspire#104 and spirekeeper#44, on the existing wire format, +and were verified against the deployed code on sintra on 2026-09-22: a zeroed machine reported +drained rather than freezing its beacon, a re-delivered operator config was dropped as stale on +eight consecutive restarts, three heartbeat republishes carried strictly increasing stamps read +back off the relay, and the machine, the relay and the operator dashboard agreed on the counts +with timestamps correlated to the second. + +Decisions 1 through 4 are the v2 operations wire and are not yet built. Decision 4a describes +what stands in for them meanwhile. Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an