Consume operator cassette operations instead of counts #106
1 changed files with 36 additions and 18 deletions
docs(adr): record that the operations wire shipped
Decisions 1 through 4 are built on both sides, so the status section says what landed rather than what is planned, and decision 4a is marked superseded. 4a is kept rather than deleted. It is the calibration for what a warning dialog is worth: it depended on an operator reading it at the end of a refill round, and it could not help at all when the stale value was already in the form. The dialog and the endpoint behind it are both gone now, which is a stronger guarantee than any wording could be. Also records the one gap left open on purpose. An operation recorded while the relay is unreachable waits for the operator's next action, because only an operator action triggers a publish.
commit
65f5f982ec
|
|
@ -65,21 +65,23 @@ The state document carries `applied_ops`, so the dashboard can render each publi
|
|||
as applied or pending. This supplies the feedback leg a replaceable event cannot, without
|
||||
needing the transport to report failures.
|
||||
|
||||
### 4a. Until decisions 1 to 4 ship, the overwrite is warned about, not prevented.
|
||||
### 4a. Superseded. Before decisions 1 to 4 shipped, the overwrite was warned about.
|
||||
|
||||
The dashboard's publish dialog already states the failure plainly — that the publish will
|
||||
overwrite the ATM's tracked counts, that decrements since the last baseline will be lost, and
|
||||
that it should follow a physical refill rather than a mid-day tweak. It also says v2
|
||||
reconciliation will replace it.
|
||||
The dashboard's publish dialog stated the failure plainly — that the publish would overwrite
|
||||
the ATM's tracked counts, that decrements since the last baseline would be lost, and that it
|
||||
should follow a physical refill rather than a mid-day tweak.
|
||||
|
||||
Recording this because it changes how the gap should be read. It is a known, deliberately
|
||||
accepted risk carrying a human-factors mitigation, not an oversight, and the product had
|
||||
already reached the same conclusion these decisions formalise. It is worth keeping in mind
|
||||
that a warning is the weakest control available: it depends on an operator reading a dialog
|
||||
at the end of a refill round, and it cannot help at all when the stale value is the one
|
||||
already in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53
|
||||
while a form loaded at 54 stayed open, and nothing but that dialog stood between the operator
|
||||
and discarding the decrement.
|
||||
Kept here rather than deleted, because it is the calibration for how much a warning is worth.
|
||||
It was a known, deliberately accepted risk carrying a human-factors mitigation, not an
|
||||
oversight, and the product had already reached the same conclusion these decisions formalise.
|
||||
It was also the weakest control available: it depended on an operator reading a dialog at the
|
||||
end of a refill round, and it could not help at all when the stale value was the one already
|
||||
in the form. Confirmed live on 2026-09-22 — a dispense moved a bay from 54 to 53 while a form
|
||||
loaded at 54 stayed open, and nothing but that dialog stood between the operator and
|
||||
discarding the decrement.
|
||||
|
||||
The dialog and the endpoint behind it are both gone. The operator dashboard no longer has a
|
||||
field that accepts a count, which is a stronger guarantee than any wording could be.
|
||||
|
||||
### 5. Ordering is decided by `created_at`, never by arrival order, on both sides.
|
||||
|
||||
|
|
@ -138,13 +140,29 @@ eight consecutive restarts, three heartbeat republishes carried strictly increas
|
|||
back off the relay, and the machine, the relay and the operator dashboard agreed on the counts
|
||||
with timestamps correlated to the second.
|
||||
|
||||
Decisions 1 through 4 are the v2 operations wire and are not yet built. Decision 4a describes
|
||||
what stands in for them meanwhile.
|
||||
Decisions 1 through 4 are the v2 operations wire and shipped in spirekeeper#46 and
|
||||
bitspire#106.
|
||||
|
||||
On the operator side there is no longer any endpoint that accepts a count: the absolute
|
||||
publish, its CRUD write and its request model were removed rather than deprecated. The
|
||||
dashboard records operations and renders each as applied or pending from the machine's
|
||||
`applied_ops` echo. On the machine side, schema v13 adds a `cassette_ops` dedup ledger, the
|
||||
`created_at` watermark on this path is retired in favour of per-op ids, and the state document
|
||||
carries `schema_version`, `seq` and `applied_ops`.
|
||||
|
||||
The wire shapes are those given under decisions 1 and 4 above.
|
||||
|
||||
Cutover for v2 is strict, no compatibility code: spirekeeper deploys first, machines follow on
|
||||
their nightly pull. During that window a not-yet-updated ATM ignores an ops payload, so an
|
||||
operator refill does not land until it updates — which fails safe, since the machine
|
||||
under-counts and will not dispense bills it believes it lacks.
|
||||
under-counts and will not dispense bills it believes it lacks. In the other direction an
|
||||
updated machine drops a v1 absolute-count payload on the missing `ops` array, which is the
|
||||
same safe direction: the machine keeps the counts it is now the only writer of.
|
||||
|
||||
One gap stays open deliberately. An operation recorded while the relay is unreachable waits
|
||||
for the operator's next action to be published, because only an operator action triggers a
|
||||
publish. The window makes that self-healing once anything is published, but nothing on the
|
||||
operator side republishes on its own. An operator-side heartbeat is the fix; it is not built.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue