A web service rollback is an undo – route the traffic back to the old version and, once the error rate drops, you are done. A robot rollback is disenfranchising old decisions – you flip the manifest pointer back, but that version still has a string of in-flight chunks, a queued async inference, and un-zeroed hidden state clutching control and refusing to let go. This article is about exactly that gap.
The previous article stood the skeleton up: the six-layer runtime stack, the three core contracts, and a minimal closed loop that runs and tests, with the conclusion that “swapping the route, the sensor, or the robot is demoted from open surgery to plugging in a component.” But between “it can be plugged and unplugged” and “you dare to pull it out and plug it back in” there is still a whole unwritten layer – the moment it comes out, who catches you?
Swap a board and a bad solder joint can be reworked; swap a set of policy weights while the robot is carrying a glass across the room. It is told in the same style as 9/17 – get the concepts to a discussable state first, then a sister piece actually runs the release state machine with a pure-stdlib fake implementation (seventeen invariants; 17 passed) – see Running the Release State Machine. What sets this apart from ordinary DevOps can be said in one line: it treats deployment as part of the robot’s control safety boundary, so the whole article is really one chain –
release_id -> compatibility -> shadow evidence -> canary -> epoch barrier -> rollback
identity can it install does it move once installed whom to run old decisions lose power retreat if it breaks
If identity doesn’t line up, there is no reviewable ledger behind it; if compatibility is never tested, installing is running with a fault; without a shadow ledger, promotion runs on nerve; without the epoch barrier, a rollback is the old and new versions fighting over the same robot. Each section takes one link.
One word-discipline that runs through the whole piece – set the ruler up front, because this article keeps crossing three layers. I split every claim into three strengths: what the demo implementation guarantees (behavior actually produced by this stdlib state machine’s current code path), what the architecture requires (must hold by design, not necessarily realized in this version), and what production must enforce (only holds once you cross into the registry / type system / concurrency). The seventeen pytest going green in the sister piece prove local invariants are self-consistent on this minimal model – it does not mean production deployment safety has been established. Reading “design intent” as “code guarantee” is the easiest trap in this kind of article, so I keep them apart wherever I can.
Set One Principle First: Deployment Is Not an Action, It Is a State Machine with Evidence
Deploying a web backend is roughly “put the new code up, watch the error rate.” Transplanting that intuition onto a robot is dangerous, and dangerous on three levels. First, the rollback window is not symmetric: a web rollback undoes; a robot rollback recovers from a physical state in which a wrong action has already been executed – the glass has already fallen. Second, the observation surface is not symmetric: error rate and latency curves are centralized, but whether a robot is “right” is scattered across every episode on every machine, and by the time you see the on-site incident the samples were long since consumed. Third – and this is what the article keeps returning to – one deployment changes more than code: bump the ckpt by one version and the normalization statistics, the ObsTransform, the calibration, and the control thresholds are often all in motion; these “invisible versions” raise not a single line of error, they just let behavior quietly drift.
So the correct model of deployment is not an action but a state machine plus an evidence chain: every version move must leave a reviewable credential (who built it, against which schema, which checks it passed, whether it was compared head-to-head with the old version and diverged), and state may only move along legal edges – illegal moves are structurally rejected rather than left to procedural self-discipline. Rings a bell? This is exactly the rule 9/17 set for the runtime – command authority has a single writer, the failure state machine has a single owner, the epoch barrier – now carried, unchanged, onto the release channel. The better the seams the architecture article designs, the cheaper the operations in this one; that is the real coupling between the two.
Release Identity: Every Field in the Manifest Is a Fool-Proof Switch
9/17 gave a deploy/artifact_manifest.yaml: checkpoint, code_commit, schema version, normalization statistics, calibration file, container digest. Back then its job was “reproduce one deployment”; here it takes on a second job – serving as the input to the release state machine. The fields themselves are nothing new; what’s new is the gate configured for each of them:
# deploy/release_manifest.yaml -- the ops-hardened form of artifact_manifest
release:
release_id: 2026.09-r7 # release-event identity: the alignment axis of the event stream (monotonicity enforced by the registry/CI issuer, not validated on this box)
checkpoint: ckpt/diffusion_v17.pt
artifact_sha256: 9b41... # weight hash: tags change, hashes don't
code_commit: 8f3a2c1
signature: minisig:r7.ok # provenance signature: covers canonical(manifest - signature), not just the artifact
compat:
state_schema: [v3, v4] # declared supported schema interval: must lie entirely within the runtime's accepted range (containment, not intersection)
action_schema: [v2]
obs_fingerprint: fp-v7 # combined hash of the semantic fingerprint + golden vectors (9/17): an evidence identifier, not a correctness proof
config_range: [3.1.0, 3.2.0] # a version-level compatibility gate: both min/max take a stance, too-old can also be incompatible; field-level schema/semantics is the config validator's job
runtime:
container_digest: sha256:71c0...
python: "3.11"
policy: # release policy enters the manifest: promotion gates are data, not constants scattered in code
baseline_release: 2026.09-r6 # compare head-to-head with whom
divergence_budget: 0.02 # promotion gate: upper bound on divergence rate
min_ticks: 10000 # promotion gate: minimum sample size
In code this manifest is not one flat pile of fields but split into three layers, a one-to-one match with the YAML’s three sections: ArtifactIdentity (hash + signature) handles “which part is this and who built it”, CompatibilityContract (schema range + config range + obs fingerprint) handles “can it run together with the current runtime”, and ReleasePolicy (min_ticks + divergence_budget) handles “how much evidence must accumulate before promotion is allowed.” The layering is not fastidiousness – it lets boot_check, promote_allowed, and rollback each read only the section they should, and the promotion gates are taken from the manifest’s policy fields rather than three magic constants scattered through the state machine.
The three identity fields each block a different class of accident, and it matters to see that they are not the same thing. artifact_sha256 is the content identity, guarding against “right name, wrong thing”: diffusion_v17.pt is a name, not an identity, and same-name-different-content ckpts are a regular at every “what on earth, the behavior is different” crime scene. signature is the provenance identity (authenticity / provenance), guarding against “the content wasn’t swapped, but it’s unclear who signed it” – the key point is that what the signature should cover is the entire canonical byte string canonical(manifest - signature), not merely the artifact’s hash; otherwise you get the bypass where “the hash didn’t change, but someone quietly edited the schema range in compat.” release_id is the release-event identity, guarding against misaligned event streams: operational events (deployment, comparison, canary, rollback) span three clock domains – CI, the registry, and every robot – and without a monotonic sequence number there is no reviewable ledger. obs_fingerprint guards against definition drift: in 9/17 it watches train/serve skew, and here it watches release/serve skew – a new ckpt wired to old preprocessing, not one field missing yet the whole behavior inverted, is caught on the spot by per-element numerical comparison. The traceability requirement of the evaluation protocol (9/12), when it lands in the deployment layer, looks exactly like this one manifest plus these gates. One sentence nails the boundary: what is described above is what the architecture requires the gate to look like; in the demo state machine in the sister piece, artifact_sha256 and signature are only carried manifest metadata – boot_check neither recomputes sha256(actual artifact bytes) to re-verify the hash nor does canonicalization + signature verification. In production both of those must happen before the manifest enters the state machine, at the registry / artifact-loader boundary, not inside this machine. release_id’s “monotonic” is the same story: it is an external ordering key for aligning event streams, monotonicity enforced by the registry/CI issuer – this state machine does not even check r7 > r6.
There is another set of easily-confused identity scales worth nailing down once here, because each solves a different alignment problem: release_id (which release), robot_id (which machine), episode_id (which job), epoch (which barrier-delimited segment of this job’s lifecycle), sequence_id (which decision within the segment). In particular, be clear that release_id != epoch: one release can span many episodes, and a single episode may bump the epoch several times because of a rollback. During an incident review they are nested coordinates – release r7 / episode 42 / epoch 8 / sequence 183 – and once you roll back it jumps to release r6 / episode 43 / epoch 9 / sequence 0. Miss any one layer and you cannot precisely point in the logs at “which decision knocked the glass off.” Push one more step toward production and the event stream should also carry config_version and policy_instance_id: the same release_id may be running with different thresholds (thresholds are release artifacts too), and under shadow / async inference a single release may more likely hold several candidate policy processes at once – release r7 / epoch 8 cannot tell apart “was it instance A or instance B that made this decision.” release_id is not enough to uniquely identify a running policy instance; if an incident review must pin down to the instance, these two fields have to enter the schema up front, not be bolted on after the fact.
One more step: put the three scales the command admission actually reads side by side – they are often lumped together as “just sequence numbers,” but each minds a different job and none substitutes for another:
| Field | Meaning | What it carries in command admission |
|---|---|---|
release_id | who produced it (provenance) | the release-identity barrier: a stale / missing release loses power |
epoch | which lifecycle it belongs to (authority generation) | the lifecycle barrier: an old epoch’s in-flight decisions lose power |
sequence_id | which decision within that lifecycle (ordering) | an ordering constraint; compares ordering only within one segment |
This answers a question reviewers keep asking: why does admission key on epoch + release and not on sequence? Because sequence_id is policy-local – in a shadow comparison what you diff is delta(sequence), not the absolute value, and after a rollback it restarts from 0. It only guarantees “decision #183 comes after #182 within the same segment”; it does not authenticate cross-release provenance: an old-release decision with a big sequence number whose epoch happens to line up is exactly the fish that gets in if you look only at the sequence (that is precisely what I16 plays out). So ordering belongs to sequence, provenance to release, lifecycle authority to epoch – the three do not replace one another. And while we are at it, nail down what epoch actually is: it is a lifecycle / authority generation, not a version number. The schema-level invariant is that a release_id change does not force epoch + 1 (one release spans many episodes, each with its own epoch), and epoch + 1 does not force a release change (a single release may bump the epoch twice during a rollback). Coordinates like r7 / epoch 8, r7 / epoch 9, r6 / epoch 10 are all legal; treat epoch as a version pointer and you will lose the entire class of history where “the same release rebuilt its lifecycle several times.”
The Compatibility Grid: Not Backward Compatibility, but “Which Cells Can Run Together”
The most common wrong assumption in version management is linearity: new code is compatible with old data, old code with new data, backward all the way. What really moves independently and is not orthogonal to the rest is the three runtime compatibility axes – schema (the state/action contract), config (thresholds/parameters), and, from 9/17’s toolkit, obs_fingerprint (semantic fingerprint + golden vectors, guarding against the preprocessing definition quietly drifting). As for code_version: it moves too, of course, but it fits better as release-artifact identity (artifact_sha256 + code_commit) kept by the registry than as an axis judged on continuous intervals – code compatibility is very rarely expressible as “fully compatible within some version range,” so this article’s boot_check carries code_version yet never compares on it (see the code in the sister piece). The grid below is spread out along those three axes:
schema v3 v4 v5
runtime ────────────────────────────── ← each cell = one real boot check that actually ran
1.4.x ✓ ✓ ✗ ✗
1.5.x ✗ ✓ ✓ ✗
config ──────────────────────────────
3.1 ✓ ✓ ✗ ✗
3.2 ✗ ✓ ✓ ✗
↑ an untested cell defaults to ✗
↑ (this grid's "✓ + evidence reference" is held by the registry/pipeline; this minimal state machine does not persist it)
Three disciplines. One, the grid records only combinations that genuinely existed. Teams that insist on maintaining a full compatibility matrix end up testing none of it. The typical truth of a robot fleet is: on any one machine, the runtime lagging the registry by a version is the norm – so the compatibility interval must be modeled explicitly (the [v3, v4] under compat), not left to “should be fine, right?” Two, an untested cell defaults to ✗. Compatibility is a property you test into existence, not one you declare; behind every ✓ there must be a real boot_check that ran plus a round of shadow comparison. This sentence needs a follow-up about a real-world pit: whose head does that ✓ actually sit on? If the ✓ is just a boolean someone hand-typed into YAML in the registry, it will sooner or later drift from reality. The trustworthy approach is to have the ✓ carry an evidence reference (which CI run, which shadow round, which test suite ran), written by the pipeline and never edited by hand – and this is exactly the interface the next article on fleet version governance will pick up. Three, config too must be forward-compatible. During a canary rollback, the old runtime often has to load config the new version already changed (episode 42); a field renamed or its semantics shifted crashes the old code on the spot – the easiest path to patching in a second fault the very night you roll back. So config changes follow the standard expand-migrate-contract: only add, never delete; give new fields defaults; put deletion after every machine has upgraded; the config_range in the manifest, with both a min and a max, is the enforcing gate on this discipline – too old and out of bounds are equally suspect.
One modeling confession: compat_within treats versions as monotonic continuous semantic intervals and does a containment check (the entire declared range must lie within the runtime’s accepted range), which holds when the schema really evolves linearly. But in the real world versions are often discrete capabilities – supporting {v3, v4} does not mean supporting some pseudoversion in between with different fields. For that case you would swap this for an intersection over capability sets, not intervals. This article chose the interval model for a minimal closed loop and left that seam open in a code comment; if your schema is a set of discrete IDs, please replace “interval containment” with “set containment.” Confess one layer deeper: this uses a custom numeric-tuple version (int.int.int) and deliberately does not call it SemVer – 1.9 < 1.10 it can compare, and parsing accepts 1~3 segments and right-pads to three, so "2.0" equals "2.0.0" and "1"/"1.2" normalize to 1.0.0/1.2.0; the ordering no longer scrambles when the segment count changes, so Python tuple comparison will not wrongly declare "1.2" < "1.2.0". But it is strict about the format: the moment there are more than three segments ("1.2.3.4") or any non-digit segment ("-1.2.0", "1.2.x"), it raises ValueError – the earlier parts + [0]*(3-len(parts)) would silently let "1.2.3.4" through, because [0]*negative == [] (this is exactly what review §3 fixed). It still does not handle -rc1/+build7 pre-release/build metadata. To actually judge SemVer you must handle that metadata or hand it to a mature version library; the interval model here is only trustworthy when the “all versions are plain numeric and at most three segments” guarantee holds externally. Note that I deliberately used containment, not intersection: the manifest declares support for [1.2, 2.0], the runtime only eats [1.8, 3.0], they intersect but must never be judged compatible – because the runtime doesn’t recognize the 1.2~1.8 half the manifest promises, and an intersection test would let it through the gate.
The startup validation order is also worth nailing down, from cheap to expensive: first structure – are the manifest fields complete, is there a release_id, is the schema range fully covered by the runtime, does the config version cross the bounds; pure comparison, milliseconds (note this layer does not even recompute the artifact hash, let alone verify a signature – that is the loader’s job before the manifest enters the machine). Then the fingerprint – compare whether the obs_fingerprint matches. Say this one precisely: what it proves is only that “the release end and the runtime end are using the same fingerprint definition / the same evidence identifier” – it is an evidence identifier, not a correctness proof; equal fingerprints do not mean ObsTransform(x) was actually computed right, only that both sides agree on “which hash represents the preprocessing.” Then the numbers – this is the real correctness gate: take the golden vectors in the release manifest and actually run the shared ObsTransform on them, comparing each real output against its expected output line by line (golden vectors -> actual transform -> expected output -> compare), and refuse on any mismatch. The fingerprint gate and the golden-vector gate are two different gates – don’t mistake the former for the latter. Last, behavior – enter shadow mode for a head-to-head comparison (next section). Rather refuse to boot than run with a fault: fail-before-activation has been said ten thousand times, and every defeat comes down to “let’s run it and see.” What this article’s fake can actually prove is only the “refuse to load” step (boot_check returns a reason and does not proceed); the production version must push the same discipline all the way to “refuse to take effect” – an illegal config never takes over at a control boundary. (The earlier wording “fail-before-motion” I put too heavily on; strictly the evidence only reaches before activation – see I13 in the sister piece.)
Shadow Mode: Compare Decisions, Never Touch the Ground
The cheapest insurance in a release system is to let the new policy see everything and touch nothing. On the same StateContract stream, the old and new policies each produce an Action: the old one goes, as usual, through SafetyGate -> CommandSink to real execution; the new one goes into the shadow ledger:
StateContract stream ──→ active policy ──→ ActionBuffer → SafetyGate → CommandSink → real robot
└──────────→ candidate policy ─→ divergence ledger (bookkeeping only, no path of its own in the sink)
The implementation discipline of shadow mode is a reuse of two of 9/17’s old rules. First, authority is unique: the candidate’s output has no path to the RobotInterface – precisely, under a dynamic type system like Python this is not “it can’t get in at the type level” but “there is no such call path structurally”: inside ShadowRunner.tick() the return value of candidate.act() flows only into the divergence ledger, and the only place that calls sink.submit() is fed the product of active through safety.check(). Type-level enforcement (splitting CandidateAction and Action into distinct types and letting the sink accept only the latter) is what the production version should add; we do not claim it is already done here. Second, comparison must be same-state, same-clock: both are fed the same state object from latest_valid(now) on the same tick, under the same injected clock; divergence computed by a shadow run on a stale state is all noise.
The ledger should not record just one number. Divergence is a matter of more than one dimension: identical numbers != identical behavior. The real danger is often not 0.50 vs 0.52 (that may be noise) but the same action’s valid_from being 100ms late – that is another chunk, taking over at the wrong physical moment. So Divergence is split into five mutually independent dimensions: value (numerical value over ε), validity (valid_from takeover-moment misaligned – governs “when it starts to take over”), horizon (chunk coverage horizon×dt inconsistent – governs “how far one action reaches”), sequence (takeover rhythm, whether the sequence_id step matches), and safety (clamping difference). validity and horizon are two different things – do not merge them into one axis, or any horizon change will incidentally light up validity too and the dimensions collapse back to four. The sequence dimension has a contract to nail down first, too: this article picks a policy-local sequence – active and candidate each hold their own SequenceAllocator, the absolute values naturally never align, so only the step delta is comparable; if your semantics is a globally shared sequence, the check should instead be an absolute a.sequence_id != c.sequence_id. The two contracts cannot be mixed – prose and code must pick the same side. The one most easily missed is the last, and the one most worth watching: the clamping difference – for the same command, SafetyGate releases the old policy but clips the new one; even if the task success rate shows nothing yet, that is direct evidence the new policy is more aggressive. In implementation, both active and candidate pass through safety.preview() (judge only, don’t send), and by comparing the two clamped flags clamp_diffs really counts; the candidate sees the safety gate’s verdict yet still cannot reach the sink – which demonstrates “can observe, cannot execute” at the same time. This invariant can no longer be held by luck: the old writing had preview() directly reuse check(), and “shadow produces no side effect” only held by accident, because check() happened to be a pure function – a real SafetyGate carries state mutation, counters, watchdogs, and rate-limit bookkeeping, and once check() has side effects a candidate’s preview quietly mutates the gate’s active-side state, breaking the authority boundary on the spot. The fix collapses the pure judgment into a single source evaluate(): preview() reads only evaluate() and never touches commit state, while check() = evaluate() then commit (writes last_command, bumps clamp_count). That is how “candidate never gets authority” is upgraded from a code-path property to an interface property – the PreviewSafety in the sister piece is split exactly this way. Divergence rate is the risk reading; a canary that ignores it is running bare – but say it clearly too: divergence_rate is an aggregate count (any of the five dims firing records one divergence), collapsing “one clamp” and “one 1e-5 numeric difference” into an equal-weight vote loses risk weighting; a production criterion should at least split safety/clamp into a hard gate and numeric into a soft gate, or better still compute a rate per dimension.
The shadow stage has two failure conditions to guard against in advance as well. The candidate must not touch state: the shadow policy is forbidden from writing anything back to the estimator / StateBuffer, otherwise “observe only” becomes co-decision. Comparison must pick the boundary cases: on a uniform timeline the two policies agree most of the time, of course – the divergence ledger must be stratified by the task labels declared in the manifest (contact, occlusion, calibration drift); a comparison window that fails to cover the boundary cases is a comparison that didn’t happen.
Canary and Promotion Gates: The Divergence Ledger Has the Say
Only after the head-to-head comparison passes does real execution come into play. A canary in the robotics setting is a different thing from a percentage of web traffic – the natural slices of traffic are not users but tasks, robots, and time windows: shadow first, then a single task on a single machine, then a same-model fleet, then cross-model; every step is “real execution + a fault budget,” differing only in blast radius.
The rules of promotion must be set up front: the gate is data, not courage. The two fields in release_manifest.yaml – min_ticks and divergence_budget – are the hinge of the promotion gate: not enough samples, no promotion (even if the divergence rate is zero); divergence over budget, no promotion (even if the sample is large enough). The most taboo shape is “let’s keep watching”: observation without a quantified gate always slides toward Friday-afternoon “looks fine, ship it.” The rollback criterion works the same way, and moreover must be graded by hardness: frequent safety-gate intervention, divergence crossing a hard threshold, is immediate rollback, no discussion; slowly degrading coverage or success rate is in-budget rollback, counted per window. Both kinds of criterion must be written as computable expressions before the release, not arrived at by meeting for consensus afterward. Another class easily lumped into one pot is “observability failure” – this article’s Fault.OBSERVABILITY keeps only one cell for the minimal closed loop, but production should split it into at least two: a metrics gap (OBSERVABILITY_DEGRADED: the eyes are temporarily blind, the control loop is still running safely, freezing promotion suffices) and a ledger write failure (EVIDENCE_LOSS: the system loses the ability to prove itself safe, the promotion evidence is broken, “freeze promotion” alone may not be enough – the candidate might have to be quarantined outright). The two are not the same severity; folding ledger-write-failure into “observability” is a simplification of this article’s fake, worth nailing down explicitly in a real deployment policy.
One scope note, so readers aren’t misled: the pytest suite in this article verifies the phase constraints and evidence structure of the release state machine – which states may legally transition between, whether promotion requires enough evidence, whether rollback must pass the gate again. It does not simulate a fleet-level canary scheduler: machine selection, task mix, and auto-promotion at the end of an observation window are the domain of the fleet registry and belong to the next article. The CANARY that appears here is only a phase name guarded by the state machine, not a scheduler. Likewise, do not overestimate promote_allowed(): it implements only two minimal gates – min_ticks (sample size) plus an divergence_budget over the aggregate divergence rate; the safety hard gate (clamping / e-stop frequency), task-stratified coverage, success-rate and P95 degradation, and fleet-level policy evaluation named in the body are all outside this demo, the work of the next scheduler layer. What this article can hold for you is the discipline that “promotion needs a computable gate”; as for which dimensions the gate should read, the code only opens a head.
Following this “a legal edge ≠ clearance” line, let me once and for all decompose the act of “returning to ACTIVE.” There are actually three “come-back” edges in the state machine; they look similar but eat entirely different evidence – and this is exactly the root cause for splitting ReleaseStateMachine (judges topological legality) from PromotionController (judges evidence) into two layers. In the demo request_transition() only looks at edges and promote_allowed() only looks at gates, and the two are currently unwired (the state-machine docstring earlier flags this clearly as option A, “keeping the model minimal”); a production implementation must have the PromotionController wire the evidence gate onto state transitions first, and callers must never bypass the gate to call request_transition("ACTIVE", "promote") directly. The three edges’ semantics and strictness increase in order:
| Source state | To return to ACTIVE / continue, what it needs | Where this article’s demo pins it |
|---|---|---|
CANARY | promotion evidence: enough samples + divergence within budget | promote_allowed()’s two minimal gates (I6); but not wired into the state transition |
DEGRADED | recovery evidence: health back, divergence back within budget, shadow evidence fresh | only the DEGRADED → ACTIVE legal edge; the recovery condition is not implemented, it is an architecture requirement |
SAFE_STOP | full re-verification: no going straight back to ACTIVE, must return to VERIFYING and rewalk from the start of the evidence chain | SAFE_STOP → VERIFYING, the sole out-edge (I4) – the strictest of the three |
To align the three in one sentence: canary is promoted by data, degraded comes back by recovery evidence, and after SAFE_STOP there simply is no “straight back to ACTIVE” path. Writing these three as layered guards (promote_allowed / recover_allowed / SAFE_STOP→VERIFYING), rather than a single if state_ok: go ACTIVE, is what actually honors the phrase “order is semantics.”
Rollback: 9/17’s Epoch Barrier Goes On Duty a Second Time
Rollback is hard, and not because of switching the config. The config flips in a second; what floats out afterward is the real trouble – in-flight things: chunks produced by the old version still in the ActionBuffer (a copy still lying in the scheduled slot), an async inference queued in the policy service, a trajectory segment mid-execution, and the policy’s hidden state. A rollback that only flips the manifest pointer is equivalent to letting the old and new versions co-drive the same robot – the old version’s last decision is still queuing in the buffer while the new version’s first decision has already entered, and their sequence_ids are each counting on their own.
So rollback is not a one-line assignment but a small transaction in which order is semantics, and not one of the five steps may be swapped: ① stop sending new commands, close command authority; ② epoch + 1, which structurally disenfranchises the old version’s in-flight decisions – without waiting for “them to finish,” and without letting them override the post-rollback version; ③ enter SAFE_STOP, issuing a safe-stop command so the robot is never left in an undefined intermediate software state (note this is only software issuing a stop order, not “the physics has actually come to rest” – requested / acknowledged / physically-confirmed are three different things; see the “Three Boundaries” subsection below); ④ the old manifest must pass boot_check once more; ⑤ only then flip the pointer + run a full reset_episode, and go back to VERIFYING to rewalk the evidence chain. On step ⑤ let me be honest: the demo’s reset_episode() does only two things – set an assertable reset_called flag, and via sync_epoch clear the ActionBuffer’s active/scheduled slots together with the sequence numbers (I8 uses exactly these two to verify “a full reset” is not empty talk); the production version must additionally clear, inside it, the policy’s hidden state, the estimator filters, controller tracking, and the random seed – these the fake does not implement, and that is an explicitly stated boundary, so do not read “cleared the buffer” as “cleared the whole state.” The epoch barrier 9/17 designed was prepared for step ②: after the epoch is incremented and synced to both the policy and the ActionBuffer, every in-flight decision of the old version – however large its sequence_id, however far from taking effect, including the one already queued in the buffer but not yet dispatched – is rejected uniformly at put()’s barrier and cannot be retrieved at current() dispatch either; the sequence space restarts, so there is no cross-version sequence entanglement. “Wait for the old chunk to finish, then switch” is neither needed nor allowed: the barrier disenfranchises the old version, and this is not about being fast, it is about being clean.
Step ④ is the single most easily missed and most lethal line in the article: “an old version” is not “a rollback-able version.” The release you roll back to – its schema, config, and obs_fingerprint may no longer be compatible with the current runtime, especially if you already upgraded the runtime first. Re-run a compatibility check against the old manifest before going back; if it fails, stop in SAFE_STOP and ask for help, never hard-flip: force-switching to an incompatible old version is just replacing an old bug with a new one. In the code this step raises RollbackRejected on failure, the state machine rests firmly in SAFE_STOP, and the pointer does not move a single word.
This also draws out an architectural rule running through both articles: the sole publisher of epoch is RuntimeCore. The supervisor, policy, and ActionBuffer only accept epoch and sync a mirror; no one may do epoch += 1 themselves. Otherwise you get the half-synchronized hell where the supervisor sees 7, the policy is still at 7, and the buffer is at 6 – and rollback is precisely the moment most prone to crash inside such inconsistency. In rollback() the only thing that increments epoch is runtime_core; the supervisor takes a mirror afterward via note_epoch() and writes it into the event, which is this single-owner discipline made concrete. Along the way, a detail easily mistaken for a slip of the hand: in one rollback the epoch goes +2, and that is not a repeated increment. This article defines epoch at each lifecycle boundary, and a rollback happens to cross two of them – epoch N -> N+1 is the invalidation barrier (step ② bump, disenfranchising the old world), epoch N+1 -> N+2 is the new-world start barrier (step ⑤, at install, bumps once more so the sequence restarts from 0). A production event log had better mark the two bumps with different reasons (say rollback_invalidate / rollback_install), or a review will wrongly assume someone pushed one cell too many.
For auditability, I also added a ROLLING_BACK state to the machine: ACTIVE -> SAFE_STOP alone cannot tell apart “a real fault,” “a manual E-stop,” “a rollback,” “a watchdog,” and “a controller timeout.” With an explicit rollback state (and a mid-way failure that can fall back to SAFE_STOP), plus an event on every transition carrying a reason, you can afterward ask “what exactly did this machine go through last night.” And the counterintuitive discipline that the rollback path itself must be tested – a release plan whose rollback has never been executed is equivalent to having no rollback – is, in the sister piece, just one line of pytest. Here too I owe a correction to myself: in a few places in the body I casually said “evidence chain,” but the sister piece’s supervisor.events is, strictly speaking, only an append-only event log – each record carries an evidence reference (which CI run, which shadow round, which reason), not the evidence itself. A real tamper-evident evidence chain has four segments – event log → evidence references → external evidence store → tamper-evident audit (hash chain / signature / WORM storage) – and the last three all live on the registry side, not in this stdlib state machine; the demo neither signs nor hashes, and anyone can mutate the in-memory events. So when you read “evidence chain,” please understand it as “event log + evidence references” and reserve “tamper-evident” for the hardened version – this is one more self-audit of the article’s three-layer word discipline.
After Rollback: Three Boundaries People Keep Missing
Those five steps round out “rollback,” but if you really want them to catch you, there are three boundaries that are easy to mistake “the code has no such path” for “structurally impossible” – and a demo is exactly where the latter should be made explicit.
Boundary one: the command-admission gate – the door every safety invariant hangs on. Rollback step ① says “stop sending new commands” and step ② says “old-epoch decisions are disenfranchised,” yet if the authority check, the epoch check, the release check, and the sequence check are scattered across several ifs before and after put(), then any future path (an async callback, a candidate promoted to active, a debug backdoor) that skips one of those ifs leaks the barrier. So collapse them into a single entry point: decisions no longer call buffer.put() directly but first pass an admission gate, and only a fully green verdict lets them into the buffer. There is a trade-off here that is easy to overlook: release identity must be a required admission input, not optional metadata – if the gate reads release = getattr(action, "release_id", current_version), then an old command that carries no release_id gets silently relabeled as “the current release” and slips through the identity gate. So this version pulls the identity into an explicit ActionContext(release_id, epoch, sequence_id) that enters the door alongside the decision; a missing release is rejected outright, with no fallback. And it is exactly the “sync active, async late arrival, and candidate-promoted-to-active – all three kinds of decision admit through the same gate” model – active no longer shortcuts straight to sink.submit either; it passes the same barrier a late result does.
active, produced sync by Policy ┐
async late-arriving inference ├──► Command Admission ← single entry; three kinds of decision admit through the same gate
candidate's first run after promote ┘ │ admission input = ActionContext(release_id, epoch, sequence_id)
│ release required: missing release is rejected outright, no getattr fallback to "current version"
▼
authority_open ? rollback ① close the gate: stop new commands
release match ? release-identity barrier: old / missing release disenfranchised
epoch match ? rollback ② lifecycle barrier: old epoch disenfranchised
ctx ↔ action consistent ? sequence / time window handed to the ActionBuffer
│ prepare() checks the gates → commit() re-checks generation, then put()
│ between the two steps it does not pretend "check == commit" is atomic; the generation barrier closes TOCTOU
▼
ActionBuffer
▼
SafetyGate (evaluate = pure judgement / check = evaluate + commit)
▼
CommandSink → real robot
── in shadow, a candidate only calls preview(): it enters neither admission nor sink
(not having authority before promotion is an interface property, not "there just happens to be no such call path")
In the sister piece this gate lands as a CommandAdmission shared by RollbackCore and ShadowRunner: the admission input is an explicit ActionContext(release_id, epoch, sequence_id), a missing release is rejected outright with no fallback; after the four steps authority_open → release → epoch → ctx↔action consistent, prepare() only checks the gates and commit() re-checks the generation before it finally calls action_buffer.put(). The key point is that even active landing now enters through this door – the active decision in shadow now goes through admission.admit() and only then sink.submit, no longer taking a safety.check → sink shortcut. So “a candidate can’t reach the sink” no longer leans on “there just happens to be no such call path,” and “no side door other than active” is no longer a documentation promise either – both become an interface property: you get gated the moment you try the door.
Boundary two: async late arrivals – the epoch barrier holds, but you still need release. The policy service is asynchronous: an inference request can go out at r7 / epoch 1 and only return after the whole rollback has run and the robot has moved to r6 / epoch 3. Is the epoch barrier alone enough? Mostly yes – an old epoch won’t match the new epoch, so it is rejected outright. But there is one fish that slips through: a late result whose epoch happens to equal the current epoch (queue skew, or a rollback that bumped back to the same number). Only the release-identity barrier can stop it – so a decision must carry a release_id alongside its epoch and be compared to the current authoritative release at the door. epoch is the lifecycle barrier, release is the release-identity barrier; their duties do not overlap, and neither can be dropped. The I16 in the sister piece demonstrates exactly this: in the ActionContext that a late decision carries into the door, one whose epoch matches but whose release_id is still r7 is firmly rejected by the release-identity gate; add a decision whose release_id is empty (no identity, hoping getattr falls it back to “the current version”) – it too is rejected, because this version has no such fallback at all. This is where the earlier “identity scales” paragraph’s insistence that release_id belongs in the action schema up front finally gets cashed.
Boundary three: the physical boundary – software disenfranchisement != the world has stopped. The epoch barrier governs three kinds of software-side in-flight: queued (not yet in the buffer), buffered (in the buffer but not dispatched), late-produced (async late arrivals) – it can void all three. But it cannot reach the one already dispatched into the physical actuator: once a command has been accepted by the MCU and a joint is already moving, that is history software cannot “undo.” To actually stop you need controller-level cancel / brake / hold / trajectory abort – an extension of 9/17’s “safety is a separate layer, hardware is the backstop.” Likewise, step ③’s SAFE_STOP means software issued a stop order; keep the three moments requested / acknowledged / physically-confirmed apart – supervisor.state == SAFE_STOP only means the command channel is closed, not that the glass has been set safely back on the table. This boundary is not an epoch bug, it is physics: it is precisely why “robot rollback != web rollback” holds – web undoes traffic, but a robot, after disenfranchisement, still has to wait for the world to genuinely come to rest. The sister piece takes all three boundaries only as far as “state it clearly and pin one assertion each” (I8 pins authority + dispatch-disenfranchisement, I16 pins async late arrival, I17 pins that there is no race window between admission and rollback – using the deterministic prepare() → rollback() → commit() interleave to prove that an old action which passed the gate before the barrier but only tries to commit afterward is turned back by the generation barrier); the physically-confirmed closed loop, concurrent hardening of the admission gate, and passing the release barrier across processes are all explicitly declared production debt.
Hot Reload of Config and Thresholds: Changing a Threshold Is Not Changing Code, but It’s More Dangerous Than Changing Code
Deployment swaps artifacts; what operations swap is often just thresholds: velocity ceilings, validity windows, control frequency. Hot config reload is tempting because it bypasses the entire release channel; it is dangerous because the thresholds are the safety parameters. Three rules: the config version goes into the manifest, and changing a threshold counts as a release too (through the same state machine, no shortcuts); hot reload goes through prepare -> validate -> commit, and on validation failure keeps the old value and records an event – a half-new, half-old state must never be read by the control loop; every threshold is annotated with its effective point (next tick / next episode), and a parameter switch mid-control must happen only at a boundary, like an ActionBuffer takeover. The words “small change” show up absurdly often in incident postmortems. This one, too, is turned into a runnable assertion in the sister piece with a ConfigManager: after an illegal threshold’s commit, current must stay byte-for-byte unchanged. But to be fair – the line self.current = staged inside commit() is, in this demo, only a single-threaded object reference swap. What it demonstrates is the commit discipline of “if validation fails, don’t touch current,” not concurrency atomicity. To truly deliver “the control loop never reads a half-new, half-old state,” production needs an immutable snapshot + atomic pointer swap / generation number / control-boundary latch (or to directly reuse the article’s epoch barrier) so that read and write cannot each catch it halfway – that is architecture-required, and this stdlib fake neither proves it nor pretends to.
Patch It Up to Runnable: A Minimal Closed Loop for a Release State Machine (moved to the sister piece)
Everything above is design. Turning it into a minimal runnable-and-testable form – the full deploy_fakes.py, a test_deploy.py that pins I1-I17 one assertion at a time, and a real run (17 passed) – is long enough to stand as its own piece, so I split it into the sister article Running the Release State Machine: A Pure-stdlib Minimal Closed Loop and Seventeen Invariants. That piece reuses 9/19’s fakes.py (ActionBuffer, the epoch barrier, injected clock, single-writer CommandSink all present and unchanged); this piece keeps only the design and boundaries, while the code and assertions live there. Want to see “which assertion pins each invariant, and how CommandAdmission’s prepare/commit actually drops the gate in code”? Jump straight to that piece; want to first be clear on “why there has to be a gate at all, and why robot rollback != web rollback”? Finish reading this one, then go.
Six Anti-Patterns at the Deployment Layer
Continuing 9/17’s list, here only the pits on the release channel, still ordered by frequency of appearance:
- Same-name, different-thing artifacts: matching versions by filename;
v17_final_v2.pt-style naming is the number-one breeding ground for skew. The fix: a hash for content identity, a signature for provenance identity (and the signature must cover the entire canonicalized manifest), a release_id for event identity – do not conflate the three. One more boundary: this article’s demo only carriesartifact_sha256/signatureinto the state machine as fields; the real hash recomputation, canonicalization + verification happen in the production registry / artifact-loader and never enter this state machine – do not read “the field is present” as “verification done.” - Rollback that only flips config: the pointer flipped, but the old version’s in-flight chunks and async inference are still in the air. The fix: rollback = stop commands + epoch disenfranchisement + pass boot_check again + flip the pointer + a full reset – the five-step order is the semantics.
- “Promote because it looks fine”: observation without a quantified gate always slides toward luck. The fix: write
min_ticks+divergence_budgetinto the manifest’s policy section; the promotion gate recognizes only these two numbers. Mark the boundary likewise: the demo’spromote_allowedeats only these two numbers, and the divergence rate it reads is a one-dimensional aggregate; splitting rates per dimension, a hard gate on clamping/e-stop, and task-stratified coverage belong to production / the next scheduler layer – this article only stands up the discipline that “promotion must have a computable gate.” - The rollback path never tested, the rollback target never checked for compatibility: running the rollback script for the first time on the night of the incident, only to find after switching back that the old version is incompatible with the current runtime. The fix: put the rollback drill in the pipeline, and re-run boot_check on the old manifest before rolling back – an old version is not a rollback-able version.
- Hot-changing thresholds by shortcut: editing on-site parameters directly, bypassing the release channel. The fix: config version into the manifest, hot reload through prepare -> validate -> commit, old value stays byte-for-byte unchanged if validation fails, and changing a threshold counts as a release too.
- A compatibility matrix built on declarations: “supports v3 and above” with v5 never tested, and the ✓ hand-typed into YAML. The fix: a cell that has not run a boot check defaults to ✗, every ✓ carries an evidence reference and is written by the pipeline; compatibility is a property you test, not one you declare.
Summary
9/17 said the skeleton must “be able to swap components”; what this article adds is the layer for the swapping: release identity (a layered manifest – content hash, provenance signature, event sequence, plus the identity scales of release_id and epoch; in this demo the hash and signature are only carried metadata, verification/recomputation is the production loader’s boundary), the compatibility grid (the runtime’s schema × config × obs-fingerprint, containment rather than intersection, code as a release-artifact identity rather than a compatibility axis, with untested cells defaulting to ✗ and every ✓ carrying evidence held in the registry; the fingerprint gate and the golden-vector gate are two different gates), shadow comparison (bookkeeping only, no landing; divergence split into the five independent dimensions of value/time/amplitude/sequence/safety, with clamping differences really counted, and preview collapsed onto a pure evaluate() with side effects kept only in check’s commit branch), canary promotion with gates (data has the say, and it is made explicit that this article does not simulate a fleet scheduler, that promotion eats only two gates, and that a legal edge is not the same as satisfying the promotion/recovery condition; the three “ways back” – CANARY / DEGRADED / SAFE_STOP – each eat different evidence: promotion / recovery / full re-verification), rollback backed by the epoch barrier (all decisions – sync active, async late arrival, candidate-promoted – admit through one shared CommandAdmission, with release_id in the ActionContext required and no getattr fallback; the authority / release / epoch gates are collapsed in one place; stop commands first, then disenfranchise the old decisions – including those queued but not yet dispatched and async late arrivals of an old release or with no identity – pass the compatibility gate again, only then flip the pointer, and the rollback path itself must be drilled; and between prepare() → rollback() → commit() a generation re-check blocks the TOCTOU window, while true concurrent linearization remains production debt), and fail-before-activation hot config reload (commit is atomic, an illegal candidate is rejected before it takes effect; it does not claim concurrency atomicity). Seventeen invariants pin this state machine to a “runs, and is locally self-consistent” degree – but remember the whole article’s three-layer word discipline: what the demo implementation guarantees ≠ what the architecture requires ≠ what production must fill in; the green lights only prove this minimal model is self-consistent, not that production deployment safety is established.
Looking back at the division of labor in this series: the evaluation protocol (9/12) defined what counts as evidence, the architecture article (9/17) defined which seam the evidence grows out of, and this article defined how evidence gates the next release. Put the three layers together, and “swap the route, the sensor, the robot” for the first time becomes a sentence with engineering meaning – pull it out and there is a ledger, plug it back and there is a barrier, every step with its evidence.
For the next step I lean toward the second direction: version governance for a multi-machine fleet. Because this article has already naturally grown the basic primitives a fleet registry needs – release_id + compatibility matrix + evidence + epoch + rollback. Promoting them into a registry / desired-state / reconciliation-loop (N robots × the three axes of schema/config/obs-fingerprint, with code hanging on the edge as a release-artifact identity rather than a compatibility axis; how the registry records the evidence behind each ✓ and how it converges drift) closes the architectural line of the whole series. The first direction (expanding shadow into a full release pipeline: comparison sample mix, boundary cases, the ledger schema) is still on the candidate list. Whichever one you want, tell me in the comments.
Comments