Deployment Release

Deployment Release

Isometric blueprint showing deployment release and safety cases: Goal Structuring Notation argument pyramid in transparent glass tablets, emerald physical evidence pipelines, golden deployment authorization seal, and crimson certified operational envelope threshold with revocation interlocks.

Purpose

Under what auditable chain of physical evidence is an autonomous machine permitted to move among humans?

Releasing an autonomous machine grants software the authority to exert physical forces in unconstrained environments alongside human beings. That deployment authority demands an auditable safety case connecting operational claims directly to empirical evidence from the integrated physical plant. High validation accuracy cannot guarantee that a robot stops within available clearance envelopes, nor can an incident-free test log confirm that emergency fallbacks remain dependable when critical sensors experience silent hardware faults.

A release decision must identify precisely which empirical measurements support each warrant and explicitly delineate the environmental assumptions bounding their validity. Unmodeled payload shifts, floor wear, or thermal drift can instantly invalidate a safety case even when neural weights remain unaltered; where evidence cannot settle an operating premise, refusal is the only acceptable engineering verdict. Within the governance envelope, deployment release consolidates the Brain’s proposals, the Nervous System’s permission path, and the Body’s physical limits into an auditable manifest that revokes motion authority the instant evidence expires.

Learning Objectives
  • Construct a claim-argument-evidence case linking a bounded operating envelope to auditable physical evidence
  • Evaluate whether each claim’s premises are measured, illustrative, or unknown, and select the verdict they support
  • Identify the measurements that would lift a refusal and the restrictions a conditional verdict would keep
  • Specify standing conditions that withdraw operating authority when wear, drift, or configuration changes invalidate a release
  • Diagnose unsupported warrants and compound defeaters through adversarial review of a release argument
  • Design a release manifest that binds its upstream records by digest and refuses motion on any mismatch

The Decision to Deploy

A directory of passing test records and hardware fault traces does not by itself decide whether an autonomous machine may move among people. In a web service, a regression in a new build costs a rollback and some corrupted records. In physical AI, release delegates authority over mass, force, and velocity. Verification has already stressed the machine’s timing, contention, and sensing assumptions under injected faults, and what it hands over is its fault records, each with a predeclared oracle, and the gaps they leave. The engineering task therefore shifts from finding defects to bounding what the evidence licenses. Neither validation accuracy (Closed-Loop Dynamics) nor a run of clean demonstrations (Why Operation Is Not Evidence) meets that task, and deciding to deploy is not a statistical prediction. It is an engineering adjudication of whether the evidence bounds the physical risk of operation in a named domain.

↰ Prerequisite: The empirical verification evidence backing release claims is produced in Fault Injection on Hardware.

The release this chapter works through is the warehouse mobile manipulator’s service at one site: driving the aisles past rack ends where a person may be standing, opening the spring-latched cage door, and handing the ceramic mug it takes from the takeaway conveyor to the coworker at the packing station. Each job carries its own criterion, a clear distance for the base and a contact force for the arm, and model accuracy establishes none of them. Release therefore depends on pre-contact clearance, measured brake capability, a bounded full response path, and the signed fault records of The Fault Manifest.

A release verdict is a bounded, revocable judgment. The chapter builds that judgment from the operating envelope, a claim-argument-evidence tree, three signed verdicts, the standing conditions whose failure revokes a verdict, a worked adjudication and the adversarial review it must survive, and a release manifest that refuses motion when the running machine no longer matches its signed evidence.

The Release Question

For the mobile manipulator, the release question is whether the evidence meets the physical criterion of each of its three jobs inside a scoped envelope. The argument names the physical bounds, specified fault set, and validated permission and fallback path. It can support a revocable grant of tracking authority while those conditions hold; it cannot infer protection against every possible fault from finite tests (1.1).

Definition 1.1: Claim-argument-evidence safety case

Claim-argument-evidence safety case is the auditable, structured assurance argument for a specified physical machine, operating envelope, hazard criteria, and fault model, linking top-level safety claims to concrete empirical evidence through explicit engineering warrants.

  1. Significance: Transforms subjective, informal deployment confidence into an auditable, falsifiable engineering argument that connects high-level safety goals to concrete empirical evidence across the entire physical stack.
  2. Distinction: Unlike software unit test suites or static model validation scorecards, a physical safety case explicitly links operational design domain boundaries, hardware fault injection logs, and runtime enforcer tripwires to demonstrate bounded residual risk.
  3. Common pitfall: Treating the safety case as a retrospective compliance checklist completed after development, rather than a living architectural contract that governs feature integration and release gating.

A release case is scoped by the operating envelope, the target ODD paired with the machine’s own internal conditions. An unconditional (UNCOND) verdict authorizes the whole operating envelope (The Operating Envelope), a conditional (COND) verdict a named part of it, and a refusal (REFUSE) none. For the mobile manipulator, the operating envelope admits aisle travel at up to 1.3 m/s (chosen; see the Reader Guide) with the tote rack loaded to at most 50 kg, on floors surveyed to a friction coefficient of at least 0.204, past rack ends with at least 1.10 m of clear distance. It admits door approaches at 0.03 m/s and item handovers to the coworker at 0.10 m/s.

↰ Prerequisite: The operational design domain and its known absences are defined in Collection Policy Coverage.

Inside these boundaries the learned policy may propose actions across the nominal task envelope, but each action still requires independent permission. Conditional states of the monitored dimensions (table 1) narrow that envelope; for example, the machine slows to 1.2 m/s before it requests an in-loop takeover. The restricted mode of a COND verdict narrows it further, so on a floor known only through the site’s inspection regime the machine runs at 1 m/s, below the 1.09 m/s inspected-floor ceiling.

As formalized in table 1, each monitored dimension is partitioned into nominal, conditional, and refusal states. A conditional state restricts admitted proposals. A refusal state withdraws tracking permission and selects a validated fallback (The Fallback Ladder). The table lists the dimensions within which a runtime monitor or a site procedure keeps the machine; the internal conditions of The Operating Envelope, such as ambient temperature, rail voltage, and compute load, enter through the fault records trialed at them.

Table 1: Operating envelope of the warehouse mobile manipulator: Each listed dimension has disjoint nominal, conditional, and refusal intervals. Conditional operation requires every dimension to remain in a signed admissible interval; missing or invalid measurements refuse. Sensor rates and stopping response are plant validation obligations.
Envelope Dimension / Invariant Nominal Envelope (\(\mathcal{S}_{\mathrm{nom}}\)) Conditional Boundary (\(\mathcal{S}_{\mathrm{cond}}\)) Hard Refusal (\(\mathcal{S}_{\mathrm{refuse}}\)) Runtime Monitor & Rate Enforcer Action & Latency
Aisle Speed \(v\le\) 1.3 m/s 1.2 m/s before an in-loop takeover request; 0.3 m/s before an out-of-loop one Measured speed above the signed 1.3 m/s (ceiling 1.39 m/s), or invalid speed estimate Encoders and IMU, every 1 ms permission tick Clamp admitted speed; refuse above the signed speed
Floor Friction Surveyed \(\mu\ge\) 0.204 (dry concrete near 0.60) Inspection regime only (\(\mu\ge\) 0.12): restricted 1 m/s, below the 1.09 m/s ceiling Unsurveyed floor or reported film; a clear oil film (\(\mu\approx\) 0.05) is invisible to the cameras Floor survey and site spill procedure; no on-board measurement Restrict speed by floor zone; refuse the aisle
Tote-Rack Payload \(0\le m\le\) 50 kg, valid load estimate (total at most 368 kg) Estimate stale since the last pick: stationary arm tasks only Load above capacity or invalid estimate Rack load cells, at each pick Refuse aisle transit until re-weighed
Rack-End Clearance Surveyed \(D_{\text{clear}}\ge\) 1.10 m at every rack end Shorter surveyed clearance: lower signed speed from the same budget Unsurveyed rack end, or obstacle inside the clear distance Site survey; lidar and navigation camera at frame rate Speed limit per rack end; stop for a detected obstacle
Door Latch Approach \(v\le\) 0.03 m/s, strike plate at its nominal position None; demonstrations cover no other latch condition 0.10 m/s approach or displaced plate Arm contact-force sensing in the 2 ms contact loop Trip at 15 N; refuse the unsupported approach
Coworker Handover TCP speed \(\le\) 0.10 m/s under accept_item Coworker present without accept_item: hold short of the handover point TCP speed above the signed 0.10 m/s (ceiling 0.118 m/s), or contact outside the assessed case Joint encoders every permission tick; arm contact-force sensing Clamp TCP speed; refuse above the signed speed

A defensible release claim states its plant, contact geometry, and fault set.1 The machine’s case uses three criteria. The base stops inside the 1.10 m clear distance before it reaches a person who has stepped out at a rack end, the arm opens the door without exceeding the 100 N latch limit, and a handover never presses on the coworker with more than a chosen 50 N peak force for an assessed contact case. The base’s claim protects a stationary person; a person still walking toward the machine is a separate premise, carried by a site crossing rule (Belief Through Occlusion). The release argument must demonstrate which human contacts are excluded before impact, which detected contacts the model covers, and which single faults the independent fallback handles. A stop after contact does not by itself cap the first force peak.2

Consider the arm at its free-space speed limit of 1 m/s, with a robot-side effective mass at the tool center point of 12 kg. In an unbraked elastic contact with the coworker’s hand, modeled with an effective tissue stiffness \(k_h\) of 15,000 N/m, the peak contact force follows equation 1: \[F_{\text{unbraked}} = v\,\sqrt{k_h\,m_{\text{eff}}} \tag{1}\] At the free-space limit this ideal elastic model predicts 424 N, or 8.5× the 50 N criterion, so the arm cannot enter an unseparated contact at that speed without a stronger, measured protection argument. The same equation bounds the item-handover speed that the enforcement record caps (The Enforcement Record, Human Authority Over Actuation), since the criterion admits at most 0.118 m/s. At the chosen 0.10 m/s, the model predicts 42.4 N, inside the criterion.

The base’s claim is a budget rather than a single force, and the delay components this chapter must evidence are the lease path and the observation age that its claims name. At the 1.3 m/s aisle speed, the finished budget (The warehouse mobile manipulator's stopping budget) is 997.4 mm against the 1.10 m clear distance, a spare of 102.6 mm. That spare is the whole allowance for error in the budget’s inputs. An unbudgeted delay of 78.9 ms, the handover deadline \(T_{\text{latest}}\) of Authority Transitions, spends it, and so does a loaded deceleration 14 percent below the 2 m/s² the budget assumes. A stopping claim whose inputs are unmeasured cannot show that its spare survives their errors.

The release question for the mobile manipulator has become three bounded criteria: the base stops inside the clear distance, the arm opens the door under the latch limit, and an item handover stays under the contact criterion. No calculation settles any of them alone, because each inherits premises about deceleration, delay, and contact that must themselves be evidenced, and each premise can lapse in service as brakes wear, floors soil, and sensors drift (section 1.5). Connecting every claim to its premises, and every premise to its evidence, is the work of a safety case.

Safety Cases and Claims

No inspection of the mobile manipulator can confirm that its loaded base will come to rest before it reaches a person at a rack end. That claim has to be decomposed until every branch ends in something measured, and a safety case is the structured argument that records the decomposition. Building on Stephen Toulmin’s argument structure (Toulmin 1958), Goal Structuring Notation (GSN)3 (Kelly 1998; Kelly and Weaver 2004), and the UL 4600 autonomy evaluation standard (Koopman 2020), this chapter constructs the claim-argument-evidence safety case of 1.1. Its root is a claim that asserts a measurable physical property, such as the base’s stop within the clear distance. Because that claim cannot be validated directly, strategies partition it by operating regime, hazard class, or subsystem, and each sub-claim ends in further sub-claims or in evidence, bounded by the context and assumptions under which its links hold.

Toulmin, Stephen E. 1958. The Uses of Argument. Cambridge University Press.
Koopman, Philip. 2020. How Safe Is Safe Enough? Measuring and Predicting Autonomous Vehicle Safety. BookBaby.

The structure’s validity hinges on distinguishing evidence from demonstration. A demonstration is a record of successful task execution over an uncharacterized test distribution, such as a long run of clean door openings with the strike plate always at its nominal position. It measures mean performance but bounds neither the tail nor the failure modes outside the sampled trajectory. Evidence is a falsifiable, calibrated artifact with quantified bounds, such as a worst-case timing distribution from hardware execution traces, loaded stopping trials across deceleration regimes, or a fault-injection trial that drives the machine to the boundary of its envelope. An artifact without quantified uncertainty, or one that never tests the boundary, remains a demonstration and cannot support an inferential link.

Connecting evidence to a claim requires a warrant, the inferential rule stating why the data support the assertion. For a learned policy the warrant is inductive, reasoning from failure-free test episodes to future service, and the exposure wall of Physical Trial Limits shows why that link cannot carry a high-consequence claim alone. No zero-event campaign, on one machine or a fleet, reaches the hazard rates such claims name within any practical schedule. The argument must therefore add an architectural warrant, grounding the claim in an independent permission path (Safety Enforcement) that bounds actuator torque and withdraws motion whenever a learned command leaves the admitted set. That warrant narrows the claim rather than escaping the wall, since the permission path changes the claim only for the hazards it detects and stops in time (What the Method Cannot Establish).

This tension between inductive fleet evidence and safety-critical assurance is illustrated by a decade of autonomous vehicle disengagement data reported to the California Department of Motor Vehicles (figure 1). Between 2015 and 2025, developers logged tens of millions of public testing miles across California roads. Over this decade, Miles Between Disengagements (MPI) expanded across three distinct eras:

  1. Era I: Early Highway & Suburban Testing (2015–2018): Systems operated primarily on well-marked arterial corridors and highways. Developers established baseline autonomy cadences, scaling from hundreds of miles between human interventions to several thousand miles.
  2. Era II: Dense Urban Core Transition (2018–2021): Fleets migrated into dense metropolitan environments, most notably San Francisco’s complex street grid. This environmental domain shift exposed uncharacterized edge cases—pedestrians stepping between double-parked vehicles, construction detour gestures, and dynamic occlusions—temporarily stalling or depressing reported MPI (e.g., Waymo’s metric dropped from \(29{,}945\) miles in 2020 to \(7{,}965\) miles in 2021 before recovering).
  3. Era III: Commercial Driverless Scale & Divergence (2021–2025): Removal of the in-vehicle safety driver led to sharp strategic divergence. Waymo scaled rider-only commercial operations, exceeding \(63{,}000\) miles per disengagement in 2025 across more than five million testing miles. Aurora pivoted its California passenger-car effort toward commercial Class 8 freight corridors in Texas. Apple cancelled its Project Titan initiative in early 2024 after a multi-billion-dollar R&D effort that struggled with early intervention sensitivity. Most instructively, Cruise crossed \(80{,}000\) miles per disengagement in California filings shortly before an October 2023 pedestrian-dragging collision in San Francisco resulted in the immediate regulatory revocation of its autonomous testing and deployment permits.

The fundamental systems lesson of figure 1 is that an inductive warrant derived solely from aggregate fleet mileage cannot establish a catastrophic hazard claim. A high MPI indicates operational maturity under nominal conditions, but it does not bound the severity or probability of tail-event failures under unmodeled physical interactions. Safety cases for physical AI must combine empirical exposure with deterministic architectural enforcers that remain valid when learned policies encounter novel distributional shifts.

An engineering safety case does not exist in an institutional vacuum. Inside the development team, the claim-argument-evidence tree establishes technical confidence that physical risks are bounded; outside the team, it forms the evidential substrate for statutory regulatory authorization and legal accountability. When an autonomous system moves from the laboratory into commercial deployment, its safety claims intersect three primary governance pillars:

  1. The European Union AI Act (High-Risk Classification): Under the EU Artificial Intelligence Act, AI systems deployed as safety components in physical machinery, collaborative robots, or autonomous transport are classified as High-Risk AI Systems. Statutory compliance mandates an auditable risk management system that spans the operational lifecycle, verified data governance documenting training and evaluation distributions (The Dataset Schema), human oversight interfaces capable of real-time intervention (Human Authority Over Actuation), and tamper-evident event logging (Trustworthy Operational Logs) to ensure forensic reconstructibility after any anomalous contact.
  2. Domain-Specific Statutory Certification: Beyond horizontal AI regulations, embodied systems must satisfy sector-specific safety mandates: the U.S. Food and Drug Administration (FDA) pre-market pathways (510(k) and De Novo) for autonomous surgical robotics, requiring human-factors validation under ANSI/AAMI HE75 and software lifecycle controls under IEC 62304; Federal Aviation Administration (FAA) airworthiness certifications (Part 107/135) for autonomous aerial logistics; and UNECE/NHTSA standards (such as UN R157 for automated vehicle steering) governing automated driving systems.
  3. The Redistribution of Legal Liability: In classical factory automation where robots operate within interlocked cages (Industrial spot-welding workcell), physical accidents are adjudicated primarily under the legal doctrine of operator negligence—evaluating whether floor personnel violated safety protocols or bypassed interlocks. When an autonomous machine operates without physical fencing in shared human spaces, legal responsibility shifts fundamentally toward strict manufacturer product liability. An unmodeled perceptual failure or tracking overshoot is no longer an operator error; it is an alleged defect in design, manufacturing, or failure to instruct. In this legal regime, a cryptographically signed safety case and release manifest is not merely good engineering practice—it is the definitive evidentiary record establishing that the engineering team bounded residual risk to state-of-the-art standards prior to public release.
Figure 1: A decade of autonomous vehicle disengagement trends reported to the California DMV (2015–2025): Annual Miles Between Disengagements (MPI, log scale) across major autonomous vehicle developers (Waymo, Cruise, Zoox, Apple Project Titan, and Aurora) spanning three historical operating eras. The horizontal dotted line marks the human driver baseline of approximately 40,000 miles between police-reported crashes (NHTSA). Milestones highlight Waymo’s 2021 dense urban transition dip, Cruise’s October 2023 permit revocation following a high-severity collision, Zoox’s consistent urban robotaxi regime, Apple’s program cancellation, and Waymo reaching 63,415 miles per disengagement in 2025. (Data source: California DMV Annual Autonomous Vehicle Disengagement Reports, 2015–2025).

Every link also rests on contextual assumptions, and an unstated one is guarded only by whatever check happens to sit downstream of it. The stopping warrant assumes, among other conditions, that the aisle stays bright enough for the navigation camera’s nominal exposure. Suppose a dim aisle doubles that exposure from 16 ms to 32 ms. Because an image’s age is timed from mid-exposure, each observation arrives 8 ms older than the budget charges, which unchecked would carry the base another 10.4 mm at the 1.3 m/s aisle speed, about 10 percent of the budget’s 102.6 mm spare. The permission path’s age check refuses those frames instead, and the base stops in an aisle the case assumed it could run. The team meets its lighting assumption as unexplained lost service, and only a check that the case never tied to lighting kept it from becoming a silent loss of clearance. An assumption that no runtime check sees has no such backstop, as the localization allowance in the compound failure of section 1.7 shows. A sound argument therefore states every physical dependency as a bounded context element with a runtime monitor that detects its violation.

↳ Downstream: Epistemic assumptions embedded within safety case warrants are critically interrogated in Epistemic Limits.

The weakest link governs every claim-argument-evidence tree. A safety case is not an aggregate in which over-performance in one branch compensates for a deficit in another. High accuracy on a recognition benchmark does not offset an uncharacterized latency tail on the CAN bus, nor does extensive policy simulation offset missing brake-wear data. If the base’s stop decomposes into sub-claims for braking, thermal protection of the drives (Thermal Duty Cycles), and trajectory bounding, a failed warrant in any one branch invalidates the root claim. The reviewer who adjudicates the case therefore looks not for the number of supporting artifacts but for the single weakest link whose failure would permit harm.

Match the safety argument to the machine

For the mobile manipulator’s arm, ISO 10218-1 and ISO 10218-2 frame the robot and integration assessment, and ISO/TS 15066 informs the collaborative-contact case once the argument names the contacted body region, pressure, geometry, and measurement method (ISO/TS 15066 2016). The base’s claims, which make up most of the case, are argued from its measured stopping budget and the site survey. The road-vehicle standards ISO 26262 and ISO 21448 lend the argument their distinction between faults and limits of intended function but certify nothing about this machine (ISO 26262 2018; ISO 21448 2022).

ISO 26262: Road Vehicles, Functional Safety. 2018. International Organization for Standardization.
ISO 21448: Road Vehicles, Safety of the Intended Functionality. 2022. International Organization for Standardization.

The claim-argument-evidence scorecard in table 2 states six claims and pairs each physical warrant with its adjudication on this book’s evidence. Claim C3 shows the decision: an oil film can defeat braking traction, so nominal model accuracy cannot justify an unrestricted verdict. A physically realizable condition that invalidates a warrant in this way is its defeater, and section 1.7 turns the search for defeaters into a review method. Claim C6 stands only on the fault records that test the four authority rules The Authority Log hands over (A1–A4).

Table 2: Claim-argument-evidence review for the warehouse mobile manipulator: Six claims and their adjudication on this book’s evidence. A failed or unknown branch cannot be compensated by nominal model accuracy.
Claim Warrant and operating assumptions on the machine Adjudication on the book’s evidence
C1. Contact and stopping Base: stop of 997.4 mm at 1.3 m/s (The warehouse mobile manipulator's stopping budget) inside the 1.10 m clear distance to a person standing at the rack end, with the payload untested at 368 kg an open premise; Door: latch force of 39 N at brake command, with the arm stopping within a further 0.15 mm (about 3 m/s²) under the 100 N limit; Mug: picked from the conveyor under the 1 m/s TCP limit on tracked intent and handed to the coworker at 0.10 m/s, under the derived 0.118 m/s ceiling, with a contact force of 42.4 N under the 50 N criterion, and the grasped object taken to be the mug the task names REFUSE: credible deceleration, brake onset, tracking and localization error, arm deceleration, and tissue stiffness are illustrative; the grounding premise is unknown
C2. Permission timing Brake onset no later than 82 ms after the last valid renewal; enforcer WCET of 250 μs inside the 400 μs deadline; declared assumption: the claim covers faults the permission path survives, lockstep diagnostics cover the path’s own failure (Two Paths on One Die), and the inhibit rung of the fallback ladder carries no budget (The Fallback Ladder) REFUSE: the WCET is declared, not measured under load
C3. Braking traction Floor friction of at least 0.204 on every aisle; a clear oil film is the defeater REFUSE; the open premise passes to A Residual-Claims Register
C4. Observation age Navigation-camera age of 51.6 ms inside the budget, with a stale-data fallback REFUSE: the age chain is illustrative
C5. Permission independence The placement record’s loaded maxima stay within the declared 250 μs WCET, and so below the deadline, and common causes are tested (Hardware Allocation) REFUSE until the placement test ledger exists
C6. Authority transfer One writer per actuator channel; no source bypasses a higher authority tier or the permission check; an unacknowledged handover falls back by the clearance-derived \(T_{\text{timeout}}\) of Authority Transitions (174.7 ms at the 1.2 m/s slow-down before an in-loop takeover; 78.9 ms for a request at 1.3 m/s); invalid authentication locks out the channel (The Authority Log) REFUSE if any of A1–A4 lacks a PASS fault record (The Fault Manifest); restrict to local operation if only the remote-channel cases fail
Checkpoint 1.1: Claim-argument-evidence decomposition and inductive warrants

Before presenting a safety case for adjudication, test its structure:

The Three Verdicts

Every row of table 2 ends in an adjudication, and the release decision resolves the whole tree into one of three verdicts, unconditional operation, conditional operation, and refusal to operate. The first, unconditional operation, grants the full envelope named in a signed release case while its standing conditions remain true. Each hazard claim in that case needs evidence for the specified plant, load, operational design domain, and fault set, with brake thermal capacity, stopping clearance, and tracking-error bounds qualified for that envelope (Thermal Duty Cycles) and independent permission checks wherever learned proposals can reach actuators (The Causal Boundary). The verdict does not eliminate residual or newly discovered hazards; it authorizes operation under the assessed claims and requires revocation when their premises fail.

The second verdict, conditional operation, applies when the evidence shows the machine safe within a restricted operating subset but lacks the warrants for the full envelope. Release then enforces physical, kinematic, or environmental boundaries that exclude the unverified hazard states, such as de-rated joint velocity, a restricted illumination range, or a barrier that keeps people outside the manipulator’s reach. Where contact force cannot be bounded for unseparated co-presence, a restricted case may require an assessed fence and validated access interlocks (figure 2), tested so the protective stop completes before a person reaches the hazardous workspace.

Yellow industrial robot arm and metal equipment inside a factory bay bounded by yellow mesh fence panels and gates. A blue lift is outside the foreground fence. Wiring, interlock rating, and stopping behavior are not visible.
Figure 2: Physical separation example: A yellow industrial arm works behind yellow mesh fencing. Fencing can be one condition of a restricted release when a validated access interlock and stopping test support it. Source credit: U.S. Army Chemical Materials Activity / PEO ACWA.

Every condition attached to a release verdict is a binding contract, and each names a number the permission path enforces rather than the ceiling the physics allows. The item handover shows the difference. The contract signs a TCP speed below the ceiling that the contact criterion admits under equation 1, and the gap between the two speeds is the allowance for what the monitor cannot see at the instant it samples: speed growth between samples, encoder error, and the drive’s response to a clamp. A contract signed at the ceiling leaves no room for those terms, and the release case needs a measured bound on each before the signed speed can stand.

↰ Prerequisite: Deterministic fallback coverage supporting conditional release verdicts builds on The Fallback Ladder.

The inspected-floor condition follows the same rule. Where the site’s inspection regime certifies a friction coefficient of only 0.12, the credible deceleration falls to what that friction supports, and the budget that admits 1.39 m/s on surveyed floors admits only 1.09 m/s. The contract binds only while its premises hold.

The third verdict, refusal to operate, is the mandatory outcome whenever the safety case contains missing warrants, uncharacterized tail latencies, incomplete fault-injection coverage, or unmonitored failure modes. If the enforcer’s worst case under the placement’s load has never been measured against its 400 μs deadline, or the fault set omitted the common-cause tests the placement names, the hazard remains unquantified and cannot support release. Refusal is likewise required when a measurement contradicts the model behind a budget term, such as loaded stopping trials that decelerate below the 2 m/s² the budget assumes.

A refusal verdict is an engineering deliverable, not an administrative failure. It moves a dispute between schedule and safety onto evidence by isolating the broken link in the argument tree. If release is refused because loaded stopping trials contradict the braking term, the team does not debate risk tolerance but repeats the trials at the loaded mass until the low tail of the deceleration is known, re-derives the stopping term from that tail, and returns for re-adjudication only when the aisle budget closes on the measured value. The refusal verdict defines the exact evidentiary test that must pass before re-adjudication can occur.

Accountability for the release decision must reside with an identifiable lead systems adjudicator rather than being diffused across development teams. When a vision team signs off on detection scores, an actuator supplier on torque curves, and a learning team on policy convergence, each assumes that another layer provides the margin. The adjudicator alone attests that the integrated system satisfies the safety case, and the attestation binds the verdict to a set of immutable technical artifacts, including the cryptographic hash of the verified policy weights, the calibration logs of the arm’s contact-force sensing, the brake-onset trials together with a separately justified worst-response bound, and the stopping trials at the loaded mass that establish the credible deceleration. Signing records the adjudicator’s bounded judgment about that evidence and the specified operating domain.

The adjudicator signs one of the three decisions in table 3. A conditional verdict needs a separately evidenced restricted envelope; a known hole in the nominal argument cannot simply be relabeled as a safe restriction. No verdict moves the proposal boundary. Under UNCOND and COND alike the policy only proposes and the enforcer still admits each action, as principle \(\ref{pri-vol4-proposal-permission}\) requires, so a release sets the envelope the permission path enforces and never grants the policy permission of its own.

Table 3: Three signed release verdicts for the mobile manipulator: Full operation and restricted mode both retain independent per-action permission. Refusal is an actionable engineering result.
Signed verdict Evidence required Permission and runtime limits Expiry or refusal trigger
UNCOND Full operating-envelope claim, measured plant and timing bounds, independently validated permission and fallback Policy proposes across full signed task envelope; every action passes the enforcer Evidence, configuration, calibration, or standing premise fails
COND Restricted operating-envelope claim and its own evidence, limits, and fault coverage Policy proposes only within signed reduced limits; every action passes the enforcer Any restricted premise becomes false or unknown
REFUSE Missing, contradictory, expired, or unauthenticated critical evidence No autonomous tracking permit; execute the assessed stop or hold, the fallback-ladder rung the release case validates for the plant (The Fallback Ladder) Re-adjudication with repaired evidence

None of the three verdicts is a permanent grant. Linkages develop backlash, friction surfaces wear, lenses accumulate film, and ambient distributions change, so operating under UNCOND or COND carries obligations for the life of the deployment. Brake wear that lowers the loaded deceleration by the shortfall that spends the aisle budget’s spare (section 1.2) leaves the budget no allowance for any other input error. That wear must fail a standing condition (a scheduled stopping trial or a validated brake self-test) before the machine re-enters the aisle, since a stop in service would first reveal it at a rack end. Each such obligation pairs a premise with a detector and a response, which makes it one of the standing conditions the verdict carries.

The Verdict’s Standing Conditions

A release verdict depends on physical premises with different observation schedules. The mobile manipulator’s braking claim relies on floor friction of at least 0.204, which dry concrete at about 0.60 exceeds, but a one-time survey cannot cover spills or surface wear. The case needs a floor inspection schedule and a stop or exclusion rule after contamination, with any on-board traction monitor specified separately. The arm’s contact-force calibration is re-established at a documented interval. Each premise must remain valid at the cadence required by its physical failure mode. Under principle \(\ref{pri-vol4-evidence-bounds-authority}\), the verdict grants authority per operating condition, and that authority lasts only while the evidence behind each premise stays current.

Every monitored premise requires both a deterministic runtime detector and a pre-allocated response. If the case claims that the navigation camera’s observation age stays inside the 51.6 ms the budget charges, but the software has no hardware timestamp comparator to measure that age on each frame, the premise has not been monitored. It has silently degraded into an assumption. The tote-rack load, bounded by its 50 kg capacity, changes only when the arm places or removes a tote, so a check at each pick suffices, while speed and lease validity are checked on every 1 ms permission tick.

Evaluating monitored premises at runtime exposes a trichotomy in envelope membership (1.2). A binary check presumes normal execution whenever no fault flag is set, but envelope membership has three states, known-true, known-false, and unknown. Membership is known-true when telemetry is verified, timestamps meet their deadlines, and estimators converge with bounded covariance. When a physical limit is breached, such as base tracking error exceeding the 40 mm bound that the stopping envelope insets, the state is known-false. When an optical encoder drops communication, a camera frame arrives with a cyclic redundancy check (CRC) mismatch, or an extended Kalman filter diverges, the machine cannot determine whether it is inside or outside its signed operating envelope. With mass in motion, an unknown state is treated as outside the envelope, since presuming safety during an epistemic gap, while the machine moves blind, lets kinetic energy accumulate unmeasured.

Definition 1.2: Envelope trichotomy

Three evidence states: known true has validated fresh measurements, unknown has insufficient or stale measurements, and known false has a measured boundary breach. Unknown and known false remain distinct logged causes but both refuse tracking permission and select a validated stop or hold.

Envelope trichotomy is the three-state epistemic classification of cyber-physical system health relative to its certified operating envelope: known-true (verified state with bounded telemetry covariance inside envelope bounds), known-false (verified boundary violation triggering deterministic hardware refusal), and unknown (epistemic state uncertainty due to sensor dropout, frame corruption, or estimator divergence).

  1. Significance: An unknown state is not evidence of a boundary breach, but neither supports continued tracking. The permission decision is refusal in both unknown and known-false states; the logged cause remains distinct and the response depends on the plant’s validated stop or hold capability.
  2. Distinction: Unlike a binary boolean status check (is_safe == true), the envelope trichotomy isolates epistemic blindness from active violations, ensuring telemetry loss is never mistaken for safe clearance.
  3. Common pitfall: Defaulting to “fail-open” or holding last-known commands when communication fails. Frozen commands in the presence of unmodeled dynamics rapidly result in mechanical collisions.

The same three states govern the evidence behind each premise. A trial marked UNOBSERVABLE or INVALID in Fault Oracles and Acceptance supplies no evidence, so the premise it was meant to test stays unknown, and an unknown premise refuses the claim that depends on it; only PASS records count toward a verdict, and a FAIL record makes the premise known-false.

The response to an unevaluable premise follows from the physical consequence of continuing motion without telemetry, not from a severity label. When the mobile manipulator runs the aisle at 1.3 m/s with its tote rack full, its 368 kg loaded mass carries 311 J of kinetic energy. If its lidar stops transmitting updates, each millisecond of blind travel carries the machine 1.3 mm closer to an unmapped obstacle. The safety case fixes three responses in advance. Two of them exist only where the signed case covers them: a degraded mode that continues when redundant sensors still bound positional uncertainty, such as falling back to wheel odometry and halving the aisle speed to 0.65 m/s, and the restricted mode of a COND verdict that clamps torque or confines proposals to a precomputed corridor. The third, the STOP rung of the fallback ladder (The Fallback Ladder), needs no such coverage and uses the configured drive and brake sequence when observability is lost; its ability to stop within clearance depends on the current speed, response delay, and measured braking capability.

Even when every monitored premise remains known-true and periodic recalibrations pass, the release verdict carries an expiration timestamp and cycle budget. Operating hours and mechanical cycles can change the plant in ways that runtime telemetry does not fully observe. Gear teeth may develop surface pitting,4 elastomer seals stiffen, and structural fasteners lose clamp preload. The load-specific fatigue data and inspection plan set the signed operating-hour and cycle limits; reaching either limit requires re-adjudication before further motion is permitted.

Standing conditions that affect motion need an enforcement path with a measured detection and response budget. In the architecture of Safety Enforcement, the independent enforcer checks proposals and monitored premises before issuing actuator permission.5 The machine’s permission path decides on a 1 ms tick. Its 82 ms bound from the last valid renewal to brake onset covers a proposer that stops renewing. A premise the enforcer monitors itself needs only the tick, the bus transfer, and brake onset, 22 ms in the budget, and the release case must measure each of those terms. When a premise is unknown or false, the enforcer revokes tracking and commands the rung of the fallback ladder that suits the plant’s state (The Fallback Ladder), and the signed release case must include that stop choice and its measured timing.

A verdict therefore stands only while each of its standing conditions is known-true, checked at the cadence its failure mode sets and backed by a response whose timing the case has measured. With claims, verdicts, and the conditions that revoke them defined, the adjudicator can take the mobile manipulator’s six claims one at a time and decide its release on the evidence this book holds.

A Case Worked in Full

The adjudicator takes each of the six claims in table 2, finds the premise that carries its warrant, and asks what kind of number supports that premise. The door branch of C1 shows the procedure in full.

At the guarded approach speed of 0.03 m/s against the 4.0 × 10⁵ N/m strike plate, contact force rises as the plate compresses until the 15 N tripwire fires, and the 2 ms contact-loop response adds further compression before the brake command begins (What Intent Hands Over): \[F_{\text{cmd}} = F_{\text{trip}} + k_{\text{latch}}\,v_{\text{latch}}\,T_c \tag{2}\] Evaluating equation 2 gives 39 N at brake command, well inside the 100 N latch limit. The limit still binds the stop that follows, because the plate keeps compressing while the arm decelerates: \[x_{\max} = \frac{F_{\text{cmd}}}{k_{\text{latch}}} + \frac{v_{\text{latch}}^2}{2a_{\text{arm}}} \le \frac{F_{\text{lim}}}{k_{\text{latch}}} \tag{3}\] The braking term in equation 3 ignores the spring work that would shorten the arm’s travel, which keeps the bound conservative. Below the limit the plate can compress only 0.15 mm further, so the arm must decelerate at about 3 m/s² and stop within about 10 ms.6 At the unsupported 0.10 m/s approach, the same two equations give 95 N at brake command and demand 400 m/s². That demand is consistent with the two reasons the operating envelope (table 1) admits only the guarded approach. At the faster speed the contact deadline barely exceeds the contact loop’s response (Policy Synthesis), and the demonstrations hold no example of it (How Datasets Go Wrong).

Neither result is in doubt as arithmetic. The door branch fails on its premise. Whether the arm, carrying its approach momentum, reaches that deceleration within a fraction of a millimeter is a property of its drives and brakes, and no record in this book measures it. The deceleration is illustrative, so the premise is unknown and the branch is refused.

The other branches fail the same way. The base’s 997.4 mm budget inherits illustrative values for credible deceleration, brake onset, tracking error, and localization, and its 102.6 mm spare is small against a plausible error in any one of them. The mug branch of C1 is refused on two premises. Its handover force sits under the criterion only through an illustrative tissue stiffness, and its grounding premise, that the object the arm grasps is the mug the task names, is unknown, because a confident grounding to a same-colored glass jar passes every geometric check (From Request to Expiring Geometric Proposal) and no record in this book settles which object the arm holds. The permission timing of C2 rests on a WCET declared at 250 μs beside an unloaded 135 μs, never measured under the placement’s worst load. The 51.6 ms observation age of C4 adds datasheet readout to illustrative pipeline terms. C5 and C6 cite records that this book describes but does not contain: the placement test ledger and a PASS record for each of A1–A4. C3 differs in kind. Its premise, floor friction of at least 0.204 on every aisle, can be surveyed on a dry floor, but a clear oil film looks like dry concrete to the cameras, and no measurement on the machine can exclude it.

Every claim therefore rests on at least one premise that the book’s evidence leaves unknown, and the adjudicator signs REFUSE. The verdict does not find the machine unsafe. Its inputs are illustrative, so the decision is illustrative too. It shows how this release would be decided and cannot license the machine. Under the fourth law (The Four Bedrock Laws), a premise the evidence cannot settle must restrict operation, and on this evidence every claim carries one. The refusal is also the chapter’s most useful result, because it names the measurements that would change it. Table 4 lists them, with the premise each settles and the restriction a supporting result would lift.

Table 4: Measurements that would move the verdict toward COND: Each row names the premise a measurement settles on the warehouse mobile manipulator and the restriction a supporting result would lift; the inference trace bears on service availability only.
Measurement Premise it settles (value in this book) Claims What a supporting result lifts
Stopping trials at the loaded mass Low-tail credible deceleration (2 m/s²) C1 The base’s budget and its 1.39 m/s ceiling, so an aisle speed can be signed
Brake-onset trials Brake onset (20 ms) C1, C2 The 82 ms bound from the last valid renewal to brake onset
End-to-end age trace Observation age (51.6 ms) C4 The age term of the budget and the stale-data threshold
Tracking-error and localization tails Tracking bound (40 mm) and localization bound (50 mm) C1 The inset and overhead terms of the budget
Site survey Rack-end clear distance (1.10 m) C1 A signed speed at each surveyed rack end
Floor survey Dry-floor friction (0.60) against the 0.204 needed C3 Surveyed aisles only; the film stays unobservable, so the aisle claim can reach at most COND on inspected floors below the 1.09 m/s ceiling
Inference trace Chunk inference at the 99th percentile (40 ms) None Service availability: renewal of each chunk inside the 60 ms lease; the C2 bound holds without it
Loaded enforcer timing Enforcer WCET (250 μs, declared) C2, C5 The 400 μs deadline under the placement’s worst load
Arm stopping trials Arm deceleration from the guarded approach (about 3 m/s² needed) C1 Door opening at the guarded approach
Handover force trials Peak contact force at the handover speed (42.4 N derived; tissue stiffness illustrative) C1 Handing the mug to the coworker at 0.10 m/s under accept_item
Operator trials Takeover and authority-handover timing C6 Remote takeover, once A1–A4 each have a PASS record; otherwise local operation only

A COND case would need every row of the table except the availability-only inference trace and the operator trials, and it could forgo the operator trials only by keeping authority local, where the remote channel’s authority cases do not arise. Local operation would still need a PASS record for each authority case it exercises, and C5 would still need the placement test ledger with its common-cause tests. Neither is a row of the table.

Even with all of that evidence supporting its premises, the most the results could license is a named part of the operating envelope. The base could travel the aisles at a restricted 1 m/s, below the 1.09 m/s inspected-floor ceiling, on inspected floors only, with the enforcer in its restricted configuration (The Enforcement Record) and the crossing rule in force at every rack end. The arm could open the door at the guarded approach with the plate at its nominal position and the base stopped, and it could take the mug and hand it over under accept_item at 0.10 m/s, with the station’s feed held to its declared item list and a singulated pick window. The grounding premise has no row at all, since no trial with the machine’s present sensors can settle it, and the feed list and singulation restrict it rather than close it. The 1.3 m/s normal mode stays unsigned whatever the results, because the oil film leaves its friction premise open. The restrictions that carry these premises, and the evidence that would close them, belong to the residual-claims register of A Residual-Claims Register (1.2).

Checkpoint 1.2: Refusal and the path to conditional release

Before approving a conditional release of the machine, verify your understanding of how a premise, rather than a calculation, decides a claim:

Adversarial Review

Suppose the measurements of table 4 arrive and support a COND case for the mobile manipulator. Before the adjudicator signs it, the case must survive a reviewer whose task is to break it. An adversarial review is a structured audit that treats the safety case as an argument to be attacked rather than a brief for deployment. Confirmation bias steers the case’s own data collection toward conditions where the learned models succeed, and the reviewer inverts that posture, searching for unsupported warrants, uncharacterized physical transitions, and edge conditions that break the release contract. An argument that fails on paper against an adversary trying to force a physical violation gives no reason to expect the machine to withstand service.

The review evaluates every evidentiary node in the argument tree with four probes. The first challenges empirical sufficiency, whether the trials covered the physical extremes of the operating envelope. Stopping trials with the loaded base on dry concrete, at a friction near 0.60, support the warrant that the base reaches its credible deceleration there. The restricted mode, however, admits floors that the site’s inspection regime certifies only to 0.12, such as a lane polished by traffic, where the base may run at the restricted 1 m/s that The Operating Envelope declares, below the 1.09 m/s ceiling. If the trial set holds no loaded stops on such floors at that speed, the warrant for the restricted mode is unsupported, and a coverage audit finds in review the gap that service would otherwise find at a rack end. The oil film is no part of this probe, because the envelope refuses it outright and C3 treats it as the defeater. The second probe examines domain transfer, claims where simulation substitutes for hardware measurement. A person detector may score well on rendered aisles and still miss a person partly hidden behind stretch-wrapped pallets under the site’s own lighting. Its benchmark recall is a proxy that scores curated frames, not the chain from a missed person to a commanded stop, so the reviewer asks for evidence on that chain from paired hardware-in-the-loop trials at the real rack ends.

The third probe attacks structural independence, asking whether the learned model and safety enforcer share a resource7 whose contention could delay sensor delivery. On the mobile manipulator such a delay would enter the 51.6 ms observation-age chain of C4, where no single component’s timing test would show it. Stress tests and timing analysis under the placement’s worst load must establish the interference bound; idle-bus unit tests cannot. The fourth probe identifies unmonitored state drift. Wear can change gearbox compliance8 and the response between motor current and link motion. The reviewer therefore asks for paired wear-state braking and contact traces before accepting the arm-stop premise of the door branch of C1 (section 1.6). A motor-current trip alone does not establish that the worn plant remains inside the case’s measured force and stopping limits.

Reviewers operationalize these vulnerabilities by constructing explicit defeaters for each warrant. The hardest to see combine minor anomalies that each stay inside their own bands, the compound faults that Deriving the Fault List derives from shared causes.

Consider the mobile manipulator’s navigation camera with a light film of dust on its lens, a mount that has shifted within tolerance after its fasteners relaxed and now shakes as the base crosses floor joints, and a harness whose flexing at those joints raises the error rate on its serializer/deserializer (SerDes) link. Each check reads normal. The camera reports adequate visibility, the mount check registers within tolerance, and the SerDes error counters stay below their alarm. The anomalies cannot combine in age, because the permission path refuses any observation whose upper age exceeds the budgeted age plus the clock-conversion bound (The Observation Contract), so a frame that retransmissions delay past that threshold costs availability, not clearance. They combine in position. Dimmed contrast leaves the feature matcher on weaker matches. The mount blurs the frames it matches as it shakes and offsets every match by an angle the calibration still accepts. The frames the harness loses to retransmission lengthen the gaps between fixes, over which odometry drift accumulates. Each error stays inside its own tolerance, but all three land in the one localization allowance of the stopping budget, which no runtime check measures directly. That allowance is sized from the measured tail of localization error (table 4), not as the worst-case sum of its contributors’ tolerances, which would give up signed speed to cover every contributor at the edge of its band at once. Trials run with a clean lens, a tight mount, and a sound harness never sample that coincidence, so the three errors together can exceed the allowance. The excess comes out of the spare (section 1.2), and the base travels on toward the rack end with less margin than the signed case assumes while every diagnostic reports normal operation.

⇄ Contrast: Algorithmic capability claims contrast with physical actuator saturation limits established in Actuator Transmission Limits.

The review also audits the execution integrity of the permission path, whether inference or background services can bypass, corrupt, or starve the enforcer (Safety Enforcement) or human overrides (Supervisory Intervention), the latter covered by C6 of table 2. Below the proposal boundary (The Machine in Five Levels), both keep non-preemptible priority over learned models, and the adversary asks whether a runaway policy thread, an accelerator-runtime exception, or a telemetry flood can saturate memory channels, lock a kernel spinlock, or brown out a rail. The reviewer accepts permission path independence, and the held-up rail that carries it on the mobile manipulator, only on physical fault injection, including bus saturation, feed removal, and a killed proposer task (F1), because C2 claims only the faults the permission path survives. Those trials are entries of the fault ledger, the set of signed fault records of The Fault Manifest. If the permission path is implemented in user-space software rather than isolated hard real-time threads and dedicated hardware interrupt lines, the reviewer marks the link as broken.

When an attack succeeds, the case cannot be salvaged by editing claim descriptions or inserting arbitrary margins. The broken warrant is traced to its systems failure, an uncharacterized physical boundary, an inadequate dataset, or an unmodeled latency, and the defect returns to the owning subsystem. A motor-current observer that fails under gearbox compliance returns to the mechanical specification as wear limits and runtime compliance estimation, and bus contention that delays the enforcer returns to the compute architecture as memory partitioning and real-time scheduling. The case returns to the engineering team with its broken warrants named, and only updated boundaries, re-derived stopping distances, and renewed verification reopen adjudication.

Review must also recur, because a case accepted once can erode in service without any single change that would force a new one. The normalization of deviance names that erosion, in which anomalies that cause no harm come to be read as acceptable operating noise rather than as evidence against a warrant (Leveson 2011; Vaughan 1996) (1.1).9 On the mobile manipulator it would look like the compound defeater above spread over months: localization residuals that creep toward the edge of their tolerance, age refusals dismissed as nuisance stops, stops that end a little longer than the trials did, and enforcer trips cleared without a trace review. When thresholds are widened, margins trimmed for throughput, and alarms suppressed because nothing has yet gone wrong, the gap between the signed envelope and the machine’s behavior widens until a routine disturbance exhausts what remains.

Leveson, Nancy. 2011. Engineering a Safer World: Systems Thinking Applied to Safety. MIT Press.
Vaughan, Diane. 1996. The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. University of Chicago Press.

War Story 1.1: Challenger and experience as a boundary (1986)
Context: The Space Shuttle’s solid rocket motors were assembled from segments whose field joints were sealed by a primary and a secondary O-ring (Presidential Commission on the Space Shuttle Challenger Accident 1986). No learned model appears in this chain. The adjudication failure it records is the one this chapter’s verdicts exist to prevent, a launch approved for conditions that the evidence did not cover.

Mechanism: The ambient temperature at launch was \(36^\circ\text{F}\), 15 degrees colder than any previous launch. O-ring resiliency falls with temperature, and a cold ring returns to shape too slowly to follow the joint gap as ignition pressure opens it. On the eve of launch, engineers at Morton Thiokol, the motor’s contractor, recommended against launching below \(53^\circ\text{F}\), the O-ring temperature of the coldest previous flight, whose joints had shown the worst blow-by to date. That limit was the edge of flight experience, not a qualified bound. NASA managers countered with a qualification temperature that applied to the propellant’s bulk temperature rather than to the seals.

Impact: A combustion gas leak through the right booster’s aft field joint, beginning at or shortly after ignition, weakened or penetrated the External Tank. The vehicle broke up about 73 seconds after launch, and all seven crew members died.

Response: The Commission attributed the failure to “a faulty design unacceptably sensitive to a number of factors,” and temperature heads its list of those factors. It recommended that the joint be changed and that the new design be certified by tests “over the full range of operating conditions, including temperature.”

Systems lesson: Thiokol management reversed its engineers’ recommendation at the urging of NASA’s Marshall Space Flight Center. Thiokol’s vice president of engineering testified, “We had to prove to them that we weren’t ready.” The burden of proof had moved from showing the joint safe inside its evidence to showing it unsafe outside the evidence, where no data existed. A release verdict keeps the burden where the envelope trichotomy puts it. A condition outside the evidence, such as a seal temperature below all flight experience, is unknown, and the verdict refuses it.

Tempe’s takeover premise (When the Boundary Fails) was a condition of the same kind, and How Handover Goes Wrong posed the operator-monitoring question it raises. A release that relies on such a premise must bind its response, refusal or a restricted case, together with its tested physical limits.

The Release Manifest

A release verdict needs a signed artifact that records the decision, physical conditions, and expiration. That artifact is the release manifest, the immutable record that binds the verification ledger to the signed verdict of the lead systems adjudicator and refuses motion whenever the running images or evidence no longer match its digests.10 It consumes the fault ledger through its root digest, and it binds three further records as evidence leaves of the signed argument graph: the authority record of The Authority Log, the placement record of Hardware Allocation, and the evaluation record of Evaluation Logs. To these it adds the claim root, the argument digest, the image digests, calibration references, the verdict, its standing conditions, a drift budget, and an expiry. The image digests cover the policy weights, the enforcer’s logic in both of its enforcement-record configurations, the normal one and the restricted one a COND case would load (The Enforcement Record), and the real-time kernel. A lifecycle budget of operating hours and cycles sits beside the calendar expiry, and the adjudicator’s signature covers the whole record. Evidence links must resolve to the reviewed record versions before permission is granted.

↳ Byte layout: The release manifest’s fields, types, and digests are given in Release manifest.

Figure 3 places authentication and physical-standing checks before any operating permit, and the gate runs them in a fixed order before the motor power stages are energized. Each check can only refuse. It first verifies the adjudicator’s key chain, which the hardware root anchors, and its revocation state, then the manifest signature, the expiry and lifecycle budgets, and every active image digest, so a changed weight file or enforcer bitstream refuses permission. It next authenticates the complete fault ledger, whose individually signed records enter the manifest as one root digest, together with the calibration references. A digest alone establishes neither record validity nor test adequacy, so the gate also checks each record’s signature and the ledger’s completeness against the reviewed evidence index. The standing conditions come last, and a known-false or unknown premise refuses as firmly as a bad signature while keeping its own logged cause. Only when every check passes does the gate load the signed envelope, and even then the enforcer still checks each proposed action. Table 5 gives the outcomes. Stale calibration is the case that separates them, since it refuses the full case and supports COND only where a separately signed restricted case covers exactly that state.

↳ Protocol: The gate’s step-by-step authorization protocol is given in Release-gate protocol.

Left-to-right release path: signed case, authentication of key chain, images, and fault ledger, then physical standing checks. Failed authentication or a false or unknown premise leads to REFUSE, permit off, and assessed fallback. A valid case leads to UNCOND or COND with a signed envelope and independent check of each proposed action. A stale calibration needs a separately authenticated restricted case.
Figure 3: Fail-closed release pipeline: The signed verdict and evidence root pass key, image, and complete fault-ledger authentication, then current calibration, brake, and operating-envelope checks. Either check can refuse permission. Only a valid signed full or separately signed restricted case reaches per-action enforcement. Attribution: Textbook production team (CC BY-NC-ND 4.0).
Table 5: Release authorization truth table: A restricted mode is a distinct authenticated claim, never an unsigned downgrade of the full case.
Signed case Evidence and standing checks Restricted case covering stale calibration Gate result
UNCOND All valid and current Not needed Full proposal envelope under enforcer; UNCOND
COND All restricted evidence valid and current Included if calibration stale Signed restricted envelope under enforcer; COND
UNCOND or COND Fault-ledger root/signature mismatch, expiry, image mismatch, or unknown physical premise Any Permit off; REFUSE
UNCOND Calibration stale No Permit off; REFUSE
Separately signed COND Only the named calibration is stale; restricted evidence and limits valid Yes Restricted envelope; COND
REFUSE Any Any Permit off; REFUSE

Every evidence leaf needs an identifier for its raw trace, instrument, calibration, configuration, and test script. A measured timing tail enters as a percentile, and a hard response claim needs the separately justified bound of Moving Commands on Time. The coworker’s handover criterion needs load-cell time series for the specified body-region surrogate and geometry. Linked records let an auditor repeat the test and identify where the claim stops applying.

The release manifest names the configuration and physical limits for which its verdict applies. A changed policy weight file, real-time priority, driver, or unverified component revision requires a new case or an explicitly approved equivalent. For the mobile manipulator, the signed wear limits on joint backlash and brake lining thickness come from measured stopping and load evidence. Refusal conditions of the operating envelope include a tote-rack load above its capacity, which the rack load cells monitor, and an unsurveyed floor, which the site procedure enforces (table 1). The manifest also signs the permission path’s isolated, held-up rail as a standing condition, sized in The Fallback Ladder, with a capacity self-test of its hold-up store gating the permit at every power-up. When a measured condition crosses a signed limit, the permission path revokes tracking at its specified detection and response time and commands the assessed fallback.

Field telemetry records torque residuals, loop timing, and enforcer trips for drift review, while the permission path checks the signed live thresholds at the cadence the plant requires. A brake response measured at scheduled inspection beyond the signed clearance budget, or stale calibration, revokes the current verdict, and reinstalling an earlier signed image cannot restore worn brakes. Unless a separately signed COND case covers that physical state, the plant uses its assessed stop and awaits maintenance and re-adjudication (table 5).

↳ Downstream: The premises a release manifest leaves open become entries in the residual-claims register of A Residual-Claims Register.

Every refusal in this chapter ends in the assessed stop, which presumes that a stopped machine is safe, a presumption that fails for a machine that falls when it stops (1.1).

Systems Perspective 1.1: When removing torque causes the harm

Question: Which verdict applies when removing torque causes the harm?

A braked base at rest stores no energy that can reach the coworker, but an upright humanoid stores its hazard in its own posture. No verdict, not even REFUSE, stands for the humanoid without a claim that its fall-arrest response works.

  • Safe torque off on an upright biped is a fall. The drive function that removes motor torque (Actuation Authority) leaves the robot’s mass to drop from standing height, so no signed verdict can name torque removal as the humanoid’s assessed stop.
  • The fall-arrest response is itself an actuation. Crouching, stepping, or lowering the robot to the floor under power is a commanded motion that needs its own claim, its own evidence, and a place in the fault set. A REFUSE verdict still depends on the part of the machine that keeps it standing.
  • Wear revokes the running permit. Sole friction and joint backlash drift out of the limits the fall-arrest evidence assumed, and crossing a signed wear limit withdraws the permit to walk before the arrest response that the permit relies on stops working.
Checkpoint 1.3: Deriving a verdict from your own measurements

Table 4 carries no measured values, because this book has none. Fill it with values of your own, from trials on a machine you can test or from numbers you choose, and derive the verdict they support:

Fallacies and Pitfalls

Fallacy: Statistical coverage substitutes for deterministic architectural warrants.

A team presents millions of simulated hours and thousands of incident-free cycles as its release argument. The record supplies evidence about the tested conditions, but demonstrating the catastrophic failure rate a release claim names can require more physical exposure than any schedule allows (Physical Trial Limits). Simulation also retains the assumptions of its physical models, such as idealized Coulomb friction or infinitely rigid gearboxes. The release case must explain how architectural enforcement, such as an independent permission path that revokes motion and commands a validated stop, constrains harmful policy behavior, and it must provide evidence for that mechanism under its stated assumptions. Empirical performance and architectural warrants answer different questions; a large trial count cannot replace the missing argument that connects actuator authority to physical limits.

Fallacy: A release verdict is a permanent certification of safety.

A machine passes validation and enters production under a signed safety case. Over time, brake pad wear lengthens stopping distances, kinematic calibration drifts due to thermal cycling, and optical contamination degrades sensing. The software can remain unchanged while the physical assumptions supporting the release fail. The release manifest needs maintenance requirements and invalidation triggers tied to relevant lifecycle counters or measured drift, with defined restrictions when those conditions fail. A signature records responsibility for a bounded release decision; it does not make the tested body’s hardware capabilities permanent.

Pitfall: Relying on unstated environmental assumptions to support a safety claim.

A team verifies the base’s stopping performance in well-lit aisles on clean floors, and its case states the floor’s friction but never its lighting. A dim aisle then costs service through the image age that a runtime check happens to see (section 1.3) and costs clearance through weaker feature matches whose localization error no check sees, against an allowance the well-lit trials sized (section 1.7). The mechanical brake performs exactly as tested and every diagnostic reads normal while the spare in the stopping distance shrinks. The safety case must state the environmental conditions on which its claims depend and explain how operation remains within them. Where monitoring can detect a violation, it must trigger the defined restriction or fallback; an unstated assumption supplies neither a testable release condition nor a usable runtime boundary.

Summary

Production release binds a claim-argument-evidence case to a measured operating envelope, a signed configuration, and a permission path that can revoke motion. The signature establishes the authorized case; runtime checks determine whether its physical premises still hold.

Key Takeaways: Deployment release and defensible safety cases
  • The envelope bounds the claim: A release claim covers only the speeds, payloads, floors, clearances, and contact cases for which evidence was gathered, and every dimension of the envelope needs a runtime monitor or a site procedure that keeps the machine inside it.
  • The weakest branch decides: A claim-argument-evidence tree is not an average. High accuracy in one branch cannot offset a missing warrant in another, so the adjudicator looks for the single premise whose failure would permit harm.
  • An unknown premise refuses: Unknown envelope membership is treated as outside the envelope, and a trial marked INVALID or UNOBSERVABLE supplies no evidence. On this book’s illustrative inputs every claim carries such a premise, so the mobile manipulator’s verdict is REFUSE.
  • Refusal is a deliverable: A refusal names the measurements that would lift it, such as loaded stopping trials and enforcer timing under the worst load, and the restriction each result would leave, including the oil film, which no measurement on the machine closes.
  • No verdict moves the proposal boundary: Under UNCOND and COND alike the policy only proposes and the enforcer admits each action. A restricted mode needs its own signed case, never an unsigned downgrade.
  • The manifest expires with its evidence: The release manifest binds the evaluation, placement, authority, and fault records by digest and refuses motion on any mismatch. Wear and drift revoke it, and a software rollback cannot restore a changed plant.

The fourth law (The Four Bedrock Laws), stated as principle \(\ref{pri-vol4-evidence-bounds-authority}\), grants authority on recorded evidence and lets it expire with that evidence, and the standing conditions are where that expiry becomes mechanism. Each pairs a premise with a detector, a cadence, and a response, so the loaded stopping trials that would support a conditional verdict for the mobile manipulator become, once signed, the evidence whose lapse ends it. On this book’s records that grant never arrives, because each of the six claims rests on a premise left illustrative or unknown. The manual lease limited an operator’s authority to the age of its evidence (Authority Transitions), and precommitment tied a test’s verdict to a rule fixed before the evidence arrived (Fault Oracles and Acceptance). Release is the widest grant the book makes, and even there admission stays below the proposal boundary.

What’s Next: From a refused release to what the evidence supports
Which premises can no evidence close? Claim C3 holds the plainest case. No measurement on the machine can exclude the oil film behind it, and more trials with the same sensors cannot change that; only a new transducer, a floor-film detector, could. A conditional release therefore keeps that condition outside the signed envelope rather than settling it. The conclusion, The Epistemic Frontier, returns to the question the book opened with, answers it with what the evidence supports, and enters each premise this release left open, the oil film first, in a residual-claims register with the test that would close it.

Back to top

Footnotes

  1. ISO/TS 15066 Biomechanical Limits: The specification gives body-region- and contact-type-dependent force and pressure data (ISO/TS 15066 2016). A site assessment must select the applicable body region, contact geometry, and measurement procedure. The handover criterion used here is a conservative case criterion, not a universal ISO/TS 15066 threshold or an injury boundary.↩︎

  2. Contact Mechanics Force Model: The elastodynamic model couples translational kinetic energy, effective moving inertia, and contact interface stiffness to determine transient peak impact force during an unmitigated collision. Impacting a rigid boundary concentrates stored kinetic energy into a sub-millisecond impulse with severe force amplification, whereas elastic compliance extends deformation time and dissipates energy over a longer stroke. For the formal elastodynamic derivation and unilateral contact formulations, see Rigid Body Contact Mechanics.↩︎

  3. Goal Structuring Notation: Developed by Tim Kelly at the University of York, GSN provides an explicit graphical grammar for structuring and auditing safety assurance cases (Kelly 1998; Kelly and Weaver 2004). The notation decomposes top-level safety goals into contextual sub-claims, argumentation strategies, and solutions grounded in test artifacts.↩︎

  4. Mechanical Fatigue Budget: A cumulative-damage model such as Miner’s rule (\(\sum_i n_i/N_i\)) can organize a measured load spectrum against component-specific fatigue data; it is an approximation, not a universal failure deadline. The assessed loads, uncertainty, inspection evidence, and stopping consequences determine any signed payload derating and cycle limit. An expired limit revokes the applicable release case rather than silently extending service.↩︎

  5. Hardware Watchdog Interlocks: A watchdog can detect a missed configured heartbeat. An independently wired supervisor then requests the validated drive stop; the detector, wiring, drive, brake, and load response all enter the timing and fault analysis.↩︎

  6. Reflected Rotor Inertia in Impact Dynamics: A geared joint can reflect rotor inertia approximately as \(N^2J_{\rm rotor}\) into the joint coordinate. Its contribution at the contact point also depends on linkage geometry and compliance. The ideal one-dimensional calculation excludes that contribution and brake torque rise; the release case needs measured or modeled bounds before treating its force result as a physical maximum.↩︎

  7. System-on-Chip Resource Sharing: A system-on-chip integrates processor cores, accelerators, memory controllers, and peripheral buses on one die, so neural inference and a real-time monitor that share its interconnect and caches contend for them unless the placement partitions them.↩︎

  8. Cycloidal Drive Dynamics: Cycloidal reducers use an eccentric cam and rolling pins to reach high reduction ratios with little backlash. Rolling-contact wear raises their torsional compliance and adds phase lag between motor and link, which delays braking response and raises peak contact force.↩︎

  9. Recurring Erosion as Accepted Risk: The Presidential Commission on the Space Shuttle Challenger Accident (Presidential Commission on the Space Shuttle Challenger Accident 1986) found that O-ring erosion and blow-by had recurred on earlier flights and that NASA and the motor contractor “accepted escalating risk apparently because they ‘got away with it last time.’”↩︎

  10. Cryptographic Manifest and Silicon Attestation: OTP/eFuse storage can anchor a verification key or its digest. Mutable signed manifests and image hashes reside in protected storage; TPM PCRs accumulate measured boot state rather than directly storing a verdict. On failed validation, revoke permission and execute the assessed stop. A second software bank is an eligible fallback only if it has its own valid signature, release case, and current physical standing conditions.↩︎

ISO/TS 15066: Robots and Robotic Devices, Collaborative Robots. 2016. International Organization for Standardization.
Kelly, Timothy Patrick. 1998. “Arguing Safety: A Systematic Approach to Managing Safety Cases.” PhD thesis, Department of Computer Science, University of York.
Kelly, Tim, and Rob Weaver. 2004. “The Goal Structuring Notation: A Safety Argument Notation.” Proceedings of the Dependable Systems and Networks Workshop on Assurance Cases.
Presidential Commission on the Space Shuttle Challenger Accident. 1986. Report of the Presidential Commission on the Space Shuttle Challenger Accident. U.S. Government Printing Office.