Adversarial Verification

Adversarial Verification

Isometric blueprint showing adversarial cyber-physical verification: four-rung qualification ladder with virtual simulation crucible, processor-in-the-loop board, hardware-in-the-loop dynamometer, crimson adversarial fault injection needles, and surveillance telemetry towers.

Purpose

Why is a robot that has operated for weeks without a single crash not evidence that the machine is safe?

Weeks of incident-free operation establish behavior solely under encountered benign conditions; they provide no empirical evidence that the machine can survive edge-case anomalies. An autonomous robot appears flawlessly reliable until an unmodeled disturbance—a frozen camera frame, an overdue trajectory chunk, or an inductive supply-rail droop—challenges unverified software assumptions. Waiting for rare corner cases in passive operation is unviable: a tenfold rarer event requires ten times the operating exposure to witness, and encountering it during deployment risks catastrophic physical damage.

Whole-system verification transforms safety assumptions into systematic, repeatable fault injection experiments executed across simulation, processor-in-the-loop testbenches, and physical dynamometers. Each experiment challenges timing, data freshness, and kinetic authority boundaries, verifying whether the machine detects faults and enforces safe fallbacks within declared budgets. Within the governance envelope, rigorous verification binds the Brain’s cognitive models, the Nervous System’s deterministic monitors, and the Body’s physical dynamics into empirical evidence capable of justifying, restricting, or denying deployment.

↰ Prerequisite: Physical failure modes and measurement limits providing verification targets originate in Measuring a Machine's Own Limits.

Learning Objectives
  • Define an operating envelope and derive its pass conditions from the machine’s stopping budget
  • Derive a fault list from design records and forward physical analysis, including authority and compound faults
  • Evaluate which simulation and hardware tests support each claim and where their evidence ends
  • Apply zero-miss sizing to a declared, weighted fault population and state what it leaves uncovered
  • Design fault-injection tests that challenge timing, data validity, and authority claims on physical hardware
  • Construct fault records that keep failed, invalid, and unobservable trials distinct from passes

Why Operation Is Not Evidence

An incident-free operating log answers only the conditions it encountered. It does not exercise a frozen camera frame, a brownout, or a bus occupied by DMA traffic. Waiting for those conditions is slow, because a tenfold rarer event takes on average ten times the exposure to observe (the exposure wall of Physical Trial Limits), and on a machine it is also unsafe, because when the event finally occurs in operation it arrives as motion that no software action can recall. Whole-system verification therefore turns each safety claim into a deliberate fault, injected under containment and scored against a test oracle and limits declared before the run. The claims come from the records the earlier chapters produced, ending with the authority record of The Authority Log, which names who may command each actuator channel and what must happen when a handover goes unanswered.

Consider the warehouse mobile manipulator running its aisle at 1.3 m/s (chosen; see the Reader Guide) toward a rack end where a person may step out. Weeks of clean runs at that speed say little about four faults the design claims to contain: a proposer that renews its lease and then stalls (F1), an observation that reaches the permission path with a stale or corrupt capture time (F2), a burst of lost EtherCAT frames between the permission path and the drives (F3), and a sag of the 24 V control rail when an inference burst coincides with the drive motors (F4). None of them shows in an operating log, because each leaves the machine running normally until the moment a stop is needed. The lease catches a stall, the evidence-epoch check a stale capture time, the drive’s fieldbus watchdog a lost frame, and the permission path’s undervoltage monitor, on its own held-up rail, a sag of the shared one. Each fault is injected at the rack end, together with the four authority faults that The Authority Log hands over (A1–A4), and every stopping trial is scored against the onset and distance bounds that section 1.2 derives from the stopping budget.

An injected fault turns a physical run into a test result that can support a safety case only when the run carries five ingredients:

  1. A bounded claim: stating the exact invariant the machine must maintain (for the base at the rack end, a stop that ends inside the 1.10 m clear distance);
  2. Defined test conditions: specifying the initial speed, payload, floor friction, ambient temperature, and supply-voltage range;
  3. A precommitted oracle: an automated evaluator that inspects recorded telemetry against the claim without human intervention or post-hoc adjustments;
  4. A stated sampling and injection procedure: deliberately driving the machine toward its performance limits and injecting faults into sensors, buses, and actuators; and
  5. Provenance: the complete metadata record of model weights, calibration tables, sensor serial numbers, and raw bus logs permitting independent reproduction.

The base’s stopping claim and the arm’s latch-force claim require separate envelopes, injection points, and pass conditions; this chapter follows the stopping claim. The first decision is where nominal regulation must yield to an independent containment response.

The Operating Envelope

The operating envelope is the target ODD (What a Success Rate Measures) extended inward to the machine’s own conditions: compute load, supply voltage, bus jitter, thermal state, and operator response time. The sets nest: the ODD contains the declared ODD, which contains the target ODD, and the operating envelope pairs the target ODD with the machine’s internal conditions. Defining this envelope requires precise numerical boundaries for an automated test framework to drive the system deterministically across each constraint. For the mobile manipulator, the envelope is a set of declared points rather than one nominal condition: the 1.3 m/s aisle speed; payloads of zero and the 50 kg rack capacity; floor friction of at least 0.204, the value the 2 m/s² credible deceleration needs, on dry aisles, and of at least 0.12 on floors under the site’s inspection regime; and ambient temperatures from 5 °C to 35 °C. For inspected floors the envelope declares a restricted point at 1 m/s, the speed of the restricted configuration that the enforcement record declares valid down to that friction (The Enforcement Record), and trials there use a test floor whose measured friction sits at the 0.12 qualification value, so that the stop they exercise is the one friction bounds.

Each point is paired with the machine’s internal conditions: application-processor load up to its peak inference burst, the control rail anywhere from its 21.5 V battery-low point to its 24 V nominal, EtherCAT jitter, and the safety microcontroller’s thermal state. One further point is the 1.2 m/s to which the machine slows before it requests a remote assistant’s in-loop takeover, and only that point carries the 150 ms takeover time (Authority Transitions).

An operating envelope is distinct from the larger survival envelope that encloses it. Inside the operating envelope the learned policy and the tracking controllers regulate the process. When a disturbance, a sensor failure, or a compute stall pushes the machine across that boundary, nominal regulation is conceded, and the objective becomes bringing the body to a bounded, safe state before structural limits are breached. For the mobile manipulator, the survival envelope is the broader region of speed, floor friction, and machine state within which the fallback ladder’s stop rung, and its inhibit rung as the last resort (The Fallback Ladder), can still bring the base to rest short of the rack end and keep the arm’s contact force below its limit. Beyond it lies structural failure, where stored energy exceeds what the hardware can dissipate.

Defining the boundary of the operating envelope requires deriving limits from first principles rather than asserting convenient software timeouts. At the rack end the limit is the stopping budget. Stopping distance grows linearly with the pre-brake delay and quadratically with speed (equation), and the finished budget of Stopping Envelopes (The warehouse mobile manipulator's stopping budget), with its 133.6 ms pre-brake delay, the smooth resident stop that Behavior at the Seam commits to, and the tracking inset, sets a ceiling of 1.39 m/s on a dry floor. At the chosen 1.3 m/s the machine needs 997.4 mm of the 1.10 m clear distance and keeps 102.6 mm in reserve, a distance it covers in 78.9 ms.

Two oracle bounds follow, and every stopping trial in this chapter is scored against both. Brake onset must come no later than 82 ms after the last valid renewal, the lease path of Multi-Rate Cadences: the 60 ms lease, one 1 ms tick of the permission loop, one 1 ms bus cycle, and 20 ms of brake onset. The rail faults of F4 take their onset from the undervoltage flag instead. The stop, measured by instruments independent of the machine, must end within the distance allowed from the same event the fault is referenced to. For a fault that follows a target’s appearance, that distance is the 997.4 mm the budget allows. A stalled proposer offers no target appearance, so its stop, like its onset, is measured from the last valid renewal, and the travel during the 82 ms before brake onset, the resident stop, and the tracking inset give 780.4 mm at the aisle speed.

Neither bound depends on the proposer. A chunk policy that stalls cannot delay brake onset, because the stop begins when the lease expires rather than when the proposer notices, so a test of F1 measures the permission path and not the policy. The budget protects a static obstacle at the clear distance, and a trial against a static target says nothing about a person still walking toward the machine, a premise the site’s crossing rule carries (Belief Through Occlusion).

A flow regulator that keeps a buffer tank below its overflow weir, the Class 3 contrast for this part, verifies a different quantity. Its oracle is volume, not distance. With the level at an 88 L trip point, 7 L below a 95 L weir, and a net inflow of 1.40 L each second after an upstream valve fails and a discharge pump trips, the tank overflows in 5 s. A protective path that detects in 210 ms and seats its isolation valve in 1.20 s finishes 3.59 s early. The same 150 ms bus delay that would cost the mobile manipulator 195 mm, 1.9 times its reserve, adds 0.21 L to the tank. The regulator’s fault list therefore injects sustained delays and failed valves, while the machine’s must resolve tens of milliseconds.

On the mobile manipulator, separate minimum and maximum values for individual variables fail to capture the true operating boundary, because physical and computational dimensions are coupled. The 1.3 m/s aisle speed and an inspected floor with friction 0.12 each lie inside their declared ranges, yet together they fall outside the envelope. On that floor friction, not the brake, sets the credible deceleration, which drops to 1.18 m/s². The resident stop lengthens, and the budget at 1.3 m/s grows to 1.44 m, past the 1.10 m clear distance; the speed boundary on that floor is 1.09 m/s, which is why the envelope’s restricted point sits below it. Payload couples less visibly. It enters no term of the budget, but the credible deceleration was declared for the loaded machine, so a campaign run only unloaded never tests the value the budget rests on. The true operating boundary is a surface in state, parameter, and latency space, and a valid verification suite must probe that combined surface rather than isolated scalar limits.

A boundary stays honest only if every parameter carries its unit, its provenance, its calibrated uncertainty, and an engineer accountable for it. A credible deceleration of 2 m/s² is incomplete until the record says whether it came from stopping trials at the loaded mass, a drive datasheet, or a friction model, and with what resolution the external distance reference measured each stop. A new wheel tread compound or new EtherCAT slave firmware obliges that owner to re-evaluate whether the boundary has moved.

Testing the envelope means attacking its boundary surface. Interior points, where the base crawls unloaded across a dry floor on fresh camera frames, show only that the machine works under ideal conditions. The test framework places points on, just inside, and just outside each consequential boundary. Points just inside check that the policy keeps closed-loop stability under the largest permitted stress; points on the boundary check that margins survive nominal sensor noise; points just outside check that the enforcer takes authority from the policy predictably and halts the base within the survival envelope before the clear distance is gone.

Probing this boundary surface formalizes the critical distinction between classical Functional Safety (governed by ISO 26262 and IEC 61508) and the Safety of the Intended Functionality (SOTIF) (governed by ISO 21448). Classical functional safety frameworks assume that operational hazards arise from system faults: broken wiring harnesses, electrical shorts, bit flips in memory, or software syntax bugs. If every hardware and software component operates strictly according to its engineering specification, classical functional safety considers the system non-hazardous. In physical AI systems, this premise fails. A vision transformer or learned depth estimator can execute with zero memory faults, zero numerical exceptions, and bit-exact instruction execution, yet fail to detect an obstacle due to unmodeled sunlight glare, retroreflective surfaces, or novel semantic objects outside its training distribution. This hazard arises from functional insufficiency—a fundamental performance limitation of learned perception in open environments, occurring in the total absence of an electrical or software fault. Verification of the operating envelope under SOTIF requires systematically identifying these unknown-unsafe conditions through adversarial boundary search, mapping the envelope within which perception is statistically valid, and demonstrating that the deterministic enforcer catches functional insufficiencies before physical clearance is breached. Yet while SOTIF verifies that the system responds safely to sensory ambiguity in the world, the enforcer’s ultimate ability to arrest physical momentum still rests on the deterministic integrity of the underlying hardware under electrical and mechanical stress.

Systems Perspective 1.1: The hardware falsification principle
If a safety property cannot be deterministically falsified by injecting a physical or electrical fault on target hardware, the property is an unverified assumption rather than an engineering invariant.

Once every constraint, latency budget, and environmental limit is stated as a number, the design claims of the preceding chapters become targets for fault injection, each with a threshold that must trigger protective intervention.

Deriving the Fault List

The first source of faults is the design records, inverted. Physical Data through Closed-Loop Evaluation state the learned policy’s properties: its training distribution, its state-action assumptions, and its inference latency ceilings. Sensor Perception through Silicon Placement state interface and execution properties such as bus allocations, observation freshness, and estimation bounds, and Supervisory Intervention states how quickly authority can leave a failing policy and who holds the override. If the machine’s records state that the chunk policy renews its proposal every 50 ms at every speed up to the 1.5 m/s drive limit, the verification suite inverts that claim by injecting stalled renewals, delayed observations, and out-of-range encoder readings to test whether the permission path detects the violation within its budget.

For a single fault, the bounded claim and the injection procedure of that list become a five-element chain comprising the negated condition, the physical injection point, the expected detector, the bounded response, and an observable pass condition, and a recorded design property becomes a testable fault only when all five are defined. For F2, the navigation camera stamps each frame at capture, and the frame’s age at dispatch is budgeted at 51.6 ms (Sensor Transduction and Calibration). Converting that stamp to the safety microcontroller’s clock adds the 0.2 ms bound of the observation contract (The Observation Contract), so the permission path refuses evidence whose upper age exceeds 51.8 ms. The negated property holds a frame at the image signal processor while preserving its capture stamp, so that it arrives past that threshold. The expected detector is the permission path’s check of each proposal’s evidence epoch, the required response is refusal of any proposal built on that evidence followed by the resident stop, and the observable pass condition is the pair of oracle bounds, brake onset within 82 ms of the last valid renewal and a measured stop within 997.4 mm. When any element of this chain cannot be defined or observed, the system contains a verification gap. A verification gap indicates that the claim cannot be falsified by test. Such gaps cannot be ignored; they require dedicated hardware instrumentation, a narrower operational claim, or a redesign of the execution path.

Inversion also reaches the learned proposer’s assumptions, because distribution shift is a system fault even when every frame arrives intact. A chunk policy trained in simulation on dry concrete with friction 0.60 proposes decelerations that a floor at 0.12 turns into wheel slip. Inverting that evaluation claim means injecting the unmodeled traction. Sensor bias and data perturbations are injected for the same reason, because a policy with a high nominal reward can still command jittering actuators once its inputs are perturbed. The machine then fails on valid, uncorrupted commands, because software and mechanics stay coupled across the causal boundary (The Causal Boundary).

At each inter-module contract, the fault list negates the property at its physical point of handoff: observation freshness, state validity, intent lifetime, planning deadlines, enforcement independence, and resource allocation. For the 60 ms chunk lease of Multi-Rate Cadences, the test environment injects an otherwise valid proposal 65 ms after its last renewal and confirms that the permission path rejects it as expired. For the enforcer’s claimed independence from the policy computer, the test saturates the shared memory bus with DMA traffic and confirms that the enforcer still meets its 400 μs deadline inside each 1 ms tick. The discipline cuts both ways. Bus starvation, clock skew, and supply fluctuation are injected where the design claims tolerance, because a fault drawn from a generic menu invents a requirement the machine was never specified to meet.

Inversion has a structural limit. It produces only the faults the design team already modeled, so the list is only as complete as the design claims. If no record says how the base’s spring-applied brake behaves when its lining glazes after repeated stops, there is nothing to invert, and the failure never reaches the suite.

A second, independent source works forward from the physical machine, enumerating what the body, the power network, the floor, and the environment can do regardless of what the software believes. Wheels skid when a commanded deceleration exceeds what the floor can transmit; brake linings glaze under repeated stops; tires wear and change the rolling radius that odometry assumes; drive vibration loosens a camera mount until its calibration drifts. These hazards originate in mechanics and tribology, outside every abstraction the software reasons over. Leveson’s Systems-Theoretic Accident Model and Processes (STAMP) framework1 (Leveson 2011) brings them into hazard analysis.

Leveson, Nancy. 2011. Engineering a Safer World: Systems Thinking Applied to Safety. MIT Press.

For example, start with the physical hazard of a base that keeps driving after the permission path has commanded a stop at the rack end. The unsafe control action is continued drive torque after the stop command. The test injects a welded drive-enable contactor or a stuck brake-release solenoid on the test aisle, observes wheel speed and brake current independently of the commanded stop bit, and requires the permission path to remove drive power through its second channel, safe torque off. The precommitted pass condition is a measured stop within 997.4 mm across the declared speed, payload, and floor-friction points; an unmeasured wheel speed makes the trial unobservable for the braking claim. This forward chain yields an injection and oracle even if the original software requirements omitted a welded contactor.

Vertical ladder showing fault injection climbing from SRAM bit-flips up to physical dyno shaft jamming.

Fault injection climbs from software bit-flips to physical dyno shaft jamming.

Working forward exposes fault classes that no software record implies: clock loss, brownouts that leave logic levels between valid thresholds, drive-switching interference coupling into encoder lines, corrupted configuration memory, and reset lines that bridge domains meant to be independent. If the safety microcontroller and the application processor shared one power-management IC, an undervoltage transient during an inference burst could reset the microcontroller at the moment the proposer loses control. No software record contains that failure, because each module assumes its supply is ideal.

Forward analysis also tests the common-cause independence claims of The Nervous System and Silicon Placement. Two paths isolated on a diagram often share one physical vulnerability, as when the lidar and a redundant bumper switch are routed through one cable carrier, share an analog reference inside one converter, or are built on timer silicon with the same errata. A transient that latches up both channels at once defeats any argument that treats their failures as independent. Common-cause analysis traces routing, power paths, and clock trees so that no single physical event can disable both the policy and the intervention mechanism.

The physical source extends to operator error and foreseeable misuse (Supervisory Intervention). Floor staff jumper a rack-end presence sensor during maintenance, park a pallet jack across a marked crossing, or unplug a warning beacon to silence it, while remote assistants request speeds the aisle does not allow. When a remote assistant commands the 1.5 m/s drive limit in the aisle, the verification framework must verify that the command is denied by capability, or clamped under a manual_drive entry to the 1.3 m/s aisle speed, and that the permission check runs on it like any other proposal.

Working forward from the physical machine adds one representative mechanical class, the actuator jam, to the record-derived list in table 1. Completeness still requires a separate hazard-analysis argument for the specified plant and operating envelope.

Working forward from physical transducers also requires testing against transduction-layer cyber-physical attacks. In connected autonomous systems, physical AI introduces attack surfaces that bypass traditional network firewall abstractions by exploiting the physics of sensory transduction. Resonant acoustic injection (targeting the micro-mechanical proof masses of MEMS gyroscopes with ultrasonic frequencies) can induce false angular rates in inertial navigation without corrupting a single bus packet. Similarly, optical laser injection and rolling-shutter illumination strobes can saturate camera photodiodes to blind range estimators, while adversarial physical patches (high-contrast geometric patterns placed on warehouse walls or floor surfaces) can exploit the gradient landscape of neural perception backbones to cause false negative obstacle detections. Under ISO/SAE 21434 (Road vehicles — Cybersecurity engineering), these threats are co-engineered alongside functional safety faults: a cyber-physical intrusion is evaluated by its ability to cross the causal boundary into hazardous actuation. The verification suite must confirm that the enforcer’s cross-sensor plausibility checks (such as comparing wheel odometry against optical flow, or checking camera evidence epochs against IMU integration bounds) detect transduction anomalies and engage the fallback ladder before false perceptual claims can command torque.

Besides inverted claims and forward hazards, the fault list must test the authority cases A1–A4 of The Authority Log, each a row of table 1, tested with the machine in the state that makes intervention urgent, the loaded base approaching the rack end. The handover deadline is tested at both speeds where Authority Transitions sets it. Once the machine has slowed to 1.2 m/s for an in-loop request, the 209.7 mm reserve gives an unacknowledged handover 174.7 ms to fall back; at the 1.3 m/s aisle speed, for a handover the policy requests without slowing, the reserve’s travel time shrinks to 78.9 ms. A handover tested with the base parked tests none of the deadlines that matter.

Table 1: Representative hardware and cyber-physical fault classes: The rack-end faults F1–F4 and the four authority fault cases (A1–A4) that The Authority Log hands over, each mapped to its injection point, detector, response, and observable pass criterion on the warehouse mobile manipulator, plus one forward-derived mechanical fault.
Failure Domain & Fault Mode Physical Mechanism & Injection Point Target Manifestation & Causal Impact Detection Primitive & Latency Budget Enforced Safe Response & Bounded Recovery Observable Pass Criterion
Proposer Stall (F1) Chunk policy renews its lease once, then its task is killed or its output frozen on the application processor No new proposals; the last admitted chunk keeps supplying setpoints while the base cruises Lease expiry on the permission path (60 ms); host heartbeat timeout (30 ms) Revoke proposal authority; run the resident stop Brake onset \(\le\) 82 ms after the last valid renewal; travel from that renewal to rest \(\le\) 780.4 mm
Stale Observation Time (F2) Navigation-camera frame held at the image signal processor with its capture stamp preserved, or the stamp corrupted, via inline frame interceptor Proposals built on evidence whose upper age exceeds 51.8 ms, the budgeted age plus the clock-conversion bound; the base acts on an old view of the rack end Evidence-epoch age check on every permission-loop tick Refuse proposals built on the stale evidence; run the resident stop Brake onset \(\le\) 82 ms after the last valid renewal; measured stop \(\le\) 997.4 mm
Lost EtherCAT Frames (F3) Frames dropped or corrupted at a drive’s slave port, injected via fieldbus disturbance node Encoder state and drive commands miss cycles; the delivered command ages by one 1 ms cycle per lost frame Working-counter and frame-arrival check; drive communication watchdog Drive enters its configured stop; permission path commands the resident stop Brake onset \(\le\) 82 ms after the last valid renewal; measured stop \(\le\) 997.4 mm; zero unhandled bus exceptions
Control-Rail Droop (F4a) Programmable DC load reproduces the inference-burst-plus-drive transient on the shared 24 V control rail at its battery-low point Shared rail sags to 17.3 V, below the 18 V regulator dropout; the application processor browns out and its sensors may freeze at plausible values Shared-rail undervoltage flag, raised by the permission path on its isolated, held-up rail Command STOP; run the resident stop Brake onset \(\le\) 22 ms after the flag; measured stop \(\le\) 997.4 mm; permission path never resets (continuous heartbeat on an independent digitizer)
Permission-Rail Feed Loss (F4b) Feed to the permission rail’s hold-up store opened during the longest stop the enforcer can admit on an inspected floor, from its 1.09 m/s ceiling Permission path, drive logic, and brake coils run on the store alone Feed-loss flag on the permission rail Complete the resident stop; set every spring brake at standstill Stop completes and every brake sets before the rail reaches dropout (1497.35 ms of the 2 s hold-up); any reset is FAIL
Actuator Jam & Mechanical Stall Wheel bearing seizure or dragging brake, injected via wheel clamp or dynamometer Locked rotor; one wheel blocked yaws the base; rapid stator winding heating (\(I^2 R\)) Motor phase overcurrent & Hall-effect velocity stall monitor, within its detection bound De-energize the power stage; engage the spring-applied brake Stator temperature \(T_{\text{coil}} < 155^\circ\text{C}\) (Class F limit, Thermal Duty Cycles); base stays inside the 0.15 m side clearance
Authority Transfer (A1): Duplicate Writer Second bus master writes a base-drive channel (duplicate EtherCAT identifier or DMA write) while policy and remote assistant both command Two writers on one channel; the delivered command is no longer the admitted one Bus or DMA write permissions; per-tick forensic authority record Reject the foreign write; the enforcer path remains the only writer Write rejected; every forensic authority record in the window shows one writer
Authority Transfer (A2): Tier Bypass Policy proposal injected while MAN is active; remote-assistant command that opposes the enforcer’s admitted limit A lower tier displaces an active higher tier, or an assistant command skips the permission check Arbiter tier check; permission check applied to every source Ignore the policy proposal; clamp or refuse the assistant command like any proposal Arbiter ignores the proposal; the permission check still runs on the assistant’s command
Authority Transfer (A3): Handover Timeout Assistant input frozen during PEND; acknowledgment frames dropped, at the takeover speed and at the aisle speed The machine waits on a handover that no one completes Handover deadline timer in the arbiter Start the validated fallback Fallback starts by \(T_{\text{timeout}}=\max(0,\min(T_{\text{config}},T_{\text{latest}}))\): 174.7 ms at 1.2 m/s, 78.9 ms at 1.3 m/s
Authority Transfer (A4): Authentication Replayed captured veto_path, manual_drive, or accept_item packet; packet carrying a bad tag An unauthenticated source gains authority Tag and freshness check on every assistant and coworker command Lock out the channel; leave authority unchanged Channel locked, authority unchanged, event recorded

Compound faults

Each row of table 1 injects one disturbance, yet a machine in operation rarely meets one at a time. Combining rows at random would multiply the list without finding the combinations that matter, so coupled disturbances are enumerated from shared causes instead. The test ledger of the placement record (Hardware Allocation) names every resource that the permission path shares with the learned workloads, and each shared rail, clock, or bus becomes one compound fault, injected at the cause rather than at its symptoms. A droop on a shared rail can slow the enforcer, freeze a sensor at a plausible value, and reset the logger in the same interval. A burst on a shared bus ages observations while it delays the permission path’s own reads and the heartbeat that would reveal the delay. A fault in a shared clock skews the timestamps of every record that crosses the proposal boundary at once, so freshness checks pass on data that is already old. The oracle for a compound fault therefore checks the protective response against its deadline and also checks that the evidence needed to judge that response survived the same cause; the pitfalls in section 1.8 show both failing together. A shared cause that the ledger does not name cannot be enumerated this way, which is why the forward analysis of the physical machine remains a separate source of faults.

The fault list now holds inverted design claims, forward-derived physical fault classes, the four authority cases, and compound faults drawn from named shared causes. It says what to inject but not where an injection is real, and that choice decides what each trial can prove.

The Qualification Ladder

A simulator can stall the proposer of F1 at every sampled speed and payload overnight, but it cannot sag the control rail of F4, because the simulated rail has no regulator to drop out and no microcontroller to reset. Each evaluation environment closes, breaks, or substitutes different links of the loop that runs from sensing through inference, buses, drives, and plant, so the environments do not form a hierarchy. A higher rung cannot subsume a lower one; each yields a distinct category of evidence and stays blind to the rest (figure 1).

Definition 1.1: Four-stage qualification ladder

Four-stage qualification ladder is the structured progression of cyber-physical evaluation environments that closes, breaks, or substitutes specific arcs in the causal loop to balance combinatorial test throughput against physical and temporal execution fidelity. Its four stages are Software-in-the-Loop (SIL), Processor-in-the-Loop (PIL), Hardware-in-the-Loop (HIL), and in-situ physical fault rigs.

  1. Significance: No single test environment can validate a physical AI system. Software simulators allow massive combinatorial sweeps (\(>10^8\text{ cycles/day}\)) but omit target silicon bus contention and material wear; physical rigs expose authentic mechanics and electrical transients but cannot execute destructive fault sweeps at scale.
  2. Distinction: Unlike a cumulative testing staircase where higher rungs subsume lower ones, each qualification rung exercises complementary, non-overlapping causal mechanisms while remaining blind to others.
  3. Common pitfall: Assuming that passing extensive Software-in-the-Loop (SIL) simulation implies safety on embedded hardware. SIL abstracts away microcontroller clock drift, DMA bus lockups, interrupt jitter, and power supply voltage sag.

Table 2 gives each rung’s substrate, the faults it can inject, its throughput, and where its evidence ends. For the mobile manipulator, software-in-the-loop evidence ends at the plant model, where a brake modeled as a first-order lag has no lining temperature, tire slip, or torque limit. Processor-in-the-loop adds the cross-compiled binary on target silicon and so tests quantization and stack headroom; hardware-in-the-loop benches put the two-processor implementation of Where the Permission Path Runs, its fieldbus, supplies, and drive stages against a field-programmable gate array (FPGA) plant or a dynamometer; and physical fault rigs add the production base, brakes, harness, and floor. Two sources beside the ladder open the loop altogether. Trace replay keeps real sensor noise but supplies the frames that followed the original command rather than the new one. Shadow operation runs the candidate beside an incumbent controller, so if the candidate proposes a velocity that would carry the base past a rack end, its log records the command while the incumbent keeps the base where it belongs. Neither shows the physical consequence of a command.

A shadow deployment cannot let a stalled proposer drive the base toward a rack end that may hold a person, a simulator cannot reveal a race in transceiver silicon, and a hardware bench cannot run the millions of randomized runs that tail statistics need. Each fault injection therefore goes to the rung that exercises its causal mechanism.

Four-stage qualification ladder schematic charting throughput (runs per day, logarithmic from 10^6 down to 10^0) against hardware fidelity (0 to 100 percent). Four semantic stages descending along the Pareto trade-off: Software-in-the-Loop (SIL, cloud cluster, high throughput), Processor-in-the-Loop (PIL, target silicon), Hardware-in-the-Loop (HIL, dyno testbed), and In-Situ (physical fault rig, field plant). Attribution: Original textbook graphic, CC BY 4.0.
Figure 1: Four-stage qualification ladder: Four rungs trading throughput for fidelity, from roughly \(10^6\) software-only runs a day at zero hardware realism down to a handful of in-situ physical fault injections at full realism. The ladder is not a progression toward accuracy but a sequence of different blind spots, because each stage establishes evidence only within its own causal loop and stays blind to every mechanism outside it.
Table 2: The four-stage qualification ladder across cyber-physical boundaries: The four qualification environments compared by computing substrate, exercised causal-loop elements, fault injection primitives, execution throughput, physical and temporal fidelity, and test oracle.
Qualification Rung & Substrate Causal Loop Realism (Compute, Bus, Actuator, Plant) Fault Injection Mechanism & Scope Throughput & Sample Scaling Physical & Temporal Fidelity Primary Test Oracle & Evidentiary Limit
Software-in-the-Loop (SIL) (Host Workstation / Cloud Cluster) Compute: Synthetic / Host x86; Bus: Idealized zero-copy buffer; Actuator: Linear / saturation model; Plant: Numerical ODE/PDE solver Tensor bit-flips, synthetic sensor dropout, floor-friction, payload, and latency sweeps \(10^3\text{--}10^5\times\) real-time parallel sweeps (\(>10^8\text{ cycles/day}\)) High mathematical flexibility; zero target silicon, bus timing, or electrical realism Oracle: Mathematical assertion corridors and state invariants. Limit: Cannot validate target execution timing or unmodeled physical dynamics.
Processor-in-the-Loop (PIL) (Target SoC / MCU Core Emulation) Compute: Real target silicon (ARM/RISC-V); Bus: Simulated bus interconnect; Actuator: Emulated command register; Plant: Real-time host model Instruction cache evictions, register corruption via JTAG/BIST, interrupt jitter injection \(0.1\text{--}1.0\times\) real-time; bound by cross-platform synchronization Instruction-set and compiler accurate; abstracts analog bus lines and actuator back-EMF Oracle: Target register state and WCET execution bounds. Limit: Ignores physical transceiver bus contention, voltage droop, and thermal rise.
Hardware-in-the-Loop (HIL) (Target ECUs + Real-Time Plant Rig) Compute: Production ECU / NPU; Bus: Real physical CAN/Ethernet; Actuator: Real motor drives / dyno; Plant: Hard real-time FPGA plant rig Pin break/short, CAN frame stuffing / babbling idiot, supply rail droop, load steps \(1.0\times\) strict real-time (\(24\text{ h/day}\) wall-clock limit) High electrical, timing, and protocol fidelity; bounded by FPGA plant model accuracy Oracle: Protocol CRC, bus deadline monitors, and interlock timings. Limit: Cannot execute destructive physical overload or simulate tire and lining wear.
Physical Fault Rig (Dynamometer Cell / Field Plant) Compute: Production ECU / NPU; Bus: Production harness & sensors; Actuator: Real physical mechanisms; Plant: Real base, arm, and floor Wheel clamps, magnetic brake stall, thermal heating, transducer bias drift \(1.0\times\) wall-clock; constrained by mechanical wear and thermal reset (\(<10^3\text{ cycles/day}\)) Full authentic physics (Coulomb friction, wheel slip, thermal dissipation, wear) Oracle: Independent physical containment instrumentation and yield transducers. Limit: High financial cost per trial; destructive tests cannot scale to statistical tail bounds.

The fidelity boundary of a test environment is where a model, an emulated bus, or a replayed trace severs the exchange of energy and information between computation and mechanics (The Causal Boundary). A mechanism the model omits cannot fail in it. A joint with unmodeled Coulomb breakaway friction and stator lag sits still for a dead time the simulator never shows and then trails the simulated trajectory, an instance of the omitted-mechanism class of What the Gap Is Made Of that only a rung with the real joint can expose.

To prevent false confidence, every reported verification result must explicitly state which elements of the causal loop were physically real, which were simulated, which were replayed from static logs, which were bypassed, and which remained unobserved. A claim that the permission path stopped the base within 997.4 mm means nothing without specifying whether the wheel-speed signal was generated by a differential equation, read from a physical encoder over the real fieldbus, or replayed from a digital buffer. Specifying this causal mapping exposes the assumptions underlying each test, so that unmodeled physics cannot hide a latent hazard.

Each rung returns a trace, and the causal mapping says which links in that trace were real. Neither says whether the trial passed. The two rack-end bounds of section 1.2 have so far been numbers, and scoring a trial requires turning them into a predicate that is fixed before the run and applied without judgment afterward.

Fault Oracles and Acceptance

An injected fault always produces a reaction, in telemetry transients and in actuator motion, and sometimes a burnt component. Without an oracle fixed in advance, engineers judge time-series plots after the run from gross symptoms, such as whether Joule heating melted a phase wire or a controller threw an unhandled exception, and the test becomes an anecdote. An oracle is the predicate, defined in full before the run, that scores a trial against its claim.

Physical systems fail through transients rather than at isolated instants, so the oracle is a temporal specification. Signal Temporal Logic (STL) combines physical predicates, such as a stopping distance below a limit, with logical connectives and time windows, so that a condition must always hold over an interval, eventually happen within a deadline, or persist until another event occurs (syntax in Signal Temporal Logic and Continuous Robustness).

STL also equips verification with quantitative robustness semantics. Instead of a simple pass/fail flag, robustness is a continuous numerical score measuring how far the system remained from violating its safety margins.2

The rack-end oracle has this form. Brake onset must eventually occur within 82 ms of the last valid renewal, and the base’s travel must always stay within the distance allowed from the event the fault is referenced to, the 997.4 mm the budget allows at 1.3 m/s from the moment a target appears, or 780.4 mm from the last valid renewal for a stall. If an injected fault delays brake onset 30 ms past its bound, the base travels 39 mm farther before it begins to slow and stops after 1,036.4 mm, still short of the 1.10 m clear distance. The distance predicate’s robustness is nonetheless negative by 39 mm, and the onset predicate fails by 30 ms. Each score quantifies the failure in its own unit without manual inspection.

Robustness turns verification from random testing into temporal logic falsification as adversarial optimization. Uniform sampling of a high-dimensional state space needs an intractable number of trials to reach a rare failure, so a falsification engine searches instead for the disturbances, such as sensor noise, voltage drops, and friction, that drive the robustness score lowest. It uses black-box optimizers such as evolution strategies or Bayesian optimization, or gradients through a differentiable simulator where one exists, and a negative score yields a concrete counterexample trajectory.

Formal verification of the network itself, such as the Reluplex solver (Katz et al. 2017), can prove that a small piecewise-linear network never emits an unsafe command over an admissible state set, but its cost grows exponentially with network depth and it does not reach closed-loop dynamics. Verification therefore combines such bounds with falsification, hardware fault injection, and runtime enforcement.

Katz, Guy, Clark Barrett, David L. Dill, Kyle Julian, and Mykel J. Kochenderfer. 2017. “Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks.” International Conference on Computer Aided Verification (CAV), 97–117. https://doi.org/10.1007/978-3-319-63387-9_5.

Scoring also needs an invalid-test rule for a run the environment failed to deliver. If a traction-control derate on a floor joint slows the base before the injection, so that it meets the fault below the declared 1.3 m/s, the trial did not test the envelope point it claims. If the frame injector overflows its own queue and drops the fault packet before it reaches the transceiver, the machine never met the disturbance. Scoring such a run as a pass manufactures confidence, and scoring it as a failure blames the machine for the fixture. The oracle therefore returns one of four verdicts. PASS means every precommitted predicate held; FAIL means a measured predicate was violated; INVALID means the injection or a precondition failed; and UNOBSERVABLE means the truth data needed to score the trial were lost, with the partial trace retained. Under either of these last two verdicts a field the trial did not measure is marked unavailable, never zero. Neither verdict ever counts as PASS, and neither is dropped from the coverage report.

Precommitment separates evidence from post-hoc storytelling. Fault-injection behavior is noisy and surprising, and criteria negotiated after the trace is seen drift to fit it. An overrun is explained away as floor variation, or an onset later than the 82 ms bound is excused because the base still stopped short of the rack end.

The precommitted record fixes the STL predicates, physical units and any normalization scales, response-time bounds, and invalidation triggers before the first injection. The evaluator retains a signed margin for each predicate and returns PASS only if every required margin and timing bound passes. If the stop exceeds its bound by a few millimeters or brake onset misses its deadline by \(2\text{ ms}\), the record states which measured predicate failed; the trace and follow-up analysis determine whether floor friction, brake lag, derating, or a diagnostic delay caused it. Precommitment is where principle \(\ref{pri-vol4-evidence-bounds-authority}\) first binds a test. Release will grant authority on these records, so a verdict scored by a rule revised after the trace was seen would grant authority on the strength of an outcome rather than a test.

Checkpoint 1.1: Temporal logic oracles and test precommitment

Before executing automated falsification against physical AI policies, verify your understanding of temporal logic specifications and test validity criteria:

With the oracle fixed before the first injection, each trial receives one of the four verdicts before anyone argues about its trace, and a signed margin states how far it sat from each bound. The oracle is only as good as the run it scores, however. The ladder can falsify a contract in simulation, but only an injection on the target hardware can show whether a sag on the shared rail or a burst on the shared bus delays the permission path, and whether the target’s own timers meet the deadlines the oracle checks.

Fault Injection on Hardware

A fault injected in simulation tests the simulator’s response. Simulations assume ideal clocks, isolated memory, deterministic context switches, and linear interconnect delay, whereas on the machine, contention on a shared bus, clock drift between microcontrollers, DMA stalls, and voltage drop across inductive drivers decide whether an intervention executes in time. Injection therefore happens on target hardware, through the real transceivers, timers, shared buses, and enforcement path, while an independent containment rig protects the test facility.

In-situ fault injection is the deliberate, deterministic introduction of physical, electrical, and computational disturbances (including sensor register latches, clock drift, bus contention, supply brownouts, and actuator load stalls) directly into target production silicon and mechanics while operating under authentic electrical, thermal, and dynamic loads.

Consider F2 on the hardware. At 1.3 m/s the base covers 1.3 mm in every millisecond, so staleness that goes undetected past the 51.8 ms threshold spends the reserve directly. Left undetected for 78.9 ms, it consumes all 102.6 mm. Neither the camera’s datasheet readout time nor the timestamp on a single frame establishes how quickly the permission path notices. A measured \(P_{99}\) detection time leaves the worst case unbounded (Moving Commands on Time), so the containment claim needs a detection bound justified over the declared compute-load, bus-load, and thermal envelope.

Testing the two-processor implementation (Where the Permission Path Runs) demands physical fault injection at the boundaries its claims name. Table 1 gives each case’s detector, response, and pass condition; the list adds how each fault is injected and, where it matters, where its events are timed:

  • Stale Observation Injection (F2): An inline FPGA interceptor latches a navigation-camera frame at the image signal processor, or suppresses encoder updates, while preserving the original hardware timestamp. Detector assertion, brake onset, and wheel speed are measured on independent instruments.
  • Expired Lease Revocation (F1): The test framework lets the proposer renew once and then suppresses every later renewal. Lease expiry, revocation on the enforcer’s next 1 ms tick, and brake onset are each measured on their own wire and clock.
  • Lost-Frame Injection (F3): A fieldbus disturbance node drops or corrupts EtherCAT frames at a drive’s slave port while the base cruises. Each lost 1 ms cycle costs 1.3 mm of travel at the aisle speed.
  • Control-Rail Droop and Feed-Loss Injection (F4a, F4b): A programmable load draws the 70 A transient of an inference burst coinciding with the drive motors (the brownout of Fallacies and Pitfalls) from the shared control rail at its 21.5 V battery-low point, while the permission path runs on the isolated, held-up rail that The Fallback Ladder sizes and the placement record carries (Hardware Allocation). An independent digitizer records the permission-path heartbeat, and brake onset is timed from the undervoltage flag. F4b opens the permission rail’s feed during the longest stop the enforcer can admit on an inspected floor, the resident stop from its 1.09 m/s ceiling. A reset of the permission path is FAIL in either trial, whatever the stopping distance, because a stop that power loss hands to the spring brakes mid-motion is not the stop the claim budgets.
  • Duplicate Writer Injection (A1): A second bus master writes a base-drive channel, either as a frame under the enforcer’s EtherCAT identifier or as a DMA write to the drive’s command register, while the policy and a remote assistant both command.
  • Tier Bypass Injection (A2): The testbed submits a policy proposal while MAN is active and, in a separate trial, a remote-assistant command that opposes the enforcer’s admitted limit.
  • Handover Stall Injection (A3): The testbed freezes the remote assistant’s input during PEND and drops the acknowledgment frames at both speeds, and the start of the fallback is timed at the drive rather than in the arbiter’s log.
  • Replay and Forgery Injection (A4): The testbed replays a captured veto_path, manual_drive, or accept_item packet and sends another with a corrupted authentication tag.

Recording the state trajectory, the detector identity, the refusal timestamp, and the time to restore the bounded state requires instrumentation electrically and logically independent of the machine, such as optical shaft encoders and high-speed cameras on isolated digitizers. An onboard logging daemon is corrupted by the very failure it tries to capture, because a faulted processor or a saturated bus delays log serialization.3 Independent measurement also rejects false passes, in which the machine survived for reasons absent from the safety claim. An injected lost-frame fault can trip an unmonitored thermal breaker or an auxiliary ground loop that drops power to the drives, so that the spring-applied brake stops the base inside the budget while the software enforcer has deadlocked in a priority inversion. An oracle that inspects only stopping distance scores that trial a pass. The test framework must verify that the enforcer path named in the safety claim executed the corrective action.

A liveness signal is another source of false passes, so the test must target what it can miss. Hardware Isolation recounts the mechanism a plaintiff expert argued in the Toyota unintended-acceleration litigation (Barr 2013), a timer interrupt that kept servicing the watchdog while control tasks could die, and draws the rule that a timer feed proves only that the feed path ran (1.2, 1.1). The test vector for that rule kills the control task, or freezes its output, while the timer interrupt stays alive, and measures the actuator command and the watchdog output on instruments independent of the controller. On the mobile manipulator, F1 is this vector applied to the proposer. The host heartbeat can keep reporting a live process after the chunk-policy task has died, so F1 passes on brake onset measured at the drive, never on a heartbeat in a log. Because the historical mechanism is attributed rather than established, a bench result speaks for the design under test, not for any vehicle.

Systems Perspective 1.2: The watchdog liveness invariant
A watchdog fed by a periodic timer interrupt shows only that the interrupt handler executed, not that safety-critical application tasks met their deadlines or executed without corruption. A qualified fault injection campaign explicitly isolates task progress from timer liveness, testing physical shutdown under task death, thread starvation, and frozen control registers.

War Story 1.1: Testing a task-blind watchdog hypothesis
Context: The Toyota electronic throttle control (ETCS-i) litigation (Bookout v. Toyota, 2013) raised questions about whether a critical throttle-control task could die while a timer-tick interrupt continued to service the hardware watchdog (Barr 2013).

Mechanism to test: Plaintiff expert Michael Barr testified regarding a task-death scenario associated with bit-flips and stack overflow memory corruption. The test vector kills the real-time throttle control task while leaving the hardware timer interrupt active, then independently measures throttle position and watchdog reset pins. Because the NHTSA/NASA investigation did not establish an electronic cause of unintended acceleration, this is an attributed test hypothesis rather than a confirmed vehicle fault.

Oracle: A running watchdog heartbeat is insufficient evidence if the safety claim requires timely, task-specific closed-loop control. The verification oracle must record whether the plant independently returns to a validated safe state (such as engine idle or mechanical brake override) within its physical deadline when task execution freezes.

Systems lesson: Inject the failure mode that liveness signals might miss, then measure the physical response on independent instrumentation. Memory protection units (MPUs), independent supervisory cores, and smart windowed watchdogs that require explicit cryptographic task signatures are options whose fault coverage must be demonstrated on the target platform.

Barr, Michael. 2013. Bookout v. Toyota Motor Corp.: 2005 Camry L4 Software Analysis.

The intervention mechanism must be tested under the conditions that make intervention necessary. An enforcer that takes over cleanly on an idle bench provides no evidence for production, where in a crisis the policy runs at peak throughput, memory buses carry high-resolution sensor streams, DMA channels move frame buffers, and the base carries its loaded rack at the aisle speed. The framework injects faults under the declared thermal and electrical load and measures both lease-expiry revocation and, separately, the time from a diagnostic flag to brake onset at the drive. If the intervention path shares a bus arbiter with the neural accelerator, the takeover can stall behind a burst transfer that cannot be preempted, and a safety case that assumed instantaneous takeover fails in the field if the takeover was verified only on a quiet bus.

↰ Prerequisite: Microarchitectural stress patterns build directly on the memory contention benchmarks from Contention for Shared Resources.

A campaign runs every injection under one fixed protocol, so that a verdict follows from the procedure rather than from whoever reads the trace. The test orchestrator first checks every precondition, such as temperature, rail voltage, and wheel speed, against its declared tolerance, in that quantity’s unit, and a trial that starts outside tolerance ends as INVALID before any fault is injected. It then brings the logic analyzer, the plant, and the truth digitizers onto one reference clock, lets the base settle at its envelope point, and fires the injection on a trigger line that stamps the fault’s start on that clock. The detector’s flag, the takeover at the drive, and the arrival at the bounded state are each timed from that stamp on an isolated analyzer, while truth sensors wired to a crowbar relay can cut actuator power if the base leaves the survival envelope. The precommitted oracle scores the trace from the independent truth instruments under the verdicts and pass rule of section 1.5, and incomplete truth frames surface only at this stage, after the injection has run, as UNOBSERVABLE. The orchestrator signs the record’s header and binds it to the evidence bundle by digest.

↳ Protocol: Fault-injection protocol gives the protocol stage by stage, with its error codes, for a reader building the test orchestrator.

Testing distributed loops exposes clock-conversion risk even when each processor meets its local schedule. F2’s refusal threshold rests on the 0.2 ms bound for converting a capture stamp to the safety microcontroller’s clock, and that bound is itself a claim to test. Take illustrative oscillator limits of ±100 ppm on the application processor and ±50 ppm on the safety microcontroller, plus a separately bounded 50 ppm thermal contribution that those specifications do not include. The worst relative rate is then 200 ppm, and drift alone would consume the 0.2 ms bound after 1 s without resynchronization. A healthy 100 ms sync period limits that drift contribution to 20 μs; network and conversion error must be budgeted separately.

Suppose two good sync frames arrive 400 ms apart while the three scheduled frames between them are dropped. The blackout lasts four sync periods, and the bounded drift contribution grows to 80 μs. If the observed synchronization jitter has standard deviation \(\sigma_{\text{jitter}}\) = 15 μs, adding one sigma to the drift rate \(r_{\text{drift}}\) times the blackout \(T_{\text{blackout}}\) gives \[\Delta t_{\text{one-sigma scenario}}=r_{\text{drift}}\,T_{\text{blackout}}+\sigma_{\text{jitter}} \tag{1}\] or 95 μs. The one-sigma value in equation 1 is not a hard bound; a deterministic deadline needs a justified jitter envelope and a clock-conversion test. A fifth missed period gives 100 μs of drift alone, before jitter, half the 0.2 ms bound on which F2’s threshold rests.

Every injection on the hardware ends with the same obligation. The detector that fired, the moment of refusal, the takeover at the drive, and the physical outcome are measured on instruments the fault cannot reach, on one reference clock, and a trial that loses any of them has not shown what it set out to show. The hardware campaign therefore produces records rather than a pass count, each binding one trial to its claim, its envelope point, its builds, and its evidence.

↳ Downstream: Hardware-in-the-loop stress results furnish the evidentiary warrants for safety cases in Safety Cases and Claims.

The Fault Manifest

When an F4a trial ends, the base has completed its resident stop while the application processor browned out, and the only witness to the order of those events, and to a permission-path heartbeat that never paused, is a trace on an independent digitizer. Release can count that trial only if the trace stays bound to the claim, the envelope point, and the builds it tested. The fault record links a small fixed header to a retained evidence bundle by digest.

Alongside the design contracts, the authority cases A1–A4, and the placement ledger that the fault list already draws on, the record consumes two sources not yet named, the omission log of the policy manifest (The Policy Manifest) and the coverage-gap catalog of the evaluation record (Evaluation Logs), which name the conditions the learned proposer was never shown. Together the header and bundle identify the claim, operating point, injection, hardware and software builds, raw trace, detector, response timing, physical outcome, and anomalies. The fixed header holds the fields every trial shares. Its identity fields name the trial and its claim, the envelope point, the injected fault by mode, magnitude, and target, and a digest of the precommitted oracle. Its outcome fields carry the measured detection and takeover times, the signed margin for the claim, the verdict status, the digest of the evidence bundle, and a signature over the header. The variable trace stays in the bundle.

One F1 record, taken at the envelope’s restricted point, 1 m/s on an inspected floor, with the rack loaded to 50 kg, shows those fields filled in (table 3). As for every stall, the distance oracle is referenced to the last valid renewal. The travel during the 82 ms before brake onset, the resident stop on a floor of friction 0.12, and the tracking inset together give 759.3 mm, shorter than the 960.9 mm the budget allows at that speed from the moment a target appears. The outcome field holds no value and is marked unavailable. No trial in this book has been run, so the field names the measurement that will fill it, and until that measurement exists the record carries no verdict.

Table 3: One F1 fault record at the restricted envelope point: The header fields of a proposer-stall trial on the warehouse mobile manipulator, with the oracle fixed before the run and the outcome left as the measurement that must fill it.
Trial and claim Envelope point Oracle, predeclared from the last valid renewal Detector and witness Outcome to be measured Verdict
F1, proposer stall: chunk-policy task killed one tick after a renewal; rack-end stopping claim 1 m/s; 50 kg payload; test floor measured at \(\mu =\) 0.12; control rail at 24 V nominal Brake onset \(\le\) 82 ms; travel to rest \(\le\) 759.3 mm Lease timer on the permission path; expiry, revocation, and onset each on its own wire and clock Trial count; maximum onset and maximum travel over the trials, each with its signed margin; unavailable until measured None until scored; PASS only if both margins are nonnegative

↳ Byte layout: Fault record lays out the header field by field, including how an unavailable value is encoded, for a reader building the ledger.

The record keeps the four verdicts of section 1.5 distinct field by field. A breaker that prevents the requested injection makes the trial INVALID, and for a trial aborted before injection the detection time, takeover time, and margin are unavailable rather than zero; the bundle records a validity flag for each measured field and why it is absent, and the oracle checks those flags and the verdict before doing any arithmetic, so an unavailable field never enters a pass count or a latency or margin statistic.

A single aggregate percentage hides where the campaign never went. A machine that survived 98 percent of five thousand fault trials has shown little if all five thousand ran at room temperature on a quiet memory bus. Coverage is therefore reported along separate dimensions: functional safety requirements, physical envelope boundaries, physical and logical injection sites, timing tails, hardware build revisions, and multi-fault combinations. For the mobile manipulator, boundary coverage tracks whether faults were injected with the rack empty and full, at the aisle speed on dry floors and at the restricted speed on inspected ones, and with the control rail at its battery-low point. Timing tail coverage isolates injections targeting the \(P_{99}\) scheduling latency of the permission loop and peak direct memory access utilization. A machine achieves acceptable verification not when an average coverage index exceeds an arbitrary threshold, but when the minimum coverage across every individual dimension satisfies the requirements of the safety case.4

Repeated trials count only where something varies. A deterministic fault script run five hundred times on a digital bench, with fixed initial conditions, register states, clock dividers, and inputs, executes the same branch sequence every time and says no more than one run. On the test aisle, repetition samples floor friction, asynchronous interrupt arrivals, bus arbitration, and the phase between the sensor clock and the inference cycle, so it measures the spread and tail of the response-time distribution.

When \(n\) independent Bernoulli trials sampled from the declared fault distribution have zero missed mitigations, the result sets a one-sided upper confidence limit on the miss probability for that distribution. The count is the zero-failure sizing of Physical Trial Limits, applied to an injected fault population rather than to runs of a policy. At 95 percent confidence, a 1 percent limit requires 299 zero-miss injections and a 0.1 percent limit requires 2,995. To infer a FIT-weighted hardware coverage (1.1 defines the FIT rate and the weights \(\lambda_i\)), injections must represent the raw-fault weights \(\lambda_i\) or use a justified stratified weighting and per-stratum analysis.5

The bound applies only when the injection outcomes meet the declared independence and sampling assumptions. Repeated stops heat the drive windings (\(I^2R\)) and the brake lining, change their resistance and friction, and couple later trials to earlier ones. A campaign run on a dry aisle at room temperature cannot establish a limit for an inspected floor or for the envelope’s 5 °C end.

Napkin Math 1.1: Fault coverage under declared assumptions
Scenario: Assume component fault rates \(\lambda_i\) sum to \(\lambda_{\text{raw}}\) = 9,800 FIT, where one FIT is one fault per \(10^9\) operating hours. The four illustrative contributions are sensors 1,500, compute 1,000, transceivers and power management 2,500, and gate drivers 4,800 FIT. Let \(C_{\text{diag}}\) be the fraction of these faults whose hazardous effect is successfully mitigated, weighted by the applicable fault rates. Then \[\lambda_{\text{unmitigated}}=\lambda_{\text{raw}}(1-C_{\text{diag}}).\] A target rate \(\lambda_{\text{target}}\) therefore needs \(C_{\text{diag}}\ge1-\lambda_{\text{target}}/\lambda_{\text{raw}}\), which is 98.9796 percent for a chosen 100 FIT target and 99.8980 percent for a 10 FIT target.

Zero-miss counts such as 299 injections for a 1 percent limit hold only for a fault population sampled in proportion to its \(\lambda_i\) weights. Uniform injections across unequal-rate components estimate a different distribution unless outcomes are weighted and each stratum has adequate evidence. A test must record which mode was injected, whether it was detected, and whether the plant actually reached its allowed state.

Operating hours cannot close what the injections leave open. A zero-event test cannot reach the hazard rates a release claim needs (the exposure wall of Physical Trial Limits), so the fault record carries what was never injected as prominently as what passed. It lists every parameter combination omitted because of bench limits, sensor bandwidth, or destructive-test cost, and an untested region therefore cannot reach the safety case of Deployment Release as a verified margin.

This mathematical boundary between empirical road exposure and safety qualification is formalized as the astronomical exposure wall (figure 2). As established by Butler and Finelli (1993), demonstrating an ultra-reliable failure rate \(\lambda\) with statistical confidence requires zero failures over an exposure horizon scaling inversely with the target: \(N \approx -\ln(1 - C) / \lambda\), which yields \(N \approx 3.0 / \lambda\) at the standard 95 percent confidence level. For safety-critical automotive systems governed by ISO 26262 ASIL D (\(\lambda \le 10^{-8}\ \text{h}^{-1}\), or 10 FIT) or civil aviation systems governed by FAA regulations and IEC 61508 SIL 4 (\(\lambda \le 10^{-9}\ \text{h}^{-1}\), or 1 FIT), validating life-critical dependability purely through statistical black-box operation requires 300 million to 3 billion consecutive zero-failure hours—spanning 34,000 to over 340,000 continuous vehicle-years. Even the largest commercial autonomous fleets in history, such as Waymo’s driverless operations logging over 200 million miles (approximately 7 million operating hours) or Tesla’s customer-supervised L2 fleet logging over 8 billion miles (approximately 250 million hours), face an insurmountable verification deficit of one to three orders of magnitude against ASIL D and SIL 4 standards. Consequently, statistical black-box fleet testing cannot qualify life-critical autonomy; system safety must instead be mathematically certified through formal runtime architecture, deterministic tripwires, and targeted hardware-in-the-loop fault injection.

Butler, Ricky W., and George B. Finelli. 1993. “The Infeasibility of Quantifying the Reliability of Life-Critical Real-Time Software.” IEEE Transactions on Software Engineering 19 (1): 3–12. https://doi.org/10.1109/32.210303.
Figure 2: The astronomical exposure wall: target failure rates versus required zero-failure operating hours: Log-log scaling of required continuous zero-failure operational hours (\(N = 3.0 / \lambda\) at 95 percent confidence, dark blue; \(N = 4.61 / \lambda\) at 99 percent confidence, dashed slate) across target hourly failure rates \(\lambda\). Benchmark thresholds mark police-reported crash rates, industrial standards (IEC 61508 SIL 1 and SIL 2), human driver fatality baselines (NHTSA), automotive functional safety (ISO 26262 ASIL B and ASIL D), and commercial aviation catastrophe ceilings (1 FIT / SIL 4). Horizontal dashed lines display real-world empirical operating hours across leading autonomous fleets (Cruise pre-suspension, Waymo rider-only driverless, and Tesla customer-supervised L2). The vertical red span highlights the 43-fold verification deficit between Waymo’s 200 million driverless miles (~7 million hours) and the 300 million zero-failure hours required to statistically demonstrate ISO 26262 ASIL D compliance.

Fallacies and Pitfalls

A machine can pass every isolated fault test and still fail when a shared rail, bus, or clock couples the faults or disables their detectors. These are the compound faults that section 1.3 enumerates from shared causes. Compound tests catch such failures only when enough galvanically or logically independent instrumentation survives the fault to establish the sequence of events.

Fallacy: Independent local watchdog timers guarantee end-to-end physical deadline compliance.

The rack-end stop must begin within the 133.6 ms pre-brake delay that the budget reserves. Each stage passes its local watchdog: the camera meets its exposure and readout times, the perception backbone its inference time, the lease trips at 60 ms, and the brake begins within 20 ms of its command. A queue between the image signal processor and the backbone that holds frames an extra 100 ms under load lies between those checks, so no individual timeout detects it. At 1.3 m/s it adds 130 mm of travel against a reserve of 102.6 mm, and the stop overruns the clear distance by 27.4 mm. Only a check that compares each observation’s capture time with the moment the permission path acts on it sees the queue. Verification must trace the complete sensing-to-actuation interval under coincident load and include communication jitter in the budget. Local watchdogs remain useful, but their bounds support an end-to-end claim only when the composition, including every handoff and buffer delay, fits the physical response envelope.

Pitfall: Allowing diagnostic logging and high-throughput telemetry to share fieldbuses with real-time safety loops.

Suppose the mobile manipulator shares one EtherCAT segment among sensing, actuation, enforcement, and evidence logging. A diagnostic memory dump raises bus utilization to 98 percent, inducing network transmission jitter that delays both the drive commands and the telemetry needed to judge their physical effect. The same bandwidth contention also delays safety heartbeats and the logs that would explain a failure. Testing each stream separately misses this common dependency. Critical traffic requires bounded service through validated time-sensitive networking (TSN) scheduling and hardware priority, or separate physical channels, while diagnostic traffic must remain within its assigned budget. Compound-load testing must confirm that logging cannot prevent either the safety response or the observation of that response.

Pitfall: Relying on software anomaly detectors that share electrical power or failure modes with monitored actuators.

A control-rail brownout freezes a wheel-encoder channel at a plausible 1.3 m/s while a drive fault carries the base at its 1.5 m/s drive limit, a speed at which the finished budget no longer fits the clear distance (Stopping Envelopes). The encoder frames remain well formed and the value stays within nominal limits, so the permission path, which reads that channel, does not recognize the overspeed. An isolated drive-fault test would not reveal this common-cause failure if the encoder’s power remained healthy. Verification must examine whether one fault can disable the evidence needed to detect another. Galvanically isolated sensing and explicit checks for specified stuck-signal behavior must support the response, rather than assuming a plausible measurement is a functioning measurement.

Fallacy: Treating an unobservable test outcome where diagnostic loggers dropped packets as evidence of safety.

During a compound-fault test, bus contention drops logger packets and a watchdog reset erases the volatile SRAM trace. The machine does not strike the rack end, but the team can no longer reconstruct its peak impact forces, its \(\mathrm{SE}(3)\) trajectory, or its response timing. That observation cannot establish the force and stopping limits the acceptance criteria specify, because they were not measured. The trial must remain inconclusive for those claims, with the recording failure treated as an instrumentation defect. Before repeating the test, the team must make the required evidence path survive the injected electrical conditions. A visually favorable outcome cannot substitute for the quantitative measurements specified by the safety acceptance criteria.

Summary

A mobile manipulator that runs its aisle for weeks without incident has shown that its nominal loop worked under the conditions it met. To test its stopping claim, the team must also stall its proposer, age its observations, drop its fieldbus frames, and sag its control rail at the rack end, measuring whether brake onset follows the event each fault is referenced to within its bound and the base comes to rest inside the distance its oracle allows from that event. The four authority rules are attacked on the same loaded approach, each against its own pass condition, from a rejected foreign write to a fallback that starts before the handover deadline. Each test links a design claim, injection, detector, response, and measured plant outcome. Passing one condition narrows uncertainty; it does not prove every combination of faults and aging states safe.

The nominal operating envelope and the separately declared survival envelope give these tests meaning. Crossing the nominal boundary can be a valid survival test if the injected state, instrumentation, and acceptance oracle remain inside the test plan. A trial becomes INVALID when the injector or preconditions fail; it becomes UNOBSERVABLE when required truth data are lost. A valid failure beyond the nominal envelope can expose an inadequate fallback or containment limit. These verdicts preserve evidence that a simple pass percentage would hide.

Simulation can sample many friction and timing scenarios but cannot establish unmodeled lining wear or electrical response. A physical test samples those effects under its particular hardware and environmental conditions. Zero-miss injections bound a miss probability only for their declared fault distribution; zero-event operating hours bound an incident rate only under the declared exposure model. Neither substitutes for measured actuator response and an independently admitted fallback.

↳ Downstream: Untested operational regimes identified during falsification feed the epistemic boundary analysis in Epistemic Limits.

Key Takeaways: Hardware-in-the-loop verification and fault injection
  • Trace the complete response: Precommit the fault, detector, deadline, physical oracle, and evidence path. Local watchdog passes do not compose into an end-to-end limit.
  • Challenge shared resources: Coincident bus, power, and sensing faults can disable both a protective response and its witness.
  • Keep verdicts distinct: A valid violation is a failure; broken preconditions are INVALID; lost required telemetry is UNOBSERVABLE.
  • State the evidence population: Report tested operating and survival conditions, fault weights, timing tails, builds, and gaps. A zero-event count is a conditional confidence limit.
  • Carry the record forward: The versioned fault record supplies one part of the release safety case, alongside analysis, architecture, and field evidence.

What’s Next: From fault records to a release decision
Verification hands release the form of the fault records and the gaps they leave. For each claim, that form fixes the envelope points and builds its trials must reach, the oracle set before each run, and the verdicts a trial can return, and it names the regions no trial reaches. No record yet carries a measured verdict, so none says whether the machine may drive the aisle. Deployment Release takes that record form, together with the authority, placement, and evaluation records, and asks which claims they support, which verdict follows, and which standing conditions must keep holding for that verdict to stay valid.

Back to top

Footnotes

  1. STAMP/STPA Hazard Analysis: Developed at MIT by Nancy Leveson, STAMP treats safety as a control problem rather than a component-reliability problem, so hazards appear as unsafe control actions at interfaces, such as a seam where bus latency or filter lag erodes a stability margin. STPA is the hazard-analysis method built on it.↩︎

  2. Quantitative Robustness Semantics: Each predicate has a signed margin in its own physical unit: clearance may use meters, brake onset milliseconds, and contact force newtons. A negative margin falsifies that predicate. A minimum across unlike units is meaningless; a combined dimensionless STL score requires each margin to be divided by a predeclared positive scale of the same unit. The record retains raw margins and scales so a pass or failure remains physically interpretable.↩︎

  3. Non-Intrusive Hardware Tracing: Non-intrusive tracing utilizes dedicated silicon hardware debug blocks (such as ARM CoreSight Embedded Trace Macrocell) and external FPGA bus sniffers that snoop internal bus transactions without stealing CPU execution cycles. By buffering execution traces in dedicated high-speed SRAM decoupled from the system memory bus, hardware analyzers capture cycle-accurate fault-reaction timestamps without perturbing real-time thread schedules. Relying on software logging daemons introduces Heisenberg observer effects, where logging overhead masks the exact race conditions under test.↩︎

  4. Scope of assurance standards: ISO 26262-5 addresses road-vehicle electrical and electronic hardware. Its single-point and latent fault metrics concern the architecture and its allocated safety goals; they are not the aggregate \(C_{\text{diag}}\) of 1.1. Fault injection is one way to test diagnostic assumptions, alongside analysis and other evidence. ANSI/UL 4600 frames an autonomous-vehicle safety case from multiple kinds of evidence. Neither standard makes a fixed number of HIL trials sufficient for release.↩︎

  5. Trial sampling and aging: The binomial count assumes independent, identically distributed Bernoulli outcomes under the declared injection distribution; it does not require Poisson event arrivals. Repeated heating, wear, or contact changes can break independence and require a state-stratified test design.↩︎