Safety Enforcement

Safety Enforcement

Isometric blueprint showing runtime safety enforcement for an autonomous mobile robot: omnidirectional mecanum base, translucent mint-green forward-invariant safe set boundary, candidate cyan neural proposal vector pointing toward a prohibited crimson hazard wall, amber CBF-QP projected admissible vector, and crimson emergency brake interlock.

Purpose

How does an unprivileged software proposal become physical torque, and what lets the gatekeeper refuse it in time?

A neural policy may propose commands that violate geometric clearance, exceed joint velocity limits, or saturate actuator torque. Once an unverified command is converted into motor phase currents, detecting the neural network’s hallucination comes too late: kinetic energy is already committed. An independent real-time enforcement path must evaluate candidate proposals before they cross the causal boundary and retain the uncompromising hardware authority to reject or project them within strict, sub-millisecond deadlines.

Timely refusal alone is insufficient; the selected fallback response must remain dynamically feasible under actual physical state. A mechanical brake cannot deliver idealized deceleration if thermal saturation or supply-rail voltage drops degrade torque generation. Runtime enforcement pairs control barrier functions with real-time feedback loops to verify that stopping envelopes remain strictly inside verified clearance margins. In the physical AI stack, enforcement embodies the core duty of the deterministic Nervous System: maintaining an unyielding permission path between statistical proposals from the Brain and physical execution by the Body.

Learning Objectives
  • Explain the relationship between tracking accuracy, command permission, and independent enforcement at the actuator boundary
  • Diagnose how update rate, feedback delay, and actuator saturation limit corrective control
  • Calculate the finished stopping budget and the speed ceiling it sets from response delay, braking, localization, and tracking error
  • Evaluate with a velocity-dependent barrier whether a tracker command computed from an admitted proposal must be accepted, projected, or refused
  • Select fallback responses when sensing, computation, or actuator limits invalidate normal enforcement
  • Construct an enforcement record that binds each check to its detector, its fallback response, and the records it consumes

The Deterministic Gatekeeper

The safety microcontroller (MCU) admits the warehouse mobile manipulator’s run down the aisle only after the matched stop suffix named in its trajectory record (The Trajectory Contract) is resident in the MCU’s memory and feasible from every state reachable before commitment. That admission checks a plan once, before the motion starts. The machine then executes the plan one millisecond tick at a time, and on each tick the live state can drift from the one the planner assumed. The base lags its commanded reference, a person steps out at the rack end, a payload loads the arm, or the chunk policy proposes a setpoint the admitted plan never contained. What follows is the check made on every tick, and the response it selects when even the resident stop no longer fits.

The learned proposers behind those commands can emit infeasible paths or joint commands beyond the actuators’ torque limits, and by the irreversibility law of The Four Bedrock Laws no exception handler can recall the momentum that current in a winding imparts, so a bad command has to be stopped before it reaches the drive.

Enforcement therefore keeps the task competence of learned policies while denying them authority over geometric clearance, thermal limits, and human-safety invariants. The enforcer of Multi-Rate Cadences, the permission path’s per-tick gatekeeper, sits on the permission side of the proposal boundary. It is a hard real-time component that evaluates candidate proposals, verifies invariant state bounds, and holds the exclusive hardware authority to permit current flow or command a fallback stop.

In the design this chapter follows, the enforcer runs on bare-metal microcontroller hardware wired directly to the motor drives and joint encoders and evaluates incoming setpoints each millisecond control cycle against validated constraints using Control Barrier Functions (CBFs)1 (Ames et al. 2019). When a proposal is safe, it passes to the motor drives unaltered. When a proposal marginally threatens a boundary, an active-set Quadratic Program (CBF-QP) minimally modifies the command, minimizing torque deviation while enforcing the derived constraint.2 When a proposal is structurally invalid or the proposer falls silent, the gatekeeper severs the command channel and triggers an autonomous fallback deceleration sequence.

However, an enforcer cannot guarantee safety beyond the tracking fidelity of the control loop beneath it, because every barrier certificate assumes that the tracking controller follows commanded references within a bounded error.

What the Tracker Hands Over

The tracker, the feedback controller that turns each admitted reference into a torque command on every tick (What Planning Hands Over), isolates the enforcer from the plant’s internal mechanics, and only one of its guarantees crosses the interface upward. That guarantee is an explicit bound on the worst-case tracking error (\(\epsilon_{\text{track}}\)) between where the machine was commanded to be and where its mass actually is. Within a validated envelope of reference acceleration, load, temperature, and disturbance, the tracker may provide a transient bound \(\|x(t) - r(t)\| \le \epsilon_{\text{track}}\) over the complete maneuver. The enforcer does not inspect transfer functions, root loci, or controller gains; to the software above the loop, the tracker is a black box that holds physical motion within the error bound \(\epsilon_{\text{track}}\) of the commanded path.

This bound is a physical distance set by rotor inertia, transmission compliance, and actuator torque limits. A steady-state stiffness calculation balances unmodeled disturbance against the restoring effort of the closed-loop stiffness, but it does not bound transient overshoot or settling, so the bound also needs validated dynamics or full-maneuver testing.3

On the warehouse mobile manipulator’s base, full-maneuver validation establishes \(\epsilon_{\text{track}} =\) 40 mm (illustrative; see the Reader Guide). The controller holds that bound only while three conditions hold. The reference cannot demand accelerations beyond what the drives can deliver, external disturbances must remain below \(d_{\text{max}}\), and the actuators must operate below the continuous thermal derating limits of Thermal Duty Cycles and the limits of their voltage rail. A reference that demands more acceleration than peak torque allows saturates the loop, the integrator winds up, and the base falls behind its reference by an amount that grows with every millisecond of saturation. Outside its rated envelope, the controller provides no bound at all.

Because the base can lag its reference by up to \(\epsilon_{\text{track}}\), every geometric and dynamic margin the enforcer evaluates is inset by that distance. A reference that keeps clearance \(c\) from the rack end guarantees the base only \(c - \epsilon_{\text{track}}\), so the enforcer checks the reference against the required clearance plus 40 mm, and section 1.5 carries that inset into the stopping distance, where delay and localization error also consume clearance. If worn drive bearings, tire wear, or altered motor parameters doubled the true error to 80 mm without the record changing, the base would spend another 40 mm of clearance that the enforcer believes it still holds. A degraded tracking bound silently erodes every margin above it.

↰ Prerequisite: Spatial clearance insetting from sensor covariance and transport delay originates in Spatial Error and Clearance Bounds.

A tracking bound never checked at runtime is an assumption, not a contract. The Nervous System must publish \(\epsilon_{\text{track}}\) with the operating conditions that validate it and monitor the instantaneous error \(e(t) = \|x(t) - r(t)\|\) on every tick. When \(e(t)\) exceeds the bound, under a disturbance spike or thermal throttling, the contract has broken, and the permission path must invalidate the spatial assumptions of the filters above it before unmodeled deflection causes damage.

Systems Perspective 1.1: The inset margin invariant
Every spatial and dynamic safety boundary evaluated by an enforcer must be inset by the underlying tracking loop’s verified worst-case tracking error, and the stopping envelope of equation 1 carries that inset, together with the travel during transport lag, into every admission check. An enforcer that evaluates barrier certificates against nominal geometric boundaries without insetting will permit trajectories that physically collide with obstacles whenever tracking error reaches its certified bound.

The synthesis of tracking controllers, observers, and disturbance rejection belongs to the control literature (Slotine and Li 1991; Åström and Murray 2021). In modern physical AI implementations, the “tracker” is rarely an elementary single-input single-output feedback loop; it is typically structured as a cascaded control hierarchy (translating Cartesian path targets into joint velocity references, and velocity into field-oriented phase currents at \(10\text{--}25\text{ kHz}\)) or as a high-rate tracking Model Predictive Controller (MPC) running at \(100\text{--}500\text{ Hz}\) on bare-metal firmware to solve a localized quadratic program for dynamic disturbance rejection. Yet regardless of its internal mathematical sophistication, an optimal tracker remains semantically blind: it possesses no internal model of obstacle clearances, task intent, or human proximity, and it will execute a collision path into a steel upright with the same dynamic fidelity it brings to unobstructed free space. Furthermore, when unmodeled physical disturbances occur—such as surface oil causing gross tire slip or sudden thermal current derating—even an optimal tracking controller can guarantee tracking only within its qualified error tube \(\epsilon_{\text{track}}\). The enforcer therefore treats the tracker strictly as a contract boundary: it encapsulates the entire feedback control stack inside the certified bound \(\epsilon_{\text{track}}\), evaluating candidate proposals against forward-invariant barrier certificates before any command reaches the tracking layer.

Slotine, Jean-Jacques E, and Weiping Li. 1991. Applied Nonlinear Control. Prentice Hall.
Åström, Karl Johan, and Richard M Murray. 2021. Feedback Systems: An Introduction for Scientists and Engineers. Princeton university press.

The Authority to Refuse

A chunk arrives from the application processor on time, well formed, and carrying a fresh evidence time, and one of its setpoints would drive the arm into a rack upright. Nothing in the packet marks that setpoint as wrong. The permission path therefore grants no authority to an incoming learned proposal until that proposal passes the configured state, input, freshness, and stopping-clearance checks. This is the third law, proposal is not permission (principle \(\ref{pri-vol4-proposal-permission}\)), enforced one tick at a time. A learned policy carries no stability guarantee, no bound on its output variance, and no representation of actuator current limits or mechanical stress, so its outputs are advisory.

The Simplex architecture (Sha 2001) is one realization of the permission path. It resolves the tension between task performance and safety by decomposing the control system into four structural components:

Sha, Lui. 2001. “Using Simplicity to Control Complexity.” IEEE Software 18 (4): 20–28.
  1. The advanced controller: An unverified, high-capacity learned model (such as a Vision-Language-Action network, a residual reinforcement learning policy (Johannink et al. 2019), or a diffusion policy executing on the application processor) that optimizes for complex, open-ended task performance.
  2. The baseline safety controller: A certified, high-assurance deterministic controller (on this permission path, the matched stop suffix admitted with the trajectory in Behavior at the Seam, resident in bare-metal MCU firmware) that acts as the trusted low-level stabilization baseline of the physical plant, providing tested stability and stopping behavior under stated conditions.
  3. The safety manager: Deterministic arbitration logic that continuously evaluates the state of the machine against forward-invariant barrier sets, tracking error bounds, and execution deadlines.
  4. The output selector: An atomic hardware or real-time firmware multiplexer that routes actuation signals to the motor bridge. The selector admits a checked proposal when conditions hold and transfers to a resident baseline within its qualified response delay when those conditions fail.
Johannink, Tobias, Shikhar Bahl, Ashvin Nair, Jianlan Luo, Avinash Kumar, Matthias Loskyll, Juan Aparicio Ojea, Eugen Solowjow, and Sergey Levine. 2019. “Residual Reinforcement Learning for Robot Control.” IEEE International Conference on Robotics and Automation (ICRA), 6023–29. https://doi.org/10.1109/ICRA.2019.8794127.

Under this contract, the upstream policy never commands the plant directly. Instead, communication across the causal boundary operates through the chunk lease of Multi-Rate Cadences, a dead-man permission that expires unless renewed by a fresh proposal. The application processor transmits candidate control vectors as structured packets over shared SRAM mailboxes through a configured interprocessor channel, such as RPMsg.4 Each proposal opens with the proposal header (Multi-Rate Cadences), and the MCU checks its evidence time and expiry on each control cycle, adding the header’s clock-conversion error bound. If the proposal is fresh, the state remains in the validated viable region, and tracking error remains bounded, the output selector permits the proposal (or its minimally projected variant) to drive the actuators. If the lease expires without renewal, tracking diverges, or the proposal threatens physical boundaries, the output selector revokes permission and engages the baseline safety controller unilaterally.5 The upstream policy is not consulted, does not negotiate, and possesses no hardware mechanism to override the selector.

When a proposal conflicts with physical constraints, the enforcer chooses between filtering and veto. Filtering applies when the proposal is structurally sound and a nearby admissible command still serves its intent. The spring-latched cage door that the arm opens supplies the case. The chunk policy’s proposal to approach the door’s strike plate at 0.10 m/s is well formed, but the door zone admits only the 0.03 m/s guarded approach of the target ODD, the one speed its demonstrations and trials cover (Evaluation Logs). That ceiling marks where the evidence stops, not where the physics of the plate does. The enforcer therefore keeps the approach’s direction and target and projects only its speed onto the 0.03 m/s ceiling, a constraint that the least-change quadratic program of section 1.8 enforces beside the half-spaces a Control Barrier Function derives.

Filtering must respect momentum as well as position, because the actuators must still be able to stop the machine before clearance vanishes (section 1.7). Projection is permissible only when the corrected command keeps the proposal’s direction and target, changes it continuously, and stays well clear of a hazard. A structurally invalid proposal, such as a discontinuous jump in commanded position, an acceleration beyond the available torque, or a trajectory aimed into an immovable obstacle, indicates that perception or inference has entered an unmodeled failure mode, and projecting it onto the nearest safe boundary would execute a motion unrelated to any task intent. The enforcer vetoes such a proposal, and the output selector drops authority to a local fallback.

Filtering and veto both happen at runtime, and training the policy cannot replace either. As Fallacies and Pitfalls argues, penalty terms and barrier losses shape the distribution a policy learns but do not bound what the deployed policy emits, and constrained reinforcement learning bounds only the expected cost of violations, which permits a rare high-momentum collision provided the average stays below its limit. The permission path therefore checks each candidate at its actuation boundary whatever objective trained the proposer, and when the proposal is safe, the check is transparent and the learned competence passes to the drive unchanged.

The chunk lease, 60 ms on this machine, also keeps the enforcer’s authority when the upstream host stalls. A separate 100 \(\text{Hz}\) host heartbeat has a 30 ms silence timeout. The 1 kHz MCU loop feeds its own 5 ms self-watchdog only after a successful cycle; a separate external trip covers MCU failure. At each timeout the MCU revokes the proposal and executes a resident, state-matched fallback, provided the reserved clearance and available braking still suffice. On the warehouse mobile manipulator, if a proposer renews its lease and then stalls, the brake still begins within 82 ms of the last valid renewal. This lease-expiry case assumes the separate heartbeat remains live while fresh admissible proposals stop arriving; a total host stall triggers the 30 ms heartbeat timeout earlier. When the danger is a person stepping out at the rack end, the age of the frame that first showed that person adds to the renewal-to-onset delay, and section 1.5 carries the sum into the admission check. Enforcing the lease at the drive, rather than trusting the host to withdraw its own proposal, is what bounds the delay at all.

Severing upstream authority requires a defined physical response at the lowest hardware interface. The zero-command floor is the configured response when active software control disappears, and writing zero to a drive produces different mechanics on different actuators. Zero current in a low-friction direct drive lets a gravity-loaded arm descend unless a separate restraint carries the load, while zero velocity in a stiff position loop can request a large transient. Electrical Power Integrity describes the three unpowered states a drive can fall into (free coasting, phase-short dynamic braking, or a spring-applied brake); which one each drive enters is a choice this chapter makes, rung by rung.

Refusal is a steady-state decision, not only an emergency exception. On each millisecond control tick the enforcer checks the proposal’s freshness against the lease and the evidence-age threshold, the host heartbeat, the tracking error, the actuator and kinetic limits, and the feasibility of the resident fallback, and if one check fails, it revokes the proposal and selects the resident response appropriate to the current state and available clearance. That path must remain able to act when the upstream proposal source fails, and where it sits in the signal chain decides what it can see and how quickly it can act.

Where the Enforcer Sits

The enforcer’s position in the signal pipeline is set by the dynamics’ timescale and the latencies of the channels that feed it, and every machine that turns learned proposals into motion offers two natural positions. Upstream of the feedback loop, a kinematic reference checker validates waypoints, coordinate targets, or velocity vectors before they enter the tracking loop; at planning rates of \(10\) to \(50\text{ Hz}\) it can check a \(300\text{ ms}\) trajectory segment against workspace boundaries and self-collision models. Its strength is spatial foresight and its limitation is blindness to the plant’s actual state, so a reference that clears a rack upright by less than the base’s transient overshoot passes every upstream check while the base still strikes the upright.

↳ Downstream: Heterogeneous silicon partitioning of the enforcer onto isolated microcontroller enclaves is mapped in Two Paths on One Die.

Downstream of the feedback controller, near the power electronics, a command checker inspects phase currents, duty cycles, and joint torque references immediately before they are written to hardware registers. Running at the machine’s 1 kHz inner-loop rate, it prevents overcurrent, over-torque, and thermal runaway by clipping or zeroing signals beyond the plant’s ratings, but it surrenders foresight. A torque well inside its continuous rating looks safe to a scalar clamp even while it accelerates a descending arm joint toward the packing station faster than the remaining travel can absorb, and the clamp registers no fault until no allowable reverse torque can prevent the collision.

The split trades latency for visibility. A reference checker at those planning rates sees several steps ahead but cannot respond to resonance, noise amplification, or instability above its update rate. A command checker acts on every tick, but its one-step horizon creates dynamic traps in which every instantaneous torque is legal until an irrecoverable state is reached.

Placement also decides how the enforcer joins the computing fabric. An in-line enforcer inspects every command before the next stage reads it, which leaves no command uninspected but puts its worst-case execution time on the critical path of the 1000 μs loop, where an overrun faults the whole cycle. A parallel monitor on a separate core preserves loop timing and opens a safety contactor or switches the drive to dynamic braking when it detects a violation, but its detection and signaling lag, bounded by the monitor cycle and bus jitter, must be paid for with wider margins or lower speed.

Because neither single location provides both dynamic foresight and instantaneous protection, physical AI architectures resolve the tension through a dual-stage enforcement pattern. This pattern maps onto the two-processor implementation of the machine model (Where the Permission Path Runs):

  • Upstream Stage (Advisory Corridor Pre-Check on the Application Processor): Operating at the chunk rate (20 \(\text{Hz}\) on this machine) on an unprivileged application microprocessor core, this stage screens multi-step trajectory chunks emitted by the learned policy against a dynamic stopping corridor under worst-case braking parameters and against static workspace envelopes before packing the waypoints into a leased chunk proposal sent over RPMsg. The stage is advisory. It spares the MCU proposals that would fail, but it runs on the proposer’s side of the proposal boundary, so the binding admission of the trajectory and its stop suffix stays on the MCU (Behavior at the Seam).
  • Downstream Stage (Instantaneous CBF-QP Safety Shield on Bare-Metal MCU): Operating inside a strict 1 kHz (1000 μs) timer interrupt loop on the isolated MCU, this stage receives the proposed setpoint from shared SRAM, verifies tracking error bounds, solves an active-set Control Barrier Function (CBF-QP) quadratic program in pre-allocated static SRAM, and clamps raw current and torque commands directly before writing to PWM timer registers. The barrier it enforces reduces to a single-step convex projection (section 1.7), whose execution time must be bounded on the selected hardware.

The advisory pre-check screens for a feasible stop before a chunk is sent, and the drive-side stage checks current observations and commands on every tick, but neither stage can recover a state that has already exhausted braking clearance.

Correct placement does not help if the enforcer’s inputs share failure modes with the component it checks. An enforcer that computes collision distances from a depth map produced by the same neural perception model that planned the motion is deceived by the same out-of-distribution artifact, so independent enforcement requires independent sensing and isolated computation. The MCU reads encoders and limit switches wired to its own GPIO and timer-capture registers, outside the Linux virtual memory space and scheduler, and a collision check reads a safety lidar or proximity sensor whose conditioning uses deterministic threshold logic rather than learned inference. Where no independent sensor exists, as for the friction under the base’s wheels or the mass of a payload that only an estimator infers, no independent physical guarantee can be built, and the claim is only as strong as the estimator’s bounded uncertainty or the premise the record states. Every invariant the enforcer checks, upstream or at the motor terminals, still assumes that the plant retains the authority to arrest its own motion before a forbidden boundary.

Systems Perspective 1.2: The monitor independence rule
The permission path must never evaluate a state estimate produced by the component it is designed to constrain, nor share an un-preemptible communication bus, memory matrix, or clock tree with unverified processes. Where independent sensing is physically impossible, the safety claim rests on the estimator’s bounded model uncertainty, not on hardware.

Stopping Envelopes

The chapters that own a stopping term have each added it to the warehouse mobile manipulator’s budget at the rack end, and one term remains. The stopping envelope is the set of states from which a validated braking response can arrest motion before a boundary, and the irreversibility law (principle \(\ref{pri-vol4-irreversibility}\)) sets its test, \(d_{\text{stop}} \le D_{\text{clear}}\) with every term of \(d_{\text{stop}}\) taken from a measured limit of this machine. Reserving that response costs achievable speed, because velocity increases blind travel linearly and braking travel quadratically. The machine must remain inside this envelope before a fault or unsafe proposal occurs.

The pre-brake delay \(\tau_{\text{delay}}\) in the stopping distance of Kinetic Momentum already includes the enforcer’s own permission-loop tick, which Multi-Rate Cadences budgets with the lease path. What remains is the gap between the reference and the plant that follows it.

The permission path closes that gap with the tracking bound \(\epsilon_{\text{track}}\) of section 1.2 and defends the stopping distance of equation inset by it. A commanded reference can be admitted only if the plant, which may lag that reference by up to \(\epsilon_{\text{track}}\), can still stop inside \(D_{\text{clear}}\). The admission condition is the stopping envelope:6 \[v\,\tau_{\text{delay}} + \frac{v^2}{2a_{\text{eff}}} + \delta_{\text{loc}} + \delta_{\text{margin}} + \epsilon_{\text{track}} \le D_{\text{clear}} \tag{1}\] Here \(a_{\text{eff}}\) is the deceleration the resident stop actually uses, equal to the credible deceleration \(a_{\text{brake}}\) for a constant-deceleration stop and smaller when the stop is shaped as in Behavior at the Seam; \(\delta_{\text{loc}}\) is the localization bound and \(\delta_{\text{margin}}\) the fixed protective clearance. The left side is the stopping distance \(d_{\text{stop}}(v)\) of equation plus the tracking inset \(\epsilon_{\text{track}}\).

Under those delay, estimate, and braking assumptions, any trajectory proposal that places the machine at a distance less than \(d_{\text{stop}}(v) + \epsilon_{\text{track}}\) in equation 1 from a physical constraint along the direction of motion is unsafe and physically inadmissible. The stopping envelope is fundamentally anisotropic; an obstacle behind the instantaneous velocity vector requires clearance only for the positional tracking error \(\epsilon_{\text{track}}\).

For the warehouse mobile manipulator, the inset is the base’s tracking bound of 40 mm, and it is the last distance term the budget receives. At the 1.5 m/s drive limit, the matched stop suffix of Behavior at the Seam had already carried the running total to 1,194.15 mm, past the 1.10 m clear distance at the rack end. The inset brings it to 1,234.15 mm, 134.15 mm more than the clear distance.

Every other term in equation 1 is now fixed by the chapter that owns it: the credible deceleration and the fixed overheads by Kinetic Momentum, the lease path by Multi-Rate Cadences, the observation age by Sensor Transduction and Calibration, the stop’s shape by Behavior at the Seam, and the inset here. Speed is the one term the machine can still choose. With \(\tau_{\text{delay}} =\) 133.6 ms and the suffix’s \(a_{\text{eff}} =\) 1.33 m/s² (equation), equation 1 holds up to a ceiling of 1.39 m/s. The finished budget therefore sets the aisle speed at 1.3 m/s, below that ceiling. At the aisle speed the stop needs 997.4 mm and leaves 102.6 mm of the clear distance unspent. Table 1 lists every row at both speeds, and figure 1 draws them.

Table 1: The warehouse mobile manipulator’s stopping budget: Each numbered row is the distance one chapter adds, loaded and on a dry floor, at the 1.5 m/s drive limit and at the 1.3 m/s aisle speed; the running total and remainder are at the drive limit, against the 1.10 m clear distance. Row 6 replaces constant-deceleration braking with the resident \(C^2\) stop, so its entry is the added distance. Time rows add no distance.
Row Added in Term At drive limit (mm) Running total (mm) Left of \(D_{\text{clear}}\) (mm) At aisle speed (mm)
1 Kinetic Momentum Braking, \(v^2/2a_{\text{brake}}\) 562.5 562.5 537.5 422.5
2 Kinetic Momentum Overhead, \(\delta_{\text{loc}}+\delta_{\text{margin}}\) 150 712.5 387.5 150
3 Kinetic Momentum Brake onset, \(vT_{\text{act}}\) 30 742.5 357.5 26
— Supply, Freshness, and the Memory Wall Time only: the chunk policy can renew a lease; the intent model’s 98.0 ms weight sweep cannot — — — —
4 Multi-Rate Cadences Lease path, \(v(T_{\text{lease}}+T_{\text{tick}}+T_{\text{bus}})\) over 62 ms 93 835.5 264.5 80.6
5 Sensor Transduction and Calibration Observation age, \(v\,t_{\text{age}}\) over 51.6 ms 77.4 912.9 187.1 67.08
— Belief Through Occlusion Time only: belief that an occluded rack end is clear expires in 31.25 ms, under two camera frames, so memory cannot certify it — — — —
6 Behavior at the Seam \(C^2\) stop suffix, \(0.75v^2/a_{\text{brake}}\) in place of \(0.5v^2/a_{\text{brake}}\) 281.25 1,194.15 \(-94.15\) 211.25
7 section 1.5 Tracking inset, \(\epsilon_{\text{track}}\) 40 1,234.15 \(-134.15\) 40
Total 1,234.15 997.43
Left of \(D_{\text{clear}}\) \(-134.15\) 102.57

Stacked bar chart of stopping distance in millimeters. Seven bars at the drive limit grow from braking alone to the full budget, adding overhead, brake onset, lease path, observation age, the matched stop suffix, and the tracking inset; the sixth and seventh bars cross a dashed horizontal line marking the clear distance at the rack end. An eighth bar at the aisle speed, built from the same terms, ends below the line.

Figure 1: The stopping budget, row by row: Each bar is the running total after one chapter adds its term, at the drive limit, loaded, on a dry floor. Braking, overhead, brake onset, the lease path, and the observation age stay inside the clear distance; the matched stop suffix carries the total past it, and the tracking inset widens the overrun. The final bar is the finished budget at the chosen aisle speed, which fits inside the clear distance with the remainder that later chapters spend.

The finished budget protects a static obstacle at \(D_{\text{clear}}\), such as a person who has stepped out from the rack and stopped or a dropped tote, so it assumes that the obstacle stays where it was when the stop began. A person who keeps walking toward the base through the whole stop closes the gap from the far side as well (Belief Through Occlusion), and counting that approach on top of the finished budget would cap the base at 0.46 m/s at this rack end. An aisle held to that speed would slow every run for an encounter that a crossing rule can keep out of the stop, so the site answers the walking person with the rule rather than with speed, and A Residual-Claims Register keeps open the premise the rule rests on, that no person approaches the rack end during a stop. The enforcer therefore guarantees the budget against a stopped obstacle, not the safety of a person who keeps walking into the stop.

Later chapters spend the remainder or restrict the premise beneath it. At the aisle speed the unspent distance lasts 78.9 ms (Authority Transitions). The credible deceleration binds only on floors with friction of at least 0.204, since \(a=\min(a_{\text{brake}},\mu g)\), so on a floor inspected to 0.12 the ceiling falls to 1.09 m/s (A Residual-Claims Register).

A stalled proposer would spend more than the whole remainder. Were the chunk lease of section 1.3 left to the host to honor, a last proposal that kept its authority for 200 ms instead of the 60 ms lease would add 182 mm to the stop and carry the base 79.4 mm past the clear distance at the rack end.

The same budget also sizes any protective sensing that stops the base independently of the proposer. A mobile base can sense the rack end through a risk-assessed protective field designed under the applicable vehicle and machinery standards, including ISO 3691-4 and ISO 13849, such as the time-of-flight safety laser scanner shown in figure 2, and that protective field must be at least as long as the stopping budget.

Industrial safety laser scanner with outer warning and inner protective fields. The protective field initiates a safety-controller response whose total sensing, control, drive, braking, and uncertainty budget determines the required field length; drive torque removal alone may allow coasting.
Figure 2: Industrial safety laser scanner: A scanner can provide an early warning field and a protective field. The warning field requests earlier deceleration; the protective field’s safety output feeds a controller that selects a risk-assessed stop. Protective-field length must include scanner response, controller and drive delay, braking travel, and uncertainty. A moving AMR does not become safe merely because motor torque is removed.

Phase portrait with obstacle distance on the horizontal axis and approach speed on the vertical axis. A braking-only parabola and an inset stopping boundary show the cost of pre-brake delay, the matched stop suffix, and the fixed allowances; a marker at the rack-end clear distance shows the speed ceiling, with the chosen aisle speed just below it.

Figure 3: Conditional stopping envelope: For the warehouse mobile manipulator’s base, a parabolic boundary separates states with modeled braking clearance from states that have already exhausted it. The inset includes the pre-brake delay, the matched stop suffix’s gentler deceleration, and the localization, protective, and tracking allowances. The curves are valid only for the stated credible deceleration and operating envelope.

In the distance–velocity plane of figure 3, the stopping equation draws a parabolic boundary between states that can still stop and states that have exhausted their clearance, and the gap to the braking-only curve is the clearance consumed by pre-brake delay, the matched stop’s gentler deceleration, and the fixed allowances. As the roofline model (Williams et al. 2009) caps attainable throughput by memory bandwidth, the stopping envelope caps legal speed by clear distance, however quickly the chunk policy proposes. To run faster, the engineer must raise the credible deceleration \(a_{\text{brake}}\), shorten \(\tau_{\text{delay}}\), or tighten \(\epsilon_{\text{track}}\) and \(\delta_{\text{loc}}\).

Williams, Samuel, Andrew Waterman, and David Patterson. 2009. “Roofline: An Insightful Visual Performance Model for Multicore Architectures.” Communications of the ACM 52 (4): 65–76.

These parameters are not stationary. An item in the gripper raises the arm’s inertia and lowers the deceleration its joints can produce within peak torque, oil on the floor lowers the traction beneath the base’s credible deceleration, and repeated hard stops heat rotors and inverters into thermal derating (Thermal Duty Cycles). An enforcer that treats \(a_{\text{brake}}\) as a datasheet constant underestimates the envelope precisely under heavy load, so it must use independently bounded braking capability and reduce permitted speed when validated friction, payload, or thermal conditions change; an uncertain estimate cannot by itself restore a guarantee.

Adaptation works only while the enforcer’s state estimates are accurate and the actuators can still supply the force. If contamination takes traction toward zero or a sustained duty cycle pushes the windings past their ceilings, the assumed deceleration vanishes, and once the required torque exceeds what the hardware can produce, the enforcer can detect the violation but can no longer prevent it.

When the Enforcer Runs Out of Authority

In control engineering, actuator saturation is a bounded nonlinearity that adds phase lag and causes windup in feedback controllers. For the permission path it is more immediate, the instant a motor is delivering all it can and the enforcer loses the capacity its safety warranty assumes. When an actuator reaches its mechanical limit (peak torque), its electrical limit (bus voltage or current clipping), or its thermal limit (Thermal Duty Cycles), the corrective action \(u_{\text{corr}}\) cannot be delivered at the commanded magnitude, and no boundary can be protected with force the body cannot generate.

A stopping envelope computed under nominal assumptions becomes invalid as soon as authority is consumed, because torque spent holding a payload against gravity or climbing a grade is unavailable for braking.7 Under a steady-state bias torque \(\tau_{\text{bias}}\), the effective deceleration available to arrest angular motion is: \[\alpha_{\text{eff}} = \frac{\tau_{\max} - \tau_{\text{bias}}}{J} \tag{2}\] where \(J\) is the mechanism inertia and \(\tau_{\max}\) is the peak dynamic torque limit. In linear translational systems with mass \(m\) and bias force \(u_{\text{bias}}\), equation 2 takes the equivalent form \(a_{\text{eff}} = (u_{\max} - |u_{\text{bias}}|) / m\). A drive that spends three-quarters of its torque holding steady speed on an incline keeps one-quarter for braking, which quadruples its braking distance.

Consider the warehouse mobile manipulator’s arm carrying the mug from the takeaway conveyor down toward the coworker at the packing station, with the loaded joint descending at an angular velocity \(\omega_0 =\) 1.2 \(\text{rad/s}\) (the scenario’s values are illustrative; see the Reader Guide). The effective link length is \(\ell =\) 0.80 m, so the tool point moves at 0.96 m/s, inside the arm’s 1 m/s free-space limit; only the final approach to the coworker’s hand runs at the far lower handover speed that section 1.11 records. The joint inertia with the mug is \(J =\) 3 kg·m², and the joint’s peak torque is \(\tau_{\max} =\) 87 N·m. With no gravitational bias (\(\tau_{\text{bias}} = 0\)), the whole peak torque is available for braking, and the angular deceleration is \(\alpha_{\text{nom}} = \tau_{\max}/J =\) 29 \(\text{rad/s}^2\). The stopping angle from \(\omega_0\) is \(\Delta \theta_{\text{nom}} = \omega_0^2 / (2 \alpha_{\text{nom}}) =\) 0.0248 \(\text{rad}\) (1.42\(^\circ\)), which moves the tool point \(d_{\text{tool,nom}} = \ell \Delta \theta_{\text{nom}} =\) 19.86 mm along its arc.

During the descent, however, the arm’s own geometry and the mug load the joint with a downward bias of \(\tau_{\text{bias}} =\) 65.25 N·m, 75 percent of its peak torque. By equation 2, the net torque left for braking is \(\tau_{\max} - \tau_{\text{bias}} =\) 21.75 N·m, and the angular deceleration falls fourfold to \(\alpha_{\text{eff}} =\) 7.25 \(\text{rad/s}^2\). The stopping angle grows fourfold to 0.0993 \(\text{rad}\) (5.69\(^\circ\)), and the tool travels 79.45 mm, an unbudgeted overrun of 59.59 mm. If the enforcer evaluates barrier envelopes with the datasheet torque \(\tau_{\max}\) rather than the net headroom \(\tau_{\max} - \tau_{\text{bias}}\), the real stop runs 59.59 mm past the predicted one. Whether that overrun ends on the station surface depends on the clearance at which braking begins, which the barrier of section 1.7 evaluates. The coworker’s hand must never come within the arm’s biased stop of 79.45 mm plus its blind travel before the handover channel engages.

Integral action beneath the enforcer adds a latency penalty during saturation. While the actuator is clipped and tracking error persists, the integrator keeps accumulating a command the plant cannot execute, so when the enforcer commands full braking the output stays pinned at the forward limit until that integral unwinds. If error accumulates for 200 ms, unwinding delays net braking force by \(\tau_{\text{unwind}} =\) 150 ms. On the mobile manipulator’s base at the 1.3 m/s aisle speed, this delay adds \(v\,\tau_{\text{unwind}} =\) 195 mm of unbraked travel, almost twice the 102.6 mm that the finished stopping budget of section 1.5 leaves spare. Anti-windup clamping, back-calculation, and conditional integration address the loop (Åström and Hägglund 2006; Franklin et al. 1998); the enforcement requirement is to measure the unwinding delay that remains and budget its unbraked distance inside the stopping horizon.

Åström, Karl Johan, and Tore Hägglund. 2006. Advanced PID Control. ISA-The Instrumentation, Systems,; Automation Society.
Franklin, Gene F, J David Powell, and Michael L Workman. 1998. Digital Control of Dynamic Systems. 3rd ed. Addison-Wesley.

Saturation is one of the few hardware failures with a direct, deterministic detector inside the permission path. On each \(1\text{ ms}\) tick the permission path compares the commanded effort \(u_{\text{cmd}}\) with the realizable effort \(u_{\text{act}} = \text{clip}(u_{\text{cmd}}, u_{\text{min}}, u_{\text{max}})\). A nonzero difference, or an inverter status register reporting current-limit saturation, shows that the actuator has exhausted its authority, and the duration and magnitude of the clamping track the remaining authority without noisy external state estimation.

Because consumed authority changes the stopping distance at once, the enforcer folds it into the envelope in real time, recomputing \(a_{\text{eff}} = (u_{\text{max}} - |u_{\text{bias}}|) / m\) before verifying the next trajectory segment. The allowable speed falls as authority is consumed, and when the computed stop no longer fits inside the clear distance, the enforcer must command deceleration while enough authority remains.

Authority depletion must also reach the planner. A 20 \(\text{Hz}\) policy that keeps commanding accelerations to a motor the 1 kHz loop has reported saturated is planning against a body that cannot execute them, so the permission path reports saturation as an active constraint violation, revokes the active trajectory, and requires a re-plan from the machine’s authority-constrained state.

Systems Perspective 1.3: The actuator authority invariant
A Control Barrier Function, as formulated in Safety and Control Barrier Functions, is mathematically valid only while physical actuators retain sufficient current, voltage, and thermal headroom to deliver the computed corrective force: \(u^* \in [u_{\min}, u_{\max}]\). When steady-state gravity or friction consumes the actuator’s operating envelope, the enforcer must dynamically contract the safe set before control authority at the margin vanishes.

Checkpoint 1.1: Stopping envelope and actuator headroom

Consider how the stopping envelope, and the actuator authority beneath it, bound permissible motion:

Safe Sets as Conditional Permission

Once its bias torque is counted, the descending arm of section 1.6 needs 79.45 mm of tool travel to stop, and that travel grows with the square of approach speed. A clearance that suffices at one speed is therefore too short at a higher one, and the enforcer needs a predicate over position and speed together that it can evaluate on every tick of the control clock. In control theory, a safe set \(\mathcal{C}\) is defined as the zero-superlevel set of a continuously differentiable scalar certificate \(h(x): \mathcal{X} \to \mathbb{R}\), such that \[\mathcal{C} = \{x \in \mathcal{X} \mid h(x) \ge 0\}, \quad \partial \mathcal{C} = \{x \in \mathcal{X} \mid h(x) = 0\}.\] The mathematical foundation of forward invariance, established by Blanchini (Blanchini 1999), formulated through Hamilton-Jacobi reachability analysis and viability kernels by Mitchell, Bayen, and Tomlin (Mitchell et al. 2005), and extended through Control Barrier Functions by Ames et al. (2014) and Ames et al. (2019),8 characterizes conditions under which an admissible controller can keep a state within a viable set. Geometric clearance alone is not that set when stopping torque is finite.

Blanchini, Franco. 1999. “Set Invariance in Control.” Automatica 35 (11): 1747–67.
Mitchell, Ian M., Alexandre M. Bayen, and Claire J. Tomlin. 2005. “A Time-Dependent Hamilton-Jacobi Formulation of Reachable Sets for Continuous Dynamic Games.” IEEE Transactions on Automatic Control 50 (7): 947–57. https://doi.org/10.1109/TAC.2005.851439.
Ames, Aaron D, Jessy W Grizzle, and Paulo Tabuada. 2014. “Control Barrier Function Based Quadratic Programs with Application to Adaptive Cruise Control.” 53rd IEEE Conference on Decision and Control, 6271–78.
Ames, Aaron D, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada. 2019. “Control Barrier Functions: Theory and Applications.” European Control Conference (ECC), 3420–31.
Definition 1.1: Control barrier function safety filter

Control barrier function safety filter is the constrained optimization filter that projects unverified nominal control proposals \(\mathbf{u}_{\text{nom}}\) onto a forward-invariant safe set \(\mathcal{C} = \{\mathbf{x} \mid h(\mathbf{x}) \ge 0\}\) by solving the quadratic program: \[\min_{\mathbf{u}} \frac{1}{2}\|\mathbf{u} - \mathbf{u}_{\text{nom}}\|^2 \quad \text{s.t.} \quad \mathbf{a}(\mathbf{x})^\top \mathbf{u} \le b(\mathbf{x}), \quad \mathbf{u}_{\min} \le \mathbf{u} \le \mathbf{u}_{\max}\] that enforces linear braking half-space constraints (\(\mathbf{a}(\mathbf{x})^\top \mathbf{u} \le b(\mathbf{x})\)), leaving candidate inputs that satisfy the safe boundary untouched and minimally projecting violating inputs in control space.

  1. Significance: Rather than attempting the impossible task of proving formal safety across billions of uninterpretable neural network weights, a CBF-QP filter decouples high-level policy innovation from low-level safety enforcement. The learned model is free to explore complex manipulation or navigation strategies, while the deterministic enforcer can preserve a validated viable set while its physical assumptions remain true.
  2. Distinction: Unlike a heuristic clamp that saturates control channels independently (destroying multi-axis coordination), a CBF-QP filter performs an optimal orthogonal projection in control space, minimizing the chosen torque-space distance; it does not automatically preserve a Cartesian path.
  3. Common pitfall: Formulating barrier constraints using raw geometric coordinates without accounting for relative degree (\(r = 2\)) or braking inertia. If the barrier does not incorporate stopping distance (\(\Delta d_{\text{stop}} = v\,\tau_{\text{delay}} + v^2 / 2a_{\text{brake}}\)), the filter intervenes too late, after the vehicle’s momentum has already made a boundary collision inevitable.

When the tracker turns a learned policy’s unverified setpoints into a nominal torque \(\mathbf{u}_{\text{nom}}\), the enforcer does not forecast the machine’s entire future path. It evaluates a single half-space inequality at the current state \(\mathbf{x}\), the linear constraint that bounds the safe approach toward the boundary. If the proposed action satisfies \(\mathbf{a}(\mathbf{x})^\top \mathbf{u}_{\text{nom}} \le b(\mathbf{x})\), the enforcer passes it to the motor untouched; if it violates a feasible half-space, the enforcer projects its torque vector.

A barrier on the arm’s distance to the station surface alone would register the danger only once momentum had already made contact inevitable,9 so the barrier must carry the kinetic stopping angle (\(\omega^2 / 2\alpha\) for approach rate \(\omega\) and braking bound \(\alpha\)) as well as position, a requirement that follows from relative degree. Actuators command current and torque, torque produces acceleration, and position therefore lies two integrals from the input: \[\text{Actuation } u \xrightarrow{\int \frac{1}{m} dt} \text{Velocity } v \xrightarrow{\int dt} \text{Position } x\] This two-step cascade defines relative degree \(r = 2\). Differentiating clearance once yields velocity: \[\dot{d} = -v\] The input \(u\) does not yet appear, because a motor cannot step velocity instantaneously. Only the second derivative brings in Newton’s second law (\(m \ddot{x} = u - d_{\text{ext}}\)): \[\ddot{d} = -\dot{v} = -\left( \frac{u - d_{\text{ext}}}{m} \right)\]

A raw-distance rule \(\dot d\ge-\gamma d\) cannot directly constrain this torque input: \(\dot d=-v\) contains no \(u\). For an approaching arm, let \(D=d_{\text{tool}}/\ell\) be the remaining angular clearance, \(\omega>0\) its approach speed, \(J\dot\omega=\tau+\tau_{\text{bias}}\), and \(\alpha\) a validated lower bound on available angular braking. Define the kinetic barrier \[h(D,\omega)=D-\frac{\omega^2}{2\alpha},\qquad \dot h=-\omega-\frac{\omega(\tau+\tau_{\text{bias}})}{J\alpha}. \tag{3}\] The raw distance \(D\) has relative degree two; this velocity-dependent \(h\) has relative degree one for \(\omega>0\). Requiring \(\dot h+\gamma h\ge0\) produces the actual affine torque half-space \[\tau\le-\tau_{\text{bias}}+J\alpha\left(\frac{\gamma h}{\omega}-1\right),\qquad -\tau_{\max}\le\tau\le\tau_{\max}. \tag{4}\] At standstill (\(\omega=0\)), the controller uses a separate no-approach and acceleration admission check; it does not divide by zero.

Carry the descending-arm scenario of section 1.6 into this barrier: \(J=\) 3 kg·m², \(\tau_{\text{bias}}=\) 65.25 N·m toward the station, \(|\tau|\le\) 87 N·m, \(\alpha=(\tau_{\max}-\tau_{\text{bias}})/J=\) 7.25 \(\text{rad/s}^2\), \(\ell=\) 0.80 m, \(\omega=\) 1.2 \(\text{rad/s}\), and \(\gamma=\) 10 \(\text{s}^{-1}\). These are stated joint-level scenario inputs; the bias is not inferred from the payload mass alone. At 110 mm of tool clearance to the station surface, \(D=\) 0.1375 \(\text{rad}\) and \(h=\) 0.0382 \(\text{rad}\), so equation 4 requires the joint to brake with at least 80.1 N·m. At 79.45 mm, \(h=0\) and the requirement reaches 87 N·m, the joint’s full peak torque. At 70 mm, it demands 89.1 N·m, more than the joint can produce, so this state is already outside the modeled viable set. The enforcer must prevent entry into it; a fallback started there cannot restore the lost clearance.

A real machine realizes neither exact state feedback nor instantaneous torque. Stator current rise, bus jitter, and gearbox compliance delay and soften the corrective action, so the enforcer evaluates the barrier against a boundary inset by the worst-case stopping deficit they cause and intervenes earlier in the trajectory.

The safety warrant provided by the barrier calculation is strictly conditional. A forward-invariance claim requires a feasible barrier and the following bounded conditions; the list is necessary for this design, not an exhaustive if-and-only-if test:

  1. Model Fidelity: The nominal plant dynamics \(f(x)\) and \(g(x)\) must capture the true system dynamics within a certified parameter error bound.
  2. Disturbance Bounds: External forces and unmodeled friction acting on the mass must remain strictly within the modeled disturbance envelope \(|d(t)| \le d_{\max}\).
  3. Actuation Authority: Physical actuators must possess sufficient voltage, current, and thermal headroom to realize the commanded corrective force without saturation (\(u^* \in [u_{\min}, u_{\max}]\)).
  4. Solver Determinism: The optimization solver or projection algorithm must terminate strictly within its allocated execution budget (\(T_{\text{solve}} \le T_{\text{budget}} < T_{\text{loop}}\)).
  5. Discretization Invariance: The discrete sampling interval \(T_{\text{ctrl}}\) must be sufficiently small that inter-sample state drift does not breach the continuous-time invariance boundary.

Each condition maps to a budget. Model fidelity and disturbance bounds come from system identification and load testing on the bench, actuation authority from dynamometer curves and thermal limits (Thermal Duty Cycles), solver determinism from static worst-case execution time analysis on the MCU, and discretization from the loop period, which at \(T_{\text{ctrl}} =\) 1 ms must keep the zero-order-hold drift between samples to a fraction of \(\delta_{\text{margin}}\).

A condition that no detector watches is an assumption, so the permission path attaches one to each: a disturbance observer compares predicted with measured acceleration on every cycle, inverter status registers and current shunts report actuator saturation, a hardware watchdog and cycle counter time the solver against its deadline, and timestamp counters on the sensor bus catch jitter and missed samples.

When any detector trips, the premise of the safe set has failed, and the enforcer must not recompute the barrier with optimistic parameters. It transfers authority through the output selector to a state-matched fallback whose response time and braking margin were reserved before the fault; later rungs can only limit harm once the viable set is lost. The safe set is a valid permission only while its operating contract is intact, and the permission path enforces safety by knowing when that contract has failed.

Minimal Intervention

Suppose the chunk policy proposes setpoints from which the tracker computes a torque that the kinetic barrier of equation 4 forbids. Replacing the learned commands with a hand-crafted conservative controller discards the competence that justifies the model, and braking hard at the first boundary encounter creates acceleration spikes and interrupts the task. The principle of minimal intervention instead leaves an admitted proposal unchanged when it satisfies every active constraint and changes it only when projection is both needed and feasible. Enforcement then acts as a safety shield (projection operator), accepting the proposal as the default intent and computing the smallest correction that keeps the system within the invariant set.

↰ Prerequisite: Nominal trajectory setpoints and feedforward torques are synthesized by the planner in From Goal to Trajectory.

Ames, Aaron D, Xiangru Xu, Jessy W Grizzle, and Paulo Tabuada. 2016. “Control Barrier Function Based Quadratic Programs for Safety Critical Systems.” IEEE Transactions on Automatic Control 62 (8): 3861–76.

This filter is the CBF-QP of 1.1 (Ames et al. 2016). At each discrete control interval \(T_{\text{ctrl}} = 1.0\text{ ms}\) on the bare-metal microcontroller, the tracker computes the nominal torque command \(\mathbf{u}_{\text{nom}} \in \mathbb{R}^m\) each tick from the admitted reference (What Planning Hands Over), and the enforcer solves the convex problem10 \[\min_{\mathbf{u}} \frac{1}{2}\|\mathbf{u} - \mathbf{u}_{\text{nom}}\|^2 \quad \text{s.t.} \quad \mathbf{a}(\mathbf{x})^\top \mathbf{u} \le b(\mathbf{x}), \quad \mathbf{u}_{\min} \le \mathbf{u} \le \mathbf{u}_{\max} \tag{5}\] where \(\mathbf{a}(\mathbf{x})^\top \mathbf{u} \le b(\mathbf{x})\) represents a torque constraint derived from a velocity-dependent barrier such as equation 4,11 and \([\mathbf{u}_{\min}, \mathbf{u}_{\max}]\) represents the verified actuator torque envelope.

When enforcing a single active linear barrier constraint \(\mathbf{a}^\top \mathbf{u} \le b\) in equation 5, the optimization admits an exact closed-form orthogonal projection, derived in Minimal-intervention quadratic program: \[\mathbf{u}^* = \mathbf{u}_{\text{nom}} - \frac{\mathbf{a}^\top \mathbf{u}_{\text{nom}} - b}{\|\mathbf{a}\|^2} \mathbf{a} \tag{6}\] where the optimal Lagrange multiplier \(\lambda^* = (\mathbf{a}^\top \mathbf{u}_{\text{nom}} - b) / \|\mathbf{a}\|^2\) scales the minimal intervention vector \(\Delta \mathbf{u} = -\lambda^* \mathbf{a}\).

The projection arithmetic is clearest on a two-axis manipulator with illustrative values. Its tracker, following setpoints from an unconstrained neural policy, computes nominal joint torques \(\mathbf{u}_{\text{nom}} = [\) 3.5, 3 \(]^\top\text{ N}\cdot\text{m}\). Suppose an abstract safety half-space has already been derived and validated, \(\mathbf{a}^\top \mathbf{u} \le b\) with \(\mathbf{a} = [\) 2, 1 \(]^\top\) and \(b =\) 5 \(\text{N}\cdot\text{m}\), with per-joint actuator limits of $$5 \(\text{N}\cdot\text{m}\). The candidate gives \(\mathbf{a}^\top \mathbf{u}_{\text{nom}} =\) 10 \(\text{N}\cdot\text{m}\) and breaches the boundary by 5 \(\text{N}\cdot\text{m}\).

By equation 6, \(\lambda^* =\) 1, and the minimal-intervention command is \(\mathbf{u}^* = [\) 1.5, 2 \(]^\top\text{ N}\cdot\text{m}\), which meets the half-space exactly while both joints stay within their limits, with an intervention norm of 2.236 \(\text{N}\cdot\text{m}\).

Per-axis clamping shows why the projection is needed. Clamping Joint 1 to 2 \(\text{N}\cdot\text{m}\) while passing Joint 2 at 3 \(\text{N}\cdot\text{m}\) gives \(\mathbf{a}^\top \mathbf{u}_{\text{clip}} =\) 7 \(\text{N}\cdot\text{m}\), still 2 \(\text{N}\cdot\text{m}\) outside the constraint, and it skews the torque vector from 40.6\(^\circ\) to 56.3\(^\circ\). The clamp therefore both violates the constraint and distorts the command. The projection satisfies the constraint with the smallest possible torque change, although it too rotates the torque vector, so the end-effector path still needs a check against the coupled plant model.

Figure 4 shows both views. In state space (panel a), proposals that would cross the barrier are deflected along the inset boundary \(\partial \mathcal{C}\); in control space (panel b), the enforcer projects \(\mathbf{u}_{\text{nom}}\) onto the half-space \(\mathbf{a}^\top \mathbf{u} \le b\), while the independently clamped vector stays outside it.

Two panels: a schematic inset stopping boundary in state space, and a two-axis torque plane in which a nominal command is projected onto a feasible half-space while a per-axis clamped command remains outside it.
Figure 4: Kinetic barrier and torque projection: Panel (a) schematically shows a velocity-dependent viable boundary inset from geometric clearance; the raw distance alone has relative degree two for torque. Panel (b) illustrates Euclidean projection onto a supplied torque half-space. The \([2,1]\) normal and \(5\text{ N}\cdot\text{m}\) threshold are abstract arithmetic inputs, not coefficients derived from the arm example.

Filter choice joins a mathematical condition to a measured execution budget and the machine’s available braking authority. As summarized in table 2, architectural choices differ in their handling of relative degree, solver worst-case execution time (WCET), forward invariance guarantees, and directional fidelity. Independent clamping can miss coupled constraints; a CBF-QP can enforce a derived half-space if it is feasible and finishes within the verified cycle. Predictive filters (Wabersich and Zeilinger 2021) trade a longer horizon for configuration-dependent compute cost.

Wabersich, Kim P., and Melanie N. Zeilinger. 2021. “A Predictive Safety Filter for Learning-Based Control of Constrained Nonlinear Dynamical Systems.” Automatica 129: 109647.
Table 2: Safety-filter decisions: A filter’s physical claim depends on the chosen barrier and validated plant, not on the algorithm name. Nominal solve measurements are not worst-case deadlines.
Filter Input and condition What it can establish Main validation task
Independent torque clamp Each axis checked separately Hardware amplitude limit, not coupled stopping clearance Test coupled wrench and stopping behavior
First-order CBF-QP Barrier whose first derivative contains the input Conditional invariance from a viable initial state Validate model, feasible torque set, and cycle bound
Kinetic / higher-order barrier Position constraint with speed and braking authority included Conditional stopping margin before geometric contact Bound bias torque, delay, estimate error, and discretization
Precomputed viable region State-to-permitted-input lookup Containment for the modeled region Verify lookup coverage and interpolation
Predictive safety filter Finite-horizon trajectory and terminal set Recursive feasibility only under terminal and timing assumptions Bound full solve time and fallback on timeout

Operating near an active safe-set boundary can produce chattering. If a learned policy repeatedly proposes an unsafe command, sampled enforcement may switch the active constraint across successive 1 ms ticks. Depending on the plant and controller, rapid torque changes can excite a structural mode or increase motor heating and transmission wear. A tested boundary layer or hysteresis can reduce switching, but its added tracking error must fit inside the validated stopping margin.

The intervention vector \(\Delta u = u^* - u_{\text{nom}}\) also measures upstream policy health. A policy operating inside its training distribution produces zero intervention on most cycles, so a rising intervention frequency or magnitude can indicate proposal drift, a changed task, or a changed plant. Summed over an illustrative rolling window of \(W = 500\text{ ms}\) as \(\sum_{k} \|u_k^* - u_{\text{nom}, k}\| \Delta t\), the metric lets the permission path flag divergence for logging or operator notification while the per-tick limits still protect the plant.

In this heterogeneous design, the shield executes inside a 1 kHz timer interrupt routine on the MCU. An unloaded cycle of acquisition, validation, and a typical active-set solve completes in about 135 \(\mu\text{s}\), an illustrative estimate rather than a bound. The qualified path instead declares a worst-case execution time of 250 \(\mu\text{s}\) for the whole cycle (Real-Time Timers, Watchdogs, and Schedulability), divided among the stages below, and the whole sequence must finish inside the 400 \(\mu\text{s}\) deadline:

  1. Interrupt Entry & State Acquisition (30 \(\mu\text{s}\)): The hardware timer triggers the ISR, and firmware latches the wheel and joint encoder registers and the latest IMU sample.
  2. Proposal & Lease Verification (15 \(\mu\text{s}\)): The MCU inspects the shared SRAM ring buffer populated by the application processor over RPMsg or shared memory. It validates the checksum, confirms sequence order, reads the proposal’s evidence epoch, and makes two age checks against it and the lease. The upper age of the evidence the proposal was built on, stamped at dispatch, may not exceed the 51.8 ms refusal threshold, which is the 51.6 ms observation age the stopping budget charged plus the 0.2 ms clock-conversion bound (The Observation Contract). The proposal’s own age on the MCU clock may not exceed the 60 ms chunk lease. The MCU also reads the host heartbeat timer, which marks the host silent once no heartbeat has arrived within the 30 ms timeout.
  3. Tracking Contract Verification (10 \(\mu\text{s}\)): Firmware evaluates instantaneous tracking error \(e(t) = \|x(t) - r(t)\|\) against the certified bound \(\epsilon_{\text{track}}\).
  4. CBF-QP Safety Shield Solve (150 \(\mu\text{s}\)): An active-set QP solver running in pre-allocated static DTCM SRAM (Where the Permission Path Runs) evaluates active barrier inequalities and computes \(u^* = \arg\min_u \frac{1}{2}\|u - u_{\text{nom}}\|^2\). A fixed iteration cap holds the solve inside its share even for ill-conditioned constraint geometry; nominal solves for one or two active constraints finish well within it. Dynamic memory allocation is forbidden, because a heap search has no bounded duration.
  5. Output Selection & Setpoint Staging (15 \(\mu\text{s}\)): If the QP is feasible and constraints hold, \(u^*\) is staged as the admitted setpoints for the next EtherCAT frame to the drives, whose current loops apply it and whose torque accuracy is tested rather than inferred from their rate; a proposal that fails validation leaves the previously admitted setpoints in force. If the chunk lease has expired, the host heartbeat has fallen silent, the shared control rail reports undervoltage, tracking error exceeds \(\epsilon_{\text{track}}\), actuator saturation renders the feasible control set empty (\(\mathcal{U}_{\text{safe}} = \emptyset\)), or an operating-contract detector such as the disturbance observer or the solver deadline trips, the output selector revokes the policy and selects the state-matched stop. At standstill, the same lease, heartbeat, or undervoltage trigger selects a powered hold instead.
  6. Completion & Reserve (30 \(\mu\text{s}\)): The MCU checks the register write and services its own watchdog only after the cycle succeeds. The reserve absorbs variation in the stages above; a cycle that has not completed by the deadline transfers authority to the resident fallback before the next tick. A missed tick therefore costs the base 1.3 mm of travel at the aisle speed. Testing the Separation tests whether the declared WCET holds under load, so that ticks are not missed.

Every stage of this cycle assumes that the solve returns a feasible command. When saturation empties the admissible torque set, the lease lapses, or the cycle overruns its deadline, projection has nothing to return, and the permission path must already hold a response it can execute without the proposer.

The Fallback Ladder

The fallback must be selected before the admissible set becomes empty. Once it is empty, a fallback can contain a fault or limit harm but cannot promise to avoid a boundary crossing. Nor can the response be a single switch. Cutting bus power whenever a proposal brushes a barrier margin produces uncoordinated deceleration and continuous mission aborts, while relying on software optimization during an inverter shoot-through (Actuator Transmission Limits) burns the gate drivers before an interrupt handler can run. The permission path therefore holds several resident responses and selects among them deterministically, from feasible projection to plant-specific controlled stopping or torque removal.

Much as the arbitration of behavior-based robotics and the subsumption architecture (Brooks 1986) lets a condition select which layer drives the actuators, the fallback ladder is a set of four deterministic responses (project: minimal-intervention QP projection; hold: active position hold; stop: controlled dynamic stop; inhibit: drive torque inhibit), each running on the execution substrate that survives its trigger, from which the permission path selects by the check that failed and the state of the plant when it failed, rather than climbing them in order.

Brooks, Rodney. 1986. “A Robust Layered Control System for a Mobile Robot.” IEEE Journal on Robotics and Automation 2 (1): 14–23.

The first rung, least-squares projection, runs on every cycle while the feasible control set is nonempty (\(\mathcal{U}_{\text{safe}} \neq \emptyset\)), computing \(u^* = \arg\min_u \|u - u_{\text{nom}}\|^2\) subject to the active barrier inequalities within its 150 \(\mu\text{s}\) share of the cycle. It succeeds only while the barrier constraints and the achievable torques intersect, so the admission path must preserve viability before saturation empties \(\mathcal{U}_{\text{safe}}\). If infeasibility occurs anyway, the MCU revokes the proposal and invokes the best available state-matched response.

Where the plant can be held rather than stopped, the response is a controlled deceleration into an energized position hold, analogous to a Category 2 stop.12 The permission path arrests motion and holds the current coordinate \(x_{\text{hold}}\) under closed-loop feedback, with the drives still modulating current against gravity and payload. The rung spends electrical energy and accumulates thermal debt in the windings (Thermal Duty Cycles), but it wears no brake and preserves the kinematic state, so after a fresh proposal and state check the enforcer may resume without homing. On the warehouse mobile manipulator this rung answers a lapsed lease or a silent heartbeat only at standstill. The same expiry in motion selects the stop rung, because the resident stop is the only deceleration whose time and distance the stopping budget has admitted. If tracking exceeds its bound, the MCU selects a validated stop, and holding remains appropriate only while available torque and feedback support it.

When a body in motion loses its proposal’s authority or a tracking or model assumption fails, the response is a controlled dynamic stop, analogous to a Category 1 stop under IEC 60204-1. On the warehouse mobile manipulator’s base, this rung executes the resident stop suffix admitted with the trajectory (Behavior at the Seam) rather than composing a stop at the moment of the fault. From the 1.3 m/s aisle speed, the \(C^2\) suffix peaks at the 2 m/s² credible deceleration, lasts 975 ms, and covers 633.75 mm, the braking term already reserved in the stopping budget of section 1.5. The kinetic energy the stop removes regenerates through the inverter into the DC bus (Electrical Power Integrity), so bus capacitance and a brake chopper must absorb the surge without an overvoltage fault.13 Power removal, if specified, follows the measured standstill condition.

A separate last-resort drive function can inhibit motor-generated torque. Drive Safe Torque Off (STO) does not itself disconnect the DC bus, apply a mechanical brake, or stop a moving motor; the motor may coast. In the illustrated architecture, an external safety controller may also command a separately rated contactor and spring-applied brake. Their response and stopping ability must be measured as one plant-specific path, because removing torque while a gravity-loaded arm or moving vehicle still depends on powered braking can worsen the hazard. On the warehouse mobile manipulator that path, STO plus the nine spring brakes setting 30 ms to 80 ms after coil decay, is the inhibit rung, and it is unbudgeted. It answers only the loss of the MCU’s own cycle, and no stopping or release claim rests on it.

Loss of supply does not route there. The permission path runs on an isolated rail held up for 2 s (Which Budget Binds First), so a sag on the shared control rail arrives as an undervoltage flag, and the enforcer answers it with the stop rung. The hold-up outlasts the longest stop the enforcer can command inside the envelope. Stop onset follows the flag within 22 ms, the resident stop from the 1.09 m/s ceiling on a floor inspected to a friction of 0.12 takes 1395.4 ms, and the brakes set at standstill within 80 ms, so every brake is set 1497.35 ms after the flag, with 502.65 ms of hold-up to spare.

↰ Prerequisite: Hardware trip zones and low-side MOSFET dynamic shunt braking are introduced in Actuation Authority.

Physical AI systems engineers must distinguish between fail-safe and fail-operational plants when configuring terminal fallbacks. In a machine already at a mechanically restrained standstill, removing drive torque may be safe after the restraint has been verified. In underactuated, open-loop unstable, or high-momentum machines, however, suddenly de-energizing actuators is itself a hazard:

  1. Aerial Platforms (Multirotors and eVTOLs): Cutting motor torque drops aerodynamic thrust to zero, transforming an airborne robot into a ballistic projectile in uncontrolled free fall.
  2. Dynamic Legged Systems (Bipedal Humanoids): During dynamic locomotion at \(1.5\text{ m/s}\), the center of mass is outside the base of support. Abruptly removing the torque needed for balance can cause a fall; the direction and consequence depend on the gait phase and available restraint.
  3. High-Speed Autonomous Vehicles: Removing powered steering or braking at highway speed can eliminate the authority needed for a controlled stop; the outcome depends on vehicle dynamics and backup actuation.

For such dynamic systems, the terminal fallback cannot simply be an uncoordinated power cut. Instead, the permission path must execute a deterministic Minimal Risk Maneuver (MRM) running on an independent, qualified controller with the required sensing and actuator authority. For aerial and legged robots, any controlled descent or balance recovery must be validated for remaining propulsion, contact, energy, and state-estimate authority. A road vehicle needs a risk-assessed powered braking and steering path. The wiring that carries a torque-removal request is the out-of-band interlock of Actuation Authority (figure 5), which keeps the spring-applied brake on its own path because removing torque and arresting a moving load are different functions carrying different validation evidence.

Illustrative circuit with two emergency-stop contacts and safety relays feeding a drive torque-inhibit path. A separate spring brake is shown as an optional function requiring its own timing and load test; any supply contactor is separate; STO alone does not brake a moving motor.
Figure 5: Illustrative emergency-stop circuit: Two emergency-stop contacts feed safety relays and drive Safe Torque Off (STO) inputs. A spring-applied brake is shown as a separate path; an external supply contactor, if fitted, provides galvanic disconnection. The diagram illustrates physical isolation; a moving motor can coast after STO, so physical stopping distance must be independently validated for the load and state.

The architectural necessity of independent, hardware-level trip circuits is demonstrated by the historical consequences of eliminating physical interlocks in favor of software decision logic (1.1).

War Story 1.1: Therac-25 software interlock elimination
Context: The AECL Therac-25 was a dual-mode (\(25\text{ MeV}\) electron / X-ray) medical linear accelerator and automated radiotherapy patient table, controlled by a single PDP-11 minicomputer running uniprocessor assembly firmware (Leveson and Turner 1993).

Mechanism: In preceding models (Therac-6 and Therac-20), physical safety was guaranteed by independent hardware interlocks: electromechanical microswitches, mechanical turntable detents, and analog diode steering circuits that physically prevented the accelerator from firing high-current beams unless the tungsten beam-flattening target was mechanically locked in place. In the Therac-25, engineers removed all physical hardware interlocks to reduce manufacturing costs, delegating exclusive safety enforcement to software. Concurrently, two latent software bugs defeated the software checks: an 8-bit shared variable (Class3) rolled over from 255 to 0 (\(0\text{xFF} \to 0\text{x00}\)) during setup, bypassing interlock evaluation; and a race condition allowed fast terminal keystrokes (< 8 seconds) to update operating modes without reconciling the physical turntable position with the commanded high-current beam energy.

Impact: Between 1985 and 1987, six patients received massive radiation overdoses up to \(25{,}000\text{ rad}\) (more than 100 times the intended therapeutic dose), resulting in catastrophic radiation burns, severe physical trauma, and several patient fatalities.

Response: Regulatory agencies halted machine operations, prompting comprehensive investigations that led to modern medical device software engineering standards, mandatory independent hardware safety interlocks, and formal verification of safety-critical concurrent states.

Systems lesson: Software-only safety filters executing on complex, non-deterministic compute stacks cannot replace independent bare-metal hardware trip zones. The Therac-25 disaster is the canonical historical proof that removing physical interlocks in favor of software decision logic transforms transient concurrency bugs and arithmetic overflows directly into catastrophic physical energy releases. In physical AI, software Control Barrier Functions and runtime projection shields are advisory safety filters; the terminal inhibit rung must always terminate in device-qualified hardware interlocks such as Safe Torque Off and mechanical brakes.

The rungs trade intervention time, available control authority, and recovery cost. A controlled stop requires reserved clearance and a working drive, and each rung’s plant-specific response must be validated for the states from which it can be entered. Figure 6 arranges the four rungs by the fault condition that selects each one.

Definition 1.2: Dynamic fallback ladder

Dynamic fallback ladder is the set of resident protective responses (\(u \in \{\mathcal{P}_{\mathcal{C}}(u_{\text{nom}}), u_{\text{hold}}, u_{\text{brake}}, u_{\text{STO}}\}\)) from which a deterministic arbiter selects one on every tick according to the check that failed (an invariant violation, a timing overrun, or a hardware fault) and the measured state of the plant, so that one trigger can select different rungs in motion and at standstill.

  1. Significance: Physical AI systems cannot treat all faults identically; cutting actuator power on dynamic, underactuated, or high-momentum plants (e.g., legged robots at speed, multirotors in flight) can cause a fall or tumble, requiring state-matched deceleration before terminal torque removal.
  2. Distinction: Unlike software exception handlers that unwind the call stack or kill a process thread, a dynamic fallback ladder operates against physical inertia, matching each rung to the remaining valid sensing, control authority, and reserved stopping clearance.
  3. Common pitfall: Reading the rungs as an escalation in which a lower one, such as Safe Torque Off (STO), provides faster or safer arrest. Each rung has its own entry conditions, and the inhibit rung removes torque without arresting anything by itself.
Policy proposals reach a local arbiter with four possible outputs: a feasible torque projection, powered position hold, controlled braking stop, or drive torque inhibit. Each output requires a state-specific clearance and hardware response check; torque inhibit alone may allow coasting.
Figure 6: Available enforcement responses: A deterministic arbiter can select feasible torque projection, controlled hold, controlled stop, or drive torque inhibit according to the current fault and measured plant state. The responses are alternatives with different prerequisites; descending the diagram does not imply increasing physical safety or faster arrest.

War Story 1.2: Cruise post-impact fallback (2023)
Context: In October 2023, a human-driven vehicle struck a pedestrian into the path of a Cruise robotaxi in San Francisco (Exponent, Inc. 2023).

Mechanism: The Cruise vehicle initially braked and stopped. Its system then misclassified the event as a side impact and, having tracked the pedestrian only intermittently in the seconds before impact, initiated a pullover about \(1.83\text{ s}\) after contact. The vehicle dragged the pedestrian roughly \(20\text{ ft}\) at a maximum reported speed of \(7.7\text{ mph}\).

Response: An independent impact and underbody-occupancy response should revoke ordinary policy authority, retain braking and steering needed for a controlled stop, and hold the vehicle once stationary.

Systems lesson: A post-contact maneuver cannot infer clear space from a missing tracked object. The fallback must use independent evidence and preserve the ability to brake and hold.

Exponent, Inc. 2023. Cruise AV SF Incident—Pedestrian Collision: Technical Root Cause Analysis. Exponent Project 2310645.000. Exponent, Inc.

The IEC 60204-1 stop categories name drive functions, not plant responses, and table 3 separates the function each category names from the physical response that must still be tested.

Table 3: Stop functions and physical response: Each stop category mapped to the drive function it names and the plant response it leaves to be tested. No stop category alone supplies a safety integrity rating, a brake, galvanic isolation, or a guaranteed stopping distance.
Response Motor torque and power Physical stop obligation Appropriate use
Admitted filtering Drive remains active Recheck the feasible kinetic barrier and tracking bound every tick Nominal operation inside the validated envelope
Controlled hold (Category 2 / SS2-like) Power remains for braking and hold Verify stopping distance, thermal headroom, and static load support Hold when feedback and torque authority remain available
Controlled stop then torque removal (Category 1 / SS1-like) Power remains through deceleration; torque is removed afterward Verify the full braking and transition time under load Moving machine with enough reserved clearance
Torque removal (Category 0 / STO-like) Drive cannot produce commanded torque; DC bus isolation is a separate function A moving motor may coast; a separate brake or mechanical restraint needs its own timing and rating State for which loss of torque has been shown safe
Checkpoint 1.2: Choosing the physical fallback

Before selecting a defensive fallback strategy, verify your understanding of state-dependent stopping modes, safety functions, and timing reserves:

How Enforcement Goes Wrong

Every rung of the fallback ladder assumes that the hardware executing it stays independent of the fault that selected it. A secondary processor beside the primary compute module looks independent, yet a shared regulator, crystal, or ground plane can reset, halt, or corrupt both chips in the same instant, as when a host-side short drags the rail below the microcontroller’s threshold or motor-driver return currents bounce the ground beneath its buses. Software on either chip cannot see these faults, which is why the permission path keeps its own supply, clock, and ground (Hardware Isolation).

An isolated enforcer still fails when its computation falls out of phase with the plant, because a solve that overruns its tick acts on a state the machine has already left, and in an underdamped assembly a braking torque computed for that earlier state can feed a structural resonance instead of damping it. A watchdog that aborts the late cycle cannot recall the motion accrued during the overrun, so the timing claim rests on the declared WCET and on the loaded measurements of Testing the Separation.

The enforcer sees the plant only through state estimates, so filter group delay, transmission compliance, encoder slip, and thermal bias drift can hold the reported state inside its limits while the physical link has crossed them. Cross-checks against independent encoders or torque sensors catch large discrepancies, but slow sub-threshold drift stays hidden until contact.

Determinism guarantees that a computation finishes in a fixed number of cycles with bit-identical outputs, but it says nothing about whether the equations match the plant. An enforcer that assumes the base’s friction premise of 0.204 on a floor that oil has made slicker plays the resident stop exactly on schedule while the tires cannot deliver its deceleration, so the base runs past the rack end with every deadline met. No timing monitor can see that failure, and A Residual-Claims Register keeps its friction premise open.

Contention on the shared system-on-chip (SoC) interconnect can defeat an enforcer that meets its own deadline, because a camera direct memory access (DMA) burst or bulk flash write that holds the bus matrix after the solve adds directly to the command’s age at the drive while the inverter holds its last setpoint. The machine’s placement keeps the permission path off that bus (Contention for Shared Resources), so a burst on the application processor reaches it only as evidence age, and cycle stage 2 refuses any proposal the burst has aged past the refusal threshold.

An enforcer hosted as a real-time operating system task rather than a bare-metal interrupt adds priority inversion at the operating-system layer, where a low-priority telemetry thread holds a mutex the 1 kHz enforcer needs and a medium-priority thread keeps the holder from running, so the check waits without bound. The machine’s timer interrupt and lock-free mailbox exclude that case by design. On a task-hosted enforcer, priority inheritance (PTHREAD_PRIO_INHERIT) or ceiling locking bounds the wait, and Moving Commands on Time examines the failure in flight hardware.

The Therac-25 radiotherapy accelerator shows a concurrency failure with nothing independent behind it. Where its predecessors had independent protective circuits and mechanical interlocks, the Therac-25 relied chiefly on software running on the same computer that ran the treatment, where a race in operator data entry and a one-byte counter that rolled over to zero each let the machine fire a high-current beam without its target in place (Leveson and Turner 1993). Between 1985 and 1987 six patients received massive overdoses, and several died.

Leveson, Nancy G., and Clark S. Turner. 1993. “An Investigation of the Therac-25 Accidents.” Computer 26 (7): 18–41.

Every failure mode in the enforcement layer traces back to a breakdown of independence or a divergence between the internal plant model and physical reality. A barrier certificate, a dual-core package, and a zero-jitter scheduler establish safety only within the envelope where supplies stay isolated, estimates match the mechanics, and deadlines are met, so an enforcer that is to be audited against these failures has to state that envelope in advance: what it assumed, what it checked on each tick, and which response each failed check selects.

The Enforcement Record

When the warehouse mobile manipulator’s base refuses a proposal at the rack end, a later reader of the log must be able to say which threshold failed, which detector measured it, and which rung it selected. The enforcement record is the specification contract and runtime audit ledger under which, on every tick, the drive receives either the projection of \(u_{\text{nom}}\) onto the safe set or a fallback-ladder response, and nothing else.

The record consumes four predecessors, drawn from the two strands of records that meet here. The offline strand records what the plant can do and what the proposer has shown it can do, and the runtime strand records what the proposer proposes now. The trajectory record of The Trajectory Contract supplies the admitted reference and its matched stop suffix. The evaluation record of Evaluation Logs supplies the runtime monitor specifications and the coverage-gap catalog, which become the record’s monitor channels. The policy manifest of The Policy Manifest supplies the declared ODD, and the validity region stays inside the sampled part of it. The limit records of Measuring a Machine's Own Limits supply the braking, onset, and actuator limits with their conditions. To these the enforcement record adds what the permission path guarantees: the verified tracking bound \(\epsilon_{\text{track}}\), the stopping time and distance (\(t_{\text{stop}}, d_{\text{stop}}\), bounding the duration and displacement required to arrest plant momentum), Control Barrier Function coefficients, the loop period, per-check thresholds with their fallback-ladder rungs, and the validity region.

↳ Downstream: High-rate enforcer refusal events populate the forensic evidence logs analyzed in The Authority Log.

Filled in for the machine, the record fixes the base’s stops at the rack end and the arm’s validity region and channels at the door and the packing station:

  • The tracking bound \(\epsilon_{\text{track}}\) is 40 mm, the inset of row 7 in table 1.
  • The resident stop is the \(C^2\) suffix, which lasts 975 ms and covers 633.75 mm from the 1.3 m/s aisle speed, or 1125 ms and 843.75 mm from the 1.5 m/s drive limit.
  • The loop period is 1 ms, with a declared WCET of 250 \(\mu\text{s}\) inside a 400 \(\mu\text{s}\) deadline. Three timers run beside it, the 60 ms chunk lease, the 100 \(\text{Hz}\) host heartbeat with its 30 ms silence timeout, and the 5 ms self-watchdog.
  • The validity region for the arm covers the door and the packing station. At the door it is the target ODD of Evaluation Logs, which admits only the nominal plate, the door’s own load, a stopped base, and the guarded approach. At the packing station it admits the mug pick from the conveyor under a live intent lease (The Intent Lease), and the handover that follows.
  • The arm’s channels are ceilings the enforcer applies to the tool point. The approach-speed channel of the evaluation record’s runtime_monitor_specs caps it at 0.03 m/s inside the door zone. The mug pick runs under the 1 m/s free-space TCP limit. While the coworker holds accept_item (Human Authority Over Actuation), a handover channel caps the approach to the coworker’s hand at 0.10 m/s, below the 0.118 m/s contact-force ceiling, and it engages before the hand comes within the arm’s biased stop plus its blind travel.
  • For the base, the normal configuration is valid on floors with friction of at least 0.204, where the credible deceleration binds. A restricted configuration is valid down to the inspected-floor friction of 0.12; its suffix is sized for the 1.177 m/s² peak that floor allows, and it takes 1274.6 ms and 637.3 mm to stop from 1 m/s, under the 1.09 m/s inspected-floor ceiling. Each configuration carries its own digest, and the release manifest (The Release Manifest) names both.

Each numerical entry in the record carries the evidence category of its limit record (Measuring a Machine's Own Limits), whether measured, derived, datasheet, or illustrative, and the categories carry very different real-world confidence. The base’s 975 ms stopping time is derived. It follows from the shape of the \(C^2\) suffix and from a credible deceleration that is itself illustrative, and it assumes that the drive delivers the commanded current and that the tires hold the floor throughout the stop. Only loaded stopping trials on the floor itself make it a measured value, and those trials expose what no derivation contains, among them the scatter in brake onset and the friction of the actual floor surface. Even then the measured sample stands for the bound only with a justified worst-case fault and stated operating conditions. Recording the category prevents an integrator from mistaking a derived stop for a measured hardware limit.

The record maps separately qualified thresholds to protective actions, and on the machine the routing depends on whether the body is moving. Routing does not rely on coarse error strings or unquantified health flags. The base stays on the project rung, where the CBF-QP may adjust commands, as long as tracking error satisfies \(e(t) \le \epsilon_{\text{track}} =\) 40 mm, the kinetic barrier stays feasible, and the solver stays within its declared 150 \(\mu\text{s}\) share of the cycle budget. A new proposal that fails validation while the active lease holds is refused, and the previously admitted setpoints stay in force until that lease lapses. In motion, a lapsed lease, a silent heartbeat, an undervoltage flag on the shared control rail, tracking error beyond \(\epsilon_{\text{track}}\), saturation that empties the feasible control set, or a tripped operating-contract detector (the disturbance observer or the solver deadline) selects the stop rung, which plays the resident stop on the base and the state-matched stop on the arm. At standstill the same lease or heartbeat expiry, or the same undervoltage flag, selects the hold rung, since no motion remains to arrest. The 5 ms MCU self-watchdog detects missed successful cycles; this design also assumes an independently wired supervisor output or external trip that requests the inhibit rung if the MCU fails. Its total trip response must be measured. A contactor, if fitted, is an independently timed supply-isolation function.

Each field of the record is either an integrity field or a threshold that one of these checks reads: the proposal’s sequence number and checksum, its evidence age against the 51.8 ms refusal threshold (the budgeted observation age plus the clock-conversion bound), its own age against the lease, the tracking bound, the joint-velocity and torque-slew ceilings, the barrier margins, the phase-current and bus-voltage limits, and the watchdog timeout. The record binds each threshold to the independent hardware that measures it and to the rung that a failed check selects. Table 4 groups those rungs by the authority path that remains active after a failed check, together with the timer that triggers each one.

↳ Byte layout: The enforcement record’s fields, types, and units are given in Enforcement record.

Table 4: Enforcement state contract: The 20 \(\text{Hz}\) proposal, 60 ms chunk lease, 100 \(\text{Hz}\) host heartbeat, and 1 kHz MCU loop have distinct renewal and failure paths.
State Trigger and authority Timing check Physical response to validate
PROJECT Fresh proposal, healthy heartbeat, feasible kinetic barrier; a proposal that fails validation under an unexpired lease is refused and the admitted setpoints stay in force \(1\text{ ms}\) cycle; declared full-cycle WCET 250 \(\mu\text{s}\) inside a 400 \(\mu\text{s}\) deadline Track admitted torque within qualified error and brake reserve
HOLD At standstill, proposal age above the 60 ms lease, host heartbeat silent beyond 30 ms, or shared-rail undervoltage Expiry detected on the MCU clock within one tick Powered position hold against gravity and payload; resume only on a fresh admitted proposal
STOP In motion, lease or heartbeat expiry, shared-rail undervoltage, tracking error beyond \(\epsilon_{\text{track}}\), an empty feasible control set, or a tripped disturbance observer or solver deadline, while powered braking remains available Include detection, dispatch, drive rise, and brake response Resident stop on the base, state-matched stop on the arm; measured stop distance under load; no claim after viability is lost
INHIBIT Independent wired supervisor or external trip detects MCU cycle loss Assumed 5 ms detection plus measured drive, contactor, or brake response STO removes motor torque and the spring brakes set after coil decay; unbudgeted, so no stopping claim rests on it

No enforcement guarantee holds when the machine operates outside its designed physical environment, so the record bounds the plant dynamics within an explicit validity region. The arm’s door ceiling rests on evidence. The arithmetic alone would have admitted the chunk policy’s 0.10 m/s approach of section 1.3, whose 2.5 ms contact deadline still exceeds the 2 ms contact loop, but neither the demonstrations nor the evaluation trials behind the policy include that speed (Evaluation Logs), so the enforcer projects the proposal to 0.03 m/s and logs the channel that bound it. The handover ceiling rests on the contact model instead, sitting below the contact-force ceiling, the speed at which an unplanned touch of the coworker’s hand would reach the contact criterion (Human Authority Over Actuation), and a faster proposal is projected and logged in the same way.

A live state outside the validity region that the machine’s estimators detect, such as a payload the manifest never declared, is refused like any other failed check, and the supervisory monitors that flag it revoke policy dispatch and select the state-matched stop. A floor below the base’s friction premise is not such a state, because no sensor on the machine observes the friction under its wheels. The normal configuration therefore rests on that friction as a premise, and the restricted configuration rests on inspection of the floor (A Residual-Claims Register).

An enforcement record connects proposed motion to validated assumptions, thresholds, and observed fallback response. The execution platform must still meet its timing and physical braking contract under load. An enforcer placed on a shared memory bus can be mathematically verified and still stall when that bus saturates during a high-speed transit maneuver, and then the failure occurs in time rather than in logic.

Fallacies and Pitfalls

Physical enforcement stops a machine before a collision only while its execution is timely and its state feedback usable. Fault tests therefore exercise the response when either assumption fails, not the filter’s nominal equations alone.

Pitfall: Assuming high-level policy computers can be relied upon to execute graceful shutdown.

The warehouse mobile manipulator’s application processor suffers a kernel panic while the base runs at 1.3 m/s. The failed host cannot generate a deceleration trajectory, and dropping the drive enable to zero torque would leave the base coasting toward the rack end. The MCU instead detects the silent heartbeat within its 30 ms timeout, well inside the lease path the budget reserves, and plays the resident \(C^2\) suffix, which stops the base in 975 ms over 633.75 mm. A host shutdown handler may coordinate an orderly stop in normal operation, but protection against host failure requires a stopping path, clearance, and torque authority that remain after the host process disappears.

Fallacy: A mathematically correct safety filter guarantees safety even when underlying state feedback has drifted.

The base’s tracking bound of 40 mm limits the gap between the reference and the state its wheel encoders report, not the gap between that state and the floor. Wheel slip or a biased encoder can make the base appear farther from the rack end than it is, and the CBF-QP solver then satisfies its inequality on the drifted estimate while the physical base spends clearance the budget never granted. Only an independent channel, such as the IMU or the safety scanner’s range to the rack, exposes the divergence, and only while the fault is observable through it. Qualification must bound how far the estimate can drift before detection, and the budget must carry that error in the localization allowance, because the tracking inset cannot absorb it.

Pitfall: Coupling safety braking authority to stochastic semantic classification rather than deterministic spatial clearance.

When braking depends on upstream classification and path prediction, a detected object may still fail to trigger an effective stop. In the March 2018 Tempe collision (figure 7), the developmental automated driving system first detected the pedestrian \(5.6\text{ s}\) before impact but did not predict her path. It recognized an imminent collision \(1.2\text{ s}\) before impact, suppressed action for one second, then planned gradual slowing rather than emergency braking; the NTSB does not establish a Kalman-filter reset on each classification change. The vehicle’s original Volvo collision-warning and emergency-braking functions were disabled during ADS operation (When the Boundary Fails). An independent geometric clearance check needs its own validated sensing, stopping authority, and timing. It refuses on distance and closing speed, whatever label the object carries.

Timeline from first ADS detection 5.6 seconds before impact to imminent-collision recognition 1.2 seconds before impact, one-second action suppression, gradual-slowdown plan, and impact. Separate note says the Volvo collision-warning and emergency-braking functions were disabled during ADS operation; no object-history reset is asserted.
Figure 7: Tempe detection and braking timeline: The NTSB reports first detection \(5.6\text{ s}\) before impact and recognition of imminent collision \(1.2\text{ s}\) before impact. The ADS suppressed its motion plan for one second and then planned gradual slowing; it did not perform emergency braking. Classification and path prediction failed. The Volvo forward-collision warning and automated emergency braking functions were disabled while the developmental ADS operated. The report does not establish a Kalman-filter reset at each label change. (Adapted from NTSB/HAR-19/03 (National Transportation Safety Board 2019).)
National Transportation Safety Board. 2019. Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian, Tempe, Arizona, March 18, 2018. HAR-19/03. National Transportation Safety Board.

Summary

A physical AI system assigns actuator permission to a path independent of learned proposals. A kinetic barrier can help preserve a validated viable set when state error, model mismatch, actuation, and response delay stay inside their declared bounds. If braking authority or clearance is lost, the enforcer can revoke the proposal but cannot claim that a later fallback restores the original safe set.

The physical clearance protecting a machine is never a static geometric measurement. It is a composite margin that compounds pre-brake delay, braking under the credible deceleration, the localization and protective allowances, and the tracking bound. The warehouse mobile manipulator at its 1.3 m/s aisle speed shows the sum. Blind travel over \(\tau_{\text{delay}}\) takes 173.68 mm, the \(C^2\) stop suffix 633.75 mm, the localization bound and protective clearance 150 mm, and the tracking inset 40 mm. Together they need 997.43 mm of the 1.10 m clear distance at the rack end and leave 102.57 mm. A budget that counted only constant-deceleration braking would have admitted the drive limit, where the finished budget runs past the clear distance.

A planner checks a future trajectory; a local kinetic barrier checks the current proposal against the modeled ability to stop. For torque-controlled motion, raw clearance has relative degree two. The velocity-dependent barrier in equation 3 produces the affine torque condition in equation 4 for positive approach speed. This local calculation still depends on a reserved stop, accurate state, and timely torque delivery.

The permission path needs an independent response when the proposal source stalls. On the warehouse mobile manipulator, a 60 ms chunk lease, a 30 ms host-heartbeat timeout, and a 5 ms MCU self-watchdog detect different failures. The MCU revokes proposal authority and chooses a prevalidated state-matched response; torque removal answers only the loss of the MCU’s own cycle, since every other trigger the record routes selects a budgeted stop in motion or a hold at standstill.

The enforcement record is where the two strands of records meet. Each entry keeps its evidence category, so the record states what the permission path assumed as plainly as what it checked.

Key Takeaways: Conditional enforcement and physical stopping
  • Tracking is a bounded premise: A steady-state error does not bound the transient, so a clearance inset needs a transient bound validated over the full maneuver in the operating envelope.
  • The finished budget sets the speed: Blind travel, braking under the resident stop, transient tracking error, and position-estimate error all enter one admission check, so the clear distance, not the drive limit, decides how fast the machine may move.
  • The controlled barrier includes speed: Raw distance has relative degree two for torque. A kinetic stopping barrier yields a torque half-space for positive approach speed, within the barrier’s operating contract.
  • Fallback is state dependent: Feasible projection, powered hold, controlled braking, and drive torque inhibit have different prerequisites. An empty admissible set cannot guarantee a safe stop.
  • Independent clocks detect different failures: Evidence age, proposal age, host heartbeat, MCU cycle progress, and external trips are checked separately; no one timeout is a universal braking guarantee.
  • Refusal must leave a record: The enforcement record binds every threshold to the detector that measures it and the rung a failed check selects, so the premises of each guarantee can be tested on silicon and judged at release.

The fallback ladder turns the law that proposal is not permission into a check the MCU executes on every tick, since each failed check selects a rung reserved before the fault, and the proposer can neither delay that selection nor override it. The kinetic barrier ties this law to irreversibility by making permission depend on speed as well as position, because a state from which the resident stop no longer fits is lost before any boundary is touched.

What’s Next: From barrier certificates to heterogeneous SoC placement
Can the learned proposers share a chassis and a memory bus with the permission path without delaying it? The enforcement record declares a worst-case cycle inside the permission deadline on every tick. That deadline is a claim about silicon, and a barrier result holds only while the permission path meets it. In Silicon Placement, we partition the learned proposers and the permission path across compute fabrics and name the loaded tests that would show whether memory starvation, clock throttling, or bus contention can delay permission or fallback, recording them in the placement record.

Back to top

Footnotes

  1. Control barrier functions and forward invariance: Control Barrier Functions define a continuously differentiable certificate \(h(\mathbf{x}) \ge 0\) over an admissible safe set \(\mathcal{C}\), enforcing the directional clearance derivative constraint \(\dot{h}(\mathbf{x}, \mathbf{u}) \ge -\gamma(h(\mathbf{x}))\). Enforcing this inequality renders \(\mathcal{C}\) forward-invariant: states initialized in a viable set can remain there while the barrier is feasible and the stated model and timing assumptions hold. When approaching the boundary, the constraint forces opposing actuation inputs that arrest momentum before physical contact can occur.↩︎

  2. Active-set quadratic programming for safety shields: On a bare-metal microcontroller, an active-set solver that restricts its linear constraints to active boundary contacts and allocates no dynamic memory can bound termination with a fixed iteration limit on a tested platform; no universal \(150\,\mu\text{s}\) limit follows from the solver type. If numerical ill-conditioning prevents convergence before the loop deadline, the solver aborts gracefully to an autonomous deceleration fallback.↩︎

  3. Closed-loop tracking stiffness and disturbance bounds: Sustained force and effective closed-loop stiffness determine static deflection. Transient tracking also depends on inertia, damping, command shape, and saturation. Disturbances outside the validated envelope invalidate the claimed tube; they need not produce exponential divergence.↩︎

  4. Remote Processor Messaging: Linux RPMsg provides virtio-based messaging between processors. Buffer ownership, copies, scheduling delay, and worst-case handover time depend on the platform; OpenAMP offers a separate no-copy API. The lease requires a measured, bounded delivery path.↩︎

  5. Hardware break circuits and PWM trip zones: A configured external timer break input such as TIMx_BKIN can force PWM outputs to a specified state without waiting for an interrupt. EPWM_TZFRC is a software force mechanism, not the external trip input. Output state and response time depend on the timer, gate drive, wiring, and qualification tests; the trip still needs a physically suitable stop path.↩︎

  6. Kinematic stopping distance and latency components: The inset stopping distance in this model compounds five terms: blind travel, braking distance under the credible deceleration, the localization bound, the fixed protective clearance, and the transient tracking bound. Omitting computational transport delay or tracking error results in premature encroachment on physical obstacles before peak braking can engage. See Kinematic Stopping Envelopes and Information Age Lag for formal kinematic derivations.↩︎

  7. Actuator bias torque and stopping margin loss: When an actuator allocates continuous baseline torque to counter gravitational loading or steady-state friction, available emergency braking headroom collapses proportionally. In vertical and heavily loaded axes, this depletion of net dynamic torque causes stopping distances to quadruple under seemingly modest bias loads. For a detailed derivation of saturated authority and stopping margin contraction, see Safety and Control Barrier Functions.↩︎

  8. Nagumo viability theorem and boundary tangent cones: Nagumo’s theorem establishes that a closed set is forward-invariant if and only if the system velocity vector belongs to the contingent cone of the set at every boundary point. Control Barrier Functions synthesize this geometric condition into real-time affine inequality constraints on actuator torque. Violating this subtangentiality condition causes immediate set exit, rendering future boundary enforcement unachievable; see Safety and Control Barrier Functions.↩︎

  9. Relative degree and inertia-aware barrier certificates: When a safety barrier monitors position while the actuator commands torque, the control input is separated from the constraint by two time integrals (\(r=2\)). Enforcing a zero-order spatial barrier without embedding kinetic stopping distance produces an optimization problem with zero control authority on the boundary. A velocity-dependent barrier can support a conditional forward-invariance argument while braking remains feasible and the state, delay, and tracking margins hold; see Safety and Control Barrier Functions.↩︎

  10. Worst-case execution time and active-set QP complexity: On bare-metal safety microcontrollers, QP safety filters execute within static memory arrays using fixed iteration limits. While typical active-set solve times range between \(60\text{--}120\ \mu\text{s}\), an uncapped solve can run to \(600\ \mu\text{s}\) under ill-conditioned or nearly degenerate constraint geometries, longer than the whole permission deadline. The fixed iteration cap exists to rule out that case. It bounds the solve to its declared share of the cycle, so the timing budget covers the capped solver rather than the uncapped worst case.↩︎

  11. Control barrier functions and Lie derivatives: In nonlinear control theory, the unforced drift and actuator control authority are formalized using Lie derivatives (\(L_f h(x)\) and \(L_g h(x)\)). When the safety boundary has relative degree one, the control input appears directly in the first time derivative of the barrier function, yielding an instantaneous affine constraint on commanded torque. If the relative degree is two, the enforcer must evaluate higher-order derivatives to prevent actuator saturation from rendering boundary violations inevitable; see Safety and Control Barrier Functions.↩︎

  12. Machine stop functions: IEC 60204-1 distinguishes uncontrolled power removal (Category 0), controlled stopping followed by power removal (Category 1), and controlled stopping with power retained (Category 2). IEC 61800-5-2 drive functions such as STO, SS1, and SS2 may implement parts of those responses. The required performance level, circuit architecture, and suitable stop are determined by the machine-specific risk assessment; ISO 13849-1 does not impose one universal PL or channel count.↩︎

  13. Dynamic brake choppers and DC-bus overvoltage clamping: In a regenerative Category 1 stop, mechanical kinetic energy flows through the inverter bridge into DC bus capacitors (\(\Delta E = \frac{1}{2} m v^2\)). A suitably rated chopper and resistor can dissipate energy when the DC bus rises; switching thresholds, pulse-energy rating, and thermal behavior must be validated for the stop profile. Other designs use mechanical dissipation or a battery that can accept regeneration.↩︎