Safety Enforcement
Safety Enforcement
Purpose
How does an unprivileged software proposal become physical torque, and what lets the gatekeeper refuse it in time?
A neural policy may propose commands that violate geometric clearance, exceed joint velocity limits, or saturate actuator torque. Once an unverified command is converted into motor phase currents, detecting the neural network’s hallucination comes too late: kinetic energy is already committed. An independent real-time enforcement path must evaluate candidate proposals before they cross the causal boundary and retain the uncompromising hardware authority to reject or project them within strict, sub-millisecond deadlines.
Timely refusal alone is insufficient; the selected fallback response must remain dynamically feasible under actual physical state. A mechanical brake cannot deliver idealized deceleration if thermal saturation or supply-rail voltage drops degrade torque generation. Runtime enforcement pairs control barrier functions with real-time feedback loops to verify that stopping envelopes remain strictly inside verified clearance margins. In the physical AI stack, enforcement embodies the core duty of the deterministic Nervous System: maintaining an unyielding permission path between statistical proposals from the Brain and physical execution by the Body.
Learning Objectives
- Explain the relationship between tracking accuracy, command permission, and independent enforcement at the actuator boundary
- Diagnose how update rate, feedback delay, and actuator saturation limit corrective control
- Calculate the finished stopping budget and the speed ceiling it sets from response delay, braking, localization, and tracking error
- Evaluate with a velocity-dependent barrier whether a tracker command computed from an admitted proposal must be accepted, projected, or refused
- Select fallback responses when sensing, computation, or actuator limits invalidate normal enforcement
- Construct an enforcement record that binds each check to its detector, its fallback response, and the records it consumes
The Deterministic Gatekeeper
The safety microcontroller (MCU) admits the warehouse mobile manipulator’s run down the aisle only after the matched stop suffix named in its trajectory record (The Trajectory Contract) is resident in the MCU’s memory and feasible from every state reachable before commitment. That admission checks a plan once, before the motion starts. The machine then executes the plan one millisecond tick at a time, and on each tick the live state can drift from the one the planner assumed. The base lags its commanded reference, a person steps out at the rack end, a payload loads the arm, or the chunk policy proposes a setpoint the admitted plan never contained. What follows is the check made on every tick, and the response it selects when even the resident stop no longer fits.
The learned proposers behind those commands can emit infeasible paths or joint commands beyond the actuators’ torque limits, and by the irreversibility law of The Four Bedrock Laws no exception handler can recall the momentum that current in a winding imparts, so a bad command has to be stopped before it reaches the drive.
Enforcement therefore keeps the task competence of learned policies while denying them authority over geometric clearance, thermal limits, and human-safety invariants. The enforcer of Multi-Rate Cadences, the permission path’s per-tick gatekeeper, sits on the permission side of the proposal boundary. It is a hard real-time component that evaluates candidate proposals, verifies invariant state bounds, and holds the exclusive hardware authority to permit current flow or command a fallback stop.
In the design this chapter follows, the enforcer runs on bare-metal microcontroller hardware wired directly to the motor drives and joint encoders and evaluates incoming setpoints each millisecond control cycle against validated constraints using Control Barrier Functions (CBFs)1 (Ames et al. 2019). When a proposal is safe, it passes to the motor drives unaltered. When a proposal marginally threatens a boundary, an active-set Quadratic Program (CBF-QP) minimally modifies the command, minimizing torque deviation while enforcing the derived constraint.2 When a proposal is structurally invalid or the proposer falls silent, the gatekeeper severs the command channel and triggers an autonomous fallback deceleration sequence.
However, an enforcer cannot guarantee safety beyond the tracking fidelity of the control loop beneath it, because every barrier certificate assumes that the tracking controller follows commanded references within a bounded error.
What the Tracker Hands Over
The tracker, the feedback controller that turns each admitted reference into a torque command on every tick (What Planning Hands Over), isolates the enforcer from the plant’s internal mechanics, and only one of its guarantees crosses the interface upward. That guarantee is an explicit bound on the worst-case tracking error (\(\epsilon_{\text{track}}\)) between where the machine was commanded to be and where its mass actually is. Within a validated envelope of reference acceleration, load, temperature, and disturbance, the tracker may provide a transient bound \(\|x(t) - r(t)\| \le \epsilon_{\text{track}}\) over the complete maneuver. The enforcer does not inspect transfer functions, root loci, or controller gains; to the software above the loop, the tracker is a black box that holds physical motion within the error bound \(\epsilon_{\text{track}}\) of the commanded path.
This bound is a physical distance set by rotor inertia, transmission compliance, and actuator torque limits. A steady-state stiffness calculation balances unmodeled disturbance against the restoring effort of the closed-loop stiffness, but it does not bound transient overshoot or settling, so the bound also needs validated dynamics or full-maneuver testing.3
On the warehouse mobile manipulator’s base, full-maneuver validation establishes \(\epsilon_{\text{track}} =\) 40 mm (illustrative; see the Reader Guide). The controller holds that bound only while three conditions hold. The reference cannot demand accelerations beyond what the drives can deliver, external disturbances must remain below \(d_{\text{max}}\), and the actuators must operate below the continuous thermal derating limits of Thermal Duty Cycles and the limits of their voltage rail. A reference that demands more acceleration than peak torque allows saturates the loop, the integrator winds up, and the base falls behind its reference by an amount that grows with every millisecond of saturation. Outside its rated envelope, the controller provides no bound at all.
Because the base can lag its reference by up to \(\epsilon_{\text{track}}\), every geometric and dynamic margin the enforcer evaluates is inset by that distance. A reference that keeps clearance \(c\) from the rack end guarantees the base only \(c - \epsilon_{\text{track}}\), so the enforcer checks the reference against the required clearance plus 40 mm, and section 1.5 carries that inset into the stopping distance, where delay and localization error also consume clearance. If worn drive bearings, tire wear, or altered motor parameters doubled the true error to 80 mm without the record changing, the base would spend another 40 mm of clearance that the enforcer believes it still holds. A degraded tracking bound silently erodes every margin above it.
↰ Prerequisite: Spatial clearance insetting from sensor covariance and transport delay originates in Spatial Error and Clearance Bounds.
A tracking bound never checked at runtime is an assumption, not a contract. The Nervous System must publish \(\epsilon_{\text{track}}\) with the operating conditions that validate it and monitor the instantaneous error \(e(t) = \|x(t) - r(t)\|\) on every tick. When \(e(t)\) exceeds the bound, under a disturbance spike or thermal throttling, the contract has broken, and the permission path must invalidate the spatial assumptions of the filters above it before unmodeled deflection causes damage.
Systems Perspective 1.1: The inset margin invariant
The synthesis of tracking controllers, observers, and disturbance rejection belongs to the control literature (Slotine and Li 1991; Åström and Murray 2021). In modern physical AI implementations, the “tracker” is rarely an elementary single-input single-output feedback loop; it is typically structured as a cascaded control hierarchy (translating Cartesian path targets into joint velocity references, and velocity into field-oriented phase currents at \(10\text{--}25\text{ kHz}\)) or as a high-rate tracking Model Predictive Controller (MPC) running at \(100\text{--}500\text{ Hz}\) on bare-metal firmware to solve a localized quadratic program for dynamic disturbance rejection. Yet regardless of its internal mathematical sophistication, an optimal tracker remains semantically blind: it possesses no internal model of obstacle clearances, task intent, or human proximity, and it will execute a collision path into a steel upright with the same dynamic fidelity it brings to unobstructed free space. Furthermore, when unmodeled physical disturbances occur—such as surface oil causing gross tire slip or sudden thermal current derating—even an optimal tracking controller can guarantee tracking only within its qualified error tube \(\epsilon_{\text{track}}\). The enforcer therefore treats the tracker strictly as a contract boundary: it encapsulates the entire feedback control stack inside the certified bound \(\epsilon_{\text{track}}\), evaluating candidate proposals against forward-invariant barrier certificates before any command reaches the tracking layer.
Where the Enforcer Sits
The enforcer’s position in the signal pipeline is set by the dynamics’ timescale and the latencies of the channels that feed it, and every machine that turns learned proposals into motion offers two natural positions. Upstream of the feedback loop, a kinematic reference checker validates waypoints, coordinate targets, or velocity vectors before they enter the tracking loop; at planning rates of \(10\) to \(50\text{ Hz}\) it can check a \(300\text{ ms}\) trajectory segment against workspace boundaries and self-collision models. Its strength is spatial foresight and its limitation is blindness to the plant’s actual state, so a reference that clears a rack upright by less than the base’s transient overshoot passes every upstream check while the base still strikes the upright.
↳ Downstream: Heterogeneous silicon partitioning of the enforcer onto isolated microcontroller enclaves is mapped in Two Paths on One Die.
Downstream of the feedback controller, near the power electronics, a command checker inspects phase currents, duty cycles, and joint torque references immediately before they are written to hardware registers. Running at the machine’s 1 kHz inner-loop rate, it prevents overcurrent, over-torque, and thermal runaway by clipping or zeroing signals beyond the plant’s ratings, but it surrenders foresight. A torque well inside its continuous rating looks safe to a scalar clamp even while it accelerates a descending arm joint toward the packing station faster than the remaining travel can absorb, and the clamp registers no fault until no allowable reverse torque can prevent the collision.
The split trades latency for visibility. A reference checker at those planning rates sees several steps ahead but cannot respond to resonance, noise amplification, or instability above its update rate. A command checker acts on every tick, but its one-step horizon creates dynamic traps in which every instantaneous torque is legal until an irrecoverable state is reached.
Placement also decides how the enforcer joins the computing fabric. An in-line enforcer inspects every command before the next stage reads it, which leaves no command uninspected but puts its worst-case execution time on the critical path of the 1000 μs loop, where an overrun faults the whole cycle. A parallel monitor on a separate core preserves loop timing and opens a safety contactor or switches the drive to dynamic braking when it detects a violation, but its detection and signaling lag, bounded by the monitor cycle and bus jitter, must be paid for with wider margins or lower speed.
Because neither single location provides both dynamic foresight and instantaneous protection, physical AI architectures resolve the tension through a dual-stage enforcement pattern. This pattern maps onto the two-processor implementation of the machine model (Where the Permission Path Runs):
- Upstream Stage (Advisory Corridor Pre-Check on the Application Processor): Operating at the chunk rate (20 \(\text{Hz}\) on this machine) on an unprivileged application microprocessor core, this stage screens multi-step trajectory chunks emitted by the learned policy against a dynamic stopping corridor under worst-case braking parameters and against static workspace envelopes before packing the waypoints into a leased chunk proposal sent over RPMsg. The stage is advisory. It spares the MCU proposals that would fail, but it runs on the proposer’s side of the proposal boundary, so the binding admission of the trajectory and its stop suffix stays on the MCU (Behavior at the Seam).
- Downstream Stage (Instantaneous CBF-QP Safety Shield on Bare-Metal MCU): Operating inside a strict 1 kHz (1000 μs) timer interrupt loop on the isolated MCU, this stage receives the proposed setpoint from shared SRAM, verifies tracking error bounds, solves an active-set Control Barrier Function (CBF-QP) quadratic program in pre-allocated static SRAM, and clamps raw current and torque commands directly before writing to PWM timer registers. The barrier it enforces reduces to a single-step convex projection (section 1.7), whose execution time must be bounded on the selected hardware.
The advisory pre-check screens for a feasible stop before a chunk is sent, and the drive-side stage checks current observations and commands on every tick, but neither stage can recover a state that has already exhausted braking clearance.
Correct placement does not help if the enforcer’s inputs share failure modes with the component it checks. An enforcer that computes collision distances from a depth map produced by the same neural perception model that planned the motion is deceived by the same out-of-distribution artifact, so independent enforcement requires independent sensing and isolated computation. The MCU reads encoders and limit switches wired to its own GPIO and timer-capture registers, outside the Linux virtual memory space and scheduler, and a collision check reads a safety lidar or proximity sensor whose conditioning uses deterministic threshold logic rather than learned inference. Where no independent sensor exists, as for the friction under the base’s wheels or the mass of a payload that only an estimator infers, no independent physical guarantee can be built, and the claim is only as strong as the estimator’s bounded uncertainty or the premise the record states. Every invariant the enforcer checks, upstream or at the motor terminals, still assumes that the plant retains the authority to arrest its own motion before a forbidden boundary.
Systems Perspective 1.2: The monitor independence rule
Stopping Envelopes
The chapters that own a stopping term have each added it to the warehouse mobile manipulator’s budget at the rack end, and one term remains. The stopping envelope is the set of states from which a validated braking response can arrest motion before a boundary, and the irreversibility law (principle \(\ref{pri-vol4-irreversibility}\)) sets its test, \(d_{\text{stop}} \le D_{\text{clear}}\) with every term of \(d_{\text{stop}}\) taken from a measured limit of this machine. Reserving that response costs achievable speed, because velocity increases blind travel linearly and braking travel quadratically. The machine must remain inside this envelope before a fault or unsafe proposal occurs.
The pre-brake delay \(\tau_{\text{delay}}\) in the stopping distance of Kinetic Momentum already includes the enforcer’s own permission-loop tick, which Multi-Rate Cadences budgets with the lease path. What remains is the gap between the reference and the plant that follows it.
The permission path closes that gap with the tracking bound \(\epsilon_{\text{track}}\) of section 1.2 and defends the stopping distance of equation inset by it. A commanded reference can be admitted only if the plant, which may lag that reference by up to \(\epsilon_{\text{track}}\), can still stop inside \(D_{\text{clear}}\). The admission condition is the stopping envelope:6 \[v\,\tau_{\text{delay}} + \frac{v^2}{2a_{\text{eff}}} + \delta_{\text{loc}} + \delta_{\text{margin}} + \epsilon_{\text{track}} \le D_{\text{clear}} \tag{1}\] Here \(a_{\text{eff}}\) is the deceleration the resident stop actually uses, equal to the credible deceleration \(a_{\text{brake}}\) for a constant-deceleration stop and smaller when the stop is shaped as in Behavior at the Seam; \(\delta_{\text{loc}}\) is the localization bound and \(\delta_{\text{margin}}\) the fixed protective clearance. The left side is the stopping distance \(d_{\text{stop}}(v)\) of equation plus the tracking inset \(\epsilon_{\text{track}}\).
Under those delay, estimate, and braking assumptions, any trajectory proposal that places the machine at a distance less than \(d_{\text{stop}}(v) + \epsilon_{\text{track}}\) in equation 1 from a physical constraint along the direction of motion is unsafe and physically inadmissible. The stopping envelope is fundamentally anisotropic; an obstacle behind the instantaneous velocity vector requires clearance only for the positional tracking error \(\epsilon_{\text{track}}\).
For the warehouse mobile manipulator, the inset is the base’s tracking bound of 40 mm, and it is the last distance term the budget receives. At the 1.5 m/s drive limit, the matched stop suffix of Behavior at the Seam had already carried the running total to 1,194.15 mm, past the 1.10 m clear distance at the rack end. The inset brings it to 1,234.15 mm, 134.15 mm more than the clear distance.
Every other term in equation 1 is now fixed by the chapter that owns it: the credible deceleration and the fixed overheads by Kinetic Momentum, the lease path by Multi-Rate Cadences, the observation age by Sensor Transduction and Calibration, the stop’s shape by Behavior at the Seam, and the inset here. Speed is the one term the machine can still choose. With \(\tau_{\text{delay}} =\) 133.6 ms and the suffix’s \(a_{\text{eff}} =\) 1.33 m/s² (equation), equation 1 holds up to a ceiling of 1.39 m/s. The finished budget therefore sets the aisle speed at 1.3 m/s, below that ceiling. At the aisle speed the stop needs 997.4 mm and leaves 102.6 mm of the clear distance unspent. Table 1 lists every row at both speeds, and figure 1 draws them.
| Row | Added in | Term | At drive limit (mm) | Running total (mm) | Left of \(D_{\text{clear}}\) (mm) | At aisle speed (mm) |
|---|---|---|---|---|---|---|
| 1 | Kinetic Momentum | Braking, \(v^2/2a_{\text{brake}}\) | 562.5 | 562.5 | 537.5 | 422.5 |
| 2 | Kinetic Momentum | Overhead, \(\delta_{\text{loc}}+\delta_{\text{margin}}\) | 150 | 712.5 | 387.5 | 150 |
| 3 | Kinetic Momentum | Brake onset, \(vT_{\text{act}}\) | 30 | 742.5 | 357.5 | 26 |
| — | Supply, Freshness, and the Memory Wall | Time only: the chunk policy can renew a lease; the intent model’s 98.0 ms weight sweep cannot | — | — | — | — |
| 4 | Multi-Rate Cadences | Lease path, \(v(T_{\text{lease}}+T_{\text{tick}}+T_{\text{bus}})\) over 62 ms | 93 | 835.5 | 264.5 | 80.6 |
| 5 | Sensor Transduction and Calibration | Observation age, \(v\,t_{\text{age}}\) over 51.6 ms | 77.4 | 912.9 | 187.1 | 67.08 |
| — | Belief Through Occlusion | Time only: belief that an occluded rack end is clear expires in 31.25 ms, under two camera frames, so memory cannot certify it | — | — | — | — |
| 6 | Behavior at the Seam | \(C^2\) stop suffix, \(0.75v^2/a_{\text{brake}}\) in place of \(0.5v^2/a_{\text{brake}}\) | 281.25 | 1,194.15 | \(-94.15\) | 211.25 |
| 7 | section 1.5 | Tracking inset, \(\epsilon_{\text{track}}\) | 40 | 1,234.15 | \(-134.15\) | 40 |
| Total | 1,234.15 | 997.43 | ||||
| Left of \(D_{\text{clear}}\) | \(-134.15\) | 102.57 |
The finished budget protects a static obstacle at \(D_{\text{clear}}\), such as a person who has stepped out from the rack and stopped or a dropped tote, so it assumes that the obstacle stays where it was when the stop began. A person who keeps walking toward the base through the whole stop closes the gap from the far side as well (Belief Through Occlusion), and counting that approach on top of the finished budget would cap the base at 0.46 m/s at this rack end. An aisle held to that speed would slow every run for an encounter that a crossing rule can keep out of the stop, so the site answers the walking person with the rule rather than with speed, and A Residual-Claims Register keeps open the premise the rule rests on, that no person approaches the rack end during a stop. The enforcer therefore guarantees the budget against a stopped obstacle, not the safety of a person who keeps walking into the stop.
Later chapters spend the remainder or restrict the premise beneath it. At the aisle speed the unspent distance lasts 78.9 ms (Authority Transitions). The credible deceleration binds only on floors with friction of at least 0.204, since \(a=\min(a_{\text{brake}},\mu g)\), so on a floor inspected to 0.12 the ceiling falls to 1.09 m/s (A Residual-Claims Register).
A stalled proposer would spend more than the whole remainder. Were the chunk lease of section 1.3 left to the host to honor, a last proposal that kept its authority for 200 ms instead of the 60 ms lease would add 182 mm to the stop and carry the base 79.4 mm past the clear distance at the rack end.
The same budget also sizes any protective sensing that stops the base independently of the proposer. A mobile base can sense the rack end through a risk-assessed protective field designed under the applicable vehicle and machinery standards, including ISO 3691-4 and ISO 13849, such as the time-of-flight safety laser scanner shown in figure 2, and that protective field must be at least as long as the stopping budget.
In the distance–velocity plane of figure 3, the stopping equation draws a parabolic boundary between states that can still stop and states that have exhausted their clearance, and the gap to the braking-only curve is the clearance consumed by pre-brake delay, the matched stop’s gentler deceleration, and the fixed allowances. As the roofline model (Williams et al. 2009) caps attainable throughput by memory bandwidth, the stopping envelope caps legal speed by clear distance, however quickly the chunk policy proposes. To run faster, the engineer must raise the credible deceleration \(a_{\text{brake}}\), shorten \(\tau_{\text{delay}}\), or tighten \(\epsilon_{\text{track}}\) and \(\delta_{\text{loc}}\).
These parameters are not stationary. An item in the gripper raises the arm’s inertia and lowers the deceleration its joints can produce within peak torque, oil on the floor lowers the traction beneath the base’s credible deceleration, and repeated hard stops heat rotors and inverters into thermal derating (Thermal Duty Cycles). An enforcer that treats \(a_{\text{brake}}\) as a datasheet constant underestimates the envelope precisely under heavy load, so it must use independently bounded braking capability and reduce permitted speed when validated friction, payload, or thermal conditions change; an uncertain estimate cannot by itself restore a guarantee.
Adaptation works only while the enforcer’s state estimates are accurate and the actuators can still supply the force. If contamination takes traction toward zero or a sustained duty cycle pushes the windings past their ceilings, the assumed deceleration vanishes, and once the required torque exceeds what the hardware can produce, the enforcer can detect the violation but can no longer prevent it.
Safe Sets as Conditional Permission
Once its bias torque is counted, the descending arm of section 1.6 needs 79.45 mm of tool travel to stop, and that travel grows with the square of approach speed. A clearance that suffices at one speed is therefore too short at a higher one, and the enforcer needs a predicate over position and speed together that it can evaluate on every tick of the control clock. In control theory, a safe set \(\mathcal{C}\) is defined as the zero-superlevel set of a continuously differentiable scalar certificate \(h(x): \mathcal{X} \to \mathbb{R}\), such that \[\mathcal{C} = \{x \in \mathcal{X} \mid h(x) \ge 0\}, \quad \partial \mathcal{C} = \{x \in \mathcal{X} \mid h(x) = 0\}.\] The mathematical foundation of forward invariance, established by Blanchini (Blanchini 1999), formulated through Hamilton-Jacobi reachability analysis and viability kernels by Mitchell, Bayen, and Tomlin (Mitchell et al. 2005), and extended through Control Barrier Functions by Ames et al. (2014) and Ames et al. (2019),8 characterizes conditions under which an admissible controller can keep a state within a viable set. Geometric clearance alone is not that set when stopping torque is finite.
Definition 1.1: Control barrier function safety filter
Control barrier function safety filter is the constrained optimization filter that projects unverified nominal control proposals \(\mathbf{u}_{\text{nom}}\) onto a forward-invariant safe set \(\mathcal{C} = \{\mathbf{x} \mid h(\mathbf{x}) \ge 0\}\) by solving the quadratic program: \[\min_{\mathbf{u}} \frac{1}{2}\|\mathbf{u} - \mathbf{u}_{\text{nom}}\|^2 \quad \text{s.t.} \quad \mathbf{a}(\mathbf{x})^\top \mathbf{u} \le b(\mathbf{x}), \quad \mathbf{u}_{\min} \le \mathbf{u} \le \mathbf{u}_{\max}\] that enforces linear braking half-space constraints (\(\mathbf{a}(\mathbf{x})^\top \mathbf{u} \le b(\mathbf{x})\)), leaving candidate inputs that satisfy the safe boundary untouched and minimally projecting violating inputs in control space.
- Significance: Rather than attempting the impossible task of proving formal safety across billions of uninterpretable neural network weights, a CBF-QP filter decouples high-level policy innovation from low-level safety enforcement. The learned model is free to explore complex manipulation or navigation strategies, while the deterministic enforcer can preserve a validated viable set while its physical assumptions remain true.
- Distinction: Unlike a heuristic clamp that saturates control channels independently (destroying multi-axis coordination), a CBF-QP filter performs an optimal orthogonal projection in control space, minimizing the chosen torque-space distance; it does not automatically preserve a Cartesian path.
- Common pitfall: Formulating barrier constraints using raw geometric coordinates without accounting for relative degree (\(r = 2\)) or braking inertia. If the barrier does not incorporate stopping distance (\(\Delta d_{\text{stop}} = v\,\tau_{\text{delay}} + v^2 / 2a_{\text{brake}}\)), the filter intervenes too late, after the vehicle’s momentum has already made a boundary collision inevitable.
When the tracker turns a learned policy’s unverified setpoints into a nominal torque \(\mathbf{u}_{\text{nom}}\), the enforcer does not forecast the machine’s entire future path. It evaluates a single half-space inequality at the current state \(\mathbf{x}\), the linear constraint that bounds the safe approach toward the boundary. If the proposed action satisfies \(\mathbf{a}(\mathbf{x})^\top \mathbf{u}_{\text{nom}} \le b(\mathbf{x})\), the enforcer passes it to the motor untouched; if it violates a feasible half-space, the enforcer projects its torque vector.
A barrier on the arm’s distance to the station surface alone would register the danger only once momentum had already made contact inevitable,9 so the barrier must carry the kinetic stopping angle (\(\omega^2 / 2\alpha\) for approach rate \(\omega\) and braking bound \(\alpha\)) as well as position, a requirement that follows from relative degree. Actuators command current and torque, torque produces acceleration, and position therefore lies two integrals from the input: \[\text{Actuation } u \xrightarrow{\int \frac{1}{m} dt} \text{Velocity } v \xrightarrow{\int dt} \text{Position } x\] This two-step cascade defines relative degree \(r = 2\). Differentiating clearance once yields velocity: \[\dot{d} = -v\] The input \(u\) does not yet appear, because a motor cannot step velocity instantaneously. Only the second derivative brings in Newton’s second law (\(m \ddot{x} = u - d_{\text{ext}}\)): \[\ddot{d} = -\dot{v} = -\left( \frac{u - d_{\text{ext}}}{m} \right)\]
A raw-distance rule \(\dot d\ge-\gamma d\) cannot directly constrain this torque input: \(\dot d=-v\) contains no \(u\). For an approaching arm, let \(D=d_{\text{tool}}/\ell\) be the remaining angular clearance, \(\omega>0\) its approach speed, \(J\dot\omega=\tau+\tau_{\text{bias}}\), and \(\alpha\) a validated lower bound on available angular braking. Define the kinetic barrier \[h(D,\omega)=D-\frac{\omega^2}{2\alpha},\qquad \dot h=-\omega-\frac{\omega(\tau+\tau_{\text{bias}})}{J\alpha}. \tag{3}\] The raw distance \(D\) has relative degree two; this velocity-dependent \(h\) has relative degree one for \(\omega>0\). Requiring \(\dot h+\gamma h\ge0\) produces the actual affine torque half-space \[\tau\le-\tau_{\text{bias}}+J\alpha\left(\frac{\gamma h}{\omega}-1\right),\qquad -\tau_{\max}\le\tau\le\tau_{\max}. \tag{4}\] At standstill (\(\omega=0\)), the controller uses a separate no-approach and acceleration admission check; it does not divide by zero.
Carry the descending-arm scenario of section 1.6 into this barrier: \(J=\) 3 kg·m², \(\tau_{\text{bias}}=\) 65.25 N·m toward the station, \(|\tau|\le\) 87 N·m, \(\alpha=(\tau_{\max}-\tau_{\text{bias}})/J=\) 7.25 \(\text{rad/s}^2\), \(\ell=\) 0.80 m, \(\omega=\) 1.2 \(\text{rad/s}\), and \(\gamma=\) 10 \(\text{s}^{-1}\). These are stated joint-level scenario inputs; the bias is not inferred from the payload mass alone. At 110 mm of tool clearance to the station surface, \(D=\) 0.1375 \(\text{rad}\) and \(h=\) 0.0382 \(\text{rad}\), so equation 4 requires the joint to brake with at least 80.1 N·m. At 79.45 mm, \(h=0\) and the requirement reaches 87 N·m, the joint’s full peak torque. At 70 mm, it demands 89.1 N·m, more than the joint can produce, so this state is already outside the modeled viable set. The enforcer must prevent entry into it; a fallback started there cannot restore the lost clearance.
A real machine realizes neither exact state feedback nor instantaneous torque. Stator current rise, bus jitter, and gearbox compliance delay and soften the corrective action, so the enforcer evaluates the barrier against a boundary inset by the worst-case stopping deficit they cause and intervenes earlier in the trajectory.
The safety warrant provided by the barrier calculation is strictly conditional. A forward-invariance claim requires a feasible barrier and the following bounded conditions; the list is necessary for this design, not an exhaustive if-and-only-if test:
- Model Fidelity: The nominal plant dynamics \(f(x)\) and \(g(x)\) must capture the true system dynamics within a certified parameter error bound.
- Disturbance Bounds: External forces and unmodeled friction acting on the mass must remain strictly within the modeled disturbance envelope \(|d(t)| \le d_{\max}\).
- Actuation Authority: Physical actuators must possess sufficient voltage, current, and thermal headroom to realize the commanded corrective force without saturation (\(u^* \in [u_{\min}, u_{\max}]\)).
- Solver Determinism: The optimization solver or projection algorithm must terminate strictly within its allocated execution budget (\(T_{\text{solve}} \le T_{\text{budget}} < T_{\text{loop}}\)).
- Discretization Invariance: The discrete sampling interval \(T_{\text{ctrl}}\) must be sufficiently small that inter-sample state drift does not breach the continuous-time invariance boundary.
Each condition maps to a budget. Model fidelity and disturbance bounds come from system identification and load testing on the bench, actuation authority from dynamometer curves and thermal limits (Thermal Duty Cycles), solver determinism from static worst-case execution time analysis on the MCU, and discretization from the loop period, which at \(T_{\text{ctrl}} =\) 1 ms must keep the zero-order-hold drift between samples to a fraction of \(\delta_{\text{margin}}\).
A condition that no detector watches is an assumption, so the permission path attaches one to each: a disturbance observer compares predicted with measured acceleration on every cycle, inverter status registers and current shunts report actuator saturation, a hardware watchdog and cycle counter time the solver against its deadline, and timestamp counters on the sensor bus catch jitter and missed samples.
When any detector trips, the premise of the safe set has failed, and the enforcer must not recompute the barrier with optimistic parameters. It transfers authority through the output selector to a state-matched fallback whose response time and braking margin were reserved before the fault; later rungs can only limit harm once the viable set is lost. The safe set is a valid permission only while its operating contract is intact, and the permission path enforces safety by knowing when that contract has failed.
Minimal Intervention
Suppose the chunk policy proposes setpoints from which the tracker computes a torque that the kinetic barrier of equation 4 forbids. Replacing the learned commands with a hand-crafted conservative controller discards the competence that justifies the model, and braking hard at the first boundary encounter creates acceleration spikes and interrupts the task. The principle of minimal intervention instead leaves an admitted proposal unchanged when it satisfies every active constraint and changes it only when projection is both needed and feasible. Enforcement then acts as a safety shield (projection operator), accepting the proposal as the default intent and computing the smallest correction that keeps the system within the invariant set.
↰ Prerequisite: Nominal trajectory setpoints and feedforward torques are synthesized by the planner in From Goal to Trajectory.
This filter is the CBF-QP of 1.1 (Ames et al. 2016). At each discrete control interval \(T_{\text{ctrl}} = 1.0\text{ ms}\) on the bare-metal microcontroller, the tracker computes the nominal torque command \(\mathbf{u}_{\text{nom}} \in \mathbb{R}^m\) each tick from the admitted reference (What Planning Hands Over), and the enforcer solves the convex problem10 \[\min_{\mathbf{u}} \frac{1}{2}\|\mathbf{u} - \mathbf{u}_{\text{nom}}\|^2 \quad \text{s.t.} \quad \mathbf{a}(\mathbf{x})^\top \mathbf{u} \le b(\mathbf{x}), \quad \mathbf{u}_{\min} \le \mathbf{u} \le \mathbf{u}_{\max} \tag{5}\] where \(\mathbf{a}(\mathbf{x})^\top \mathbf{u} \le b(\mathbf{x})\) represents a torque constraint derived from a velocity-dependent barrier such as equation 4,11 and \([\mathbf{u}_{\min}, \mathbf{u}_{\max}]\) represents the verified actuator torque envelope.
When enforcing a single active linear barrier constraint \(\mathbf{a}^\top \mathbf{u} \le b\) in equation 5, the optimization admits an exact closed-form orthogonal projection, derived in Minimal-intervention quadratic program: \[\mathbf{u}^* = \mathbf{u}_{\text{nom}} - \frac{\mathbf{a}^\top \mathbf{u}_{\text{nom}} - b}{\|\mathbf{a}\|^2} \mathbf{a} \tag{6}\] where the optimal Lagrange multiplier \(\lambda^* = (\mathbf{a}^\top \mathbf{u}_{\text{nom}} - b) / \|\mathbf{a}\|^2\) scales the minimal intervention vector \(\Delta \mathbf{u} = -\lambda^* \mathbf{a}\).
The projection arithmetic is clearest on a two-axis manipulator with illustrative values. Its tracker, following setpoints from an unconstrained neural policy, computes nominal joint torques \(\mathbf{u}_{\text{nom}} = [\) 3.5, 3 \(]^\top\text{ N}\cdot\text{m}\). Suppose an abstract safety half-space has already been derived and validated, \(\mathbf{a}^\top \mathbf{u} \le b\) with \(\mathbf{a} = [\) 2, 1 \(]^\top\) and \(b =\) 5 \(\text{N}\cdot\text{m}\), with per-joint actuator limits of $$5 \(\text{N}\cdot\text{m}\). The candidate gives \(\mathbf{a}^\top \mathbf{u}_{\text{nom}} =\) 10 \(\text{N}\cdot\text{m}\) and breaches the boundary by 5 \(\text{N}\cdot\text{m}\).
By equation 6, \(\lambda^* =\) 1, and the minimal-intervention command is \(\mathbf{u}^* = [\) 1.5, 2 \(]^\top\text{ N}\cdot\text{m}\), which meets the half-space exactly while both joints stay within their limits, with an intervention norm of 2.236 \(\text{N}\cdot\text{m}\).
Per-axis clamping shows why the projection is needed. Clamping Joint 1 to 2 \(\text{N}\cdot\text{m}\) while passing Joint 2 at 3 \(\text{N}\cdot\text{m}\) gives \(\mathbf{a}^\top \mathbf{u}_{\text{clip}} =\) 7 \(\text{N}\cdot\text{m}\), still 2 \(\text{N}\cdot\text{m}\) outside the constraint, and it skews the torque vector from 40.6\(^\circ\) to 56.3\(^\circ\). The clamp therefore both violates the constraint and distorts the command. The projection satisfies the constraint with the smallest possible torque change, although it too rotates the torque vector, so the end-effector path still needs a check against the coupled plant model.
Figure 4 shows both views. In state space (panel a), proposals that would cross the barrier are deflected along the inset boundary \(\partial \mathcal{C}\); in control space (panel b), the enforcer projects \(\mathbf{u}_{\text{nom}}\) onto the half-space \(\mathbf{a}^\top \mathbf{u} \le b\), while the independently clamped vector stays outside it.
Filter choice joins a mathematical condition to a measured execution budget and the machine’s available braking authority. As summarized in table 2, architectural choices differ in their handling of relative degree, solver worst-case execution time (WCET), forward invariance guarantees, and directional fidelity. Independent clamping can miss coupled constraints; a CBF-QP can enforce a derived half-space if it is feasible and finishes within the verified cycle. Predictive filters (Wabersich and Zeilinger 2021) trade a longer horizon for configuration-dependent compute cost.
| Filter | Input and condition | What it can establish | Main validation task |
|---|---|---|---|
| Independent torque clamp | Each axis checked separately | Hardware amplitude limit, not coupled stopping clearance | Test coupled wrench and stopping behavior |
| First-order CBF-QP | Barrier whose first derivative contains the input | Conditional invariance from a viable initial state | Validate model, feasible torque set, and cycle bound |
| Kinetic / higher-order barrier | Position constraint with speed and braking authority included | Conditional stopping margin before geometric contact | Bound bias torque, delay, estimate error, and discretization |
| Precomputed viable region | State-to-permitted-input lookup | Containment for the modeled region | Verify lookup coverage and interpolation |
| Predictive safety filter | Finite-horizon trajectory and terminal set | Recursive feasibility only under terminal and timing assumptions | Bound full solve time and fallback on timeout |
Operating near an active safe-set boundary can produce chattering. If a learned policy repeatedly proposes an unsafe command, sampled enforcement may switch the active constraint across successive 1 ms ticks. Depending on the plant and controller, rapid torque changes can excite a structural mode or increase motor heating and transmission wear. A tested boundary layer or hysteresis can reduce switching, but its added tracking error must fit inside the validated stopping margin.
The intervention vector \(\Delta u = u^* - u_{\text{nom}}\) also measures upstream policy health. A policy operating inside its training distribution produces zero intervention on most cycles, so a rising intervention frequency or magnitude can indicate proposal drift, a changed task, or a changed plant. Summed over an illustrative rolling window of \(W = 500\text{ ms}\) as \(\sum_{k} \|u_k^* - u_{\text{nom}, k}\| \Delta t\), the metric lets the permission path flag divergence for logging or operator notification while the per-tick limits still protect the plant.
In this heterogeneous design, the shield executes inside a 1 kHz timer interrupt routine on the MCU. An unloaded cycle of acquisition, validation, and a typical active-set solve completes in about 135 \(\mu\text{s}\), an illustrative estimate rather than a bound. The qualified path instead declares a worst-case execution time of 250 \(\mu\text{s}\) for the whole cycle (Real-Time Timers, Watchdogs, and Schedulability), divided among the stages below, and the whole sequence must finish inside the 400 \(\mu\text{s}\) deadline:
- Interrupt Entry & State Acquisition (30 \(\mu\text{s}\)): The hardware timer triggers the ISR, and firmware latches the wheel and joint encoder registers and the latest IMU sample.
- Proposal & Lease Verification (15 \(\mu\text{s}\)): The MCU inspects the shared SRAM ring buffer populated by the application processor over RPMsg or shared memory. It validates the checksum, confirms sequence order, reads the proposal’s evidence epoch, and makes two age checks against it and the lease. The upper age of the evidence the proposal was built on, stamped at dispatch, may not exceed the 51.8 ms refusal threshold, which is the 51.6 ms observation age the stopping budget charged plus the 0.2 ms clock-conversion bound (The Observation Contract). The proposal’s own age on the MCU clock may not exceed the 60 ms chunk lease. The MCU also reads the host heartbeat timer, which marks the host silent once no heartbeat has arrived within the 30 ms timeout.
- Tracking Contract Verification (10 \(\mu\text{s}\)): Firmware evaluates instantaneous tracking error \(e(t) = \|x(t) - r(t)\|\) against the certified bound \(\epsilon_{\text{track}}\).
- CBF-QP Safety Shield Solve (150 \(\mu\text{s}\)): An active-set QP solver running in pre-allocated static DTCM SRAM (Where the Permission Path Runs) evaluates active barrier inequalities and computes \(u^* = \arg\min_u \frac{1}{2}\|u - u_{\text{nom}}\|^2\). A fixed iteration cap holds the solve inside its share even for ill-conditioned constraint geometry; nominal solves for one or two active constraints finish well within it. Dynamic memory allocation is forbidden, because a heap search has no bounded duration.
- Output Selection & Setpoint Staging (15 \(\mu\text{s}\)): If the QP is feasible and constraints hold, \(u^*\) is staged as the admitted setpoints for the next EtherCAT frame to the drives, whose current loops apply it and whose torque accuracy is tested rather than inferred from their rate; a proposal that fails validation leaves the previously admitted setpoints in force. If the chunk lease has expired, the host heartbeat has fallen silent, the shared control rail reports undervoltage, tracking error exceeds \(\epsilon_{\text{track}}\), actuator saturation renders the feasible control set empty (\(\mathcal{U}_{\text{safe}} = \emptyset\)), or an operating-contract detector such as the disturbance observer or the solver deadline trips, the output selector revokes the policy and selects the state-matched stop. At standstill, the same lease, heartbeat, or undervoltage trigger selects a powered hold instead.
- Completion & Reserve (30 \(\mu\text{s}\)): The MCU checks the register write and services its own watchdog only after the cycle succeeds. The reserve absorbs variation in the stages above; a cycle that has not completed by the deadline transfers authority to the resident fallback before the next tick. A missed tick therefore costs the base 1.3 mm of travel at the aisle speed. Testing the Separation tests whether the declared WCET holds under load, so that ticks are not missed.
Every stage of this cycle assumes that the solve returns a feasible command. When saturation empties the admissible torque set, the lease lapses, or the cycle overruns its deadline, projection has nothing to return, and the permission path must already hold a response it can execute without the proposer.
The Fallback Ladder
The fallback must be selected before the admissible set becomes empty. Once it is empty, a fallback can contain a fault or limit harm but cannot promise to avoid a boundary crossing. Nor can the response be a single switch. Cutting bus power whenever a proposal brushes a barrier margin produces uncoordinated deceleration and continuous mission aborts, while relying on software optimization during an inverter shoot-through (Actuator Transmission Limits) burns the gate drivers before an interrupt handler can run. The permission path therefore holds several resident responses and selects among them deterministically, from feasible projection to plant-specific controlled stopping or torque removal.
Much as the arbitration of behavior-based robotics and the subsumption architecture (Brooks 1986) lets a condition select which layer drives the actuators, the fallback ladder is a set of four deterministic responses (project: minimal-intervention QP projection; hold: active position hold; stop: controlled dynamic stop; inhibit: drive torque inhibit), each running on the execution substrate that survives its trigger, from which the permission path selects by the check that failed and the state of the plant when it failed, rather than climbing them in order.
The first rung, least-squares projection, runs on every cycle while the feasible control set is nonempty (\(\mathcal{U}_{\text{safe}} \neq \emptyset\)), computing \(u^* = \arg\min_u \|u - u_{\text{nom}}\|^2\) subject to the active barrier inequalities within its 150 \(\mu\text{s}\) share of the cycle. It succeeds only while the barrier constraints and the achievable torques intersect, so the admission path must preserve viability before saturation empties \(\mathcal{U}_{\text{safe}}\). If infeasibility occurs anyway, the MCU revokes the proposal and invokes the best available state-matched response.
Where the plant can be held rather than stopped, the response is a controlled deceleration into an energized position hold, analogous to a Category 2 stop.12 The permission path arrests motion and holds the current coordinate \(x_{\text{hold}}\) under closed-loop feedback, with the drives still modulating current against gravity and payload. The rung spends electrical energy and accumulates thermal debt in the windings (Thermal Duty Cycles), but it wears no brake and preserves the kinematic state, so after a fresh proposal and state check the enforcer may resume without homing. On the warehouse mobile manipulator this rung answers a lapsed lease or a silent heartbeat only at standstill. The same expiry in motion selects the stop rung, because the resident stop is the only deceleration whose time and distance the stopping budget has admitted. If tracking exceeds its bound, the MCU selects a validated stop, and holding remains appropriate only while available torque and feedback support it.
When a body in motion loses its proposal’s authority or a tracking or model assumption fails, the response is a controlled dynamic stop, analogous to a Category 1 stop under IEC 60204-1. On the warehouse mobile manipulator’s base, this rung executes the resident stop suffix admitted with the trajectory (Behavior at the Seam) rather than composing a stop at the moment of the fault. From the 1.3 m/s aisle speed, the \(C^2\) suffix peaks at the 2 m/s² credible deceleration, lasts 975 ms, and covers 633.75 mm, the braking term already reserved in the stopping budget of section 1.5. The kinetic energy the stop removes regenerates through the inverter into the DC bus (Electrical Power Integrity), so bus capacitance and a brake chopper must absorb the surge without an overvoltage fault.13 Power removal, if specified, follows the measured standstill condition.
A separate last-resort drive function can inhibit motor-generated torque. Drive Safe Torque Off (STO) does not itself disconnect the DC bus, apply a mechanical brake, or stop a moving motor; the motor may coast. In the illustrated architecture, an external safety controller may also command a separately rated contactor and spring-applied brake. Their response and stopping ability must be measured as one plant-specific path, because removing torque while a gravity-loaded arm or moving vehicle still depends on powered braking can worsen the hazard. On the warehouse mobile manipulator that path, STO plus the nine spring brakes setting 30 ms to 80 ms after coil decay, is the inhibit rung, and it is unbudgeted. It answers only the loss of the MCU’s own cycle, and no stopping or release claim rests on it.
Loss of supply does not route there. The permission path runs on an isolated rail held up for 2 s (Which Budget Binds First), so a sag on the shared control rail arrives as an undervoltage flag, and the enforcer answers it with the stop rung. The hold-up outlasts the longest stop the enforcer can command inside the envelope. Stop onset follows the flag within 22 ms, the resident stop from the 1.09 m/s ceiling on a floor inspected to a friction of 0.12 takes 1395.4 ms, and the brakes set at standstill within 80 ms, so every brake is set 1497.35 ms after the flag, with 502.65 ms of hold-up to spare.
↰ Prerequisite: Hardware trip zones and low-side MOSFET dynamic shunt braking are introduced in Actuation Authority.
Physical AI systems engineers must distinguish between fail-safe and fail-operational plants when configuring terminal fallbacks. In a machine already at a mechanically restrained standstill, removing drive torque may be safe after the restraint has been verified. In underactuated, open-loop unstable, or high-momentum machines, however, suddenly de-energizing actuators is itself a hazard:
- Aerial Platforms (Multirotors and eVTOLs): Cutting motor torque drops aerodynamic thrust to zero, transforming an airborne robot into a ballistic projectile in uncontrolled free fall.
- Dynamic Legged Systems (Bipedal Humanoids): During dynamic locomotion at \(1.5\text{ m/s}\), the center of mass is outside the base of support. Abruptly removing the torque needed for balance can cause a fall; the direction and consequence depend on the gait phase and available restraint.
- High-Speed Autonomous Vehicles: Removing powered steering or braking at highway speed can eliminate the authority needed for a controlled stop; the outcome depends on vehicle dynamics and backup actuation.
For such dynamic systems, the terminal fallback cannot simply be an uncoordinated power cut. Instead, the permission path must execute a deterministic Minimal Risk Maneuver (MRM) running on an independent, qualified controller with the required sensing and actuator authority. For aerial and legged robots, any controlled descent or balance recovery must be validated for remaining propulsion, contact, energy, and state-estimate authority. A road vehicle needs a risk-assessed powered braking and steering path. The wiring that carries a torque-removal request is the out-of-band interlock of Actuation Authority (figure 5), which keeps the spring-applied brake on its own path because removing torque and arresting a moving load are different functions carrying different validation evidence.
The architectural necessity of independent, hardware-level trip circuits is demonstrated by the historical consequences of eliminating physical interlocks in favor of software decision logic (1.1).
War Story 1.1: Therac-25 software interlock elimination
Mechanism: In preceding models (Therac-6 and Therac-20), physical safety was guaranteed by independent hardware interlocks: electromechanical microswitches, mechanical turntable detents, and analog diode steering circuits that physically prevented the accelerator from firing high-current beams unless the tungsten beam-flattening target was mechanically locked in place. In the Therac-25, engineers removed all physical hardware interlocks to reduce manufacturing costs, delegating exclusive safety enforcement to software. Concurrently, two latent software bugs defeated the software checks: an 8-bit shared variable (Class3) rolled over from 255 to 0 (\(0\text{xFF} \to 0\text{x00}\)) during setup, bypassing interlock evaluation; and a race condition allowed fast terminal keystrokes (< 8 seconds) to update operating modes without reconciling the physical turntable position with the commanded high-current beam energy.
Impact: Between 1985 and 1987, six patients received massive radiation overdoses up to \(25{,}000\text{ rad}\) (more than 100 times the intended therapeutic dose), resulting in catastrophic radiation burns, severe physical trauma, and several patient fatalities.
Response: Regulatory agencies halted machine operations, prompting comprehensive investigations that led to modern medical device software engineering standards, mandatory independent hardware safety interlocks, and formal verification of safety-critical concurrent states.
Systems lesson: Software-only safety filters executing on complex, non-deterministic compute stacks cannot replace independent bare-metal hardware trip zones. The Therac-25 disaster is the canonical historical proof that removing physical interlocks in favor of software decision logic transforms transient concurrency bugs and arithmetic overflows directly into catastrophic physical energy releases. In physical AI, software Control Barrier Functions and runtime projection shields are advisory safety filters; the terminal inhibit rung must always terminate in device-qualified hardware interlocks such as Safe Torque Off and mechanical brakes.
The rungs trade intervention time, available control authority, and recovery cost. A controlled stop requires reserved clearance and a working drive, and each rung’s plant-specific response must be validated for the states from which it can be entered. Figure 6 arranges the four rungs by the fault condition that selects each one.
Definition 1.2: Dynamic fallback ladder
Dynamic fallback ladder is the set of resident protective responses (\(u \in \{\mathcal{P}_{\mathcal{C}}(u_{\text{nom}}), u_{\text{hold}}, u_{\text{brake}}, u_{\text{STO}}\}\)) from which a deterministic arbiter selects one on every tick according to the check that failed (an invariant violation, a timing overrun, or a hardware fault) and the measured state of the plant, so that one trigger can select different rungs in motion and at standstill.
- Significance: Physical AI systems cannot treat all faults identically; cutting actuator power on dynamic, underactuated, or high-momentum plants (e.g., legged robots at speed, multirotors in flight) can cause a fall or tumble, requiring state-matched deceleration before terminal torque removal.
- Distinction: Unlike software exception handlers that unwind the call stack or kill a process thread, a dynamic fallback ladder operates against physical inertia, matching each rung to the remaining valid sensing, control authority, and reserved stopping clearance.
- Common pitfall: Reading the rungs as an escalation in which a lower one, such as Safe Torque Off (STO), provides faster or safer arrest. Each rung has its own entry conditions, and the inhibit rung removes torque without arresting anything by itself.
War Story 1.2: Cruise post-impact fallback (2023)
Mechanism: The Cruise vehicle initially braked and stopped. Its system then misclassified the event as a side impact and, having tracked the pedestrian only intermittently in the seconds before impact, initiated a pullover about \(1.83\text{ s}\) after contact. The vehicle dragged the pedestrian roughly \(20\text{ ft}\) at a maximum reported speed of \(7.7\text{ mph}\).
Response: An independent impact and underbody-occupancy response should revoke ordinary policy authority, retain braking and steering needed for a controlled stop, and hold the vehicle once stationary.
Systems lesson: A post-contact maneuver cannot infer clear space from a missing tracked object. The fallback must use independent evidence and preserve the ability to brake and hold.
The IEC 60204-1 stop categories name drive functions, not plant responses, and table 3 separates the function each category names from the physical response that must still be tested.
| Response | Motor torque and power | Physical stop obligation | Appropriate use |
|---|---|---|---|
| Admitted filtering | Drive remains active | Recheck the feasible kinetic barrier and tracking bound every tick | Nominal operation inside the validated envelope |
| Controlled hold (Category 2 / SS2-like) | Power remains for braking and hold | Verify stopping distance, thermal headroom, and static load support | Hold when feedback and torque authority remain available |
| Controlled stop then torque removal (Category 1 / SS1-like) | Power remains through deceleration; torque is removed afterward | Verify the full braking and transition time under load | Moving machine with enough reserved clearance |
| Torque removal (Category 0 / STO-like) | Drive cannot produce commanded torque; DC bus isolation is a separate function | A moving motor may coast; a separate brake or mechanical restraint needs its own timing and rating | State for which loss of torque has been shown safe |
Checkpoint 1.2: Choosing the physical fallback
Before selecting a defensive fallback strategy, verify your understanding of state-dependent stopping modes, safety functions, and timing reserves:
How Enforcement Goes Wrong
Every rung of the fallback ladder assumes that the hardware executing it stays independent of the fault that selected it. A secondary processor beside the primary compute module looks independent, yet a shared regulator, crystal, or ground plane can reset, halt, or corrupt both chips in the same instant, as when a host-side short drags the rail below the microcontroller’s threshold or motor-driver return currents bounce the ground beneath its buses. Software on either chip cannot see these faults, which is why the permission path keeps its own supply, clock, and ground (Hardware Isolation).
An isolated enforcer still fails when its computation falls out of phase with the plant, because a solve that overruns its tick acts on a state the machine has already left, and in an underdamped assembly a braking torque computed for that earlier state can feed a structural resonance instead of damping it. A watchdog that aborts the late cycle cannot recall the motion accrued during the overrun, so the timing claim rests on the declared WCET and on the loaded measurements of Testing the Separation.
The enforcer sees the plant only through state estimates, so filter group delay, transmission compliance, encoder slip, and thermal bias drift can hold the reported state inside its limits while the physical link has crossed them. Cross-checks against independent encoders or torque sensors catch large discrepancies, but slow sub-threshold drift stays hidden until contact.
Determinism guarantees that a computation finishes in a fixed number of cycles with bit-identical outputs, but it says nothing about whether the equations match the plant. An enforcer that assumes the base’s friction premise of 0.204 on a floor that oil has made slicker plays the resident stop exactly on schedule while the tires cannot deliver its deceleration, so the base runs past the rack end with every deadline met. No timing monitor can see that failure, and A Residual-Claims Register keeps its friction premise open.
Contention on the shared system-on-chip (SoC) interconnect can defeat an enforcer that meets its own deadline, because a camera direct memory access (DMA) burst or bulk flash write that holds the bus matrix after the solve adds directly to the command’s age at the drive while the inverter holds its last setpoint. The machine’s placement keeps the permission path off that bus (Contention for Shared Resources), so a burst on the application processor reaches it only as evidence age, and cycle stage 2 refuses any proposal the burst has aged past the refusal threshold.
An enforcer hosted as a real-time operating system task rather than a bare-metal interrupt adds priority inversion at the operating-system layer, where a low-priority telemetry thread holds a mutex the 1 kHz enforcer needs and a medium-priority thread keeps the holder from running, so the check waits without bound. The machine’s timer interrupt and lock-free mailbox exclude that case by design. On a task-hosted enforcer, priority inheritance (PTHREAD_PRIO_INHERIT) or ceiling locking bounds the wait, and Moving Commands on Time examines the failure in flight hardware.
The Therac-25 radiotherapy accelerator shows a concurrency failure with nothing independent behind it. Where its predecessors had independent protective circuits and mechanical interlocks, the Therac-25 relied chiefly on software running on the same computer that ran the treatment, where a race in operator data entry and a one-byte counter that rolled over to zero each let the machine fire a high-current beam without its target in place (Leveson and Turner 1993). Between 1985 and 1987 six patients received massive overdoses, and several died.
Every failure mode in the enforcement layer traces back to a breakdown of independence or a divergence between the internal plant model and physical reality. A barrier certificate, a dual-core package, and a zero-jitter scheduler establish safety only within the envelope where supplies stay isolated, estimates match the mechanics, and deadlines are met, so an enforcer that is to be audited against these failures has to state that envelope in advance: what it assumed, what it checked on each tick, and which response each failed check selects.
The Enforcement Record
When the warehouse mobile manipulator’s base refuses a proposal at the rack end, a later reader of the log must be able to say which threshold failed, which detector measured it, and which rung it selected. The enforcement record is the specification contract and runtime audit ledger under which, on every tick, the drive receives either the projection of \(u_{\text{nom}}\) onto the safe set or a fallback-ladder response, and nothing else.
The record consumes four predecessors, drawn from the two strands of records that meet here. The offline strand records what the plant can do and what the proposer has shown it can do, and the runtime strand records what the proposer proposes now. The trajectory record of The Trajectory Contract supplies the admitted reference and its matched stop suffix. The evaluation record of Evaluation Logs supplies the runtime monitor specifications and the coverage-gap catalog, which become the record’s monitor channels. The policy manifest of The Policy Manifest supplies the declared ODD, and the validity region stays inside the sampled part of it. The limit records of Measuring a Machine's Own Limits supply the braking, onset, and actuator limits with their conditions. To these the enforcement record adds what the permission path guarantees: the verified tracking bound \(\epsilon_{\text{track}}\), the stopping time and distance (\(t_{\text{stop}}, d_{\text{stop}}\), bounding the duration and displacement required to arrest plant momentum), Control Barrier Function coefficients, the loop period, per-check thresholds with their fallback-ladder rungs, and the validity region.
↳ Downstream: High-rate enforcer refusal events populate the forensic evidence logs analyzed in The Authority Log.
Filled in for the machine, the record fixes the base’s stops at the rack end and the arm’s validity region and channels at the door and the packing station:
- The tracking bound \(\epsilon_{\text{track}}\) is 40 mm, the inset of row 7 in table 1.
- The resident stop is the \(C^2\) suffix, which lasts 975 ms and covers 633.75 mm from the 1.3 m/s aisle speed, or 1125 ms and 843.75 mm from the 1.5 m/s drive limit.
- The loop period is 1 ms, with a declared WCET of 250 \(\mu\text{s}\) inside a 400 \(\mu\text{s}\) deadline. Three timers run beside it, the 60 ms chunk lease, the 100 \(\text{Hz}\) host heartbeat with its 30 ms silence timeout, and the 5 ms self-watchdog.
- The validity region for the arm covers the door and the packing station. At the door it is the target ODD of Evaluation Logs, which admits only the nominal plate, the door’s own load, a stopped base, and the guarded approach. At the packing station it admits the mug pick from the conveyor under a live intent lease (The Intent Lease), and the handover that follows.
- The arm’s channels are ceilings the enforcer applies to the tool point. The approach-speed channel of the evaluation record’s
runtime_monitor_specscaps it at 0.03 m/s inside the door zone. The mug pick runs under the 1 m/s free-space TCP limit. While the coworker holdsaccept_item(Human Authority Over Actuation), a handover channel caps the approach to the coworker’s hand at 0.10 m/s, below the 0.118 m/s contact-force ceiling, and it engages before the hand comes within the arm’s biased stop plus its blind travel. - For the base, the normal configuration is valid on floors with friction of at least 0.204, where the credible deceleration binds. A restricted configuration is valid down to the inspected-floor friction of 0.12; its suffix is sized for the 1.177 m/s² peak that floor allows, and it takes 1274.6 ms and 637.3 mm to stop from 1 m/s, under the 1.09 m/s inspected-floor ceiling. Each configuration carries its own digest, and the release manifest (The Release Manifest) names both.
Each numerical entry in the record carries the evidence category of its limit record (Measuring a Machine's Own Limits), whether measured, derived, datasheet, or illustrative, and the categories carry very different real-world confidence. The base’s 975 ms stopping time is derived. It follows from the shape of the \(C^2\) suffix and from a credible deceleration that is itself illustrative, and it assumes that the drive delivers the commanded current and that the tires hold the floor throughout the stop. Only loaded stopping trials on the floor itself make it a measured value, and those trials expose what no derivation contains, among them the scatter in brake onset and the friction of the actual floor surface. Even then the measured sample stands for the bound only with a justified worst-case fault and stated operating conditions. Recording the category prevents an integrator from mistaking a derived stop for a measured hardware limit.
The record maps separately qualified thresholds to protective actions, and on the machine the routing depends on whether the body is moving. Routing does not rely on coarse error strings or unquantified health flags. The base stays on the project rung, where the CBF-QP may adjust commands, as long as tracking error satisfies \(e(t) \le \epsilon_{\text{track}} =\) 40 mm, the kinetic barrier stays feasible, and the solver stays within its declared 150 \(\mu\text{s}\) share of the cycle budget. A new proposal that fails validation while the active lease holds is refused, and the previously admitted setpoints stay in force until that lease lapses. In motion, a lapsed lease, a silent heartbeat, an undervoltage flag on the shared control rail, tracking error beyond \(\epsilon_{\text{track}}\), saturation that empties the feasible control set, or a tripped operating-contract detector (the disturbance observer or the solver deadline) selects the stop rung, which plays the resident stop on the base and the state-matched stop on the arm. At standstill the same lease or heartbeat expiry, or the same undervoltage flag, selects the hold rung, since no motion remains to arrest. The 5 ms MCU self-watchdog detects missed successful cycles; this design also assumes an independently wired supervisor output or external trip that requests the inhibit rung if the MCU fails. Its total trip response must be measured. A contactor, if fitted, is an independently timed supply-isolation function.
Each field of the record is either an integrity field or a threshold that one of these checks reads: the proposal’s sequence number and checksum, its evidence age against the 51.8 ms refusal threshold (the budgeted observation age plus the clock-conversion bound), its own age against the lease, the tracking bound, the joint-velocity and torque-slew ceilings, the barrier margins, the phase-current and bus-voltage limits, and the watchdog timeout. The record binds each threshold to the independent hardware that measures it and to the rung that a failed check selects. Table 4 groups those rungs by the authority path that remains active after a failed check, together with the timer that triggers each one.
↳ Byte layout: The enforcement record’s fields, types, and units are given in Enforcement record.
| State | Trigger and authority | Timing check | Physical response to validate |
|---|---|---|---|
PROJECT |
Fresh proposal, healthy heartbeat, feasible kinetic barrier; a proposal that fails validation under an unexpired lease is refused and the admitted setpoints stay in force | \(1\text{ ms}\) cycle; declared full-cycle WCET 250 \(\mu\text{s}\) inside a 400 \(\mu\text{s}\) deadline | Track admitted torque within qualified error and brake reserve |
HOLD |
At standstill, proposal age above the 60 ms lease, host heartbeat silent beyond 30 ms, or shared-rail undervoltage | Expiry detected on the MCU clock within one tick | Powered position hold against gravity and payload; resume only on a fresh admitted proposal |
STOP |
In motion, lease or heartbeat expiry, shared-rail undervoltage, tracking error beyond \(\epsilon_{\text{track}}\), an empty feasible control set, or a tripped disturbance observer or solver deadline, while powered braking remains available | Include detection, dispatch, drive rise, and brake response | Resident stop on the base, state-matched stop on the arm; measured stop distance under load; no claim after viability is lost |
INHIBIT |
Independent wired supervisor or external trip detects MCU cycle loss | Assumed 5 ms detection plus measured drive, contactor, or brake response | STO removes motor torque and the spring brakes set after coil decay; unbudgeted, so no stopping claim rests on it |
No enforcement guarantee holds when the machine operates outside its designed physical environment, so the record bounds the plant dynamics within an explicit validity region. The arm’s door ceiling rests on evidence. The arithmetic alone would have admitted the chunk policy’s 0.10 m/s approach of section 1.3, whose 2.5 ms contact deadline still exceeds the 2 ms contact loop, but neither the demonstrations nor the evaluation trials behind the policy include that speed (Evaluation Logs), so the enforcer projects the proposal to 0.03 m/s and logs the channel that bound it. The handover ceiling rests on the contact model instead, sitting below the contact-force ceiling, the speed at which an unplanned touch of the coworker’s hand would reach the contact criterion (Human Authority Over Actuation), and a faster proposal is projected and logged in the same way.
A live state outside the validity region that the machine’s estimators detect, such as a payload the manifest never declared, is refused like any other failed check, and the supervisory monitors that flag it revoke policy dispatch and select the state-matched stop. A floor below the base’s friction premise is not such a state, because no sensor on the machine observes the friction under its wheels. The normal configuration therefore rests on that friction as a premise, and the restricted configuration rests on inspection of the floor (A Residual-Claims Register).
An enforcement record connects proposed motion to validated assumptions, thresholds, and observed fallback response. The execution platform must still meet its timing and physical braking contract under load. An enforcer placed on a shared memory bus can be mathematically verified and still stall when that bus saturates during a high-speed transit maneuver, and then the failure occurs in time rather than in logic.
Fallacies and Pitfalls
Physical enforcement stops a machine before a collision only while its execution is timely and its state feedback usable. Fault tests therefore exercise the response when either assumption fails, not the filter’s nominal equations alone.
Pitfall: Assuming high-level policy computers can be relied upon to execute graceful shutdown.
The warehouse mobile manipulator’s application processor suffers a kernel panic while the base runs at 1.3 m/s. The failed host cannot generate a deceleration trajectory, and dropping the drive enable to zero torque would leave the base coasting toward the rack end. The MCU instead detects the silent heartbeat within its 30 ms timeout, well inside the lease path the budget reserves, and plays the resident \(C^2\) suffix, which stops the base in 975 ms over 633.75 mm. A host shutdown handler may coordinate an orderly stop in normal operation, but protection against host failure requires a stopping path, clearance, and torque authority that remain after the host process disappears.
Fallacy: A mathematically correct safety filter guarantees safety even when underlying state feedback has drifted.
The base’s tracking bound of 40 mm limits the gap between the reference and the state its wheel encoders report, not the gap between that state and the floor. Wheel slip or a biased encoder can make the base appear farther from the rack end than it is, and the CBF-QP solver then satisfies its inequality on the drifted estimate while the physical base spends clearance the budget never granted. Only an independent channel, such as the IMU or the safety scanner’s range to the rack, exposes the divergence, and only while the fault is observable through it. Qualification must bound how far the estimate can drift before detection, and the budget must carry that error in the localization allowance, because the tracking inset cannot absorb it.
Pitfall: Coupling safety braking authority to stochastic semantic classification rather than deterministic spatial clearance.
When braking depends on upstream classification and path prediction, a detected object may still fail to trigger an effective stop. In the March 2018 Tempe collision (figure 7), the developmental automated driving system first detected the pedestrian \(5.6\text{ s}\) before impact but did not predict her path. It recognized an imminent collision \(1.2\text{ s}\) before impact, suppressed action for one second, then planned gradual slowing rather than emergency braking; the NTSB does not establish a Kalman-filter reset on each classification change. The vehicle’s original Volvo collision-warning and emergency-braking functions were disabled during ADS operation (When the Boundary Fails). An independent geometric clearance check needs its own validated sensing, stopping authority, and timing. It refuses on distance and closing speed, whatever label the object carries.
Summary
A physical AI system assigns actuator permission to a path independent of learned proposals. A kinetic barrier can help preserve a validated viable set when state error, model mismatch, actuation, and response delay stay inside their declared bounds. If braking authority or clearance is lost, the enforcer can revoke the proposal but cannot claim that a later fallback restores the original safe set.
The physical clearance protecting a machine is never a static geometric measurement. It is a composite margin that compounds pre-brake delay, braking under the credible deceleration, the localization and protective allowances, and the tracking bound. The warehouse mobile manipulator at its 1.3 m/s aisle speed shows the sum. Blind travel over \(\tau_{\text{delay}}\) takes 173.68 mm, the \(C^2\) stop suffix 633.75 mm, the localization bound and protective clearance 150 mm, and the tracking inset 40 mm. Together they need 997.43 mm of the 1.10 m clear distance at the rack end and leave 102.57 mm. A budget that counted only constant-deceleration braking would have admitted the drive limit, where the finished budget runs past the clear distance.
A planner checks a future trajectory; a local kinetic barrier checks the current proposal against the modeled ability to stop. For torque-controlled motion, raw clearance has relative degree two. The velocity-dependent barrier in equation 3 produces the affine torque condition in equation 4 for positive approach speed. This local calculation still depends on a reserved stop, accurate state, and timely torque delivery.
The permission path needs an independent response when the proposal source stalls. On the warehouse mobile manipulator, a 60 ms chunk lease, a 30 ms host-heartbeat timeout, and a 5 ms MCU self-watchdog detect different failures. The MCU revokes proposal authority and chooses a prevalidated state-matched response; torque removal answers only the loss of the MCU’s own cycle, since every other trigger the record routes selects a budgeted stop in motion or a hold at standstill.
The enforcement record is where the two strands of records meet. Each entry keeps its evidence category, so the record states what the permission path assumed as plainly as what it checked.
Key Takeaways: Conditional enforcement and physical stopping
- Tracking is a bounded premise: A steady-state error does not bound the transient, so a clearance inset needs a transient bound validated over the full maneuver in the operating envelope.
- The finished budget sets the speed: Blind travel, braking under the resident stop, transient tracking error, and position-estimate error all enter one admission check, so the clear distance, not the drive limit, decides how fast the machine may move.
- The controlled barrier includes speed: Raw distance has relative degree two for torque. A kinetic stopping barrier yields a torque half-space for positive approach speed, within the barrier’s operating contract.
- Fallback is state dependent: Feasible projection, powered hold, controlled braking, and drive torque inhibit have different prerequisites. An empty admissible set cannot guarantee a safe stop.
- Independent clocks detect different failures: Evidence age, proposal age, host heartbeat, MCU cycle progress, and external trips are checked separately; no one timeout is a universal braking guarantee.
- Refusal must leave a record: The enforcement record binds every threshold to the detector that measures it and the rung a failed check selects, so the premises of each guarantee can be tested on silicon and judged at release.
The fallback ladder turns the law that proposal is not permission into a check the MCU executes on every tick, since each failed check selects a rung reserved before the fault, and the proposer can neither delay that selection nor override it. The kinetic barrier ties this law to irreversibility by making permission depend on speed as well as position, because a state from which the resident stop no longer fits is lost before any boundary is touched.
What’s Next: From barrier certificates to heterogeneous SoC placement
Footnotes
Control barrier functions and forward invariance: Control Barrier Functions define a continuously differentiable certificate \(h(\mathbf{x}) \ge 0\) over an admissible safe set \(\mathcal{C}\), enforcing the directional clearance derivative constraint \(\dot{h}(\mathbf{x}, \mathbf{u}) \ge -\gamma(h(\mathbf{x}))\). Enforcing this inequality renders \(\mathcal{C}\) forward-invariant: states initialized in a viable set can remain there while the barrier is feasible and the stated model and timing assumptions hold. When approaching the boundary, the constraint forces opposing actuation inputs that arrest momentum before physical contact can occur.↩︎
Active-set quadratic programming for safety shields: On a bare-metal microcontroller, an active-set solver that restricts its linear constraints to active boundary contacts and allocates no dynamic memory can bound termination with a fixed iteration limit on a tested platform; no universal \(150\,\mu\text{s}\) limit follows from the solver type. If numerical ill-conditioning prevents convergence before the loop deadline, the solver aborts gracefully to an autonomous deceleration fallback.↩︎
Closed-loop tracking stiffness and disturbance bounds: Sustained force and effective closed-loop stiffness determine static deflection. Transient tracking also depends on inertia, damping, command shape, and saturation. Disturbances outside the validated envelope invalidate the claimed tube; they need not produce exponential divergence.↩︎
Remote Processor Messaging: Linux RPMsg provides virtio-based messaging between processors. Buffer ownership, copies, scheduling delay, and worst-case handover time depend on the platform; OpenAMP offers a separate no-copy API. The lease requires a measured, bounded delivery path.↩︎
Hardware break circuits and PWM trip zones: A configured external timer break input such as
TIMx_BKINcan force PWM outputs to a specified state without waiting for an interrupt.EPWM_TZFRCis a software force mechanism, not the external trip input. Output state and response time depend on the timer, gate drive, wiring, and qualification tests; the trip still needs a physically suitable stop path.↩︎Kinematic stopping distance and latency components: The inset stopping distance in this model compounds five terms: blind travel, braking distance under the credible deceleration, the localization bound, the fixed protective clearance, and the transient tracking bound. Omitting computational transport delay or tracking error results in premature encroachment on physical obstacles before peak braking can engage. See Kinematic Stopping Envelopes and Information Age Lag for formal kinematic derivations.↩︎
Actuator bias torque and stopping margin loss: When an actuator allocates continuous baseline torque to counter gravitational loading or steady-state friction, available emergency braking headroom collapses proportionally. In vertical and heavily loaded axes, this depletion of net dynamic torque causes stopping distances to quadruple under seemingly modest bias loads. For a detailed derivation of saturated authority and stopping margin contraction, see Safety and Control Barrier Functions.↩︎
Nagumo viability theorem and boundary tangent cones: Nagumo’s theorem establishes that a closed set is forward-invariant if and only if the system velocity vector belongs to the contingent cone of the set at every boundary point. Control Barrier Functions synthesize this geometric condition into real-time affine inequality constraints on actuator torque. Violating this subtangentiality condition causes immediate set exit, rendering future boundary enforcement unachievable; see Safety and Control Barrier Functions.↩︎
Relative degree and inertia-aware barrier certificates: When a safety barrier monitors position while the actuator commands torque, the control input is separated from the constraint by two time integrals (\(r=2\)). Enforcing a zero-order spatial barrier without embedding kinetic stopping distance produces an optimization problem with zero control authority on the boundary. A velocity-dependent barrier can support a conditional forward-invariance argument while braking remains feasible and the state, delay, and tracking margins hold; see Safety and Control Barrier Functions.↩︎
Worst-case execution time and active-set QP complexity: On bare-metal safety microcontrollers, QP safety filters execute within static memory arrays using fixed iteration limits. While typical active-set solve times range between \(60\text{--}120\ \mu\text{s}\), an uncapped solve can run to \(600\ \mu\text{s}\) under ill-conditioned or nearly degenerate constraint geometries, longer than the whole permission deadline. The fixed iteration cap exists to rule out that case. It bounds the solve to its declared share of the cycle, so the timing budget covers the capped solver rather than the uncapped worst case.↩︎
Control barrier functions and Lie derivatives: In nonlinear control theory, the unforced drift and actuator control authority are formalized using Lie derivatives (\(L_f h(x)\) and \(L_g h(x)\)). When the safety boundary has relative degree one, the control input appears directly in the first time derivative of the barrier function, yielding an instantaneous affine constraint on commanded torque. If the relative degree is two, the enforcer must evaluate higher-order derivatives to prevent actuator saturation from rendering boundary violations inevitable; see Safety and Control Barrier Functions.↩︎
Machine stop functions: IEC 60204-1 distinguishes uncontrolled power removal (Category 0), controlled stopping followed by power removal (Category 1), and controlled stopping with power retained (Category 2). IEC 61800-5-2 drive functions such as STO, SS1, and SS2 may implement parts of those responses. The required performance level, circuit architecture, and suitable stop are determined by the machine-specific risk assessment; ISO 13849-1 does not impose one universal PL or channel count.↩︎
Dynamic brake choppers and DC-bus overvoltage clamping: In a regenerative Category 1 stop, mechanical kinetic energy flows through the inverter bridge into DC bus capacitors (\(\Delta E = \frac{1}{2} m v^2\)). A suitably rated chopper and resistor can dissipate energy when the DC bus rises; switching thresholds, pulse-energy rating, and thermal behavior must be validated for the stop profile. Other designs use mechanical dissipation or a battery that can accept regeneration.↩︎



