The Nervous System
The Nervous System
Purpose
What forces an independent Nervous System to stand between deliberative neural policies and physical actuator power?
A moving machine’s kinetic momentum cannot pause while a neural policy completes deliberation. Direct connections between high-capacity foundation models and motor power stages expose actuators to operating system scheduling jitter, garbage-collection stalls, and memory bus contention. While software threads await DMA transfers, mechanical linkages coast under unmanaged inertia, magnetic flux collapses, and dynamic balance degrades before a software runtime can raise an exception. Software task priorities cannot preempt hardware memory saturation or mechanical velocity; physical stability requires cycle-deterministic real-time control.
The Nervous System resolves this mismatch by decoupling deliberative planning from high-frequency actuator commutation. Standing as an independent, cycle-deterministic gatekeeper, it continuously verifies observation freshness and holds exclusive write authority over inverter power stages. By translating the physical plant’s stopping budget into rigid temporal leases and control barrier constraints, the Nervous System bounds how long any proposal may guide motion before expiring. If a neural policy delays, fails, or emits an unsafe trajectory, the Nervous System intervenes to arrest motion before the machine violates its safe envelope.
↰ Prerequisite: Silicon domain isolation physically enforces the causal boundary defined in The Causal Boundary.
Learning Objectives
- Determine control update rates from physical dynamics and allowable response delay
- Evaluate proposal exchanges for freshness and bounded transfer time between the learned proposer and the permission path
- Distinguish measured latency percentiles from defensible worst-case bounds on real-time execution
- Evaluate hardware isolation that reserves actuator authority when learned policies or communication fail
- Size a proposal’s lease between the proposer’s renewal period and the pre-brake time the stopping budget leaves
- Diagnose an interface failure as an over-budget limit at a specific handoff, using like-for-like evidence
- Construct handoff-record entries that give each interface limit its conditions, evidence category, and bound kind
The Multi-Rate Bridge
The warehouse mobile manipulator that this book follows regulates the current in its motor windings thousands of times per second, while the learned models that proposed its grasp on the conveyor mug in The Cognitive Brain need orders of magnitude longer to digest each camera frame. If deliberative reasoning and motor regulation share a blocking execution path or an unpartitioned memory bus, any neural network delay starves physical control. The physics of moving mass and electromagnetic coils create a strict architectural requirement: the Brain needs time to deliberate, while the Body needs continuous control to survive.
In standard computer systems, workloads share processors, caches, and memory interconnects under general-purpose operating systems. When a neural network streams multi-gigabyte weight tensors across a shared memory bus, it saturates direct memory access crossbars,1 and no thread priority can preempt a hardware transaction already occupying the bus. If a garbage-collection pause, page fault, or thermal throttle stalls the neural engine, an unpartitioned system stops sending setpoints to the motor inverters.
The Cognitive Brain established how long each learned proposer on the machine needs to turn camera frames into a proposal. An electromagnetic motor coil has an electrical time constant (\(\tau_e = L/R\)) measured in fractions of a millisecond, while moving linkages carry kinetic energy (\(E_k = \frac{1}{2} m v^2\)). Delaying or miscalculating the current command can therefore exhaust a physical margin before deliberation finishes. On the mobile manipulator, the mismatch takes two forms:
- The base (Class 1) cruises at up to its 1.5 m/s drive limit (illustrative; see the Reader Guide), so it travels 60 mm during one 40 ms chunk-policy inference and 240 mm during one 160 ms intent-model inference, and nothing brakes it in that time unless a process other than the proposer is watching.
- The arm (Class 2) carries reflected rotor inertia (\(N^2 J\)) through its gearboxes; if its high-rate impedance loops are interrupted while the gripper is on a door latch or a mug, that inertia drives the contact force past its limit.
Software discipline, OS thread priorities, and runtime exception handlers cannot by themselves keep the plant safe. Safety requires an explicit, deterministic computing, memory, and communication fabric. That fabric is the Nervous System of the five-level machine (The Machine in Five Levels), and it holds the permission path.
Every signal crossing between deliberation and physical actuation is constrained by five interface limits, which section 1.8 lays out handoff by handoff. Reconciling them requires an architecture that delivers microsecond cycle determinism (\(1\text{ kHz}\) to \(20\text{ kHz}\)) to preserve controller phase margin, maintains lock-free asynchronous decoupling to prevent slow neural inference from blocking inner control loops, isolates actuator writes from host crashes, restricts actuation privilege to the permission path, and requests a validated local response when a proposal’s chunk lease expires.
The first decision is where the two workloads execute. The separation in section 1.2 keeps inference off the timed actuator path. The design in section 1.3 then budgets their exchange and sizes the lease that ends a stale proposal’s authority; section 1.4 bounds the time a command takes to reach the drives. The remaining sections locate actuator authority, fault isolation, and invariant checks, and close Part I with the handoff record and one budget that couples this chapter’s lease to the limits of The Physical Body and The Cognitive Brain. Each mechanism must leave the physical plant a measured, feasible response when the proposer pauses.
Where the Permission Path Runs
Two regimes of computation meet in a physical AI system, and they have almost nothing in common. The first regime is deliberation, where learned models evaluate sensory context, infer environmental state, and propose future action plans. The second regime is reactive enforcement, where deterministic control logic samples joints, validates kinematic limits, interpolates motion commands, and regulates electrical currents to the motors.
Deliberative work features deep computational graphs, large memory footprints, and variable execution times. A vision-language-action model processing multi-camera video streams, tactile histories, and natural language prompts requires billions of multiply-accumulate operations across hundreds of neural network layers. Storing the parameters and intermediate activation tensors demands gigabytes of high-bandwidth memory. To achieve acceptable arithmetic throughput, deliberative computation executes on massively parallel accelerators such as graphics processing units (GPUs) or dedicated neural processing units (NPUs). These processors maximize aggregate throughput across wide vector registers and batched memory pipelines at the direct expense of deterministic latency. Memory access stalls, dynamic tensor kernel selection, bus transfers, and runtime memory management scatter inference execution times across a wide distribution. If a deliberative model experiences a temporary delay, the failure is soft, because the planning pipeline produces a delayed proposal or a stale trajectory that can be held or discarded while the permission path maintains stable equilibrium.
Reactive work operates under the inverse constraints. An inner control loop computing inverse dynamics, joint limit monitoring, or field-oriented motor control requires a few thousand arithmetic operations per update. Its memory footprint is static, measured in kilobytes rather than gigabytes, with every variable allocated at compile time in tightly-coupled on-chip static memory. Reactive routines execute on cycle-counted microcontrollers or field-programmable gate arrays (FPGAs) designed for bounded execution rather than maximum parallel throughput. The deadlines governing this work are hard real-time constraints dictated by the mechanical inertia and electrical inductance of the physical system, typically running at rates between \(1\text{ kHz}\) and \(10\text{ kHz}\) on update budgets of \(100\,\mu\text{s}\) to \(1000\,\mu\text{s}\). If a reactive loop misses its execution deadline, the failure is hard, because lost current regulation induces severe torque ripple, actuator saturation, or mechanical collision.
Learned inference runtimes depend on memory access patterns, cache residency, dynamic tensor shapes, and hardware prefetching, making execution times stochastic. An inference pipeline can meet its average deadline while still missing individual deadlines in its latency tail. At tens of requests per second, even a sub-percent exceedance probability yields many misses an hour, each an interval in which the motor actuators receive no updated command unless an independent, deterministic process bridges the gap.2
General-purpose operating systems and real-time kernel patches (PREEMPT_RT) cannot reconcile these two execution profiles on a shared processor. Although real-time extensions allow static task priorities and preemptible kernels, they cannot eliminate shared microarchitectural resources. When a thread executing a learned model streams gigabytes of weight matrices across the memory bus, it evicts cache lines belonging to high-priority control tasks. If the deliberative runtime triggers a translation lookaside buffer (TLB) shootdown,3 enters a driver-level spinlock, or saturates the memory bus crossbar with direct memory access (DMA) requests, the hardware pipeline stalls.4 A high-priority real-time interrupt can preempt software execution, but it cannot preempt an in-flight hardware burst transaction already occupying the shared memory crossbar. Furthermore, dynamic voltage and frequency scaling (DVFS) governors adjust core clock frequencies based on thermal dissipation and processor utilization, altering instruction cycle times dynamically (Fallacies and Pitfalls prices the travel this costs). The resulting jitter ranges from tens of microseconds to several milliseconds, easily exceeding the entire time budget of a \(1\text{ kHz}\) control loop. Contention for Shared Resources measures how far each of these channels reaches on a shared die.
The interface hardware on each side makes the separation concrete. A host processor running an accelerator communicates across high-latency bus topologies such as PCIe or system fabric crossbars, where DMA transfers require ring-buffer synchronization, host interrupt service routines, and cache coherency handshakes that consume \(10\,\mu\text{s}\) to \(100\,\mu\text{s}\) per transaction. In contrast, an embedded microcontroller or dedicated real-time motor controller IC (figure 1) interfaces directly with sensor transceivers and power stages through memory-mapped registers, single-cycle input and output pins, and tightly-coupled static memory with zero wait states. Reading an absolute joint encoder over a serial peripheral interface (SPI) or updating a field-oriented pulse-width modulation counter takes less than \(1\,\mu\text{s}\) of bounded execution time,5 isolated from system bus contention.
In the two-processor implementation (figure 2), deliberative planning and reactive enforcement do not share a processor core, a cache hierarchy, or a common memory bus. The slow, stochastic path runs on parallel hardware (Linux host application processors and tensor accelerators) optimized for raw throughput across wide data streams, while the fast, deterministic path runs on isolated hardware (bare-metal microcontrollers or FPGA cores) optimized for worst-case execution time and cycle-counted predictability. On the mobile manipulator, the permission path runs on the safety microcontroller (MCU). Separate chips are one realization; a single die can also host the permission path if its response bound and fault containment are shown to hold (Two Paths on One Die).
Safety in physical AI is an architectural partition in silicon, not an operating system scheduling priority. The high-throughput deliberative engine acts strictly as an untrusted client, while dedicated real-time silicon acts as the hardware gatekeeper. The partition settles where each workload runs, but it leaves the two sides on separate clocks at rates that differ by orders of magnitude, and every proposal must still cross between them without either side waiting for the other.
Multi-Rate Cadences
The mobile manipulator’s drives regulate winding current every \(50\ \mu\text{s}\), while its chunk policy needs 40 ms at P99 to turn camera frames into one action chunk. No single execution frequency serves both the body’s physical dynamics and the cost of perception and inference, so a physical AI machine runs several coupled cadences (figure 3), each set by a different physical process. The mobile manipulator runs four cadences. The current loop in the drives runs at \(20\text{ kHz}\); the permission tick on the safety microcontroller runs at \(1\text{ kHz}\) and carries the joint loop and the evaluation of the admitted spline; the chunk policy runs at 20 Hz, fed by perception at its sensors’ rates; and the intent model revises goals at 5 Hz. These last two cadences lie within the Brain’s \(1\text{--}50\text{ Hz}\) inference rates (Learned Representations).
↰ Prerequisite: Microsecond interrupt latency requirements build upon the sensor freshness bounds introduced in The Five Physical Budgets.
The inner loop rate is governed by the physical time constants of the actuator and the structural resonance of the machine. For an actuator with electrical time constant \(\tau_e=L/R\), stator winding current responds to applied voltage over that interval. For an illustrative winding with \(L/R=0.5\text{ ms}\), choosing ten current samples per electrical time constant gives \(T_s=50\,\mu\text{s}\), or \(20\text{ kHz}\). For the winding \(L/R\) derivation, see Systems and Hardware. At the joint level, the mechanical time constant and structural stiffness dictate the required closed-loop bandwidth, and a discrete sample-and-hold loop introduces phase delay proportional to computation latency and sampling interval.6 At an assumed joint-loop crossover of \(20\text{ Hz}\), a \(1\text{ ms}\) sample-and-hold interval contributes \(\Delta\phi=-(2\pi\cdot20)(0.001)/2\approx-3.6^\circ\) of phase lag. If the allowed sampling erosion is \(5^\circ\), the interval must be below about \(1.39\text{ ms}\); \(1\text{ kHz}\) is one feasible choice before computation and transport delays are added.
Slow control loops erode the phase margin of the feedback controller. Accumulated phase lag causes corrective torques to arrive out of phase with physical link motion, converting negative feedback stabilization into positive feedback amplification. At the same \(20\text{ Hz}\) crossover, a \(500\text{ Hz}\) loop adds about \(7.2^\circ\) of sample-and-hold lag, above the \(5^\circ\) allowance.
Perception, which feeds the chunk policy, executes at an intermediate rate dictated by photon collection physics, exposure windows, and target velocity relative to the sensor field of view. In an illustrative \(300\text{ lux}\) scene with a 15 ms exposure, the base at its 1.5 m/s drive limit travels \(\Delta x = v\,T_{\text{exp}} =\) 22.5 mm during the exposure; image-plane blur additionally depends on depth, optics, and camera motion. Exposure signal-to-noise, this displacement, and processing time jointly bound a useful perception rate for the task.
The chunk-policy rate is chosen within the renewal latency that Supply, Freshness, and the Memory Wall measures on the application processor. On the mobile manipulator, a 20 Hz renewal and a 16-waypoint chunk, the chunk horizon \(H\) of action-chunking policies (Zhao et al. 2023), bridge observation and actuation rates.
Because these cadences execute on isolated hardware domains at different frequencies, data exchange across rate boundaries must be asynchronous, non-blocking, and zero-copy (figure 3). If a \(1\text{ kHz}\) control loop waits on a mutex lock held by a \(20\text{ Hz}\) neural network runtime, a single scheduling delay in the operating system immediately causes the real-time controller to miss its deadline.
A baseline nonblocking exchange uses a single-producer single-consumer (SPSC) seqlock mailbox in shared static memory.
Definition 1.1: Lock-free seqlock buffer
Lock-free seqlock buffer is a nonblocking shared-memory synchronization mechanism that couples an atomic sequence counter with explicit hardware memory barriers, allowing an unprivileged writer to publish multi-word trajectory payloads without blocking and enabling concurrent readers to verify payload integrity via counter parity checks: \[c_{\text{start}} \equiv c_{\text{end}} \pmod 2 \quad \text{and} \quad c_{\text{start}} = c_{\text{end}}\] where an odd counter value signals active writer mutation, and an identical even counter pair confirms an uncorrupted, race-free payload read.
- Significance: Decouples high-rate real-time control loops (\(1\text{ kHz}\)) from variable-latency neural policy processes (\(10\text{--}50\text{ Hz}\)) without mutual exclusion locks, preventing operating system priority inversion from causing hard real-time deadline misses.
- Distinction: Unlike mutexes or semaphores that suspend thread execution during contention, a seqlock reader never blocks; if a collision or torn read is detected, the reader retains its previously validated trajectory spline within a bounded execution budget (\(O(1)\) time).
- Common pitfall: Omitting bidirectional hardware memory barriers (
DMB OSH/DMB OSHLDon ARM) around payload accesses, which allows out-of-order processor pipelines to speculate or hoist memory loads ahead of the initial sequence counter validation.
The writer, the untrusted application processor, makes the counter odd, copies the complete proposal into the shared slot, makes the payload visible across the interconnect, and then makes the counter even with release ordering. The reader, the permission path, makes one bounded copy attempt per tick. It rejects the attempt if the counter was odd or changed during the copy, and otherwise promotes the copy to its active spline only after the checksum and freshness checks pass, so a rejected read leaves the active spline unchanged and the reader never waits for the writer. Because the writer is a preemptible user-space process rather than a kernel thread, a writer stopped mid-write can leave the counter odd indefinitely; the reader then keeps its last validated spline only until that spline’s lease expires. A three-slot publication variant, sketched in figure 3, keeps a complete published slot available while the writer stages another, at the cost of atomic slot publication and reader-slot ownership. The complete single-slot protocol, including the memory barriers a weakly ordered processor requires, appears in Real-Time Fieldbuses and Lock-Free Synchronization.
The mailbox’s own collisions are rare and cheap. A read collides only when it overlaps a write, and a write occupies the slot for microseconds once per chunk period. If the writer’s and reader’s clocks drift against each other, only a small fraction of reads collide; if they lock at an unfavorable phase, one read may collide with every write. Either way, a write far shorter than one reader tick cannot overlap two adjacent reads, so a detected odd or changed counter makes the reader discard one attempt and keep its last validated spline. A collision therefore costs one tick of freshness.
How long a spline’s lease may run is a question about the plant, not the mailbox. Under the irreversibility law (principle \(\ref{pri-vol4-irreversibility}\)), each millisecond a stalled proposal keeps its authority is travel that no later check can recall, so the lease must be paid for out of the stopping budget.
Napkin Math 1.1: Lease sizing from the stopping budget
1. What Kinetic Momentum left:
- Braking at the credible deceleration \(a_{\text{brake}} =\) 2 m/s² takes \(v^2/2a_{\text{brake}} =\) 562.5 mm, the braking term of equation; localization and protective clearance, \(\delta_{\text{loc}} + \delta_{\text{margin}}\), take 150 mm; brake onset, \(T_{\text{act}} =\) 20 ms, takes 30 mm.
- The stopping budget therefore stands at 742.5 mm of the 1.10 m clear distance.
2. The pre-brake ceiling:
- Everything that happens before braking force begins must fit in the time that the distance left by the braking term and the overheads allows at speed \(v\): \[\tau_{\text{delay}} \le \frac{D_{\text{clear}} - \delta_{\text{loc}} - \delta_{\text{margin}} - v^2/2a_{\text{brake}}}{v}.\] Here the right-hand side is 258.3 ms, and brake onset has already spent 20 ms of it.
3. The lease path (row 4 of the stopping budget):
- The stalled proposal keeps its authority for the chosen chunk lease, \(T_{\text{lease}} =\) 60 ms. The permission path notices the expiry on its next tick, \(T_{\text{tick}} =\) 1 ms, and the stop command reaches the drives one bus cycle later, \(T_{\text{bus}} =\) 1 ms. Each of the three enters \(\tau_{\text{delay}}\). The joint loop’s phase margin sets the tick; at 1.5 mm per millisecond the stopping budget can afford it.
- The row adds \(v\,(T_{\text{lease}} + T_{\text{tick}} + T_{\text{bus}})\), the distance covered in 62 ms at 1.5 m/s, which is 93 mm. The running total rises from 742.5 mm to 835.5 mm, leaving 264.5 mm.
- From the last valid renewal to brake onset, the machine spends 82 ms, and 176.3 ms of the ceiling remains for the terms that later chapters add.
- The lease has a floor as well as a ceiling. The chunk policy publishes a fresh chunk every 50 ms, and with a 5 ms allowance for renewal jitter any lease shorter than 55 ms would let a healthy proposer’s own chunks expire. The chosen 60 ms sits just above that floor.
Systems insight: A chunk lease is a term in the stopping budget, not a timeout chosen for convenience. Every millisecond added to it costs 1.5 mm of clear distance at the drive limit, distance that the observation age and the planner’s stop profile will need.
The worst case is a proposer that renews its lease on the frame before the person appears and then stalls while its heartbeat stays live. Nothing reads the frame showing the person, and the brake begins only when the lease expires, plus one tick, one bus cycle, and brake onset; the frame’s own age, which Sensor Perception measures, adds to that. The permission path can bound this case without seeing the person, because the enforcer, the permission path’s checking loop, reads the encoders, the IMU, and the evidence time stamped in the proposal’s header rather than a raw observation.
Three timers guard the exchange, and each answers a different question. The chunk lease bounds how long an admitted proposal may drive the actuators; only a fresh proposal renews it, and its length comes from the stopping budget, here the chosen 60 ms between the renewal floor and the pre-brake ceiling of 1.1. The heartbeat reports that the application processor is alive. It runs at 100 Hz with a 30 ms timeout, faster than the proposer, and its silence revokes the proposal without waiting for the lease. It never renews the chunk lease, because a live process can still publish nothing admissible, so the lease is what bounds a host that stays alive but stops publishing admissible chunks. The self-watchdog certifies that the permission path itself is completing its checks; the permission loop feeds it only after a completed tick, and it trips after 5 ms without a feed, so a stalled enforcer trips it even while its timer interrupts keep firing (section 1.6).
The lease is enforced through the proposal header, the part of every proposal that lets the permission path judge it without trusting its author. By meaning, the header carries a sequence number, the evidence time of the observation the proposal was computed from together with that time’s conversion error, the issue time, the expiry, a reference to the parent record the proposal was derived from, the coordinate frame, the payload kind, and a checksum. The minimal payload behind it, a chunk, carries the number of waypoints and, for each waypoint, joint positions and joint velocities. It carries no torques, because a learned proposal specifies kinematic setpoints (Learned Representations) and torque is what the tracker on the permission side computes from them. The proposal’s evidence deadline is its evidence time plus the largest evidence age the permission path accepts, which Sensor Perception budgets. A proposal drives the actuators only until the earlier of its expiry and its evidence deadline, whether or not any cancellation arrives.
↳ Byte layout: The proposal header’s and chunk payload’s fields, types, and offsets are given in Proposal header and chunk payload.
Every record that crosses the proposal boundary carries its times on the safety microcontroller’s clock. A producer on another clock converts with a versioned mapping and carries that mapping’s error bound, which the consumer adds to every age it checks. Calendar time appears only in logs and in a release verdict’s expiry, never in an age test.
A record crossing the boundary is protected against three different failures by three different fields. A checksum over header and payload detects accidental corruption; it proves nothing about who wrote the bytes. A digest of the record it was derived from establishes lineage; it is not an authenticator either. A keyed tag appears only where the sender’s identity itself confers authority: on operator commands that arrive over a network (Authenticating an Intervention), on the forensic log (Trustworthy Operational Logs), and on offline evidence and release artifacts (The Fault Manifest, The Release Manifest). Proposals from the learned model carry no keyed tag, because the permission path admits by content, not origin, and a key held by a compromised proposer would authenticate nothing.
The reader’s possible failures fall into four classes (table 1), and none of them lets a failed producer make the permission path wait. The classes keep the exchange from blocking; they do not by themselves contain the plant, because the fallback each one selects still needs bounded timing and a validated stop.
| Failure Class | Reader Detection Condition | Physical Systems Fallback Action | Bounded Response |
|---|---|---|---|
| Torn / In-Flight | Odd \(c_{\text{start}}\) or changed \(c_{\text{end}}\) | Reject current snapshot; use admissible prior chunk or fallback | One bounded attempt; measure duration |
| Malformed | Checksum or configured command bound fails | Reject proposal; use independently checked fallback | Measure check and response |
| Stale | Past proposal expiry or evidence deadline (local clock) | End tracking authority; enter the state-matched fallback | Deadline from plant budget |
| Missing | Sequence gap | Record audit gap; evaluate current record on its own merits | One tick for flag |
Rate mismatches introduce three failure modes: phase lag accumulation, aliasing, and buffer starvation. In a discrete-to-continuous conversion, a Zero-Order Hold (ZOH) maintains setpoints constant across the planning period \(T_{\text{plan}}\), delivering piecewise-constant position steps rather than a smooth trajectory. This step quantization introduces an intrinsic half-sample phase lag \(\Delta \phi = -\frac{\omega T_{\text{plan}}}{2}\). Directly feeding raw planner waypoints to inner motor loops at \(10\text{ Hz}\) incurs \(50\text{ ms}\) of uncompensated phase delay, degrading tracking margins and exciting unmodeled mechanical resonances. Across the latched waypoint seam, the real-time enforcer can use a quintic bridge that matches position, velocity, and acceleration at both ends. Those six endpoint conditions, not the tick rate, give the seam \(C^2\) continuity, and evaluating the admitted bridge at \(1\text{ kHz}\) supplies the corresponding velocity and feedforward acceleration setpoints.
↰ Prerequisite: High-rate joint interpolation of action chunks builds on the proposal generator in Supply, Freshness, and the Memory Wall.
Aliasing occurs when high-frequency sensor noise folds into low-rate estimators, which the system prevents through analog anti-aliasing filters and digital decimation7 within the high-rate microcontroller domain. Classical stability guarantees (Slotine and Li 1991) assume inner-loop execution jitter remains strictly bounded and setpoints arrive before the trajectory buffer drains. The permission path enforces these conditions through cycle-deterministic microcontroller interrupts and buffer depth tracking. If the planner stalls and the buffer reaches its final valid waypoint, the permission path rejects stale proposals and requests its validated local deceleration or hold.
An overrun means something different at each of the four cadences. The current loop runs in the drives on a jitter budget of microseconds, and a detected fault there trips torque in hardware. At the 1 ms permission tick, an overrun is a fault that selects the validated local stop, and a spline stays in force only within its lease and stopping margin. The chunk policy’s 40 ms inference time on the mobile manipulator is a P99, a statistical figure rather than a worst-case bound, and this cadence holds the 60 ms chunk lease only because it renews every 50 ms, inside the lease; an overrun that delays a renewal past the lease lets it lapse. When the intent model overruns, its targets expire, and the machine enters its state-matched fallback, holding or stopping from its current plant state.
Every timed cadence is a deadline, and the lease budget assumes that the tick, the bus cycle, and brake onset meet theirs on every cycle. That assumption needs a worst-case bound on every step between the enforcer’s check and the drives, and neither a faster processor nor a longer latency measurement supplies one.
Moving Commands on Time
Observed latency distributions do not establish worst-case timing bounds. In statistical performance analysis, an engineer measures ten million executions of a task, notes that the ninety-ninth percentile (\(P_{99}\)) latency is \(135\,\mu\text{s}\) and the maximum observed latency is \(150\,\mu\text{s}\), and infers that the computation fits comfortably inside a \(1.0\text{ ms}\) control deadline. The inference does not follow. An empirical distribution records what happened during finite testing, whereas a hard real-time bound must hold in every hardware state the analysis admits. When an execution platform includes out-of-order execution pipelines, shared cache hierarchies, speculative prefetchers, and concurrent direct memory access traffic, a finite latency sample cannot establish an upper bound. The missing tail events do not vanish; they emerge under rare microarchitectural collisions, such as concurrent DMA burst transfers and cache line evictions coinciding with a critical interrupt, that manifest under production stress.
Establishing a worst-case response time requires bounding every term that contributes to task response time: the base computation time, the blocking time spent waiting for shared critical sections, and the interference time from high-priority interrupts or memory bus contention.8 If any one of these three terms lacks an analyzed upper bound, the system has no bound on its worst-case response time, however small the mean latency appears in benchmarks.
↳ Downstream: Deterministic fieldbus jitter bounds provide the execution baseline for heterogeneous placement in Contention for Shared Resources.
Bounding those terms starts by sorting every source of execution variance by whether its worst case can be bounded at all. Bounded sources admit exact upper bounds, among them fixed-pipeline instructions, round-robin bus arbitration, zero-wait-state tightly coupled memory, and vectored interrupt dispatch on a bare-metal microcontroller. Experimentally characterized sources, such as last-level cache misses under line locking, write-buffer drains, DRAM refresh, and thermal throttling under active cooling, cannot be bounded from first principles but can be enveloped by stress profiling under a locked hardware configuration. Unbounded sources have no finite worst case. They include page faults under demand paging, preemptive thread scheduling without priority inheritance (war story 1.1), heap allocation over fragmented free lists, and unscheduled DMA traffic, and a hard real-time safety path must eliminate every one of them.
War Story 1.1: The Mars Pathfinder priority inversion (1997)
Mechanism: A high-priority thread (bc_dist) distributed critical spacecraft bus telemetry, while a low-priority thread (ASI/MET) gathered meteorological sensor data through a shared mutex-protected pipe. When ASI/MET held the mutex, medium-priority tasks preempted it. Because they outranked the meteorological thread, they kept it from running while bc_dist blocked waiting for the mutex it held, an unbounded priority inversion.
Impact: The high-priority bus distribution task missed its deadline. When the bc_sched task started the next bus cycle and found bc_dist unfinished, the flight software reset the computer. Data already collected survived each reset, but the resets recurred until the software was changed.
Response: Engineers reproduced the reset on a ground replica of the spacecraft with VxWorks event tracing, then uploaded a patch that enabled priority inheritance (Sha et al. 1990) on the mutex semaphore.
Systems lesson: The medium-priority tasks kept the low-priority mutex holder from running, so the high-priority bus task could not progress. Priority inheritance let the holder finish and release the mutex. In a motor-control path, engineers must bound such blocking time (\(B\)) or remove the shared lock from the cycle-critical exchange.
The same classification settles dynamic memory, which is not categorically forbidden on a hard timing path. A general-purpose heap allocator (malloc/free) that walks fragmented free lists has a search time that depends on allocation history, so it is unbounded. A fixed-size block pool or a Two-Level Segregated Fit (TLSF) allocator has constant-time \(\mathcal{O}(1)\) bounds and can sit on a real-time path. The strictest design for the permission path is a zero-allocation static architecture, in which every buffer, ring queue, and state variable is placed at compile time and no memory management runs at all.
The bus that carries the command contributes to the interference term \(I\) and the blocking term \(B\) of the worst-case response time. Hermann Kopetz draws the distinction that decides whether those terms have a bound, between event-triggered and time-triggered communication (Kopetz 2011). In event-triggered communication, such as standard CAN, nodes transmit when their local state changes, so bus load is unpredictable and a single malfunctioning node can flood the shared medium with spurious traffic, the “babbling idiot” failure. A time-triggered architecture (TTA) synchronizes a global clock and gives each node its own slot in a time-division schedule (TDMA), which bounds transmission interference once clock synchronization, fault behavior, and the full schedule are verified. Industrial robots implement this deterministic layer in dedicated fieldbus hardware (figure 4) that couples real-time controllers to I/O and drive terminals over Ethernet frames synchronized by hardware Distributed Clocks (DC).9
On the mobile manipulator the candidate links sort by where they sit. SPI carries the board-level traffic of joint encoders, current-sense converters, and gate drivers, and its timing is bounded under a fixed transfer schedule. CAN-FD suits chassis and battery-management traffic, where arbitration and retransmission must enter the response-time bound under a stated traffic and fault model. EtherCAT forwards one cyclic frame through every servo node, and TSN Ethernet fits the backbone that joins camera streams to control traffic. None of these bounds belongs to a protocol name. Each depends on the configured devices, wiring, and schedule, so the question to answer is whether the permission path’s servo axes fit within its bus cycle.
Napkin Math 1.2: Fieldbus serialization before schedulability
1. CAN-FD Serialization and Arbitration Latency (9 Axes, 64-byte payload per axis): - Bitrate: \(1.0\text{ Mbps}\) arbitration phase, \(5.0\text{ Mbps}\) data phase. - Per-frame bit count: 32 bits arbitration/CRC overhead at \(1.0\text{ Mbps}\) (\(32\,\mu\text{s}\)), 512 bits payload at \(5.0\text{ Mbps}\) (\(102.4\,\mu\text{s}\)), plus inter-frame spacing \(\approx 6.0\,\mu\text{s}\). - Transmission time per axis: \(t_{\text{tx}} \approx\) 140.4 μs. - For 9 axes polling cyclically: 9 \(\times\) 140.4 μs \(=\) 1263.6 μs (\(=\) 1.264 ms). - Result: Bus utilization exceeds 100 percent (\(\rho \approx\) 126.4 \(> 1.0\)), so this schedule of 9 frames cannot fit a 1 ms bus cycle even before processing. If payload is reduced to 32 bytes (\(t_{\text{tx}} \approx\) 89.2 μs), total bus time is 802.8 μs (80.3 utilization). This reduced payload still needs arbitration, error/retransmission allowances, competing frames, and response-time analysis; no universal 70 percent utilization threshold establishes schedulability.
2. EtherCAT On-The-Fly Processing (9 Axes, Daisy-Chain Ring): - Bitrate: \(100\text{ Mbps}\) full-duplex Fast Ethernet (\(10\text{ ns/bit}\)). - A single standard Ethernet frame carries all 9 joint commands and sensor returns in one pass. - Frame serialization time: 28.2 μs. - Node forwarding delay: 9 slave application-specific integrated circuit (ASIC) controller hops at \(0.6\,\mu\text{s}\) per hop \(=\) 5.4 μs. - Modeled frame serialization plus 9 forwarding hops: \(T_{\text{bus}} =\) 28.2 μs \(+\) 5.4 μs \(=\) 33.6 μs. - Result: These modeled bus terms occupy 3.36 of the cycle, leaving 966.4 μs unallocated by this bus calculation. Master/slave processing, propagation, clock synchronization, retries, application computation, and actuator response still consume the end-to-end deadline.
Systems insight: For the machine’s 9 axes at 64 bytes per axis, CAN-FD serialization alone misses the 1 ms bus cycle, while EtherCAT frame and forwarding occupancy leaves room for other work. TSN Ethernet can carry gigabit camera feeds and microsecond servo frames on one physical link, but only if its time-gated queues are configured and verified across every switch so that video bursts cannot collide with time-critical control slots.
Real-time dependability cannot be achieved by buying faster clock speeds or measuring statistical latency percentiles on an idle bench. Dependability requires structural analysis of the execution path: eliminating all unbounded variance sources (\(I_{\text{unbounded}} = 0\)), bounding shared hardware blocking (\(B\)), enforcing zero-allocation static memory layouts, and selecting fieldbuses with cycle-deterministic distributed clocks (figure 5).
While bandwidth determines how many bytes can cross the bus in a given interval, synchronization jitter (\(\sigma_{\text{jitter}}\))—the worst-case temporal variation in when distributed nodes sample sensors and apply actuator setpoints—determines whether multi-axis physical systems can maintain stable dynamical contact. In multi-axis robotic arms and legged humanoids, field-oriented control (FOC) switches inverter pulse-width modulation (PWM) gates at \(20\text{--}40\,\text{kHz}\), demanding sub-microsecond phase synchronization (\(\le 1\,\mu\text{s}\)) across all joint drives to prevent phase distortion, audible acoustic resonance, and destructive torque ripple.
As mapped across four decades of communication standards in figure 5, legacy serial protocols (Modbus RTU) and carrier-sense priority-arbitrated buses (CAN 2.0B, CAN-FD) exhibit milliseconds of non-deterministic queuing jitter under heavy bus loads, restricting them to low-rate telemetry and supervisory commands. Achieving the sub-microsecond synchronization required for dynamic physical AI necessitates hardware-scheduled time-triggered architectures: either processing-on-the-fly industrial Ethernet with hardware Distributed Clocks (EtherCAT, EtherCAT G) or IEEE 802.1AS/Qbv Time-Sensitive Networking (TSN).10 These architectures anchor clock alignment in silicon PHY timestamps, achieving sub-100 nanosecond jitter across dozens of coordinated actuators and bridging high-bandwidth neural policies directly to hard real-time motor inverters.
A bounded path delivers each command to the drives within its cycle, but it is indifferent to which commands it carries. Which processor may place a command on that path, and what that processor must check first, is a question of authority rather than timing.
Hardware Isolation
Independence between the application processor and the permission path is a claim that often fails its first verification. System architectures routinely label the two as isolated because software tasks execute in separate processes, containers, or hypervisor partitions. Verifying the claim means checking four common-cause paths: independent execution silicon with private instruction pipelines, separate timing references from segregated crystal oscillators, isolated power rails from independent regulators, and segregated memory domains without shared dynamic random-access memory chips or interconnect arbitration buses.14 Any shared substrate needs a fault analysis that shows it cannot defeat the independent permission path under the stated failure model.
Shared electrical and physical infrastructure creates cross-domain failure modes that bypass all software access controls. On the mobile manipulator the application processor and the permission path share a battery but not a rail. The permission path runs on the isolated rail of Which Budget Binds First, held up for 2 s, so a sag of the shared 24 V control rail below its regulators’ 18 V dropout (Electrical Power Integrity) reaches it only as an undervoltage flag. On a shared silicon package or common thermal heat sink, sustained computational load on the application processor elevates junction temperatures beyond \(105^\circ\text{C}\), triggering thermal throttling or silicon shutdown in adjacent real-time cores. When communication relies on an arbitrated interface such as a shared peripheral interconnect or memory bus, a high-bandwidth visual pipeline or a misbehaving direct memory access controller can saturate bus bandwidth, starving the permission path of encoder feedback and sensor telemetry. Even a common clock oscillator creates an unbudgeted single point of failure, where thermal drift, mechanical vibration, or crystal cracking desynchronizes both domains at the same instant.
The permission path must distinguish two failure modes of the proposer, unresponsiveness and corruption. Unresponsiveness occurs when an inference pipeline hangs, a scheduling deadline is missed, or operating system memory fragmentation blocks execution, producing no output. The permission path detects it through the heartbeat of section 1.3, which the application processor must deliver across an isolated channel before its timeout.
Corruption occurs when the proposer remains fully active and passes its heartbeat deadline, but generates kinematically impossible setpoints, discontinuous velocity requests, or numerical infinities resulting from out-of-distribution visual inputs or matrix overflow. The enforcer answers the two failures through separate mechanisms. A missing fresh chunk or a missed heartbeat revokes the proposal when its timer expires, whereas corruption triggers an immediate invariant rejection on the arrival of the malformed proposal packet.
The chunk lease is distinct from the host heartbeat. When the application processor stays alive but stops publishing admissible chunks, the drives would otherwise keep tracking the last admitted setpoints, and the machine would keep moving under a proposal that nothing has renewed, until the enforcer detects the expiry and begins deceleration. Enforcing the lease at the causal boundary (The Causal Boundary) bounds that interval. On the mobile manipulator, section 1.3 sized that interval as a term of the stopping budget and charged it against the pre-brake ceiling.15 The heartbeat can report that the process is alive, but it cannot renew an expired chunk, and its period is independent of the proposer’s. Only the chunk policy renews often enough to hold the lease, and section 1.8.1 shows why the intent model cannot.
A timer feed proves only that the feed path ran. In Bookout v. Toyota (2013), a plaintiff expert argued that a timer-tick interrupt kept servicing the watchdog while other tasks could die (Barr 2013); this is a reported trial mechanism, not a finding that electronic faults caused unintended acceleration in every investigated vehicle.
When a lease expires or an enforcer rejects a corrupted proposal, the permission path executes an explicit degradation strategy rather than an unconditioned power cut, because cutting power hands a moving actuator to whichever unpowered state its drive falls into (Electrical Power Integrity), and none of those states is a controlled stop. Degradation follows a deterministic state hierarchy calculated within the permission path. An optional short continuation of the last validated trajectory is permissible only while its chunk lease remains valid and its stopping margin is feasible. After expiry, any further hold or dead reckoning \(T_{\text{hold}}\) enters \(\tau_{\text{delay}}\) and must fit in the 176.3 ms of pre-brake ceiling that the lease path and brake onset leave at the drive limit (1.1); the enforcer must begin braking sooner if the measured bounds of the remaining terms do not fit. If communication does not resume within that budget, the permission path requests controlled deceleration using the available torque and thermal budget established in Actuator Transmission Limits. It reaches rest only if clearance and braking capability remain sufficient. A gravity-loaded arm may require a rated mechanical holding brake after motion has been arrested.
Once the physical plant is brought to a complete, stationary halt, the architecture executes a recovery handshake before restoring autonomous proposal authority. The safety microcontroller asserts a persistent FAULT_HELD latch in a status word it owns and reports the exact physical rest pose \((\mathbf{q}_{\text{rest}}, \dot{\mathbf{q}} = \mathbf{0})\). The application processor must flush its stale action chunks, clear its autoregressive key-value (KV) attention cache, re-ingest fresh optical observations tagged with new hardware timestamps, and publish a re-synchronized chunk proposal. The microcontroller verifies that the initial waypoint of the new proposal matches the physical rest state within a strict spatial epsilon (\(||\mathbf{a}_0 - \mathbf{q}_{\text{rest}}|| \le \epsilon_{\text{sync}}\)) before releasing the fault latch and re-energizing pulse-width modulation (PWM) to the motor drives.
↳ Downstream: Real-time fieldbus fault-recovery state machines are verified via hardware fault-injection in Fault Injection on Hardware.
Verifying that these isolation boundaries function as designed requires physical falsification on the test bench. Four specific fault-injection regimes must be executed against the physical prototype: power-rail droop testing (injecting maximum-slew current steps to verify less than \(50\text{ mV}\) of rail cross-coupling noise), clock glitching (verifying independent timing domains maintain stable cycle counts under frequency drift), communication bus flooding (saturating links to verify heartbeat telemetry retains deterministic transmission slots),16 and numerical fault injection (streaming synthetic proposals containing NaNs and infinite floats to verify single-cycle rejection without execution stalls).
Independence is therefore validated against identified common-cause faults rather than read off a block diagram, and the lease that bounds a stopped proposer comes from the stopping margin that remains after detection and actuation delays. Both defenses bound a proposer that stops or fails, which leaves open whether a proposer running at the enforcer’s own rate still needs a separate check.
Invariant Checking
Locomotion policies and compact visuomotor models compiled for edge accelerators can emit a new actuator command every \(2.0\text{ ms}\), which tempts designers to let the policy write the motor registers directly. Rate matching does not confer authority. A neural network emitting torques at \(500\text{ Hz}\) is governed by the same statistical uncertainty as a large model generating waypoints at \(5\text{ Hz}\), and even when policy and enforcer share one \(2.0\text{ ms}\) period, the network’s output remains a proposal that must pass the proposal boundary of The Machine in Five Levels before it changes current in a winding. Placing an unverified policy inside the fast loop exposes the plant to three failure modes. Compute underrun occurs when jitter pushes an inference past the actuation deadline and leaves the power stage without a valid setpoint. Out-of-distribution torque spikes occur when an input the network never saw produces a discontinuous torque reversal within a single cycle. Structural limit cycling occurs when unmodeled compliance or sensor lag couples with the policy into resonance. The permission path meets each with a check that fits inside its own deadline: a timeout for the underrun, a torque and rate-of-change limit against the previous cycle for the spike, and an energy budget for the limit cycle (Passivity and Energy-Bounded Interaction). The projection that admits or modifies a candidate command belongs to Safe Sets as Conditional Permission, where it can be built on a control barrier function.
The same reasoning covers the one proposer that is not learned. A remote operator is a third proposer. Its commands pass the same permission path, and network delay spends the same pre-brake budget (Human Authority Over Actuation). Neither speed nor a human origin gives a proposal authority; only passing the permission path’s checks does, and each check spends part of a budget that the next section lays out handoff by handoff.
Checkpoint 1.1: Hardware isolation and real-time invariant checking
Before turning to the handoffs between proposer and actuator, verify your understanding of the exchange, hardware isolation, and pre-actuation invariant defenses:
The Handoffs Matrix
When the mobile manipulator overruns a stop, the diagnosis starts by locating the fault at one of five canonical handoffs between the photons its cameras collect and the current its drives deliver. Upstream, sensing transfers raw transducer signals to perception pipelines. Perception yields intermediate representations that fuse into an estimated world state. The state feeds the learned policy to produce an action plan. The plan submits candidate trajectories to the safety enforcer. Finally, the approved command passes from the enforcer to the physical actuators. Across each boundary, five interface limits govern feasibility: latency (\(\tau\)), bandwidth (\(\beta\)), energy and heat (\(\mathcal{E}\)), determinism (\(\sigma_t\)), and fault domain isolation (\(\mathcal{F}\)). The five handoffs and the five interface limits form a five-by-five matrix that maps every signal path to the physical budgets of the machine. The limits are distinct from the body’s five budgets of The Five Physical Budgets, which enter the matrix chiefly through its latency and energy columns. The record this section produces, the handoff record, consumes the limit records of Measuring a Machine's Own Limits, including the proposer limits that The Cognitive Brain states in the same format.
The example allocation in table 2 marks which limits are design-limiting at each handoff, and its numbers are chosen budgets, not measurements. The matrix gathers the budgets of all three Part I chapters, together with the sensor capture that The Machine in Five Levels places at the Boundary. In the Boundary row (\(\mathcal{H}_1\)), where the world becomes measurement, bandwidth and latency bind first, because raw multi-camera pixels stream over SerDes lanes and the exposure window sets a floor on image age before serialization begins. The Brain rows (\(\mathcal{H}_2\), \(\mathcal{H}_3\)) bind first on energy and latency, the inference costs that The Cognitive Brain established. The Body row (\(\mathcal{H}_5\)) binds first on determinism and heat, because the current loops of The Physical Body must run on time and every ampere they deliver warms the windings and the power stage.
The Nervous System row (\(\mathcal{H}_4\)) is the architectural bridge that ties the Brain and the Body together. The proposal boundary (The Machine in Five Levels) between the learned planner and the enforcer carries an action proposal (\(p_t\)), which consists of a compact spline parameterization or waypoint chunk. For the compact payload assumed in this grid, \(\beta < 100\text{ kB/s}\) is budgeted, so bandwidth does not bind at this boundary; latency, determinism (\(\sigma_t\)), and fault domain isolation (\(\mathcal{F}\)) do. The permission path isolates stochastic proposals with nonblocking mailboxes and expiring leases. On the mobile manipulator the latency limit is the chunk lease, which takes 60 ms of the 258.3 ms pre-brake ceiling at the drive limit (section 1.3), and each chunk is admitted or refused within one 1 ms permission tick. Isolation rests on a private bus, a separate oscillator, and the isolated, held-up rail of section 1.6.
From the causal boundary perspective framed in The Causal Boundary, when a physical AI system fails, the breakdown manifests as an unbudgeted limit crossing at a specific handoff cell rather than an abstract deficiency in learned intelligence. A handoff diagnosis must compare like evidence. For the \(\mathcal{H}_2\) state-estimation path, suppose engineers allocate a \(30.0\text{ ms}\) latency budget and measure \(P_{99}=24.5\text{ ms}\) on a baseline hardware-in-the-loop run, a sampled margin of \(+5.5\text{ ms}\). After thermal throttling, a measured \(P_{99}=85.0\text{ ms}\) gives \(-55.0\text{ ms}\) against that design budget. The decision is to reduce or reallocate workload and retain independent enforcement while revalidating. A \(P_{99}\) sample diagnoses the regression; it is not a worst-case timing proof.
| Canonical Handoff Boundary | Latency Limit (\(\tau\)) | Bandwidth Limit (\(\beta\)) | Energy & Heat Limit (\(\mathcal{E}\)) | Determinism Limit (\(\sigma_t\)) | Fault Domain Isolation (\(\mathcal{F}\)) |
|---|---|---|---|---|---|
| \(\mathcal{H}_1\): Sensing \(\to\) Perception (Boundary) | BINDING \(\tau \le 16.7\text{ ms}\) (exposure & readout sets sensor age) | BINDING \(\beta \ge 1.2\text{ GB/s}\) (multi-camera 4K uncompressed streams) | Non-binding \(P < 3\text{ W}\) (sensor head dissipation isolated) | BINDING \(\sigma_t \le 1.0\text{ ms}\) (frame-sync triggers prevent skew) | Non-binding Defended by SerDes hardware & DMA ring buffers |
| \(\mathcal{H}_2\): Perception \(\to\) State (Brain) | BINDING \(\tau \le 30.0\text{ ms}\) (visual-inertial odometry & feature window) | Non-binding \(\beta \approx 10\text{--}50\text{ MB/s}\) (compressed keypoints & voxels) | BINDING \(P \approx 15\text{--}35\text{ W}\) (edge vision backbone accelerator envelope) | Non-binding Defended by sliding-window state estimator buffers | BINDING Spatial memory isolation: MPUs prevent memory corruption |
| \(\mathcal{H}_3\): State \(\to\) Plan (Brain) | BINDING \(\tau \le 100.0\text{ ms}\) (autoregressive tokenization & diffusion horizon) | Non-binding \(\beta < 5\text{ MB/s}\) (low-dimensional state & prompt embeddings) | BINDING \(P \approx 30\text{--}75\text{ W}\) (transformer/diffusion matrix accelerators) | Non-binding Defended by receding-horizon chunking & spline buffers | BINDING Sandboxed policy container prevents host panics halting telemetry |
| \(\mathcal{H}_4\): Plan \(\to\) Enforcer (Nervous System) | BINDING chunk lease 60 ms (chosen, inside limits derived from the stopping budget); per-tick admission within each 1 ms tick | Non-binding Defended by compact parameterized splines (\(\beta < 100\text{ kB/s}\)) | Non-binding \(P < 0.5\text{ W}\) (fixed-point constraint evaluation on MCU) | BINDING \(\sigma_t \le 100\,\mu\text{s}\) (cycle-deterministic packet reception on MCU) | BINDING Hardware boundary: private bus, separate oscillator, isolated held-up rail |
| \(\mathcal{H}_5\): Enforcer \(\to\) Actuator (Body) | BINDING \(\tau \le 1.0\text{ ms}\) (servo interval \(1.0\text{ ms}\); PWM loop \(50\,\mu\text{s}\)) | Non-binding \(\beta \approx 1\text{--}5\text{ Mbps}\) (cyclic PDOs over deterministic fieldbus) | BINDING \(I^2 R\) copper winding dissipation, MOSFET switching losses | BINDING \(\sigma_t < 10\,\mu\text{s}\) (hard real-time PWM phase alignment & FOC) | BINDING Galvanic isolation, out-of-band analog E-stops, STO circuits |
To ensure these budgets survive the lifecycle of the machine, every boundary condition in the grid is recorded as a handoff record, the verifiable contract between hardware designers, systems engineers, and model developers. Each entry carries the six elements of a limit record (Measuring a Machine's Own Limits) and adds the fields that locate the limit at a handoff and say what kind of bound it is (table 3).
| Field Name | Physical SI Units | Structural Invariant & Description |
|---|---|---|
signal_id |
dimensionless | Unique alphanumeric identifier (e.g. SIG_VIO_01, SIG_TRQ_AX1). |
source_subsystem |
dimensionless | Originating hardware or software module (Camera_SerDes, Planner_VLA). |
destination_subsystem |
dimensionless | Receiving hardware or software module (State_Estimator, Safety_MCU). |
handoff |
dimensionless | The handoff the signal crosses (\(\mathcal{H}_1\) to \(\mathcal{H}_5\)). |
physical_dimension |
dimensionless | Target limit category (\(\tau\), \(\beta\), \(\mathcal{E}\), \(\sigma_t\), \(\mathcal{F}\)). |
value |
SI units (\(\text{ms}\), \(\text{MB/s}\), \(\text{W}\), \(\mu\text{s}\)) | Allocated budget, or the measured distribution it is checked against; for \(\mathcal{H}_4\), the chunk lease derived from the stopping budget. |
conditions |
as recorded | Payload, temperature, supply voltage, firmware build, and concurrent workload under which the value holds; the value lapses when they change. |
uncertainty |
same as value |
Error bound of the instrument or model that produced the value. |
margin |
same as value |
Engineering margin between the operating target and the breakdown threshold. |
detector |
dimensionless | Runtime monitor that catches an excursion (for example, a lease timer, a rail-voltage monitor, or a phase-current sensor). |
provenance |
dimensionless | Source of the value, including its evidence category (measured, derived, datasheet, or illustrative). |
bound_kind |
dimensionless | requirement, proven bound, or measured distribution; only a proven bound may stand in for a worst case. |
clock_domain |
dimensionless | Clock on which every time-valued field is expressed; a handoff that feeds the permission path uses the safety microcontroller’s clock. |
verification_timestamp |
UTC timestamp | Date of bench falsification and sign-off; a log field, never an input to an age test. |
One coupled budget for the anatomy
The three chapters of Part I each set terms of one budget, and the chunk lease is where they meet. Kinetic Momentum fixed the distance side. At its 1.5 m/s drive limit the loaded machine must stop inside the 1.10 m clear distance at a rack end, and the braking term and the fixed overheads leave 258.3 ms for everything that happens before braking force begins. Section 1.3 spent part of that time. A proposer that stalls keeps its authority for one lease, the stop reaches the drives one tick and one bus cycle later, and the stopping budget stands at 835.5 mm.
The lease cannot be made as short as the budget would like, because a proposer must renew it. Supply, Freshness, and the Memory Wall established what each learned proposer can supply. The intent model’s weights take 98.0 ms to stream from memory once, before any vision work, and a full inference takes 160 ms, so at 5 Hz it cannot renew a proposal every few tens of milliseconds. The chunk policy’s weights stream in 0.70 ms, and at 20 Hz it publishes a fresh chunk every 50 ms, each within its 40 ms P99 inference time, so it can.
These facts fix the lease rule for this machine. The lease must outlast one chunk period plus an allowance for renewal jitter, which together set the 55 ms floor of 1.1. It must also fit inside the pre-brake ceiling once the tick, the bus cycle, and brake onset are reserved, which leaves 236.3 ms less the observation age that Sensor Perception measures. The chosen 60 ms sits just above the lower limit, because every millisecond of lease costs 1.5 mm at the drive limit, and later chapters add terms that need that distance.
Each chunk carries the base’s velocity and the arm’s joint targets together, so one lease governs both bodies, and when it lapses each body enters its own stop. The lease is sized on the base’s 258.3 ms pre-brake ceiling because that ceiling is tighter than the limit the arm’s commitment sets. The same lease bounds the arm, and What Planning Hands Over derives the arm’s limit and checks its lease path and ceiling against its own clearance. The arm’s evidence horizon bounds a different lease, the intent lease, which limits how long an admitted goal from the intent model stays valid, and the intent model’s slower targets expire under it (The Intent Lease).
The \(\mathcal{H}_4\) row of table 2 is this chapter’s handoff record filled in, the first link after the limit record in the chain of records that the rest of the book extends. Its value is the 60 ms chunk lease handed from the chunk policy to the permission path, a requirement chosen inside limits derived from the stopping budget and valid up to the 1.5 m/s drive limit, loaded, on a dry floor, at the chunk period above; its margin is the span between the renewal floor and the ceiling-derived upper limit less the observation age, and its detector is the lease timer on the safety microcontroller, counted on that processor’s own clock.
Checkpoint 1.2: The handoffs matrix and physical budgets
Before testing the lease on a humanoid and reviewing the engineering fallacies, verify your understanding of system-wide interface contracts:
Systems Perspective 1.2: A deadline set by a fall
Question: Does the split among proposer, lease, and permission path still hold when the deadline comes from a fall rather than a stop?
An upright biped diverges from balance on the timescale \(\tau_{\text{fall}} = \sqrt{z_c/g}\), where \(z_c\) is its center-of-mass height, about 303 ms for a center of mass near hip height. The split holds, and what changes is the deadline its terms answer to:
- The lease is sized from the step deadline. The lease, the tick, the bus cycle, and the observation age must fit inside the time before a recovery step must begin, and the fallback that an expired lease selects is a validated step rather than a brake.
- The intent model is still disqualified. Its renewal period alone spends most of the fall timescale and leaves no room for the tick, the bus cycle, and the step itself, so only the chunk policy can renew a lease on this deadline.
- Heat and supply sag limit the step. The recovery step draws peak current in both legs at once, heating windings and sagging the supply (Electrical Power Integrity), so the permission path must admit only a step that the supply and the windings can deliver at that moment.
Fallacies and Pitfalls
A block diagram draws a logical line between high-level deliberation and low-level enforcement, but physical hardware often connects them invisibly through shared power rails, interrupt lines, or back-channel diagnostic interfaces. These physical couplings dictate whether the permission loop remains standing when a complex subsystem fails.
Pitfall: Servicing a safety watchdog inside a high-priority timer interrupt rather than an application-verified safety loop.
The mobile manipulator’s self-watchdog removes drive power through the Safe Torque Off path of section 1.5 if it receives no feed within 5 ms, because an enforcer that has stopped completing its checks can no longer command a controlled stop. Suppose engineers refresh it from a high-priority timer interrupt and the lower-priority enforcement task deadlocks. The interrupt continues firing every 1 ms tick even though the task has frozen and stopped checking torque limits, clearance, and chunk leases. The watchdog therefore stays fed while the enforcer has silently lost its safety function. Its service pulse must depend on successful, sequential completion of the required enforcement checks (a windowed watchdog pet conditioned on a task completion token). An independent timer tick establishes that hardware interrupts still cycle, not that the controller has safely evaluated the machine’s physical state (section 1.6).
Pitfall: Leaving auxiliary diagnostic or calibration interfaces unmediated in production operation.
Suppose the mobile manipulator validates base velocity proposals through its permission path, yet its drive-wheel motor drives retain an unmonitored serial calibration interface. A diagnostic daemon uses that interface to write a \(35.0\text{ A}\) test value directly into an inverter register. The proposal mailbox still contains zero velocity, so checking that mailbox does not reveal the bypassed command. The drive wheels can then produce torque that no motion check ever saw or approved. True actuation authority includes every physical wire or digital interface capable of changing a power-stage register. Service interfaces therefore need operational interlocks or equivalent enforcement that prevents calibration writes from bypassing the controller during normal operation.
Fallacy: Evaluating safety barrier invariants on commanded trajectory setpoints is equivalent to evaluating physical plant state.
Suppose the machine’s arm evaluates a control barrier function (Ames et al. 2019) using commanded joint positions instead of load-side encoder measurements to reduce bus traffic. When a gear fails and the link moves away from its commanded trajectory, the interpolator continues producing valid setpoints. The enforcer still calculates that all is safe (\(h(\mathbf{q}_{\text{cmd}}) > 0\)), even as the physical link enters a forbidden region. Computing the wrong input faster cannot correct it. Safety enforcement belongs at the causal boundary (The Causal Boundary), so it must evaluate physical state measurements through a dependable sensory channel and account for their limitations; inspecting the command stream merely establishes what the planner intended, not where the heavy machinery actually is.
Fallacy: Nominal, fault-free operating logs are sufficient to prove physical architectural isolation.
The mobile manipulator completes normal trials without a missed enforcement deadline. Those logs show behavior under a fault-free workload; they do not expose the hidden physical couplings that emerge only under faults. Qualification therefore injects power transients, floods communication buses, corrupts policy memory, and locks a shaft on a dynamometer. Each experiment must check an explicit response bound, such as whether locked-rotor detection clamps current within \(2.0\text{ ms}\), rather than treating the absence of a crash as success. Evidence for isolation comes from exercising the claimed fault boundaries and measuring continued enforcement or the required safe transition.
Summary
The Physical Body measured how far the loaded machine travels before it can stop, and The Cognitive Brain measured what each learned proposer can supply and how quickly its evidence ages. This chapter set the permission path’s check rate from the joint loop’s phase margin, and it converted those two measurements into the time for which a proposal may drive the actuators. The result on the mobile manipulator is a chunk lease long enough to outlast one renewal of the chunk policy and short enough to fit inside the pre-brake ceiling, and the handoff record carries that lease, with its conditions and its evidence, to every later chapter. The lease and every other mechanism of this chapter remain claims until they have been exercised under compute delays, bus faults, bad proposals, sensor errors, and loss of braking authority. Part I opened on the irreversibility law (principle \(\ref{pri-vol4-irreversibility}\)), and the lease closes Part I as a rule about authority, since motion that cannot be recalled may be commanded only for as long as the stopping budget can pay for its travel.
Key Takeaways: The permission path and determinism
- Proposal is not permission: High-level learned models act strictly as unprivileged proposal engines (\(p_t\)). Physical write privilege over motor drive registers is reserved exclusively for the deterministic real-time enforcer, on silicon whose isolation from the proposer has been demonstrated.
- Cadence is dictated by physics, not software: The example rates (\(20\text{ kHz}\) current regulation, \(1\text{ kHz}\) joint checks, \(10\text{--}50\text{ Hz}\) action chunking, and slower deliberation) are selected against winding time constants, phase margin, exposure, workload, and response budgets.
- The exchange never waits, and every proposal expires: The permission path reads the mailbox in one bounded attempt per tick and keeps its last validated chunk on a collision, and the proposal header’s evidence time and expiry end that chunk’s authority whether or not any cancellation arrives.
- Worst-case bounds outrank empirical distributions: A statistical latency percentile (\(P_{99}\)) cannot prove real-time safety. Deterministic certification requires analytical Response Time Analysis (\(R_i = C_i + B_i + I_i \le D_i\)), eliminating unbounded bus contention and memory allocations from the safety path.
- The veto is only as independent as what it shares: A permission path that shares a clock, a power rail, a memory bus, or silicon with the proposer can be silenced by the proposer’s faults, so independence is a claim that fault injection must demonstrate, not a property of separate processes.
- Stopping dynamics dimension the chunk lease: On the mobile manipulator at its drive limit, the pre-brake ceiling is 258.3 ms. The chosen 60 ms lease sits just above the shortest lease the chunk policy can renew, and the path from its last renewal to brake onset spends 82 ms, bringing the stopping budget to 835.5 mm of the 1.10 m clear distance.
- A failure is an over-budget limit at one handoff: The handoff record places each limit at one of five handoffs with its conditions, evidence category, and bound kind, so a regression is diagnosed against its allocated budget with like-for-like evidence.
What’s Next: From machine anatomy to closed-loop physical learning
Footnotes
Interconnect Crossbar Contention: When an accelerator streams multi-gigabyte foundation model weights across shared AXI or PCIe crossbars, dense burst transactions saturate dynamic RAM banks and arbitration queues. Co-located real-time processor cores attempting to read motor encoder registers face unbudgeted hardware wait states that software priorities cannot preempt. Maintaining microsecond-scale control stability therefore requires dedicated physical memory channels or isolated on-chip SRAM for the real-time domain.↩︎
Inference Latency Distribution: Log-normal and shifted Weibull models are common descriptions of measured inference tails. The median and \(P_{99}\) alone do not fix the exceedance probability at a given deadline, which must be measured or modeled separately. For the formal derivation of latency exceedance probabilities in real-time inference, see Machine Learning Background.↩︎
TLB Shootdown: When an operating system modifies page tables, it broadcasts inter-processor interrupts (IPIs) to invalidate stale Translation Lookaside Buffer entries across all active cores. Receiving cores must suspend thread execution and flush affected entries, introducing unbudgeted multi-microsecond stalls. In shared-silicon architectures, these unpredictable synchronization pauses destroy the cycle determinism demanded by inner motor regulation loops.↩︎
DMA Cache Coherence: Direct memory access transfers bypass CPU cache hierarchies, requiring explicit cache line invalidation or write-back operations to maintain memory consistency between peripherals and cores. Without hardware-enforced snoop filtering or explicit barriers, processors read stale sensor frames or overwrite pending actuator commands. These maintenance operations inject variable latency penalties into high-rate control routines unless static, unbuffered SRAM regions are dedicated to peripheral transfers.↩︎
Dead-Time Delay: Dedicated pulse-width modulation timers (TI
ePWM, STM32TIM1) inject hardware-enforced dead-time delays (\(100\text{--}500\text{ ns}\)) between complementary high-side and low-side gate signals (TIMx_BDTR,DBRED). This delay ensures that the conducting power switch fully turns off before its opposing switch conducts, preventing shoot-through short circuits across high-voltage DC rails. Because dead time slightly distorts the applied phase voltage, motor control firmware incorporates online volt-second compensation to maintain smooth torque generation.↩︎Zero-Order Hold Phase Lag: Holding a discrete control command constant across sampling interval \(T_s\) introduces an effective delay of \(T_s/2\), producing a linear phase reduction \(\Delta \phi = -\omega T_s / 2\). In high-bandwidth robotic actuators, this phase erosion directly reduces open-loop phase margin, driving lightly damped mechanical modes into resonance. For the formal frequency-domain derivation of discrete sampling phase erosion, see Control and Dynamics.↩︎
Phase Lag in Moving Average Smoothing: Averaging noisy actuator feedback using an \(N\)-tap finite impulse response (FIR) filter introduces a group delay of \((N-1)/2\) sample periods. In high-gain position control, this artificial transport delay consumes phase margin, turning stabilizing feedback into positive feedback that drives the actuator into self-excited limit cycles.↩︎
Rate-Monotonic Scheduling: Liu and Layland (1973) established that static-priority rate-monotonic scheduling guarantees periodic task schedulability when total processor utilization satisfies \(U \le n(2^{1/n} - 1)\), converging to \(\ln(2) \approx 0.693\) for large \(n\). In distributed robotic architectures with arbitrary execution deadlines and shared resource blocking, systems engineers deploy exact Response Time Analysis (\(R_i = C_i + B_i + I_i \le D_i\)) to verify deterministic schedulability. For the formal mathematical derivation of response time recurrences, see Systems and Hardware.↩︎
Distributed Clock: Under IEC 61158-4-12, EtherCAT Distributed Clocks (DC) synchronize distributed servo drive nodes to a master reference clock by measuring propagation delays across the physical ring topology. Distributed clock hardware can generate synchronized
SYNC0pulses. Achieved clock error and application sampling jitter depend on the selected devices, configuration, and measurement; clock alignment alone does not bound the complete control cycle.↩︎Time-Sensitive Networking: IEEE 802.1Qbv Time-Sensitive Networking establishes deterministic Ethernet communication by configuring time-aware traffic shapers across network switches. High-priority control frames execute within dedicated, pre-scheduled transmission windows, while asynchronous high-bandwidth payloads (such as uncompressed camera streams) are temporarily gated. This scheduled time-division multiplexing allows vision telemetry and microsecond servo commands to coexist over a unified physical link without packet collision or buffer bloat.↩︎
Hardware Trip Zone: Microcontroller hardware trip zones (such as TI
EPWM_TZFRCor STM32TIMx_BDTRBreak inputs) couple analog overcurrent and overvoltage comparators directly to timer output logic. Upon fault detection, configured trip logic may force gate signals high-impedance or activate a validated braking mode. These have different mechanical consequences: torque removal can leave the plant coasting, while dynamic braking requires a rated energy path.↩︎Braking Dynamics: Arresting a high-inertia manipulator joint transfers kinetic energy into the implemented braking path: winding and power electronics losses, a rated brake resistor, or mechanical friction. The trajectory must satisfy stopping clearance (\(d_{\text{stop}} \le D_{\text{clear}}\)) and each component’s pulse-energy and thermal limits; \(\int I^2R\,dt\) accounts for the winding share when resistive motor braking is used. For the complete derivation of kinematic stopping envelopes and transient thermal dissipation, see Control and Dynamics.↩︎
Safe Torque Off (STO): An STO implementation removes torque-producing drive signals through a hardware interlock. An IEC 61800-5-2 integrity rating depends on the actual devices, diagnostic coverage, wiring, and validation, not on the circuit pattern alone.↩︎
Freedom from Interference (FFI): ISO 26262-9 and IEC 61508 formalize Freedom from Interference (FFI) as the absence of cascading faults between software elements possessing different safety integrity levels. Unverified neural policies sharing silicon or bus fabrics with safety-critical tasks require architects to demonstrate spatial (memory protection), temporal (independent scheduling and timers), and electrical isolation (independent power rails). Shared resources require evidence that interference cannot defeat the specified enforcement deadlines and fault response.↩︎
Chunk Lease Bound: For a constant-speed approach, the chunk lease, the tick, the bus cycle, and brake onset must satisfy the pre-brake inequality of section 1.3 once the observation age \(t_{\text{age}}\) joins them on the left side, and a post-expiry hold \(T_{\text{hold}}\) adds to that side as well. The right-hand side is the pre-brake ceiling, not a lease to spend in full. For the complete algebraic derivation of lease bounds under bounded deceleration, see Watchdog validity leases and teleoperation latency bounds.↩︎
CAN Arbitration: Controller Area Network protocols resolve bus contention at the physical layer using wired-AND dominant (‘0’) and recessive (‘1’) signaling. If a defective transceiver continuously asserts a dominant state (a “babbling idiot” fault), it monopolizes the shared physical bus and blocks all higher-level communication. Hardware bus guardians and dominant-timeout clamp circuits monitor transmission duration, physically disconnecting jabbering nodes to preserve network availability for remaining safety participants.↩︎


