Silicon Placement
Silicon Placement
Purpose
Why can a learned model delay a safety check that runs on a different processor core?
Software architecture diagrams often depict proposal and permission paths as cleanly decoupled functional blocks. On physical silicon, however, disparate software components inevitably contend for shared hardware resources: memory buses, power distribution networks, clock trees, and thermal dissipation paths. A high-throughput burst of neural inference can saturate the shared memory interconnect or induce inductive voltage droop across power rails, delaying a safety check running on an adjacent processor core while the physical chassis continues to coast.
Software privilege rings govern address space permissions, not transaction scheduling or DRAM queue latencies. Consequently, logical isolation cannot guarantee worst-case execution time under peak contention. Only empirical stress measurement under adversarial workloads can verify timing deadlines. Enforcing real-time guarantees requires hardware-level isolation, dedicated memory channels, or physically separate microcontroller enclaves. In the physical AI stack, hardware placement establishes the computational foundation of the Nervous System, ensuring deterministic enforcement checks complete before commands cross the causal boundary.
Learning Objectives
- Inventory the memory, interconnect, power, clock, reset, and thermal resources that proposal and permission paths share
- Explain why average bus utilization cannot bound a deadline, using transaction size, queue policy, and bank timing
- Size a windowed watchdog’s service window from measured completion jitter for a non-blocking record exchange between the two paths
- Analyze how rail droop and thermal throttling stretch the permission path’s execution time on different timescales
- Design a loaded stress test whose falsification threshold is the permission path’s declared worst-case execution time
- Select among reservation, partitioning, workload reduction, and relocation from measured interference and the lease’s timing allowance
- Construct a placement record that scopes loaded observed maxima, mitigations, and untested conditions for release review
The Silicon Substrate
The warehouse mobile manipulator carries its application processor and its permission path in one chassis. On its software architecture diagram, the proposal boundary between the Brain and the Nervous System is a clean dashed line. The barrier filter and fallback ladder of Safety Enforcement keep their guarantees only while the permission path finishes its check inside its deadline on every tick, and that condition depends on where the check executes, which the diagram does not show. On physical silicon, the dashed line can collapse onto a single shared die. Suppose both paths ran on one heterogeneous System-on-Chip (SoC), the placement this machine rejects. The chunk policy and the intent model would execute on the embedded accelerator while the permission path executed on an adjacent real-time core. When the intent model begins an inference pass across its multi-billion-parameter weights, billions of transistors activate simultaneously. This current surge (\(dI/dt\)) induces an inductive and resistive voltage droop (\(\Delta V = L \cdot dI/dt + IR\)) across the chip’s internal Power Distribution Network (PDN), risking clock instability and setup-time violations in adjacent real-time logic.1 Concurrently, dense tensor matrix multiplications saturate the shared LPDDR memory bus, streaming gigabytes of activation maps and evicting the safety enforcer’s pinned sensor tables from the last-level cache.
When the permission path reads a sensor buffer while the proposal workload occupies the shared memory controller, its requests queue behind direct memory access traffic, and the wait can consume the share of the deadline that the enforcer’s computation needed. Logical privilege does nothing to prevent that wait, because it governs which addresses a task may touch, not where its requests wait in the memory queue.
↰ Prerequisite: Hardware execution cadences and microcontroller timing loops were defined in Multi-Rate Cadences.
Both of the machine’s jobs would feel shared-die coupling. For the Class 1 base, a memory stall during feature extraction delays the permission path’s refusal while the loaded base closes on a rack end, and every tick of delay beyond the one the stopping budget counts is distance it never reserved. For the Class 2 arm, an accelerator burst that evicts the permission path’s tables from a shared cache delays the contact check while the gripper presses on the spring-latched cage door.
Following the end-to-end argument in system design (Saltzer et al. 1984), memory protection and operating system process isolation do not by themselves establish operational independence. On a heterogeneous die,2 operational independence requires that proposal-side activity cannot push enforcement execution beyond its hard physical deadline, and that proposal-side crashes or thermal overruns cannot disable the permission path’s authority to refuse.
Because these couplings are governed by physical solid-state phenomena rather than software abstractions, architectural claims of isolation must be treated as hypotheses to be falsified through measurement. Testing such a claim takes four engineering steps: an inventory of every shared silicon resource and its claimants; a stress campaign that drives the proposal path to worst-case compute, memory, and bus contention while recording the permission path’s latency distribution and observed maximum; measurements of voltage droop and junction temperature under sustained load; and a placement record that states the observed maxima against the deadline together with the conditions left untested. Those steps recur through the rest of the chapter, which examines the shared die as the placement the machine rejects and the discrete controller as the one it uses. It begins with what each path must compute, where it runs, and how much slack it keeps when nothing competes with it.
The historical evolution of physical AI architectures is fundamentally rooted in the physical limits of semiconductor scaling (figure 1). Over fifty years of microprocessor progress (1975–2025), computing systems transitioned across two major architectural boundaries. During the classical scaling era (1975–2004), Moore’s Law and Dennard scaling allowed clock frequencies to surge exponentially from \(2\,\text{MHz}\) to \(3.8\,\text{GHz}\) without increasing thermal power density. Autonomous robotics rode this wave directly, transitioning from offboard minicomputer tethers (Shakey) to onboard VMEbus racks (CMU Navlab 1) and embedded x86 motherboards (Stanford Stanley).
When Dennard scaling collapsed in 2004, static leakage power capped single-core clock frequencies near \(3.8\,\text{GHz}\), bringing the era of frequency scaling to a halt. Industrial robotics adapted by decoupling architectures: deterministic microcontrollers assumed responsibility for cycle-counted \(1\,\text{kHz}\) fieldbus motor commutation loops, while application processors handled multi-threaded trajectory generation. When high-capacity deep learning and vision-language-action (VLA) foundation models arrived, physical AI encountered a second, severe hardware constraint: the mobile robot payload thermal and battery power ceiling (\(15\text{--}60\,\text{W}\)). Server-grade accelerators consuming \(300\text{--}700\,\text{W}\) cannot be thermally dissipated or battery-powered on an untethered mobile manipulator or humanoid. This constraint forced the deployment of Domain-Specific Accelerators (DSAs) and heterogeneous edge Systems-on-Chip (such as NVIDIA Jetson AGX Orin and Drive Thor), expanding onboard compute density by nine orders of magnitude (\(10^9\times\)) while staying strictly within mobile power limits. Most recently, sequential token decoding in autoregressive VLAs hit the cognitive deliberation wall (\(1\text{--}3\,\text{Hz}\)), prompting the shift to flow-matching and diffusion policies with temporal action chunking that restore closed-loop control cadences to \(30\text{--}50\,\text{Hz}\).
/tmp/ipykernel_753/979283785.py:180: UserWarning: This figure includes Axes that are not compatible with tight_layout, so results might be incorrect.
plt.tight_layout()
Bridging high-capacity neural deliberation to edge silicon requires fundamentally different compilation principles than general-purpose cloud computing. Cloud inference engines optimize for average throughput (tokens per second) through dynamic batching, runtime graph tracing, and heap memory allocation (malloc). In an embodied physical AI system, these practices are hazardous. Dynamic memory allocation introduces heap fragmentation and allocator lock contention, while just-in-time compilation and paging faults inject multi-millisecond tail jitter (\(P_{99.9}\)) that directly steals stopping clearance (Stopping Envelopes). Edge compilers for physical AI therefore enforce ahead-of-time (AOT) static graph compilation: the execution graph’s topology is permanently frozen, tensor dimensions are fixed, and every intermediate activation buffer is mapped to a static, pre-allocated physical memory offset before actuation begins, guaranteeing zero runtime dynamic allocations.
Beyond eliminating tail jitter, compilation on edge SoCs is governed by thermodynamics and the memory wall. As established in Supply, Freshness, and the Memory Wall, moving data across an off-chip dynamic RAM (DRAM) bus consumes roughly \(10\text{--}20\times\) more energy per bit than performing arithmetic operations inside on-chip registers or static RAM (SRAM). On an untethered mobile robot operating within a \(15\text{--}60\,\text{W}\) payload thermal budget, passing intermediate activation tensors back and forth across the DDR bus rapidly exhausts the thermal dissipation capacity of the chassis. Operator fusion compiles multi-layer subgraphs (such as Convolution + Bias + Activation, or Attention Projection + LayerNorm) into unified fused kernels that execute entirely within on-chip register files and local SRAM caches, eliminating off-chip memory traffic. In parallel, numerical precision reduction (quantizing weights and activations from FP32/FP16 down to INT8 or FP4) acts as a direct multiplier on memory bus capacity. Halving the operand bitwidth doubles the effective transfer throughput of the memory crossbar, directly reducing memory bus occupancy and preserving bus bandwidth for concurrent real-time safety monitoring.
Two Paths on One Die
The mobile manipulator partitions its control across two paths with conflicting requirements. The proposal path runs the chunk policy on the application processor at 20 Hz, taking the navigation and wrist camera streams and emitting base velocity and arm joint targets together. Its P99 inference, vision included, takes 40 ms (illustrative; see the Reader Guide) of each 50 ms chunk period; it holds no direct hardware authority over the drives and runs in unprivileged user space. The permission path runs on the discrete safety microcontroller (MCU) at 1 kHz. Each tick it checks the latest proposal against the stopping envelope of Stopping Envelopes and stages the admitted setpoints for the EtherCAT frame to the drives. It must finish inside a chosen 400 μs deadline within its 1 ms tick, and it runs in a privileged real-time execution domain on a dual-core lockstep controller,3 whose comparator diagnostics cover the permission path’s own failure, because an envelope check cannot detect a fault in the core that executes it. Only the permission path holds memory-mapped input-output (MMIO) access to the fieldbus master that carries drive commands.
Definition 1.1: Heterogeneous silicon partitioning
Heterogeneous silicon partitioning is the architectural allocation of stochastic, high-capacity cognitive workloads (perception, learned policy evaluation) and deterministic safety loops (safety invariant checks, admission and fallback) across specialized execution domains isolated in hardware.
- Significance: Placing high-capacity neural deliberation and real-time safety enforcement on the same general-purpose operating system or compute engine creates common-mode failures. Kernel thread preemption, dynamic memory allocation, and GPU warp scheduling inject multi-millisecond tail jitter that causes the permission path to miss physical control deadlines.
- Distinction: Unlike logical software process isolation (such as Linux containers or POSIX process priorities), silicon partitioning enforces isolation across independent hardware registers, tightly-coupled static memories (TCM), and dedicated peripheral buses.
- Common pitfall: Assuming multi-core CPU affinity provides true failure-domain separation, when the cores still share caches, memory controllers, clock trees, and power rails.
↰ Prerequisite: The sub-millisecond execution deadline of the CBF-QP enforcer was established in Minimal Intervention.
The two processors run different kinds of software, following the two-processor implementation of Where the Permission Path Runs. Proposal inference, telemetry logging, network communication, and fault recording run on the application processor under Linux, whose scheduling tail stays exposed to kernel and driver jitter even in a real-time configuration.4 Enforcement and sensor validation run on the discrete safety microcontroller as bare-metal firmware behind a hardware memory protection unit,5 with the encoder interfaces, the IMU, and the timer channels wired directly to it. Each proposal reaches the microcontroller over a board-level link and is published there through the seqlock mailbox of Multi-Rate Cadences, so no dynamic random-access memory (DRAM) is shared between the paths. On the supply side the two paths share only the battery. The microcontroller, the drive logic, and the spring-brake coils sit on the isolated rail of Which Budget Binds First, which carries no proposal-side load and is held up by a local store for 2 s, longer than any stop the enforcer can command inside the envelope (The Fallback Ladder). A sag on the shared control rail therefore cannot reset the permission path.
On the unloaded machine the computational budgets appear generous. Illustrative sub-costs on the safety microcontroller put reading the encoders, the IMU, and the proposal’s evidence epoch at 40 μs, evaluating the constraints at 80 μs, and staging the admitted setpoints for the drive frame at 15 μs. The enforcer’s unloaded cycle totals 135 μs, which at the microcontroller’s 400 MHz clock is 54,000 cycles. Against the 400 μs deadline, the unloaded path leaves 265 μs of slack, and the declared worst-case execution time of 250 μs, the qualified budget of Minimal Intervention, still leaves 150 μs. On the application processor, the chunk policy’s 40 ms inference leaves 10 ms of its 50 ms period. Evaluated in isolation, both loops operate with comfortable margins.
Assigning tasks to separate processor cores does not create independent failure domains if the silicon infrastructure is shared.6 Hardware Isolation traced these common-cause paths on the machine’s board, where separate packages and the isolated rail cut the memory and supply paths. On a shared die, the rejected option for this machine, they move inside one package. The real-time core executes instructions independently from the application processor, yet both cores rely on common hardware services:
- Shared interrupt routing: The Generic Interrupt Controller (GIC) routes peripheral and inter-processor notifications across a shared fabric, where a flood of Ethernet packet interrupts on the application processor can delay the timer interrupt that triggers the 1 kHz permission loop.
- Shared memory controllers: Both paths access the same dynamic memory controller and interconnect crossbar, where large matrix tile loads generated by the neural accelerator saturate DRAM command queues, turning an expected \(20\text{ ns}\) level-two cache refill on the real-time core into a \(350\text{ ns}\) round-trip to main memory.7
- Shared clock trees: The cores share phase-locked loops (PLLs) for clock generation, where dynamic voltage and frequency scaling (DVFS) transients on the application processor introduce clock jitter across adjacent real-time domains.
- Shared power rails: A single Power Distribution Network (PDN) delivers current to both domains, where switching surges from the neural accelerator drop the internal voltage rail, risking logic state corruption or brownout resets.
- Shared reset domains: A kernel panic on the application processor that triggers a hardware watchdog reset can pull down the shared peripheral bus or system interconnect, disabling the drive outputs and leaving the base without a permitted command.
Shared hardware contention across these domains is not merely a passive scheduling inefficiency; it establishes an active cyber-physical attack vector. In a heterogeneous SoC where third-party neural workloads or network-exposed software containers share silicon with the permission path, a compromised application process can execute deliberate denial-of-service (DoS) attacks against physical safety. By generating pathological unaligned memory traffic or triggering cache-line thrashing (exploiting cache side-channel mechanisms or Rowhammer-style disturbances), an attacker on the unprivileged proposal path can deliberately saturate memory controller queues, delaying the safety microcontroller’s memory access and forcing deadline misses. Similarly, an adversarial workload can trigger high-frequency tensor core switching patterns designed to maximize \(L \cdot dI/dt\) inductive voltage droop, driving the shared power distribution network into brownout and inducing common-mode hardware resets. In physical AI architectures, software process isolation cannot defend against electrical and microarchitectural coupling; hardware-enforced quality-of-service partitions (such as ARM MPAM) and physical silicon separation are fundamental cyber-physical security controls.
Because software process isolation cannot mitigate these hardware-level timing and security couplings, physical silicon placement becomes the primary architectural defense. For the machine’s permission loop, compare three placements. Shared application cores expose the permission loop to the kernel and driver scheduling tail described earlier, a tail that can push it past not only the 400 μs deadline but the 1 ms tick itself. A real-time cluster on the application SoC keeps its own core but shares the die’s memory controller, rails, and thermal path; section 1.5 shows the clock throttle that rejects it. The discrete safety microcontroller that the machine uses pays for its separation in transport instead, a board-level hop from the application processor that neither shared-die option needs and that section 1.7 charges to the lease. The choice among the three rests on each option’s measured sensor-to-drive response under interference, not on how many packages it uses.
Principle \(\ref{pri-vol4-proposal-permission}\) requires a permission path that learned workloads cannot delay, and the three placements differ precisely in which resources the two paths still share. Even the discrete microcontroller’s separation is a claim about timing under load, which only a measurement of its tick under that load can support. Every resource the two paths share has claimants on that tick, and the memory system is where they collide first.
Contention for Shared Resources
Every shared physical resource has multiple claimants; a utilization metric that ignores concurrent competition cannot bound a deadline. Had the mobile manipulator placed its permission path on the application SoC, the proposal path, the permission path, and the system infrastructure would compete directly for the same memory controllers, crossbar switches, and power distribution rails. Digital claimants on that shared memory bus would include camera and lidar DMA streams depositing frames into DRAM, the neural accelerator fetching model weight matrices and intermediate activation tensors, the real-time core reading its constraint tables and sensor history, actuator write buffers for drive commands, cache controllers executing dirty line writebacks, and diagnostic logging processes writing state records to nonvolatile storage. On the machine as built, the same bursts stay on the application processor, where they delay the observation and the chunk rather than the permission tick. The permission path meets them as evidence age (The Cost of Ingestion), refusing evidence older than its refusal threshold and entering its fallback if no admitted chunk renews the lease.
On a shared die, crossbar bus and DRAM bank contention is queuing delay when memory and DMA transactions share an interconnect and controller, and its bound depends on transaction size, queue policy, and bank timing.8
Each claimant’s average demand follows from its transfer size and event rate, and on a shared LPDDR channel the camera streams, weight fetches at the policy rate, the enforcer’s reads, drive writes, writebacks, and logging together occupy a small fraction of its capacity on average. An engineer who inspects only that average concludes that most of the bus remains free. Memory arbiters do not schedule average rates. They arbitrate discrete bursts, and two workloads with identical average throughput interfere differently if one streams in small packets while the other arrives in long, saturating blocks.
Arithmetic intensity (Williams et al. 2009) says whether a routine is bound by execution units or by memory bandwidth in steady state; it says nothing about which request waits when two paths reach the bus together. While weights stream in from DRAM, the accelerator’s arithmetic intensity is zero, and it acts as a high-bandwidth direct memory access engine that saturates the shared bus. Weight streaming is also what caps the rate of fresh proposals in the memory wall of Supply, Freshness, and the Memory Wall. Placement sees the other side of that limit, because every weight transfer that bounds the proposer is traffic a co-located permission path would wait behind.
How long a safety read waits depends on transaction arbitration and DRAM scheduling. An accelerator that moves weights in \(262{,}144\)-byte tiles does not issue one burst: Arm’s AXI4 specification permits at most 256 beats per burst and forbids a burst from crossing a \(4{,}096\)-byte address boundary, so on an assumed 128-bit internal path a tile becomes 64 aligned transactions. An in-flight transaction may be non-preemptible, and whether the remaining tile transactions block a later safety read depends on whether the controller implements priority bypass. DRAM adds its own delay, because only one row per bank is open at a time: like a clerk who must shut one drawer before opening another, a read to a closed row waits for precharge (\(t_{\text{RP}}=14\text{ ns}\)) and activation (\(t_{\text{RCD}}=14\text{ ns}\)) before column service.
Controllers that use First-Ready First-Come-First-Served (FR-FCFS) scheduling let ready row hits overtake older requests, and a streaming accelerator supplies row hits in abundance, so that policy needs its own starvation bound. Under a FIFO queue that lets no later arrival pass the safety read, the idealized wait is \[t_{\text{queue}} = \sum_{k \in \mathcal{Q}} \frac{B_k}{\text{BW}_{\text{peak}}} + N_{\text{conflicts}} (t_{\text{RP}} + t_{\text{RCD}}) \tag{1}\] where \(\mathcal{Q}\) holds the transactions ahead of the read, \(B_k\) is each one’s byte count, \(N_{\text{conflicts}}\) counts bank switches, and \(\text{BW}_{\text{peak}}\) is an optimistic service rate, so arbitration overhead and refresh add to the result.
Napkin Math 1.1: The permission path's loaded read on a shared die
Variables: The tick allots 40 μs to sensing and 95 μs to constraint evaluation and the drive-frame write, inside a 400 μs deadline. The snapshot is 4,096 bytes on a 12.8 GB/s channel. Ahead of it the FIFO holds one weight tile (64 transactions), a camera transfer, and a telemetry record: 88 transactions, 360,448 bytes, 3 bank switches.
Math: By equation 1, serving the backlog’s 360,448 bytes at 12.8 GB/s, plus 3 switches of 28 ns, takes 28.244 \(\mu\text{s}\). The snapshot itself reads in 0.32 \(\mu\text{s}\), so the loaded read takes 28.564 \(\mu\text{s}\) and leaves 11.436 \(\mu\text{s}\) of the sense allowance. A second backlog would overrun it. The declared worst-case execution time of 250 μs is crossed once the read exceeds 155 μs, which takes more than 5 backlogs, and that crossing already refutes independence even though the tick still meets its deadline. The deadline itself holds until the read exceeds 305 μs, which takes more than 10 backlogs, so a tick that meets its deadline is no evidence that the placement is sound. With validated priority bypass, only the one in-flight transaction waits ahead, and the read takes 0.64 \(\mu\text{s}\).
Systems insight: The read’s bound belongs to the memory controller, not to the workload. It exists only if the controller limits how much traffic it admits ahead of a safety read or lets that read bypass the queue, and an FR-FCFS controller serving a streaming accelerator provides neither unless its configuration is qualified to do so. The machine’s discrete microcontroller removes the question by reading its snapshot from tightly coupled memory on its own bus.
Contention extends to the power distribution network, where claimants compete for charge rather than bandwidth. The droop of section 1.1, \(\Delta V = L\,dI/dt + IR\), grows with both the rate and the size of the current step. An off-chip voltage regulator responds far too slowly to catch a nanosecond current step, so the first charge comes from on-die decoupling capacitance, and once that is depleted the shared rail sags under every core it feeds.
On a shared \(V_{\text{core}} =\) 0.85 \(\text{V}\) supply with package inductance \(L =\) 30 \(\text{pH}\) and resistance \(R =\) 8 \(\text{m}\Omega\), an accelerator current step of \(\Delta I =\) 6 \(\text{A}\) in 2 \(\text{ns}\) produces 90 \(\text{mV}\) of inductive and 48 \(\text{mV}\) of resistive drop, 138 \(\text{mV}\) in all, a 16.2 percent sag to 0.712 \(\text{V}\). The alpha-power gate model9 turns that sag into a 39 percent increase in propagation delay (Inductive Voltage Droop and Gate Delay Stretch derives both steps). That is a gate delay, not a tick. Whether it consumes the permission path’s slack, trips a brownout reset, or does nothing depends on the device’s minimum operating voltage and clock policy, and only rail and timing measurements under the declared current step can settle it. Both interference channels—interconnect FIFO queueing and power rail voltage droop—are structured in figure 2. A shared die leaves those questions open for every accelerator burst; the discrete microcontroller, on its own bus and rail, does not inherit them.
Detecting these interactions requires instrumentation outside the affected core, because software on a stalled or slowed core cannot time its own delay. Memory-controller counters, rail monitors, thermal diodes, and reset-cause registers record the coupling.
A shared die can meter these couplings in hardware, with a crossbar arbiter that routes tagged safety transactions through their own queue while metering accelerator bursts with bandwidth credits, and a power-management controller that throttles only the accelerator’s clock at a thermal trip. Each is a property of the controller that the placement test must confirm, not a scheduler setting the block diagram can assume. Arbitration decides when a transaction moves, not what it carries. What crosses the physical boundary between the two paths must be defined by explicit guarantees on production time, validity, and failure response.
The Boundary Contract
What crosses between the proposal and permission paths under a boundary contract is a record, not an unqualified pointer into shared memory, because passing a raw memory address across the trust boundary couples the reader directly to the writer’s internal memory layout, execution timing, and failure domain. If the chunk policy running on the accelerator provides only a pointer to its private output buffer, the permission path has no mechanism to determine whether the writer is still modifying the buffer, has stalled mid-inference, or has encountered an unhandled fault. The boundary payload must therefore be an immutable, self-contained record. The record that crosses is the proposal header of Multi-Rate Cadences, published through its seqlock mailbox; placement adds no fields, only the evidence that the exchange meets its deadline under load. The permission path reads the record as an isolated snapshot, evaluating its temporal freshness and physical validity before any setpoint reaches a drive.
Preserving temporal isolation, whether the exchange crosses a shared die or the machine’s board-level link, requires eliminating blocking synchronization at the exchange interface. If the permission loop acquires a lock held by the proposal generator, an inference stall can delay a physical decision. A \(25\text{ ms}\) host stall must therefore leave the 1 kHz enforcer running each tick: it may use the last validated chunk only while its lease and current state remain admissible, then transfer to its independently validated fallback. Allocation and logging on the permission path likewise need measured time bounds. The producer may stall or drop updates; it cannot backpressure the plant’s enforcement cadence.
The mailbox rejection rules of Boundary contract failure classes and deterministic fallback matrix keep a failed producer from forcing the enforcer to wait. A rejected proposal leaves the admitted setpoints in force until the lease lapses; if no admissible proposal renews it and the machine is moving, the MCU executes the stop rung of The Fallback Ladder, the base’s resident stop and the arm’s state-matched stop, neither of which waits for the proposer. That fallback must still complete within a bounded time, and the placement test measures that bound under load on the silicon that hosts the path.
Windowed watchdog timers and execution liveness
Software memory boundaries and lock-free seqlocks protect against torn reads, but cannot guarantee that the safety enforcer code is alive and progressing through its intended operations. Liveness of the enforcer belongs to the self-watchdog of Multi-Rate Cadences, which is fed only after a completed tick. A countdown watchdog serviced from a timer interrupt stays blind to a dead task, the failure argued in Bookout v. Toyota. On shared silicon, the enforcer can also be starved by interconnect contention or run too fast after clock corruption, so placement needs a stronger form of that self-watchdog, one that constrains the feed from both sides.
High-integrity physical AI architectures therefore enforce windowed watchdog timers. A windowed watchdog rejects any refresh that occurs either too early or too late. The timer hardware defines a service window parameterized by a lower bound (the closed window, \(T_{\text{lower}}\)) and an upper bound (the open window, \(T_{\text{upper}}\)):
- Early-kick violation (\(t_{\text{service}} < T_{\text{lower}}\)): If software attempts to refresh the watchdog before the lower threshold has elapsed, the hardware immediately triggers a fault reset. An early kick can indicate that the software has bypassed required algorithmic stages, branched into a pathological short loop, or suffered execution speedup from clock corruption.
- Late-kick violation (\(t_{\text{service}} > T_{\text{upper}}\)): If the timer counts down past \(T_{\text{upper}}\) without receiving a valid refresh, the counter underflows, asserting a hardware reset. A late kick can indicate that the enforcer was starved of compute cycles by interconnect contention, deadlocked on an unhandled exception, or delayed by excessive cache evictions.
- Valid service window (\(T_{\text{lower}} \le t_{\text{service}} \le T_{\text{upper}}\)): A refresh inside the window shows the feed path met its timing condition; task progress still needs a suitable challenge or independent supervision.
If the controller feeds once after each successful tick, the watchdog measures the interval between feeds. Let completion offset \(c_n\) from tick \(n\) start lie in \([c_{\min},c_{\max}]\), and bound tick-phase variation by \(J_{\text{phase}}\). A conservative range for the inter-feed interval is \[T_{\text{feed}}\in[T_{\text{period}}-(c_{\max}-c_{\min})-J_{\text{phase}},\;T_{\text{period}}+(c_{\max}-c_{\min})+J_{\text{phase}}].\] Configure \(T_{\text{lower}}\) and \(T_{\text{upper}}\) around that measured or justified range, while keeping the late-trip response within the plant’s separately derived fault budget. A one-tick execution time alone does not set a watchdog service interval. On the machine, \(T_{\text{period}}\) is the 1 ms tick and the declared 250 μs worst case supplies \(c_{\max}\). The unloaded 135 μs path does not supply \(c_{\min}\), because the enforcer’s shortest branch can complete sooner than its nominal one, so the lower bound comes from the shortest completion the stress campaign records. \(J_{\text{phase}}\) is the jitter of the microcontroller’s tick timer against its reference, measured under the same load. That range would give the machine’s timeout-only self-watchdog the \(T_{\text{lower}}\) it lacks, while \(T_{\text{upper}}\) stays at the chosen 5 ms timeout, so the window tolerates several missed ticks before it trips.
The tolerance is deliberate. One missed tick costs millimeters of travel at aisle speed (section 1.6), well inside the fault budget, and a window that closed on a single late feed would halt a healthy enforcer to save that displacement. A lower bound taken from estimates rather than recorded completions trips on healthy ticks, and an upper bound wider than the fault budget hides the starvation the window exists to catch.
Some designs add challenge-response checkpointing: the feed depends on fresh sensor validation, permission evaluation, and output-state checks rather than a timer ISR alone. Its fault coverage depends on where the challenge is generated, who can compute or replay the response, and whether failed enforcer states can still execute the feed path. The scheme’s coverage is therefore a fault-injection result, not a property of the token.
However, establishing clean boundaries in memory, interconnect arbitration, and watchdog execution liveness solves only the logical and timing half of co-location. When the accelerator and the real-time processor share the same silicon die, the high power dissipated during dense matrix operations conducts directly through the substrate, elevating local junction temperatures and altering transistor switching speeds across adjacent cores.
Thermal and Timing Coupling
Every watt dissipated on a silicon die flows into a single thermal domain. On a shared SoC the matrix units, the application cores, the memory interface, and the integrated regulators all heat one package, and while the permission loop draws a small, steady baseline, the proposal path swings with its workload. The package’s finite conductance to ambient lets peak accelerator activity raise the junction temperature of every circuit on the die.
In physical AI systems engineering, thermal clock throttling is the automatic, hardware-enforced reduction of processor clock frequency triggered when sustained high-power cognitive inference elevates silicon junction temperature (\(T_j\)) past thermal trip thresholds (\(T_{\text{trip}}\)). To prevent damage to the die, the hardware abruptly cuts clock speed. This introduces a lingering, temperature-dependent slowdown (hysteresis) that inflates execution times even after the cognitive workload subsides and can push co-located real-time safety loops past their hard physical deadlines.
The electronics bay in the machine’s base sits at 45\(\,^{\circ}\text{C}\) ambient. Figure 3 shows the developer kit of the module class the machine uses as its application processor; its enclosure reveals neither the heat path nor the clock domains on which the following estimate depends. Suppose the shared SoC were conduction-cooled with a junction-to-ambient thermal resistance of \(\theta_{\text{JA}} =\) 2 \(\text{K/W}\) and a thermal time constant of \(\tau_{\text{th}} =\) 8 \(\text{s}\). The permission loop with idle and memory power draws 4 \(\text{W}\) and holds the junction at 53\(\,^{\circ}\text{C}\). Continuous proposal inference adds accelerator, memory-traffic, and regulator losses, and a total of 26 \(\text{W}\) raises the steady state by \(P\,\theta_{\text{JA}} =\) 52 \(\text{K}\) to \(T_j =\) 97\(\,^{\circ}\text{C}\).
Thermal management alone rejects the shared-die placement for the mobile manipulator. Suppose its permission path ran on a real-time cluster of the application SoC at \(f_{\text{nom}} =\) 1.2 GHz, costing \(N_{\text{cycles}} =\) 300,000 cycles per tick, several times the 54,000 cycles the discrete microcontroller needs unloaded, because a cluster without tightly coupled memory pays cache-refill and coherence cycles on every snapshot and table read. When the junction crosses \(T_{\text{trip}} =\) 95\(\,^{\circ}\text{C}\), the governor drops the cluster to \(f_{\text{throt}} =\) 600 MHz and the enforcer’s cycle stretches to \(t_{\text{exec,throt}} =\) 500 μs, missing the 400 μs deadline by 100 μs (a 25 percent overrun) and far exceeding the declared bound on which every margin above the enforcer was sized. The discrete microcontroller runs on its own clock and rail, beyond the reach of that governor.
The overrun outlasts the load because the governor applies state-dependent hysteresis, restoring 1.2 GHz only after the junction cools to \(T_{\text{clear}} =\) 85\(\,^{\circ}\text{C}\) so that the clock does not oscillate around the trip line. With the policy halted and power back at 4 \(\text{W}\), the junction relaxes toward its 53\(\,^{\circ}\text{C}\) baseline: \[T_j(t) = T_{j,\text{target}} + (T_{j,\text{start}} - T_{j,\text{target}}) e^{-t/\tau_{\text{th}}} \tag{2}\] where \(T_{j,\text{start}}\) is the junction temperature when throttling begins, \(T_{j,\text{target}}\) is the baseline equilibrium, and \(\tau_{\text{th}}\) is the package thermal time constant. Solving equation 2 for the fall from 95 to 85\(\,^{\circ}\text{C}\) gives \(t_{\text{recov}} = \tau_{\text{th}} \ln[(T_{j,\text{start}} - T_{j,\text{target}})/(T_{\text{clear}} - T_{j,\text{target}})] =\) 2.18 s. At the machine’s 1 ms tick, 2,175 scheduled ticks fall inside that cooldown, and a shared-die design would need a separately timed fallback active for every one of them.
A stretched cycle on a shared die can also come from electrical power-rail coupling, on a different timescale. The rail droop of section 1.3 arrives within nanoseconds of an accelerator burst, while thermal coupling builds over seconds as heat diffuses through the package, so a timing fault that tracks junction temperature \(T_j(t)\) rather than transients in \(V_{\text{core}}(t)\) is throttling, not rail collapse.
Testing the Separation
An independence claim, on any placement, is a testable bound on enforcement latency and refusal availability across declared proposal workloads, chip temperatures, and modes. If contention delays a refusal past the physical deadline, the claimed operational separation fails.
↳ Downstream: Adversarial microarchitectural contention tests inform the hardware fault injection suite in Fault Injection on Hardware.
The pass threshold for this experiment comes from the claim under test. The machine’s permission path declares a worst-case execution time \(C_{\text{wcet}} =\) 250 μs and must finish each tick inside its chosen deadline \(D_{\text{enf}} =\) 400 μs, which sits inside the tick \(T_{\text{tick}} =\) 1 ms. Every measured enforcement time \(t_{\text{enf}}\) must respect that chain, as equation 3 states: \[t_{\text{enf}} \le C_{\text{wcet}} < D_{\text{enf}} < T_{\text{tick}} \tag{3}\] The declared bound \(C_{\text{wcet}}\) is the falsification threshold. If any enforcement cycle under test exceeds it, the independence hypothesis is refuted, even when the cycle still meets the deadline, because every margin above the enforcer, from its deadline to the tick that the stopping budget counts, was sized on the declared number. A measured maximum can refute the declared bound but cannot establish it, because a campaign samples only the states it reaches, and on speculative cores timing anomalies can place the true worst case above every cycle recorded.10
Definition 1.2: Worst-case execution time
Worst-case execution time is the formal upper bound on the physical wall-clock duration required for a computational task to execute to completion on a target hardware platform under all admissible inputs, architectural contention states, and operating conditions: \[t_{\text{exec}} \le t_{\text{WCET}} \le t_{\text{deadline}}\]
- Significance: In cyber-physical control loops, correctness depends not only on the numerical validity of the output but also on the instant of its delivery; an optimal control command delivered after the hard physical deadline is functionally equivalent to a total system failure.
- Distinction: Unlike average execution time (\(\bar{t}\)) or high percentiles (\(P_{99}\)), which characterize throughput in soft real-time software, WCET is an absolute deterministic bound that must hold under concurrent memory-bus saturation, cache thrashing, and peripheral interrupt storms.
- Common pitfall: Relying on empirical execution profiling or benchmark averages to establish WCET on modern superscalar or out-of-order processors. Microarchitectural timing anomalies (where a local cache hit or branch prediction success paradoxically increases total execution time) mean empirical testing alone cannot prove the true global upper bound without formal static analysis and hardware contention bounds.
The distance at stake in one tick is small, which makes the bound easy to undervalue. The finished budget in The warehouse mobile manipulator's stopping budget stops the machine in 997.4 mm at the 1.3 m/s aisle speed, 102.6 mm inside the 1.10 m clear distance, and one missed tick adds only \(v\,T_{\text{tick}} =\) 1.3 mm. What placement protects is the bound itself, because \(\tau_{\text{delay}}\) counts exactly one \(T_{\text{tick}}\), and contention without a bound is a stall of unknown length.
Refuting an independence claim requires synthetic workloads engineered to induce worst-case resource contention rather than standard benchmarks that report average throughput. High average utilization schedules memory accesses evenly across clock cycles, making the system look safely underloaded, whereas actual physical interference peaks during synchronized bursts when multiple components demand the same wires at the exact same instant. The test framework must deliberately flood shared memory channels with unaligned direct memory access transfers, thrash shared cache capacity with dirty line writebacks from transformer weight tensors, trigger accelerator matrix multiplications that induce rapid supply rail voltage droops (\(L \cdot di/dt\)), and stream peak-rate sensor telemetry while diagnostic logging writes to non-volatile flash storage. These stress profiles reproduce the transient contention that occurs when a proposal model initiates a large attention pass at the exact microsecond the permission loop reads its sensor snapshot.
Instrumentation records enforcement entry after the sensor snapshot is coherent (\(t_0\)), permission completion (\(t_1\)), and publication of the admitted setpoints to the drive frame (\(t_2\)). These diagnose compute and dispatch, but \(t_2\) is not braking. The frame’s arrival at the drive (\(t_3\)) and measured brake onset (\(t_4\)) complete the physical chain. The endpoint that matters is \(t_4\) measured from the last valid lease renewal, with clock error included. Brake onset must follow within 82 ms, the sum of the lease, one tick, one bus cycle, and the brake-onset delay in Multi-Rate Cadences. The 250 μs computation bound alone is insufficient.
Coupling mechanisms compound across physical, electrical, and logical domains, requiring a full factorial experimental matrix. The evaluation sweeps across cold junction conditions (\(T_j = 35^\circ\text{C}\)) and a warm junction held at or above the throttle trip \(T_{\text{trip}}\) of section 1.5, an idle proposal path versus a saturated proposal engine executing maximum-batch inference, diagnostic circular logging disabled versus streaming at full non-volatile memory bandwidth, and nominal tracking commands versus boundary-violating commands that trigger refusal routines. Testing refusal paths under full stress is mandatory because executing a safe clamping trajectory or entering emergency shutdown involves cold branch paths, non-cached error-handling code, and secondary bus transactions that never execute during steady-state aisle driving.
On the rejected shared-die placement, the pair the matrix exists to catch is specific. A person steps out at a rack end, and the permission path takes its refusal branch toward the stop rung, a cold path through uncached code. At the same instant the application processor, having flagged the same anomaly, starts a diagnostic memory dump over the shared direct memory access fabric. A test suite that exercises refusal and diagnostic logging separately passes both and never measures the pair, because the two subsystems contend for one memory controller only when they run together.
Consider an illustrative \(10^7\)-tick stress trace for the rejected placement, the permission path on a real-time cluster of the application SoC. At one tick per 1 ms, the trace spans 10,000 s, or 2 h 46 min 40 s, before reset or setup time. The cold trace has \(P_{50}=\) 95 μs, \(P_{99}=\) 128 μs, and observed maximum 152 μs; the warm, loaded trace has \(P_{50}=\) 210 μs, \(P_{99}=\) 330 μs, \(P_{99.9}=\) 395 μs, and observed maximum 468 μs. The loaded \(P_{99}\) already exceeds the 250 μs declared bound, and the maximum crosses the 400 μs deadline, so the trace refutes that placement.
Had the trace passed, its ticks would show only that observed interference stayed below the threshold in the tested states. A report gives the latency distribution, observed maximum, sample count, duration, and unexercised conditions, including supply degradation, electromagnetic interference, and aging.11
War Story 1.1: Apollo 11 radar load and computer restarts (1969)
Mechanism: With the rendezvous-radar mode switch out of its LGC position, the radar resolvers were excited by an \(800\text{ Hz}\) supply that was locked in frequency but not in phase with the guidance system’s reference. The coupling data units read the phase difference as antenna motion and sent a continuous stream of counter pulses to the computer. NASA’s mission report, section 16 gives a maximum of 12,800 pulses per second from the two data units, each taking one \(11.7\,\mu\text{s}\) memory cycle, or 15 percent of computer time, while the computer normally ran at about 90 percent of capacity during powered descent. Repetitive jobs that could not finish were scheduled again, each reserving another memory area, until the Executive ran out of them.
Impact: Five Executive overflow alarms (1201 and 1202) occurred before the low-gate phase of the descent. The mission report states that the performance of guidance and control functions was not affected.
Response: Each alarm initiated a software restart, and the AGC resumed its essential guidance and display work according to priority while the crew and mission control assessed the alarms and descent continued. Later missions zeroed the data units whenever the switch was not in LGC, removing the spurious pulses at their source.
Systems lesson: A peripheral can consume execution capacity at the same moment guidance needs it. Modern placement tests should account for asynchronous peripheral load and verify that the critical path and its recovery behavior remain within their declared timing bounds.
Apollo 11’s Executive alarms and the Mars Pathfinder priority inversion (examined in Moving Commands on Time) illustrate two distinct forms of real-time resource contention. In Apollo 11 a peripheral consumed execution capacity; in Pathfinder a low-priority task held a resource that a high-priority task needed.
A failed interference test identifies where the assumed separation may collapse: cache, memory arbitration, thermal clocks, power, or output queues. Further instrumentation localizes the mechanism before the team reserves, partitions, or relocates the shared resource.
Silicon Isolation Decisions
A refuted independence claim leaves four remedies, each priced in a different currency (table 1): static reservation, spatial partitioning, workload reduction, and physical relocation to separate silicon.
| Remediation Strategy | Target Coupling Layer | Implementation Mechanism | Systems Engineering Trade-Off | Residual Vulnerability | Physical Silicon Migration Trigger |
|---|---|---|---|---|---|
| Static Reservation | CPU execution time, OS scheduling preemption | Fixed-priority preemptive scheduling (SCHED_FIFO), static time-slot slicing (ARINC 653) |
Idle capacity (\(20\text{--}40\%\) throughput loss); unallocated cycles wasted during quiescent states | Micro-architectural contention beneath OS visibility (shared L3, DRAM bank conflicts) | Memory bus or cache stalls violate real-time deadlines despite \(100\%\) CPU core allocation |
| Spatial Partitioning | Cache capacity thrashing, DRAM bus contention | Cache way-partitioning (ARM MPAM, Intel CAT), MPU/IOMMU domains, AXI QoS rate limiters | Structural fragmentation; reduced L3 working set for perception (\(+15\text{--}30\text{ ns}\) DRAM access penalty) | Unpartitionable DRAM controller queues, row-buffer locality reordering, PDN voltage droop | Matrix engine current steps (\(\Delta I > 10\text{ A}\)) trigger rail droop, or DRAM queue stalls exceed slack |
| Workload Reduction | Memory throughput saturation, peak thermal power | Model quantization (INT8/FP8), context downsampling, halved policy rate | Degraded policy tracking fidelity; increased control phase lag near dynamic stability boundaries | Common clock distribution trees, shared Power Management Units (PMU), unified reset domains | Model accuracy drops below operational limits, or shared clock/reset faults propagate across domains |
| Physical Relocation | Shared power, clock, reset, or unbounded controller interference | Separate enforcer MCU with independently assessed power, clock, memory, and sensor/actuator path | More board area and an external communication budget | Harness, transceiver, and common-supply faults remain | Choose when the shared die cannot provide evidence for required timing and fault containment |
Static reservation preserves the physical placement of workloads on the shared chip while allocating a guaranteed fraction of a shared resource to the critical path. Reservation assigns dedicated execution slots or fixed memory bus bandwidth to the critical loop. The cost of reservation is idle capacity. Had the machine’s permission path shared the die, an arbiter slot sized to its snapshot read’s worst case would stay reserved in every 1 ms tick, unavailable to the policy model even while the base is parked and no stop is under way.
Spatial partitioning divides physical hardware structures so that workloads no longer compete for the same physical cells.12 Cache way-partitioning, core pinning, and dedicated memory channels establish hardware barriers within a single die. The cost of partitioning is lost flexibility and structural fragmentation. Splitting an eight-megabyte shared cache into a six-megabyte partition for the policy model and a two-megabyte partition for the permission path’s tables prevents cache line evictions across the boundary. However, it permanently restricts the working set of the learned model, forcing higher DRAM access rates and increasing mean memory latency by tens of nanoseconds on every inference step.
Workload reduction resolves contention by shrinking the computational or memory footprint of the non-critical workload until its peak demand falls within the uncontended margins of the shared system. An engineer might downsample telemetry logging from \(1\text{ kHz}\) to \(100\text{ Hz}\), quantize policy weights from sixteen-bit floating point to eight-bit integers to cut memory bus traffic in half, or lower the chunk policy’s rate. The cost of reduction is degraded functional authority. On the mobile manipulator the lease bounds the last lever. A chunk period longer than the 60 ms lease of Multi-Rate Cadences lets proposals expire between renewals, so the policy’s rate must stay above 16.7 Hz even before renewal jitter is counted, which leaves little room below its 20 Hz. Weight quantization can degrade the policy’s accuracy near the grasp tolerance of the conveyor pick.
If reservations, partitioning, and workload reduction cannot establish the required response bound, physical relocation is a candidate remedy. It reaches couplings that reservation and partitioning cannot, such as a shared reset or clock domain that lets one fault halt both paths; a partitioned SoC that claims to contain such a fault has to show it by design and injection tests. A separate controller removes these common-cause paths only as far as its power, clock, reset, sensor, and actuator independence are designed and tested, and the integrity target it serves rests on the application’s safety case, not on a count of packages. An external controller also adds transport and common-supply obligations, so its sensor and actuator path must be measured as a whole.
Relocating a workload to separate silicon is rarely a localized modification. It triggers a cascade of placement and communication adjustments across the entire machine, and the mobile manipulator shows where that cost lands. The \(T_{\text{bus}}\) term of Multi-Rate Cadences does not change, because drive commands cross the EtherCAT fieldbus wherever the permission path runs. What relocation adds is the proposal’s crossing from the application processor to the microcontroller over a board-level link instead of an on-die mailbox. That delay enters no new term of \(\tau_{\text{delay}}\). A proposal that arrives late spends part of the 5 ms renewal-jitter allowance that the lease rule of One coupled budget for the anatomy places in the gap between the 50 ms chunk period and the 60 ms lease. Host memory contention draws on the same allowance, so the most it may stretch the chunk policy’s 40 ms inference is 12.5 percent (The Cost of Ingestion). The link’s worst case, measured under load, must therefore fit inside what that contention and the proposer’s own jitter leave of the allowance.
The harder decision is where each sensor attaches. The machine keeps the encoders and the IMU on the microcontroller’s side of the link, because those are what the permission path reads every tick, and it leaves the cameras and the lidar on the application processor, because the learned models consume them and the permission path reads only the proposal’s evidence epoch. Moving a camera interface to the microcontroller would offload the main chip, but it would force the learned models to receive their observation stream over the slower external link and add to the observation age that Sensor Perception measured. Every re-placement decision propagates through the physical interconnects, shifting bottlenecks from on-die memory arbiters to board-level pin counts, wiring runs, and serial buses.
Hardware selection in physical AI cannot be decoupled from the mechanical dynamics of the embodiment. While conventional benchmark suites evaluate edge accelerators on isolated metrics such as theoretical peak TOPS or raw batch inference throughput, embodied systems operate under closed-loop energy constraints where compute latency and physical motion are tightly coupled (figure 4). Evaluating edge accelerator configurations across a continuum—from ultra-low-power edge TPUs to multi-hundred-watt workstation GPUs—reveals an inverted-U Pareto efficiency frontier.
Under-provisioned accelerators (such as a \(4\,\text{TOPS}\) edge TPU or an unaccelerated CPU) incur long vision-language-action inference latencies (\(>200\,\text{ms}\) per step), stretching a pick-and-place task to over three minutes (\(T_{\text{mission}} = 215\,\text{s}\)). During this entire duration, the robot’s base electronics, sensors, motor holding currents, and thermal management systems continuously draw static standby power (\(P_{\text{static}}\)). Consequently, the static energy penalty (\(P_{\text{static}} \times T_{\text{mission}} = 23.65\,\text{kJ}\)) dominates total mission consumption, yielding poor net system efficiency (\(40.8\,\text{tasks/MJ}\)). Conversely, over-provisioning edge hardware with desktop-grade GPUs (\(300\,\text{W}\) TDP) reduces model execution time to \(55\,\text{ms}\), but the system hits the robot’s kinematic velocity ceiling (\(v_{\text{max}} = 0.8\,\text{m/s}\)). The marginal time savings (\(27.5\,\text{s} \to 24.0\,\text{s}\)) fail to compensate for the massive dynamic chip dissipation (\(P_{\text{chip}} \times T_{\text{mission}}\)), causing total energy per task to surge to \(9.84\,\text{kJ}\). The optimal co-design window emerges at the balance point (exemplified by the \(50\,\text{W}\) Jetson AGX Orin), where model inference is sufficiently amortized to prevent static chassis drain without inducing unnecessary dynamic power penalties.
Hardware Allocation
The mobile manipulator now has a placement for its permission path, a deadline, an unloaded execution estimate, and a test that could refute the separation. None of it reaches the release decision unless it is written down with the conditions under which it holds. The placement record consumes the enforcement record of The Enforcement Record, from which it takes the permission path’s declared worst-case execution time and deadline, and adds the evidence that the declared bound holds under load on the silicon where the path runs: execution domain, privilege, period, unloaded and loaded observed maxima, arbitration settings, droop and thermal trips, mitigations, and a digest of the test ledger. The loaded maxima decide the independence claim, which stands only while every one of them stays at or below the declared worst-case execution time under the declared load; as Moving Commands on Time argued for any observed latency distribution, a sample maximum is evidence about that bound and never replaces it.
The record has two parts. A compact header holds the configuration values and the observed maxima, and it carries the digest of an external test ledger. The ledger holds what does not fit a fixed layout: inventories, the latency distributions and sample counts behind each maximum, test conditions, mitigations, and open limits.
↳ Byte layout: The placement record’s field offsets and sizes are given in Placement record.
For the mobile manipulator, the header records the 1 ms tick, the 400 μs deadline, and an illustrative 135 μs unloaded estimate that is not yet a measured maximum. The header’s mitigation flags include the permission path’s isolated rail. Its loaded field stays empty, because no loaded test ledger exists for this machine yet. Filling it takes the campaign of section 1.6 run on the machine itself, with the discrete microcontroller ticking while the application processor runs peak proposal inference with full logging, at warm junction, with the refusal branch exercised in the same trials. Every loaded maximum that campaign records must stay at or below the declared 250 μs, and brake onset, timed from the last valid renewal, must follow within 82 ms. Until those measurements exist, the permission-independence claim that Deployment Release weighs rests on a declared bound with nothing behind it, and a verdict that needs the claim cannot be granted on it.
The external ledger inventories memory channels, cache partitions, DMA engines, power rails, and their competing claimants. For each it records the arbitration configuration and measured interference. For the permission rail it records the 2 s hold-up, which the feed-loss injection of Fault Injection on Hardware exercises, and a claimant list that carries the isolation claim itself, naming the microcontroller, the drive logic, and the spring-brake coils and no proposal-side load.
An interference measurement needs its operating state: model and firmware hashes, clock policy, enclosure and thermal interface, ambient and junction temperatures, power rail, arbitration settings, traffic generators, sampling rates, instrumentation, sample count, and test duration. The ledger carries those fields beside the latency distributions and unexercised states they qualify, and the header’s observed maxima summarize those distributions. The \(10^7\)-tick trace of section 1.6 spans only 2 h 46 min 40 s of one machine’s operation, too short to reveal any interference rarer than about one tick in \(10^7\), and thin evidence for a fleet.
Every placement entry requires a predeclared falsification criterion written before the test executes. This criterion specifies the physical threshold that constitutes architectural failure, such as any enforcement cycle longer than the declared 250 μs worst case during simultaneous neural accelerator bursts and DMA transfers. The record logs the measured distribution against this criterion and states whether the tested claim survived or was refuted. When interference violates the target budget and forces an engineering change, the record details the exact mitigation applied, such as enabling memory bus bandwidth throttling on the neural accelerator. The entry must also document the boundary conditions that were not exercised during the test run, such as operating states above the 45\(\,^{\circ}\text{C}\) bay ambient, multi-bit memory error injection, or simultaneous sensor transceiver fault states.
Adversarial Verification consumes the recorded arbitration points and shared power rails as the target map for hardware and software fault injection, deliberately corrupting shared structures to test whether fault propagation matches the predicted failure domains. Deployment Release incorporates the recorded loaded execution bounds and unexercised test limits into the formal safety case, transforming empirical measurements into evidentiary arguments for operational deployment.
↳ Downstream: Hardware placement safety cases provide evidentiary warrants for production deployment in Safety Cases and Claims.
Systems Perspective 1.1: When the fallback is a step
Question: Which rows of the stopping budget survive when the fallback is a step rather than a stop?
With placement protecting the one tick that \(\tau_{\text{delay}}\) counts, Part III’s stopping budget is complete for a machine whose fallback brakes it to rest, as the mobile manipulator’s does. A walking humanoid has no rest state to brake into mid-stride, so its fallback is a recovery step, and each row of the budget has to be rechecked against that maneuver.
- Observation age survives. Age still becomes displacement (Photons to Spatial Claims), and on a turning head the error is angular. An uncalibrated 10 ms camera–IMU offset during a 1.5 rad/s head turn misplaces a point 5 m ahead by 7.5 cm.
- Commitment changes. The latched stop suffix of Behavior at the Seam becomes a validated step, admitted with the trajectory as before but ending on a new foothold rather than at rest.
- The inset changes. A balance constraint can have relative degree above one, the case that Safe Sets as Conditional Permission resolves for raw clearance by folding velocity into the barrier. The balance barrier needs that same construction, and the tracking inset then applies to the barrier it produces.
Fallacies and Pitfalls
Placement determines the physical resources shared by complex proposal algorithms and critical enforcement routines. Traces of normal operation cannot prove these systems are truly independent when rare physical events, such as power voltage transients, clock faults, memory contention, or a stalled producer, stress those shared hardware pathways.
Fallacy: Logical privilege rings protect real-time enforcers from shared-rail electrical droop.
Suppose the mobile manipulator placed its lockstep permission core on the accelerator’s supply in one SoC. The droop of section 1.3 then stretches the propagation delay of every gate on that shared rail by 39 percent, the permission core’s gates included, whatever priority its software holds. A core whose clock leaves less timing slack than that stretch misses setup time on its critical paths. Both lockstep cores see the same sag, so they can latch the same wrong value and the comparator stays silent, or the brownout detector resets the enforcer while the base is moving. Access control decides which addresses the enforcer may write; it has no term for the voltage at its flip-flops. Placement qualification must bound the droop with decoupling capacitance or \(di/dt\) slew limits and show the timing margin under the declared current step, or remove the shared supply. The machine removes it, running its permission path on the isolated, held-up rail of Which Budget Binds First, which the accelerator does not draw on.
Pitfall: Sharing Phase-Locked Loops or reset controllers across proposal and permission domains.
Had the machine’s application processor and permission domains shared one PLL and one reset controller (Hardware Isolation), a thermal clock fault or a watchdog reset of the stalled policy cluster could have halted the permission domain as well. The nominal block diagram shows separate processors, but both depend on fragile silicon infrastructure that a single event can silently disable. Trace clock and reset trees through the physical implementation to ensure independent resources for continuous enforcement. The base and arm drives also need a defined hardware transition to a safe state when permitted commands stop arriving; restarting the application processor cannot be the only recovery mechanism.
Fallacy: Low average memory bus utilization guarantees bounded read latency for real-time safety tasks.
Suppose the machine’s permission path read its sensor snapshot from shared DRAM, where a modest average bus use hides camera and logging bursts that coincide with its reads. The loaded read of 1.1 fits the sense allowance only while a single backlog precedes it, and nothing in the average bounds how many backlogs do. An assumed 320 μs read tail leaves 80 μs of the 400 μs deadline, less than the 95 μs that constraint evaluation and the drive-frame write need, so this configuration fails. A coherent TCM snapshot may remove the shared-DRAM read, provided its DMA completion, capacity, and local transfer time are qualified. Average utilization alone cannot establish the deadline.
Pitfall: Using blocking software mutexes across inter-processor communication boundaries with untrusted neural policies.
The chunk policy crashes halfway through publishing a chunk of 16 setpoints. An enforcer that had to acquire a lock held by that producer would freeze with it, which is why the boundary contract of section 1.4 admits no blocking synchronization. Its seqlock lets the enforcer reject a torn record without waiting and proceed with its fallback, but only if the retry behavior itself has a bound. The communication contract must specify timing behavior for interrupted publications as rigorously as it specifies the payload format.
Summary
Physical AI platforms place learned proposers on application processors and accelerators and the permission path on a bare-metal microcontroller, but that allocation alone does not separate them, because privilege boundaries govern none of the physical resources the two paths still share. Partitioning is therefore a design hypothesis until a loaded test fails to refute it.
Physical isolation must be supported by measured bounds on memory queues, crossbar arbitration, power transients, and the actuator path. A placement result is an empirical claim scoped to the measured conditions. A passing placement record shows that under a declared stress workload executing at a specific ambient temperature, with fixed clock governance policies and documented fault domains, the shared physical substrates of the die did not cause the untrusted path to violate the timing constraints of the safety path. On the warehouse mobile manipulator, drawing the chunk policy and the permission path in separate boxes on an architectural diagram would not have isolated them from a shared voltage regulator, a common memory crossbar, or a heat sink clamped across both processor clusters. Had the accelerator executing the learned proposal drawn peak current while every memory channel serviced direct memory access transfers, the resulting voltage drop would have reached a co-located permission core within nanoseconds and the heat within seconds; the machine places that core on separate silicon for this reason. The boundaries drawn during system design exist on the silicon only when verified against these shared electrical, thermal, and architectural limits.
The placement record carries that evidence forward. It takes the permission path’s declared worst-case execution time and deadline from the enforcement record and adds the loaded observed maxima for the declared test conditions, the shared-resource queue and power headroom, the common power, clock, reset, and memory failure paths, and the conditions left untested. The machine’s own record holds a 135 μs illustrative unloaded estimate against the declared 250 μs worst case inside the 400 μs deadline and still awaits a loaded ledger, which must name the arbitration policy, clocks, sample count, and physical endpoint before any release claim can rest on it.
With the completion of this placement analysis, the stages of Part III form a single machine that must be measured as one. The observation contract, the belief record, the intent lease, the trajectory record, the enforcement record, and the placement record form one chain, each a claim with conditions, so the stages that produce them cannot be treated as independent software modules communicating across an idealized channel. The logical boundaries established across earlier chapters provide the structure needed to reason about authority and failure propagation, but physical reality governs their execution. A memory controller follows configured arbitration and row scheduling, which need not reflect software privilege, and an interrupt can arrive during either workload. When the permission path admits a base velocity setpoint, the action is the last step of a pipeline whose timing, power, and memory footprints must be measured together on one chassis. The earlier chapters assumed that learned workloads could not delay the permission path that principle \(\ref{pri-vol4-proposal-permission}\) demands, and placement turns that assumption into a falsifiable threshold with a stress test and a record behind it. Part IV inherits the three things the chain leaves open: the spare of the finished stopping budget, the unsettled premise of the site crossing rule, and the placement record’s empty loaded field.
Key Takeaways: Heterogeneous silicon placement and memory contention
- Logical separation is a hypothesis about hardware: Privilege rings govern which addresses a task may touch, not where its requests wait, which rail feeds it, or which clock and reset tree it depends on. An independence claim stands only after a loaded test fails to refute it.
- Averages cannot bound a deadline: On the rejected shared die, one backlog of weight-tile, camera, and telemetry transactions stretches the permission path’s snapshot read from 0.32 to 28.564 \(\mu\text{s}\). More than 5 such backlogs cross the declared worst-case execution time, which nothing in the average forbids. Transaction size, queue policy, and bypass behavior decide the wait, so the bound belongs to the memory controller.
- Coupling runs on three timescales: Rail droop acts in nanoseconds, queueing in microseconds, and thermal throttling persists for seconds after the load stops. In the rejected shared-die placement, a halved clock stretches the enforcer’s 250 μs cycle to 500 μs for 2,175 ticks.
- The declared bound is the falsification threshold: Any loaded cycle above the declared worst-case execution time refutes independence even when the deadline still holds, because every margin above the enforcer was sized on that number. A passing trace supports only the states it tested.
- A refuted claim forces a priced choice: Reservation costs idle capacity, partitioning costs working set, workload reduction costs policy authority, and relocation adds a board-level link whose worst case must fit inside the lease’s renewal-jitter allowance.
- The placement record scopes what was shown: It takes its declared worst-case execution time and deadline from the enforcement record and adds loaded observed maxima, mitigations, and untested conditions. On the mobile manipulator the loaded field is still empty, so the independence claim reaches release review without evidence.
What’s Next: From compute placement to supervisory intervention
Footnotes
Inductive Voltage Droop Mechanics: In high-performance VLSI design, inductive voltage droop (\(\Delta V = L \cdot dI/dt\)) occurs across the parasitic inductance (\(L\)) of power distribution package pins and on-chip power grids when large numbers of logic gates switch simultaneously. If the local rail voltage drops below the minimum operating threshold (\(V_{\min}\)), digital flip-flops suffer setup-time violations, metastabilities, or brownout resets that disrupt deterministic real-time execution.↩︎
Heterogeneous System-on-Chip Architecture: Heterogeneous SoCs integrate multiple distinct compute domains onto a single silicon substrate, combining high-throughput application processors and neural accelerators with deterministic real-time safety islands. Although memory management units isolate software address spaces across domains, hardware subcomponents continue to share power distribution networks, clock distribution trees, and DRAM controllers. Under peak cognitive burst workloads, uncoordinated resource contention across these shared physical structures can degrade real-time response times beyond critical physical deadlines.↩︎
Dual-core lockstep: Two cores can execute matching instruction streams while comparator logic detects selected divergences. Detection coverage, delay, shared-rail vulnerability, and diagnostic response depend on the device and the application safety case, so lockstep is one element of an ASIL D argument, not the whole of it.↩︎
PREEMPT_RT: A Linux real-time configuration reduces some scheduler latencies, but its tail depends on kernel, drivers, workload, and shared hardware. No portable sub-\(50\,\mu\text{s}\) or bare-metal comparison follows without measurement.↩︎
Hardware memory protection and IOMMU domains: Physical silicon structures that enforce spatial address isolation. While an on-chip Memory Protection Unit (MPU) confines CPU thread execution to bounded memory segments, an Input-Output Memory Management Unit (IOMMU) translates and polices direct memory access (DMA) requests from peripheral bus masters (such as PCIe neural accelerators or camera deserializers), physically preventing an untrusted device from corrupting safety state buffers.↩︎
Avionic Multicore Certification: FAA Advisory Circular AC 20-193 and CAST-32A formalize multicore interference analysis for safety-critical avionics. The guidance requires bounding worst-case execution times by demonstrating that shared caches, interconnect crossbars, and memory controllers cannot cause critical tasks to overrun execution deadlines. Without hardware-enforced quality-of-service partitions or dedicated interconnect channels, cross-core memory traffic invalidates a single-core timing bound.↩︎
Cache coloring and way partitioning: Techniques that divide shared cache capacity among competing tasks. Software cache coloring partitions physical page allocations so that addresses from distinct processes map to non-overlapping cache sets, while hardware partitioning mechanisms (such as ARM Memory System Resource Partitioning and Monitoring, MPAM, or Intel Cache Allocation Technology, CAT) dynamically restrict the cache ways an untrusted neural accelerator can allocate, preventing line evictions in critical safety threads.↩︎
DRAM bank conflicts: A different row in the same active bank may require precharge and activation before column service. Actual bank mapping, overlap, and scheduling determine the added delay.↩︎
Alpha-Power CMOS Delay Model: The model uses \(t_{\text{prop}} \propto V/(V-V_{\text{th}})^\alpha\) with \(V_{\text{th}}=0.42\text{ V}\) and \(\alpha=1.3\). It relates supply voltage to the switching delay of one gate; setup failures and brownout behavior come from device characterization.↩︎
Microarchitectural Timing Anomalies: In out-of-order and speculative processing architectures, a local performance improvement, such as a cache hit or accelerated branch evaluation, can alter instruction scheduling and resource contention, resulting in a net increase in global execution time. This counterintuitive behavior invalidates simple compositional worst-case execution time analysis. Safety verification needs a justified bound for the declared workload, arbitration, temperature, and fault conditions; measurements alone and a non-speculative core alone do not establish a hard end-to-end bound.↩︎
Silicon Substrate Aging and Timing Margins: Long-term substrate degradation mechanisms, such as Negative Bias Temperature Instability (NBTI) and Hot Carrier Injection (HCI), gradually increase transistor switching delays over years of operation. When coupled with transient supply-voltage droop under sudden multi-core current steps (\(\frac{di}{dt}\)), aged silicon experiences propagation delays that exceed design margins. In hard real-time systems, these physical timing degradations cause intermittent race conditions and missed scheduler deadlines that never manifest during clean bench qualification.↩︎
Automotive Freedom From Interference: ISO 26262 Part 9 formalizes Freedom From Interference (FFI) as the absence of cascading faults between software elements with different safety integrity levels. In heterogeneous robotics SoCs, achieving FFI requires hardware-enforced spatial partitioning (memory protection units, IOMMU page tables) and temporal partitioning (fixed-priority bus arbitration, TDMA interconnects) to prevent unprivileged perception models from degrading safety-critical enforcement tasks.↩︎

