The Cognitive Brain

The Cognitive Brain

Isometric blueprint of the Brain level, the application processor that runs the learned proposers: heterogeneous SoC, NPU tensor systolic array, stacked HBM memory dies, high-speed weights streaming bus, and 3D action chunking token lattice.

Purpose

Why can the most capable model on a robot revise its plan only a few times a second?

A robot operating among humans must interpret dynamic environments and propose actions far beyond the rigid scripts of classical automation. High-capacity learned models supply this semantic reasoning by grounding multimodal sensory streams in vast representations acquired during offline training. Yet this expressive power is paid for in high memory traffic and severe computational latency. Every forward evaluation of a multi-billion-parameter network streams weights across a memory interconnect of finite bandwidth. Consequently, the deliberative model that understands the open-world scene best is inevitably the component that revises its decisions least frequently.

To keep actuators continuously supplied between deliberative updates, the Brain must emit prospective action chunks—temporal sequences of future setpoints. However, these proposed setpoints age with the stale observation that generated them, and a proposal that arrives on time can still be physically disastrous. High prediction confidence does not entitle a statistical model to command the physical machine. The Brain generates candidate behaviors across the proposal boundary, while the deterministic Nervous System evaluates each chunk against current sensor telemetry and safety envelopes before allowing low-level commands to cross the causal boundary.

↰ Prerequisite: A learned proposal stays unprivileged until the permission path admits it, the third law of The Four Bedrock Laws.

Learning Objectives
  • Compare classical analytical pipelines and learned representations on coverage, robustness, and verifiability
  • Calculate the bandwidth bound on a model’s fresh-evaluation rate from its weight footprint and sustained memory bandwidth
  • Estimate key-value cache growth from camera count, tokens per step, and retained history
  • Explain why action chunking multiplies setpoint supply without making proposals any fresher
  • Apply replanning period, inference latency, and target drift to size an action chunk for a moving target
  • Diagnose proposer failures caused by hallucination, miscalibrated confidence, perceptual chattering, and expired intent
  • Evaluate why these proposer limits require a permission path that can refuse any proposal

The Embodied Brain

The warehouse mobile manipulator can meet all electrical and thermal budgets yet still fail at a task as ordinary as taking an unfamiliar ceramic mug from the moving conveyor at the pick station and handing it to the coworker at the packing station. The motor drives remain well within their continuous torque ratings, the control rail holds its voltage at full charge, and the joint velocities remain below physical limits. Yet the machine drops the object or crushes the ceramic rim because low-level motor physics cannot determine what the object is, where its structural affordances lie, or how to reach around clutter. Unlike an advisory display, whose output can be reviewed before it affects the world, a cognitive proposal here can reach moving mass with no human in between. A mistaken proposal can lead to unsafe kinetic energy if the execution path accepts it, so permission must be checked independently. The central dilemma of embodied cognition is that high-capacity neural models are inherently stochastic, computationally expensive, and prone to inventing objects that are not there, yet they must orchestrate physical bodies whose collisions cannot be recalled (the first law, The Four Bedrock Laws).

The Brain is therefore an unprivileged deliberation engine that may only propose (\(\ref{dfn-boundary-proposal-boundary}\)). This architectural separation organizes the machine into a two-speed brain architecture (figure 1). At the deliberative tier, high-capacity foundation models perform slow cognition at \(1\text{--}50\text{ Hz}\), parsing open-world multimodal streams to emit structured proposals. At the real-time tier, deterministic safety monitors execute at \(1000\text{ Hz}\) on isolated safety silicon, evaluating candidate proposals against physical limits before granting execution permission (\(u^*\)), or asserting an asynchronous refusal that triggers deliberative state resynchronization. On the mobile manipulator it holds two learned proposers, an intent model that revises goals at \(1\text{--}5\text{ Hz}\) and a chunk policy that emits short sequences of future setpoints at \(10\text{--}50\text{ Hz}\), and the permission path may refuse anything either one proposes. That protection holds only while the limits, the state estimate, and the fallback it relies on remain valid.

Two-speed brain architecture showing unprivileged deliberative planning at 1 to 50 hertz crossing the proposal-permission privilege boundary to 1000 hertz real-time safety enforcement with refusal feedback.
Figure 1: Two-speed brain architecture: An unprivileged deliberative realm running at \(1\text{--}50\text{ Hz}\) and a real-time realm running at \(1000\text{ Hz}\), separated by the proposal-permission privilege boundary. The independent safety enforcer checks sensed state and modeled limits on each control cycle before granting execution permission; a refusal path feeds state resynchronization back to the deliberative policy.

A learned proposer earns its place by doing what hand-authored geometry cannot. It recognizes an object no one modeled from an open-vocabulary instruction, tolerates glare and clutter that break hand-written filters, and anchors what it recognizes to metric coordinates the arm can reach. Each of these capabilities comes from a model whose size also sets how often it can run. Figure 2 shows the tension across workload classes, with the models that carry the most parameters updating least often, far below the rates at which the Body’s loops close. The chapter’s question is therefore what a learned proposer can deliver on time, and at what cost in memory, latency, and freshness.

Log-log scatter of parameter counts and example update rates for different workload classes, with shaded control and deliberation rate bands.

Figure 2: Cadence regimes: Parameter counts and example update rates for different control, vision, and embodied workloads. The points are illustrative workload regimes rather than runs on a common device, and parameter count alone does not determine update rate.

Section 1.2 shows why pipelines built on hand-authored geometry and state machines lose coverage outside the factory cell, and section 1.3 shows what learned representations recover and what they cost in tokens and memory traffic. Section 1.4 turns that cost into the machine’s numbers: the memory wall that bounds how often its intent model can run, the action chunking that keeps the arm supplied between evaluations, and the drift that bounds how long each chunk stays true. Section 1.5 then takes up the ways a proposal can arrive on time and still be wrong, each of which leaves the permission path with a proposal it must be able to refuse.

The Classical Wall

For nearly a century, automated control and robotics approached cognition as an explicit geometric and analytical physics problem (Wiener 1948; Kalman 1960; Craig 2005). Under this classical paradigm, engineers constructed mathematical representations from first principles.

Wiener, Norbert. 1948. Cybernetics: Or Control and Communication in the Animal and the Machine. John Wiley & Sons.
Kalman, Rudolph Emil. 1960. “A New Approach to Linear Filtering and Prediction Problems.” Journal of Basic Engineering 82 (1): 35–45.
Craig, John J. 2005. Introduction to Robotics: Mechanics and Control. 3rd ed. Pearson Prentice Hall.

Engineers used analytical kinematics and differential dynamics to map desired end-effector motions into actuator torques using closed-form equations.1 Perception relied on matching hand-crafted 3D Computer-Aided Design (CAD) models against sensor point clouds using geometric alignment algorithms, such as iterative closest point (ICP) or random sample consensus (RANSAC). High-level task logic was authored as explicit Boolean state machines that dictated deterministic transitions between operational phases, such as approaching, aligning, grasping, and retracting.

Articulated industrial robots executing automated spot welding on an automotive chassis within a safety-caged factory cell.
Figure 3: Industrial spot-welding workcell: Multiple six-degree-of-freedom articulated robots execute preprogrammed spot welding paths on stamped automotive chassis at the BMW Leipzig plant. Operating in a calibrated factory environment with submillimeter repeatable fixtures and safety interlock cages, classical controllers provide deterministic cycle timing and formal stability proofs. (Credit: BMW Group / Wikimedia Commons, CC BY-SA 2.0 de).

Where the operational environment is strictly structured and bounded (figure 3), classical methods can provide dependable behavior within their modeled conditions. On an automated automotive assembly line, an industrial robot welds stamped steel panels where every workpiece arrives at an exact coordinate within sub-millimeter tolerances. Ambient lighting is tightly regulated by industrial fixtures, and human workers are excluded by physical interlock safety fences. Within the causal boundary defined in The Causal Boundary, this controlled regime permits explicit models, bounded servo timing, and stability analysis under stated assumptions. Engineers can verify selected trajectory and torque constraints and assess controller software coverage against applicable standards without collecting task demonstrations.

Yet when an embodied machine leaves the structured factory cage to operate in unstructured environments (figure 4), this paradigm hits the classical wall (Brooks 1991; Levine et al. 2018). The physical world is not a sterile CAD drawing; it is continuous, deformable, and endlessly variable.

Brooks, Rodney A. 1991. “Intelligence Without Representation.” Artificial Intelligence 47 (1–3): 139–59. https://doi.org/10.1016/0004-3702(91)90053-M.
Levine, Sergey, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen. 2018. “Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection.” The International Journal of Robotics Research 37 (4-5): 421–36.
Tracked mobile robot traversing irregular rubble and wooden obstacles in an unstructured search-and-rescue test arena.
Figure 4: Unstructured search-and-rescue robot: A tracked mobile robot navigates an obstacle course with irregular inclines, loose debris, and visual occlusions. In such unmapped environments, preauthored geometric models and rigid state machines can lose coverage; learned policies may help infer spatial affordances and traversability from sensory streams. (Credit: TU Darmstadt Rescue Robotics / Wikimedia Commons, CC BY-SA 3.0).

The first breakdown is the combinatorial explosion of geometry. In a commercial logistics warehouse handling hundreds of thousands of retail goods, authoring CAD models, mesh templates, and grasp points for every bottle, wrapped package, and flexible garment is impossible. Discretizing orientations, frictional properties, and contact permutations across overlapping, deformable items creates a combinatorial space that hand-crafted templates cannot cover.

The second breakdown arises from sensory brittleness. A hand-crafted edge-and-segmentation pipeline may rely on filters, Harris corners, or surface-normal thresholds. Under changed sunlight, reflections, or dust, a required feature can disappear. In that pipeline, a missing edge can make planar segmentation bridge two objects and propose a grasp in empty space.

The third breakdown is the engineering bottleneck of hand-authored logic. The machine’s operational capability is bounded by the engineer’s ability to anticipate every physical edge case. When an unexpected variation occurs, such as a bent bracket, a scuffed label, or an unfamiliar obstacle, the hand-authored rulebook contains no matching branch, causing the machine to halt, drop its payload, or collide.

Learned representations answer the first two breakdowns directly and ease the third, because they draw coverage from data rather than from hand-authored models.

Learned Representations

Asked to take the unfamiliar ceramic mug from the conveyor, a classical pipeline needs a CAD model, a feature threshold, and a state-machine branch that nobody wrote for this object. A learned model trained on many graspable objects can propose the handle as a grasp region from pixels and a sentence of instruction. Physical AI extends classical pipelines with such neural representations wherever hand-authored models lose coverage, and four pillars support the shift.

The first pillar is open-world semantic generalizability. Classical controllers are geometrically precise but semantically blind: an inverse kinematics solver can expertly drive an end-effector to coordinate \((x, y, z)\), yet it possesses zero understanding of whether that location contains a sturdy mug or a fragile glass, or how surrounding objects relate to human tasks. Modern vision-language-action (VLA) foundation models project open-vocabulary natural language instructions and high-resolution camera streams into shared semantic embeddings (Driess et al. 2023; Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023; Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023; Kim et al. 2024). When commanded to “wipe the coffee spill from the counter and stow the mug in the drying rack,” a trained model may propose the mug handle and spill boundary as task-relevant regions without a new CAD model.

Driess, Danny, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, et al. 2023. “PaLM-e: An Embodied Multimodal Language Model.” International Conference on Machine Learning (ICML), 8469–88.

The second pillar is empirical sensory robustness. Hand-authored visual thresholds can flip under changing illumination or texture. Learned representations can tolerate some of these changes when training and testing cover them, but their responses are neither necessarily smooth nor bounded under shift.

The third pillar is the ability to handle complex multi-contact and deformable dynamics. Many physical manipulation tasks (folding textiles, routing flexible wire harnesses through automotive door assemblies, or scooping granular media) involve non-smooth contact mechanics, variable friction, and infinite-dimensional deformations that cannot be modeled in closed-form differential equations. With suitable contact-rich demonstrations and feedback, neural policies can learn compliance-aware responses that reduce force against unexpected resistance.

The fourth pillar is empirical scaling with data and compute, the bitter lesson that The Four Bedrock Laws sets against the physical wall of onboard energy, heat, and memory. Training on diverse robotic demonstrations (figure 5) and multimodal data can improve performance on tested tasks (Zhao et al. 2023; Chi et al. 2024).

Each of these gains holds only where training and test data cover the operating conditions, and none of them bounds what the model proposes outside those conditions. That limit is why a learned output reaches the actuators only as a candidate that the permission path may refuse.

Dual-arm robotic teleoperation workstation with leader and follower arms, overhead cameras, and grippers mounted on an aluminum frame.
Figure 5: Bimanual teleoperation workcell: Dual-arm teleoperation setup featuring two six-degree-of-freedom follower arms, two leader arms for kinesthetic teleoperation, multi-view cameras (overhead and wrist-mounted), and parallel grippers mounted to an extruded aluminum frame. The physical AI policy ingests synchronized multi-camera streams to predict continuous end-effector trajectories, and the permission path must refuse any proposal that would cause self-collision or motor over-torquing during closed-loop execution. (Credit: Stanford University & Google DeepMind, CC-BY / Apache 2.0).
Taxonomic diagram categorizing embodied foundation models into inputs, multimodal neural backbones, and navigation or manipulation action spaces.
Figure 6: Embodied model taxonomy: Embodied foundation models grouped by perceptual frontend (visual tokens, point clouds, language prompts), neural backbone (transformers, diffusion heads, world models), and target action space. The taxonomy contrasts Vision-Language-Navigation policies that decode topological waypoints and semantic frontiers with Vision-Language-Action policies that emit joint and Cartesian trajectory chunks for articulated manipulators. (Adapted from Ma et al. (Ma et al. 2024)).
Ma, Yingdong, Yifan Song, Ziyuan Cao, Yu Zhang, et al. 2024. “A Survey on Vision-Language-Action Models for Embodied AI.” arXiv Preprint arXiv:2405.14093.
Comprehensive architectural schematic of a Vision-Language-Action foundation model showing spatial patch slicing, 2D positional embeddings, multimodal self-attention weights in an illustrative image grid, discrete versus continuous action decoders, and an independent proposal boundary.
Figure 7: Embodied Vision-Language-Action (VLA) architecture: A representative dataflow. (1) Camera patches receive 2D image-position and camera identifiers alongside language and robot-state tokens; metric depth still requires calibrated evidence. (2) Self-attention mixes the combined token sequence; the colored grid illustrates internal attention weights, not a verified affordance or grasp pose. (3) Discrete autoregressive decoders emit tokens sequentially, whereas continuous decoders generate candidate multi-step chunks (\(\mathbf{A}_{t:t+H-1}\)) through one pass or iterative denoising. (4) An independent permission path checks candidate motion against sensed state and modeled limits before actuation. (Synthesized from Brohan et al. (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023; Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023), Kim et al. (Kim et al. 2024), and Octo Model Team (Octo Model Team et al. 2024)).

Rather than an opaque monolith, an embodied foundation model is a three-stage neural pipeline (figure 6, figure 7) that transforms sensory inputs into candidate motion proposals, and it ends at a boundary where those proposals still need kinematic checks.

  1. Sensory Ingestion and Spatial Tokenization (The Inputs): A vision transformer frontend cuts each camera frame into a grid of fixed-size pixel patches and projects each patch into one token, tagged with its image position and, when several cameras are mounted, a learned camera identifier. Language directives and the robot’s joint state are projected into tokens of the same width, and the three streams concatenate into one input sequence. Every patch becomes a token that the backbone must attend over and the key-value (KV) cache must hold, so a robot with three cameras (left wrist, right wrist, overhead) contributes hundreds of visual tokens per frame before any text or proprioception arrives (1.1). Image position is not metric depth; placing the gripper in 3D still requires depth evidence and calibrated geometry.

  2. Multimodal Alignment and Attention (The Backbone): The combined sequence passes through tens of transformer layers whose self-attention weighs every token against every other, so attention work grows with the square of sequence length while the KV cache that stores prior token states grows linearly with it. Every added camera or retained frame is therefore paid for in latency and in memory traffic. The same pairwise weighting does not establish geometry, so a text token such as “handle” can attend strongly to a visual patch without that weight being a verified grasp location or metric pose. Nor does the KV cache track object permanence; an explicit state estimator must associate observations across frames and grow its uncertainty while the mug is occluded. The patch-embedding, position-encoding, and attention equations behind both stages are in Visual patch tokenization, attention, and 3D metric lifting.

  3. Specialized Action Decoders (The Outputs): Instead of generating conversational text, the final hidden layer \(\mathbf{h}_{\text{context}}\) feeds task-specific action decoders. While vision-language navigation (VLN) models emit topological waypoints or heading rates \((v_t, \omega_t)\) (figure 8), manipulation (VLA) models diverge sharply in how they bridge semantic representations to physical motion:

    • Discrete Categorical Tokenization (RT-1, RT-2, OpenVLA): Early VLAs adapt autoregressive language modeling directly to robotics by discretizing continuous joint or Cartesian offsets into \(B = 256\) uniform bins per dimension (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023; Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023; Kim et al. 2024): \[\text{bin}(a_k) = \left\lfloor \frac{a_k - a_{\min}}{a_{\max} - a_{\min}} \times 255 \right\rfloor \in [0, 255]\] Each bin is assigned a discrete token (<act_000> to <act_255>) in the language vocabulary. While this natively leverages web-scale vision-language pretraining, it introduces a systems bottleneck: a decoder that emits one scalar token at a time needs seven serial token predictions for a 7-DoF action. Its latency depends on caching, hardware, and whether each prediction revisits the full backbone; seven serial steps can miss a control deadline. Scalar binning has a resolution of roughly \(\Delta a = (a_{\max} - a_{\min}) / 256\); whether that produces contact chatter depends on smoothing, the action range, and the downstream controller.
    • Continuous Generative Action Chunking (Octo, \(\pi_0\), ACT, Diffusion Policy): To avoid paying a full weight sweep for every scalar action, modern generalist policies couple the transformer backbone with continuous generative decoders, such as diffusion policy heads (Chi et al. 2024; Octo Model Team et al. 2024) or flow-matching action experts (Black et al. 2024). Rather than predicting a single scalar per forward pass, the model predicts an entire continuous trajectory chunk: \[\mathbf{A}_{t:t+H-1} = [\mathbf{a}_t, \mathbf{a}_{t+1}, \dots, \mathbf{a}_{t+H-1}] \in \mathbb{R}^{H \times d_a} \quad (H = 16\text{--}64\text{ steps})\] A backbone may evaluate one new camera observation, then a smaller decoder may perform several denoising iterations to propose a whole chunk of waypoints. That chunk is one candidate horizon, not a sequence of fresh observations. The next backbone evaluation must arrive before the accepted execution window expires; downstream interpolation and feedback run at their own rates.
  4. The Proposal Boundary (Isolation from Actuators): Whether the model is a VLN mobility planner or a VLA manipulation policy, its output is a candidate proposal (\(p_t\)) on the unprivileged side of the proposal boundary (\(\ref{dfn-boundary-proposal-boundary}\)). Under the machine model of The Machine in Five Levels, the foundation model runs in user space on the application processor and has no path to the pulse-width modulation (PWM) registers or to the current in the motor coils.

    The design rule for this deliberative boundary is that learned models operating at the Brain’s inference cadences (\(1\text{--}50\text{ Hz}\)) propose kinematic setpoints (\(\mathbf{q}, \dot{\mathbf{q}}, \ddot{\mathbf{q}}\) or Cartesian poses), never unbuffered, open-loop motor torques (\(\boldsymbol{\tau}\)) directly to inverter switches. Motor torque is coupled to the plant through the manipulator equation of the classical-dynamics note in section 1.2, plus a contact term \(\boldsymbol{\tau}_{\text{contact}}\). Because inertia \(\mathbf{M}(\mathbf{q})\) and Coriolis forces \(\mathbf{C}(\mathbf{q}, \dot{\mathbf{q}})\) vary non-linearly with instantaneous joint configuration and velocity, commanding open-loop torques, without high-frequency feedback, from a neural model that updates at tens of hertz at most can destabilize the limb when model error, friction, or payload variation is large enough. Kinematic setpoints, by contrast, specify an intended spatial path and leave high-gain closed-loop torque synthesis to the Body’s drives, whose field-oriented current loops run at \(10\text{--}25\text{ kHz}\), once the permission path has admitted each setpoint. (Where specialized high-rate reinforcement learning policies operating at \(100\text{--}500\text{ Hz}\) output torques directly, they still require an independently validated high-rate permission path, as Invariant Checking argues; passivity checks for such torque streams are derived in Passivity and Energy-Bounded Interaction.)

Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees Nair, et al. 2024. “Octo: An Open-Source Generalist Robot Policy.” arXiv Preprint arXiv:2404.08839.

↳ Downstream: Zero-copy proposal passing between the Brain and the Nervous System relies on the lock-free buffers analyzed in Multi-Rate Cadences.

Proposals reach the permission path through the handoff of The Nervous System, and the permission path can refuse a hallucinated proposal when independent measurements, modeled limits, and a feasible fallback support that decision.2

Schematic of topological vision-language navigation showing panoramic observation tokens, topological graph matching, and decoded waypoint offsets and heading setpoints.
Figure 8: Topological vision-language navigation: Embodied navigation architecture executing natural-language spatial instructions. Panoramic visual observations are encoded into perceptual tokens, matched against prompt embeddings, and cross-referenced with a dynamically updated topological navigation graph. The action decoder prunes unreachable nodes and emits metric waypoint offsets \((\Delta x, \Delta y)\) and heading setpoints to guide the chassis along traversable semantic frontiers without requiring an exact geometric prior map. (Adapted from An et al. (An et al. 2023)).
An, Dong, Yuankai Qi, Yang Huang, Yan Wei, Liang Wang, Tieniu Tan, et al. 2023. “ETPNav: Evolving Topological Planner for Vision-Language Navigation in Continuous Environments.” arXiv Preprint arXiv:2304.03047.

Each neural stage of this pipeline is paid for in memory traffic. Every patch token must be attended over and held in the KV cache, and every evaluation of the backbone streams its weights from DRAM, so the capabilities this section described arrive at a rate the application processor sets rather than one the task chooses. On the mobile manipulator, that rate for its largest model falls well below the rate at which its arm consumes setpoints.

Supply, Freshness, and the Memory Wall

At the pick station the mobile manipulator’s arm consumes a new setpoint every few tens of milliseconds, while the mug it is reaching for keeps moving along the conveyor. Whatever model proposes those setpoints must settle two questions. The first is how often it can produce a proposal on the application processor the machine actually has; the second is how long each proposal stays true once the mug has moved. The answers explain why the machine carries two learned proposers rather than one.

After ingestion (the tokenization of section 1.3), four stages turn those tokens into proposals (figure 9), and Part III develops each one. Here they matter only for what each hands downstream and how quickly that goes stale. Perception grounds what it recognizes in metric coordinates relative to the base, so that the mug becomes a pose and a graspable handle rather than a label;3 its output ages from the moment of capture, and Sensor Perception turns that age into a contract. Memory carries those estimates through occlusion and must let each one expire when the evidence behind it does (Spatial Memory). Intent decomposes a task such as “take the ceramic mug off the conveyor and hand it to the coworker at the packing station” into subgoals, and because the mug moves while the intent model deliberates, it hands the planner a target that expires, an intent lease (The Intent Lease). Planning converts that target into candidate setpoints for the base and the arm, the proposals the permission path will judge (Trajectory Planning). These stages, and the ingestion before them, all run on the application processor and draw on the same memory bus, which is where the machine’s cadence is decided.

Horizontal flowchart of the five-stage cognitive pipeline: Ingestion, Perception, Memory, Intent, and Planning, culminating at the proposal boundary.
Figure 9: The physical AI cognitive flow: Information flows sequentially from physical sensors through Ingestion (tokenization), Perception (spatial geometry and affordances), Memory (persistent world model belief), Intent (deliberative goal leases), and Planning (multi-step action chunking). The flow ends at the proposal boundary, where the permission path evaluates each proposed action against real-time constraints and resources.

↳ Downstream: High-throughput neural inference accelerators are isolated from deterministic cores in Two Paths on One Die.

The memory wall

Consider what happens if a robot attempts to generate continuous motor actions one discrete step at a time through an autoregressive foundation model (such as a multi-billion-parameter VLA model). A model whose weights cannot remain on chip must stream them across the external DRAM bus once per evaluation, so the bus bandwidth \(B_{\text{mem}}\) bounds its evaluation rate at \(f_{\max} = B_{\text{mem}} / M_{\text{weights}}\) before tokenization, attention, or decoding adds any time. In single-step autoregressive decoding, each sequential token may require another such evaluation; the actual traffic depends on caching and the decoder architecture. For detailed memory wall derivations and roofline hardware models, see Systems and Hardware.

This memory wall4 places single-token streaming deep in the memory-bound region of the roofline (figure 10) (Williams et al. 2009). On the mobile manipulator’s application processor it caps the intent model’s fresh-evaluation rate, before any vision or decoding work, far below the rate of a joint controller that asks for a new command every millisecond (1.1). A model renewed that slowly can say where the arm should go, but it cannot close the loop on moving mass, so it proposes waypoints and a faster loop below it tracks them.

Williams, Samuel, Andrew Waterman, and David Patterson. 2009. “Roofline: An Insightful Visual Performance Model for Multicore Architectures.” Communications of the ACM 52 (4): 65–76.

This continuous data movement also carries an energy penalty. Fetching weights from off-chip DRAM consumes orders of magnitude more energy than the arithmetic computation itself. That energy drains battery reserves, and the heat it leaves in the chassis raises the local ambient of the joint actuators, shrinking the winding margins of Thermal Duty Cycles.

Illustrative roofline diagram plotting arithmetic intensity against attainable throughput, with example points for token decoding and action chunking.

Figure 10: Illustrative embodied roofline: Arithmetic intensity against attainable throughput for an example edge accelerator. Single-token weight streaming can be memory-bound; reuse across a chunk may increase work per byte. A separate real-time loop executes accepted waypoints between fresh model updates.

Action chunking

To supply dense waypoints despite slow full-model inference, modern physical AI policies employ action chunking (1.1) (Zhao et al. 2023; Chi et al. 2024). Each chunk then crosses the proposal boundary as one packet, the chunk payload of Multi-Rate Cadences.

A chunking policy (such as Action Chunking with Transformers [ACT], Diffusion Policy, or Flow Matching) generates its whole chunk of \(H\) setpoints in one backbone pass followed, for an iterative decoder, by an assumed \(K = 10\text{--}16\) refinement steps. Each refinement re-evaluates the smaller action head, so full latency is \(t_{\text{vision}}+t_{\text{backbone}}+K t_{\text{head}}+t_{\text{transfer}}\), not just one weight sweep. A new observation changes the plan only after another backbone and decoder pass; see Machine Learning Background for the generative mechanism.

Definition 1.1: Temporal action chunking

Temporal action chunking is the policy formulation that predicts an entire horizon of \(H\) future continuous actions \(\mathbf{A}_{t:t+H-1} = [\mathbf{a}_t, \dots, \mathbf{a}_{t+H-1}]\) in a single neural evaluation or diffusion rollout, amortizing parameter memory transfers across extended physical execution windows.

  1. Significance: On a bandwidth-limited device, repeating a full model evaluation for every action can be too slow for millisecond control. Temporal action chunking can reuse a backbone evaluation across multiple waypoints; it does not make fresh observations or policy decisions arrive at the waypoint rate. A downstream controller tracks the accepted trajectory between model updates.
  2. Distinction: Unlike single-step Markovian policies \(\mathbf{a}_t \sim \pi(\cdot \mid \mathbf{o}_t)\) that re-evaluate the full neural network at every control step, action chunking deliberately generates multi-step feedforward trajectories, utilizing receding-horizon execution and blending across overlapping chunks to reconcile slow neural deliberation with fast closed-loop control.
  3. Common pitfall: Executing action chunks open-loop without continuous temporal smoothing or runtime safety filtering. Without smooth blending across overlapping chunk boundaries, switching between independently generated horizons injects acceleration jumps and torque shocks that excite mechanical gearbox resonances.

For example, when a canonical bimanual manipulator threads a flexible USB cable (figure 5), it executes a 32-step action chunk, a longer chunk than the mobile manipulator uses. The policy emits coordinated 14-DoF joint setpoints that smoothly align both grippers, flex the cable, and insert the connector without pausing between individual control cycles.

Sequence strip illustrating overlapping action chunk horizons of 32 steps with 12-step replanning and temporal ensembling.

Overlapping action chunks blend predictions across receding temporal execution windows.

However, action chunking introduces the problem of chunk seam continuity. Because each action chunk is generated independently from rolling sensory observations, the tail of chunk \(k\) and the head of chunk \(k+1\) may exhibit slight positional and velocity discrepancies. If a controller concatenated these chunks directly, the resulting step-discontinuities in acceleration would inject large jerk (\(\dddot{\mathbf{q}}\)) into the physical linkages, inducing mechanical shock, acoustic vibration, cycloidal pin wear, and motor overcurrent faults.

To reduce discontinuities between overlapping chunks, systems can implement temporal ensembling, blending overlapping trajectory predictions via exponentially weighted averaging:

\[\mathbf{a}_t = \frac{\sum_{i=0}^{\min(t, H-1)} w_i \, \mathbf{a}_{t \mid t-i}}{\sum_{i=0}^{\min(t, H-1)} w_i}, \quad w_i = \exp(-m \cdot i)\]

where \(\mathbf{a}_{t \mid t-i}\) represents the action predicted for time step \(t\) by the action chunk initiated at time \(t-i\), and \(m > 0\) is a tunable temporal discount factor. The permission path interpolates the blended waypoints below the proposal boundary (Multi-Rate Cadences). Ensembling averages setpoints without enforcing derivative continuity, so a seam can still need the \(C^2\) bridge that Kinodynamic Feasibility builds.

This operational tension between neural deliberation capacity and physical control latency defines the central scaling frontier of modern robotic brains (figure 11). Between 2020 and 2023, pioneering vision-language-action (VLA) models such as RT-1 (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023), RT-2 (Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023), and OpenVLA (Kim et al. 2024) tokenized continuous actions into discrete bins, predicting one action step per autoregressive forward pass. However, because each forward pass requires streaming every model parameter from DRAM across the silicon memory bus, policy update rates collapsed as model capacity scaled (\(f_{\text{action}} \propto 1 / M_{\text{weights}}\))—trapping multi-billion-parameter models behind an insurmountable “deliberation wall” (\(1\text{--}3\,\text{Hz}\)) that fell far below the \(50\,\text{Hz}\) closed-loop threshold required for dynamic contact and reactive manipulation.

Brohan, Anthony, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, et al. 2023. “RT-1: Robotics Transformer for Real-World Control at Scale.” Robotics: Science and Systems (RSS). https://doi.org/10.15607/RSS.2023.XIX.025.
Brohan, Anthony, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, et al. 2023. “Rt-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.” arXiv Preprint arXiv:2307.15818.
Kim, Moo Jin, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, et al. 2024. “OpenVLA: An Open-Source Vision-Language-Action Model.” arXiv Preprint arXiv:2406.09246.
Zhao, Tony Z., Vikash Kumar, Sergey Levine, and Chelsea Finn. 2023. “Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.” Robotics: Science and Systems (RSS). https://doi.org/10.15607/RSS.2023.XIX.016.
Chi, Cheng, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. 2024. “Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.” The International Journal of Robotics Research 44 (10-11): 1684–704. https://doi.org/10.1177/02783649241273668.
Black, Kevin, Homer Walke, Kevin Bishop, Lee Zhang, Albert Fang, Brent Chen, Steven Morad, et al. 2024. “\(\pi_0\): A Flow-Matching Foundation Model for Generalist Robots.” arXiv Preprint arXiv:2410.24164.

Action chunking fundamentally decoupled deliberation from actuation frequency. Instead of streaming billions of parameters for every low-level setpoint, architectures such as ACT (Zhao et al. 2023), Diffusion Policy (Chi et al. 2024), and modern flow-matching foundation models (e.g., Physical Intelligence \(\pi_0\) (Black et al. 2024) and NVIDIA Project GR00T) evaluate their large perceptual backbones periodically to generate entire multi-step action trajectories (\(H = 16\text{--}64\)). Compact, highly optimized generative action heads then denoise or integrate these trajectories at \(50\text{--}100\,\text{Hz}\), shattering the autoregressive memory wall and reconciling multi-billion parameter semantic reasoning with millisecond closed-loop physical control.

Figure 11: The embodied brain frontier: model parameter scaling versus physical action frequency across robotic policies (2020–2026): Trade-off between foundation policy parameter capacity (\(M_{\text{weights}}\), log scale, \(20\,\text{M}\) to \(100\,\text{B}\) parameters) and onboard closed-loop action frequency (\(f_{\text{action}}\), log scale, \(0.5\,\text{Hz}\) to \(200\,\text{Hz}\)). Single-token discrete autoregressive policies (RT-1, RT-2-PaLM-E, RT-2-PaLI-X, OpenVLA Base) hit an insurmountable memory wall (\(f \propto 1/M_{\text{weights}}\)) bounded by DRAM streaming bandwidth (\(B_{\text{mem}} / M_{\text{weights}}\)), stranding large models at \(1\text{--}3\,\text{Hz}\). Continuous action chunking (ACT, Diffusion Policy) and flow-matching dual-system architectures (Physical Intelligence \(\pi_0\), NVIDIA Project GR00T, 2026 dual-brain SoCs) shatter this bottleneck: decoupling multi-step chunk generation (\(H = 16\text{--}64\)) from low-latency generative action heads enables multi-billion parameter foundation models to achieve \(50\text{--}100\,\text{Hz}\) real-time reactive control.

The machine’s two proposers

On the mobile manipulator’s own application processor, the memory wall and the amortization that chunking offers decide how the machine divides proposal work between its models.

Napkin Math 1.1: Memory walls, supply, and freshness
Can action chunking let a multi-billion-parameter model on an embedded compute module renew proposals as fast as the arm consumes them, or does the machine need a second, smaller proposer?

1. The single-token memory bottleneck. Consider the mobile manipulator’s intent model, a 7B VLA model stored in 16-bit precision (FP16) in the LPDDR5 memory of its application processor, an embedded edge System-on-Chip:

  • Weight footprint: \(M_{\text{weights}} =\) 7B params \(\times\) 2 bytes \(\approx\) 14 GB.
  • Theoretical memory bandwidth: \(B_{\text{peak}} =\) 204 GB/s.
  • Assumed sustained DRAM bandwidth (70 percent of peak to model bus scheduling and memory overhead): \(B_{\text{sustained}} \approx\) 142.8 GB/s.

Minimum streaming latency to fetch the weights from DRAM across the silicon bus once: \[t_{\text{stream}} = \frac{14 GB}{142.8 GB/s} \approx 98.0 ms\]

Bandwidth-only upper bound on fresh evaluations when each evaluation streams the full weights: \[f_{\text{single}} = \frac{1}{t_{\text{stream}}} \approx \frac{1}{0.0980 s} \approx 10.2 Hz\]

Beyond model weights, autoregressive foundation models also grow their KV cache. When multi-camera sensory feeds (e.g., 3 \(224 \times 224\) RGB streams) are ingested, each video frame contributes hundreds of visual patch tokens (3 \(\times\) 256 \(=\) 768 tokens). In the intent model’s 32-layer transformer with hidden dimension \(d =\) 4096, storing the FP16 Key and Value states consumes 0.52 MB per token: \[M_{\text{KV}} = 2 \times L \times d \times S \times \text{bytes} = 2 \times 32 \times 4096 \times S \times 2 \approx 0.52 MB \cdot S\] Across a 768-token visual prompt, the KV cache expands by 402.7 MB per frame.

2. The action chunking amortization. Suppose the intent model itself emitted an action chunk of \(H =\) 16 future setpoints (\(\Delta t =\) 20 ms, spanning 320 ms) from each forward pass, instead of one action step per DRAM weight sweep.

  • Number of weight memory sweeps: 1 sweep (98.0 ms).
  • A chunk can reuse one backbone evaluation across \(H\) future setpoints.
  • Amortized weight-sweep time per setpoint (excluding all other inference work): \[\tau_{\text{step}} = \frac{t_{\text{stream}}}{H} = \frac{98.0 ms}{16\text{ setpoints}} \approx 6.1 ms\text{ per setpoint}\]

3. The machine’s split. Amortization spreads one sweep across more setpoints, but it does not make a fresh proposal arrive sooner. Each new chunk from the intent model would still wait for at least one 98.0 ms sweep, plus tokenization, attention, and decoding. The mobile manipulator therefore divides the work between two models. Its intent model revises the goal at 5 Hz with 160 ms of inference. A separate 50M-parameter chunk policy, whose 0.10 GB of weights sweep in 0.70 ms, proposes base velocity and arm joint targets together at 20 Hz, within a 40 ms P99 inference budget that includes vision. The intent model’s sweep is too slow to renew a proposal every few tens of milliseconds; the chunk policy’s is not. The Nervous System turns that difference into a lease.

The 98.0 ms sweep in 1.1 is a floor reached before vision, attention, or decoder work begins. Two alternatives to the machine’s split, lower weight precision and a single model with a fast head, would attack it, and neither removes it cheaply. INT4 storage would reduce the nominal weight footprint to 3.5 GB but needs accuracy validation for contact tasks. A single-model design could instead trade cadence for density without changing weight precision. It would evaluate the 7B backbone periodically while a lightweight diffusion head denoises candidate trajectories on-chip, supplying dense candidate waypoints from each slower full-model evaluation. Its new-observation rate would still be limited by full inference latency, including vision and decoder work, which is why the mobile manipulator gives renewal to a separate chunk policy instead.

Retained context adds to the pressure the notebook began to measure. Counting language and proprioception tokens along with the visual patches raises each step to \(S \approx\) 800 tokens, so at 0.52 MB of FP16 key-value state per token each observation step adds 419.4 MB of KV cache. Retaining an unpruned 5 s historical context window at 10 Hz yields 21.0 GB of KV cache. Combined with the intent model’s 14 GB of weights, this example claims 35.0 GB, about half of the application processor’s 68.7 GB before runtime overhead, memory that perception, mapping, and the chunk policy also need. Even where memory capacity permits, reading a large cache is further traffic on that same bus (Contention for Shared Resources). Embodied architectures must therefore avoid unpruned autoregressive visual histories, relying instead on spatial pooling, cross-attention feature projections, or fixed-horizon action chunking.

At the pick station, the mobile manipulator’s arm takes the mug from a conveyor moving at 0.20 m/s, and each chunk must be executed without letting the mug drift out of reach. Untracked, the belt’s steady motion would use up the margin \(r_{\text{tol}} - e_0\) between the grasp tolerance and the handle estimate’s initial error \(e_0\) within the freshness deadline \(e_{\max}/v_{\max}\) of Measurement Freshness, with \(e_{\max} = r_{\text{tol}} - e_0\). A tracker that follows the mug removes that steady motion, but residual slip can still accelerate the target at 0.50 m/s², so a segment executed open-loop for a time \(T\) drifts by \(\frac{1}{2} a_{\text{slip}} T^2\). Added to the 3 mm initial error \(e_0\), that drift reaches the 15 mm grasp tolerance after the tracked evidence horizon \(\tau_{\text{ev}} = \sqrt{2 (r_{\text{tol}} - e_0) / a_{\text{slip}}} \approx\) 219 ms. The chunk policy replans every 50 ms and needs up to 40 ms to do so, so while renewals arrive on time the arm executes no more than 90 ms of motion from one observation, which drifts only 2.0 mm and sits well inside the tracked evidence horizon.

That open-loop span also sets the minimum supply, because a chunk shorter than 5 setpoints would leave the arm waiting for its successor. The machine’s chunk carries 16 setpoints at 20 ms spacing, 320 ms of motion, longer than the tracked evidence horizon; executed open-loop to its end, it would drift 25.6 mm, past the tolerance. The setpoints beyond the next replanning instant are therefore a supply buffer, consumed only when renewals stop, and the lease derived in Multi-Rate Cadences ends execution long before that buffer runs out. Overlapping chunks and temporal ensembling then blend the executed motion across each renewal.

↳ Downstream: Bounded action chunk proposals are filtered against forward-invariant barriers in Safe Sets as Conditional Permission.

These are the proposer limits Part I hands forward: one chunk covers 320 ms of motion, a new chunk takes up to 40 ms, and the evidence behind it stays within tolerance for about 219 ms. They bound when a proposal arrives and how long it remains true, and each enters the handoff record of The Nervous System as a limit record in the form of Measuring a Machine's Own Limits. None of them says whether the proposal was right when it was made.

Checkpoint 1.1: Edge SoC deliberation limits: Memory walls, KV cache, and dynamic drift

Before analyzing runtime failure modes, verify your understanding of edge deliberation limits across memory and time:

Cognitive Vulnerabilities

A chunk the mobile manipulator’s policy proposes for the mug can arrive inside its 40 ms budget, with every setpoint finite and within joint limits, and still close the gripper on the rim. Its timing is sound and its content is not, and the failure modes of learned models live in that gap. Learned foundation models extend task coverage but make a general proof of integrated closed-loop behavior difficult; specified components and bounded operating regions can still be analyzed.5

Without a verified input domain and sensitivity bound, small sensing changes to a learned proposer can yield unexpectedly large changes in proposed motion. Benchmark success does not establish closed-loop stability, timing bounds, or safety for unseen conditions; those properties require separate evidence for the model, runtime, and physical plant.6

A learned proposer’s failure modes also differ from traditional software defects. A classical program fails by throwing an exception, hanging in an infinite loop, or crashing. A neural network fails silently instead, as the mug chunk did, with output that passes every format check. An independent permission path must reject a detected violation before those values can become actuator commands. Systems architects must design around four primary cognitive vulnerabilities.

Semantic hallucinations

A shipping carton behind the conveyor carries a printed photograph of a mug, and the mobile manipulator’s policy proposes a grasp on the printed handle. The trajectory tensor is finite and within joint limits, yet the object it targets is not there. Its plausibility cannot substitute for an independent geometric check, which would find a flat surface where the policy saw a handle.

The Softmax illusion and out-of-distribution overconfidence

Deployment interfaces often confuse mathematical normalization with physical probability. In deep neural networks, output logits are passed through a softmax function,7 whose outputs sum to one whatever the input, so embodied machines encountering out-of-distribution (OOD) observations (novel optical textures, deep shadows, or lens water droplets) can assign confidence scores exceeding \(0.95\) to erroneous classifications.

Evaluating the true relationship between reported model confidence and empirical physical success requires measuring Expected Calibration Error (ECE), the gap between how sure the model thinks it is and how frequently the physical task actually succeeds.8

The regimes in table 1 show how high reported confidence can coexist with low task success under shift. The single-row differences are confidence–success gaps, not ECE; computing ECE requires bins over a specified evaluation set. Raw scores cannot grant execution permission without validation and independent physical checks.

Table 1: Illustrative Confidence–Success gaps under distribution shift: Confidence and task-success rates for four operating regimes, with the supervisory response each gap calls for.
Distribution Shift Regime Physical & Environmental Trigger Conf. vs Success (\(\text{conf} \text{ vs } \text{success}\)) Confidence–Success Gap (\(\text{conf}-\text{acc}\)) Downstream Supervisory Action
In-Distribution (Nominal Baseline) Clean optics, structured laboratory lighting, nominal workpiece textures \(\text{conf}=0.91, \, \text{acc}=0.89\) \(+0.02\) Apply independent limits before nominal execution
Near-OOD Shift (Covariate Shift) Specular reflections, lens water droplets, dynamic shadow transitions \(\text{conf}=0.89, \, \text{acc}=0.52\) \(+0.37\) Selective abstention triggered: clamp velocity, initiate exploratory tactile probe
Far-OOD Shift (Semantic Shift) Uncataloged deformable obstacles, human limb intrusion into workspace \(\text{conf}=0.78, \, \text{acc}=0.08\) \(+0.70\) Reject proposal; use feasible controlled response
Sensor Degradation (Hardware Drift) Photodetector thermal noise (\(>75^\circ\text{C}\)), focal blur from vibration \(\text{conf}=0.84, \, \text{acc}=0.31\) \(+0.53\) Trigger operational downgrade; restrict kinematics to certified limits

Perceptual chattering and covariate shift

Ambiguous images can make successive object classifications change rapidly. If a motion planner treats each new label as a fresh physical state, it may discard useful motion history, as each reclassification did in the Tempe collision (When the Boundary Fails). In physical manipulation, perceptual chattering injects high-frequency direction reversals into trajectory planners, and the commanded reversals demand torque steps that the drivetrain must absorb, oscillations that damage gearboxes. Temporal smoothing can help motion continuity, but it cannot repair a false obstacle estimate; independent detection and a feasible response remain necessary.

Temporal expiration of intent

A proposal stays valid only as long as the evidence behind it, a span the conveyor-pick example computes as the evidence horizon \(\tau_{\text{ev}}\). If the mobile manipulator’s intent model needs 160 ms to deliberate, the people and carts in the aisle, the arm itself, and the mug on the conveyor have already moved by the time the plan is ready. Delayed downstream communication or execution leads to acting on an expired intent lease (The Intent Lease), commanding the physical body based on an obsolete mental representation. A proposer that stalls outright leaves the plant moving with no plan at all, so recovery must live in a local controller below it, a requirement the 2015 DARPA Robotics Challenge Finals imposed on legged entrants by degrading communications deliberately (Krotkov et al. 2017).

Krotkov, Eric, David Hackett, Lawrence Jackel, Michael Perschbacher, James Pippine, Craig Strauss, Gill Pratt, and Charles Orillat. 2017. “The DARPA Robotics Challenge Finals: Results and Perspectives.” The DARPA Robotics Challenge Finals: Humanoid Robots to the Rescue, Springer tracts in advanced robotics, vol. 121: 1–26.

The four vulnerabilities, with the memory wall of section 1.4, give three independent reasons against granting a learned proposer actuator authority. The memory wall means the most capable model cannot renew a proposal as fast as the machine consumes one. Hallucination, overconfidence, and chattering mean that a proposal can be wrong or unstable while its reported confidence stays high, so confidence cannot certify it. Expiry means that even a correct proposal stops being correct once the world has moved. Table 2 sets the learned proposer’s limits beside those of the classical methods the chapter began with and shows which side of the proposal boundary each capability occupies.

Table 2: Classical analytical methods vs. learned foundation models: Trade-offs across generalizability, verifiability, pacing, and fallback, with the synthesis that partitions proposal capability from execution authority.
Systems Dimension Classical Analytical Methods Learned Foundation Models Physical AI Architectural Synthesis
Foundational Principle Analytical differential equations, geometry, CAD models, Lyapunov stability High-capacity neural representations fitted to empirical multimodal data Hybrid Hierarchy: Learned intent proposals gated by the permission path
Generalizability & Perception Model-dependent: Explicit geometry and features may lose coverage under shift Data-dependent: Open-vocabulary grounding can improve coverage but may fail under shift Unprivileged Deliberation: Foundation models synthesize continuous semantic goals
Mathematical Verifiability Conditional: Selected properties can be proved within model assumptions Limited: Empirical behavior needs separate bounds and runtime evidence Proposal Boundary: Gate checks modeled limits using independent evidence
Temporal Pacing Hard Real-Time: Deterministic millisecond execution loops (\(1\text{ kHz}\)) Multi-Rate: Full inference, chunk refresh, and waypoint execution have distinct rates Split Proposers: A small chunk policy renews while the intent model revises goals
Failure Modes & Fallback Boolean branch crashes; failure on uncataloged objects; rule exhaustion Semantic hallucinations; out-of-distribution overconfidence; expired intent Conditional Fallback: Reject detected violations and execute a feasible response

Fallacies and Pitfalls

Learned capability is useful only when the machine can interpret and constrain its proposals. The following mistakes confuse a model’s success, confidence, or output format with permission to act.

Fallacy: An engineering team must choose entirely between formal classical safety proofs and learned empirical capability.

An assembly plant compares a classical wiring-harness controller with a learned diffusion policy. The classical design exposes explicit motion constraints but struggles with cable deformation; the learned policy completes 96 percent of trial insertions without establishing worst-case force or velocity bounds. These strengths address different engineering obligations. The learned policy can propose adaptive motions while a deterministic supervisor checks them against the physical limits, using a control barrier function where its assumptions apply (Ames et al. 2019) (detailed in Safety Enforcement). Partitioning capability from enforcement lets the team evaluate task performance and the protection mechanism separately. High empirical success does not establish the supervisor’s adequacy or the safety of the integrated machine.

Ames, Aaron D, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada. 2019. “Control Barrier Functions: Theory and Applications.” European Control Conference (ECC), 3420–31.

Pitfall: Allowing a neural model’s high confidence to bypass deterministic geometric boundaries.

A cleanroom manipulator mistakes glare on a transparent barrier for open space and proposes reaching through it while carrying reagent vials. The proposal contains finite, in-range coordinates and arrives within its deadline. Those checks establish that the packet is usable, not that its interpretation of the scene is correct. High confidence cannot resolve the same perceptual error that produced the trajectory. Before granting permission, the permission path must compare the proposal with independently maintained geometric boundaries (such as hard-coded stay-out zones or real-time depth maps).

Fallacy: Raw softmax probabilities or predicted confidence scores reflect true physical success rates.

A bin-picking policy reports \(\hat{p} \approx 0.90\) for a group of grasps, but only 64 of 100 physical attempts succeed. An abstention threshold based on the reported score would therefore accept more failures than its numerical value suggests. The discrepancy concerns calibration: confidence does not yet match observed success in the operating conditions (as illustrated in table 1). It does not, by itself, establish whether the policy ranks candidates correctly. Engineers must validate the relationship between scores and outcomes across the intended domain before using confidence to choose between execution, another observation, and operator assistance. Independent physical limits still apply to accepted grasps.

Pitfall: Coupling high-frequency semantic classifications directly to motor setpoints without continuous filtering.

A delivery robot observes reflective tarps over traffic cones. At every perception update, its pipeline changes the classification among obstacle, movable drape, and free path. If each change immediately resets motor setpoints, uncertainty in the scene becomes repeated requests for direction reversal. The drive motors cannot follow those requests beyond their torque slew limits, and the chassis chatters and shudders. The architecture needs continuity between changing interpretations and physical commands. Filtering or action chunking can support that continuity, but the resulting trajectory must still respect the body’s motion limits when the semantic estimate changes.

Fallacy: A numerically valid and smooth action chunk can safely be executed whenever it arrives.

The mobile manipulator’s chunk policy proposes an arm trajectory toward the mug on the conveyor. Host scheduling delays delivery, while residual belt slip has carried the mug toward the edge of the grasp tolerance. The trajectory remains smooth and numerically valid, but the observation that justified it no longer describes where the mug is. Executing it could close the gripper on the rim or knock the mug from the belt. Once the delay outlasts the chunk lease of Multi-Rate Cadences, the lease ends the stalled policy’s authority and the arm enters its validated fallback stop instead of waiting, and when the delayed trajectory finally arrives, the observation it was computed from is too old for the permission path to admit it. Smoothness describes the proposed command; temporal validity determines whether the machine may still act on it.

Summary

Learned representations give the mobile manipulator what hand-authored geometry could not, the ability to find the handle of a mug it has never seen. The price is paid in memory traffic and in time. The largest model evaluates least often, and every setpoint either proposer supplies ages with the observation that produced it.

Key Takeaways: Learned proposers and their deliberation limits
  • Weight traffic bounds how often a model sees: One sweep of the intent model’s 14 GB of weights takes 98.0 ms on the application processor before vision or decoding begins, so its fresh-evaluation rate cannot exceed about 10.2 Hz whatever its compute peak.
  • Chunking multiplies supply, not freshness: A 16-setpoint chunk spreads one evaluation across 320 ms of motion, but every setpoint is as old as the observation behind it. The machine therefore gives renewal to a small chunk policy at 20 Hz within 40 ms.
  • Size a chunk between starvation and drift: The replanning period plus inference latency (90 ms) sets the minimum supply, and target drift sets the evidence horizon (219 ms). Setpoints beyond the next renewal are a buffer that the lease ends, not a plan.
  • Retained context is paid for twice: Keeping 5 s of multi-camera history adds 21.0 GB of key-value cache, which with the weights claims about half the processor’s memory, and every read of it is traffic on the bus that real-time transfers share.
  • Confidence is not calibration: Softmax scores sum to one by construction, so reported confidence can stay high while task success collapses under distribution shift. Only calibration measured over the operating domain can justify an abstention threshold.
  • A proposal is true only for a while: Intent and chunks expire with the evidence behind them, and a smooth, in-range, on-time proposal can still be stale or wrong.

Every limit this chapter derived is a limit of supply or of truth, never of authority. None of them decides whether a proposal may reach the motors; the permission path takes them as inputs when it makes that decision.

What’s Next: From learned proposals to the permission path
How long may one proposal keep driving the machine if its successor never arrives? The mobile manipulator’s chunk policy renews every 50 ms and its intent model far less often, and either can stall, arrive late, or be wrong about the mug. The Nervous System sets how often the permission path checks each proposal and derives from the stopping budget of The Physical Body how long one may drive the actuators before the machine falls back, and it records both in the handoff record alongside the proposer limits established here.

Back to top

Footnotes

  1. Rigid-Body Dynamics Formulation: Classical control computes generalized torques via the standard manipulator equation \(\mathbf{M}(\mathbf{q})\ddot{\mathbf{q}} + \mathbf{C}(\mathbf{q}, \dot{\mathbf{q}})\dot{\mathbf{q}} + \mathbf{g}(\mathbf{q}) = \boldsymbol{\tau}\), where \(\mathbf{M}(\mathbf{q})\) represents the positive-definite inertia matrix, \(\mathbf{C}(\mathbf{q}, \dot{\mathbf{q}})\) captures Coriolis and centripetal terms, and \(\mathbf{g}(\mathbf{q})\) accounts for gravitational loading. This analytical formulation becomes computationally intractable and physically inaccurate when machines interact with unmodeled frictional contacts, fluid drag, or deformable objects. Derivations are provided in Control and Dynamics.↩︎

  2. Proposal Boundary: Treating the neural policy as an unprivileged proposal generator isolates stochastic inference from actuator registers. Candidate trajectory setpoints \(\mathbf{q}_{\text{ref}}(t)\) are checked on dedicated safety silicon, so proposals that violate modeled actuator limits or keep-out regions can be rejected before any reaches a drive (Actuation Authority).↩︎

  3. Metric Spatial Grounding: While disembodied AI operates over discrete semantic tokens, physical AI must ground those semantics in the Special Euclidean group \(SE(3)\), mapping entities to metric distances (meters) and spatial orientations (quaternions) relative to the robot’s base coordinate frame. Failing to maintain this metric grounding causes neural policies to emit kinematically impossible Cartesian paths that breach actuator workspace boundaries.↩︎

  4. Autoregressive DRAM Bandwidth Saturation: Sequential action tokens can trigger repeated full-model weight traffic when weights are nonresident and the decoder revisits the backbone; cache behavior and hardware determine the actual bandwidth ceiling. Action chunking is one way to reduce the number of full policy evaluations per unit of physical time.↩︎

  5. Connectionist and Symbolic Duality: Physical AI architectures reconcile statistical connectionism with deterministic symbolic computation: high-capacity neural networks synthesize open-world behavioral proposals, while formal symbolic logic and invariant monitors verify kinematic validity before setpoints are granted execution authority.↩︎

  6. Lipschitz Continuity in Neural Policies: Although neural networks with smooth activations are continuous on compact domains, a bound \(L = \sup_{\mathbf{x}_1 \ne \mathbf{x}_2} \frac{\lVert \pi(\mathbf{x}_1) - \pi(\mathbf{x}_2) \rVert}{\lVert \mathbf{x}_1 - \mathbf{x}_2 \rVert}\) must be established over a specified operating domain. If a valid bound exists, sensor error \(\lVert \delta \mathbf{x} \rVert \le \epsilon\) implies \(\lVert \delta \mathbf{a} \rVert \le L\epsilon\); an unverified \(L\) supplies no usable action-error guarantee. Mathematical formulations appear in Machine Learning Background.↩︎

  7. Softmax Normalization and Epistemic Uncertainty: The softmax operator \(\sigma(\mathbf{z})_i = \frac{e^{z_i}}{\sum_j e^{z_j}}\) enforces normalization across the categorical simplex \(\Delta^K\) by algebraic construction rather than Bayesian epistemic confidence. High output confidence reflects only relative logit separation, meaning models assign near-unity probabilities even when evaluating inputs completely outside their training distribution. Derivations appear in Machine Learning Background.↩︎

  8. Expected Calibration Error Formulation: Partition predicted confidence into bins. In each bin, compare observed success frequency with mean reported confidence; average the absolute gaps, weighted by the bin’s share of samples. Elevated ECE shows that reported confidence is a poor basis for granting physical permission. The equation and bin definitions appear in Softmax overconfidence and calibration error.↩︎