The Cognitive Brain
The Cognitive Brain
Purpose
Why can the most capable model on a robot revise its plan only a few times a second?
A robot operating among humans must interpret dynamic environments and propose actions far beyond the rigid scripts of classical automation. High-capacity learned models supply this semantic reasoning by grounding multimodal sensory streams in vast representations acquired during offline training. Yet this expressive power is paid for in high memory traffic and severe computational latency. Every forward evaluation of a multi-billion-parameter network streams weights across a memory interconnect of finite bandwidth. Consequently, the deliberative model that understands the open-world scene best is inevitably the component that revises its decisions least frequently.
To keep actuators continuously supplied between deliberative updates, the Brain must emit prospective action chunks—temporal sequences of future setpoints. However, these proposed setpoints age with the stale observation that generated them, and a proposal that arrives on time can still be physically disastrous. High prediction confidence does not entitle a statistical model to command the physical machine. The Brain generates candidate behaviors across the proposal boundary, while the deterministic Nervous System evaluates each chunk against current sensor telemetry and safety envelopes before allowing low-level commands to cross the causal boundary.
↰ Prerequisite: A learned proposal stays unprivileged until the permission path admits it, the third law of The Four Bedrock Laws.
Learning Objectives
- Compare classical analytical pipelines and learned representations on coverage, robustness, and verifiability
- Calculate the bandwidth bound on a model’s fresh-evaluation rate from its weight footprint and sustained memory bandwidth
- Estimate key-value cache growth from camera count, tokens per step, and retained history
- Explain why action chunking multiplies setpoint supply without making proposals any fresher
- Apply replanning period, inference latency, and target drift to size an action chunk for a moving target
- Diagnose proposer failures caused by hallucination, miscalibrated confidence, perceptual chattering, and expired intent
- Evaluate why these proposer limits require a permission path that can refuse any proposal
The Embodied Brain
The warehouse mobile manipulator can meet all electrical and thermal budgets yet still fail at a task as ordinary as taking an unfamiliar ceramic mug from the moving conveyor at the pick station and handing it to the coworker at the packing station. The motor drives remain well within their continuous torque ratings, the control rail holds its voltage at full charge, and the joint velocities remain below physical limits. Yet the machine drops the object or crushes the ceramic rim because low-level motor physics cannot determine what the object is, where its structural affordances lie, or how to reach around clutter. Unlike an advisory display, whose output can be reviewed before it affects the world, a cognitive proposal here can reach moving mass with no human in between. A mistaken proposal can lead to unsafe kinetic energy if the execution path accepts it, so permission must be checked independently. The central dilemma of embodied cognition is that high-capacity neural models are inherently stochastic, computationally expensive, and prone to inventing objects that are not there, yet they must orchestrate physical bodies whose collisions cannot be recalled (the first law, The Four Bedrock Laws).
The Brain is therefore an unprivileged deliberation engine that may only propose (\(\ref{dfn-boundary-proposal-boundary}\)). This architectural separation organizes the machine into a two-speed brain architecture (figure 1). At the deliberative tier, high-capacity foundation models perform slow cognition at \(1\text{--}50\text{ Hz}\), parsing open-world multimodal streams to emit structured proposals. At the real-time tier, deterministic safety monitors execute at \(1000\text{ Hz}\) on isolated safety silicon, evaluating candidate proposals against physical limits before granting execution permission (\(u^*\)), or asserting an asynchronous refusal that triggers deliberative state resynchronization. On the mobile manipulator it holds two learned proposers, an intent model that revises goals at \(1\text{--}5\text{ Hz}\) and a chunk policy that emits short sequences of future setpoints at \(10\text{--}50\text{ Hz}\), and the permission path may refuse anything either one proposes. That protection holds only while the limits, the state estimate, and the fallback it relies on remain valid.
A learned proposer earns its place by doing what hand-authored geometry cannot. It recognizes an object no one modeled from an open-vocabulary instruction, tolerates glare and clutter that break hand-written filters, and anchors what it recognizes to metric coordinates the arm can reach. Each of these capabilities comes from a model whose size also sets how often it can run. Figure 2 shows the tension across workload classes, with the models that carry the most parameters updating least often, far below the rates at which the Body’s loops close. The chapter’s question is therefore what a learned proposer can deliver on time, and at what cost in memory, latency, and freshness.
Section 1.2 shows why pipelines built on hand-authored geometry and state machines lose coverage outside the factory cell, and section 1.3 shows what learned representations recover and what they cost in tokens and memory traffic. Section 1.4 turns that cost into the machine’s numbers: the memory wall that bounds how often its intent model can run, the action chunking that keeps the arm supplied between evaluations, and the drift that bounds how long each chunk stays true. Section 1.5 then takes up the ways a proposal can arrive on time and still be wrong, each of which leaves the permission path with a proposal it must be able to refuse.
The Classical Wall
For nearly a century, automated control and robotics approached cognition as an explicit geometric and analytical physics problem (Wiener 1948; Kalman 1960; Craig 2005). Under this classical paradigm, engineers constructed mathematical representations from first principles.
Engineers used analytical kinematics and differential dynamics to map desired end-effector motions into actuator torques using closed-form equations.1 Perception relied on matching hand-crafted 3D Computer-Aided Design (CAD) models against sensor point clouds using geometric alignment algorithms, such as iterative closest point (ICP) or random sample consensus (RANSAC). High-level task logic was authored as explicit Boolean state machines that dictated deterministic transitions between operational phases, such as approaching, aligning, grasping, and retracting.
Where the operational environment is strictly structured and bounded (figure 3), classical methods can provide dependable behavior within their modeled conditions. On an automated automotive assembly line, an industrial robot welds stamped steel panels where every workpiece arrives at an exact coordinate within sub-millimeter tolerances. Ambient lighting is tightly regulated by industrial fixtures, and human workers are excluded by physical interlock safety fences. Within the causal boundary defined in The Causal Boundary, this controlled regime permits explicit models, bounded servo timing, and stability analysis under stated assumptions. Engineers can verify selected trajectory and torque constraints and assess controller software coverage against applicable standards without collecting task demonstrations.
Yet when an embodied machine leaves the structured factory cage to operate in unstructured environments (figure 4), this paradigm hits the classical wall (Brooks 1991; Levine et al. 2018). The physical world is not a sterile CAD drawing; it is continuous, deformable, and endlessly variable.
The first breakdown is the combinatorial explosion of geometry. In a commercial logistics warehouse handling hundreds of thousands of retail goods, authoring CAD models, mesh templates, and grasp points for every bottle, wrapped package, and flexible garment is impossible. Discretizing orientations, frictional properties, and contact permutations across overlapping, deformable items creates a combinatorial space that hand-crafted templates cannot cover.
The second breakdown arises from sensory brittleness. A hand-crafted edge-and-segmentation pipeline may rely on filters, Harris corners, or surface-normal thresholds. Under changed sunlight, reflections, or dust, a required feature can disappear. In that pipeline, a missing edge can make planar segmentation bridge two objects and propose a grasp in empty space.
The third breakdown is the engineering bottleneck of hand-authored logic. The machine’s operational capability is bounded by the engineer’s ability to anticipate every physical edge case. When an unexpected variation occurs, such as a bent bracket, a scuffed label, or an unfamiliar obstacle, the hand-authored rulebook contains no matching branch, causing the machine to halt, drop its payload, or collide.
Learned representations answer the first two breakdowns directly and ease the third, because they draw coverage from data rather than from hand-authored models.
Learned Representations
Asked to take the unfamiliar ceramic mug from the conveyor, a classical pipeline needs a CAD model, a feature threshold, and a state-machine branch that nobody wrote for this object. A learned model trained on many graspable objects can propose the handle as a grasp region from pixels and a sentence of instruction. Physical AI extends classical pipelines with such neural representations wherever hand-authored models lose coverage, and four pillars support the shift.
The first pillar is open-world semantic generalizability. Classical controllers are geometrically precise but semantically blind: an inverse kinematics solver can expertly drive an end-effector to coordinate \((x, y, z)\), yet it possesses zero understanding of whether that location contains a sturdy mug or a fragile glass, or how surrounding objects relate to human tasks. Modern vision-language-action (VLA) foundation models project open-vocabulary natural language instructions and high-resolution camera streams into shared semantic embeddings (Driess et al. 2023; Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023; Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023; Kim et al. 2024). When commanded to “wipe the coffee spill from the counter and stow the mug in the drying rack,” a trained model may propose the mug handle and spill boundary as task-relevant regions without a new CAD model.
The second pillar is empirical sensory robustness. Hand-authored visual thresholds can flip under changing illumination or texture. Learned representations can tolerate some of these changes when training and testing cover them, but their responses are neither necessarily smooth nor bounded under shift.
The third pillar is the ability to handle complex multi-contact and deformable dynamics. Many physical manipulation tasks (folding textiles, routing flexible wire harnesses through automotive door assemblies, or scooping granular media) involve non-smooth contact mechanics, variable friction, and infinite-dimensional deformations that cannot be modeled in closed-form differential equations. With suitable contact-rich demonstrations and feedback, neural policies can learn compliance-aware responses that reduce force against unexpected resistance.
The fourth pillar is empirical scaling with data and compute, the bitter lesson that The Four Bedrock Laws sets against the physical wall of onboard energy, heat, and memory. Training on diverse robotic demonstrations (figure 5) and multimodal data can improve performance on tested tasks (Zhao et al. 2023; Chi et al. 2024).
Each of these gains holds only where training and test data cover the operating conditions, and none of them bounds what the model proposes outside those conditions. That limit is why a learned output reaches the actuators only as a candidate that the permission path may refuse.
Rather than an opaque monolith, an embodied foundation model is a three-stage neural pipeline (figure 6, figure 7) that transforms sensory inputs into candidate motion proposals, and it ends at a boundary where those proposals still need kinematic checks.
Sensory Ingestion and Spatial Tokenization (The Inputs): A vision transformer frontend cuts each camera frame into a grid of fixed-size pixel patches and projects each patch into one token, tagged with its image position and, when several cameras are mounted, a learned camera identifier. Language directives and the robot’s joint state are projected into tokens of the same width, and the three streams concatenate into one input sequence. Every patch becomes a token that the backbone must attend over and the key-value (KV) cache must hold, so a robot with three cameras (left wrist, right wrist, overhead) contributes hundreds of visual tokens per frame before any text or proprioception arrives (1.1). Image position is not metric depth; placing the gripper in 3D still requires depth evidence and calibrated geometry.
Multimodal Alignment and Attention (The Backbone): The combined sequence passes through tens of transformer layers whose self-attention weighs every token against every other, so attention work grows with the square of sequence length while the KV cache that stores prior token states grows linearly with it. Every added camera or retained frame is therefore paid for in latency and in memory traffic. The same pairwise weighting does not establish geometry, so a text token such as “handle” can attend strongly to a visual patch without that weight being a verified grasp location or metric pose. Nor does the KV cache track object permanence; an explicit state estimator must associate observations across frames and grow its uncertainty while the mug is occluded. The patch-embedding, position-encoding, and attention equations behind both stages are in Visual patch tokenization, attention, and 3D metric lifting.
Specialized Action Decoders (The Outputs): Instead of generating conversational text, the final hidden layer \(\mathbf{h}_{\text{context}}\) feeds task-specific action decoders. While vision-language navigation (VLN) models emit topological waypoints or heading rates \((v_t, \omega_t)\) (figure 8), manipulation (VLA) models diverge sharply in how they bridge semantic representations to physical motion:
- Discrete Categorical Tokenization (RT-1, RT-2, OpenVLA): Early VLAs adapt autoregressive language modeling directly to robotics by discretizing continuous joint or Cartesian offsets into \(B = 256\) uniform bins per dimension (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023; Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023; Kim et al. 2024): \[\text{bin}(a_k) = \left\lfloor \frac{a_k - a_{\min}}{a_{\max} - a_{\min}} \times 255 \right\rfloor \in [0, 255]\] Each bin is assigned a discrete token (
<act_000>to<act_255>) in the language vocabulary. While this natively leverages web-scale vision-language pretraining, it introduces a systems bottleneck: a decoder that emits one scalar token at a time needs seven serial token predictions for a 7-DoF action. Its latency depends on caching, hardware, and whether each prediction revisits the full backbone; seven serial steps can miss a control deadline. Scalar binning has a resolution of roughly \(\Delta a = (a_{\max} - a_{\min}) / 256\); whether that produces contact chatter depends on smoothing, the action range, and the downstream controller. - Continuous Generative Action Chunking (Octo, \(\pi_0\), ACT, Diffusion Policy): To avoid paying a full weight sweep for every scalar action, modern generalist policies couple the transformer backbone with continuous generative decoders, such as diffusion policy heads (Chi et al. 2024; Octo Model Team et al. 2024) or flow-matching action experts (Black et al. 2024). Rather than predicting a single scalar per forward pass, the model predicts an entire continuous trajectory chunk: \[\mathbf{A}_{t:t+H-1} = [\mathbf{a}_t, \mathbf{a}_{t+1}, \dots, \mathbf{a}_{t+H-1}] \in \mathbb{R}^{H \times d_a} \quad (H = 16\text{--}64\text{ steps})\] A backbone may evaluate one new camera observation, then a smaller decoder may perform several denoising iterations to propose a whole chunk of waypoints. That chunk is one candidate horizon, not a sequence of fresh observations. The next backbone evaluation must arrive before the accepted execution window expires; downstream interpolation and feedback run at their own rates.
- Discrete Categorical Tokenization (RT-1, RT-2, OpenVLA): Early VLAs adapt autoregressive language modeling directly to robotics by discretizing continuous joint or Cartesian offsets into \(B = 256\) uniform bins per dimension (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023; Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023; Kim et al. 2024): \[\text{bin}(a_k) = \left\lfloor \frac{a_k - a_{\min}}{a_{\max} - a_{\min}} \times 255 \right\rfloor \in [0, 255]\] Each bin is assigned a discrete token (
The Proposal Boundary (Isolation from Actuators): Whether the model is a VLN mobility planner or a VLA manipulation policy, its output is a candidate proposal (\(p_t\)) on the unprivileged side of the proposal boundary (\(\ref{dfn-boundary-proposal-boundary}\)). Under the machine model of The Machine in Five Levels, the foundation model runs in user space on the application processor and has no path to the pulse-width modulation (PWM) registers or to the current in the motor coils.
The design rule for this deliberative boundary is that learned models operating at the Brain’s inference cadences (\(1\text{--}50\text{ Hz}\)) propose kinematic setpoints (\(\mathbf{q}, \dot{\mathbf{q}}, \ddot{\mathbf{q}}\) or Cartesian poses), never unbuffered, open-loop motor torques (\(\boldsymbol{\tau}\)) directly to inverter switches. Motor torque is coupled to the plant through the manipulator equation of the classical-dynamics note in section 1.2, plus a contact term \(\boldsymbol{\tau}_{\text{contact}}\). Because inertia \(\mathbf{M}(\mathbf{q})\) and Coriolis forces \(\mathbf{C}(\mathbf{q}, \dot{\mathbf{q}})\) vary non-linearly with instantaneous joint configuration and velocity, commanding open-loop torques, without high-frequency feedback, from a neural model that updates at tens of hertz at most can destabilize the limb when model error, friction, or payload variation is large enough. Kinematic setpoints, by contrast, specify an intended spatial path and leave high-gain closed-loop torque synthesis to the Body’s drives, whose field-oriented current loops run at \(10\text{--}25\text{ kHz}\), once the permission path has admitted each setpoint. (Where specialized high-rate reinforcement learning policies operating at \(100\text{--}500\text{ Hz}\) output torques directly, they still require an independently validated high-rate permission path, as Invariant Checking argues; passivity checks for such torque streams are derived in Passivity and Energy-Bounded Interaction.)
↳ Downstream: Zero-copy proposal passing between the Brain and the Nervous System relies on the lock-free buffers analyzed in Multi-Rate Cadences.
Proposals reach the permission path through the handoff of The Nervous System, and the permission path can refuse a hallucinated proposal when independent measurements, modeled limits, and a feasible fallback support that decision.2
Each neural stage of this pipeline is paid for in memory traffic. Every patch token must be attended over and held in the KV cache, and every evaluation of the backbone streams its weights from DRAM, so the capabilities this section described arrive at a rate the application processor sets rather than one the task chooses. On the mobile manipulator, that rate for its largest model falls well below the rate at which its arm consumes setpoints.
Supply, Freshness, and the Memory Wall
At the pick station the mobile manipulator’s arm consumes a new setpoint every few tens of milliseconds, while the mug it is reaching for keeps moving along the conveyor. Whatever model proposes those setpoints must settle two questions. The first is how often it can produce a proposal on the application processor the machine actually has; the second is how long each proposal stays true once the mug has moved. The answers explain why the machine carries two learned proposers rather than one.
After ingestion (the tokenization of section 1.3), four stages turn those tokens into proposals (figure 9), and Part III develops each one. Here they matter only for what each hands downstream and how quickly that goes stale. Perception grounds what it recognizes in metric coordinates relative to the base, so that the mug becomes a pose and a graspable handle rather than a label;3 its output ages from the moment of capture, and Sensor Perception turns that age into a contract. Memory carries those estimates through occlusion and must let each one expire when the evidence behind it does (Spatial Memory). Intent decomposes a task such as “take the ceramic mug off the conveyor and hand it to the coworker at the packing station” into subgoals, and because the mug moves while the intent model deliberates, it hands the planner a target that expires, an intent lease (The Intent Lease). Planning converts that target into candidate setpoints for the base and the arm, the proposals the permission path will judge (Trajectory Planning). These stages, and the ingestion before them, all run on the application processor and draw on the same memory bus, which is where the machine’s cadence is decided.
↳ Downstream: High-throughput neural inference accelerators are isolated from deterministic cores in Two Paths on One Die.
The memory wall
Consider what happens if a robot attempts to generate continuous motor actions one discrete step at a time through an autoregressive foundation model (such as a multi-billion-parameter VLA model). A model whose weights cannot remain on chip must stream them across the external DRAM bus once per evaluation, so the bus bandwidth \(B_{\text{mem}}\) bounds its evaluation rate at \(f_{\max} = B_{\text{mem}} / M_{\text{weights}}\) before tokenization, attention, or decoding adds any time. In single-step autoregressive decoding, each sequential token may require another such evaluation; the actual traffic depends on caching and the decoder architecture. For detailed memory wall derivations and roofline hardware models, see Systems and Hardware.
This memory wall4 places single-token streaming deep in the memory-bound region of the roofline (figure 10) (Williams et al. 2009). On the mobile manipulator’s application processor it caps the intent model’s fresh-evaluation rate, before any vision or decoding work, far below the rate of a joint controller that asks for a new command every millisecond (1.1). A model renewed that slowly can say where the arm should go, but it cannot close the loop on moving mass, so it proposes waypoints and a faster loop below it tracks them.
This continuous data movement also carries an energy penalty. Fetching weights from off-chip DRAM consumes orders of magnitude more energy than the arithmetic computation itself. That energy drains battery reserves, and the heat it leaves in the chassis raises the local ambient of the joint actuators, shrinking the winding margins of Thermal Duty Cycles.
Action chunking
To supply dense waypoints despite slow full-model inference, modern physical AI policies employ action chunking (1.1) (Zhao et al. 2023; Chi et al. 2024). Each chunk then crosses the proposal boundary as one packet, the chunk payload of Multi-Rate Cadences.
A chunking policy (such as Action Chunking with Transformers [ACT], Diffusion Policy, or Flow Matching) generates its whole chunk of \(H\) setpoints in one backbone pass followed, for an iterative decoder, by an assumed \(K = 10\text{--}16\) refinement steps. Each refinement re-evaluates the smaller action head, so full latency is \(t_{\text{vision}}+t_{\text{backbone}}+K t_{\text{head}}+t_{\text{transfer}}\), not just one weight sweep. A new observation changes the plan only after another backbone and decoder pass; see Machine Learning Background for the generative mechanism.
Definition 1.1: Temporal action chunking
Temporal action chunking is the policy formulation that predicts an entire horizon of \(H\) future continuous actions \(\mathbf{A}_{t:t+H-1} = [\mathbf{a}_t, \dots, \mathbf{a}_{t+H-1}]\) in a single neural evaluation or diffusion rollout, amortizing parameter memory transfers across extended physical execution windows.
- Significance: On a bandwidth-limited device, repeating a full model evaluation for every action can be too slow for millisecond control. Temporal action chunking can reuse a backbone evaluation across multiple waypoints; it does not make fresh observations or policy decisions arrive at the waypoint rate. A downstream controller tracks the accepted trajectory between model updates.
- Distinction: Unlike single-step Markovian policies \(\mathbf{a}_t \sim \pi(\cdot \mid \mathbf{o}_t)\) that re-evaluate the full neural network at every control step, action chunking deliberately generates multi-step feedforward trajectories, utilizing receding-horizon execution and blending across overlapping chunks to reconcile slow neural deliberation with fast closed-loop control.
- Common pitfall: Executing action chunks open-loop without continuous temporal smoothing or runtime safety filtering. Without smooth blending across overlapping chunk boundaries, switching between independently generated horizons injects acceleration jumps and torque shocks that excite mechanical gearbox resonances.
For example, when a canonical bimanual manipulator threads a flexible USB cable (figure 5), it executes a 32-step action chunk, a longer chunk than the mobile manipulator uses. The policy emits coordinated 14-DoF joint setpoints that smoothly align both grippers, flex the cable, and insert the connector without pausing between individual control cycles.
However, action chunking introduces the problem of chunk seam continuity. Because each action chunk is generated independently from rolling sensory observations, the tail of chunk \(k\) and the head of chunk \(k+1\) may exhibit slight positional and velocity discrepancies. If a controller concatenated these chunks directly, the resulting step-discontinuities in acceleration would inject large jerk (\(\dddot{\mathbf{q}}\)) into the physical linkages, inducing mechanical shock, acoustic vibration, cycloidal pin wear, and motor overcurrent faults.
To reduce discontinuities between overlapping chunks, systems can implement temporal ensembling, blending overlapping trajectory predictions via exponentially weighted averaging:
\[\mathbf{a}_t = \frac{\sum_{i=0}^{\min(t, H-1)} w_i \, \mathbf{a}_{t \mid t-i}}{\sum_{i=0}^{\min(t, H-1)} w_i}, \quad w_i = \exp(-m \cdot i)\]
where \(\mathbf{a}_{t \mid t-i}\) represents the action predicted for time step \(t\) by the action chunk initiated at time \(t-i\), and \(m > 0\) is a tunable temporal discount factor. The permission path interpolates the blended waypoints below the proposal boundary (Multi-Rate Cadences). Ensembling averages setpoints without enforcing derivative continuity, so a seam can still need the \(C^2\) bridge that Kinodynamic Feasibility builds.
This operational tension between neural deliberation capacity and physical control latency defines the central scaling frontier of modern robotic brains (figure 11). Between 2020 and 2023, pioneering vision-language-action (VLA) models such as RT-1 (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023), RT-2 (Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023), and OpenVLA (Kim et al. 2024) tokenized continuous actions into discrete bins, predicting one action step per autoregressive forward pass. However, because each forward pass requires streaming every model parameter from DRAM across the silicon memory bus, policy update rates collapsed as model capacity scaled (\(f_{\text{action}} \propto 1 / M_{\text{weights}}\))—trapping multi-billion-parameter models behind an insurmountable “deliberation wall” (\(1\text{--}3\,\text{Hz}\)) that fell far below the \(50\,\text{Hz}\) closed-loop threshold required for dynamic contact and reactive manipulation.
Action chunking fundamentally decoupled deliberation from actuation frequency. Instead of streaming billions of parameters for every low-level setpoint, architectures such as ACT (Zhao et al. 2023), Diffusion Policy (Chi et al. 2024), and modern flow-matching foundation models (e.g., Physical Intelligence \(\pi_0\) (Black et al. 2024) and NVIDIA Project GR00T) evaluate their large perceptual backbones periodically to generate entire multi-step action trajectories (\(H = 16\text{--}64\)). Compact, highly optimized generative action heads then denoise or integrate these trajectories at \(50\text{--}100\,\text{Hz}\), shattering the autoregressive memory wall and reconciling multi-billion parameter semantic reasoning with millisecond closed-loop physical control.
The machine’s two proposers
On the mobile manipulator’s own application processor, the memory wall and the amortization that chunking offers decide how the machine divides proposal work between its models.
Napkin Math 1.1: Memory walls, supply, and freshness
1. The single-token memory bottleneck. Consider the mobile manipulator’s intent model, a 7B VLA model stored in 16-bit precision (FP16) in the LPDDR5 memory of its application processor, an embedded edge System-on-Chip:
- Weight footprint: \(M_{\text{weights}} =\) 7B params \(\times\) 2 bytes \(\approx\) 14 GB.
- Theoretical memory bandwidth: \(B_{\text{peak}} =\) 204 GB/s.
- Assumed sustained DRAM bandwidth (70 percent of peak to model bus scheduling and memory overhead): \(B_{\text{sustained}} \approx\) 142.8 GB/s.
Minimum streaming latency to fetch the weights from DRAM across the silicon bus once: \[t_{\text{stream}} = \frac{14 GB}{142.8 GB/s} \approx 98.0 ms\]
Bandwidth-only upper bound on fresh evaluations when each evaluation streams the full weights: \[f_{\text{single}} = \frac{1}{t_{\text{stream}}} \approx \frac{1}{0.0980 s} \approx 10.2 Hz\]
Beyond model weights, autoregressive foundation models also grow their KV cache. When multi-camera sensory feeds (e.g., 3 \(224 \times 224\) RGB streams) are ingested, each video frame contributes hundreds of visual patch tokens (3 \(\times\) 256 \(=\) 768 tokens). In the intent model’s 32-layer transformer with hidden dimension \(d =\) 4096, storing the FP16 Key and Value states consumes 0.52 MB per token: \[M_{\text{KV}} = 2 \times L \times d \times S \times \text{bytes} = 2 \times 32 \times 4096 \times S \times 2 \approx 0.52 MB \cdot S\] Across a 768-token visual prompt, the KV cache expands by 402.7 MB per frame.
2. The action chunking amortization. Suppose the intent model itself emitted an action chunk of \(H =\) 16 future setpoints (\(\Delta t =\) 20 ms, spanning 320 ms) from each forward pass, instead of one action step per DRAM weight sweep.
- Number of weight memory sweeps: 1 sweep (98.0 ms).
- A chunk can reuse one backbone evaluation across \(H\) future setpoints.
- Amortized weight-sweep time per setpoint (excluding all other inference work): \[\tau_{\text{step}} = \frac{t_{\text{stream}}}{H} = \frac{98.0 ms}{16\text{ setpoints}} \approx 6.1 ms\text{ per setpoint}\]
3. The machine’s split. Amortization spreads one sweep across more setpoints, but it does not make a fresh proposal arrive sooner. Each new chunk from the intent model would still wait for at least one 98.0 ms sweep, plus tokenization, attention, and decoding. The mobile manipulator therefore divides the work between two models. Its intent model revises the goal at 5 Hz with 160 ms of inference. A separate 50M-parameter chunk policy, whose 0.10 GB of weights sweep in 0.70 ms, proposes base velocity and arm joint targets together at 20 Hz, within a 40 ms P99 inference budget that includes vision. The intent model’s sweep is too slow to renew a proposal every few tens of milliseconds; the chunk policy’s is not. The Nervous System turns that difference into a lease.
The 98.0 ms sweep in 1.1 is a floor reached before vision, attention, or decoder work begins. Two alternatives to the machine’s split, lower weight precision and a single model with a fast head, would attack it, and neither removes it cheaply. INT4 storage would reduce the nominal weight footprint to 3.5 GB but needs accuracy validation for contact tasks. A single-model design could instead trade cadence for density without changing weight precision. It would evaluate the 7B backbone periodically while a lightweight diffusion head denoises candidate trajectories on-chip, supplying dense candidate waypoints from each slower full-model evaluation. Its new-observation rate would still be limited by full inference latency, including vision and decoder work, which is why the mobile manipulator gives renewal to a separate chunk policy instead.
Retained context adds to the pressure the notebook began to measure. Counting language and proprioception tokens along with the visual patches raises each step to \(S \approx\) 800 tokens, so at 0.52 MB of FP16 key-value state per token each observation step adds 419.4 MB of KV cache. Retaining an unpruned 5 s historical context window at 10 Hz yields 21.0 GB of KV cache. Combined with the intent model’s 14 GB of weights, this example claims 35.0 GB, about half of the application processor’s 68.7 GB before runtime overhead, memory that perception, mapping, and the chunk policy also need. Even where memory capacity permits, reading a large cache is further traffic on that same bus (Contention for Shared Resources). Embodied architectures must therefore avoid unpruned autoregressive visual histories, relying instead on spatial pooling, cross-attention feature projections, or fixed-horizon action chunking.
At the pick station, the mobile manipulator’s arm takes the mug from a conveyor moving at 0.20 m/s, and each chunk must be executed without letting the mug drift out of reach. Untracked, the belt’s steady motion would use up the margin \(r_{\text{tol}} - e_0\) between the grasp tolerance and the handle estimate’s initial error \(e_0\) within the freshness deadline \(e_{\max}/v_{\max}\) of Measurement Freshness, with \(e_{\max} = r_{\text{tol}} - e_0\). A tracker that follows the mug removes that steady motion, but residual slip can still accelerate the target at 0.50 m/s², so a segment executed open-loop for a time \(T\) drifts by \(\frac{1}{2} a_{\text{slip}} T^2\). Added to the 3 mm initial error \(e_0\), that drift reaches the 15 mm grasp tolerance after the tracked evidence horizon \(\tau_{\text{ev}} = \sqrt{2 (r_{\text{tol}} - e_0) / a_{\text{slip}}} \approx\) 219 ms. The chunk policy replans every 50 ms and needs up to 40 ms to do so, so while renewals arrive on time the arm executes no more than 90 ms of motion from one observation, which drifts only 2.0 mm and sits well inside the tracked evidence horizon.
That open-loop span also sets the minimum supply, because a chunk shorter than 5 setpoints would leave the arm waiting for its successor. The machine’s chunk carries 16 setpoints at 20 ms spacing, 320 ms of motion, longer than the tracked evidence horizon; executed open-loop to its end, it would drift 25.6 mm, past the tolerance. The setpoints beyond the next replanning instant are therefore a supply buffer, consumed only when renewals stop, and the lease derived in Multi-Rate Cadences ends execution long before that buffer runs out. Overlapping chunks and temporal ensembling then blend the executed motion across each renewal.
↳ Downstream: Bounded action chunk proposals are filtered against forward-invariant barriers in Safe Sets as Conditional Permission.
These are the proposer limits Part I hands forward: one chunk covers 320 ms of motion, a new chunk takes up to 40 ms, and the evidence behind it stays within tolerance for about 219 ms. They bound when a proposal arrives and how long it remains true, and each enters the handoff record of The Nervous System as a limit record in the form of Measuring a Machine's Own Limits. None of them says whether the proposal was right when it was made.
Checkpoint 1.1: Edge SoC deliberation limits: Memory walls, KV cache, and dynamic drift
Before analyzing runtime failure modes, verify your understanding of edge deliberation limits across memory and time:
Cognitive Vulnerabilities
A chunk the mobile manipulator’s policy proposes for the mug can arrive inside its 40 ms budget, with every setpoint finite and within joint limits, and still close the gripper on the rim. Its timing is sound and its content is not, and the failure modes of learned models live in that gap. Learned foundation models extend task coverage but make a general proof of integrated closed-loop behavior difficult; specified components and bounded operating regions can still be analyzed.5
Without a verified input domain and sensitivity bound, small sensing changes to a learned proposer can yield unexpectedly large changes in proposed motion. Benchmark success does not establish closed-loop stability, timing bounds, or safety for unseen conditions; those properties require separate evidence for the model, runtime, and physical plant.6
A learned proposer’s failure modes also differ from traditional software defects. A classical program fails by throwing an exception, hanging in an infinite loop, or crashing. A neural network fails silently instead, as the mug chunk did, with output that passes every format check. An independent permission path must reject a detected violation before those values can become actuator commands. Systems architects must design around four primary cognitive vulnerabilities.
Semantic hallucinations
A shipping carton behind the conveyor carries a printed photograph of a mug, and the mobile manipulator’s policy proposes a grasp on the printed handle. The trajectory tensor is finite and within joint limits, yet the object it targets is not there. Its plausibility cannot substitute for an independent geometric check, which would find a flat surface where the policy saw a handle.
The Softmax illusion and out-of-distribution overconfidence
Deployment interfaces often confuse mathematical normalization with physical probability. In deep neural networks, output logits are passed through a softmax function,7 whose outputs sum to one whatever the input, so embodied machines encountering out-of-distribution (OOD) observations (novel optical textures, deep shadows, or lens water droplets) can assign confidence scores exceeding \(0.95\) to erroneous classifications.
Evaluating the true relationship between reported model confidence and empirical physical success requires measuring Expected Calibration Error (ECE), the gap between how sure the model thinks it is and how frequently the physical task actually succeeds.8
The regimes in table 1 show how high reported confidence can coexist with low task success under shift. The single-row differences are confidence–success gaps, not ECE; computing ECE requires bins over a specified evaluation set. Raw scores cannot grant execution permission without validation and independent physical checks.
| Distribution Shift Regime | Physical & Environmental Trigger | Conf. vs Success (\(\text{conf} \text{ vs } \text{success}\)) | Confidence–Success Gap (\(\text{conf}-\text{acc}\)) | Downstream Supervisory Action |
|---|---|---|---|---|
| In-Distribution (Nominal Baseline) | Clean optics, structured laboratory lighting, nominal workpiece textures | \(\text{conf}=0.91, \, \text{acc}=0.89\) | \(+0.02\) | Apply independent limits before nominal execution |
| Near-OOD Shift (Covariate Shift) | Specular reflections, lens water droplets, dynamic shadow transitions | \(\text{conf}=0.89, \, \text{acc}=0.52\) | \(+0.37\) | Selective abstention triggered: clamp velocity, initiate exploratory tactile probe |
| Far-OOD Shift (Semantic Shift) | Uncataloged deformable obstacles, human limb intrusion into workspace | \(\text{conf}=0.78, \, \text{acc}=0.08\) | \(+0.70\) | Reject proposal; use feasible controlled response |
| Sensor Degradation (Hardware Drift) | Photodetector thermal noise (\(>75^\circ\text{C}\)), focal blur from vibration | \(\text{conf}=0.84, \, \text{acc}=0.31\) | \(+0.53\) | Trigger operational downgrade; restrict kinematics to certified limits |
Perceptual chattering and covariate shift
Ambiguous images can make successive object classifications change rapidly. If a motion planner treats each new label as a fresh physical state, it may discard useful motion history, as each reclassification did in the Tempe collision (When the Boundary Fails). In physical manipulation, perceptual chattering injects high-frequency direction reversals into trajectory planners, and the commanded reversals demand torque steps that the drivetrain must absorb, oscillations that damage gearboxes. Temporal smoothing can help motion continuity, but it cannot repair a false obstacle estimate; independent detection and a feasible response remain necessary.
Temporal expiration of intent
A proposal stays valid only as long as the evidence behind it, a span the conveyor-pick example computes as the evidence horizon \(\tau_{\text{ev}}\). If the mobile manipulator’s intent model needs 160 ms to deliberate, the people and carts in the aisle, the arm itself, and the mug on the conveyor have already moved by the time the plan is ready. Delayed downstream communication or execution leads to acting on an expired intent lease (The Intent Lease), commanding the physical body based on an obsolete mental representation. A proposer that stalls outright leaves the plant moving with no plan at all, so recovery must live in a local controller below it, a requirement the 2015 DARPA Robotics Challenge Finals imposed on legged entrants by degrading communications deliberately (Krotkov et al. 2017).
The four vulnerabilities, with the memory wall of section 1.4, give three independent reasons against granting a learned proposer actuator authority. The memory wall means the most capable model cannot renew a proposal as fast as the machine consumes one. Hallucination, overconfidence, and chattering mean that a proposal can be wrong or unstable while its reported confidence stays high, so confidence cannot certify it. Expiry means that even a correct proposal stops being correct once the world has moved. Table 2 sets the learned proposer’s limits beside those of the classical methods the chapter began with and shows which side of the proposal boundary each capability occupies.
| Systems Dimension | Classical Analytical Methods | Learned Foundation Models | Physical AI Architectural Synthesis |
|---|---|---|---|
| Foundational Principle | Analytical differential equations, geometry, CAD models, Lyapunov stability | High-capacity neural representations fitted to empirical multimodal data | Hybrid Hierarchy: Learned intent proposals gated by the permission path |
| Generalizability & Perception | Model-dependent: Explicit geometry and features may lose coverage under shift | Data-dependent: Open-vocabulary grounding can improve coverage but may fail under shift | Unprivileged Deliberation: Foundation models synthesize continuous semantic goals |
| Mathematical Verifiability | Conditional: Selected properties can be proved within model assumptions | Limited: Empirical behavior needs separate bounds and runtime evidence | Proposal Boundary: Gate checks modeled limits using independent evidence |
| Temporal Pacing | Hard Real-Time: Deterministic millisecond execution loops (\(1\text{ kHz}\)) | Multi-Rate: Full inference, chunk refresh, and waypoint execution have distinct rates | Split Proposers: A small chunk policy renews while the intent model revises goals |
| Failure Modes & Fallback | Boolean branch crashes; failure on uncataloged objects; rule exhaustion | Semantic hallucinations; out-of-distribution overconfidence; expired intent | Conditional Fallback: Reject detected violations and execute a feasible response |
Fallacies and Pitfalls
Learned capability is useful only when the machine can interpret and constrain its proposals. The following mistakes confuse a model’s success, confidence, or output format with permission to act.
Fallacy: An engineering team must choose entirely between formal classical safety proofs and learned empirical capability.
An assembly plant compares a classical wiring-harness controller with a learned diffusion policy. The classical design exposes explicit motion constraints but struggles with cable deformation; the learned policy completes 96 percent of trial insertions without establishing worst-case force or velocity bounds. These strengths address different engineering obligations. The learned policy can propose adaptive motions while a deterministic supervisor checks them against the physical limits, using a control barrier function where its assumptions apply (Ames et al. 2019) (detailed in Safety Enforcement). Partitioning capability from enforcement lets the team evaluate task performance and the protection mechanism separately. High empirical success does not establish the supervisor’s adequacy or the safety of the integrated machine.
Pitfall: Allowing a neural model’s high confidence to bypass deterministic geometric boundaries.
A cleanroom manipulator mistakes glare on a transparent barrier for open space and proposes reaching through it while carrying reagent vials. The proposal contains finite, in-range coordinates and arrives within its deadline. Those checks establish that the packet is usable, not that its interpretation of the scene is correct. High confidence cannot resolve the same perceptual error that produced the trajectory. Before granting permission, the permission path must compare the proposal with independently maintained geometric boundaries (such as hard-coded stay-out zones or real-time depth maps).
Fallacy: Raw softmax probabilities or predicted confidence scores reflect true physical success rates.
A bin-picking policy reports \(\hat{p} \approx 0.90\) for a group of grasps, but only 64 of 100 physical attempts succeed. An abstention threshold based on the reported score would therefore accept more failures than its numerical value suggests. The discrepancy concerns calibration: confidence does not yet match observed success in the operating conditions (as illustrated in table 1). It does not, by itself, establish whether the policy ranks candidates correctly. Engineers must validate the relationship between scores and outcomes across the intended domain before using confidence to choose between execution, another observation, and operator assistance. Independent physical limits still apply to accepted grasps.
Pitfall: Coupling high-frequency semantic classifications directly to motor setpoints without continuous filtering.
A delivery robot observes reflective tarps over traffic cones. At every perception update, its pipeline changes the classification among obstacle, movable drape, and free path. If each change immediately resets motor setpoints, uncertainty in the scene becomes repeated requests for direction reversal. The drive motors cannot follow those requests beyond their torque slew limits, and the chassis chatters and shudders. The architecture needs continuity between changing interpretations and physical commands. Filtering or action chunking can support that continuity, but the resulting trajectory must still respect the body’s motion limits when the semantic estimate changes.
Fallacy: A numerically valid and smooth action chunk can safely be executed whenever it arrives.
The mobile manipulator’s chunk policy proposes an arm trajectory toward the mug on the conveyor. Host scheduling delays delivery, while residual belt slip has carried the mug toward the edge of the grasp tolerance. The trajectory remains smooth and numerically valid, but the observation that justified it no longer describes where the mug is. Executing it could close the gripper on the rim or knock the mug from the belt. Once the delay outlasts the chunk lease of Multi-Rate Cadences, the lease ends the stalled policy’s authority and the arm enters its validated fallback stop instead of waiting, and when the delayed trajectory finally arrives, the observation it was computed from is too old for the permission path to admit it. Smoothness describes the proposed command; temporal validity determines whether the machine may still act on it.
Summary
Learned representations give the mobile manipulator what hand-authored geometry could not, the ability to find the handle of a mug it has never seen. The price is paid in memory traffic and in time. The largest model evaluates least often, and every setpoint either proposer supplies ages with the observation that produced it.
Key Takeaways: Learned proposers and their deliberation limits
- Weight traffic bounds how often a model sees: One sweep of the intent model’s 14 GB of weights takes 98.0 ms on the application processor before vision or decoding begins, so its fresh-evaluation rate cannot exceed about 10.2 Hz whatever its compute peak.
- Chunking multiplies supply, not freshness: A 16-setpoint chunk spreads one evaluation across 320 ms of motion, but every setpoint is as old as the observation behind it. The machine therefore gives renewal to a small chunk policy at 20 Hz within 40 ms.
- Size a chunk between starvation and drift: The replanning period plus inference latency (90 ms) sets the minimum supply, and target drift sets the evidence horizon (219 ms). Setpoints beyond the next renewal are a buffer that the lease ends, not a plan.
- Retained context is paid for twice: Keeping 5 s of multi-camera history adds 21.0 GB of key-value cache, which with the weights claims about half the processor’s memory, and every read of it is traffic on the bus that real-time transfers share.
- Confidence is not calibration: Softmax scores sum to one by construction, so reported confidence can stay high while task success collapses under distribution shift. Only calibration measured over the operating domain can justify an abstention threshold.
- A proposal is true only for a while: Intent and chunks expire with the evidence behind them, and a smooth, in-range, on-time proposal can still be stale or wrong.
Every limit this chapter derived is a limit of supply or of truth, never of authority. None of them decides whether a proposal may reach the motors; the permission path takes them as inputs when it makes that decision.
What’s Next: From learned proposals to the permission path
Footnotes
Rigid-Body Dynamics Formulation: Classical control computes generalized torques via the standard manipulator equation \(\mathbf{M}(\mathbf{q})\ddot{\mathbf{q}} + \mathbf{C}(\mathbf{q}, \dot{\mathbf{q}})\dot{\mathbf{q}} + \mathbf{g}(\mathbf{q}) = \boldsymbol{\tau}\), where \(\mathbf{M}(\mathbf{q})\) represents the positive-definite inertia matrix, \(\mathbf{C}(\mathbf{q}, \dot{\mathbf{q}})\) captures Coriolis and centripetal terms, and \(\mathbf{g}(\mathbf{q})\) accounts for gravitational loading. This analytical formulation becomes computationally intractable and physically inaccurate when machines interact with unmodeled frictional contacts, fluid drag, or deformable objects. Derivations are provided in Control and Dynamics.↩︎
Proposal Boundary: Treating the neural policy as an unprivileged proposal generator isolates stochastic inference from actuator registers. Candidate trajectory setpoints \(\mathbf{q}_{\text{ref}}(t)\) are checked on dedicated safety silicon, so proposals that violate modeled actuator limits or keep-out regions can be rejected before any reaches a drive (Actuation Authority).↩︎
Metric Spatial Grounding: While disembodied AI operates over discrete semantic tokens, physical AI must ground those semantics in the Special Euclidean group \(SE(3)\), mapping entities to metric distances (meters) and spatial orientations (quaternions) relative to the robot’s base coordinate frame. Failing to maintain this metric grounding causes neural policies to emit kinematically impossible Cartesian paths that breach actuator workspace boundaries.↩︎
Autoregressive DRAM Bandwidth Saturation: Sequential action tokens can trigger repeated full-model weight traffic when weights are nonresident and the decoder revisits the backbone; cache behavior and hardware determine the actual bandwidth ceiling. Action chunking is one way to reduce the number of full policy evaluations per unit of physical time.↩︎
Connectionist and Symbolic Duality: Physical AI architectures reconcile statistical connectionism with deterministic symbolic computation: high-capacity neural networks synthesize open-world behavioral proposals, while formal symbolic logic and invariant monitors verify kinematic validity before setpoints are granted execution authority.↩︎
Lipschitz Continuity in Neural Policies: Although neural networks with smooth activations are continuous on compact domains, a bound \(L = \sup_{\mathbf{x}_1 \ne \mathbf{x}_2} \frac{\lVert \pi(\mathbf{x}_1) - \pi(\mathbf{x}_2) \rVert}{\lVert \mathbf{x}_1 - \mathbf{x}_2 \rVert}\) must be established over a specified operating domain. If a valid bound exists, sensor error \(\lVert \delta \mathbf{x} \rVert \le \epsilon\) implies \(\lVert \delta \mathbf{a} \rVert \le L\epsilon\); an unverified \(L\) supplies no usable action-error guarantee. Mathematical formulations appear in Machine Learning Background.↩︎
Softmax Normalization and Epistemic Uncertainty: The softmax operator \(\sigma(\mathbf{z})_i = \frac{e^{z_i}}{\sum_j e^{z_j}}\) enforces normalization across the categorical simplex \(\Delta^K\) by algebraic construction rather than Bayesian epistemic confidence. High output confidence reflects only relative logit separation, meaning models assign near-unity probabilities even when evaluating inputs completely outside their training distribution. Derivations appear in Machine Learning Background.↩︎
Expected Calibration Error Formulation: Partition predicted confidence into bins. In each bin, compare observed success frequency with mean reported confidence; average the absolute gaps, weighted by the bin’s share of samples. Elevated ECE shows that reported confidence is a poor basis for granting physical permission. The equation and bin definitions appear in Softmax overconfidence and calibration error.↩︎







