Multi-Rate Cadences; the machine’s intent model runs faster than the deliberative rate shown, and the chunk policy is omitted here.">
Grounded Intent
Grounded Intent
Purpose
What must an autonomous goal carry to prevent a slow reasoning model from driving a machine into a situation that no longer exists?
A high-level semantic command such as ‘retrieve the pallet’ specifies a desired outcome while leaving essential physical parameters—trajectory, contact force, and temporal bounds—completely unresolved. Because the physical scene continuously evolves while large reasoning or vision-language models deliberate, raw model outputs cannot be executed without physical grounding. A target that was reachable when deliberative inference began may become fully obstructed before execution commences, even if the model’s high-level task decomposition was logically flawless.
Grounded intent translates abstract directives into physically admissible regions constrained by explicit leases, spatial tolerances, and hard deadlines. When an environmental lease expires or the target drifts beyond certified bounds, safe deceleration cannot wait for a high-latency deliberative loop to acknowledge the timeout. The real-time execution layer must independently revoke motion authority and execute an assured fallback. In the physical AI stack, intent bridges the Brain’s semantic goals to the Body’s physical limits, granting the deterministic Nervous System bounded leases that expire safely when unrenewed.
Learning Objectives
- Construct an intent lease from a belief record, with a goal region, error envelope, task tolerance, and effort ceilings
- Derive a target-evidence horizon from task tolerance and target drift, and distinguish it from the stopping deadline
- Decide when grounding ambiguity or a wrong-object premise requires refusal rather than admission
- Evaluate whether a proposed goal is reachable within its remaining evidence window
- Design lease expiry that revokes motion authority even when the issuing process cannot cancel it
Grounded Affordances
When the warehouse mobile manipulator is told to “pick the red mug” from the takeaway conveyor, the directive specifies an outcome but leaves path, force, and deadline unresolved. A motor inverter accepts electrical commands, not language. Vision-Language-Action (VLA) models (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023; Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023) can propose actions or targets, but their update cadence depends on the model and hardware and may be much slower than local control. If a target moves during inference, a stale proposal can consume clearance. The system therefore needs dated geometric evidence, bounded requests, and an independent permission path that can withdraw authority while the model is silent.
↰ Prerequisite: The belief record a lease consumes (evidence epoch, initial error, growth law) is specified in The Spatial State Schema.
Memory can supply a spatial belief with an expiry (Belief Through Occlusion); it cannot decide what action that belief authorizes. Intent closes that gap by connecting high-level task semantics to a bounded physical goal. It specifies an admissible, time-bounded target before the trajectory planner of Trajectory Planning constructs a continuous motion profile to reach it.
Connecting words to physical targets is the robotics form of what Harnad termed the symbol grounding problem (Harnad 1990): purely symbolic or linguistic tokens within a computational system have no intrinsic physical meaning unless causally grounded in non-symbolic, sensorimotor interactions with the physical environment.1 Whereas early artificial intelligence systems such as Winograd’s SHRDLU (Winograd 1972) operated inside closed-world blocks domains with synthetic, analytic geometries,2 modern physical AI systems must ground open-vocabulary instructions into continuous, dynamic, and uncertain physical spaces.3
Definition 1.1: Intent grounding refusal
Intent grounding refusal is the admission rule whereby an embodied system explicitly rejects semantically ambiguous natural-language directives, returning a typed failure status (ERR_GROUNDING_AMBIGUITY) rather than executing ungrounded maximum-likelihood guesses across physical clearance margins.
- Significance: Rejects a proposal when a task-specific ambiguity test (section 1.6) fails, before that proposal reaches trajectory generation.
- Distinction: Unlike a trajectory planning failure that cannot find a collision-free path to a valid goal, grounding refusal rejects the goal itself as physically ungrounded or ambiguous before the planner is invoked.
- Common pitfall: Forcing an argmax token decode or nearest-neighbor guess when candidate referents remain unresolved, driving the physical plant toward an arbitrary or hazardous target.
Whatever the model proposes must also remain bounded in time. Figure 1 gives a schedule in which a slow deliberative model, local tracker, permission loop, and drive run at different rates. A conveyor target whose evidence expires within tens of milliseconds cannot be renewed by a model updating every few hundred milliseconds, so it requires a separate validated local observation path or a task pause. The figure’s rates are illustrative, and section 1.4 sets the machine’s own intent cadence against the conveyor horizon.
↰ Prerequisite: The proposal boundary and the permission path are defined in The Machine in Five Levels.
Grounding an open-ended directive therefore ends in a dated intent proposal, which the application processor sends across the proposal boundary. It requests bounds in space, time, and effort, and the independent permission path tests each against configured limits, current state, and stopping clearance before motion begins.
What Intent Hands Over
When the arm reaches for the red mug, the Brain sends neither joint trajectories nor winding voltages across the proposal boundary. It sends a request for bounded authority, and the permission path must be able to check every part of it before it grants any of it. The checks are target tolerance, evidence validity, requested effort, and a feasible fallback, and each exists because a delayed proposal can fail in a different way. Section 1.7 names the record that carries the request and its belief-record lineage.
In physical manipulation, an isolated coordinate is incomplete. In Gibson’s sense, an affordance is an action-relative relation among geometry, robot reach, and contact limits (Gibson 1979). The intent proposal therefore requests a target pose in \(SE(3)\), the space of three-dimensional positions and orientations, and a task tolerance region within which the task counts as complete. This tolerance differs from the estimator’s covariance, which describes uncertainty about the target. For the machine’s mug grasp, the approved task profile allows 15 mm of translational error. A controller can then terminate within that region rather than integrating forever toward an exact noisy point; it still checks that estimated target error fits inside the tolerance.
Spatial validity in the physical world is transient. A physical goal therefore carries an evidence epoch and a finite requested expiry rather than relying on an asynchronous cancellation message4 from the reasoning layer. A model that produces a new result only once per period cannot, even before jitter, renew a target whose evidence stays valid for less than that period. An independent permission controller evaluates the admitted deadline on each control cycle and drops tracking when the evidence or lease expires. It selects a configured hold or stop only when the measured state and remaining clearance make that response feasible.5 The local timer removes dependence on a host abort packet; the stopping distance remains a separate physical obligation.
A target bounded in space and time is hazardous if the controller is allowed unlimited physical effort. The intent proposal therefore requests effort ceilings on speed, acceleration, and contact wrench. A model may place a target slightly behind the surface it must touch because of stereo disparity error; a stiff position controller can then command excessive restoring torque. For the arm’s guarded approach to the spring-latched cage door, the proposal might request the 15 N tripwire as its contact-force ceiling, well below the 100 N latch limit, and the 0.03 m/s approach speed as its speed cap. The permission path compares each request with independently configured actuator and task ceilings, then admits only a feasible subset. The low-level impedance controller6 enforces the admitted limits.
Omitting any of these three bounds leaves a machine exposed during communication dropouts or perception stalls. Stopping distance grows linearly with the delay before braking and quadratically with speed (equation). On the moving base, whose stopping budget is a distance, a stalled proposer lengthens the delay term by the time its last proposal keeps authority. That time is the chunk lease of Multi-Rate Cadences, and it enters \(\tau_{\text{delay}}\). Clearance holds only if an independently checked obstacle boundary lies beyond the stopping distance, with model and sensing margins added.
The intent lease is a different object. It bounds how long a target stays justified, and its expiry withdraws the target and ends tracking at once, without adding a term to the base’s stopping budget.
Napkin Math 1.1: Lapsed authority and physical response
The three models below answer different questions. A stopping calculation needs the total delay to brake onset and the deceleration available from the current state.
- The mug intent lapse: The arm reaches for the red mug under an intent admitted without local tracking. At 0.20 m/s the conveyor consumes the 12 mm between the 3 mm initial error and the 15 mm tolerance in 60 ms, so the target’s evidence expires then, as the mug-handle belief record of The Spatial State Schema states. Suppose the application processor stalls. A target held for 200 ms by the stalled proposer would sit 43 mm from the mug. The lease does not wait for the proposer. At 60 ms the intent expires, the arm’s tracking authority ends, and the arm enters a stop matched to its current state, with onset within 22 ms (one permission tick, one bus cycle, and the brake onset). The chunk lease of Multi-Rate Cadences is the backstop if that expiry is never observed. Under local tracking the same 200 ms leaves the target 13 mm off, inside the tolerance, because the tracked horizon derived in section 1.4 is 219 ms; tracking is what makes the 5 Hz intent refresh viable.
- The arm at the cage door: At the guarded approach speed of 0.03 m/s against the 4.0 × 10⁵ N/m strike plate of the cage door, modeled force rises at 12 N/ms. A 15 N tripwire and 2 ms contact-loop response yield 39 N when the brake command begins; force keeps rising while the arm decelerates. Staying under the 100 N latch limit leaves only 0.15 mm of further compression, so the arm must decelerate at about 3 m/s² and stop within about 10 ms. Whether the arm can brake that hard from the approach is an open premise, which only arm stopping trials can settle. Unbraked motion for 200 ms advances 6 mm and gives 2,400 N in the linear-spring model, far past the latch limit.
- A process-plant contrast: A heated process has no braking distance, but its setpoint’s authority lapses the same way, spent as temperature rather than travel. For an illustrative lumped capacitance of 900 J/K and uncompensated 3600 W input, temperature rises at 4 K/s. A chosen 2 K task overshoot allowance is consumed in 500 ms, so a heater setpoint whose proposer stalls must lose authority well inside that interval; 10 s at the same net heat input yields a 40 K modeled rise.
These three examples clarify the division of authority. High-level reasoning can propose a target, time window, and effort ceiling. The independent permission path determines which limits are admissible from measured state, configured ceilings, and a feasible fallback. Local control then tracks only an admitted proposal and withdraws authority on invalid evidence or expiry. This resembles the priority separation of the subsumption architecture7 (Brooks 1986).
Grounding an intent proposal into these three bounds requires a formal transformation from ambiguous sensory inputs into structured, verifiable proposals. When a system processes a high-level directive or a sequence of visual tokens, the initial proposal is rarely a single, unambiguous set of coordinates. It begins as an uncertain geometric hypothesis that must carry its own spatial ambiguity and its own temporal limits before it can be admitted to the execution pipeline.
From Request to Expiring Geometric Proposal
A natural language instruction becomes dangerous when an unverified semantic hypothesis commands motion. When an operator or high-level supervisor asks for the red mug, an onboard vision-language model parses the character stream, projects image features into a shared embedding space, and selects a candidate visual bounding region. From those pixels, the network extracts an estimated pose in \(SE(3)\) and writes coordinate values into a memory buffer. At the proposal boundary (The Machine in Five Levels), an ungrounded string of tokens becomes a requested destination; if admitted, it reaches the causal boundary as mass, velocity, and current.8
If the visual backbone associates the language token with an adjacent fixture or an unlabeled calibration target, the resulting coordinate vector \(\mathbf{p} \in \mathbb{R}^3\) enters the trajectory pipeline as an authoritative physical goal. Because a low-level motor drive measures only encoder pulses, winding temperatures, and phase currents, it cannot evaluate whether the coordinate it tracks corresponds to the intended mug or to an empty stretch of belt.9 The grounding process must produce an explicit geometric proposal that includes spatial uncertainty, selection evidence, and temporal expiration before any motion planner accepts it.
Open-vocabulary affordance grounding uses language and sensor observations to propose a task-relevant skill, spatial field, or program, and none of its outputs is a verified target. SayCan (Ahn et al. 2023) multiplies a language model’s preference by a learned estimate of each skill’s success (Value Functions and Affordance Grounding), which yields a ranking conditioned on training and observation, not a veto on obstructed or impossible motion. VoxPoser (Huang et al. 2023) (figure 2) has a language model write code that composes 3D attraction and repulsion value maps in the robot’s workspace for a trajectory optimizer to use as costs, and gradients over a cost field certify neither collision freedom nor torque limits. Code as Policies (Liang et al. 2023) generates programs that call perception APIs and parameterized motion primitives, and a generated program can stall as surely as a model can, so the countdown and the permission path keep control whatever the program does. OK-Robot (figure 3) filters candidate grasps by language-conditioned object selection, and the segmentation mask behind such a selection is a candidate image region, not a verified 3D hull. Table 1 sets out what each kind of representation sends downstream and what the permission path still needs before any of them can become an admitted request. Each needs at least a named frame, SI units, dated evidence, and the permission path’s own checks.
| Proposal source | Representation sent downstream | Early check | Remaining evidence needed |
|---|---|---|---|
| 2D box or mask | Image region plus measured depth and camera calibration | Reject absent depth or a target outside a validated region | 3D pose error, occlusion, object identity, and swept-path clearance |
| 3D feature or value field | Candidate target or cost field in a stated frame | Reject out-of-workspace targets and stale map cells | Conservative occupancy, unknown-space policy, actuator and timing bounds |
| Action-token VLA | Decoded action in its trained action range | Reject out-of-range or inadmissible proposals | State-dependent physical limit and independent permission; token bins provide no clearance proof |
| Language-selected skill | Skill identifier with estimated success | Rank or decline low-confidence candidates | Current preconditions, path feasibility, and separate physical enforcement |
Discretized outputs carry no metric bound, and the two common kinds lack it for different reasons. RT-1 (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023) and OpenVLA (Kim et al. 2024) quantize each robot action dimension into 256 bins over a trained range, so a decoded bin is an action and encodes no clearance to any target. PaliGemma’s location tokens quantize normalized image coordinates, which become a base-frame pose only with measured depth, a calibration chain, and an error model whose extrinsic and depth terms can dominate the quantization itself.
To prevent a learned output from masquerading as a deterministic setpoint, the producer must retain the referent identity and the dated evidence that selected it. Its observation record may contain bounding boxes, attention weights, class probabilities, and a full oriented covariance estimate. The intent proposal (section 1.7) links to the belief record built from that observation (The Spatial State Schema) by digest and evidence epoch; it carries a candidate pose, conservative base-frame error bounds, task tolerances, and requested expiry, not those full producer-side fields. The independent permission path checks the linked evidence and request against its own limits. An object can shift or become occluded after camera exposure, so a coordinate without a valid evidence link and deadline cannot remain an authorized target.
When a vision-language model sees multiple candidate objects, an argmax can hide an unresolved choice. A digital recommendation can still mislead a user, but it has no direct actuator path; in a robot, an admitted spatial proposal can move mass. If a red glass jar sits beside the red mug on the takeaway conveyor, a poorly grounded guess may select the jar. The system should preserve candidate identity and calibrated uncertainty, refusing a proposal when a task-specific ambiguity threshold is exceeded. The permission path then maintains a feasible hold or stop while the producer seeks a clearer observation or instruction.
Formally, open-vocabulary grounding ambiguity is the condition where multimodal feature projections yield multi-peak posterior distributions, overlapping geometric hypotheses, or unresolvable evidential ties across candidate physical referents, requiring an explicit grounding refusal (ERR_GROUNDING_AMBIGUITY) rather than optimistic argmax tie-breaking. A high entropy over the candidate-referent distribution \(\mathbf{z}\), \(H(\mathbf{z}) > H_{\text{thresh}}\), is one signal of it, and the producer-side filter of figure 4 can refuse on that signal before the proposal crosses the proposal boundary. The test that must hold for every proposal is an independent count of the candidates in the pick window (section 1.6), because a confidently wrong model can report low entropy.
Consider the grounding model’s estimate of the red mug’s handle from a single wrist-camera view taken at a distance, before the arm begins its approach. In an illustrative calibrated Gaussian localization model, the principal sensor-frame standard deviations are 4 mm, 5 mm, and 12 mm, the largest along the camera’s depth axis, which is also the gripper’s approach axis. A 95 percent statistical region around the estimated handle position \(\mathbf{p}_0\) uses Mahalanobis squared distance \((\mathbf{p}-\mathbf{p}_0)^T\mathbf{\Sigma}^{-1}(\mathbf{p}-\mathbf{p}_0)\le\chi^2_{3,0.95}\), where the three-degree-of-freedom quantile is 7.815, yielding principal semi-axes 11.2 mm, 14.0 mm, and 33.5 mm. Figure 5 contrasts that uncertain target with the mug’s 15 mm grasp tolerance. Both lateral semi-axes fit inside the tolerance, but the depth semi-axis does not, so this estimate fails admission and needs better sensing, a different approach, or refusal.
The pick admitted in section 1.7 rests on better sensing. Its belief record (The Spatial State Schema) comes from a closer wrist-camera capture on the approach, which bounds the initial error at 3 mm. Even that admitted estimate’s 95 percent region leaves a residual tail and cannot by itself certify clearance, so the intent proposal carries conservative base-frame error bounds rather than claiming to encode this covariance.
A detection with a large uncertainty estimate can be rejected by a configured statistical gate. A confident grounding to the wrong physical entity can pass that gate. Suppose instead that a thin-walled glass jar arrives alone where the mug should be, and the vision-language model assigns it a high confidence score when asked for the red mug. The model emits a sharp, low-variance geometric proposal centered on the glass. The trajectory planner receives the coordinate, computes an \(SE(3)\) inverse kinematics solution, verifies joint limits, and confirms that the path clears all obstacles in the static environment map. The permission path may find the proposed motion within its configured geometric and force bounds. If those bounds allow the grip force the ceramic mug needs, a smooth, geometrically valid trajectory can crush the glass despite passing the checks.
This failure reveals the boundary between what the permission path can verify and what it must accept as an unvalidated premise. The permission path checks coordinate bounds, kinematic reachability, joint torque limits, and geometric collisions against known occupancy grids. Those physical checks do not establish that the coordinate names the intended object, so the premise that the grasped object is the named object lies outside the permission path’s reach. It needs its own task evidence, and it travels with the mug into the handover branch of the release case (A Case Worked in Full) and into the register of claims the machine’s evidence leaves open (A Residual-Claims Register).
Multimodal feature extraction, cross-attention alignment, and open-vocabulary phrase grounding belong to the computer vision and language literature (Kamath et al. 2021; Gu et al. 2022; Radford et al. 2021; Liu et al. 2023). A physical AI architecture consumes these grounding models without relying on their internal activations for safety, and even a filtered proposal stream still needs the independent motion-permission check. AutoRT’s language-model task filter raised the fraction of acceptable proposed tasks yet still passed a small fraction of unsafe ones, and its deployment relied on joint-force pausing, emergency stops, and human supervision rather than on the prompt (1.1) (Ahn et al. 2024).
War Story 1.1: AutoRT task filtering and supervision (2024)
Evidence: In a labeled sample of 259 proposed tasks, 228 were initially judged acceptable. After an LLM affordance filter, 200 of 214 retained tasks were acceptable, raising the acceptable fraction from about 88 to 93 percent while leaving some unsuitable tasks. The paper explicitly says constitutional prompting cannot guarantee compliance.
Response: The reported deployment supplemented prompts with joint-force pausing, physical emergency stops, and human line-of-sight supervision unless a work area was barricaded. These are distinct controls with different coverage; the prompt itself was not a certified geometric or physical safety supervisor.
Systems lesson: A learned task filter can improve the proposal stream, while an independent permission and fallback path must still validate each physical action for its stated operating region.
Whatever its source, a grounded target is also a claim about a moment. The red mug that the model selected correctly at capture keeps moving with the conveyor, so a proposal that was right when issued becomes wrong with no change to its contents. It must therefore carry a deadline that the permission path can enforce without the model.
Goals That Expire
Suppose the application processor hangs while the arm reaches for the red mug. The fault that makes cancellation necessary is the same fault that stops the cancellation from being sent, because an explicit abort requires a working sender, scheduler, and communication path. An intent lease must therefore expire on its own. This is principle \(\ref{pri-vol4-proposal-permission}\) applied to a goal. The permission path owns a fallback it can execute without the proposer, so an independent countdown withdraws the stalled proposal’s tracking authority without waiting for the model. This adapts the lease primitive of Gray and Cheriton (1989) to a physical controller: expiry triggers a separately validated fallback, whose ability to stop before a hazard depends on current state, clearance, actuator authority, and total response delay.10
↰ Prerequisite: Kinetic momentum budgets and stopping distance constraints originate in Kinetic Momentum.
The target-evidence deadline follows the rate at which the physical target can diverge from its last measurement. Suppose validated bounds cap relative target drift at \(v_{\text{drift}}\) and initial position error at \(e_0\) within the belief record’s declared validity region, and the task allows a target error of \(r_{\text{tol}}\). The horizon is then equation applied to equation with \(a_{\text{dist}}=0\), writing \(e_0\) for \(E_0\) and \(r_{\text{tol}}\) for \(E_{\max}\): \[\tau_{\text{ev}} \le \frac{r_{\text{tol}} - e_0}{v_{\text{drift}}} \tag{1}\]
When \(r_{\text{tol}}\le e_0\), the target is inadmissible at capture. The controller refuses dependent motion. Even when equation 1 gives a positive horizon, a separate motion-permission check must verify current clearance against total sensing, permission, communication, and brake-onset delay plus physical braking distance, as in The Nervous System.
If a validated disturbance acceleration bound \(a_{\text{dist}}\) is also needed, the quadratic term of equation returns, and its positive root at \(r_{\text{tol}}\) shortens the evidence horizon.
For the red mug on the takeaway conveyor, the grounded target moves at 0.20 m/s. A 15 mm grasp tolerance and 3 mm initial error leave a 60 ms target-evidence horizon by equation 1. Stretching the horizon to the 200 ms refresh period of the machine’s 5 Hz intent model would let modeled target error reach 43 mm, or 28 mm beyond the tolerance. The intent model, whose inference alone takes 160 ms, cannot by itself refresh a 60 ms target for continuous motion. Without a qualified tracker the conveyor must slow or pause, or the task must be refused. At 20 mm/s drift the same bound gives a 600 ms evidence horizon (figure 6).
A qualified local tracker changes the arithmetic. It re-measures the mug on every frame and follows the conveyor’s velocity, so the conveyor speed no longer drives the error. The same 3 mm initial error is still charged against the tolerance, which leaves 12 mm for the residual slip of 0.50 m/s² to consume. Against that budget, residual slip alone sets the tracked evidence horizon that Supply, Freshness, and the Memory Wall solves, 219 ms. The tracker renews that evidence at its own rate, so the tracker’s period plus latency is what must fit inside the horizon. The intent model only re-confirms which mug is meant, and each renewal it sends carries the tracker’s fresh evidence rather than the frame the model last read. Whichever horizon applies caps the intent lease on this target (section 1.7).
Coupling target-evidence validity to the reasoning model’s cadence creates a drift hazard. Suppose a design ties the mug’s evidence to the intent model’s output rather than to a tracker, so the untracked 60 ms horizon applies, and a single inference runs to a high-tail latency of 650 ms. Extending the target record to 650 ms solely to hide that delay would leave its estimate outside the declared error envelope for 590 ms. The 60 ms limit is an evidence-validity deadline, and only new evidence, checked again against the motion, can extend it; faster compute cannot.
Continuous tracking requires fresh physical evidence, not merely a new packet from the same stale frame. A renewal based on a new capture must arrive before the current evidence and admitted authority expire, with capture, processing, transfer, clock uncertainty, and permission time included in its budget. The simple condition \(T_{\text{period}}+t_{\text{latency}}<\tau_{\text{ev}}\) is only a necessary scheduling check. When a fresh proposal arrives, the planner may compute a jerk-limited splice from the measured state. It admits that splice only if the swept path, new evidence, and stopping reserve remain feasible; an abrupt target change can instead require refusal or a checked stop.
When the countdown the MCU armed at admission runs out, whether or not the host is still running, the independent controller withdraws the proposal’s authority and selects a previously validated response for the measured load and contact mode. An active position hold needs torque and can protect a suspended payload; controlled deceleration needs brake capability and clearance; torque disable can be suitable for some drives but can release a load or leave momentum uncontrolled. The proposer may request a mode, but it cannot choose the terminal action for the MCU. Admission fails when no response remains feasible from the current state.
Executing an expiring lease safely assumes that the commanded geometric proposal is actually achievable by the physical body. If a proposed coordinate lies beyond kinematic reach or demands joint velocities that exceed actuator limits during the lease window, granting the lease wastes control authority and can trip the drive’s overcurrent protection as a motor draws excessive current toward an impossible speed. Determining whether a goal is kinematically and dynamically admissible belongs before the trajectory generator accepts the lease, establishing a physical filter between high-level intent and motion execution.
Admissible Goal Sets
If the red mug rides past on the far side of the belt, beyond the arm’s reach from where the base is parked, no sequence of joint torques can bring the gripper to its handle. That impossibility should be caught at the proposal boundary, before the target enters the trajectory optimization queue. When an upstream vision-language model or neural policy emits a coordinate, treating that proposal as an unverified candidate protects the real-time motion pipeline from executing or even evaluating impossible motions. Kinematic workspace analysis, imported from classical formulations (Craig 2005; Lynch and Park 2017), defines the closed spatial manifold of reachable end-effector configurations (the physical 3D volume the robot arm can reach) given link dimensions, joint range limits, and mounting geometry.11 An analytical reachability check evaluates whether the target \(SE(3)\) pose \(\mathbf{p}_{\text{target}}\) resides inside this workspace manifold \(\mathcal{W}\), while an occupancy check verifies that the bounding volume of the target does not intersect known persistent obstacles, and both are cheap enough to run before trajectory optimization (figure 6, panel b). Passing an out-of-reach goal to a trajectory optimizer forces the solver to search an empty solution space, converting geometric impossibility into computing failure. Passing these cheap tests is necessary but never sufficient, because joint orientation, swept obstacles, torque, thermal state, unknown space, and stopping clearance still require a checked path and a permission decision.
↳ Downstream: Admissible goal regions specify the boundary conditions for spline generation in Continuous Trajectories.
A target inside the static workspace may still be out of reach within the time its evidence stays valid. Dynamic reachability is the condition that the end effector can move from its measured state \((\mathbf{p}_0, \mathbf{v}_0)\) to the target pose within the remaining window \(t_{\text{exp}} - t_{\text{now}}\) while respecting its velocity, acceleration, jerk, and thermal torque limits. Whatever profile the actuators follow, no move covers a distance \(d\) faster than its speed ceiling allows, so admission requires at least \[\frac{d}{v_{\max}} \le t_{\text{exp}} - t_{\text{now}}.\] Acceleration and jerk limits only lengthen the transit, so a proposal that fails this inequality fails every fuller test. Kinematic Reachability and Intent Envelopes derives the triangular, trapezoidal, and jerk-limited profiles those limits produce.
On the arm, the mug’s evidence under the qualified tracker stays valid for 219 ms, and at the 1 m/s tool-center-point speed limit the gripper covers at most 219 mm in that window. A proposal issued with the gripper 300 mm from the handle needs at least 300 ms of transit, so the MCU rejects it with ERR_DYNAMIC_TIME_INSUFFICIENT before any path is solved. Without the tracker the window shrinks to 60 ms and the reach to 60 mm, so an untracked reach can close only the last few centimeters of an approach. Passing the inequality only permits the fuller path, torque, and stopping check.
This bound and the fuller checks behind it rest on imported rigid-body mechanics under three assumptions, namely that actuator torque-speed curves remain within nominal thermal ratings, joint encoders accurately report the initial state \((\mathbf{p}_0, \mathbf{v}_0)\), and the payload mass matches the inertial model behind the actuator limits. A representative permission-path implementation monitors joint current, motor temperature, and communication delay at rates chosen from the plant and sensor budgets. These observations can detect some departures from the model but cannot verify every assumption, such as payload inertia, on every tick.
If motor winding temperatures exceed the thermal thresholds of Thermal Duty Cycles, the permission path imposes thermal derating, scaling down the maximum allowed acceleration (\(a_{\max}\)) and velocity (\(v_{\max}\)) in the parameter registry. An external force measurement that indicates a payload mass deviation lowers the same limits. The intent filter immediately evaluates subsequent incoming leases against these reduced physical limits. A derating event can make a previously admitted path infeasible; the controller must recheck it and select a feasible fallback if the margin is lost.
When the reachability filter rejects an intent proposal, it must not simply discard the packet. Discarding without feedback leaves the upstream reasoning model unaware of why motion failed to occur, forcing it to wait for an end-to-end task timeout. The intent filter generates an immediate, structured diagnostic packet returned directly to the reasoning layer across the proposal boundary. This diagnostic payload contains the rejection classification, such as a kinematic workspace violation (ERR_WORKSPACE_OUT_OF_BOUNDS), dynamic time insufficiency (ERR_DYNAMIC_TIME_INSUFFICIENT), or swept-volume collision intersection (ERR_OCCUPANCY_BLOCKED), along with the numerical margins of failure. For a dynamic reachability failure, the feedback returns the shortfall between the transit lower bound and the remaining window, and the maximum reachable distance along the commanded vector within the allotted lease. For a static workspace violation, the feedback provides the projection of the target onto the boundary of \(\mathcal{W}\) (the closest physical point the arm can actually reach). The reasoning layer uses this structured signal to adjust its proposal on the subsequent inference cycle, by obtaining fresher evidence, choosing an alternate approach, slowing or pausing the process, or proposing base relocation. It cannot extend the evidence horizon without a new justification.
Reachability filters remove targets the body cannot reach, or cannot reach in time, and leave the complete motion to the planner and the permission path.
How Intent Goes Wrong
When a vision-language model outputs an \(SE(3)\) target pose for a physical manipulator, that pose can fail in several distinct ways. It may be a hallucinated target generated from visual attention patterns on reflective background textures. It may be a stale target representing an object that a human coworker bumped or displaced moments earlier while the camera pipeline was buffering frames. It may be a reference ambiguity where the model arbitrarily selects one of two identical red mugs side by side on the conveyor without resolving which one the task requires. Or it may stem from an out-of-distribution scene shift, such as unexpected lighting or dust on an optical lens, that degrades the spatial prediction accuracy of the network. Despite their different causes, these failures arrive downstream in the same form.
The planner receives a dated target proposal, not an actuator command. A pose on an empty location may lead to a futile reach; a pose inside an occupied fixture may lead to contact if the later path and permission checks fail. In the wrong-object case, by contrast, the geometry can be valid while the selected object is semantically wrong. The independent permission path therefore checks what it can measure (freshness, coverage, geometry, current state, and actuator limits) without claiming to verify the task’s meaning.
A fresh depth observation can reject a proposed grasp whose required object evidence is absent in a covered, visible region. No return in an occluded or unobserved region leaves the region unknown. A task-specific ambiguity test runs on every proposal, confident or not, because a high score describes the model’s certainty and cannot show that no second candidate shares the pick window. The test refuses a window that holds two candidates, and the item goes to the coworker for a manual pick. The permission path runs it on its side of the proposal boundary, as the ambiguity gate of figure 6 (panel b), on evidence the grounding model does not produce; the producer’s entropy filter can refuse earlier but cannot replace it.
Table 2 groups these failures by the check that catches them, adds two that the model’s output cannot reveal until contact or timing exposes them (a misclassified affordance and a target unreachable within its window), and separates admission checks from responses after motion or contact has begun. At the cage-door latch of 1.1, the 2 ms contact-loop response already leaves 39 N at brake-command onset, and the peak comes later, so precontact speed and force limits, actuator response, compliance, and available clearance determine whether the full contact trace stays within the validated contact limits.
| Failure mode | Evidence or check | Admission or response | Physical limit to verify |
|---|---|---|---|
| Phantom target: image artifact in empty space | Fresh depth evidence and target-region occupancy, with coverage and occlusion checked | Reject the proposal when the required object evidence is absent; request another view if the region is unknown | Rejection precedes motion only if the proposal has not already been admitted. Missing returns alone do not prove empty space. |
| Stale target: workpiece moved after capture | Converted capture time, bounded clock error, local tracking, and target-error envelope | Revoke dependent tracking when the evidence limit fails; select the validated fallback | Total detection, communication, actuation, and braking delay must fit the measured clearance. |
| Semantic ambiguity: two candidate objects in the pick window | Independent count of candidates in the pick window from fresh depth evidence; model entropy as an earlier producer-side signal | Refuse admission and route the item to the coworker for a manual pick | A single wrong object passes the count, and a confident model can report low entropy; independent physical checks remain necessary. |
| Misclassified affordance: bolted fixture selected as graspable | Contact-force and motion residuals after contact; precontact task/geometry checks where possible | Limit precontact speed and force, then invoke the validated contact response on a threshold crossing | A reactive trip does not cap the first impact peak; test the full force trace, response delay, and stopping energy. |
| Dynamic unreachability: target cannot be reached in its valid window | Speed-ceiling transit lower bound followed by a complete path and stopping check | Reject a proposal whose lower bound exceeds the remaining evidence window | Passing the lower bound proves neither collision-free motion nor a feasible stop. |
Every row of the table therefore requires the proposal to carry its spatial bounds, the provenance of its evidence, and its expiry, in a form the permission path can check without the model. A rejected proposal returns a reason to the learned layer; an admitted proposal remains subject to independent permission checks as the plant moves. The proposal carries those fields across the proposal boundary, and the record the permission path grants from it is the intent lease.
The Intent Lease
The intent lease takes its lineage from memory. Its parent is the belief record of The Spatial State Schema, named by digest, and that record’s evidence epoch, converted to the MCU’s clock, becomes the lease’s evidence time. The lease itself carries a goal region, an error envelope and task tolerance, a requested wrench, a requested terminal mode, an evidence horizon, and an expiry. It is a lease in the sense of Multi-Rate Cadences, here on a target, and the proposal that requests it travels under the proposal header with an intent payload in place of a chunk. An admitted intent grants a target and effort ceilings, never a path.
↳ Byte layout: The intent payload’s fields, types, and offsets are given in Intent payload.
On the running machine the proposal travels from the application processor to the MCU as a fixed-size, statically parsed message in a bounded shared mailbox,12 so its parse time can be measured against the control deadline. The message only transports a request; the independent permission controller owns the configured force ceilings, deadline ceilings, and state-dependent fallback.
The covariance ellipsoid of the mug-handle estimate in section 1.3 is a producer-side localization estimate. The proposal does not try to encode a general oriented \(3\times3\) covariance; its bounds carry three conservative position-error envelopes \((e_x,e_y,e_z)\) and three requested task tolerances \((r_x,r_y,r_z)\) in the named robot-base frame instead. Envelope and tolerance keep the distinction drawn in section 1.2, so admission requires \(0\le e_i\le r_i\) on each axis. The producer must justify any conversion from covariance to an envelope, including calibration bias and declared tail risk. Orientation tolerances and contact-mode rules come from an independently approved task profile; the proposal carries only the target orientation itself.
The proposal’s evidence time is the belief record’s evidence epoch, never its estimate epoch, converted to the MCU’s clock under the clock rule of Multi-Rate Cadences. The MCU matches both the parent digest and that epoch to a retained belief record and adds the mapping’s error bound to its age test. The requested expiry uses the same clock, so the MCU can compare it directly with the evidence and stopping deadlines that bound it. The parent digest is lineage, not message authentication, under the integrity policy of Multi-Rate Cadences. It does not prove who chose the target, limits, or requested terminal mode, and the intent proposal carries no keyed tag because the permission path admits it by content, not origin.
The admitted intent moves through Pending, Active, Renewed, Expired, and Aborted states. The MCU enters Active only after checking the request against independently configured ceilings and a feasible response from the measured state. A renewal needs a newer sequence, a fresh belief record, and a new permission check; it does not automatically extend authority. Meeting the approved task tolerance completes the goal. Unexpected contact or loss of a required belief aborts the dependent motion.
Definition 1.2: Expiring intent lease
Expiring intent lease (\(\mathcal{L}_{\text{intent}}\)) is a time-bounded execution authorization issued by the independent permission path that binds an admitted target pose and effort ceilings to a monotonic hardware deadline \(t_{\text{exp}}\): \[t_{\text{exp}} = \min(t_{\text{evidence}} + \tau_{\text{ev}}, \; t_{\text{issue}} + \tau_{\text{lease}}, \; t_{\text{stop\_deadline}})\] upon whose expiration motion tracking authority is automatically revoked without requiring an explicit cancellation packet.
- Significance: Withdraws tracking authority on a hardware deadline when upstream deliberation stalls, so the validated fallback starts without waiting for the proposer.
- Distinction: Unlike a network keepalive or software ping, an intent lease couples the horizon of physical sensor evidence (\(\tau_{\text{ev}}\)) with dynamic braking feasibility, and its expiry falls no later than the stop deadline at which a validated fallback still fits the remaining clearance.
- Common pitfall: Sizing lease durations generously to mask tail latency (\(P_{99}\)) in neural inference, which creates an unmonitored open-loop window where the robot moves through an environment that may have changed.
At expiration, the timer revokes tracking authority and the MCU starts whichever of its approved hold, controlled-deceleration, and torque-disable responses the rule of section 1.4 selects for the measured load and contact mode. The requested terminal mode is advisory, so a request for torque disable over a suspended payload cannot override the powered hold that such a load requires. Neither a 1 kHz check period nor a hardware interrupt establishes a physical stop within one period.
The proposal’s fields support layered checks, cheapest first, and one implementation orders them in this bounded admission sequence:
- Evidence and sequence: Read a complete proposal packet. Reject an old sequence or a parent digest and evidence epoch that do not match a retained belief record. Compute evidence age in the MCU monotonic domain and add the bounded clock-conversion error. Reject when that age exceeds the target-evidence horizon.
- Proposal limits: Check finite pose values, quaternion normalization, base-frame error and task-tolerance axes, and requested wrench against approved task and actuator limits. Check known occupancy and unknown-space policy.
- Dynamic admission: Use current velocity, clearance, and actuator state to reject proposals that already fail a lower bound on transit time. Require that the total sensing-to-brake-onset travel plus braking distance fit inside the current clearance before granting motion. If no feasible fallback exists from this state, refuse the proposal. Refuse also when fresh evidence shows more than one candidate in the pick window, and route the item to the coworker for a manual pick. This stage admits a target and effort ceilings, never a path; admission of the whole trajectory belongs to the planner’s record (Behavior at the Seam).
- Authority latch: Store only the MCU-approved target, effort ceilings, expiry no later than the evidence deadline, and selected fallback. Arm the local timer for the admitted interval. The untrusted requested expiry and terminal mode do not control the drive directly.
- Runtime check: On each configured control tick and relevant hardware fault event, compare measured state with the admitted limits. On evidence expiry, missed renewal, or an observed violation, revoke the proposal and start the previously validated response. The ensuing braking and hold have their own measured delay and physical margin.
Filled in for the red mug, the lease admitted for the pick carries six entries. Its parent is the mug-handle belief record of The Spatial State Schema, matched by digest and evidence epoch. Its goal is the handle pose, and its error envelope is that record’s 3 mm initial error, inside a 15 mm task tolerance. Its evidence horizon is the age at which the governing growth law reaches that tolerance, 60 ms under the record’s belt-drift law and 219 ms under the residual-slip law of the qualified tracker’s renewed record. Its speed ceiling is the arm’s 1 m/s tool-center-point speed limit, and it requests no item-specific wrench ceiling, because none is registered for the mug. The configured task and actuator ceilings still bound the wrench, and fragility is handled by the site’s control of which items reach the conveyor, not by the controller. Its requested terminal mode is grasp, then hold for the handover, although the MCU still selects the fallback from its approved responses. Its expiry is the earliest of that evidence horizon, the configured lease ceiling, and the stop deadline.
This record gives Adversarial Verification a concrete interface for fault-injection tests in which corrupted quaternions, stale capture records, oversize force requests, and incompatible terminal requests should each be rejected under the tested state and timing conditions.
Checkpoint 1.1: Expiring intent and fallback authority
Verify your understanding of how a proposal is refused or admitted, and how the lease granted from it is sized and revoked:
Fallacies and Pitfalls
Each mistake in this section trusts an upstream message to mean more than it does. Local checks must refuse a target, or withdraw tracking authority over one, when its evidence expires or its physical justification dissolves, and they must decide on evidence the proposer does not control, even while its messages remain well formed and confident.
Fallacy: A renewal computed from the same stale frame is a valid renewal.
The machine’s intent model re-reads its last wrist-camera frame of the red mug and sends a renewal with a newer sequence number, a later issue time, and the same pose. Every field a keepalive would check has advanced, yet no image has been captured since the evidence that justified the target, and the mug has kept moving at 0.20 m/s. Because the lease dates its evidence by the belief record’s evidence epoch, not by the renewal’s issue time, the MCU matches the renewal’s parent digest and epoch to the retained record, finds the epoch unchanged, and keeps counting age from the original capture. The target still expires 60 ms after that capture. A controller that dated evidence by message arrival would let a stream of well-formed renewals hold an obsolete target indefinitely, turning the drift hazard of section 1.4 from a single late inference into a standing condition.
Pitfall: Allowing an unexpired temporal lease to continue commanding motion when target spatial bounds are violated.
The machine’s arm lowers its gripper toward the red mug under an intent with the 15 mm grasp tolerance. A slip on the takeaway conveyor carries the mug beyond that tolerance while time still remains on the lease. The timestamp still permits execution, but the target has left the \(SE(3)\) region that justified the approach. Continuing until expiry would preserve an obsolete, potentially colliding command. The local perception and permission loops must compare the observed target with the intent’s spatial bounds and withdraw permission when those bounds fail. Encoding that geometric condition in the lease allows the controller to interrupt the descent without waiting for a new high-level plan.
Fallacy: A confident grounding needs no refusal path.
A confidence score cannot certify that the scene holds only one candidate. A design that sends only low-confidence proposals to the ambiguity test treats the model’s score as evidence about the scene, when it is the model’s report about itself. The glass jar beside the red mug in section 1.3 arrives with a high score and a tight geometric envelope, so a score-gated design skips the test exactly where the test was needed. The refusal path must run on every proposal and must draw on evidence the model does not produce. A check that the pick window holds exactly one object refuses the jar-beside-mug scene whatever the score. What such a check cannot see, a lone jar arriving where the mug should be, remains a premise the permission path names rather than settles, and the site’s feed procedure, which admits to the conveyor only items on the station’s declared list, is what contains it.
Fallacy: Neural networks can resolve bimodal spatial ambiguity by averaging candidate target locations.
The machine sees two identical red mugs set a few centimeters apart on the takeaway conveyor and receives their spatial mean as a grasp target. The midpoint lies in empty space, so averaging has destroyed the \(SE(3)\) distinction the task requires. Both mugs are real, so the innovation check of Belief Through Occlusion rejects neither measurement before fusion. Each hypothesis alone is sharp, and a gate applied to either one passes. The window check of section 1.6 still catches the case, because it counts two candidates in the pick window and the admission gate refuses the proposal and maintains a hold, whatever the fused estimate reports. A covariance gate would catch it only for an estimator that reports the bimodal ambiguity, which an averaging estimator does not. The intent interface must preserve statistical uncertainty rather than collapse competing object hypotheses into an apparently precise Cartesian coordinate, because only a representation that keeps both hypotheses lets a check other than the window count see the ambiguity.
Summary
Grounded intent meets a request whose meaning the permission path cannot check, and it keeps the third law, that proposal is not permission, because the path grants only what it can check. A learned model proposes a dated target and an uncertainty description. The record the permission path grants from that proposal is the intent lease, which takes its evidence epoch and parent from the belief record of Spatial Memory and grants a target and effort ceilings with an expiry no later than the evidence horizon that target drift and task tolerance set. An independent countdown then withdraws authority that a stalled model could never cancel. The independent permission path checks the target’s evidence age, configured motion bounds, current state, path clearance, and feasible fallback before applying motion. The target-evidence horizon \(\tau_{\text{ev}}\le(r_{\text{tol}}-e_0)/v_{\text{drift}}\) limits how long the target estimate may be used under its stated drift bound; it grants no time to stop, which a separate total-delay and braking check must find within the available physical clearance.
The red mug on the conveyor makes the cadence constraint concrete. Its untracked evidence lasts 60 ms, and the machine’s 5 Hz intent model, refreshing every 200 ms, cannot maintain that target by itself. A qualified local tracker can refresh it from new evidence and lengthen the horizon to 219 ms, leaving the intent model only to re-confirm which mug is meant; otherwise the machine slows, pauses, or refuses the task. Expiry revokes the proposal’s tracking authority and invokes a response validated for the measured state. It cannot undo motion already executed.
What grounded intent adds to the law is evidence age on a moving target, and refusal as the permission path’s answer to what it can detect but not resolve. A target the arm cannot reach inside its evidence window fails the transit lower bound before any path is solved, a pick window holding two candidates fails the ambiguity test, and each rejection returns a reason the proposer can act on. The wrong-object case marks the limit of that machinery. A confident grounding to the wrong entity passes every geometric and timing check, so the premise that the grasped object is the named one stays open rather than being settled by the controller. After contact, a force trip bounds later loading only through the measured response, and precontact limits are what bound the first impact.
Key Takeaways: Grounded intent at the proposal boundary
- A lease grants a target, never a path: The Brain supplies a target, evidence epoch, error envelope, task tolerances, and requested expiry, and it writes no actuator setpoints. The permission path admits a target and effort ceilings that expire at the earliest of the evidence horizon, the lease ceiling, and the stop deadline.
- Evidence expires with drift: Target drift and task tolerance set a finite validity horizon under a declared validity region. A fresh timestamp does not make a target correct, and only fresh evidence, never a faster model, renews it.
- The evidence deadline is not the motion deadline: The permission path separately budgets sensing, communication, enforcement, actuator delay, and braking distance for the current state and clearance.
- Geometry cannot verify meaning: A confident grounding to the wrong object can pass every coordinate, reachability, and force check. That premise needs its own task evidence, because the permission path cannot settle it.
- Refusal is an admissible outcome: Ambiguous identity, stale evidence, insufficient reachability, or infeasible stopping can each reject a proposal, and a structured rejection tells the proposer what to change. Each check has its own evidence limit.
What’s Next: From intent leases to trajectory planning
Footnotes
Symbol Grounding Problem: Stevan Harnad’s 1990 formulation argues that syntactic manipulation alone does not give symbols intrinsic meaning. In physical AI, grounding binds a proposed label to dated sensor evidence, calibrated metric frames, and task-specific physical checks. Without that binding, a planner may optimize toward the wrong physical target.↩︎
Classical STRIPS planning: Classical planning represents goals as logical predicates in a modeled world. Dynamic obstacles and sensor dropouts can invalidate a physical target before its planned use.↩︎
Affordance space parameterization: In robot learning, affordances are parameterized as continuous spatial probability fields or reachability manifolds over the Lie group \(SE(3)\). Parameterizing reachability as an explicit manifold enables low-level controllers to evaluate physical feasibility before invoking computationally intensive trajectory optimizers.↩︎
Asynchronous Safety Event Dispatch: Relying on asynchronous cancellation messages or network event dispatch to abort motion introduces non-deterministic queue serialization delays, and every millisecond of that delay is unbraked travel at the base’s drive limit. A local countdown can revoke tracking authority on missed renewal.↩︎
Hardware watchdog safety monitors: A safety design can use an independently clocked watchdog to detect missed renewals and initiate a validated response without the host. Safe Torque Off removes drive torque; it does not by itself hold a suspended load or stop moving mass within a specified clearance.↩︎
Mechanical impedance regulation: Impedance control regulates the relation between contact force and motion through effective mass, damping, and stiffness. It can reduce loads within a validated contact regime; the first impact and peak force still depend on approach speed, passive compliance, sensing, and actuator response.↩︎
Subsumption architecture principles: Rodney Brooks’s subsumption architecture decomposed autonomous robotics into asynchronous, priority-arbitrated behavioral layers operating on sensory feedback. Here, the learned layer proposes a bounded target while an independent controller checks permission and executes a validated response on expiry.↩︎
Semantic-to-metric transformation: Projecting discrete linguistic embeddings into continuous \(SE(3)\) coordinates requires an unbroken coordinate calibration chain from image pixels to actuator base frames. Unmodeled camera mounting compliance or lens distortion maps semantic attention peaks into systematic metric offsets. If the transformation chain lacks rigid extrinsic calibration, high semantic confidence produces large physical tracking errors at the end effector.↩︎
Learned-output sensitivity: A learned pose estimate can change substantially when the image or scene changes, and an empirical confidence score does not bound that change. The permission path limits the admitted target and motion using independently checked geometry, state, and actuator constraints.↩︎
Hardware watchdog down-counters: Detecting a missed renewal is only the first term of the response budget. Drive command propagation, brake engagement, and physical deceleration each add a separately measured term.↩︎
Kinematic workspace boundaries and singularities: The kinematic workspace defines the continuous manifold of task-space poses attainable by the manipulator given link geometry and joint travel limits. Near workspace boundaries and internal singular configurations, the kinematic Jacobian \(\mathbf{J}(\mathbf{q})\) loses rank, requiring infinite joint velocities to maintain finite Cartesian end-effector motion. Enforcing workspace manifold filters at the proposal boundary prevents trajectory optimizers from driving actuators into kinematic lockup.↩︎
Inter-core communication across the proposal boundary: RPMSG over a statically bounded shared SRAM ring is one transport. Copy count, cache maintenance, coherency, and worst-case handoff time depend on the SoC and driver, and the MCU’s independent timer detects lost renewals even through a host kernel panic. See Heterogeneous SoC Mailboxes and Memory Barriers.↩︎


