Spatial Memory

Spatial Memory

Isometric blueprint showing persistent spatial memory: tiered octree 3D voxel grid with fresh observation updates, frosted glass occlusion divider, floating spatial belief decay uncertainty halos, cylindrical temporal lease ledger, and crimson stale token eviction basin.

Purpose

How fast does a spatial belief turn into dangerous fiction once a sensor loses sight of an object?

A dynamic obstacle remains a critical physical constraint long after it moves behind an occluding barrier. A machine must remember the obstacle’s presence, yet its recorded state rapidly degrades in reliability as time elapses. Treating a stale observation as ground truth risks planning trajectories through occupied physical space, while instantly discarding the record erases genuine hazards. Spatial memory must therefore track both a state estimate and an explicit model of growing uncertainty. Sensor capture age, odometry drift, and coordinate frame calibration dictate whether a remembered entity remains valid to authorize motion.

Software failures such as frozen frame transforms or silent tracking loss can deceptively present stale estimates as fresh data. A robust memory system must enforce bounded belief expiration, withdrawing proposal authority before uncertainty expands beyond available clearance envelopes without mistakenly assuming unobserved regions are free space. In the physical AI architecture, memory populates the Brain’s world model above the proposal boundary, while the deterministic permission path below verifies that every trajectory relies strictly on unexpired, certified spatial bounds.

Learning Objectives
  • Specify belief records that keep evidence age separate from estimate age, with frame, uncertainty, and validity region
  • Set belief expiry from bounded uncertainty growth and the task’s physical tolerance
  • Compare spatial representations by what each can certify about clearance
  • Design spatial updates that preserve evidence age and distinguish missing observations from verified free space
  • Diagnose frozen transforms, phantom tracks, and unobservable drift that pass freshness and innovation checks

Preserving Spatial Belief

A memory register can preserve a record exactly while the world it describes moves. As the warehouse mobile manipulator closes on the end of a storage rack, the rack hides the cross aisle from the navigation camera, and whatever the camera last saw there now survives only in memory. A person out of sight past the rack end can walk into the space that record calls empty, and republishing the record with a new software timestamp cannot restore the lost physical evidence. Spatial memory must therefore carry the last measurement time and an error-growth rule into the motion-permission decision.

↰ Prerequisite: Observation age and calibrated spatial covariances are generated in The Observation Contract.

A physical machine cannot rely solely on raw sensory observations or static memory registers. It requires a belief, an estimate of physical state carried forward from historical sensory evidence, coupled explicitly to an expanding model of uncertainty.1 Perception supplies a dated, calibrated observation (as established in What Perception Hands Over). Spatial memory must determine how long that observation, combined with dynamic motion models, can continue to authorize physical action before uncertainty consumes the machine’s remaining clearance.

As the second of the four proposal-side stages of Part III, between perception and intent, spatial memory carries geometric and physical state estimates forward across sensor dropouts, occlusions, and variable processing delays. It sits above the proposal boundary (The Machine in Five Levels), so a remembered state can shape a proposal but cannot certify its own freshness; the permission path checks the age of the evidence against its own clock.

In software systems, data age is frequently measured from the instant a message was serialized, published over an inter-process bus, or updated in a shared memory table. In a physical AI system, this convention is dangerous. Propagating a state estimate through an extended Kalman filter at \(1\text{ kHz}\) or republishing an \(SE(3)\) kinematic transform at \(200\text{ Hz}\) creates a stream of fresh software timestamps, but it does not refresh the underlying physical information. The true age of a belief is determined strictly by the evidence epoch \(t_0\) of the last physical measurement that constrained it. A downstream tracking loop that reads a position estimate generated by dead reckoning or policy rollouts may observe a publish timestamp only \(2\text{ ms}\) old, yet the physical error bound continues to widen at the velocity of the unobserved drift. Timestamping an inference output or a cache lookup records when compute occurred, but it tells the receiver nothing about when reality was last checked.

Memory integrity alone cannot establish geometric validity (Durrant-Whyte and Bailey 2006). To prevent a machine from acting on stale data, memory systems must treat expiry as an explicit interface event. Each belief carries a deadline, its invalidation horizon (derived in section 1.6), at which its declared error bound reaches the task’s physical margin. At that deadline the belief stops supporting motion, and the permission path refuses any proposal that still depends on it and selects a feasible fallback.

Durrant-Whyte, Hugh, and Tim Bailey. 2006. “Simultaneous Localization and Mapping: Part I.” IEEE Robotics & Automation Magazine 13 (2): 99–110. https://doi.org/10.1109/MRA.2006.1638022.

A belief that must declare its own expiry has to travel with more than its value, so the first handoff question is what a consumer needs to evaluate a stored value after its sensor loses sight of the object.

What Memory Hands Over

An unannotated vector such as \([0.142, -0.051, 0.812]\) arriving over a real-time bus conveys nothing about what physical entity is being described, what physical dimension is represented, or how that quantity is scaled. A belief handed to a consumer therefore cannot be a naked array of floating-point numbers. It is a decaying interpretation of physical inputs and must carry its epoch, physical uncertainty, and reference frame to remain usable. A downstream controller receiving raw coordinates without metadata must rely on implicit conventions established at compile time. If an estimator produces positions in meters while an actuator controller expects millimeters, or if an angular rate is mistaken for an absolute orientation, the resulting command registers as a thousandfold tracking error, driving motor current to the actuator thermal limits detailed in Thermal Duty Cycles. To prevent misinterpretation across the causal boundary established in The Causal Boundary, a belief record must first specify the entity identity, the physical quantity, and the measurement units. The consumer must be able to verify that the value represents the center of mass of a specific workpiece in meters, or the normal contact force on a fingertip in newtons, before routing that number into a control law.

Specifying the quantity and units remains incomplete without binding the value to a concrete reference frame. Identical spatial coordinates describe entirely different physical realities depending on the origin and orientation of the coordinate system. A Cartesian position \([0.142, -0.051, 0.812]\text{ m}\) designates a point suspended in free space when evaluated in the robot base frame, but represents a point several hundred millimeters inside a rigid fixture when interpreted in the tool flange frame. Because kinematic links flex, joints move, and coordinate transforms update at varying rates, a belief record must carry an explicit frame identifier. Carrying the frame allows the receiving node to look up the corresponding spatial transformation tree and project the stored geometry into the local frame required by the actuator, preventing the controller from acting on geometric misalignments caused by moving kinematic chains.

Geometry and identity describe the physical state at a single moment, but memory must also capture the temporal anchor of that state. The record carries the evidence epoch \(t_0\) of the physical interaction or optical exposure that last constrained the estimate. In distributed architectures, intermediate nodes often republish data, execute Kalman filter prediction steps, or run forward kinematic propagations to maintain high-rate control streams, and none of these adds physical evidence. The field must therefore pass unchanged through every republish, filter, and prediction stage, so that consumers measure the age of the information against the physical world rather than against the scheduling frequency of the local computer.

Generating this anchored belief relies on multi-sensor state estimation—whether implemented through recursive Extended Kalman Filters (EKF), Unscented Kalman Filters (UKF) (Kalman 1960), or nonlinear factor-graph optimization (Smoothing and Mapping, SAM)—an established mathematical discipline whose operational role is to map asynchronous, heterogeneous sensor observations (such as high-rate IMU accelerations, wheel encoder ticks, and low-rate vision claims) into a minimal sufficient statistic of the plant’s latent physical state. The physical AI architecture abstracts the internal numerical mechanics of this fusion pipeline behind the spatial belief contract. Downstream trajectory planning and safety enforcement do not recompute sensor innovation covariances or re-solve factor graphs; they require only that the state estimator provide an unbiased state vector \(\hat{\mathbf{x}}\), a physically honest uncertainty bound \(\mathbf{\Sigma}\), and the true evidence epoch \(t_0\) of the measurements that constrained it. The foundational systems requirement is not the choice of filtering algorithm, but the obligation that the estimator model process noise expansion honestly across blind intervals (Continuous Covariance Propagation and Observability derives continuous error covariance propagation) and revoke its validity status whenever physical observability rank collapses (section 1.7).

Knowing the evidence epoch \(t_0\) establishes elapsed time \(\Delta t = t - t_0\), but elapsed time alone cannot determine whether a belief is valid for a given action. Two estimates of identical age diverge at different rates depending on physical constraints. The rack end itself is bolted to the floor, so its remembered position stays within the machine’s localization error however long the camera is blind, whereas a person hidden behind it can move at walking speed and spend the same clearance in a few tens of milliseconds.

The minimum record fields necessary for a consumer to evaluate a stored state can be derived directly from the requirement to bound current error. Let \(E(\Delta t)\) represent the upper bound of the discrepancy between the stored value and true physical state after an elapsed time \(\Delta t = t - t_0\). If an estimator establishes a state at time \(t_0\) with an initial uncertainty bound \(E_0\), and external disturbances or unmodeled velocities cause the true state to diverge at a maximum rate \(r\), the upper bound on error at current time \(t\) obeys the linear expansion \(E(\Delta t) = E_0 + r\,\Delta t\). At the rack end, \(E_0\) is the machine’s localization error and \(r\) is the walking speed of the person the rack may hide; section 1.6 evaluates that law. Omitting either term leaves the consumer unable to compute \(E(\Delta t)\), so an initial bound and an expansion rate are both structurally required fields of the record.

Expressing uncertainty through physical bounds and expansion rates distinguishes a physical AI belief from a probabilistic confidence score. An object detection network might output a class confidence of \(0.94\), but that softmax score is not calibrated evidence (Cognitive Vulnerabilities), and a downstream position controller cannot convert a unitless probability into a safe actuator displacement margin (i.e., how far the arm can physically move before risking a collision). A confidence score of \(0.94\) provides no information about whether an obstacle boundary is uncertain by \(2.0\text{ mm}\) or by \(200\text{ mm}\), nor does it indicate how rapidly that boundary moves when occlusion occurs. A physical memory record must express its initial bound \(E_0\) and expansion rate \(r\) in concrete physical dimensions such as meters, radians, or newtons per second. Furthermore, these rates are valid only within a declared validity region, such as a maximum assumed acceleration or a verified contact friction coefficient. If physical operating conditions leave that region, the expansion rate no longer holds, and the belief record must be treated as invalid.

When memory hands over a tuple containing value, reference frame, evidence epoch \(t_0\), initial bound \(E_0\), and growth rate \(r\), the consumer executes a deterministic verification step before applying the data to physical control. The consumer possesses a task tolerance \(E_{\max}\) dictated by the physical task. When the arm guides a gripper finger into the latch-release slot of the cage door, for instance, the slot’s clearance sets \(E_{\max}\), and positional uncertainty must stay below it. Upon reading the belief record at current time \(t_{\text{now}}\), the consumer calculates the elapsed time \(\Delta t = t_{\text{now}} - t_0\), evaluates the present error bound \(E(\Delta t) = E_0 + r \Delta t\), and compares the result directly against \(E_{\max}\). If \(E(\Delta t) \le E_{\max}\), the belief passes this expiry check and can enter the independent motion-permission check. Otherwise the dependent proposal is refused; the controller selects a validated hold, search, or stop response only when it remains feasible.

Definition 1.1: Spatial belief record

Spatial belief record is a typed record that binds an identified physical quantity and SI unit to a frame, evidence and estimate epochs, clock conversion, initial error, growth law, validity region, expiry, provenance, and lifecycle status. These fields let a consumer evaluate freshness before a separate motion-permission check.

  1. Significance: A state without an error bound and growth rate cannot be safely evaluated by downstream controllers. As time elapses (\(\Delta t = t_{\text{now}} - t_0\)), unobserved motion or sensor drift widens the spatial error bound (\(E(\Delta t) = E_0 + r \Delta t\)). When that bound exceeds the task tolerance (\(E(\Delta t) > E_{\max}\)), the record expires, requiring the permission path to reject dependent proposals and choose a feasible validated response.
  2. Distinction: Unlike a unitless classification probability (such as \(P(\text{object}) = 0.94\)), a spatial belief record provides physically grounded metric units (meters, radians) and explicit time-decay mechanics.
  3. Common pitfall: Relying on external publishers to broadcast invalidation signals upon occlusion. If the communication channel drops or the sensor driver hangs, the downstream consumer continues executing open-loop motions on frozen beliefs unless it independently evaluates the record’s expiry.

Evaluating the spatial belief record depends on calculating the elapsed time \(\Delta t = t_{\text{now}} - t_0\) against physical reality. If the consumer cannot verify the exact physical instant when energy transduction occurred, the calculated error bound \(E(\Delta t)\) becomes an unanchored fiction, mistaking software compute cycles for physical evidence. The estimate epoch and clock conversion named in the definition exist to protect that anchor, while provenance and lifecycle status, the record’s remaining fields, are developed with the schema in section 1.9.

A State Has an Epoch

A pose that leaves an object detector describes where a workpiece was when the camera exposed the frame, not where it is when the message arrives. Every state therefore refers to a specific instant, its epoch, and a state crossing a physical pipeline carries three of them before a controller can act. The acquisition epoch (\(t_0\)) is the capture epoch of the observation contract (The Observation Contract), the instant of energy transduction at the causal boundary, such as photons striking a photodiode array or a digital register latching an encoder count, and the belief record carries it as its evidence epoch.2 Memory adds the two epochs that follow it. The estimate epoch (\(t_{\text{est}}\)) marks the instant a state estimator or neural network completed its inference over those measurements and resolved them into a concrete state claim. The publication epoch (\(t_{\text{pub}}\)) marks the instant the serialized message was committed to the transport bus for dispatch to downstream consumers.

Collapsing these three timestamps into a single receipt or transmission timestamp destroys the physical validity of the state. Tagging a pose with its publication time hides every millisecond of exposure, readout, inference, and scheduling that preceded it, while gravity, momentum, and contact forces keep acting on the physical system without pause. That hidden age is displacement at speed, as Photons to Spatial Claims established; for the navigation camera at the base’s drive limit, the age at dispatch alone amounts to several centimeters of travel, all of it invisible to a consumer that reads only the publication stamp. A state that records only when data became available tells the consumer nothing about the physical instant the measurement represents, making temporal alignment impossible.

↰ Prerequisite: Synchronized hardware timestamping across fieldbuses is established in Moving Commands on Time.

Even when a system records true acquisition epochs, physical AI architectures rely on distributed processing nodes whose local oscillator clocks drift independently. Synchronizing these distributed clocks over standard bus networks introduces a bounded synchronization error, and that temporal uncertainty becomes spatial uncertainty at the relative speed of whatever the record describes.3 The observation record of The Observation Contract already carries this bound as its clock-conversion error, and a belief built from that observation keeps it. Every consumer adds the bound to the evidence age before judging freshness, and the error budget allocates the distance that bound covers at the record’s relative speed alongside sensor resolution and structural deflection.

Real physical systems never sample all relevant state variables simultaneously. In an illustrative cycle, a tactile array might latch normal force readings at \(t_1 =\) 12 ms, an optical tracking camera might finish sensor integration at \(t_2 =\) 18.5 ms, and joint optical encoders might sample angular positions at \(t_3 =\) 21 ms. These observations occur at different epochs and cannot be combined through direct algebraic operations. Fusing a camera observation at \(t_2\) with an encoder observation at \(t_3\) requires the consumer to project one state across the temporal gap \(\Delta t = |t_3 - t_2| =\) 2.5 ms using explicit kinematic models and bounded velocities (for example, propagating the historical camera observation along the joint velocity vector to reconcile it with the instantaneous encoder reading). Without the true acquisition epoch for each individual signal, the alignment calculation cannot determine the sign or magnitude of \(\Delta t\). Extrapolating state without an accurate epoch causes the fusion algorithm to compute spatial offsets from fictitious temporal baselines, compounding errors across every fusion cycle.

With these epochs and the clock-synchronization bound carried in the record of section 1.2, the consumer computes the elapsed duration between the evidence epoch and the execution deadline, evaluates the worst-case expansion of the error bound, and determines whether the state remains within task tolerances. Handing over a state without its acquisition epoch leaves the consumer blind to transport delays and scheduling jitter, turning deterministic control into open-loop guesswork. Because physical geometry also changes as bodies move through space, spatial transforms between reference frames are subject to the same temporal decay, carrying their own distinct evidence epochs.

How a Transform Goes Stale

A frame binds a coordinate, but in a machine the frame itself moves. What Perception Hands Over gave every observation a named frame and related the sensor, body, odometry, and map frames through a transform tree. What the tree does not show is how a transform goes stale. Each edge of the tree in figure 1 stores a rigid-body pose in \(SE(3)\), parameterized by joint angles (Lynch and Park 2017; Craig 2005), and that pose describes the machine only at the epoch when its geometry was measured. When an onboard computer chains transforms measured at different times, through transport delay, asynchronous sampling, or dropped frames, the product is a hybrid pose the machine never occupied (Forward kinematics and singularities derives the kinematic chain).

Lynch, Kevin M, and Frank C Park. 2017. Modern Robotics: Mechanics, Planning, and Control. Cambridge University Press.

A transform also spans several machine words, so a reader racing a writer can see an old rotation paired with a new translation. The seqlock mailbox of Multi-Rate Cadences closes that race without blocking the control thread, because each reader makes one bounded attempt per tick and never waits; a torn read is discarded and the last validated transform ages in its place.

Manipulator Kinematic Frame Tree and Temporal Lever-Arm Dynamics. (a) Physical geometry of an articulated manipulator with base link, joint linkages, wrist camera, and tool center point (tool0), indicating lever arm ℓ and angular deflection error e_tool ≈ ℓ · Δθ. (b) Spatial transformation tree (DAG) connecting world reference, base link, kinematic chain link1 through link6, wrist frame, wrist camera, tool0 TCP, and target object, highlighting an asynchronous stale transform edge. (Adapted from Craig 2005 and Foote 2013).
Figure 1: Manipulator kinematic frame tree and temporal dynamics: Articulated manipulator geometry (a) and corresponding directed acyclic coordinate transform tree in ROS (b). Each directed edge stores a homogeneous transformation matrix parameterized by joint state. Composing transforms across disparate acquisition epochs injects lever-arm positioning error \(e_{\text{tool}} \approx \ell\,\Delta\theta\), while an uncompensated asynchronous edge references stale coordinates (adapted from Craig 2005 and Foote 2013 (Craig 2005; Foote 2013)).
Craig, John J. 2005. Introduction to Robotics: Mechanics and Control. 3rd ed. Pearson Prentice Hall.
Foote, Tully. 2013. “Tf: The Transform Library.” IEEE International Conference on Technologies for Practical Robot Applications (TePRA), 1–6.

The cost of an expired transform is direct spatial displacement (figure 2). Applying the age-as-displacement relation of Photons to Spatial Claims to kinematics, uncompensated latency translates directly into geometric positioning error through the lever arm of the linkages. For the rigorous cross-product derivation of this kinematic error propagation, see Spatial coordinate frames and lever-arm kinematics.

How little staleness it takes to spend a task’s tolerance shows at the latch-release slot, while the wrist camera can still see it. Each wrist-camera frame, one every 16.7 ms, renews the slot’s pose in the camera frame and with it the slot-to-finger edge the controller aligns on, and the finger must stay within the slot tolerance \(E_{\max}\) of 1 mm. Suppose that edge misses a single frame. Starting from an initial error of 0.1 mm, the finger’s lateral offset drifts at most \(e_{\text{trans}} = v_{\text{drift}}\,\Delta t\) under an illustrative 12 mm/s bound on its unobserved lateral velocity, validated for the arm’s operating region, or 0.20 mm. With the fingertip an illustrative \(\ell\) = 0.15 m beyond the wrist and the wrist turning at an illustrative \(\omega\) = 0.30 rad/s as it aligns the finger, the lever arm adds \(e_{\text{rot}} \approx \ell\,\omega\,\Delta t\), or 0.75 mm. The joint encoders measure that turn, but the stale edge still dates the slot to the previous frame while the wrist pose it is composed with is current, so the tree returns a hybrid pose, a finger position the arm never held. Only chaining the slot against the wrist pose of its own capture epoch, or expiring the edge, removes that arc. Left in the chain, the stale edge misplaces the finger by 1.05 mm, 5 percent more than the slot allows, and more of that error comes from the lever arm than from the drift. Because the transform tree evaluates without error, nothing flags the stale pose, and the controller keeps aligning the finger on a slot position that is off by more than the clearance, which on an unfavorable approach puts the finger against the strike plate beside the slot.

Transform staleness at the latch-release slot. A dashed ghost arm shows the wrist and finger as dated by the stale edge at t0, with a dashed circle marking the slot tolerance E_max around the finger. The current arm at t0 + Δt shows the wrist shifted laterally by e_trans = v_drift Δt and the link of length ℓ turned by Δθ = ω Δt, which swings the finger by e_rot ≈ ℓ Δθ. A red vector from the stale finger position to the current finger, labeled e_tool ≈ E_0 + e_trans + ℓ Δθ, ends well outside the E_max circle.
Figure 2: Transform staleness and lever-arm error at the latch-release slot: Finger error when the slot-to-finger edge misses one wrist-camera frame of duration \(\Delta t\). Lateral drift at the velocity bound \(v_{\text{drift}}\) moves the finger by \(e_{\text{trans}} = v_{\text{drift}}\,\Delta t\), while the wrist turning at \(\omega\) swings the fingertip, a lever arm \(\ell\) beyond it, by \(e_{\text{rot}} \approx \ell\,\omega\,\Delta t\). Composing the stale edge with the current wrist pose adds both to the initial error \(E_0\), so the finger error \(e_{\text{tool}} \approx E_0 + e_{\text{trans}} + \ell\,\Delta\theta\) exceeds the slot tolerance \(E_{\max}\).

Systems Perspective 1.1: The leverage of staleness
A transform’s admissible age is set at the point it serves, not at the joint it measures. The tool point drifts at roughly \(v + \ell\omega\), so the longer the lever arm, the sooner an edge expires, and an edge still fresh enough to locate the wrist can be too old to locate the tool center point.

A transform tree dates the machine’s own geometry, and a stale edge misplaces the tool by the motion it missed, amplified by the lever arm. The world around the machine ages in the same way, but it has no joint encoders to report its motion, so memory must also hold the surrounding surfaces and the space between them, and it must record which of that space a sensor has actually seen.

What Scene Memory Can Certify

A robot that has swept its range sensor down an aisle holds three kinds of space afterward: space it saw occupied, space it saw empty, and space it never saw. Only the second can ever vouch that a path is clear, and only while its evidence is fresh. Every scene representation stores some version of these three states, and they differ in what they store and what it costs. An occupancy grid (Elfes 1989), or its sparse octree form (Hornung et al. 2013), keeps an occupancy estimate per cell. A truncated signed distance field (TSDF) (Curless and Levoy 1996) keeps a truncated projective distance to the observed surface along each sensing ray. 3D Gaussian splatting (Kerbl et al. 2023) keeps optimized primitives that render new camera views. A dense grid’s storage grows with the cube of its cells per dimension, while the octree, the surface-band blocks of a TSDF, and a Gaussian map store only allocated structure, whose size depends on the observed geometry and the accuracy target; Log-odds occupancy and signed-distance fusion and 3D Gaussian splatting parameters give the updates behind these costs. For motion permission they share one limit. A projective distance is not the nearest-surface Euclidean distance a clearance decision needs, and a Gaussian density does not identify occupied space, so none of them certifies clearance by itself. Each must first be converted into a validated occupancy layer or Euclidean signed distance field (ESDF), with an explicit policy for unknown space, before the permission path can query it.

Elfes, Alberto. 1989. “Using Occupancy Grids for Mobile Robot Perception and Navigation.” Computer 22 (6): 46–57. https://doi.org/10.1109/2.30720.
Hornung, Armin, Kai M Wurm, Maren Bennewitz, Cyrill Stachniss, and Wolfram Burgard. 2013. “OctoMap: An Efficient Probabilistic 3D Mapping Framework Based on Octrees.” Autonomous Robots 34 (3): 189–206.
Curless, Brian, and Marc Levoy. 1996. “A Volumetric Method for Building Complex Models from Range Images.” Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’96), 303–12.
Kerbl, Bernhard, Georgios Kopanas, Thomas Leimk"uhler, and George Drettakis. 2023. “3D Gaussian Splatting for Real-Time Radiance Field Rendering.” ACM Transactions on Graphics 42 (4): 139:1–14.

The choice of metric representation defines an empirical Pareto trade-off between volumetric memory footprint and spatial collision query latency (figure 3). Continuous trajectory optimization and real-time reflex controllers running in 1 kHz control loops (\(1{,}000\ \mu\text{s}\) periods) require spatial clearance and gradient queries to execute in under \(10\ \mu\text{s}\) across dozens of robot body articulation points. While implicit neural representations such as coordinate MLPs (NeRF, iMAP) achieve minimal memory density (\(\sim 0.08\ \text{MB/m}^3\)), querying signed distance or occupancy requires evaluating deep neural network forward passes, imposing massive computational latency (\(>650\ \mu\text{s}\) per query point) that confines them to deliberative offline planners. Hierarchical octrees (OctoMap (Hornung et al. 2013)) compress unobserved free space down to \(0.18\text{--}1.25\ \text{MB/m}^3\), but pointer indirection during recursive tree traversals incurs CPU cache misses that limit query speeds to \(24\text{--}38\ \mu\text{s}\). Conversely, dynamically hashed truncated signed distance fields (Voxblox TSDF (Oleynikova et al. 2017)) allocate memory only around observed surface bands (\(4.1\ \text{MB/m}^3\)), enabling \(1.8\ \mu\text{s}\) ray queries, while propagating Euclidean signed distance fields (ESDF) increases memory overhead to \(28.5\ \text{MB/m}^3\) to deliver sub-microsecond (\(0.85\ \mu\text{s}\)) distance evaluations with analytic continuous gradients. Meanwhile, 3D Gaussian splatting primitives (Kerbl et al. 2023) occupy a middle regime (\(1.6\text{--}5.2\ \text{MB/m}^3\), \(5.8\text{--}8.2\ \mu\text{s}\)), enabling photographic radiance rendering alongside bounding-volume-hierarchy accelerated collision checks.

Figure 3: 3D scene representation frontier: memory footprint versus spatial query latency: Empirical Pareto trade-off across implicit neural fields, hierarchical octrees, 3D Gaussian splatting, and hashed voxel grids. While implicit coordinate MLPs compress rooms into megabytes of weights, multi-layer forward passes exceed real-time collision query limits (\(>100\ \mu\text{s}\)). Hashed voxel grids (Voxblox TSDF/ESDF (Oleynikova et al. 2017)) consume \(4.1\text{--}28.5\ \text{MB/m}^3\) but provide the sub-microsecond query latencies (\(<2\ \mu\text{s}\)) necessary to service 1 kHz reactive control loops.
Oleynikova, Helen, Zachary Taylor, Marius Fehr, Roland Siegwart, and Juan Nieto. 2017. “Voxblox: Incremental 3D Euclidean Signed Distance Fields for on-Board MAV Planning.” IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1366–73.

What a map can certify comes from how a range return updates it. Traversing the ray from the sensor to the return (Amanatides and Woo 1987), the map asserts negative evidence in every cell the ray crossed and positive evidence in the cell where it ended; Log-odds occupancy and signed-distance fusion gives the log-odds update. Negative evidence lets the map retire a phantom obstacle as soon as the same region is seen empty, instead of waiting for a timeout, and it also marks the limit of what the map can say. It covers only cells a valid ray actually traversed. It says nothing about the volume behind the return, and a new scan down the aisle cannot renew a cell hidden behind the rack, so the occluded cross aisle stays unknown and still needs expiry. Every return also touches every cell along its ray, so sparse allocation saves capacity but not traversal, and that traffic reaches the permission path only as added evidence age (Contention for Shared Resources).

Amanatides, John, and Andrew Woo. 1987. “A Fast Voxel Traversal Algorithm for Ray Tracing.” Eurographics ’87, 3–10.

Causal chain showing raycast through free space to surface hit, evicting ghost obstacles.

Active negative evidence along sightlines clears phantom obstacles without waiting for timeout decay.

Because negative evidence renews only the cells a ray crossed, a map’s publication time dates none of its cells, so the manifest in table 1 carries evidence epochs per cell. Those epochs are not free, because every allocated cell stores its own timestamps beside its occupancy estimate, and they belong in the map’s memory budget for the same reason the value does. A motion consumer checks the evidence age and state of every cell its path uses, including cells marked unknown, and derives conservative clearance from a separately validated layer.

Table 1: Volumetric belief manifest and cell evidence: The map header describes layout and update time, while cell-local acquisition epochs determine whether the particular path region is fresh. The consumer treats unobserved or expired cells as unknown unless a separately validated policy permits traversal.
Map field Unit or type Meaning and consumer check
voxel_grid_origin, voxel_resolution_m m; m Spatial anchor and pitch; match the active frame and calibration
map_update_epoch capture-clock ns Last map publication/update; not freshness of every cell
cell_evidence_epoch capture-clock ns per cell Last physical observation for that cell; never refreshed by prediction or unrelated map updates
free_space_sweep_epoch capture-clock ns per occupancy cell Last valid observed free ray through that cell
clock_conversion, sync_error_ns mapping; ns Convert capture epochs to the consumer’s monotonic clock with bounded error

Semantic memory sits above the metric layer and inherits its limits. Beyond raw metric volumes and radiance fields, foundation-model-driven physical AI requires spatial memory to organize space into semantic hierarchies. The 3D dynamic scene graph (DSG), formulated by Rosinol et al. (2021) in Kimera and extended by Hughes et al. (2022) in Hydra, bridges metric SLAM with symbolic reasoning by structuring spatial memory into a five-layer hierarchical graph (figure 4):

  1. Metric-Semantic Mesh (Layer 1): High-resolution triangular mesh capturing fine surface geometry with vertex semantic labels.
  2. Object and Agent Nodes (Layer 2): Bounded 3D geometric entities with oriented bounding boxes, visual embeddings, and tracked dynamic trajectories.
  3. Places and Freespace Topology (Layer 3): 3D Voronoi topology and navigable freespace clearance spheres (\(p_i\)) for planning collision-free paths.
  4. Rooms and Structural Enclosures (Layer 4): Functional architectural subdivisions partitioned by semantic walls and portals.
  5. Buildings and Global Topology (Layer 5): Top-level semantic graph governing long-horizon mission allocation.
Hierarchical five-layer 3D Dynamic Scene Graph architecture. From bottom to top: Layer 1 Metric Mesh surface geometry, Layer 2 Objects and Agents with support relations, Layer 3 Places with Voronoi freespace clearance spheres, Layer 4 Rooms connected by portals, and Layer 5 Buildings with global topology. Side arrows depict upward Abstraction and downward Grounding across tiers.
Figure 4: Hierarchical 3D dynamic scene graph architecture: Five-layer spatio-temporal memory organization bridging metric SLAM to symbolic foundation reasoning. Layer 1 captures triangular metric mesh geometry; Layer 2 extracts 3D objects and dynamic agents; Layer 3 models topological places via Voronoi freespace clearance spheres (\(p_i\)); Layer 4 structures architectural rooms and portals; and Layer 5 establishes global building mission topology. Vertical inter-layer edges coordinate upward abstraction (compressing metric surfaces into symbolic entities) and downward grounding (translating task tokens into metric coordinates) (adapted from Rosinol et al. (Rosinol et al. 2021) and Hughes et al. (Hughes et al. 2022)).
Rosinol, Antoni, Andrew Violette, Marcus Abate, Nathan Hughes, Yun Chang, and Luca Carlone. 2021. “Kimera: From SLAM to Hierarchical 3D Dynamic Scene Graphs.” The International Journal of Robotics Research 40 (12–14): 1410–44. https://doi.org/10.1177/02783649211056674.
Hughes, Nathan, Yun Chang, and Luca Carlone. 2022. “Hydra: A Real-Time Spatial Perception Engine for 3D Scene Graph Construction and Optimization.” Proceedings of Robotics: Science and Systems (RSS).

The computational advantage of this bidirectional hierarchy lies in how it resolves the memory-compute tension on edge SoCs. Upward abstraction collapses gigabytes of dense triangular surface meshes into kilobyte-scale symbolic subgraphs, allowing foundation deliberators to evaluate long-horizon task plans without exhausting LPDDR5 bandwidth or choking KV caches with millions of raw geometric tokens. Conversely, downward grounding lets high-level spatial targets reach local geometric evidence. A deliberator queries space without searching every cell, but each node is only as fresh as the metric cells beneath it, and grounding a symbolic target still ends in a clearance query against the validated metric layer.

A learned policy’s context is memory of a different kind. An autoregressive policy holds its recent observations in a key-value cache whose growth Supply, Freshness, and the Memory Wall derives, and a streaming window bounds that cache by evicting its oldest tokens (Xiao et al. 2024), whether or not their physical consequences have ended. A person seen stepping behind the rack before the window began is gone from the context but not from the aisle. A physical fact the machine must still honor therefore lives in a belief record with its own evidence epoch, never only in the policy’s context.

Xiao, Guangxuan, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. “Efficient Streaming Language Models with Attention Sinks.” International Conference on Learning Representations (ICLR).

For motion permission, then, a representation without an epoch is an unverified assertion, and an estimate whose error bound exceeds task tolerance must be excluded from motion proposals before the permission path considers actuator commands. The machine cannot halt every operation at the first missed frame, however. Extending action across brief sensing gaps requires estimators that propagate belief forward until new measurements arrive to confirm or correct the state.

Belief Through Occlusion

When direct sensing fails, a physical AI system uses an internal transition model to update its state estimate. A Kalman filter updates state estimates using plant dynamics,4 while a learned policy might predict the next relative pose using an autoregressive network. Either way, prediction extrapolates old evidence forward without renewing it, as section 1.1 established; it preserves the utility of an observation by compensating for expected physical motion, but it does not reset the age of the underlying measurement.

The difference between predicting a state and observing it decides what the machine may do at the rack end. A person can step out there into the illustrative 1.10 m clear distance \(D_{\text{clear}}\) that the stopping budget defends, and once the rack hides the cross aisle, the question is how long the remembered view may vouch that the cross aisle is still empty.

Every physical model simplifies reality, causing error between the projected state and the physical world to expand continuously during the blind interval. In state estimation, this dispersion is governed by occlusion covariance expansion, the continuous, physics-governed widening of uncertainty during sensor line-of-sight dropouts. Mathematically formalized by stochastic models (such as the continuous Lyapunov/Riccati covariance propagation equation, detailed in Continuous Covariance Propagation and Observability), this expansion quantifies the growing doubt about where an object actually is, bounding the unobserved divergence caused by unknown drift velocities and unexpected physical disturbances before the sensor regains visibility. For a deterministic worked model, suppose sensor characterization and disturbance limits establish an initial position-error bound \(E_0\) at the moment of occlusion (\(\Delta t = 0\)). Two bounded terms then degrade the prediction. Motion the model does not know about introduces a drift velocity bounded by \(v_{\text{drift}}\), and forces it does not model create a disturbance acceleration bounded by \(a_{\text{dist}}\). Over an occlusion duration \(\Delta t\), the error bound \(E(\Delta t)\) then follows the quadratic expansion:

\[E(\Delta t) = E_0 + v_{\text{drift}}\Delta t + \frac{1}{2}a_{\text{dist}}(\Delta t)^2 \tag{1}\]

At the rack end the terms are the machine’s own. The initial error is the localization bound \(\delta_{\text{loc}}\), 50 mm, because the machine knows where the rack end lies only to within that bound. The tolerance is the protective clearance \(\delta_{\text{margin}}\), 100 mm, which the budget reserves between the stopped machine and a person. A person hidden behind the rack can close on that clearance at up to the walking-speed bound \(v_h\) of 1.6 m/s that safety-distance calculations assume for an approaching person. Because \(v_h\) bounds speed directly, equation 1 keeps only its linear term, with \(v_{\text{drift}} = v_h\) and \(a_{\text{dist}} = 0\), and the belief that the rack end is clear expires after \((\delta_{\text{margin}} - \delta_{\text{loc}})/v_h\), which comes to 31.25 ms. The navigation camera delivers a frame every 16.7 ms, so the belief lapses in under two frames, and the frame that showed the cross aisle empty is already 51.6 ms old when perception dispatches it (Sensor Transduction and Calibration).

Nothing in the software marks the moment of expiry. The map still shows the cross aisle free, and the estimator still publishes a well-formed record, yet the prediction has stopped functioning as evidence. Memory therefore cannot certify an occluded rack end as clear, nor can the frame that feeds it, and the region stays unknown. The stopping budget protects a static obstacle at \(D_{\text{clear}}\), such as a person who has stepped out and stands in the aisle, or a dropped tote, and memory adds a deadline to that budget but no distance. The running total at the drive limit stays at the 912.9 mm that Sensor Transduction and Calibration left.

A person who keeps walking toward the machine during the whole stop is a different obstacle. A safety-distance calculation counts that approach as \(v_h(\tau_{\text{delay}} + T_{\text{stop}})\), the walking speed applied over the entire time from detection to standstill. At the 1.5 m/s drive limit, where a constant-deceleration stop lasts \(T_{\text{stop}} = v/a_{\text{brake}}\), that term alone comes to 1.41 m, more than the entire 1.10 m clear distance before any other distance in the budget is counted. The site therefore handles the approaching person with a crossing rule that site operations owns, a floor-marked crossing and a rack-end presence sensor. A Residual-Claims Register records the premise the rule rests on. Whatever crossing rule the site adopts, the permission path must refuse any proposal whose clearance relies on remembered emptiness past the belief’s deadline.

Definition 1.2: Belief state invalidation horizon

Belief state invalidation horizon (\(t_{\mathrm{exp}}\)) is the elapsed time since the last physical observation at which the belief’s validated error bound \(E(\Delta t)\) first reaches the task tolerance \(E_{\max}\): \[t_{\mathrm{exp}} = \inf \{ \Delta t \ge 0 : E(\Delta t) \ge E_{\max} \} \tag{2}\] where \(E(\Delta t)\) is a bounded-disturbance growth law such as equation 1 or a calibrated error quantile, and a belief whose evidence epoch is \(t_0\) stops supporting motion at \(t_0 + t_{\mathrm{exp}}\).

  1. Significance: Formalizes the hard mathematical boundary where forward-propagated state estimation ceases to function as valid physical evidence, requiring the permission path to refuse dependent motion proposals before geometric clearance is breached.
  2. Distinction: Unlike open-loop timeout watchdogs that monitor software process heartbeat intervals, the invalidation horizon is derived from continuous physical kinematics, bounded disturbance dynamics, and mechanical clearance constraints.
  3. Common pitfall: Allowing a neural planner or Kalman filter to continuously extrapolate state estimates during extended sensor dropouts without bounding error expansion, mistaking smooth mathematical extrapolation for physical ground truth.

When the sensor pipeline reacquires line of sight, the system cannot simply overwrite the propagated state with the new measurement, nor can it blindly pass the raw measurement into the downstream controller. Sensor readings obtained immediately after occlusion often suffer from transient lighting reflections, partial edge visibility, or correspondence errors. Before blending or resetting the state, the estimator must calculate the innovation, a mathematical reality check defined as the difference between the new measurement and where the internal model predicted the pose would be. If the new measurement falls inside the calculated error bound \(E(\Delta t)\), the measurement confirms the model’s drift assumptions, allowing the estimator to collapse the uncertainty and reset the evidence clock. If the incoming measurement falls outside that bound, either the physical object moved unexpectedly or the sensor returned an outlier. Snapping the coordinate frame directly to an erroneous outlier tells the planner that the robot has jumped, introducing an instantaneous position step into the trajectory. The step demands an acceleration no drive can supply, which pushes the actuators against the current ceiling of Actuator Transmission Limits and toward the thermal limits of Thermal Duty Cycles. The estimator must isolate the discrepancy, evaluate whether the measurement is valid, and reject discontinuous jumps before updating actuation targets.

Model prediction serves as a temporary bridge across missing or delayed observations, and the bridge holds only while the accumulated model error remains smaller than the mechanical tolerance of the task. How soon it fails depends on what drives the error. At the rack end a speed bound alone set the rate. At the cage door, the width of the latch-release slot sets the tolerance \(E_{\max}\) for the gripper finger that must enter it, and unmodeled forces add terms that grow faster than any speed bound. Between sensor updates the estimator propagates the state by integrating commanded forces through its model of the plant, without feedback confirmation.

The stochastic model behind that propagation separates three mechanisms of decay. An unmeasured velocity error at the onset of occlusion grows position variance as \(\sigma_v^2(0)(\Delta t)^2\), and unmodeled stochastic forces such as torque ripple or structural vibration, integrated twice, grow it as \(\frac{1}{3}q_c(\Delta t)^3\); these variance terms are what a calibrated quantile must absorb. A steady unmodeled force bounded by \(|a_{\text{dist}}|\), such as a dragging cable or persistent sliding friction, instead deflects the body by at most \(\frac{1}{2}a_{\text{dist}}(\Delta t)^2\), the deterministic term of equation 1.5

At the cage door, the finger now enters the slot blind. In the final approach the finger hides the slot from the wrist camera, and the finger must stay within 1 mm of the slot’s centerline, the task tolerance \(E_{\max}\) (figure 5). The arm’s validated operating region supplies the bounds already charged at the slot’s missed frame in section 1.4, initial position error \(E_0\) at most 0.1 mm and unobserved velocity error \(v_{\text{drift}}\) at most 12 mm/s, and occlusion adds one more, an illustrative acceleration disturbance \(a_{\text{dist}}\) at most 50 mm/s². The growth law of equation 1 carries no lever-arm arc, because the belief chains the slot against the wrist pose of its own capture epoch, the remedy for the missed frame, so position error grows only through drift and disturbance.

Evaluating this growth bound with the stated parameters yields \(E(\Delta t) = 0.1 + 12\,\Delta t + 25\,(\Delta t)^2\), with \(\Delta t\) in seconds and \(E(\Delta t)\) in millimeters. Setting \(E(\Delta t) = E_{\max}\) gives the quadratic equation \(25\,(\Delta t)^2 + 12\,\Delta t - 0.9 = 0\). Its single positive real root is 66 ms, so a proposal depending on this belief must expire by then. If the slot were machined to a clearance of 0.30 mm, the allowable unobserved duration shrinks to 16.1 ms (\(25\,(\Delta t)^2 + 12\,\Delta t - 0.2 = 0\)). Tightening mechanical tolerance by 3.3× shrinks the allowable unobserved duration by 4.1×, requiring sensor re-acquisition or compliant impedance holding within that shorter window. In contrast, if an estimator naively assumes constant velocity (\(a_{\text{dist}} = 0\)), it predicts a valid horizon of \((E_{\max} - E_0)/v_{\text{drift}}\), or 75 ms. At that elapsed time the modeled bound reaches 1.141 mm, exceeding the clearance by 14.1 percent. A constant-velocity-only admission rule would keep accepting the estimate for 9 ms beyond the bounded-model deadline; whether contact occurs depends on the path and geometry.

Figure 5: Belief uncertainty and invalidation horizon: Spatial error bound \(E(\Delta t)\) of a gripper finger entering the cage door’s latch-release slot, expanding quadratically under unobserved forward propagation until intersecting the slot clearance \(E_{\max} = 1\text{ mm}\) at invalidation horizon \(t_{\text{exp}} = 66\text{ ms}\). A dashed constant-velocity estimate grows more slowly, and the red marker shows that estimate still admitting the belief after the bounded error has passed the clearance. Belief has a shelf life governed by physical disturbance dynamics, and only sensor re-acquisition resets the evidence clock.

An estimator’s internal covariance matrix reflects its stochastic model, not a measured hard maximum. A safety argument can rest on one of two kinds of bound. A bounded-disturbance growth law needs justified limits on initial error, velocity, acceleration, clock conversion, and mode changes; then the 31.25 ms rack-end root and the 66 ms latch-slot root are conditional worst-case deadlines within that region. Alternatively, staged dropouts and external ground truth can calibrate a specified residual quantile (for example \(P_{99}\) or \(P_{99.9}\)) under a named test distribution. That yields a probabilistic admission limit with a one-percent or 0.1-percent tail before other model and coverage errors. The record and consumer must declare which regime applies, monitor its validity region, and refuse dependent motion or choose a validated feasible fallback when evidence exceeds it.

Either regime yields the same kind of deadline, because the invalidation horizon of 1.2 asks only when the declared bound first reaches the task tolerance, whatever model produced the bound. Expiry is therefore neither an arbitrary communication watchdog timer nor a fixed \(100\text{ ms}\) heuristic chosen for scheduling convenience. It is a physical deadline dictated by task geometry, actuator authority, and disturbance dynamics, and past \(t_0 + t_{\mathrm{exp}}\) no downstream loop may assume the estimated pose lies within the required clearance.

The horizon of equation 2 is not one scalar for the whole machine. Payload, speed, and contact state change the disturbance bounds, so the state interface selects the active expiry for the current operating regime, and a transition from free motion into contact, which changes velocity within milliseconds, invalidates a smooth drift model outright. Three beliefs on the same machine expire on different clocks (table 2). The latch-slot finger is bounded by drift and disturbance acting against a millimeter tolerance. The person at the rack end and a coworker beside the arm are both bounded by the same walking speed, yet their beliefs lapse at different times because each starts from a different initial error and closes on a different tolerance.

Table 2: Belief-expiry roots on the machine: Each root solves \(E(t)=E_{\max}\) for that belief’s error model and task limit. The coworker’s initial error is illustrative and its tolerance chosen, while its speed is the same walking-speed bound as the rack end, so both person rows grow linearly. The last column names the response the root would trigger. The record for the mug-handle belief that Grounded Intent consumes is filled in section 1.9.
Belief Error bound \(E(t)\) (mm; \(t\) in s) Task limit (mm) Root Response at expiry
Person past the rack end \(50+1600t\) 100 31.25 ms Treat the cross aisle as unknown; refuse dependent advance
Latch-slot finger \(0.1+12t+25t^2\) 1 66 ms Refuse the insertion; hold compliantly or reacquire the slot
Coworker beside the arm \(15+1600t\) 150 84 ms Revoke dependent arm motion; stop only if braking clearance remains

Every root in this section trusts the growth law it was computed from. A growth law is only as good as the machine’s ability to catch its failures, through a residual at reacquisition or a staged dropout test, and some errors produce no residual at all because no sensor on the machine can see them.

When the Estimator Cannot Know

An illustrative strain gauge on a robot’s end-effector reads a steady 5 N, a reading that fits any true contact force once an unknown bias is allowed. An estimator, or a learned model, that predicts that force with confidence has not shown the force is observable. In state estimation, a state trajectory is observable if and only if the history of available sensor outputs over a finite time window uniquely determines the initial state of the machine, so that distinct physical trajectories produce distinct sensor trajectories rather than indistinguishable measurement sequences. When two distinct physical state histories produce the exact same sequence of sensor readings, no mathematical operation, filter, or neural network running anywhere in the machine can determine which physical trajectory the body actually followed. The machine receives identical bitstreams from its sensor interfaces across both realities.

To see why predictability differs from observability, follow that end-effector as it regulates contact force against a rigid surface. Let the true normal contact force be \(x(t)\), and let the strain-gauge conditioning electronics introduce an unknown, slowly changing sensor bias \(b(t)\). The raw measurement delivered to the estimator at each sample interval is \(y(t) = x(t) + b(t)\). Suppose the sensor outputs a constant stream \(y(t) =\) 5 N. This observation is satisfied by a true physical force \(x(t) =\) 5 N with zero offset \(b(t) =\) 0 N. It is equally satisfied by a true force \(x(t) =\) 2 N with a sensor bias \(b(t) =\) 3 N, or by \(x(t) =\) 0 N with a bias \(b(t) =\) 5 N. For any arbitrary offset trajectory \(\Delta(t)\), the state pair \((x(t) - \Delta(t), b(t) + \Delta(t))\) produces the exact same measurement sequence \(y(t)\). The difference between the true force and the sensor offset lies in an unobservable subspace of the measurement system, a structural blind spot where physical variations perfectly cancel each other out in the sensor data.

Because the measurement stream cannot distinguish between physical load and sensor drift, an estimator must rely on prior assumptions about how each state evolves. If a Kalman filter or a learned recurrent estimator is configured with a belief that the sensor offset \(b(t)\) is static with near-zero variance, it attributes every variation in \(y(t)\) to changes in true contact force \(x(t)\). As new samples arrive at 1,000 \(\text{Hz}\), the filter incorporates the incoming data, repeatedly updating its internal state. Because the incoming measurements match the forward projection of the state dynamics, the calculated state covariance contracts, and the estimator reports an increasingly narrow uncertainty bound, publishing an estimate of 5 N \(\pm\) 0.05 N at \(P_{99}\). Meanwhile, Joule heating (\(I^2 R\)) in the joint stator coils (Thermal Duty Cycles) warms the sensor bridge (altering the metal’s electrical resistance as it expands), causing the physical bias \(b(t)\) to drift at a thermoelastic rate of 0.2 \(\text{N/min}\). Over 15 \(\text{min}\), \(b(t)\) shifts by 3 N. The true force on the surface has fallen from 5 N to 2 N, yet the estimator reports high confidence in a state that has drifted by 60 percent of its nominal value.

The downstream control policy responds to this false certainty by commanding motor torques to maintain what it believes is a 5 N preload. Because the true force is only 2 N, the friction force at the contact patch drops below the shear load of the workpiece, causing the part to slip from the gripper and fall into the tooling fixture. This failure occurs without triggering a single internal software diagnostic. The estimator monitors its innovation sequence, the innovation of section 1.6 evaluated at every sample between incoming measurements \(y(t)\) and predicted measurements \(\hat{y}(t)\). In this drift scenario, the estimator predicted \(\hat{y}(t) = \hat{x}(t) + \hat{b}(t) =\) 5 N, and the physical sensor returned \(y(t) =\) 5 N. The innovation is zero at every time step. An innovation check measures only whether the physical sensors disagree with the forward projection of the internal state. It cannot expose an unobservable error mode, because along the unobservable manifold the sensor output is invariant to the drift.

Recovering observability requires the system architecture to provide an additional constraint that breaks the symmetry along the unobservable subspace. An engineering design accomplishes this in one of three ways. First, independent sensing can introduce a second measurement channel with different failure physics, such as monitoring joint motor current or an optical deflection sensor that does not share the strain gauge’s thermal sensitivity. Second, the machine can perform deliberate physical excitation, commanding the gripper to briefly break contact with the workpiece for 40 \(\text{ms}\), creating a known boundary condition of \(x(t) =\) 0 N where the measured sensor value directly exposes \(b(t)\). Third, the machine can rely on periodic external anchoring, bringing the tool into contact with a calibrated reference load cell during scheduled tool-change cycles to reset the bias estimate.

When a physical state cannot be constrained by complementary sensing or active excitation, the estimator must treat that state as unobservable and refuse to publish false precision. The estimation layer must mark the corresponding state dimensions as unavailable, or widen their error bound according to the worst-case physical drift rate of the underlying hardware rather than the contracted covariance of a filtering algorithm. Formal observability analysis for linear and nonlinear dynamic systems is an established foundation of control theory, treated rigorously by Kalman (1960) for linear state spaces and by Hermann and Krener (1977) for nonlinear systems via Lie derivatives (see Continuous Covariance Propagation and Observability). A systems engineer does not rederive these rank conditions at runtime. The engineering task is to identify which physical states lack full observability rank under the current sensor suite, and to enforce that the state interface never reports statistical convergence on an unconstrained degree of freedom.

Hermann, Robert, and Arthur J. Krener. 1977. “Nonlinear Controllability and Observability.” IEEE Transactions on Automatic Control 22 (5): 728–40.

The observability and convergence results of state estimation come from control theory, and they hold only under four explicit assumptions. First, the sensor forward model must capture all physical couplings without unmodeled algebraic dependencies. Second, physical parameters such as component stiffness and thermal coefficients must remain within pre-calibrated ranges. Third, sensor noise must remain zero-mean and uncorrelated with actuator commands. Fourth, the operating trajectory of the machine must avoid singular configurations that cause the observability matrix to lose rank, a mathematical collapse indicating that certain physical states have temporarily become invisible to the sensors. The estimator validates these assumptions during operation by cross-checking redundant measurement channels, checking environmental temperature sensors, and enforcing active re-zeroing routines before the belief’s expiry. When an operating condition violates these assumptions, such as when an optical sensor becomes occluded or a kinematic configuration reaches a singular alignment, the estimator withdraws the estimate’s valid status, and downstream controllers fall back to conservative velocity and torque limits. An estimator that withdraws its own claim fails in the open; the harder failures are the ones in which nothing withdraws.

How Belief Goes Wrong

A camera stops delivering frames, yet a transport node keeps republishing its last transform with the current time, and a downstream controller’s 5 \(\text{ms}\) publish-time freshness check passes every copy. That failure differs in kind from a camera that simply falls silent, and the difference sorts every estimation failure into one of two forms. In a loud failure, a component halts execution, drops a communication packet, clears an output buffer, or asserts an explicit error bit. Downstream consumers immediately observe the absence of data or the cleared validity flag, enabling a validated fallback, such as a hold, stop, or compliant mode, when one remains feasible. In a silent failure, the memory layer continues to publish a complete, numerically well-formed state record with valid status bits, low variance indicators, and continuous timestamps, while the numbers inside the record diverge from the physical configuration of the machine. Downstream control loops consume these corrupted records without raising an exception, computing feedback corrections against a phantom geometry. Because physical actuators execute commands derived from these valid-looking records, such as driving an arm toward a target that now intersects a wall, silent state corruption can deliver force across the causal boundary of The Causal Boundary before any software monitor registers an anomaly.6

The republishing hazard of section 1.1 has a name, the frozen transform, and it is the most deceptive silent failure because a stale spatial snapshot, republished with current timestamps, masks a dead sensor. In an illustrative timeline, an optical tracking node computes a transformation matrix \(T_{\text{camera}}^{\text{object}}\) at timestamp \(t_0 = 0\text{ ms}\). If the optical sensor disconnects, drops frames, or suffers thread starvation at \(t =\) 10 \(\text{ms}\), the underlying tracking algorithm ceases to produce new spatial evidence. If the message-passing middleware or an intermediate transport node retains the last computed matrix in an unexpired buffer and republishes it at a loop rate of 1,000 \(\text{Hz}\), generating a new transmission timestamp \(t_k = t_0 + k\,T_{\text{rep}}\) with \(T_{\text{rep}} =\) 1 \(\text{ms}\), the temporal provenance of the measurement is erased. A downstream tracking controller verifies freshness by computing data age as \(\Delta t = t_{\text{now}} - t_{\text{msg}}\). Because the publisher stamped the message with the current system clock, \(\Delta t\) evaluates to less than 1 \(\text{ms}\), satisfying a strict freshness threshold of \(\tau_{\text{max}} =\) 5 \(\text{ms}\). The controller accepts the transform as current evidence. If the physical object moves at 0.5 \(\text{m/s}\) while the transform remains static, the controller accumulates 50 \(\text{mm}\) of spatial tracking error in 100 \(\text{ms}\), yet reports zero estimated tracking divergence.

Phantom-object tracks arise when the estimator decouples the lifetime of spatial coordinates from the lifetime of existence evidence. A perception module maintains an object track comprising a kinematic state vector \(\mathbf{x} = [p, v]^T\) (encoding spatial position and velocity) and an existence probability \(p_{\text{exist}} \in [0, 1]\) (encoding confidence that the object is physically real). When the physical object is removed from the workspace or an occlusion begins, raw detector hits drop to zero. If the estimation layer continues to propagate \(\mathbf{x}\) forward using process dynamics while existence probability decay is governed by a long time constant (for instance, an illustrative decay under which \(p_{\text{exist}}\) remains above an operating threshold for \(1500\text{ ms}\) without reinforcement), the tracker maintains a persistent phantom entity. The path planner routes the manipulator around empty space, or an automated grasping routine closes its fingers on an empty coordinate. Detecting a phantom track cannot rely on passive measurement absence. It requires negative evidence (section 1.5), where a line-of-sight sensor sweeps the volume and confirms that the expected occupancy is empty, or an unfulfilled contact expectation when a force sensor registers zero load at the predicted contact boundary.7

Coordinate frame misattributions divide along a similar boundary between syntax and semantics. A schema-detectable frame mismatch occurs when a publisher labels a transform with an invalid or undefined frame identifier in the transformation tree graph. A message bus with static schema enforcement rejects the packet at the serialization boundary because the named frame cannot be resolved to a known parent link. In contrast, a semantically wrong frame assignment occurs when an engineer or configuration file assigns a valid, type-correct frame identifier that represents the wrong physical link. For example, commanding a tool center point trajectory using the transform of the bare mounting flange \(T_{\text{flange}}\) instead of the extended tool tip \(T_{\text{tool}}\) satisfies every schema validation check, type assertion, and tree traversal test. In an illustrative setup with a tool tip \(150\text{ mm}\) beyond the flange along the tool \(z\)-axis, a flange-as-tip error shifts the intended contact point by that offset. Whether contact occurs depends on the fixture geometry and admitted path. Detecting a semantic frame error requires physical consistency checks, such as verifying measured motor torques against an analytical dynamic model of the arm, or cross-referencing multi-camera optical tracker observations against joint encoder kinematics.

When sensor updates drop out, an estimator relies entirely on its predictive forward model to advance state estimates across the blind interval. A well-formulated filter updates the mean state estimate \(\hat{\mathbf{x}}(t)\) while monotonically expanding the estimation covariance \(\mathbf{\Sigma}(t)\) to reflect the growth of unconstrained process noise. An overconfident predictor understates process variance \(\mathbf{Q}\), holding \(\mathbf{\Sigma}(t)\) artificially compact over an extended prediction duration exceeding \(200\text{ ms}\). When fresh sensor measurements eventually arrive, the artificially small covariance causes the filter to discount incoming observations; it weights its internal dead-reckoning so heavily that it classifies valid measurements as statistical outliers and rejects them. Engineers identify overconfidence during system qualification by executing staged synthetic occlusions, systematically dropping sensor frames for intervals between \(50\text{ ms}\) and \(500\text{ ms}\) and evaluating the innovation at reacquisition, \(y(t) - \hat{y}(t)\). If that residual consistently exceeds the predicted \(3\sigma\) covariance bounds across fifty staged trials, the process noise model is invalid. Drift along an unobservable degree of freedom, such as workpiece slippage inside compliant gripper pads, produces no reacquisition residual at all (section 1.7). At vehicle scale, a terrain belief stands until the wheels themselves contradict it.

War Story 1.1: A rover trapped in sand (2009)
Context: Spirit’s wheels started to sink into soft soil at Troy on Sol 1886, and the rover was embedded by Sol 1899, as the JPL navigation-camera record dates the approach and embedding. A darker top layer covered the site; the wheels broke through it into lighter material beneath.

Mechanism: Wheel rotations did not produce the expected translation as the rover sank. JPL reported that a rock beneath the vehicle might have contacted its underside; contact was a possibility, not an established mechanism. Surface appearance alone could not establish the subsurface load-bearing state.

Impact: Recovery teams tested extraction maneuvers, but Spirit did not regain mobility. The event illustrates a terrain-belief limit: the physical interaction supplied evidence that the previous traversability estimate could not safely preserve.

Response and lesson: JPL had created the Slip Check mode after Opportunity embedded itself in Purgatory Ripple (JPL technical account), so slip checks based on visual odometry already existed before Troy. A controller should compare wheel rotation with verified motion, expire a terrain assumption when that evidence conflicts, and admit further driving only through a feasible permission check.

Frozen transforms, phantom tracks, misassigned frames, and overconfident predictors all cross the boundary between estimation and control, and the spatial belief record of section 1.2 is the structure that must catch them there, because an ad hoc dictionary of values and matrices cannot tell active evidence from cached momentum. Every element of spatial memory must therefore reside within an explicit state record that couples the numerical value to its base coordinate frame, its sensor provenance, its growth bound, and a hard invalidation deadline that forces downstream consumers to refuse the record once that deadline passes.

The Spatial State Schema

The belief about the latch-slot finger in section 1.6 now carries everything a consumer needs to judge it, and the spatial belief record of 1.1 is the structure in which those quantities travel together. From the observation contracts of The Observation Contract the record keeps each observation’s capture epoch, clock identifier, and conversion bound, so the permission controller can convert the capture clock to its own monotonic clock with a bounded error before it evaluates age. Memory contributes four fields of its own. An estimate epoch dates the latest computation and stays apart from the evidence epoch, which dates the last physical measurement; an uncertainty-growth bound states how fast error accumulates without new evidence; a validity region states where that bound holds; and an expiry states when the belief stops being admissible. A record is VALID while fresh observations constrain it. Once observations stop, prediction can only carry it forward, to DEGRADED while its bound stays inside tolerance and to EXPIRED when that bound reaches tolerance or the machine leaves the validity region; only a fresh observation returns it to VALID. A record is UNOBSERVABLE when the sensors cannot constrain the quantity at all (section 1.7). Table 3 fills the record for the finger as a bounded-disturbance belief.

Table 3: Complete spatial-belief record: Value, unit, entity, frame, two epochs, clock conversion, error model, validity region, expiry, provenance, and lifecycle travel together. A later estimate epoch dates a computation, not a measurement, and cannot extend expiry past \(t_0 + t_{\text{exp}}\).
Field Example record value Consumer check
Entity and quantity Gripper finger relative to the latch-release slot, lateral offset from the slot centerline Match the target entity and task quantity
Value and SI unit Lateral offset, stored in meters; the bounds below are shown in millimeters for reading Decode as meters, not millimeters or force
Frame latch_slot frame, with the transform revision of the wrist-camera calibration Resolve a time-valid transform chain
Evidence epoch \(t_0\), the last wrist-camera capture before the finger occludes the slot, on the wrist camera’s capture clock Preserve across prediction and republishing
Estimate epoch Latest propagation of the offset, on the same capture clock Date the propagated estimate; do not renew evidence
Clock conversion Versioned \(t_{\text{MCU}}=a t_{\text{capture}}+b\), error at most 0.2 ms (The Observation Contract) Refuse if conversion expires; add error to upper age
Initial error \(E_0\) = 0.1 mm Check the initial bound against calibration
Growth law \(E(t) = 0.1\text{ mm} + (12\text{ mm/s})\,t + \tfrac12(50\text{ mm/s}^2)\,t^2\) Compare at evidence age with the slot tolerance
Validity region Arm acceleration and contact mode unchanged since \(t_0\); drift and disturbance bounds above validated within this region Expire immediately on a mode or residual-bound violation
Expiry \(t_0\) + 66 ms, where the growth law reaches the 1 mm slot tolerance Refuse dependent proposals once expiry passes; a later estimate epoch does not extend it
Provenance Wrist-camera observation ID, calibration revision, estimator build hash Match approved sources; hash is version provenance, not payload integrity
Lifecycle VALID at capture, DEGRADED once the finger occludes the slot, EXPIRED at \(t_0\) + 66 ms Never restore VALID without fresh physical evidence
Map cell evidence, when applicable Not used by this record; a map belief such as the rack-end cross aisle carries a per-cell capture epoch and free/occupied/unknown state Check every cell used by the candidate path; unknown stays unknown

The same form carries the belief that Grounded Intent builds on, the pose of the mug’s handle as the mug rides the takeaway conveyor toward the arm. Its frame is the robot base frame, its evidence epoch is the last wrist-camera capture of the handle, converted under the same 0.2 ms clock bound, and its initial error is 3 mm. Its growth law differs in kind from the finger’s. The belt carries the mug at no more than 0.20 m/s, a bound on speed rather than on force, so the error grows linearly from its initial value and reaches the 15 mm grasp tolerance at \(t_0\) + 60 ms. Its validity region is the belt running forward within that speed with the mug seated on it, so a knock that tips or lifts the mug, or a belt that overspeeds or reverses, expires the record at once, whatever time remains. A qualified tracker that renews the record on every frame replaces that belt-drift law with one bounded by the belt’s residual slip (Goals That Expire).

A consumer of either record converts the capture epoch to its own clock, adds the clock-error margin to the age, checks the lifecycle status and the validity region, and only then evaluates \(E(t)\) against the proposed motion’s tolerance. Passing these checks admits the belief to the separate permission decision; it does not actuate the robot. On expiry or unobservability, the controller rejects dependent proposals and selects a validated fallback whose braking, hold, or search behavior remains physically feasible.

Checkpoint 1.1: Belief expiry, observability, and the belief record

Before turning to the fallacies that arise when valid records are trusted too far, verify your understanding of evidence epochs, belief expiry, observability limits, and the belief record:

Fallacies and Pitfalls

Individually valid records need not form a consistent account of the workspace. A controller must reconcile their times and the evidence behind them before using them together to authorize motion.

Pitfall: Evaluating spatial clearance between moving bodies using asynchronous timestamps.

The base drives toward a rack end at its 1.5 m/s drive limit while a coworker walks toward the base at the 1.6 m/s walking-speed bound, so the two close at 3.1 m/s. Their position records differ in acquisition time by 20 ms, and both remain inside the 31.25 ms expiry window. Comparing them directly combines positions that never described the same instant. Across the gap the two bodies close by 62 mm, more than the 50 mm between the localization bound and the protective clearance. Individual freshness checks cannot establish joint clearance. The controller must propagate both estimates to a common evaluation time, carrying each growth bound to that instant, or bound the allowable skew from the closing speed and the clearance.

Fallacy: Bounded spatial covariance proves that a physical entity still exists in the workspace.

This is the phantom track of section 1.8 at the rack end. A person steps out, is detected, and steps back behind the rack, and the tracker keeps predicting the person along the cross aisle with a covariance below the rejection threshold. The planner then routes the base around the predicted person and treats the rest of the cross aisle as accounted for. That second conclusion fails even if the person is really there, because the absence of detections in an occluded region is evidence neither of presence nor of emptiness, and a precise extrapolation of one track certifies nothing about the space around it. A tracker must keep existence evidence separately from spatial covariance, and only negative evidence from a sensor that could have seen the cross aisle can clear the space around the track.

Summary

Spatial memory carries an estimate across the intervals when no sensor constrains it. Carrying an estimate forward is not renewing it. A predictor, a filter, or a republishing bridge advances a state’s timestamp to the current cycle, but the evidence behind that state stays rooted at the capture epoch of the last physical measurement. The spatial belief record therefore keeps two epochs apart, the evidence epoch that dates the measurement and the estimate epoch that dates the computation, and binds both to a frame, an initial error, a growth law, a validity region, and an expiry. However often the estimate is recomputed, only a fresh observation restores the record’s valid status. The frame binding matters because a transform is itself a belief with its own evidence epoch, and a stale edge passes every check the tree can run while the tool point it locates drifts out of clearance.

The lifetime of a state estimate cannot be governed by convenient software timeouts, such as an operating system watchdog or a transport buffer set to a round \(100\text{ ms}\). Expiry is a physical calculation dictated by the task’s tolerance and the validated growth rate of unobserved error. At a rack end, the belief that the cross aisle is empty expires 31.25 ms after the last frame that saw it, before perception has even dispatched that frame, regardless of whether the software message queue remains healthy. Memory therefore cannot certify the occluded cross aisle as clear. It adds a deadline to the machine’s stopping budget but no distance, the budget protects only a static obstacle, and a person still walking toward the machine is carried by a site crossing rule instead. Negative evidence sorts spatial representations. Only space that a valid ray has seen empty vouches for clearance, and only while that evidence is fresh, so neither a rendered view nor a map’s publication time is a clearance bound.

↳ Downstream: Forensic recording of expired spatial transforms populates the authority logs in The Authority Log.

Two kinds of error hide from the checks a consumer can run. A frozen transform or a phantom track publishes complete, fresh-looking records, which makes it more hazardous than an explicit missing frame, because a missing frame triggers a validated fallback while a stale record disguised as fresh passes an inadequate freshness check. Unobservable error hides deeper still. When the available sensors cannot distinguish two physical states, the innovation stays at zero and the covariance contracts while the true error grows, as it did for the force estimate whose sensor bias drifted with winding heat. Only a second channel with different failure physics, deliberate excitation, or an external reference restores what the sensors cannot see.

Key Takeaways: Evidence age and belief expiry
  • Prediction advances state, never evidence: Forward propagation moves an estimate’s timestamp to the present, but its evidence epoch stays at the last physical measurement. Age is measured from that epoch, so a republished or extrapolated state is exactly as old as the observation beneath it.
  • Expiry is a physical deadline: A belief expires when its validated error bound \(E(\Delta t)\) reaches the task tolerance \(E_{\max}\), whether that bound is a bounded-disturbance growth law or a calibrated quantile. A software timeout chosen for scheduling convenience says nothing about either.
  • Memory adds a deadline, not distance: The rack-end belief lapses in 31.25 ms, under two camera frames, so memory cannot certify an occluded region as clear. The stopping budget gains a deadline from memory but no distance term, and still protects only a static obstacle. Unknown space stays unknown, and a map’s publication time never refreshes a cell no sensor has revisited.
  • Stale records look healthier than missing ones: A frozen transform, a phantom track, or two records compared across mismatched epochs all pass freshness checks that a dropped frame would fail. Existence, cross-record skew, and transform epochs each need their own check against evidence time.
  • Unobservable error hides from innovation checks: Along a degree of freedom the sensors cannot distinguish, residuals stay at zero and covariance contracts while the error grows. Independent sensing, deliberate excitation, or external anchoring must supply the missing constraint before the estimator may report convergence.

The third law requires the permission path to decide each action on state whose age is known (principle \(\ref{pri-vol4-proposal-permission}\)), and spatial memory gives that age its definition, counting it from the evidence epoch of the last physical measurement rather than from the estimate epoch of the latest computation. Space that no valid ray has crossed has no evidence epoch at all, so it reaches the permission path as unknown, with no age that could make it clear. A known age is still not a license. It means something only through the growth law that turns it into an error bound, and however exactly the age is known, the permission path must refuse the state at the deadline where that bound reaches the task tolerance.

What’s Next: From valid beliefs to bounded goals
What may a valid belief authorize? Memory rules only on validity. It can say that a remembered object lies within a stated bound until a stated deadline, but it cannot say what the machine should do about that object. In Grounded Intent, a human goal, taking the mug from the conveyor, becomes an intent lease that consumes the belief record of the mug’s handle, carries that belief’s evidence epoch forward as its own evidence time, caps its expiry at that belief’s horizon, and grants a bounded target with effort limits, never a path, for no longer than the evidence beneath it supports.

Back to top

Footnotes

  1. Belief state formulation: In a Partially Observable Markov Decision Process (POMDP), a belief represents a probability distribution over latent states. A physical state interface additionally carries an error bound. Without the value, frame, evidence epoch, and applicable uncertainty model, a downstream optimizer cannot evaluate clearance during a sensor dropout.↩︎

  2. Hardware timer input-capture: Microcontroller timer input-capture channels latch counter values into dedicated shadow registers upon external signal transitions, bypassing operating system interrupt latency. Reading capture registers without atomic synchronization against timer overflow flags causes 16-bit counter aliasing, injecting artificial timestamp jumps equal to the rollover period (\(65.5\text{ ms}\) on a \(1\text{ MHz}\) timer). In downstream visual-inertial odometry, this temporal discontinuity introduces false acceleration spikes that destabilize platform attitude filters.↩︎

  3. Fieldbus clock synchronization jitter: Fieldbus architectures exhibit distinct hardware clock synchronization bounds: IEEE 1588 PTP over Ethernet achieves \(\delta t < 1.0\,\mu\text{s}\), EtherCAT Distributed Clocks achieve \(<100\text{ ns}\) hardware SYNC0 jitter, and CAN-FD bit-phase synchronization yields \(\pm 10\text{--}50\,\mu\text{s}\). Because spatial uncertainty scales with target relative velocity (\(e_{\text{sync}} = v_{\text{rel}} \delta t\)), microsecond-level synchronization is essential to keep temporal errors within micrometer mechanical budgets. Uncalibrated clock drift between distributed sensor and motor nodes degrades multi-axis tool coordination, inducing high-frequency chatter in joint torque loops.↩︎

  4. Bayesian state estimation provenance: Kalman (1960) introduced recursive minimum mean-square error state estimation for linear dynamic systems. Moravec and Elfes (1985) generalized these recursive Bayesian updates to spatial occupancy grids, modeling sensor beam uncertainty through explicit range-bearing distributions. In embodied systems, uncompensated propagation across extended sensor dropouts causes Bayesian belief distributions to disperse, requiring explicit invalidation horizons to prevent open-loop control divergence.↩︎

  5. Non-smooth contact dynamics: Physical surface contact introduces set-valued Coulomb friction and non-smooth stick-slip transitions that break linear covariance propagation models. Transitioning between stick and slip modes injects discontinuous velocity steps that invalidate continuous kinematic drift bounds. Managing these contact regime transitions requires hybrid state estimators that switch between distinct error growth matrices upon detecting normal force thresholds.↩︎

  6. AMR Safety Standards: ISO 3691-4 addresses safety requirements for driverless industrial trucks and their systems. Occlusion and stopping distance are assessed for each design’s operating domain and sensor and actuator configuration, not taken as universal values.↩︎

  7. Autonomous Safety Standards: UL 4600 addresses safety evaluation of autonomous products, and ISO 26262 addresses automotive functional safety. The data-age, track-invalidation, and transform checks in this chapter are design requirements of its worked architecture, not clauses of either standard.↩︎

Kalman, Rudolph Emil. 1960. “A New Approach to Linear Filtering and Prediction Problems.” Journal of Basic Engineering 82 (1): 35–45.
Moravec, Hans, and Alberto Elfes. 1985. “High Resolution Maps from Wide Angle Sonar.” IEEE International Conference on Robotics and Automation 2: 116–21.