Sensor Perception
Sensor Perception
Purpose
Why does an obstacle detection with 99 percent recognition accuracy still cause a collision if its capture timestamp, coordinate frame, and spatial uncertainty are ignored?
An obstacle detection aids safe motion only when its exposure timestamp, reference coordinate frame, and spatial covariance are explicitly known. The physical world advances continuously while an image sensor integrates photons, transfers pixels across direct memory access buses, and submits tensors to an inference model. An accurate bounding box estimate describes geometry that may no longer exist by the time an actuator applies torque. Mechanical vibration, optical distortion, and thermal camera drift introduce spatial uncertainty, altering the relationship between raw pixel coordinates and physical clearance boundaries.
Reporting high-confidence object semantics without these spatial and temporal qualifications strips downstream controllers of the information required to guarantee safe separation. Real-time perception must preserve the provenance of every observation—tracking capture timestamps across clock domains and bounding measurement error ellipsoids. In physical AI architecture, perception serves as the bridge between raw sensory hardware at the causal boundary and cognitive reasoning in the Brain, transforming noisy signals into certified spatial claims that the deterministic permission path can verify before committing kinetic energy.
Learning Objectives
- Specify an observation contract that carries capture time, a bounded clock conversion, coordinate frame, and uncertainty
- Calculate observation age across the sensing, transport, preprocessing, and inference pipeline, and convert it into its stopping-budget term
- Diagnose spatial errors caused by sensor timing, motion during capture, and misaligned coordinate frames
- Explain why average capture bandwidth does not bound a frame’s age or a renewal’s lateness, and size the host contention a renewal allowance tolerates
- Evaluate whether observation uncertainty leaves enough physical clearance for a proposed action
- Distinguish perception failures that runtime checks can detect from errors that remain silent
Photons to Spatial Claims
The warehouse mobile manipulator’s vision model can recognize a person stepping out from a rack end almost every time and still hand the planner a position that is already out of date. The world keeps moving while the navigation camera integrates and reads out its image and while the pipeline processes it. On this machine, an illustrative 51.6 ms passes between the middle of the exposure and the moment the observation reaches any consumer, and at the base’s 1.5 m/s drive limit the machine covers 77.4 mm in that time. Recognition accuracy does not say whether that observation is fresh enough for the proposed motion. In physical AI, an observation has an expiration date.
Perception runs in the Brain, above the proposal boundary of The Machine in Five Levels. It turns photons, time-of-flight pulses, and reflected radio waves into dated spatial claims and proposes them, but it holds no authority over the actuators; that authority stays with the permission path below the boundary.
The work of turning photons into dated claims does not depend on how the architecture is drawn. An end-to-end vision-to-action policy (Brohan, Brown, Carbajal, Chebotar, Dabis, et al. 2023; Brohan, Brown, Carbajal, Chebotar, Chen, et al. 2023; Zhao et al. 2023) maps camera observations to proposed actions within one network, while a classical modular pipeline splits sensor ingestion, state estimation, mapping, and trajectory generation across separate processes. Marr’s analysis of vision (Marr 1982) separates what a visual computation must achieve from the representation that carries it and the hardware that executes it, and in physical AI each of these levels must be realized somewhere between the stimulus and the response.
When an engineering team combines these responsibilities in an end-to-end model, the responsibilities remain, but their intermediate representations become harder to inspect independently. An explicit interface can expose a coordinate transform for checking, carry an expiration deadline, or let another component reject an inconsistent estimate. A hidden representation requires additional instrumentation or external checks to provide comparable evidence. If perception mistakes a wet-floor reflection for open clearance, the engineer needs a way to detect or contain the resulting error regardless of how many software modules implement the pipeline.
That evidence begins where physical energy converts into digital data. From transduction at the detector array to consumption by a state estimator, every processing stage adds observation age. Photons striking the photodiode array of a CMOS camera do not instantly appear as feature embeddings in an inference engine; exposure, readout, transport, memory transfer, preprocessing, and the vision backbone each take a share of the delay. If the perception stage hands downstream components an object detection bounding box without an explicit timestamp and a calibrated spatial transform, a downstream planner will evaluate that obstacle relative to the current position of the chassis, misplacing it by the distance the machine has traveled since capture.
↰ Prerequisite: Transduction freshness and sensor measurement aging are derived in Measurement Freshness.
Definition 1.1: Spatial information age
Spatial information age (\(t_{\text{age}}\)) is the total elapsed physical time between the physical energy integration midpoint (\(t_{\text{exp}}/2\)) at the transducer and the downstream consumption of the resulting observation tensor by a state estimator or control policy. It is the capture-to-consumption part of the measurement age that Measurement Freshness counts from transduction, and running text calls it observation age.
- Significance: Because physical mass moves while digital computation executes, sensory latency translates directly into unmodeled spatial displacement (equation 1): \[\Delta \mathbf{x} = \int_{t_0}^{t_0+t_{\text{age}}} \mathbf{v}(t)\,dt \tag{1}\] This continuous physical shift depletes geometric safety buffers before an obstacle is even classified.
- Distinction: Unlike software pipeline latency (wall-clock duration of a forward pass), observation age accounts for the full physical journey: half the exposure, readout, transport, DMA and queueing, ISP processing, the vision backbone, and dispatch.
- Common pitfall: Treating the frame arrival timestamp (\(t_{\text{recv}}\)) in userspace memory as the observation time (\(t_0\)). Using arrival time hides upstream capture and transport delays and adds unmodeled phase lag to the feedback controllers that consume the observation.
Temporal batching and phase margin erosion
Batching improves accelerator throughput by amortizing work across images, but collecting a batch across successive capture times ages its earliest member. At the navigation camera’s 60 Hz cadence (\(T_{\text{sample}} \approx\) 16.7 ms), a batch of four consecutive frames holds the first for 50 ms, the second and third for 33.3 ms and 16.7 ms, and the last not at all. Adding the oldest frame’s wait to the 51.6 ms observation age roughly doubles it, to 101.6 ms. That lies far above the refusal threshold, the budgeted age plus the clock-conversion bound that section 1.7 sets, so on this machine the permission path refuses every claim built from a batch collected across capture times. Images captured together by several cameras may instead be batched only when the batch adds no age beyond the budgeted stages and meets its deadline.
Age spends a feedback loop’s phase margin as well, because a pure delay \(t_{\text{age}}\) adds phase lag \(\omega t_{\text{age}}\) at frequency \(\omega\) without changing gain. A loop with nominal margin \(\text{PM}_{\text{nom}}\) at gain crossover \(\omega_c\) therefore exhausts it once \(t_{\text{age}} \ge \text{PM}_{\text{nom}}/\omega_c\), and exposure, processing, transport, permission, and actuator delay all draw on that one ceiling (Discrete-time sampling and zero-order hold phase lag).
The same delay is paid twice, as phase lag in the loop and as travel toward the obstacle, and a consumer can budget only the delay it can measure. What perception hands over must therefore carry a timestamp, a coordinate frame, and a usable error bound, and transduction, ingress, and geometric estimation can each invalidate them before the permission path admits a motion.
What Perception Hands Over
Raw sensor hardware yields electrical measurements such as photodiode voltages, time-of-flight counter values, or capacitive displacements. These unstructured signals differ significantly from the state variables needed for downstream autonomy. An actuator controller cannot track a desired force using a two-dimensional grid of raw pixel intensities, and a trajectory optimizer cannot plan collision-free paths using unorganized time-of-flight returns. Downstream estimation, planning, and memory subsystems require structured representations of physical state, including metric positions, velocities, bounding volumes, surface normals, and semantic classifications. Translating raw sensor streams into these representations is the operational responsibility of the perception pipeline. Each observation it hands over must carry three properties for downstream software: a capture timestamp, a reference frame, and an uncertainty envelope.
The first mandatory property is the capture timestamp \(t_0\), which specifies the temporal midpoint of the physical energy integration window (the effective moment the camera shutter was open or the laser pulse fired). Sensors do not take instantaneous snapshots; a camera integrates photons across an exposure duration, and a mechanical lidar sweeps laser pulses over an azimuth arc across several milliseconds. Downstream state estimators require \(t_0\) to integrate the measurement into the machine’s dynamic equations of motion. The arrival stamp that the definition warns against, \(t_{\text{recv}}\), marks the moment the operating system kernel completes the direct memory access transfer. Half the exposure, readout, transport, and direct memory access (DMA) all elapse before that moment, so the arrival offset \(\tau_{\text{recv}} = t_{\text{recv}} - t_0\) is never zero, and substituting arrival time for capture time corrupts the state estimator.
The magnitude of this corruption follows from equation 1 with the arrival offset in place of the full age, and at constant velocity \(v\) it reduces to \(\Delta x = v\,\tau_{\text{recv}}\). On the mobile manipulator, the navigation camera’s frame completes its DMA transfer 26.6 ms after the exposure midpoint. An estimator that stamps the frame on arrival therefore dates it that much late, and at the base’s 1.5 m/s drive limit it pairs the observation with a pose 39.9 mm beyond the one the camera occupied at capture.
↳ Downstream: How a Transform Goes Stale adds how these sensor-to-body transforms go stale over time.
The second mandatory property is the reference frame identifier \(\mathcal{F}_{\text{sensor}}\). Every spatial measurement is natively resolved in the local coordinate system of the physical transducer. A camera measures bearing angles relative to its optical center, while a radar measures range and Doppler velocity relative to its antenna array. Downstream trajectory planners and obstacle maps operate in a chassis body frame \(\mathcal{F}_{\text{body}}\) or an inertial navigation frame \(\mathcal{F}_{\text{world}}\). Transforming an observation from \(\mathcal{F}_{\text{sensor}}\) to \(\mathcal{F}_{\text{body}}\) requires applying a rigid-body spatial transformation consisting of a rotation matrix and a translation vector that define the sensor mounting extrinsics (the exact physical location and orientation where the sensor is bolted onto the machine). The frames nest from the globe down to the transducer, \(\mathcal{F}_{\text{earth}} \to \mathcal{F}_{\text{map}} \to \mathcal{F}_{\text{odom}} \to \mathcal{F}_{\text{body}} \to \mathcal{F}_{\text{sensor}}\), and every link in that tree carries its own transform and uncertainty (figure 1).
When perception outputs omit the frame identifier or rely on nominal CAD dimensions rather than calibrated physical extrinsics, small angular offsets create large spatial errors that scale with distance. An uncalibrated pitch or yaw error of \(\theta =\) 1° on a sensor detecting an object at range \(d =\) 15 m induces a transverse position error \(\Delta y \approx d \sin\theta\) of 0.26 m, about 1.7× the 0.15 m side clearance the base keeps from the racks, so the planner either steers into the racking or rejects a valid aisle as obstructed.
The REP 105 frame convention names these frames and separates a locally continuous odom pose from a globally corrected map pose. On the mobile manipulator, \(\mathcal{F}_{\text{body}}\) is the base’s base_link, tracked through the odometric frame odom, and the arm’s serial chain maps its mounting flange tool0 through the joint angles back to base_link. Table 1 lists the roles and the rule each consumer follows. A local controller consumes a continuous estimate and rejects unexpected transform jumps, and an implementation still has to check transform timestamps and estimator uncertainty.
odom-relative pose to remain continuous but does not prescribe \(C^1\) smoothness, a covariance growth rate, or fixed accuracy; these require estimator-specific evidence.
| Frame | Parent in this example | Meaning | Consumer rule |
|---|---|---|---|
earth |
root | Optional geodetic datum | Verify datum before global routing |
map |
earth if present |
Global localization; corrections may jump | Keep relocalization jumps out of local feedback |
odom |
map |
Locally continuous pose estimate; may drift | Track motion locally; monitor uncertainty and discontinuities |
base_link |
odom |
Body-fixed origin | Body pose still inherits odometry uncertainty |
sensor_link |
base_link |
Calibrated sensor mounting frame | Apply time-valid extrinsics and their uncertainty |
The third mandatory property is an uncertainty representation, such as covariance \(\mathbf{\Sigma}_z\) together with a validated tail or hard error bound when clearance depends on it. Low illumination, atmospheric scattering, specular surfaces, and model error change the residual distribution. A covariance describes second moments, not a region guaranteed to contain the true state. Estimators can use it to weight innovations against a motion model; an underestimated covariance can pull a state estimate toward a spurious measurement and lead to poor downstream commands. Safety decisions must also account for bias, missed detections, and validation limits.
Systems Perspective 1.1: The three invariants of observation
Together, these properties are the core of the observation contract that section 1.7 turns into a record that every downstream consumer checks before use. Dropping an incomplete or unbounded observation allows a downstream estimator to propagate its state forward using physical dynamics, degrading uncertainty smoothly over time. Accepting an ungrounded packet, by contrast, injects unmodeled delays, spatial distortions, and false confidence directly into the control loop. Enforcing this contract requires tracing how an observation originates, beginning at the hardware interface where physical phenomena are first sampled.
Sensor Transduction and Calibration
Lag accumulates before software receives anything at all, beginning at the photosite where physical phenomena are converted into electrical charge. A camera does not capture an instantaneous snapshot. Photons collect on a photodiode over a finite exposure duration \(t_{\text{exp}}\), which means the recorded pixel intensity represents the temporal integral of scene radiance \(I(t)\) over that interval: \[I_{\text{pixel}} = \int_{0}^{t_{\text{exp}}} I(t)\,dt\]
When the sensor or the scene moves at velocity \(v\) during exposure, a sharp boundary such as a rack upright smears across several photosites, a motion blur of width \(\Delta x_{\text{blur}} = v\,t_{\text{exp}}\). Because charge accumulates across the whole window, the time the accumulated image represents is the exposure midpoint, a lag of \(t_{\text{exp}}/2\) behind the start of exposure. With the navigation camera’s 16 ms exposure at the base’s 1.5 m/s drive limit, the image does not record where the robot is. It records where the robot was 8 ms earlier, smeared across 24 mm of motion blur.
Rolling shutter distortion and motion compensation
The exposure window schedule across the pixel array determines the spatial geometry of the image. In a global shutter sensor, every pixel photodiode accumulates charge during the identical time interval and transfers the accumulated electrons to shielded storage nodes simultaneously. In a rolling shutter sensor, silicon area and cost constraints dictate that rows are reset, exposed, and read out sequentially from the top scanline \(y = 0\) to the bottom scanline \(y = H - 1\). Each successive scanline begins exposure after a row period \(t_{\text{row}}\), creating a progressive temporal offset across the image plane:1 \[t(y) = t_{\text{start}} + y \cdot t_{\text{row}}\] where \(t_{\text{start}}\) is the start of exposure for the top row. When the body translates at linear velocity \(v\) perpendicular to a vertical feature, scanline \(y\) captures it at a spatial displacement: \[\Delta x(y) = v \cdot y \cdot t_{\text{row}}\]
Under rotational motion, rolling shutter distortion becomes non-linear. If the camera undergoes angular rotation at angular velocity vector \(\boldsymbol{\omega} = [\omega_x, \omega_y, \omega_z]^T\) (e.g., rapid yaw or pitch during aggressive robot maneuvers), the effective camera attitude changes continuously as scanlines are read. Physically, the camera points in a slightly different direction for every single row. For a scanline at vertical coordinate \(y\), the differential rotation relative to the frame start is \(\Delta \boldsymbol{\theta}(y) \approx \boldsymbol{\omega} \cdot y \cdot t_{\text{row}}\). This scanline-dependent rotation warps straight physical lines into curves, induces focal-length dilation (scaling distortion during pitch), and shears vertical walls into slanted surfaces, as demonstrated by the propeller blade curvature in figure 2.
Consider the mobile manipulator’s navigation camera while the base passes a vertical rack upright at its 1.5 m/s drive limit. The rolling shutter reads the frame’s 3,040 rows in \(t_{\text{readout}} =\) 16.6 ms, so the inter-row interval is \(t_{\text{row}} =\) 5.46 μs. An upright spanning the full height of the frame does not appear vertical in the captured array. The bottom scanline (\(y =\) 3,039) is sampled 16.6 ms after the top scanline (\(y = 0\)), producing a total spatial skew \(\Delta x = v\,t_{\text{readout}}\) of 24.9 mm.
If downstream geometry reconstruction treats this image as an instantaneous rigid projection, the 24.9 mm tilt enters the obstacle map uncorrected. That is about a sixth of the 0.15 m side clearance the base keeps from the racks, and the motion planner can read the slanted upright as an artificial intrusion into the free corridor. The stopping budget’s localization allowance \(\delta_{\text{loc}}\) assumes this skew has been removed before the estimator sees the frame. Correcting this distortion requires rolling-shutter motion compensation, the continuous-time geometric unwarping algorithm that utilizes high-rate (1,000 Hz) inertial angular velocities (\(\boldsymbol{\omega}\)) and linear velocities (\(\mathbf{v}\)) to interpolate the instantaneous 6-DoF camera pose \(\mathbf{T}_{wb}(t(y))\) for each individual scanline \(y\), restoring rigid epipolar geometry prior to 3D triangulation.
Camera calibration
Rolling-shutter compensation assumes that the camera’s geometry is already known, and every metric claim the camera supports rests on that calibration. Under the pinhole camera model (Hartley and Zisserman 2003), the intrinsic matrix \(\mathbf{K}\) maps a point \(\mathbf{P}_c = [X_c, Y_c, Z_c]^\top\) in the camera’s optical frame to the pixel \([u, v]^\top\): \[\begin{bmatrix} u \\ v \\ 1 \end{bmatrix} \sim \frac{1}{Z_c} \mathbf{K} \mathbf{P}_c = \frac{1}{Z_c} \begin{bmatrix} f_x & 0 & c_x \\ 0 & f_y & c_y \\ 0 & 0 & 1 \end{bmatrix} \begin{bmatrix} X_c \\ Y_c \\ Z_c \end{bmatrix} \tag{2}\] where \(f_x, f_y\) are focal lengths in pixels and \((c_x, c_y)\) is the principal point. Inverting the map recovers only a bearing ray, along which every depth is consistent with the pixel: \[\mathbf{r}(u, v) = \mathbf{K}^{-1} \begin{bmatrix} u \\ v \\ 1 \end{bmatrix} = \begin{bmatrix} \frac{u - c_x}{f_x} \\ \frac{v - c_y}{f_y} \\ 1 \end{bmatrix} \tag{3}\] Metric depth must therefore come from stereo disparity, a range sensor, or a learned prior, each with its own error, and section 1.5 charges that error against the clearance. Lens distortion is rectified in the image signal processor stage of table 2. The projection’s derivation, the distortion model, and Zhang’s planar calibration (Zhang 2000) are developed in Camera Projective Geometry and Covariance Propagation.
The camera’s relation to the inertial measurement unit must be calibrated in space and in time together (Furgale et al. 2013). An extrinsic rotation error \(\theta\) misplaces a point at range \(z\) by about \(z\,\theta\), and the camera’s offset from the inertial unit adds lever-arm accelerations (Spatial coordinate frames and lever-arm kinematics). An uncalibrated offset \(\Delta t\) between exposure and the inertial sample misorients the camera by \(\omega\,\Delta t\) while the base turns at angular velocity \(\omega\), misplacing the same point by about \(z\,\omega\,\Delta t\), so the timing error grows with both turn rate and range. Latching exposure starts and inertial samples on one hardware trigger bounds that offset, leaving a residual that online calibration estimates. Each of these calibrations holds only inside the conditions under which it was fitted, which is why the observation record carries a calibration revision and its validity range.
The latency waterfall
Exposure and readout are only the first stages a frame passes through. Summing the stages of the latency waterfall of Measurement Freshness up to dispatch gives the observation age: \[t_{\text{age}} = \frac{t_{\text{exp}}}{2} + t_{\text{readout}} + t_{\text{transport}} + t_{\text{DMA}} + t_{\text{ISP}} + t_{\text{backbone}} + t_{\text{IPC}}\]
Napkin Math 1.1: Transduction latency and information age
Stage-by-Stage Latency Waterfall:
- Exposure Integration Midpoint (\(t_{\text{exp}}/2\)): The navigation camera’s exposure \(t_{\text{exp}} =\) 16 ms (datasheet) \(\implies \Delta t_1 =\) 8 ms.
- Rolling Shutter Line Readout (\(t_{\text{readout}}\)): The rolling readout of the frame contributes \(\Delta t_2 =\) 16.6 ms (datasheet).
- MIPI CSI-2 Serialization & Transport (\(t_{\text{transport}}\)): The Sony IMX477 flyer specifies a four-lane option up to \(2.1\text{ Gbps/lane}\). The budget assigns \(\Delta t_3 =\) 0.8 ms (illustrative) to residual transport after streaming readout; the value is not frame size divided by peak lane rate.
- DMA Ingress & Kernel V4L2 Buffer Queue (\(t_{\text{DMA}}\)): Scatter-gather DRAM write and driver interrupt \(\implies \Delta t_4 =\) 1.2 ms (illustrative).
- Hardware ISP Debayering & Rectification (\(t_{\text{ISP}}\)): Dedicated on-chip ISP hardware pass \(\implies \Delta t_5 =\) 2.5 ms (illustrative).
- Vision Backbone Tensor Forward Pass (\(t_{\text{backbone}}\)): INT8 quantized Vision Transformer on edge NPU \(\implies \Delta t_6 =\) 22 ms (illustrative).
- Covariance Packaging & Zero-Copy IPC (\(t_{\text{IPC}}\)): Shared-memory circular ring buffer dispatch \(\implies \Delta t_7 =\) 0.5 ms (illustrative).
Total Observation Age: \[t_{\text{age}} = 8 ms + 16.6 ms + 0.8 ms + 1.2 ms + 2.5 ms + 22 ms + 0.5 ms = 51.6 ms\]
Displacement at the Drive Limit:
At the base’s drive limit \(v =\) 1.5 m/s, holding speed through the age window: \[\Delta x = (1.5 m/s) \times (0.0516 s) = 0.0774 m = 77.4 mm\]
Systems Impact: The base covers this distance before any consumer can read the frame that shows the person. Timestamped pose propagation can place the old observation correctly in the current frame, but no later stage returns the distance already traveled.
The seven stages of table 2 sum to the observation age \(t_{\text{age}}\) of 51.6 ms, which enters \(\tau_{\text{delay}}\). This age completes the worst case that Multi-Rate Cadences left open, a proposer that renews its lease and then stalls. The frame that shows the person reaches a proposer only \(t_{\text{age}}\) after the person appears. The last renewal the proposer sends before it could have seen the person can leave as late as \(t_{\text{age}}\) after the person appears, and if the proposer then stalls, the lease runs its full length from that renewal. The delay from the person’s appearance to the first retarding force is therefore the age plus the 60 ms lease, a 1 ms tick, a 1 ms bus cycle, and 20 ms of brake onset, a total \(\tau_{\text{delay}}\) of 133.6 ms.
The permission path never reads the image, so it cannot shorten the age term; the age reaches it only through the proposal’s evidence epoch. At the 1.5 m/s drive limit, the observation-age term \(v\,t_{\text{age}}\) adds 77.4 mm to the stopping distance of equation. The running total rises from the 835.5 mm that the rows through the lease path produced to 912.9 mm, which leaves 187.1 mm of the 1.10 m rack-end clear distance.
| Stage | Increment | Serial age after stage | What to verify |
|---|---|---|---|
| Exposure midpoint | 8 ms | 8 ms | Exposure and row timestamp convention |
| Rolling readout | 16.6 ms | 24.6 ms | Selected row and overlap with streaming |
| Residual transport | 0.8 ms | 25.4 ms | Selected sensor mode and host receiver |
| DMA and queue | 1.2 ms | 26.6 ms | Arbitration and queue tail |
| ISP | 2.5 ms | 29.1 ms | Mode and competing jobs |
| Vision backbone | 22 ms | 51.1 ms | Workload, clock, thermal state |
| Dispatch | 0.5 ms | 51.6 ms | End-to-end capture-to-consumer age |
Checkpoint 1.1: Spatial information age and the stopping budget
Before propagating visual perception into spatial planning, verify your understanding of how pipeline latency creates unmodeled spatial displacement:
The serial budget treats transport, DMA, and queueing as fixed increments, yet each depends on what else the host is moving at the same moment. Those terms are where camera traffic can add age that no single stage owner measures.
The Cost of Ingestion
Raw pixel streams arriving over MIPI CSI-2 or Ethernet consume DMA and memory service before inference can read them. On a shared application processor, camera writes can also delay other host reads, or overrun a receiver and drop a frame.
MIPI CSI-2 DMA ingress and zero-copy buffers
Once analog-to-digital converters on the sensor die digitize the pixel charges, the frame must travel across a physical interconnect to host memory. Camera streams enter the application processor over the MIPI CSI-2 bus.2
Systems Perspective 1.2: Silicon ingestion dataflow architecture
- D-PHY / C-PHY physical layer: Physical lanes serialize pixel packets at the selected sensor and receiver’s supported rate; packet headers carry error checks.
- System-on-chip (SoC) CSI-2 host controller: Deserializes byte packets, strips headers, and validates CRC integrity.
- Scatter-gather DMA controller: A dedicated DMA engine transfers pixel byte words directly into DRAM over the internal system interconnect (e.g., AXI crossbar).
- Kernel DMA ring buffers: A Linux
v4l2driver can map capture buffers into user space.
Once DMA write completion occurs, the kernel driver issues an interrupt.3 A driver may export a DMA buffer through dma_buf so compatible ISP or accelerator paths can share storage without an extra CPU copy. Format conversion, layout changes, or device constraints can still require a copy or transform.
The physical interconnect choice dictates the lower bound on sensor transport and host ingestion overhead (figure 3). In robotic systems, designers frequently select camera interfaces based on mechanical convenience (such as USB 3.0 cabling) without accounting for driver queueing and protocol packetization taxes. Measuring the full photon-to-actuator latency waterfall across four standard embodied vision interfaces reveals the cumulative cost of these transport choices. While standard industrial USB 3.0 introduces a substantial 58.4 ms closed-loop latency due to USB Video Class (UVC) driver packetization, userspace kernel copies, and variable USB frame-interval polling, embedded MIPI CSI-2 achieves a deterministic 34.6 ms round trip via direct SoC D-PHY direct memory access (DMA). Automotive GMSL2 serializers/deserializers (SerDes) match the latency performance of MIPI CSI-2 (34.8 ms) while extending physical signal reach to 15 meters over lightweight coaxial cable with only ~50 µs link serialization delay. Finally, neuromorphic event cameras (such as Prophesee Metavision DVS) bypass conventional exposure integration and rolling shutter readout altogether, streaming asynchronous polarity changes within microsecond intervals to enable an ultra-responsive 11.9 ms closed-loop reaction time.
Memory bus contention and observation age
The raw rates of the machine’s navigation and wrist cameras are sized in \(\ref{nbk-data-ingestion-bandwidth-storage-budgets-multi-camera}\). When the perception stack reads those frames, copies or transforms intermediate buffers, and the accelerator reads weights and writes activations, the memory controller carries several times that ingress rate. Even that multiple sits well inside the memory system’s peak bandwidth when averaged over a second.
Average bandwidth does not determine the latency of one read. Synchronized cameras concentrate their DMA writes into bursts, and arbitration granularity, read priority, and bank placement decide how long any other request waits behind them. On the mobile manipulator the burst lands only on the application processor. The permission path runs on its own microcontroller and reads its state from its own memory, so a burst cannot stretch its tick, and the shared-die placement that would expose those reads is the one Contention for Shared Resources rejects. The burst reaches the permission path only as the age of the frame a proposal is built on or the lateness of the renewal the chunk policy sends, and the two have different allowances. A frame that waits behind camera and accelerator traffic reaches dispatch older than the budgeted stages, and the refusal threshold of section 1.7 leaves no slack above them, so every proposal built on that frame is refused.
A chunk whose inference waits behind the same traffic renews the lease late, and there the lease rule of One coupled budget for the anatomy leaves 5 ms of renewal jitter above the 50 ms chunk period. Host contention may therefore stretch the chunk policy’s 40 ms inference by at most 12.5 percent, and by less in practice, because the same allowance also holds the proposer’s own jitter and the jitter of the board-level link to the microcontroller (Silicon Isolation Decisions). A renewal 10 ms late lets the lease lapse and stops the machine. Contention on this placement costs availability, never distance, because the permission path refuses what the stopping budget did not charge.
How long a burst holds the bus also matters to the camera itself. If the added delay outlasts what a receiver FIFO can buffer, the receiver may overrun and drop the frame, with a threshold that depends on its buffer and burst pattern. A dropped frame doubles the interval between two otherwise periodic samples, and the consumer treats the missing update as an age or sequence fault rather than assuming that a new image will arrive on schedule.
The hardware mechanisms governing frame delivery (DMA rings, IOMMU mappings, and DRAM arbitration) are treated by Hennessy and Patterson (2019) and Bryant and O’Hallaron (2015). For this chapter’s decision, the quantity that matters is capture-to-consumer age and its tail under the intended camera workload.
An on-time frame answers only when the scene was sampled. Where the objects in it are, and how far that estimate can be trusted, is a separate question, and the error in that answer is spent from the same clearance.
Spatial Error and Clearance Bounds
When perception places a rack upright a few centimeters farther from the base’s path than it actually stands, the base plans its pass with those centimeters missing from the 0.15 m side clearance it keeps from the racks, and no stage of the pipeline decided to spend them. Once downstream control logic uses an estimated coordinate to authorize motion, a spatial error becomes a physical displacement, spent out of the same clearance that keeps the machine off the obstacle. Transforming a sensor observation into a clearance claim requires metric grounding, which maps raw detector measurements into geometric space. This grounding relies on unprojecting two-dimensional sensor coordinates into three dimensions, an inversion that assumes known optics, rigid mounting, and predictable surface properties.
The mathematical principles of projection and triangulation, detailed by Hartley and Zisserman (2003) as well as Szeliski (2022), govern how measurement noise converts into spatial displacement. In passive stereo vision, depth \(z\) is computed from baseline \(b\), focal length \(f\), and disparity \(d\) through the relationship \(z = \frac{f b}{d}\). Differentiating this unprojection with respect to disparity yields the depth uncertainty (equation 4): \[\delta z \approx \frac{z^2}{f b} \delta d \tag{4}\] where \(\delta d\) is a small disparity error.4 If \(\delta d\) denotes a disparity standard deviation \(\sigma_d\), first-order propagation gives \(\sigma_z\approx z^2\sigma_d/(fb)\). Doubling range quadruples this modeled standard deviation at fixed \(f\), \(b\), and \(\sigma_d\). For distant landmarks, an estimator instead carries depth as inverse depth \(\rho = 1/z\), whose error stays roughly linear, and this inverse-depth metric grounding keeps an extended Kalman filter from diverging on the quadratic tail. A pulsed time-of-flight sensor instead estimates \(z=c\Delta t/2\); its range error depends on its timing and detection mechanism and may be approximately flat over a specified operating region.
To quantify the geometry, assume an illustrative stereo pair (the navigation camera alone recovers only the bearing ray of equation 3) with \(b =\) 0.20 m, \(f =\) 800 px, and disparity standard deviation \(\sigma_d =\) 0.5 px over the tested range. First-order propagation gives \(\sigma_z =\) 12.5 mm at 2 m and \(\sigma_z =\) 312.5 mm at 10 m (figure 4). Those are one-standard-deviation scales, not clearance insets or 99th-percentile bounds. A usable inset requires a tail quantile of the full spatial residual, calibrated on the intended sensor, range, texture, motion, and lighting conditions, with systematic bias and missed detections handled separately. A 150 mm clearance kept along the viewing ray to an obstacle therefore cannot be widened to a valid inset of 162.5 mm or 462.5 mm merely by adding one \(\sigma_z\). Speed must be reduced or the proposal refused whenever the validated range-error and stopping margins do not fit.
Clearance inset margins and transform chain propagation
Spatial error extends beyond the sensor optical frame. To govern motion, the estimated point must propagate through coordinate transformations to the robot base frame and the end effector. The mechanics of rigid body transformations and SE(3) kinematic frame composition are standard results treated by Craig (Craig 2005) as well as Lynch and Park (2017). Each link in this transform chain contributes scale errors, mounting calibration drift, joint encoder quantization, and mechanical compliance under load (such as a robot arm sagging slightly under gravity). Lever-arm effects amplify a small angular twist at the sensor mount linearly over distance, sweeping projected coordinates across a wide arc. If the navigation camera’s extrinsic calibration shifts by \(\theta =\) 0.005 rad against the chassis, the lateral error \(\delta x \approx z\theta\) at \(z =\) 10 m is 50 mm, equal in size to the whole 50 mm localization allowance \(\delta_{\text{loc}}\) that the stopping budget holds along the direction of travel. The comparison is one of scale, not allocation. The allowance carries small validated residuals, such as the clock-conversion travel of section 1.7, and a shift that matches it alone is a calibration failure the consumer must refuse through the calibration validity check rather than absorb. When combined with quadratic depth uncertainty, the total spatial error vector \(\boldsymbol{\epsilon}_{\text{spatial}}\) at long range is dominated by depth along the optical axis and angular lever-arm effects orthogonal to it.
↳ Downstream: Spatial covariance envelopes inflate the forward-invariant safety barrier sets derived in Safe Sets as Conditional Permission.
Every physical safety barrier operates by defending a geometric boundary around the machine. For position error vector \(\boldsymbol{\epsilon}_{\text{spatial}}\), the triangle inequality gives a conservative relation between nominal and true clearance: \[C_{\text{true}} \ge C_{\text{nominal}} - \|\boldsymbol{\epsilon}_{\text{spatial}}\|\] For a validated 99th-percentile norm-error quantile, the permission path can choose an inset clearance \(C_{\text{inset}}\) (equation 5): \[C_{\text{inset}} = C_{\text{safe}} + \|\boldsymbol{\epsilon}_{\text{spatial}}\|_{P_{99}} \tag{5}\] Here \(C_{\text{safe}}\) is the physical clearance the motion must keep, such as the base’s side clearance, and the quantile must come from a specified, checked residual distribution for the full spatial error, including calibration bias. Even then it leaves a one-percent tail under the test distribution; only a justified hard bound on the relevant error, a feasible stopping path, and a permission response within the remaining delay budget would close it. Figure 4 shows how stereo’s growing depth-error scale can require a larger inset at longer range.
Because clearance depends on propagated error, downstream planners and the permission path cannot consume raw coordinate tuples. A covariance \(\mathbf{\Sigma}\in\mathbb{R}^{3\times3}\) alone does not say whether a \(10\text{ mm}\) or \(500\text{ mm}\) clearance inset is warranted; that takes an independently validated tail or hard bound.
When spatial error is qualified, the permission path can expand the clearance inset, trading speed or workspace for separation. A textureless wall, glare, or lost calibration can invalidate that qualification without making a numerical variance literally infinite. The permission path must then refuse motion that depends on the affected claim and select a feasible fallback.
On every frame the producer builds a spatial claim in a fixed order, and any step can end it. It first converts the capture timestamp to the consumer’s clock with a bounded synchronization error and computes \(t_{\text{age}}\); if age plus clock uncertainty exceeds the refusal threshold, it marks the observation stale, and proposals that need it are refused. The permission path repeats this check at its own tick on the upper age the dispatched record carries (section 1.7). For a fresh observation, the producer then propagates position and uncertainty to the decision time with a validated motion model, unprojects each pixel measurement into optical coordinates with a sensor covariance,5 and carries point and covariance into the body frame by first-order SE(3) error propagation through the calibrated extrinsics. Validated residuals, including bias, then set the inset radius, and when no valid bound remains, the affected motion is refused. What leaves perception is a claim \((\mathbf{p}_{\text{body}}, \mathbf{\Sigma}_{\text{body}}, C_{\text{inset}}, t_{\text{age}})\), not a bare coordinate.
Learned scene features and depth
Open-vocabulary perception adds a learned semantic layer above metric sensing, built from 3D vision foundation models. A self-supervised backbone such as DINOv2 (Oquab et al. 2023) yields dense patch features, a promptable segmenter such as SAM (Kirillov et al. 2023) yields class-agnostic instance masks, and systems such as ConceptGraphs (Kuwajerwala et al. 2024; Jatavallabhula et al. 2023) and Hydra (Hughes et al. 2022) lift those masks and features through calibrated depth, extrinsics, and odometry into an object-centric 3D scene graph whose nodes carry a centroid, a spatial covariance, and a text-aligned embedding.
These features spend a compute and age budget (Williams et al. 2009). An assumed 3 TFLOP encoder at the navigation camera’s 60 Hz frame rate would require 180 TFLOP/s of sustained useful throughput before preprocessing or memory traffic, while NVIDIA lists up to 43 dense FP16 TFLOP/s for the Jetson AGX Orin in its technical brief, a peak arithmetic rate rather than sustained model throughput. The encoder cannot run at camera rate on that processor even at the arithmetic ceiling, so perception splits into a multi-rate hierarchy. Metric ingress on the application processor turns the camera’s DMA buffers into obstacle points and occupancy grids at \(30\text{--}60\text{ Hz}\) without the heavy backbone, and the foundation models update the scene graph’s semantic nodes on keyframes at \(1\text{--}5\text{ Hz}\). A slow semantic update is admissible only when fresh metric sensing and the permission path hold the physical margin between updates.
A camera-to-bird’s-eye-view lifting stage such as Lift-Splat-Shoot (Philion and Fidler 2020), developed in Camera-to-bird's-eye-view lifting (Camera ray frustum lifting and BEV splatting), places learned depth hypotheses from 2D vision backbones in the planner’s metric grid, but its depth-bin weights are model scores rather than calibrated metric error, so the clearance inset of equation 5 still rests on validated residuals.
Every bound in this section, metric or learned, assumes that the sensor behaves as calibrated. Propagation bounds the error of a sensor working as modeled; it says nothing about a sensor that has stopped working that way while still delivering frames at its nominal rate.
How Observations Go Wrong
The navigation camera can fail in two ways that look nothing alike to its consumer. If its serializer link loses lock, the driver raises an error and the frame never arrives. If light through an open dock door saturates the lens as the base turns into an aisle, the camera delivers a complete frame at its nominal rate in which the rack end is a patch of saturated white. Perception failures divide along exactly this line into two operational classes. Detectable faults occur when the physical sensor or the communication transport reports an explicit error condition. A serializer-deserializer link drops a lock (losing clock synchronization), a cyclical redundancy check detects a corrupted payload, a frame buffer overflows, or a transport bus times out. These failures are straightforward to handle since the hardware driver asserts an error flag at the ingestion boundary. Downstream consumers register the missing frame and invalidate stale state estimates, and once the evidence age exceeds its budget the permission path refuses and moves the machine to a validated fallback. In contrast, silent failures occur when the sensor hardware and transport pipeline function without error, delivering a fully populated, structurally valid tensor whose numerical values misrepresent the physical world.
Physical degradation of the optical channel is a common source of silent error.6 When direct sunlight enters an optical aperture, incident photons exceed the physical charge capacity (the “full-well capacity”) of the CMOS photodiodes, clamping pixel intensities at their maximum quantization limit (clipping the image to pure white) across large patches of the array. This dynamic range saturation destroys local spatial gradients. Feature extractors, optical flow estimators, and disparity matching algorithms fail because mathematical optimization on a uniform flat field yields zero information, yet the driver emits a valid frame at nominal line rate. Low-light environments introduce the opposite failure mode. As ambient illumination drops, the discrete, random arrival of individual photons (photon shot noise) and the intrinsic electrical fluctuations of the sensor’s amplifier (sensor read noise) dominate the measurement. The signal-to-noise ratio collapses, causing downstream neural networks to extract geometric features from random noise fluctuations rather than physical scene boundaries. Environmental contaminants on the protective lens or dome window (such as water droplets, airborne dust, or mud splatters) act as irregular refractive elements and optical low-pass filters. They attenuate contrast, distort spatial geometry, and obscure obstacles while the transport bus continues to operate without asserting a single fault flag.
Timing and transport anomalies can also corrupt an observation silently. If an unshielded physical interface experiences electromagnetic interference, deserializer bit slips (where the hardware drops or duplicates a bit, misaligning the entire subsequent data stream) can alter payload semantics without invalidating standard packet lengths. More dangerous than a lost packet, because nothing flags it, is an observation delivered with a timestamp that misreports its true capture epoch. If an image sensor captures photons at epoch \(t_0\), but a late interrupt handler or an unsynchronized system clock timestamps the resulting memory buffer at \(t_0 + \delta t\), the downstream state estimator integrates spatial features under an incorrect body pose. Consider a machine traveling with instantaneous velocity \(v\) and linear acceleration \(a\) along a given trajectory. When perception suffers an undetected timing desynchronization \(\delta t\), the state estimator assigns the detected position of an obstacle to the body coordinates at \(t_0 + \delta t\) rather than \(t_0\). Over the unmodeled time interval \(\delta t\), the physical state estimation error \(\Delta p\) between the asserted position and the true position is governed by the kinematic drift relation (equation 6): \[\Delta p = v \cdot \delta t + \frac{1}{2} a \cdot (\delta t)^2 \tag{6}\]
The arrival stamp of section 1.2 is a silent desynchronization of exactly this kind. At a constant drive speed the acceleration term vanishes, leaving the \(v\,\tau_{\text{recv}}\) offset sized there, and nothing raises a fault flag. Because the pipeline delivers a continuous stream of valid tensors, the state estimator assumes a fresh observation, computing safe actuator trajectories from stale physics. A missing frame causes a known information loss that a Kalman filter can handle by expanding its state covariance. A frame that arrives on time with a corrupt timestamp causes the filter to contract its covariance around an incorrect mean, driving actuators toward an obstacle that the model believes has already been cleared.
Defending against silent failure requires dedicated runtime health monitors at the sensor ingestion boundary. A pixel intensity histogram monitor measures the distribution of raw digitized values immediately after direct memory access transfer, identifying over-exposure saturation when values concentrate at the maximum bit depth or identifying lens occlusion when spatial variance drops below an empirical noise floor. A hardware watchdog timer tracks inter-arrival times across the direct memory access channel, flagging transport jitter when inter-frame intervals exceed \(T_{\text{nominal}} + 3\sigma_{\text{jitter}}\). Higher up the stack, an optical flow plausibility monitor compares computed pixel displacement vectors against the physical velocity envelope of the body, rejecting observations that imply accelerations beyond mechanical capability. In multi-modal configurations, a cross-sensor consistency check compares high-rate (1,000 Hz) inertial measurement unit integrations against visual odometry. If the visual pipeline reports that the machine is stationary while the accelerometer measures a sustained \(2.0\text{ m/s}^2\) specific force (the physical acceleration felt by the sensor body), the monitor detects the discrepancy and trips a perception fault.
Every runtime health monitor sees the world only through what the transducer reported at the causal boundary of The Causal Boundary, so a single sensing channel cannot validate its own semantic interpretations. A camera assembly with clean optics, balanced exposure, jitter-free transport, and synchronized microsecond clocking can still present an observation that leads a learned model to hallucinate a clear path or ignore an obstacle. Because deep vision models project high-dimensional inputs through complex non-linear manifolds, in-distribution optical anomalies (such as a high-contrast shadow across an aisle, a person printed on a carton, or specular glare from shrink-wrap or a wet floor) generate plausible pixel statistics, valid optical flow fields, and high model confidence. A single-sensor monitor evaluates whether the data collection process operated within nominal limits, not whether the inferred semantics match physical reality. A second modality, such as lidar or radar, can corroborate some optical claims when its coverage and failure dependence have been checked. A separately validated permission path can enforce clearance only within its sensing, model, timing, and actuator assumptions; neither route establishes unrestricted semantic truth.
The comparison in table 3 shows what each sensing mechanism can claim and what must be checked before combining claims. A second modality can reduce a shared failure mode only when its own coverage, timing, calibration, and fault correlation are established.
| Modality | Direct physical information | Main limitation | Permission-relevant check |
|---|---|---|---|
| RGB camera | Bearing, appearance, and image motion | Monocular depth scale is ambiguous; glare, darkness, and occlusion can hide objects | Fresh capture time and validated metric corroboration |
| Calibrated stereo | Disparity-derived range over a qualified baseline and texture region | Range error grows with distance for fixed disparity error; calibration and matching can fail | Range-dependent residual and bias validation |
| Pulsed lidar | Time-of-flight range at sampled directions | Precision, angular coverage, and weather response depend on the device and scene | Qualified range, return quality, and scan timing |
| FMCW radar | Range and radial velocity of detected returns | Multipath, angular ambiguity, and detection thresholds vary by configuration | Track association and coverage of crossing targets |
| Event camera | Asynchronous brightness-change events and motion cues | Static objects may yield few events; output bandwidth depends on event rate and encoding | Event-rate health and independent metric depth source |
War Story 1.1: Silent perception omission (2016)
Mechanism: The crossing-path collision lay outside the front-to-rear crash modes for which the 2015 warning and automatic braking systems were designed. NHTSA reported that automatic braking required agreement between camera and radar; unusual vehicle shapes could impede camera classification (NHTSA ODI PE16-007).
Impact: The NTSB found no recorded indication that forward collision warning or automatic emergency braking detected an object, and no pre-impact braking occurred. The car passed under the trailer, which sheared off its roof (figure 5).
Response: The NTSB examined operational design domain limits and driver engagement as well as the collision-avoidance systems.
Systems lesson: An absent detection is not evidence that a crossing path is clear. A permission design must account for detection coverage and unsupported scenarios; a health flag cannot certify that every relevant object was perceived.
Whatever record the producer attaches to an observation can expose only the faults the producer can check. The detectable faults and the timing and calibration conditions fit in such a record; silent semantic errors, like the omission at Williston, still require separate evidence.
The Observation Contract
A downstream tracker or the permission path cannot reconstruct exposure timing, mounting calibration, or record integrity from a tensor alone. The perception producer therefore wraps each observation in a record, and the observation is unusable without a capture epoch, a clock identifier, a bounded conversion to the permission path’s clock, and a calibration inside its validity range, whose revision the frame identifier carries; the consumer refuses a record that lacks any of them before reading the payload. The clock domain and the \(\mathcal{H}_1\) age budget come from the handoff record of The Handoffs Matrix, which this record consumes, and perception supplies the capture-side fields that let a consumer check an observation against them.
Definition 1.2: Observation contract
Observation contract is the formal cyber-physical specification governing sensory data structures traversing computational boundaries, binding a raw observation tensor to its exposure midpoint timestamp (\(t_{\text{capture}}\)), synchronization error bound (\(\epsilon_{\text{sync}}\)), sensor coordinate frame identifier, positive-semidefinite spatial covariance matrix (\(\boldsymbol{\Sigma}\)), and hardware health bitmask: \[\mathcal{O} = \langle t_{\text{capture}}, \epsilon_{\text{sync}}, \text{frame\_id}, \mathbf{T}_{\text{payload}}, \boldsymbol{\Sigma}, \text{mask}_{\text{health}}, \text{crc32} \rangle\]
- Significance: Lets a downstream state estimator or the permission path compute an upper age for each claim, select its valid transform, and bound its spatial error before the claim influences a motion proposal.
- Distinction: Unlike a generic network message header or serialization schema (e.g., Protocol Buffers), an observation contract couples serialization integrity and build provenance with physical time (\(t_{\text{MCU}} = a t_{\text{capture}} + b\)) and spatial transformation validity.
- Common pitfall: Treating perception output as a pure geometric coordinate \([x, y, z]^\top\) without propagating its spatial covariance or sensor health bitmask, blinding downstream safety filters to sensor occlusion or degraded lighting.
| Record field | Meaning | Consumer check | Navigation camera |
|---|---|---|---|
capture_timestamp_ns, capture_clock_id |
Exposure midpoint in one shared, monotonic capture clock | Convert to safety microcontroller (MCU) time with a valid bounded-skew mapping before checking age | Exposure midpoint on clock cam_nav; dispatched 51.6 ms later (table 2) |
sync_error_ns |
Bound on capture-to-MCU clock conversion error | Add to worst-case age; refuse if missing or expired | 0.2 ms, so the upper age at dispatch is 51.8 ms |
sequence_index |
Monotonic camera frame counter | Detect gaps and reordering | One count per frame at 60 Hz |
sensor_frame_id |
Transform-tree key and calibration revision | Select the valid extrinsic transform | nav_cam with its calibration revision |
tensor_payload |
Format, dimensions, length, and buffer | Validate bounds and record checksum before reading | |
spatial_covariance |
Positive-semidefinite spatial error scale | Use with separately calibrated bias and tail bounds | |
calibration_metadata |
Intrinsics, extrinsics, coefficients, temperature, validity range | Reject calibration outside its validated region | |
provenance_digest |
SHA-256 of firmware and model revisions | Match the approved build; this is not payload integrity | |
payload_crc32 |
CRC over serialized header (excluding this field) and payload | Detect accidental record corruption; not authentication | |
health_bitmask |
Hardware and perception validity flags | Refuse claims with safety-critical faults | Synchronization-loss, link, transport, and saturation bits; input-support bits from What Evaluation Cannot Establish |
The fields in table 4, the full set behind the definition’s tuple, date each observation and its generating revision. All cameras in this example latch exposure midpoints in a shared monotonic capture domain; the timestamp is not nanoseconds since the UNIX epoch. A rolling-shutter claim must also identify the row or region whose acquisition time it represents, and sequence_index detects missing frames independently of clock conversion. If the MCU has a different clock, a synchronization service supplies a versioned conversion \(t_{\text{MCU}}=a t_{\text{capture}}+b\) (with \(t_{\text{capture}}\) the exposure midpoint \(t_0\) of section 1.2), its validity interval, and a measured error bound \(\epsilon_{\text{sync}}\). Dispatch stamps the upper age estimate \(t_{\text{MCU,dispatch}}-t_{\text{MCU,capture}}+\epsilon_{\text{sync}}\) on the MCU clock at the instant the record goes to the proposer. The proposal carries that age to the permission path, which refuses the observation if the conversion is invalid or that upper age exceeds its refusal threshold. The interval from dispatch to the permission tick, which holds the chunk inference and the link, is not evidence age; the lease of One coupled budget for the anatomy charges it.
For the navigation camera the threshold is the age the stopping budget charged plus the conversion bound, 51.6 ms plus 0.2 ms, or 51.8 ms, and a record leaving dispatch already stands at it, so any further delay, from batching frames across capture times or from host contention, is refused. The budget charges only the 51.6 ms age as distance; the bound’s 0.3 mm of travel at the drive limit is carried inside the localization allowance \(\delta_{\text{loc}}\).
The transform key and calibration revision anchor metric coordinates.7 The firmware/model digest supports version provenance; payload_crc32 detects accidental changes to the actual header and payload after serialization. Neither a checksum nor a build hash proves semantic truth, and neither authenticates an adversarial producer without a keyed check.
In the health bitmask, the lower bits report hardware faults such as synchronization loss or transport errors; the upper bits can report perceptual validity concerns such as occlusion or an input outside the support the evaluation tested, and those input-support bits implement the runtime monitor specified in What Evaluation Cannot Establish. A set safety-critical bit requires refusal of the affected claim.
The explicit record layout also supports fault injection. A test can alter the capture time, omit a sequence number, inflate reported covariance, or set an occlusion flag, then check whether the consumer rejects or degrades that claim within its budget. Such mutations verify response to the injected record faults; physical glare, miscalibration, and actuator behavior still require separate tests in Adversarial Verification.
↳ Downstream: Observation contract mutations feed the hardware-in-the-loop verification harness in Fault Injection on Hardware.
Checkpoint 1.2: Observation record checks
Before passing raw sensor observations into downstream state estimators, verify your understanding of observation record contracts:
A passed record check shows only that a claim is well formed and in date, and the recurring errors in perception systems begin when such a pass is read as evidence that the claim is true.
Fallacies and Pitfalls
Perception must describe the world within the strict temporal deadlines required for control actuation. Semantic confidence and high spatial accuracy cannot compensate for pipelines that deliver stale observations, fuse asynchronous exposures, or read average bandwidth as a bound on one frame’s age.
Fallacy: Accepting an observation because its semantic classification confidence is high, regardless of its age.
The mobile manipulator receives a high-confidence observation that the rack end ahead is clear, but a stall in the vision pipeline delivers it several frames late. A person who stepped out after the exposure is absent from that frame. The confidence score describes the earlier image, not present clearance. Accepting the packet as current can therefore authorize motion toward an occupied rack end. The perception interface must check measurement age against its refusal threshold (section 1.7) and reject observations that exceed it. Any continued state propagation without new data must carry expanding kinematic uncertainty covariances that bound the reachable sets of dynamic obstacles, allowing the controller to enforce safety envelopes as historical evidence ages.
Pitfall: Fusing multi-sensor spatial observations without sub-millisecond hardware clock synchronization.
Two cameras on the mobile manipulator watch the same rack upright while the base drives at its 1.5 m/s drive limit, but their clocks differ by \(\Delta t_{\text{skew}}\). Treating the two exposure times as equal places the upright \(v\,\Delta t_{\text{skew}}\) apart in the two views, and track association can then split one physical object into two conflicting tracks (“ghosting”). Spatial calibration cannot correct an unknown temporal offset. The synchronization requirement must follow the tolerated position error and maximum relevant speed, expressed by \(\Delta t_{\text{skew}} \le \epsilon_{\text{pos}} / v_{\text{max}}\). Hardware synchronization at exposure time makes that bound enforceable; selecting a timing technology (like IEEE 1588 PTP) without first checking the spatial error budget leaves the fusion assumption unverified.
Pitfall: Inferring a critical read’s deadline from average camera bandwidth.
An acceptable average utilization says nothing about how long one read waits behind a synchronized capture and inference burst. On the mobile manipulator no permission-path read waits behind the burst, because the microcontroller reads from its own memory; the burst instead ages the frame or delays the chunk that the application processor is preparing (section 1.4), so the quantities to verify under the synchronized burst are the host’s capture-to-dispatch age and its renewal interval, not its average bandwidth.
Summary
A perception result is usable only when its capture time, clock conversion, coordinate frame, and error qualification fit the proposed motion. Age belongs to the measurement, not to the pipeline that carries it. It starts at the exposure midpoint, and on the mobile manipulator the navigation camera’s capture-to-dispatch age is a row of the stopping budget in its own right and the first interval of the pre-brake delay \(\tau_{\text{delay}}\), counted from the moment the person appears. An estimator that dates a frame by its arrival hides part of that age rather than removing it, and host memory contention can only lengthen it, so the age that binds a design is the one measured under the intended camera load. The permission path refuses any record whose upper age exceeds the budgeted age plus the clock-conversion bound, so delay beyond the budget costs availability, never distance.
Spatial error is spent from the same clearance. Stereo depth error grows with squared range under fixed disparity noise, an extrinsic shift adds a lever-arm error that grows with distance, and neither a covariance nor a learned depth score supplies a 99th-percentile clearance bound by itself. A clearance inset rests on validated residuals, checked bias, and a stated tail risk or hard bound.
The failures that matter most are the silent ones, from a misdated frame to a saturated image or a hallucinated clear path, because none of them raises a fault flag. The observation contract lets the consumer refuse whatever it can check before reading the payload (capture epoch, clock conversion, frame, calibration, integrity, and health), and semantic truth still requires independent evidence.
Key Takeaways: Spatial perception and age of information
- Observation Age Begins at Capture: Hardware-latched exposure time, row timing when relevant, and bounded clock conversion let the consumer compute an upper age before admitting motion.
- A Wrong Timestamp Is Worse Than a Lost Frame: A dropped frame widens the estimator’s uncertainty; a plausible but misdated frame contracts it around the wrong pose. Clock bounds, sequence counters, and age checks must precede any use of the payload.
- A Model Score Is Not Geometric Uncertainty: A classification or depth-bin score does not supply a calibrated spatial tail. Metric covariance, bias, and residual validation are separate inputs to a clearance decision.
- Calibration Can Change: Shock or temperature can alter the sensor-to-body extrinsic transform. The chapter’s 0.005 rad extrinsic shift at 10 m projects to 50 mm laterally, as large as the whole \(\delta_{\text{loc}}\) allowance. A shift of that size is a failure to refuse, not an error to budget, so every calibration carries a validity interval that the consumer checks.
- Health Flags Cannot Certify Semantics: A structurally valid record from a healthy sensor can still describe a clear path that is not there. A second modality narrows that gap only when its coverage and failure dependence have been checked.
- Perception Provides Dated Snapshots: Persistent spatial memory propagates state and uncertainty between observations and through occlusion; the permission path still checks whether that propagated claim fits the motion margin.
Perception is where the part’s principle, proposal is not permission, meets its freshness clause in measurable form. The proposal’s evidence epoch is only as good as the capture epoch and bounded clock conversion the observation contract declares. Sensors record physical energy interactions, not geometric truth, and a claim that cannot show when and where it was measured cannot support motion.
What’s Next: From sensory photons to spatial memory
Footnotes
Rolling shutter epipolar geometry: Rolling shutter CMOS image sensors reset, integrate, and read pixel scanlines sequentially across the array rather than exposing all photosites simultaneously. When the camera rotates during this readout interval, straight physical edges project into curved geometric arcs across the image plane. This progressive scanline skew violates rigid multi-view epipolar constraints, causing the factor graphs of visual-inertial odometry, which estimates motion by fusing camera features with inertial measurements, to compute spurious feature depths unless continuous-time B-spline camera trajectories interpolate the instantaneous pose of every scanline.↩︎
MIPI CSI-2 bus protocol: CSI-2 transmits serialized pixel streams over D-PHY or C-PHY lanes to a host receiver and DMA controller. It can avoid a network stack, but transport and queueing delay still depend on sensor mode, lane rate, controller, and competing traffic; see Heterogeneous SoC Mailboxes and Memory Barriers.↩︎
Video4Linux2 direct memory access: A V4L2 driver may use scatter-gather or physically contiguous DMA storage and expose a capture buffer through
V4L2_MEMORY_MMAP, which a compatibledma_bufimporter can share. See Heterogeneous SoC Mailboxes and Memory Barriers.↩︎Inverse-depth parameterization: In visual state estimation, parameterizing 3D landmarks by inverse depth (\(\rho = 1/z\)) converts the non-linear, heavy-tailed depth error distribution (\(\delta z \propto z^2\)) into an approximately Gaussian error space over the image plane. Under inverse depth, points at optical infinity (\(z \to \infty\)) exhibit finite variance (\(\rho \to 0\)). For the rigorous derivation and factor graph formulation, see Spatial Representations and Scene Parameterization.↩︎
Metric unprojection and covariance propagation: First-order propagation maps measurement covariance to \(\mathbf{\Sigma}_c = \mathbf{J}_{\text{unproj}}\mathbf{\Sigma}_{\text{meas}}\mathbf{J}_{\text{unproj}}^\top\) and then to \(\mathbf{\Sigma}_b = \mathbf{R}_{bc}\mathbf{\Sigma}_c\mathbf{R}_{bc}^\top + \mathbf{\Sigma}_{\text{ext}}\). Its largest eigenvalue gives a principal variance scale, not a 99th-percentile error radius. A quantile requires an assumed and calibrated error distribution, including bias; see Camera Projective Geometry and Covariance Propagation.↩︎
Lidar range walk and detector blooming: Analogous silent degradation occurs in active time-of-flight lidar receivers utilizing Avalanche Photodiodes (APDs) or Single-Photon Avalanche Diodes (SPADs). When an emitted laser pulse reflects off a retroreflective surface, intense return energy drives the photodetector into non-linear charge saturation, accelerating leading-edge discriminator threshold crossing via range walk (detector blooming). The resulting timing distortion artificially shortens the calculated round-trip time (\(t_{\text{TOF}}\)), shifting the apparent obstacle boundary several centimeters toward the sensor without asserting a hardware fault.↩︎
Mechanical thermal expansion: For a stereo baseline \(b =\) 0.20 m and an unmodeled change \(|\delta b| =\) 0.1 mm, first-order geometry gives \(|\delta z|/z\approx|\delta b|/b\). At \(z =\) 20 m, the corresponding depth bias is about 10 mm, assuming focal length and disparity are otherwise unchanged. Temperature compensation must be checked against the actual mount and calibration.↩︎


