Physical Data
Physical Data
Purpose
What does the physical machine that collected a dataset leave permanently stamped into the policy that trains on it?
A robot trajectory reflects the physical machine that produced it. Joint geometry, gearbox backlash, and contact friction shape every motion, while sensor timing decides which observation accompanies each command. Crucially, raw actions change between proposal and execution—first mapped into discrete setpoints, then admitted, rate-limited, or clamped by the real-time permission path beneath the proposal boundary. A policy trained on those records inevitably learns these collection conditions alongside the nominal task. Successful demonstrations also omit the recovery behavior needed once an error diverts the machine from its path, leaving the policy unprepared for the failure states its own mistakes produce.
Every physical demonstration consumes finite actuator life, mechanical wear, and operational time on real hardware. Because identical commands yield divergent dynamics across different physical plants, embodied datasets transfer poorly without explicit calibration. High-capacity models cannot overcome systematic recording gaps through volume alone; an autonomous policy requires data that exposes the boundary between nominal performance and unmodeled disturbance. A physical dataset is therefore evidence about one specific plant’s interaction across the causal boundary. It can be trusted only when its provenance captures what was requested, what was admitted, what the body executed, and which operating envelopes were never explored.
Learning Objectives
- Estimate usable demonstration yield from collection time and reset overhead, and budget the mechanical wear a campaign consumes
- Analyze how the collecting policy limits dataset coverage and downstream recovery behavior
- Separate failed autonomous actions from human recovery trajectories in timestamped intervention records
- Evaluate whether a dataset’s mechanical and sensing assumptions transfer to another machine
- Design a provenance record linking each observation to its requested, mapped, enforced, and measured actions and to the collecting plant
- Assess whether a dataset supports a claimed operating domain by joining its occupancy against a scenario ledger
Endogenous Experience
A policy trained on the retained demonstrations of a door-opening campaign, clean teleoperated openings of the warehouse mobile manipulator’s spring-latched cage door, can still fail at the latch on its first autonomous attempt. One possible cause is the shift the machine’s own actions create.1 In a fixed offline vision dataset, the model’s prediction does not change the next stored example. Here, an executed action changes the body’s trajectory, camera pose, and contacts, thereby changing later observations (\(s_{t+1} \sim P(s \mid s_t, a_t)\)).
Definition 1.1: Endogenous covariate shift
Endogenous covariate shift is the closed-loop distribution discrepancy wherein an embodied agent’s executed actions alter its subsequent physical states and sensory observations (\(s_{t+1} \sim P(s \mid s_t, a_t)\)), driving the system into state-space regimes unsupported by its training demonstrations.
- Significance: Explains how offline behavioral cloning can fail during closed-loop physical execution: a perturbation may drive the system into states where the policy lacks training support.
- Distinction: Unlike exogenous dataset shift caused by changing ambient lighting or seasonal weather, endogenous shift is actively generated by the agent’s own closed-loop actions altering its subsequent observation distribution.
- Common pitfall: Attempting to fix closed-loop policy failure by collecting more nominal expert demonstrations, which only increases data density along the success manifold without providing the corrective recovery data needed to escape boundary excursions.
This shift is the collection-side face of principle \(\ref{pri-vol4-endogenous-drift}\). It makes an archive of flawless executions without recovery trajectories a liability, because the policy’s first mistake moves the robot outside the demonstrated states. Let \(\pi\) be the learned policy, \(\pi^*\) the expert, \(J\) expected cumulative task cost over \(T\) steps, and \(\epsilon\) the probability of disagreement on expert-visited states under 0–1 action loss. If per-step task cost lies in \([0,1]\), behavioral cloning admits the worst-case bound (equation 1):2 \[J(\pi)-J(\pi^*) \leq T^2\epsilon. \tag{1}\]
Gathering physical experience is also expensive. Every recorded second consumes actuator fatigue life, heats motor windings, cycles gearbox teeth, and drains the battery.3 Operator rest, calibration, scene resets, and failed or rejected attempts leave less than a quarter of scheduled cell time as validated demonstration data (section 1.4). Physical demonstration data is therefore scarce, and it misleads training if logged naively.
↰ Prerequisite: Actuator thermal dissipation (\(dT/dt\)) and winding damage limits are derived in Thermal Duty Cycles.
A physical command also changes meaning on its way to the actuators. In a machine divided by the proposal boundary,4 the operator’s or policy’s request is mapped into a setpoint, admitted or modified by the permission path, and answered by the plant with a measured response. Logging only the request trains the policy on demands the plant could never execute, and logging only the response teaches it structural deflection and gear backlash as though they were deliberate human skill.
A collection campaign therefore makes two decisions before it records anything: which source will choose the actions, and which of those four forms of each action the dataset will keep as its label.
Where Physical Data Comes From
A physical dataset is not found or scraped. Every transition in the archive required moving mass against friction on a real machine. The difficulty of learning from such records was visible in Dean Pomerleau’s ALVINN (Autonomous Land Vehicle in a Neural Network) (Pomerleau 1989), which learned to steer from camera images of human driving. Because the human drivers rarely erred, the dataset held no recoveries from the lane edge, and small execution errors compounded in closed loop.
Physical data sources separate according to who chooses the action:
- In kinesthetic teaching, a colocated demonstrator physically grasps and moves the unpowered or gravity-compensated joints of the mechanism.
- In teleoperation, a remote operator guides an input device such as a leader arm, a multi-axis space mouse, or a tracked hand controller, transmitting commands across a software interface to the follower robot.5
- In scripted collection, an engineer executes a deterministic control routine, such as a state machine tracking interpolated splines or force-thresholded primitives.
- In exploratory collection, an active policy selects actions to optimize an information-theoretic objective, sample state space, or maximize a reward signal.
- In autonomous logging during deployment, the production policy commands the actuators as it performs work in a factory or warehouse, storing incoming sensor streams in situ.
Because the source decides how the machine approaches the world, these sources occupy disjoint state-action regions even on the same task. Consider releasing the cage-door latch. A colocated demonstrator guiding the arm feels the latch bind and corrects at the wrist. A teleoperator watching video has no such feel and pecks at the strike plate in short moves and pauses. A scripted controller approaches along one programmed path at constant speed, producing the same force spike against the plate on every attempt. A partially trained policy wanders, binding the latch at oblique angles until the contact tripwire ends the attempt. These logs are not noisy samples of one trajectory. They are outputs of distinct dynamical systems with different bandwidths, delays, and compliances.
⇄ Contrast: The display and operator-response delay that paces a teleoperator’s step-and-wait motion (section 1.4) contrasts with the \(1\text{ ms}\) permission tick on which each command is checked (Multi-Rate Cadences).
Treating each recorded control step as an independent experience is a common error. A logger running at a high control rate produces enormous step counts, but adjacent steps are nearly identical, since two frames a few tens of milliseconds apart differ only by sub-millimeter displacements and sensor noise. Coverage scales with the number of independent episodes and distinct physical configurations, not with the sampling clock, so a modest archive spread across hundreds of workpiece positions and lighting conditions supports far more than a recording ten times longer at one bench location.
Simulation avoids the wear bill and is treated in Policy Training. When synthetic trajectories are mixed with physical logs, each record must carry its source. A model trained on an undifferentiated mixture inherits the physics engine’s contact errors where synthetic data dominates and the hardware’s backlash and noise where physical data dominates, and no one can later tell which caused an on-robot failure. A record that names its source has settled the first of the campaign’s two decisions, and it still has to say which form of each action it kept.
Which Action Becomes the Label
By the time the gripper meets the strike plate on the cage-door latch, the operator’s hand motion has become four different signals, and the one the dataset keeps as its label decides which the learned model reproduces. The machine splits its control path across two processors (figure 1), so the label can be tapped in two computational domains:
- The unprivileged application processor captures camera streams, runs teleoperation input mapping or the policy’s forward pass, and maps coordinates.
- On the permission side of the proposal boundary (The Machine in Five Levels), the safety microcontroller (MCU) runs at \(1\text{ kHz}\) on bare metal, executing impedance control, safety barrier functions, velocity clamps, and current limits.6
The four taps along this control path capture fundamentally different physical semantics:
- The requested action (\(a_{\text{req}}\)): the raw operator input or policy proposal, tagged with its representation, units, frame, and issue time.
- The mapped command (\(a_{\text{cmd}}\)): the retargeted Cartesian or joint setpoint, tagged with the new representation, units, frame, and mapping time.
- The enforced command (\(a_{\text{enf}}\)): the same typed setpoint after the independent safety path has admitted, modified, or rejected the proposal.
- The measured response (\(a_{\text{meas}}\)): the corresponding measured position or velocity over a stated sampling interval; encoder readings and motor current are separate sensor channels, not interchangeable action coordinates.
Only the full tuple, logged across the boundary between application processor and MCU, separates what the operator asked for from what the permission path allowed.
Definition 1.2: Four-stream action tap
Four-stream action tap \((a_{\text{req}}, a_{\text{cmd}}, a_{\text{enf}}, a_{\text{meas}})\) is the time-synchronized multi-tier telemetry record that captures an action signal across its four causal transformation stages: unprivileged operator or policy intent (\(a_{\text{req}}\)), kinematically retargeted setpoint (\(a_{\text{cmd}}\)), safety-governed command (\(a_{\text{enf}}\)), and physical plant response (\(a_{\text{meas}}\)).
- Significance: Training imitation learning policies on raw requested actions (\(a_{\text{req}}\)) injects dynamically infeasible commands that trigger safety trips in deployment. Conversely, training on measured responses (\(a_{\text{meas}}\)) causes the model to imitate mechanical compliance, backlash, and external disturbances rather than intentional control. Logging all four streams disambiguates control intent from permission-path intervention.
- Distinction: Unlike classical robot telemetry logs that capture only final motor encoder states or raw teleoperation joystick inputs, the four-stream tap preserves the exact causal filtering stages applied by intermediate software and safety hardware.
- Common pitfall: Logging streams without timestamps synchronized in hardware over the IEEE 1588 Precision Time Protocol (PTP), written \(t_{\text{PTP}}\). Even a few milliseconds of clock skew between the application processor and the MCU corrupts the temporal alignment between visual observations and corresponding torque actions.
Recording the right tap does not settle when each recorded value was true. Cameras integrate photons over an exposure window of tens of milliseconds before readout and transfer, force sensors sample at around a hundred hertz, joint encoders and current loops run at the MCU’s kilohertz tick, and the action loop dispatches at tens of hertz. A data loader that latches the latest available sample pairs each action with a camera frame whose exposure was centered \(25\text{--}40\text{ ms}\) before the frame reached host memory, and the policy learns to map past states to present actions, with the phase lag baked into its weights. Logging engines prevent this in three steps:
- Exposure midpoint timestamp: Each frame is stamped at its exposure midpoint (Measurement Freshness), latched on hardware triggers synchronized over PTP or on external TTL sync lines, so that every camera’s observation carries the time at which it was true.
- Kinematic interpolation: The joint configuration at the camera’s optical timestamp is interpolated between the two nearest MCU encoder records (the spline form is in Smooth Trajectories and Bumpless Transfer).
- Alignment before serialization: Containers such as MCAP, LeRobot (Cadene et al. 2024), and RLDS (Open X-Embodiment Collaboration et al. 2024) store multi-rate streams with whatever timestamps they are given, so the midpoint alignment must happen before the state-action pairs are written.
↰ Prerequisite: Sub-microsecond PTP clock synchronization across distributed bus controllers is established in Moving Commands on Time.
A tap and a timestamp are only as good as the loop that produced them. In teleoperation, still the standard way to gather high-quality physical trajectories, that loop runs through a human, and the interface between human and machine shapes the recorded dynamics long before the data reaches a model.
Teleoperation Demonstration Costs
Teleoperation closes a feedback loop through the operator, the input device, the bus, the motor controller, and the environment (figure 2). A camera captures the workspace and streams it to a display; the operator sees the scene, decides on a correction, and moves an input device; the workstation maps that motion into a setpoint and sends it to the motor drives; the actuators accelerate the arm against inertia, friction, and contact, and the changed scene reaches the camera on the next cycle.
Each stage in this loop introduces latency that affects the timing of the correction. In a reference setup with a high-speed camera and a wired local network (illustrative; see the Reader Guide), display processing requires \(t_{\text{disp}} = 50\text{ ms}\) across exposure, H.264 decoding, and monitor refresh. Human visual perception and neuromuscular transmission introduce an operator response latency of \(t_{\text{hum}} = 200\text{ ms}\) for continuous path corrections. Command quantization, packetization, and bus transport to the robot controller add \(t_{\text{cmd}} = 10\text{ ms}\). Actuator torque ramp-up and structural inertia impose a mechanical response latency of \(t_{\text{mech}} = 80\text{ ms}\) before the robot achieves the commanded acceleration. The teleoperation round-trip latency \(t_{\text{loop}}\), the time from a contact event to the corrective actuator response, is their sum: \[t_{\text{loop}} = t_{\text{disp}} + t_{\text{hum}} + t_{\text{cmd}} + t_{\text{mech}} = 50\text{ ms} + 200\text{ ms} + 10\text{ ms} + 80\text{ ms} = 340\text{ ms}.\]
When a contact misalignment occurs at time \(t\), the operator sees it at \(t + 50\text{ ms}\), commands a correction at \(t + 250\text{ ms}\), and the arm pushes back at \(t + 340\text{ ms}\), while the arm keeps moving under momentum and contact forces. A logger that pairs the observation \(s_t\) with the command issued at the same instant pairs a state with an action chosen for a state \(250\text{ ms}\) older. The policy learns that delay as deliberate behavior and over-corrects or oscillates when it runs autonomously at the control rate. The round-trip latency is therefore a property of the dataset, not only of the rig, because it enters every demonstration as phase lag dominated by the operator’s own reaction time.
When contact force is reflected back to the operator’s hand in bilateral teleoperation, the same delay becomes a stability hazard. Anderson and Spong showed that transmission delay turns a passive link into an energy source (Anderson and Spong 1989). The delayed reaction force arrives out of phase with the operator’s hand, pushing it in the direction of motion and driving leader and follower into limit cycles against rigid contact (Passivity and Energy-Bounded Interaction). Time-domain passivity controllers and wave-variable transformations bleed that energy off with virtual damping, and the damping low-pass filters the operator’s motion, attenuating the very contact cues the demonstration was meant to capture.
The interface thus filters what a demonstration can express. A force-reflecting leader lets the operator feel binding; a visual-only rate controller encourages short moves and pauses while the operator checks the video. A backdriven joint-space leader such as ALOHA, recorded at \(50\text{ Hz}\) (Zhao et al. 2023), sits between them, conveying the leader’s mechanical feel without follower force. With any of these interfaces, operators slow down near contact, because a high-gain loop closed around \(t_{\text{loop}}\) would lose its phase margin. Table 1 lists these behaviors as artifacts to check on a given rig rather than as universal ranges.
| Interface and feedback | Command mapping | Contact information available to operator | Possible collection artifact to check |
|---|---|---|---|
| Force-reflecting bilateral leader | Leader motion mapped to follower joints; measured follower force reflected to leader | Vision and physical contact force, subject to link delay and fidelity | Contact corrections and oscillation when delayed force reflection loses passivity |
| Backdriven joint-space leader (ALOHA) (Zhao et al. 2023) | Leader joint positions copied to follower; demonstration recorded at \(50\text{ Hz}\) | Vision and leader mechanism feel, without active follower force reflection | Joint-space posture preferences and unobserved follower contact forces |
| Rate-controlled 3D mouse | Device deflection mapped to Cartesian velocity, then inverse kinematics | Primarily visual feedback unless a separate haptic channel exists | Step-and-wait motion and zero-velocity pauses under delayed feedback |
| Tracked spatial controller | Tracked hand pose mapped to tool pose with clutching and scale | Visual feedback; optional vibration is not directional force reflection | Clutch discontinuities and pose-tracking gaps |
| Instrumented exoskeleton | Wearable joint motion retargeted to robot joints | Depends on whether active force feedback is fitted | Biomechanical calibration and retargeting bias |
Whatever interface the operator holds, the rig must record every stream the session produces, including all four action taps of 1.2, and that recording has a bandwidth and storage budget of its own.
Napkin Math 1.1: Multi-camera ingestion and storage budgets
Parameters (camera geometry is datasheet; logging formats are chosen):
- Navigation camera: 4056 by 3040 px at 60 Hz, logged as 24-bit RGB (\(3\text{ bytes/px}\)).
- Wrist camera: 1440 by 1080 px at 60 Hz, logged the same way.
- External cameras: 2 RGB-D cameras viewing the door, color at 1920 by 1080 px and 30 Hz, with depth aligned to the color frame as raw
uint16millimeters (\(2\text{ bytes/px}\)). - Kinematics: 9 axes (two base wheels, seven arm joints) logged at 1 kHz. Each step stores, for every axis, position and velocity (the measured response \(a_{\text{meas}}\)), measured torque, and the requested, mapped, and enforced commands as 32-bit values, plus one 64-bit timestamp, for 224 bytes per step.
Calculation:
- Raw video bandwidth: the navigation camera streams 2219.44 MB/s and the wrist camera 279.94 MB/s. The external pair adds 373.25 MB/s of color and 248.83 MB/s of depth, 622.08 MB/s in all.
- Kinematic bandwidth: 224 bytes per step at 1 kHz is 0.22 MB/s.
- Total sustained ingestion rate: the streams sum to 3121.68 MB/s, about 3.12 GB/s.
- Hourly and shift storage consumption: uncompressed telemetry accumulates at 11.24 TB per hour, or 89.9 TB across a 8-hour collection shift on one machine.
Systems insight: The navigation camera alone writes 4.0× what a SATA SSD sustains (550 MB/s). The full session fits within a Gen4 NVMe drive’s 7 GB/s, but only by consuming storage at 11.24 TB per hour. Intra-frame video compression at an illustrative 8× brings the rate to about 390.2 MB/s (1.40 TB per hour) at the cost of encoder latency and artifacts in the frames the policy will train on. The kinematic stream, which carries all four action taps, is 0.007 percent of the bytes. The record that preserves the causal chain is the cheapest stream to log, so a logger that drops data under buffer pressure must never drop it first.
Storage bounds how much a session can record; yield bounds how much of it is worth training on. Demonstration yield (\(Y_{\text{usable}} = \lfloor T_{\text{motion}} / t_{\text{cyc}} \rfloor \cdot R\)) is the number of retained trajectory records a scheduled session produces after startup, calibration, operator rest, resets, and quality filtering have taken their share. Consider one illustrative session of the latch campaign, declared to last \(T_{\text{sched}} = 60.0\text{ min}\). Startup, calibration, and operator rest leave an active window of \(T_{\text{motion}} = 35.0\text{ min}\). Each attempt is one teleoperated opening of \(t_{\text{att}} = 25.0\text{ s}\), followed by a re-latch of the door by its fixture of \(t_{\text{rst}} = 15.0\text{ s}\) (\(t_{\text{cyc}} = 40.0\text{ s}\)), so the session fits 52 attempts. In-trial failures (20 percent) and post-collection filtering (15 percent) give a retention rate \(R = (1 - 0.20)(1 - 0.15) =\) 0.68, and the session yields 35.36 retained episodes, 14.73 min of validated data. That is a net session fraction, validated trajectory time over scheduled cell time, of 24.6 percent. A target of 5,000 retained demonstrations therefore costs 7,353 attempts and 141.4 h of scheduled cell time, 4.07× the demonstration time the campaign keeps (table 2). A campaign is budgeted in validated trajectory time, never in facility hours.
Wear accumulates per cycle, not per hour. In an illustrative wear model of the arm’s wrist, the strain-wave gear is rated for \(L_{\text{gear}} = 5{,}000{,}000\) stress cycles and the internal cable bundle for \(L_{\text{cable}} = 200{,}000\) bending cycles. Finding the handle, releasing the latch, and swinging the door add about \(n_{\text{gear}} = 14\) torque reversals and \(n_{\text{cable}} = 6\) flex cycles per opening, and the return stroke during the re-latch adds two more flex cycles. Failed attempts wear the joint as much as retained ones, so the bill is paid on all 7,353 attempts: 102,942 gear stress cycles, 2.1 percent of rated life, and 58,824 cable bends, 29.4 percent of rated life, for one campaign. The cable, not the gear, sets the replacement schedule. Three campaigns on one machine would consume most of the cable’s rated life, and an intermittent conductor break corrupts the very sensor streams being recorded.
| Session Stage / Activity | Allocated Duration (\(\Delta t\)) | Wall-Clock Share (%) | Physical Operation & Control State | Hardware Stress & Wear Mechanism | Net Trajectory Output (\(N_{\text{episodes}}\) / \(T_{\text{yield}}\)) |
|---|---|---|---|---|---|
| Cell Startup & Bus Sync | \(8.0\text{ min}\) (\(480\text{ s}\)) | \(13.3\%\) | Bus synchronization (EtherCAT/CANopen), safety watchdog verification, optical homing | Power-on electrical inrush; thermal warm-up in motor windings and sensor bridges | \(N = 0\text{ ep}\), \(0.0\text{ min}\) yield |
| Sensor Calibration & Nulling | \(7.0\text{ min}\) (\(420\text{ s}\)) | \(11.7\%\) | Camera intrinsic/extrinsic sweep, joint-torque bias nulling, encoder homing | Static holding torque; sensor bridge thermal stabilization; zero dynamic gear wear | \(N = 0\text{ ep}\), \(0.0\text{ min}\) yield |
| Operator Rest & Ergonomics | \(10.0\text{ min}\) (\(600\text{ s}\)) | \(16.7\%\) | Mandatory operator breaks to mitigate visual fatigue, cognitive saturation, and tremor drift | Passive standby; electromagnetic brakes engaged or zero-velocity holding current | \(N = 0\text{ ep}\), \(0.0\text{ min}\) yield |
| Physical Scene Resets | \(13.0\text{ min}\) (\(780\text{ s}\)) | \(21.7\%\) | Door closed and re-latched by its fixture, arm return stroke (\(15.0\text{ s/attempt}\)) | \(52\text{ return strokes} \times 2\text{ bends} = 104\text{ cable flex cycles}\); zero contact payload | \(N = 0\text{ ep}\), \(0.0\text{ min}\) yield |
| Active Teleoperated Attempts | \(21.67\text{ min}\) (\(1300\text{ s}\)) | \(36.1\%\) | Real-time human teleoperation guiding handle approach, latch release, and door swing (\(25.0\text{ s/attempt}\)) | \(52 \times 14 = 728\text{ torque reversals}\) on harmonic drive; \(52 \times 6 = 312\text{ cable flex cycles}\); contact loads below 15 N | \(N = 52\text{ raw attempts}\), \(21.67\text{ min}\) gross motion |
| In-Trial Failures (Labeled) | \(-4.33\text{ min}\) (\(-260\text{ s}\)) | \(-7.2\%\) | Handle slip, latch bind, contact tripwire above 15 N (\(20\%\) loss) | High-force contact transients; abrasive wear on gripper silicone pads | \(-10.4\text{ ep}\) from success yield; valid logs kept |
| Post-Collection QA Filtering | \(-2.60\text{ min}\) (\(-156\text{ s}\)) | \(-4.3\%\) | Offline filtering: link latency \(>100\text{ ms}\), visual occlusion, operator hesitation (\(15\%\) loss) | Zero hardware wear (offline computational validation pass) | \(-6.24\text{ ep}\) discarded |
| Net Validated Dataset Yield | \(14.73\text{ min}\) (\(884\text{ s}\)) | \(24.6\%\) | High-quality synchronized multi-modal trajectories (\(s_t, a_t\)) admitted into training buffer | Consumes \(2.1\%\) of wrist gear life and \(29.4\%\) of internal flex cable life per \(5000\text{ episode}\) campaign | \(N_{\text{usable}} = 35.36\) episodes |
Teleoperated demonstrations are bounded by what an operator can perceive through a display, coordinate through an interface, and stabilize against the loop delay, so the dataset reflects the human-interface composite rather than an unconstrained sample of the plant’s dynamics, and that bias sets the dataset’s coverage (section 1.5).
Collection Policy Coverage
A physical dataset is generated by executing actions on hardware and logging the resulting transitions. The logged samples define an empirical state-action occupancy distribution (Spencer et al. 2021) \(d^{\pi_{\text{col}}}(s, a)\), and imitation learning minimizes prediction error weighted by it, so the learned policy \(\pi_\theta\) is trained only on states where \(d^{\pi_{\text{col}}}(s) > 0\). The dataset is not a sample of the physical state space. It is the region traced by the collecting policy \(\pi_{\text{col}}\) under its control biases, interface delays, and risk tolerance.
Consider an illustrative contact task, a gripper aligning with a workpiece edge, whose deployment requires contact angles over \(\theta \in [-15.0^\circ, +15.0^\circ]\) and normal forces up to \(20.0\text{ N}\). Binned at \(1.0^\circ\) and \(1.0\text{ N}\), that is a grid of 600 cells. Two collectors each record \(1{,}000{,}000\) transitions with identical sensors. A cautious teleoperator holds contact within \(\pm 2.0^\circ\) of perpendicular under \(4.0\) to \(6.0\text{ N}\) and fills 8 cells, leaving 98.7 percent of the grid empty. An exploratory policy approaches across the full angle range, slips, and loads the contact between \(1.0\) and \(18.0\text{ N}\), filling 510 cells and leaving 15.0 percent empty. The byte counts are identical; the questions the two datasets can answer are not.
An action label exists only because the collecting policy visited that state and issued a command. At a state with zero visits, the dataset holds neither a preferred action nor evidence that none exists, and the loss function cannot tell an unstable configuration from a safe but unvisited one or a damaging one. A policy may generalize there, fail, or be stopped by a separate safeguard, and the dataset alone cannot say which; only targeted collection, expert relabeling, or a refusal of autonomous authority settles the cell. Suppose a policy trained on the cautious data is perturbed by a friction variation to \(\theta = +4.5^\circ\), outside its 8 demonstrated cells. Its network now evaluates an input no demonstration constrains, and the unvalidated action carries it to \(+6.0^\circ\) at \(11.0\text{ N}\) one control step later, a state from which the archive holds no correction. Each further action moves it farther from the corridor until the gripper jams against the fixture. That is endogenous covariate shift (1.1) acting on a coverage gap.
Measuring coverage against the extremes of the collected values is circular. A collector that logs forces only between \(4.0\) and \(6.0\text{ N}\) reports complete coverage of that interval. Coverage must instead be evaluated against an external scenario ledger, a table of the conditions the deployment requires, divided into cells. The scenario ledger is the operational design domain (ODD) written down: the object poses, surface conditions, payloads, speeds, and disturbances in which the machine is meant to operate. A dataset supports the part of the ODD that its occupancy join covers; the rest is its known absence.
Unsupported regions are found by a support envelope relational join, which cross-references the empirical occupancy \(d^{\pi_{\text{col}}}(s,a)\) (where the machine went) against the scenario ledger \(\mathcal{S}_{\text{req}}\) (where it needs to go). For each ledger cell the join reports the transition count \(N_{\text{cell}}\) and so partitions the ODD into supported regimes (\(N_{\text{cell}} \ge N_{\text{min}}\)), high-variance transition zones, and zero-support blind spots (\(N_{\text{cell}} = 0\)). A blind spot has no training supervision, so the permission path must decide whether the policy may operate there at all; low-count cells need uncertainty and scenario-specific checks.
Validation loss cannot detect the deficit. A held-out split drawn from the same trajectories shares the occupancy \(d^{\pi_{\text{col}}}\), so the cautious dataset can reach near-zero validation loss while 98.7 percent of the cells the ODD requires hold no examples (table 3, figure 3). Only the ratio of supported ledger cells to required cells, computed before deployment, measures coverage.
| Metric | Cautious Teleoperation (\(d^{\pi_{\text{teleop}}}\)) | Exploratory Policy (\(d^{\pi_{\text{expl}}}\)) | Systems Consequence |
|---|---|---|---|
| Recorded Transitions | \(1{,}000{,}000\) transitions (\(100\text{ Hz}\)) | \(1{,}000{,}000\) transitions (\(100\text{ Hz}\)) | Equivalent dataset scale |
| Operational Grid Coverage | \(8\,/\,600\) bins (\(1.3\%\)) | \(510\,/\,600\) bins (\(85.0\%\)) | Supported cells of the ODD |
| Zero-Count Blind Spots (\(N_{\text{cell}} = 0\)) | \(592\,/\,600\) bins (\(98.7\%\)) | \(90\,/\,600\) bins (\(15.0\%\)) | Unmonitored risk volume |
| Held-Out Validation Loss (\(L_{\text{val}}\)) | \(0.001\text{ rad}\) | \(0.012\text{ rad}\) | Validation loss is non-diagnostic |
| Closed-Loop Compounding Drift | Divergence beyond \(\pm 2.5^\circ\) | Bounded recovery across \(\pm 15^\circ\) | Physical policy stability |
| Training Admission | Refuse (require targeted collection) | Admit (blind spots stay cataloged) | Scenario ledger audit gate |
The latch campaign shows the join on the running machine. Every retained episode approaches the strike plate at the guarded 0.03 m/s with the plate within $$0.5 mm of nominal, and every retained success stays below the 15 N tripwire. Against a ledger holding only the nominal plate at the guarded speed, the dataset looks complete. Against the ledger the door needs, three plate positions (nominal and $$3 mm) crossed with two approach speeds (0.03 m/s and 0.10 m/s), the same data fills one cell of six, and every displaced plate and every approach at 0.10 m/s holds zero demonstrations. Nothing inside the dataset differs between the two answers; only the external ledger does. The five empty cells therefore travel with the dataset into Policy Training as its known absences, and until targeted collection fills them the machine owes those conditions a refusal rather than a guess.
Checkpoint 1.1: Empirical occupancy and scenario ledger joins
Before evaluating policy coverage across an operational design domain, verify your understanding of dataset occupancy and scenario ledger accounting:
Recording more nominal demonstrations along the corridor cannot fill a blind spot. Reaching boundary states takes supervisory takeovers or deliberate perturbation, and the data they produce raises causal problems of its own (section 1.6).
Intervention and Recovery Data
When a hazardous or boundary state does not appear in a training dataset, an engineer cannot conclude that the physical system avoids that state naturally. An absent state in the empirical log corresponds to one of four distinct physical mechanisms during data collection:
- Unvisited nominal path: The collecting policy never visited the state because nominal rollouts remained entirely within the primary task corridor. This absence represents unmonitored nominal operation, detected by cross-referencing trajectory logs against the scenario ledger.
- Permission-path interception: An active safety barrier or deterministic reflex in the permission path detected an impending boundary violation and truncated the motion before the state was reached, as the 15 N contact tripwire does to a latch attempt. This absence is detected directly in the permission path’s logs, which record the barrier trip.
- Human supervisory takeover: A human safety operator perceived a developing hazard and disengaged the autonomous policy before contact occurred. This absence is detected through the takeover disengagement flag on the vehicle or manipulator bus.
- Post-collection censoring: The machine actually entered the hazardous state and collided with an obstacle or jammed a joint, but an automated post-processing script discarded the episode as corrupt or incomplete. This absence is not detected by dataset inspection, because the failure transitions were purged before the training arrays were written to disk.
Each mechanism produces an identical zero count in a scenario cell, yet treating them identically leads to incorrect conclusions about policy stability.
Human takeovers and automated disengagements introduce a selection trap into fleet data. An intervention occurs precisely when the autonomous policy has entered a region of state space that it handles poorly, causing tracking error or model uncertainty to exceed an operational threshold. Because disengagements are triggered by failing rollouts, the collected intervention data is not drawn from the nominal state distribution of the task. A trajectory recorded after such a takeover captures an expert recovering from an edge case disturbance. The distribution of states visited during these rescues, denoted \(d^{\text{takeover}}\), is conditioned on policy failure rather than standard task execution. Treating intervention logs as if they were drawn from the nominal distribution \(d^{\pi^*}\) distorts the training distribution toward states that only exist because the prior policy failed.
Appending a rescued episode directly to a behavioral cloning dataset can make unsafe policy commands look like expert demonstrations. Before takeover, the learned policy may issue commands that increase heading error or contact force; after takeover, the expert issues corrective commands from states the learner visited. Those drift states are valuable for corrective aggregation when an expert can label a safe action there.
Separating these labels requires locating the divergence point where the autonomous policy departed the nominal path (extending The Multi-Rate Bridge), rather than treating takeover as the first error. An out-of-loop supervisory takeover (Human Authority Over Actuation budgets its latency) takes longer than the teleoperation loop’s continuous correction (section 1.4); at an illustrative \(100\text{ Hz}\) that delay spans many control cycles, all logged under the learner’s authority.
Detecting \(t_{\text{div}}\) in offline ingestion pipelines requires searching backwards in time from the physical override timestamp across a synchronized pre-trigger telemetry buffer: \[\begin{aligned} \hat{t}_{\text{div}} = \min \Bigl\{ t \in [t_{\text{takeover}} - \Delta t_{\text{lookback}}, t_{\text{takeover}}] \;\Big|\; &\|\mathbf{x}_{\text{plant}}(t) - \mathbf{x}_{\text{ref}}(t)\|_2 > \delta_{\text{tol}} \\ &\lor\; \text{Tr}(\boldsymbol{\Sigma}_{\pi}(t)) > \sigma_{\text{thresh}}^2 \Bigr\}. \end{aligned}\] The pipeline scans a stated lookback window (here \(\Delta t_{\text{lookback}}=1.0\text{--}2.0\text{ s}\)) for the first threshold crossing, using calibrated tracking residuals and, if available, policy uncertainty. The thresholds are scenario-specific.
Intervention curation separates learner commands, expert labels on learner-visited states, and post-takeover recovery at the divergence and takeover timestamps. It keeps the learner commands between divergence and takeover in the audit log with their true authority and safety flags but never as expert imitation targets, admits the states they reached to corrective aggregation only when a qualified expert can label a safe action there, and records the recovery after takeover as a separate expert sequence. The states that reveal the policy’s coverage gap are never erased.
The states between divergence and takeover are where principle \(\ref{pri-vol4-endogenous-drift}\) becomes concrete. The learner’s own actions chose them, no demonstration visits them, and they are the states a behavioral-cloning policy reaches after its first error, which is why their absence is expensive. Cloning on expert states alone admits only the quadratic worst-case bound of equation 1, \(O(T^2\epsilon)\).
In corrective aggregation, such as DAgger (Ross et al. 2011), an expert labels states visited by the learner, including eligible pre-takeover drift states. The learned policy can attain a cost bound linear in horizon and in its loss on learner-visited states, \(O(T\epsilon)\).7
For \(T=500\) steps (\(5.0\text{ s}\) at \(100\text{ Hz}\)) and \(\epsilon=0.01\), the expressions \(T^2\epsilon=2500\) and \(T\epsilon=5\) illustrate different horizon dependence. Because both terms omit constants and rest on different error assumptions, their ratio is not a \(500\times\) improvement. With per-step cost at most one, cumulative task cost itself cannot exceed 500 in this example (figure 4).
Training successive policy generations on drift-window learner commands mislabeled as expert actions can create a retraining cascade:
- Generation 1: A policy operates with a slight calibration error, causing a robotic arm to drift toward a joint torque limit until a human operator grabs the override pendant.
- Generation 2: If the drift-window commands are presented as expert labels, the policy learns to repeat them at similar off-nominal states. On this rig, that could increase excursions toward the torque limit before a correction.
- Generation 3: Operators may then intervene earlier. Intervention rate can rise while validation loss on the mislabeled archive falls, making the offline metric look better as physical performance worsens.
To detect such a cascade, the permission path emits an intervention record for each takeover or barrier trip,8 the training-side view of the forensic authority record that The Authority Log builds. The intervention record needs more than a binary flag on video: a pre-trigger buffer long enough for the declared lookback (two seconds in this example), synchronized observations and action taps, a monotonic override timestamp and authority ID, and the subsequent recovery sequence. Without the pre-trigger evidence, the pipeline cannot locate \(t_{\text{div}}\) reliably or separate unsafe learner commands from states eligible for expert relabeling.
Checkpoint 1.2: Intervention curation and compounding error bounds
Before curating human intervention logs for policy retraining, verify your understanding of compounding error dynamics and timestamp attribution:
Even with expert labels on learner-visited states, the resulting demonstration archive remains tethered to the kinematics and dynamics of the collecting machine. Pooling demonstrations across robot arms, bases, and end-effectors raises the cross-embodiment questions in section 1.7.
Cross-Embodiment Transfer
Data from one physical machine does not represent the task. Every recorded trajectory binds the physical characteristics of the collecting hardware into its numerical values, including link lengths, joint velocity limits, sensor optical paths, communication delays, and contact surface compliances. When an engineer saves an observation-action trace \((o_t, a_t)\), the record does not capture an abstract manipulation skill. It captures the embodiment contract established between that specific machine and its environment. This contract encompasses the sensor geometry and optical calibration, the sampling period and transport latency of the perception bus, the coordinate frames and kinematic limits of the actuators, the torque-speed saturation curves of the motors (including reflected rotor inertia \(N^2 J_{\text{rotor}}\) in geared transmissions, as established in Actuator Transmission Limits), and the friction and compliance of the contact interface. Altering any parameter in this contract changes the mapping between commanded inputs and physical state transitions.
Where the embodiment contract describes one machine, the cross-embodiment compatibility contract decides whether records made under one embodiment contract may drive another. It is the formal geometric, dynamical, and temporal specification (such as those underpinning the Open X-Embodiment (Open X-Embodiment Collaboration et al. 2024) and DROID (Khazatsky et al. 2024) datasets) that keeps a force profile or spatial velocity learned on a heavy industrial arm from reaching a lightweight desktop robot unscaled. It defines the forward kinematics mapping, control rate normalization (\(\Delta t\)), and actuator effort bounds required before multi-robot trajectories can be transferred to a target physical plant.
Engineering a data reuse pipeline requires separating task-level invariants from embodiment-specific realizations. A task-level invariant is a geometric or dynamical condition required by the environment, such as bringing a compliant probe into contact with a terminal pad at an angle within \(5^\circ\) of the surface normal while maintaining a normal force between \(2.0\text{ N}\) and \(5.0\text{ N}\). An embodiment-specific realization is the sequence of joint encoder counts, pulse-width modulation duty cycles, bus transport delays, and camera pixel rasters that achieved that condition on a specific test rig. While the contact target exists in world coordinates, the recorded dataset contains only the realization. A downstream model trained directly on those raw records learns the idiosyncratic dynamics of the source rig, including its structural flex, drive backlash, and camera mount vibration.
In an illustrative example, consider transferring a dataset recorded on Machine A to Machine B for an identical touch-and-grasp task:
- Machine A uses a parallel-jaw gripper with a \(0\text{ to }100\text{ mm}\) stroke range, a \(50\text{ Hz}\) command rate (\(20\text{ ms}\) period), a continuous force limit of \(40\text{ N}\), and a wrist camera offset \(45\text{ mm}\) behind the tool center point with a \(15\text{ ms}\) transport delay.
- Machine B uses a compact gripper with a \(0\text{ to }60\text{ mm}\) stroke range, a \(20\text{ Hz}\) command rate (\(50\text{ ms}\) period), a continuous force limit of \(15\text{ N}\), and a wrist camera offset \(20\text{ mm}\) behind the tool center point with a \(35\text{ ms}\) transport delay.
Mapping the dataset from Machine A through the limits of Machine B reveals the points where transfer fails mechanically. For an object width of \(75\text{ mm}\), Machine A opens to \(80\text{ mm}\), a command that exceeds the mechanical travel of Machine B by \(20\text{ mm}\) and drives its leadscrew against the physical endstop. During the hold phase, Machine A applies \(30\text{ N}\) of clamping force. That same command exceeds Machine B’s continuous rating of \(15\text{ N}\) by a factor of two, driving the motor amplifier into current saturation. In the time domain, subsampling the \(50\text{ Hz}\) trajectory onto Machine B’s \(20\text{ Hz}\) bus aliases high-frequency contact transitions and introduces a zero-order hold phase lag of \(15\text{ ms}\) (Discrete-time sampling and zero-order hold phase lag). Combined with the \(20\text{ ms}\) increase in camera transport delay (\(35\text{ ms} - 15\text{ ms}\)), Machine B perceives contact events \(35\text{ ms}\) later than Machine A does. At an approach velocity of \(0.10\text{ m/s}\), this added sensor-to-actuation latency carries the gripper \(3.5\text{ mm}\) past the intended contact plane before the first corrective command can be dispatched, while the \(25\text{ mm}\) camera mounting disparity introduces visual parallax that shifts the apparent position of the workpiece boundary in the image frame.
Standard machine learning pipelines reconcile these physical discrepancies by normalizing observation and action channels to \([0, 1]\) or \([-1, 1]\). Normalization creates a cosmetic alignment of numerical ranges, leaving physical incompatibility intact. If an opening command is normalized by dividing by the maximum stroke, a normalized command \(u = 0.80\) corresponds to \(80\text{ mm}\) on Machine A, which clears the \(75\text{ mm}\) object. On Machine B, that same normalized command \(u = 0.80\) commands an opening of \(48\text{ mm}\) (\(0.80 \times 60\text{ mm}\)), causing the fingers to collide with the object during the approach trajectory. Normalization rescales numbers, but it cannot expand actuator stroke, cannot lower electrical resistance, and cannot eliminate transport latency. When the data distributions of both embodiments are plotted in normalized space, the demonstration trajectories appear identical, yet only the subset lying within the physical intersection (openings under \(60\text{ mm}\), forces under \(15\text{ N}\), and command bandwidths under \(20\text{ Hz}\)) can execute on the target hardware without saturation or impact damage.
Pooled corpora raise these questions at scale. While scaling demonstration volume within a single narrow workcell yields rapid in-distribution saturation (\(>85\text{ percent}\) success within 15 hours on ALOHA (Zhao et al. 2023)), out-of-distribution (OOD) generalization across novel scenes, visual distractor textures, and unmodeled object variations remains poor (\(22\text{ percent}\), creating a 66 percentage point brittleness envelope). Overcoming this failure requires expanding beyond single-cell capture into in-the-wild multi-scene collection (such as the 564 kitchen and household environments in DROID (Khazatsky et al. 2024) or the multi-institution BridgeData v2 (Walke et al. 2023)) and pooling heterogeneous kinematics across embodiments (such as the 22 robot platforms in Open X-Embodiment (Open X-Embodiment Collaboration et al. 2024) and foundation VLAs like \(\pi_0\) (Black et al. 2024)). As mapped in figure 5, scaling multi-scene and cross-embodiment demonstration hours systematically contracts this generalization gap from 66 percentage points down to 13 percentage points.
Stroke, force, rate, and viewpoint mismatches arise in any pool exactly as in the two-gripper example. Two further mismatches arise only when the pool mixes kinematic topologies:
- Action space incompatibility: A 7-DoF arm records joint positions \(\mathbf{q} \in \mathbb{R}^7\), a bimanual workstation records \(\mathbf{q} \in \mathbb{R}^{14}\), and a mobile base logs planar velocities \((\dot{x}, \dot{y}, \dot{\theta})\). Because joint-space labels do not transfer across topologies, multi-embodiment models commonly project actions into a Cartesian end-effector delta: \[\Delta \mathbf{x}_t = (\Delta x, \Delta y, \Delta z, \Delta \phi_{\text{roll}}, \Delta \phi_{\text{pitch}}, \Delta \phi_{\text{yaw}}) \in SE(3), \quad a_{\text{gripper}} \in [0, 1].\] The delta standardizes geometry and discards actuator dynamics. A displacement that is trivial for an industrial arm near its nominal configuration can carry a lightweight arm through a kinematic singularity (Forward kinematics and singularities), where the required joint velocities grow without bound and overcurrent protection trips.
- Null-space projection: Mapping a 7-DoF arm into a 6-DoF Cartesian delta projects out the one-dimensional null space, the elbow swivel angle \(\psi\). On the target, the inverse kinematics solver resolves that redundancy by its own criterion, such as joint-limit avoidance or manipulability. Without the demonstrator’s elbow posture, the target arm can drift into joint limits or sweep its elbow into fixtures during motions that were collision-free in the source data.
Reusing physical data across embodiments requires a formal compatibility check at the ingestion boundary before any sample enters the training buffer. The ingestion pipeline must transform the recorded trace through forward kinematics to evaluate coordinate consistency (Lynch and Park 2017), verify that all commanded joint velocities and accelerations lie within the target machine’s dynamic limits (Tedrake 2023), and verify that perceptual features match the target sensor’s field of view and resolution. Actions that exceed kinematic reach or actuator effort limits cannot simply be clamped at execution time; clamping distorts the velocity profile and breaks the causal relationship between observed errors and corrective motions. Samples that pass the compatibility check are admitted; samples that can be geometrically re-targeted through inverse kinematics are relabeled; and samples that demand unavailable physical authority or unobservable states are rejected. An unverified ingestion pipeline across an embodiment boundary injects invalid dynamics into the model’s training distribution. The resulting policy does not fail because the neural network underfitted the training loss. It fails because the training dataset contained actions that were physically sound only on the machine that collected them.
A compatibility check guards one boundary, but data collected on the machine itself can go wrong without crossing an embodiment. Operator drift, calibration shifts, rig mismatch, and misclassified failures corrupt physical data while leaving the tensor dimensions untouched (section 1.8).
How Datasets Go Wrong
Beyond the absences the latch ledger declares (the faster approach and the displaced plate), the harder failures lie inside the one supported cell, where the retained demonstrations can be wrong in ways no tensor check sees, because every metric computed on a dataset evaluates only the states it contains.
The first is operator drift. Over the campaign’s 141.4 h of scheduled cell time, an operator who began by moving deliberately, with corrective sub-movements and pauses at the strike plate, develops muscle memory and starts rounding corners and swinging the door on momentum. The late demonstrations rely on timing the machine cannot reproduce under closed-loop delay at deployment. Drift is detected by partitioning the dataset into temporal windows and comparing episode duration, corrective jerk count, command magnitude, and intervention rate across them. Pooling drifted windows as one identically distributed sample teaches a policy that hesitates or lunges depending on which stage of the operator’s learning curve it happens to imitate.
The second is a change in the machine between recording runs. File containers such as HDF5, Zarr, or Parquet enforce tabular schema consistency and conceal everything else. An arm shut down overnight, recalibrated, or bumped by a technician changes its encoder zero offsets or a camera’s extrinsics, and a camera mount shifted by \(3\text{ mm}\) moves the workpiece in the image while the array keeps its shape and type. Clock synchronization can drift across controllers in the same way, shifting the pairing between what the arm sees and what it feels by tens of milliseconds (Moving Commands on Time).9 figure 6 shows the visual half of the check, in which the robot’s kinematic model, rendered from its joint encoders, is projected over each camera stream so that an extrinsic shift shows up as a gap between wireframe and arm. The recorded half joins every episode against an audit trail of calibration identifiers and clock-health metrics, since a dataset whose trajectories do not share one valid calibration forces the learner to fit conflicting mappings in the same coordinates.
The third is a mismatch between nominally identical rigs, which the fleet meets as soon as a second mobile manipulator records latch demonstrations for the first to learn from. Two arms of the same model can differ in gearbox wear, thermal behavior, or firmware filters. In an illustrative comparison, a new arm shows \(0.02^\circ\) of backlash and a worn one \(0.35^\circ\) with a \(22\text{ ms}\) torque-response lag, and the worn arm’s demonstrations carry compensating lead and gain that oscillate on the new one. A record of each arm’s embodiment contract (kinematics, actuator serial numbers, firmware versions, and a dynamic model) bound to every file lets ingestion exclude trajectories that fail the target’s compatibility contract.
The fourth is conflating failed attempts with failed recordings. A failed attempt is a complete physical sequence that missed the objective; discarding it creates survivorship bias and strips the dataset of the error states a stabilizing policy needs. A failed recording is a logging fault, such as a dropped packet, a corrupt frame, or a buffer overrun; keeping it without accounting presents irregular sample periods as a constant control interval. Ingestion therefore reconciles every attempt as a retained success, a retained failure, or an excluded recording with its error code. In the latch campaign, an attempt ended by the contact tripwire is a failed attempt and stays in the archive, while a recording whose force channel dropped packets leaves it with its error code.
A unit contract can fail just as silently, and no format check will notice (1.1).
War Story 1.1: Mars Climate Orbiter unit mismatch (1999)
Mechanism: A ground software file (called “Small Forces” in the investigation) reported thruster impulse in pound-force-seconds (\(\text{lbf}\cdot\text{s}\)), while the trajectory models that consumed it assumed newton-seconds (\(\text{N}\cdot\text{s}\)) as the interface specification required. Because \(1\text{ lbf}\cdot\text{s} = 4.45\text{ N}\cdot\text{s}\), every desaturation event entered the models too small by a factor of \(4.45\). The values were well formed, in range, and of the expected type, so no schema check could flag them.
Impact: The accumulated error was never resolved during cruise. At orbit insertion the spacecraft passed about \(57\text{ km}\) above the surface instead of the planned \(226\text{ km}\) and was lost.
Response: The Mishap Investigation Board traced the loss to the missing unit conversion in the ground file and recommended that the Mars Polar Lander project verify consistent units throughout design and operations and audit all data transferred between JPL and Lockheed Martin Astronautics for compliance with the interface specification, along with stronger navigation review and anomaly escalation.
Systems lesson: A number carries its meaning only together with its units and its frame. A robot dataset runs the same risk on every action tap and every calibration. A record that stores a value without the units, frame, and calibration identifier it was produced under lets a mismatched rig pass every tensor check, which is why the provenance record of section 1.9 binds each tap to them.
Beyond silent coordinate and unit mismatches, an even more perilous failure occurs when an autonomous policy encounters a situation completely absent from its training distribution. While a software unit mismatch accumulates numerical error, an unvisited state forces a learned policy to interpolate or select nominal primitives outside its support envelope. When sensory occlusion renders the true physical state unobservable, the resulting execution can convert a survivable incident into physical catastrophe (1.2).
War Story 1.2: Cruise robotaxi pedestrian dragging (2023)
Mechanism: The robotaxi executed a rapid emergency stop within \(0.46\text{ s}\) at \(1.4g\), coming to rest over the pedestrian pinned beneath the chassis. However, the collision detection system misclassified the event as a lateral side-swipe rather than a collision with an obstacle under the vehicle. Because roof-mounted LiDAR and perimeter cameras had an acute physical blind spot beneath the floorpan, the pedestrian became unobservable. The learned driving policy, trained on nominal fleet demonstrations where pulling out of traffic after an incident is standard practice, selected a nominal pull-over primitive to clear the travel lane. The training distribution contained zero demonstrations of post-collision states with humans lodged under the chassis (\(d^{\pi_{\text{col}}}(s) = 0\)).
Impact: Executing the ungrounded pull-over primitive outside its training support envelope, the robotaxi accelerated forward and dragged the trapped pedestrian for \(20\text{ feet}\) (\(6.1\text{ meters}\)) at up to \(7\text{ mph}\) (\(11\text{ km/h}\)), inflicting severe secondary trauma before coming to a final halt.
Response: Regulatory agencies immediately suspended commercial deployment permits. The company grounded its entire fleet, overhauled its perception and safety software, and instituted hard physical governors requiring positive post-collision interlocks and teleoperation confirmation before any vehicle re-engagement.
Systems lesson: The Cruise incident (figure 7) demonstrates the mortal danger of missing support in physical AI datasets. Outside the empirical support envelope of demonstrated states, policies do not fail into safe passivity—they interpolate nominal primitives that convert a survivable incident into a fatal interaction. Datasets must bound their explicit scenario support hulls, and architectures must enforce deterministic control refusal when sensor occlusions and post-impact uncertainty violate support boundaries.
When a coordinate failure or unobserved scenario blind spot passes ingestion, loss curves do not reveal it. The network averages conflicting demonstrations or hallucinates actions outside its training support, converging on training loss while generating dangerous physical behaviors at deployment. No operation on a tensor can verify that the machine was calibrated, that the operator did not drift, or that unvisited regions are safe, so the collection history must be recorded as rigorously as the features, in the schema of section 1.9.
The Dataset Schema
Consider an illustrative audit of two files from the latch campaign, each holding the same number of state-action pairs at the same logging rate. File A comes from four authorized operators across several days, reports no permission-path clamps, and shows the same fingertip pads throughout. File B comes from one operator, logs 620 steps where the joint velocity limit clamped the requested motion, and spans a force-sensor re-zeroing midway through the session without a recalibration of the camera extrinsics. A validation script that checks tensor shapes and value ranges accepts both, yet File A can support a claim about the supported cell and File B mixes operator intent with permission-path action across an unrecorded change in the body. Nothing in the arrays tells the two apart, so the distinction has to live in a record of how each file was collected.
The demonstration provenance record is that record, binding each demonstration to the body and conditions that produced it and stating coverage as the dataset’s join against the scenario ledger, never as the range of values the collector happened to log. It consumes the handoff record of The Handoffs Matrix, keeping the operating conditions and clock domains under which the collecting machine’s budgets hold, and adds its own fields: the identity, calibration, and contact wear of the individual plant that produced each demonstration, the session and its timing, the four action taps, the collection denominators, a digest of the scenario ledger the data was collected against, and the ledger’s supported cells and known absences.
Every trajectory in the training corpus must identify its collection source, collector policy, and execution authority. When data originates from human teleoperation, the record logs the operator identifier, the interface hardware, and the transport timing, such as a 6-DoF leader arm sampled at \(100\text{ Hz}\) with \(0.8\text{ ms}\) of input jitter. When data originates from an autonomous or mixed controller, the record stores the exact policy checkpoint, the planner version, and the active enforcement configuration, including Cartesian bounding boxes, joint velocity limits, and contact-force clamping thresholds. Trajectories must also record the active retention rule, documenting whether an episode was committed automatically upon task completion, retained after an explicit operator review, or flagged for inclusion as an instructive failure recovery.
A physical machine cannot be identified by a generic model name. The data record specifies the individual plant through its extrinsic and intrinsic sensor calibrations, action contract, motor drive firmware revisions, and component wear state. Replaceable contact surfaces alter the governing dynamics. A gripper fitted with compliant \(30\text{ A}\) durometer silicone pads produces different contact compliance and friction from the same gripper fitted with worn \(60\text{ A}\) polyurethane pads under identical joint commands. Maintenance milestones, such as cable retensioning, drive belt replacements, gearbox swaps, and camera remounts, are logged directly into the machine state record. If an arm underwent joint recalibration between morning and afternoon collection shifts, the dataset must expose that boundary so downstream training does not attempt to resolve two conflicting kinematic mappings into a single policy.
Accounting for dataset volume requires explicit denominators across both time and trial counts, because a count of retained frames or successful episodes hides the yield of the process that produced them. The record therefore tracks scheduled cell time, active motion time, attempted episodes, retained successes, retained task failures, excluded recordings, and the count of independent physical task configurations, so that every attempt reconciles against a retained or an excluded episode and the time spent resetting the physical scene stays visible.
A single action column in a dataset conceals the filtering between request and response, so the record keeps all four action taps of 1.2 at every control step. Logged alongside boolean override flags, truncation markers, and numeric exclusion codes, the taps show where the permission path intervened and where tracking lag separated command from execution. Discarding the intermediate representations leaves the training pipeline unable to distinguish between an intentional human demonstration and an action heavily modified by the permission path.
Every physical dataset supports only part of the ODD, so the record catalogs both its supported cells and its known absences, and stating what is absent keeps an engineer from treating the absence of training failures as evidence of policy competence. Downstream verification treats any absence the record leaves uncataloged as unsafe. Filled in for the latch campaign, with each field group and the admission check it supports listed in table 4, the record reads:
- Denominators: 52 attempts per scheduled session at a net retention rate of 0.68, so 5,000 retained successful demonstrations cost 7,353 attempts over 141.4 h of scheduled cell time.
- Retention rule: an attempt whose contact force crosses the 15 N tripwire ends there and is kept as a labeled failure, outside the demonstration count, when its telemetry is valid; a recording that fails quality filtering is excluded with its exclusion code.
- Scenario ledger: plate offset (nominal, +3 mm, −3 mm) crossed with approach speed (0.03 m/s, 0.10 m/s), six cells, held in the record as a digest rather than as a copy.
supported_cells: the nominal plate at 0.03 m/s, the only cell the retained demonstrations occupy.known_absences: the other five cells, every displaced plate and every approach at 0.10 m/s.
| Provenance record | Evidence retained | Admission decision |
|---|---|---|
| Session | Collector, operator, policy version, sample rate, and timing jitter. | Check operator authorization and the session jitter budget. |
| Embodiment | Robot identity, geometry, gearing, backlash, and firmware. | Check reach, dynamics, and joint limits on the target plant. |
| Sensor calibration | Camera geometry, clock synchronization, and force/torque bias. | Check calibration age and timebase error. |
| Contact wear | End-effector material, hardness, friction, flex cycles, and maintenance history. | Revalidate contact data after material or compliance changes. |
| Four action taps | Requested, mapped, enforced, and measured values, each with its units and frame, plus override and clamp codes. | Compare like-typed, aligned values; exclude permission-path actions from labels. |
| Transition timing | Times of exposure midpoint, action issue, mapping, enforcement, response start, and response end, with clock IDs and synchronization errors. | Check causal order and the declared response window. |
| Collection denominators | Scheduled cell time, motion time, attempts, retained successes and failures, excluded recordings, and task configurations. | Reconcile attempts against retained and excluded episodes. |
| Scenario support | Scenario-ledger digest, the supported cells of the join, and known absences. | Refuse deployment queries outside the supported cells pending evidence. |
The next record in the chain is the policy manifest of The Policy Manifest, which cites this record by digest and may declare an ODD no larger than the supported cells listed here. Beyond training, downstream lifecycle stages consume the data record for distinct engineering decisions:
- The human and override provenance feeds directly into the supervisory auditing of Supervisory Intervention, where safety engineers evaluate demonstrator consistency, teleoperation latency tails, and the frequency of permission-path interventions to assess human reliability.
- The coverage boundaries, exclusion ledgers, and embodiment calibration histories pass forward to Deployment Release, where they form the empirical baseline for the safety case, verifying that the training distribution covers the ODD the release claims before the learned policy is granted execution authority on a physical machine.
By coupling physical movement trajectories with verifiable embodiment metadata, multi-stream action tuples, and support boundaries, the provenance schema makes visible any extrapolation into regimes the plant never visited, so downstream stages can refuse it.
↳ Downstream: The known absences and ODD boundary schema feed the formal safety case claims in Safety Cases and Claims.
Fallacies and Pitfalls
A clean dataset can omit states the deployed machine must handle. Coverage depends on the collecting policy (the human operator or automated script generating the motions), the physical rig, and the curation decisions that determine which experiences survive in the archive.
Fallacy: Massive sample volume compensates for lack of operational state-space coverage.
A surface-tracking task requires data across 500 combinations of force, friction, and contact angle, but the archive covers only 45. Another session adds \(72{,}000\) control steps by repeating nominal motions in populated state space cells. The pipeline’s ingestion report shows massive data growth with zero dropped packets, but physical coverage remains unchanged. Closely spaced samples describe the same limited physical experience more densely; they add no evidence about the missing friction or contact conditions. Dataset acceptance must therefore measure added coverage against the operational scenario ledger. More demonstrations earn their collection cost when they address a missing regime or an identified weakness, not merely when they enlarge the archive.
Pitfall: Assuming physical demonstration labels survive mechanical or contact hardware maintenance.
A maintenance technician replaces a gripper’s compliant fingertips with stiffer pads and adjusts its camera bracket. The archived demonstrations still pass structural schema checks (the data types and array shapes remain perfectly valid), but the physical reality has shifted. A known camera-pose adjustment can support coordinate reprojection for the vision data. However, that geometric correction cannot establish how the new, stiffer fingertips behave under contact, where they slip at lower force thresholds than the compliant ones did. Historical force and tactile measurements remain records of the old body, so their suitability as training targets requires revalidation. Dataset provenance must identify the mechanical and calibration state of the collection rig, allowing maintenance changes to trigger the appropriate review or recollection.
Pitfall: Filtering out aborted safety interventions to sanitize the demonstration dataset.
An insertion dataset contains 1000 attempts, including 180 non-nominal runs where contact forces reached or exceeded threshold limits (such as a \(15.0\text{ N}\) safety trip), leaving 820 nominal successes (an 82 percent task success yield). A curation script strips out the 180 aborted or recovery trials, leaving only the 820 flawless episodes whose contact forces stayed below \(8.0\text{ N}\). When deployment disturbances later push the robot’s contact force to \(9.5\text{ N}\), the policy finds itself immediately outside the retained experience, even though the hardware limit has not yet been reached. The superficially cleaner training set has erased the crucial moments of approaching the physical boundary and the recovery interventions that followed. Retaining and identifying those episodes preserves evidence about failure and recovery; deleting them because they are unsuccessful creates a misleading view of the task.
Pitfall: Merging demonstrations collected across kinematically asymmetric leader interfaces.
Two operators collect snap-fit demonstrations using different leader devices (the hand-held controllers used to puppet the robot). One interface permits only two wrist-roll positions (\(\theta_{\text{roll}} \in \{0.0^\circ, 90.0^\circ\}\)); the other mechanically limits roll to \([-15.0^\circ, +15.0^\circ]\). Merging their 4000 episodes increases volume, but neither interface supplies the required intermediate orientations between \(20.0^\circ\) and \(70.0^\circ\). A policy trained on the combined archive still lacks demonstrations for those perturbations. Aggregation cannot remove a blind spot shared by the collection interfaces. The team must map recorded orientations against task requirements and check whether the leader devices can express the missing motions before assigning more collection hours.
Fallacy: Physical data collection scales linearly with labor hours just like internet scraping scales with compute.
A team projects its bimanual dataset budget from an initial rate of 50 successful episodes per operator-hour. As collection continues, repeated physical impacts damage gearboxes, sensor drift invalidates some demonstrations (as uncalibrated joint angles gradually decouple from physical ground truth), and operator fatigue alters the timing of human motions. The rate of retained episodes falls even though the scheduled labor hours remain unchanged. A short nominal trial therefore overstates sustained collection throughput. Planning must include maintenance, calibration checks, rest, and rejected data alongside operator time. The useful production measure is accepted experience under documented conditions, since mechanically degraded or poorly aligned recordings can increase storage without increasing the dataset’s value.
Summary
Physical data collection converts mechanical work and hardware fatigue into the supported cells of a robot’s ODD. A session keeps less than a quarter of its scheduled time as validated demonstrations, and every attempt, retained or not, draws on the wrist’s rated fatigue life, so the archive is scarce before any question of its quality arises. Because the collecting policy chooses its own next observations (principle \(\ref{pri-vol4-endogenous-drift}\)), coverage is a fact about how the archive was gathered rather than about what it holds. Validation loss, drawn from the same occupancy, cannot expose that fact after the campaign ends, so it has to be written down while the gathering is still observable, as a join against a declared ledger and as provenance fields recorded at every step. The supported cells, not the size of the archive, bound what a policy trained on the data may be trusted to do, because a loss function cannot account for physical responses the robot never encountered. The demonstration provenance record makes that bound explicit by binding each demonstration to its plant, its action taps, and the cells it covers and misses.
Key Takeaways: Physical data collection and provenance
- Coverage counts configurations, not samples: Repeated nominal motions add rows without adding scenario cells, and adjacent control steps are near duplicates. A dataset’s reach is its join against a scenario ledger spanning the ODD, never its sample count or the extremes it happened to log.
- Validation loss cannot see a coverage gap: A held-out split shares the collector’s occupancy, so near-zero validation loss can coexist with most required cells empty. Only the join against an external ledger exposes the gap.
- Every action tap is a different signal: The request, the mapped setpoint, the enforced command, and the measured response differ wherever the permission path or the body intervenes. A dataset that keeps one of them trains on infeasible demands or on backlash mistaken for skill, and hides every intervention.
- Keep failures and takeovers, correctly attributed: Censoring aborted runs erases the states nearest the boundary. Learner commands before a takeover stay in the record for audit and expert relabeling but never as expert labels, and the recovery after it is a separate expert sequence.
- Data stays bound to the body that collected it: Camera geometry can be reprojected, but contact wear, compliance, and backlash are baked into force and tactile labels. Data from a changed or different body requires revalidation, and range normalization hides the mismatch rather than removing it.
- Absences travel with the dataset: Training cannot synthesize physics that collection omitted, so the provenance record lists its supported cells and its known absences. A policy trained on the data may later declare an operating domain no larger than those supported cells.
What’s Next: From demonstrations to policies
Footnotes
Causal Boundary: The causal boundary between computation and physics is established in The Causal Boundary. Because state evolution depends on motor actuation, policy actions perturb the distribution of future observations, which breaks the independent and identically distributed (i.i.d.) assumption of disembodied machine learning.↩︎
Behavioral Cloning Bound: The bound derived by Ross et al. (2011) is a worst case on expected cumulative cost, not a prediction of tracking error; actual drift depends on the plant, task, policy, and safeguards. Imitation Learning and Compounding Covariate Shift derives the bound and states its loss and bounded-cost assumptions.↩︎
Hardware Fatigue Life: Actuator thermal limits and gearbox wear, detailed in Actuator Transmission Limits and Thermal Duty Cycles, change joint friction over a campaign, so a pipeline that assumes stationary plant dynamics meets drift as the collection machine ages.↩︎
Proposal Boundary: The proposal boundary, defined in The Machine in Five Levels, separates learned proposals from the permission path that alone may admit them to the actuators (Safety Enforcement). Logging the action on both sides of it keeps permission-path interventions out of the demonstrator’s labels.↩︎
Bilateral Teleoperation Lineage: Raymond Goertz built the first bilateral master-slave telemanipulators at Argonne National Laboratory in the late 1940s and 1950s for handling radioactive materials. His linkages reflected force across a physical barrier, because operators without kinesthetic resistance jammed the remote mechanism.↩︎
Hardware Timer Interrupt: On real-time MCUs, encoder counts and current-shunt readings are latched by hardware timers and Direct Memory Access (DMA), which holds sampling jitter to microseconds and keeps timestamp jitter out of numerical differentiation for velocity and acceleration.↩︎
Corrective Aggregation Error Bound: DAgger (Ross et al. 2011) queries an expert on learner-visited states and aggregates those labeled states. Its linear-horizon result depends on the expert oracle, online-learning regret, bounded cost, and the measured loss of the resulting policy. See Imitation Learning and Compounding Covariate Shift.↩︎
Fault-Event Logging: This collection design preserves pre-trigger state, action authority, fault codes, and actuator-current samples across an override, with integrity checks on the stored record. Such telemetry helps distinguish policy drift from hardware or environment faults.↩︎
Host-Assigned Timestamps: Unlike PTP-synchronized Ethernet, buses without hardware timestamping, such as USB and classic CAN, leave the host to assign arrival times, which adds driver latency and scheduling jitter to every sample.↩︎

![Kinematically Isomorphic Leader-Follower Teleoperation Architecture across Heterogeneous Manipulators. Demonstration acquisition on physical robot platforms using low-cost, kinematically isomorphic leader controllers (canonical architecture [@wu2023gello]). (A) Bimanual teleoperation setup with twin 6-DoF leader arms mapped directly to dual heavy-payload follower manipulators during a compliant liquid-pouring task. (B) Single-arm leader-follower configuration controlling a torque-controlled manipulator with an overhead task camera. (C) Precision pick-and-place manipulation with passive gravity compensation and joint-angle direct drive. Because the leader device mirrors follower kinematics without requiring numerical inverse kinematics (IK), the interface eliminates IK singularities while preserving the operator's direct joint-level coordination. However, the operator's continuous-correction latency and communication delays still introduce phase lag into the recorded (st, at) demonstration tuples. Source: Adapted from Wu et al. [@wu2023gello], CC BY 4.0.](images/png/fig05_real_gello_teleop.png)
![Multi-Camera-to-Robot Extrinsic Calibration and Visual Verification in Diverse Demonstration Environments. Data collection across heterogeneous physical environments requires continuous validation of the extrinsic spatial transformation Tcam -> base between camera optical frames and the robot base frame. Six representative real-world collection scenes from the DROID platform depict autonomous extrinsic calibration verification using 3D kinematic reprojection (CtRNet-X / PyTorch3D) [@khazatsky2024droid]. Forward kinematics computed from optical joint encoders render the robot's URDF mesh directly over the camera video stream. If mechanical vibration or physical disturbances displace camera brackets by even a few millimeters, the visual discrepancy between the kinematic wireframe and the physical arm exposes extrinsic drift before corrupted trajectories enter the training pipeline. Documenting immutable calibration hashes in the dataset record prevents learning algorithms from attempting to fit contradictory spatial mappings within identical numerical coordinates. Source: Adapted from Khazatsky et al. [@khazatsky2024droid], CC BY 4.0.](images/jpg/fig05_real_camera_calibration.jpg)