Policy Training
Policy Training
Purpose
Why can a policy that succeeds in every simulated trial break the hardware on its first real contact?
Physical trial and error consumes hardware as well as computation. An exploratory action can overheat a motor coil, collapse battery voltage, or overload a transmission before a policy receives meaningful reward feedback. Consequently, physical policy synthesis shifts predominantly to offline demonstrations and high-throughput physics simulators. Yet each experience source leaves distinct, unmodeled inductive biases in the learned neural weights. Demonstrations cover only nominal visited trajectories, naive regression losses average distinct multimodal solutions into unviable interpolated actions, and simulation solvers reward non-physical contact shortcuts that hardware instantly punishes.
None of these discrepancies appear in offline training metrics, yet all of them accompany the weights onto the machine. Deploying learned policies safely requires selecting training regimes matched to dataset support, measuring where numerical solvers depart from real-world plant dynamics across friction, compliance, and latency, and rigorously bounding where the policy remains valid. Any unmodeled assumption left unchecked in software is ultimately settled at the causal boundary, where commands deliver electromagnetic energy and irreversible mechanical work that software can never recall.
Learning Objectives
- Compare training regimes by the dependency each leaves in the weights and the failure it produces on hardware
- Explain why a cloned policy’s small errors compound once it leaves the demonstrated states
- Explain how a squared-error loss turns two valid demonstrations into an action neither demonstrator took
- Diagnose a transfer failure by locating it in the dynamics, sensing, appearance, or timing channel of the sim-to-real gap
- Evaluate a real-machine fine-tuning run by the wear it incurs and the prior evidence it invalidates
- Construct a policy manifest whose declared operating domain excludes every state the data does not support or the simulator does not model
Policy Synthesis
Two learned controllers can succeed at the same task under nominal conditions and still fail on the same machine for different reasons. The cage-door latch task of Physical Data has two such controllers. The first latch policy, latch-bc-01, was cloned from the latch campaign’s teleoperated door-opening demonstrations. The second, latch-sim-02, started from latch-bc-01 and was refined by reinforcement learning in a physics simulator.
At the nominal strike plate and the guarded approach speed of 0.03 m/s (illustrative; see the Reader Guide), both policies open the door. Their failure modes diverge once conditions leave that point. When the door has sagged and its strike plate sits a few millimeters off nominal, the cloned policy’s first correction lands in a state its demonstrations never visited, because the demonstration campaign recorded no displaced plate. Each later command, issued at the chosen 20 ms setpoint period, compounds the error until the gripper binds on the edge of the plate. The simulated policy fails on the nominal plate. Its simulator modeled contact as a soft penalty spring, so the optimizer found that approaching at 0.10 m/s costs nothing. On the real plate, contact force rises in proportion to approach speed and time, and at that speed it reaches the latch’s 100 N limit 2.5 ms after touch, barely longer than the contact loop’s 2 ms response (the speed, limit, and response time are illustrative). That leaves too little time to shed the approach speed, so the plunge overloads the latch. The data collection regime defined what the policy saw, while the simulator shaped its assumptions about the physical world.
The two latch failures share one structure. Each training regime left a dependency in the weights. The cloned policy depended on the states its demonstrations happened to cover and the simulated policy on a contact model softer than the plate, and in each case a deployment condition that violated the dependency produced a failure no training score had shown. The practical choice is therefore which experience source teaches the needed recovery without hiding actuator dynamics, and what physical evidence must precede actuator authority. The answer travels with the checkpoint as a record of the states it may enter. Whatever the regime, the checkpoint only proposes chunks to the permission path (The Machine in Five Levels), and training that never modeled that path yields a policy whose torque spikes the permission path must refuse, or whose proposals trip its limits continuously.
Policy Synthesis Regimes
The latch task could draw on four sources of experience: the teleoperated door openings of Physical Data, a log of imperfect or scripted door runs, a simulator of the cage door, and the arm itself. Each buys a different capability at a different physical price, and table 1 sets them side by side with the failure each is known for and the check that intercepts it.
| Policy Synthesis Regime | Core Purchased Capability | Data & Compute Budget | Sample Efficiency & Scale | Hardware Safety & Wear Risk | Dominant Failure Mode & Interceptor |
|---|---|---|---|---|---|
| Behavioral Cloning (BC) | Direct imitation of demonstrated trajectories without reward engineering or system identification | \(10^2\)–\(10^3\) teleop trajectories (\(\sim 10\)–\(50\text{ h}\) data); \(<50\text{ GPU-h}\) compute | Supervised \(\mathcal{O}(N)\) over expert demonstrations; fixed \(1\times\) real-time data throughput | Minimal exploration risk during logging; zero hardware exploration during optimization | Quadratic compounding error (\(\mathcal{O}(T^2 \epsilon)\)) via covariate shift. Interceptor: Kernel/Mahalanobis density monitors, convex hull boundary clamps |
| Offline / Batch RL | Value-guided policy optimization over fixed, heterogeneous sub-optimal logs; stitches multi-trajectory transitions | \(10^5\)–\(10^7\) logged transitions with state-action coverage; \(100\)–\(500\text{ GPU-h}\) compute | Offline \(\mathcal{O}(N)\) over static datasets; sample-efficient but demands high state-action overlap | Zero active exploration hazards during offline training; logging requires caged rigs or safe scripted controllers | Out-of-distribution action extrapolation and Q-value overestimation. Interceptor: Conservative Q-learning (CQL) penalty bounds, action support filters |
| Sim-to-Real RL | Discovery of dynamic, non-prehensile manipulation and agile locomotion via billions of unconstrained rollouts | Zero physical data; \(10^7\)–\(10^9\) simulated transitions across \(10^3\)–\(10^4\) GPU environments; \(500\)–\(5000\text{ GPU-h}\) | Low algorithmic sample efficiency (\(10^7\)–\(10^9\) steps), offset by \(10^3\times\)–\(10^5\times\) simulation acceleration | Zero hardware wear or collision damage during training; severe deployment risk if plant dynamics diverge | Exploitation of numerical solver artifacts (contact interpenetration, zero latency). Interceptor: Paired-trace discrepancy budgets, force-derivative (\(dF/dt\)) clamps |
| Residual RL & Real Fine-Tuning | Direct policy correction against physical contact compliance, gearbox hysteresis, and true sensor noise | \(10^3\)–\(10^5\) physical steps (\(0.5\)–\(10\text{ h}\) wall-clock time); \(<20\text{ GPU-h}\) compute | High sample efficiency required (SAC/TD3 or residual delta-action formulation); strictly \(1\times\) real-time | High physical exposure; repetitive impact shocks degrade gearboxes; exploratory current dither elevates winding temperature | Catastrophic forgetting of nominal transit stability, local policy collapse. Interceptor: Parameter drift bounds (\(\Vert\Delta \theta\Vert_2 \le \delta\)), frozen regression suites, thermal limiters |
Of the four, behavioral cloning is the cheapest route to a working policy and the one whose failure mode is best understood. It frames policy synthesis as supervised regression or classification over a collected dataset of expert state-action pairs \(\mathcal{D} = \{(\mathbf{o}_t, \mathbf{a}_t)\}\). Following the action taps of Which Action Becomes the Label, a demonstration dataset pairs observations with either the operator’s mapped setpoints or the setpoints the permission path admitted: \[\mathcal{D}_{\text{intent}} = \left\{ \left(\mathbf{o}_t, \mathbf{a}_{\text{cmd}, t:t+H-1}\right) \right\} \quad \text{versus} \quad \mathcal{D}_{\text{verified}} = \left\{ \left(\mathbf{o}_t, \mathbf{a}_{\text{enf}, t:t+H-1}\right) \right\}\] Cloning the mapped setpoints (\(\mathbf{a}_{\text{cmd}}\)), the operator’s intent in the robot’s own coordinates before the permission path acts, trains the model on demonstration intent, whereas cloning the setpoints the permission path admitted (\(\mathbf{a}_{\text{enf}}\)) bakes its barrier deceleration profiles directly into the network weights, so the policy inherits decelerations tuned to that particular path’s limits. In its standard formulation, the optimization objective minimizes a multi-step prediction loss over continuous action sequences drawn from demonstration trajectories (see Mode averaging and the decoders that avoid it for classical regression formulations).1 Because this loss penalizes only the discrepancy between the commanded action and the recorded demonstration at observed states, the training objective does not evaluate task completion, terminal cost, or the stability of the closed-loop dynamical system. It fits actions to observations only.
↰ Prerequisite: The four-stream action tap tuple separating intent from enforcement is formulated in Which Action Becomes the Label.
The multimodality breakdown and mode averaging
When human demonstrators perform a physical task, their strategy is frequently multimodal. Consider an obstacle avoidance scenario where a mobile base or robotic arm must navigate around a central pillar. A human demonstrator may choose to pass around the pillar to the left in half of the demonstrations, or to the right in the remaining half. The true demonstration data contains distinct valid choices, but zero examples of driving straight down the middle.
A deterministic point-regression behavioral-cloning policy cannot represent both choices for the same observation. Under squared-error loss its optimal output is the conditional mean (for the derivation, see Mode averaging and the decoders that avoid it),2 here the midpoint between the “steer left” and “steer right” commands. The policy outputs zero lateral velocity, commanding the machine to drive directly into the obstacle, an action that never occurred in the demonstration dataset.
A process plant shows that mode averaging needs no geometry. Suppose a flow controller’s demonstrations alternate between two illustrative valve regimes, a bypass regime at 10 percent opening and a main-feed regime at 90 percent, in equal numbers. Squared-error cloning outputs their mean, 50 percent opening, a flow that neither regime commands and for which the plant was never calibrated. That intermediate flow can cavitate the supply pump and miss the cooling deadline each regime was tuned to meet. No obstacle is involved, so no geometric check on the proposal catches it; only a check against the demonstrated regimes does.
A mode boundary also produces a timing failure, temporal inconsistency in a high-frequency single-step controller. In an illustrative \(50\text{--}100\text{ Hz}\) controller, sensor noise moves consecutive observations across that boundary, and a policy mapping only \(o_t \to a_t\) alternates between \(+10\text{ N}\cdot\text{m}\) and \(-10\text{ N}\cdot\text{m}\) requests. Whether the alternation excites resonance, wear, or winding heat depends on the drive and the load, so it is a failure mode to test for rather than a property of every single-step policy.
Mode averaging is a failure at a single observation. Behavioral cloning has a second failure that appears only when the policy runs in closed loop. Its own errors move the machine off the states its demonstrations cover, where the loss supplied no training signal, which is how latch-bc-01 binds on the sagged door. Under principle \(\ref{pri-vol4-endogenous-drift}\), the policy’s own actions choose which states it visits next, so a fixed archive cannot anticipate them, and for expert-only cloning the worst-case cost grows with the square of the horizon (equation, developed in Endogenous Experience). A low validation loss on held-out demonstrations therefore measures accuracy only on states the demonstrator visited; it says nothing about how the policy returns once displaced. Corrective aggregation labels the learner’s own states and, under its expert-labeling assumptions, reduces that growth to linear (Intervention and Recovery Data), but on a physical machine those states are reached by running an unverified policy, which can strike a fixture or damage a transmission before a supervisor intervenes.
Definition 1.1: Compounding rollout error
Compounding rollout error is the quadratic accumulation of closed-loop tracking drift incurred when an open-loop supervised policy executes sequentially on a physical dynamical plant, bounded in expectation by: \[J(\pi) - J(\pi^*) \le \mathcal{O}(T^2 \epsilon)\] where \(T\) is the task execution horizon and \(\epsilon = \mathbb{E}_{\mathbf{s} \sim d_{\pi^*}} [\mathbb{I}(\pi(\mathbf{s}) \ne \pi^*(\mathbf{s}))]\) is the single-step policy error rate on the demonstration distribution \(d_{\pi^*}\).
- Significance: Proves that minimizing single-step prediction error on offline demonstrations is fundamentally insufficient for closed-loop stability; any nonzero error rate \(\epsilon > 0\) eventually drives physical state outside demonstrated support.
- Distinction: Unlike static supervised learning where validation error scales linearly with sample count, compounding rollout error couples policy mistakes with environmental physics, turning small initial trajectory deviations into irreversible state divergence.
- Common pitfall: Assuming low offline validation loss on held-out demonstration splits translates to high closed-loop task success, ignoring that offline datasets contain zero examples of returning to nominal support once perturbed.
What action chunking changes—and what it leaves unresolved
Action chunking (\(\ref{dfn-brain-action-chunking}\)) can improve temporal coherence and reduce proposal frequency in generative architectures such as diffusion policies and Action Chunking with Transformers (ACT). It can also address the alternation in the mode-boundary example above. A chunking policy that predicts \(H\) future actions \(\mathbf{a}_{t:t+H-1}\) at once improves execution through three mechanisms:
- Temporal and Dynamic Consistency: Rather than generating disjoint single-step predictions that chatter, the model predicts an entire coordinated trajectory.
- Maintaining a Chosen Mode: Single-step Markov policies are prone to causal confusion, mistaking momentary spurious correlations for task intent. Conditioning on trajectory history and proposing a chunk helps preserve a chosen mode across adjacent actions.
- Receding-Horizon Replanning and Temporal Ensembling: The deployed policy queries a fresh chunk every few steps and blends the overlapping predictions (Supply, Freshness, and the Memory Wall), which permits periodic feedback while retaining a multi-step proposal.
None of these mechanisms changes the behavioral-cloning support bound \(\mathcal{O}(T^2 \epsilon)\). As demonstrated theoretically and empirically in figure 1, single-step behavior cloning experiences a steep drop in task survival as rollout horizon \(T\) grows, falling into the catastrophic intervention envelope (\(<18\text{ percent}\) survival) within 250 control steps at \(50\text{ Hz}\). While temporal architectures like BC-RNN delay the onset of divergence, their quadratic error compounding remains asymptotically identical. In contrast, on-policy interactive aggregation (DAgger (Ross et al. 2011)) bounds error growth linearly to \(\mathcal{O}(T \epsilon)\) by collecting expert feedback on learner-visited states. Without interactive supervisors, modern generative policies counteract compounding drift through action chunking (\(K=16\)), which contracts the effective execution horizon to \(T_{\text{eff}} = T/K\) (ACT (Zhao et al. 2023)), and continuous receding-horizon diffusion replanning (Chi et al. 2024), maintaining over \(80\text{ percent}\) task survival across extended 500-step manipulation horizons.
Increasing dataset size, model capacity, or tuning can reduce the one-step error rate on evaluated observations and therefore lower the chance of a first policy error during a fixed horizon. That improvement does not characterize what happens after an error. An external disturbance, uncalibrated payload, or voltage sag can also move the machine into weakly supported states without a policy disagreement. Recovery data, measured dynamics, and independent runtime limits determine whether either event remains within the physical envelope.
Drift off the demonstrated states raises no alarm of its own, so detecting it before physical contact cannot rely on the offline training loss. It requires runtime telemetry that tracks state space support directly, measuring Mahalanobis distance in feature space, estimating model epistemic uncertainty, or enforcing an explicit envelope monitor on the permission path that trips when state trajectories leave the convex hull of the demonstration dataset.
What the decoder choice costs at run time
A point-regression decoder averages distinct valid actions into an invalid command, and several decoder families avoid the mean: mixture density networks, implicit energy-based cloning, conditional diffusion and flow matching, latent-variable chunking such as ACT, and discrete action tokens in vision-language-action models. Mode averaging and the decoders that avoid it and Continuous trajectory diffusion and flow matching develop their losses and sampling procedures. None of them supplies missing recovery states. What the choice changes for the machine is its cost at run time, because the decoder fixed at training time sets how many network evaluations each chunk requires.
The step count is the cost that reaches the control deadline. A diffusion decoder produces each chunk by repeated denoising, and every step is a full evaluation of the noise-prediction action head, so inference time scales with the number of steps (the refinement term of the chunk latency in Supply, Freshness, and the Memory Wall). Under the step counts and per-step times assumed in Continuous trajectory diffusion and flow matching, a many-step denoising diffusion implicit model (DDIM) sampler overruns the 20 ms setpoint period when it runs synchronously, while a few-step flow-matching sampler fits within it, provided those few steps meet the task’s action-quality criterion. A decoder that cannot meet the period synchronously runs asynchronously. The application processor refreshes a chunk buffer at a lower rate, and the permission path reads the newest chunk at its own rate, checks its freshness, and falls back when it expires, so actuation never waits on the proposer. The machine’s own chunk policy is this asynchronous case. It renews its chunk every 50 ms, so its inference deadline is that chunk period, and the 20 ms setpoint period only paces playback of the setpoints inside the chunk.
As detailed in Diffusion policy architecture (Continuous trajectory diffusion and flow matching), the policy conditions on an observation history of images and joint states to iteratively denoise an action chunk \(\mathbf{A}^0 \in \mathbb{R}^{K \times D_a}\) from pure Gaussian noise. Rather than executing the entire prediction horizon, the machine dispatches an execution horizon to the real-time microcontroller before receding-horizon re-planning occurs. The formal dual-rate pipeline—coupling asynchronous reverse-diffusion on the host with deterministic spline interpolation, rate-limiting, and safety clamping on the microcontroller—is developed in (algo-diffusion-policy-trajectory?) (Continuous trajectory diffusion and flow matching).
↳ Downstream: Hardware placement of multi-rate action chunk buffers across the application processor and the safety microcontroller is engineered in Two Paths on One Die.
Offline reinforcement learning and action distributional shift
When physical demonstration datasets \(\mathcal{D} = \{(s_t, a_t, r_t, s_{t+1})\}\) contain heterogeneous or sub-optimal trajectories (e.g., mixtures of teleoperation, scripted routines, and exploratory human play), behavioral cloning struggles because it treats all demonstrations with equal supervised weight. In classical reinforcement learning formulated on Markov decision processes (MDPs) (Sutton and Barto 2018), an agent optimizes a policy to maximize expected discounted cumulative returns \(J = \mathbb{E}\left[\sum_{t=0}^\infty \gamma^t r(s_t, a_t)\right]\). Offline reinforcement learning (Batch RL) aims to optimize a policy directly from logged transition archives without active hardware exploration, using reward signals to stitch together high-performing composite trajectories from disjoint, imperfect demonstrations (Kumar et al. 2020).
However, applying standard Q-learning to static offline datasets triggers a failure mode known as out-of-distribution (OOD) action overestimation. In standard reinforcement learning, the algorithm estimates the value of future states by asking its value network to predict the highest possible return it could achieve from the next step (for formal Bellman equations, see Value Functions and Affordance Grounding). That maximum is taken over the entire continuous action space. For actions that were never executed in the logging dataset, the network has received zero training signal and often outputs erroneously large value estimates due to generalization variance, predicting a high return for an untried physical maneuver. The algorithm greedily selects these spurious peaks, propagating over-optimistic value errors backwards. When deployed, the policy selects these ungrounded actions, and it can command motions that stall a joint or strike a boundary.
To reduce value overestimation without querying the physical environment, Conservative Q-learning (CQL) adds a regularizer to the offline Bellman objective.3 It penalizes high values for actions sampled by the learned policy relative to actions in the logged data. Limited state-action coverage and approximation error remain, so the permission path still checks what the resulting policy proposes.
Cloning, chunking, and offline value learning all remain inside the support of the logged data, and Conservative Q-learning makes an offline policy more cautious without widening that support. When the logged transitions lack the high-performance trajectories or recovery maneuvers that dynamic manipulation or agile locomotion needs, the system must collect new experience through closed-loop trial.
Learning by Trial
The cloned latch policy failed on the displaced plate because no demonstration had ever shown it one. Learning by trial closes that blind spot by letting the policy generate its own state-action trajectories in closed loop. Algorithms such as Q-learning or policy gradients adjust the policy from the rewards its own actions earn, so an exploratory action that pushes the gripper off its nominal path forces the optimizer to evaluate recovery from exactly that perturbed state, including the near-misses and unstable contacts an operator never exhibits. By sampling transitions from the state distribution the current weights induce, trial-based learning trains on the states the policy actually reaches, which removes the covariate shift behind the quadratic worst-case growth of cloning error.
Closed-loop exploration on hardware is billed first in calendar time, which depends on how many trials the algorithm needs, how long each trial lasts, how long a reset takes, and how many machines run in parallel. Consider the door-opening task itself, in which the arm regulates contact forces over an active horizon of \(t_{\text{rollout}}\) = 8 s, the autonomous contact phase of an opening rather than a full teleoperated attempt. Reinforcement learning on high-dimensional sensory inputs can require on the order of \(N = 10^6\) trials to discover coordinated force-velocity profiles, the scale of the grasping campaign in figure 2. On a single physical machine with \(P = 1\), the raw execution time alone accounts for 8.0 × 10⁶ s, roughly 92.6 days of uninterrupted round-the-clock operation before accounting for any reset overhead.
Factoring in the reset phase reveals the true temporal bottleneck of physical experimentation. If a dedicated automated fixture, faster than the collection rig’s re-latch (Teleoperation Demonstration Costs), re-latches the door and returns the arm to its starting pose in \(t_{\text{reset,auto}}\) = 4 s, the total collection time for one million trials expands to 1.2 × 10⁷ s, or 138.9 days. When exploratory trajectories jam the latch cam or trigger joint overcurrent trips, automated resetting becomes impossible, requiring a human operator to free the latch, re-seat the door, and re-index the arm. If every trial instead needs a manual reset averaging \(t_{\text{reset,manual}}\) = 30 s, the wall-clock requirement grows to 439.8 days. Scaling physical parallelism to clusters of \(P = 14\) synchronized rigs (figure 2) compresses calendar duration to several weeks (Levine et al. 2018), but it multiplies the physical footprint, capital investment, electrical power delivery, and human supervisory labor proportionally.
Time is often not the binding constraint; wear is. Each door opening loads each arm joint through about \(n_{\text{rev}} = 4\) bidirectional load reversals, so \(N = 10^6\) trials accumulate 4.0 × 10⁶ of them in the drivetrain.4 The same million contact events wear the latch cam and strike plate, crush the gripper’s compliant pads, and cycle the arm’s wiring toward fatigue fracture.
In a purely computational environment, an exploratory branch that goes numerically unstable costs nothing beyond a discarded buffer. On the running machine, exploration is irreversible. Unconstrained exploratory torque drives the gripper into rigid fixtures, bending load-cell flexures, stripping harmonic drive teeth, and heating the motor stators past their \(155^\circ\text{C}\) Class F insulation threshold.5
↰ Prerequisite: Motor winding thermal limits and torque-speed derating curves are derived in Thermal Duty Cycles.
Furthermore, physical degradation alters the plant during the training process itself. As contact surfaces burnish, gear backlash expands (increasing the mechanical slop or dead-zone between gear teeth), and structural fasteners loosen under repetitive vibration, the physical transition distribution \(p(s_{t+1} \mid s_t, a_t)\) at trial \(n = 900{,}000\) differs measurably from the pristine dynamics at trial \(n = 1\). The machine being controlled mutates under the physical cost of collecting its own training data.
The reward function adds a second fragility. In trial-based learning the scalar reward is the interface specification between the engineer’s intent and the optimizer, and any constraint it does not measure and penalize is a free degree of freedom the optimizer will exploit. A reward for rapid setpoint tracking with no penalty on jerk or winding temperature converges toward bang-bang switching between maximum forward and reverse effort, which vibrates brackets and overheats drive electronics. With contact force unmeasured, the policy may learn to brace a joint at maximum current against a hard stop because doing so removes tracking jitter. The network satisfies the formula presented to it, not the designer’s unstated common sense.
For low-dimensional systems with trivial resets, negligible wear, and wide margins, reinforcement learning directly on hardware remains viable. Simulation becomes mandatory where the required sample count exceeds the time, budget, structural life, or risk tolerance of the apparatus, and the \(10^6\) to \(10^9\) transitions a complex motor policy needs move into a physics engine run in parallel across a computing cluster.
The substitution changes the contract between training and reality. A policy trained in a simulator bends no linkage during exploration, but it inherits the omissions of the reward and the simulator’s model of mechanics, and every unmodeled friction discontinuity, idealized contact impulse, and uncharacterized latency becomes an assumption it relies on. On the physical machine those inaccuracies return as unmodeled disturbances and tracking failures.
The Simulator as a System Component
Treating the simulator as an external development tool is an architectural error. A compiler produces machine instructions from source code, but the resulting binary does not retain a hidden dependency on the compiler’s internal memory allocator once executing on standard hardware. A simulator behaves differently. It is an upstream policy-production component whose physical assumptions, numerical approximations, and boundary simplifications become permanently embedded in the policy parameters. When the trained neural network is copied onto the target machine, the simulator is absent from the execution path, yet the policy acts as though every unmodeled dynamic in that simulator remains true. The simulator is an active component of the system architecture, with its own specification, tolerances, and failure modes.
Definition 1.2: Sim-to-real reality gap vector
Sim-to-real reality gap vector \(\boldsymbol{\Delta}_{\text{sim2real}} = [\Delta_{\text{dyn}}, \Delta_{\text{sens}}, \Delta_{\text{vis}}, \Delta_{\text{lat}}]\) is the coupled, multi-channel decomposition of physical discrepancies between simulated and hardware environments spanning contact mechanics (\(\Delta_{\text{dyn}}\)), sensor transduction and noise (\(\Delta_{\text{sens}}\)), visual appearance (\(\Delta_{\text{vis}}\)), and asynchronous transport latency and jitter (\(\Delta_{\text{lat}}\)).
- Significance: Treating the “reality gap” as a scalar difficulty metric or pure visual domain mismatch causes engineers to overlook dynamical and temporal divergences. Discrepancies in actuator back-EMF, gearbox compliance, or bus jitter deplete phase margin before the policy completes its first physical control step.
- Distinction: An inaccurate parameter value (such as a 10 percent error in link mass) preserves the causal differential equations of motion, whereas an omitted mechanism (such as zero-delay communication or unmodeled Coulomb stick-slip) entirely severs physical feedback paths, leading policies to exploit non-physical simulator artifacts.
- Common pitfall: Relying exclusively on extreme domain randomization to bridge large gaps. Randomizing parameters across unphysical ranges dilutes policy capacity, resulting in overly conservative, sluggish behaviors that fail to achieve precision tasks on real hardware.
Like any hardware subsystem, the simulator requires a formal interface contract, stated in quantities that can be measured and budgeted. This contract specifies observation dimensions, action bounds, simulation time step \(\Delta t\), transport latency between command issuance and physical actuation, actuator torque and current limits, sensor noise distributions, reset distributions, and termination conditions. If the simulator specifies an observation loop executing at \(100\text{ Hz}\) with an assumed zero-delay sensor channel and Gaussian noise with variance \(\sigma^2 = 1.0 \times 10^{-4}\text{ (rad/s)}^2\), the policy learns feedback gains tuned to those precise signal statistics, and an unmodeled \(15\text{ ms}\) bus transport delay on the physical machine invalidates them. Dactyl’s twin (figure 3) could not match its physical hand exactly, so its simulator was calibrated against the real hand and then randomized around the calibrated values, and the policy met the unmatched remainder as variation it had already trained against (Andrychowicz et al. 2020).
The interface contract matters more as simulation’s role grows. Figure 4 shows the rollout throughput some parallel accelerators reach, and for policies requiring many trials, simulation can become the primary source of training experience.
War Story 1.1: Sim-to-real hardware breakage (2018)
Mechanism: The simulated hand applied torque directly at its joints, while the physical hand is tendon-actuated and has substantial backlash. The authors smoothed every action with an exponential moving average, in simulation and on the robot, to avoid abrupt changes in the action signal that could harm the physical hand.
Impact: The authors report robot breakages during the physical experiments, most often at the wrist’s vertical joint, probably because it carries the largest load. Each repair took time and often changed aspects of the system, so the physical results were collected at different times, and the authors count hardware breakage among the key challenges of the work.
Response: The team also trained a second policy with the wrist pitch joint locked, and it transferred better to the physical hand.
Systems lesson: A simulator that neither wears nor breaks cannot price hardware exposure into the reward. The limits that protected the hand, an action filter and a locked joint, were imposed outside the learned policy, and every repair changed the plant on which the policy was being evaluated.
When a simulator uses an inaccurate rotor inertia, assigning \(J = 2.0 \times 10^{-3}\text{ kg}\cdot\text{m}^2\) instead of a true \(2.5 \times 10^{-3}\text{ kg}\cdot\text{m}^2\), the equations of motion still compute an acceleration-dependent inertial torque. The parameter error alters the magnitude of the response, but the causal structure remains intact. When a simulator omits an entire physical mechanism, such as drivetrain backlash, structural compliance, or inter-channel communication jitter, it does not merely set the corresponding parameter to zero. It provides no causal mechanism at all. An omitted delay or compliance represents an unmodeled physical degree of freedom.6 A simulator that omits joint compliance permits instantaneous torque reversals, which a real drivetrain converts into gear-tooth loading and excited structural resonances.
The most consequential omissions come from how simulators handle contact. Rigid contact is non-smooth. The normal force is zero while bodies separate and takes whatever compressive value prevents interpenetration once they touch, and velocity jumps at impact. Friction obeys a Coulomb cone, and rigid bodies under Coulomb friction can have no unique forward solution (the Painlevé paradox). Physics engines take one of two numerical paths (Rigid Body Contact Mechanics). Complementarity solvers keep strict non-penetration at a cost that grows steeply with contact count and can stall when geometries pinch. Penalty and convex relaxations replace the hard constraint with a virtual spring-damper, which converges quickly at high frame rates but lets solid bodies interpenetrate slightly. The latch simulator takes the second path, and its softness is what latch-sim-02 learned to use.
The sim-to-real contact exploit
Reinforcement learning optimizes against the physics the solver presents, not the physics of the plant. In simulators that implement soft contact relaxations or penalty springs, the optimizer discovers that “sinking” into the surface by \(1.0\text{--}3.0\text{ mm}\) (a negative signed gap \(\phi(q) < 0\) between the contacting surfaces) provides two unphysical affordances:
- The artificial spring-damper reaction acts as an instantaneous energy sink, arresting downward momentum in a single integration step without bouncing or plastic deformation.
- Interpenetration increases the effective contact surface area in the solver, generating unphysical lateral traction forces that prevent slipping during aggressive maneuvers.
The arm drives the latch into its strike plate with a robot-side effective mass of 12 kg at the tool center point. Suppose the simulator’s penalty contact lets the gripper sink an illustrative 2.4 mm into the plate at the fast approach of 0.10 m/s. Equating the approach kinetic energy with the spring energy at peak compression gives a peak force \(F_{\text{peak}} = v\sqrt{k\,m_{\text{eff}}}\), which for the simulated spring reduces to \(m_{\text{eff}}v^2/\delta_{\text{pen}}\), or 50 N, under the latch’s 100 N limit. The reinforcement learning optimizer records zero damage penalty, high task speed, and maximum reward, and it learns to plunge into the plate without braking.
When the policy moves to the machine, the real strike plate provides no such compliance. Its contact stiffness is an illustrative 4.0 × 10⁵ N/m, and the same plunge peaks at 219.1 N, 4.4× the simulated force and 2.2× the latch limit. Whether the permission path can prevent that peak is a question of time. Contact force follows \(F = k\,\Delta x\) (Actuator Transmission Limits), and because the plate compresses at the approach speed, \(\Delta x = vt\) and the force rises as \(F = kvt\) from the moment of touch. At the fast approach the latch limit arrives 2.5 ms after touch, barely more than the 2 ms the contact loop needs to respond, while at the guarded approach of 0.03 m/s it arrives after 8.3 ms. A response that begins with 0.5 ms to spare cannot decelerate the arm before the force crosses the limit, and the plunge runs on toward that peak. Prevention must therefore act before contact. An independently enforced approach-speed envelope, such as \(v_{\text{max}}(z) \le \sqrt{2 a_{\text{brake}} z}\) with \(z\) the remaining distance to the plate and \(a_{\text{brake}}\) the validated deceleration, limits the speed at touch, and a force tripwire can limit subsequent loading only if its sensing, brake response, and remaining compression have been validated. The contact deadline, not the training reward, is what confines the latch policy’s declared ODD to the guarded approach (section 1.8).
Simulator fidelity for contact is therefore budgeted in observable mechanical quantities, namely contact onset timing, peak normal force, transferred impulse, slip distance, and release delay, not in the internals of the solver. The mismatch can also take the opposite sign. A penalty engine that resolves a collision within one integration step concentrates the momentum transfer into a sharp spike, while pad deformation and structural compliance on hardware spread it over many milliseconds. The integrated impulse can then agree closely while the peak force differs severalfold, and a policy that learned to detect contact onset from the simulated spike never sees that spike on the machine, so it keeps pushing until the drive reaches its current limit.
Systems Perspective 1.1: The contact fidelity invariant
Which omissions matter depends on the task. A policy that moves only at high continuous velocity may tolerate a coarse friction model because it never lingers near the stiction threshold, while a policy that regulates force while holding a static contact will enter uncontrolled limit-cycle oscillations if static friction, stick-slip transitions, and actuator deadbands are omitted from the simulator.
⇄ Contrast: The 2.4 mm of penalty-spring penetration is compliance the solver invents, whereas the backlash and torsional compliance of the real drivetrain, analyzed in Actuator Transmission Limits, are compliance a rigid-body solver omits.
Establishing the validity of a simulator component requires paired probing. The engineer applies identical open-loop command sequences to both the simulator and the physical hardware, measuring the difference between the resulting state trajectories against explicit error budgets. For the latch, that means driving the same approach command into the simulated and the real plate and comparing the contact quantities above, so the discrepancy is budgeted in milliseconds and newtons. When a measured difference exceeds the tolerance for the operating envelope, the simulator is defective for that training task.
Because the structural assumptions of the simulator are frozen into the neural network weights during optimization, the simulator specification forms an immutable element of the policy’s configuration and identity. Every trained policy artifact must document its complete simulator provenance, the simulation_provenance field of the policy manifest (section 1.8). Deploying a policy without it is like shipping firmware without recording the target architecture it was compiled for. Auditing this provenance takes more than an aggregate pass rate; it requires decomposing the reality gap into its dynamics, sensing, appearance, and timing channels, which section 1.5 sizes one at a time.
What the Gap Is Made Of
The latch simulator departs from the real cage door in two unrelated ways, a penalty spring far softer than the strike plate and a command path with no transport lag, and a single distance between simulated and real trajectories would blend them into a number that names neither. Errors in different physical channels carry different units, propagate through different paths in the closed loop, and produce different failures, so the mismatch is a vector of coupled components spanning dynamics, sensing, appearance, and timing. Sizing it is an accounting of physical errors against the task’s tolerances, from paired command-response traces recorded under identical excitations, and each channel demands its own interceptor. Table 2 lists the interceptors for dynamics, which it splits into actuator and contact rows, and for sensing and timing; appearance is sized in the prose below and in panel (d) of figure 5.
Dynamics mismatch measures differences between commanded actuator effort and the resulting physical motion or contact force across the operational envelope. A simulator that models a rigid link with zero backlash and instantaneous contact stiffness predicts a faster force rise than a drivetrain whose gearbox compliance and drive delay add phase lag, and a policy that keyed contact onset to the simulated rise keeps pushing on the machine, the opposite-sign contact mismatch of section 1.4.
Sensing mismatch measures discrepancies between the true physical state and the observation vector delivered to the policy network. A simulated range sensor is clean and uniformly quantized; a physical one carries calibration bias, noise, thermal drift, and dropout bursts during specular reflections. A bias moves the point at which the controller switches from approach to force regulation, and a dropout burst leaves the policy blind to the surface during descent.
Appearance mismatch is not a mean squared error computed over raw camera pixels. Image differences between photorealistic rendering and physical camera frames are dominated by high-frequency background textures and minor sensor noise patterns that may have no effect on policy action. Appearance mismatch is sized by measuring the deviation in policy-relevant perceptual outputs, such as estimated surface normals, object keypoints, or spatial feature activations, under matched physical geometry, illumination, materials, and camera exposure settings.
Timing mismatch governs the temporal alignment of observation capture, policy evaluation, and actuation. Identical state transitions occurring at different moments in physical time do not produce equivalent closed-loop dynamics. A simulator that executes observation, inference, and actuation instantaneously at the start of each discrete step omits camera exposure, bus serialization, inference tail latency, and motor bus transport, and it omits their cycle-to-cycle jitter. Every millisecond of that omitted delay is distance the mechanism travels unobserved before corrective action begins, and it erodes the damping and phase margin of the closed loop.
Figure 5 shows all four channels diverging on an illustrative linear drive, separate from the mobile manipulator, that approaches a rigid surface at 0.20 m/s and regulates contact force to 10 N. An unmodeled \(8\text{ ms}\) drive delay and a softer surface cut the force rise from 500 N/s to 250 N/s, so the target force arrives in 40 ms rather than 20 ms and the drive overtravels \(3.0\text{ mm}\) into a current-limit shutdown. A \(+3.2\text{ mm}\) range bias starts force regulation early, and a four-frame specular dropout blinds the descent. A 5.5 percent raw pixel difference hides a \(6.0\text{ mm}\) shift in the contact-patch estimate, larger than the \(2.0\text{ mm}\) chamfer tolerance, and actuation arriving as late as 49.5 ms at \(P_{99}\) carries the drive 5.9 mm beyond what the simulator’s 20 ms lockstep accounted for before correction begins. None of these divergences would register in a scalar loss.
| Reality Gap Dimension | Unmodeled Physical Mechanism | Nominal Simulation Spec vs. Real Hardware | Latent Policy Exploitation / Behavioral Failure | Paired-Trace Diagnostic Metric & Tolerance | Hardware & Permission-Path Interceptor |
|---|---|---|---|---|---|
| Actuator & Drivetrain Dynamics | Gearbox harmonic compliance, Coulomb stiction, velocity-dependent Stribeck friction, and back-EMF voltage saturation | Idealized direct torque source (\(J_{\text{rotor}}=0\), \(\tau_{\text{latency}}=0\text{ ms}\)) vs. physical harmonic drive (\(k_c=6.7\times 10^2\text{ N}\cdot\text{m/rad}\), backlash \(\theta_b=0.08^\circ\), bus delay \(\tau=8\text{ ms}\)) | Policy commands high-frequency bang-bang torque reversals (\(>50\text{ Hz}\)), inducing mechanical chatter, gear fatigue, and windings driven past their derating threshold | Step-torque response discrepancy: phase lag \(\Delta \phi \le 5^\circ\) at \(10\text{ Hz}\); settling time discrepancy \(\Delta t_{\text{settle}} \le 15\text{ ms}\); overshoot \(\le 5\%\) | Low-pass action filtering (\(f_{\text{cutoff}} \le 20\text{ Hz}\)), motor current RMS thermal estimators, and actuator slew-rate limiters (\(d\tau/dt \le \tau_{\text{max}}/\Delta t\)) |
| Contact Mechanics & Compliance | Viscoelastic surface deformation, contact hysteresis, non-smooth stick-slip transitions, and finite restitution phase | Point-mass rigid contact via LCP/penalty spring (\(k_p=10^6\text{ N/m}\), restitution \(\Delta t_{\text{res}}=2\text{ ms}\)) vs. physical contact pad (\(k=8.0\times 10^4\text{ N/m}\), deformation span \(\Delta t_{\text{def}}=18\text{ ms}\)) | Policy keys contact onset to the sharp simulated force spike, never sees that spike on the slower real rise, and keeps pushing until the drive reaches its current limit | Contact impulse error \(\Delta I \le 0.05\text{ N}\cdot\text{s}\); peak contact force discrepancy \(\Delta F_{\text{peak}} \le 15\text{ N}\); contact onset rise-time error \(\Delta t_{\text{rise}} \le 5\text{ ms}\) | Approach velocity envelope governor (\(v_{\text{safe}}(z) \propto \sqrt{2 a_{\text{brake}} z}\)), force-derivative clamp (\(dF/dt \le 1.0\times 10^4\text{ N/s}\)), compliant mechanical flexures |
| Sensor Transduction & Noise Profiles | Optical specular dropout, CMOS rolling-shutter distortion, ambient lux-dependent noise, and force-torque gauge thermal drift | Synthetic RGB-D with zero-bias Gaussian noise (\(\sigma = 0.5\text{ mm}\), zero dropout) vs. physical structured-light depth (\(+3.2\text{ mm}\) bias, 4-frame dropout bursts, drift \(0.4\text{ N}/10\,^\circ\text{C}\)) | Policy triggers premature deceleration above surface or remains blind during multi-frame dropouts, driving end-effector past hard fixture boundaries into stall overcurrent | Sensor distribution Kolmogorov-Smirnov distance \(D_{\text{KS}} \le 0.05\); spatial feature activation centroid error \(\le 2.0\text{ mm}\); zero sustained dropout frames (\(N_{\text{drop}} = 0\)) | Multi-sensor fusion (IMU/encoders + vision), optical dropout holdover filters, thermal drift autocalibration, independent proximity break-beams |
| Temporal Latency & Bus Asynchrony | Variable camera exposure times, PCIe/EtherCAT bus serialization, neural inference tail latency (\(P_{99}\)), and non-deterministic OS scheduling | Synchronous discrete-time lockstep (\(\Delta t = 20.0\text{ ms}\), zero transport latency) vs. asynchronous physical pipeline (\(\tau_{\text{total}} = 35.0\text{ ms}\) at \(P_{50}\), \(49.5\text{ ms}\) at \(P_{99}\), jitter \(\pm 6.2\text{ ms}\)) | Severe phase margin erosion (\(>30^\circ\) loss); policy experiences underdamped oscillations, hunting limit-cycles, and unobserved travel during latency spikes (\(5.9\text{ mm}\) at \(P_{99}\) beyond the simulated lockstep) | End-to-end command-response latency jitter \(\Delta \tau_{\text{jitter}} \le 2.0\text{ ms}\); worst-case observation age \(\tau_{\text{obs}} \le 25\text{ ms}\); zero deadline overruns at \(50\text{ Hz}\) | Real-time PREEMPT_RT / Xenomai scheduling, action buffer delay queue injection in simulation, kinematic extrapolation (Smith predictor) |
Domain randomization is often treated as a method for erasing the sim-to-real gap, but parameter randomization operates within strict structural limits.7 Sampling friction coefficients across a uniform distribution from \(\mu = 0.2\) to \(\mu = 0.8\), or link masses across \(\pm 20\) percent, broadens the coverage of training over mechanisms that the simulator already represents. It cannot synthesize an omitted physical mechanism, a missing state correlation, or an unmodeled rare event. If the simulator differential equations omit gearbox backlash, thermal resistance in motor windings, or thread preemption in the operating system, no distribution over link mass will force the policy to learn compensation for those dynamics. Adding unstructured Gaussian noise to simulated observations to mask unmodeled effects can degrade performance, training the policy to lower its control gains and narrow its operational bandwidth to avoid actions that excite the unmodeled modes.
Within the mechanisms the simulator does model, randomization has a useful width, a domain randomization corridor. Too narrow a distribution leaves the policy brittle against the variation the plant actually shows, which risks mechanical shock or a thermal trip, and too wide a distribution drives it toward conservative, low-gain control that gives up throughput. Keeping randomization inside that corridor rather than tuning it by trial takes a fixed order of work. System identification on the physical bench measures the nominal masses, torque constants, delays, and friction first. The randomization bounds are then grounded in the mechanisms that move those parameters, such as winding temperature for the torque constant and surface condition for friction, rather than in heuristic intervals. Training begins in a narrow band around the nominal values and widens it only as the policy meets its declared success threshold, the schedule known as adaptive domain randomization. Finally, paired traces audit every channel against its tolerance, and a discrepancy past budget sends the simulator back for revision before any policy trained on it is deployed.
What randomization cannot reach, the permission path must bound. Every proposal passes the permission path (The Machine in Five Levels) before it reaches an actuator, and there force-rate clamps, torque slew limits, and velocity governors hold it to what its validated sensing and stopping assumptions cover.
↳ Downstream: The sub-millisecond Control Barrier Function filters bounding unverified policy proposals are formulated in Minimal Intervention.
Decomposing the gap into its constitutive physical channels yields two concrete engineering artifacts, namely a supported region and an explicit residual mismatch list. The supported region is the set of conditions within which physical hardware matches simulation within validated tolerances. The residual mismatch list itemizes every unclosed difference that remains outside the training distribution. For the latch simulator, paired probes have not yet closed the strike-plate contact stiffness, which its penalty spring sets far softer than the plate, or the actuator transport lag, so both head its residual list, and the supported region excludes the contact states those two departures govern. Each entry on the residual list defines an explicit requirement for runtime monitors and safety filters, which must detect when physical execution drifts past the boundary where the simulated policy remains valid. Both artifacts pass into the policy manifest of section 1.8: the residual list enters its omission log, and every state outside the supported region is entered there as an exclusion.
Even under careful domain randomization and contact calibration, residual discrepancies persist in high-precision, contact-rich regimes, and the latch simulator’s residual list is not yet empty. Some local mismatches are cheaper to absorb on the deployed plant than to model, at a price in wear and in lost evidence.
Fine-Tuning on the Real Machine
The real cage door carries its strike plate at a small structural mounting offset from where the simulator placed it, a departure the simulator does not model. A local, measurable mismatch of this kind is what fine-tuning on physical hardware can absorb. Gradient updates computed from physical rollouts adapt the policy to the true compliance, contact dynamics, and bus delays of the machine, but they measure none of those mechanisms, so the simulator’s other two departures, the soft contact spring and the missing transport lag, stay in the omission log until paired probes close them. The adaptation also has a physical price. In a simulator, sampling states near the boundary of stability costs only compute cycles, while on hardware every sample expends component life and risks driving the machine into an unrecoverable contact state.
The cost of updating policy parameters on physical hardware compounds, because the gradient steps that solve localized contact errors also consume finite mechanical fatigue life, heat the motor windings, and risk catastrophic forgetting of pre-verified nominal transit stability.
Consider the mobile manipulator’s arm adapting latch-sim-02 into latch-res-03, a residual candidate that adds a gated learned correction to the frozen latch-sim-02 output, because the real plate sits 1.5 mm from where the simulator placed it. The gripper guides the spring-loaded latch past its strike plate, where non-smooth Coulomb friction, cam surface geometry, and spring preload produce sharp stick-slip transitions. The policy’s setpoints arrive every 20 ms (50 Hz). To adapt its parameters to the contact dynamics of the strike plate, an on-robot fine-tuning routine collects an illustrative 20,000 interaction steps, representing 400 s of active contact and near-contact operation. Because the policy must explore to improve, stochastic action perturbations generate unpredictable force transients before low-level safety controllers can intervene. If 1.5 percent of exploratory steps push past the 15 N contact tripwire and strike the rigid shoulder of the strike plate before the permission path clamps the command, the fine-tuning run injects 300 impact shocks into the drivetrain.
Each of those impacts dissipates the arm’s kinetic energy in the gear teeth and bearings, and assigning them a fraction of service life requires a drive-specific load-spectrum and life model (Actuator Transmission Limits). Heat is the second constraint. Exploration dither raises the root-mean-square current through the motor coils, and that Joule heating can bring the winding to its derating threshold (Thermal Duty Cycles) within minutes unless the training schedule budgets cooling intervals.
Beyond mechanical wear, real-world adaptation undermines the statistical verification established during predeployment testing. In simulation, a policy can be subjected to \(1.0\times 10^7\) validation episodes across randomized perturbations, establishing bounded tracking errors and tail behavior out to the 99.9th percentile (\(P_{99.9}\)). Once weight updates \(\Delta\theta\) are applied on hardware, every statistical claim conditioned on the original parameter vector \(\theta_0\) is void. The updated policy may resolve the steady-state force error against the target fixture, yet the gradient steps can simultaneously degrade free-space tracking accuracy, lower phase margins (reducing stability reserves) during high-speed transit, or distort response profiles during emergency stops. Because hardware throughput is bounded by physical time, the machine cannot execute millions of physical trials to re-establish those confidence bounds. The adapted policy enters production with a narrower body of empirical evidence than the simulated baseline it replaced.
Fine-tuning on the deployed machine therefore converts the policy from a qualified, frozen component into an evolving state. Qualification depends on configuration control, where an exact parameter hash corresponds to a documented safety case, so every online weight update reopens that record. The burden of preventing failure shifts to the deterministic monitors of the permission path, which must intercept the errors of the baseline policy and the intermediate behaviors the optimizer produces while it updates against physical feedback.
How Training Choices Go Wrong
A physical failure in a learned policy is rarely arbitrary. The command that drives a gripper into a strike plate was produced by optimization over a specific dataset, loss, and simulator, and the assumptions of that optimization leave a structural footprint that can be traced before the machine runs. The policy vulnerability inference chain makes the trace explicit in five links (\(\text{Exposure} \to \text{Dependency} \to \text{Violated Assumption} \to \text{Observable Signature} \to \text{Permission-Path Interceptor}\)). The training exposure establishes a learned dependency, a deployment condition violates it, the violation produces a signature in the commands or the plant, and the signature names the interceptor the permission path requires. Each of the three latch checkpoints runs the chain once.
For latch-bc-01, the exposure is the teleoperated door openings of Physical Data, and the dependency is the support of those demonstrations, which covers only the nominal strike plate at the guarded approach and none of the off-nominal states a failed attempt reached, because the policy’s imitation targets were the retained successful openings (the tripwire failures stay in the archive as labeled failures, not targets) and the campaign never displaced the plate. A sagged door violates it. The signature is compounding covariate shift. Each correction, issued every 20 ms, lands farther from the demonstrated states, yet every command stays kinematically plausible, so the proposal looks normal until the gripper binds on the edge of the plate. The policy cannot detect its own departure, so the interceptor must be a support monitor at the permission path’s input that compares the measured state against the supported cells and refuses policy authority outside them. Without that monitor, the failure first becomes visible at contact. The predeployment test follows from the dependency. Bench trials on a deliberately displaced plate, a state family the latch campaign’s occupancy join already lists as a known absence (Collection Policy Coverage), show whether the monitor trips before the gripper binds.
For latch-sim-02, the exposure is simulated door openings against a penalty contact spring with no command lag, and the dependency is that a fast plunge costs nothing, because the soft spring absorbs the approach at a simulated peak of 50 N. The real plate violates it. At the 0.10 m/s approach the stiff plate drives the force past the 15 N contact tripwire almost at once and on to the latch limit inside the contact deadline derived in section 1.4, toward the 219.1 N peak. The signature is an overforce transient that the task reward never recorded. The tripwire detects it but fires inside a window the arm cannot use, so the interceptor is the precontact approach-speed envelope, and the dependency is revealed before deployment by paired contact tests rather than by a trip on the machine.
For latch-res-03, the exposure is a narrow slice of contact steps collected on one rig, and the dependency is an over-representation of those recent transitions. Any state the slice left out violates it, such as the free-space transit before the door or an approach off the rig’s axis. The signature is regression, lost stability in states adjacent to the tuned contact, and because the network assigns valid outputs everywhere, nothing in its commands announces the loss. The interceptor is partly procedural and partly at run time. The frozen contact and free-space regression suite must pass before the new weights receive authority, and the residual gate holds the correction at zero outside the measured contact phase. Exploration on the rig leaves a second signature that no task reward records, the shock and winding heat of section 1.6, so the permission path must also watch observables the reward omitted, such as winding temperature, current spectra, and accelerometer energy.
The lineage from latch-bc-01 to latch-res-03 does not retire these failures; it moves the dominant one. Refining the cloned policy in simulation traded drift for solver exploitation, and tuning the result on the rig suppresses the exploit at contact while risking forgetting upstream. The earlier mechanisms stay latent in the weights and return whenever conditions leave the envelope the last stage tuned.
Each chain therefore leaves three obligations. Every dependency needs a predeployment test that stresses the boundary where training departed from the plant. Every signature that appears in a measurable mechanical quantity needs a runtime detector on the permission path at loop rate, whether on tracking deviation, force rate, or current spectra. Where the policy cannot detect its own failure, as with latch-bc-01 off its demonstrated states, the permission path needs a support monitor, the manifest must exclude from its declared ODD the states that monitor cannot cover, and deterministic interlocks must refuse policy authority there. Meeting these obligations requires an exact record of how the checkpoint was built, from the source data and simulator dynamics to the reward terms and every later modification. That record is the policy manifest.
The Policy Manifest
A learned policy fails at the seams where its training regime, source demonstrations, or simulator departs from physical reality, and the three latch chains have just located one seam per checkpoint. The record a trained checkpoint carries into evaluation must capture all three.
The policy manifest records how a checkpoint was produced: its training stages, the data support it inherits from the provenance records, its simulator versions and the omission log of mechanisms they leave out and mismatches paired probes have not closed, the wear incurred during adaptation, and the operating domain it declares (table 3).
The manifest separates optimization results from the conditions the policy may be asked to operate in. The declared ODD is the part of the ODD from Collection Policy Coverage that the policy manifest claims: the supported cells of the data join, minus every state the omission log excludes. High training reward or low offline loss does not show that the policy is safe outside it, so the declared ODD is computed from the join and the log, never from the training curve, and a convergent metric cannot stand in for clearance over regimes nobody characterized.
| Record Schema Field | Physical / Computational Data Modality | Provenance & Immutability Guarantee | Downstream Verifier & Usage |
|---|---|---|---|
regime_sequence |
Array of optimization stages (\(\text{BC} \to \text{SimRL} \to \text{Residual}\)), loss formulas, reward weights, and stopping criteria hashes | Git commit SHA, PyTorch/JAX checkpoint UUIDs, parameter hash \(H(\theta_k)\) | Tracing which regime introduced latent failure modes or high-gain chatter |
data_support_ledger |
Digests of the consumed provenance records and of the scenario ledger they were measured against; the supported_cells and known_absences of their occupancy join |
Cites each provenance record by digest; no dataset field is copied | Upper bound on the declared ODD; parameterizes out-of-distribution support monitors |
negative_event_registry |
Intervention records (Intervention and Recovery Data) from training and adaptation rollouts | Cites each intervention record by digest; its pre-trigger trace is not copied | Isolating physical failure boundaries and preventing corrupted drift rollouts from training |
simulation_provenance |
Physics engine build, integrator method, timestep \(\Delta t\), contact solver penalty parameters, and domain randomization ranges \(\mathcal{P}(\Xi)\) | Physics engine binary hash, URDF/MJCF asset tree hash | Reproducing the training physics and auditing paired-trace budgets |
omission_log |
Mechanisms each simulator leaves out (backlash, compliance, transport lag) and the residual mismatches paired probes did not close, with the states each one affects | Bound to the simulator build it describes; an entry closes only on paired-trace evidence | Removes the states it affects from the declared ODD; carried into evaluation’s coverage gaps and the fault list |
adaptation_wear_ledger |
Serialized hardware IDs, contact shock counts, motor current RMS thermal logs, modified layer indices, and regression suite hash | Calibration certificate hashes, hardware lifetime fatigue logs | Re-opening formal qualification cases and preventing silent regression of nominal capabilities |
declared_odd_boundary |
Declared ODD: the supported cells minus the states the omission log excludes (for the latch policy, the nominal strike plate at the guarded approach speed) | Never exceeds the supported cells of data_support_ledger |
Bounds the target ODD that evaluation may sample and the region in which the permission path admits policy proposals |
For the mobile manipulator’s cage-door latch, a compact manifest links the fields into one decision:
regime_sequence: demonstration checkpointlatch-bc-01→ simulated contact refinementlatch-sim-02→ gated residual candidatelatch-res-03; retain each checkpoint hash, loss or reward specification, and stop rule.data_support_ledger: cite the latch campaign’s provenance records and its scenario ledger by digest; their join supports the nominal strike plate at the guarded approach, and lists the displaced plate and the faster approach, whose cells hold no demonstrations, as known absences with no recovery label.simulation_provenance: record the actual solver build and contact parameters.omission_log: the example simulator models strike-plate contact with a penalty spring far softer than the plate and omits actuator transport lag, so the contact states those two departures govern are unverified for the checkpoints trained in simulation.negative_event_registry: cite by digest the intervention record, with its pre-trigger ring-buffer trace, for every fixture-shoulder strike, operator veto, or permission-path clamp as labeled failure evidence, without treating unsafe commands as expert targets, and mark episodes an operator truncated so that an update does not score an aborted command as a success.adaptation_wear_ledger: associatelatch-res-03with the specific rig, cumulative contact shocks, motor-current history, modified residual weights, and free-space regression result.declared_odd_boundary: the supported cells minus the states the omission log excludes, here the nominal plate at the guarded approach of 0.03 m/s. The contact states governed by the two omissions stay outside the declared ODD for every simulation-refined checkpoint,latch-sim-02andlatch-res-03alike, until paired probes close them. The displaced plate and the 0.10 m/s approach lie outside it, and the fast approach would stay outside even with demonstrations, because its contact deadline (section 1.4) leaves almost nothing beyond the contact-loop response; there the permission path must impose the guarded approach or refuse policy authority.
A release review would reject latch-res-03 until paired contact probes quantify the contact-stiffness error and the omitted lag, and the combined controller passes contact, free-space, stopping, and wear-budget tests. The same rule leaves latch-bc-01 as the one checkpoint evaluation may sample, provided the support monitor of its chain refuses policy authority off the nominal plate (section 1.7). Trained without a simulator, it carries no omission-log exclusions, so its declared ODD keeps the door’s contact states at the nominal plate and the guarded approach. The manifest records that decision and its evidence rather than converting a training score into permission.
For data support, the manifest leaves the collector, plant instance, calibration, and action taps in the provenance record of The Dataset Schema and cites each record by digest rather than copying it. It also cites the digest of the scenario ledger the coverage was measured against and carries forward the result of their occupancy join (Collection Policy Coverage), namely the supported cells and the known absences. Because the citation is by digest, a change to any dataset changes the join, and a manifest that names a superseded digest no longer describes its checkpoint. A state outside the supported cells, such as the latch’s displaced plate, is then visible in the record as ungrounded extrapolation rather than demonstrated competence.
When policy weights are optimized in simulation, the manifest captures the physics engine build hash, the interface contract between the policy and the simulator, the reward formulation, the numerical integration method, and the solver timestep,8 such as a fixed \(2.0\text{ ms}\) step at \(500\text{ Hz}\). It documents the domain randomization distributions applied to link masses, surface friction coefficients, actuator response latencies, and sensor noise profiles. Known physical omissions, most of them left out of the simulation graph to maintain training throughput, go in the omission log alongside any mismatch paired probes have not yet closed (section 1.5). Building on the actuator dynamics established in Actuator Transmission Limits, if a rigid-body simulator omits harmonic drive compliance, structural deflection under a \(5.0\text{ kg}\) payload, or thermal resistance changes in motor windings, the record lists those omissions as unverified dynamic assumptions. When the resulting policy exhibits limit-cycle oscillations during high-load contact on the physical bench, the omission log points the diagnosis to joint compliance rather than ungrounded parameter drift. The log is kept apart from the simulator description because each entry has a downstream consequence, removing the states it affects from the declared ODD and becoming a coverage gap that evaluation and fault injection must address.
Downstream engineering teams consume the manifest throughout the design, evaluation, and release lifecycle. The evaluation record of Evaluation Logs consumes it directly. Its target ODD is sampled from inside the declared ODD, and its catalog of coverage gaps inherits the known absences and the omission log, so evaluation suites probe the specific failure modes each training regime predicts. The permission path of Safety Enforcement takes the declared ODD as the outer limit of the region in which it admits policy proposals, and uses the recorded support to parameterize out-of-distribution detectors. Verification in Adversarial Verification and the safety case in Deployment Release draw on the manifest to establish the provenance of every parameter update and to confirm that each omission is bounded by an independent physical interlock.
↳ Downstream: The policy manifest’s omission log parameterizes adversarial fault-injection benches in Deriving the Fault List.
Checkpoint 1.1: Policy manifests and declared ODDs
Before releasing a trained policy checkpoint to physical verification, verify your understanding of the policy manifest, its omission log, and the declared ODD:
The manifest turns a training history into a claim with stated limits, and the latch record shows the decision that claim supports, sending latch-bc-01 to evaluation and keeping latch-res-03 out of service until paired probes close its omissions. The common ways that decision goes wrong all let a training score stand in for that evidence.
Fallacies and Pitfalls
Minimizing loss on an offline demonstration split, or maximizing simulated return, measures compliance with the training environment rather than competence on the machine. When the objective rewards behavior that exploits data curation or simulator approximations, the ranking of candidate policies established in training inverts on deployment (figure 6), because the virtual environment forgives maneuvers that real hardware does not. Before a trained policy receives actuator authority, the causal assumptions linking its training metrics to physical outcomes must be stress-tested.
Pitfall: Sanitizing recovery episodes to achieve lower offline validation loss.
Suppose the latch archive also held recovery runs, and two cage-door latch policies train on different versions of it. One retains the recovery runs, in which the operator backed the gripper off and re-seated it after the latch cam caught on the strike plate; the other removes them and achieves a lower loss on its own held-out split. When the cam catches on the machine, the first policy backs off and re-seats, while the second keeps pushing along the demonstrated path until the permission path’s contact tripwire halts the arm. Because the sanitized dataset lacks examples of correcting from errors, the lower loss merely reflects an artificially simplified test set, not better physical disturbance handling. Scrubbing the recovery data has made the mathematical prediction task easier at the direct expense of removing the corrective behavior the deployed controller physically needs. Candidate policies must face the same operational disturbances before their training scores can support a deployment choice.
Fallacy: Transitioning from human demonstrations to simulation reinforcement learning removes the failures training left in the policy.
An assembly team replaces behavioral cloning with simulation reinforcement learning to correct lateral drift during gasket seating. The new policy eliminates the drift but seats the mechanism with \(65\text{ Hz}\) torque chatter, which the simulator’s massless-rotor assumption permits because it ignores the inertia the motor must reverse. On the physical machine the chatter excites drivetrain resonance and heats the stator windings to their derating threshold through \(I^2R\) losses (Thermal Duty Cycles), so the drives derate and most seatings fail to complete. The training change addressed one failure mechanism and exposed another, so evaluation must follow the complete task through contact and repeated operation and check the assumptions the new regime introduced as carefully as the data limitations it replaced.
Summary
Every training regime buys a capability and leaves a dependency in the weights. Behavioral cloning depends on the states its demonstrators happened to visit and, with a point decoder, averages the choices they made; simulation depends on its contact model, its timing, and the mechanisms it omits; fine-tuning on the machine depends on a narrow slice of recent experience and voids the evidence gathered for the weights it replaced. None of these dependencies appears in a training score. Each surfaces as a predictable failure signature when deployment violates it, which is why training ends in a record rather than a metric, a policy manifest whose declared ODD no training curve can enlarge. Principle \(\ref{pri-vol4-endogenous-drift}\) already requires that record to mark the states a cloned policy’s own corrections reach without demonstrations, and simulation adds a second kind of absence, the mechanisms a solver omits, which no optimization against that solver can supply.
Key Takeaways: Policy synthesis and sim-to-real transfer
- Every regime leaves a dependency no training score shows: Offline loss and simulated return reward compliance with the data and the solver. Candidate rankings can therefore invert on hardware, and a deployment condition that violates a dependency produces a predictable failure signature that names the detector the permission path needs.
- Cloning fails off the demonstrated states and between the demonstrated modes: Small errors carry the machine into states the demonstrations never visited, where recovery data or a verified support monitor, not action chunking, addresses the drift. A squared-error loss outputs the conditional mean, which can steer between two valid paths into the obstacle both avoided; decoders that represent the modes avoid the mean, but a diffusion decoder’s step count sets its inference time against the deadline its chunks must meet.
- Randomization covers parameters, not missing mechanisms: Domain randomization broadens training over dynamics the simulator already models, and paired probes size the remaining gap channel by channel. An omitted backlash, compliance, or transport delay stays omitted and enters the omission log.
- Fine-tuning on hardware spends wear and evidence: Gradient steps on physical rollouts absorb local contact errors while consuming component life and voiding every statistical claim made for the earlier weights. They do not close an omission, and a frozen regression suite must pass before the new checkpoint receives authority.
- The declared ODD never exceeds the data’s support: The policy manifest cites its data by digest, keeps the omission log as its own field, and declares only the supported cells minus the log’s exclusions, however well the policy scored in training.
What’s Next: From the policy manifest to closed-loop evidence
latch-bc-01, the checkpoint the manifest admits to evaluation, then completes twenty consecutive door runs, and the question becomes what those runs show about the machine, and how many more a stated failure rate would require.
Footnotes
ALVINN road perspective warping: Pomerleau’s Autonomous Land Vehicle in a Neural Network (ALVINN) (Pomerleau 1989), told in Where Physical Data Comes From, drifted toward the lane edge because its human drivers never demonstrated a recovery from it. To supply recovery examples without unsafe driving, Pomerleau used synthetic perspective warping, rendering laterally shifted and rotated road images paired with transformed steering corrections, an early case of augmenting demonstrations with recovery data off the nominal trajectory.↩︎
Mode averaging under L2 loss: Under an \(L_2\) regression loss, the expected risk \(\mathcal{R}(\hat{a}) = \int \|\hat{a} - a\|_2^2 p(a \mid o) da\) is strictly convex with respect to \(\hat{a}\). Differentiating with respect to \(\hat{a}\) and setting the gradient to zero yields \(2 \int (\hat{a} - a) p(a \mid o) da = 0 \implies \hat{a}^* = \int a p(a \mid o) da = \mathbb{E}[a \mid o]\). For a nonlinear plant \(f_{\text{plant}}\), \(\mathbb{E}[f_{\text{plant}}(a)]\) generally differs from \(f_{\text{plant}}(\mathbb{E}[a])\); equality can hold in particular distributions. For a convex action cost \(c(a)\), Jensen’s inequality gives \(c(\mathbb{E}[a]) \le \mathbb{E}[c(a)]\), but it does not establish the physical safety of the mean action. In the obstacle example, the mean lies outside the demonstrated set of valid paths.↩︎
Scope of the CQL bound: Under its data-coverage, optimization, and function-approximation assumptions, CQL constructs a conservative estimate whose expected value for a policy lower-bounds that policy’s true value (Kumar et al. 2020). This is not a pointwise bound for every state-action pair or every dataset, and expected return does not certify physical safety.↩︎
Bearing life versus command reversals: ISO 281 estimates rolling-bearing fatigue life under specified load and speed. A gear-unit catalog may express its wave-generator-bearing \(L_{10}\) in operating hours at rated conditions; neither figure converts commanded reversals into consumed gearheads without a unit-specific duty-cycle model.↩︎
Winding temperature limit: Thermal Duty Cycles gives the IEC 60085 Class F rating behind this threshold. Bang-bang torque commands from unregularized reward functions can hold the current high enough to age the insulation quickly.↩︎
OpenAI Dactyl domain randomization: On the physical hand, Andrychowicz et al. (2020) found that policies trained without the physics randomizations, or without the added models of effects the simulator omits, such as backlash and action delay, achieved far fewer consecutive rotations than the policy trained with all of them.↩︎
Bounded support of domain randomization: Domain randomization optimizes expected return over parameter variations \(\xi \sim \mathcal{P}(\Xi)\). Any generalization claim extends only over the support of \(\mathcal{P}(\Xi)\) and says nothing about dynamics omitted from the simulator’s equations of motion. On a mechanism with unmodeled transmission compliance or bus jitter, the policy treats those dynamics as unexplained noise and can fall into a high-frequency limit cycle.↩︎
Choosing a penalty-contact timestep: The needed step size depends on effective contact mass, penalty stiffness and damping, the integrator, and solver tolerances. The characteristic contact timescale changes with roughly \(\sqrt{m_{\text{eff}}/k_{\text{eff}}}\); refine the timestep until contact onset, peak force, and impulse converge to the accuracy the task requires.↩︎

![Physical Policy Training Farm and Hardware Parallelism. Large-scale robotic data collection cluster comprising 14 custom articulated manipulators operating concurrently over months to collect over 800,000 physical grasp attempts for vision-based hand-eye coordination [@levine2018learning]. Each robotic station consists of a 7-DoF arm, a two-finger compliant gripper, an over-the-shoulder RGB camera, and a dedicated bin containing diverse household objects. Scaling hardware parallelism to P=14 compresses the wall-clock calendar time Twall = N(trollout + treset)/P, but incurs heavy capital expenditure, continuous mechanical wear across gearboxes and linkage bearings, thermal buildup in motor windings, and substantial supervisory overhead to clear dropped objects and mechanical jams. Attribution: @levine2018learning.](images/png/fig06_real_robot_arm_farm.png)
![Sim-to-Real Robotic Transfer Setup with Digital Twin Simulation. Authentic experimental apparatus for vision-based dexterous in-hand manipulation from OpenAI's Dactyl project [@andrychowicz2020learning]. (Left) The physical experimental cage housing a 24-DoF anthropomorphic Shadow Dexterous Hand, surrounded by 16 PhaseSpace infrared motion-tracking cameras and 3 Basler RGB cameras for ground-truth state validation and multi-view vision. (Right) The high-throughput MuJoCo simulated digital twin environment used to train reinforcement learning policies across billions of synthetic transitions. Robust sim-to-real transfer is achieved by applying extensive domain randomization over surface friction, joint damping, actuator force gains, link and object masses, object dimensions, visual lighting, and camera parameters, together with modeled motor backlash and action delay, forcing the neural policy to treat unmodeled physical variations as noise while operating within the supported deployment envelope. Attribution: @andrychowicz2020learning.](images/jpg/fig06_real_sim2real_dactyl.jpg)