Closed-Loop Evaluation
Closed-Loop Evaluation
Purpose
Why does passing twenty consecutive physical trials without failure tell you surprisingly little about whether the machine is safe to deploy?
Demonstrating twenty consecutive error-free trials reveals what occurred during a specific test protocol, not how the system behaves under open-world variability. Bounded trial counts lack statistical power to observe rare tail failures, while testing under nominal conditions reveals nothing about behavior under shifted friction, sensor degradation, or unexpected human interference. Offline benchmark scores compound this blind spot by scoring static logs. In physical deployment, closed-loop actions alter future states and subsequent sensor observations, creating cascading drift where an initial micro-error drives the system into unmodeled dynamics.
Rigorous physical evaluation must track closed-loop performance on the integrated machine and quantify the empirical confidence supporting every reliability claim. Crucially, evaluation must also delineate where statistical evidence ends. Because catastrophic physical failures cannot be rolled back, test campaigns are necessarily censored at the boundary of safe operation. Testing alone cannot demonstrate high-consequence failure rates; an empirical evaluation serves the system by bounding claims strictly to sampled operating domains, leaving unverified edge states to the deterministic runtime monitors guarding the causal boundary.
Learning Objectives
- Interpret physical success rates as conditional evidence about a specified machine, environment, and trial protocol
- Compare open-loop prediction accuracy with closed-loop behavior shaped by the policy’s own actions
- Select the evidence regime (replay, simulation, shadow operation, or constrained physical testing) that can answer a given evaluation question
- Calculate trial counts and test time needed to support a stated reliability target
- Evaluate which rare failures remain unsupported by feasible testing and the available statistical assumptions
- Design a predeclared trial protocol with fixed outcome rules and full-denominator accounting of interventions and aborts
- Construct an evaluation record that confines its claim to a target operating domain and assigns uncovered conditions to runtime monitors
What Twenty Clean Runs Establish
Fresh from the offline training pipeline, latch-bc-01 is flashed onto the warehouse mobile manipulator; it is the one checkpoint the policy manifest of The Policy Manifest declares evaluable. In initial validation, the machine completes twenty consecutive runs of its door task. Each run starts with the base 2 m (illustrative; see the Reader Guide) from the spring-latched cage door; the base drives in and stops, and with the base at rest the arm unlatches and opens the door at the 0.03 m/s guarded approach, without a contact-tripwire crossing or a human intervention. Its empirical sample success rate is \(\hat{p} = 1.00\) (100 percent). In software testing, a clean pass on a fixed test suite is often sufficient evidence for release. In physical AI systems engineering, that record leaves the deployment decision unresolved, because twenty clean runs establish only a bounded statistical lower limit and leave substantial uncertainty about unobserved failures.
The familiar Wald interval (\(p \pm z \sqrt{\hat{p}(1-\hat{p})/n}\)) cannot describe that uncertainty, because at zero failures its variance term vanishes and the interval collapses to \([1.0, 1.0]\), a claim of certain reliability. The exact bound inverts the binomial distribution instead. This one-sided Clopper-Pearson bound (Clopper and Pearson 1934) holds at any sample size, and its approximate form is the Rule of Three (Hanley and Lippman-Hand 1983): when \(n\) independent trials pass with zero failures, the upper 95 percent confidence bound on the failure probability \(p_{\text{fail}}\) is approximately \(3/n\). For the twenty door runs (equation 1), \[p_{\text{fail, upper}}^{95\%} \approx \frac{3}{20} = 0.15. \tag{1}\] The exact computation in 1.1 tightens this to a 13.9 percent failure upper limit, which still leaves the machine short of 95 percent demonstrated reliability. A policy that fails once in every eight runs still completes twenty consecutive runs 6.9 percent of the time.
Napkin Math 1.1: Derivation of the 20-run lower bound
Mathematical Derivation: For independent, identically distributed Bernoulli trials from the declared operational test distribution, \(p_{\text{lower}}\) is the success probability at which 20 consecutive successes have probability \(\alpha=0.05\). Thus the one-sided 95 percent lower confidence limit is \(p_{\text{lower}} = \alpha^{1/n} = (0.05)^{1/20} \approx\) 0.861 (86.1 percent); the corresponding failure-probability upper confidence limit is 13.9 percent.
Systems insight: The observation “20 out of 20 passed” supplies a confidence limit under those sampling assumptions, not an observed failure rate, a mean interval between failures, or a deployment permission. Inverting the bound gives the trials a target needs, \(n = \lceil \ln\alpha / \ln p_{\text{target}} \rceil\), so a \(99.9\) percent success lower limit requires 2,995 consecutive zero-failure trials drawn from the target distribution.
↰ Prerequisite: Physical failure boundaries for autonomous mobile robots and manipulators are established in Which Budget Binds First.
Evaluation closes the learning loop of Part II. Evaluating the twenty clean runs that opened this chapter turns the experience logged in Physical Data and the weights trained in Policy Training into a concrete decision about the physical claim those trials can support. Each trial measures the whole machine rather than the policy alone. It records the policy’s proposals, the permission path’s decisions, and the Body’s response, including which out-of-distribution proposals the permission path rejects, which it misses, and how much stopping margin remains.1
The twenty-run result leaves three questions open: what outcome was sampled, what each test regime can expose, and which hazards remain outside the evidence. Finite trials bound only the distribution they sampled, so an evaluation is complete only when it also names the conditions it did not cover and the runtime monitors the permission path needs for them. A success rate means nothing until the population it was drawn from is written down.
What a Success Rate Measures
An evaluation benchmark routinely reports a single summary statistic, such as an 85 percent success rate on an obstacle avoidance task. In machine learning, that number is treated as an intrinsic score of the learned model weights, similar to top-one accuracy on static image classification datasets. In a physical machine, that interpretation is invalid. A learned policy has no independent physical performance. What the benchmark measured was the closed-loop execution of an entire physical system across a specific sequence of days and environmental states. Success rate is an empirical property of the coupled system, comprising the policy, mechanical plant, computational hardware, operating environment, and trial protocol.
To evaluate a policy, an engineering team executes a sequence of discrete physical runs. A single trial is fully determined only when every variable that can alter the physical state trajectory is fixed or explicitly sampled. Formally, let a single trial condition be defined by a tuple \(c = (s_0, e, m, \pi_\theta, \kappa)\), drawn from an operational testing distribution \(\mathcal{D}\). The vector \(s_0\) represents the initial state of the machine, including SE(3) pose, linear and angular velocities, and joint encoder positions. The environment vector \(e\) captures external variables such as floor friction, ambient illuminance, obstacle geometry, and ambient air temperature. The machine state \(m\) captures physical and hardware properties, including battery state of charge, payload mass, tire wear, mechanical backlash in the drivetrain, and sensor calibration parameters. The computational configuration \(\pi_\theta\) encompasses the frozen policy weights \(\theta\), inference quantization, scheduling deadlines, and software versions. Finally, the termination protocol \(\kappa = (T_{\max}, g, \phi)\) defines the execution time limit \(T_{\max}\), the spatial goal region \(g\), and the binary evaluation function \(\phi\) that maps the resulting trajectory to a terminal outcome \(Y \in \{0, 1\}\).
The true success probability under this evaluation protocol is the expected value of the binary task outcome averaged over the joint distribution of trial conditions, \(p(\mathcal{D}) = \mathbb{E}_{c \sim \mathcal{D}}[\phi(\text{Traj}(c))]\) (Closed-loop trajectory distributions and evaluation estimands gives the formal integral). Here \(\text{Traj}(c)\) is the continuous physical trajectory produced when the policy \(\pi_\theta\) executes within the physical environment and hardware state defined by \(c\). This dependency demonstrates why success probability cannot be treated as a measurement of \(\theta\) alone. The model weights \(\theta\) constitute only one parameter inside the condition vector \(c\). If the distribution over floor friction or the distribution over initial battery charge changes, the expected success rate changes while \(\theta\) remains identical. The sampling protocol is not an external apparatus measuring an isolated model; it is part of the mathematical quantity being estimated.
Definition 1.1: Closed-loop evaluation estimand
Closed-loop evaluation estimand is the expected physical task outcome \(p(\mathcal{D}) = \mathbb{E}_{c \sim \mathcal{D}}[\phi(\text{Traj}(c))]\) defined over the joint distribution of initial states, mechanical plant wear, environmental physics, frozen policy weights, and evaluation predicates: \[c = (s_0, e, m, \pi_\theta, \kappa) \sim \mathcal{D}\] rather than an intrinsic property of the neural policy weights evaluated in isolation.
- Significance: In static machine learning, held-out accuracy is treated as a property of the model \(\pi_\theta\). In physical AI, substituting a worn tire, letting battery voltage sag, or shifting ambient illuminance alters the physical state transition function without changing a single weight. The sampling distribution of operational conditions \(\mathcal{D}\) is an inseparable part of the quantity being estimated.
- Distinction: Unlike open-loop validation metrics (such as validation loss or mean squared trajectory error) which score predictions against frozen recorded states, the closed-loop estimand integrates over the endogenous feedback loop where past actions physically move the machine, directly changing the future sensor observations it will receive.
- Common pitfall: Reporting a single success rate percentage without predeclaring the support, density, and environmental bounds of the test distribution \(\mathcal{D}\). Any physical policy can be made arbitrarily successful by testing exclusively in benign, unperturbed states.
The physical mechanisms that alter \(c\) between trials operate independently of the policy, and the mobile manipulator exhibits each of them. As its battery discharges, the drives lose voltage headroom against back-EMF, so the torque available at speed falls along the torque-speed curve. Daylight shifting across the site lengthens the navigation camera’s exposure and blurs frames during turns. Dust on the aisle floor lowers the friction the brakes can use. In the arm, repeated door openings heat the joint windings (\(I^2R\)) and warm the harmonic drive gearboxes, which changes their damping and, as the links expand, shifts the tool-center-point calibration, so the gripper meets the latch at a slightly different point on an afternoon run than on a cold morning one. The weights are identical in every case, yet the trajectory can move from success to failure.
The reset procedure moves \(c\) as well. An operator who repositions the base after a difficult run may favor an easy start, and an automated reset that repeats one path heats the windings cyclically. An operator who catches a failing run and discards it as a reset erases the hardest failure from the sample.
Statistical evaluation must therefore distinguish sampling variation from a change in the estimand. When \(n\) trials are drawn from a fixed \(\mathcal{D}\), the sample proportion \(\hat{p} = \frac{1}{n} \sum_{i=1}^n Y_i\) fluctuates around \(p(\mathcal{D})\), and standard estimators bound that noise. Loading the tote rack or moving the door’s strike plate changes the experiment to a different distribution \(\mathcal{D}'\) with its own \(p(\mathcal{D}')\), and pooling the two samples combines observations from two different machines.
Because success probability is defined only with respect to a specific joint distribution \(\mathcal{D}\), the intended operating region must be declared before trials are sampled. The target ODD is the part of the declared ODD (The Policy Manifest) that the evaluation samples; its sampling weights define \(\mathcal{D}\), and no evaluation claim reaches beyond it. If an engineering team executes fifty physical runs across ad hoc trajectories and reports a 90 percent pass rate, the denominator corresponds to an undefined population. A measured success rate supports a target-distribution claim only when the support and sampling weights of \(\mathcal{D}\) are specified in advance, establishing explicit bounds on illuminance, payload ranges, surface friction coefficients, and allowable initial state offsets.
Extracting valid confidence bounds from empirical trials requires three operational conditions:
- Independent and Identically Distributed Sampling: Trials must be independent realizations drawn from \(\mathcal{D}\). If thermal accumulation in motor windings or progressive mechanical wear degrades hardware efficiency across consecutive runs, outcomes drift with run order, so successive trials are neither independent nor identically distributed.
- Fixed Outcome Predicates: The binary evaluation function \(\phi\) and the termination criteria \(\kappa\) must remain fixed throughout the campaign. Modifying the time horizon \(T_{\max}\) or adjusting the spatial tolerance of the goal region \(g\) midway through testing alters the definition of success.
- Full-Denominator Accounting: Every initiated run that aborts due to hardware faults, communication timeouts, or safety-rated emergency stops (e.g., exceeding the collaborative-robot power and force limits of ISO 10218 / ISO/TS 15066) must be recorded as a failure (\(Y = 0\)). Dropping aborted trials substitutes a measure of conditional execution on surviving hardware for the true operational reliability of the machine.
When these conditions hold, physical trials yield an unbiased sample mean for the closed-loop system under \(\mathcal{D}\). Physical trials are also slow, so evaluation workflows often substitute open-loop scoring on recorded datasets, which measures how accurately a candidate reproduces logged expert actions. That substitution assumes that accurate prediction of demonstrated actions carries over when the policy’s own actions choose its next state, and only the closed loop can test it.
Closed-Loop Dynamics
A navigation policy can reproduce recorded aisle runs almost exactly and still spend the base’s clearance to a rack before it is a third of the way down the aisle. The difference lies in what each kind of test feeds back to the policy. In an open-loop evaluation, a candidate policy \(\pi_\theta\) is scored across a static dataset of recorded demonstrations \(\mathcal{D}_{\mathrm{demo}} = \{(s_t^*, a_t^*)\}_{t=1}^T\). At each discrete timestep \(t\), the policy receives a recorded demonstrator state \(s_t^*\) and outputs a predicted action \(\hat{a}_t = \pi_\theta(s_t^*)\). An error metric, such as mean squared error or cross-entropy, measures the distance between \(\hat{a}_t\) and the recorded expert action \(a_t^*\). At step \(t+1\), the evaluation pipeline supplies the next recorded state \(s_{t+1}^*\) from the dataset. The action \(\hat{a}_t\) produced by the policy is discarded without executing on hardware, keeping the sequence of input states independent of the policy’s past outputs. In a closed-loop evaluation, the policy proposes \(a_t = \pi_\theta(s_t)\); the permission path checks it against measured state and constraints, applies \(u_t\) (or a fallback), and the plant evolves as \(s_{t+1} = f(s_t,u_t)\). The applied command changes the physical state and later observations, crossing the causal boundary in The Causal Boundary.
Open-loop evaluation tests actions along the demonstrator’s state trajectory; a policy with low offline prediction error can fail when its biased proposals are admitted to physical motion, because its own actions carry it off the states it was trained on. This is the endogenous-data principle (\(\ref{pri-vol4-endogenous-drift}\)); Endogenous Experience developed its consequence for data, and evaluation must measure it on the machine. The door campaign cannot show the effect, because its base stands still while the arm works, so the example comes from the same machine’s navigation stack rather than from the latch policy under evaluation. The base drives a 30 m aisle at its 1.5 m/s drive limit (both illustrative), and the nominal expert demonstration follows the centerline with heading \(\theta^* = 0\text{ rad}\) and lateral offset \(y^* = 0\text{ m}\). Suppose the candidate navigation policy has a systematic heading bias of \(\epsilon_\theta = 1.0^\circ\) (0.01745 rad) on every step, and the permission path admits each of these proposals because none is infeasible on its own while clearance remains.
In an open-loop evaluation, the policy receives the demonstrator state \(s_t^* = (x_t^*, 0, 0)\) at each step and predicts a forward velocity with a \(1.0^\circ\) steering offset. The offline scoring pipeline records a mean absolute angular error of \(1.0^\circ\) (0.01745 rad) and an extremely low mean squared error of \(\text{MSE}_\theta = (0.01745)^2 \approx 3.05 \times 10^{-4}\text{ rad}^2\).2 In closed-loop execution, however, because the machine physically moves based on previously admitted commands, uncorrected heading errors compound continuously over time. Under planar unicycle kinematics (\(\dot{y} = v \sin \theta \approx v \epsilon_\theta\)), this persistent angular offset integrates into a cumulative lateral displacement: \[y(t) = \int_0^t v \sin \epsilon_\theta \, d\tau \approx v \epsilon_\theta t\] The base is 100 cm wide in a 130 cm aisle, which leaves \(W_{\text{clear}} =\) 15 cm of side clearance to the rack on each side, and the drift spends that clearance at \[t_{\text{impact}} = \frac{W_{\text{clear}}}{v \sin \epsilon_\theta},\] which at the drive limit is 5.7 s, after 8.6 m of travel and less than a third of the way down the aisle (figure 1).3 Either the permission path’s clearance check stops the base short of the rack or, without that check, the base strikes it. Both outcomes are a failed trial that open-loop scoring rated near-perfect.
Systems Perspective 1.1: Open-loop blindness
The aisle example shows both reasons that an open-loop score decouples from closed-loop survival. First, static evaluation cannot integrate a persistent error into displacement. It resets the machine to an expert state at every clock cycle, masking the compounding of integration errors, whereas in closed loop the state at index \(t\) is the expanded composition of all previous actions: \[s_t = f(f(\dots f(s_0, u_0)\dots, u_{t-2}), u_{t-1})\] A persistent heading bias needs no software bug, since differential tire wear, camera-mast vibration, or floor camber produces one. Second, static evaluation cannot follow the machine into the unfamiliar states that early errors steer it toward. The first admitted error moves the base from the demonstrated state \(s_1^*\) to an unvisited state \(s_1\), and once it drifts a few centimeters off the centerline it occupies states absent from \(\mathcal{D}_{\mathrm{demo}}\). The offline metric, computed only where \(y = 0\text{ m}\), holds no evidence about \(\pi_\theta(s)\) for \(y \neq 0\text{ m}\), and a policy never shown recoveries from off-center states can make larger errors there, accelerating the divergence that Policy Synthesis Regimes traced for behavioral cloning.
The arm shows the same disconnect in contact. An end-effector position error of a fraction of a millimeter per step scores as negligible Cartesian error offline, but on the latch that offset acts as a wedge. The latch’s stiffness turns the lateral displacement into a reaction force (\(F_{xy} = k_{\text{env}} \Delta_{xy}\)) that climbs past the 15 N contact tripwire, which ends the attempt before the door unlatches.4
Offline scores can also mislead in the pessimistic direction when the task admits several valid actions. If the demonstrator passed an obstacle on the left, an open-loop score penalizes a candidate that passes on the right (the two-mode case of Policy Synthesis Regimes), although in closed loop both reach the goal. An action-level error measures divergence from one recording, not satisfaction of the task.
⇄ Contrast: Open-loop prediction loss contrasts with the photon-to-actuation age-of-information bounds analyzed in The Cost of Ingestion.
The same gap separates each offline metric commonly reported for a learned policy from the machine failure it hides, and table 1 pairs each with the closed-loop measure that exposes it, among them the time to intervention \(T_{\text{TTI}}\), measured from the start of policy authority to the first permission-path override (the permission path refuses a proposal and applies its fallback) or operator takeover.
| Offline metric | Open-loop assumption | Why it breaks in closed loop | Failure on the machine | Closed-loop measure |
|---|---|---|---|---|
| Action mean squared error | Small residuals give small path deviations | A persistent bias integrates into drift and carries the machine off the demonstrated states | The base spends its rack clearance (stop or strike) | Peak deviation from the reference path; time to intervention (\(T_{\text{TTI}}\)) |
| Discrete action accuracy | Correct labels on logged states give identical execution | Benign alternatives and boundary excursions cost the same; errors correlated in time are invisible | Command chatter and joint over-torque | Peak jerk; adherence to actuator torque limits |
| Cosine similarity | Collinear proposals give stable paths | Magnitude is ignored, so under-actuation stalls against static friction and over-actuation slips | The base stalls or slips its wheels; the arm stalls short of the latch | Slip ratio and damping (Wheel slip ratio and friction boundaries) |
| Negative log-likelihood | Likelihood at the expert action captures the action distribution | Valid alternatives, such as passing an obstacle on the other side, are penalized | Valid policies rejected; demonstrator noise imitated as jitter | Closed-loop task success over randomized starts |
| Reconstruction or Chamfer loss | Accurate scene prediction gives safe plans | Background pixels outweigh thin obstacles and contact geometry | Phantom braking; clipping a cable or a glass partition | Measured clearance; barrier-condition invariance |
The disconnect also runs across the proposal boundary of The Machine in Five Levels. In the two-processor implementation, offline trace replay on the application processor measures only whether neural inference executes within its memory and compute budgets. In physical closed-loop execution, each candidate proposal crosses to the microcontroller that runs the permission path. If the policy’s proposals drift off target or chatter, the permission path detects the spatial boundary breach or the expired lease, refuses the proposals, and commands a controlled stop. Evaluating a policy offline, without the physical plant and the permission path in the loop, evaluates tensor arithmetic, not the survival of the machine; only the closed loop measures the path excursion, the intervention timing, and the permission path’s measured performance.
↳ Downstream: Deterministic gatekeeping and stopping envelopes for unverified neural policies are derived in Stopping Envelopes.
Open-loop scores cannot certify a closed loop, so the evidence about the proposer must come from executions whose conditions are known and recorded. Every setting in which a policy can be executed changes some part of the loop, whether the plant, the sensor stream, or the authority over the actuators, and section 1.4 sorts those settings by what each can and cannot show.
Evidence Regimes
Before its twenty door runs, the latch policy could have been scored against its logged demonstrations, run against a simulated door, run in shadow while a validated controller held the arm, or run on the real door with its contact tripwire armed. These four settings are the evidence regimes of figure 2 and table 2, namely recorded replay, closed-loop simulation, shadow operation, and constrained physical testing. Each establishes a different relationship between the policy, its sensory inputs, and the state of the machine, so each supports a bounded claim. None of the four establishes how the unassisted policy behaves during unconstrained physical deployment, because each regime alters either the causal feedback of the control loop or the distribution of states the machine encounters.
In recorded replay, the policy executes as a feedforward computation on stored sensor buffers. The environment does not respond to the policy outputs, and the state sequence remains fixed to the trajectory recorded during data collection. This regime supports claims about computational throughput, memory allocation, and loss against a logged dataset, such as whether inference fits its latency budget on the onboard accelerator and whether its numerical outputs match the recorded actions. It cannot support any claim about closed-loop stability, because the predicted action never alters the next observation; if the policy outputs a hard left turn, the next recorded frame still shows the base driving straight.
Closed-loop simulation restores the causal feedback loop by substituting an explicit mathematical model, \(f_{\text{sim}}\), for the physical machine and its surroundings. When the policy proposes an action and the modeled permission path selects \(u_t\), the simulation integrates the equations of motion, updates the scene, and renders synthetic sensor inputs for the next cycle at index \(t+1\). This regime supports claims about the policy’s behavior under the modeled dynamics, such as whether the base recovers from an initial offset or the arm avoids self-collision across its workspace, and thousands of instances can run in parallel. The evidence holds only for \(f_{\text{sim}}\); unmodeled backlash, tire slip, contact friction transitions, and lighting are excluded by construction, so stabilizing \(f_{\text{sim}}\) says nothing yet about the hardware.
Shadow operation mounts the candidate policy onto an active physical machine while a validated baseline controller, or a human operator, retains authority over the actuators. The candidate policy receives live sensor frames, runs onboard inference under real mechanical vibration and bus traffic, and computes candidate actions \(a_t^{\text{cand}}\).5 These candidate actions are logged but never dispatched to the motor drives. Shadow operation supports claims about execution timing, such as whether the latch policy holds its chosen 20 Hz chunk rate under operational thermal load, and about how often \(a_t^{\text{cand}}\) diverges from the baseline action \(a_t^{\text{base}}\) by more than a stated threshold. It shares the limitation of replay, because the baseline keeps the machine on the baseline’s trajectory; an action that would have driven the base into a rack never moves it, and the sensor stream never shows the candidate’s failure.
Constrained physical testing closes the causal loop on real hardware by letting candidate proposals reach the actuators within a bounded envelope. Boundaries take the form of physical tethers, velocity clamps, padded test cells, or the permission path itself (The Machine in Five Levels), which refuses unsafe commands at a validated response rate. This regime supports claims about contact dynamics, actuator saturation, and sensor noise within the tested envelope, which is why the twenty door runs belong to it. It cannot establish the policy’s tail behavior in open deployment, because every intervention by a tether or a permission-path override truncates the trajectory. The data measure the combined policy and permission path inside the constraint boundary, not the unassisted policy.
Some properties transfer across regime boundaries. Numerical outputs for identical input tensors transfer, as do algorithmic invariants, so an overflow that saturates an output in simulation saturates it on hardware given the same input; measured timing and memory pressure must be retested under each regime’s bus, thermal, and scheduling load. The state distribution does not transfer. The moment the environment model changes from \(f_{\text{sim}}\) to \(f_{\text{real}}\), or the tested permission path changes, the sequence of states \(s_0, s_1, \dots, s_t\) diverges.
Viewing these regimes as a ladder of fidelity is a trap. Advancing from replay to simulation to shadow mode to constrained testing does not strengthen one confidence claim; each step trades one distortion for another. Simulation buys closed-loop causality at the price of real sensing and contact, shadow mode trades causality back for real sensor streams and hardware timing, and constrained testing restores both but censors the states beyond its safety boundary. Zero interventions across a hundred constrained trials show that the trajectories stayed where the constraints never engaged, not that the failure rate outside the envelope is low. Each question belongs in the regime able to answer it (table 2). For the latch policy, whether inference fits the chunk period belongs in replay and then in shadow operation on the target hardware; whether the arm recovers from a misplaced first contact belongs in simulation, where thousands of offsets cost no wear; and whether the door’s spring and the gripper’s compliance keep contact force below the limit belongs only in constrained physical trials. Each regime is blind to the failure modes of the others.
| Evaluation Evidence Regime | Causal Feedback Coupling | Environmental Dynamics Grounding | Supported Empirical Evidence | Masked Failure Mode / Boundary Distortion |
|---|---|---|---|---|
| Recorded Replay | Open Loop (\(\partial s_{t+1}/\partial \hat{a}_t = 0\)) | Static Demonstration Buffers \(\mathcal{D}_{\text{demo}}\) | Latency budget, DRAM footprint, numerical accuracy | Covariate shift; errors never alter subsequent observations |
| Closed-Loop Simulation | Closed Loop (\(s_{t+1} = f_{\text{sim}}(s_t, u_t)\)) | Synthetic Physics Operator \(f_{\text{sim}}\) | Kinematic reachability, large parallel test suites | Sim-to-real gap; unmodeled contact non-linearities and tire slip |
| Shadow Mode Operation | Open Loop for candidate (actions muted) | Physical machine \(f_{\text{real}}\) | Bus contention, thermal throttling, real sensor noise | Trajectory confinement; states locked to baseline controller \(d^{\pi_{\text{base}}}\) |
| Constrained Physical Testing | Closed Loop with runtime veto (\(u_{\text{safe}}\)) | Physical machine \(f_{\text{real}}\) within envelope | Contact dynamics, compliance, friction, actuator limits | State censorship; permission-path overrides and tethers truncate failure tails |
The quantitative relationship between simulated proxy benchmarks and real-world robot performance is governed by the sim-to-real proxy transfer law (figure 3). While offline open-loop metrics (such as action Mean Squared Error) fail to predict closed-loop outcomes (\(r = 0.308\), suffering severe ranking inversions as detailed in table 1), closed-loop simulation benchmarks such as SIMPLER (Li et al. 2024) demonstrate a strong linear correspondence with physical manipulation success across embodiments (\(R^2 = 0.836\), Pearson \(r = 0.914\), \(p < 0.0001\); Spearman rank correlation \(\rho = 0.942\)).
Critically, the linear transfer law (\(y = 1.04 x + 5.8\%\)) holds reliably only within nominal visual and kinematic envelopes. The transfer breaks down catastrophically across three distinct physical gap regimes:
- Visual domain gap: Specular lighting miscalibrations or rendering discrepancies can disrupt vision transformer tokenizers. For example, untuned visual textures in Octo-Base collapsed simulated solve rate to \(0.0\%\) despite achieving \(29.3\%\) real-world task success.
- Contact mechanics gap: Rigid-body physics engines (such as standard ODE or Bullet solvers) idealize multi-body contact dynamics and omit non-linear Coulomb stiction. An RT-1-X policy evaluated on top-drawer opening reached \(89.1\%\) in simulation but dropped to \(40.7\%\) on real hardware due to mechanical rail jamming.
- Tactile and force gap: Fine-grained manipulation (such as sub-millimeter peg insertion) yields high visual solve rates in simulation (\(82.0\%\)), but collapses in reality (\(14.0\%\)) where vision alone cannot detect contact force wedging.
Consequently, while closed-loop simulation serves as an effective screening filter for policy candidates, physical trial limits remain the ultimate arbiter of embodied competence.
Checkpoint 1.1: Trade-offs across evidence regimes
Before evaluating physical trial limits, verify your understanding of evidence regimes and fidelity trade-offs:
Physical Trial Limits
A reliability claim is paid for in trials. Choosing a regime settles what a test can show, but on a physical machine the plant, not the compute budget, fixes what each closed-loop trial costs, and statistics fixes how many trials the claim demands. Those two quantities together separate a qualification argument from an anecdote.
What a physical trial costs
A physical trial consumes more than the simulated step or log-replay time. It consumes wall-clock time, stresses mechanical linkages, heats motor windings, and risks physical damage to the hardware. The door campaign shows how the cost accumulates. In one run the base drives in from its 2 m standoff and stops, and the arm then unlatches at the 0.03 m/s guarded approach and swings the door open, which together occupy 25 s of the test cell (illustrative). The reset that makes the next run an independent draw re-latches the door, returns the base to its standoff, and flushes the run’s telemetry, adding 15 s (illustrative), so each trial costs 40 s.6 At that cycle the twenty clean runs that opened the chapter took 13.3 min, and the campaign that a 99.9 percent success lower limit requires (1.2) runs far longer, even before the arm’s windings take any thermal rest (Thermal Duty Cycles).
The plant also changes while the trials accumulate. The confidence bound assumes independent draws from a stationary process, and wear breaks stationarity: gearhead backlash, wheel radius, tire friction, and joint damping drift at rates that must be measured on the test hardware. Pooling a thousand runs then treats a progression of degrading machines as repeated tests of one, and a late failure may reflect the worn plant rather than the policy.
The censoring built into constrained testing (section 1.4) compounds the drift. Because a failed physical trial cannot be recalled (the first law of The Four Bedrock Laws), safety mechanisms intervene before a failure completes, and the evidence stops where they intervene. The states a team most needs to evaluate, such as loss of traction at speed or a collision the base cannot avoid, are the states it cannot let the machine enter, because a skid into a rack damages the chassis and the sensing rig and takes the test cell offline for repair. Cages, geofences, and operators with emergency stops cut each run short at a conservative envelope. The trial records the intervention while the downstream consequence of the policy’s output stays unobserved; the intervention stays in the denominator, and the consequence stays unknown.
For physical survival, a mean is an inadequate description of a peak. What threatens the latch is a run’s peak contact force against the 100 N damage limit, however modest the average, and a long bus delay can consume stopping margin however small the mean latency. Extreme value theory (EVT) offers candidate models for maxima, but a Gumbel, Fréchet, or Weibull family must be fitted to a defined sample of block maxima or threshold exceedances, checked against alternatives, and reported with its uncertainty; the physical mechanism alone does not select the family.7 Without those data, table 3 states what to measure and what protection to validate rather than predicting rare-event frequencies.
| Disturbance channel | Peak quantity and sampling record | Physical consequence to test | Candidate protection to validate |
|---|---|---|---|
| Compute and bus latency | Per-cycle end-to-end age and deadline misses | Stale proposals erode braking margin | Freshness lease and independently timed fallback |
| Contact shock | Force-time trace, impulse, and peak at the load path | Yield or overload before a reactive command takes effect | Passive compliance and precontact speed limit; reactive limits address later load |
| Low traction | Local friction estimate and stopping response | Slip extends stopping distance | Conservative speed ceiling and tested fallback on low-grip surfaces |
| Optical glare | Saturation duration and perception misses | False clearance or lost state estimate | Redundant sensing and independent physical constraints |
The latch shows why the contact-shock row pairs a measurement with a protection. Each door run contributes one block maximum, its peak latch force, and twenty maxima are far too few to fit a tail that would predict how often a run crosses 100 N. A force monitor that reacts only after contact cannot undo the first peak either; its measured detection and actuator response can limit the loading that follows, while the 0.03 m/s guarded approach, a precontact speed limit, reserves the margin for the initial impact.
Distributing trials across a fleet does not buy independence. Ten machines tested concurrently in one warehouse share a production batch, a floor finish, a ventilation system, and operators trained under one protocol, so an anomaly absent from the facility, such as floor glazing or low-angle glare, reaches none of them. Ten thousand fleet hours in one facility are ten parallel observations of one narrow distribution, not ten thousand hours of independent exposure. Standardized test fixtures such as those in figure 4 buy repeatability by bounding the state space to the fixture.
Trials per claim
A zero-failure campaign is sized by inverting the exact bound of section 1.1 for the trial count, so the claim, not the test budget, fixes how many runs the test cell must deliver (figure 6 and table 4; derivations in Zero-Failure Testing and the Exposure Wall).
Napkin Math 1.2: Statistical scaling of assurance
Calculations: The inversion of 1.1 gives \(n = \lceil\ln(0.05)/\ln(0.999)\rceil = 2{,}995\) consecutive zero-failure trials. At 40 s per trial, this demands \(T_{\text{total}} =\) 2,995 \(\times\) 40 s \(\approx\) 33.3 h (1.4 days) of continuous testing. With a two-operator safety protocol, this consumes 66.6 h of operator time, or over eight eight-hour shifts.
Systems insight: Even on a cycle as short as the door’s, a 99.9 percent lower limit costs 1.4 days of continuous test-cell occupancy, and a longer cycle scales it in proportion. A component replacement during the campaign may require a new population definition rather than pooling incompatible runs.
When a reliability requirement is stated as a failure rate \(\lambda\) per operating hour, the trial count becomes an exposure time. The exposure wall is the failure-free exposure a zero-event test needs before it can bound a hazard rate. Under a constant-rate model at 95 percent confidence, \(T \ge -\ln(0.05)/\lambda \approx 3.0/\lambda\): a target of \(\lambda=10^{-6}\) per hour needs \(3.0 \times 10^{6}\) failure-free hours (342 machine-years), and \(\lambda=10^{-9}\) per hour needs \(3.0 \times 10^{9}\).
Butler and Finelli drew the design conclusion from this arithmetic (Butler and Finelli 1993) (1.2). A fleet does not rescue the test, because 100 rigs running continuously would still need about 3,420 calendar years at \(10^{-9}\) per hour. Road-vehicle mileage meets the same wall (figure 5), and simulated miles add no physical exposure to it. A high-assurance claim therefore rests on independently validated runtime constraints and structured fault injection (Deriving the Fault List).8
Systems Perspective 1.2: Limits of testing for ultra-reliability
The exposure wall binds even at a site requirement far looser than those rates. Suppose the warehouse mobile manipulator must demonstrate a critical failure rate \(\lambda \le 1.0\times 10^{-4}\text{ failures/hour}\) (Mean Time Between Critical Failures \(\text{MTBF} \ge 10{,}000\text{ hours}\)) at 95 percent statistical confidence (\(\alpha = 0.05\)). Each retrieval mission takes \(t_{\text{mission}} =\) 360 s (0.10 h) and covers 540 m at the 1.5 m/s drive limit. Automated dock recharge and telemetry reset add \(t_{\text{reset}} =\) 120 s, for 480 s per trial cycle. The drive-wheel diameter is \(D_{\text{wheel}} =\) 0.20 m, and the drive gearhead carries an illustrative \(L_{10}\) bearing fatigue rating of \(5.0 \times 10^{7}\) revolutions.
By the exposure wall, this target needs 29,957 failure-free hours (3.42 years) of continuous exposure. Discretizing this requirement into missions with per-trial failure budget \(p_{\text{fail}} \le \lambda \cdot t_{\text{mission}} =\) \(1.0 \times 10^{-5}\) requires \(N_{\text{req}} = \ln(\alpha)/\ln(1-p_{\text{fail}})\), or 299,572 consecutive zero-failure trials. This inverts the zero-failure bound of section 1.1 at a target far beyond practical budgets of tens to thousands of runs. On a single physical rig, executing these trials at 480 s per cycle consumes 39,943 hours (4.56 years) and accumulates \(N_{\text{rev}} = (N_{\text{req}} \cdot v \cdot t_{\text{mission}}) / (\pi D_{\text{wheel}}) \approx\) \(2.57 \times 10^{8}\) wheel revolutions.
The wheel-rotation total is 5.15 times the gearhead’s \(L_{10}\) rating, but the comparison holds only after the drivetrain ratio maps wheel revolutions to bearing revolutions, and \(L_{10}\) is a population-life statistic at specified load and speed, not a destruction deadline. A campaign of this length must therefore plan maintenance and component replacement, and record each one, before it may pool trials.
| Zero-failure trials \(n\) | 90% lower success limit | 95% lower success limit | 99% lower success limit | 95% upper failure limit | Test time at door cycle | Interpretation under iid test draws |
|---|---|---|---|---|---|---|
| 10 | 79.43% | 74.11% | 63.10% | 25.89% | 0.1 h | Wide uncertainty; no observed failures |
| 20 | 89.13% | 86.09% | 79.43% | 13.91% | 0.2 h | Bound applies only to the sampled distribution |
| 50 | 95.50% | 94.18% | 91.20% | 5.82% | 0.6 h | More evidence; untested conditions remain |
| 100 | 97.72% | 97.05% | 95.50% | 2.95% | 1.1 h | No deployment permission follows from the interval |
| 250 | 99.08% | 98.81% | 98.17% | 1.19% | 2.8 h | Check hardware and sampling stationarity |
| 500 | 99.54% | 99.40% | 99.08% | 0.60% | 5.6 h | Bound remains tied to declared conditions |
| 1,000 | 99.77% | 99.70% | 99.54% | 0.30% | 11.1 h | Below a 99.9% lower-limit target |
| 2,995 | 99.92% | 99.90% | 99.85% | 0.10% | 33.3 h | Meets that statistical target if assumptions hold |
These ceilings force an architectural conclusion. Physical AI safety cannot rest on statistical confidence in the learned weights, and classical control offers no analytical substitute for the proposer.
A classical controller is certified by a Lyapunov function \(V(\mathbf{x}) > 0\) with \(\dot{V}(\mathbf{x}) \le -\gamma V(\mathbf{x})\) over a stated region, not by counting runs, and with disturbances an input-to-state stability argument bounds the state error by the size of the input under its own model assumptions. A learned proposer carries no such stability certificate. Checking the decrease condition through its non-smooth activations is a non-convex search over millions of parameters, its latent inputs carry no physical units on which to build \(V\), and contact discontinuities break the smoothness the proofs assume. Outside its training distribution a learned policy can also change its output abruptly, so that a shadow or a camera glint flips the direction of a proposed motion. The burden of proof therefore moves off the proposer and onto a low-dimensional check on an isolated microcontroller. Two strategies carry that burden, one at run time and one on the bench:
- Runtime Containment: Bounding the physical authority of the learned policy by delegating the collaborative-robot force limits, kinematic acceleration limits, and emergency stops to the permission path, an independent, deterministic mechanism below the proposal boundary of The Machine in Five Levels. The protection claim then rests on the permission path and its fallback, whose state estimates, plant and actuator models, response timing, and reserved stopping margin must be validated separately from the policy.
- Fault Injection on Physical Benches: Rather than waiting for passive disturbances to strike, engineers inject them. Injected trials follow the fault-injection method of Deriving the Fault List; this chapter requires only that they enter the same full-denominator ledger.
Checkpoint 1.2: Sample size and MTBF evidence target
Before moving on, check your understanding of sample size scaling and physical test limits:
Even under rigorous qualification protocols and automated test cells, statistical sampling alone cannot establish ultra-reliable autonomy. What remains outside any finite campaign, and which runtime monitors the record must name for it, decides how far the evidence can carry the release decision.
What Evaluation Cannot Establish
The limits of evaluation follow from the endogenous loop of section 1.3. The states a physical campaign visits depend on the commands the permission path admitted, so no fixed dataset matches the candidate’s closed-loop state distribution, and a simulator substitutes its own transition model for the plant’s (formal trajectory distributions in Closed-loop trajectory distributions and evaluation estimands). A simulator that models the aisle floor with the friction of clean, inspected concrete cannot show what the base does on a clear oil film, and one whose actuators respond instantly cannot show the lag of a real stator’s electrical time constant; the physical trajectory departs from the simulated one, and the departures compound into states no simulated run visited. A test therefore bounds only the distribution it actually sampled.
The strongest conclusion an empirical evaluation can support is therefore narrow and conditional. The door campaign establishes that one software build on one hardware revision, from the sampled initial conditions and inside the environment present during the runs, failed with probability below 13.9 percent at 95 percent confidence. It says nothing about a colder morning that derates the joints, a gearbox whose compliance and backlash have grown with wear, or a displaced strike plate. More runs on the nominal plate reduce sampling uncertainty and tighten the bound (table 4), but they cannot reduce structural uncertainty about unmodeled physics or unvisited conditions; two thousand further runs would measure the same cell with finer precision and leave the displaced-plate cells, the loaded machine, and unmodeled dynamics such as bus jitter exactly as unsupported as before.
Since evaluation cannot establish how the policy behaves outside the sampled distribution, each monitorable boundary of the claim needs a runtime signal, threshold, and tested response, and the permission path must enforce independent physical constraints. Distribution support is only partially observable and an OOD detector has measurable false negatives, so conditions that cannot be sensed reliably require ODD restrictions, passive protection, or other independent barriers. The evaluation record carries these specifications in its runtime_monitor_specs field (section 1.8). The field names five monitoring channels whose coverage and response must be measured:
- Input Support / Out-of-Distribution (OOD) Monitor: Estimates whether incoming observations depart from the tested distribution using a calibrated score and reports false negatives on held-out shifted conditions.
- Execution Latency Monitor: Measures each inference age against the declared control deadline (for example, the latch policy’s 50 ms chunk period at 20 Hz), records misses, and checks that lease expiry invokes the validated fallback. A \(P_{99}\) summary alone is not a hard deadline bound.
- Constraint Margin Monitor: Tracks the distance between the current physical state and hardware safety limits, including joint torque limits, motor winding temperature, and obstacle clearance against limits specified for the hardware and stopping model.
- Intervention Rate Monitor: Calculates the frequency of permission-path overrides, detecting when policy instability causes excessive corrections.
- Environmental Assumption Monitor: Measures external operating conditions, such as surface traction estimated from wheel slip and ambient illumination measured by photometric sensors.
A monitor that only writes an anomaly to a log protects nothing. When a monitor reports that a condition has left the target ODD, or that an inference missed its deadline, the permission path discards the proposal and acts on the actuators through a validated fallback, a speed restriction or a controlled stop, provided the sensed state, actuator capability, and reserved clearance make that response feasible (fallback logic in Safety and Control Barrier Functions). A timely, confident proposal still passes the independent state, clearance, and actuator checks before any \(u_t\) is applied.
Some failures no runtime monitor detects. A proposal can meet every deadline, respect every torque limit, and arise from familiar inputs, yet steer the base toward a glass partition that the vision-language backbone (Brohan et al. 2023) took for free space; the measured clearance is wrong until contact. Safety against such failures cannot come from evaluation or software monitors. It requires physical zoning, mechanical current limiters, and independent safety barriers whose protective capability is validated independently of the learned model under specified physical assumptions.
↳ Downstream: Physical failures undetectable by sensory telemetry are formally cataloged in Failures Without Detectors.
These limits bind even a flawless campaign. The evidence that survives them is only as good as the campaign that produced it, and procedural and statistical errors can corrupt a campaign long before its result reaches the record.
How Evaluation Goes Wrong
Systematic errors undermining policy evaluation rarely stem from deliberate fraud. They arise when a protocol answers an easier question than the claim asks, and six recurring modes do so: environment selection, metric drift, test-set contamination, optional stopping, denominator filtering, and a demonstration offered as a measurement.
Environment selection occurs whenever the test bench is narrower than the claim. The twenty door runs all met the nominal strike plate at the 0.03 m/s guarded approach with the base stopped and the tote rack empty. If the report states their success rate as the machine’s ability to open the site’s cage doors, the claim silently expands to the five unsampled cells of the ledger and the loaded machine (section 1.6). Environment selection is detected by comparing a condition ledger, which records the measured strike-plate position, approach speed, payload, and floor condition of every trial, against the boundaries of the claimed domain; a ledger that covers one cell establishes nothing about the others.
Metric drift occurs when the evaluation metric slides from physical task completion toward an offline surrogate such as action agreement, the mean squared error or cosine distance between predicted commands and recorded expert actions. Section 1.3 showed why such a score cannot certify the closed loop. Metric drift is detected by requiring the primary evaluation metric to be a closed-loop terminal outcome, such as an open door with no contact-tripwire crossing and zero permission-path overrides, rather than open-loop prediction error.
Even with closed-loop physical trials, repeated tuning against a fixed evaluation setup degrades result independence. When a team adjusts the latch policy’s hyperparameters, observation filtering, or checkpoint after each batch of runs on the same cage door, that door ceases to be an independent holdout. Each change absorbs information about that door’s strike-plate position, spring, and lighting, and after enough rounds the final checkpoint reflects implicit training on the test door itself, masking how the policy responds to any other door. This test set contamination cannot be detected by inspecting the final model weights or running a single verification trial on the same door. It is detected only through provenance tracking, analyzing the chronological history of policy versions, tuning commits, and test dates to confirm that the reported evaluation was executed on a fresh, quarantined configuration that was never exposed to iterative tuning.
Evaluation protocols are equally vulnerable to unintended cherry-picking through optional stopping. When an engineer monitors a continuous sequence of physical runs and halts the experiment as soon as a target performance threshold or statistical significance criterion is reached, the reported confidence is distorted. If an evaluation team conducts an experiment where each look carries a nominal false-positive threshold of 5 percent, stopping early upon observing an apparent success across multiple evaluation checkpoints inflates the true error rate. Five independent 5 percent looks would have a 22.6 percent chance of at least one false positive. Repeated looks at one accumulating run are dependent, so the inflation for a real stopping rule must be computed for that rule. Optional stopping cannot be detected from summary tables or final pass rates. It is detected only by auditing the pre-registered trial protocol, verifying that sample sizes, stopping conditions, and evaluation milestones were fixed before the first physical run began.
A long campaign still has reason to stop as early as the evidence allows, because every run wears the testbed and spends the arm’s thermal budget (The Five Physical Budgets). To stop early without inflating Type I error, evaluation teams employ Abraham Wald’s sequential probability ratio test (SPRT)9 (Wald 1945). Rather than fixing sample size in advance or naively peeking at intermediate results, SPRT evaluates an accumulation of evidence after each physical rollout, computing the cumulative likelihood ratio between an acceptable reliability hypothesis (\(H_1\)) and an unacceptable failure hypothesis (\(H_0\)). Testing halts dynamically the moment the accumulated evidence crosses predeclared likelihood-ratio thresholds. For two specified simple hypotheses and a declared sequential design, SPRT can reduce expected sample count compared with a fixed-horizon test. The common Wald boundary formulas are approximations affected by likelihood-ratio overshoot; actual Type I and Type II errors and testbed savings must be calculated or simulated for the chosen hypotheses and outcome model. For the exact log-likelihood martingale formulations, Bernoulli updates, and Wald boundary derivations, see Sequential Probability Ratio Test.
A related statistical distortion occurs when the denominator of an evaluation is quietly filtered, breaking the full-denominator condition of section 1.2. A team that discards runs ending on software timeouts, emergency stops, sensor communication drops, or permission-path overrides as invalid test artifacts reports success over the surviving trials only. If the arm opens the door on nine of ten attempts and the tenth needs an operator takeover to stop the gripper wedging against the latch, removing that attempt raises the apparent success rate from 90 percent to 100 percent. When physical constraints or safety limits truncate a run before its time limit \(T_{\max}\) expires, the trial is recorded as a censored failure rather than omitted from the ledger. Reading survival on truncated trajectories as operational competence is the censored trial fallacy. Denominator filtering is detected by cross-referencing low-level actuator telemetry and permission-path logs against the trial ledger, confirming that every power-on cycle and autonomy engagement corresponds to a recorded outcome.
War Story 1.1: The 2017 NHTSA Autopilot crash rate mirage
Mechanism: The rate divided airbag deployments by miles driven before and after installation, and the report’s own footnote noted that the rates “are for all miles travelled before and after Autopilot installation and are not limited to actual Autopilot use.” The denominator therefore never measured exposure to Autosteer itself, and it was incomplete as well. A replication by Quality Control Systems Corporation found that the odometer reading at installation appears to have been reported for fewer than half of the vehicles, and that exact exposure before and after installation could be computed for only 13 percent of them (Quality Control Systems Corp. 2019). The mileage NHTSA left uncounted, because it could be assigned to neither period, was concentrated among the vehicles with the least exposure before installation.
Impact: In the one cohort with complete exposure in both periods, the airbag-deployment crash rate rose 59 percent after installation, from 0.76 to 1.21 per million miles, a result the replication qualified with “if these data are to be believed.” Its authors concluded that the overall 40 percent reduction was “an artifact of the Agency’s treatment of mileage information that is actually missing in the underlying dataset.” The replication does not measure Autosteer’s effect either. It shows that the published denominator could not support the reported reduction.
Response: The correction came from outside the agency and two years after the closing report, because the per-vehicle mileage records that exposed the denominator were released only after a Freedom of Information Act lawsuit against the Department of Transportation (Quality Control Systems Corp. 2019).
Systems lesson: In physical AI evaluation, an unverified denominator produces a statistical mirage. Aggregating open-world mileage without engagement telemetry, operational design domain boundaries, and complete exposure records turns missing data into a safety claim. An evaluation audits the integrity of its denominator with the same rigor it applies to the failure events in its numerator.
The temptation to substitute visual plausibility for statistical sampling represents a pervasive vulnerability in physical systems engineering. A single video recording of a robotic arm grasping an irregular workpiece or a mobile robot navigating a cluttered room is a demonstration, not a measurement. A visually convincing demonstration proves that the system state trajectory intersects the success set for at least one sequence of initial conditions and environmental disturbances, but it provides no information about the volume of that success set. Because human perception is tuned to smooth motion and coordinated SE(3) kinematics, an observer easily mistakes kinematic fluidity for policy competence. An execution that appears deliberate and smooth can operate within \(10\text{ mm}\) of a kinematic singularity or carry an unmeasured \(P_{95}\) tracking error that produces a collision in nine out of ten repeats. A representative demonstration has no statistical standing. A representative campaign needs predeclared sampling from the target ODD or separate per-stratum analysis, full-denominator accounting, and a sample size tied to its stated claim.
The same discipline governs comparisons. As Jacob Cohen formalized in his statistical power analysis10 (Cohen 1988), deciding whether latch-bc-01 opens the cage door more reliably than another checkpoint requires the Type I error (\(\alpha\)), the Type II error (\(\beta\)), and the effect size, the smallest difference in success rate worth detecting. Twenty runs per checkpoint carry no stated power until those quantities and the baseline success rate are known, so the protocol fixes \(\alpha\), \(\beta\), and the effect size before the first run.
Each of these pathologies widens the claim beyond the evidence while leaving the report looking cleaner: a benign bench, a surrogate metric, a contaminated test door, a well-timed stop, a filtered denominator, or a persuasive video. Understanding the mechanisms by which evaluation fails clarifies the purpose of physical testing. An empirical trial is not a demonstration to convince an observer that a model works. It is an instrument to measure where policy authority fails, what physical conditions trigger the failure, and how much statistical confidence the evidence supports. Documenting those boundaries requires an immutable record that binds every trial, condition, and intervention to a verifiable physical quantity, and section 1.8 specifies that record.
Evaluation Logs
When the twenty door runs are over, the record they leave behind is the only form in which the rest of the machine will ever see them, so what that record keeps decides which claims survive. An evaluation record consumes the policy manifest of The Policy Manifest, citing it by digest, and adds the fields required to audit the learned policy that the manifest describes; it makes no claim beyond the target ODD and names a runtime monitor for every monitorable boundary. The record links measured outcomes to declared conditions and preserves human observations separately with their source. It serves as a self-contained evidence record that states what was measured, under what physical conditions, with what confidence, and which operational claims the resulting evidence explicitly does not support.11
| Record Schema Field | Physical / Statistical Modality | Cryptographic & Provenance Guarantee | Downstream Verifier & Usage |
|---|---|---|---|
target_odd_boundary |
Target ODD: the sampled cells of the declared ODD, here the nominal strike plate at the 0.03 m/s guarded approach under the door’s own latch load, base stopped | Signed ODD specification; must lie inside the policy manifest’s declared_odd_boundary |
Defining tested conditions and calibrated OOD alerts; physical gating remains independent |
component_hashes |
Policy manifest digest, frozen neural network weights (\(H(\theta)\)), inference engine binary, chassis serial number, tire batch, sensor calibration matrices, and firmware build | SHA-256 binary digests, hardware UUID, ISO 9001 calibration certificate references | Supporting configuration traceability; preventing unrecorded hardware or driver substitutions |
sampling_protocol |
Target-distribution draw weights or per-stratum quotas, PRNG seed, reset procedure, and independence audit | Automated test harness configuration script hash, telemetry sequence log | Auditing against optional stopping and testing against hidden autocorrelation across trials |
complete_outcome_ledger |
Full-denominator raw outcome counts: \(n_{\text{attempted}}\), \(n_{\text{success}}\), \(n_{\text{estop}}\), \(n_{\text{timeout}}\), \(n_{\text{override}}\) (permission-path overrides), and \(T_{\text{TTI}}\) time-series | Monotonic microsecond hardware timestamps, uninterrupted ring-buffer telemetry | Preventing censored trial fallacy; calculating uncorrupted empirical reliability |
statistical_evidence |
Exact binomial limit with sampling assumptions; Poisson rate only for separately modeled physical-hour exposure | Cryptographic verification signature, reproducible Python evaluation script | Passing bounded evidence and explicit assumptions to Adversarial Verification |
unsupported_gap_catalog |
Explicit catalog of unvisited state spaces and omitted dynamics (surface water, dynamic crowds); contains the provenance record’s known absences and the manifest’s omission log | Declared non-support manifest, signed safety engineering sign-off | Parameterizing the permission path’s checks, CBF constraints, and fallback policies in Part III |
runtime_monitor_specs |
For each monitorable boundary of the claim: signal, threshold, and tested response, across the five channels of section 1.6 | Measured detection coverage and false-negative rate per channel | Implemented by the permission path and the observation contract’s health field; read by the enforcement record |
The evaluation record begins with the operational claim and its formal predicates, as specified in the predeployment evaluation record schema (table 5). It defines the mathematical success criterion, such as unlatching the cage door and swinging it open without crossing the 15 N contact tripwire and without a permission-path override of the policy. It bounds the target ODD, which lies inside the declared ODD of the policy manifest, across the physical parameters that manifest names, including strike-plate position, approach speed, and the load the door presents. Finally, it names the target population of physical environments to which the result is intended to generalize, distinguishing the one cage door tested from cage doors with other strike plates or spring rates. Defining these boundaries in the record before trials begin prevents post-hoc adjustments to the pass criteria when anomalous behaviors appear.
The record captures immutable cryptographic identifiers for every software and hardware component to enable independent replication of the physical execution. These identifiers include the content hashes of the neural network weights, sensor preprocessing pipelines, policy configuration files, simulation engine builds when synthetic runs are incorporated, and the evaluation execution scripts. On the physical system, the record logs the machine chassis serial number, wheel tire compound batch, sensor calibration matrices, motor controller firmware revisions, and communication bus configurations. Swapping a camera lens, altering a driver firmware version, or replacing worn drive belts changes the physical transfer function of the machine, invalidating the historical data if the change is unrecorded.
The record preserves the trial protocol, which was fixed before the first run so that the campaign cannot be halted at a favorable moment (the optional-stopping hazard of section 1.7). The protocol predeclares the sample size \(N_{\text{req}}\) from the claim’s target and confidence, the outcome classes, and the stopping rule: a zero-failure campaign stops at its first failure and reports the exact lower bound its successes support, or completes \(N_{\text{req}}\) runs and reports \(\alpha^{1/N_{\text{req}}}\). A pooled binomial bound needs each condition drawn independently from the target distribution, whereas a fixed quota grid needs separate per-stratum inference and declared weights before any claim spans the target ODD. The record documents how the runs were drawn from the cells of the target ODD, the pseudorandom seed sequence governing where the base stops in front of the door, and the reset sequence executed between consecutive runs, which re-latches the door and returns the base to its 2 m standoff. When the base returns to that standoff under autonomous control, battery thermal rise and odometry drift introduce statistical dependence between adjacent attempts. The record specifies whether trials were statistically independent or chained in series, and it registers every manual reset, hardware abort, or deviation from the pre-specified protocol.
Aggregate percentages obscure the physical failure modes that determine system risk. Each attempted run lands in exactly one predeclared class: success, reaching the goal within \(T_{\max}\) with no override; censored failure, when the permission path overrode the policy; timeout; or hardware failure, such as a dropped bus or a faulted drive. The record reports the raw count in every class, and it reports kinematic boundary violations and manual emergency stops as subcounts of the class in which each affected run ended. In addition to summary tallies, the record preserves per-trial time-series metadata, including peak motor winding temperatures, maximum trajectory tracking errors, and task completion durations. Logging this telemetry ensures that nominal successes operating near thermal limits or within \(5\text{ mm}\) of a hard boundary are visible during audit. Every trial initiated under policy authority remains in the denominator.
The statistical specification defines the point estimator, confidence method, nominal confidence level, sampling distribution, and underlying mathematical assumptions. Alongside the supported claim, the record documents unsupported claims and known coverage gaps; the catalog begins from the known absences of the demonstration data and the omission log of the policy manifest, and adds what the trials themselves left unvisited. For each gap that can be sensed, the runtime monitor specifications name the signal, threshold, and tested response of section 1.6; a gap that cannot be sensed stays in the catalog as a restriction on operation.
Consider the twenty-run door campaign that opened this chapter. The latch policy opened the cage door twenty times from the 2 m standoff at the 0.03 m/s guarded approach, with the base stopped and no permission-path override. An auditor reads the record’s statistical evidence as the exact bound of section 1.1, a 95 percent lower confidence limit of 86.1 percent, never as the degenerate interval \([1.0, 1.0]\).
The completed record exposes the structural limits of the claim, and it shows how far the evidence narrowed on its way through Part II. The scenario ledger of Collection Policy Coverage crossed three strike-plate positions with two approach speeds, six cells in all, and the retained demonstrations supported one, the nominal plate at the guarded approach. The policy manifest made that cell the declared ODD of latch-bc-01, the target ODD took in the whole cell, the twenty clean runs were drawn from it, and they bound the failure probability there at 13.9 percent, the one-sided upper confidence limit under the tested iid distribution (table 4). Everything that narrowing left out enters the coverage gap catalog as a gap the campaign inherited rather than closed. The five cells that held zero demonstrations hold zero trials as well, and the catalog inherits the manifest’s omission log, whose two unclosed simulator omissions, the soft penalty contact and the missing transport lag, bar this record from being reused for any checkpoint refined in simulation. The trials add a gap of their own, because the base carried nothing in its 50 kg tote rack, so the record says nothing about the loaded machine. The record shows why twenty successful runs cannot support a high-reliability claim, and it gives a fixed basis for refusing any assertion that the policy is ready for unmonitored deployment.
An evaluation record does not certify that an autonomous machine is safe. It establishes the empirical boundaries and statistical uncertainty of a single learned component operating under controlled conditions. In Adversarial Verification, the verification argument ingests these evaluation records as formal evidence nodes, combining policy bounds with evidence for the physical guardrails and the permission path’s latencies to evaluate whether the integrated machine satisfies its operational safety case.
↳ Downstream: Non-asymptotic Clopper-Pearson trial bounds supply the empirical evidence nodes for safety cases in Safety Cases and Claims.
Systems Perspective 1.3: Evaluating a machine that falls
Question: What does a coverage record say when every failure in an unsupported cell is a fall that damages the test article?
The mobile manipulator’s failed latch trial ends at the contact tripwire with an intact arm, so its campaign can keep sampling a hard cell; a humanoid’s failed trial can end the campaign’s hardware.
- Fall count is its own estimand. A humanoid’s failed trial can end on the floor with the robot under test damaged, so the record reports falls with their own denominator and confidence limit beside task success rather than folding them into one success rate.
- The denominator is stratified before collection. The disturbance strata (slippery floor, a change in floor compliance, a curb, a shove) and their weights are fixed in the sampling protocol before the first trial, because a campaign on a flat floor bounds nothing about the strata it skipped.
- A demonstrator’s balance does not transfer. A human operator keeps balance with a human body, not the robot’s, so a cell such as a balanced heavy reach has no valid demonstrations; it enters the coverage gap catalog as a known absence, and the campaign cannot close it with trials it cannot afford to fail.
A log that fixes its denominator, strata, and coverage gaps before collection makes the record auditable, but it cannot stop a reader of that record from drawing conclusions it does not support. Practitioners still misread benchmark metrics and carry evidence beyond the target ODD in a few recurring ways, and those misreadings are the last hazard to the claim.
Fallacies and Pitfalls
An evaluation supports claims only within its tested conditions and limits. These mistakes arise when teams extend a result beyond those limits or attribute an architectural safeguard to the learned policy.
Fallacy: Repeated success in a single nominal condition proves broad operational readiness.
The door campaign repeats one condition twenty times: the nominal strike plate, the guarded approach, the base stopped, the rack empty. Under the binomial model, those runs give the 86.1 percent lower bound of 1.1 for that condition, and nothing for the other cells of the ledger or a loaded tote rack that moves the arm’s torque margins. Repeating the condition sharpens the nominal estimate while the deployment-critical envelopes stay untested. A readiness claim must name its target ODD, allocate trials across the physical variations inside it, and report uncertainty within each tested cell.
Fallacy: Zero failures in 1,000 trials proves compliance with a 99.9 percent reliability specification.
Suppose the latch policy opens the door 1,000 times without a trip against a requirement of 99.9 percent success at 95 percent confidence. The observed success fraction is perfect, yet the one-sided lower bound \(0.05^{1/n}\) is only 0.9970, short of the requirement (table 4); meeting it takes 2,995 clean runs. The gap is an evidence deficit, not a policy failure. The trial budget follows from the claim and confidence level, fixed in advance, and the report states the lower bound rather than the perfect fraction.
Fallacy: High offline action agreement on static logs guarantees closed-loop physical stability.
An evaluation report that cites close agreement with expert actions on recorded runs as release evidence has measured imitation on those records and nothing more. The aisle policy of section 1.3 matched its recordings almost exactly and still spent its side clearance to covariate shift. Agreement can screen candidates before physical trials, but a claim of closed-loop stability rests only on closed-loop outcomes counted against the full denominator, the metric-drift guard of section 1.7.
Pitfall: Crediting the learned policy with surviving transport stalls that the enforcement architecture handled.
The mobile manipulator completes 50 nominal aisle runs, and the team credits its learned policy with tolerance to SerDes bus jitter. A qualification test then injects a mailbox stall that outlasts the chosen 60 ms chunk lease of Multi-Rate Cadences. The lease expires, and the permission path detects the stale lease and initiates a controlled dynamic stop, analogous to a Category 1 stop, from the fallback ladder of The Fallback Ladder. That response comes from the enforcement architecture; nominal completion records do not establish that the neural weights detect or compensate for the stall. Fault injection must measure the permission path’s detection and stopping behavior against the available stopping distance margin, and the evaluation report must attribute the protection to the hardware mechanism that supplied it.
Summary
Closed-loop evaluation measures the coupled policy, permission path, plant, environment, and trial protocol, so a success rate describes a declared population rather than a weight file. Because the machine’s own admitted actions choose the states it reaches, which is Part II’s endogenous-data principle, only trials that close the loop through the permission path measure what the proposer does; open-loop scores cannot see the drift those actions produce. A zero-failure confidence limit then bounds the sampled population under its sampling assumptions, and censoring at the edge of safe operation keeps that population inside the envelope the campaign could afford to enter. More trials narrow that bound without widening the population, and the exposure wall places rare-hazard rates beyond any campaign a physical machine can run.
The door campaign shows the shape of the claim that survives. Twenty clean runs support a lower confidence limit for the nominal strike plate at the guarded approach, with the base stopped and unloaded, and nothing about the displaced-plate cells the demonstrations never covered or about the loaded machine. The coverage gap catalog records the states no trial visited, which the endogenous-data principle predicts no dataset holds, and assigns them to independent physical constraints and a feasible fallback on the permission path rather than to the proposer, because a runtime OOD detector can miss what it has never seen.
The chapter’s product is the evaluation record, and four of its parts travel forward:
- Bounded operational claim: The target ODD, inside the policy manifest’s declared ODD, where performance was empirically characterized.
- Reproducible configuration and ledger: Digests of weights, firmware, and calibration, with every attempted, aborted, and overridden trial kept in the denominator.
- Coverage gap catalog: The known absences of the data and the manifest’s omission log, plus the conditions the trials left unvisited.
- Runtime monitor specifications: For each boundary that can be sensed, the signal, threshold, and tested response that the permission path must implement at runtime.
Key Takeaways: Closed-loop evaluation and survival reliability
- A success rate belongs to a declared population: The rate reflects the coupled policy, permission path, plant, environment, and trial protocol. Changing the payload, the floor, or the reset procedure changes the estimand, so the target ODD and its sampling weights are fixed before the first trial.
- Only a closed loop measures a proposer: Open-loop scoring returns the machine to a demonstrated state at every step. A persistent \(1.0^\circ\) heading error scores as negligible offline yet spends the aisle’s 0.15 m side clearance within 8.6 m, because the machine’s admitted actions choose its next states.
- Zero failures bound a rate without measuring it: \(n\) clean iid trials give a one-sided 95 percent failure upper limit near \(3/n\). Twenty clean runs leave room for a 13.9 percent failure probability, and a 99.9 percent success lower limit needs 2,995 clean runs.
- The exposure wall moves rare hazards to architecture: Bounding a rate \(\lambda\) with zero events takes about \(3/\lambda\) of failure-free exposure, and a fleet’s pooled hours share floors, builds, and conditions. Protection against rare hazards therefore rests on the permission path, a feasible fallback, and physical barriers, never on the learned weights.
- An evaluation record states where its evidence stops: Its claim stays inside the target ODD, its ledger keeps every aborted and overridden run in the denominator, and its gap catalog and runtime monitor specifications tell the runtime what the permission path must cover.
What’s Next: From the evaluation record to runtime perception
Footnotes
Runtime safety enforcement: A separately validated permission path can reject proposals that violate measured state, freshness, and actuator constraints. Its detection coverage, worst-case response time, and available stopping clearance must be tested for the implementation; a 1 kHz sampling rate alone cannot guarantee interception before contact.↩︎
Metric scale masking in regression loss: Aggregating squared steering errors across thousands of validation frames obscures severe, intermittent directional mistakes. A single localized prediction spike of \(15^\circ\) is mathematically washed out when averaged across hours of straight aisle driving. A brief angular spike can matter if its integrated motion exhausts clearance before the independent permission path responds; the outcome requires a plant and timing model.↩︎
Kinematic integration of heading bias: This example holds heading at a constant \(1^\circ\) offset, so \(y(t)=v\sin(\epsilon_\theta)t\) is linear. A different example with constant yaw-rate bias \(\omega\) begins with \(\theta(t)=\omega t\) and has the small-angle approximation \(y(t)\approx \tfrac12v\omega t^2\).↩︎
Non-Lipschitz contact bifurcation: Rigid-body contact is not Lipschitz continuous, because friction cones are set-valued and impacts produce velocity jumps (\(v^+ \ne v^-\)). A sub-millimeter deviation in gripper contact location can switch the trajectory from stable force regulation to slip or jamming, a discontinuity that open-loop scoring on recorded states cannot expose.↩︎
Hardware-in-the-loop bus contention: While neural accelerator inference latency may exhibit low variance in isolation, shared PCIe backplanes, direct memory access (DMA) transfers, and CAN bus arbitration introduce variable transport jitter under production workloads. A nominal inference cycle can stretch under memory bus saturation, depleting the phase margin of high-rate control loops. In shadow mode, this latency jitter is logged across real bus traffic without exposing the physical chassis to the resulting trajectory oscillations.↩︎
Physical trial cycle budget: The duration of a physical trial includes the reset as well as the run, \(T_{\text{trial}} = t_{\text{run}} + t_{\text{reset}} =\) 25 s \(+\) 15 s \(=\) 40 s for the door. Budgeting by active run time alone produces unrealistic testing timelines and tempts teams to truncate reset and cooldown intervals, which causes thermal non-stationarity across runs.↩︎
Extreme value disturbance modeling: The Fisher–Tippett–Gnedenko theorem describes the possible limiting forms of block maxima under stated assumptions; it does not assign a tail family to impact, glare, or latency by physical mechanism alone.↩︎
Functional safety scope: Functional safety standards define context-specific hardware metrics and assurance processes; a single failure-rate number cannot be assigned to all physical AI systems or certified by mileage alone. The Poisson examples here are analytical targets, not claims of compliance with IEC 61508 or ISO 26262.↩︎
Sequential Probability Ratio Test (SPRT): Predeclared likelihood-ratio boundaries allow valid sequential stopping for specified hypotheses when their actual error probabilities are calibrated, including boundary overshoot. The number of saved runs is path- and hypothesis-dependent; no universal wear saving follows.↩︎
Statistical Power in Physical Trials: The probability \(1 - \beta\) of correctly rejecting a false null hypothesis. For a two-policy comparison, compute power from the baseline rate, effect size, sample allocation, and chosen test before collecting physical trials (Cohen 1988).↩︎
Safety of the intended functionality (SOTIF): ISO 21448 and ANSI/UL 4600 provide frameworks for arguing about intended-functionality and autonomous-system hazards. This chapter’s immutable record is a proposed engineering evidence artifact, not a specific form mandated by either standard.↩︎


