Teaching the Machine

an offline machine learning benchmark, the dataset is fixed before evaluation. A classifier reads held-out examples drawn from a specified data distribution (\(x \sim P_{\text{data}}\)); its prediction does not change the next stored example. This useful testing assumption says nothing about a deployed model whose outputs affect people, later data, or physical state.

Principle 1: Endogenous data
Invariant: A physical agent’s actions choose its next observations (\(s_{t+1} \sim P(s \mid s_t, a_t)\)), so the states a deployed policy visits need not be the states it was trained or tested on. For behavioral cloning, the worst-case cost of that drift can grow with the square of the horizon (Ross et al. 2011), as equation states.

Implication: A learned proposer’s competence is a claim about a declared region of states, established in closed loop with the permission path in place. States the data never visited are recorded as absent rather than assumed covered, and the finite trials behind a claim bound it rather than prove it.

Ross, Stéphane, Geoffrey Gordon, and Drew Bagnell. 2011. “A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning.” Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, 627–35.

Part II follows a learned proposer, the policy that opens the warehouse mobile manipulator’s spring-latched cage door, from demonstrations recorded through the permission path under the clock domains of the handoff record that Part I built, to the evidence that it works. Each chapter narrows what the next may claim.

  • Physical Data (Physical Data): Decides which conditions the demonstrations support and what they cost to collect, and produces the demonstration provenance record, which lists the conditions the data does not cover as known absences.
  • Policy Training (Policy Training): Decides which training regime fits the data’s support and which operational design domain (ODD) the resulting policy may declare, and produces the policy manifest.
  • Closed-Loop Evaluation (Closed-Loop Evaluation): Decides what a finite set of closed-loop trials supports about the proposer, and produces the evaluation record, whose claim stays inside a target ODD.

Part II hands forward an evaluation record: a claim about the learned proposer confined to a target ODD, the finite-sample bound that supports it, a catalog of the conditions it does not cover, and the runtime monitors the rest of the machine must supply. That record rests on the policy manifest from Policy Training, whose declared ODD cannot exceed the support that the provenance record from Physical Data shows. Part III trusts the proposer no further than that record, and its permission path is built for the states the evaluation left uncovered.

Back to top