Supervisory Intervention
Supervisory Intervention
Purpose
When a human grabs the controls of a running machine, how do you transfer kinetic authority without causing the very collision you meant to prevent?
When a human supervisor asserts control over an autonomous system, the takeover request reaches hardware already in motion under neural control. Naively blending human and automated commands without strict arbitration yields conflicting control inputs, while instantaneous switching risks severe mechanical shock or sudden loss of traction. A safe intervention architecture must explicitly designate kinetic authority, enforce bumpless command transfer, and verify that handovers preserve dynamic feasibility throughout the transition window.
Crucially, human cognitive reaction time and telemetry transmission latency consume the same physical stopping budget as algorithmic delay; an operator cannot serve as a zero-latency safety fallback. When supervisory heartbeats drop or handovers stall, the machine must independently engage deterministic fallbacks. Every intervention must be immutably recorded with microsecond timestamps to separate autonomous failures from operator recoveries. Within the governance envelope, supervisory intervention treats human input as an external proposal that must pass through the deterministic permission path before crossing the causal boundary.
Learning Objectives
- Specify operator permissions that bind authorized actions to scope, deadlines, and authentication requirements
- Calculate the latest fallback time and takeover speed ceilings from reaction latency and remaining stopping clearance
- Design authority transfers that keep one actuator writer and preserve command continuity within physical limits
- Distinguish policy failure, authority transfer, and human recovery in timestamped authority records
- Evaluate whether an authority record can reconstruct command authority and physical outcomes after an incident
Transfer Without Discontinuity
At commit time \(t_0\), the enforcer latches the last admitted policy command \(\mathbf u_0\) and a checked human target \(\mathbf u_1\). A direct switch would request \(\Delta\mathbf u=\mathbf u_1-\mathbf u_0\) in one tick. Its ideal command derivative is discontinuous, but the actual plant filters the command through finite drive, brake, and mechanical dynamics. A large change may excite the arm’s gearbox, exceed the drive wheels’ available traction, or shift the tote on its rack. The transfer therefore checks both rate and full stopping clearance rather than assuming a software tick maps instantaneously into physical force (The Causal Boundary).
Eliminating this step discontinuity requires smoothly mixing the old and new commands across a finite time window \(\tau_{\text{blend}}\), a technique known as bumpless transfer. A linear blend changes slope at its endpoints. For a fixed-endpoint change in commanded acceleration, a \(C^2\) bumpless transfer ramp eases in and out, removing endpoint jumps in the command and its first two derivatives. A standard curve used for this is the quintic smoothstep.
The enforcer blends with the quintic smoothstep, whose coefficients follow from requiring zero first and second derivatives at both ends of the window (Smooth Trajectories and Bumpless Transfer). It is the same \(C^2\) property the planner relies on at its seam, where no acceleration step appears at the switch (Behavior at the Seam); here the switch is between proposers rather than between trajectory segments. The appendix also derives the shortest window that keeps a fixed change in commanded acceleration within a chosen command-jerk limit \(j_{\text{allowable}}\). The same bound applies to a torque command, so an illustrative 15 N·m change in one of the arm’s joint torques under a 750 N·m/s rate ceiling takes at least 37.5 ms.
Increasing \(\tau_{\text{blend}}\) reduces the profile’s peak commanded jerk, but delays the full human target and can lengthen the stopping trajectory. Consider the remote assistant’s takeover of the mobile manipulator’s base, committed while the machine cruises at the \(v =\) 1.2 m/s it slowed to before requesting it (section 1.3), and suppose the assistant commands a stop at the scenario’s braking deceleration, a fixed step of \(|\Delta a| =\) 2 m/s² from the latched cruise command. The enforcer holds the blend to \(j_{\text{allowable}} =\) 8.89 m/s³, the peak jerk \(6v/T_{\text{stop}}^2\) of the resident stop from the same speed (Behavior at the Seam), so the transfer loads the tote rack no more abruptly than the machine’s own stop. The smoothstep bound then sets a minimum blending window of 421.88 ms. The smoothstep weight averages one half over the window, so the base would leave the blend at 0.78 m/s after traveling 455.4 mm, and braking from there at the full target would take another 151.4 mm, for 606.8 mm to rest.
↰ Prerequisite: Acceleration-bounded bumpless transfer profiles build upon the physical stopping constraints enforced in Stopping Envelopes.
An instantaneous step takeover would stop in 360 mm, so bumpless blending imposes a spatial penalty of 246.8 mm.4 The blended stop is also 66.8 mm longer than the 540 mm resident stop from the same speed. A blend therefore spends stopping reserve beyond the stop the budget holds. The enforcer admits each blended tick only while the resident stop from the resulting state still fits, and where no blend fits, the base receives the resident stop itself, the suffix the fallback ladder’s stop rung executes (The Fallback Ladder). Bypassing the blend window by applying the acceleration step within a single 1 ms permission tick would instead request a command jerk of about 2,000 m/s³ over that tick. The idealized step is only a comparison baseline.
Command blending removes the chosen output discontinuity, but a standby human feedback controller can still accumulate integral error while disconnected from the plant. Initial condition matching pre-biases that controller against the measured state before commit, avoiding an idle-integrator transient. It does not force the deliberate human target to equal the last policy command. The enforcer latches both endpoints, sizes the blend for their actual difference, and checks the resulting command and stopping reserve throughout transfer.
To verify a transfer, record the command profile, motor current, encoder motion, and chassis inertial response on a shared timebase with bounded alignment error. A step command has a derivative discontinuity in the ideal model; ringing, high jerk, and slip are possible observations, not inevitable waveforms. A fixed-endpoint quintic bounds the commanded rate, and figure 4 contrasts the two profiles schematically. The measured plant response establishes whether the handover met its physical limits.
The enforcer’s \(1\text{ kHz}\) loop gathers the four-phase handshake, integrator pre-biasing, and the \(C^2\) quintic blend into one protocol on the safety microcontroller. Each tick it samples plant state, stopping reserve, protective inputs, and the authority lease. It commits an ownership change only on the qualified commit tick that follows a verified acknowledgment, blends only between latched endpoints and checks every blended candidate before dispatch, and falls back on a lapsed lease or a failed clearance check.
↳ Protocol: The enforcer’s cadence, memory and write invariants, and per-state execution steps are given in Authority-transfer protocol.
Once the blend completes and the remote assistant holds the arm in Human-Manual, driving it through a force-reflecting leader, the arm must respond either to the operator’s motion or to the operator’s force. Impedance control (Hogan 1985) renders a force from measured motion through the virtual spring-damper of \(\ref{psp-body-impedance-control}\) and suits backdrivable joints. Admittance control measures the operator’s force at a wrist force-torque sensor and yields in motion under stiff position tracking, which suits the high-ratio drives whose reflected rotor inertia \(N^2 J\) resists being moved by hand (Actuator Transmission Limits). Either mode closes a loop through the operator, the mechanism, and the transport delay, and the stability of that loop depends on grip, gains, and delay, so each mode must be tested for passivity over its specified operating range. A passive teleoperator, meaning the leader, link, and arm taken together as seen from the operator’s hand, never returns more energy to the operator than the operator put in. The enforcer keeps its clearance and torque limits whichever mode the leader uses.
Blending, pre-biasing, and compliant interaction all assume that the proposals they shape come from an operator entitled to control the arm, and a force measured at a sensor carries no credential of its own. If the permission path accepts a transfer request without validating identity, an unauthenticated command injection or a malfunctioning subsystem can seize authority and override the machine while it is in motion. The mechanisms that make an intervention smooth must therefore be paired with mechanisms that establish who is allowed to intervene.
Authenticating an Intervention
The remote assistant’s veto_path packet and a forged wheel drive-torque packet can reach the mobile manipulator over the same wireless link in the same control cycle. The permission path must keep the forger from seizing physical authority while still letting an authorized operator halt the machine before collision. These requirements conflict, so intervention operations are classified by the mechanical work they induce. Removing power or asserting a mechanical brake is physically distinct from injecting active motion commands, and the two actions do not warrant identical verification mechanisms. A protective stop is a braking or holding rung of the fallback ladder, not torque removal alone (The Fallback Ladder). In contrast, injecting base velocity, wheel torque, or joint velocity setpoints introduces fresh mechanical power into the system. An unauthorized or corrupted motion command can accelerate a moving mass into an obstacle, whereas a trusted out-of-band protective input requests the plant-specific minimum-risk stop.
Motion commands introduce active kinetic hazards; thus, the verification mechanism must meet both latency and freshness bounds from the kinematics of the maneuver. Verifying the signature and checking the replay window of a network motion or override command is one more delay before braking can begin. It adds to the actuator onset \(T_{\text{act}}\), the bus delay \(T_{\text{bus}}\), and the other terms of \(\tau_{\text{delay}}\) (equation), so for a static obstacle at clear distance \(D_{\text{clear}}\) and forward speed \(v\) the maximum permissible verification latency \(T_{\text{auth,max}}\) is the handover deadline itself, \(T_{\text{latest}}\) of equation 2.
On the mobile manipulator at its 1.3 m/s aisle speed, that bound is 78.9 ms. Assume two illustrative verifiers on a microcontroller, one slow and one fast:
- Verifier A (slow): Verification takes \(t_{\text{auth,A}} =\) 180 ms, more than the deadline. The machine travels 234 mm while the signature is checked, 131.4 mm more than the spare, so its stop ends that far past the rack-end clear distance.
- Verifier B (fast): Verification takes \(t_{\text{auth,B}} =\) 15 ms, well inside the deadline. The machine travels 19.5 mm during the check and keeps 83.1 mm of the spare.
CAN bus priority: A transmitting CAN frame is non-preemptive. A higher-priority identifier wins only at the next arbitration opportunity; controller buffer policy and worst-case current-frame serialization must be included in the measured delivery bound. An independent protective input avoids this shared queue where the safety case requires it.
Freshness requirements arise from the same kinematics. If an authentication window permits a timestamp drift or message age threshold of \(\tau_{\text{fresh}} =\) 100 ms, an authenticated packet replayed from the boundary of that window commands an action based on a base position displaced by \(\Delta x = v\,\tau_{\text{fresh}} =\) 130 mm at the aisle speed, more than the whole 102.6 mm spare.
A trusted, independently wired protective input is evaluated by its safety circuit without waiting for a network signature, and its response is the stop or hold selected and validated for the plant. Remote motion or takeover packets follow the capability and authentication checks. A corrupted checksum, stale nonce, or failed signature on that network is rejected, logged, and handled as rule A4 of section 1.8 specifies; malformed traffic does not acquire the authority to command a brake. If valid command renewal is lost, the separately monitored lease expires and the enforcer invokes its validated fallback. This separation preserves an available local protective path while preventing arbitrary network noise from forcing an unsafe stop.
Authentication and transport delay also reach training. A training simulator that always supplies immediate human rescue can underprice risky states. If deployed intervention instead incurs authentication, transport, and actuation delay, a policy trained on that mismatch may visit states from which rescue is no longer feasible. The remedy is to model the qualified latency and fallback path in training and hardware-in-the-loop evaluation, retain failed learner-visited states, and test the actual clearance distribution.
↰ Prerequisite: Microsecond override masking registers on real-time fieldbuses are implemented in Actuation Authority.
The authority inventory must include diagnostic, debug, bootloader, and wiring paths that can affect actuation. A production design may disable debug access, authenticate maintenance commands, protect boot updates, and enforce bus permissions according to its threat and fault model. The verification task is to show that each reachable actuator-write path is either gated by the permission authority or independently protected. The argument of this section covers remote packet injection over wireless links, replay on internal buses, and commands from a compromised application processor. Physical tampering lies outside it, and whether the safety case must answer tampering depends on the threat model, inspection access, and packaging.
Authentication establishes who is entitled to take over. It does not ensure that the handover leaves the machine under coherent control. When authority transfers while the base carries momentum and the arm’s drives deliver torque, the policy, the operator, and the enforcer must agree on who holds each channel. Handover goes wrong when that agreement breaks.
How Handover Goes Wrong
A handover can fail while every component works. When an autonomous policy and a human operator share authority, it can fail through mode confusion and authority ambiguity, where the machine assumes the human is managing an operational axis while the human assumes the machine retains jurisdiction. In multi-axis physical systems, authority is frequently decomposed across channels, and on the mobile manipulator it splits between the base and the arm. If the learned policy encounters an unmodeled surface reflection and quietly yields the arm to the remote assistant without releasing the base, an assistant who guides the arm may assume the base is holding still while the policy is still driving it toward the rack end. Conversely, if the assistant commands the base to stop, firmware that transfers only the base leaves the arm policy still reaching for the tote against the assistant’s intent. In telemetry logs, this failure presents as disjointed control flags where authority bits toggle asynchronously across execution threads. At the human-machine interface, however, this state is undetected until kinematic error exceeds the physical clearance margin, because neither claimant receives an immediate, unambiguous signal of the other’s partial withdrawal. Per-channel ownership permits such a partial transfer, so the protocol must make it explicit. Every request names the channels it covers in its channel mask, the state machine of section 1.3 commits owner states only for those channels, and the Confirm phase announces the owner of each channel to the operator, so a base that stays with the policy is shown as the policy’s.
Even with defined authority transfer protocols, the human nervous system cannot match the response deadlines of moving mass. A remote assistant in direct control of the mobile manipulator answers within the 150 ms in-loop takeover of section 1.2. When the same assistant supervises several machines passively while their policies run, situational awareness decays, and an unexpected handover request meets the 2 s out-of-loop takeover instead. As section 1.2 showed, a machine left at aisle speed covers more than twice the rack-end clear distance during this gap, before the assistant has perceived the obstacle, re-established spatial orientation, and applied input. Operator-facing cameras can track gaze direction and eye closure. They cannot measure comprehension of the scene geometry, so the growth in latency is detected only after the response deadline has passed.
At Tempe, vigilance decay met a disabled stop path (When the Boundary Fails) (National Transportation Safety Board 2019). In the Cruise incident, a vehicle that had stopped after striking a pedestrian began a pull-over and dragged her beneath it (The Fallback Ladder) (Quinn Emanuel Urquhart and Sullivan, LLP 2024). The first asks whether an operator-monitoring contract backed by an independent stop would have stopped the vehicle within the remaining clearance, and the second whether contact uncertainty should revoke motion authority until a validated hold is established.
When an intervention occurs while an actuator applies force, manual input and the enforcer can oppose each other. For example, the remote assistant driving the arm under manual_drive through a force-reflecting leader may command a joint toward the position limit at which the enforcer is holding it. The torque the assistant’s command requests and the drive current holding that joint then act against each other on the same link. If the human-device loop has insufficient phase margin, repeated corrections may oscillate at a frequency set by the gains, delay, grip, and joint mechanics. The enforcer must retain the physical limit while the leader makes the opposing force and active mode legible to the assistant.
Uncoordinated handovers also corrupt the data drawn from fleet logs. If intervention segments enter behavioral cloning or diffusion training without intervention masking, which uses the logged commit instant to separate the policy’s own states from the recovery, the policy learns emergency recovery as nominal behavior, and validation loss on trajectory chunks does not reveal the corruption.
The arbiter of section 1.3 exists to prevent these failures, because the control architecture cannot rely on implicit handoffs, human attentiveness, or ad hoc torque overrides. Showing afterward that it did requires a record kept at every tick, channel by channel. Deciding afterward whether a stop answered a genuine boundary violation, an invalid token, or noise on a safety line requires that per-tick record to survive power cycles and scheduling stalls.
Trustworthy Operational Logs
A record supporting an argument months later has requirements that a debug log lacks. A debug log may be buffered in volatile user space and dropped under backpressure. If an asynchronous disk write drops three lines of trace data during an operating system stall, an engineer loses minor diagnostic context while the application continues running. Missing records weaken reconstruction across the causal boundary established in The Causal Boundary. When the arm strikes the tote rack, establishing whether the learned policy commanded an illegal torque, whether the enforcer failed to clamp it, or whether an operator intervened too late requires a bounded, integrity-checked record plus independent plant observations.5 A forensic log therefore does not exist to assist a developer with line-by-line code tracing. Its purpose is to preserve what was sampled, proposed, admitted, and dispatched. Motor current, encoder motion, and contact-force sensors provide separate physical observations; commands alone do not prove delivered force or collision cause.
The chosen \(1\text{ kHz}\) control design records a structured authority and action tuple each tick, with declared loss handling. The minimal evidence tuple takes the form \(\langle t_{\text{MCU}}, S_{\text{auth}}, \mathbf{x}_{\text{state}}, \mathbf{u}_{\text{prop}}, \mathbf{u}_{\text{enf}}, \mathbf{u}_{\text{act}}, \text{ID}_{\text{operator}}\rangle\), which records both sides of the proposal boundary (The Machine in Five Levels). In this structure, \(t_{\text{MCU}}\) is the event time in nanoseconds on the safety microcontroller’s clock, converted from any other clock with a recorded error bound, \(S_{\text{auth}}\) records the active authority state indicating whether autonomous policy, manual override, or fault mitigation holds control, and \(\text{ID}_{\text{operator}}\) identifies the authenticated source of any active intervention. The three action vectors map onto the action taps of Which Action Becomes the Label, with \(\mathbf{u}_{\text{prop}}\) as the requested action \(a_{\text{req}}\) and \(\mathbf{u}_{\text{enf}}\) as the enforced command \(a_{\text{enf}}\). The mapped command \(a_{\text{cmd}}\) lies between them and needs no field of its own when the proposal already arrives in actuator coordinates. The field \(\mathbf{u}_{\text{act}}\) is the command dispatched to the drives, and the measured response \(a_{\text{meas}}\) comes from the plant observations in \(\mathbf{x}_{\text{state}}\). Omitting one element can limit a particular causal question, so the record declares what each field actually measures. If the log records \(\mathbf{u}_{\text{act}}\) without \(\mathbf{u}_{\text{prop}}\), investigators lose the policy-command comparison. If \(\mathbf{u}_{\text{enf}}\) is omitted, the log loses the proposed-versus-admitted comparison. Chained with integrity and clock-conversion metadata, these tuples form the forensic authority record that section 1.8 specifies.
Ordering these tuples requires physical timestamp integrity rather than software clock queries. Software timestamps derived from operating system calls inherit scheduling jitter, context-switch delays, and variable interrupt servicing latencies that can exceed \(5\text{ ms}\) on loaded application processors. Over a \(5\text{ ms}\) interval, an internal control loop running at \(1\text{ kHz}\) executes five complete control cycles. If an operating system logs the arrival of a manual intervention command with a \(5\text{ ms}\) delay, the record places the operator’s command after the actuator response, inverting physical cause and effect. A hardware timestamp records a defined electrical event, such as frame ingress or a protective-input edge, after clock conversion with bounded error. Investigators can order command and detected contact only when their timestamp uncertainty, sampling windows, and sensor delays do not overlap.
Persisting full sensor streams and high-rate telemetry continuously to non-volatile storage exceeds both the bandwidth and the endurance limits of embedded hardware. The four cameras of a mobile-manipulator teleoperation session stream raw frames at the ingestion rate that Teleoperation Demonstration Costs sizes, a rate only a Gen4 NVMe drive sustains, so a pre-trigger and a post-trigger segment of those frames far exceed any microcontroller’s memory. Raw frames do not belong in an MCU SRAM ring. The \(1\text{ kHz}\) control tuple is a small fraction of that stream and can use a bounded small-memory ring. If full raw frames are required, provision a separate DRAM or storage tier sized to the measured stream budget, or retain selected or compressed frames with a declared loss of detail. At a trigger, seal the pre-trigger segment and continue writing post-trigger samples into a different allocated segment; freezing the sole ring would lose the promised post-event evidence. Frame records retain their native capture times, while control tuples retain their \(1\text{ kHz}\) issue and sample times. Background persistence must not block the enforcer.
The pre-trigger window must capture the moment of policy divergence rather than merely the moment of physical takeover. When a learned policy commands an erroneous trajectory, human operators and external monitoring processes exhibit reaction latencies ranging from hundreds of milliseconds to multiple seconds before recognizing the deviation and asserting authority. If the ring buffer only recorded data from the instant an operator pulled an override switch, the telemetry would capture the rescue while omitting the anomalous commands that made the rescue necessary. Capturing a declared pre-trigger interval lets investigators inspect policy proposals and states before intervention; whether a proposal was unsafe needs the contemporaneous model, state estimate, and physical evidence. Without the pre-divergence history, evaluation cannot separate a necessary rescue from a premature interruption, and the learner-visited states that corrective relabeling needs are gone.
Integrity and retention can use a protected hash chain and independently tested power-loss storage. Within the Nervous System, logging routines batch evidence tuples into fixed blocks corresponding to \(100\text{ ms}\) of execution. Each block is hashed using SHA-256, and its hash is concatenated with the hash of the preceding block to form an append-only cryptographic chain. With protected key and anchor custody, changing a committed block is detectable against the retained chain; loss before commit and compromised keys remain separate failure modes. One implementation periodically anchors root hashes in protected nonvolatile storage. A hash chain supplies neither persistence nor complete event capture, so the safety case tests key custody, which committed blocks survive power loss, how interrupted writes are marked, and where intervals are missing.
A retained, integrity-checked log can establish recorded authority and command values for its sampled intervals. It cannot say whether those values were permitted. That judgment needs the manifest the log is read against, which names who may hold each channel, under which conditions, and within which deadlines, and it needs the log’s entries to carry the fields the manifest’s rules refer to.
Fallacies and Pitfalls
Changing an authority register does not complete a safe handover. The machine must preserve its stopping margin and enforcement functions through the transfer, while the resulting logs must distinguish intervention from expert recovery.
Pitfall: Treating human takeover timeout as a static time delay rather than a dynamic spatial stopping budget.
The mobile manipulator requests a takeover at its 1.3 m/s aisle speed under a configured 500 ms acknowledgment window. Waiting out the window before braking covers 0.65 m, and with the rest of the finished stopping budget the stop ends 0.55 m past the 1.10 m rack-end clear distance. The same window fits at the 0.3 m/s crawl, where the budget leaves 2.8 s before fallback must begin. A window that is safe in one state overruns in another, so handover timing must follow current speed, clearance, and credible deceleration, with autonomous fallback beginning while the stop still fits.
Fallacy: Human manual takeover should bypass automated safety barriers and grant unconstrained actuator authority.
The remote assistant takes over the mobile manipulator’s base to route it around a pallet, and the chosen path leads toward a rack upright. Switching to human control only changes who is sending the command; it does not change the physical realities of momentum, friction, or nearby obstacles. Within this architecture, the enforcer must continue checking the manual proposal against the admissible motion set and constrain it when a kinodynamically feasible alternative exists (building on the feasibility principles from Kinodynamic Feasibility). If the request cannot be admitted, the defined autonomous fallback still applies. Takeover lets the operator guide the task planning; it does not make force, clearance, and stopping limits optional at the joint actuator boundary.
Pitfall: Relying on cryptographic authentication without rate limiting the traffic that reaches the verifier.
The mobile manipulator’s safety microcontroller checks each remote takeover packet with the fast verifier of section 1.5, at 15 ms per signature. Verification rejects a forged packet only after spending that time on it. If a faulty or compromised node on the link sends manual_drive packets that fail verification faster than the verifier retires them, as few as 5 of them queued ahead of the remote assistant’s genuine veto_path delay its admission to 90 ms, past the 78.9 ms fallback deadline \(T_{\text{latest}}\) at the aisle speed. The independently wired protective input does not wait in that queue, but every network command does. Authentication preserves command integrity while leaving an availability hazard. Mailbox filtering and per-source rate limiting must bound the traffic before it reaches the verifier, and qualification must test both command rejection and enforcement timing under the admitted load.
Fallacy: Training directly on raw human intervention logs captures expert demonstrations.
A learning pipeline treats every remote-assistant takeover of the mobile manipulator as an expert demonstration and trains on the preceding action window. The logs combine genuine policy failures, cautious early interventions, and operator kinematic mistakes, while the control transition itself often includes jerky physical corrections, delayed neuromuscular responses, and hesitant corrections as the human orients to the scene. Cloning those segments can reproduce the unstable behavior that made recovery difficult rather than the recovery the team wants to teach. Dataset curation must identify why intervention occurred, who held the authority register, and which commands actually restored acceptable operation. Reviewed recovery trajectories can supply useful demonstrations, but neither an intervention timestamp nor full manual authority alone establishes that the recorded behavior is suitable for imitation.
Summary
Human authority over an embodied physical AI system needs an engineered real-time contract that arbitrates takeover, revokes autonomous policy authority at commit, and selects a validated fallback within the available physical budget. Supervisory intervention governs the transition between proposers while the independent enforcer checks the admitted command and executes the selected stop when a lease lapses. Network commands reach the enforcer only after capability and authentication checks, whose own cost makes rate limiting ahead of the verifier part of the design, while trusted protective wiring bypasses the network entirely. On the mobile manipulator at its 1.3 m/s aisle speed, the finished stopping budget leaves 102.6 mm, which buys 78.9 ms of waiting before the machine must dispatch its own stop. The machine therefore slows to 1.2 m/s before it requests an in-loop takeover, and to a 0.3 m/s crawl or a standstill before it requests an out-of-loop one.
Transferring control from an active policy takes time, because replacing an admitted policy setpoint with a different human target in one \(1\text{ ms}\) cycle requests a torque step. The \(C^2\) bumpless authority transfer ramp instead varies an authority parameter \(\alpha(t)\) smoothly from zero to one between latched command endpoints. Its sampled implementation must respect command-rate limits and the stopping reserve of Stopping Envelopes.
On the mobile manipulator the enforcer blends the base only from the 0.3 m/s crawl, over 106 ms. At the 1.2 m/s in-loop takeover speed no blended base command fits the clearance that remains, so the base receives the resident stop while the arm blends over 38 ms. The enforcer admits each blended tick only with a validated stopping reserve and monitors the plant response.
The authority record makes each intervention reviewable. Its static manifest binds every actuator channel to its actors, tiers, and deadlines, and its per-tick entries log who held authority and what was proposed, admitted, and delivered, naming fallback rungs from the enforcement record and operator-path timing from the placement record. A logged commit marks the recorded boundary where autonomous policy authority was revoked, subject to the logger’s clock and sampling limits, and hash chaining with protected keys can reveal later modification. Command and authority records support reconstruction only when paired with independent sensor, current, and contact evidence and their timing uncertainty.
For model training, the pre-commit interval is a policy-owned segment even when a human is deciding to intervene. Retain its learner-visited states, unsafe proposals, and authority flags as failure evidence and candidates for expert relabeling; exclude unsafe policy commands from expert imitation targets. The post-commit blend is enforcer-owned transition data. Only reviewed post-transfer human recovery states and actions that meet the task’s validity criteria should become demonstrations.
Key Takeaways: Supervisory intervention and bounded authority transfer
- Clearance governs waiting: At the aisle speed the machine’s spare margin buys 78.9 ms; at the 1.2 m/s slow-down it buys the 174.7 ms that governs every in-loop takeover. A 2 s out-of-loop takeover needs a speed below 0.40 m/s. Dispatch fallback by the earlier of the configuration ceiling and the state-dependent stopping deadline.
- One actuator writer: Commit revokes the policy’s token on the committed channels and latches its last admitted command. The enforcer owns
BLENDand checks human proposals against the same physical limits. - Smooth commands cost stopping distance: A fixed-endpoint \(C^2\) acceleration profile bounds commanded jerk. Its extra stopping distance must be reserved.
- Distinct stop paths: Trusted protective wiring invokes a risk-assessed plant response. Invalid network packets are rejected; lost valid renewal invokes the lease fallback.
- Causal intervention data: The logged commit separates the policy-owned states before takeover, which corrective relabeling needs (Intervention and Recovery Data), from the human recovery after it. Use only reviewed post-transfer recovery actions as imitation targets.
- Authority becomes evidence: The per-tick record counts a missing entry as a gap, never a pass. The manifest’s four authority rules become the fault cases verification injects, and their passing records, with the authority record, are the evidence for release claim C6.
A person’s authority over the machine is granted on evidence and expires with it, as principle \(\ref{pri-vol4-evidence-bounds-authority}\) requires of every authority the machine honors. The credential and the committed handshake are the evidence that grants it; the manual lease is that evidence’s age limit, and when the operator’s input stops renewing it, the authority lapses to a fallback the permission path owns. A preemption tier does not survive the lapse, and the remote assistant who held the base before it holds no authority after it. Returning control to any proposer waits on the reset conditions the state machine specifies, and on fresh evidence.
What’s Next: From authority revocation to physical verification
Footnotes
Supervisory Control Levels: Sheridan and Verplank (1978) formulated a ten-level taxonomy of supervisory control, running from a human who performs the whole task without computer assistance to a computer that decides and acts while ignoring the human. A capability matrix draws such distinctions per operation rather than once for the whole machine.↩︎
Three-Position Enabling Devices: The cited ABB operating manual documents a three-position enabling grip and a \(250\text{ mm/s}\) manual-mode limit for that robot. Releasing or over-squeezing the grip is a request into the configured safety circuit, and the machine’s risk assessment under ISO 13849-1 sets the required performance level and the stop the circuit selects.↩︎
Fieldbus Write Exclusivity: Authority arbitration requires hardware-enforced exclusivity across distributed fieldbuses. Dual-writing to CAN-FD or EtherCAT segments from autonomous and manual controllers induces non-deterministic actuator chatter and inverter current oscillations. A qualified arbiter and bus permissions must prevent two software sources from writing one actuator channel.↩︎
Bumpless Transfer Braking Penalty: The penalty belongs to the smoothstep profile, its fixed endpoints, and the stated acceleration step; another profile or endpoint pair gives a different penalty.↩︎
Precision Time Protocol Logging: Unlike legacy event data recorders logging at coarse \(10\text{ Hz}\) intervals, physical AI incident reconstruction demands sub-millisecond precision. A qualified PTP installation can align the named device clocks with a measured conversion-error bound. Event ordering requires the intervals around timestamps, sensor exposure and sampling, and actuator response not to overlap.↩︎

