Supervisory Intervention

Supervisory Intervention

Isometric blueprint showing supervisory intervention on an autonomous vehicle steer-by-wire chassis: independent double-wishbone suspension, vehicle perception mast, steer-by-wire rack and pinion, human steering wheel override, glowing kinetic authority clutch, timestamped intervention relay, and tire contact friction ellipses.

Purpose

When a human grabs the controls of a running machine, how do you transfer kinetic authority without causing the very collision you meant to prevent?

When a human supervisor asserts control over an autonomous system, the takeover request reaches hardware already in motion under neural control. Naively blending human and automated commands without strict arbitration yields conflicting control inputs, while instantaneous switching risks severe mechanical shock or sudden loss of traction. A safe intervention architecture must explicitly designate kinetic authority, enforce bumpless command transfer, and verify that handovers preserve dynamic feasibility throughout the transition window.

Crucially, human cognitive reaction time and telemetry transmission latency consume the same physical stopping budget as algorithmic delay; an operator cannot serve as a zero-latency safety fallback. When supervisory heartbeats drop or handovers stall, the machine must independently engage deterministic fallbacks. Every intervention must be immutably recorded with microsecond timestamps to separate autonomous failures from operator recoveries. Within the governance envelope, supervisory intervention treats human input as an external proposal that must pass through the deterministic permission path before crossing the causal boundary.

Learning Objectives
  • Specify operator permissions that bind authorized actions to scope, deadlines, and authentication requirements
  • Calculate the latest fallback time and takeover speed ceilings from reaction latency and remaining stopping clearance
  • Design authority transfers that keep one actuator writer and preserve command continuity within physical limits
  • Distinguish policy failure, authority transfer, and human recovery in timestamped authority records
  • Evaluate whether an authority record can reconstruct command authority and physical outcomes after an incident

Shared Authority

A human supervisor is often drawn as one arrow into the plant. When a pallet blocks an aisle, the warehouse mobile manipulator needs a remote assistant to take over, yet at its chosen 1.3 m/s aisle speed a 2 s pause (illustrative; see the Reader Guide) before braking would carry it 2.6 m, more than twice its 1.10 m rack-end clear distance. An operator request reaches a machine that is already moving under an admitted command; it cannot erase momentum or make an actuator respond instantaneously. The arbiter, the permission path’s source-selection step, must decide who may propose the next command, and the enforcer, its admission check, must admit or reject it; both run on the safety microcontroller (MCU). The plant must still have enough clearance to execute the chosen stop. An abrupt setpoint switch requests a torque step that the drives and the load must absorb. A delayed switch can exhaust stopping room even when both controllers are individually well behaved.

↰ Prerequisite: The proposal boundary of The Machine in Five Levels, and the third law in The Four Bedrock Laws that it enforces, set the baseline for supervisory preemption.

Bainbridge, Lisanne. 1983. “Ironies of Automation.” Automatica 19 (6): 775–79. https://doi.org/10.1016/0005-1098(83)90046-8.

Bainbridge’s classic ironies of automation (Bainbridge 1983) observe that an autonomous system tends to leave human operators responsible for the rarest, most complex edge cases while stripping them of the continuous manual engagement necessary to maintain situational awareness. When a remote assistant takes over the mobile manipulator in a blocked aisle, supervisory authority is effective only if it maps to an executable physical response bounded by actuator torque and thermal derating limits. Taking control of a machine already in motion forces the system to resolve two tightly coupled challenges simultaneously: discrete state arbitration without ambiguity in software, and continuous trajectory transition without mechanical shock in the plant. In discrete software, the architecture must prevent split-brain execution where the learned policy and the human operator issue concurrent, conflicting torque setpoints. In this chapter’s \(1\text{ kHz}\) controller, the independent enforcer revokes the policy’s execution token, its right to have proposals admitted, and latches its last approved reference at commit, then remains the sole actuator writer during the subsequent blend. In continuous dynamics, the permission path must execute a bumpless transfer, blending between latched command endpoints, matching initial velocities, ramping each command within its rate ceiling, and checking measured plant limits during the transition.

Intervention events also reach the machine learning lifecycle. A rescued episode logged without its handover instant looks like an expert demonstration of the very near-failure states that made the rescue necessary. Intervention and Recovery Data shows how corrective aggregation instead labels the learner’s own states, which requires the intervention record to keep the trace from before the takeover. The forensic authority record therefore timestamps every authority transition on the safety microcontroller’s clock. The commit instant marks the boundary between policy-owned failure evidence and human recovery.

The first design decision is operational. It fixes which named intervention may reach which physical channel, and how long the machine can wait for it.

Human Authority Over Actuation

The remote assistant who clears a blocked aisle and the coworker who takes a mug from the gripper act on the same machine, and neither should hold the other’s authority. Authority cannot attach to a generic human role or abstract supervision block. It must attach to specific named operations with explicit timing budgets, physical degrees of freedom, and deterministic preemption rules. Rather than granting an operator broad administrative authority, the permission path maintains an explicit capability matrix, a permissions table mapping individual operations (such as an emergency power cutoff, a velocity trim adjustment, a trajectory path veto, or manual drive of the base and arm) to authenticated roles, reflecting classic taxonomies of human supervisory control1 (Sheridan 1992) and multi-level automation (Parasuraman et al. 2000). Each entry in this matrix defines the exact actuator channels the operation governs, the preemption tier it carries over the learned policy, and the maximum propagation latency permitted from physical input to actuator response.

Sheridan, Thomas B. 1992. Telerobotics, Automation, and Human Supervisory Control. Wiley.
Parasuraman, Raja, Thomas B. Sheridan, and Christopher D. Wickens. 2000. “A Model for Types and Levels of Human Interaction with Automation.” IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans 30 (3): 286–97. https://doi.org/10.1109/3468.844354.
Ames, Aaron D, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada. 2019. “Control Barrier Functions: Theory and Applications.” European Control Conference (ECC), 3420–31.

Naming the operations separates intervention from safety enforcement, two mechanisms that system architectures routinely conflate. Enforcement withholds permission from a candidate action proposed by a learned model, modifying, clamping, or rejecting the setpoint while leaving the machine under autonomous control. Intervention transfers authority away from the model to an external entity, terminating autonomous operation and giving the supervisor the right to propose, while every command still passes the permission path. An enforcement filter sits in the fast control path to prevent the policy from demanding an over-current maneuver or violating a barrier certificate (Ames et al. 2019) during nominal tracking, respecting the actuator thermal limits established in The Five Physical Budgets and the kinodynamic feasibility boundaries from Kinodynamic Feasibility. An intervention mechanism strips the policy of its execution token when an operational boundary is breached or an operator demands control. A system equipped with enforcement but lacking intervention cannot be redirected when human intent changes or external conditions fall outside the policy design envelope. Conversely, a system equipped with intervention but lacking enforcement permits a lagging operator to drive the plant into structural failure during nominal manual operation.

For example, the remote assistant who supervises the mobile manipulator may hold a signed trajectory veto capability for the base’s planner channel: role remote_assistant, operation veto_path, channel planner_proposal, preemption tier above the learned planner but below the local enforcer, authenticated session and nonce, and a clearance-derived acknowledgment deadline. The enforcer responds to a valid veto with a checked local hold or stop. The same session carries a second entry, manual_drive, on the base-velocity and arm-joint channels, which becomes active only through a handover that commits within the clearance-derived deadline, and lapses with the manual lease of section 1.3; every command it carries still passes the permission check. A wheel drive-torque packet from that credential is denied, because neither entry grants the base’s wheel drive-torque channel. The coworker at the packing station holds a different entry, accept_item, which lets the coworker take from the gripper the mug the arm picked from the takeaway conveyor (Grounded Intent, Trajectory Planning). That entry grants no motion authority. The item-handover channel of the enforcement record (The Enforcement Record) caps the arm’s approach to the coworker’s hand at 0.10 m/s, under the 0.118 m/s contact ceiling, and no capability lifts that cap. Each capability entry binds actor, operation, channel, deadline, authentication, preemption tier, and fallback; a valid signature alone grants no motion authority.

When an operator attempts an intervention, human biology imposes a hard real-time latency budget before any physical corrective action can occur. Human response time is not a single point event, but a cascaded pipeline of physiological and mechanical stages. For an operator monitoring a machine’s telemetry stream, the sensory recognition stage covers the time for light to strike the retina, travel through the visual pathway, and register as an optical anomaly. The situational assessment stage follows, in which the operator recognizes that the learned policy is drifting toward a hazard and selects an appropriate counter-action. Finally, the motor actuation stage covers the transmission of motor signals down the spinal cord, muscle recruitment in the hand, and the physical displacement of an override switch. Summing the durations of these sequential stages establishes the physiological takeover reaction floor. Accounting for network transport and communication handshake latency \(T_{\text{comm}}\), total supervisory takeover latency \(T_{\Sigma}\) decomposes into: \[T_{\Sigma} = T_{\text{sensory}} + T_{\text{cognitive}} + T_{\text{motor}} + T_{\text{comm}} \tag{1}\]

During this reaction interval the machine keeps moving under the command it was last admitted, before the intervention signal reaches the arbiter, so adding the takeover latency \(T_\Sigma\) to \(\tau_{\text{delay}}\) in equation gives the distance a handover consumes. Because reaction travel grows with both the delay and the speed, a delay that fits at a crawl can consume the whole clearance at aisle speed.

↰ Prerequisite: The stopping-distance equation whose pre-brake delay a takeover lengthens is derived in Kinetic Momentum.

The finished stopping budget of Stopping Envelopes (The warehouse mobile manipulator's stopping budget) shows what equation 1 costs on the mobile manipulator. At the 1.3 m/s aisle speed it leaves 102.6 mm of the 1.10 m rack-end clear distance. A takeover delay adds to the machine’s own pre-brake delay \(\tau_{\text{delay}}\), and every millisecond of it becomes travel at the aisle speed:

  1. Mode 1: Direct Bilateral Teleoperation: An in-loop assistant with a haptic channel and an active proprioceptive loop responds within an illustrative \(T_{\Sigma,1} =\) 150 ms, bus transit included. The machine travels 195 mm during that delay, already more than the spare margin, so at the aisle speed even with this takeover the machine’s stop would end past the rack-end clear distance. Solving the budget for speed with \(T_{\Sigma,1}\) added to \(\tau_{\text{delay}}\) bounds the machine to 1.22 m/s.
  2. Mode 2: Passive Out-of-the-Loop Supervisory Takeover: When an assistant monitors the telemetry of several machines passively, sensory recognition consumes \(T_{\text{sensory}} =\) 300 ms, cognitive situational assessment requires \(T_{\text{cognitive}} =\) 1,200 ms, motor switch actuation takes \(T_{\text{motor}} =\) 300 ms, and network transport and handshake protocol overhead adds \(T_{\text{comm}} =\) 200 ms. By equation 1, the takeover delay expands to \(T_{\Sigma,2} =\) 2 s, during which the machine covers 2.60 m, more than twice the clear distance, before any brake command. The same solution bounds the machine to 0.40 m/s.

In Mode 2 at the aisle speed, the machine would reach a static obstacle at the rack end at full speed, because contact occurs before braking ever begins. Human re-engagement latency therefore sets a ceiling on operating speed that no higher-gain actuator or brake pad can lift, and section 1.3 turns these two ceilings into the speeds at which the machine may request each kind of takeover.

Napkin Math 1.1: Takeover budget and clearance margin
Concept: The stopping distance during a human intervention adds every source of delay and distance before the machine comes to rest. On the mobile manipulator there are four terms:

  1. Machine blind travel: The distance covered during the machine’s own pre-brake delay (\(v\,\tau_{\text{delay}}\)), from The warehouse mobile manipulator's stopping budget.
  2. Takeover travel: The distance covered while the operator perceives the obstacle, decides, acts, and the command crosses the network (\(v\,T_{\Sigma}\)).
  3. Resident stop: The \(C^2\) stop suffix of Behavior at the Seam, \(v^2/(2a_{\text{eff}})\) with \(a_{\text{eff}} = a_{\text{brake}}/1.5\).
  4. Fixed overheads: \(\delta = \delta_{\text{loc}} + \delta_{\text{margin}} + \epsilon_{\text{track}}\).

Worked Comparison: At the aisle speed \(v =\) 1.3 m/s against a rack-end clear distance of 1.10 m:

  • Direct Bilateral Teleoperation (\(T_{\Sigma,1} =\) 150 ms): \[d_{\text{stop},1} = v\,\tau_{\text{delay}} + v\,T_{\Sigma,1} + \frac{v^2}{2a_{\text{eff}}} + \delta = 173.68 + 195 + 633.75 + 190 = 1192.43\text{ mm}\]
  • Out-of-the-Loop Takeover (\(T_{\Sigma,2} =\) 2 s): \[d_{\text{stop},2} = v\,\tau_{\text{delay}} + v\,T_{\Sigma,2} + \frac{v^2}{2a_{\text{eff}}} + \delta = 173.68 + 2600 + 633.75 + 190 = 3597.43\text{ mm}\]

Systems Insight: Passive monitoring inflates the takeover delay from 150 ms to 2 s and grows the stop from 1.19 m to 3.60 m, 3.0× longer. Neither fits the 1.10 m clear distance at the aisle speed, which is why the machine slows before it asks for either one. For the kinematic derivation of these stopping distances, see Kinematic Stopping Envelopes and Information Age Lag.

The physical tension between human reaction latency and stopping distance extends across every mobile embodied system, from indoor warehouse manipulators to highway-speed autonomous vehicles (figure 1). Stopping distance decomposes into two fundamentally distinct kinematic regimes: linear reaction drift distance \(d_{\mathrm{drift}} = v \cdot (T_{\Sigma} + \tau_{\mathrm{delay}})\) during which the machine travels unbraked under sensory, cognitive, and network delay, followed by quadratic kinetic dissipation \(d_{\mathrm{brake}} = v^2 / (2 a_{\mathrm{eff}})\). Because reaction drift scales linearly with velocity while braking scales quadratically, an out-of-loop supervisor monitoring passively from a remote console (\(T_{\Sigma} \approx 2.5\ \text{s}\)) accumulates over 72 meters of blind drift at 100 km/h before the physical brake pads engage. When evaluated against a fixed forward visibility clearance (\(d_{\mathrm{clear}} = 60\ \text{m}\)), the remaining clearance margin \(M = d_{\mathrm{clear}} - d_{\mathrm{stop}}\) defines a hard operational ceiling. Beyond critical velocity thresholds—ranging from 33.7 km/h for remote teleoperation to 78.3 km/h for an alert in-loop operator—the clearance margin becomes negative (\(M < 0\)), rendering high-speed collision physically inevitable regardless of braking force.

Figure 1: Human supervisory takeover reaction budget versus stopping clearance margin: Two-panel kinematic analysis across vehicle velocity under four supervisory engagement modes (\(T_\Sigma \in [0.8, 1.5, 2.5, 4.0]\ \text{s}\)) with 100 ms onboard bus delay and \(a_{\mathrm{eff}} = 6.0\ \text{m/s}^2\) braking. Left: Stopping distance decomposition into linear reaction drift and quadratic braking distance against urban (30 m), suburban (60 m), and highway (100 m) headways. Right: Remaining clearance margin \(M = d_{\mathrm{clear}} - d_{\mathrm{stop}}\) under a 60 m forward sightline. Intercepts denote critical speed ceilings beyond which collision is physically inevitable before coming to rest.

Table 1 extends the two takeover modes to the other ways a site might let a person intervene, each with an illustrative delay. Each row adds its delay to \(\tau_{\text{delay}}\) and solves the finished stopping budget for the largest speed at which the machine still stops inside the rack-end clear distance. The ceiling falls from 1.22 m/s for direct bilateral teleoperation to 0.40 m/s for an out-of-loop takeover, and no mode fits at the 1.3 m/s aisle speed. The last column names the local protection that must hold the machine safe while that delay runs.

Table 1: Takeover delay and speed ceiling: Each row adds its selected delay to the machine’s pre-brake delay \(\tau_{\text{delay}}\) and solves the finished stopping budget of The warehouse mobile manipulator's stopping budget for the largest speed that still stops inside the rack-end clear distance.
Intervention mode Selected total delay Added travel at aisle speed Speed ceiling Local protection
Direct bilateral teleoperation 150 ms 195 mm 1.22 m/s Qualified passivity/rate checks
Continuous velocity guidance 450 ms 585 mm 0.96 m/s Onboard velocity/clearance gate
Supervisory waypoint dispatch 1,150 ms 1,495 mm 0.60 m/s Local obstacle veto and fallback
Out-of-loop takeover 2,000 ms 2,600 mm 0.40 m/s Independent stop while waiting
Physical protective input 250 ms human-plus-device response 325 mm 1.13 m/s Risk-assessed protective stop

Human response is slow and prone to interruption; thus, manual authority cannot be an open-ended state. If an operator assumes control and subsequently experiences wireless packet loss, physical incapacitation, or cognitive distraction, the machine cannot remain adrift under stale manual setpoints. The architecture protects against operator dropouts by bounding manual authority with configured and validated temporal leases (countdown timers that expire without active, periodic renewal) and hardware enabling devices, which require continuous physical engagement to maintain actuation channels (figure 2). In a lease-based communication model, the delegation window is sized from the machine’s communication and stopping budgets, as the mobile manipulator’s manual lease in section 1.3 shows. The operator console must continuously stream authenticated keep-alive packets to renew the lease before the deadline expires. If a single lease window lapses without renewal, or if an operator releases a spring-loaded physical deadman switch on the controller, manual proposals are revoked at the qualified detection deadline, and the enforcer commands the plant-specific validated stop or hold of the fallback ladder (The Fallback Ladder).2

Front of an ABB industrial robot teach pendant with a display, keys, red emergency-stop actuator, and attached cable; rear enabling grip and safety wiring are not visible.
Figure 2: Industrial teach pendant: Front view of an ABB pendant showing the display, buttons, emergency-stop actuator, and cable. Its rear enabling grip, internal channel wiring, response time, and stop behavior are not visible here. The ABB manual documents a three-position grip and a 250 mm/s manual-mode limit for its specified configuration. Photograph by Auledas, CC BY-SA 4.0, via Wikimedia Commons; local copy is re-encoded.

Even while an operator holds an active lease, the enforcer evaluates human setpoints against the same barrier certificates, current limits, and kinematic limits applied to the learned model. If an operator panics and commands a base turn that would tip the loaded tote rack, or requests a joint torque that would shear the gear teeth of the arm’s actuator, the enforcer rejects the command. A human command therefore remains a proposal, and principle \(\ref{pri-vol4-proposal-permission}\) applies to it unchanged. Intervention transfers the right to propose away from the learned policy, while the right to admit stays with the permission path. What remains to design is how the right to propose changes hands, at which tick, and what the machine does when the handover stalls.

Authority Transitions

Control ownership must be a discrete state, never an inference from a blended command. When an engineering team attempts to soften transitions by continuously averaging the control vectors of an autonomous policy and a human operator, the resulting command may represent neither intent. If the remote assistant steers the base left around a pallet while the learned planner steers right to pass it on the other side, a linear interpolation between the two commands holds the heading straight and drives the base into the pallet.3 To prevent this ambiguity, control authority over each actuator channel, such as the base drives or the arm joints, must resolve at every discrete time tick \(t\) to a single, mutually exclusive state \(S_t \in \{\text{Autonomous}, \text{Handover-Pending}, \text{Blend}, \text{Human-Manual}, \text{Fallback}, \text{Protective}\}\). A fallback state records the rung of the fallback ladder (The Fallback Ladder) that the permission path is executing (hold, stop, or inhibit), and the protective state is entered only from an independent protective input or a qualified critical fault. The permission path evaluates arbitration logic, determines the active state, and commits that identity to the bus before dispatching setpoints to low-level motor controllers. At no point do two masters write the actuator. During BLEND, the enforcer alone owns the actuator channel and interpolates between a latched last-approved policy reference and a checked human target.

Each edge in table 2 names its guard and deadline, its single actuator owner, and the fallback taken when the guard fails or an active lease lapses.

↳ Downstream: These takeover transitions become the authority fault cases derived in Deriving the Fault List.

Table 2: Authority state machine transition matrix: Each row names the single actuator owner and the state-dependent fallback. The example uses a \(1\text{ ms}\) arbiter tick. The BLEND state interpolates only latched, checked endpoints, under enforcer ownership.
From → to Guard and deadline Actuator owner / command Failure response
Autonomous → Handover-Pending Policy requests supervision; deadline computed from clearance Enforcer admits policy proposals while feasible Validated fallback if deadline reached
Autonomous or Handover-Pending → Blend Authenticated operator ACK, fresh state, feasible target and stopping reserve; commit on a control tick At commit, revoke policy token, latch last approved command and checked manual target; enforcer alone writes actuator Deny invalid target; use resident fallback
Blend → Human-Manual Profile completes and manual lease remains valid Enforcer admits live human proposals within physical limits Validated fallback on lease or enabling-device loss
Handover-Pending, Blend, or Human-Manual → Fallback Timeout, stale lease, or invalid state The permission path executes the validated HOLD or STOP rung and records it Re-arm only after validated recovery procedure
Any state → Protective Independent protective input or qualified critical fault Safety circuit/controller selects plant-specific protective stop Validated brake or hold, not torque removal alone

Shifting authority between these discrete states requires a deterministic four-phase handshake protocol (Request \(\to\) Acknowledge \(\to\) Commit \(\to\) Confirm). Executed by the enforcer, this four-stage transaction enforces mutual exclusion, validates cryptographic freshness to ensure commands are not stale network echoes, and atomically shifts actuation authority between discrete holders across a synchronous tick boundary. The sequencing excludes split-brain command contention, where two masters attempt to drive the hardware in conflicting directions during the same compute cycle. In the first phase, the Request phase, one entity signals an intent to transfer control. This occurs either when an operator asserts an override through a physical input like a button press or console force exceeding a calibrated threshold, or when the autonomous policy detects that sensor confidence has degraded and requests human assumption of control. In the second phase, the Acknowledge phase, the enforcer inspects the target state prerequisites, verifying that sensor data is valid, communications are active, and the recipient is physically capable of receiving command authority. In the third phase, the Commit phase, the enforcer executes the state transition at a synchronous control tick boundary, revoking the policy’s token on the committed channels, latching its last admitted endpoint and the checked human target, and assigning the enforcer sole actuator ownership throughout BLEND. In the final phase, the Confirm phase, the permission path broadcasts a state confirmation packet across the network and triggers physical feedback, such as haptic vibration or optical indicators, to tell the operator, channel by channel, whether blending or manual control is active; lost confirmation does not change the committed owner.

Sequence strip showing four-phase handshake timing over a 31-millisecond total budget.

The illustrative four-phase budget totals 31 ms; the enforcer owns the blend after commit.

Each phase in table 3 carries a budgeted bound, a guard and owner, and its failure handling. The transaction rejects spoofed override injections and routes interrupted handovers to the separately validated fallback path.

Table 3: Four-phase handover timing: The phase budgets sum to \(10+5+1+15=31\text{ ms}\). Confirmation and physical blending are distinct from the atomic ownership commit.
Phase Budgeted bound Guard and owner Failure handling
1. Request 10 ms Authenticated role and fresh, clock-converted nonce; no owner change Reject invalid packet; keep current qualified mode
2. Acknowledge 5 ms Check state, manual target, and full stop reserve Reject request; invoke fallback if current mode cannot continue
3. Commit 1 ms tick Revoke policy token, latch both checked endpoints, assign enforcer sole actuator write path No partial commit within the requested channel mask; fallback on failed guard
4. Confirm 15 ms Announce BLEND, then MAN only after profile and lease check Lost notice alarms operator but does not undo commit

Because human attention is volatile, a handover request initiated by the autonomous policy cannot remain open indefinitely. If a machine encounters an operational design domain boundary at speed and requests intervention, an operator who is distracted, asleep, or incapacitated will fail to complete the handshake. The state machine enforces the earlier of a configured protocol ceiling and a clearance-derived latest fallback time. Take a static obstacle at clear distance \(D_{\text{clear}}\), the case the stopping budget protects, with forward speed \(v>0\), the resident stop’s deceleration \(a_{\text{eff}}>0\), fixed overheads \(\delta_{\text{loc}}+\delta_{\text{margin}}+\epsilon_{\text{track}}\), and the machine’s own pre-brake delay \(\tau_{\text{delay}}\). Solving the stopping envelope of equation for the delay a request may add then gives the latest time after the request at which fallback may be dispatched: \[T_{\text{latest}} = \frac{D_{\text{clear}}-\delta_{\text{loc}}-\delta_{\text{margin}}-\epsilon_{\text{track}}-v^2/(2a_{\text{eff}})}{v}-\tau_{\text{delay}},\qquad T_{\text{timeout}}=\max\!\left(0,\min(T_{\text{config}},T_{\text{latest}})\right). \tag{2}\] The safety microcontroller dispatches fallback no later than the last permission tick before this deadline; a nonpositive \(T_{\text{latest}}\) calls for immediate fallback and may mean stopping is already infeasible. On the mobile manipulator, a pallet left just past a rack end, which comes into view only at the clear distance, is such an obstacle, and the numerator is the rack-end clear distance less the finished stopping budget of The warehouse mobile manipulator's stopping budget without its blind travel (\(v\,\tau_{\text{delay}}\)), which the formula subtracts as time. \(T_{\text{latest}}\) is therefore the budget’s spare margin spent as time. If the policy requests a handover at the 1.3 m/s aisle speed, the machine has 102.6 mm to spare, so \(T_{\text{latest}}=\) 78.9 ms. On the 1 ms permission tick, fallback begins by 78 ms, long before a 500 ms configured ceiling would expire. \(T_{\text{latest}}\) is a state-dependent deadline; the configured ceiling alone is not a deployable safety rule.

The deadline decides which takeovers the machine may request at speed. The 31 ms four-phase handshake of table 3 fits inside it, so an authenticated commit can land before the fallback must begin. A person cannot respond that fast. The 150 ms in-loop takeover of section 1.2 needs a speed of 1.22 m/s or less, so before it requests one the machine slows to a chosen 1.2 m/s. There the budget leaves 209.7 mm, equation 2 gives \(T_{\text{latest}}=\) 174.7 ms, and the takeover commits with 24.7 ms in hand. The deadline at the slow-down, not the aisle-speed one, governs every in-loop takeover the machine requests. A 2 s out-of-loop takeover needs 0.40 m/s or less, so the machine requests one only at a chosen 0.3 m/s crawl or from standstill. The hidden pallet allows neither, so the machine answers it with its own stop and requests the out-of-loop takeover from standstill. The slowed approach and the crawl govern requests the machine can anticipate, and the remote assistant who routes it around an obstacle therefore takes over from one of those speeds or from a stop, never at aisle speed.

Once authority resides in \(\text{Human-Manual}\), the machine must verify the engagement signal required by its delegated-control design. A teleoperation console might sense grip, gaze, or force, and the designer chooses among them by their measured coverage and false-loss rate. The manual lease \(T_{\text{lease,man}}\) of section 1.2, a lease in the sense of Multi-Rate Cadences on the operator’s input, is set here at a chosen 100 ms. After a lapse, the state machine blocks a return to \(\text{Autonomous}\) or \(\text{Human-Manual}\) until its specified reset conditions are met. Those conditions may require a full stop, fault review, and a new handshake.

When multiple control sources propose commands, the permission path resolves source selection through a fixed arbitration hierarchy (Brooks 1986) evaluated each cycle, below the proposal boundary of The Machine in Five Levels. The independent enforcer admits, modifies, or rejects the selected candidate under its configured physical limits. In \(\text{Human-Manual}\) with a live lease, the human proposal takes precedence over a learned proposal. In \(\text{Autonomous}\), the learned policy may propose actions for admission. Fixed-state arbitration takes bounded work per tick in this design; it does not require live negotiation between the two proposers.

Brooks, Rodney. 1986. “A Robust Layered Control System for a Mobile Robot.” IEEE Journal on Robotics and Automation 2 (1): 14–23.

The runtime intervention interface joins the four-phase handshake to an arbiter with bounded timeouts and explicit fallback transitions (figure 3).

Two-panel schematic of authority handover architecture. Panel (a) shows request, acknowledgment, commit to enforcer-owned blend, and confirmation. Panel (b) shows autonomous, pending, enforcer-owned blend, manual, and fallback states. Commit occurs before blend; timeout and lease expiry lead to fallback.
Figure 3: Authority handshake and owner states: Request and acknowledgment precede atomic commit. At commit the policy token is revoked and the enforcer owns the blend between checked, latched endpoints; manual proposals begin only after profile completion with a live lease. Timeout invokes a separately validated fallback.

One writer makes authority unambiguous at each control tick, but an unblended switch may still request an abrupt command change. If the policy’s last admitted torque on an arm joint differs from the operator’s target, directly swapping those setpoints requests the whole difference as a step within one tick. The committed enforcer instead uses the latched endpoints to generate a checked transition profile with a reserved stopping margin (Continuous Trajectories).

Checkpoint 1.1: Authority transition protocols and timeout bounding
Before designing handshake protocols for human-robot authority handoffs, verify your understanding of transition state machines and timeout dynamics:

If any transition paths permit unhandled timeouts or undefined intermediate states, the authority protocol is vulnerable to mode confusion and runaway actuation. Review this section before proceeding to smooth command blending.

Transfer Without Discontinuity

At commit time \(t_0\), the enforcer latches the last admitted policy command \(\mathbf u_0\) and a checked human target \(\mathbf u_1\). A direct switch would request \(\Delta\mathbf u=\mathbf u_1-\mathbf u_0\) in one tick. Its ideal command derivative is discontinuous, but the actual plant filters the command through finite drive, brake, and mechanical dynamics. A large change may excite the arm’s gearbox, exceed the drive wheels’ available traction, or shift the tote on its rack. The transfer therefore checks both rate and full stopping clearance rather than assuming a software tick maps instantaneously into physical force (The Causal Boundary).

Eliminating this step discontinuity requires smoothly mixing the old and new commands across a finite time window \(\tau_{\text{blend}}\), a technique known as bumpless transfer. A linear blend changes slope at its endpoints. For a fixed-endpoint change in commanded acceleration, a \(C^2\) bumpless transfer ramp eases in and out, removing endpoint jumps in the command and its first two derivatives. A standard curve used for this is the quintic smoothstep.

Definition 1.1: C² bumpless authority transfer ramp

\(C^2\) bumpless authority transfer ramp is the enforcer-owned command profile between two checked, latched references \(\mathbf u_0\) and \(\mathbf u_1\), parameterized by a monotonic weighting function \(\alpha(t) \in [0,1]\) with vanishing first and second boundary derivatives (\(\dot{\alpha} = \ddot{\alpha} = 0\)) over a finite blending window \(\tau_{\text{blend}}\), ensuring continuity of commanded acceleration and bounded derivative jerk across authority handoffs.

  1. Significance: A setpoint step has an unbounded ideal derivative. A \(C^2\) command profile removes that discontinuity; measured actuator tracking and plant limits still decide physical feasibility.
  2. Distinction: Unlike a linear interpolation ramp that exhibits derivative discontinuities at its endpoints, a \(C^2\) quintic spline enforces vanishing first and second derivatives (\(\dot{\alpha} = \ddot{\alpha} = 0\)) at both entry and exit, making the commanded acceleration continuous with zero endpoint jerk for fixed endpoints.
  3. Common pitfall: Using \(\tau_{\text{blend}} < 1.875 |\Delta \mathbf u|/j_{\text{allowable}}\) for a fixed acceleration-command change breaches the stated command-jerk limit (Smooth Trajectories and Bumpless Transfer). Resonance, slip, and stopping distance require plant tests.

The enforcer blends with the quintic smoothstep, whose coefficients follow from requiring zero first and second derivatives at both ends of the window (Smooth Trajectories and Bumpless Transfer). It is the same \(C^2\) property the planner relies on at its seam, where no acceleration step appears at the switch (Behavior at the Seam); here the switch is between proposers rather than between trajectory segments. The appendix also derives the shortest window that keeps a fixed change in commanded acceleration within a chosen command-jerk limit \(j_{\text{allowable}}\). The same bound applies to a torque command, so an illustrative 15 N·m change in one of the arm’s joint torques under a 750 N·m/s rate ceiling takes at least 37.5 ms.

Increasing \(\tau_{\text{blend}}\) reduces the profile’s peak commanded jerk, but delays the full human target and can lengthen the stopping trajectory. Consider the remote assistant’s takeover of the mobile manipulator’s base, committed while the machine cruises at the \(v =\) 1.2 m/s it slowed to before requesting it (section 1.3), and suppose the assistant commands a stop at the scenario’s braking deceleration, a fixed step of \(|\Delta a| =\) 2 m/s² from the latched cruise command. The enforcer holds the blend to \(j_{\text{allowable}} =\) 8.89 m/s³, the peak jerk \(6v/T_{\text{stop}}^2\) of the resident stop from the same speed (Behavior at the Seam), so the transfer loads the tote rack no more abruptly than the machine’s own stop. The smoothstep bound then sets a minimum blending window of 421.88 ms. The smoothstep weight averages one half over the window, so the base would leave the blend at 0.78 m/s after traveling 455.4 mm, and braking from there at the full target would take another 151.4 mm, for 606.8 mm to rest.

↰ Prerequisite: Acceleration-bounded bumpless transfer profiles build upon the physical stopping constraints enforced in Stopping Envelopes.

An instantaneous step takeover would stop in 360 mm, so bumpless blending imposes a spatial penalty of 246.8 mm.4 The blended stop is also 66.8 mm longer than the 540 mm resident stop from the same speed. A blend therefore spends stopping reserve beyond the stop the budget holds. The enforcer admits each blended tick only while the resident stop from the resulting state still fits, and where no blend fits, the base receives the resident stop itself, the suffix the fallback ladder’s stop rung executes (The Fallback Ladder). Bypassing the blend window by applying the acceleration step within a single 1 ms permission tick would instead request a command jerk of about 2,000 m/s³ over that tick. The idealized step is only a comparison baseline.

Command blending removes the chosen output discontinuity, but a standby human feedback controller can still accumulate integral error while disconnected from the plant. Initial condition matching pre-biases that controller against the measured state before commit, avoiding an idle-integrator transient. It does not force the deliberate human target to equal the last policy command. The enforcer latches both endpoints, sizes the blend for their actual difference, and checks the resulting command and stopping reserve throughout transfer.

To verify a transfer, record the command profile, motor current, encoder motion, and chassis inertial response on a shared timebase with bounded alignment error. A step command has a derivative discontinuity in the ideal model; ringing, high jerk, and slip are possible observations, not inevitable waveforms. A fixed-endpoint quintic bounds the commanded rate, and figure 4 contrasts the two profiles schematically. The measured plant response establishes whether the handover met its physical limits.

Two panels compare a step command with a fixed-endpoint quintic command. The step has an ideal rate impulse and a schematic damped response; the quintic command has a bounded rate across its blend window. The plotted rate is not a measured actuator jerk.
Figure 4: Step versus quintic command transfer: A step and fixed-endpoint quintic blend are compared. Blue is the command; red is its ideal derivative, and gray ringing is schematic. The profile bounds the commanded rate for the assumed endpoints.

The enforcer’s \(1\text{ kHz}\) loop gathers the four-phase handshake, integrator pre-biasing, and the \(C^2\) quintic blend into one protocol on the safety microcontroller. Each tick it samples plant state, stopping reserve, protective inputs, and the authority lease. It commits an ownership change only on the qualified commit tick that follows a verified acknowledgment, blends only between latched endpoints and checks every blended candidate before dispatch, and falls back on a lapsed lease or a failed clearance check.

↳ Protocol: The enforcer’s cadence, memory and write invariants, and per-state execution steps are given in Authority-transfer protocol.

Hogan, Neville. 1985. “Impedance Control: An Approach to Manipulation: Part i–Theory, Part II–Implementation, Part III–Applications.” Journal of Dynamic Systems, Measurement, and Control 107 (1): 1–24.

Once the blend completes and the remote assistant holds the arm in Human-Manual, driving it through a force-reflecting leader, the arm must respond either to the operator’s motion or to the operator’s force. Impedance control (Hogan 1985) renders a force from measured motion through the virtual spring-damper of \(\ref{psp-body-impedance-control}\) and suits backdrivable joints. Admittance control measures the operator’s force at a wrist force-torque sensor and yields in motion under stiff position tracking, which suits the high-ratio drives whose reflected rotor inertia \(N^2 J\) resists being moved by hand (Actuator Transmission Limits). Either mode closes a loop through the operator, the mechanism, and the transport delay, and the stability of that loop depends on grip, gains, and delay, so each mode must be tested for passivity over its specified operating range. A passive teleoperator, meaning the leader, link, and arm taken together as seen from the operator’s hand, never returns more energy to the operator than the operator put in. The enforcer keeps its clearance and torque limits whichever mode the leader uses.

Blending, pre-biasing, and compliant interaction all assume that the proposals they shape come from an operator entitled to control the arm, and a force measured at a sensor carries no credential of its own. If the permission path accepts a transfer request without validating identity, an unauthenticated command injection or a malfunctioning subsystem can seize authority and override the machine while it is in motion. The mechanisms that make an intervention smooth must therefore be paired with mechanisms that establish who is allowed to intervene.

Authenticating an Intervention

The remote assistant’s veto_path packet and a forged wheel drive-torque packet can reach the mobile manipulator over the same wireless link in the same control cycle. The permission path must keep the forger from seizing physical authority while still letting an authorized operator halt the machine before collision. These requirements conflict, so intervention operations are classified by the mechanical work they induce. Removing power or asserting a mechanical brake is physically distinct from injecting active motion commands, and the two actions do not warrant identical verification mechanisms. A protective stop is a braking or holding rung of the fallback ladder, not torque removal alone (The Fallback Ladder). In contrast, injecting base velocity, wheel torque, or joint velocity setpoints introduces fresh mechanical power into the system. An unauthorized or corrupted motion command can accelerate a moving mass into an obstacle, whereas a trusted out-of-band protective input requests the plant-specific minimum-risk stop.

Motion commands introduce active kinetic hazards; thus, the verification mechanism must meet both latency and freshness bounds from the kinematics of the maneuver. Verifying the signature and checking the replay window of a network motion or override command is one more delay before braking can begin. It adds to the actuator onset \(T_{\text{act}}\), the bus delay \(T_{\text{bus}}\), and the other terms of \(\tau_{\text{delay}}\) (equation), so for a static obstacle at clear distance \(D_{\text{clear}}\) and forward speed \(v\) the maximum permissible verification latency \(T_{\text{auth,max}}\) is the handover deadline itself, \(T_{\text{latest}}\) of equation 2.

On the mobile manipulator at its 1.3 m/s aisle speed, that bound is 78.9 ms. Assume two illustrative verifiers on a microcontroller, one slow and one fast:

  • Verifier A (slow): Verification takes \(t_{\text{auth,A}} =\) 180 ms, more than the deadline. The machine travels 234 mm while the signature is checked, 131.4 mm more than the spare, so its stop ends that far past the rack-end clear distance.
  • Verifier B (fast): Verification takes \(t_{\text{auth,B}} =\) 15 ms, well inside the deadline. The machine travels 19.5 mm during the check and keeps 83.1 mm of the spare.

CAN bus priority: A transmitting CAN frame is non-preemptive. A higher-priority identifier wins only at the next arbitration opportunity; controller buffer policy and worst-case current-frame serialization must be included in the measured delivery bound. An independent protective input avoids this shared queue where the safety case requires it.

Freshness requirements arise from the same kinematics. If an authentication window permits a timestamp drift or message age threshold of \(\tau_{\text{fresh}} =\) 100 ms, an authenticated packet replayed from the boundary of that window commands an action based on a base position displaced by \(\Delta x = v\,\tau_{\text{fresh}} =\) 130 mm at the aisle speed, more than the whole 102.6 mm spare.

A trusted, independently wired protective input is evaluated by its safety circuit without waiting for a network signature, and its response is the stop or hold selected and validated for the plant. Remote motion or takeover packets follow the capability and authentication checks. A corrupted checksum, stale nonce, or failed signature on that network is rejected, logged, and handled as rule A4 of section 1.8 specifies; malformed traffic does not acquire the authority to command a brake. If valid command renewal is lost, the separately monitored lease expires and the enforcer invokes its validated fallback. This separation preserves an available local protective path while preventing arbitrary network noise from forcing an unsafe stop.

Authentication and transport delay also reach training. A training simulator that always supplies immediate human rescue can underprice risky states. If deployed intervention instead incurs authentication, transport, and actuation delay, a policy trained on that mismatch may visit states from which rescue is no longer feasible. The remedy is to model the qualified latency and fallback path in training and hardware-in-the-loop evaluation, retain failed learner-visited states, and test the actual clearance distribution.

↰ Prerequisite: Microsecond override masking registers on real-time fieldbuses are implemented in Actuation Authority.

The authority inventory must include diagnostic, debug, bootloader, and wiring paths that can affect actuation. A production design may disable debug access, authenticate maintenance commands, protect boot updates, and enforce bus permissions according to its threat and fault model. The verification task is to show that each reachable actuator-write path is either gated by the permission authority or independently protected. The argument of this section covers remote packet injection over wireless links, replay on internal buses, and commands from a compromised application processor. Physical tampering lies outside it, and whether the safety case must answer tampering depends on the threat model, inspection access, and packaging.

Authentication establishes who is entitled to take over. It does not ensure that the handover leaves the machine under coherent control. When authority transfers while the base carries momentum and the arm’s drives deliver torque, the policy, the operator, and the enforcer must agree on who holds each channel. Handover goes wrong when that agreement breaks.

How Handover Goes Wrong

A handover can fail while every component works. When an autonomous policy and a human operator share authority, it can fail through mode confusion and authority ambiguity, where the machine assumes the human is managing an operational axis while the human assumes the machine retains jurisdiction. In multi-axis physical systems, authority is frequently decomposed across channels, and on the mobile manipulator it splits between the base and the arm. If the learned policy encounters an unmodeled surface reflection and quietly yields the arm to the remote assistant without releasing the base, an assistant who guides the arm may assume the base is holding still while the policy is still driving it toward the rack end. Conversely, if the assistant commands the base to stop, firmware that transfers only the base leaves the arm policy still reaching for the tote against the assistant’s intent. In telemetry logs, this failure presents as disjointed control flags where authority bits toggle asynchronously across execution threads. At the human-machine interface, however, this state is undetected until kinematic error exceeds the physical clearance margin, because neither claimant receives an immediate, unambiguous signal of the other’s partial withdrawal. Per-channel ownership permits such a partial transfer, so the protocol must make it explicit. Every request names the channels it covers in its channel mask, the state machine of section 1.3 commits owner states only for those channels, and the Confirm phase announces the owner of each channel to the operator, so a base that stays with the policy is shown as the policy’s.

Even with defined authority transfer protocols, the human nervous system cannot match the response deadlines of moving mass. A remote assistant in direct control of the mobile manipulator answers within the 150 ms in-loop takeover of section 1.2. When the same assistant supervises several machines passively while their policies run, situational awareness decays, and an unexpected handover request meets the 2 s out-of-loop takeover instead. As section 1.2 showed, a machine left at aisle speed covers more than twice the rack-end clear distance during this gap, before the assistant has perceived the obstacle, re-established spatial orientation, and applied input. Operator-facing cameras can track gaze direction and eye closure. They cannot measure comprehension of the scene geometry, so the growth in latency is detected only after the response deadline has passed.

At Tempe, vigilance decay met a disabled stop path (When the Boundary Fails) (National Transportation Safety Board 2019). In the Cruise incident, a vehicle that had stopped after striking a pedestrian began a pull-over and dragged her beneath it (The Fallback Ladder) (Quinn Emanuel Urquhart and Sullivan, LLP 2024). The first asks whether an operator-monitoring contract backed by an independent stop would have stopped the vehicle within the remaining clearance, and the second whether contact uncertainty should revoke motion authority until a validated hold is established.

National Transportation Safety Board. 2019. Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian, Tempe, Arizona, March 18, 2018. HAR-19/03. National Transportation Safety Board.
Quinn Emanuel Urquhart and Sullivan, LLP. 2024. Report to the Boards of Directors of Cruise LLC, GM Cruise Holdings LLC, and General Motors Holdings LLC Regarding the October 2, 2023 Accident in San Francisco. Quinn Emanuel Urquhart; Sullivan, LLP.

When an intervention occurs while an actuator applies force, manual input and the enforcer can oppose each other. For example, the remote assistant driving the arm under manual_drive through a force-reflecting leader may command a joint toward the position limit at which the enforcer is holding it. The torque the assistant’s command requests and the drive current holding that joint then act against each other on the same link. If the human-device loop has insufficient phase margin, repeated corrections may oscillate at a frequency set by the gains, delay, grip, and joint mechanics. The enforcer must retain the physical limit while the leader makes the opposing force and active mode legible to the assistant.

Uncoordinated handovers also corrupt the data drawn from fleet logs. If intervention segments enter behavioral cloning or diffusion training without intervention masking, which uses the logged commit instant to separate the policy’s own states from the recovery, the policy learns emergency recovery as nominal behavior, and validation loss on trajectory chunks does not reveal the corruption.

War Story 1.1: Ground-contact interlocks on flight 2904
Architectural block diagram of the supervisory state machine lockout on Lufthansa Flight 2904. The right landing gear contacts the wet runway before the left. Distinct strut and wheel-speed conditions delay spoilers, reversers, and wheel braking; the diagram separates these gates.

Context: On Lufthansa Flight 2904 at Warsaw, the crew commanded deceleration that the aircraft’s ground-contact logic did not grant, because each decelerating function carried its own precondition on sensed state (Main Commission Aircraft Accident Investigation Warsaw and German Federal Bureau of Aircraft Accident Investigation 1994).

Mechanism: The right gear contacted the wet runway about nine seconds before the left. Spoiler deployment required both struts compressed or both wheel speeds above 72 knots; thrust reversers required both struts. Wheel brakes had a separate wheel-speed gate and alternate mode. Each precondition acted as a capability entry, granting a function only in a sensed state, and together they withheld the requested deceleration.

Outcome and lesson: Wheel braking began roughly four seconds after the left-gear contact, about thirteen seconds after the first right-gear contact. The aircraft overran the runway, leaving its end at 72 knots. A capability matrix that gates operations on operating preconditions must annunciate an unmet precondition to the person commanding, name the sensed state it waits on, and be tested against the states, such as one strut compressed on a wet runway, in which its preconditions disagree with the operator’s intent.

Main Commission Aircraft Accident Investigation Warsaw and German Federal Bureau of Aircraft Accident Investigation. 1994. Report on the Accident to Airbus A320-211 Aircraft in Warsaw on 14 September 1993. Polish Ministry of Transport; Maritime Economy / BFU.

The arbiter of section 1.3 exists to prevent these failures, because the control architecture cannot rely on implicit handoffs, human attentiveness, or ad hoc torque overrides. Showing afterward that it did requires a record kept at every tick, channel by channel. Deciding afterward whether a stop answered a genuine boundary violation, an invalid token, or noise on a safety line requires that per-tick record to survive power cycles and scheduling stalls.

Trustworthy Operational Logs

A record supporting an argument months later has requirements that a debug log lacks. A debug log may be buffered in volatile user space and dropped under backpressure. If an asynchronous disk write drops three lines of trace data during an operating system stall, an engineer loses minor diagnostic context while the application continues running. Missing records weaken reconstruction across the causal boundary established in The Causal Boundary. When the arm strikes the tote rack, establishing whether the learned policy commanded an illegal torque, whether the enforcer failed to clamp it, or whether an operator intervened too late requires a bounded, integrity-checked record plus independent plant observations.5 A forensic log therefore does not exist to assist a developer with line-by-line code tracing. Its purpose is to preserve what was sampled, proposed, admitted, and dispatched. Motor current, encoder motion, and contact-force sensors provide separate physical observations; commands alone do not prove delivered force or collision cause.

The chosen \(1\text{ kHz}\) control design records a structured authority and action tuple each tick, with declared loss handling. The minimal evidence tuple takes the form \(\langle t_{\text{MCU}}, S_{\text{auth}}, \mathbf{x}_{\text{state}}, \mathbf{u}_{\text{prop}}, \mathbf{u}_{\text{enf}}, \mathbf{u}_{\text{act}}, \text{ID}_{\text{operator}}\rangle\), which records both sides of the proposal boundary (The Machine in Five Levels). In this structure, \(t_{\text{MCU}}\) is the event time in nanoseconds on the safety microcontroller’s clock, converted from any other clock with a recorded error bound, \(S_{\text{auth}}\) records the active authority state indicating whether autonomous policy, manual override, or fault mitigation holds control, and \(\text{ID}_{\text{operator}}\) identifies the authenticated source of any active intervention. The three action vectors map onto the action taps of Which Action Becomes the Label, with \(\mathbf{u}_{\text{prop}}\) as the requested action \(a_{\text{req}}\) and \(\mathbf{u}_{\text{enf}}\) as the enforced command \(a_{\text{enf}}\). The mapped command \(a_{\text{cmd}}\) lies between them and needs no field of its own when the proposal already arrives in actuator coordinates. The field \(\mathbf{u}_{\text{act}}\) is the command dispatched to the drives, and the measured response \(a_{\text{meas}}\) comes from the plant observations in \(\mathbf{x}_{\text{state}}\). Omitting one element can limit a particular causal question, so the record declares what each field actually measures. If the log records \(\mathbf{u}_{\text{act}}\) without \(\mathbf{u}_{\text{prop}}\), investigators lose the policy-command comparison. If \(\mathbf{u}_{\text{enf}}\) is omitted, the log loses the proposed-versus-admitted comparison. Chained with integrity and clock-conversion metadata, these tuples form the forensic authority record that section 1.8 specifies.

Ordering these tuples requires physical timestamp integrity rather than software clock queries. Software timestamps derived from operating system calls inherit scheduling jitter, context-switch delays, and variable interrupt servicing latencies that can exceed \(5\text{ ms}\) on loaded application processors. Over a \(5\text{ ms}\) interval, an internal control loop running at \(1\text{ kHz}\) executes five complete control cycles. If an operating system logs the arrival of a manual intervention command with a \(5\text{ ms}\) delay, the record places the operator’s command after the actuator response, inverting physical cause and effect. A hardware timestamp records a defined electrical event, such as frame ingress or a protective-input edge, after clock conversion with bounded error. Investigators can order command and detected contact only when their timestamp uncertainty, sampling windows, and sensor delays do not overlap.

Persisting full sensor streams and high-rate telemetry continuously to non-volatile storage exceeds both the bandwidth and the endurance limits of embedded hardware. The four cameras of a mobile-manipulator teleoperation session stream raw frames at the ingestion rate that Teleoperation Demonstration Costs sizes, a rate only a Gen4 NVMe drive sustains, so a pre-trigger and a post-trigger segment of those frames far exceed any microcontroller’s memory. Raw frames do not belong in an MCU SRAM ring. The \(1\text{ kHz}\) control tuple is a small fraction of that stream and can use a bounded small-memory ring. If full raw frames are required, provision a separate DRAM or storage tier sized to the measured stream budget, or retain selected or compressed frames with a declared loss of detail. At a trigger, seal the pre-trigger segment and continue writing post-trigger samples into a different allocated segment; freezing the sole ring would lose the promised post-event evidence. Frame records retain their native capture times, while control tuples retain their \(1\text{ kHz}\) issue and sample times. Background persistence must not block the enforcer.

The pre-trigger window must capture the moment of policy divergence rather than merely the moment of physical takeover. When a learned policy commands an erroneous trajectory, human operators and external monitoring processes exhibit reaction latencies ranging from hundreds of milliseconds to multiple seconds before recognizing the deviation and asserting authority. If the ring buffer only recorded data from the instant an operator pulled an override switch, the telemetry would capture the rescue while omitting the anomalous commands that made the rescue necessary. Capturing a declared pre-trigger interval lets investigators inspect policy proposals and states before intervention; whether a proposal was unsafe needs the contemporaneous model, state estimate, and physical evidence. Without the pre-divergence history, evaluation cannot separate a necessary rescue from a premature interruption, and the learner-visited states that corrective relabeling needs are gone.

Integrity and retention can use a protected hash chain and independently tested power-loss storage. Within the Nervous System, logging routines batch evidence tuples into fixed blocks corresponding to \(100\text{ ms}\) of execution. Each block is hashed using SHA-256, and its hash is concatenated with the hash of the preceding block to form an append-only cryptographic chain. With protected key and anchor custody, changing a committed block is detectable against the retained chain; loss before commit and compromised keys remain separate failure modes. One implementation periodically anchors root hashes in protected nonvolatile storage. A hash chain supplies neither persistence nor complete event capture, so the safety case tests key custody, which committed blocks survive power loss, how interrupted writes are marked, and where intervals are missing.

A retained, integrity-checked log can establish recorded authority and command values for its sampled intervals. It cannot say whether those values were permitted. That judgment needs the manifest the log is read against, which names who may hold each channel, under which conditions, and within which deadlines, and it needs the log’s entries to carry the fields the manifest’s rules refer to.

The Authority Log

An investigator who asks, weeks after the blocked-aisle takeover, who was driving the base as it neared the rack end has only what the machine recorded. The authority record consumes two records from Part III. From the enforcement record of The Enforcement Record it takes the rungs of the fallback ladder, so every fallback it logs names the rung the permission path executed. From the placement record of Hardware Allocation it takes the measured timing of the operator-input path. It adds two parts of its own. The static authority manifest, the capability matrix of section 1.2 in machine-checkable form, binds every actuator channel to its authorized actors, operational preconditions, preemption tiers, and transition deadlines, and it is loaded and verified before the actuator drivers energize. The per-tick forensic authority record is a separate stream that records, for every tick, who held authority and what was proposed, admitted, and delivered, so that a tick without an entry reads as a gap, never as an implied pass; neither part substitutes for the other.

The static authority manifest specifies every permitted interaction as a hard contract between the physical platform and its potential operators. Each entry identifies a distinct physical authority domain through a hardware actuator mask, an authorized entity class (such as a local physical operator, a remote teleoperation console, an autonomous trajectory policy, or an emergency stop loop), and an integer preemption tier. An entity presenting a command vector must prove its authority through cryptographic role credentials verified by the permission path before its input can enter the actuation pipeline. A network operator’s command arrives as an OPERATOR payload under the proposal header of Multi-Rate Cadences, carrying the operator’s role, the requested channel mask, the command, a nonce, and a keyed tag. Alongside the credential, the entry specifies the mutual exclusion policy for the channel, the command rate ceiling \(r\) (how rapidly a requested command may change), the active blending duration \(\tau_{\text{blend}}\) enforcing a \(C^2\) bumpless authority transfer ramp to maintain the kinodynamic feasibility bounds established in Kinodynamic Feasibility, and the terminal fallback state that the hardware enters if an active session is severed.

Each per-tick entry of the forensic authority record carries the event time on the safety microcontroller’s clock, the actuator channels the decision applies to, the authority state (AUTO, PEND, BLEND, MAN, FALLBACK with its ladder rung, or PROTECTIVE, codes for the six states of section 1.3), the authenticated operator role, the cause of any preemption, the blend factor of the enforcer-owned profile, the proposed, admitted, and delivered commands, the sampled plant state, and a keyed link to the preceding block of the log. The versioned static manifest supplies vector units, field order, the clock mapping and its error bound, credential rules, and hardware access controls.

↳ Byte layout: The forensic authority record’s fields, types, and offsets are given in Authority record and operator payload.

Mapping these entries to the mobile manipulator requires bounding four timing parameters against its limits: the operator reaction budget (the takeover latency \(T_{\Sigma}\) of equation 1 for the takeover mode in use), the handshake protocol latency \(T_{\text{handshake}}\), the manual lease \(T_{\text{lease,man}}\), and the minimum blending window \(\tau_{\text{blend}}\). Table 4 states them for the machine’s two command channels, the base and the arm, which one remote-assistant session commands together.

Table 4: Blending windows on the mobile manipulator: One remote-assistant session commands both channels, and each window is rounded up to whole 1 ms permission ticks. The rate ceiling \(r\) bounds the derivative of each command. The base row is sized at the out-of-loop crawl, where its blend fits the stopping reserve; at the in-loop takeover speed no blended base command fits, and the base receives the resident stop. Every lease and pending deadline must fit the state-dependent stopping reserve.
Machine channel Command change and chosen rate ceiling Smoothstep window \(\tau_{\text{blend}}\ge1.875\lvert\Delta u\rvert/r\), whole ticks Human / protocol time Fallback contract
Base acceleration command Crawl at 0.3 m/s to braking at 2 m/s²; \(r=\) 35.56 m/s³, the resident stop’s peak jerk at the crawl 106 ms 2 s / 31 ms 100 ms manual lease; 500 ms pending ceiling shortened by clearance
Arm joint torque command 15 N·m; torque rate \(r=\) 750 N·m/s 38 ms 150 ms / 31 ms The same manual lease; its loss invokes the arm’s validated stop

Consider the remote assistant’s in-loop takeover of the mobile manipulator, requested after the machine has slowed to 1.2 m/s and committed under the assistant’s manual_drive entry. The operator reaction budget \(T_{\Sigma}\) sets the maximum duration the permission path permits between issuing a takeover demand and detecting the assistant’s first command; here it is the 150 ms in-loop takeover of section 1.2, which already includes transport. The handshake latency \(T_{\text{handshake}}\) defines the round-trip deadline for mutual authentication and state acknowledgment over the bus, budgeted at 31 ms per the four-phase protocol in table 3; it is part of \(T_{\Sigma}\), not an addition to it. The manual lease \(T_{\text{lease,man}}\) of section 1.3 sets the longest silence tolerated from the assistant before the permission path declares a link failure and reverts to the fallback trajectory, here 100 ms. Finally, the blending window \(\tau_{\text{blend}}\) governs the rate of authority transfer to prevent mechanical shock, and table 4 gives it for each channel. The base’s window depends on its speed, because the resident stop’s peak jerk, which caps the blend under the smoothstep bound of section 1.4, rises as speed falls; the table therefore sizes it at the 0.3 m/s crawl.

The in-loop takeover leaves 29.7 mm of the rack-end clear distance at 1.2 m/s, the 24.7 ms by which it beats \(T_{\text{latest}}\) (section 1.3). That remainder bounds the transfer itself. Any further delay before a stop could begin while authority is changing hands must fit inside the remainder, and the enforcer checks the resident stop against it at every tick of the transfer. No base blend passes that check at this speed. Blending the step to braking from 1.2 m/s grows the travel plus the resident stop from the blended state by up to 197.6 mm, far past the remainder, so in this session the base receives the resident stop and only the arm blends. After a 2 s out-of-loop takeover at the crawl, 236.2 mm remains, and the crawl’s blend uses at most 12.3 mm of it. Once the assistant holds MAN, the takeover delay is spent and the enforcer re-evaluates the budget each tick with the manual lease in the place of the 60 ms chunk lease in \(\tau_{\text{delay}}\). The 100 ms lease adds 48 mm of travel at this speed, which the 209.7 mm spare covers.

Verification draws its test inputs from the static manifest and its evidence from the event schema. The manifest’s authority_claims field hands Deriving the Fault List four authority rules by identifier, each of which becomes a hardware-in-the-loop fault case on the machine:

  • A1. One enforcer path writes each actuator channel, even when a second bus master attempts to write a base-drive channel while the policy and the remote assistant both command.
  • A2. No source bypasses an active higher-tier source or the independent permission check. A policy proposal during MAN is ignored, and an assistant command that opposes the enforcer’s admitted limit is still checked.
  • A3. An unacknowledged handover falls back by \(T_{\text{timeout}}\) of section 1.3, the earlier of the configured ceiling and the clearance-derived deadline of equation 2, which is 174.7 ms at the 1.2 m/s slow-down before an in-loop takeover and 78.9 ms for a handover the policy requests at the 1.3 m/s aisle speed.
  • A4. A replayed veto_path, manual_drive, or accept_item packet, or one with a bad authentication tag, locks out the affected channel, leaves authority unchanged, and records the event.

When assembling the safety case for operational certification, Deployment Release evaluates the authority record as the evidence for claim C6, the claim that authority transfers as these four rules specify. A safety assessment examines the stated authority constraints, implementation evidence, fault coverage, and tested operating envelope. The certification package pairs the static authority manifest with the plant measurements that section 1.4 requires of every transfer and with execution traces from the fault-injection campaign that show any gaps in the audit trail. The plant measurements are needed because a handover that is correct at the bus can still slide the base or shift the tote rack. The measured preemption latency in those traces is a percentile of the tested distribution rather than a bound (Moving Commands on Time).

↳ Downstream: Forensic authority records provide evidentiary artifacts for the deployment safety cases in Safety Cases and Claims.

Fallacies and Pitfalls

Changing an authority register does not complete a safe handover. The machine must preserve its stopping margin and enforcement functions through the transfer, while the resulting logs must distinguish intervention from expert recovery.

Pitfall: Treating human takeover timeout as a static time delay rather than a dynamic spatial stopping budget.

The mobile manipulator requests a takeover at its 1.3 m/s aisle speed under a configured 500 ms acknowledgment window. Waiting out the window before braking covers 0.65 m, and with the rest of the finished stopping budget the stop ends 0.55 m past the 1.10 m rack-end clear distance. The same window fits at the 0.3 m/s crawl, where the budget leaves 2.8 s before fallback must begin. A window that is safe in one state overruns in another, so handover timing must follow current speed, clearance, and credible deceleration, with autonomous fallback beginning while the stop still fits.

Fallacy: Human manual takeover should bypass automated safety barriers and grant unconstrained actuator authority.

The remote assistant takes over the mobile manipulator’s base to route it around a pallet, and the chosen path leads toward a rack upright. Switching to human control only changes who is sending the command; it does not change the physical realities of momentum, friction, or nearby obstacles. Within this architecture, the enforcer must continue checking the manual proposal against the admissible motion set and constrain it when a kinodynamically feasible alternative exists (building on the feasibility principles from Kinodynamic Feasibility). If the request cannot be admitted, the defined autonomous fallback still applies. Takeover lets the operator guide the task planning; it does not make force, clearance, and stopping limits optional at the joint actuator boundary.

Pitfall: Relying on cryptographic authentication without rate limiting the traffic that reaches the verifier.

The mobile manipulator’s safety microcontroller checks each remote takeover packet with the fast verifier of section 1.5, at 15 ms per signature. Verification rejects a forged packet only after spending that time on it. If a faulty or compromised node on the link sends manual_drive packets that fail verification faster than the verifier retires them, as few as 5 of them queued ahead of the remote assistant’s genuine veto_path delay its admission to 90 ms, past the 78.9 ms fallback deadline \(T_{\text{latest}}\) at the aisle speed. The independently wired protective input does not wait in that queue, but every network command does. Authentication preserves command integrity while leaving an availability hazard. Mailbox filtering and per-source rate limiting must bound the traffic before it reaches the verifier, and qualification must test both command rejection and enforcement timing under the admitted load.

Fallacy: Training directly on raw human intervention logs captures expert demonstrations.

A learning pipeline treats every remote-assistant takeover of the mobile manipulator as an expert demonstration and trains on the preceding action window. The logs combine genuine policy failures, cautious early interventions, and operator kinematic mistakes, while the control transition itself often includes jerky physical corrections, delayed neuromuscular responses, and hesitant corrections as the human orients to the scene. Cloning those segments can reproduce the unstable behavior that made recovery difficult rather than the recovery the team wants to teach. Dataset curation must identify why intervention occurred, who held the authority register, and which commands actually restored acceptable operation. Reviewed recovery trajectories can supply useful demonstrations, but neither an intervention timestamp nor full manual authority alone establishes that the recorded behavior is suitable for imitation.

Summary

Human authority over an embodied physical AI system needs an engineered real-time contract that arbitrates takeover, revokes autonomous policy authority at commit, and selects a validated fallback within the available physical budget. Supervisory intervention governs the transition between proposers while the independent enforcer checks the admitted command and executes the selected stop when a lease lapses. Network commands reach the enforcer only after capability and authentication checks, whose own cost makes rate limiting ahead of the verifier part of the design, while trusted protective wiring bypasses the network entirely. On the mobile manipulator at its 1.3 m/s aisle speed, the finished stopping budget leaves 102.6 mm, which buys 78.9 ms of waiting before the machine must dispatch its own stop. The machine therefore slows to 1.2 m/s before it requests an in-loop takeover, and to a 0.3 m/s crawl or a standstill before it requests an out-of-loop one.

Transferring control from an active policy takes time, because replacing an admitted policy setpoint with a different human target in one \(1\text{ ms}\) cycle requests a torque step. The \(C^2\) bumpless authority transfer ramp instead varies an authority parameter \(\alpha(t)\) smoothly from zero to one between latched command endpoints. Its sampled implementation must respect command-rate limits and the stopping reserve of Stopping Envelopes.

On the mobile manipulator the enforcer blends the base only from the 0.3 m/s crawl, over 106 ms. At the 1.2 m/s in-loop takeover speed no blended base command fits the clearance that remains, so the base receives the resident stop while the arm blends over 38 ms. The enforcer admits each blended tick only with a validated stopping reserve and monitors the plant response.

The authority record makes each intervention reviewable. Its static manifest binds every actuator channel to its actors, tiers, and deadlines, and its per-tick entries log who held authority and what was proposed, admitted, and delivered, naming fallback rungs from the enforcement record and operator-path timing from the placement record. A logged commit marks the recorded boundary where autonomous policy authority was revoked, subject to the logger’s clock and sampling limits, and hash chaining with protected keys can reveal later modification. Command and authority records support reconstruction only when paired with independent sensor, current, and contact evidence and their timing uncertainty.

For model training, the pre-commit interval is a policy-owned segment even when a human is deciding to intervene. Retain its learner-visited states, unsafe proposals, and authority flags as failure evidence and candidates for expert relabeling; exclude unsafe policy commands from expert imitation targets. The post-commit blend is enforcer-owned transition data. Only reviewed post-transfer human recovery states and actions that meet the task’s validity criteria should become demonstrations.

Key Takeaways: Supervisory intervention and bounded authority transfer
  • Clearance governs waiting: At the aisle speed the machine’s spare margin buys 78.9 ms; at the 1.2 m/s slow-down it buys the 174.7 ms that governs every in-loop takeover. A 2 s out-of-loop takeover needs a speed below 0.40 m/s. Dispatch fallback by the earlier of the configuration ceiling and the state-dependent stopping deadline.
  • One actuator writer: Commit revokes the policy’s token on the committed channels and latches its last admitted command. The enforcer owns BLEND and checks human proposals against the same physical limits.
  • Smooth commands cost stopping distance: A fixed-endpoint \(C^2\) acceleration profile bounds commanded jerk. Its extra stopping distance must be reserved.
  • Distinct stop paths: Trusted protective wiring invokes a risk-assessed plant response. Invalid network packets are rejected; lost valid renewal invokes the lease fallback.
  • Causal intervention data: The logged commit separates the policy-owned states before takeover, which corrective relabeling needs (Intervention and Recovery Data), from the human recovery after it. Use only reviewed post-transfer recovery actions as imitation targets.
  • Authority becomes evidence: The per-tick record counts a missing entry as a gap, never a pass. The manifest’s four authority rules become the fault cases verification injects, and their passing records, with the authority record, are the evidence for release claim C6.

A person’s authority over the machine is granted on evidence and expires with it, as principle \(\ref{pri-vol4-evidence-bounds-authority}\) requires of every authority the machine honors. The credential and the committed handshake are the evidence that grants it; the manual lease is that evidence’s age limit, and when the operator’s input stops renewing it, the authority lapses to a fallback the permission path owns. A preemption tier does not survive the lapse, and the remote assistant who held the base before it holds no authority after it. Returning control to any proposer waits on the reset conditions the state machine specifies, and on fresh evidence.

What’s Next: From authority revocation to physical verification
Does a clean authority log show that the authority rules hold? It does not, because no logged handover exercised the duplicate writer, the replayed packet, or the unacknowledged request that the rules exist to contain. Adversarial Verification turns A1–A4 into fault cases beside the rack-end faults, injects each with the loaded base approaching the rack end, and times each response with instruments independent of the machine’s own log, so that a handover fallback counts from its start at the drive rather than from the arbiter’s record of it.

Back to top

Footnotes

  1. Supervisory Control Levels: Sheridan and Verplank (1978) formulated a ten-level taxonomy of supervisory control, running from a human who performs the whole task without computer assistance to a computer that decides and acts while ignoring the human. A capability matrix draws such distinctions per operation rather than once for the whole machine.↩︎

  2. Three-Position Enabling Devices: The cited ABB operating manual documents a three-position enabling grip and a \(250\text{ mm/s}\) manual-mode limit for that robot. Releasing or over-squeezing the grip is a request into the configured safety circuit, and the machine’s risk assessment under ISO 13849-1 sets the required performance level and the stop the circuit selects.↩︎

  3. Fieldbus Write Exclusivity: Authority arbitration requires hardware-enforced exclusivity across distributed fieldbuses. Dual-writing to CAN-FD or EtherCAT segments from autonomous and manual controllers induces non-deterministic actuator chatter and inverter current oscillations. A qualified arbiter and bus permissions must prevent two software sources from writing one actuator channel.↩︎

  4. Bumpless Transfer Braking Penalty: The penalty belongs to the smoothstep profile, its fixed endpoints, and the stated acceleration step; another profile or endpoint pair gives a different penalty.↩︎

  5. Precision Time Protocol Logging: Unlike legacy event data recorders logging at coarse \(10\text{ Hz}\) intervals, physical AI incident reconstruction demands sub-millisecond precision. A qualified PTP installation can align the named device clocks with a measured conversion-error bound. Event ordering requires the intervals around timestamps, sensor exposure and sampling, and actuator response not to overlap.↩︎

Sheridan, Thomas B., and William L. Verplank. 1978. Human and Computer Control of Undersea Teleoperators. Technical Report N00014-77-C-0256. Man-Machine Systems Laboratory, Massachusetts Institute of Technology.