Preface
Who this book is for
This textbook is for computer science and engineering students, machine learning researchers, robotics engineers, and embedded systems architects who want to build embodied intelligent systems that safely sense, reason, and act in the physical world.
An advisory model ends at a display or a human decision, so its computation can be retried or canceled before anything outside the computer acts on the result. Its errors can still cause harm, but only through whoever or whatever responds to them.
In physical AI, software directs matter, momentum, and energy. When a vision-language model proposes motion for a mobile base or an articulated robotic arm, execution crosses the causal boundary into continuous Newtonian physics. If an actuation command exceeds the yield strength of a gearbox or the thermal dissipation limit of an inverter, physical hardware breaks. Once kinetic energy is released or material yields, software cannot undo the state transition.
Building systems that survive this reality requires a shared systems engineering discipline, and readers arrive at it from three directions:
- Machine Learning Practitioners: Moving beyond static test-set accuracy evaluated on stored datasets. In physical AI, actions alter future sensory observations through closed-loop feedback (\(s_{t+1} \sim P(s \mid s_t, a_t)\)), so a policy drifts off the states it was trained on, and the worst-case cost of that drift grows with the square of the horizon (The Four Bedrock Laws). Practitioners learn how to formulate continuous action policies, evaluate closed-loop survival, and submit unprivileged proposals to an independent permission path.
- Robotics and Control Engineers: Augmenting classical feedback loops with high-capacity neural foundation models. Modern vision-language-action models provide open-vocabulary semantic scene understanding and multimodal generalization, but they are learned proposers that run at tens of hertz with variable latency, so an independent permission path whose rate follows the plant’s stopping budget must check what they propose. Control engineers learn how to build that path, including safety filters such as Control Barrier Functions, on real-time silicon.
- Embedded Systems and Silicon Architects: Managing multi-rate execution across heterogeneous system-on-chip (SoC) platforms. Architects learn how to isolate high-throughput neural accelerators from the safety microcontroller that runs the permission path, preventing memory bus crossbar contention, direct memory access (DMA) starvation, and supply voltage droop from breaching hard physical deadlines.
Why a physical AI systems textbook
This book approaches physical AI from a first-principles systems perspective. Its scope is clearest beside the two neighboring kinds of book that it is not.
- Not a classical control theory book: We do not develop continuous-time Laplace transforms, contour integrals, or Lyapunov stability proofs at length. Formal continuous-control proofs and reachability derivations are collected in the volume appendices for optional depth.
- Not a traditional robotics kinematics book: We do not work through Denavit-Hartenberg parameter matrices, forward-kinematics linkage catalogs, or specialized mechanical design heuristics.
Instead, we reason from physical principles (conservation of energy and momentum, stopping kinematics, thermal dissipation, and Coulomb friction) and from computational systems architecture (multi-rate execution loops, memory bus contention, asynchronous message queues, and permission paths that check learned proposals).
Mechanisms taught to the depth that changes a budget
This book does not treat learned models or control blocks as black boxes, and it does not survey their internals either. A mechanism enters the main text at the depth at which it changes a number the machine must budget: a latency, a memory footprint, a clearance, or the evidence a claim can rest on. An action-chunk decoder appears through the horizon it supplies and the time a new chunk takes. A scene representation appears through the memory it occupies and the age of the evidence it holds. A barrier filter appears through the premises its guarantee needs and the deadline it must meet. Internals that move no such number, including the arithmetic of patch tokenization and attention, decoder loss functions, and surface-reconstruction methods, are in the machine learning and control appendices for readers who want them.
Delay is distance
A machine keeps moving while it senses, decides, and waits for its brakes. Every millisecond between an observation and the moment braking force builds is spent as travel toward whatever the machine must not reach, so operating speed, delay, clearance, and braking capacity form one budget rather than four separate specifications. The book builds that budget for one machine, adding each term in the chapter that owns it, beginning with the body’s stopping distance in Kinetic Momentum. A learned model’s inference time enters the budget only through the age of the evidence it hands forward, because the permission path, not the model, decides each tick whether motion may continue.
What this book promises
This book does not promise a machine that is provably safe. No finite test campaign, runtime filter, or certification scheme delivers that for a learned system in an open world, and the chapters say so wherever it matters. The book teaches something narrower. It shows how to build a machine whose learned proposals cannot authorize their own motion, whose permission path decides on evidence of known age within the margin that delay has not yet consumed, and whose every guarantee is written down with the premises it rests on. A barrier filter protects the plant only while its model, its state estimate, its actuator authority, and its deadline hold. A release is justified only for the conditions its evidence covers. A premise that evidence cannot settle restricts what the machine may do instead of being assumed away. The final chapter returns to the opening question with that answer. The guarantee is conditional, and the engineering lies in knowing its conditions and recording them.
The machine in five levels
The machine in this book has five levels. The Brain holds the learned proposers, which revise goals at a few hertz and emit short action sequences at tens of hertz. Below the proposal boundary, the Nervous System holds the permission path, which admits, modifies, or refuses every proposal on fresh state and owns the fallback. The Body is the plant under its drives, the Boundary is where commands become current and the world becomes measurement, and the World is everything the machine does not own. Governance surrounds all five as a lifecycle envelope, deciding who may command the machine and what evidence licenses it to operate. The Brain proposes; only the permission path permits. The Machine in Five Levels gives the full model and its rates.
The three machine classes
To keep systems concepts concrete, this textbook grounds its principles in the three machine classes of Three Machine Classes:
- Class 1 · Mass-Dominated Mobility (AMRs and Quadrupeds): Planar transport bases and legged platforms (such as Boston Dynamics’ Spot) navigating open facilities and rugged terrain. Governed by vehicle momentum, tire-road and foot-ground Coulomb friction, kinematic stopping distance envelopes, and multi-camera / lidar perception pipelines under rolling-shutter skew.
- Class 2 · Contact-Dominated Manipulation: Multi-link robotic arms performing dexterous contact tasks. Governed by non-linear arm dynamics, reflected rotor inertia (\(J_{\text{ref}} = N^2 J_{\text{rotor}}\)), cycloidal gearbox compliance, and contact force containment.
- Class 3 · Flow-Dominated Process and Energy Systems: High-pressure fluid loops, chemical batch reactors, and thermal plants. Governed by enthalpy exchange, pipe fluid dynamics, pressure transients, and irreversible thermodynamic phase changes.
One machine runs through the book, the warehouse mobile manipulator of Three Machine Classes, and each chapter adds what it owns to that machine’s budgets. Each part closes by testing its law on a humanoid, which couples all three classes on one chassis.
How to read this book: Progressive arc vs. modular drop-in
This textbook supports two reading patterns:
For sequential readers (the cumulative curriculum arc)
Readers moving cover-to-cover follow an introduction, four parts, and a conclusion, and each part is led by one of the book’s four laws. The Architecture of This Book traces the chain of records the chapters hand forward.
- Introduction (The Causal Boundary) asks what a machine must guarantee before a learned model’s proposal becomes physical work.
- Part I: The Machine Anatomy turns the first law, irreversibility, into budgets of time, distance, torque, heat, and voltage.
- Part II: Teaching the Machine asks what a learned proposer’s competence rests on under the second law, endogenous data, when its own actions choose what it observes next.
- Part III: Running the Machine builds the runtime around the third law, proposal is not permission.
- Part IV: Governing the Machine applies the fourth law, evidence bounds authority, to decide who may command the machine and under which conditions it may operate.
- Conclusion (The Epistemic Frontier) answers the opening question with what the evidence supports.
For drop-in readers and modular course instructors
Instructors and practitioners often need to study one topic in isolation, such as sensor DMA pipelines, trajectory chunking, or Control Barrier Functions. To support modular use without prerequisite confusion:
- Orientation at each chapter opening: Every chapter begins with an orientation section that frames the core problem and states its physical stakes, and the margin stack at each chapter’s opening marks the level the chapter works on.
- Explicit Foundational Context: Whenever a chapter draws upon a core concept established earlier (such as the proposal boundary, the stopping budget, or the three machine classes), the text provides an immediate inline orientation gloss and a precise section cross-reference.
- Self-Contained Worked Examples: Each chapter contains standalone analytical derivations, executable Python calculations, and industrial case studies that can be studied independently.
The margin stack is ordered by control rate and authority rather than by software dependency alone.
A stack with no level highlighted marks the governance chapters, because governance is an envelope around all five levels rather than one of them.
These three stacks are those of The Physical Body, The Nervous System, and The Cognitive Brain.
Margin reading aids
The margins carry small, unnumbered visual notes beside the passage where a relationship, a gap in scale, or a threshold is easiest to see. They are reading aids rather than figures to cite, and they come in five forms.
- Ladders span orders of magnitude (such as execution rates from inverter switching down to intent-model revision, or sensor ingestion bandwidths).
- Threshold knees mark the point beyond which a trend stops holding (such as reflected rotor inertia limiting joint acceleration, force rising with contact stiffness, or tail latency rising with memory bus contention).
- Budget envelopes set competing demands against a deadline (such as diffusion denoising times against control periods, or fieldbus cycle budgets).
- Sequence strips show multi-phase timing (such as authority handshakes, receding-horizon chunk windows, or the steps that seal a record).
- Causal chains trace how a fault propagates and how far it reaches (such as from an unmodeled simulator artifact to a damaged part).
The golden thread: An end-to-end forensic trace
To anchor abstract concepts across all seventeen chapters into a cohesive engineering reality, this book weaves a recurring real-world scenario—the Golden Thread—through the four parts and the five levels of the machine:
- Part I · The Physical Challenge (The Plant Hazard): An autonomous mobile robot (AMR) navigates an industrial aisle at \(v = 1.3\text{ m/s}\). Without warning, an obstacle enters its path with limited clearance. The plant’s physical braking capacity, tire-floor Coulomb friction, and actuator response delay dictate a hard stopping budget (\(d_{\text{stop}} \le D_{\text{clear}}\)) that no software optimization can bypass.
- Part II · The Learning Gap (The Policy Failure): The onboard vision-language-action (VLA) policy, trained on nominal demonstrations without sudden intrusions, encounters acute out-of-distribution covariate shift (\(\mathcal{O}(T^2 \epsilon)\)). Instead of braking, the generative action chunk decoder proposes an aggressive swerve that violates kinematic turning limits and would destabilize the chassis.
- Part III · The Runtime Execution (The Silicon Bottleneck): Camera exposure + SerDes transport + neural inference consumes critical milliseconds of unguided motion before an action proposal is emitted. When the proposal reaches the shared-memory mailbox, heterogeneous memory bus contention on the system-on-chip threatens the real-time deadline.
- Part IV · The Safety Shield (The Deterministic Intervention): Operating on an isolated real-time safety microcontroller, the high-rate reflex path evaluates the Control Barrier Function (\(h(x) \ge 0\)). Calculating that the proposed swerve will breach the admissible stopping envelope, the quadratic program minimally projects the proposal into an admissible braking clamp, arresting the chassis safely before contact without relying on software optimism.
Related books in the machine learning systems curriculum
This textbook is part of the Machine Learning Systems curriculum. Introduction to Machine Learning Systems examines how models use computation and memory on a single machine. Scaling Machine Learning Systems extends that engineering scope to coordinating distributed training and inference across cluster fleets.
The curriculum also addresses systems that take autonomous actions. Agentic Machine Learning Systems focuses on stateful goal-directed trajectories in digital software environments, while Physical AI Systems examines actions that produce force, motion, and thermodynamic change in the physical world. Each text offers a complementary perspective, and this book is self-contained, developing the physical, computational, and architectural principles required to engineer dependable embodied machines.