Notation

Machine Learning Systems spans machine learning (computer science and statistics) and systems (computer architecture and hardware). Each field developed its notation independently, and many symbols mean different things across communities and publications. This collision creates real confusion when the disciplines merge. The conventions below establish a single notation that eliminates this ambiguity.

Consider a simple statement: “Increasing \(B\) improves throughput.” To an ML researcher, \(B\) means batch size. To a hardware engineer, \(B\) means bandwidth. Both interpretations are correct in their respective fields, but in ML Systems we need both concepts in the same equation—hence the need for a single, consistent convention.

The Iron Law of ML Systems

For serialized execution phases, the iron-law performance decomposition is:

\[T = \underbrace{\frac{D_{\text{vol}}}{\text{BW}}}_{\text{Memory Time}} + \underbrace{\frac{O}{R_{\text{peak}} \cdot \eta_{\text{hw}}}}_{\text{Compute Time}} + \underbrace{L_{\text{lat}}}_{\text{Latency Overhead}}\]

Each variable was chosen deliberately to avoid collision with standard ML terminology.

Symbol Definition Unit Why This Symbol?
\(T\) Time seconds Unambiguous. Wall-clock time for an operation.
\(D_{\text{vol}}\) Data Volume bytes Avoids collision with \(D\) (Dataset Size). In scaling laws, \(D\) means training tokens. Here we need bytes moved through memory. The subscript disambiguates.
\(\text{BW}\) Bandwidth bytes/s Avoids collision with \(B\) (Batch Size). Systems literature often uses \(B\) for bandwidth, while ML literature often uses \(B\) for batch size. The notation preserves the ML convention.
\(O\) Operations FLOPs Total floating-point operations. Clean in equations (vs. “\(\text{Ops}\)”).
\(R_{\text{peak}}\) Peak Rate FLOP/s Avoids collision with \(P\) (Parameters). Roofline presentations often use \(P\) for performance, while ML literature commonly uses \(P\) for parameter count. The notation preserves the ML convention.
\(\eta_{\text{hw}}\) Efficiency — Hardware utilization \((0 \le \eta_{\text{hw}} \le 1)\). Avoids collision with learning rate \((\eta)\).
\(L_{\text{lat}}\) Latency seconds Avoids collision with \(\mathcal{L}\) (Loss). Fixed overhead time (kernel launch, network RTT). The subscript distinguishes from the loss function.

Why these choices matter

Without careful notation, sentences become ambiguous:

“Reducing \(D\) improves performance.”

The sentence has two possible readings:

  • Reducing dataset size (fewer training samples) can speed training but may reduce accuracy.
  • Reducing data volume moved (for example, through compression or quantization) can speed inference, with accuracy effects that depend on the technique and calibration.

With our notation, we can write precisely:

“Reducing \(D_{\text{vol}}\) through FP32-to-INT8 quantization cuts parameter memory traffic to one quarter while \(D\) (training data) remains unchanged.”

Our notation eliminates this ambiguity: “\(\text{BW}\) limits throughput” has only one reading.

Subscripted variants

An upright root names the quantity, and the subscript names which instance of it is meant. The convention holds closely related measurements apart instead of collapsing them into a single symbol:

  • \(\text{BW}_{\text{disk}}\), \(\text{BW}_{\text{network}}\), \(\text{BW}_{\text{accelerator}}\): bandwidth at three points on the data path
  • \(D_{\text{vol}}\), \(R_{\text{peak}}\), \(L_{\text{lat}}\), \(\eta_{\text{hw}}\): iron-law terms, each held clear of a common ML symbol
  • \(E_{\text{move}}\), \(E_{\text{compute}}\), \(E_{\text{total}}\): one energy budget split into tradeable parts

The subscript therefore carries part of the claim. Disk and accelerator bandwidth differ by orders of magnitude, so a figure is reproducible only when the subscript says where it was measured.

The Degradation Equation

The degradation equation is a fitted local model. Divergence alone does not determine the direction of accuracy change, so estimate \(\lambda\) from labeled outcomes. Some symbols below, such as \(\tau\), appear in chapter prose rather than in the equation. \[\text{Accuracy}(t) \approx \text{Accuracy}_0 - \lambda \cdot \mathcal{D}(P_t \lVert P_0)\]

Symbol Definition Unit/Type Notes
\(\text{Accuracy}(t)\) Accuracy at Time \(t\) Scalar Model accuracy after the model has been deployed for time \(t\).
\(\text{Accuracy}_0\) Initial Accuracy Scalar Model accuracy at deployment time.
\(\lambda\) Fitted Sensitivity Scalar Local coefficient estimated from outcomes over a stated range; not a universal constant. (Not wavelength.)
\(P_t\) Current Distribution Distribution The data distribution at time \(t\). (Not parameters—use \(P\) for parameter count.)
\(P_0\) Training Distribution Distribution The data distribution at training time.
\(\mathcal{D}(P_t \lVert P_0)\) Statistical Divergence Scalar \(\ge 0\) Measures how far \(P_t\) has drifted from \(P_0\). Common choices: KL divergence, total variation, Wasserstein. (Calligraphic to avoid collision with \(D\) = dataset size.)
\(\tau\) Response Threshold Scalar \(> 0\) Operational threshold for investigation; it does not automatically trigger retraining.

The Energy Corollary

With effective per-byte and per-operation costs, workload energy is approximated as: \[E_{\text{total}} \approx D_{\text{vol}} \times E_{\text{move}} + O \times E_{\text{compute}}\]

Symbol Definition Unit Notes
\(E_{\text{total}}\) Total Energy joules Total energy consumed by an ML workload, decomposed into data-movement and compute terms.
\(E_{\text{move}}\) Energy per Byte Moved joules/byte Effective movement cost; total energy also depends on bytes moved and operations executed.
\(E_{\text{compute}}\) Energy per Operation joules/FLOP Energy cost of a single arithmetic operation.

Deep Learning Notation

This book follows standard deep learning conventions (Goodfellow et al. 2016) with explicit disambiguation for systems variables.

Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press.
Symbol Definition Dimensions/Type
\(B\) Batch Size Integer. The number of samples processed in parallel. (Never bandwidth.)
\(P\) Parameters Integer. The total count of scalar model parameters, including weights and biases. (Never peak FLOP/s.)
\(D\) Dataset Size Integer. Number of training samples or tokens. (Never data volume in bytes—use \(D_{\text{vol}}\).)
\(S\) Sequence Length Integer. Number of tokens or time steps.
\(d\) Hidden Dimension Integer. Size of the hidden state vector.
\(d_{\text{head}}\) Attention Head Dimension Integer. Per-head hidden dimension in attention layers.
\(N_L\) Number of Layers Integer. Total number of layers in a network.
\(N_{\text{heads}}\) Number of Attention Heads Integer. Number of attention heads in a multi-head attention layer.
\(H_{\text{KV}}\) Number of Key-Value Heads Integer. Number of key-value heads in grouped-query or multi-query attention.
\(\ell\) Layer Index Integer. Index for a layer. Use instead of bare \(L\) when indexing layers, since \(L\) collides with loss and latency.
\(\mathcal{L}\) Loss Function Scalar. The objective function minimized during training.
\(\eta\) Learning Rate Scalar. Step size for the optimizer. (Never bare for hardware efficiency—use \(\eta_{\text{hw}}\).)
\(\theta\) Model Parameters Parameter collection, often treated as a vector. The set of all learnable parameters.
\(p(x)\) Distribution Probability mass/density for a random variable. Use lowercase \(p(\cdot)\) for generic distributions to avoid collision with \(P\) (parameter count).
\(p(y \mid x)\) Conditional Distribution Conditional probability/density. Use this form for generic label relationships; reserve \(P_0\) and \(P_t\) for the degradation equation’s training/current distributions.
\(\Pr(E)\) Event Probability Probability of an event \(E\). Use for event statements such as \(\Pr(\text{batch}=0)\); use \(p(x)\) and \(p(y \mid x)\) for distributions.

Latin matrices and vectors are set in bold (\(\mathbf{W}\), \(\mathbf{x}\)); scalars, dimensions, indices, and individual matrix entries stay italic (\(N\), \(d\), \(i\), \(W_{ij}\)). The generic parameter symbol \(\theta\) is the one conventional exception and stays italic. Local matrix algebra follows the same bold rule: matrix operands are bold (\(\mathbf{A}\), \(\mathbf{B}\), \(\mathbf{C}\)), as in general matrix multiply (GEMM) \(\mathbf{C}=\alpha \mathbf{A}\mathbf{B}+\beta \mathbf{C}\). The bold form marks a matrix operand and is distinct from the italic scalar batch size \(B\), which remains batch size in scalar contexts.

Performance, Serving, and Memory Notation

Reusable performance and serving quantities follow the same collision-avoidance rule as the iron law. Rates use \(R\) or descriptive Greek symbols; request counts use \(Q_{\text{req}}\) rather than overloading \(N\); queue utilization uses a subscripted \(\rho\) so bare \(\rho\) remains available for other ratio models.

Symbol Definition Unit/Type Notes
\(I\) Arithmetic Intensity FLOP/byte Workload FLOPs per byte moved. The Roofline Model uses \(I\) as the independent variable.
\(I_{\text{ridge}}\) Roofline Ridge Point FLOP/byte \(I_{\text{ridge}} = R_{\text{peak}}/\text{BW}\). Prefer this explicit form over starred shorthand so the meaning remains clear in prose.
\(R_{\text{attain}}\) Attainable Compute Rate FLOP/s Roofline bound: \(R_{\text{attain}} \leq \min(R_{\text{peak}}, I \times \text{BW})\). Uses \(R\), not \(T\), because this quantity is a rate.
\(\text{MFU}\) Model FLOPs Utilization Dimensionless Useful model FLOP/s divided by available peak FLOP/s. Text acronym avoids overloading \(\eta\).
\(r_{\text{comp}}\) Compression Ratio Dimensionless Uncompressed size divided by compressed size. A compressed payload has \(1/r_{\text{comp}}\) of the uncompressed size; the subscript avoids bare \(C\) collisions.
\(Q_{\text{req}}\) Request Concurrency requests Average in-flight requests in a stable serving system (\(Q_{\text{req}} = \lambda_{\text{arr}} \cdot T_{\text{lat}}\) via Little’s Law). Avoids collision with \(N\) as device count in distributed settings.
\(\lambda_{\text{arr}}\) Arrival Rate requests/s Request arrival rate. The subscript avoids collision with \(\lambda\) as degradation sensitivity or failure rate.
\(T_{\text{lat}}\) Request Time in System seconds End-to-end queueing/serving latency for Little’s Law. Distinct from \(L_{\text{lat}}\), the fixed-latency term in the iron law.
\(T_{\text{svc}}(B)\) Batch Service Time seconds Time to serve a batch of size \(B\).
\(\mu_{\text{eff}}(B)\) Effective Service Rate requests/s Batched service rate, typically \(\mu_{\text{eff}}(B) = B/T_{\text{svc}}(B)\).
\(\rho_{\text{serv}}\) Serving Utilization Dimensionless Queue/server utilization. Use instead of bare \(\rho\), which is reserved for communication-computation ratio in distributed contexts.
\(M_{\text{total}}\) Total Memory Footprint bytes Sum of explicit memory components; avoids bare \(M\) ambiguity.
\(M_{\text{weights}}\) Weight Memory bytes Memory occupied by model parameters.
\(M_{\text{gradients}}\) Gradient Memory bytes Memory occupied by stored gradients.
\(M_{\text{optimizer}}\) Optimizer-State Memory bytes Momentum, variance, master weights, and related optimizer buffers.
\(M_{\text{activations}}\) Activation Memory bytes Retained activations for the backward pass or serving intermediates.
\(s_{\text{elem}}\) Element Storage Size bytes/element Bytes per stored tensor element.

Units and Precision

  • Physical units: This book uses SI (metric) units throughout, including meters, kilograms, seconds, watts, and °C, consistent with standard engineering and scientific practice. Where source data was originally reported in imperial units, the book converts to SI and notes the original values parenthetically. A space always separates the number from the unit (for example, 100 ms, 2 TB/s).
  • Data and memory: This book uses decimal SI prefixes only: KB = \(10^3\) bytes, MB = \(10^6\), GB = \(10^9\), TB = \(10^{12}\). Binary-prefixed units do not appear in prose; all capacities, throughputs, and model sizes are reported in decimal units (for example, 80 GB, 2 TB/s, 102 MB).
  • Compute: The notation distinguishes operation counts from rates. Total work uses FLOPs (for example, GFLOPs, TFLOPs), while throughput uses FLOP/s with decimal prefixes (for example, GFLOP/s, TFLOP/s).
    • 1 TFLOP = \(10^{12}\) FLOPs
    • 1 TFLOP/s = \(10^{12}\) FLOPs per second
    • Arithmetic intensity conventionally uses FLOP/byte as a unit ratio (floating-point operations per byte moved). This is a ratio unit, not a total-work symbol or throughput symbol.
  • Currency: Dollar amounts use the dollar sign ($); unless otherwise noted, dollar-denominated costs are U.S. dollars (USD).
  • Precision:
    • FP64: Double precision (8 bytes)
    • FP32: Single precision (4 bytes)
    • TF32: TensorFloat-32 (19-bit Tensor Core compute mode; inputs and storage remain FP32)
    • FP16: Half precision (2 bytes, standard range)
    • BF16: Brain float (2 bytes, wide dynamic range)
    • FP8: Quarter precision (1 byte, E4M3 or E5M2 format)
    • FP4: 4-bit floating-point format
    • INT8: 8-bit integer (1 byte)
    • INT4: 4-bit integer; lower integer precisions follow the same uppercase INTn pattern (for example, INT3, INT2)

Quick Reference: Resolving Collisions

Common collision points in ML Systems literature include:

Symbol ML Meaning Systems Meaning Book Convention
\(B\) Batch Size Bandwidth Batch Size. Use \(\text{BW}\) for bandwidth.
\(P\) Parameters Peak FLOP/s Parameters. Use \(R_{\text{peak}}\) for peak rate.
\(D\) Dataset Size Data Volume Dataset Size. Use \(D_{\text{vol}}\) for bytes moved.
\(L\) Loss Latency Loss \((\mathcal{L})\). Use \(L_{\text{lat}}\) for latency.
\(\eta\) Learning Rate Efficiency Learning Rate. Use \(\eta_{\text{hw}}\) for efficiency.

As a general principle, ML conventions take precedence for single letters, while systems concepts get subscripts or multi-letter symbols. This reflects the primary audience (ML practitioners learning systems) and preserves compatibility with the vast ML literature.

Physical AI and Embodied Systems Notation

Physical AI and embodied machine learning systems merge three distinct technical traditions:

  1. Machine Learning and Deep Learning (statistical optimization, transformers, diffusion policies, imitation learning, autoregression, Action Chunking with Transformers).
  2. Robotics, Mechanics, and Classical Control (multibody dynamics, \(SE(3)\) Lie groups, rigid-body kinematics, Lyapunov stability, Control Barrier Functions, contact mechanics, actuator drive physics).
  3. Embedded Systems and Real-Time Architectures (system-on-chip interconnects, cache and memory bus contention, hardware interrupt latency, seqlocks, CAN-FD/EtherCAT, intent leases, fail-safe interlocks).

Because each discipline developed its mathematical notation independently over decades, combining them creates severe symbol collisions—the “traditional warts” of cyber-physical engineering. For example, writing \(\nabla_\theta \pi_\theta(\mathbf{a} \mid \mathbf{q})\) when \(\theta\) also denotes joint angles, or writing \(\mathbf{a}_t\) for an action when \(\mathbf{a}\) denotes Cartesian acceleration or barrier constraints, confuses spatial mechanics with parameter optimization and control execution authority.

Below, we formalize the cross-discipline disambiguation contract, state the governing relationships of embodied systems, and tabulate the canonical symbols used throughout this book.

Cross-Discipline Disambiguation and Traditional Warts

When physical plant dynamics meet deep neural networks and real-time embedded hardware, common mathematical shorthand breaks down. The conventions below resolve these cross-domain collisions systematically:

Symbol Canonical Convention Colliding Meaning Disambiguation Rule
\(\theta\) Neural policy network weights \(\pi_\theta\) Joint angle \(\theta_i\), trajectory polynomial \(\theta(t)\), thermal resistance \(\theta_{JA}\), or heading \(\theta\) Reserve \(\theta\) strictly for neural parameters. Robot joint configurations are always generalized coordinates \(\mathbf{q} \in \mathbb{R}^n\) and \(q_i\). Trajectory polynomials are \(q(t) = \sum c_k t^k\) (never \(\theta(t)\)). Thermal resistance is \(R_{\text{th}}\) or \(R_{\text{th}, JA}\) (in \(\text{K/W}\)). Heading or spatial angles use \(\phi, \psi\) or rotation matrices.
\(\mathbf{a}\) vs. \(\hat{\mathbf{u}}\) vs. \(\mathbf{u}^*\) vs. \(\boldsymbol{\tau}\) Four-stage authority chain (\(\mathbf{a}_t \to \hat{\mathbf{u}}_t \to \mathbf{u}^*_t \to \boldsymbol{\tau}_t\)) Generic “action” symbol colliding with deceleration \(a_{\text{brake}}\) or barrier normal \(\mathbf{a}(\mathbf{x})\) Disambiguate execution authority: \(\mathbf{a}_t \in \mathcal{A}\) is the unprivileged policy action (dimensionless/normalized); \(\hat{\mathbf{u}}_t \in \mathcal{U}\) is the proposed plant control setpoint; \(\mathbf{u}^*_t \in \mathcal{U}\) is the certified, CBF-filtered control command; \(\boldsymbol{\tau}_t \in \mathbb{R}^n\) is realized physical actuator torque. Credible linear deceleration is \(a_{\text{brake}}\) (\(\text{m/s}^2\)); barrier constraints use \(\mathbf{A}_{\text{cbf}}(\mathbf{x})\mathbf{u} \le \mathbf{b}_{\text{cbf}}(\mathbf{x})\).
\(\boldsymbol{\tau}\) vs. \(\tau\) Actuator torque vector \(\boldsymbol{\tau} \in \mathbb{R}^n\) (\(\text{N}\cdot\text{m}\) or \(\text{N}\)) Time constants (\(\tau_{\text{th}}\)), latencies (\(\tau_{\text{ZOH}}, \tau_{\text{lag}}\)), intent lease duration (\(\tau\)), or response threshold (\(\tau\)) Realized torque is always bold vector \(\boldsymbol{\tau} \in \mathbb{R}^n\) or scalar joint torque \(\tau_i\) / component torque (\(\tau_{\text{motor}}, \tau_{\text{hold}}\)). Time intervals and latencies use Latin \(T\) or lowercase \(t\) (\(T_{\text{lease}}, t_{\text{age}}, t_{\text{ZOH}}, T_{\text{blend}}\)). Thermal time constants use explicitly subscripted \(\tau_{\text{th}} = R_{\text{th}} C_{\text{th}}\). Operational thresholds use \(\tau_{\text{resp}}\). The total pre-brake delay \(\tau_{\text{delay}}\) keeps its Greek letter, the control-theory convention for a transport delay, and its subscript keeps it distinct from torque.
\(\mathbf{x}\) vs. \(s\) vs. \(\mathbf{p}\) vs. \(\mathbf{x}_p\) Mechanical plant state \(\mathbf{x} = [\mathbf{q}^\top, \dot{\mathbf{q}}^\top]^\top \in \mathbb{R}^{2n}\) World state \(s \in \mathcal{S}\), 3D position \(\mathbf{p} = [x, y, z]^\top\), or ViT image patch \(\mathbf{x}_p\) Continuous mechanical plant state is bold vector \(\mathbf{x} \in \mathbb{R}^{2n}\). Ground-truth latent world state is \(s \in \mathcal{S}\). 3D Cartesian position is vector \(\mathbf{p} \in \mathbb{R}^3\) or camera-frame point \(\mathbf{P}_c\). Vision transformer patch embeddings use explicitly subscripted \(\mathbf{x}_p\) or latent feature \(\mathbf{z}\).
\(P\) Model parameter count (\(P\) or \(P_{\text{params}}\)) Power (\(P_{\text{elec}}, P_{\text{cont}}\)), Riccati error covariance \(\mathbf{P}(t)\), failure rate \(p_{\text{fail}}\), or 3D point \(\mathbf{p}\) Use \(P\) for parameter count. Electrical/mechanical power is \(P_{\text{elec}}, P_{\text{cont}}, P_{\text{peak}}\) (in Watts); continuous Riccati/Kalman error covariance is bold matrix \(\mathbf{P}(t)\); statistical failure probability or rate is lowercase italic \(p\) (\(p_{\text{target}}, p_{\text{emp}}, p_{\text{fail}}\)); 3D Cartesian position is bold lowercase \(\mathbf{p} = [x, y, z]^\top\).
\(B\) Batch size (\(B\) or \(B_{\text{tokens}}\)) Bandwidth (\(\text{BW}\)), Body frame (\(\mathcal{F}_B\)), bus burst size (\(B_k\)), or input matrix \(\mathbf{B}\) Batch size is \(B\); sustained bandwidth is \(\text{BW}\) (per Iron Law); coordinate reference frames use calligraphic script (\(\mathcal{F}_W, \mathcal{F}_B, \mathcal{F}_E\)); bus transaction sizes use \(D_{\text{trans}, k}\) or \(B_k\); state-space control input matrix is bold capital \(\mathbf{B} \in \mathbb{R}^{n \times m}\).
\(D\) Dataset size (\(D\) in tokens or trajectories) Data volume (\(D_{\text{vol}}\)), clear distance (\(D_{\text{clear}}\)), duty cycle (\(D_{\text{cycle}}\)), TSDF field \(D_t(\mathbf{v})\), or damping matrix \(\mathbf{D}_{\text{virt}}\) Training dataset size is \(D\); bytes moved across buses is \(D_{\text{vol}}\); stopping distance is lowercase \(d_{\text{stop}}\), and the clear distance to a boundary is \(D_{\text{clear}}\), whose subscript keeps it distinct from dataset size; actuator duty cycle uses subscripted \(D_{\text{cycle}} \in [0, 1]\); voxel TSDF fields use \(D_t(\mathbf{v})\); virtual damping matrices use bold \(\mathbf{D}_{\text{virt}}\).
\(\alpha\) Bumpless transfer blending function \(\alpha(t) \in [0, 1]\) CBF decay (\(\gamma_{\text{cbf}}\)), significance (\(\alpha_{\text{sig}}\)), angular acceleration (\(\boldsymbol{\alpha}\)), copper tempco (\(\alpha_{\text{Cu}}\)), or noise schedule (\(\bar{\alpha}_k\)) Authority transfer uses scalar function \(\alpha(t) \in [0, 1]\); barrier linear decay uses \(\gamma_{\text{cbf}} > 0\); statistical confidence significance uses \(\alpha_{\text{sig}} = 1 - \text{CL}\); angular acceleration is bold vector \(\boldsymbol{\alpha} \in \mathbb{R}^3\) (\(\text{rad/s}^2\)); copper resistance tempco is \(\alpha_{\text{Cu}}\); diffusion noise schedules use indexed \(\bar{\alpha}_k\).
\(\mathbf{C}\) Coriolis matrix \(\mathbf{C}(\mathbf{q}, \dot{\mathbf{q}}) \in \mathbb{R}^{n \times n}\) Safe set \(\mathcal{C}\), state-space output matrix \(\mathbf{C}\), capacitance (\(C_{\text{th}}, C_{\text{bus}}\)), CoM \(\mathbf{c}\), or shield coverage \(c_{\text{shield}}\) Coriolis/centrifugal forces use matrix function \(\mathbf{C}(\mathbf{q}, \dot{\mathbf{q}})\); forward-invariant safe sets use calligraphic \(\mathcal{C} \subset \mathcal{S}\); state-space output mappings use bold \(\mathbf{C} \in \mathbb{R}^{p \times n}\); electrical/thermal capacitance uses italic \(C\); center of mass is bold vector \(\mathbf{c} \in \mathbb{R}^3\); safety shield diagnostic coverage uses \(c_{\text{shield}} \in [0, 1)\).
\(\mathbf{K}\) Camera intrinsic calibration matrix \(\mathbf{K} \in \mathbb{R}^{3 \times 3}\) Contact stiffness (\(k_{\text{contact}}\)), control gains (\(K_p, K_d\)), motor constants (\(K_t, K_e\)), ensembling window \(K\), or diffusion step \(k\) Camera intrinsics use bold matrix \(\mathbf{K}\); contact stiffness is lowercase scalar \(k\) or \(k_{\text{contact}}\) (\(\text{N/m}\)); feedback control gains use capitalized scalars \(K_p, K_d\); motor torque/back-EMF constants use \(K_t, K_e\); temporal ensembling window length uses integer \(K\); diffusion noise steps use lowercase index \(k\).
\(\mathbf{M}\) Articulated inertia matrix \(\mathbf{M}(\mathbf{q}) \in \mathbb{R}^{n \times n}\) Mass (\(m, m_{\text{eff}}\)), memory footprint (\(M_{\text{weights}}, M_{\text{KV}}\)), or fleet size (\(N_{\text{fleet}}\)) Articulated generalized mass/inertia uses bold matrix \(\mathbf{M}(\mathbf{q})\); physical body mass is lowercase italic \(m\) (\(\text{kg}\)); effective impact mass is \(m_{\text{eff}}\); hardware memory storage uses capital \(M\) with descriptive subscript (in bytes); robot fleet size uses \(N_{\text{fleet}}\) or \(M_{\text{fleet}}\).

Governing Relationships

Embodied intelligence operates subject to thermodynamic, mechanical, kinematic, and communications constraints. The relationships below state those constraints in the symbols the chapters use.

1. Sense-to-actuation delay budget

Delay converts directly into travel before any retarding force begins. The total pre-brake delay \(\tau_{\text{delay}}\) runs from detection to retarding force, and the stop it leaves must fit inside the clear distance:

\[\tau_{\text{delay}} = t_{\text{age}} + T_{\text{lease}} + T_{\text{tick}} + T_{\text{bus}} + T_{\text{act}}\]

\[v\,\tau_{\text{delay}} + \frac{v^2}{2a_{\text{eff}}} + \delta_{\text{loc}} + \delta_{\text{margin}} + \epsilon_{\text{track}} \le D_{\text{clear}}\]

Inference time is spent on the proposal side of the proposal boundary. It ages the evidence behind a proposal, which the permission path checks against the proposal’s expiry, but it is neither a term of \(\tau_{\text{delay}}\) nor part of the permission loop’s period \(T_{\text{tick}}\), because the fallback does not wait for the proposer.

Symbol Definition Unit Notes
\(\tau_{\text{delay}}\) Total Pre-Brake Delay \(\text{ms}\) Time from detection of a hazard to the onset of retarding force; the machine travels \(v\,\tau_{\text{delay}}\) before braking begins.
\(t_{\text{age}}\) Observation Age \(\text{ms}\) Age of the state the permission path acts on, from the exposure midpoint at the sensor to the decision, which enters \(\tau_{\text{delay}}\).
\(T_{\text{lease}}\) Permission-Path Lease \(\text{ms}\) Longest interval a proposal may keep driving the actuators without renewal, which enters \(\tau_{\text{delay}}\); chosen, then checked against the ceiling the stopping budget derives.
\(T_{\text{tick}}\) Permission-Loop Period \(\text{ms}\) Period of the permission loop, which enters \(\tau_{\text{delay}}\) because a hazard may wait up to one period before it is checked.
\(T_{\text{bus}}\) Command Transport Time \(\text{ms}\) Fieldbus delivery time from the permission path to the drive, which enters \(\tau_{\text{delay}}\).
\(T_{\text{act}}\) Actuator Response Time \(\text{ms}\) Time from a delivered stop command to retarding force, including gate-drive and current rise, which enters \(\tau_{\text{delay}}\).
\(t_{\text{sense}}\) Sensor Integration Time \(\text{ms}\) Half the exposure duration (its integration midpoint), ADC conversion, and sensor readout latency; contributes to \(t_{\text{age}}\).
\(t_{\text{xfer}}\) Interconnect Transfer Time \(\text{ms}\) DMA bus streaming, PCIe handoff, and serializer-deserializer (SerDes) link latency; contributes to \(t_{\text{age}}\).
\(t_{\text{infer}}\) Proposer Inference Time \(\text{ms}\) Forward-pass time of the learned proposer; spent above the proposal boundary and outside the permission loop’s period.

2. Action-chunk streaming cadence

Cognitive neural policies cannot execute at kilohertz frequencies due to memory bandwidth limits. Action chunking predicts an action trajectory of horizon \(H\) at a lower cognitive frequency \(f_{\text{policy}}\), while the chunk’s waypoints are consumed at \(f_{\text{exec}}\). The memory streaming latency must remain bounded by the chunk execution window:

\[t_{\text{stream}} = \frac{M_{\text{weights}}}{\text{BW}_{\text{mem}}} \le \frac{H}{f_{\text{exec}}} - t_{\text{compute}}\]

Symbol Definition Unit Notes
\(H\) Action Chunk Horizon Time steps Number of contiguous future action steps generated in a single policy forward pass (\(H=16\) on the mobile manipulator; 16 to 64 steps in the literature).
\(M_{\text{weights}}\) Model Parameter Footprint bytes Total memory required to store policy weights (\(P \times \text{bytes-per-parameter}\)).
\(\text{BW}_{\text{mem}}\) Sustained Memory Bandwidth bytes/s Achievable memory throughput of the SoC or accelerator memory subsystem (e.g., LPDDR5, HBM).
\(t_{\text{stream}}\) Parameter Streaming Time \(\text{ms}\) Time required to stream non-resident model weights into execution registers.
\(f_{\text{exec}}\) Action Execution Rate \(\text{Hz}\) Temporal consumption rate of chunked waypoints along the planned trajectory (\(100\text{ Hz}\) to \(1\text{ kHz}\)).

3. The dynamic stopping, braking, and safe-set invariants

Forward invariance of the safe set \(\mathcal{C}\) requires that the machine can arrest all motion before it reaches the boundary. The stopping distance combines the travel during the pre-brake delay with the braking distance, and its margins are added in the inequality of the delay budget above:

\[d_{\text{stop}}(v) = v\,\tau_{\text{delay}} + \frac{v^2}{2 a_{\text{eff}}}\]

Symbol Definition Unit Notes
\(d_{\text{stop}}\) Stopping Distance \(\text{m}\) Distance required to bring the plant to rest from speed \(v\).
\(D_{\text{clear}}\) Clear Distance \(\text{m}\) Clear distance from the machine to the boundary it must not cross; the subscript keeps it distinct from dataset size \(D\).
\(v\) Speed at Detection \(\text{m/s}\) Speed of the machine at the moment the hazard is detected.
\(a_{\text{brake}}\) Credible Deceleration \(\text{m/s}^2\) Validated low tail of the deceleration the machine achieves under its stated conditions, not the datasheet peak.
\(a_{\text{eff}}\) Effective Deceleration \(\text{m/s}^2\) Deceleration used in the braking term: \(a_{\text{brake}}\) for a constant-deceleration stop, \(a_{\text{brake}}/1.5\) when the resident stop is the \(C^2\) stopping suffix of equation.
\(\delta_{\text{loc}}\) Localization Bound \(\text{m}\) Bound on localization and state-estimation error.
\(\delta_{\text{margin}}\) Protective Clearance \(\text{m}\) Fixed, nonzero protective clearance between the machine envelope and the boundary.
\(\epsilon_{\text{track}}\) Tracking-Error Bound \(\text{m}\) Bound on the deviation between the commanded and the executed trajectory, stated by the enforcement record.
\(v_h\) Human Walking-Speed Bound \(\text{m/s}\) Walking-speed bound for a person approaching the machine, used in the form its safety standard states.
\(j_{\max}\) Jerk Limit \(\text{m/s}^3\) Maximum rate of change of deceleration; appears only in jerk-limited refinements of the braking term, which the book’s stopping equations omit.

4. The digital sampling zero-order hold (ZOH) phase margin

Holding discrete control commands constant over sampling period \(T_s = 1/f_s\) introduces an effective transport delay of half the sample period, eroding feedback stability margins:

\[\tau_{\text{ZOH}} = \frac{T_s}{2}, \qquad \Delta \Phi_{\text{total}} = \omega_c \left( \tau_{\text{comp}} + \tau_{\text{bus}} + \frac{T_s}{2} \right) \le \text{PM}_{\text{margin}}\]

Symbol Definition Unit Notes
\(\tau_{\text{ZOH}}\) ZOH Transport Delay \(\text{s}\) Equivalent continuous delay introduced by discrete Zero-Order Hold reconstruction (\(T_s / 2\)).
\(T_s, f_s\) Sampling Period and Rate \(\text{s}, \text{Hz}\) Discrete controller execution period and frequency (\(f_s = 1/T_s\)).
\(\omega_c\) Open-Loop Crossover Frequency \(\text{rad/s}\) Frequency at which open-loop gain \(\lvert L(j\omega_c) \rvert = 1\) (\(0\text{ dB}\)).
\(\Delta \Phi\) Phase Lag Degradation \(\text{rad}\) / \(\text{deg}\) Phase erosion (\(\Delta \Phi \approx 28.65^\circ \cdot (\omega_c / f_s)\)) depleting closed-loop stability margin.
\(\text{PM}_{\text{net}}\) Net Phase Margin \(\text{deg}\) Residual stability margin (\(\text{PM}_{\text{plant}} - \Delta \Phi_{\text{total}} > 0\)) preventing contact chatter.

5. Actuator electromechanical and thermal rise invariants

Actuators are constrained by copper winding heating, DC bus rail droop, and magnetic saturation. Sustainable torque output is governed by the coupled electromechanical balance:

\[V_{\text{bus}} = L \frac{di}{dt} + R(T) i + K_e \omega_m, \qquad C_{\text{th}} \frac{dT}{dt} = I^2 R(T) - \frac{T - T_{\text{amb}}}{R_{\text{th}}}, \qquad \boldsymbol{\tau} = \eta N_{\text{gear}} K_t I\]

Symbol Definition Unit Notes
\(K_t, K_e\) Torque and Back-EMF Constants \(\text{N}\cdot\text{m/A}, \text{V}\cdot\text{s/rad}\) Motor electromechanical coupling constants (\(K_t \approx K_e\) in SI units).
\(R(T)\) Temperature-Dependent Resistance \(\Omega\) Stator copper resistance scaling: \(R(T) = R_0 [1 + \alpha_{\text{Cu}}(T - T_0)]\) with \(\alpha_{\text{Cu}} = 0.00393\text{ K}^{-1}\).
\(R_{\text{th}}\) Thermal Resistance (\(\theta_{JA}\)) \(\text{K/W}\) Lumped thermal resistance from stator windings to ambient environment.
\(C_{\text{th}}\) Stator Thermal Capacitance \(\text{J/K}\) Lumped heat capacity of the motor core, windings, and potting compound.
\(\tau_{\text{th}}\) Thermal Time Constant \(\text{s}\) Actuator heating characteristic time constant (\(\tau_{\text{th}} = R_{\text{th}} C_{\text{th}}\), typically \(60\text{--}300\text{ s}\)).
\(D_{\text{cycle}}\) Actuator Duty Cycle \([0, 1]\) Ratio of active pulse duration to cycle period (\(t_{\text{on}} / (t_{\text{on}} + t_{\text{off}})\)) preventing thermal runaway.
\(\Delta V_{\text{droop}}\) DC Bus Dynamic Rail Droop \(\text{V}\) Transient voltage collapse (\(\Delta I R_{\text{bus}} + L_{\text{bus}}\frac{dI}{dt} + \frac{\Delta I \Delta t}{2 C_{\text{bus}}}\)) under step motor current loads.

6. Shielded hazard architecture and statistical exposure bound

Catastrophic failure rates in embodied systems cannot be certified through stochastic testing alone. A deterministic safety shield bounds overall system risk, while zero-failure trial counts follow the Rule of Three:

\[p_{\text{system}} = p_{\text{brain}} (1 - c_{\text{shield}}) \le p_{\text{target}}, \qquad n \ge \frac{-\ln \alpha_{\text{sig}}}{p}\]

Symbol Definition Unit Notes
\(p_{\text{system}}\) Overall System Hazard Rate \(\text{failures/hour}\) Certified dangerous failure rate of the integrated cyber-physical machine.
\(p_{\text{brain}}\) Neural Policy Defect Rate \(\text{failures/hour}\) Empirical rate of unshielded hallucinations, distribution drift, or policy blunders.
\(c_{\text{shield}}\) Shield Diagnostic Coverage Dimensionless Fraction of hazardous neural commands successfully intercepted by deterministic QP/interlocks (\(c \ge 0.9999\)).
\(p_{\text{target}}\) Target Risk Tolerance \(\text{failures/hour}\) Statutory regulatory safety requirement (e.g., \(10^{-9}\text{ failures/h}\) for SIL 4 / DAL A).
\(\alpha_{\text{sig}}\) Statistical Significance Level Dimensionless Type I error threshold (\(\alpha_{\text{sig}} = 1 - \text{CL}\), e.g., \(\alpha=0.05\) for 95% confidence where \(-\ln(0.05) \approx 3.0\)).
\(n\) Zero-Failure Sample Exposure Hours or Cycles Required zero-incident validation exposure (\(n \approx 3/p\)) to substantiate failure rate \(p\).

1. Physical Plant State, Actuation, and Kinematic Topology

Symbol Definition Unit Notes
\(s_t \in \mathcal{S}\) Physical World State Mixed Ground-truth continuous state of the machine, environment, and physical contacts at time \(t\).
\(\mathbf{x}\) Mechanical Plant State \(\text{rad, rad/s}\) Generalized state vector \(\mathbf{x} = [\mathbf{q}^\top, \dot{\mathbf{q}}^\top]^\top \in \mathbb{R}^{2n}\) containing positions and velocities.
\(\mathbf{q} \in \mathbb{R}^n\) Generalized Coordinates \(\text{rad}\) or \(\text{m}\) Degrees of freedom describing mechanical articulation (joint angles and prismatic link extensions).
\(\dot{\mathbf{q}} \in \mathbb{R}^n\) Generalized Velocities \(\text{rad/s}\) or \(\text{m/s}\) First time derivative of generalized coordinates.
\(\ddot{\mathbf{q}} \in \mathbb{R}^n\) Generalized Accelerations \(\text{rad/s}^2\) or \(\text{m/s}^2\) Second time derivative of generalized coordinates.
\(\boldsymbol{\tau} \in \mathbb{R}^n\) Actuator Torque Vector \(\text{N}\cdot\text{m}\) or \(\text{N}\) Generalized forces and torques delivered by actuators at mechanical joints.
\(\mathbf{M}(\mathbf{q})\) Generalized Inertia Matrix \(\text{kg}\cdot\text{m}^2\) Symmetric positive-definite mass/inertia matrix of the articulated multibody plant.
\(\mathbf{C}(\mathbf{q}, \dot{\mathbf{q}})\) Coriolis / Centrifugal Matrix \(\text{N}\cdot\text{m}\cdot\text{s/rad}\) Matrix satisfying \(\dot{\mathbf{M}} - 2\mathbf{C}\) skew-symmetry (passivity property).
\(\mathbf{g}(\mathbf{q})\) Gravitational Torque Vector \(\text{N}\cdot\text{m}\) Generalized gravitational load acting on links as a function of configuration.
\(\mathbf{J}(\mathbf{q})\) Manipulator Jacobian Mixed Linear mapping \(\mathbf{v}_e = \mathbf{J}(\mathbf{q})\dot{\mathbf{q}}\) from joint rates to operational task-space velocity.
\(w(\mathbf{q})\) Yoshikawa Manipulability Dimensionless Singularity proximity metric \(w(\mathbf{q}) = \sqrt{\det(\mathbf{J}\mathbf{J}^\top)}\) vanishing at kinematic singular configurations.
\(K_t, K_e\) Motor Torque / Back-EMF \(\text{N}\cdot\text{m/A}, \text{V}\cdot\text{s/rad}\) Electromechanical coupling constants linking winding current to rotor shaft torque and back-EMF voltage.
\(e_{\text{bemf}}\) Back-EMF Voltage \(\text{V}\) Opposing voltage \(e_{\text{bemf}} = K_e \omega_m\) developed across stator coils by rotor rotation.
\(J_{\text{rotor}}, J_{\text{load}}\) Rotor and Load Inertia \(\text{kg}\cdot\text{m}^2\) Moment of inertia of motor rotor shaft and coupled mechanical load.
\(J_{\text{ref}}\) Reflected Rotor Inertia \(\text{kg}\cdot\text{m}^2\) Effective inertia \(J_{\text{ref}} = N_{\text{gear}}^2 J_{\text{rotor}}\) amplified quadratically by gearbox ratio.
\(N_{\text{gear}}\) Gear Reduction Ratio Dimensionless Transmission reduction ratio (\(N_{\text{gear}} \gg 1\)) between motor rotor and joint output shaft.
\(N^*\) Optimal Inertia-Matching Ratio Dimensionless Ratio \(N^* = \sqrt{J_{\text{load}} / J_{\text{rotor}}}\) maximizing load angular acceleration under peak torque.
\(\eta_{\text{gear}}\) Transmission Efficiency Dimensionless Mechanical power transmission efficiency (\(0 < \eta_{\text{gear}} \le 1\)).
\(\tau_{\text{hold}}, \tau_{\text{peak}}\) Continuous / Peak Torque \(\text{N}\cdot\text{m}\) Thermally continuous holding torque and transient peak acceleration torque.
\(I_{\text{cont}}, I_{\text{peak}}\) Continuous / Peak Current \(\text{A}\) RMS current ratings bounded by winding insulation class and inverter MOSFET dissipation.

2. Rigid-Body Contact Mechanics, Friction Cones, and Centroidal Dynamics

Symbol Definition Unit Notes
\(\mathbf{S} \in \mathbb{R}^{n \times (n+6)}\) Actuation Selection Matrix Matrix Matrix \([\mathbf{0}_{n \times 6}, \mathbf{I}_{n \times n}]\) projecting generalized forces onto actuated joints in floating-base systems.
\(\mathbf{f}_k \in \mathbb{R}^3\) Contact Force Vector \(\text{N}\) Physical contact force exerted by environment on contact patch \(k\) (\(\mathbf{f}_k = [\mathbf{f}_{k, \parallel}^\top, f_{k, \perp}]^\top\)).
\(\mathbf{J}_k(\mathbf{q})\) Contact Point Jacobian Mixed Mapping joint velocities to contact patch Cartesian linear velocity: \(\dot{\mathbf{p}}_k = \mathbf{J}_k(\mathbf{q}) \dot{\mathbf{q}}\).
\(f_{k, \perp}\) Normal Contact Force \(\text{N}\) Unilateral compressive force (\(f_{k, \perp} = \mathbf{f}_k^\top \hat{\mathbf{n}}_k \ge 0\)) perpendicular to contact plane.
\(\mathbf{f}_{k, \parallel}\) Tangential Contact Shear \(\text{N}\) Friction traction force tangential to contact plane (\(\mathbf{f}_{k, \parallel} = \mathbf{f}_k - f_{k, \perp} \hat{\mathbf{n}}_k\)).
\(\mu\) Coulomb Friction Coefficient Dimensionless Static/kinetic coefficient bounding friction cone: \(\lVert \mathbf{f}_{k, \parallel} \rVert \le \mu f_{k, \perp}\).
\(\mathbf{c} \in \mathbb{R}^3\) Center of Mass (CoM) \(\text{m}\) Spatial coordinates of the total articulated body mass centroid (\(\mathbf{c} = \frac{1}{m} \sum m_i \mathbf{p}_i\)).
\(\mathbf{h}_{\text{lin}} \in \mathbb{R}^3\) Centroidal Linear Momentum \(\text{kg}\cdot\text{m/s}\) Total plant linear momentum (\(\mathbf{h}_{\text{lin}} = m \dot{\mathbf{c}}\)), governed by \(\dot{\mathbf{h}}_{\text{lin}} = \sum \mathbf{f}_k + m \mathbf{g}\).
\(\mathbf{h}_{\text{ang}} \in \mathbb{R}^3\) Centroidal Angular Momentum \(\text{kg}\cdot\text{m}^2/\text{s}\) Angular momentum about CoM: \(\dot{\mathbf{h}}_{\text{ang}} = \sum (\mathbf{p}_k - \mathbf{c}) \times \mathbf{f}_k + \boldsymbol{\tau}_k\).
\(\mathbf{p}_{\text{ZMP}} \in \mathbb{R}^2\) Zero Moment Point (ZMP) \(\text{m}\) Ground-plane point where net horizontal tipping moment vanishes; must reside in \(\text{Conv}(\{\mathbf{p}_k\})\).
\(e \in [0, 1]\) Coefficient of Restitution Dimensionless Kinematic velocity rebound ratio during non-smooth impact (\(\dot{q}^+ = -e \dot{q}^-\)).
\(k, k_{\text{tissue}}\) Contact Stiffness \(\text{N/m}\) Elastic stiffness of structural environment or human biomechanical tissue.
\(m_{\text{eff}}\) Effective Impact Mass \(\text{kg}\) Apparent mass of moving kinematic links along collision trajectory (\(m_{\text{eff}} = (\mathbf{u}_{\text{dir}}^\top \mathbf{M}^{-1} \mathbf{u}_{\text{dir}})^{-1}\)).
\(F_{\text{unbraked}}\) Unbraked Impact Force \(\text{N}\) Transient collision force peak (\(F_{\text{unbraked}} = v_{\max} \sqrt{k_{\text{tissue}} m_{\text{eff}}}\)) governed by ISO/TS 15066.
\(E_{k, \text{allow}}\) Allowable Collision Energy \(\text{Joules (J)}\) Statutory kinetic energy limit (\(\frac{1}{2} m_{\text{eff}} v^2 \le E_{k, \text{allow}}\)) for transient collaborative contact.

3. Geometry, Spatial Transformations, and Perception Projections

Symbol Definition Unit Notes
\(SE(3)\) Special Euclidean Group Lie Group Group of 3D rigid-body spatial transformations (translations and rotations: \(\mathbb{R}^3 \rtimes SO(3)\)).
\(SO(3)\) Special Orthogonal Group Lie Group Group of \(3 \times 3\) proper orthogonal rotation matrices (\(\mathbf{R}^\top \mathbf{R} = \mathbf{I}, \det(\mathbf{R}) = +1\)).
\(\mathcal{F}_W, \mathcal{F}_B, \mathcal{F}_E\) Coordinate Reference Frames Reference Fixed World frame (\(\mathcal{F}_W\)), mobile robot Base/body frame (\(\mathcal{F}_B\)), and End-effector frame (\(\mathcal{F}_E\)).
\(\mathbf{p} \in \mathbb{R}^3\) Cartesian 3D Translation \(\text{m}\) Translation vector \([x, y, z]^\top\) expressing relative point coordinates.
\(\mathbf{R} \in SO(3)\) Directional Cosine Matrix Matrix \(3 \times 3\) orthonormal matrix parameterizing spatial attitude and orientation.
\(\mathbf{q}_{\text{quat}} \in \mathbb{H}\) Unit Quaternion Dimensionless Hypercomplex orientation vector \([w, x, y, z]^\top\) satisfying \(\lVert \mathbf{q}_{\text{quat}} \rVert = 1\).
\(\mathbf{T}_{A B} \in SE(3)\) Homogeneous Transform Matrix \(4 \times 4\) matrix \(\begin{bmatrix} \mathbf{R}_{AB} & \mathbf{p}_{AB} \\ \mathbf{0}^\top & 1 \end{bmatrix}\) mapping coordinates from frame \(B\) into frame \(A\).
\(\mathbf{Ad}_T\) Adjoint Representation Matrix \(6 \times 6\) transformation matrix mapping spatial twists and wrenches between coordinate frames.
\(\boldsymbol{\omega}, \boldsymbol{\alpha} \in \mathbb{R}^3\) Angular Velocity & Acceleration \(\text{rad/s}, \text{rad/s}^2\) Spatial angular velocity and angular acceleration vectors of a rigid body.
\(\mathbf{r}_{P/O} \in \mathbb{R}^3\) Sensor Lever-Arm Vector \(\text{m}\) Displacement vector from body rotation origin \(O\) to displaced sensor location \(P\).
\(\mathbf{K} \in \mathbb{R}^{3 \times 3}\) Pinhole Camera Intrinsics Pixels Calibration matrix \(\begin{bmatrix} f_x & 0 & c_x \\ 0 & f_y & c_y \\ 0 & 0 & 1 \end{bmatrix}\) parameterizing focal lengths and principal center.
\(\mathbf{P}_c = [X, Y, Z]^\top\) Camera-Frame 3D Coordinate \(\text{m}\) Metric coordinates of a spatial point in the optical camera frame (\(Z_c\) is optical depth).
\([u, v]^\top\) Image Plane Pixel Coordinates Pixels Discrete sensor pixel coordinates (\(u \sim f_x \frac{X}{Z} + c_x\), \(v \sim f_y \frac{Y}{Z} + c_y\)).
\(t_{\text{row}}\) Rolling Shutter Row Readout \(\mu\text{s}\) Exposure time delta between adjacent CMOS sensor scanlines.
\(\Delta \mathbf{x}(y)\) Rolling Shutter Distortion \(\text{m}\) Kinematic skew displacement \(\Delta \mathbf{x} = \mathbf{v} \cdot y \cdot t_{\text{row}}\) induced by high-velocity platform motion.
\(\text{tsdf}(\mathbf{v})\) Truncated Signed Distance \([-1, +1]\) Normalized volumetric distance \(\text{tsdf}(\mathbf{v}) = \text{clamp}(d(\mathbf{v})/\mu_{\text{tsdf}}, -1, 1)\) of voxel \(\mathbf{v}\) to nearest surface.
\(\mu_{\text{tsdf}}\) TSDF Truncation Band \(\text{m}\) Metric thickness of the spatial surface interface band stored in voxel memory.
\(D_t(\mathbf{v}), W_t(\mathbf{v})\) TSDF Accumulated Volume Mixed Running weighted signed distance and accumulated visibility weight at voxel \(\mathbf{v}\).
\(\boldsymbol{\Sigma}_{\text{depth}}\) Perception Depth Covariance \(\text{m}^2\) Anisotropic 3D uncertainty covariance ellipsoid associated with stereo/lidar/depth lifting.
\(t_{\text{age}}\) Observation Age \(\text{ms}\) The observation age of the delay budget, split into its perception-side terms: the exposure midpoint \(\tfrac{1}{2}t_{\text{exp}}\), sensor readout, transfer, and perception processing up to the decision.

4. Classical and Modern Control (State-Space, Observers, CBF-QPs)

Symbol Definition Unit Notes
\(\mathbf{A}, \mathbf{B}, \mathbf{C}\) Linear State-Space Matrices Matrix Continuous LTI plant: \(\dot{\mathbf{x}} = \mathbf{A}\mathbf{x} + \mathbf{B}\mathbf{u} + \mathbf{w}\), observed output \(\mathbf{y} = \mathbf{C}\mathbf{x} + \mathbf{v}\).
\(\mathbf{w}(t), \mathbf{v}(t)\) Process and Measurement Noise Vector Zero-mean Gaussian noise vectors with covariance matrices \(\mathbf{Q} \succeq 0\) and \(\mathbf{R} \succ 0\).
\(\mathbf{P}(t) \in \mathbb{R}^{n \times n}\) Estimation Error Covariance Matrix Symmetric positive-definite matrix \(\mathbb{E}[(\mathbf{x} - \hat{\mathbf{x}})(\mathbf{x} - \hat{\mathbf{x}})^\top]\) evolving via matrix Riccati equation.
\(\mathcal{O}, \mathcal{O}_{\text{NL}}\) Observability Matrices Matrix Kalman observability matrix \([\mathbf{C}^\top, (\mathbf{CA})^\top, \dots]^\top\) and nonlinear Lie matrix \([d\mathbf{h}, dL_{\mathbf{f}}\mathbf{h}, \dots]^\top\).
\(\mathcal{C} \subset \mathcal{S}\) Forward-Invariant Safe Set Set Sublevel/superlevel set \(\mathcal{C} = \{\mathbf{x} \mid h(\mathbf{x}) \ge 0\}\) where safety invariants are strictly maintained.
\(\partial \mathcal{C}\) Safety Boundary Manifold Boundary Zero level set \(\{\mathbf{x} \mid h(\mathbf{x}) = 0\}\) separating nominal operation from irreversible hazard.
\(h(\mathbf{x}) \ge 0\) Control Barrier Function (CBF) Scalar Continuously differentiable function whose superlevel set defines forward-invariant set \(\mathcal{C}\).
\(L_{\mathbf{f}} h(\mathbf{x})\) Lie Derivative along Drift \(\text{s}^{-1}\) Directional gradient derivative \(\nabla h(\mathbf{x}) \mathbf{f}(\mathbf{x})\) along unforced plant dynamics.
\(L_{\mathbf{g}} h(\mathbf{x})\) Lie Derivative along Control Vector Directional control effectiveness gradient \(\nabla h(\mathbf{x}) \mathbf{g}(\mathbf{x})\).
\(r\) Barrier Relative Degree Integer Number of differentiations required before control input \(\mathbf{u}\) appears (\(L_{\mathbf{g}} L_{\mathbf{f}}^{r-1} h \neq 0\)).
\(h_{\text{kinetic}}(\mathbf{x})\) Kinetic Inset Barrier Function Scalar Relative-degree-1 barrier: \(h_{\text{kinetic}}(\mathbf{x}) = h(\mathbf{x}) - v^2 / (2 a_{\text{brake}})\) incorporating stopping distance.
\(\gamma_{\text{cbf}}\) Linear Barrier Decay Rate \(\text{s}^{-1}\) Extended class-\(\mathcal{K}\) linear decay coefficient in barrier condition \(\dot{h} + \gamma_{\text{cbf}} h \ge 0\).
\(\mathbf{u}_{\text{nom}} \in \mathcal{U}\) Nominal Control Command \(\text{N}\cdot\text{m}\) / \(\text{rad}\) Unverified control candidate proposed across the causal boundary by learned policy.
\(\mathbf{u}^* \in \mathcal{U}\) Certified Control Action \(\text{N}\cdot\text{m}\) / \(\text{rad}\) Privileged, safety-filtered control command resolved by the real-time active-set CBF-QP.
\(\alpha(\hat{t})\) Bumpless Transfer Blend \([0, 1]\) \(C^2\) quintic polynomial \(6\hat{t}^5 - 15\hat{t}^4 + 10\hat{t}^3\) ramping between prior and candidate controllers.
\(\tau_{\text{blend}}, T_{\text{blend}}\) Blending Transition Window \(\text{s}\) Duration of bumpless transfer; bounded by \(\tau_{\text{blend}, \min} = 1.875 \lVert \Delta \mathbf{u} \rVert / j_{\text{allowable}}\).
\(j(t) = \dot{\mathbf{u}}(t)\) Control Setpoint Jerk \(\text{m/s}^3, \text{N}\cdot\text{m/s}\) Time derivative of commanded plant acceleration or actuator torque during mode transitions.
\(K_p, K_d, K_i\) PID / Feedback Gains Mixed Proportional, derivative, and integral gains for joint-level or operational-space error stabilization.
\(\mathbf{D}_{\text{virt}}, \mathbf{K}_{\text{virt}}\) Virtual Damping and Stiffness Matrix Target impedance/admittance matrices in compliant interaction: \(\boldsymbol{\tau}_{\text{safe}} = -\mathbf{D}_{\text{virt}} \dot{\mathbf{q}} + \mathbf{g}(\mathbf{q})\).

5. Deep Learning for Robotics (Diffusion Policies, ACT, VLAs)

Symbol Definition Unit Notes
\(\pi_\theta(\mathbf{a} \mid \mathbf{o})\) Parameterized Neural Policy Distribution Deep policy network parameterized by weights \(\theta\), mapping observation tensors to action distributions.
\(\mathbf{o}_t \in \mathcal{O}\) Multimodal Observation Mixed Synchronized tokenized sensor packet (RGB images, depth latents, joint encoders, tactile arrays).
\(\mathbf{a}_t \in \mathcal{A}\) Cognitive Policy Action Normalized Unprivileged action emitted by neural policy \(\pi_\theta\) (normalized joint delta, target pose, or gripper command).
\(\mathbf{A}_{t:t+H}\) Action Chunk Trajectory Tensor Contiguous horizon sequence \([\mathbf{a}_t, \mathbf{a}_{t+1}, \dots, \mathbf{a}_{t+H-1}] \in \mathbb{R}^{H \times d_a}\) predicted in a single forward pass.
\(H\) Chunk Prediction Horizon Time steps Number of future action steps emitted per inference cycle (16 to 64 steps in the literature).
\(K\) Ensembling Window Length Time steps Number of overlapping past chunks blended in receding horizon execution (\(K \le H\)).
\(w_i\) Temporal Ensembling Weight Scalar Exponential decay weighting factor \(w_i = \exp(-m \cdot i)\) applied to chunk predictions at offset \(i\).
\(\mathbf{z} \in \mathbb{R}^d\) Latent Style / State Embedding Dimensionless Compressed semantic representation or CVAE latent style variable in ACT policies.
\(\bar{\alpha}_k\) Diffusion Schedule Parameter \([0, 1]\) Cumulative noise variance schedule parameter \(\bar{\alpha}_k = \prod_{s=1}^k (1 - \beta_s)\) at diffusion step \(k\).
\(\boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\) Diffusion Gaussian Noise Tensor Standard normal Gaussian noise corrupted onto clean action trajectories.
\(\boldsymbol{\epsilon}_\theta(\mathbf{o}, \mathbf{A}^k, k)\) Noise Prediction Network Tensor Neural network predicting added noise in Denoising Diffusion Probabilistic Models (DDPM).
\(\mathcal{L}_{\text{diff}}(\theta)\) Diffusion Matching Loss Scalar Mean-squared error objective: \(\mathbb{E}[\lVert \boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta(\mathbf{o}_t, \mathbf{A}_t^k, k) \rVert^2]\) optimized across denoising steps.
\(\mathcal{L}_{\text{ACT}}(\theta, \phi)\) ACT Variational Objective Scalar CVAE objective combining trajectory reconstruction \(\mathcal{L}_{\text{recon}}\) with latent KL-regularization \(\beta D_{\text{KL}}\).
\(q_\phi(\mathbf{z} \mid \mathbf{o}, \mathbf{A})\) CVAE Recognition Encoder Distribution Variational encoder in Action Chunking with Transformers parameterizing latent distribution \(\mathcal{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})\).
\(\epsilon_{\text{BC}}\) Single-Step Imitation Error \([0, 1]\) Bounded probability of an unforced policy error during an individual execution step.
\(O(T^2 \epsilon_{\text{BC}})\) Compounding Covariate Drift Error bound Quadratic compounding trajectory drift under open-loop behavioral cloning rollouts of horizon \(T\).
\(\mathbf{W}_{\text{patch}}, \mathbf{E}_{\text{pos}}\) ViT Projection & Positional Bias Matrix Convolutional/linear patch projection matrix and spatial coordinate position embedding tensors.
\(\mathbf{Z}\) Multimodal Fused Token Stream Tensor Concatenated token sequence \([\mathbf{z}_{\text{vision}}, \mathbf{z}_{\text{text}}, \mathbf{z}_{\text{proprio}}]\) ingested by embodied transformer backbones.
\(M_{\text{KV}}\) Transformer KV-Cache Footprint bytes Autoregressive key-value cache memory footprint (\(2 \times N_L \times d \times S \times s_{\text{elem}}\)).
\(\Lambda_{\text{intent}}\) Expiring Intent Lease Permit Cryptographically signed, bounded time-to-live operational lease granted to learned policies.

6. Embedded Systems, Silicon Real-Time, Thermal, and Hardware Safety

Symbol Definition Unit Notes
\(I_{\text{phase}}\) Motor Phase RMS Current \(\text{Amperes (A)}\) Electrical phase current driving stator windings in brushless PM motors (\(I = \tau / (\eta N K_t)\)).
\(R_{\text{phase}}, R(T)\) Phase Winding Resistance \(\Omega\) Stator copper DC resistance at temperature \(T\) (\(R(T) = R_0 [1 + \alpha_{\text{Cu}}(T - T_0)]\)).
\(R_{\text{th}}, \theta_{JA}\) Lumped Thermal Resistance \(\text{K/W}\) Thermal impedance between internal hot spots and heatsink/ambient casing.
\(C_{\text{th}}\) Stator Thermal Capacitance \(\text{J/K}\) Stator core heat capacity governing temperature rise dynamics (\(\tau_{\text{th}} = R_{\text{th}} C_{\text{th}}\)).
\(\Delta T\) Joulean Temperature Rise \(\text{Kelvin (K)}\) Steady-state or transient thermal rise (\(\Delta T = P_{\text{cont}} R_{\text{th}}\)).
\(L_{\text{bus}}, R_{\text{bus}}\) Power Rail Inductance / ESR \(\mu\text{H}, \text{m}\Omega\) Parasitic power wiring harness inductance and Equivalent Series Resistance feeding inverters.
\(C_{\text{bus}}\) DC Bus Bulk Capacitance \(\mu\text{F}\) Inverter decoupling capacitor bank buffering transient inductive voltage collapse.
\(\Delta V_{\text{droop}}\) Dynamic Rail Voltage Droop \(\text{Volts (V)}\) Transient voltage dip (\(\Delta I R_{\text{bus}} + L_{\text{bus}}\frac{dI}{dt} + \frac{\Delta I \Delta t}{2 C_{\text{bus}}}\)) triggering undervoltage reset.
\(t_{\text{vulnerable}}\) Seqlock Vulnerable Interval \(\mu\text{s}\) Race window (\(t_{\text{write}} + t_{\text{read}}\)) during lock-free inter-processor shared SRAM exchanges.
\(P_{\text{collision}}\) Seqlock Retry Probability Ratio Probability \(t_{\text{vulnerable}} / T_{\text{writer}}\) that reader thread observes torn telemetry and must retry.
\(t_{\text{queue}}\) Memory Bus Queue Delay \(\text{ms}\) Contention latency \(\sum B_k / \text{BW}_{\text{peak}} + N_{\text{conflict}} t_{\text{penalty}}\) across shared SoC crossbars.
\(N_{\text{conflict}}\) Interconnect Port Conflicts Count Number of simultaneous DMA / GPU / CPU memory channel access collisions.
\(\Delta t_{\text{fresh}}\) Telemetry Freshness Deadline \(\text{ms}\) Maximum allowable age of sensor telemetry before an interlock trips: \(\Delta t_{\text{fresh}} \le e_{\max} / v_{\max}\).
\(t_{\text{watchdog}}\) Hardware Watchdog Timeout \(\text{ms}\) Hardware timer period triggering fail-safe dynamic braking upon lost processor heartbeats.
\(t_{\text{NMI}}\) Non-Maskable Interrupt Time \(\mu\text{s}\) Deterministic hardware interrupt response time for motor overcurrent or gate-drive faults.
\(\delta_{\text{skew}}\) Clock Synchronization Skew \(\text{ppm}, \mu\text{s/s}\) Hardware crystal drift between the application processor and safety microcontroller clocks.
\(\Delta t_{\text{peak}}\) Peak Inter-Node Phase Jitter \(\mu\text{s}\) Cumulative timing uncertainty \(\delta_{\text{skew}} \Delta T_{\text{blackout}} + \sigma_{\text{jitter}}\) during bus sync blackouts.
\(\text{MTBUF}\) Mean Time Between Unsafe Failure Hours Operating hours between undetected hazardous safety violations: \(\text{MTBUF} = \frac{1}{\lambda_{\text{raw}}(1 - c_{\text{shield}})}\).
\(\mathcal{T}_{\text{trichotomy}}\) Trichotomy Monitor State Enum Three-valued real-time monitor evaluation: \(\{\text{Known-True}, \text{Known-False}, \text{Unknown}\}\).
Back to top