Notation
Machine Learning Systems spans machine learning (computer science and statistics) and systems (computer architecture and hardware). Each field developed its notation independently, and many symbols mean different things across communities and publications. This collision creates real confusion when the disciplines merge. The conventions below establish a single notation that eliminates this ambiguity.
Consider a simple statement: “Increasing \(B\) improves throughput.” To an ML researcher, \(B\) means batch size. To a hardware engineer, \(B\) means bandwidth. Both interpretations are correct in their respective fields, but in ML Systems we need both concepts in the same equation—hence the need for a single, consistent convention.
The Iron Law of ML Systems
For serialized execution phases, the iron-law performance decomposition is:
\[T = \underbrace{\frac{D_{\text{vol}}}{\text{BW}}}_{\text{Memory Time}} + \underbrace{\frac{O}{R_{\text{peak}} \cdot \eta_{\text{hw}}}}_{\text{Compute Time}} + \underbrace{L_{\text{lat}}}_{\text{Latency Overhead}}\]
Each variable was chosen deliberately to avoid collision with standard ML terminology.
| Symbol | Definition | Unit | Why This Symbol? |
|---|---|---|---|
| \(T\) | Time | seconds | Unambiguous. Wall-clock time for an operation. |
| \(D_{\text{vol}}\) | Data Volume | bytes | Avoids collision with \(D\) (Dataset Size). In scaling laws, \(D\) means training tokens. Here we need bytes moved through memory. The subscript disambiguates. |
| \(\text{BW}\) | Bandwidth | bytes/s | Avoids collision with \(B\) (Batch Size). Systems literature often uses \(B\) for bandwidth, while ML literature often uses \(B\) for batch size. The notation preserves the ML convention. |
| \(O\) | Operations | FLOPs | Total floating-point operations. Clean in equations (vs. “\(\text{Ops}\)”). |
| \(R_{\text{peak}}\) | Peak Rate | FLOP/s | Avoids collision with \(P\) (Parameters). Roofline presentations often use \(P\) for performance, while ML literature commonly uses \(P\) for parameter count. The notation preserves the ML convention. |
| \(\eta_{\text{hw}}\) | Efficiency | — | Hardware utilization \((0 \le \eta_{\text{hw}} \le 1)\). Avoids collision with learning rate \((\eta)\). |
| \(L_{\text{lat}}\) | Latency | seconds | Avoids collision with \(\mathcal{L}\) (Loss). Fixed overhead time (kernel launch, network RTT). The subscript distinguishes from the loss function. |
Why these choices matter
Without careful notation, sentences become ambiguous:
“Reducing \(D\) improves performance.”
The sentence has two possible readings:
- Reducing dataset size (fewer training samples) can speed training but may reduce accuracy.
- Reducing data volume moved (for example, through compression or quantization) can speed inference, with accuracy effects that depend on the technique and calibration.
With our notation, we can write precisely:
“Reducing \(D_{\text{vol}}\) through FP32-to-INT8 quantization cuts parameter memory traffic to one quarter while \(D\) (training data) remains unchanged.”
Our notation eliminates this ambiguity: “\(\text{BW}\) limits throughput” has only one reading.
Subscripted variants
An upright root names the quantity, and the subscript names which instance of it is meant. The convention holds closely related measurements apart instead of collapsing them into a single symbol:
- \(\text{BW}_{\text{disk}}\), \(\text{BW}_{\text{network}}\), \(\text{BW}_{\text{accelerator}}\): bandwidth at three points on the data path
- \(D_{\text{vol}}\), \(R_{\text{peak}}\), \(L_{\text{lat}}\), \(\eta_{\text{hw}}\): iron-law terms, each held clear of a common ML symbol
- \(E_{\text{move}}\), \(E_{\text{compute}}\), \(E_{\text{total}}\): one energy budget split into tradeable parts
The subscript therefore carries part of the claim. Disk and accelerator bandwidth differ by orders of magnitude, so a figure is reproducible only when the subscript says where it was measured.
The Degradation Equation
The degradation equation is a fitted local model. Divergence alone does not determine the direction of accuracy change, so estimate \(\lambda\) from labeled outcomes. Some symbols below, such as \(\tau\), appear in chapter prose rather than in the equation. \[\text{Accuracy}(t) \approx \text{Accuracy}_0 - \lambda \cdot \mathcal{D}(P_t \lVert P_0)\]
| Symbol | Definition | Unit/Type | Notes |
|---|---|---|---|
| \(\text{Accuracy}(t)\) | Accuracy at Time \(t\) | Scalar | Model accuracy after the model has been deployed for time \(t\). |
| \(\text{Accuracy}_0\) | Initial Accuracy | Scalar | Model accuracy at deployment time. |
| \(\lambda\) | Fitted Sensitivity | Scalar | Local coefficient estimated from outcomes over a stated range; not a universal constant. (Not wavelength.) |
| \(P_t\) | Current Distribution | Distribution | The data distribution at time \(t\). (Not parameters—use \(P\) for parameter count.) |
| \(P_0\) | Training Distribution | Distribution | The data distribution at training time. |
| \(\mathcal{D}(P_t \lVert P_0)\) | Statistical Divergence | Scalar \(\ge 0\) | Measures how far \(P_t\) has drifted from \(P_0\). Common choices: KL divergence, total variation, Wasserstein. (Calligraphic to avoid collision with \(D\) = dataset size.) |
| \(\tau\) | Response Threshold | Scalar \(> 0\) | Operational threshold for investigation; it does not automatically trigger retraining. |
The Energy Corollary
With effective per-byte and per-operation costs, workload energy is approximated as: \[E_{\text{total}} \approx D_{\text{vol}} \times E_{\text{move}} + O \times E_{\text{compute}}\]
| Symbol | Definition | Unit | Notes |
|---|---|---|---|
| \(E_{\text{total}}\) | Total Energy | joules | Total energy consumed by an ML workload, decomposed into data-movement and compute terms. |
| \(E_{\text{move}}\) | Energy per Byte Moved | joules/byte | Effective movement cost; total energy also depends on bytes moved and operations executed. |
| \(E_{\text{compute}}\) | Energy per Operation | joules/FLOP | Energy cost of a single arithmetic operation. |
Deep Learning Notation
This book follows standard deep learning conventions (Goodfellow et al. 2016) with explicit disambiguation for systems variables.
| Symbol | Definition | Dimensions/Type |
|---|---|---|
| \(B\) | Batch Size | Integer. The number of samples processed in parallel. (Never bandwidth.) |
| \(P\) | Parameters | Integer. The total count of scalar model parameters, including weights and biases. (Never peak FLOP/s.) |
| \(D\) | Dataset Size | Integer. Number of training samples or tokens. (Never data volume in bytes—use \(D_{\text{vol}}\).) |
| \(S\) | Sequence Length | Integer. Number of tokens or time steps. |
| \(d\) | Hidden Dimension | Integer. Size of the hidden state vector. |
| \(d_{\text{head}}\) | Attention Head Dimension | Integer. Per-head hidden dimension in attention layers. |
| \(N_L\) | Number of Layers | Integer. Total number of layers in a network. |
| \(N_{\text{heads}}\) | Number of Attention Heads | Integer. Number of attention heads in a multi-head attention layer. |
| \(H_{\text{KV}}\) | Number of Key-Value Heads | Integer. Number of key-value heads in grouped-query or multi-query attention. |
| \(\ell\) | Layer Index | Integer. Index for a layer. Use instead of bare \(L\) when indexing layers, since \(L\) collides with loss and latency. |
| \(\mathcal{L}\) | Loss Function | Scalar. The objective function minimized during training. |
| \(\eta\) | Learning Rate | Scalar. Step size for the optimizer. (Never bare for hardware efficiency—use \(\eta_{\text{hw}}\).) |
| \(\theta\) | Model Parameters | Parameter collection, often treated as a vector. The set of all learnable parameters. |
| \(p(x)\) | Distribution | Probability mass/density for a random variable. Use lowercase \(p(\cdot)\) for generic distributions to avoid collision with \(P\) (parameter count). |
| \(p(y \mid x)\) | Conditional Distribution | Conditional probability/density. Use this form for generic label relationships; reserve \(P_0\) and \(P_t\) for the degradation equation’s training/current distributions. |
| \(\Pr(E)\) | Event Probability | Probability of an event \(E\). Use for event statements such as \(\Pr(\text{batch}=0)\); use \(p(x)\) and \(p(y \mid x)\) for distributions. |
Latin matrices and vectors are set in bold (\(\mathbf{W}\), \(\mathbf{x}\)); scalars, dimensions, indices, and individual matrix entries stay italic (\(N\), \(d\), \(i\), \(W_{ij}\)). The generic parameter symbol \(\theta\) is the one conventional exception and stays italic. Local matrix algebra follows the same bold rule: matrix operands are bold (\(\mathbf{A}\), \(\mathbf{B}\), \(\mathbf{C}\)), as in general matrix multiply (GEMM) \(\mathbf{C}=\alpha \mathbf{A}\mathbf{B}+\beta \mathbf{C}\). The bold form marks a matrix operand and is distinct from the italic scalar batch size \(B\), which remains batch size in scalar contexts.
Performance, Serving, and Memory Notation
Reusable performance and serving quantities follow the same collision-avoidance rule as the iron law. Rates use \(R\) or descriptive Greek symbols; request counts use \(Q_{\text{req}}\) rather than overloading \(N\); queue utilization uses a subscripted \(\rho\) so bare \(\rho\) remains available for other ratio models.
| Symbol | Definition | Unit/Type | Notes |
|---|---|---|---|
| \(I\) | Arithmetic Intensity | FLOP/byte | Workload FLOPs per byte moved. The Roofline Model uses \(I\) as the independent variable. |
| \(I_{\text{ridge}}\) | Roofline Ridge Point | FLOP/byte | \(I_{\text{ridge}} = R_{\text{peak}}/\text{BW}\). Prefer this explicit form over starred shorthand so the meaning remains clear in prose. |
| \(R_{\text{attain}}\) | Attainable Compute Rate | FLOP/s | Roofline bound: \(R_{\text{attain}} \leq \min(R_{\text{peak}}, I \times \text{BW})\). Uses \(R\), not \(T\), because this quantity is a rate. |
| \(\text{MFU}\) | Model FLOPs Utilization | Dimensionless | Useful model FLOP/s divided by available peak FLOP/s. Text acronym avoids overloading \(\eta\). |
| \(r_{\text{comp}}\) | Compression Ratio | Dimensionless | Uncompressed size divided by compressed size. A compressed payload has \(1/r_{\text{comp}}\) of the uncompressed size; the subscript avoids bare \(C\) collisions. |
| \(Q_{\text{req}}\) | Request Concurrency | requests | Average in-flight requests in a stable serving system (\(Q_{\text{req}} = \lambda_{\text{arr}} \cdot T_{\text{lat}}\) via Little’s Law). Avoids collision with \(N\) as device count in distributed settings. |
| \(\lambda_{\text{arr}}\) | Arrival Rate | requests/s | Request arrival rate. The subscript avoids collision with \(\lambda\) as degradation sensitivity or failure rate. |
| \(T_{\text{lat}}\) | Request Time in System | seconds | End-to-end queueing/serving latency for Little’s Law. Distinct from \(L_{\text{lat}}\), the fixed-latency term in the iron law. |
| \(T_{\text{svc}}(B)\) | Batch Service Time | seconds | Time to serve a batch of size \(B\). |
| \(\mu_{\text{eff}}(B)\) | Effective Service Rate | requests/s | Batched service rate, typically \(\mu_{\text{eff}}(B) = B/T_{\text{svc}}(B)\). |
| \(\rho_{\text{serv}}\) | Serving Utilization | Dimensionless | Queue/server utilization. Use instead of bare \(\rho\), which is reserved for communication-computation ratio in distributed contexts. |
| \(M_{\text{total}}\) | Total Memory Footprint | bytes | Sum of explicit memory components; avoids bare \(M\) ambiguity. |
| \(M_{\text{weights}}\) | Weight Memory | bytes | Memory occupied by model parameters. |
| \(M_{\text{gradients}}\) | Gradient Memory | bytes | Memory occupied by stored gradients. |
| \(M_{\text{optimizer}}\) | Optimizer-State Memory | bytes | Momentum, variance, master weights, and related optimizer buffers. |
| \(M_{\text{activations}}\) | Activation Memory | bytes | Retained activations for the backward pass or serving intermediates. |
| \(s_{\text{elem}}\) | Element Storage Size | bytes/element | Bytes per stored tensor element. |
Units and Precision
- Physical units: This book uses SI (metric) units throughout, including meters, kilograms, seconds, watts, and °C, consistent with standard engineering and scientific practice. Where source data was originally reported in imperial units, the book converts to SI and notes the original values parenthetically. A space always separates the number from the unit (for example, 100 ms, 2 TB/s).
- Data and memory: This book uses decimal SI prefixes only: KB = \(10^3\) bytes, MB = \(10^6\), GB = \(10^9\), TB = \(10^{12}\). Binary-prefixed units do not appear in prose; all capacities, throughputs, and model sizes are reported in decimal units (for example, 80 GB, 2 TB/s, 102 MB).
- Compute: The notation distinguishes operation counts from rates. Total work uses FLOPs (for example, GFLOPs, TFLOPs), while throughput uses FLOP/s with decimal prefixes (for example, GFLOP/s, TFLOP/s).
- 1 TFLOP = \(10^{12}\) FLOPs
- 1 TFLOP/s = \(10^{12}\) FLOPs per second
- Arithmetic intensity conventionally uses FLOP/byte as a unit ratio (floating-point operations per byte moved). This is a ratio unit, not a total-work symbol or throughput symbol.
- Currency: Dollar amounts use the dollar sign (
$); unless otherwise noted, dollar-denominated costs are U.S. dollars (USD). - Precision:
- FP64: Double precision (8 bytes)
- FP32: Single precision (4 bytes)
- TF32: TensorFloat-32 (19-bit Tensor Core compute mode; inputs and storage remain FP32)
- FP16: Half precision (2 bytes, standard range)
- BF16: Brain float (2 bytes, wide dynamic range)
- FP8: Quarter precision (1 byte, E4M3 or E5M2 format)
- FP4: 4-bit floating-point format
- INT8: 8-bit integer (1 byte)
- INT4: 4-bit integer; lower integer precisions follow the same uppercase
INTnpattern (for example, INT3, INT2)
Quick Reference: Resolving Collisions
Common collision points in ML Systems literature include:
| Symbol | ML Meaning | Systems Meaning | Book Convention |
|---|---|---|---|
| \(B\) | Batch Size | Bandwidth | Batch Size. Use \(\text{BW}\) for bandwidth. |
| \(P\) | Parameters | Peak FLOP/s | Parameters. Use \(R_{\text{peak}}\) for peak rate. |
| \(D\) | Dataset Size | Data Volume | Dataset Size. Use \(D_{\text{vol}}\) for bytes moved. |
| \(L\) | Loss | Latency | Loss \((\mathcal{L})\). Use \(L_{\text{lat}}\) for latency. |
| \(\eta\) | Learning Rate | Efficiency | Learning Rate. Use \(\eta_{\text{hw}}\) for efficiency. |
As a general principle, ML conventions take precedence for single letters, while systems concepts get subscripts or multi-letter symbols. This reflects the primary audience (ML practitioners learning systems) and preserves compatibility with the vast ML literature.
Agentic Systems Notation
This book adds notation for trajectories, the task contract, the H·S·A exposures, model calls and their budgets, the time and cost of a trajectory, recovery, and learning from trajectories. The shared tables above still govern every symbol they define, and this extension only adds to them. A symbol that appears only inside one derivation is defined where it is used and is not listed here.
Five relations recur throughout the book. A trajectory, the ordered record of one delegated task, is written
\[\tau = (s_0, a_0, o_1, s_1, a_1, o_2, \dots, s_T)\]
Its wall-clock duration over a horizon of \(H\) turns is
\[T_{\text{task}} = \sum_{k=1}^{H} \left( T_{\text{model}}^{(k)} + T_{\text{tool}}^{(k)} + T_{\text{wait}}^{(k)} + T_{\text{runtime}}^{(k)} \right)\]
Trajectory goodput is the share of a resource \(R\) (turns, tokens, sandbox time, or dollars) spent on trajectories whose results passed verification,
\[\mathcal{G} = \frac{\sum_{i \in \mathcal{T}_{\text{success}}} R_i}{\sum_{j \in \mathcal{T}_{\text{all}}} R_j}\]
A trajectory that runs \(N\) sequential proposals without checking any of them, each correct with probability \(1 - \epsilon\) and independent of the others, compounds its errors and succeeds with probability
\[\Pr(\text{success}) = (1 - \epsilon)^N\]
The cost per accepted task divides everything spent on \(N\) attempts by the \(N_{\text{acc}}\) attempts that were accepted,
\[C_{\text{eff}} = \frac{\sum_{i=1}^{N} C_{\text{attempt}, i}}{N_{\text{acc}}}\]
Trajectories and the agent loop
The loop symbols describe one turn of a trajectory: the context the runtime assembles, the proposal the model returns, the decision the runtime makes about it, and what comes back.
| Symbol | Definition | Unit | Notes |
|---|---|---|---|
| \(\tau\) | Trajectory | Tuple | Ordered record of one delegated task, \(\tau = (s_0, a_0, o_1, \dots, s_T)\). Bare \(\tau\) always means a trajectory in this book. |
| \(t\) | Turn index | Integer | Counts turns within a trajectory. \(s_T\) is the state after the final turn. |
| \(g\) | Goal | Text | The delegated task and its completion rule, fixed for the trajectory. The task contract writes it \(G\). |
| \(c_t\) | Context | Tokens | What the runtime assembles and sends on turn \(t\): instructions, goal, tool definitions, and the history of actions, observations, and evidence. |
| \(s_t\) | State at turn \(t\) | Record | The context state the model saw on turn \(t\). In Part V, the state a policy conditions on. |
| \(a_t\) | Action proposal | Text or tool call | What the model proposed on turn \(t\). Nothing happens until the runtime admits it. |
| \(a_{\text{prop}}\), \(a_{\text{perm}}\) | Proposed, permitted action | Tool call | A proposal before the runtime checks it, and the action it permits after the check. |
| \(o_t\) | Observation | Typed result | What an action returned, such as an exit code, captured output, a response body, or a status. |
| \(v_t\) | Verification evidence | Pass/fail or score | Result of a check the runtime runs on the resulting state, independent of the action and of the model’s report. |
| \(\pi_\theta\) | Policy | Distribution | The model viewed as a distribution over proposals given a context, with parameters \(\theta\). |
| \(\mathcal{S}\) | Agentic system | Tuple | \(\mathcal{S} = \langle \pi_\theta, \mathcal{H}, \mathcal{E}, \mathcal{M}, \mathcal{V} \rangle\): policy, runtime harness, environment, memory, and verifiers. |
| \(\mathcal{I}\) | Closure operator | Function | \(\mathcal{I}: \mathcal{A}_{\text{prop}} \times \mathcal{S}_{\text{sys}} \to \mathcal{A}_{\text{perm}} \cup \{\bot\}\). The runtime’s check of each proposal against the current system state \(\mathcal{S}_{\text{sys}}\). |
| \(\bot\) | Rejection | Symbol | A refused proposal. It reaches the model as a typed error and leaves the environment unchanged. |
The task contract
A task contract fixes what a trajectory is for, where it runs, what it may do and see, and what ends it.
| Symbol | Definition | Unit | Notes |
|---|---|---|---|
| \(\mathcal{C}\) | Task contract | Tuple | \(\mathcal{C} = \langle G, \mathcal{E}_{\text{env}}, \mathcal{A}_{\text{perm}}, \mathcal{O}_{\text{avail}}, \mathcal{K}_{\text{comp}} \rangle\). |
| \(G\) | Goal | Text | Desired end state, with explicit negative scope. |
| \(\mathcal{E}_{\text{env}}\) | Environment | Specification | Reproducible place the work happens: base commit, image, configuration, network rules. |
| \(\mathcal{A}_{\text{perm}}\) | Permitted actions | Set | Tools the task may call and the limits on each. Least privilege stated per task. |
| \(\mathcal{O}_{\text{avail}}\) | Available observations | Specification | What the agent may see, and how much, including truncation and redaction rules. |
| \(\mathcal{K}_{\text{comp}}\) | Completion criteria | Set of checks | Checks that decide when the task is done, at a stated closure evidence level. The model’s claim of completion is never one of them. |
The H·S·A exposures
H·S·A names the three exposures an agentic task adds beyond a single model call. A task’s position says how exposed it is. The closure the runtime must supply against that exposure is governed by the invariant closure principle and is not a fourth axis. Zero ambient authority describes the model at every authority level and has no symbol.
| Symbol | Definition | Unit | Notes |
|---|---|---|---|
| \(H\) | Horizon | Turns | Number of turns a trajectory runs: \(H = 1\) (single call), \(H \approx 10\) (short loop), \(H \geq 100\) (minutes to hours). Wall-clock duration given \(H\) is \(T_{\text{task}}\). |
| \(S_0\) to \(S_3\) | State | Level | What the task carries between turns: \(S_0\) (nothing), \(S_1\) (context window and KV cache), \(S_2\) (sandboxed working files), \(S_3\) (durable state that outlives the session or is shared beyond it). Written with its level, because bare \(S\) is sequence length. |
| \(A_0\) to \(A_3\) | Authority | Level | What the task may do through its tools: \(A_0\) (read access), \(A_1\) (mutation confined to a discardable sandbox), \(A_2\) (external action that is retry-safe or compensable), \(A_3\) (irreversible external action). Covers confidentiality as well as integrity, so \(A_0\) is not a zero blast radius. |
Closure evidence levels, from weakest to strongest, are named in words and carry no symbol: self-report, static checks, visible tests, sealed tests, and formal proof.
Model calls and budgets
A model call is bounded by its window and by limits the runtime sets on each request. A trajectory is bounded by a budget the harness keeps outside the context.
| Symbol | Definition | Unit | Notes |
|---|---|---|---|
| \(S\) | Sequence length | Tokens | The shared symbol for sequence length, used here for the prompt a call reads. The State exposure is always written with a level, \(S_0\) to \(S_3\). |
| \(K\) | Output length | Tokens | Tokens a call generates. Lowercase \(k\) is reserved for sample counts, as in pass@\(k\). |
| \(S_{\max}\) | Context window | Tokens | Hard per-call limit on prompt and output together: \(S + K_{\max} \le S_{\max}\). |
| \(K_{\max}\) | Output limit | Tokens | Maximum tokens the runtime lets a call generate. |
| \(B_{\text{reason}}\) | Reasoning budget | Tokens | Limit on reasoning tokens a call may spend before it answers. |
| \(t_{\text{tok}}\) | Per-token decode time | Seconds | Time to generate one output token on a single stream. |
| \(T_{\text{call}}\) | Call latency | Seconds | \(T_{\text{call}} \approx T_{\text{queue}} + T_{\text{prefill}}(S) + K \cdot t_{\text{tok}}\). The first two terms are the time to first token. |
| \(\mathbf{B}\) | Trajectory budget | Vector | \(\mathbf{B} = \langle H_{\max}, N_{\max}, T_{\max}, C_{\max} \rangle\), ceilings on turns, tokens, wall-clock time, and dollars, checked by the harness before every call. Bold, to stay distinct from batch size \(B\). |
| \(T_{\max}\) | Deadline | Seconds | Wall-clock ceiling enforced by the runtime, on one call or on a whole trajectory. The text names the scope. |
| \(C_{\max}\) | Cost ceiling | Dollars | Spend ceiling the runtime enforces before dispatching the next call. |
| \(m_{\text{token}}\) | KV bytes per token | Bytes/token | \(m_{\text{token}} = 2 \cdot N_L \cdot H_{\text{KV}} \cdot d_{\text{head}} \cdot s_{\text{elem}}\), in the shared symbols for layers, KV heads, head dimension, and element size. |
| \(M_{\text{KV}}\) | KV footprint | Bytes | Attention state a trajectory holds in accelerator memory, \(M_{\text{KV}} = S \cdot m_{\text{token}}\). |
| \(O_{\text{mem}}\) | Memory held during a wait | Byte-seconds | \(O_{\text{mem}} = M_{\text{KV}} \cdot T_{\text{wait}}\), the accelerator memory a paused trajectory withholds while it waits on a tool. |
Time and cost of a trajectory
Agentic systems are measured per accepted task, not per request. The time terms decompose one turn, and the cost terms price attempts and the accepted results they buy.
| Symbol | Definition | Unit | Notes |
|---|---|---|---|
| \(T_{\text{task}}\) | Trajectory duration | Seconds | Sum over turns of model, tool, wait, and runtime time. |
| \(T_{\text{model}}\) | Model time | Seconds | Time spent in model calls during a turn. |
| \(T_{\text{tool}}\) | Tool time | Seconds | Time a tool takes to execute, such as a build or a test run. |
| \(T_{\text{wait}}\) | External wait | Seconds | Time spent waiting on a remote service or a person’s approval. |
| \(T_{\text{runtime}}\) | Runtime time | Seconds | The runtime’s own work: assembling context, validating proposals, running checks. |
| \(\mathcal{T}_{\text{all}}\), \(\mathcal{T}_{\text{success}}\) | Trajectory sets | Set | Every trajectory dispatched, and the subset whose results passed verification. Calligraphic \(\mathcal{T}\) with a subscript always names a set of trajectories. |
| \(\mathcal{G}\) | Trajectory goodput | Fraction \([0, 1]\) | Share of a resource spent on trajectories that passed verification. Badput is \(1 - \mathcal{G}\). |
| \(C_{\text{task}}\) | Trajectory cost | Dollars or tokens | What one trajectory spent, checked against \(C_{\max}\). |
| \(C_{\text{attempt}}\) | Attempt cost | Dollars | Full cost of one attempt: input and output tokens, tool fees, sandbox time, verification, and human review. |
| \(N_{\text{acc}}\) | Accepted attempts | Count | Attempts whose results passed the acceptance checks. |
| \(C_{\text{eff}}\) | Cost per accepted task | Dollars/task | Total spend on all attempts divided by \(N_{\text{acc}}\). Failed attempts are paid for in full. |
Recovery, evaluation, and learning
These symbols recur across the runtime, evaluation, and learning chapters.
| Symbol | Definition | Unit | Notes |
|---|---|---|---|
| \(k_{\text{idem}}\) | Idempotency key | String | Deterministic key derived from the trajectory, turn, call, tool, and canonical arguments, so a retried call cannot apply its effect twice. |
| \(K_{\text{repair}}\) | Repair cap | Count | Maximum forward repair attempts on one failed step before the harness compensates or escalates. |
| pass@\(k\) | Pass at \(k\) | Probability | Probability that at least one of \(k\) runs of a task succeeds. Measures potential, and is reachable only when a verifier selects among the runs. |
| pass\(^k\) | Pass to the \(k\) | Probability | Probability that all \(k\) runs of a task succeed. Measures reliability over repeated runs. |
| \(m_t\) | Loss mask | Binary \(\{0, 1\}\) | \(m_t = 1\) on tokens the policy wrote (reasoning, tool calls, answers) and \(m_t = 0\) on prompts and observations. |
| \(\mathcal{L}_{\text{SFT}}\) | Masked fine-tuning loss | Cross-entropy | \(\mathcal{L}_{\text{SFT}}(\theta) = -\sum_t m_t \log \pi_\theta(x_t \mid x_{<t})\). |
| \(R(\tau)\) | Trajectory reward | Scalar | Reward a verifier assigns to a whole trajectory. |
| \(\hat{V}(s_t)\) | Estimated value | Scalar | Estimated probability of success from state \(s_t\). |
| \(\hat{A}(s_t, a_t)\) | Estimated advantage | Scalar | How much an action raised the estimated value, \(\hat{V}(s_{t+1}) - \hat{V}(s_t)\). Distinct from the authority levels \(A_0\) to \(A_3\). |
Resolving collisions
Agentic systems borrow notation from machine learning, systems, and control, so several letters arrive with more than one meaning. The conventions below keep one meaning per symbol.
| Symbol | Other common meanings | Convention in this book |
|---|---|---|
| \(\tau\) | Temperature; a threshold | Bare \(\tau\) is a trajectory. Sampling temperature is named in words or written \(\tau_{\text{temp}}\). |
| \(H\) | Attention heads; hidden size | \(H\) is the horizon in turns. Heads use \(N_{\text{heads}}\) and \(H_{\text{KV}}\); hidden size uses \(d\). |
| \(S\) | The State exposure | Bare \(S\) is sequence length. State is written only with a level, \(S_0\) to \(S_3\). |
| \(A\) | Advantage; action sets | Authority is written only with a level, \(A_0\) to \(A_3\). Advantage is \(\hat{A}\); action sets are \(\mathcal{A}_{\text{prop}}\) and \(\mathcal{A}_{\text{perm}}\). |
| \(P\) | Prompt length; precision | \(P\) is parameter count, as in the shared table. Prompt length is \(S\). |
| \(B\) | A budget | \(B\) is batch size. The trajectory budget is bold \(\mathbf{B}\), and the reasoning budget is \(B_{\text{reason}}\). |
| \(C\) | A fourth taxonomy axis | There is no \(C\) axis; closure evidence levels are named in words. Costs carry subscripts (\(C_{\text{task}}\), \(C_{\text{eff}}\), \(C_{\max}\)), and the task contract is \(\mathcal{C}\). |
| \(\mathcal{G}\) | The goal | \(\mathcal{G}\) is trajectory goodput. The goal is \(g\) in the loop and \(G\) in the task contract. |
| \(\mathcal{T}\) | A sequence of steps | Calligraphic \(\mathcal{T}\) with a subscript is a set of trajectories. Time is \(T\). |
| \(K\) | Top-\(k\); sample count | \(K\) is output length in tokens. Sample counts use lowercase \(k\), as in pass@\(k\). |
Additional units
- Tokens count what a call reads and writes, and serving throughput is reported in tokens per second.
- Turns count the steps of a trajectory, and the horizon \(H\) is measured in turns.
- Dollars per accepted task normalize spend by the results that passed acceptance, not by attempts or requests.
- Trajectory goodput (\(\mathcal{G}\)) is reported as a fraction in \([0, 1]\) or as a percentage, together with the resource it counts.
- Byte-seconds measure memory held over time, such as attention state kept resident during a tool wait.