Notation

Machine Learning Systems spans machine learning (computer science and statistics) and systems (computer architecture and hardware). Each field developed its notation independently, and many symbols mean different things across communities and publications. This collision creates real confusion when the disciplines merge. The conventions below establish a single notation that eliminates this ambiguity.

Consider a simple statement: “Increasing \(B\) improves throughput.” To an ML researcher, \(B\) means batch size. To a hardware engineer, \(B\) means bandwidth. Both interpretations are correct in their respective fields, but in ML Systems we need both concepts in the same equation—hence the need for a single, consistent convention.

The Iron Law of ML Systems

For serialized execution phases, the iron-law performance decomposition is:

\[T = \underbrace{\frac{D_{\text{vol}}}{\text{BW}}}_{\text{Memory Time}} + \underbrace{\frac{O}{R_{\text{peak}} \cdot \eta_{\text{hw}}}}_{\text{Compute Time}} + \underbrace{L_{\text{lat}}}_{\text{Latency Overhead}}\]

Each variable was chosen deliberately to avoid collision with standard ML terminology.

Symbol Definition Unit Why This Symbol?
\(T\) Time seconds Unambiguous. Wall-clock time for an operation.
\(D_{\text{vol}}\) Data Volume bytes Avoids collision with \(D\) (Dataset Size). In scaling laws, \(D\) means training tokens. Here we need bytes moved through memory. The subscript disambiguates.
\(\text{BW}\) Bandwidth bytes/s Avoids collision with \(B\) (Batch Size). Systems literature often uses \(B\) for bandwidth, while ML literature often uses \(B\) for batch size. The notation preserves the ML convention.
\(O\) Operations FLOPs Total floating-point operations. Clean in equations (vs. “\(\text{Ops}\)”).
\(R_{\text{peak}}\) Peak Rate FLOP/s Avoids collision with \(P\) (Parameters). Roofline presentations often use \(P\) for performance, while ML literature commonly uses \(P\) for parameter count. The notation preserves the ML convention.
\(\eta_{\text{hw}}\) Efficiency — Hardware utilization \((0 \le \eta_{\text{hw}} \le 1)\). Avoids collision with learning rate \((\eta)\).
\(L_{\text{lat}}\) Latency seconds Avoids collision with \(\mathcal{L}\) (Loss). Fixed overhead time (kernel launch, network RTT). The subscript distinguishes from the loss function.

Why these choices matter

Without careful notation, sentences become ambiguous:

“Reducing \(D\) improves performance.”

The sentence has two possible readings:

  • Reducing dataset size (fewer training samples) can speed training but may reduce accuracy.
  • Reducing data volume moved (for example, through compression or quantization) can speed inference, with accuracy effects that depend on the technique and calibration.

With our notation, we can write precisely:

“Reducing \(D_{\text{vol}}\) through FP32-to-INT8 quantization cuts parameter memory traffic to one quarter while \(D\) (training data) remains unchanged.”

Our notation eliminates this ambiguity: “\(\text{BW}\) limits throughput” has only one reading.

Subscripted variants

An upright root names the quantity, and the subscript names which instance of it is meant. The convention holds closely related measurements apart instead of collapsing them into a single symbol:

  • \(\text{BW}_{\text{disk}}\), \(\text{BW}_{\text{network}}\), \(\text{BW}_{\text{accelerator}}\): bandwidth at three points on the data path
  • \(D_{\text{vol}}\), \(R_{\text{peak}}\), \(L_{\text{lat}}\), \(\eta_{\text{hw}}\): iron-law terms, each held clear of a common ML symbol
  • \(E_{\text{move}}\), \(E_{\text{compute}}\), \(E_{\text{total}}\): one energy budget split into tradeable parts

The subscript therefore carries part of the claim. Disk and accelerator bandwidth differ by orders of magnitude, so a figure is reproducible only when the subscript says where it was measured.

The Degradation Equation

The degradation equation is a fitted local model. Divergence alone does not determine the direction of accuracy change, so estimate \(\lambda\) from labeled outcomes. Some symbols below, such as \(\tau\), appear in chapter prose rather than in the equation. \[\text{Accuracy}(t) \approx \text{Accuracy}_0 - \lambda \cdot \mathcal{D}(P_t \lVert P_0)\]

Symbol Definition Unit/Type Notes
\(\text{Accuracy}(t)\) Accuracy at Time \(t\) Scalar Model accuracy after the model has been deployed for time \(t\).
\(\text{Accuracy}_0\) Initial Accuracy Scalar Model accuracy at deployment time.
\(\lambda\) Fitted Sensitivity Scalar Local coefficient estimated from outcomes over a stated range; not a universal constant. (Not wavelength.)
\(P_t\) Current Distribution Distribution The data distribution at time \(t\). (Not parameters—use \(P\) for parameter count.)
\(P_0\) Training Distribution Distribution The data distribution at training time.
\(\mathcal{D}(P_t \lVert P_0)\) Statistical Divergence Scalar \(\ge 0\) Measures how far \(P_t\) has drifted from \(P_0\). Common choices: KL divergence, total variation, Wasserstein. (Calligraphic to avoid collision with \(D\) = dataset size.)
\(\tau\) Response Threshold Scalar \(> 0\) Operational threshold for investigation; it does not automatically trigger retraining.

The Energy Corollary

With effective per-byte and per-operation costs, workload energy is approximated as: \[E_{\text{total}} \approx D_{\text{vol}} \times E_{\text{move}} + O \times E_{\text{compute}}\]

Symbol Definition Unit Notes
\(E_{\text{total}}\) Total Energy joules Total energy consumed by an ML workload, decomposed into data-movement and compute terms.
\(E_{\text{move}}\) Energy per Byte Moved joules/byte Effective movement cost; total energy also depends on bytes moved and operations executed.
\(E_{\text{compute}}\) Energy per Operation joules/FLOP Energy cost of a single arithmetic operation.

Deep Learning Notation

This book follows standard deep learning conventions (Goodfellow et al. 2016) with explicit disambiguation for systems variables.

Goodfellow, Ian, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press.
Symbol Definition Dimensions/Type
\(B\) Batch Size Integer. The number of samples processed in parallel. (Never bandwidth.)
\(P\) Parameters Integer. The total count of scalar model parameters, including weights and biases. (Never peak FLOP/s.)
\(D\) Dataset Size Integer. Number of training samples or tokens. (Never data volume in bytes—use \(D_{\text{vol}}\).)
\(S\) Sequence Length Integer. Number of tokens or time steps.
\(d\) Hidden Dimension Integer. Size of the hidden state vector.
\(d_{\text{head}}\) Attention Head Dimension Integer. Per-head hidden dimension in attention layers.
\(N_L\) Number of Layers Integer. Total number of layers in a network.
\(N_{\text{heads}}\) Number of Attention Heads Integer. Number of attention heads in a multi-head attention layer.
\(H_{\text{KV}}\) Number of Key-Value Heads Integer. Number of key-value heads in grouped-query or multi-query attention.
\(\ell\) Layer Index Integer. Index for a layer. Use instead of bare \(L\) when indexing layers, since \(L\) collides with loss and latency.
\(\mathcal{L}\) Loss Function Scalar. The objective function minimized during training.
\(\eta\) Learning Rate Scalar. Step size for the optimizer. (Never bare for hardware efficiency—use \(\eta_{\text{hw}}\).)
\(\theta\) Model Parameters Parameter collection, often treated as a vector. The set of all learnable parameters.
\(p(x)\) Distribution Probability mass/density for a random variable. Use lowercase \(p(\cdot)\) for generic distributions to avoid collision with \(P\) (parameter count).
\(p(y \mid x)\) Conditional Distribution Conditional probability/density. Use this form for generic label relationships; reserve \(P_0\) and \(P_t\) for the degradation equation’s training/current distributions.
\(\Pr(E)\) Event Probability Probability of an event \(E\). Use for event statements such as \(\Pr(\text{batch}=0)\); use \(p(x)\) and \(p(y \mid x)\) for distributions.

Latin matrices and vectors are set in bold (\(\mathbf{W}\), \(\mathbf{x}\)); scalars, dimensions, indices, and individual matrix entries stay italic (\(N\), \(d\), \(i\), \(W_{ij}\)). The generic parameter symbol \(\theta\) is the one conventional exception and stays italic. Local matrix algebra follows the same bold rule: matrix operands are bold (\(\mathbf{A}\), \(\mathbf{B}\), \(\mathbf{C}\)), as in general matrix multiply (GEMM) \(\mathbf{C}=\alpha \mathbf{A}\mathbf{B}+\beta \mathbf{C}\). The bold form marks a matrix operand and is distinct from the italic scalar batch size \(B\), which remains batch size in scalar contexts.

Performance, Serving, and Memory Notation

Reusable performance and serving quantities follow the same collision-avoidance rule as the iron law. Rates use \(R\) or descriptive Greek symbols; request counts use \(Q_{\text{req}}\) rather than overloading \(N\); queue utilization uses a subscripted \(\rho\) so bare \(\rho\) remains available for other ratio models.

Symbol Definition Unit/Type Notes
\(I\) Arithmetic Intensity FLOP/byte Workload FLOPs per byte moved. The Roofline Model uses \(I\) as the independent variable.
\(I_{\text{ridge}}\) Roofline Ridge Point FLOP/byte \(I_{\text{ridge}} = R_{\text{peak}}/\text{BW}\). Prefer this explicit form over starred shorthand so the meaning remains clear in prose.
\(R_{\text{attain}}\) Attainable Compute Rate FLOP/s Roofline bound: \(R_{\text{attain}} \leq \min(R_{\text{peak}}, I \times \text{BW})\). Uses \(R\), not \(T\), because this quantity is a rate.
\(\text{MFU}\) Model FLOPs Utilization Dimensionless Useful model FLOP/s divided by available peak FLOP/s. Text acronym avoids overloading \(\eta\).
\(r_{\text{comp}}\) Compression Ratio Dimensionless Uncompressed size divided by compressed size. A compressed payload has \(1/r_{\text{comp}}\) of the uncompressed size; the subscript avoids bare \(C\) collisions.
\(Q_{\text{req}}\) Request Concurrency requests Average in-flight requests in a stable serving system (\(Q_{\text{req}} = \lambda_{\text{arr}} \cdot T_{\text{lat}}\) via Little’s Law). Avoids collision with \(N\) as device count in distributed settings.
\(\lambda_{\text{arr}}\) Arrival Rate requests/s Request arrival rate. The subscript avoids collision with \(\lambda\) as degradation sensitivity or failure rate.
\(T_{\text{lat}}\) Request Time in System seconds End-to-end queueing/serving latency for Little’s Law. Distinct from \(L_{\text{lat}}\), the fixed-latency term in the iron law.
\(T_{\text{svc}}(B)\) Batch Service Time seconds Time to serve a batch of size \(B\).
\(\mu_{\text{eff}}(B)\) Effective Service Rate requests/s Batched service rate, typically \(\mu_{\text{eff}}(B) = B/T_{\text{svc}}(B)\).
\(\rho_{\text{serv}}\) Serving Utilization Dimensionless Queue/server utilization. Use instead of bare \(\rho\), which is reserved for communication-computation ratio in distributed contexts.
\(M_{\text{total}}\) Total Memory Footprint bytes Sum of explicit memory components; avoids bare \(M\) ambiguity.
\(M_{\text{weights}}\) Weight Memory bytes Memory occupied by model parameters.
\(M_{\text{gradients}}\) Gradient Memory bytes Memory occupied by stored gradients.
\(M_{\text{optimizer}}\) Optimizer-State Memory bytes Momentum, variance, master weights, and related optimizer buffers.
\(M_{\text{activations}}\) Activation Memory bytes Retained activations for the backward pass or serving intermediates.
\(s_{\text{elem}}\) Element Storage Size bytes/element Bytes per stored tensor element.

Units and Precision

  • Physical units: This book uses SI (metric) units throughout, including meters, kilograms, seconds, watts, and °C, consistent with standard engineering and scientific practice. Where source data was originally reported in imperial units, the book converts to SI and notes the original values parenthetically. A space always separates the number from the unit (for example, 100 ms, 2 TB/s).
  • Data and memory: This book uses decimal SI prefixes only: KB = \(10^3\) bytes, MB = \(10^6\), GB = \(10^9\), TB = \(10^{12}\). Binary-prefixed units do not appear in prose; all capacities, throughputs, and model sizes are reported in decimal units (for example, 80 GB, 2 TB/s, 102 MB).
  • Compute: The notation distinguishes operation counts from rates. Total work uses FLOPs (for example, GFLOPs, TFLOPs), while throughput uses FLOP/s with decimal prefixes (for example, GFLOP/s, TFLOP/s).
    • 1 TFLOP = \(10^{12}\) FLOPs
    • 1 TFLOP/s = \(10^{12}\) FLOPs per second
    • Arithmetic intensity conventionally uses FLOP/byte as a unit ratio (floating-point operations per byte moved). This is a ratio unit, not a total-work symbol or throughput symbol.
  • Currency: Dollar amounts use the dollar sign ($); unless otherwise noted, dollar-denominated costs are U.S. dollars (USD).
  • Precision:
    • FP64: Double precision (8 bytes)
    • FP32: Single precision (4 bytes)
    • TF32: TensorFloat-32 (19-bit Tensor Core compute mode; inputs and storage remain FP32)
    • FP16: Half precision (2 bytes, standard range)
    • BF16: Brain float (2 bytes, wide dynamic range)
    • FP8: Quarter precision (1 byte, E4M3 or E5M2 format)
    • FP4: 4-bit floating-point format
    • INT8: 8-bit integer (1 byte)
    • INT4: 4-bit integer; lower integer precisions follow the same uppercase INTn pattern (for example, INT3, INT2)

Quick Reference: Resolving Collisions

Common collision points in ML Systems literature include:

Symbol ML Meaning Systems Meaning Book Convention
\(B\) Batch Size Bandwidth Batch Size. Use \(\text{BW}\) for bandwidth.
\(P\) Parameters Peak FLOP/s Parameters. Use \(R_{\text{peak}}\) for peak rate.
\(D\) Dataset Size Data Volume Dataset Size. Use \(D_{\text{vol}}\) for bytes moved.
\(L\) Loss Latency Loss \((\mathcal{L})\). Use \(L_{\text{lat}}\) for latency.
\(\eta\) Learning Rate Efficiency Learning Rate. Use \(\eta_{\text{hw}}\) for efficiency.

As a general principle, ML conventions take precedence for single letters, while systems concepts get subscripts or multi-letter symbols. This reflects the primary audience (ML practitioners learning systems) and preserves compatibility with the vast ML literature.

Agentic Systems Notation

This book adds notation for trajectories, the task contract, the H·S·A exposures, model calls and their budgets, the time and cost of a trajectory, recovery, and learning from trajectories. The shared tables above still govern every symbol they define, and this extension only adds to them. A symbol that appears only inside one derivation is defined where it is used and is not listed here.

Five relations recur throughout the book. A trajectory, the ordered record of one delegated task, is written

\[\tau = (s_0, a_0, o_1, s_1, a_1, o_2, \dots, s_T)\]

Its wall-clock duration over a horizon of \(H\) turns is

\[T_{\text{task}} = \sum_{k=1}^{H} \left( T_{\text{model}}^{(k)} + T_{\text{tool}}^{(k)} + T_{\text{wait}}^{(k)} + T_{\text{runtime}}^{(k)} \right)\]

Trajectory goodput is the share of a resource \(R\) (turns, tokens, sandbox time, or dollars) spent on trajectories whose results passed verification,

\[\mathcal{G} = \frac{\sum_{i \in \mathcal{T}_{\text{success}}} R_i}{\sum_{j \in \mathcal{T}_{\text{all}}} R_j}\]

A trajectory that runs \(N\) sequential proposals without checking any of them, each correct with probability \(1 - \epsilon\) and independent of the others, compounds its errors and succeeds with probability

\[\Pr(\text{success}) = (1 - \epsilon)^N\]

The cost per accepted task divides everything spent on \(N\) attempts by the \(N_{\text{acc}}\) attempts that were accepted,

\[C_{\text{eff}} = \frac{\sum_{i=1}^{N} C_{\text{attempt}, i}}{N_{\text{acc}}}\]

Trajectories and the agent loop

The loop symbols describe one turn of a trajectory: the context the runtime assembles, the proposal the model returns, the decision the runtime makes about it, and what comes back.

Symbol Definition Unit Notes
\(\tau\) Trajectory Tuple Ordered record of one delegated task, \(\tau = (s_0, a_0, o_1, \dots, s_T)\). Bare \(\tau\) always means a trajectory in this book.
\(t\) Turn index Integer Counts turns within a trajectory. \(s_T\) is the state after the final turn.
\(g\) Goal Text The delegated task and its completion rule, fixed for the trajectory. The task contract writes it \(G\).
\(c_t\) Context Tokens What the runtime assembles and sends on turn \(t\): instructions, goal, tool definitions, and the history of actions, observations, and evidence.
\(s_t\) State at turn \(t\) Record The context state the model saw on turn \(t\). In Part V, the state a policy conditions on.
\(a_t\) Action proposal Text or tool call What the model proposed on turn \(t\). Nothing happens until the runtime admits it.
\(a_{\text{prop}}\), \(a_{\text{perm}}\) Proposed, permitted action Tool call A proposal before the runtime checks it, and the action it permits after the check.
\(o_t\) Observation Typed result What an action returned, such as an exit code, captured output, a response body, or a status.
\(v_t\) Verification evidence Pass/fail or score Result of a check the runtime runs on the resulting state, independent of the action and of the model’s report.
\(\pi_\theta\) Policy Distribution The model viewed as a distribution over proposals given a context, with parameters \(\theta\).
\(\mathcal{S}\) Agentic system Tuple \(\mathcal{S} = \langle \pi_\theta, \mathcal{H}, \mathcal{E}, \mathcal{M}, \mathcal{V} \rangle\): policy, runtime harness, environment, memory, and verifiers.
\(\mathcal{I}\) Closure operator Function \(\mathcal{I}: \mathcal{A}_{\text{prop}} \times \mathcal{S}_{\text{sys}} \to \mathcal{A}_{\text{perm}} \cup \{\bot\}\). The runtime’s check of each proposal against the current system state \(\mathcal{S}_{\text{sys}}\).
\(\bot\) Rejection Symbol A refused proposal. It reaches the model as a typed error and leaves the environment unchanged.

The task contract

A task contract fixes what a trajectory is for, where it runs, what it may do and see, and what ends it.

Symbol Definition Unit Notes
\(\mathcal{C}\) Task contract Tuple \(\mathcal{C} = \langle G, \mathcal{E}_{\text{env}}, \mathcal{A}_{\text{perm}}, \mathcal{O}_{\text{avail}}, \mathcal{K}_{\text{comp}} \rangle\).
\(G\) Goal Text Desired end state, with explicit negative scope.
\(\mathcal{E}_{\text{env}}\) Environment Specification Reproducible place the work happens: base commit, image, configuration, network rules.
\(\mathcal{A}_{\text{perm}}\) Permitted actions Set Tools the task may call and the limits on each. Least privilege stated per task.
\(\mathcal{O}_{\text{avail}}\) Available observations Specification What the agent may see, and how much, including truncation and redaction rules.
\(\mathcal{K}_{\text{comp}}\) Completion criteria Set of checks Checks that decide when the task is done, at a stated closure evidence level. The model’s claim of completion is never one of them.

The H·S·A exposures

H·S·A names the three exposures an agentic task adds beyond a single model call. A task’s position says how exposed it is. The closure the runtime must supply against that exposure is governed by the invariant closure principle and is not a fourth axis. Zero ambient authority describes the model at every authority level and has no symbol.

Symbol Definition Unit Notes
\(H\) Horizon Turns Number of turns a trajectory runs: \(H = 1\) (single call), \(H \approx 10\) (short loop), \(H \geq 100\) (minutes to hours). Wall-clock duration given \(H\) is \(T_{\text{task}}\).
\(S_0\) to \(S_3\) State Level What the task carries between turns: \(S_0\) (nothing), \(S_1\) (context window and KV cache), \(S_2\) (sandboxed working files), \(S_3\) (durable state that outlives the session or is shared beyond it). Written with its level, because bare \(S\) is sequence length.
\(A_0\) to \(A_3\) Authority Level What the task may do through its tools: \(A_0\) (read access), \(A_1\) (mutation confined to a discardable sandbox), \(A_2\) (external action that is retry-safe or compensable), \(A_3\) (irreversible external action). Covers confidentiality as well as integrity, so \(A_0\) is not a zero blast radius.

Closure evidence levels, from weakest to strongest, are named in words and carry no symbol: self-report, static checks, visible tests, sealed tests, and formal proof.

Model calls and budgets

A model call is bounded by its window and by limits the runtime sets on each request. A trajectory is bounded by a budget the harness keeps outside the context.

Symbol Definition Unit Notes
\(S\) Sequence length Tokens The shared symbol for sequence length, used here for the prompt a call reads. The State exposure is always written with a level, \(S_0\) to \(S_3\).
\(K\) Output length Tokens Tokens a call generates. Lowercase \(k\) is reserved for sample counts, as in pass@\(k\).
\(S_{\max}\) Context window Tokens Hard per-call limit on prompt and output together: \(S + K_{\max} \le S_{\max}\).
\(K_{\max}\) Output limit Tokens Maximum tokens the runtime lets a call generate.
\(B_{\text{reason}}\) Reasoning budget Tokens Limit on reasoning tokens a call may spend before it answers.
\(t_{\text{tok}}\) Per-token decode time Seconds Time to generate one output token on a single stream.
\(T_{\text{call}}\) Call latency Seconds \(T_{\text{call}} \approx T_{\text{queue}} + T_{\text{prefill}}(S) + K \cdot t_{\text{tok}}\). The first two terms are the time to first token.
\(\mathbf{B}\) Trajectory budget Vector \(\mathbf{B} = \langle H_{\max}, N_{\max}, T_{\max}, C_{\max} \rangle\), ceilings on turns, tokens, wall-clock time, and dollars, checked by the harness before every call. Bold, to stay distinct from batch size \(B\).
\(T_{\max}\) Deadline Seconds Wall-clock ceiling enforced by the runtime, on one call or on a whole trajectory. The text names the scope.
\(C_{\max}\) Cost ceiling Dollars Spend ceiling the runtime enforces before dispatching the next call.
\(m_{\text{token}}\) KV bytes per token Bytes/token \(m_{\text{token}} = 2 \cdot N_L \cdot H_{\text{KV}} \cdot d_{\text{head}} \cdot s_{\text{elem}}\), in the shared symbols for layers, KV heads, head dimension, and element size.
\(M_{\text{KV}}\) KV footprint Bytes Attention state a trajectory holds in accelerator memory, \(M_{\text{KV}} = S \cdot m_{\text{token}}\).
\(O_{\text{mem}}\) Memory held during a wait Byte-seconds \(O_{\text{mem}} = M_{\text{KV}} \cdot T_{\text{wait}}\), the accelerator memory a paused trajectory withholds while it waits on a tool.

Time and cost of a trajectory

Agentic systems are measured per accepted task, not per request. The time terms decompose one turn, and the cost terms price attempts and the accepted results they buy.

Symbol Definition Unit Notes
\(T_{\text{task}}\) Trajectory duration Seconds Sum over turns of model, tool, wait, and runtime time.
\(T_{\text{model}}\) Model time Seconds Time spent in model calls during a turn.
\(T_{\text{tool}}\) Tool time Seconds Time a tool takes to execute, such as a build or a test run.
\(T_{\text{wait}}\) External wait Seconds Time spent waiting on a remote service or a person’s approval.
\(T_{\text{runtime}}\) Runtime time Seconds The runtime’s own work: assembling context, validating proposals, running checks.
\(\mathcal{T}_{\text{all}}\), \(\mathcal{T}_{\text{success}}\) Trajectory sets Set Every trajectory dispatched, and the subset whose results passed verification. Calligraphic \(\mathcal{T}\) with a subscript always names a set of trajectories.
\(\mathcal{G}\) Trajectory goodput Fraction \([0, 1]\) Share of a resource spent on trajectories that passed verification. Badput is \(1 - \mathcal{G}\).
\(C_{\text{task}}\) Trajectory cost Dollars or tokens What one trajectory spent, checked against \(C_{\max}\).
\(C_{\text{attempt}}\) Attempt cost Dollars Full cost of one attempt: input and output tokens, tool fees, sandbox time, verification, and human review.
\(N_{\text{acc}}\) Accepted attempts Count Attempts whose results passed the acceptance checks.
\(C_{\text{eff}}\) Cost per accepted task Dollars/task Total spend on all attempts divided by \(N_{\text{acc}}\). Failed attempts are paid for in full.

Recovery, evaluation, and learning

These symbols recur across the runtime, evaluation, and learning chapters.

Symbol Definition Unit Notes
\(k_{\text{idem}}\) Idempotency key String Deterministic key derived from the trajectory, turn, call, tool, and canonical arguments, so a retried call cannot apply its effect twice.
\(K_{\text{repair}}\) Repair cap Count Maximum forward repair attempts on one failed step before the harness compensates or escalates.
pass@\(k\) Pass at \(k\) Probability Probability that at least one of \(k\) runs of a task succeeds. Measures potential, and is reachable only when a verifier selects among the runs.
pass\(^k\) Pass to the \(k\) Probability Probability that all \(k\) runs of a task succeed. Measures reliability over repeated runs.
\(m_t\) Loss mask Binary \(\{0, 1\}\) \(m_t = 1\) on tokens the policy wrote (reasoning, tool calls, answers) and \(m_t = 0\) on prompts and observations.
\(\mathcal{L}_{\text{SFT}}\) Masked fine-tuning loss Cross-entropy \(\mathcal{L}_{\text{SFT}}(\theta) = -\sum_t m_t \log \pi_\theta(x_t \mid x_{<t})\).
\(R(\tau)\) Trajectory reward Scalar Reward a verifier assigns to a whole trajectory.
\(\hat{V}(s_t)\) Estimated value Scalar Estimated probability of success from state \(s_t\).
\(\hat{A}(s_t, a_t)\) Estimated advantage Scalar How much an action raised the estimated value, \(\hat{V}(s_{t+1}) - \hat{V}(s_t)\). Distinct from the authority levels \(A_0\) to \(A_3\).

Resolving collisions

Agentic systems borrow notation from machine learning, systems, and control, so several letters arrive with more than one meaning. The conventions below keep one meaning per symbol.

Symbol Other common meanings Convention in this book
\(\tau\) Temperature; a threshold Bare \(\tau\) is a trajectory. Sampling temperature is named in words or written \(\tau_{\text{temp}}\).
\(H\) Attention heads; hidden size \(H\) is the horizon in turns. Heads use \(N_{\text{heads}}\) and \(H_{\text{KV}}\); hidden size uses \(d\).
\(S\) The State exposure Bare \(S\) is sequence length. State is written only with a level, \(S_0\) to \(S_3\).
\(A\) Advantage; action sets Authority is written only with a level, \(A_0\) to \(A_3\). Advantage is \(\hat{A}\); action sets are \(\mathcal{A}_{\text{prop}}\) and \(\mathcal{A}_{\text{perm}}\).
\(P\) Prompt length; precision \(P\) is parameter count, as in the shared table. Prompt length is \(S\).
\(B\) A budget \(B\) is batch size. The trajectory budget is bold \(\mathbf{B}\), and the reasoning budget is \(B_{\text{reason}}\).
\(C\) A fourth taxonomy axis There is no \(C\) axis; closure evidence levels are named in words. Costs carry subscripts (\(C_{\text{task}}\), \(C_{\text{eff}}\), \(C_{\max}\)), and the task contract is \(\mathcal{C}\).
\(\mathcal{G}\) The goal \(\mathcal{G}\) is trajectory goodput. The goal is \(g\) in the loop and \(G\) in the task contract.
\(\mathcal{T}\) A sequence of steps Calligraphic \(\mathcal{T}\) with a subscript is a set of trajectories. Time is \(T\).
\(K\) Top-\(k\); sample count \(K\) is output length in tokens. Sample counts use lowercase \(k\), as in pass@\(k\).

Additional units

  • Tokens count what a call reads and writes, and serving throughput is reported in tokens per second.
  • Turns count the steps of a trajectory, and the horizon \(H\) is measured in turns.
  • Dollars per accepted task normalize spend by the results that passed acceptance, not by attempts or requests.
  • Trajectory goodput (\(\mathcal{G}\)) is reported as a fraction in \([0, 1]\) or as a percentage, together with the resource it counts.
  • Byte-seconds measure memory held over time, such as attention state kept resident during a tool wait.
Back to top