Glossary
This glossary defines key terms used throughout Agentic Machine Learning Systems. Terms are organized alphabetically.
A
- accepted task
- A trajectory whose final environment state passes the task’s completion criteria at the closure evidence level the task requires, with the trajectory inside every budget ceiling its contract states. It is the unit in which agent evaluation counts success and agent economics counts cost.
- admission cascade
- The sequence of checks, ordered from cheapest to most expensive, that a logged trajectory must pass before it may enter a training set. Cheap static checks run first so that the expensive sealed tests run only on candidates that could still pass.
- agent harness
- The runtime component that owns the agent loop for one trajectory. It assembles each model call, interprets the reply, decides whether and when each proposed tool call is dispatched, and keeps the trajectory record through which anything outside the loop can inspect, pause, bound, or stop the trajectory.
- agentic machine learning system
- A closed-loop system in which a foundation model proposes actions toward a delegated goal and a deterministic runtime decides what each call sees, which proposals take effect, and what evidence completes the task.
- agentic retrieval
- Retrieval in which the model issues search, read, and lookup calls as tools, turn by turn, instead of receiving a fixed set of passages staged before the call. It trades extra turns for evidence chosen as the task unfolds.
- approval gate
- The harness mechanism that holds an irreversible tool call, undispatched, until an authorized reviewer approves it. See also escrow.
B
- benchmark contamination
- Overlap between an evaluation’s tasks, or their solutions, and the data a model was trained on, which inflates measured success without reflecting capability on unseen tasks. Split hygiene and decontamination keep training data and evaluation tasks apart.
- best-of-\(n\) sampling
- Drawing \(n\) independent candidates for one decision and selecting one with a verifier, a judge, or a vote. Its ceiling is pass@\(k\) with \(k = n\), reached only by a selector that always picks a correct candidate when one exists.
- blackboard topology
- A multi-agent coordination pattern with no coordinator, in which workers read and write a shared store of artifacts, claim subtasks whose inputs have appeared, and post results.
- blast radius
- The set of state that an action can read, change, or send outside its sandbox, as fixed by the runtime’s enforcement rather than by what the model was asked to do.
- budget reservation
- An amount carved out of a parent task’s remaining budget before a subagent or long-running call starts, so that concurrent spenders cannot together exceed the parent’s limit. Unspent reservations return to the parent when the child finishes.
C
- call surface
- The typed request and response of one model call: the messages, tool definitions, and limits the runtime sends, and the proposal, stop reason, and token usage the model service returns.
- cancellation context tree
- The runtime’s record of which agent spawned which, in which each agent carries a cancellation state and a deadline no later than its parent’s, and canceling a node cancels every descendant.
- capability
- An unforgeable token, signed by the runtime, that names the rights an action may exercise, the objects it may touch, a predicate bounding its arguments and cost, and an expiry, written \(C = (R, O, P, E)\). The sandbox checks the capability on every access.
- checkpoint
- A durable snapshot of the state a trajectory’s log cannot rebuild cheaply, such as its workspace, bound to the log position it corresponds to, so that resume need not replay from the beginning.
- chunked prefill
- A serving technique that splits a long prompt or observation into fixed-size chunks and processes one chunk per step alongside ongoing decodes, so that one large prefill does not stall every other trajectory on the server.
- circuit breaker
- A tool-gateway mechanism that stops dispatching calls to a dependency whose transport-level failures cross a threshold, answers further calls at once with a typed observation, and probes the dependency again after a cooldown. It keeps the model from acting as a retry amplifier against a failing service.
- closure evidence levels
- The ordered kinds of evidence that a task’s result is correct: the model’s self-report, static checks, visible tests, sealed tests, and formal proof. Higher H·S·A exposure demands a higher level.
- complete mediation
- The invariant that every tool call a model proposes passes through the runtime’s validation and authorization checks, on every call, before any effect reaches the environment, and that every result passes back through the runtime before it enters the model’s context.
- context compaction
- The transformation of a trajectory’s staged history into fewer tokens that keeps verbatim every identifier, path, failure, and constraint that later actions depend on.
- context engineering
- The runtime discipline of deciding what enters a model’s context on each call of a trajectory, in what layout and form, and for how long each staged item remains valid.
- context poisoning
- The failure in which a wrong assumption or false result, once in a trajectory’s context, becomes the premise of later proposals, because every later call is conditioned on the history.
- coordination tax
- The time and tokens a multi-agent design spends on handoffs, duplicated context, waiting at join points, and reconciling conflicting work. It grows with the number of agents and can make a larger fleet slower than a smaller one.
- credential brokering
- Keeping long-lived secrets out of the sandbox and the model’s context by having a broker attach scoped, short-lived credentials to outbound requests at the point of use.
D
- DAgger (dataset aggregation)
- An imitation-learning procedure that runs the current policy, has an expert or verifier label the states it actually visits, and adds those labels to the training data, which bounds the compounding of errors that pure imitation suffers.
- decode
- The phase of a model call that generates output one token at a time. For a single stream each step reads all of the model’s weights to produce one token, so decode is limited by memory bandwidth and its latency grows with output length.
- distillation
- Training a model to reproduce behavior it previously needed extra context or a larger model to produce. Context distillation moves instructions from the prompt into the weights; teacher distillation trains a smaller model on a larger model’s trajectories.
E
- escrow
- The harness’s holding of an irreversible proposed tool call, undispatched, until an authorized reviewer approves it. The held call carries an expiry, is re-checked against its recorded preconditions before dispatch, and is replaced by a refusal observation if it is denied, expires, or has gone stale. This is the only sense of the term in this book.
- exposure bias
- The mismatch between training a policy on expert histories and running it on its own, in which one early mistake leads to states the training data never covered and errors compound over the horizon.
F
- fail-plausible fault model
- The fault model of a model in an agent loop, whose output conforms to every syntactic check and exits cleanly while violating the task’s actual requirements. Unlike a fail-stop component it does not halt at the edge of its competence, and unlike a Byzantine adversary it has no intent.
- fixed workflow
- A program whose steps, branches, and tool calls are decided by code written in advance, possibly with a model inside individual steps. It is the right design whenever a task’s steps can be enumerated.
G
- grammar-constrained decoding
- A decoding method that restricts each sampling step to the tokens a formal grammar admits in the current state, by setting the logits of every other token to \(-\infty\). It guarantees that output parses, and nothing about what the output means.
- grant
- The runtime’s per-call decision of whether a tool may run now, with these arguments, under the permissions the session explicitly holds, made from the runtime’s own descriptor of the tool rather than from anything the tool’s server advertises.
- Group Relative Policy Optimization (GRPO)
- A policy-gradient method for reinforcement learning that scores each sampled response relative to the other responses to the same prompt, using the group’s mean and spread in place of a learned value model.
H
- H·S·A exposures
- The three ways an agent task can fail that a single model call cannot: Horizon (\(H\), the number of turns), State (\(S_0\) to \(S_3\), what the task carries between turns), and Authority (\(A_0\) to \(A_3\), what the task may do to the world). Closure is the runtime’s answer to them, not a fourth exposure.
- heterogeneous invariant-gated coordination
- A multi-agent discipline in which diverse model families or prompts generate and rank candidates, and a deterministic check that no agent can modify decides which candidate is committed.
I
- idempotency key
- An identifier the runtime derives from a trajectory’s logical call and attaches to a mutating tool call, so the endpoint can recognize a repeat and return the stored result instead of executing again.
- idempotent action
- A tool call whose repeated execution leaves the environment in the same state as a single execution, \(f(f(s)) = f(s)\), so resending it after an ambiguous timeout cannot duplicate its effect.
- indirect prompt injection
- Instructions embedded in content an agent reads, such as a web page, email, issue, or tool result, that steer the model’s later proposals. Model-side defenses lower its probability; only enforcement below the model bounds its damage.
- intervention ladder
- The ordering of fixes for a measured agent failure from cheapest and most reversible to most expensive: a fixed workflow, then context, tools, runtime, supervised fine-tuning, reinforcement learning, and more agents. Each rung is climbed only when a failure has been measured and attributed to the rung below.
- invariant closure principle
- The principle that an agent’s safety and resource bounds must be enforced mechanically by the runtime below the model, and its task correctness established by evidence collected at the end-to-end boundary above it, never by the model’s own compliance or self-report.
- isolation envelope
- The enforcement boundary around the code an agent’s action runs, defined by what that code shares with the host (kernel, memory, files, network) and enforced by a layer the code cannot modify. Shared-kernel containers, user-space kernels, microVMs, and bytecode sandboxes are the common envelopes.
J
- job handle
- A receipt the runtime returns within the same turn when it starts a long-running tool, so the harness worker is released and the trajectory resumes when the job’s result arrives instead of blocking on it.
K
- KV cache
- The keys and values a model computes for every token in its context, kept so that each new token attends to earlier ones without recomputing them. Its size grows linearly with context length and it must stay in serving memory while the trajectory is live, unless it is evicted and later recomputed.
L
- LLM judge
- A general large language model prompted with a task and one or more candidates and asked to score or compare them. Judges are useful where no mechanical check exists, carry measurable biases, and must be calibrated against verifiers before their scores are trusted.
M
- Model Context Protocol (MCP)
- A protocol through which tool servers advertise tools and their schemas to an agent runtime and receive calls. Discovering a tool through it does not authorize the tool; the runtime’s grant does.
- model cascade
- A routing design that sends each call to the least expensive model likely to handle it and escalates to a more capable model when a check on the result fails.
- model-directed loop
- A design in which the model chooses the next action from a permitted set after reading each result, and decides when to stop, so two runs of the same task may take different paths. It is warranted only when a task’s path cannot be enumerated and its result can be checked.
- monotonic delegation
- The rule that authority can only narrow as work is delegated: a child agent’s capability never holds a right, object, budget, or lifetime its parent lacks.
O
- observation loss masking
- A per-token training mask over a serialized trajectory that computes loss only on tokens the policy writes (its reasoning, its tool calls, and its end-of-turn token) and keeps prompts and environment observations as context.
P
- paged KV cache
- A serving-memory layout that stores each trajectory’s attention state in fixed-size blocks from a shared pool and maps its token positions to those blocks through a per-trajectory table, so memory is held only for tokens that exist and forks can share blocks.
- pass@\(k\)
- The probability that at least one of \(k\) independent attempts at a task succeeds. It measures potential, and it is the ceiling a best-of-\(k\) selector can reach.
- pass\(^k\)
- The probability that all \(k\) independent attempts at a task succeed. It measures reliability, falls as \(k\) grows, and is the metric for tasks that recur without supervision.
- pending-proposal buffer
- Runtime-owned memory that holds a model’s proposal as inert data while it is parsed, validated, and authorized, before anything acts on it.
- pivot action
- An irreversible action with no compensating action, such as sending an email or settling a payment. A trajectory allows at most one, placed after all compensable work and behind an approval gate, and after it recovery can only continue forward.
- prefill
- The phase of a model call that processes the prompt. All prompt tokens share one read of the weights, so prefill is limited by arithmetic and costs far less time per token than decode.
- prefix caching
- Reusing the attention state already computed for the unchanged beginning of a prompt, so that each turn of a trajectory processes only its new tokens. It is the dominant saving on the input an agent re-sends every turn, and any change to the stable prefix discards it.
- process reward model (PRM)
- A learned verifier that scores each intermediate step or tool action of a multi-step trajectory, rather than only its final result.
- progress detection
- Runtime monitoring that judges whether a trajectory is advancing from the environment’s state, for example by hashing repeated actions and comparing states across turns, rather than from heartbeats or the model’s own account.
Q
- quarantined reader
- A pattern in which a model with no tools reads untrusted content and returns only typed values, while a separate planner that never sees the untrusted text decides which actions to take.
R
- radix prefix cache
- An index that maps token-sequence prefixes to the attention-state blocks computed for them, so that a new request reuses the blocks for its longest cached prefix and processes only the remaining tokens.
- reasoning budget
- A limit, set in the call surface, on how many tokens a reasoning model may spend thinking before it answers.
- reciprocal rank fusion (RRF)
- A way to merge ranked lists from different retrieval methods by scoring each item by the sum of the reciprocals of its ranks, which avoids calibrating incompatible scores against each other.
- recovery demonstration
- An admitted training trajectory that contains an environment fault, tool error, or wrong hypothesis at a marked turn, followed by diagnosis and a corrective action, and that still passes every gating check.
- reinforcement learning with verifiable rewards (RLVR)
- Policy optimization in which a trajectory’s reward is computed by an automated check that runs outside the model, such as a sealed test suite, a compiler, or a proof checker, applied to the state the trajectory produced.
- replay
- Running the harness against a trajectory’s log with the model and tools stubbed to return their recorded replies and results, so the harness’s behavior can be reproduced and inspected without new model calls or effects.
- resume
- The protocol that continues a trajectory on a new worker from its log and latest checkpoint, settling every in-doubt effect before anything new is dispatched.
- retrieval contract
- The typed query and response between the runtime and durable storage that bounds scope, candidate count, latency, and freshness, and checks each result’s version and provenance against its source before it may enter context.
- reward hacking
- A policy raising its measured reward through behavior that does not accomplish the task the reward was meant to measure, by exploiting gaps in what the reward checks or by reaching and altering the mechanism that computes it.
S
- sealed test
- A test that judges an agent’s result from outside its reach, run on the resulting state in a fresh environment with the tests unreadable and unmodifiable by the agent. It is the strongest closure evidence short of proof.
- self-consistency
- Selecting among sampled candidates by majority vote over their final answers. It works when answers are short and comparable and fails when candidates share the same error.
- semantic compensation
- A forward action that returns the environment to a state equivalent to the one before a step, with respect to the invariants the task cares about, while leaving a record that the step and its reversal both happened.
- sequence packing
- Concatenating several independent training examples into one fixed-length sequence, with attention and positions partitioned so that each example is processed as if it were alone.
- settlement
- Resolving whether a tool call whose outcome is unknown, such as one that timed out, actually took effect, using its idempotency key or a probe of the environment, before any retry.
- shadow run
- A release stage in which a candidate agent receives copies of production inputs without authority to change production state, so its behavior can be compared with the baseline’s.
- specification boundary
- The line past which an agent cannot go: it can realize a design but cannot author the contract that governs it. The goal, environment, permitted actions, observations, and completion criteria stay with the engineer, and any runtime can verify a result only against what that contract states.
- speculative decoding
- A serving technique in which a small draft model proposes several tokens and the target model checks them in one step, with an acceptance rule that keeps the output distribution identical to the target model’s. It shortens decode latency by a constant factor and pays most when the server is lightly loaded.
- spend-rate limit
- A bound on how fast a task or tenant spends, computed over a sliding window, which pauses spending that is within budget but too fast.
- stop reason
- The field in a model response that reports what ended generation, such as a finished turn, emitted tool calls, the output limit, a stop sequence, or a refusal. It is the agent loop’s control signal.
- Stochastic Computer
- The single map, drawn once in the introduction, that relates the parts of an agentic system to the parts of a computer: the model to a processor, context and stores to memory, tools to input and output, the runtime to an operating system. It is a map for orientation, not a mechanism.
- structured progress notes
- Typed state, such as decisions, open hypotheses, and constraints, that the runtime extracts from a trajectory’s history before compaction and keeps outside the model’s free-form summary, so it survives the loss of the raw history.
T
- taint tracking
- Marking data derived from untrusted sources and propagating the mark through everything computed from it, so that tool calls with authority can refuse arguments that carry it.
- task contract
- The five-part specification of a delegated task: its goal with negative scope, the environment, the permitted actions, the available observations, and the completion criteria at a stated evidence level.
- task environment
- A reproducible, isolated setting in which an agent attempts an evaluation or training task, reset between trials and holding its sealed tests outside the agent’s reach.
- task fixture
- A versioned, reproducible task definition for collecting training data: the starting environment, the task statement, and the checks that decide success.
- test-tampering defense
- The admission rule that the tests judging a trajectory live outside the agent’s writable workspace and that no admitted change touches them.
- test-time compute
- Inference computation spent on a single decision beyond one model call, through longer reasoning, additional candidates, or check-and-revise rounds, before the runtime admits a result.
- time-travel replay
- Reconstructing a trajectory’s exact state at any past turn by loading the nearest earlier checkpoint and replaying the log up to that turn, with the model and tools stubbed.
- trace
- A graph of spans recording one trajectory’s model calls, tool calls, approvals, and child agents, with their durations, token counts, and outcomes, used to measure and attribute behavior rather than to recover it.
- trajectory
- The ordered record of a delegated task’s execution, interleaving the context the model saw, the action it proposed, and the observation the runtime returned at each step. It is the unit of engineering in agentic systems.
- trajectory goodput
- The fraction of the resources spent across all trajectories that went to trajectories whose results passed independent verification.
- trajectory log
- The append-only sequence of typed events (contexts built, model replies, tool intents and results, approvals, compactions) that is the authoritative record of what a trajectory did. The trajectory record is a projection of it.
- trajectory record
- The per-trajectory control state that the harness owns and the model never writes: its identity, current state, per-call limits, budgets and measured spend, grants and leases, pending work, and pointers to its context, plan, log, and workspace.
- trajectory saga
- The structure a runtime imposes on a trajectory’s external effects, in which each tool call commits immediately and has a compensating action registered before it is dispatched, so that a failure partway through can be answered by running the compensators in reverse order.
- turn boundary
- A moment in a trajectory when no model call is in flight and no tool call is half dispatched, at which the harness applies pending control actions such as pause or cancel.
- typed task envelope
- The schema-validated record a runtime passes from a delegating agent to a child, naming the task’s identity and trace context, immutable references to its inputs, the attenuated authority the child receives, its budget, and machine-checkable completion criteria.
U
- usable context
- The input length over which a model reliably retrieves and combines the several pieces of evidence a task depends on, measured on the workload rather than read from a specification. It is usually much shorter than the nominal context window.
V
- verification enclave
- A separate, fresh sandbox in which a trajectory’s result is checked, holding the tests and reward logic outside the policy’s workspace so the policy cannot read or alter them.
- verifier
- Any check that judges a candidate or a result, from deterministic checks such as compilers and sealed tests to learned scorers and judges. Deterministic verifiers can gate what is committed; learned ones only rank what the deterministic gate then decides.
W
- whole-trajectory cost
- The total spent on model tokens, tools, sandboxes, verification, and human review across every attempt, divided by the number of attempts whose results were accepted.
- write-ahead ordering
- The rule that no external effect may leave the host until the log event that authorizes it is durable, so that after a crash every effect is either known to have been intended or known not to have happened.
Z
- zero ambient authority
- The property that the model itself holds no rights to read, change, or reach anything; every effect it proposes takes place only through a grant and a capability the runtime issues and enforces.