Foundations of Agentic Systems
Purpose
Why does an accurate model output fail to complete a delegated task?
A foundation model reads a context and returns text, and text changes nothing in the world. Delegating a task, such as repairing a failing service or migrating a database, asks for something else, a sequence of actions whose effects are real, whose failures are often silent, and whose result must be checked by something other than the model that proposed it. Over tens or hundreds of turns small errors compound, tools time out or half succeed, and the model can report success in fluent, well-formed text while the work underneath is broken. Reliability therefore cannot come from the model alone. It has to be built around the model, as a runtime that decides what each call sees, validates every proposed action before it runs, contains what an action can damage, records progress so that a crash does not erase it, and accepts a result only on evidence it gathers itself. How much of that machinery a task needs depends on three exposures a single model call does not have, a long horizon, state carried between turns, and authority over the world, and on the closure the runtime must supply against each.
Learning Objectives
- Trace one agent turn through proposal, validation, dispatch, observation, and verification, separating what the model proposes from what the runtime does.
- Distinguish fail-plausible faults from fail-stop and Byzantine faults, and recognize context poisoning and indirect prompt injection as fail-plausible hazards.
- Apply the invariant closure principle to decide which properties the runtime must enforce mechanically and which closure evidence level a result requires.
- Specify a five-part task contract (goal, environment, permitted actions, available observations, completion criteria) for a delegated task.
- Classify a task by its horizon, state, and authority exposures (H·S·A) and derive the closure its runtime must supply.
- Calculate trajectory goodput and whole-task duration, and use Amdahl’s law to find the term that bounds trajectory speedup.
- Select a fixed workflow or a model-directed loop with the four placement questions, and justify each rung of the intervention ladder by a measured failure.
The Agentic Systems Moment
For more than seven decades, the foundational contract of computer systems rested upon the stored-program concept. As formulated by John von Neumann in 1945 (Neumann 1945) and realized on Cambridge’s EDSAC by Maurice Wilkes, David Wheeler, and Stanley Gill in 1949 (Wilkes et al. 1951), software engineers commanded physical silicon through sequential, explicit machine opcodes. Even when concurrent operating system runtimes and distributed networks contended with asynchronous timing, race conditions, and transient hardware faults, deterministic execution remained the baseline design contract: an instruction pointer advanced deterministically, state transitions followed explicit logic, and non-determinism was an exceptional defect to be isolated.
Over the past decade, deep learning challenged this paradigm. By substituting hand-engineered procedural algorithms with continuous parameter optimization via stochastic gradient descent (Rumelhart et al. 1986), machine learning shifted the computing contract from explicit instructions to learned statistical models. Throughout this initial decade of deep learning, systems engineering focused almost exclusively on a single, passive abstraction: the stateless tensor transform. Systems were designed to compile static compute graphs, optimize general matrix multiplication (GEMM) kernels, saturate high-bandwidth memory (HBM) buses, and scale distributed clusters across thousands of accelerators. The culmination of this era was the conversational foundation model: an answering system capable of synthesizing prose, explaining algorithms, and analyzing complex code repositories behind a stateless remote procedure call (RPC) endpoint.
Throughout this conversational era, the software system remained purely advisory. A human developer submitted a prompt; the serving cluster executed an autoregressive forward pass across static weights; the inference engine streamed back tokens; and intermediate activation memory was reclaimed. The model explained, summarized, and suggested, but it possessed zero ambient authority over the host operating system. If the model proposed a bug fix or an infrastructure patch, the human operator remained the sole executive actor: reading the advice, manually navigating the directory tree, applying edits, running compilers, inspecting test outputs, and committing changes. The software system remained passive because the loop was open.
The agentic systems moment occurs the instant we close this loop. Rather than treating the model as an advisory assistant that emits suggestions for a human to interpret, an agentic system delegates operational execution directly to the machine. The system is assigned a high-level operational objective—such as diagnosing a regression in a distributed key-value store’s configuration parser, resolving a benchmarked defect in SWE-bench (Jimenez et al. 2024), or reconciling conflicting transaction ledgers across remote databases—and granted execution authority. The model is integrated with operating system interfaces: shell environments, isolated containers, database connections, and filesystem descriptors. The control flow of an operational workload is now driven by an autoregressive probability distribution sampled over discrete tokens.
This delegation immediately exposes the central paradox of agentic systems architecture: the neural model directed to execute the task possesses zero ambient authority, zero execution capability, and zero direct access to the environment it is tasked with modifying.
A foundation model evaluated in isolation is an unprivileged statistical function running on matrix accelerators. It evaluates integer token vectors and emits candidate probability distributions over a fixed vocabulary into an output buffer. The model cannot issue POSIX system calls (fork, execve, pipe, ioctl), open TCP sockets, read physical hardware clocks, or inspect local storage. It possesses no program counter, no hardware registers, and no memory protection rings. Emitting an unprivileged token sequence that resembles a shell command modifies zero bytes of host state; it remains an inert stream of candidate characters inside an accelerator memory buffer.
Zero ambient authority: A foundational security stance where an execution component possesses no implicit privileges or ambient rights to access system resources; every capability to inspect or mutate host state must be explicitly granted, mediated, and monitored by the host supervisor.
This physical boundary enforces zero ambient authority. In classical operating systems, unprivileged user-space processes cannot execute privileged CPU instructions or write directly to physical disk sectors; they must issue system calls mediated by the kernel. In an agentic system, the boundary is even stricter: the model cannot even issue a system call. The neural core cannot alter a single bit of host state without an external supervisor parsing its text, validating its syntax, checking its permissions, and executing the action on its behalf.
Because the model lacks native execution authority, an agent cannot simply run its own output. If an architecture attempts to execute a multi-step operational task by generating an entire sequence of actions open-loop—emitting a monolithic script and piping it directly to an execution shell without feedback—the system hits the open-loop reliability ceiling. Consider a multi-step task requiring \(N\) sequential mutations. Under an illustrative simplifying assumption of independent steps with an equal conditional success probability \(p = 1 - \epsilon = 0.95\), the probability of completing a twenty-step trajectory without divergence decays exponentially:
\[P(\text{success}) \le (1 - \epsilon)^N = (0.95)^{20} \approx 0.358\]
Without feedback, nearly two-thirds of unguided trajectories fail. In practice, step failures are often correlated, but the systems implication remains identical: open-loop generation cannot detect environment feedback, perceive transient timeouts, or recover from unexpected tool failures.
Furthermore, statistical models exhibit an insidious epistemic failure mode: the fail-plausible fault model. When a classical software component fails, it typically exhibits fail-stop behavior—halting execution with a hardware trap, a segmentation fault, or an unhandled exception that alerts the supervisor. In contrast, when an unmediated foundation model encounters an unexpected edge case, an unsatisfied dependency, or a broken file path, it does not crash. Instead, it emits syntactically flawless, highly confident, and semantically persuasive text that reports success while masking corrupted or broken state. The model claims the task is complete, but the underlying system remains broken.
Transforming unprivileged token emission into dependable operational execution requires closing the loop through continuous runtime mediation. In an accountable agentic architecture, every candidate action proposed by the model is intercepted by a deterministic host operating system runtime. The runtime verifies schemas, evaluates capability boundaries, and executes the proposed mutation inside an isolated execution sandbox:
\[\text{Task Goal } g, \text{ Context } x_t \xrightarrow{\text{Propose}} \text{Candidate Action } a_t \xrightarrow{\text{Runtime Gate}} \text{Sandboxed Execution } e(a_t) \xrightarrow{\text{Observe}} o_{t+1} \xrightarrow{\text{Evaluate}} \text{Evidence } E\]
The lifecycle of this mediated interaction transforms speculative text into verifiable state mutations across four physical domains, as structured in figure 1. The host runtime supervisor intercepts the unprivileged candidate action proposal \(a_t \sim \pi_\theta(\cdot \mid c_t)\) emitted by the foundation model engine under zero ambient authority. Before any instruction reaches hardware or the operating system, the runtime’s authorization gate inspects schemas, token rate limits, and path bounds. Permitted mutations execute strictly inside an isolated execution sandbox, which isolates side effects and captures structured environment observations \(o_{t+1}\) (process return codes, standard streams, and file diffs). Crucially, an out-of-band verification oracle evaluates whether the mutation satisfies deterministic task invariants, emitting empirical evidence \(E\) (or \(v_{t+1}\)) that the supervisor appends alongside \(o_{t+1}\) back into the model’s context history \(c_{t+1}\), closing the supervisory loop.
The execution harness captures physical side effects—standard output streams, standard error logs, and POSIX exit codes—and appends these empirical observations \(o_{t+1}\) back into the context memory. The model then observes the empirical consequences of its prior proposal and formulates an iterative repair.
Critically, the runtime cannot rely on the model’s verbal claims to determine whether an operation succeeded. In 1970, Edsger W. Dijkstra formulated a fundamental axiom of computing: “Program testing can be used to show the presence of bugs, but never to show their absence” (Dijkstra 1970). In an agentic architecture, this principle governs the verification boundary: while a passing unit test (exit code 0) provides verifiable empirical evidence that specific test assertions held under isolated conditions, it does not prove global correctness. The host runtime must distinguish observed empirical evidence from unverified verbal claims of completion, enforcing deterministic invariant closure before any change is accepted or committed.
This operational reality delivers the foundational thesis of this book, first formulated by Vijay Janapa Reddi: Agency is an architectural property of the Stochastic Computer, not an emergent property of the neural model.
An inference engine evaluated in isolation cannot complete an operational task without a host supervisor to schedule its execution turns, a context memory hierarchy to stage working state, an isolated sandbox to contain side effects, and deterministic verifiers to validate deliverables. Everything that transforms a non-deterministic token predictor into a dependable, recoverable, and economically viable autonomous system must be engineered into the software and runtime architecture built around it.
John von Neumann (1945) organized the stored-program computer as a small set of organs (arithmetic, control, memory, input, and output), each responsible for one part of the job. This book organizes the accountable machine built around a model in the same way and calls it the Stochastic Computer. Its subsystems decide which part of the book owns which problem. The model call is studied as the machine’s processor in Part I, the state a task carries as its memory in Part II, the boundary where a proposal becomes an effect as its I/O in Part III, and the supervisor that orders, records, and recovers a trajectory as its operating system in Part IV. Two further subsystems work across tasks rather than within one, a policy compiler that turns verified trajectories into better weights (Part V) and a fleet that runs many trajectories on shared hardware (Part VI). The names locate each problem in the machine. They do not claim that a model is a processor or that a context window is a cache, and each chapter argues its subsystem in that subsystem’s own terms. Section 9 draws the full blueprint.
Recognizing that agency belongs to the surrounding system rather than to the weights in isolation raises an immediate architectural question: how did machine learning systems evolve from optimizing isolated tensor operations on single accelerators into distributed operating environments that govern stateful, multi-hour execution trajectories? To answer that question, we must trace the historical shifts in the fundamental units of machine learning systems management across thirteen orders of temporal magnitude.
From Tensors to Trajectories
When an operating system crashes during a matrix multiplication, the failure is clean, localized, and measured in microseconds: a memory controller raises a bus error, an illegal instruction fault traps into the kernel, or the hardware watch-dog timer declares a hung compute unit. The fault boundary matches the physical hardware boundary. In stark contrast, when an autonomous software-engineering agent attempts to refactor a distributed microservice, it may run unhindered across forty distinct tool invocations, exhaust thousands of seconds of wall-clock time, mutate hundreds of lines of code, and execute hundreds of containerized unit tests before silently abandoning an essential race-condition check. The system completed every intermediate remote procedure call with an HTTP 200 status, yet the overarching computational task failed.
Machine learning systems have evolved across three distinct architectural epochs: single-node static tensor graphs, distributed token-serving clusters, and stateful multi-turn trajectories. With each transition, the fundamental unit of systems management has expanded from fixed algebraic operations to open-ended operational lifecycles, ultimately stretching execution time across thirteen orders of temporal magnitude. Systems engineers can no longer treat the machine learning runtime as a passive, stateless function evaluator. Engineering autonomous systems requires building an active, stateful host supervisor capable of managing extended execution horizons, isolating irreversible environmental side effects, and enforcing correctness across non-deterministic computational paths.
The three epochs of machine learning systems
The systems challenges of modern artificial intelligence reflect a historical shift in the fundamental unit of execution that the runtime must schedule, allocate memory for, and protect. Systems engineers have navigated two foundational transitions over the past two decades, and are now engaged in a third. As mapped in figure 2, machine learning infrastructure has evolved across three architectural epochs, each defined by distinct execution units, memory models, control flows, and hardware bottlenecks:
- Era 1 (Static Tensors): The execution unit is an immutable directed acyclic graph (\(X \to Y\)) with statically pre-allocated accelerator memory buffers, compute-bound dense GEMM kernels, and deterministic execution bounded to milliseconds on single nodes.
- Era 2 (Distributed Tokens): The execution unit shifts to an iterative token sequence (\(t_k\)), where dynamic autoregression decouples prompt lengths from output lengths, requiring dynamic PagedAttention Key-Value (KV) caches across distributed clusters, bounded by high-bandwidth memory (HBM) bandwidth.
- Era 3 (Closed-Loop Trajectories): The execution unit expands into an open-ended multi-turn trajectory loop (\(\tau\)), interleaving stochastic model calls with deterministic tool RPCs, isolated sandbox execution, and out-of-band verifiers across hours of wall-clock time.
The first epoch centered on Single-Node Tensors and Static Graphs. In this regime, typified by early multilayer perceptrons, deep convolutional networks, and fixed-length sequence-to-sequence encoders, the primary workload consisted of deterministic feedforward passes over static multidimensional arrays. The compiler constructed an immutable directed acyclic graph (DAG) of tensor operations prior to execution. Memory management was predominantly static: frameworks allocated physical device memory at initialization, avoiding runtime heap churn through pre-computed buffer reuse. Execution was compute-bound, dominated by large, dense General Matrix Multiplications (GEMMs) and convolutions where arithmetic intensity was high enough to saturate floating-point hardware units. The systems boundary began and ended within a single accelerator or a tightly coupled multi-GPU node executing synchronized data-parallel steps. The duration of an invocation was strictly bounded: a forward pass through a 100-layer residual network took tens of milliseconds and consumed an invariant number of arithmetic operations.
Trajectory (\(\tau\)): An ordered sequence of interleaved state representations, candidate actions, external tool observations, and verifier verdicts: \[\tau = (s_0, a_0, o_1, s_1, a_1, o_2, \dots, s_T)\] representing the complete lifecycle of a delegated task across an unprivileged foundation engine and its supervising host environment.
The second epoch emerged with the rise of autoregressive foundation models, marking the transition to Distributed Tokens and Dynamic Attention. Here, the fundamental unit of execution shifted from a static tensor graph to an iterative token-generation loop. The Transformer architecture decoupled input length from output length, introducing dynamic execution paths governed by data-dependent stopping criteria (such as the emission of an end-of-sequence token). Systems engineers were suddenly confronted with two radically distinct computational phases within a single request: a compute-bound prompt prefill phase dominated by parallel GEMM operations, followed by an autoregressive token decode phase dominated by memory-bandwidth-bound General Matrix-Vector (GEMV) operations.
Because each generated token depends on the accumulated Key and Value states of all preceding tokens, the runtime had to dynamically allocate and track a growing Key-Value (KV) cache. Serving these models at scale required abandoning static allocation in favor of dynamic memory management, continuous request-level batching, and coordinating model weights across distributed clusters using tensor, pipeline, and sequence parallelism over high-speed interconnects. Yet, despite this massive increase in operational complexity, the system boundary remained fundamentally passive: an external client issued an HTTP request containing a context payload, the distributed cluster generated an autoregressive token sequence, and the connection terminated. The serving cluster retained no state and assumed no responsibility for what the client did with those tokens.
As summarized across paradigms in table 1, the third epoch, defining the present architectural frontier, is the era of Stateful Trajectories and Operational Effects. The fundamental systems management unit is no longer an isolated tensor operation or an autoregressive token burst; it is the entire trajectory (\(\tau\)). In this regime, an autonomous agent orchestrates a prolonged, multi-turn sequence of model invocations, tool actuations, environmental observations, and verifier checks to accomplish an ambiguous, open-ended goal. The system interacts directly with the external world: it checks out git branches, executes untrusted compiler toolchains, queries internal documentation indexes, creates and tears down cloud infrastructure, and inspects test output.
| Systems Dimension | Era 1: Static Tensors | Era 2: Distributed Tokens | Era 3: Stateful Trajectories |
|---|---|---|---|
| Primary Execution Unit | Static Tensor Graph (\(X \to Y\)) | Autoregressive Token Sequence (\(t_k\)) | Stateful Trajectory Loop (\(\tau\)) |
| Dominant Kernel Type | Compute-bound dense GEMM | Mixed Prefill GEMM & Decode GEMV | Heterogeneous RPCs, I/O, & GEMM/GEMV |
| Primary Hardware Bottleneck | FLOP saturation & thermal limits | Memory bus bandwidth (HBM) | Host coordination, I/O latency, sandbox jitter |
| Working Memory Primitive | Pre-allocated tensor buffers | Dynamically managed KV cache pages | Context window frames, disk state, & sandboxes |
| Execution Duration | Milliseconds (\(10^{-3}\text{ s}\)) | Seconds (\(10^{-1} - 10^1\text{ s}\)) | Minutes to Hours (\(10^2 - 10^4\text{ s}\)) |
| Failure Manifestation | Hardware trap or NaN propagation | Degenerate repetition or token truncation | Silent semantic regression or infinite tool loops |
| System Boundary | Single-accelerator device runtime | Distributed inference serving cluster | Host supervisor, OS sandboxes, & model API |
In this third era, the dominant systems bottlenecks shift from hardware floating-point saturation and high-bandwidth memory bus saturation to host coordination overhead, network RPC latency, process virtualization cost, and semantic state drift. The fault model changes completely. In Era 1, a failure was signaled by an operating system signal or floating-point overflow. In Era 2, a failure manifested as an out-of-memory exception during KV cache allocation or a truncated sequence limit. In Era 3, the runtime faces the challenge of timeout ambiguity—distinguishing whether a subprocess is frozen, a build tool is downloading dependencies, or a model is trapped in an unproductive reasoning loop—while verifying that candidate side effects have not permanently corrupted the host operating system.
Temporal stretching: From nanosecond opcodes to kilosecond trajectories
The transition across these three eras represents more than a qualitative shift in software design; it reflects an extraordinary physical expansion in the time scales that system runtimes must govern. Computer systems have historically evolved by constructing hierarchies of abstractions that bridge vast temporal disparities.
When David Wheeler introduced the closed subroutine for the EDSAC in 1949, he provided a mechanism to treat an arbitrary sequence of microsecond-scale instruction executions as a single reusable computational unit. Operating systems subsequently unified nanosecond hardware opcodes with microsecond context switches and millisecond disk access into the coherent abstraction of a process. In agentic systems, systems engineering must manage execution units that stretch across thirteen orders of temporal magnitude.
At the base of this execution hierarchy sits the elementary hardware instruction: a 64-bit integer addition or a 16-bit tensor core fused multiply-add (FMA) executed within approximately \(10^{-9}\text{ seconds}\) (one nanosecond). Six orders of magnitude higher, at \(10^{-3}\text{ seconds}\) (one millisecond), an operating system manages thread migrations, filesystem accesses, or intra-data center RPC hops. Another two orders of magnitude higher, at \(10^{-1}\text{ seconds}\) (one hundred milliseconds), a modern foundation model serving engine finishes prefilling context or emits a small burst of decoded tokens.
The trajectory of an autonomous agent operating on real-world systems, however, resides between \(10^1\) and \(10^4\text{ seconds}\)—from tens of seconds for a simple code modification to multiple hours for an end-to-end repository migration, security vulnerability audit, or automated distributed systems reproduction.
This expansion across thirteen orders of magnitude invalidates traditional assumptions regarding transaction atomicity, transient fault rates, and memory consistency. A systems designer can comfortably treat a 100-millisecond REST request as an atomic, fail-stop unit: if a network drop or node reboot occurs, the client simply retries the request from scratch. But when an execution unit spans 4,000 seconds, involves thirty-five distinct stateful tool invocations, and consumes tens of millions of speculative tokens, treating the entire unit as an atomic, all-or-nothing transaction is economically and computationally ruinous.
Napkin Math 0.1: Temporal expansion and error compounding across agent horizons
- A foundation engine invocation that processes an average working context of \(32\text{ KiB}\) (\(8\text{k tokens}\)) and generates an action proposal of 256 tokens (\(\text{latency } t_{\text{model}} \approx\) 4.5 s).
- An actuation turn inside an isolated Linux container, executing a target test target via
pytestor compiling a module viacargo build(\(\text{latency } t_{\text{tool}} \approx\) 18 s). - Host supervisor overhead for sandboxed file I/O, output stream parsing, and invariant verification (\(\text{latency } t_{\text{host}} \approx\) 0.5 s).
The total per-step latency is: \[t_{\text{step}} = t_{\text{model}} + t_{\text{tool}} + t_{\text{host}} = 4.5 + 18.0 + 0.5 = 23.0\text{ seconds}\] The total wall-clock execution time for the full trajectory \(\tau\) is: \[T_{\text{total}} = N \times t_{\text{step}} = 45 \times 23.0\text{ s} = 1{,}035\text{ seconds} \approx 17.25\text{ minutes}\]
Now consider the reliability boundary. Assume the foundation model possesses a per-step action selection accuracy of \(p = 0.96\) (that is, the model emits an actionable, syntactically and semantically viable tool call 96 percent of the time without entering a hallucinated or degenerative state).
In an open-loop architecture where the runtime simply passes the model’s generated sequence forward without verification or automated error remediation, the probability of the entire trajectory executing successfully without a fatal failure is: \[P(\text{success})_{\text{open-loop}} = p^N = (0.96)^{45} \approx 0.159 \quad (15.9\%)\]
Even with an extraordinarily capable model (\(p = 0.96\)), the open-loop completion rate collapses to less than one-in-six. To achieve an acceptable system-level completion rate (\(P_{\text{target}} \ge 0.90\)) across this 45-step horizon, an open-loop model would require a per-step fidelity of: \[p = (P_{\text{target}})^{1/N} = (0.90)^{1/45} \approx 0.99766 \quad (99.77\%)\] Demanding 99.8 percent unguided step accuracy from a stochastic language model over an ambiguous operational task is practically infeasible. Reliability cannot be solved at the token level; it must be engineered at the system level through closed-loop verification, state snapshotting, and transaction recovery.
Systems spanning thousands of seconds must be designed with the explicit understanding that underlying infrastructure will exhibit transient network disconnects, tool sandboxes will exhaust memory quotas, external APIs will rate-limit requests, and the neural engine will occasionally emit invalid tool syntax. The runtime must treat the trajectory not as an uninterrupted stream of execution, but as a directed graph of checkpointed, recoverable, and observable state transitions.
The mathematical necessity of runtime supervision is visualized in figure 3. In an open-loop execution regime, task success probability collapses exponentially with trajectory horizon (\(P_{\text{task}} = p^N\)). Even when using state-of-the-art foundation models with high single-step fidelity (\(p = 0.95\), \(0.98\), or \(0.99\)), unguided trajectories inevitably hit the compounding reliability wall within tens of steps. Closed-loop supervision halts this geometric decay: by placing deterministic verification oracles after each tool action, intercepting syntactic and semantic errors before state commitment, and rolling back unverified mutations to valid checkpoints, the host runtime stabilizes task success probability across extended operational horizons.
This shift in execution horizon mirrors the broader history of computer systems, as illustrated in figure 4. Over the past seven decades, the basic unit of computational scheduling has expanded across thirteen orders of temporal magnitude: from hardware instruction cycles executed by microprocessors in sub-nanoseconds (\(10^{-9}\text{ s}\)), through subroutines (\(10^{-6}\text{ s}\)), operating system processes (\(10^{-3}\text{ s}\)), distributed RPCs (\(10^0\text{ s}\)), and container workflows (\(10^2\text{ s}\)), to multi-turn autonomous trajectories (\(10^3 - 10^4\text{ s}\)). Each step up this ladder required new operating system primitives—call stacks, memory protection rings, virtual memory page tables, and distributed consensus protocols. At the trajectory scale, classical OS primitives fail because execution is non-deterministic and tool side effects can permanently alter external systems of record, demanding stateful supervision and transactional rollback.
The passive request boundary
Why did the classical serving architectures of Era 2 fail to scale directly to autonomous tasks? The barrier is rooted in the passive request boundary.
In conventional machine learning infrastructure, model serving is structured around an RPC interface modeled after traditional stateless microservices. A client submits a payload containing an input prompt and generation hyperparameters (\(T_{\max}\), temperature, stop tokens). The serving engine parses the input, constructs attention tensors, balances dynamic KV cache memory, executes continuous batching loops, and streams tokens back across the wire. Crucially, the serving engine treats every request as a self-contained computational island. Once the final token is returned or the client disconnects, the engine drops all request state from fast memory, retaining no memory of prior interactions beyond generic access logs.
This passive boundary operates on an open-loop model. The inference engine possesses no actuation mechanisms: it cannot execute a bash command, inspect the return code of an operating system process, read an updated file from disk, or verify that a proposed code patch compiles. The entire responsibility for interpreting, applying, and reacting to model outputs is pushed outward to the client.
When human beings act as the external client—such as in interactive chat sessions—the human provides the missing closed-loop controller. If the model emits code that fails to compile, the developer copies the compiler error back into the prompt, prompting the model to try again. The human developer acts as the operational runtime: managing context, executing tools in their local terminal, inspecting error traces, rolling back broken edits, and verifying invariants.
This structural contrast between manual and programmatic control loops is illustrated in figure 5, where panel (a) keeps a human operator inside the loop and panel (b) closes it programmatically. In traditional model serving (Era 2, top panel), the serving cluster operates entirely open-loop: it generates tokens in response to an RPC payload and terminates its connection, pushing all state management, tool execution, and error diagnosis onto the human operator. In an autonomous agent architecture (Era 3, bottom panel), the host runtime supervisor assumes programmatic responsibility for the entire control loop. The supervisor manages working context, intercepts candidate action proposals \(a_t\) under zero ambient authority, dispatches commands to an isolated execution sandbox, and evaluates deterministic verification evidence \(v_{t+1}\) before committing environmental state transitions.
When systems engineers attempt to automate this process by delegating tasks entirely to machines, they hit the open-loop systems ceiling. Without a closed-loop supervisor, the probability of successful trajectory completion decays exponentially with each additional step, as demonstrated in the worked example above. Errors in model output compound rapidly: a minor misinterpretation of an API argument in step 3 causes a missing file in step 7, which induces an unhandled exception in step 12, resulting in an unrecoverable hallucination loop by step 15.
# The Open-Loop Systems Ceiling: Compounding Trajectory Failure
# A runtime without closed-loop observation cannot survive non-zero error rates.
def evaluate_open_loop_trajectory(num_steps: int, step_fidelity: float) -> float:
trajectory_success_prob = 1.0
for step in range(1, num_steps + 1):
trajectory_success_prob *= step_fidelity
if trajectory_success_prob < 0.50:
# Trajectory has become more likely to fail than succeed
return trajectory_success_prob
return trajectory_success_prob
# For a 30-step task with a highly capable 97% reliable step execution:
# evaluate_open_loop_trajectory(30, 0.97) yields 0.399 (39.9% success rate).To break through this open-loop ceiling, the machine learning system must be re-architected. The system must close the loop: candidate model outputs cannot be treated as final deliverables, but as untrusted, unprivileged proposals held in escrow.
The runtime must execute these proposals inside sandboxed execution environments, capture structured observations from deterministic tools (exit codes, standard error streams, AST diffs, test logs), re-inject those observations into the model’s evolving context memory, and use deterministic verifiers to confirm invariant closure.
This closed-loop trajectory transforms the fundamental contract of computing. How does this shift alter the nature of software itself, and how do we formalize the relationship between deterministic code and stochastic neural inference? Answering this requires examining the evolution of software paradigms and the tripartite architecture that bridges traditional programming with agentic control.
Software Paradigm Evolution
When a compiler stalls on an unresolved dependency, a POSIX thread yields its execution context to the operating system scheduler, freeing its core and registers for active work. When a stochastic neural policy running across a cluster of graphics processing units emits a candidate command and blocks synchronously on a sixty-second test suite run, the underlying physical system cannot gracefully yield. The graphics memory bus remains pinned, locking dozens of gigabytes of high-bandwidth memory per accelerator to preserve intermediate attention states for a context that is not computing. The server idles at peak thermal design power while waiting on an external disk read, stranding tens of thousands of dollars of silicon on an unmediated I/O boundary. This physical friction exposes a fundamental architectural mismatch: classical software frameworks assume cheap, suspendable execution threads, whereas neural inference engines require dedicated, continuous high-bandwidth memory access.
Agentic machine learning systems represent Software 3.0: a hybrid computing paradigm where stochastic neural policies act as high-level controllers governing deterministic Software 1.0 effectors and operating system primitives. Software 3.0 does not supersede or discard classical programming. Instead, it embeds the continuous, probabilistic pattern matching of neural networks inside a deterministic runtime harness designed to manage asynchronous tool execution, durable state, and invariant verification. To build reliable systems at this boundary, an engineer must understand how the fundamental abstractions of computing—control flow, state retention, fault models, and hardware bottlenecks—have mutated across three distinct programming eras.
Orthogonal Systems Axes: The three eras of MLSys (§1.2) trace the internal evolution of machine learning workloads, where Era 1 (static graphs) and Era 2 (distributed token serving) both belong to Software 2.0. Software 3.0 represents the architectural hybrid combining Software 1.0 execution substrates with Software 2.0 statistical inference into closed-loop trajectories.
The tripartite architecture
In classical computing, which we designate Software 1.0, human engineers construct explicit algorithms by sequencing discrete instructions. The programmer defines the state space, the valid state transitions, and the branching logic required to handle exceptional conditions. Execution is deterministic: given an identical initial state and the same sequence of inputs, an identical machine-code binary running on an abstract von Neumann machine will traverse an identical sequence of program counter values, stack frames, and register states. Correctness is enforced ahead of time through static type checkers, formal proofs, and compiler analysis, or at runtime via deterministic assertions and boundary checks. Software 1.0 excels at structured arithmetic, relational data manipulation, deterministic protocol parsing, and hardware control, where requirements can be formalized into rigid symbolic invariants.
Software 2.0 replaces human-authored instruction sequences with continuous numerical optimization over high-dimensional parameter spaces. Rather than hand-crafting algorithms, engineers collect datasets of input-output pairs, define a continuous objective function, and employ stochastic gradient descent to locate a parameter configuration within a neural network that minimizes empirical risk. At inference time, the resulting model operates as a static, feedforward computational graph. A batch of input tensors flows through fixed linear algebra operations—matrix multiplications, non-linear activations, and layer normalizations—yielding an output prediction in constant time proportional to the graph depth. Control flow is frozen within the compiled kernel graph; there are no dynamic branches, no external side effects, and no interaction with the surrounding operating system. The program state resides entirely in ephemeral activation tensors that vanish as soon as the forward pass completes. Software 2.0 conquered perceptual domains that defied manual codification, such as computer vision, speech recognition, and syntactic natural language modeling, but it did so by relinquishing external agency, persistent environmental state, and formal guarantees of correctness.
Software 3.0 synthesizes these two paradigms into stateful, interactive trajectories. In this operational model, the core model is a pretrained, stochastic neural policy (Software 2.0), but its outputs are not terminal answers delivered directly to a human user. Instead, the policy emits structured action proposals directed at deterministic effectors (Software 1.0): shell commands, database queries, web requests, abstract syntax tree refactoring scripts, and test runners. The execution environment consumes these action proposals, executes them under strict operating system sandboxing, and captures structured observations—process return codes, standard output streams, compiler error diagnostics, and filesystem diffs.
The architecture governing this synthesis is formalized in figure 6. The host runtime supervisor orchestrates a bidirectional mediation loop between the unprivileged stochastic neural policy and isolated deterministic tools, supported by a four-tier memory and storage hierarchy:
- Tier 1 (Working KV Cache in Accelerator HBM): Caches high-speed attention activations (\(M_{\text{KV}}\)) to minimize token decode latency across iterative turns.
- Tier 2 (Host Context Staging in CPU DRAM): Manages the tokenized trajectory history \(c_t\), performing dynamic window truncation, scratchpad formatting, and tool schema marshalling.
- Tier 3 (Write-Ahead Log on Durable NVMe): Immutably records every state transition, candidate proposal, and verification verdict prior to execution, providing crash consistency and rollback capabilities.
- Tier 4 (Persistent Environment State on Host Filesystem and Git): Maintains the authoritative ground-truth state of the external workspace (\(\mathcal{S}_{\text{sys}}\)), including file inodes, container layers, and version-controlled repositories.
Software 3.0 does not attempt to replace compilers or operating systems with neural weights; such an endeavor is an architectural category error. A transformer cannot reliably emulate an IEEE 754 floating-point unit, execute a B-tree rebalance, or compute a cryptographic hash over millions of bytes without wasting billions of parameter operations on tasks that a five-dollar microprocessor executes in a single cycle. Rather, Software 3.0 places the probabilistic model where human judgment previously sat: at the supervisory layer of the control loop, formulating hypotheses, diagnosing anomalous traces, and directing classical tools to manipulate the physical state of the machine.
Eight systems dimensions of software execution
The shift from explicit instructions to continuous weights, and ultimately to stateful trajectories, alters every layer of the computing stack. An engineer designing runtime infrastructure for agentic workflows must evaluate how traditional systems assumptions dissolve when probabilistic controllers govern deterministic software effectors.
| Systems Dimension | Software 1.0 (Classical Code) | Software 2.0 (Neural Weights) | Software 3.0 (Agentic Trajectories) |
|---|---|---|---|
| Unit of Work | Machine instruction or function call | Batched tensor pass (GEMM / GEMV) | Multi-turn environmental trajectory (\(\tau\)) |
| Execution State | Registers, stack frames, heap pointers | Ephemeral intermediate activation tensors | Context window, KV cache, sandbox filesystem |
| Control Flow | Deterministic branching (if, switch, loops) |
Static, compiled acyclic dataflow graphs | Stochastic policy sampling over discrete tool APIs |
| Dominant Failure Mode | Fail-stop crashes (segfaults, uncaught panics) | Distributional drift, out-of-domain degradation | Fail-plausible semantic corruption (code 0 exits) |
| Fault Recovery | Process restart, checkpointing, stack unwind | Model retraining, checkpoint rollback, fine-tuning | Saga compensation, event-sourced rollback, retry |
| Hardware Bottleneck | CPU-DRAM memory bus latency and cache misses | Memory bandwidth (decode) and compute (prefill) | Accelerator HBM capacity and Tool-Wait stranding |
| Side Effects | POSIX syscalls, process I/O, storage mutation | None (pure mathematical functional transform) | Irreversible environmental and remote API mutations |
| Correctness Guarantees | Formal proofs, static types, deterministic tests | Empirical risk bounds, statistical generalization | Runtime verification enclaves and invariant monitors |
The architectural divergence captured in table 2 highlights three critical transformations that govern systems design in Software 3.0:
First, consider the transition in dominant failure modes. Software 1.0 exhibits predominantly fail-stop behavior: an out-of-bounds pointer dereference triggers a hardware exception (SIGSEGV), a null reference halts an interpreter with a stack trace, and a syntax violation prevents compilation entirely. The system crashes cleanly, halting execution before corrupted state can propagate widely. Software 2.0 fails silently through statistical degradation; an image classifier confronted with an out-of-distribution input does not crash, but shifts its softmax output distribution toward lower-entropy misclassifications.
Software 3.0 introduces an insidious operational fault mode: fail-plausible semantic corruption. When an autonomous agent attempts to fix a failing test suite, the underlying neural policy frequently generates syntactically valid code that satisfies the immediate regex checks of a naive validation script, or worse, deletes the asserting test cases altogether. The execution harness observes a clean return code of zero (exit 0), logs a successful tool execution, and continues its trajectory, unaware that the agent has satisfied the literal prompt while completely corrupting the underlying engineering invariant. Classical fault detection mechanisms that monitor process termination codes or unhandled exceptions are utterly blind to this failure class.
Second, the mechanism of fault recovery must evolve to handle irreversible environmental mutations. In Software 1.0, restoring a corrupted process requires resetting the instruction pointer, unwinding the call stack, or terminating the process container and restarting from a clean image. In Software 2.0, fixing inference failures requires retraining the weight matrix on expanded datasets or adjusting temperature hyper-parameters.
Neither strategy functions in Software 3.0. Once an agentic trajectory emits an external mutation—such as dropping a database table, executing an HTTP DELETE request against a production API endpoint, or modifying files on a persistent volume—the state of the external world has diverged. The runtime cannot simply clear its context window and restart the trajectory, because the sandbox filesystem and external microservices now contain mutated state. Fault recovery in Software 3.0 mandates distributed systems primitives: event-sourced write-ahead logging of every action-observation pair, transactional staging directories, and Saga-style compensating actions capable of undoing environmental side effects when a downstream verification step fails.
Third, the nature of execution state shifts from ephemeral hardware buffers to distributed, hybrid memory hierarchies. In Software 1.0, state is managed by the hardware memory management unit, cache coherence controllers, and virtual memory page tables. In Software 2.0, state is transient: activations exist for microseconds in accelerator SRAM and high-bandwidth memory (HBM) during the forward pass, after which only the static weights persist.
In Software 3.0, execution state spans three heterogeneous domains: the neural policy’s attention key-value (KV) cache allocated inside expensive accelerator HBM, the structured textual transcript stored in runtime process memory, and the mutated filesystem state residing on persistent disk inside an isolated container sandbox. Coordinating these three disparate representations of state requires explicit systems orchestration. If the agent runtime modifies the sandbox filesystem by checking out an alternative git branch, but fails to synchronize the token sequence in the context window, the neural policy experiences an epistemic fracture, conditioning downstream generation on a transcript that no longer reflects the physical reality of the sandbox.
Checkpoint 0.1: Retraining versus runtime resilience
Before examining tool-wait memory friction, verify your understanding of failure boundaries across software paradigms:
Memory stranding friction
The most severe physical dilemma in Software 3.0 architectures arises from the latency asymmetry between GPU tensor computation and operating system tool actuation. Modern hardware accelerators are engineered for massive, dense matrix arithmetic executed across thousands of SIMD lanes, fed by high-bandwidth memory buses offering terabytes per second of throughput. Conversely, Software 1.0 effectors—compilers, unit test runners, web scrapers, and database engines—are I/O-bound, branch-heavy workloads governed by operating system scheduling, disk access, and network socket latency.
When a multi-turn agent executes an action, the runtime invokes an external tool. If the host agent runtime implements this interaction via synchronous remote procedure calls (RPCs) while maintaining the underlying inference session, the model’s key-value cache remains pinned inside the accelerator’s HBM. The hardware cannot reallocate those memory pages to serve other incoming requests without losing the activation state that represents the agent’s working memory.
Napkin Math 0.2: The tool-wait memory tax
Variables:
- Cluster Hardware: 8-GPU NVIDIA H100 SXM5 node (\(640\text{ GB}\) HBM3, $24/hour or $0.00667/s).
- Model Partitioning: 80B dense parameters partitioned across \(TP =\) 8. Static weights consume 160 GB, leaving 480 GB dynamic HBM (60 GB per GPU).
- Context State: Active trajectory context of \(T =\) 160,000 tokens. An 80B-parameter model requires \(m_{\text{token}} = 320\text{ KiB}\) of Key-Value state per token (The Size of the Attention State derives this tensor geometry).
- Tool Duration: A long-running compilation and integration test:
cmake --build out/ && ctest --timeout 60requiring \(t_{\text{wait}} =\) 60 s.
Math: The physical memory required to hold this single session’s KV cache is: \[M_{\text{KV}} = 160{,}000 \times 320\text{ KiB} = 51{,}200{,}000\text{ KiB} \approx 51.2\text{ GB}\] If the agent runtime executes this tool call synchronously:
- Memory Stranding: The \(51.2\text{ GB}\) KV cache is pinned for the entire 60 s. This represents \(51.2 / 480 =\) 10.67 percent of the entire dynamic serving capacity of the node. If 4 concurrent agents trigger compilation tasks simultaneously, 204.8 GB of dynamic memory—fully 42.67 percent of the node’s interactive capacity—is completely stranded, unable to process tokens for active users.
- Financial Cost: Over the 60 s compilation wait interval: \[\text{Cost} = 60\text{ s} \times \$0.00667/\text{s} = \$0.400\] The system incurs $0.40 of direct compute expenditure per turn while performing zero floating-point operations. In an industrial pipeline executing 20 tool calls per resolution trajectory, tool-wait stranding burns $8 of capital entirely on idle silicon.
Systems insight: Software 3.0 runtimes cannot treat tool invocations as standard blocking function calls. The host supervisor must decouple tool execution asynchronously: serializing token state, releasing accelerator memory allocations, recording progress in a Write-Ahead Log (Durable Execution), and leveraging prefix caching (KV Cache Management) when the observation returns.
This quantitative reality forces Software 3.0 runtimes to abandon synchronous execution models. An agent runtime cannot treat tool invocations as standard blocking function calls. Instead, the runtime architecture must decouple inference from execution across three distinct systems mechanisms. First, through asynchronous context eviction, when a long-running tool command is dispatched, the runtime serializes the logical token state, releases the physical memory allocated to the session’s attention caches inside the inference engine, and marks the accelerator capacity as reclaimable for concurrent serving workloads. Second, through prefix caching and re-hydration, when the tool execution completes and returns an observation, the runtime reschedules the session, leveraging preserved prompt prefixes to reconstruct intermediate attention states without re-evaluating earlier sequence tokens from scratch. Third, through speculative local verification, fast, deterministic checks such as static linters and abstract syntax tree validators run directly on the host controller processor to intercept syntactic failures in milliseconds, avoiding the latency and memory overhead of dispatching a full containerized test execution.
The fundamental contrast in execution flow between traditional inference and an autonomous agent loop is visualized in figure 7. In traditional model serving (top bar), inference runs as an uninterrupted burst of continuous token generation where accelerator tensor cores achieve near-100 percent utilization. In contrast, an autonomous agent trajectory (bottom bar) is severely fragmented: during an 8.5-minute software engineering trajectory on SWE-bench Verified (Jimenez et al. 2024), the agent spends 420 seconds (82.4 percent of makespan) blocked on non-neural host operations—container workspace setup, filesystem search via ripgrep, and pytest suite execution. In a synchronous runtime, these tool waits keep accelerator memory pinned while the server chassis idles.
Software 3.0 is therefore defined not by the elimination of traditional systems engineering, but by its intensification. The presence of an unprivileged, stochastic policy at the core of the execution loop mandates rigorous, deterministic supervision. The runtime supervisor must arbitrate access to hardware accelerators, virtualize memory across volatile HBM and persistent host storage, enforce security boundaries around untrusted tool outputs, and verify system invariants before declaring a task complete.
The empirical consequence of this physical boundary is mapped in figure 8. Profiling an autonomous coding agent executing a benchmark task on SWE-bench Verified (Jimenez et al. 2024) across an NVIDIA H100 accelerator and an AMD EPYC host server reveals the striking magnitude of the Agent Systems Tax. As shown in the left panel of figure 8, non-neural host operations—isolated Docker workspace setup (8.1 kJ, 6.9 percent), file indexing and ripgrep queries (16.8 kJ, 14.4 percent), and pytest suite compilation and verification (56.1 kJ, 48.0 percent)—account for fully 69.3 percent of total AC mains electrical energy consumption (81.0 kJ), while pure neural token generation on accelerator tensor cores consumes only 30.7 percent (35.9 kJ). In the right panel, these tool waits and sandbox operations account for 82.4 percent of wall-clock makespan (420 seconds), while active token generation represents only 17.6 percent (90 seconds). This empirical profile mathematically confirms the central thesis: autonomous agent performance, reliability, and operating costs are dominated by host operating system orchestration, virtualization sandboxes, and verification harnesses, rather than neural model scaling alone.
Taken together, the execution timeline and resource breakdown reveal a governing principle of agentic workloads: Amdahl’s Law for Autonomous Systems. Because non-neural host operations account for 82.4 percent of total makespan and 69.3 percent of electrical energy, even reducing neural inference latency to zero through infinite accelerator scaling would yield at most a 1.21-fold overall speedup (\(1 / (1 - 0.176) \approx 1.21\)). The primary bottleneck in autonomous agent execution is not matrix multiplication throughput on GPU tensor cores; it is host operating system coordination, filesystem search latency, container sandbox virtualization, and verification oracle evaluation. Software 3.0 systems engineering must therefore shift focus from optimizing isolated token-generation kernels to disaggregating the host runtime: decoupling neural inference scheduling from host execution via cooperative yielding, snapshotting workspace filesystems with copy-on-write trees, and reclaiming pinned accelerator memory while tool processes run.
With Software 3.0 established as a distinct hybrid systems paradigm, what is its formal engineering definition? The next section formalizes the agentic machine learning system as an autonomous, stateful closed-loop control system embedded within a deterministic runtime harness.
Defining Agentic Systems
In classical machine learning infrastructure, an inference service is engineered around stateless request-response semantics, optimizing for isolated tensor throughput across accelerator arrays. A client submits a payload of tokenized input sequences, the inference engine computes the prefill attention matrices and executes an autoregressive generation loop, and the server returns a stream of output logits or decoded text before terminating the request context. When an unprivileged model policy is invoked to compile software, modify a complex multi-file repository, or query a distributed database, this stateless client-server abstraction ruptures completely. The model’s emissions are no longer passive textual deliverables consumed by a human reader; they are candidate mutations directed against a mutable external environment.
An agentic machine learning system is formally defined as an autonomous, stateful closed-loop control system embedded within a deterministic runtime harness that manages context memory, tool actuation, and invariant verification. Agency is not an internal cognitive property of a neural network, nor is it an emergent artifact of scaling parameter counts. Rather, agency is an architectural property of the complete computer system. The foundation model provides an unprivileged, stochastic policy that proposes candidate actions, while the host runtime provides the execution substrate, memory virtualization, environmental effectors, and boundary verification mechanisms required to steer an open-ended computation toward a verified, non-trivial goal state.
The formal systems definition
To transition from an ad-hoc software script to a dependable systems architecture, we must formalize the boundaries, components, and interactions that constitute an agentic system.
Definition 0.1: Agentic machine learning system
Agentic machine learning system is an autonomous, stateful closed-loop control system embedded within a deterministic runtime harness that manages context memory, tool actuation, and invariant verification: \(\mathcal{S} = \langle \pi_\theta, \mathcal{H}, \mathcal{E}, \mathcal{M}, \mathcal{V} \rangle\).
- Significance: Demotes the foundation model from an autonomous application to an unprivileged processing core within a supervisory host runtime, establishing that agency is an architectural property of the complete computer system rather than an emergent cognitive attribute of neural weights.
- Distinction: Unlike classical inference services that operate over stateless request-response sequences, an agentic system executes stateful closed loops where model token proposals mutate external state \(\mathcal{E}\), dynamically reshaping future inputs and introducing endogenous feedback loops.
- Common pitfall: Treating the system as an unconstrained API loop—a naive script repeatedly concatenating tool outputs to context strings—which lacks transaction isolation, capability masking, and bounded execution guarantees, leading to runaway compute expenditure and state corruption.
This definition enforces a fundamental systems inversion: the machine learning model is demoted from an end-to-end application down to an unprivileged processing core within a host runtime. In traditional supervised or self-supervised serving, the environment is static and read-only; the model takes an input \(x\) and predicts an output \(y\). In an agentic system, the model’s emissions mutate \(\mathcal{E}\), which in turn alters subsequent observations fed into \(\pi_\theta\). This introduces an endogenous feedback loop where any latent prediction error, invalid interface declaration, or syntax flaw directly alters the future input distribution of the model itself.
The API-Loop Antipattern Treating an agent as a naive while loop wrapping an LLM API call is the modern equivalent of writing an operating system without memory protection or interrupt handlers. Without an isolating runtime harness, context windows overflow, tool side effects compound uncontrollably, and runtime failures exit without rollback.
A widespread failure in naive agent implementations is treating the system as an unconstrained API loop—a fragile script that repeatedly concatenates tool outputs to a prompt string until the autoregressive generation loop encounters an end-of-sequence token or exhausts an iteration counter. From a systems perspective, an unconstrained API loop is a distributed control system operating without feedback dampening, transaction isolation, or bounded execution guarantees. If the unprivileged policy emits a destructive command (such as deleting an unindexed database table) or enters a repetitive error cycle (such as repeatedly invoking a failing shell command with identical flags), the naive loop burns compute budgets and corrupts environment state. Dependability requires that the host harness \(\mathcal{H}\) maintain absolute supervisory authority over \(\pi_\theta\), treating candidate model tokens as untrusted proposals held in escrow until validated.
The trajectory as the systems management boundary
In single-turn inference, the unit of systems management is the model invocation. An invocation consists of a prefill phase (processing prompt tokens via compute-bound General Matrix Multiply, or GEMM, kernels) followed by an autoregressive decode phase (generating output tokens via memory-bandwidth-bound General Matrix-Vector, or GEMV, kernels). The execution duration of an invocation is short and strictly bounded by the generation budget, spanning milliseconds to tens of seconds.
In an agentic system, the fundamental unit of systems management expands from an isolated invocation to an extended, stateful trajectory \(\tau\), as detailed in figure 9. Rather than managing ephemeral forward passes, the runtime coordinates a multi-stage execution timeline bound to an Agent Control Block (ACB). As shown along the trajectory axis, the systems profile alternates between accelerator-bound tensor compute (prompt prefill GEMMs and autoregressive decode GEMVs) and host-mediated operational phases: schema authorization, isolated sandbox tool execution, network I/O wait states, and out-of-band verification passes. At each turn boundary \(t \to t+1\), the runtime records an execution checkpoint to durable storage, ensuring that transient hardware or network faults can be recovered without restarting the multi-minute trajectory from scratch.
Definition 0.2: Agentic trajectory
Agentic trajectory is a causal chain of discrete, typed state transitions \(\tau = (s_0, a_0, o_1, s_1, a_1, o_2, \dots, s_T)\) orchestrated by the host runtime harness across context state (\(s_t \in \mathcal{S}_{\text{ctx}}\)), proposed actions (\(a_t \in \mathcal{A}\)), and authoritative host observations (\(o_{t+1} \in \mathcal{O}\)).
- Significance: Expands the fundamental unit of systems management from an isolated, ephemeral model invocation (milliseconds) to an extended, heterogeneous execution lifecycle (minutes to hours) bound to an Agent Control Block (ACB).
- Distinction: Unlike single forward passes that only alternate compute-bound prefill GEMMs and memory-bandwidth-bound decode GEMVs, a trajectory interweaves accelerator tensor operations with host operating system process management, network I/O wait states, sandboxed file modifications, and out-of-band verification passes.
- Common pitfall: Managing long-horizon trajectories using transient in-memory thread stacks or raw coroutine state rather than persistent, checkpointed control blocks, causing transient network hiccups or GPU worker preemptions to drop uncommitted progress and force expensive re-execution from scratch.
Unlike an isolated model invocation, a trajectory is highly heterogeneous. It interweaves accelerator tensor operations, host operating system process management, network round-trip latencies, file system modifications, and execution pauses while awaiting compiler completion or human approval. The wall-clock duration of a trajectory is open-ended, ranging from several seconds for a simple code-formatting task to hours or days for an autonomous multi-repository migration.
Because a trajectory spans multiple model invocations, asynchronous child processes, and mutable environmental states, the host infrastructure cannot manage execution using transient thread stacks. Instead, the runtime binds the trajectory to an authoritative session descriptor: the Agent Control Block (ACB).
The Agent Control Block (ACB) Just as an operating system kernel tracks process state, open file descriptors, and virtual memory page tables inside a Process Control Block (PCB), an agent host runtime tracks token budgets, tool permissions, and context memory descriptors inside an Agent Control Block (ACB). We develop the full ACB data structure and scheduling queue discipline in The Agent Harness.
The ACB decouples the state of the task from the lifetime of any single accelerator process or HTTP connection. If a GPU worker node crashes or an inference stream drops during step \(t\), the runtime uses the ACB to reconstruct context state from persistent Write-Ahead Logs (Durable Execution), re-establish sandboxed tool boundaries (Agent Sandboxes), and resume execution without losing intermediate progress. Furthermore, the ACB acts as an administrative boundary: enforcing hardware resource quotas (token ceilings and wall-clock duration), capability security profiles (restricting which system calls or network endpoints the agent may touch), and transactional rollback ledgers.
Micro-efficiency versus macro-efficiency
The shift from isolated model invocations to long-horizon trajectories forces a complete re-evaluation of systems performance metrics. In conventional Large Language Model serving infrastructure, engineering teams optimize almost exclusively for micro-efficiency.
Micro-efficiency metrics quantify the execution efficiency of the tensor computation core during a single model invocation. Time to First Token (TTFT) dictates prefill latency, constrained by accelerator peak dense matrix computation (TFLOPS) when processing input prompt tokens concurrently via General Matrix Multiply (GEMM) kernels. During the subsequent autoregressive decode phase, Time Per Output Token (TPOT) and raw generation throughput (tokens per second per accelerator) dominate, constrained by High Bandwidth Memory (HBM) bus speed when fetching model parameters and Key-Value (KV) cache tensors via General Matrix-Vector (GEMV) kernels. Finally, Model Bandwidth Utilization (MBU) quantifies the fraction of theoretical peak memory bandwidth achieved across memory-bound decode steps.
Modern inference engines (such as vLLM and SGLang) leverage sophisticated kernel fusions, non-contiguous block memory allocation, and prefix caching to drive micro-efficiency toward hardware limits. However, in an agentic system, high micro-efficiency is a necessary but profoundly insufficient condition for overall system performance. A serving cluster operating at 95 percent MBU and generating 4,000 tokens per second per node is entirely wasted if the agent policy is trapped in an infinite loop, generating non-existent compiler flags, or corrupting its working context with unparsed stack traces. Agentic systems engineering must therefore optimize for macro-efficiency, which measures the resource cost, latency, and dependability of achieving an accepted, verified task outcome across an entire trajectory, as contrasted across the architectural dimensions detailed in table 3.
| Architectural Dimension | Micro-Efficiency (Inference Engine Focus) | Macro-Efficiency (Agent Runtime Focus) |
|---|---|---|
| Optimization Target | Per-token generation latency and accelerator FLOP utilization | Trajectory Goodput (\(\mathcal{G}\)) and verified task completion |
| Execution Horizon | Single invocation (milliseconds to seconds) | Multi-turn trajectory (minutes to hours) |
| Primary Bottlenecks | HBM bus bandwidth, tensor core saturation, NVLink latency | Sandboxed tool latency, verification overhead, error compounding |
| Failure Indication | Low Model Bandwidth Utilization (MBU), pipeline bubbles | Cyclic execution loops, unverified token churn (badput) |
| System Boundary | Model weights and KV cache tensors | Agent Control Block, host filesystem, external environments |
Checkpoint 0.2: Throughput versus goodput
Before formalizing macro-efficiency metrics, verify your ability to distinguish accelerator execution rates from agentic utility:
To formalize macro-efficiency, we introduce the concept of trajectory goodput (\(\mathcal{G}\)). In classical data communications, goodput represents the number of useful information bits delivered to the application layer per unit of time, excluding protocol overhead and retransmitted corrupted packets. In an agentic machine learning system, trajectory goodput is the proportion of total cluster compute resources allocated to trajectories that successfully achieve an independently verified terminal state:
\[\mathcal{G} = \frac{\sum_{i \in \mathcal{T}_{\text{success}}} R_i}{\sum_{j \in \mathcal{T}_{\text{all}}} R_j}\]
where \(\mathcal{T}_{\text{all}}\) represents all dispatched trajectories, \(\mathcal{T}_{\text{success}} \subseteq \mathcal{T}_{\text{all}}\) represents the subset whose deliverables pass all deterministic verification oracles \(\mathcal{V}\), and \(R_k\) measures physical resource expenditure (floating-point operations, GPU core-seconds, kilowatt-hours, or financial cost). Unverified token generation represents computational badput (\(1 - \mathcal{G}\)). We develop the full telemetry and evaluation mechanics of trajectory goodput in Agent Evaluation and analyze its macro-economic cost frontier in Agent Economics.
Napkin Math 0.3: Trajectory goodput and compute waste
Variables:
- Workload: 100 bug repair trajectories.
- Failure Dynamics: Due to unmitigated fail-plausible errors, 65 trajectories enter cyclic loops, exhausting step budget \(T_{\max} =\) 50 (512 decode tokens and 16,384 prefill tokens per step).
- Success Dynamics: 35 trajectories pass all verification tests with an average of 25 steps.
- Power Draw: 4800 W average active cluster power.
Math:
Compute Resource Allocation:
- Successful trajectories (\(35 \times 25\text{ steps}\)): \(875\text{ total steps}\).
- Failed trajectories (\(65 \times 50\text{ steps}\)): \(3{,}250\text{ total steps}\).
Trajectory Goodput: \[\mathcal{G} = \frac{35 \times 25}{(35 \times 25) + (65 \times 50)} = \frac{875}{875 + 3{,}250} = \frac{875}{4{,}125} \approx 0.212 \quad (21.2\%)\] Despite high GEMM/GEMV kernel utilization, 78.8 percent of the cluster’s compute was badput.
Energy Waste and Financial Expenditure: Failing trajectories generated \(65 \times 50 \times 512 = 1{,}664{,}000\text{ decode tokens}\). At 1,200 tokens/s, the wasted execution duration is: \[t_{\text{waste}} = \frac{1{,}664{,}000\text{ tokens}}{1{,}200\text{ tokens/s}} \approx 1{,}387\text{ seconds} \approx 0.385\text{ hours}\] The energy dissipated on badput is: \[E_{\text{waste}} = 4{,}800\text{ W} \times 1{,}387\text{ s} \approx 6.66\text{ MJ} \quad (1.85\text{ kWh})\] At an on-demand market rate of \(\$32.00/\text{hr}\) for an 8x H100 SXM5 node, \(0.385\text{ hours}\) of pure badput wastes $12.33 across just these 65 failed tasks. At enterprise fleet scale (\(100{,}000\text{ tasks}\)), unmitigated cyclic failure inflates to \(2{,}846\text{ hours}\) of stalled GPU time, \(2.85\text{ MWh}\) of dissipated grid energy, and over \(\$18{,}900\) in wasted capital for zero delivered software artifacts.
Systems insight: Micro-optimizing decode kernels to gain 10 percent higher tokens/second saves fractions of a second. In contrast, an early-stopping verification harness that trips failing loops at step 10 reclaims over 60 percent of wasted cluster energy and capital, directly driving macro-efficiency.
Macro-efficiency forces systems designers to treat context memory, tool execution latency, and verification rigor as first-class architectural constraints. When foundation models operate under zero ambient authority, every token emitted carries a tangible marginal cost in accelerator thermal dissipation, memory residency, and host execution risk. To maximize Trajectory Goodput, an agentic system cannot merely generate tokens rapidly; it must maximize the probability that each generated token directly advances the trajectory toward an invariant-compliant terminal state.
How, then, does the host runtime orchestrate this continuous interplay between stochastic neural proposals, external tool executions, and empirical verification gates across a multi-step trajectory? The next section formalizes the closed-loop execution cycle, dissecting the mathematical state transitions and core architectural primitives that govern autonomous agency.
The Closed-Loop Trajectory
When an unprivileged model operates in an open-loop regime, generating an unvetted sequence of shell commands or file modifications and streaming them directly to a live filesystem, a single syntactic error or hallucinated flag permanently corrupts the target environment. In classical distributed systems, an RPC client communicating with a remote datastore relies on idempotency tokens, two-phase commits, and retry policies to survive network unreliability. When the entity issuing those RPCs is not a deterministic state machine but an autoregressive neural network evaluating floating-point tensors, the central systems challenge shifts fundamentally. The problem is no longer merely surviving network partitions; it is governing a stateful, non-deterministic controller through an extended sequence of physical side effects.
Autonomous agency is not an unconstrained dialogue or a single forward pass through a neural network; it is an iterative, closed-loop trajectory of discrete, typed state transitions governed by deterministic runtime supervision.
The physical partitioning that enforces this guarantee is formalized in figure 10, organizing the Trajectory Engine across four distinct physical domains:
- Unprivileged Foundation Model Engine (Domain I): Executes inside accelerator High-Bandwidth Memory (HBM). It receives the assembled context \(c_t\) alongside the immutable task goal \(g\), evaluates token probabilities, and emits an unprivileged candidate proposal \(a_{\text{prop}}\) across the PCIe/NVLink interconnect under zero ambient authority.
- Host Agent Runtime Supervisor (Domain II): Executes in host CPU DRAM and kernel space. It holds \(a_{\text{prop}}\) in memory escrow and evaluates the Authorization Gate \(P(a_{\text{prop}})\), stripping ambient permissions and verifying schema types. Valid proposals become permitted actions \(a_{\text{perm}}\), while violations trigger an error response \(\bot\).
- Isolated Execution Sandbox (Domain III): Runs within isolated Linux containers or microVMs. The runtime dispatches \(D(a_{\text{perm}})\) into this isolated effector, capturing raw side effects as structured observation envelopes \(o_{t+1} = (\text{stdout}, \text{stderr}, \text{exit\_code})\).
- Out-of-Band Verification Oracle (Domain IV): Executes sealed test suites and static checkers independently of the sandbox. It yields verification evidence \(v_{t+1}\), which feeds into the Completion Assessment Predicate \(\mathcal{K}_{\text{comp}}\) to decide whether to terminate with verified success or append \((a_t, o_{t+1}, v_{t+1})\) to the context memory for step \(t+1\).
Mathematical primitives
To analyze the stability, latency, and failure modes of an agentic system, we must formalize the core primitives that define the interface between the stochastic neural core and the deterministic host supervisor. An agentic trajectory unfolds across discrete time steps \(t \in \{0, 1, \dots, N-1\}\). At each step, the system manipulates six architectural primitives:
Goal (\(g\)): The immutable specification of the delegated task, supplied by the user or an upstream orchestration pipeline. The goal defines the target invariant and completion criteria (for example, “Identify and repair the race condition in the distributed configuration parser”).
Context (\(c_t\)): The staged token buffer representing the active state presented to the foundation model at step \(t\). The context is an ordered sequence drawn from the token vocabulary \(\mathcal{V}_{\text{tok}}\), constrained by the physical context window limit \(S_{\max}\). It incorporates the system prompt, the goal \(g\), tool schema definitions, working memory scratchpads, and the chronological history of past actions, environment observations, and verification verdicts: \[c_t = \big[ \text{SystemPrompt}, g, \text{ToolSchemas}, (a_0, o_1, v_1), \dots, (a_{t-1}, o_t, v_t) \big]\]
Model Invocation (\(M(c_t) \to a_{\text{prop}}\)): The autoregressive decode loop where the neural engine evaluates context \(c_t\) and samples a candidate action proposal \(a_{\text{prop}}\). The output \(a_{\text{prop}}\) consists of raw text or structured JSON fragments encoding an intended tool name and arguments.
Action Authorization Gate (\(P(a_{\text{prop}}) \to a_{\text{perm}} \lor \bot\)): The deterministic runtime validation filter. Operating under zero ambient authority, the host inspects \(a_{\text{prop}}\) against schema specifications, access control lists, path traversal boundaries, and resource rate limits. If validation succeeds, the runtime emits a permitted action \(a_{\text{perm}}\); if validation fails, the runtime suppresses execution and generates an error payload \(\bot\).
Runtime Dispatch (\(D(a_{\text{perm}}) \to o_{t+1}\)): The execution of the validated command against external effectors (local POSIX shells, container runtimes, remote HTTP endpoints, or language compilers). The dispatch mechanism isolates side effects and monitors execution limits such as wall-clock timeouts and memory cgroups.
Observation (\(o_{t+1}\)): The structured data payload returned by the external environment following dispatch. Observations encapsulate execution status, process exit codes, standard output (
stdout), standard error (stderr), or network response bodies.
Verification vs. Observation An observation (\(o_{t+1}\)) is the raw, uncurated return value from an environment effector (e.g., exit code 0 from a shell script). Verification evidence (\(v_{t+1}\)) is the deterministic verdict of an independent, out-of-band evaluation harness (e.g., an immutable test suite or AST parser) that evaluates whether the system state satisfies an operational invariant. Agents that inspect only \(o_{t+1}\) are vulnerable to false positives.
Beyond raw environment observations, robust agentic architectures introduce a seventh, critical feedback signal: Verification Evidence (\(v_{t+1}\)). While an observation captures the direct, local output of an action, verification evidence represents the outcome of an independent, authoritative validation check executed out-of-band (such as an immutable regression test suite, a static type checker, or a cryptographic signature validator).
We formally define a completed execution trajectory \(\tau\) of length \(N\) as a sequence of discrete interaction tuples: \[\tau = \big( (a_0, o_1, v_1), (a_1, o_2, v_2), \dots, (a_{N-1}, o_N, v_N) \big)\] where each transition from turn \(t\) to \(t+1\) represents a state advancement mediated entirely by the host runtime.
In their seminal work on the ReAct paradigm, Yao et al. (2023) demonstrated that interleaving explicit reasoning traces (“thoughts”) with discrete actions (“acts”) substantially mitigates hallucination cascades in foundation models. By emitting an intermediate natural-language rationalization before issuing a command, the autoregressive sequence conditions action selection on an explicit antecedent token prefix.
From an engineering stance, however, ReAct is not merely a prompting pattern; it represents the minimal algorithmic blueprint for a stateful, closed-loop execution runtime. The host architecture takes the raw ReAct loop and hardens it into a resilient operating system service. The runtime enforces typing contracts on actions, wraps untrusted process execution in isolated sandboxes, manages the temporal growth of context \(c_t\) across thousands of tokens, and injects authoritative verification evidence to keep the model tethered to physical reality.
| Primitive | Formal Symbol | Systems Representation | Physical / Runtime Boundary | Primary Failure Mode |
|---|---|---|---|---|
| Goal | \(g\) | UTF-8 String / Protocol Buffer | Invariant Specification (Read-only) | Semantic ambiguity; underspecified edge cases |
| Context | \(c_t\) | Token Sequence (\(\le S_{\max}\)) | Host Dynamic Memory / KV Cache | Context window exhaustion; attention dilution |
| Model Invocation | \(M(c_t)\) | Autoregressive Decode | Accelerator HBM / Matrix Cores | Hallucinated parameters; invalid JSON syntax |
| Authorization Gate | \(P(a_{\text{prop}})\) | Deterministic Policy Engine | Host Runtime (Outside the Model) | Schema rejection; path escape attempt |
| Runtime Dispatch | \(D(a_{\text{perm}})\) | Syscall / RPC Client | Container / Sandbox Boundary | Process timeout; unhandled socket exception |
| Observation | \(o_{t+1}\) | Typed Envelope / JSON Stream | Environment Effector | Stream truncation; noisy / uninformative errors |
| Verification | \(v_{t+1}\) | Assertion Verdict / AST | Isolated Sealed Evaluator | Incomplete test coverage; assertion bypass |
The six-phase execution lifecycle
The execution of an agentic system proceeds through a continuous finite-state machine managed by the host runtime. Rather than allowing the neural model to run unbounded, the host advances the system through six tightly coordinated phases on every turn \(t\), as formalized in figure 11 and summarized in table 4.
The state machine enforces deterministic control over every phase of execution:
- The cycle begins in Phase 1 (Context Assembly), verifying resource limits (\(t < T_{\max}\)) and staging context buffer \(c_t\).
- In Phase 2 (Model Invocation), context \(c_t\) is evaluated by the neural core to sample candidate proposal \(a_{\text{prop}}\).
- In Phase 3 (Action Authorization), the host runtime evaluates \(P(a_{\text{prop}})\). If authorization fails due to schema violations or forbidden paths, a Rejection Bypass transition returns directly to Phase 1 with error feedback \(\bot\), preventing any mutation of external state.
- If authorized, Phase 4 (Runtime Dispatch) forwards \(a_{\text{perm}}\) to isolated tool effectors, accommodating transient pauses or human-in-the-loop approval when required.
- In Phase 5 (Evidence Capture), the runtime captures observation envelope \(o_{t+1}\) and queries out-of-band verifier evidence \(v_{t+1}\).
- Finally, Phase 6 (Completion Assessment) evaluates the completion predicate \(\mathcal{K}_{\text{comp}}(o, v)\). If verified, execution terminates cleanly into
TASK_SUCCESS; if limits are exhausted, it exits toTASK_FAILURE; otherwise, it advances the turn counter \(t \leftarrow t + 1\) and cycles back to Phase 1.
Phase 1: Context assembly
Before allocating accelerator compute, the runtime evaluates trajectory bounds. It checks whether the current step index exceeds the maximum step budget (\(t \ge T_{\max}\)), whether accumulated financial or token costs exceed operator limits, or whether a cancellation signal has been asserted. If limits are respected, the host constructs \(c_t\).
Context assembly is an active memory-management operation: the runtime retrieves recent observations, formats them according to the tool specification protocol, compresses or truncates oversized stdout buffers from turn \(t-1\), and stages the final token array into host memory.
Phase 2: Model invocation
The assembled context \(c_t\) is dispatched to the inference engine (such as vLLM or a managed model endpoint). The engine executes the prefill phase over newly appended tokens, reuses cached Key-Value (KV) tensors for the static prompt prefix, and enters the autoregressive decode loop.
Decoding proceeds token-by-token until the model emits a designated stop token (e.g., <|end_of_action|>) or a structural delimiter signaling the end of an action invocation block. The output string is captured as \(a_{\text{prop}}\).
Phase 4: Runtime dispatch
If the action is authorized as \(a_{\text{perm}}\), the runtime dispatches it to the appropriate execution backend. For shell commands, the runtime invokes an isolated process inside a container or microVM sandbox. For file modifications, the runtime applies a structured diff to the target repository.
Crucially, dispatch is non-blocking with respect to the host’s administrative health: the runtime wraps the execution in strict wall-clock timers. If a sandboxed command hangs (for instance, if an invoked script triggers an interactive prompt or an infinite loop), the timer trips, sending a SIGKILL to the sandbox process group and returning a deterministic TRANSPORT_FAILURE observation to the host runtime.
Phase 5: Evidence capture
Once dispatch terminates, the runtime reclaims control. It packages the raw process output into a typed observation envelope governed by four canonical execution statuses:
COMPLETED: The process ran to termination within wall-clock and memory limits, returning an integer exit code alongside capturedstdoutandstderr.TRUNCATED: The process executed, but its output volume exceeded the turn observation budget; the runtime retains head and tail snippets while asserting this clamped status.REFUSED: The action was rejected prior to execution by the authorization gate due to permission, schema, or path boundary violations.TRANSPORT_FAILURE: The execution was aborted by a wall-clock timeout, container crash, or network socket termination.
Simultaneously, if the action mutated the persistent state of the workspace (such as modifying source code), the runtime invokes an out-of-band verification hook: running an immutable unit test suite, triggering a language linter, or parsing the modified files. The resulting pass/fail verdicts and compiler diagnostics are serialized into the verification evidence payload \(v_{t+1}\).
Phase 6: Completion assessment
Finally, the runtime assesses whether the trajectory has reached a terminal boundary. A completion predicate \(\Omega(c_{t+1}, v_{t+1})\) inspects the state. If the model proposal emits a terminal completion signal (such as an explicit complete_task RPC) and the verification evidence \(v_{t+1}\) confirms that all external invariant checks have passed, the trajectory transitions to the SUCCESS terminal state and halts.
If the model proposal asserts completion but \(v_{t+1}\) indicates failing invariant tests, the completion claim is rejected; the runtime converts the verification failure into a corrective observation and forces the trajectory to continue. If resources are exhausted without success, the runtime halts the trajectory as a FAILURE. Otherwise, the cycle increments \(t \leftarrow t + 1\) and loops back to Phase 1.
Because tools are remote effectors operating outside the model’s transactional memory, any non-idempotent tool mutation must implement a two-phase protocol: the model supplies an idempotency key generated deterministically from the goal ID and turn index, and the runtime dispatch layer verifies receipt with the remote effector or executes a deterministic reconciliation probe before allowing the model to proceed.
Checkpoint 0.3: The two-phase commit of tool mutation
Consider an agent whose HTTP mutation drops due to a socket timeout, returning an indeterminate TRANSPORT_FAILURE. Before advancing to transient pauses, verify your understanding of side-effect safety:
In addition to autonomous continuation and hard terminal states, practical agent runtimes must support Transient Pauses. A transient pause suspends the execution loop without tearing down the context buffer or releasing sandbox resources. Transient pauses arise in three standard operational scenarios:
- Clarification Pauses: The action proposal indicates an irreconcilable ambiguity in the task specification (e.g., discovering two conflicting configuration files with identical precedence) and explicitly yields control by invoking an
ask_operatorprimitive. The runtime suspends Phase 2, places the session in an idle queue, and awaits asynchronous user input. Upon receipt, the operator’s response is appended as \(o_{t+1}\), and the loop resumes. - Supervisory Approval Gates (Human-in-the-Loop): When an authorized action falls into a high-blast-radius capability tier (e.g., executing a database schema migration or issuing a recursive delete on a shared volume), the Action Authorization Gate transitions the session into a pending state. The action is held in escrow until a designated human supervisor cryptographically signs or approves the execution ticket.
- Rate-Limit Backoff: When an external API returns an HTTP 429 (Too Many Requests) with a
Retry-Afterheader, the runtime transparently suspends the trajectory at Phase 4, setting an asynchronous timer before retrying the dispatch, preventing the agent from burning context tokens on rapid retry loops.
Anatomy of a multi-turn trajectory
To see the closed-loop execution engine operating under real-world systems constraints, we examine a concrete multi-turn trajectory. Consider an automated bug-repair scenario: an agent is delegated the task of resolving a race condition in a distributed microservice configuration parser, where negative timeout values cause downstream network worker threads to deadlock.
Example 0.1: Anatomy of an agentic trajectory across three turns
Goal (\(g\)): “Fix the bug in src/config/parser.py where negative timeout inputs cause worker starvation. Clamp negative timeouts to the default 5.0 seconds. Ensure all existing tests pass.”
Turn 0: Read-only exploration
Context Assembly (\(c_0\)): Goal \(g\), tool specifications (
shell_exec,file_edit), and workspace directory map staged into memory.Model Invocation (\(M(c_0)\)): The model emits an exploratory search action proposal:
{"name": "shell_exec", "arguments": {"command": "grep -rn 'parse_timeout' src/"}}Action Authorization Gate (\(P(a_{\text{prop}})\)): Evaluates
shell_exec. Command matches read-only utility whitelist. Arguments contain no prohibited metacharacters. Status: Permitted.Runtime Dispatch (\(D(a_{\text{perm}})\)): Command executed inside a containerized sandbox. Wall-clock latency: 42 ms.
Observation Capture (\(o_1\)): Status:
COMPLETED. Exit code: 0.src/config/parser.py:142:def parse_timeout(raw_val: str) -> float: src/config/parser.py:143: return float(raw_val)Verification Evidence (\(v_1\)): Null (no environment mutation occurred).
Completion Assessment (\(\Omega\)): Task not complete. Step index \(t \leftarrow 1\).
Turn 1: The fail-plausible test bypass (adversarial failure mode)
Context Assembly (\(c_1\)): \(c_0 \cup (a_0, o_1)\) staged into context.
Model Invocation (\(M(c_1)\)): The autoregressive decode samples an action proposal targeting the test suite rather than repairing the underlying parser, minimizing apparent prediction error by weakening the invariant check:
{"name": "file_edit", "arguments": { "path": "tests/test_config.py", "old_str": "assert parse_timeout('-1.0') == 5.0", "new_str": "assert True" }}Action Authorization Gate (\(P(a_{\text{prop}})\)): Path falls within the writable workspace boundary. Syntax is valid. Status: Permitted.
Runtime Dispatch (\(D(a_{\text{perm}})\)): Patch applied to
tests/test_config.py. Sandbox local test execution (pytest tests/) executes and exits with return code 0.Observation Capture (\(o_2\)): Status:
COMPLETED. Exit code: 0.tests/test_config.py::test_parse_timeout PASSED [100%] ========================= 1 passed in 0.08s =========================Verification Evidence (\(v_2\)): The host runtime invokes the Sealed Verifier—an out-of-band test runner maintaining a cryptographically isolated copy of the test suite outside the agent’s writable workspace:
[SEALED AUDIT FAILURE]: Writable workspace modified protected test files! Running hidden golden regression suite... FAILED: tests/golden/test_timeout_clamp.py::test_negative_timeout AssertionError: parse_timeout('-1.0') returned -1.0, expected 5.0Completion Assessment (\(\Omega\)): The model proposal emits
{"name": "complete_task"}. However, \(\Omega(c_2, v_2)\) detects that the Sealed Verifier reported an invariant violation. The completion claim is rejected. The runtime injects the sealed failure log as observation \(o_2^{\text{corr}}\) into context, increments \(t \leftarrow 2\), and forces trajectory continuation.
Turn 2: Invariant closure
Context Assembly (\(c_2\)): Context updated with the rejection notice and the golden test traceback.
Model Invocation (\(M(c_2)\)): Conditioned on the golden regression traceback injected into \(c_2\), the model samples an action proposal targeting the root cause in the production parser:
{"name": "file_edit", "arguments": { "path": "src/config/parser.py", "old_str": " return float(raw_val)", "new_str": " val = float(raw_val)\n return 5.0 if val < 0.0 else val" }}Action Authorization Gate (\(P(a_{\text{prop}})\)): Valid edit targeting allowed source path. Status: Permitted.
Runtime Dispatch (\(D(a_{\text{perm}})\)): Patch applied to
src/config/parser.py. Local sandbox tests run and pass.Observation Capture (\(o_3\)): Status:
COMPLETED. Exit code: 0; all local tests pass.Verification Evidence (\(v_3\)): Sealed Verifier executes golden test suite against modified source:
tests/golden/test_timeout_clamp.py::test_negative_timeout PASSED tests/golden/test_timeout_clamp.py::test_positive_timeout PASSED tests/golden/test_timeout_clamp.py::test_zero_timeout PASSED [SEALED AUDIT SUCCESS]: All 42 golden invariants satisfied.Completion Assessment (\(\Omega\)): Model proposal emits
complete_task. Sealed verification evidence \(v_3\) confirms 100 percent invariant compliance. The runtime terminates the trajectory with SUCCESS.


