Reference Architecture

The chapters build an agent system one mechanism at a time, each where its failure first appears. This appendix assembles those mechanisms into a single reference architecture, so that a reader designing a system, reviewing one, or tracing a failure across it can check which components are present, what each one owns, which H·S·A exposure it closes, and which chapter derives it. It is organized by the components the book uses (the model, context and memory, tools and sandboxes, the harness, the durable log, recovery, evaluation, the learning loop, and the fleet) and names techniques rather than products; where an implementation is mentioned, it is one example of the technique.

How to Use This Appendix

Consult the section that matches the question in front of you.

  • When you are designing a new agent system and need a checklist of components and the interfaces between them, start with the component map in section 1.
  • When a task’s H·S·A position is known and you need to decide how much of the architecture it requires, use the closure profiles in section 2.
  • When a failure crosses components and you need to know which one owned the step that went wrong, follow one turn through the architecture in section 3.
  • When you are choosing an isolation envelope and need representative start costs, use section 6.
  • When you are laying out where durable state lives and who may write it, use section 5.
  • When you are deploying the architecture across machine pools and need the failure signals each pool emits, use section 12, and classify the failure itself with Agentic Failure Taxonomy.

The Component Map

An agent system has one component that proposes and many that decide. The model reads a context and returns a proposal (The Foundation Model). Everything else exists to decide what the model sees, which proposals take effect, how effects are made safe, how the trajectory survives failure, and how anyone knows the result is right. Functional Architecture of the Stochastic Computer draws these components as a map of the book’s parts. Table 1 lists them as an engineer would build them, with the exposure each closes, the invariant it enforces, and what goes wrong when it is missing.

Table 1: The Component Map: Each component of an agent system, what it owns, the H·S·A exposure it closes, the invariant it enforces, the failure its absence produces, and the chapter that develops it.
Component Owns Exposure closed Invariant it enforces Failure when absent Chapter
Model service The call surface, sampling, stop reasons, token accounting None (the origin) Every output is a proposal until the runtime acts on it Proposals treated as results; fluent output taken as done The Foundation Model
Context assembler What each call sees: layout, provenance labels, compaction, invalidation State (\(S_1\)) Every staged item has a source, a budget share, and a validity rule Context overflow, dilution, stale observations, poisoning Context Engineering
Serving memory The attention state of live trajectories, prefix reuse, residency during tool waits State (\(S_1\)) Attention state is derived from the token prefix and can always be rebuilt from it Memory stranded during tool waits; prefix reuse lost to layout churn KV Cache Management
Durable stores Sources, agent-written memory, retrieval indexes, the retrieval contract State (\(S_3\)) Authority stays with the source; every derived copy is checked against it Stale retrieval, self-written poisoning, contradictions with disk Long-Term Memory
Tool gateway Schema validation, the grant, idempotency keys, settlement, observation shaping Authority Complete mediation: no call reaches the environment without validation and a grant Duplicate effects, unvalidated calls, unreadable observations Tool Calling
Isolation envelope Capabilities, credentials, the workspace, egress, taint, sandbox pools Authority (\(A_1\)) Containment beneath the model, sized to the task’s authority A granted call reaching state, secrets, or hosts the task never needed Agent Sandboxes
Harness The trajectory record, the state set, control actions, approval gates, budgets Horizon The runtime can pause, bound, or stop the loop without the model’s cooperation Runaway loops, late results accepted, irreversible calls dispatched unseen The Agent Harness
Trajectory log Write-ahead intent, checkpoints, resume, replay Horizon No external effect leaves the host before the intent that authorizes it is durable Lost progress after a crash; repeated effects on resume Durable Execution
Recovery Sagas and compensation, the pivot boundary, forward repair, progress detection, circuit breakers Horizon and Authority Every compensable effect has a registered compensator before dispatch Half-finished effects, silent spinning, retry storms Failure Recovery
Evaluation Task environments, grading, repeated-trial statistics, traces, attribution, release gates All three (measured) Success is a verified change in the environment at a stated evidence level Releases gated on transcripts; failures blamed on the wrong layer Agent Evaluation
Learning loop Curation, admission cascades, fine-tuning, reinforcement learning against verifiers None moved Nothing is trained that evaluation has not verified and split hygiene has not separated Training on unverified or contaminated trajectories; reward hacking Trajectory Curation
Fleet Multi-agent envelopes and delegation, capacity, routing, spending governance All three multiplied Authority only attenuates on delegation; every accepted task has a price Correlated errors, orphaned work, budgets breached across delegation Multi-Agent Coordination

The table reads in two directions. Read down the “Owns” column, it is a checklist, and a system that cannot name the component responsible for each row has a gap. Read across the “Failure when absent” column, it is a diagnostic, and a failure that matches a row points at the component that should have prevented it. Even the model service’s invariant is enforced by the runtime rather than by the model, because every invariant in the table holds below the model (principle \(\ref{pri-invariant-closure}\)).

Closure Profiles by Exposure

Not every task needs every component. A read-only question-answering loop needs a context assembler and a harness with a turn ceiling, and adding a saga engine to it only adds cost. A production change needs nearly everything. The H·S·A position of a task (The H·S·A exposures) decides how much of the architecture it requires, and table 2 states the minimum for each level.

Table 2: Closure Profiles by Exposure Level: The minimum closure the runtime must supply at each H·S·A level and the components that supply it. A task needs the union of the rows its position selects, and the evidence level of its completion criteria rises with the same exposure.
Exposure level Minimum closure the runtime supplies Components that must be present
Horizon, short (\(H \lesssim 10\)) Turn, token, and cost ceilings; output ceiling and deadline on each call Harness (budgets)
Horizon, long (\(H \gtrsim 100\)) Durable log with write-ahead intent, checkpoints, resume, progress detection Harness, trajectory log, recovery
State \(S_1\) (context only) A context budget, provenance labels, compaction that keeps identifiers verbatim Context assembler, serving memory
State \(S_2\) (sandboxed workspace) A workspace leased to the trajectory, reset between tenants, snapshots bound to log positions Isolation envelope, trajectory log
State \(S_3\) (durable, shared) A retrieval contract that checks version and provenance against the source; a write path the runtime admits Durable stores, context assembler
Authority \(A_0\) (read) Credentials kept out of context; egress restricted, since reading can still leak Tool gateway, isolation envelope (egress)
Authority \(A_1\) (sandboxed mutation) An envelope matched to the code, capabilities scoped to the workspace, a reset contract Tool gateway, isolation envelope
Authority \(A_2\) (retriable, compensable) Idempotency keys and settlement before retry; a compensator registered before dispatch Tool gateway, trajectory log, recovery
Authority \(A_3\) (irreversible) An approval gate that holds the call with an expiry and re-checks its preconditions; the pivot placed last Harness (approval gate), recovery (pivot boundary)

A task selects one row from each exposure and needs the union of what they require. The flaky-test repair of Architectural decision matrix: Workflows versus model-directed loops sits at a long horizon, \(S_2\), and \(A_1\), so it needs the harness, the trajectory log, recovery, the isolation envelope, and the tool gateway, but no approval gate. The same agent pushing its fix to a shared branch moves to \(A_2\) for the push and \(S_3\) for the branch, which adds settlement, a compensator, and a source check on anything it later retrieves about that branch. Lowering a task’s authority to what it needs is usually cheaper than building closure for authority it never needed, which is the lesson of the incident in The H·S·A exposures.

One Turn Through the Architecture

A single turn of a long trajectory touches almost every component, and each step has one owner. Following the turn in order shows where each invariant is checked and which log event records it.

  1. Admit the turn. The harness takes the trajectory from READY, checks its remaining budgets, and moves it to CALLING_MODEL (Trajectory States).
  2. Assemble the context. The context assembler builds the call from the trajectory’s history under the context budget, stable prefix first and volatile content last, with provenance labels on untrusted text (Staging the Next Invocation). The log records that the context was built, with a hash of what went in.
  3. Call the model. The model service receives the call with an output ceiling and a deadline, reuses the cached prefix where it can (Prefix Caching Across Turns), and returns a proposal with a stop reason and token usage (The Invocation Contract). The log records the reply.
  4. Validate and grant. The tool gateway parses each proposed call against its schema, decides the grant from the runtime’s own descriptor of the tool (Interoperable Tool Discovery), and routes an irreversible call to the approval gate instead of dispatching it (Approval Gates).
  5. Write intent, then dispatch. The trajectory log makes the intent durable before the effect leaves the host (Intent Before Effect). For a compensable call, recovery registers the compensator at the same point (The Trajectory Saga Pattern). The gateway then dispatches the call with its idempotency key into the isolation envelope, under a capability scoped to the task (Capabilities and Credentials).
  6. Wait without holding more than needed. The harness moves the trajectory to WAITING_TOOL. The serving memory decides whether the trajectory’s attention state stays resident, is evicted and recomputed, or is offloaded for the length of the wait (Retain, Evict, Recompute, or Offload).
  7. Settle and shape the result. The gateway settles the call, accepting a result only for a call ID that is still pending, and shapes the output into a bounded, structured observation (Observation Stream Truncation). The log records the result.
  8. Check progress and continue. Recovery compares the new environment state with the recent history to detect a loop (Semantic Watchdog Timers), and the harness returns the trajectory to READY or, if the completion criteria are met, to FINALIZING, where evaluation’s checks decide acceptance (The multi-layer evaluation contract).

Each numbered step also emits a span to the trace (Distributed trajectory tracing), so a failed turn can be replayed from the log and attributed to the step, and therefore the component, that let it through.

The Model Service

The rest of the architecture needs five things from the model service, and none of them is the model’s accuracy. It needs a typed call surface, so a finished tool call can be told from a truncated one and an answer from a refusal (The Invocation Contract). It needs stop reasons, because they are the loop’s control signal. It needs token usage per call, because every budget in the harness is counted in tokens. It needs cancellation, so the harness can stop generation the moment a streamed proposal fails validation. And it needs prefix reuse across the turns of one trajectory, because a trajectory re-sends its context on every turn and prefix caching is the dominant saving on that re-sent input (Accelerator Serving Latency).

Behind that surface, the serving system applies a small set of techniques whose effects an agent engineer sees directly. Paged attention state lets many trajectories share one memory pool and lets forked branches share their common prefix (Paged KV Allocation). A prefix cache indexed by token sequence lets a new turn prefill only what changed (Prefix Caching Across Turns). Batching many trajectories’ decode steps together lowers the fleet’s cost per token without shortening any one trajectory’s critical path, and chunked prefill keeps a large observation from stalling every other trajectory’s decode (Why Output Costs More Than Input, Chunked Prefill Scheduling). The Size of the Attention State derives the size of the attention state per token, and Queueing for Trajectories sizes how many trajectories a server can hold.

Two design consequences follow for the rest of the architecture. Routing all turns of a trajectory to a server that still holds its prefix turns most re-sent input into cache hits, so the harness and the fleet router should carry a session identifier that the serving layer can use for affinity. And any change to the stable prefix, such as reordering tool definitions or inserting a timestamp near the top of the context, silently discards that saving, so the context assembler owns cache stability as much as the serving layer does (Staging the Next Invocation).

Context and Memory

An agent’s state lives in several places that differ in who may write them, how long they last, and what keeps them true. The architecture needs one owner for each, and it must never confuse a derived copy with its source (principle \(\ref{pri-vol3-source-authority}\)). Table 3 places the state classes of Taxonomy of State Tiers in Persistent Systems on the components of this appendix.

Table 3: State Classes by Owner: Where each class of agent state lives in the reference architecture, who may write it, and what keeps it consistent with the world.
State class Level Component that owns it Who may write How it is kept true
Working context \(S_1\) Context assembler Runtime assembles; the model only reads Rebuilt every turn; stale items invalidated against their source
Attention state \(S_1\) Serving memory Serving engine, derived from tokens Recomputed from its token prefix when evicted
Workspace artifacts \(S_2\) Isolation envelope The agent, through tools, inside the sandbox Leased per trajectory, reset between tenants, snapshotted at checkpoints
Source artifacts \(S_3\) The owning system Owners and granted tool writes Authoritative; change only through explicit writes
Trajectory log \(S_3\) Trajectory log Runtime only, append-only Never rewritten; the trajectory record is folded from it
Agent-written memory \(S_3\) Durable stores (write path) The model proposes; the runtime admits Admitted, consolidated, and expired by policy
Retrieval indexes \(S_3\) Durable stores (indexer) Indexer only, rebuilt from sources Invalidated when a source changes; disposable

Three architectural rules follow from the table. First, the trajectory log is authoritative about the past and nothing else, so it belongs beside the harness, not in the memory hierarchy, and no cache ever stands in front of it. Second, every derived row (attention state, agent-written memory, retrieval indexes) can be discarded and rebuilt, so losing one costs time, never correctness. Third, the only rows the agent writes directly are the workspace, which the envelope contains, and agent-written memory, which passes through an admission step. The context assembler reads from all of them, and the retrieval contract (The Retrieval Contract) is the interface through which durable state re-enters a call, with its version and provenance checked before it is staged.

Tools and Sandboxes

The tool gateway and the isolation envelope together turn a proposal into a contained, settled effect, and they divide the work cleanly. The gateway decides whether a call may run. It validates arguments against the tool’s schema, makes the grant from the runtime’s own descriptor rather than from what a tool server advertises, attaches an idempotency key to every mutating call, settles an ambiguous timeout before any retry, and shapes the result into a bounded observation (Tool Calling). The envelope decides what the call can reach once it runs, through the capability it carries, the credentials it never sees, the workspace it writes, and the hosts it may contact (Agent Sandboxes). A grant that the envelope does not enforce is a request, not a bound.

The envelope’s cost is dominated by how quickly a fresh one can be made, because that sets whether the runtime can afford a clean envelope per tool call, per trajectory, or only per tenant.

Table 4: Representative Envelope Costs: Orders of magnitude for making a fresh isolation envelope, with the microVM row taken from the coding-agent profile the sandbox chapter uses for pool sizing. Measurements for a particular implementation vary with image size, snapshot layout, and host load; the orders of magnitude are what drive the choice.
Envelope What untrusted code shares with the host Fresh-envelope cost Per-envelope memory Typical placement
Shared-kernel container The whole host kernel interface, filtered Hundreds of milliseconds The process’s own memory; the kernel is shared Code the operator wrote; read-only work with no secrets in reach
User-space kernel A small set of host calls behind an emulating kernel Hundreds of milliseconds, with a per-call tax on system-call-heavy work The process plus the emulating kernel Untrusted interpreted code with modest input and output
MicroVM A minimal monitor and the hypervisor interface About 5 ms from a warm pool; about 15 ms to restore from a snapshot A 512 MiB guest image, mostly shared copy-on-write across a pool Untrusted native code, package installs, multi-turn coding
Bytecode sandbox Only the functions the host imports Microseconds The module’s linear memory Pure computation, parsers, one fresh envelope per call

The microVM row explains why pools exist. A restore in tens of milliseconds is cheap next to a model call, but a burst of hundreds of trajectories that each need one at once turns it into a queue, which is why Sandbox Pools and Reset sizes a warm standby from the arrival rate and the restore time. Around the envelope, three supporting services complete the authority half of the architecture. A credential broker injects scoped credentials at the point of use so that no secret enters the model’s context (Credential brokering). An egress proxy allows only the hosts the task contract names and closes the side channels, such as name lookups, that a default-deny firewall leaves open (Egress Control). Taint tracking marks data that came from untrusted sources and keeps it away from calls with authority, down to a quarantined reader that sees untrusted text but holds no tools (Tracking Untrusted Data).

The Harness

The harness is the component that owns the loop for one trajectory, and its central data structure is the trajectory record: the trajectory’s identity, current state, per-call limits, budgets and measured spend, grants and leases, pending work, and pointers to its context, plan, log, and workspace (The Trajectory Record). The model never writes the record. The record lives in a closed set of states with guarded transitions, so a tool result that arrives after its deadline finds no pending call and is dropped, and an irreversible call can leave the approval state only through an approval that passes its re-check (Trajectory States).

The harness acts on the record through a small set of control actions (cancel, interrupt, pause, resume, kill, and budget exhaustion), each taking effect at a turn boundary or, for generation in flight, through streaming cancellation (Control Actions). Its approval gate is the one place the architecture holds a proposal in escrow. An irreversible call waits, undispatched, with an expiry, is re-checked against its recorded preconditions when approved, and becomes a refusal observation if it is denied or goes stale (Approval Gates). Its budgets bound turns, tokens, wall-clock time, and money per trajectory, and tighten progressively from a steer to read-only operation to a halt (Budgets and Ceilings). In a deployment the harness runs as a pool of stateless workers, because the record it needs to continue any trajectory can be rebuilt from the log.

The Durable Log and the Trace

Two records describe a trajectory, and the architecture keeps them separate because they answer different questions. The trajectory log is the source of truth for what happened. It is an append-only sequence of typed events (context built, model replied, tool intent, tool result, approval decided, compaction applied), and the trajectory record is a projection of it that any worker can rebuild by folding the events in order (The Trajectory Log). Its one ordering rule is write-ahead intent. The event that authorizes an external effect must be durable before the effect leaves the host, so that after a crash every effect is either known to have been intended or known not to have happened (Intent Before Effect).

Around the log sit three services. Checkpoints bind a snapshot of what the log cannot rebuild cheaply, such as the workspace, to a log position, so resume does not replay from the beginning (Checkpoints). Resume rebuilds the record, re-derives the context, settles every in-doubt effect, and pays one cold prefill unless the trajectory is routed back to a server that still holds its prefix (Resuming a Trajectory). Replay runs the real harness code against the log with the model and tools stubbed, which turns every logged failure into a reproducible test (Replay).

The trace serves measurement rather than recovery. It is a graph of spans, one per model call, tool call, approval, and child agent, each carrying token counts, durations, and outcomes, and it is sampled and redacted by policy rather than kept in full (Distributed trajectory tracing). The log must be complete and durable, and the trace must be cheap and queryable. An architecture that uses one for both jobs either pays durability costs on every metric or loses events it later needs for resume.

Recovery

The log says what happened. Recovery decides what to do when what happened was wrong (Failure Recovery). Its components sit on both sides of the model. Forward recovery goes through the model. It uses an error observation shaped so that the next proposal can fix the cause, a bounded number of repair attempts, resampling or a different model when repair stalls, and rollback of a poisoned context branch (Forward Recovery). Backward recovery goes around it. A saga registers a compensator for each compensable effect before dispatch and runs them in reverse order when the trajectory is abandoned (The Trajectory Saga Pattern), and the pivot boundary places the one irreversible step last, behind the approval gate (Pivot Action Irreversibility).

Two detectors decide when recovery starts. Progress detection watches for a trajectory that is busy but going nowhere, by hashing repeated actions and comparing environment states across turns (Semantic Watchdog Timers). A circuit breaker in the tool gateway stops dispatching to a dependency whose calls keep failing at the transport level and returns a typed observation instead, so the model does not become a retry amplifier (Tool Circuit Breakers). After a failure that may have spread, quarantine freezes the trajectory’s grants and marks what it touched until the damage is assessed (Blast Radius Quarantine).

Evaluation

Evaluation is part of the architecture, not an activity performed on it, because it needs components of its own that the agent cannot reach. A task environment starts every trial from a reproducible state with sealed tests the agent cannot read or modify (Hermetic evaluation gyms). Graders judge the resulting environment state, with deterministic checks where they exist and calibrated judges where they do not (Grading trajectories). A statistics layer turns repeated trials into intervals and separates what an agent can do on some attempt from what it does reliably on every attempt (Statistical evaluation rigor). Attribution replays a failed trajectory and ablates one layer at a time to decide whether the model, the harness, or the environment caused it (Forensic incident post-mortems). Release gates carry a candidate from an offline suite through a shadow run and a canary before it takes full traffic (Staged canary deployments).

The evaluation environment reuses the production isolation envelope and pools (Sandboxes as evaluation and rollout substrate) but must not share its credentials, its egress rules, or its verifiers’ storage with the agent under test. A verifier the agent can reach is a verifier the agent can learn to satisfy without doing the task.

The Learning Loop

The learning loop sits on the far side of evaluation and consumes only what evaluation certified. Curation turns logged, verified trajectories into training data through an admission cascade ordered by cost, keeps the tests that judged a trajectory out of its reach, and separates training tasks from evaluation tasks before anything is trained (Trajectory Curation). Supervised fine-tuning serializes admitted trajectories into the model’s call format and computes loss only on what the policy wrote (Trajectory Fine-Tuning). Reinforcement learning optimizes against verifiers isolated from the policy’s rollouts (Reinforcement Learning from Verifiable Rewards).

Architecturally, the loop adds three things. It adds a rollout fleet that reuses the sandbox pools at a scale set by training rather than by traffic. It adds a data path from the trajectory log, through redaction and provenance manifests, into training storage. And it adds a release path back into the model service that passes through the same gates as any other change. The loop moves no exposure, since a better model still only proposes, which is why it is the last rung of the intervention ladder (The intervention ladder).

Deployment and the Fleet

A deployment separates the architecture into pools whose resources, failure modes, and security boundaries differ. The model service runs on accelerator nodes and is sized by attention-state memory and decode throughput. Harness workers are stateless, run on ordinary hosts, and are sized by the number of live trajectories. Sandbox hosts run the isolation envelopes with hardware virtualization, sit behind the egress proxy, and are never on the same network as the durable stores. The durable stores hold the trajectory log, the source systems’ mirrors, agent-written memory, and retrieval indexes. Verification runs on hosts the agent’s envelopes cannot reach.

Capacity is set by trajectory lifetime, not call duration. By Little’s law, the number of live trajectories equals the rate at which they start times how long each one lasts, and a trajectory lives for minutes to hours while each of its calls lasts seconds (Capacity for Trajectories; Queueing for Trajectories). Every pool that holds something for a trajectory’s lifetime, whether a harness worker’s record, a sandbox lease, or cache memory, is sized from that product. Across agents, a typed task envelope carries attenuated authority and a budget from parent to child (Typed task envelopes), budget reservations keep a delegation tree within its spending limit (Spending Governance), and cancellation propagates down the tree (Cancellation cascades).

Each pool fails in its own way and emits its own signal. Table 5 lists the signal to watch and the component that responds.

Table 5: Deployment Failure Signals: The symptom each pool shows when a component is missing or undersized, the signal that exposes it, and the component that responds.
Symptom Pool Likely cause Signal to watch Responding component and chapter
Calls queue while accelerators sit idle Model service Attention state held by trajectories waiting on tools Resident attention state per active decode stream Serving memory: evict, recompute, or offload during waits (Retain, Evict, Recompute, or Offload)
Cost per turn rises without longer contexts Model service Prefix-cache hit rate collapsed after a change to the stable prefix Cached-token share of input tokens per trajectory Context assembler: restore a stable layout (Staging the Next Invocation)
Trajectories busy but going nowhere Harness A loop the model cannot see from inside its growing context Repeated action hashes; unchanged environment state across turns Recovery: progress detection (Semantic Watchdog Timers)
Sandbox leases time out under bursts Sandbox hosts Standby pool smaller than arrivals during one restore Standby queue empty; restore queue depth Isolation envelope: pool sizing (Sandbox Pools and Reset)
Duplicate external effects after a restart Harness and log Effects dispatched before their intent was durable, or retried unsettled Tool results without a matching durable intent; repeated idempotency keys Trajectory log and gateway: write-ahead intent and settlement (Intent Before Effect)
A failing dependency drains budgets Tool gateway The model retrying a tool whose service is down Transport-level failure rate per dependency Recovery: circuit breaker (Tool Circuit Breakers)
Retrieved facts contradict the workspace Durable stores An index or memory entry not invalidated after a source write Version mismatches rejected by the retrieval contract Durable stores: invalidation (Storage Invalidation)
Spending exceeds limits across a delegation tree Fleet Children spending from a shared budget without reservations Reserved versus settled spend per tree Fleet: spending governance (Spending Governance)

Summary

The reference architecture is the book’s argument laid out as a system. One component, the model, proposes. Every other component closes an exposure the proposal creates: the context assembler, serving memory, and durable stores close state; the tool gateway and isolation envelope close authority; the harness, trajectory log, and recovery close horizon; evaluation measures whether the closure held; the learning loop improves the proposer only after evaluation can tell; and the fleet multiplies all of it. A task’s H·S·A position selects which components it needs, and a failure’s symptom points at the component that should have caught it.

Back to top