# Model Principles {.unnumbered} Part I studies the component every agent is built around, the foundation model, through the single call from which the rest of the book measures the H·S·A exposures. One call runs one turn, keeps nothing between calls, and holds zero ambient authority, so it adds none of the three exposures. What it returns is a proposal, and a proposal changes nothing until the runtime admits it (principle \ref{pri-invariant-closure}). Part I therefore fixes what every later part must govern. @Sec-vol3-foundation-model takes one call apart: the tokens and window that bound it, the sampling that makes it nondeterministic, the call surface through which the runtime offers tools and reads stop reasons, and the price of a call and of a trajectory of calls. @Sec-vol3-test-time-compute builds on that price and asks when spending more inference on one decision, through longer reasoning, more candidates, or check-and-revise rounds, raises the evidence behind the proposal the runtime finally admits. The first principle prices a single proposal. ::: {#pri-vol3-memory-bandwidth-decoding .callout-principle title="Output-bound latency"} **Invariant**: A model writes its output one token at a time, and each step waits for the one before it, so the latency of a call grows linearly with the tokens it generates. On an agent's own trajectory, which runs as a single stream because each turn waits on the previous observation, an output token costs about two orders of magnitude more time than a prompt token. The full treatment appears in @sec-vol3-foundation-model-cost, and @sec-vol3-appendix-inference-roofline derives why. **Implication**: The runtime controls a call's latency through its output: interfaces that express the same action in fewer tokens, output limits that stop runaway generations, and streaming validation that cancels a doomed generation early. Output length is not the lever for a whole trajectory's cost, because the context re-sent on every turn dominates the tokens processed, and there the lever is a context that grows only by extension so the serving system can reuse its prefix. ::: Output-bound latency prices a proposal that finishes. A call can also stop short, at its output limit, on a cancel, or on a dropped connection, and what it returns then is a fragment rather than a proposal. @Sec-vol3-foundation-model reduces every outcome of a call to one closed set of statuses and states the rule that governs them, the quarantining invariant (principle \ref{pri-02-quarantining-invariant}), under which only a call that completed cleanly may reach verification or dispatch. Test-time compute buys more finished proposals, through longer reasoning, more candidates, or feedback from checks, and the question becomes when that extra spend pays. ::: {#pri-vol3-test-time-scaling .callout-principle title="Evidence-bounded deliberation"} **Invariant**: Extra inference compute raises verified task success only when it acquires evidence that separates valid candidates from invalid ones. Longer generation and unguided resampling reuse the same parameters and the same prompt, so without a discriminating check they add cost faster than they add evidence. The full treatment appears in @sec-vol3-test-time-compute-insufficient-response. **Implication**: The runtime should favor allocations that consult a check outside the model, meter generation, verification, and tool cost in one ledger with hard ceilings, stop when marginal gain falls below marginal cost, and judge every test-time strategy against a single-candidate baseline at a matched budget. ::: Whether a candidate counts as evidence depends on the check that judges it, and checks differ sharply in what they cost and what they can establish. ::: {#pri-vol3-verification-asymmetry .callout-principle title="The verification asymmetry"} **Invariant**: Generating candidates is cheap, and so is running a stated mechanical check on one, such as a compiler pass, a schema, or a sealed test suite. Sound verification is not cheap. A check establishes only what it covers, deciding that it covers what the task requires is undecidable in general, and that decision cannot be delegated to the model whose output is being checked. When valid candidates are rare, a check that seldom errs still approves more invalid candidates than valid ones, and searching harder against an imperfect check selects for what it fails to test. The full treatment appears in @sec-vol3-test-time-compute-verification. **Implication**: Search scales only as far as sound checks bound it. Learned verifiers may rank candidates to steer search, but the commit gate must be a deterministic check whose coverage matches the task's completion criteria, written and held where the agent cannot modify it. A task whose completion criteria have no mechanical form cannot be accepted autonomously, and a second model acting as judge does not supply one. ::: Together, these principles fix what the rest of the book builds around. A call yields an unverified, nondeterministic proposal whose latency is set by what it writes and which counts only if the call completed, more inference helps only when it brings in evidence, and only a deterministic check the agent cannot reach may accept the result. @Sec-vol3-foundation-model establishes the first, states the quarantining invariant that keeps an unfinished call from being treated as a proposal, and shows why the model cannot check itself, and @sec-vol3-test-time-compute works the second and third through selection, verification, and stopping. Every later part surrounds the model with machinery that supplies what a single call lacks, beginning with the state a trajectory must carry once it runs for more than one turn.