# Agents at Scale Principles {.unnumbered} Parts I through V built one agent and then made it better: a model call, the memory and tools around it, a runtime that carries it across time and measures it, and training that makes its proposals pass more often. Part VI changes the count, in two directions. One task may be split across several agents, so a single trajectory becomes a tree of trajectories joined by handoffs. Many tasks also run at once on shared models, quotas, and sandboxes, so every trajectory has a price and competes with the others for capacity. Scale adds no new exposure. It multiplies all three H·S·A exposures. Delegation stretches the horizon across agents, so an error at one handoff becomes the premise of the next. State is copied into every child's context and read after it has gone stale. Authority and budget pass down every delegation and must not grow on the way. Every closure that Parts II through IV built for one agent must therefore hold across agent boundaries and across a fleet, and none of it is worth running unless it pays for itself per accepted task. The two chapters take these questions in order. @Sec-vol3-multi-agent decides when splitting a task beats one agent given the same total budget, judged on accepted-task rate, makespan, and cost per accepted task, and builds what the runtime must then carry across every agent boundary: typed handoffs with bounded returns, isolated writes, gated commits, cancellation, and narrowing authority. Cost per accepted task, the axis on which a fleet most often loses, is the quantity @sec-vol3-agent-economics then prices for one agent or many. That chapter works through caching and reasoning budgets, critical-path latency, model routing, capacity for trajectories, and money budgets reserved across the delegation trees of the first chapter, and it ends by pricing whether a task deserves an agent at all. Five principles govern the part. ::: {#pri-vol3-coordination-tax .callout-principle title="The coordination tax"} **Invariant**: Splitting a task across agents places handoffs, duplicated context, waits at joins, and reconciliation on the critical path, and agents that share weights share errors. Speedup is capped by the serial fraction of the task and falls once pairwise coordination grows faster than the parallel work shrinks. The full treatment appears in @sec-vol3-multiagent-need. **Implication**: A fleet earns its tax only through context isolation, authority separation, or independent exploration, and it must then run as an explicit task graph whose edges carry typed envelopes and bounded returns, never as open conversation among agents. ::: Whether a fleet earns its tax is settled by measurement, and a measurement is only as fair as its baseline. Extra agents are extra inference spent in parallel, so the baseline is the same one that evidence-bounded deliberation (principle \ref{pri-vol3-test-time-scaling}) sets for any extra inference, a single agent given the same total budget. ::: {#pri-vol3-tri-axial-evaluation .callout-principle title="Tri-axial multi-agent evaluation frontier"} **Invariant**: A multi-agent topology has positive systems value only if, against a single agent given the same total budget, it is no worse on accepted-task rate, critical-path makespan, and cost per accepted task, and strictly better on at least one of them. The full treatment appears in @sec-vol3-multiagent-evaluation. **Implication**: Claims for a multi-agent design must hold the budget equal, compare against a single agent that uses test-time revision and tools, and count coordination overhead in both makespan and cost. ::: A fleet that clears this frontier must still stay bounded, because no handoff between its agents may become a way to acquire authority or budget. ::: {#pri-vol3-monotonic-delegation .callout-principle title="Monotonic delegation"} **Invariant**: When one agent delegates to another, the child's authority, time, and budget must be a subset of what the parent still holds, and whatever the child receives is withheld from the parent until the child returns it. No sequence of delegations can create authority or spending that the root did not have, and money once spent cannot be recovered. The full treatment appears in @sec-vol3-multiagent-authority. **Implication**: The runtime, not the agents, enforces the subset rule at every handoff, using credentials that each holder can narrow but never widen, budget reservations subtracted from the parent before any child spends them, and cancellation that reaches every descendant. ::: Bounding what a fleet may spend leaves the question of what its spending buys. The answer is the accepted task, the unit that agent evaluation already counts as success (@sec-vol3-evaluation-contract). ::: {#pri-vol3-trajectory-goodput .callout-principle title="Trajectory goodput and the accepted task"} **Invariant**: Only trajectories that pass independent verification deliver value, so every resource spent on failed or unverified attempts is charged to the accepted ones. Cost per accepted task, not cost per token, is the quantity a fleet minimizes, and trajectory goodput is the share of resources that reaches accepted tasks. The full treatment appears in @sec-vol3-agent-economics-accounting. **Implication**: Model choice, caching, routing, acceleration, and capacity are judged by their effect on cost and time per accepted task across the whole trajectory, which often makes tool execution, sandboxes, and verification, rather than model inference, the term worth optimizing. ::: Cost per accepted task says what to minimize. Capacity must also absorb how trajectories arrive and how long they hold what they are given. ::: {#pri-vol3-heavy-tailed-scheduling .callout-principle title="Trajectory-lifetime capacity and tail-aware scheduling"} **Invariant**: A trajectory holds capacity for its whole lifetime, which is orders of magnitude longer than any one model call, and the calls it issues arrive in correlated bursts with service times that differ by orders of magnitude. Queue delay grows with that variability and rises steeply as utilization approaches saturation, and a trajectory of many serial calls is likely to meet at least one tail delay on its critical path. The full treatment appears in @sec-vol3-agent-economics-capacity. **Implication**: Capacity is sized from trajectory concurrency, the arrival rate of trajectories times their lifetime, and from tokens per trajectory against any provider quota, and it is run below the knee of the utilization curve. Because a scheduler cannot know a generation's length in advance, it demotes calls that exceed a token quantum so that short calls never wait behind long ones. ::: The first three principles decide whether a fleet may run, how it is judged, and how it stays bounded; the last two decide what it costs and how much capacity it needs. With every subsystem built and priced, Part VII runs them together on one trajectory.