Preface
Who This Book Is For
This book is for senior undergraduates and beginning graduate students in computer science and computer engineering who are taking a course on agentic machine learning systems, and for practicing engineers who build, operate, or evaluate agents. It is written for the engineer who has watched an agent demonstration succeed and now has to ship a system that carries delegated work to an accepted result, day after day, without deleting the wrong directory, looping until a budget runs out, or reporting success on a task it never finished.
The reader should know how a neural network is trained and served and be comfortable with the core ideas of computer systems. The book assumes no prior volume of this series. Every agentic concept, and every serving or systems idea an agent engineer needs, is developed from first principles where it is first used.
Why an Agentic Systems Textbook
A reasonable engineer might look at the current wave of agent demonstrations and ask a skeptical question. Is agentic machine learning a systems discipline with durable principles, or a collection of framework wrappers, prompt templates, and API glue that will be rewritten next year?
The skepticism is fair. Much of what is written about agents concerns surface choices, such as a framework’s abstractions, a persona in a system prompt, or a ten-line loop that looks impressive in a demonstration and fails under real workloads. Those choices change every few months. What does not change is the engineering problem underneath them.
A single model call reads a context and returns a proposal. It keeps nothing between calls, takes no action, and is judged by a person who reads the answer. An agent puts that call inside a loop. The loop runs for many turns, carries state from one turn to the next, and turns the model’s proposals into edits, commands, queries, and messages that change the world. Each of those departures is an exposure a single call does not have, and each fails in its own way. A long loop compounds small errors and can run without advancing. Carried state overflows the context, goes stale, or contradicts itself. Actions leak data, repeat when a timeout hides whether they ran, or cannot be undone. The model cannot close any of these exposures by itself, because it holds no authority and cannot be trusted to judge its own results. The closure has to come from the machine built around it.
This book teaches how to design, measure, and improve that machine. Its central question is how much machine a delegated task needs around a foundation model before the result can be trusted, and how to build, measure, improve, and scale that machine. Readers should leave able to decide when a fixed workflow suffices and when a model-directed loop earns its added cost, what memory, tools, and runtime that loop requires, how to show that it works, and when changing the model or adding agents is worth the price.
The Trajectory and Its Exposures
The unit of engineering in this book is the trajectory, the ordered record of one delegated task’s execution. It holds what the model saw on each turn, what it proposed, what the runtime did with the proposal, and what came back. Accuracy on a single call is not the measure of an agentic system. The measure is whether trajectories reach results that pass an independent check, and at what cost in turns, tokens, seconds, and dollars.
Three exposures describe how far a task departs from a single model call, and the book calls them H·S·A.
- Horizon is how many turns the trajectory runs, and so how long errors have to compound and how much can go wrong between the start and the result.
- State is what the task carries from one turn to the next, from nothing, through the context window and a sandboxed workspace, to durable state that outlives the session or is shared beyond it.
- Authority is what the task may do to the world through its tools, from read access, which can still leak, to irreversible external actions such as a payment or a production migration.
Closure is not a fourth exposure. It is the runtime’s answer to the three. The invariant closure principle states that the guarantees an agentic system offers come from mechanical checks below the model, such as schemas, budgets, sandboxes, logs, and approval gates, together with evidence of completion that the model cannot edit. Measures applied to the model itself, such as instructions, fine-tuning, or a judge, lower the probability of a bad outcome; they do not bound it. Higher exposure demands more closure. The introduction develops the exposures, the principle, and the levels of evidence that can stand behind a claim that a task is done.
What This Book Covers
The book builds an agentic system in the order its dependencies impose, which is model, memory, tools, runtime, measure, learn, scale. Each part answers a failure that the part before it exposes, and each closes one exposure or climbs one rung of the intervention ladder, the discipline of fixing a measured failure with the cheapest change that addresses it before reaching for a more expensive one.
Foundations of Agentic Systems opens the book. It defines the agentic system and the trajectory, names the fail-plausible faults that make agents hard to trust, states the invariant closure principle, turns a task into a contract whose horizon, state, and authority set how much runtime it needs, and decides between a fixed workflow and a model-directed loop. It draws the whole system once, as a single map, and then leaves the analogy behind.
- Part I: The Model studies the component every agent is built around through a single call. The Foundation Model establishes what a call reads, returns, and costs, and why its output is a proposal the runtime must check rather than a result. Test-Time Compute asks when spending more inference on one decision, through longer reasoning, more candidates, or check-and-revise rounds, raises the evidence behind it.
- Part II: Agent Memory closes the state exposure. Context Engineering decides what each call sees under a fixed budget, KV Cache Management prices what holding that context costs in the serving system, and Long-Term Memory governs what persists across sessions and returns only through retrieval.
- Part III: Tool Use closes the authority exposure. Tool Calling turns a model’s tool call into a validated, authorized effect that happens exactly once, and Agent Sandboxes contains the code behind that call inside an envelope enforced below the model.
- Part IV: The Agent Runtime closes the horizon exposure. The Agent Harness runs the loop under budgets and approvals, Durable Execution lets a trajectory survive a crash, and Failure Recovery repairs or undoes bad actions. The part ends with Agent Evaluation, which measures the whole single-agent system, because nothing should be trained or scaled before it can be shown to work.
- Part V: Learning from Trajectories changes the model itself, the last and most expensive rung for a single agent. It moves no exposure. Trajectory Curation selects the trajectories evaluation verified, Trajectory Fine-Tuning trains the model to imitate them, and Reinforcement Learning from Verifiable Rewards optimizes the model against verifiers it cannot tamper with.
- Part VI: Agents at Scale multiplies all three exposures. Multi-Agent Coordination splits a task across agents only when the coordination pays for itself, and Agent Economics prices and provisions agents per accepted task.
- Part VII: Synthesis runs one trajectory through every part in Conclusion and ends at the specification boundary, where writing the check that defines success, and answering for the result, stays with the engineer.
Reference appendices collect material that supports the chapters without interrupting them: a reference architecture, tool design patterns, a failure taxonomy, the inference cost derivations behind the book’s serving estimates, the mathematical background, and a glossary.
How to Read This Book
The book is written to be read in order. Each chapter uses only ideas introduced earlier in the book or defined where they appear, and each owns its concepts, so a mechanism is taught once and cited afterward. Every chapter opens with the agent problem it solves, introduces the mechanism as the answer to that problem, and closes with a summary and a handoff to the next chapter. Classical systems ideas, such as write-ahead logs, compensating transactions, idempotency keys, and capabilities, appear where they are the real engineering, introduced as the fix for an agent failure rather than as a lesson in their own right.
Readers with a particular role can weight their reading after the introduction. Engineers who build serving and execution infrastructure will lean on the chapters on the model, the KV cache, sandboxes, durable execution, and economics. Engineers who build agent platforms will lean on context engineering, tool calling, the harness, failure recovery, evaluation, and multi-agent coordination. Researchers and post-training teams will lean on the model, test-time compute, evaluation, and all of Part V. The introduction describes these paths in more detail. Every path begins with Foundations of Agentic Systems and ends with Conclusion.
What This Book Is Not
This is not a prompt-engineering guide. Prompts matter, but the book treats them as one input the runtime assembles, and it never relies on an instruction to the model where a mechanical check is possible. It is not a tutorial for any agent framework, protocol implementation, or model provider, and it does not rank this year’s models or benchmarks. Products and benchmarks appear only as examples of a technique or a class of evaluation, with citations. It is not a book about the internals of accelerators. It states the few serving facts an agent engineer needs, such as why output tokens cost far more than input tokens and why reusing a shared prefix across turns is the largest saving, once, with the derivations collected in an appendix. Finally, it covers software agents that act through tools on digital systems. Agents that act on the physical world raise problems of their own and are outside its scope.
A Note on Durability
Frameworks, protocols, and model weights will change, some of them before this book reaches print. The text therefore emphasizes the ideas that survive those changes, such as the trajectory as the unit of work, the three exposures and the closure against them, the separation between a model that proposes and a runtime that decides, and the discipline of measuring before changing anything. Current protocols and systems appear throughout to make the ideas concrete, but the principles stand independently of any vendor or tool. Where a question is still open, such as how best to evaluate systems that act or how to structure an agent’s long-term memory, the book says so and presents the question as an open problem rather than a settled method.
Prerequisites
Readers should be comfortable with:
- machine learning fundamentals, including how transformer models are trained and how they generate text;
- computer systems fundamentals, including processes, memory, file systems, and networking;
- basic distributed systems concepts, including remote calls, failure, and consistency;
- undergraduate probability; and
- programming in Python or another modern language.
Familiarity with operating systems, databases, or computer security helps, but the book develops each idea it needs. Readers who want more background on how machine learning systems are built and served can find it in Introduction to Machine Learning Systems, which this book does not assume.