Author’s Note

When an AI answers a question, the interaction is over. When an AI takes an action, the real engineering has only just begun.

A prompt-and-response model is an oracle behind glass: it answers when spoken to, incurs no permanent consequence, and resets its state the moment the session ends. But when you give that intelligence tools, memory, and a delegated goal, it ceases to be an oracle. It becomes an actor.

My curiosity about agentic systems began during a National Science Foundation (NSF) workshop on AI efficiency. As someone whose work has long focused on benchmarking and quantitative systems evaluation, including the initialization of MLPerf, I found myself asking a fundamental systems question: how do we benchmark and evaluate the efficiency of these actors?

At the algorithmic and application layers, the community has introduced a vibrant lexicon of techniques: “prompt engineering,” “prompt chaining,” “in-context reflection,” and “agentic workflows.” These ideas are genuinely creative and impactful. They show us what is possible when models interact iteratively with task environments. But as a computer systems architect, I found myself asking a complementary question: what is the fundamental machine required to build, benchmark, and scale these ideas? When you hear these exotic-sounding terms at the application level, what is the actual system you have to construct underneath?

To measure and engineer these workflows, we have to break them down to their fundamentals. And when you look at them from first principles, an agentic system is not an alien artifact. It is a machine built around a foundation model, and its parts are ones a systems engineer recognizes.

  • The foundation model reads a context and proposes the next action.
  • Agent memory decides what each call sees and keeps what must outlast the session.
  • Tools turn an approved proposal into an effect on the world, inside a sandbox.
  • The agent runtime validates every proposal, records every step, and recovers when something goes wrong.

The introduction draws this machine once, as a single map that I call the Stochastic Computer. After that, the book argues each part in its own terms, because a model is not a processor and a context window is not a cache.

Yet this machine carries a profound systems twist: it is stochastic. A classical program executed a thousand times produces the same result a thousand times. A foundation model samples from a distribution. Ninety-nine times out of a hundred its proposal is the one you intended, and then, governed by the long tail of probability, it proposes something else entirely.

I experienced this firsthand while drafting these chapters. Working alongside an autonomous coding assistant with direct terminal access, the agent took an unexpected turn and executed an rm -f directly on the directory containing my active edits. To compound the disaster, in an unthinking moment of late-night muscle memory, I committed and pushed the changes upstream. In an instant, days of work across a holiday weekend evaporated into git history.

You would think a computer systems professor would know better than to give an unconstrained actor raw effector privileges without an isolated sandbox, or to blindly push a tree without auditing the diff. But that is the precise nature of the trap: autonomy is seductive right up until an irreversible mutation occurs.

The introduction returns to this incident and classifies it (\(\ref{exmp-01-hsa-incident}\)). The task needed only reversible authority over a copy of my files, and the environment granted irreversible authority over the only one.

In a deterministic world, reliability is primarily a software debugging problem. When the component proposing every action is probabilistic, reliability becomes an architectural imperative. We cannot rely on human vigilance or prompt cleverness alone to guarantee safety. The runtime around the model must bound what each call sees, contain what each action can reach, record every step durably, and recover when an action goes wrong.

I write these chapters as a student of a discipline in motion. Putting this book together, shuttling between Cambridge and Zurich, has been my way of distilling the durable first principles that will still govern these systems a decade from now, the three exposures of horizon, state, and authority and the closure the runtime builds against them.

I hope you find it helpful.

— Vijay Janapa Reddi
Cambridge, Massachusetts & Zurich, Switzerland
2026

Back to top