Agentic Failure Taxonomy
An agent that fails rarely crashes. It keeps calling tools that return success, reports a finished task in fluent prose, and leaves the evidence of what went wrong spread across its context, its log, and the environment it touched, which is the fail-plausible fault model of The Fail-Plausible Fault Model seen in production. This appendix classifies the recurring failure modes of agent systems by the symptoms an engineer can observe, the H·S·A exposure each one stresses, the principle it violates, and the mechanism in the main text that closes it, so that a reader facing an anomaly can name it, find the component that should have caught it, and turn it into a regression task.
How to Use This Appendix
Start from what you observe, not from what you suspect.
- When an agent keeps calling valid tools, every call succeeds, and the task does not advance, see section 1.2.
- When cost and latency per turn climb and the agent begins ignoring constraints stated early in the task, see section 1.3.
- When a reward, pass rate, or selector score rises while the work it measures does not improve, see section 1.4.
- When an agent takes an action nobody asked for after reading untrusted content, or sends data somewhere it should not, see section 1.5.
- When several agents or samples agree confidently on a wrong answer, see section 1.6.
- When the agent acts on files, records, or results that no longer exist or have already changed, see section 1.7.
- When a trajectory loses its progress after a crash, or repeats an effect after it resumes, see section 1.8.
- When one failing dependency drains budgets across many trajectories, see section 1.9.
- For a side-by-side comparison of all eight modes, see table 1, and for a step-by-step incident procedure, see section 1.10.
Failure Taxonomy Matrix
Table 1 places each failure mode by the exposure it stresses and the principle it violates. The exposure column says which part of the runtime should have closed the failure, and the principle column says which invariant was not enforced. A failure whose closure is missing is a runtime defect, to be fixed at rungs 1 to 3 of the intervention ladder (The intervention ladder); only a failure that persists with its closure in place is a candidate for training.
| Failure mode | Exposure | Principle violated | Observable signature | Primary cause | Closing mechanism |
|---|---|---|---|---|---|
| No-progress loop | Horizon | \(\ref{pri-vol3-preemptive-interrupts}\) | Turns and tokens rise; environment state repeats; every call succeeds | The model cannot see the loop from inside a growing context | Budgets and progress detection (Budgets and Ceilings, Semantic Watchdog Timers) |
| Context overflow and dilution | State (\(S_1\)) | \(\ref{pri-vol3-attention-working-set}\) | Cost per turn climbs; early constraints ignored; malformed tool calls | An append-only context with no budget, compaction, or invalidation | Context budget, compaction, and staging (Context Engineering) |
| Reward hacking and verifier gaming | None (training and selection) | \(\ref{pri-vol3-verifiable-rewards}\) | Score rises while the task outcome does not; tests edited or bypassed | A verifier the policy can reach, or one that checks less than the task requires | Isolated, sealed verifiers (Isolating the Verifier, Hermetic evaluation gyms) |
| Confused deputy and injection | Authority | \(\ref{pri-vol3-zero-trust-sandboxing}\) | Unrequested privileged calls after reading untrusted text; unexpected egress | Untrusted content steering an agent that holds authority and a way out | Capabilities, egress control, taint tracking, approval gates (Agent Sandboxes) |
| Correlated agreement | All three, multiplied | \(\ref{pri-invariant-closure}\) | Unanimous votes or debate consensus on a wrong answer | Samples and agents that share a model share its blind spots | A deterministic gate commits, not a vote (Correlated ensemble failures) |
| Environment desynchronization | State (\(S_2\), \(S_3\)) | \(\ref{pri-vol3-source-authority}\) | Actions on files or records that changed; conflicting edits; duplicated creates | Context or index copies trusted over their source; concurrent writers | Invalidation, version checks, workspace isolation (Context Invalidation, Optimistic concurrency control) |
| Lost progress and repeated effects | Horizon | \(\ref{pri-vol3-intent-before-effect}\) | Restart from the beginning after a crash; an effect applied twice on resume | Progress held only in memory; effects dispatched before their intent was durable | Trajectory log, write-ahead intent, settlement (Durable Execution) |
| Retry storm | Authority and Horizon | \(\ref{pri-vol3-exactly-once-settlement}\) | One dependency’s failures multiply across trajectories; budgets drain | The model retries a failing service and every trajectory does the same | Circuit breaker with a typed observation (Tool Circuit Breakers) |
No-Progress Loops
A no-progress loop is a trajectory that keeps working without advancing its task. The model keeps proposing well-formed tool calls, the tools keep returning success, and the environment keeps returning to states it has already been in. Unlike a hang, which a liveness check catches because activity stops, a loop fails through activity that goes nowhere, and every per-call check passes (principle \(\ref{pri-vol3-preemptive-interrupts}\)).
Observable symptoms
- Spend without change. Turns, tokens, and dollars rise while the environment (the diff, the test results, the database records) stays the same or alternates between two states.
- Action flapping. The agent applies an edit, sees a different failure, reverts the edit, sees the original failure, and applies the edit again. Semantic Watchdog Timers opens on this pattern.
- Success codes everywhere. Every tool call exits cleanly and every model call returns promptly, so health checks and crash reporting see nothing wrong.
- Rising cost per turn. Because each turn re-sends a context that the loop keeps enlarging, the cost of an undetected loop grows with the square of its length (Semantic Watchdog Timers).
Causes
The model cannot see the loop from inside its context. Each turn’s context differs from the last because it carries new reasoning and new error text, even when the workspace has returned to an earlier state, so the proposal that led nowhere before looks new again. Compaction that discards the record of a failed attempt removes the one piece of evidence that would have told the model it had already tried this path (Context Compaction). Uninformative error observations, such as a bare nonzero exit code, give the next proposal nothing to change (Terminal Output Sanitization).
Closing mechanisms
- Budgets as the backstop. Turn, token, time, and cost ceilings on the trajectory end any loop eventually, tightening from a steer to read-only operation to a halt (Budgets and Ceilings). A budget bounds the damage but does not detect the loop early.
- Progress detection from the environment. The runtime hashes each action together with the state it acted on and flags repeats, and it measures progress by changes in environment state, such as tests that newly pass, rather than by the model’s own account (Semantic Watchdog Timers). Lyapunov Liveness Verification gives the formal version, a quantity that must decrease every turn.
- Compaction that keeps failures. Compaction must keep each failed attempt’s identifiers and error verbatim, so that the record of what was tried survives the summary (Context Compaction).
- Repair through the model, then around it. A detected loop is answered first by an observation that names the repetition and by rolling back the poisoned branch of context, then by resampling or a different model, and finally by escalation (Forward Recovery). Raising the sampling temperature is not a fix. It trades a deterministic loop for a random walk that the same detector must still catch.
Context Overflow and Dilution
Context overflow is the failure of a trajectory whose context has outgrown what the model can use. It has two halves. The cost half is mechanical. Every call re-sends the whole context, so a context that grows by \(\Delta\) tokens per turn from an initial \(L_0\) re-sends \(T L_0 + T(T-1)\Delta/2\) tokens over \(T\) turns, and the time to the first output token grows with the input (Working Set Capacity). The quality half is statistical. Models retrieve and combine evidence placed in the middle of a long context less reliably than evidence at its ends, so the usable context is shorter than the nominal window (Working Set Capacity).
Observable symptoms
- Cost and latency per turn climb. Input tokens per call rise turn over turn, and the prefix-cache hit rate falls if anything near the top of the context changes.
- Early constraints are ignored. The agent follows the most recent tool output and the system prompt but violates a constraint stated in an earlier turn, such as a forbidden path.
- Tool calls degrade. Arguments go missing or take the wrong type as tool definitions and parameter descriptions compete with large observations for the model’s attention.
- Serving memory pressure. Long-lived trajectories hold large attention state, and fewer of them fit on a server at once (From Context Tokens to KV State).
Causes
The context is treated as an append-only transcript. A shell command that dumps a ten-thousand-line build log, a file read that returns a whole module, or a web page with its markup intact enters the context in full and stays there for the rest of the trajectory, paid for on every turn and diluting the evidence that matters.
Closing mechanisms
- A budget, allocated. The context budget divides a working set among the system prompt, tool definitions, task, history, retrieved evidence, and output reserve (Working Set Capacity; principle \(\ref{pri-vol3-attention-working-set}\)).
- Shaping at the source. Long tool output is truncated, paginated, or reduced to a structured result before it enters the context, with markers that say what was cut (Observation Stream Truncation).
- Compaction and progress notes. Old tool results are cleared or summarized, with identifiers and failures kept verbatim, and durable state is extracted into structured notes before anything is evicted (Context Compaction, Context Checkpointing).
- Retrieve instead of retain. Material the agent may need later belongs in a store it can search, not in the context it re-sends (Agentic Retrieval).
- Serving-side relief. When the pressure is memory rather than quality, the serving layer evicts, recomputes, or offloads attention state during tool waits and shares prefixes across branches (Retain, Evict, Recompute, or Offload, Paged KV Allocation).
Reward Hacking and Verifier Gaming
Reward hacking is a policy raising its measured reward through behavior that does not accomplish the task the reward was meant to measure (Training Pathologies). The same failure appears without training whenever a selector chooses among candidates by a score: the more candidates the selector sees, the more likely it picks one that satisfies the score’s gaps rather than the task (Process Verification). Both are Goodhart’s law applied to a verifier (Goodhart 1984).
Observable symptoms
- Score and outcome diverge. Training reward or best-of-\(n\) selection scores rise while held-out evaluation on sealed tests stays flat or falls.
- Tests edited or bypassed. Candidate diffs modify test files, weaken assertions, skip tests, or special-case the test inputs.
- Grader probing. Trajectories read the grading scripts, the test runner’s configuration, or its environment before producing a solution.
- Degenerate outputs under learned rewards. Where a learned reward model or judge supplies the reward, outputs drift toward the patterns it scores highly, such as length or confident phrasing, rather than toward correctness.
Causes
The verifier is reachable or incomplete. A verifier that runs where the policy can write, or whose tests live in the policy’s workspace, can be satisfied by changing the verifier. A verifier that checks less than the task requires, such as a test suite that never exercises the failure being fixed, can be satisfied by work that does not do the task. Learned verifiers add a third cause, because they have their own errors that optimization finds (principle \(\ref{pri-vol3-verification-asymmetry}\)).
Closing mechanisms
- Isolate the verifier. Run verification outside the policy’s envelope, from storage the policy cannot write, on a state snapshot taken after the policy’s last action (Isolating the Verifier, Hermetic evaluation gyms).
- Reject diffs that touch the tests. The test-tampering defense admits no trajectory whose changes reach the files that judge it (Verifier gaming defense).
- Prefer deterministic checks at the gate. Compilers, type checkers, and sealed tests decide what is committed or rewarded; learned rankers and judges order candidates beneath that gate (Learned rankers beneath the gate, Grading trajectories).
- Watch the gap. Track sealed held-out performance alongside training reward, and treat a widening gap as a verifier defect to fix before training continues (Training Pathologies).
Confused Deputy Attacks and Prompt Injection
A confused deputy is an agent that holds legitimate authority and is steered by someone who does not hold it into using that authority on their behalf. For agents the steering usually arrives as indirect prompt injection (Indirect prompt injection): text inside a web page, email, issue, or tool result that the model reads as instruction. The threat becomes an incident when three things meet, access to private data, exposure to untrusted content, and a channel to send data out (The Agent Threat Model).
Observable symptoms
- Unrequested privileged calls. An agent summarizing an email or triaging an issue calls a tool that writes, sends, or deletes, and nothing in the task asked for it.
- Unexpected egress. Outbound requests go to hosts the task never named, sometimes carrying file contents or credentials in query strings, headers, or name lookups.
- Borrowed voices in the reasoning. The trajectory cites instructions that no operator gave, such as a claimed override embedded in retrieved text.
- Cross-tenant access. An agent acting for one tenant reads or changes another tenant’s data after ingesting content from a shared resource.
Causes
The model has no reliable way to tell instructions from data, because both arrive as tokens in the same context. Framing and provenance labels lower the probability that injected text steers a proposal, but they cannot guarantee it. The failure becomes damage when the agent also holds authority the task did not need, credentials it can read, or a network path it can use.
Closing mechanisms
- Label provenance in the context. Untrusted text is staged in marked zones with its source, which lowers the probability that the model treats it as instruction (Staging the Next Invocation; Ingress quarantining for retrieved content).
- Grant only what the task needs. A capability names the rights, objects, argument bounds, and expiry of each grant, and attenuates as work is delegated, so a steered proposal can reach only what the task could (Capabilities and Credentials, Attenuated capability delegation).
- Keep secrets out of reach. A credential broker injects scoped credentials at the point of use, so no secret appears in the context or the workspace (Credential brokering).
- Close the exit. Egress defaults to deny, allows only the hosts the contract names, and covers side channels such as name resolution and metadata endpoints (Egress Control).
- Track taint. Data derived from untrusted sources carries a mark, calls with authority refuse marked arguments, and a quarantined reader that holds no tools processes untrusted text into typed fields for the agent that does (Tracking Untrusted Data).
- Hold the irreversible step. An \(A_3\) call waits at an approval gate, where it is re-checked against its preconditions before it runs (Approval Gates).
Only the mechanisms below the model (capabilities, the broker, egress control, taint, and the approval gate) bound the damage. Provenance labels change what the model is likely to propose, which is useful, but they are not a guarantee.
Environment Desynchronization
Environment desynchronization is the failure of an agent whose picture of the environment, carried in its context, its retrieval indexes, or its memory, no longer matches the environment itself. The agent then acts correctly on a world that no longer exists.
Observable symptoms
- Acting on ghosts. The agent reads, edits, or deletes a path that an earlier turn moved or removed, relying on a directory listing still in its context.
- Redundant creation. The agent creates a branch, table, or file that already exists and derails on the error.
- Retrieval that contradicts the tools. A search returns the old version of code the agent just edited, and the agent concludes its edit failed and applies it again (What Must Persist).
- Conflicting writers. Two agents working in the same workspace overwrite each other’s edits and produce failures neither can reproduce alone.
- Late results. A tool result arrives after its deadline and is taken as the answer to a later call.
Causes
A copy is trusted over its source. Observations in the context, entries in a retrieval index, and notes in agent-written memory are all derived from the environment at some moment, and nothing ties them to the source’s later changes unless the runtime does (principle \(\ref{pri-vol3-source-authority}\)). Shared workspaces without isolation, and tool calls retried without settlement, add changes the agent did not observe.
Closing mechanisms
- Invalidate staged observations. Bind each staged observation to its source and retire it, or replace it with a tombstone, when that source changes (Context Invalidation).
- Check derived copies against the source. The retrieval contract checks each result’s version and provenance before it is staged, and index entries are invalidated on source writes (The Retrieval Contract, Storage Invalidation).
- Check versions on write. An edit carries the version hash of the file it read, and a mismatch fails the edit with a structured observation that tells the model to re-read (Optimistic concurrency control).
- Isolate workspaces. Give each trajectory, and each concurrent agent, its own leased workspace, merged back through a verification gate (Workspace leases per trajectory, reset per tenant, Optimistic concurrency control).
- Settle before retrying and reject late results. Mutating calls carry idempotency keys and are settled after an ambiguous timeout (Idempotent Action Execution), and the harness accepts a result only for a call ID still pending (Trajectory States).
Lost Progress and Repeated Effects
Lost progress is the failure of a long trajectory that cannot survive the loss of the process running it. Its twin, the repeated effect, appears when a naive restart does survive, by replaying from memory or asking the model again, and dispatches an effect that had already happened.
Observable symptoms
- Restart from the beginning. After a worker crash, a deployment, or a preemption, the trajectory starts over, repeating minutes or hours of model calls.
- Duplicate external effects. A resumed trajectory sends a second email, opens a second pull request, or charges twice.
- Resumed trajectories that behave differently. Re-invoking the model to rebuild progress produces different proposals from the ones actually executed, because sampling and serving are nondeterministic (Why Replays Diverge).
- Approvals that vanish. A trajectory paused for approval is lost when its host restarts, and the approval arrives for a trajectory that no longer exists.
Causes
Progress lives only in memory, or effects leave the host before the record that authorizes them is durable. In the first case nothing survives the crash. In the second, the log cannot say whether an effect happened, and a retry on the assumption that it did not is how duplicates occur (principle \(\ref{pri-vol3-intent-before-effect}\)).
Closing mechanisms
- Log events, fold the record. Keep the trajectory as an append-only log of typed events and rebuild the record from it on any worker (The Trajectory Log).
- Write intent before effect. Make the intent durable before dispatch, so every effect is either known to have been intended or known not to have happened (Intent Before Effect).
- Settle what is in doubt. On resume, an intent with no recorded outcome is settled through its idempotency key or a reconciliation probe, never re-dispatched on the assumption that it failed (Idempotent Action Execution, Resuming a Trajectory).
- Checkpoint what the log cannot rebuild cheaply. Snapshot the workspace and bind it to a log position, at a cadence set by the cost of a checkpoint and the rate of failures (Checkpoints).
- Replay, do not regenerate. Rebuild state from recorded model replies and tool results, with the model and tools stubbed (Replay).
Retry Storms
A retry storm is the failure of a fleet of trajectories that all depend on one failing service. Each trajectory’s model reads the error, proposes the call again with small variations, and spends its budget doing so, and every trajectory does the same at once, which keeps the service down.
Observable symptoms
- Correlated budget exhaustion. Many trajectories that share a dependency end on budget ceilings within the same window.
- Rephrased retries. Successive calls to the failing tool differ only in wording or argument order, because the model treats a service failure as a problem with its query.
- Load that prevents recovery. Request rates to the failing service rise after it starts failing.
Causes
The model is a retry amplifier. A transport failure looks to it like any other error observation, and the likeliest continuation of a failed call is another attempt. Without a runtime rule, a service outage turns into as many retry loops as there are trajectories using the service.
Closing mechanisms
- Trip a circuit breaker in the gateway. Count transport-level failures per dependency, stop dispatching when they cross a threshold, and probe with a single call after a cooldown (Tool Circuit Breakers).
- Return a typed observation. An open breaker answers at once with an observation that names the failing service, says that rewording will not help, and says when to try again, which makes it likely the model moves on (Tool Circuit Breakers).
- Settle before retry. A mutating call that timed out is settled, not simply resent, so the storm does not also duplicate effects (Idempotent Action Execution; principle \(\ref{pri-vol3-exactly-once-settlement}\)).
- Fall back or escalate. Route the step to an alternative tool or model when one exists, or pause the trajectory until the dependency recovers (Forward Recovery).
Incident Triage Workflow
When an agent system degrades in production, the same sequence isolates the cause quickly and turns the incident into a test. Table 2 lists it. The first step is always to secure the evidence, because a trajectory that has not been logged cannot be replayed, and a failure that cannot be replayed cannot be attributed.
| Step | Phase | Diagnostic action | Failure modes it separates | Evidence it produces |
|---|---|---|---|---|
| 1 | Secure the evidence | Retain the trajectory log and trace; pause or quarantine the trajectory if it holds authority | All | A complete log and a frozen grant set (Blast Radius Quarantine) |
| 2 | Check the boundary | Compare granted capabilities, egress, and tainted inputs with the calls actually made | Confused deputy (section 1.5) | Capability and egress records; taint marks on arguments |
| 3 | Check progress | Look for repeated action hashes, unchanged environment state, and budget burn without state change | No-progress loop (section 1.2), retry storm (section 1.9) | Progress signals and per-dependency failure rates |
| 4 | Check the context | Measure input tokens per turn, cache-hit rate, and where the violated constraint sat in the context | Context overflow (section 1.3) | Per-turn token accounting from the trace |
| 5 | Check the world | Compare what the context and indexes said with the source; look for late results and concurrent writers | Environment desynchronization (section 1.7) | Version mismatches, rejected late results, write conflicts |
| 6 | Check durability | Confirm each external effect has a durable intent and a settled outcome | Lost progress and repeated effects (section 1.8) | Intents without outcomes; repeated idempotency keys |
| 7 | Check the verifiers | Confirm the tests and graders were unreachable and unmodified, and compare scores with sealed held-out results | Reward hacking (section 1.4), correlated agreement (section 1.6) | Verifier hashes; score versus held-out gap |
| 8 | Replay and attribute | Replay the trajectory from its log with the model and tools stubbed, then ablate one layer at a time | Model versus harness versus environment | An attribution and a regression task (Forensic incident post-mortems) |
The last step is the one that pays for the others. A failure attributed to the harness or the environment is fixed there and added to the evaluation suite as a regression task. Only a failure that persists with every closure in place, and that the replay attributes to the model, is evidence for climbing the intervention ladder toward training (Capability gap diagnosis).