Typed Tool Interfaces
Tool Calling explains why every tool call must pass through the runtime’s validation and grant before it takes effect, and Agent Sandboxes explains how a granted call is contained once it runs. This appendix collects the reusable artifacts an engineer needs to build that boundary: schema rules and a reference tool definition, the break-even for constraining output to a schema, reference code for idempotency keys and settlement at the endpoint, reference code for the capability check and its attenuation, and the arithmetic behind a small catalog of general tools. It is a workbench for building tool interfaces, and each section points back to the chapter that argues for the design.
How to Use This Appendix
Reach for the section that matches the failure you are seeing at the tool boundary.
- When tool calls fail validation, invent arguments, or confuse units, use the schema rules and reference definitions in section 1.1.
- When the model emits malformed calls often enough that repair turns cost more than prevention, use the break-even in section 1.2 to decide whether to constrain decoding.
- When timeouts, dropped connections, or model retries risk duplicate side effects, use the key derivation and endpoint record in section 1.3.
- When a tool call or a delegated subagent needs bounded authority, use the capability check and attenuation code in section 1.3.
- When the tool catalog consumes a large share of every context or the model picks the wrong tool, use the token arithmetic in section 1.4.
Schema Rules for Tool Definitions
A tool schema is read twice. The model reads its names and descriptions as part of the context and proposes arguments from them, and the runtime reads its types and bounds to validate every proposed call before the grant (Tool Interface Schemas). A schema that serves only one reader fails the other. Five rules make a schema serve both.
- Close the property set (
additionalProperties: false). A model will propose plausible arguments the tool does not accept, such asdry_run,verbose, oruser_id, if the schema permits them. A closed schema makes the runtime reject undeclared inputs instead of silently passing them through, and it lets a grammar-constrained decoder exclude them before they are generated (Grammar-Guided Decoding). - List every required field in
required. A description that says an argument is required is a suggestion to the model; therequiredarray is a rule the validator enforces. - Bound every type. Constrain numbers with
minimumandmaximum, strings withpattern,enum, ormaxLength, and arrays withminItemsandmaxItems. Unbounded inputs let one call carry an argument larger than the tool, or the context that records it, can hold. - Put units and conventions in names and descriptions. Write
timeout_msrather thantimeout, state whether line numbers start at 1, and say what a missing optional field means. - Tag every union. An
anyOforoneOfwithout a discriminator field leaves the model to guess which branch it is filling and leaves the validator unable to report which branch failed.
The same schema also carries the side-effect annotation that tells the runtime how to treat the call (read-only, idempotent, or destructive), which maps onto the authority levels \(A_0\) to \(A_3\) and decides whether the call needs an idempotency key, a compensator, or an approval (Tool Interface Schemas).
A reference tool definition
Listing 1 defines an exact-replacement edit tool that follows all five rules.
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"name": "file_edit",
"description": "Replaces an exact block of text in an existing file. old_str must match exactly one contiguous block unless allow_multiple is true; if it matches zero blocks, or more than one when allow_multiple is false, the file is left unchanged and the call returns an error naming the match count.",
"parameters": {
"type": "object",
"properties": {
"path": {
"type": "string",
"description": "Absolute path to the target file, inside the workspace root.",
"pattern": "^/([a-zA-Z0-9_.-]+/)*[a-zA-Z0-9_.-]+$",
"maxLength": 1024
},
"old_str": {
"type": "string",
"description": "The exact block to replace, including indentation and newlines.",
"minLength": 1
},
"new_str": {
"type": "string",
"description": "The replacement text. May be empty to delete old_str."
},
"allow_multiple": {
"type": "boolean",
"description": "If true, replace every match; if false, fail unless exactly one match exists.",
"default": false
}
},
"required": ["path", "old_str", "new_str"],
"additionalProperties": false
}
}Writing schemas by hand invites drift between the schema the model sees and the code that runs. Generating the schema from a typed model of the arguments keeps them in step. Listing 2 shows one way, using a Python data-validation library as the example; any typed schema generator serves the same purpose.
from typing import Annotated
from pydantic import BaseModel, ConfigDict, Field, StringConstraints
class FileEditInput(BaseModel):
"""Arguments for the exact-replacement edit tool."""
model_config = ConfigDict(extra="forbid", frozen=True)
path: Annotated[
str,
StringConstraints(strip_whitespace=True, max_length=1024,
pattern=r"^/([a-zA-Z0-9_.-]+/)*[a-zA-Z0-9_.-]+$"),
Field(description="Absolute path to the target file, inside the workspace root."),
]
old_str: Annotated[str, StringConstraints(min_length=1),
Field(description="Exact block to replace.")]
new_str: Annotated[str, Field(description="Replacement text; may be empty.")]
allow_multiple: Annotated[bool, Field(default=False,
description="Replace every match instead of exactly one.")]
def tool_descriptor() -> dict:
"""The descriptor the runtime publishes to the model and validates against."""
return {
"name": "file_edit",
"description": "Replaces an exact block of text in an existing file.",
"inputSchema": FileEditInput.model_json_schema(),
}Schema anti-patterns
Tools adapted from existing service interfaces often carry patterns that suit a programmer calling from code but mislead a model proposing from a description. Table 1 lists the common ones.
| Anti-pattern | What goes wrong | Hardened form | Effect |
|---|---|---|---|
Untyped argument bag (args: dict) |
The model invents keys; the validator has nothing to check | Explicit typed fields with additionalProperties: false |
Invalid calls are rejected, or never generated under constrained decoding |
Unbounded inputs (string with no limit) |
One call carries an argument larger than the tool or the context can hold | maxLength, maxItems, and explicit windowing parameters |
Argument size is bounded before dispatch |
Untagged union (anyOf with no discriminator) |
The model mixes fields from several branches; errors cannot name the branch | A discriminator field, such as "kind" with values "file" or "dir" |
Each call selects one branch explicitly |
Implicit units (timeout: 30) |
Seconds, milliseconds, and minutes are confused | Units in the name (timeout_ms) with range bounds |
Unit errors become validation errors |
Whole-file rewrite (content: str) |
Output length grows with the file, long files are truncated, and unrelated lines drift | Exact replacement of a block (old_str, new_str) |
Output tokens scale with the change, not the file |
The last row matters most for latency, because output tokens are the expensive ones on an agent’s critical path (Accelerator Serving Latency). A whole-file rewrite pays decode time for every unchanged line, and Interface Benchmarking works an example in which the rewrite misses a per-call deadline that an anchored diff meets.
When to Constrain Decoding
A model that samples freely can emit a malformed call: an unescaped newline, a trailing comma, a missing required field. Grammar-constrained decoding prevents that by restricting each sampling step to the tokens the schema’s grammar admits in its current state (Grammar-Guided Decoding), Decode-Loop Mechanics sizes the mask tables and explains why the mask must run beside the sampler, and Deterministic Finite Automata and Grammar-Constrained Logit Decoding proves that the result always parses. What this section adds is the decision of when constraining pays, because the alternative, letting the malformed call fail validation and asking the model to repair it, has a cost of its own.
Let a tool call carry \(L_{\text{gen}}\) output tokens at \(t_{\text{decode}}\) per token, let masking add \(t_{\text{mask}}\) per token, let \(p_{\text{err}}\) be the rate at which unconstrained calls fail to parse, and let a repair turn cost \(T_{\text{repair}}\) of latency, including the re-sent context and the regenerated call. The expected latency without constraints, allowing one repair, is
\[\mathbb{E}[T_{\text{free}}] = L_{\text{gen}} \, t_{\text{decode}} + p_{\text{err}} \, T_{\text{repair}}\]
and with constraints it is
\[T_{\text{constrained}} = L_{\text{gen}} \, (t_{\text{decode}} + t_{\text{mask}})\]
Setting the two equal gives the break-even error rate
\[p_{\text{break}} = \frac{L_{\text{gen}} \, t_{\text{mask}}}{T_{\text{repair}}} \tag{1}\]
Constraining pays whenever the observed malformation rate exceeds \(p_{\text{break}}\). The rate is a property of a particular model, schema, and prompt, and must be measured on the tool’s own traffic. As an illustration, with an 80-token call, a mask overhead of 0.5 ms per token, and a 1.8-second repair turn, equation 1 gives \(p_{\text{break}} = 40\ \text{ms} / 1{,}800\ \text{ms} \approx 2.2\) percent, so a tool whose calls fail to parse more often than about one time in forty is cheaper to constrain. Latency is only part of the case. A repair turn also adds error text to the context, which makes the next proposal about the same object more likely to repeat the mistake (Context poisoning dynamics).
Constraining guarantees syntax and nothing else. A call that parses can still name a file that does not exist, pass an argument outside the task’s authority, or do the wrong thing correctly, which is why validation, the grant, and the envelope stay in place behind it (Grammar-Guided Decoding).
Idempotency Keys and Capabilities
Two artifacts make a granted call safe to send. An idempotency key makes a repeated call harmless, so a runtime that cannot tell whether a timed-out call ran can settle it instead of guessing (Idempotent Action Execution). A capability bounds what the call can reach once it runs, so a steered or mistaken proposal can do no more than the task allows (Capabilities and Credentials). The chapters argue for both. This section gives the reference form of each.
Deriving the key
The key must identify the logical call, not the attempt, so the runtime derives it deterministically from the trajectory. Idempotent Action Execution uses
\[k_{\text{idem}} = \mathcal{H}\big(\text{trajectory\_id} \,\|\, \text{turn} \,\|\, \text{call\_id} \,\|\, \text{tool} \,\|\, \text{canonical\_args}\big)\]
where \(\mathcal{H}\) is a cryptographic hash, or a keyed hash when keys must not be guessable by the endpoint’s other clients. The arguments must be serialized canonically before hashing. Ordinary JSON serialization does not fix key order or whitespace, so {"path": "/a", "force": true} and {"force": true, "path": "/a"} would hash to different keys and defeat deduplication. A canonical form, such as the JSON Canonicalization Scheme of RFC 8785, sorts keys and fixes number and string encoding. Because the key is derived rather than drawn at random, a runtime that recovers from a crash and re-sends a call regenerates the same key, and the endpoint recognizes the repeat (Intent Before Effect).
The endpoint record
At the endpoint, each key maps to a record with one of four states, and the transitions between them are what make a retry safe.
- Unseen. The endpoint claims the key atomically, with a set-if-absent operation and a lease that outlives the worst-case execution time plus the caller’s timeout and any clock skew (Idempotent Action Execution), and then executes the call.
- In flight. A repeat arriving while the first execution runs is refused or made to wait. It is never executed in parallel.
- Completed. A repeat of a completed call returns the stored result without executing again.
- In doubt. If execution fails in a way that leaves its effect unknown, such as a crash of the executor partway through, the record stays claimed and is marked in doubt. Repeats are refused until a reconciliation probe inspects the environment and settles the record as completed or not executed.
The fourth state is the one naive implementations omit. Releasing the key after an exception invites the next retry to execute again, which is exactly the duplicate the key exists to prevent, because an exception does not prove that the effect did not happen. Listing 3 implements all four states.
import hashlib, json, time
from typing import Any, Callable, Dict
def canonical(args: Dict[str, Any]) -> str:
"""Stable serialization: sorted keys, fixed separators."""
return json.dumps(args, sort_keys=True, separators=(",", ":"), ensure_ascii=True)
def derive_key(trajectory_id: str, turn: int, call_id: str, tool: str,
args: Dict[str, Any]) -> str:
material = f"{trajectory_id}|{turn}|{call_id}|{tool}|{canonical(args)}"
return hashlib.sha256(material.encode("utf-8")).hexdigest()
class DedupEndpoint:
"""Executes each logical call at most once per idempotency key."""
def __init__(self, store, lease_s: float):
self.store, self.lease_s = store, lease_s # store offers set_if_absent/get/put
def handle(self, key: str, execute: Callable[[], Dict[str, Any]]) -> Dict[str, Any]:
claimed = self.store.set_if_absent(
key, {"state": "IN_FLIGHT", "until": time.time() + self.lease_s})
if not claimed:
record = self.store.get(key)
if record["state"] == "COMPLETED":
return record["result"] # repeat: stored result
return {"status": record["state"], "retry": True} # IN_FLIGHT or IN_DOUBT
try:
result = execute()
except Exception as exc: # effect unknown
self.store.put(key, {"state": "IN_DOUBT", "error": str(exc)})
return {"status": "IN_DOUBT", "retry": False} # settle by probe first
self.store.put(key, {"state": "COMPLETED", "result": result})
return resultA tool whose endpoint cannot keep such a record needs the other settlement path, a reconciliation probe that reads the environment to decide whether the effect happened before any retry (Idempotent Action Execution).
The capability check
Capabilities and Credentials writes a capability as \(C = (R, O, P, E)\), signed with a key only the runtime holds: \(R\) is the set of rights (read, write, execute, connect), \(O\) the objects the capability designates (path prefixes, host names, repositories), \(P\) a predicate that bounds arguments and cost, and \(E\) the expiry. A capability issued for one trajectory and bounded by \(E\) is that trajectory’s lease on the resources it names (Workspace leases per trajectory, reset per tenant). The runtime approves a proposal only if the signature verifies, the operation is among the rights, the target lies within the objects, the cost satisfies the predicate, and the capability has not expired.
When work is delegated, a child’s capability must attenuate from its parent’s. Writing \(C_c \preceq C_p\) for “the child holds no more than the parent”,
\[C_c \preceq C_p \iff R_c \subseteq R_p \;\land\; O_c \sqsubseteq O_p \;\land\; P_c \Rightarrow P_p \;\land\; E_c \le E_p\]
so a child can never hold a right its parent lacks, reach an object outside its parent’s scope, spend beyond its parent’s bound, or outlive its parent’s grant (Attenuated capability delegation; principle \(\ref{pri-vol3-monotonic-delegation}\)). Listing 4 implements the check and the attenuation, with the predicate reduced to a cost ceiling for brevity.
import hashlib, hmac, time
from dataclasses import dataclass
from typing import FrozenSet
@dataclass(frozen=True)
class Capability:
rights: FrozenSet[str] # R: e.g. {"read", "write"}
objects: FrozenSet[str] # O: path or host prefixes
max_cost: float # P: reduced here to a cost ceiling
expires_at: float # E: epoch seconds
sig: str = ""
def body(self) -> bytes:
return repr((sorted(self.rights), sorted(self.objects),
self.max_cost, self.expires_at)).encode()
def sign(cap: Capability, key: bytes) -> Capability:
return Capability(cap.rights, cap.objects, cap.max_cost, cap.expires_at,
hmac.new(key, cap.body(), hashlib.sha256).hexdigest())
def within(target: str, objects: FrozenSet[str]) -> bool:
return any(target == o or target.startswith(o.rstrip("/") + "/") for o in objects)
def approve(cap: Capability, key: bytes, op: str, target: str, cost: float) -> bool:
good_sig = hmac.compare_digest(
cap.sig, hmac.new(key, cap.body(), hashlib.sha256).hexdigest())
return (good_sig and op in cap.rights and within(target, cap.objects)
and cost <= cap.max_cost and time.time() < cap.expires_at)
def attenuate(parent: Capability, key: bytes, *, rights, objects,
max_cost, expires_at) -> Capability:
"""Issue a child capability; refuse any component wider than the parent's."""
if not (set(rights) <= parent.rights
and all(within(o, parent.objects) for o in objects)
and max_cost <= parent.max_cost and expires_at <= parent.expires_at):
raise PermissionError("attenuation may only narrow a capability")
return sign(Capability(frozenset(rights), frozenset(objects),
max_cost, expires_at), key)The check runs below the model, in the runtime and the sandbox’s drivers, on every access. Nothing the model writes can widen a capability, because the model never holds the signing key, and a proposal that names a target outside the capability fails before any host operation runs.
A Small Catalog of General Tools
A natural first catalog wraps each endpoint of an existing service as its own tool: git_status, git_diff, git_add, git_commit, and so on across every system the agent touches. The result is dozens or hundreds of narrow tools. Ousterhout calls a module shallow when its interface is large relative to the functionality behind it (Ousterhout 2018), and a catalog of narrow tools is shallow in exactly that sense, with a cost an agent pays on every turn. Toolkit Granularity Partitioning argues the design choice. This section gives the arithmetic.
The token cost of a catalog
Every tool’s definition is sent with every call that may use it. With \(K\) tools whose definitions average \(S\) tokens, a trajectory of \(N\) turns sends
\[T_{\text{catalog}} = N \cdot K \cdot S \tag{2}\]
tokens of tool definitions. As an illustration, a catalog of 80 narrow tools at about 350 tokens each costs 28,000 tokens per turn and 1,120,000 over a 40-turn trajectory. Four general tools at about 250 tokens each cost 1,000 tokens per turn and 40,000 over the same trajectory, a reduction of about 96 percent. Prefix caching lowers the price of the re-sent definitions as long as the catalog stays byte-identical and in a fixed position, but it does not lower their share of the context budget, and any change to the catalog mid-trajectory discards the cached prefix (Staging the Next Invocation). Tool selection also degrades as the catalog grows and definitions overlap, a trend that function-calling benchmarks measure by varying catalog size (Toolkit Granularity Partitioning).
Table 2 sets the two catalog shapes side by side on these costs and on where authority is enforced.
| Dimension | Many narrow tools | A few general tools | Consequence |
|---|---|---|---|
| Interface surface | Tens to hundreds of endpoint wrappers | A handful of general operations | Fewer definitions for the model to read and choose among |
| Definition tokens | Grows with \(K\) on every turn (equation 2) | Small and stable | More of the context budget left for the task |
| Selection | Overlapping names and arguments invite the wrong tool | Few, distinct choices | Fewer wrong-tool calls; more depends on each tool’s description |
| Observations | A different result shape per tool | One result shape (output, error, exit status) | One observation-shaping path in the gateway |
| Enforcement | Authorization logic scattered across many handlers | Authority enforced by the envelope and capability | One place to enforce and audit, at the price of broader tools |
The last row is the cost of the general design. A general shell tool can do far more than a narrow one, so it moves the burden of bounding each action from the tool’s code to the isolation envelope and the capability it runs under (Agent Sandboxes). A general catalog is safe only inside an envelope sized for the broadest thing its tools can do.
A reference catalog for coding agents
Many coding agents converge on four general operations. A shell runs builds, tests, version control, and package managers inside the sandbox. A windowed file viewer returns a bounded slice of a file, so one read cannot flood the context. An exact-replacement editor changes one block at a time, following listing 1. A search tool returns structured file and line matches without shell quoting hazards. Listing 5 gives their definitions. Other domains have their own small general sets, such as query, read, and a bounded write for a data agent, and the same arithmetic applies to them.
{
"tools": [
{
"name": "bash_execute",
"description": "Runs a shell command in the task's sandbox and returns stdout, stderr, and the exit status, truncated with a marker if long.",
"parameters": {
"type": "object",
"properties": {
"command": {"type": "string", "maxLength": 8192, "description": "Shell command line to run."},
"timeout_ms": {"type": "integer", "default": 30000, "minimum": 1000, "maximum": 300000}
},
"required": ["command"], "additionalProperties": false
}
},
{
"name": "file_view",
"description": "Returns a window of lines from a text file.",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string", "description": "Absolute path inside the workspace."},
"offset": {"type": "integer", "minimum": 1, "description": "First line to return, 1-indexed."},
"limit": {"type": "integer", "minimum": 1, "maximum": 500, "description": "Maximum lines to return."}
},
"required": ["path", "offset", "limit"], "additionalProperties": false
}
},
{
"name": "file_edit",
"description": "Replaces exactly one matching block of text in a file.",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string", "description": "Absolute path inside the workspace."},
"old_str": {"type": "string", "minLength": 1, "description": "Exact block to replace."},
"new_str": {"type": "string", "description": "Replacement text."}
},
"required": ["path", "old_str", "new_str"], "additionalProperties": false
}
},
{
"name": "grep_search",
"description": "Searches files for a regular expression and returns matching paths and line numbers.",
"parameters": {
"type": "object",
"properties": {
"pattern": {"type": "string", "maxLength": 512, "description": "Regular expression to match."},
"path": {"type": "string", "description": "Directory or file to search."},
"glob": {"type": "string", "description": "Optional file filter, such as '*.py'."}
},
"required": ["pattern", "path"], "additionalProperties": false
}
}
]
}A small catalog of general tools keeps the context for the task, gives the model fewer and more distinct choices, and moves enforcement to one place. It also makes each tool more powerful, which is why the envelope and the capability, not the tool list, must carry the bound on what the agent can do.