Typed Tool Interfaces

Tool Calling explains why every tool call must pass through the runtime’s validation and grant before it takes effect, and Agent Sandboxes explains how a granted call is contained once it runs. This appendix collects the reusable artifacts an engineer needs to build that boundary: schema rules and a reference tool definition, the break-even for constraining output to a schema, reference code for idempotency keys and settlement at the endpoint, reference code for the capability check and its attenuation, and the arithmetic behind a small catalog of general tools. It is a workbench for building tool interfaces, and each section points back to the chapter that argues for the design.

How to Use This Appendix

Reach for the section that matches the failure you are seeing at the tool boundary.

  • When tool calls fail validation, invent arguments, or confuse units, use the schema rules and reference definitions in section 1.1.
  • When the model emits malformed calls often enough that repair turns cost more than prevention, use the break-even in section 1.2 to decide whether to constrain decoding.
  • When timeouts, dropped connections, or model retries risk duplicate side effects, use the key derivation and endpoint record in section 1.3.
  • When a tool call or a delegated subagent needs bounded authority, use the capability check and attenuation code in section 1.3.
  • When the tool catalog consumes a large share of every context or the model picks the wrong tool, use the token arithmetic in section 1.4.

Schema Rules for Tool Definitions

A tool schema is read twice. The model reads its names and descriptions as part of the context and proposes arguments from them, and the runtime reads its types and bounds to validate every proposed call before the grant (Tool Interface Schemas). A schema that serves only one reader fails the other. Five rules make a schema serve both.

  1. Close the property set (additionalProperties: false). A model will propose plausible arguments the tool does not accept, such as dry_run, verbose, or user_id, if the schema permits them. A closed schema makes the runtime reject undeclared inputs instead of silently passing them through, and it lets a grammar-constrained decoder exclude them before they are generated (Grammar-Guided Decoding).
  2. List every required field in required. A description that says an argument is required is a suggestion to the model; the required array is a rule the validator enforces.
  3. Bound every type. Constrain numbers with minimum and maximum, strings with pattern, enum, or maxLength, and arrays with minItems and maxItems. Unbounded inputs let one call carry an argument larger than the tool, or the context that records it, can hold.
  4. Put units and conventions in names and descriptions. Write timeout_ms rather than timeout, state whether line numbers start at 1, and say what a missing optional field means.
  5. Tag every union. An anyOf or oneOf without a discriminator field leaves the model to guess which branch it is filling and leaves the validator unable to report which branch failed.

The same schema also carries the side-effect annotation that tells the runtime how to treat the call (read-only, idempotent, or destructive), which maps onto the authority levels \(A_0\) to \(A_3\) and decides whether the call needs an idempotency key, a compensator, or an approval (Tool Interface Schemas).

A reference tool definition

Listing 1 defines an exact-replacement edit tool that follows all five rules.

Listing 1: Reference Tool Definition: A JSON Schema (Draft 2020-12) for an exact string replacement tool with a closed property set, bounded strings, and failure behavior stated in the description.
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "name": "file_edit",
  "description": "Replaces an exact block of text in an existing file. old_str must match exactly one contiguous block unless allow_multiple is true; if it matches zero blocks, or more than one when allow_multiple is false, the file is left unchanged and the call returns an error naming the match count.",
  "parameters": {
    "type": "object",
    "properties": {
      "path": {
        "type": "string",
        "description": "Absolute path to the target file, inside the workspace root.",
        "pattern": "^/([a-zA-Z0-9_.-]+/)*[a-zA-Z0-9_.-]+$",
        "maxLength": 1024
      },
      "old_str": {
        "type": "string",
        "description": "The exact block to replace, including indentation and newlines.",
        "minLength": 1
      },
      "new_str": {
        "type": "string",
        "description": "The replacement text. May be empty to delete old_str."
      },
      "allow_multiple": {
        "type": "boolean",
        "description": "If true, replace every match; if false, fail unless exactly one match exists.",
        "default": false
      }
    },
    "required": ["path", "old_str", "new_str"],
    "additionalProperties": false
  }
}

Writing schemas by hand invites drift between the schema the model sees and the code that runs. Generating the schema from a typed model of the arguments keeps them in step. Listing 2 shows one way, using a Python data-validation library as the example; any typed schema generator serves the same purpose.

Listing 2: Generating a Schema from Typed Arguments: A typed argument model that forbids extra fields and exports the JSON Schema the runtime publishes to the model and validates against.
from typing import Annotated
from pydantic import BaseModel, ConfigDict, Field, StringConstraints

class FileEditInput(BaseModel):
    """Arguments for the exact-replacement edit tool."""
    model_config = ConfigDict(extra="forbid", frozen=True)

    path: Annotated[
        str,
        StringConstraints(strip_whitespace=True, max_length=1024,
                          pattern=r"^/([a-zA-Z0-9_.-]+/)*[a-zA-Z0-9_.-]+$"),
        Field(description="Absolute path to the target file, inside the workspace root."),
    ]
    old_str: Annotated[str, StringConstraints(min_length=1),
                       Field(description="Exact block to replace.")]
    new_str: Annotated[str, Field(description="Replacement text; may be empty.")]
    allow_multiple: Annotated[bool, Field(default=False,
                              description="Replace every match instead of exactly one.")]

def tool_descriptor() -> dict:
    """The descriptor the runtime publishes to the model and validates against."""
    return {
        "name": "file_edit",
        "description": "Replaces an exact block of text in an existing file.",
        "inputSchema": FileEditInput.model_json_schema(),
    }

Schema anti-patterns

Tools adapted from existing service interfaces often carry patterns that suit a programmer calling from code but mislead a model proposing from a description. Table 1 lists the common ones.

Table 1: Tool Schema Anti-Patterns: Common schema patterns that mislead a model or defeat the validator, the hardened form of each, and its effect on the boundary.
Anti-pattern What goes wrong Hardened form Effect
Untyped argument bag (args: dict) The model invents keys; the validator has nothing to check Explicit typed fields with additionalProperties: false Invalid calls are rejected, or never generated under constrained decoding
Unbounded inputs (string with no limit) One call carries an argument larger than the tool or the context can hold maxLength, maxItems, and explicit windowing parameters Argument size is bounded before dispatch
Untagged union (anyOf with no discriminator) The model mixes fields from several branches; errors cannot name the branch A discriminator field, such as "kind" with values "file" or "dir" Each call selects one branch explicitly
Implicit units (timeout: 30) Seconds, milliseconds, and minutes are confused Units in the name (timeout_ms) with range bounds Unit errors become validation errors
Whole-file rewrite (content: str) Output length grows with the file, long files are truncated, and unrelated lines drift Exact replacement of a block (old_str, new_str) Output tokens scale with the change, not the file

The last row matters most for latency, because output tokens are the expensive ones on an agent’s critical path (Accelerator Serving Latency). A whole-file rewrite pays decode time for every unchanged line, and Interface Benchmarking works an example in which the rewrite misses a per-call deadline that an anchored diff meets.

When to Constrain Decoding

A model that samples freely can emit a malformed call: an unescaped newline, a trailing comma, a missing required field. Grammar-constrained decoding prevents that by restricting each sampling step to the tokens the schema’s grammar admits in its current state (Grammar-Guided Decoding), Decode-Loop Mechanics sizes the mask tables and explains why the mask must run beside the sampler, and Deterministic Finite Automata and Grammar-Constrained Logit Decoding proves that the result always parses. What this section adds is the decision of when constraining pays, because the alternative, letting the malformed call fail validation and asking the model to repair it, has a cost of its own.

Let a tool call carry \(L_{\text{gen}}\) output tokens at \(t_{\text{decode}}\) per token, let masking add \(t_{\text{mask}}\) per token, let \(p_{\text{err}}\) be the rate at which unconstrained calls fail to parse, and let a repair turn cost \(T_{\text{repair}}\) of latency, including the re-sent context and the regenerated call. The expected latency without constraints, allowing one repair, is

\[\mathbb{E}[T_{\text{free}}] = L_{\text{gen}} \, t_{\text{decode}} + p_{\text{err}} \, T_{\text{repair}}\]

and with constraints it is

\[T_{\text{constrained}} = L_{\text{gen}} \, (t_{\text{decode}} + t_{\text{mask}})\]

Setting the two equal gives the break-even error rate

\[p_{\text{break}} = \frac{L_{\text{gen}} \, t_{\text{mask}}}{T_{\text{repair}}} \tag{1}\]

Constraining pays whenever the observed malformation rate exceeds \(p_{\text{break}}\). The rate is a property of a particular model, schema, and prompt, and must be measured on the tool’s own traffic. As an illustration, with an 80-token call, a mask overhead of 0.5 ms per token, and a 1.8-second repair turn, equation 1 gives \(p_{\text{break}} = 40\ \text{ms} / 1{,}800\ \text{ms} \approx 2.2\) percent, so a tool whose calls fail to parse more often than about one time in forty is cheaper to constrain. Latency is only part of the case. A repair turn also adds error text to the context, which makes the next proposal about the same object more likely to repeat the mistake (Context poisoning dynamics).

Constraining guarantees syntax and nothing else. A call that parses can still name a file that does not exist, pass an argument outside the task’s authority, or do the wrong thing correctly, which is why validation, the grant, and the envelope stay in place behind it (Grammar-Guided Decoding).

Idempotency Keys and Capabilities

Two artifacts make a granted call safe to send. An idempotency key makes a repeated call harmless, so a runtime that cannot tell whether a timed-out call ran can settle it instead of guessing (Idempotent Action Execution). A capability bounds what the call can reach once it runs, so a steered or mistaken proposal can do no more than the task allows (Capabilities and Credentials). The chapters argue for both. This section gives the reference form of each.

Deriving the key

The key must identify the logical call, not the attempt, so the runtime derives it deterministically from the trajectory. Idempotent Action Execution uses

\[k_{\text{idem}} = \mathcal{H}\big(\text{trajectory\_id} \,\|\, \text{turn} \,\|\, \text{call\_id} \,\|\, \text{tool} \,\|\, \text{canonical\_args}\big)\]

where \(\mathcal{H}\) is a cryptographic hash, or a keyed hash when keys must not be guessable by the endpoint’s other clients. The arguments must be serialized canonically before hashing. Ordinary JSON serialization does not fix key order or whitespace, so {"path": "/a", "force": true} and {"force": true, "path": "/a"} would hash to different keys and defeat deduplication. A canonical form, such as the JSON Canonicalization Scheme of RFC 8785, sorts keys and fixes number and string encoding. Because the key is derived rather than drawn at random, a runtime that recovers from a crash and re-sends a call regenerates the same key, and the endpoint recognizes the repeat (Intent Before Effect).

The endpoint record

At the endpoint, each key maps to a record with one of four states, and the transitions between them are what make a retry safe.

  1. Unseen. The endpoint claims the key atomically, with a set-if-absent operation and a lease that outlives the worst-case execution time plus the caller’s timeout and any clock skew (Idempotent Action Execution), and then executes the call.
  2. In flight. A repeat arriving while the first execution runs is refused or made to wait. It is never executed in parallel.
  3. Completed. A repeat of a completed call returns the stored result without executing again.
  4. In doubt. If execution fails in a way that leaves its effect unknown, such as a crash of the executor partway through, the record stays claimed and is marked in doubt. Repeats are refused until a reconciliation probe inspects the environment and settles the record as completed or not executed.

The fourth state is the one naive implementations omit. Releasing the key after an exception invites the next retry to execute again, which is exactly the duplicate the key exists to prevent, because an exception does not prove that the effect did not happen. Listing 3 implements all four states.

Listing 3: Endpoint Deduplication with Settlement: A reference endpoint that claims each idempotency key atomically, returns the stored result for completed keys, refuses repeats while a call is in flight, and marks a failed execution as in doubt rather than releasing its key.
import hashlib, json, time
from typing import Any, Callable, Dict

def canonical(args: Dict[str, Any]) -> str:
    """Stable serialization: sorted keys, fixed separators."""
    return json.dumps(args, sort_keys=True, separators=(",", ":"), ensure_ascii=True)

def derive_key(trajectory_id: str, turn: int, call_id: str, tool: str,
               args: Dict[str, Any]) -> str:
    material = f"{trajectory_id}|{turn}|{call_id}|{tool}|{canonical(args)}"
    return hashlib.sha256(material.encode("utf-8")).hexdigest()

class DedupEndpoint:
    """Executes each logical call at most once per idempotency key."""

    def __init__(self, store, lease_s: float):
        self.store, self.lease_s = store, lease_s   # store offers set_if_absent/get/put

    def handle(self, key: str, execute: Callable[[], Dict[str, Any]]) -> Dict[str, Any]:
        claimed = self.store.set_if_absent(
            key, {"state": "IN_FLIGHT", "until": time.time() + self.lease_s})
        if not claimed:
            record = self.store.get(key)
            if record["state"] == "COMPLETED":
                return record["result"]                       # repeat: stored result
            return {"status": record["state"], "retry": True}  # IN_FLIGHT or IN_DOUBT
        try:
            result = execute()
        except Exception as exc:                              # effect unknown
            self.store.put(key, {"state": "IN_DOUBT", "error": str(exc)})
            return {"status": "IN_DOUBT", "retry": False}     # settle by probe first
        self.store.put(key, {"state": "COMPLETED", "result": result})
        return result

A tool whose endpoint cannot keep such a record needs the other settlement path, a reconciliation probe that reads the environment to decide whether the effect happened before any retry (Idempotent Action Execution).

The capability check

Capabilities and Credentials writes a capability as \(C = (R, O, P, E)\), signed with a key only the runtime holds: \(R\) is the set of rights (read, write, execute, connect), \(O\) the objects the capability designates (path prefixes, host names, repositories), \(P\) a predicate that bounds arguments and cost, and \(E\) the expiry. A capability issued for one trajectory and bounded by \(E\) is that trajectory’s lease on the resources it names (Workspace leases per trajectory, reset per tenant). The runtime approves a proposal only if the signature verifies, the operation is among the rights, the target lies within the objects, the cost satisfies the predicate, and the capability has not expired.

When work is delegated, a child’s capability must attenuate from its parent’s. Writing \(C_c \preceq C_p\) for “the child holds no more than the parent”,

\[C_c \preceq C_p \iff R_c \subseteq R_p \;\land\; O_c \sqsubseteq O_p \;\land\; P_c \Rightarrow P_p \;\land\; E_c \le E_p\]

so a child can never hold a right its parent lacks, reach an object outside its parent’s scope, spend beyond its parent’s bound, or outlive its parent’s grant (Attenuated capability delegation; principle \(\ref{pri-vol3-monotonic-delegation}\)). Listing 4 implements the check and the attenuation, with the predicate reduced to a cost ceiling for brevity.

Listing 4: Capability Check and Attenuation: A reference implementation of the capability \(C = (R, O, P, E)\), with the predicate reduced to a cost ceiling, a runtime-held signing key, and an attenuation step that refuses to widen any component.
import hashlib, hmac, time
from dataclasses import dataclass
from typing import FrozenSet

@dataclass(frozen=True)
class Capability:
    rights: FrozenSet[str]     # R: e.g. {"read", "write"}
    objects: FrozenSet[str]    # O: path or host prefixes
    max_cost: float            # P: reduced here to a cost ceiling
    expires_at: float          # E: epoch seconds
    sig: str = ""

    def body(self) -> bytes:
        return repr((sorted(self.rights), sorted(self.objects),
                     self.max_cost, self.expires_at)).encode()

def sign(cap: Capability, key: bytes) -> Capability:
    return Capability(cap.rights, cap.objects, cap.max_cost, cap.expires_at,
                      hmac.new(key, cap.body(), hashlib.sha256).hexdigest())

def within(target: str, objects: FrozenSet[str]) -> bool:
    return any(target == o or target.startswith(o.rstrip("/") + "/") for o in objects)

def approve(cap: Capability, key: bytes, op: str, target: str, cost: float) -> bool:
    good_sig = hmac.compare_digest(
        cap.sig, hmac.new(key, cap.body(), hashlib.sha256).hexdigest())
    return (good_sig and op in cap.rights and within(target, cap.objects)
            and cost <= cap.max_cost and time.time() < cap.expires_at)

def attenuate(parent: Capability, key: bytes, *, rights, objects,
              max_cost, expires_at) -> Capability:
    """Issue a child capability; refuse any component wider than the parent's."""
    if not (set(rights) <= parent.rights
            and all(within(o, parent.objects) for o in objects)
            and max_cost <= parent.max_cost and expires_at <= parent.expires_at):
        raise PermissionError("attenuation may only narrow a capability")
    return sign(Capability(frozenset(rights), frozenset(objects),
                           max_cost, expires_at), key)

The check runs below the model, in the runtime and the sandbox’s drivers, on every access. Nothing the model writes can widen a capability, because the model never holds the signing key, and a proposal that names a target outside the capability fails before any host operation runs.

A Small Catalog of General Tools

A natural first catalog wraps each endpoint of an existing service as its own tool: git_status, git_diff, git_add, git_commit, and so on across every system the agent touches. The result is dozens or hundreds of narrow tools. Ousterhout calls a module shallow when its interface is large relative to the functionality behind it (Ousterhout 2018), and a catalog of narrow tools is shallow in exactly that sense, with a cost an agent pays on every turn. Toolkit Granularity Partitioning argues the design choice. This section gives the arithmetic.

Ousterhout, John. 2018. A Philosophy of Software Design. Yaknyam Press.

The token cost of a catalog

Every tool’s definition is sent with every call that may use it. With \(K\) tools whose definitions average \(S\) tokens, a trajectory of \(N\) turns sends

\[T_{\text{catalog}} = N \cdot K \cdot S \tag{2}\]

tokens of tool definitions. As an illustration, a catalog of 80 narrow tools at about 350 tokens each costs 28,000 tokens per turn and 1,120,000 over a 40-turn trajectory. Four general tools at about 250 tokens each cost 1,000 tokens per turn and 40,000 over the same trajectory, a reduction of about 96 percent. Prefix caching lowers the price of the re-sent definitions as long as the catalog stays byte-identical and in a fixed position, but it does not lower their share of the context budget, and any change to the catalog mid-trajectory discards the cached prefix (Staging the Next Invocation). Tool selection also degrades as the catalog grows and definitions overlap, a trend that function-calling benchmarks measure by varying catalog size (Toolkit Granularity Partitioning).

Table 2 sets the two catalog shapes side by side on these costs and on where authority is enforced.

Table 2: Narrow Versus General Tool Catalogs: How catalog shape affects definition tokens, tool selection, observations, and where authority is enforced.
Dimension Many narrow tools A few general tools Consequence
Interface surface Tens to hundreds of endpoint wrappers A handful of general operations Fewer definitions for the model to read and choose among
Definition tokens Grows with \(K\) on every turn (equation 2) Small and stable More of the context budget left for the task
Selection Overlapping names and arguments invite the wrong tool Few, distinct choices Fewer wrong-tool calls; more depends on each tool’s description
Observations A different result shape per tool One result shape (output, error, exit status) One observation-shaping path in the gateway
Enforcement Authorization logic scattered across many handlers Authority enforced by the envelope and capability One place to enforce and audit, at the price of broader tools

The last row is the cost of the general design. A general shell tool can do far more than a narrow one, so it moves the burden of bounding each action from the tool’s code to the isolation envelope and the capability it runs under (Agent Sandboxes). A general catalog is safe only inside an envelope sized for the broadest thing its tools can do.

A reference catalog for coding agents

Many coding agents converge on four general operations. A shell runs builds, tests, version control, and package managers inside the sandbox. A windowed file viewer returns a bounded slice of a file, so one read cannot flood the context. An exact-replacement editor changes one block at a time, following listing 1. A search tool returns structured file and line matches without shell quoting hazards. Listing 5 gives their definitions. Other domains have their own small general sets, such as query, read, and a bounded write for a data agent, and the same arithmetic applies to them.

Listing 5: A Four-Tool Catalog for Coding Agents: Definitions for a sandboxed shell, a windowed file viewer, an exact-replacement editor, and a structured search, each with a closed property set and bounded inputs.
{
  "tools": [
    {
      "name": "bash_execute",
      "description": "Runs a shell command in the task's sandbox and returns stdout, stderr, and the exit status, truncated with a marker if long.",
      "parameters": {
        "type": "object",
        "properties": {
          "command": {"type": "string", "maxLength": 8192, "description": "Shell command line to run."},
          "timeout_ms": {"type": "integer", "default": 30000, "minimum": 1000, "maximum": 300000}
        },
        "required": ["command"], "additionalProperties": false
      }
    },
    {
      "name": "file_view",
      "description": "Returns a window of lines from a text file.",
      "parameters": {
        "type": "object",
        "properties": {
          "path": {"type": "string", "description": "Absolute path inside the workspace."},
          "offset": {"type": "integer", "minimum": 1, "description": "First line to return, 1-indexed."},
          "limit": {"type": "integer", "minimum": 1, "maximum": 500, "description": "Maximum lines to return."}
        },
        "required": ["path", "offset", "limit"], "additionalProperties": false
      }
    },
    {
      "name": "file_edit",
      "description": "Replaces exactly one matching block of text in a file.",
      "parameters": {
        "type": "object",
        "properties": {
          "path": {"type": "string", "description": "Absolute path inside the workspace."},
          "old_str": {"type": "string", "minLength": 1, "description": "Exact block to replace."},
          "new_str": {"type": "string", "description": "Replacement text."}
        },
        "required": ["path", "old_str", "new_str"], "additionalProperties": false
      }
    },
    {
      "name": "grep_search",
      "description": "Searches files for a regular expression and returns matching paths and line numbers.",
      "parameters": {
        "type": "object",
        "properties": {
          "pattern": {"type": "string", "maxLength": 512, "description": "Regular expression to match."},
          "path": {"type": "string", "description": "Directory or file to search."},
          "glob": {"type": "string", "description": "Optional file filter, such as '*.py'."}
        },
        "required": ["pattern", "path"], "additionalProperties": false
      }
    }
  ]
}

A small catalog of general tools keeps the context for the task, gives the model fewer and more distinct choices, and moves enforcement to one place. It also makes each tool more powerful, which is why the envelope and the capability, not the tool list, must carry the bound on what the agent can do.

Back to top