The Agents Zoo

Vetted Specifications for Agentic ML Systems, Deliberation Profiles, and Harnesses

The Agents Zoo is the authoritative registry for autonomous agent system configurations in mlsysim. It captures operational execution boundaries, context window allocations, deliberation search profiles, and tool-use execution environments across Volume III (Agentic ML Systems).

TipHow to use this page

Reference these profiles when designing multi-turn agent systems, budgeting KV-cache memory for deliberation trees, sizing multi-agent worker fleets, or evaluating tool-use sandbox latency. Load any agent architecture directly in Python: agent = mlsysim.Agents.Deliberation.TreeSearch.

Agent Architectures & Execution Harnesses

Agent Profile Paradigm Context Window Working Memory Sandbox Boot Snapshot Restore Sandbox Footprint Reference Model Serving Node
SWE-bench Coding Agent Harness ReAct / Tool-Use 128,000 tokens 32,000 tokens 5.0 ms 15.0 ms 512 MiB Llama-3.1-70B HGX H100 (Dual AMD EPYC 9654)
SWE-bench Workstation Developer Harness ReAct / Tool-Use 128,000 tokens 32,000 tokens 25.0 ms 45.0 ms 512 MiB Llama-3.1-8B MacBook Pro M3 Max Workstation
Deliberative MCTS Reasoning Agent TreeSearch / PRM Verifier 64,000 tokens 16,000 tokens — — — Llama-3.1-70B DGX H100
Supervisor-Worker Multi-Agent Fleet Hierarchical Orchestrator-Worker 32,000 tokens 8,000 tokens — — — Llama-3.1-70B DGX H100
Real-Time Streaming Voice Agent Audio-to-Audio Streaming Pipeline 8,000 tokens 2,000 tokens 0.0 ms — — Llama-3.1-70B DGX H100

Agentic Systems Architecture (Volume III)

Agent systems transition foundation models from static advisory oracles into active processors:

  1. Coding & Software Engineering Agents (Agents.Coding.SWE_Bench_Runner)
    • ReAct and Plan-and-Solve execution patterns executing inside isolated, hermetic microVM sandboxes (Firecracker CoW fork latency \(< 5\text{ ms}\), snapshot restore \(< 15\text{ ms}\), \(512\text{ MiB}\) memory footprint).
    • Served on high-density execution nodes (Systems.Nodes.HGX_H100_EPYC) with 192 AMD EPYC CPU cores to prevent host core contention during parallel microVM execution.
    • Bounded by token context limits (\(128\text{k}\) tokens) and deterministic AST / unit-test verification oracles across trajectory steps (\(\le 30\) steps).
  2. Developer Workstation Baseline (Agents.Coding.SWE_Bench_Workstation)
    • Local unified memory developer baseline running on Systems.Nodes.Workstation_M3Max (16-core CPU, 128 GiB unified LPDDR5X @ 400 GB/s) paired with Models.Language.Llama3_8B.
  3. Reference Platforms (Agents.Platforms / AgentPlatforms)
    • Canonical whole-system envelopes binding compute nodes, foundation models, and sandboxes: Agents.Platforms.Cloud_DGX_H100 for enterprise clusters and Agents.Platforms.Workstation_Apple for local developer iteration.
  4. Deliberative Reasoning Agents (Agents.Deliberation.TreeSearch)
    • Allocates test-time inference compute through tree search / MCTS (\(N=8\) branches) scored by Process Reward Models (PRMs).
    • Multiplies branch capacity via PagedAttention Copy-on-Write KV-cache page sharing (\(38.75\times\) expansion).
  5. Multi-Agent Fleets (Agents.MultiAgent.SupervisorWorker)
    • Coordinates specialized workers (Coder, Reviewer, Tester) via asynchronous message buses with bounded coordination overhead (\(\beta = 0.030\)).
  6. Real-Time Interactive Agents (Agents.Interactive.StreamingVoice)
    • Streaming speech-to-speech pipelines bound by human conversational latency budgets (\(\text{TTFT} \le 200\text{ ms}\)).

Python Access

import mlsysim

# Load agent architectures and canonical platforms
swe = mlsysim.Agents.Coding.SWE_Bench_Runner
swe_local = mlsysim.Agents.Coding.SWE_Bench_Workstation
cloud_platform = mlsysim.Agents.Platforms.Cloud_DGX_H100

# Access vetted parameters
print(swe.context_window)                   # 128,000 token
print(swe.sandbox_startup_latency)          # 5.0 ms
print(swe.sandbox_snapshot_restore_latency) # 15.0 ms
print(swe.sandbox_memory_footprint)         # 512 MiB
print(swe.max_trajectory_steps)             # 30
print(swe.serving_node.host_cpu)            # Dual AMD EPYC 9654
print(swe.serving_node.host_cpu_cores)      # 192

# Deliberation and multi-agent coordination
tree = mlsysim.Agents.Deliberation.TreeSearch
fleet = mlsysim.Agents.MultiAgent.SupervisorWorker
print(tree.deliberation_branches)           # 8
print(fleet.coordination_overhead_beta)     # 0.03
Back to top