Bonus dive
Agent Harnesses: A Guided Deep Dive
A hands-on playground for the part of agent work the Agents dive left off. Once you've hand-written the loop, most real agent engineering happens on a harness, the layer that runs the loop for you and adds subagents, hooks, permission policies, sandboxed tool execution, headless automation, durable checkpointed and resumable runs, and orchestration through parallel workers, mid-run steering, and graph control flow. You'll build a small harness from scratch and watch each of those pieces appear as a thin wrapper around the loop you already know. No framework magic, just enough code to see what a harness gives you, and to answer the interview question: "you have a working agent loop, so when do you throw it away for the SDK, and what does the SDK actually give you?"
Here's what makes this repo work. It runs completely offline on a mock provider, with no API key. Hooks, policies, sandboxing, subagents, and event streams are all provider-neutral, so a deterministic rule-based "model" is all you need to see every one of them work. Flip one env var and the same harness drives a real OpenAI or Claude model.
This is a bonus dive and the direct sequel to Agents, #6. That dive builds the loop. This one builds the layer above it. It also connects to Prompt Injection, since hooks and the sandbox are where guardrails live, and to Context Engineering, since subagents give each agent its own window. Its code depends on none of them.
Like its siblings, walk through it. Each section ends with something to run, and every section runs offline and free on the mock. EXERCISES.md has a predict-then-run prompt for each one.
0. The one big idea
A harness is the agent loop, wrapped, so instead of writing the loop you configure it and consume its event stream. That wrapper is where subagents, hooks, permission policies, the sandbox, durable checkpoints, and orchestration live. In 2026, most agent work happens on a harness rather than in a hand-rolled loop.
That's the whole repo. The Agents dive proved the loop is about 20 lines. But a production
loop needs a place to gate a dangerous call, a place to redact a secret, a boundary on
where tools act, a way to delegate, structured output you can log and test, and a way to
survive a crash mid-run. Bolt all of that into a bare while loop and it stops being
readable. A harness lifts each concern out into its own join. Everything below is one of
those joins, a small addition to the loop rather than a new concept. Hold onto that and
none of this feels complicated.
1. Setup (5 minutes)
# 1. Create an isolated Python environment
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 2. Install dependencies (the default mock stack needs only python-dotenv + rich)
pip install -r requirements.txt
# 3. Copy the env file: the default runs keyless (no API key needed)
cp .env.example .env
# (Real provider instead of the mock? Its key goes in your OS keychain,
# not .env: see ../docs/SECRETS.md, then run scripts as `secrun python ...`.)
# 4. Confirm everything is wired up (makes no API call, costs nothing)
python check_setup.py
No key required. The default PROVIDER=mock is a deterministic in-process tool-calling
"model". Pick your stack with PROVIDER in .env.
PROVIDER |
What runs the model | Key needed | Cost |
|---|---|---|---|
mock (default) |
a deterministic offline planner | none | $0 |
openai |
OpenAI gpt-6-luna |
OPENAI_API_KEY |
tiny |
claude |
Claude claude-haiku-4-5 |
ANTHROPIC_API_KEY |
tiny |
The only file that knows which one you picked is harness/providers.py. Everything above it, meaning the harness, hooks, policy, sandbox, and subagents, is provider-neutral.
Why a mock is the right call here. The subject is the harness rather than the model. A rule-based planner that reliably asks for the tool each example is about lets you watch hooks fire, policies gate, and sandboxes refuse, deterministically, offline, for $0. Real models make the same kinds of requests, just less predictably.
2. The bare loop, and what it's missing
python examples/01_bare_loop_recap.py # offline
Here's the Agents-dive loop again, in about 15 lines, driving the mock. It works, and
that's the point. It works and it's naked. There's no structured way to observe it beyond
print, nowhere to gate a write_file, nowhere to block a call or redact a result, no
boundary on where a tool acts, and no way to delegate. The example runs the loop, then
names those five gaps. The rest of the dive fills them.
3. The harness, which is the loop wrapped
python examples/02_harness_events.py
Hand the same task to a Harness and stop writing loops. You configure it, then iterate
its event stream, one typed event per thing that happens
(harness/events.py). The while is gone from your code and lives
inside Harness.run(). That inversion is the reason a harness is worth adopting. Every
event is a place to observe, react, or record. The next three examples slot capabilities
into those places without touching the loop.
4. Hooks, intercepting without editing the loop
python examples/03_hooks.py
A hook is a function the harness calls at a fixed point in every tool cycle. Pre-tool runs
before a tool, and can return a substitute result or raise HookBlock to refuse. Post-tool
runs after, and transforms the result before it re-enters the model's context, which makes
it the natural home for redaction. The example blocks reads of credential-looking files
with a pre-tool hook and redacts an API key that slips through a file's contents with a
post-tool hook, so neither the model nor your logs ever see the raw secret. This is where
the Prompt Injection dive's
defenses live in a real system: at the harness boundary, rather than re-implemented in
every tool.
5. Permission policies: allow, ask, deny
python examples/04_permissions.py
The Agents dive gated dangerous tools with an approve callback tangled into the loop. A
harness makes the policy a declarative object, harness/policy.py, that
you read, diff, version, and swap per environment. Three verdicts: allow runs it, ask
pauses for a human, deny never runs. The example allows the calculator, asks before
write_file, and denies run_command, and the agent adapts when a call is denied, because
a denial comes back as another tool result. This is the shape of Claude Agent SDK
permission modes and Managed Agents' per-tool always_allow and always_ask config.
A fourth answer to ask: let a reviewer decide
python examples/16_model_approval.py # rule reviewer, offline
PROVIDER=openai secrun python examples/16_model_approval.py # a model reviews
ask assumes a person reads each prompt. Mostly they don't. Anthropic's study of Claude
Code's auto mode found testers caught a deliberately inserted dangerous command 13.6% of
the time, while its classifier blocked 89%, and users approved 97% of the prompts they
saw. harness/review.py puts a reviewer where the person was. It's
just an approve callback, so the loop doesn't change. A block goes back to the agent
with its reason, and after 3 blocks in a row or 20 in a session the next decision goes
to a person, whose approval hands control back. Those are Claude Code's own thresholds.
The example shows two things worth more than the mechanism. Reviewers disagree: asked
whether to save a file whose text contains curl ... | sh, the rule reviewer and
claude-haiku-4-5 blocked it and gpt-6-luna allowed it, since saving isn't running.
And a reviewer only covers what you route to it. Blocked on run_command, the agent
tried read_file on the same ~/.ssh path, a tool the policy allows, and only §6's
sandbox stopped it.
6. The sandbox, the boundary tools execute inside
python examples/05_sandbox.py
The model chose the arguments, and it may be acting on attacker-controlled text. So tool
execution needs a boundary the model can't argue past,
harness/sandbox.py. Two reject-by-default boundaries. A path jail
resolves every file path and requires it to stay under the workspace root, so
../../etc/passwd is refused. A command allowlist runs only named executables. The example
reads a legitimate file, then watches a directory-traversal escape and a non-allowlisted
command get refused as tool results, so the agent adapts. Real harnesses sandbox far harder
with containers, seccomp, egress rules, and a provider-hosted per-session workspace, and it's
the same contract. The model proposes, the sandbox disposes.
7. Subagents, delegating with an isolated context
python examples/06_subagents.py
The Agents dive showed a subagent as a tool whose function runs its own loop. A harness
makes it first-class. Register a Subagent with its own persona and toolset and it appears
to the model as an ordinary tool. When called, the harness spawns a nested harness, a fresh
context window holding only that subagent's tools, runs it, and returns the final answer.
What you get is context isolation. The orchestrator's window never fills with the
subagent's intermediate steps. The example has an arithmetic-only orchestrator delegate a
lookup to a research subagent that owns the knowledge-base tool, and the nested run
appears indented in the stream. Scale this up and it's how large agent systems get
built.
8. Headless automation, one-shot and scriptable and structured
python examples/07_headless.py
python examples/07_headless.py "What is 19 * 21?"
The other way to run an agent, the one job descriptions call agentic automation, is
headless. No human, kicked off by cron or CI, emitting structured output another program
consumes. Because everything is events, you fold a run into a machine-readable record as it
happens. The example runs a task with no interaction and prints a JSON summary, the shape
you'd write to a log, post to a webhook, or assert on in CI, failing the build if a
blocked tool shows up.
9. Computer use and hosted sandboxes
python examples/08_computer_use.py # offline simulation of the pattern
Computer use is the same loop with a different set of tools. The tools are screenshot,
click, and type, and the observation fed back each step is an image of a screen.
Observe, act, observe, exactly as before. The example is a self-contained simulation of
that loop, with a mock login form and a scripted planner standing in for a vision model, so
you see the shape offline. A harness adds two things here. The same permission and hook
points, so you can gate a click on a payment page or redact a typed password. And a hosted
sandbox, a provider-run VM or browser, so the agent drives an isolated machine rather than
your laptop. Reach for computer use only when the task lives in a GUI with no API. A real
tool or API is cheaper and far more reliable than driving pixels.
10. Durable runs: checkpoint, crash, resume
python examples/09_checkpoint_resume.py # offline
A long-horizon agent runs for minutes or hours. If the process dies mid-run, through a deploy, an OOM kill, a timeout, or a reboot, an in-memory loop loses everything and starts over, re-paying for every step it already finished. A harness checkpoints instead. It persists its state after each step and can resume in a fresh process, redoing nothing. The elegant part is that the harness's own transcript is the checkpoint. Because every tool result gets fed back into it, persisting the transcript is all it takes, as harness/checkpoint.py shows. Reload it, keep looping, and the model, seeing the results already there, moves on. The example runs a two-step task, lets "process 1" crash after the first tool, and has a brand-new "process 2" resume, with a counter proving each tool runs exactly once across both. Real systems do the same with a database instead of a JSON file: LangGraph checkpointers, Temporal-style durable workflows, Managed Agents' server-side sessions.
Read the counter for exactly what it claims. The crash here lands between steps, after the tool returned and its result was written down, which is the case checkpointing solves completely. The case it doesn't solve is a crash inside the call, where the request reached the outside world and the response never came back. The checkpoint has no record, so resuming runs the tool again, and whether that's harmless or a second charge on a customer's card depends on the tool, not on the harness. This is why the boring advice is the load-bearing one: give every tool with an external effect an idempotency key, and make the retry safe at the thing being retried. No amount of durability upstream can fix an effect that isn't repeatable.
On the newest Claude models a checkpoint has one more job: replay exactly what was sent.
Claude Fable 5.1, Opus 5.5, and Sonnet 5.5 sign each thinking block with everything before
it, so a resume that rebuilds the request from templates, re-renders the system prompt, or
reorders the tools invalidates every block, and accounts created since 2026-08-31 get a
400 for it. The Context Engineering dive's example 11 shows the check live. This harness
sidesteps it the blunt way: its transcript keeps text and tool calls but not thinking
blocks, so there's nothing to fail the check, and also nothing of the model's earlier
reasoning survives a resume. Fine on Haiku 4.5, which is what it runs. A harness for those
newer models would store the provider's content blocks as returned, keep system and
tools frozen for the run, and append, never edit.
11. Durable task state, a queryable run log
python examples/10_run_records.py # offline
The same persisted state gives you the other half for free, a task-state log you can query.
Each run carries a status through its lifecycle: queued, then running, then done. Or
failed, meaning it gave up. Or stuck in running, meaning it crashed mid-run. Because
every run is a file on disk, you can list them all and see which finished, which are still
going, and which crashed and need resuming. That's exactly what a job queue, a cron
dashboard, or Managed Agents' deployment-run records give you. The example runs three jobs,
one that completes, one capped so it fails, and one that crashes, prints the durable log,
and resumes the crashed one straight from it. That status column is the difference between
an agent you hope finished and one you can prove did.
12. Parallel subagents, fanning out and joining
python examples/11_parallel_subagents.py # offline
Example 07 delegated to one subagent, and they run one at a time. But a lot of agent work
is independent, like researching five topics or reviewing ten files, and running that
serially wastes wall-clock. The batch should cost the slowest worker rather than the sum.
fan_out in harness/orchestrate.py is the coordinator's map
step. Hand it a list of (subagent, task) workers and it runs them concurrently, each in
its own harness and context window, returning every result for you to aggregate, which is
the reduce. The example times the same batch serially and concurrently and shows the
concurrent run finishing in about one worker's time, roughly 3× here. Keep parallel workers
independent and read-mostly, since they share the sandbox. For dependent steps, delegate
serially. This is the from-scratch shape of LangGraph parallel branches or a Managed Agents
multiagent coordinator.
13. Steering a running agent, injecting and interrupting
python examples/12_steering.py # offline
The permission policy in §5 gates a tool before it runs. Steering is the other half of
operator control, acting on a run while it's in flight. The harness polls a controller,
harness/steer.py, at each step boundary, so you can inject a message
that changes the next step ("actually, only Pro") without restarting, queue follow-ups
processed in order, and interrupt, stopping the run at a safe boundary instead of killing
it mid-tool. The example injects a follow-up that redirects the agent, then interrupts it
cleanly. An interrupted run gets checkpointed as interrupted, which is resumable rather
than lost. A real app drives the live QueueController, calling steer() and interrupt()
from a UI or chat bridge. This is the from-scratch shape of Managed Agents' message queue
plus user.interrupt.
Notice that steering appends a message. That's the only safe way on the newest Claude
models: editing the system prompt to redirect a run would invalidate the thinking already
in it (§10). Anthropic's API has a purpose-built version, a role: "system" message
appended mid-conversation, which carries system-prompt authority without touching
anything before it.
14. Orchestration as a graph, with routing, branching, and cycles
python examples/13_orchestration_graph.py # offline
The loop lets the model choose the next step. When the path is knowable you want code to
choose it, which means a graph of nodes wired by conditional edges,
harness/graph.py. The example builds a support-ticket workflow.
classify routes each ticket to a per-category handler, which is branching, since a
billing ticket and a technical ticket visit different nodes. A review gate loops back
through revise until the draft passes, which is a cycle. Then it routes to send. A node
is state -> state, so it can run plain code or a whole Harness. The graph owns the control
flow, not what's inside a node. This is the workflow-against-agent call from the Agents
dive made concrete, and it's the model behind LangGraph. If you can draw the flowchart,
build a graph, because it's cheaper, predictable, and testable. Reach for the model-driven
loop only when the path genuinely can't be known up front.
15. Managed Agents, when you own none of it
python examples/14_managed_agents.py # explain only, free
secrun python examples/14_managed_agents.py --real # provisions, then cleans up
Sections 3 through 14 built a harness. Managed Agents is Anthropic running that whole layer. You don't write the loop, host the container, or persist the run. It's the far end of the axis this dive walks.
your own loop -> your own harness -> a harness you host -> hosted entirely
(§2) (§3-14) (Claude Agent SDK) (Managed Agents)
Almost nothing in it is a new idea. It's this dive's problem list with someone else's answers plugged in.
| This dive | Managed Agents |
|---|---|
| your event stream (§3) | sessions.events.stream() |
| permission policies (§5) | always_allow / always_ask per tool |
| the sandbox (§6) | the session's container, Anthropic-run |
| subagents (§7) | a multiagent roster on the agent |
| checkpoint/resume (§10) | the session, durable by construction |
| run records (§11) | deployment runs |
| steering (§13) | user.message / user.interrupt events |
One structural rule is worth memorizing. An Agent is a persisted, versioned config holding the model, system prompt, and tools. A Session is one run of it. Those fields live on the agent, never on the session. Creating an agent per run is the classic mistake. It orphans objects, pays creation latency every time, and discards the versioning that is the whole reason agents are separate objects. Create once, store the id, reuse. That's the hosted version of not re-instantiating your harness inside the request handler.
What you give up is real. You can't reach into the loop the way §4's hooks let you, tools run in a container you don't own, and it's one vendor. Good trade when the alternative is maintaining §3 through §14 yourself. Bad trade when that layer is where your value lives, which is precisely the judgement this dive exists to give you.
The same idea from OpenAI: the Agents API
OpenAI shipped its own hosted harness on 2026-09-10, the Agents API, in public beta. It
runs OpenAI's Codex harness for you and handles sessions, orchestration, context
compaction, and recovery. The shape is close enough that the table above mostly carries
over: an agent (model, instructions, tools, MCP servers), an optional environment (an
OpenAI-hosted sandbox, or one you run), a durable session, and a stream of events, with
webhooks as the other way to hear when a turn finishes. Subagents are a multi_agent
setting with a concurrency cap, and you can steer a session mid-turn or hand it the next
task.
The difference worth noticing is the structural rule above. OpenAI's quickstart passes the agent's model and instructions inline when it creates each session:
client.beta.agents.sessions.create(
agent={"model": "gpt-6-astra", "instructions": "..."},
environment={"type": "openai_hosted"},
input="...",
)
That's fine for a first run. The SDK also has client.beta.agents.create(...), a persisted
agent you create once and point sessions at, and for anything you deploy the same advice
holds as for Managed Agents: create it once, keep the id, and get versioning for free. Two
vendors, one lesson. The convenient call isn't the one you want in production. (Checked
against openai 3.24.0 and OpenAI's docs on 2026-10-07; it's a beta, so expect the surface
to move.)
16. Agent Skills, instructions that load on demand
python examples/15_skills.py # explain + list skills, free
secrun python examples/15_skills.py --real # actually builds a spreadsheet
A capable agent needs more instructions than fit comfortably in one system prompt, and
stuffing them all in makes every request slower, dearer, and, as the
Context Engineering dive
shows, measurably worse. A skill is the answer on the instruction side. A folder with a
SKILL.md, whose one-line description sits in context always while the body gets read only
when the task calls for it.
| Mechanism | What stays in context |
|---|---|
| system prompt | all of it, every request, forever |
| skill | one line; the rest loads on demand |
| subagent (§7) | nothing; it gets its own window |
Three things have to travel together or the request fails: the code-execution-2025-08-25
beta, a container naming the skills, and the code_execution tool, because skills
execute in the container. Anthropic ships xlsx, pptx, docx, and pdf, and you can
register your own.
Skills itself went GA, so the skills-2025-10-02 header it used to need is gone and the
namespace is client.skills rather than client.beta.skills. Worth noticing that one
half of this request graduated and the other didn't: a request straddling a GA feature
and a beta one carries headers for the beta half only, and "it's beta" is a property of
each feature rather than of the call.
This sits in a harness dive rather than an API one because a skill is configuration your harness owns, exactly like §5's permission policy or §4's set of tools. Which skills to attach, and whether the model may reach for one unprompted, are your decisions. And since a skill can carry scripts, and those scripts run, a skill you didn't write deserves the same suspicion as a tool you didn't write.
The capstone: agent_harness.py
Everything assembled into a harness you can drive. A real permission policy where
write_file asks and run_command is denied, a sandboxed workspace with a command
allowlist, a redaction post-tool hook, a research subagent, durable checkpointing, and a
choice of live event trace or headless JSON.
# One-off task with a live event trace (offline on the mock): the agent delegates
# the lookup to the research subagent, THEN computes with the calculator: a real
# two-step chain, and the final answer reports both.
python hands_on/agent_harness.py "Look up the plans and prices, then compute a year of Pro (30 * 12)."
# Auto-approve the `ask` tools (non-interactive):
python hands_on/agent_harness.py "write file todo.txt containing: ship it" --yes
# Headless: emit a JSON record instead of a trace (for CI / cron):
python hands_on/agent_harness.py "What is (23 * 47) + 100?" --json
# Durable: checkpoint under an id; re-run with the same id to RESUME a crashed run.
python hands_on/agent_harness.py "read the file plan.txt and compute (2 + 2)." --run-id job1
Read hands_on/agent_harness.py. It's the library composed.
build_agent() wires policy, sandbox, hook, and subagent together, and the main loop
consumes the event stream. Suggested exercise: add a second subagent, say a math
specialist, or tighten the policy to deny write_file outright, and watch the trace
change. Adding a capability is one step: register it, and the harness routes to it.
When do you throw away your loop for the SDK?
Here's the honest answer, and the one to give in an interview.
- Write the loop by hand when the agent is simple, with a few tools and one context, or you need to understand exactly what happens, or you're learning. The loop is about 20 lines, and a dependency plus its concepts cost more than that.
- Adopt a harness or SDK when you need any of the pieces this dive built: gated tools, hooks and guardrails, a real sandbox, subagents, structured headless output, durable resumable runs, or orchestration through parallel workers, mid-run steering, and graph control flow. Especially when you'd otherwise reimplement them badly. A harness is a pile of hard, security-sensitive code, covering sandboxing, permission prompts, event plumbing, reconnection, and streaming, that someone else has already hardened.
- What the SDK gives you that this toy doesn't: a real sandbox with containers rather than a path check, provider-hosted per-session workspaces, streaming and reconnection that hold up, subagent orchestration, permission UIs, and headless run records. The productionized version of every piece here.
Two named options exist in the Claude world. The Claude Agent SDK, where you host the compute and the SDK runs the loop and gives you hooks, subagents, permission modes, and sandboxing. And Managed Agents, where Anthropic hosts the loop and a per-session container where tools execute. OpenAI's Agents SDK is the equivalent on that stack. All of them are this dive's harness, hardened and hosted.
Where to go next
You've built a harness from scratch. What comes next is the same pieces, harder.
-
A real sandbox. Swap the path jail for a container or microVM with seccomp, read-only mounts, and network egress rules, or use a provider-hosted sandbox.
-
Richer permission policies. Per-argument rules (allow
read_fileanywhere butwrite_fileonly under/tmp), rate limits, and budgets per run. -
Harder durable execution. §10 and §11 checkpoint to a JSON file and resume. Next comes a DB-backed durable-workflow engine, plus reconnecting a dropped event stream without losing events. Note what that does and doesn't buy you: engines like Temporal make the workflow durable and replayable, while the activities inside it still run at-least-once. Nothing at the engine layer can make an external effect happen exactly once. If the process dies after
charge_cardsucceeded but before its result reached the checkpoint, resuming replays the call, and only an idempotency key the payment provider honors stops the second charge. Exactly-once effects are something you build at the effect, not something a workflow engine hands you. -
Deeper orchestration. §12 through §14 fan out to parallel workers, steer a run mid-flight, and route with a graph. Next comes hierarchical multi-level delegation, backpressure and concurrency limits across many workers, and graph engines with persistence and streaming built in, such as LangGraph and Managed Agents' multiagent coordinator.
-
Agents talking to agents you don't own. Everything above coordinates workers inside one system, where you wrote both ends and share a process, a queue, or a database. The harder version is an agent calling one that belongs to someone else: different owner, different model, no shared state, and no ability to read the other side's code. That needs discovery (what can you do?), a message format, identity and authorization, and a way to attribute cost. Interop protocols in this space, A2A among them, are trying to standardize exactly that layer, roughly where MCP sits for tools.
Worth being clear about what's settled and what isn't. MCP gave an agent a standard way to reach tools and data, and it converged fast. The agent-to-agent layer hasn't, and the reason is instructive: two tool calls are comparable because a tool is a function with a schema, while two agents differ in autonomy, cost per call, failure modes, and what they're allowed to do on your behalf. That's a harder thing to put behind one interface, and it's why the security half, delegated authority in particular, is the part to read first. Everything the Prompt Injection dive says about untrusted input applies to another agent's output, with the twist that this one can be persuaded to act.
-
Provider-hosted tools and agents. Web search, code execution, and computer use run by the provider; and fully managed agents where you never run the loop.
-
Evaluating harness behavior. Score trajectories (right tools, right order, no denied-then-retried loops) with the Evals dive, not just final answers.
From teaching code to production
The shortcuts that make this repo readable and free are exactly what a real harness replaces:
| This repo's teaching shortcut | In production |
|---|---|
| Sandbox is a path check + command allowlist | A container / microVM with seccomp, read-only mounts, and egress rules, or a provider-hosted sandbox |
| Hooks are in-process Python functions | A guardrail pipeline (input/output classifiers, PII redaction, injection detection) wired at the same seam |
| Permission policy is a dict of verdicts | A policy engine with per-argument rules, budgets, rate limits, and an audit log of every decision |
Subagents share one process; fan_out uses a thread pool |
Isolated workers with their own resource limits, backpressure/concurrency caps, and a coordinator that survives a crash |
| Steering is an in-process controller; the graph is plain Python | A durable message queue (steer/interrupt across processes) and a graph engine with persistence, streaming, and observability baked in |
| Events are printed | A structured trace (a span per step) shipped to observability, plus durable run records |
| Checkpoint is a JSON file per run | A durable-execution engine: DB- or workflow-backed state that survives a mid-tool crash (Temporal-style), or a provider's server-side sessions, with idempotency keys on every tool that touches the outside world, because the engine replays at-least-once |
| The mock (or one model) is hard-wired | A model router with fallbacks, retries, and cost/latency budgets per run |
| Headless run is a script | A queue/worker with retries, idempotency, and a webhook or eval gate on the result |
These are right for learning and wrong for production. The general ops machinery observability, cost, reliability, caching, guardrails, prompt versioning, eval gates) is built from scratch and wired into one running app in Production (#8), which also runs offline on a mock provider.
File map
check_setup.py ← run first: verifies Python, packages, provider
README.md ← this guide
EXERCISES.md ← predict-then-run prompts, one per section
harness/ ← the from-scratch harness library (read it!)
providers.py ← the ONLY provider file: mock (default) + openai + claude
tools.py ← what a tool is + a sandboxed toolbox
sandbox.py ← the boundary tools run inside (path jail + command allowlist)
policy.py ← declarative allow / ask / deny permission policy
review.py ← a reviewer that answers ask prompts, then hands back to a person
events.py ← the typed event stream the harness emits
checkpoint.py ← durable run state: persist the transcript, resume after a crash
steer.py ← steering controllers: inject / queue / interrupt a running run
orchestrate.py ← fan out to many subagents concurrently, then join (map-reduce)
graph.py ← orchestration as a graph: nodes, conditional routing, cycles
core.py ← the Harness: loop + hooks + policy + sandbox + subagents + checkpointing + steering
hands_on/
agent_harness.py ← capstone: a configured harness CLI (trace / headless JSON / --run-id resume)
examples/
01_bare_loop_recap.py ← the bare loop and its five missing pieces (offline)
02_harness_events.py ← the same task via the harness's event stream (offline)
03_hooks.py ← pre-tool block + post-tool redaction (offline)
04_permissions.py ← allow / ask / deny policy (offline)
05_sandbox.py ← path jail + command allowlist, escapes refused (offline)
06_subagents.py ← delegate to a nested harness with its own context (offline)
07_headless.py ← one-shot scriptable run → JSON record (offline)
08_computer_use.py ← the loop pointed at a screen; hosted sandboxes (offline sim)
09_checkpoint_resume.py ← durable runs: checkpoint, crash, resume without redoing work (offline)
10_run_records.py ← durable task state: a queryable queued/running/done/failed log (offline)
11_parallel_subagents.py ← fan out to many workers concurrently, then join (offline)
12_steering.py ← inject / interrupt a running agent mid-run (offline)
13_orchestration_graph.py ← routing, branching, and cycles as a graph (offline)
14_managed_agents.py ← the hosted end of the axis: Anthropic runs the harness
15_skills.py ← progressive-disclosure instructions (SKILL.md)
16_model_approval.py ← a reviewer instead of a person on ask, with a hand-off
(workspace/ and runs/ are created by the examples and are git-ignored.)
Troubleshooting
Run python check_setup.py first; it catches most problems. Then, by symptom:
| What you see | What it means / the fix |
|---|---|
ModuleNotFoundError (dotenv / rich) |
Deps aren't installed or the venv isn't active. source .venv/bin/activate then pip install -r requirements.txt. |
PROVIDER=... needs ... in the environment |
You switched to a real provider without a key. Load it from your keychain with secrun (see SECRETS.md), or go back to PROVIDER=mock. |
| A tool ran that I expected to be blocked | Check the policy verdict and your hooks. deny blocks outright; ask runs if your approve callback returns True (the capstone's --yes auto-approves ask, but never overrides deny). |
| "escapes the sandbox" on a path I meant | Working as intended: the jail resolves .. and symlinks and refuses anything outside the root. Use a relative path inside workspace/. |
| The mock takes one step where I expected several | The deterministic planner does one tool per turn for clarity; a real model may chain more. Switch PROVIDER to see it. |
SyntaxError / odd type errors on startup |
You're likely on Python 3.10 or older; this repo needs 3.11+. check_setup.py confirms your version. |
Still stuck? Every file is small and self-contained. Open it, read the docstring at the top, and run it directly. harness/core.py is the whole story: the loop, wrapped.
The series
This is one of the standalone, hands-on deep dives into building with LLM APIs eight core, plus the bonus dives. Each stands on its own, with its own setup, examples, and capstone, and they share one house style: provider-agnostic where it makes sense, built from scratch (no frameworks), offline-first examples, and a real capstone. Do them in any order; this sequence builds naturally:
- OpenAI API: the API from zero
- Claude API: the same ideas, the Anthropic way
- Prompt Engineering: shape model behavior with better prompts
- RAG: answer questions over your own documents
- Evals: measure whether a change actually helps
- Agents: give a model tools and a loop so it can act
- Prompt Injection & Guardrails: attack and defend all of the above
- Production: operate one app end to end
Bonus dives, standalone and slotting in where they're most useful:
- Agent Harnesses: build on the loop, adding hooks, permissions, sandboxing, subagents, and headless runs
- Context Engineering: manage what's in the window
- AI Data Engineering: the corpus behind the index, with versions, lineage, ACLs, and deletes
- Multimodal: images and audio as well as text
- Realtime Voice: low-latency speech-to-speech agents
- Fine-tuning: teach a model new behavior by example
- MCP: serve tools, data, and prompts over a standard protocol
- Local Models: run open-weight models on your own machine
- Observability: watch a running app over time, covering drift, quality, alerting, and the feedback loop
- Architecture: the seams between the components, each decision measured rather than asserted
- GenAI Security: treat the model as an untrusted principal, and put identity, supply chain, isolation, budgets, and release gates around it
- Inference Platform Engineering: turn finite GPU memory and a request queue into latency, throughput, and a fleet size you can defend
- Testing & Delivery: decide whether a build is fit to promote, using evidence, gates, staged rollout, and rollback
- Professional Tools: rebuild each hand-written piece with the tool professionals reach for, and measure both
And the whole series lands in one codebase in the capstone: a codebase Q&A tool built step by step, one tag per dive.
Agent Harnesses is a bonus dive. It slots directly after Agents (#6), since that dive builds the loop; this one builds the layer you run it on.