av / dives /Agents: Exercises
source about me

Core path - 6 of 8

Exercises: make the learning stick

Reading code teaches you less than predicting what it'll do and then checking. This file turns each section of the README into a few quick active-recall prompts.

How to use it: work the section first, then come back. Commit to an answer before you run or reveal. The prediction is where the learning happens. Answers are hidden behind ▸ toggles.

Examples 01, 07, 10, and 18 are (offline): no API call, no cost. The rest make small, cheap calls.


Section 2: What a tool is (offline)

Recall. When the model "uses a tool," does it run your function? What does it actually do?

▸ Answer

No: the model never runs anything. It emits a request: a tool name plus arguments. Your code decides whether and how to execute it. That gap is where all of an agent's safety lives (approval, sandboxing, validation).

Do (offline). In examples/01_tools.py, the model only ever sees a tool's name, description, and parameter schema, not the function body. Why does that make the description a piece of prompt engineering?

▸ Answer

Because the description is the model's only basis for deciding when and how to call the tool. A vague description ("does stuff") leads to misuse; a precise one ("use this for arithmetic; expression like '2+2'") leads to correct calls. You're programming the model's behavior through that text.


Section 3: One tool call

Predict, then run. In examples/02_one_tool_call.py, you ask "What is 23 * 47?" with the calculator available. Will the model reply with the number, or with something else?

▸ Answer

With a tool call request (calculator, expression="23 * 47"), not the number. It defers the math to the tool. You then run it, and to turn that result into a final answer you'd feed it back and ask again, which is the loop.


Section 4: The agent loop

Recall. Describe the agent loop in one sentence. What makes it stop?

▸ Answer

Ask the model; if it requests tools, run them, append the results, and ask again; repeat until it responds with no tool calls; that final, tool-free response is the answer. (A step cap is the backstop in case it never stops.)

Do. In examples/03_agent_loop.py, give it a question needing two or three calculations and watch the trace. Does it make all the calls up front, or use each result to decide the next?

▸ Answer

Usually one (or a few) at a time, using earlier results to inform later steps that step-by-step, result-dependent reasoning is exactly what the loop enables and a single call can't do.


Section 5: Multiple tools

Predict. Give the agent calculator and search_notes and ask "what does the Plus plan cost per year?" Which tool runs first, and why?

▸ Answer

search_notes first (to find the monthly price, a product fact), then calculator (to multiply by 12). The model routes each sub-task to the tool whose description fits. Nothing hard-codes that order; it's the model's choice, guided by the tool descriptions.


Section 6: Limits & error recovery

Recall. Why feed a tool's error back to the model instead of crashing? And what does max_steps protect against?

▸ Answer

Feeding the error back lets the model see what went wrong and adapt (retry with different inputs, or explain): robustness instead of a dead program. max_steps protects against an infinite loop: a model that keeps calling tools and never finishes would otherwise run forever (and run up a bill).

Do. Set max_steps=1 on a task that clearly needs several steps. What does the agent return, and what does stopped_early tell you?

▸ Answer

It returns a "stopped early" message and stopped_early=True, because it hit the ceiling before reaching a final answer. That flag is how your code knows the result is incomplete rather than a real answer.


Section 7: Human-in-the-loop

Recall. What makes a tool require approval, and what happens to the agent when you deny one?

▸ Answer

You mark it dangerous=True; the loop then calls your approve callback before running it. With no callback it fails closed; an explicit denial is returned to the model as a normal tool result ("permission denied"), so the agent acknowledges it and adapts instead of forcing the action.


Section 7A: Tool contracts (offline)

Predict, then run. Before running examples/18_tool_contracts.py, write down which of these proposals should cause a refund: valid billing request, model-supplied tenant, amount above the schema maximum, support-role request, and a replay of the valid request. Then run it.

▸ Answer

Only the first proposal crosses the effect boundary. The tenant attempt is rejected as trusted-context forgery before ordinary schema validation; the oversized amount fails schema validation; the support principal fails authorization; and the replay returns the first completed result. The final effect count is one.

Recall. Why do OpenAI strict mode and local JSON Schema validation both exist? Which one authorizes the caller?

▸ Answer

Strict mode constrains what the provider generates and reduces malformed calls. Local validation treats the resulting request as untrusted at the execution seam and works regardless of provider or transport. Neither authorizes the caller; authorization uses roles and identity from authenticated ExecutionContext, not the model's arguments.

Counterfactual. In tests/test_tool_contracts.py, change one thing at a time: raise amount_cents from 5,000 to 5,001; change the context role from billing to support; add tenant_id to the proposal; replay with a different request ID. Predict which decision code changes and whether the callable count changes before running the test.

▸ Answer

The amount becomes schema_validation, the role becomes not_authorized, and the tenant becomes trusted_context_forgery; none invokes the callable. A new request ID changes the replay scope, so an otherwise allowed call executes a new effect. These independent perturbations prove the decision didn't derive its expected result from the proposal itself.

Predict. A mutating call is replayed by (request_id, call.id, tool). Suppose the same call.id comes back inside the same request, but this time asking to refund a different order for ten times the money. Should the executor return the stored result, run the new one, or neither?

▸ Answer

Neither: it denies with idempotency_key_reuse. Returning the stored result would answer a question nobody asked and bury the second attempt (the audit record would carry the first call's argument digest, so the log wouldn't even show it happened). Running it would defeat the point of the key. A settled key is a promise about one specific payload, which is why the executor compares argument digests and not just the key. Stripe's API behaves the same way for the same reason.

Design. The teaching executor records a timeout and replays it for a repeated mutating call. Why not automatically retry, and what must replace this cache in a multi-worker production service?

▸ Answer

A timeout says the caller stopped waiting, not that the effect failed; the tool may have committed just before the connection disappeared. An automatic retry could duplicate it. Production needs a durable idempotency key enforced transactionally by the sink (plus coordination across workers), because an in-process bounded cache disappears on restart and doesn't stop concurrent workers from racing.


Section 8: Observability (offline)

Predict, then run. In examples/07_observability.py, the same mutating call is proposed twice with one trusted request context. Which fields should differ between the two Step records, and how many events should reach the sink?

▸ Answer

Only replayed changes: it's False for the dispatch and True for the cached retry. Both steps keep status=ok, approval=approved, and the same argument and output digests because they describe the same settled operation. Exactly one event reaches the sink.

Recall. Why is a step trace essential for agents specifically, more than for a single LLM call?

▸ Answer

Because an agent makes its own multi-step decisions (which tools, in what order, with what arguments. When it goes wrong, the why is in that sequence. Without a trace you're debugging a black box; with one you can see (and later eval) exactly what it did.


Section 9: Memory

Predict, then run. In examples/08_memory.py, ask it to "search the plans," then ask "which is the cheapest paid one?" Does the second question work? Why?

▸ Answer

It works, because the same history list is passed back each turn, so the earlier search and its results are still in the conversation the model sees. The API is stateless; the memory is the growing list you choose to resend.


Section 10: Multi-agent

Recall. In examples/09_multi_agent.py, what is the research sub-agent, from the orchestrator's point of view?

▸ Answer

Just another tool. Calling research(question=...) looks identical to calling any tool, but its function body runs a second agent loop with its own prompt and tools. Sub-agents are agents wrapped as tools; that's how large systems decompose.


Capstone: agent_cli.py

Do. Run secrun python hands_on/agent_cli.py "Save a note titled 'todo' with body 'x'" and deny the approval prompt. Then run it again with --yes. Then add --trace to either. You've now exercised the loop, the approval gate, and observability in one tool.

Stretch. Add a new tool to agent/tools.py (say, a word_count(text) tool), include it in default_tools(), and ask the capstone to use it. Adding a capability is just: write a function, describe it, register it. That's the whole extensibility story.


Going further: four more agent patterns

Recall (11). You can solve a support task with a hard-coded workflow or with the agent loop. When should you not reach for an agent?

▸ Answer

When you can already draw the flowchart. If the steps are known, a workflow (classify → route → handle, orchestrated in code) is cheaper, more predictable, and easier to test. An agent is worth its cost and unpredictability only when the path is genuinely open-ended.

Recall (12). Planning and reflection are both extra LLM passes around the loop. What does each buy you, and what's the weakness of self-reflection?

▸ Answer

Plan (before) keeps a multi-step task from dropping a sub-goal. Reflect (after) catches half-answers and slips, then revises. The weakness: a model grading itself is fallible; the strong version grounds the critic in a real check (tests, a schema, a verifier), as in the prompt-engineering "reflexion" lesson.

Predict (13). The model asks for three independent get_weather calls in one turn. Why does running them concurrently help, and roughly how much?

▸ Answer

Independent calls don't depend on each other, so there's no reason to wait. Run them concurrently and the turn finishes in the time of the slowest call, not the sum. Three ~0.8s calls: ~2.4s sequential vs ~0.8s parallel (~3×). The final answer is a normal completion, so you can stream it on top.

Recall (14). Example 13 streams the final answer; example 14 streams every turn. What has to change in the loop, and what's the one fiddly part of streaming a turn that also calls a tool?

▸ Answer

Almost nothing changes. You swap run_turn for stream_turn and print the text deltas via a callback; the rest of the loop (run tools, feed results back, repeat) is identical. The fiddly part is reassembling the tool call: OpenAI streams the arguments as JSON fragments you must concatenate by index before you can parse them. That lives in agent/providers.py so the loop stays clean.


Section: Provider-hosted tools

Predict (15). You run examples/15_hosted_tools.py with the hosted web-search tool declared. How many client-side tool rounds does your loop handle, and does that mean no search happened?

▸ Answer

Zero client-side rounds, and search did happen. A hosted tool is run by the provider inside the turn, not by your loop. You send one request and get one final answer already grounded in search; there's no tool_use/tool_result round-trip for your code, even though the provider ran search one or more times (the example prints that count). Client-executed and hosted tools are different in kind.

Recall. What do you give up by using a hosted tool instead of a client-executed one, and what do you gain?

▸ Answer

You give up control: the call runs on the provider's side, so Section 7's approval gate can't reach it, and you can't custom-log, validate, or sandbox it. You gain zero plumbing: no schema, no execution, no result round-trip. Real agents mix both: client tools for actions you must govern, hosted tools (web search, code execution) for capability you're happy to rent.


Bonus: MCP: a tool you didn't ship with (offline)

Predict (10). The client turns each remote tool descriptor into an ordinary Tool object and hands it to the loop. When the agent later "runs" one, what does that tool's func actually do, and could run_agent from example 03 tell the difference from a local tool?

▸ Answer

The func is a closure that sends a tools/call request over the protocol (JSON lines over the server's stdin/stdout) and reads the text content blocks back. The loop can't tell the difference; the converted tool has the same name, description, schema, and callable shape as a local one, so tracing and approval gates apply unchanged. That's the payoff of the conversion step: the transport is invisible to the loop.

Recall. Example 01 said a tool is "a name, a description, and a JSON Schema." What does MCP add to that picture, and what still never happens, whether the tool is local or served across the world?

▸ Answer

MCP adds discovery and a wire format: tools/list advertises those same three things from another process (another team, another language, another machine), and tools/call invokes one by name with JSON arguments, with no hand-written glue per tool. What never changes: the model still never runs anything. It only ever sees the menu entry and emits a request. Now even your client doesn't hold the implementation; the server does. (For the protocol's other primitives, like resources, prompts, HTTP transport, and security, see the MCP deep dive.)

Predict. You point the client at a different team's MCP server. Its schema is valid JSON Schema and describes every field it wants, but never mentions additionalProperties. You pass the discovered tool to run_agent. What happens?

▸ Answer

ToolExecutor raises ValueError before a single model call, because it refuses any schema that hasn't closed the door on undeclared fields. That's deliberate: a schema this loose would let the model slip in an argument nobody described, and the repo would rather fail loudly at wiring time than silently forward it. seal_schema() is the fix, and as_tools() applies it to every descriptor as it comes off the wire. Note the cost, which is real: if that server accepts a field it never declared and says nothing, sealing rejects a call it would have honored.

Recall. The server in this repo runs every tools/call through the same ToolExecutor the agent loop uses. The client already validates. Why check twice?

▸ Answer

Because neither end can verify the other did it. The server has no idea whose model is on the other end of the pipe, which prompt it was given, or whether that prompt came off an injected web page; "my client validated" isn't a fact a server can check. Symmetrically, the client can't audit the server's code. Each side enforces the contract at its own trust boundary, which is what makes them boundaries. Skipping either is how an internal service ends up trusting arguments that originated in text a stranger wrote.


Where to take it next

Invent your own. Wire one of your real functions in as a tool, something that reads a file, hits an API you own, or queries a database. Mark it dangerous if it writes, and let the agent drive it. The first time an agent completes a multi-step task you didn't script step-by-step, the idea has landed.