av / dives /Production: Exercises
source about me

Core path - 8 of 8

Exercises: make the learning stick

Reading code teaches you less than predicting what it'll do and then checking. This file turns each section of the README into a few quick active-recall prompts: a thing to predict, a thing to change, and a question to answer from memory.

How to use it: work the section first, then come back. Commit to an answer before you run or reveal. The prediction is where the learning happens, even (especially) when you're wrong. Answers are hidden behind ▸ toggles.

Everything here is offline: the default mock provider needs no key and costs nothing. Run as much as you like.


Section 2: The app and the mock

Predict, then run. Run examples/00_mock_provider.py twice. Will the token counts differ between runs? Why does that matter for the rest of the repo?

▸ Answer

No: the mock is deterministic, so the same question yields the same answer and the same token counts every time. That determinism is what makes caching, cost, and evals demonstrable: a cache hit is provably identical, a cost is reproducible, and a gold answer is stable to grade against.


Section 3: Observability

Recall. A user reports "the assistant was slow an hour ago." You have a print() in your code. Why doesn't that help, and what does a trace give you that it doesn't?

▸ Answer

The print is gone. It scrolled past on a screen you weren't watching, for a request that already finished. A trace is a stored record with a unique id and per-span timings, so you can find that request later and see exactly which span (cache, model call, guardrails) ate the time.

Change it. In examples/01_observability.py, the spans sleep for fixed durations. Predict which span dominates the trace summary, then change the sleeps and confirm the spans dict tracks your change.


Section 4: Cost

Predict, then run. examples/02_cost.py sets a $0.00008 budget. Before running, guess how many of the five questions get answered before BudgetExceeded. Then run it.

▸ Answer

Three. Each mock call costs about $0.000024 at the gpt-6-luna rate, so the fourth would take the total to about $0.000094 and the budget refuses it. The exact cutoff depends on token counts. The lesson isn't the number: it's that the budget refuses the call rather than spending past the limit. check() runs before the spend; record() after.


Section 5: Reliability

Predict, then run. In examples/03_reliability.py, the circuit breaker has fail_threshold=3. On which call does it flip from "closed" to "open", and what changes about calls after that?

▸ Answer

It opens on the 3rd failure. From call 4 on, it fails fast, raising "circuit open" without even calling the provider, until the cooldown elapses. That's the point: a sick provider stops getting hammered, and your users fail fast instead of waiting through the full backoff each time.

Recall. Why retry a 503 but not a 400?

▸ Answer

A 503 is transient: the server is briefly unavailable, and the same request often succeeds on retry. A 400 means the request itself is malformed; retrying it unchanged just wastes time and quota. Only retry what a retry could fix.

Predict, then run. Demo 4 in examples/03_reliability.py sends two 429s. One has the code slow_down and a Retry-After of 0.4 seconds; the other has the code credit_balance_exhausted. How many attempts does each get, and how long does the first one wait before retrying, given a 0.05-second base delay?

▸ Answer

The slow_down waits 0.4 seconds, not 0.05: Retry-After is a floor on the backoff, because retrying sooner just earns another slow_down. Then it succeeds on attempt 2. The credit_balance_exhausted gets exactly one attempt. It's a 429 too, but it's about money rather than speed, so classify_error() makes it a PermanentProviderError and the retry layer passes it straight up. The status code alone would have retried it four times and still failed.


Section 6: Caching

Predict, then run. examples/04_caching.py asks the password question under v1, then again under v2. The text of both answers is nearly identical. Will the second be a cache HIT? Why or why not?

▸ Answer

MISS. The cache key hashes the prompt version too, not just the question, and v1 ≠ v2. That's deliberate: a prompt change must invalidate the cache, or you'd serve answers shaped by the old prompt after shipping a new one.


Section 7: Guardrails

Predict, then run. In examples/05_guardrails.py, two outputs contain an email: one to jane.doe@gmail.com, one to support@acme.example. Predict what the output guard does to each.

▸ Answer

It redacts jane.doe@gmail.com (PII) but leaves support@acme.example alone it's on the allowlist of addresses the app is supposed to surface. A blunt PII filter that scrubs your own help desk email is worse than useless, so every real redactor needs an allowlist.


Section 8: Prompt versioning

Recall. The prompt lives in prompts/v2.txt, not inline in the code. Name two things that becomes possible because of that.

▸ Answer

(1) A rollout/rollback is a config change (PROMPT_VERSION) or a one-line file revert, not a code edit. (2) The eval gate can score a new version against the old one on the same dataset before you promote it. Bonus: prompt changes show up in git diff like any other code.


Section 9: Eval gates

Predict, then run. examples/07_eval_gate.py scores v1 and v2. One fails the gate. Which, and on which case? Predict before running.

▸ Answer

v1 fails the cite-source case. Its prompt doesn't ask the model to cite, so the answer lacks "Source," and the gold set requires it. v2 (constrained, cites sources) passes all cases. The gate's non-zero exit is what would block the merge in CI.


Section 10: The capstone

Predict, then run. Start the server (python hands_on/serve.py --server --port 8099) and curl the same question twice, then hit /metrics. What will cache_hit_rate be, and what will the second request's cost_usd be?

▸ Answer

The second request is a cache HIT, so its cost_usd is 0 and cache_hit_rate climbs toward 0.5 (one hit out of two asks). Every layer you built (trace, guard, cache, budget, prompt version) is visible in the JSON response and the /metrics summary. That's one operable service, offline, no key.


Going further: five more production concerns (offline)

Predict (semantic caching, 09). "How do I reset my password?" is cached. A new query "How can I reset my password if I forgot it?" arrives. Exact-match cache: hit or miss? Semantic cache: hit or miss? What's the danger of setting the threshold too low?

▸ Answer

Exact-match misses (different bytes); the semantic cache hits (high similarity) and saves a call. Too low a threshold and you serve a similar-but-wrong cached answer to a genuinely different question, so you tune it, and keep exact-match for things that must never be confused.

Recall (fallback/routing, 10). Name the two ways a second model helps, and how each differs from the retries in Section 5.

▸ Answer

Failover: when the primary is down even after retries, serve from a backup (cheaper model or canned answer) instead of erroring. Cost routing: send easy questions to a cheap model, hard ones to the expensive model, to cut the bill. Retries re-call the same model; these reach for a different one.

Recall (rate limiting & feedback, 11). What does a per-tenant token bucket protect against, and why is a thumbs-down the most valuable signal you can log?

▸ Answer

It stops any one client/tenant from starving a shared, costly backend (fairness, cost control, multi-tenancy), so one tenant's burst is capped without affecting others. A thumbs down is a labelled example of something your system got wrong, exactly the regression test (evals dive) and fine-tuning data that makes the next version better.

Predict (unit economics, 12). Your workflow costs $0.024 a call and a person spends about a minute checking each result. Someone proposes a quarter of engineering work to cut the model bill by 90%. Roughly how much does that move the cost of a finished task, and what would you propose instead?

▸ Answer

Almost nothing, around 3%. At $45/hour loaded, one minute of review is $0.75, which is twenty times the model call, so the model is a rounding error in the total and cutting it in ten still leaves the review. Halving review time moves the same number by about 48%, and that's the work worth funding: better retrieval so there's less to correct, a confidence signal that routes only the uncertain cases to a human, or an interface that makes an edit take twenty seconds instead of sixty.

The general lesson is about which numbers are visible. The model bill has a dashboard and an invoice, so it gets the attention. Review time has neither, so it gets assumed away, and it's usually the larger number. Run the example and check the assumptions at the top of the file against your own workflow before you trust the conclusion.

Refusals (13_refusal.py)

Predict, then run. A safety classifier declines one of your requests. Your code is wrapped in try/except and sits behind the retry layer from Section 5. Which of your defenses catches it: the except, the retry, the error-rate alert, or none?

▸ Answer

None. A refusal is not an exception. The HTTP call succeeds with a 200, the response carries stop_reason: "refusal", and text is empty. There is nothing for except to catch, so the retry never fires and the error rate never moves. The request is counted as a success by every meter you have. That's what makes it worse than an outage: an outage is loud.

Do. Run examples/13_refusal.py and look at part 2. One refusal happened. How many users got an empty answer, and why?

▸ Answer

All of them, indefinitely. The naive path stores resp.text in the cache without asking why generation stopped, so the empty string becomes the cached answer for that question and every later request is served from it. The refusal was transient; the cache made it permanent.

Two layers that each look correct alone combine into a bug, which is the general shape worth remembering. The cache's contract is "store what the model returned." The refusal's contract is "the model returned nothing, and told you why in a field you didn't read." Anything that persists or scores a model response, a cache, an eval harness, a golden-output file, needs to know what a refusal is.

Recall. Where should the refusal count go, and why isn't "log it and move on" enough?

▸ Answer

Next to successes and errors on the Section 1 dashboard, as its own counter. A refusal rate that moves means either a provider policy changed under you or one of your prompts started tripping a classifier, and those have completely different fixes. You can only tell them apart if you were counting before it moved.

"Log it and move on" leaves the user with a blank reply. A refusal needs a decided outcome: a different model, a canned response, or a human. Anthropic's server-side fallbacks will route by refusal category if you opt in, which is a reasonable default precisely because it forces the decision to exist.


Done? You've operated one app end to end. The "Where to go next" section of the README maps each from-scratch layer here to its industrial counterpart, same interfaces, bigger machinery.