Bonus dive
Exercises
For each lesson, write your prediction before running the example. Then make the smallest change requested, rerun the relevant test, and explain the observed decision in terms of policy, stimuli, evidence, units, or state.
1. Test portfolios
Run:
python examples/01_test_portfolio.py
Before running, predict the exact missing-evidence reason. Then:
- Add a green
loadobservation without adding it torequired_kinds. Explain why it's extra evidence rather than a release requirement. - Add
loadto the policy but remove its observation. Confirm the decision changes. - Supply two unit results, one passing and one failing. Explain why choosing either one silently would hide flakiness.
- Name one failure that an eval catches but a unit test usually doesn't, and one in the opposite direction.
2. Recorded SDK contracts
Run:
python examples/02_contract_fixtures.py
Predict whether the extra expected_id_type field repairs the missing id. Then:
- Put a synthetic
Bearervalue three levels deep in the fixture and locate the path in the violation. - Set
allow_unknown_response_fields=Falseand observe the forward-compatibility tradeoff. - Add an optional field to the independent contract, then omit it from the fixture.
- Design a fixture-refresh review that can't automatically approve the shape it just recorded.
3. Property testing
Run:
python examples/03_property_testing.py
Predict the minimal counterexample for the broken lower clamp. Then:
- Repair
buggy_clampand confirm all generated cases pass. - Change the invariant so zero itself fails. Trace the already-minimal path.
- Use a very large failing value and a tiny shrink budget. Confirm the report doesn't claim complete shrinking.
- Write one property for tenant filtering or context-window packing. State which inputs are generated and which requirement stays independent.
4. Deterministic doubles
Run:
python examples/04_deterministic_doubles.py
Predict which configured fake fails the fixed product requirement. Then:
- Add a one-shot error and prove the following call uses reusable behavior.
- Make five calls with
max_history=2; compare retained history with total count. - Write the circular anti-pattern
expected = fake.generate(input)and explain whyfake.generate(input) == expectedproves nothing useful. - Identify a real SDK interface that should use autospeccing in addition to a wire fixture.
5. Load tests and units
Run:
python examples/05_load_testing.py
Predict whether shifting timestamps by 3,600 seconds changes throughput. Then:
- Hand-calculate span, throughput, p95 milliseconds, and error ratio.
- Add exactly enough latency to meet the maximum and confirm equality passes. Add one more millisecond and confirm it fails.
- Replace
span_swith the absolute maximum finish timestamp temporarily. Show why the time-origin metamorphic test catches the mistake. - List workload dimensions the synthetic test doesn't model.
6. Faults and idempotency
Run:
python examples/06_fault_testing.py
Predict the side-effect count for each writer. Then:
- Move the fault before commit and compare both writers.
- Exhaust all four attempts and verify the delay schedule is 25, 50, then 100 ms.
- Raise
lost_responsesabove the attempt budget. Explain why the idempotent writer now fails without ever committing a second effect. - Set
max_keys=2, write keys A, B, C, then retry A. Explain the fourth effect. - Choose an idempotency retention duration for a queued job system and justify it from actual retry and redelivery horizons.
7. Artifact compatibility
Run:
python examples/07_compatibility.py
Predict the index-schema violation. Then:
- Set model context to exactly prompt input plus headroom, then one token below.
- Remove the
jsonmodel feature and inspect the independent feature requirement. - Try
allowed_models={"model-*"}and then{"*"}. Explain the deliberate wildcard semantics. - Add an embedding-model revision to both candidate and policy. Write a test where dimensions match but the semantic embedding revision doesn't.
8. Dependency locking
Run:
python examples/08_dependency_locking.py
Predict the unhashed artifact violation. Then:
- Add a SHA-256 fixture value and confirm the teaching audit passes.
- Add a VCS source with
requested-revision="main"and a full immutablecommit-id. Explain which value an installer must use. - Add two legal marker variants for the same normalized name, then add an exact duplicate entry.
- Explain what the audit can't prove without downloading bytes and evaluating marker expressions.
9. CI matrix coverage
Run:
python examples/09_ci_matrix.py
Predict why one green current-runtime job is insufficient. Then:
- Fail only the Python 3.11 cell and inspect the run-specific reason.
- Add an optional failing macOS live-provider cell. Explain why it doesn't alter this policy, then decide whether your production policy should require it.
- Duplicate a required cell with one green and one red observation. Explain why the teaching gate refuses to cherry-pick.
- Find the declared Python minimum in
pyproject.tomland the matching real workflow cell in.github/workflows/ci.yml.
10. Security scanning
Run:
python examples/10_security_scanning.py
Predict whether a completed scan with a high finding passes. Then:
- Upgrade
unsafe-parserbeyond the synthetic affected set and rerun the join. - Lower the finding to medium without changing policy, then lower the policy bar. Identify which is observation and which is requirement.
- Add a low-severity finding with a blocked license and confirm license policy acts independently from severity.
- Write the narrow claim a clean result supports without saying the application is vulnerability-free.
11. Staged rollouts
Run:
python examples/11_staged_rollout.py
Predict the decisions at 500 and 1,000 requests. Then:
- Put every metric exactly on its boundary and confirm promotion.
- Use ten requests with a severe quality regression. Confirm safety rollback wins over the insufficient-volume hold.
- Mark rollback unverified and observe the blocked state.
- Add one external side effect in shadow and compare the same observation at canary. Explain the deliberate stage asymmetry.
12. Evidence lineage and freshness
Run:
python examples/12_release_evidence.py
Predict why prompt v6 unit evidence can't release prompt v7. Then:
- Test ages 3,599, 3,600, and 3,601 seconds against a 3,600-second maximum.
- Put the production time one second in the future and inspect the reason.
- Reverse record input order and verify deterministic bundle order.
- Change one decision payload field without updating its record. Recompute its digest and explain why a signed production verifier should reject the mismatch.
13. Capstone challenge
Run:
python hands_on/release_candidate.py
python -m unittest tests.test_capstone -v
Before reading the tests, predict the full rollout path and evidence count. Then:
- Trace the load summary from request events through its decision payload, payload digest, evidence record, portfolio, and final release flag.
- Remove the minimum-runtime CI observation. Confirm the current-runtime job stays green while the release fails.
- Change only the index schema. Confirm every evidence record remains bound to the newly derived candidate subject while compatibility fails.
- Change the writer to non-idempotent. Confirm eventual retry success is reported but the real side-effect requirement fails.
- Set canary observations between the passing and failing examples. Find a value that holds, one that promotes, and one that rolls back without relying on a tie.
- Make all decisions pass but age all records beyond policy. Explain why build-time success and releasable current evidence are separate claims.
- Add an
artifact_signaturefield and a verifier interface. Keep the course offline by using a deterministic teaching signer, and state why it isn't a production trust root.
Senior review checklist
Review a real AI delivery pipeline and answer with file paths or workflow links:
- Where are release requirements declared independently from results?
- Which quality claims are evals, and which deterministic code paths have unit tests?
- Which external SDK boundaries have schema contracts and sanitized fixtures?
- Which seeds or replay tokens reproduce generated failures?
- What mutable fake or idempotency state is retained, expired, and bounded?
- Which load metrics name their units and percentile definition?
- Which fault test observes side effects rather than only response status?
- What exact prompt/model/index/SDK/dependency tuple was tested?
- Does CI execute the minimum declared runtime and every supported platform path?
- Which scanners are required, and what narrow claim does a clean result support?
- What production evidence promotes a canary, and what independently tested path reverses it?
- Can every release headline be traced to a source observation and subject digest?