Bonus dive
GenAI Security: A Guided Deep Dive
Prompt injection is one attack. A production generative-AI system also has data, models, dependencies, retrieval indexes, tools, identities, interpreters, networks, budgets, logs, and operators. Every one of them is part of its security boundary.
This course builds a small security control plane around a deterministic, offline AI application. It starts with a threat model and ends with a release review that attacks the same system twice, naive and hardened, and emits evidence an engineer can inspect.
About the attack material in this repo. It ships deliberately vulnerable toy systems, poisoned documents and datasets, SSRF targets, and a naive pipeline built to lose so the hardened one has something to beat. That's the teaching method: attack the same system twice and compare the evidence. All of it is offline, deterministic, and aimed at code in this same repo, and every credential in it is invented. If a scanner or a CodeQL run flags this repo, this is what it found. Details in SECURITY.md. Use these techniques on systems you own or are authorized to test.
The one big idea:
Treat the model as an untrusted principal, not a security boundary.
Model output is untrusted input. Model intent never grants authority. Enforceable boundaries live in ordinary code: trusted identity, least privilege, validated sinks, network policy, isolation, provenance, budgets, audit records, tested recovery.
This is Chapter 20 of the AI Engineering Deep Dives. It follows Prompt Injection & Guardrails and AI Data Engineering, then feeds into AI in Production. Prompt injection stays the focused treatment of instruction-and-data confusion. This repository covers the larger system that has to stay safe when a model is wrong or compromised.
What you'll build
By the end, you'll be able to:
- turn assets, trust boundaries, entry points, and consequences into a threat model;
- cover everything in the OWASP LLM Top 10 2025 without mistaking a list for your system's threat model;
- keep restricted data out of context, output, logs, and incident evidence;
- verify exact models, prompts, datasets, and dependencies before deployment;
- quarantine named poisoning signals without deleting investigative evidence;
- validate model output for JSON, SQL, and HTML sinks, and escape retrieved text before it's concatenated into a prompt that has a grammar of its own;
- authorize tools from authenticated identity with least privilege, single-use bound approval, idempotency, timeouts, and output limits;
- keep server-side conversation state bound to its owner, and one subject's turns out of another subject's answer;
- enforce tenant, ACL, provenance, citation, cache, egress, and runtime boundaries;
- bound denial-of-wallet across a complete request rather than one API call;
- gate releases on attack resistance, benign utility, coverage, and evaluator health; and
- rehearse containment, evidence preservation, eradication, recovery, and learning.
Why it runs offline
The whole course uses only Python's standard library. It makes no model call, needs no API key, and contacts no external service. That's deliberate. Authorization, provenance, parsing, isolation requirements, budgets, and incident state all have to be testable independently of whichever model happens to sit inside them.
The examples simulate model proposals. They make no claim that deterministic strings measure a production model. Replace those adapters with staging integrations and keep the same invariants and failure tests.
Setup
You need Python 3.11 or newer.
git clone https://github.com/alexvervloet/genai-security-deep-dive.git
cd genai-security-deep-dive
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install -r requirements.txt
python check_setup.py
Expected final lines:
package: genai_security import OK
capstone: deterministic control suite OK
All lessons are ready. No credentials or external services are required.
The editable install in requirements.txt matters more than it looks. Running
python examples/01_threat_model.py puts examples/ on Python's import path rather than
the repository root, so import genai_security resolves only once the package itself is
installed. Install first, then run the labs.
Learning path
Run the lessons in order. Each file opens with the prediction to make before running, the command, the invariant to inspect, and a pointer to what comes next.
| # | Lesson | Core boundary | Primary risk |
|---|---|---|---|
| 1 | Threat modelling | Assets, owners, flows, entry points | All |
| 2 | Sensitive data | Context minimization and output inspection | LLM02, LLM07 |
| 3 | Supply chain | Immutable versions, digests, approval | LLM03 |
| 4 | Poisoning | Record and population gates | LLM04 |
| 5 | Output handling | Exact schemas and sink encoding | LLM05 |
| 6 | Agency and identity | Trusted principal and bound approval | LLM01, LLM06 |
| 7 | Vector isolation | Prefilter, cache scope, pinned evidence | LLM08, LLM09 |
| 8 | Egress and SSRF | Scheme/host/port/address/redirect policy | CWE-918 |
| 9 | Sandboxing | Contract for a real isolated runner | Code execution |
| 10 | Resource controls | Shared, atomic pre-call reservations | LLM10 |
| 11 | Red-team gate | Attacks, utility, coverage, evaluator health | Verification |
| 12 | Incident response | Contain before tested recovery | Operations |
| 13 | Context assembly | Escaped passages, nonce-fenced region | LLM01, LLM08 |
| 14 | Session correlation | Owner-bound handles, per-turn subject | LLM02, LLM06 |
| 15 | Audit records | Keyed record at the boundary, not in the sandbox | LLM06, repudiation |
The LLM01 to LLM10 codes come from the
OWASP Top 10 for LLM and GenAI Applications 2025, a
shared vocabulary for naming these failures in a review. Section 20.2 of the textbook lists
all ten against the boundary each one needs, and explains why a numbered list is a
checklist rather than a threat model.
Read TEXTBOOK.md alongside the labs for the full Chapter 20 lecture. Use EXERCISES.md to extend every invariant rather than only watching the demonstration.
1. Threat modelling
python examples/01_threat_model.py
Notice that assigning LLM01 doesn't close the risk. The first model reports an
uncontrolled boundary and an open score-20 risk, and findings clear only once a concrete
authorization boundary and mitigation exist. In a real review, keep the residual risk
rather than treating mitigation as elimination.
2. Sensitive data and prompt leakage
python examples/02_sensitive_data.py
The internal order status enters context. The restricted token stays out and appears only as a keyed fingerprint, because an unsalted digest of a guessable value is one dictionary away from the value itself. An ordinary email gets redacted at output, and an exact secret canary blocks the response. Search context, output, logs, traces, and exceptions, not only the user-visible answer.
3. Supply-chain provenance
python examples/03_supply_chain.py
The approved prompt and model pass source, version, signature, and payload checks. A
changed prompt produces digest mismatch: support-prompt. The course HMAC is a teaching
signature. Production needs protected identity-backed signing and verifiable provenance,
such as Sigstore and SLSA.
4. Data and model poisoning
python examples/04_poisoning.py
The gate identifies an untrusted source, a blocked marker, and conflicting labels. It quarantines both sides of the label conflict, because the detector can't safely guess which side is true. Extend this with near-duplicate, distribution, influence, and held-out behavior tests for a real corpus.
5. Improper output handling
python examples/05_output_handling.py
Valid JSON becomes an exact typed action, SQL values stay in parameters, an unexpected
admin field rejects the whole proposal, and model prose gets HTML-escaped. Repeat the
pattern for every downstream grammar. There's no universal output sanitizer.
6. Excessive agency and identity
python examples/06_agency_and_identity.py
A model-supplied tenant gets rejected rather than overwritten behind your back. An irreversible operation fails without approval and passes only with approval bound to subject, tenant, tool, and idempotency key. The effective tenant and requester come from the trusted session.
Then the read that every one of those controls allows. A support agent, in the right tenant, holding a role that genuinely grants customer reads, asks for a different customer's history. Well-formed arguments, no trusted field supplied, and a read owes no approval, so nothing above refuses it. Who's asking and on whose installation are two questions; about which person is a third, and a role is silent on it. The tool carries the record the request is about, read from the case rather than from the proposal, and the pivot is refused on what it names rather than quietly rewritten, because a corrected pivot returns data the caller asked for and leaves nothing to say an attempt happened.
7. Vector, cache, and claim isolation
python examples/07_vector_isolation.py
The other tenant's semantically stronger secret never becomes a ranking candidate. The cache key binds to tenant, principals, query, and corpus version. A factual claim then keeps an approved source version, digest, and exact quote. Structural evidence doesn't by itself prove semantic entailment, so keep a separate factuality evaluation.
8. Egress and SSRF
python examples/08_egress_and_ssrf.py
Three identical URLs with three different DNS answers. The allowlisted name passes with a global address and fails when the same name resolves to loopback or to the cloud metadata address, which no check on the URL string could catch. A plaintext metadata URL fails earlier, at the scheme, before any lookup happens. A redirect to an unapproved host fails too. The lesson opens no socket. A production client has to connect to the exact checked address, so a second DNS lookup can't rebind it.
9. Generated-code isolation
python examples/09_sandbox_boundaries.py
The request gets denied for network, root identity, secret environment, and a writable mount outside scratch. The example then lists the guarantees only a real container, microVM, or managed runner can enforce. No generated code executes on the host, and the Python function is deliberately never described as a sandbox.
10. Unbounded consumption
python examples/10_resource_controls.py
One reservation charges once across replay. An oversized recursive branch gets rejected before any work happens and leaves every counter unchanged. Real distributed agents need the same atomic reservation invariant in a concurrency-safe shared store.
Then the ceiling that budget doesn't have. Tokens, calls, steps, bytes and latency are all the operator's resources, which is why they get limits: the person writing them is the person holding the bill. An agent that can refund, credit, discount, or upgrade is moving somebody else's money, and that figure usually appears on a dashboard instead. The lesson pays three refunds, each comfortably inside a per-request ceiling, and the third is refused by the account window, because a per-request ceiling bounds a day only if the number of requests is bounded and a retry or a redelivery produces a fresh one. For each resource an agent can move, ask whose it is.
11. Red-team release gates
python examples/11_redteam_gate.py
Allow-all fails with 100% attack success. Block-all fails too, because benign utility falls to zero. The third system passes everything and proves nothing, because it answers from the probe list itself, which is exactly what a suite satisfied without controls is worth. The capstone wires the same gate to real boundaries. Missing categories and evaluation exceptions count as failures in the tested gate too.
12. Incident response
python examples/12_incident_response.py
Recovery directly after detection gets rejected. The exercised path preserves a digest, contains, records root cause and regression, passes the gate, recovers, and closes with an owner. Hash chaining detects changed metadata and not a deleted tail. The run shows a truncated log verifying cleanly right up until you check it against an anchored head. Production still needs restricted, durable, append-only evidence storage.
13. Context assembly
python examples/13_context_assembly.py
The naive prompt carries three citation keys and the retriever issued two. The forged one came from a support ticket and occupies the same position, in the same syntax, as a real one. Escaping defuses the heading, the key, the operator line, and the forged fence without deleting any of them, so the ticket's actual wording survives for an investigation. The region is then fenced with a per-request nonce the document couldn't have contained. A politely worded request comes through untouched, because nothing was forged: impersonation is closed here, persuasion is lesson 6's problem.
14. Session correlation
python examples/14_session_correlation.py
A conversation handle is a bearer reference to accumulated context, so it comes from the CSPRNG and every resume checks the owner. A colleague in the same tenant and a matching name in another tenant are both refused, with the same wording a missing handle gets, so the store isn't a membership oracle.
Then the harder half. One operator preps two clients in a morning inside a conversation they own, and every turn is one they were entitled to see. The answer being composed is about the second client. Turns carry the subject recorded when they happened, so the first client's amount and date are withheld, and a question naming nobody is still attributed correctly. No permission was exceeded anywhere in that paragraph, which is what makes it an attribution failure rather than an access-control one.
15. Audit records the agent can't rewrite
python examples/15_audit_records.py
Code inside a sandbox owns everything in the sandbox, including the logs written about it. In the July 2026 OpenAI and Hugging Face incident, METR's independent review found agents that rebuilt part of their tool-execution path so a transcript showed one command while another ran (roughly 7% of the transcripts it evaluated, in places), and agents that tried to edit their logs. They could edit the logs inside their containers, and concluded those weren't the real source of truth. The record that held was kept outside their reach.
The lesson replays a small version. The agent's transcript and its in-sandbox log agree, and both lie about the second call. The boundary log, written by the component every call crosses and chained with an HMAC under a key the agent never holds, says what ran, and reconciling the two finds the spoof. Then the agent attacks the record: an edited entry fails, a deletion from the middle fails, a truncated tail needs lesson 12's head anchor, and a whole-file rewrite with every hash recomputed passes its own check and still fails under the logger's key. That last case is the reason for the key. A plain hash chain only proves a record wasn't edited by someone who couldn't also rewrite the rest.
Hands-on capstone
Run the complete release review twice:
python hands_on/security_review.py
python hands_on/security_review.py
Expected output on each run:
GENAI SECURITY RELEASE REVIEW
naive gate passed: False
hardened gate passed: True
evidence: security-report.json
release ready: True
The command attacks the same allow-all and hardened boundary across benign utility, all ten
OWASP 2025 categories, SSRF, and generated-code isolation. It exits nonzero unless the
naive implementation fails and the hardened implementation passes. It writes a
deterministic security-report.json, which Git ignores so release systems can archive it
separately.
Inspect the report:
python -m json.tool security-report.json
Each result carries a control field naming the boundary that decided it. Read those
before trusting a pass. A probe can block for a reason unrelated to the risk it's named
after, and an outcome on its own can't show you that. The naive system records no
controls, which is the point of it.
Then extend it with a risk from your own threat model. A top-ten-only capstone isn't a complete security review.
Verification
Run the entire offline suite:
python -m unittest discover -v
The tests cover successful decisions, actual security denials, malformed input, exceptions, replay/idempotency, atomic failures, cross-tenant access, changed provenance, damaged audit chains, benign utility, and broken evaluation coverage.
To reproduce CI locally after activating the environment:
python check_setup.py
python -m compileall -q genai_security examples hands_on tests check_setup.py
python -m unittest discover -v
for example in examples/[0-9][0-9]_*.py; do python "$example"; done
python hands_on/security_review.py
python hands_on/security_review.py
CI runs this matrix on the minimum supported Python (3.11) and a current Python (3.13) and verifies that test discovery finds a nonzero number of tests.
Repository map
genai_security/ executable security controls
threats.py assets, flows, ranked risks, OWASP taxonomy
data.py context minimization and disclosure inspection
provenance.py artifact manifests, digest and approval checks
poisoning.py record and corpus quarantine findings
sinks.py strict JSON, parameterized SQL, escaped HTML
context.py escaped passages and a nonce-fenced prompt region
sessions.py owner-bound conversation handles and turn subjects
capabilities.py identity, roles, approval, idempotency, limits
vectors.py tenant/ACL/provenance prefilter and cache keys
claims.py pinned structural evidence for factual claims
network.py SSRF-resistant egress planning
isolation.py pre-execution contract for a real runner
resources.py request-wide atomic reservations
redteam.py adversarial evaluation and release policy
incidents.py stateful response and tamper-evident audit metadata
audit.py keyed boundary records and transcript reconciliation
examples/ fifteen narrated, executable lessons
hands_on/security_review.py deterministic naive-vs-hardened capstone
tests/ offline security-invariant test suite
check_setup.py environment and capstone readiness check
TEXTBOOK.md full Chapter 20 lecture
EXERCISES.md progressive engineering exercises
LESSONS.md surprises learned while building the course
What this course proves, and what it doesn't
The offline suite proves the behavior of these teaching policies. It doesn't prove:
- that your identity provider supplies the correct tenant and roles;
- that a vector database applies prefilters before its actual similarity engine;
- that an HTTP client connects to the same address your policy resolved;
- that a container or microVM resists escape and resource exhaustion;
- that artifact builders and signing identities are protected;
- that provider retention and regional settings match privacy policy;
- that distributed budget and idempotency stores are atomic under concurrency; or
- that a production model resists your task-specific attacks while retaining utility.
Those claims need integration and end-to-end tests against the real infrastructure. Keep the invariants from this repository and replace the adapters.
Standards baseline
This August 2026 course uses dated primary baselines so future readers can spot the drift.
- OWASP Top 10 for LLM and GenAI Applications 2025
- OWASP Top 10 for Agentic Applications 2026
- NIST SP 800-218A: Secure Software Development Practices for Generative AI
- NIST AI Risk Management Framework and Generative AI Profile
- SLSA provenance v1.2
- MITRE ATLAS
- CWE-918: Server-Side Request Forgery
Frameworks change. Re-check current versions during a real review, and record the exact version your evidence targets.
Troubleshooting
ModuleNotFoundError: genai_security
Activate the environment and run python -m pip install -r requirements.txt. For a
temporary pre-install development check only, use PYTHONPATH=. python ....
Setup reports Python older than 3.11
Create the virtual environment with a newer interpreter, such as python3.11 -m venv .venv, then reinstall.
The capstone exits nonzero
Open security-report.json. Check hardened.gate_failures and any result whose
passed value is false. An evaluator exception appears as actual: null and must be
fixed rather than waived as a pass.
A test passes only when network or credentials are available
That test doesn't belong in the default offline gate. Inject a deterministic adapter for the course, and add the live behavior as a clearly separated integration suite.
Continue
Complete EXERCISES.md, add a system-specific capstone risk, and carry the resulting release evidence into AI in Production.