av / dives /GenAI Security
source about me

Bonus dive

GenAI Security: A Guided Deep Dive

Prompt injection is one attack. A production generative-AI system also has data, models, dependencies, retrieval indexes, tools, identities, interpreters, networks, budgets, logs, and operators. Every one of them is part of its security boundary.

This course builds a small security control plane around a deterministic, offline AI application. It starts with a threat model and ends with a release review that attacks the same system twice, naive and hardened, and emits evidence an engineer can inspect.

About the attack material in this repo. It ships deliberately vulnerable toy systems, poisoned documents and datasets, SSRF targets, and a naive pipeline built to lose so the hardened one has something to beat. That's the teaching method: attack the same system twice and compare the evidence. All of it is offline, deterministic, and aimed at code in this same repo, and every credential in it is invented. If a scanner or a CodeQL run flags this repo, this is what it found. Details in SECURITY.md. Use these techniques on systems you own or are authorized to test.

The one big idea:

Treat the model as an untrusted principal, not a security boundary.

Model output is untrusted input. Model intent never grants authority. Enforceable boundaries live in ordinary code: trusted identity, least privilege, validated sinks, network policy, isolation, provenance, budgets, audit records, tested recovery.

This is Chapter 20 of the AI Engineering Deep Dives. It follows Prompt Injection & Guardrails and AI Data Engineering, then feeds into AI in Production. Prompt injection stays the focused treatment of instruction-and-data confusion. This repository covers the larger system that has to stay safe when a model is wrong or compromised.

What you'll build

By the end, you'll be able to:

  • turn assets, trust boundaries, entry points, and consequences into a threat model;
  • cover everything in the OWASP LLM Top 10 2025 without mistaking a list for your system's threat model;
  • keep restricted data out of context, output, logs, and incident evidence;
  • verify exact models, prompts, datasets, and dependencies before deployment;
  • quarantine named poisoning signals without deleting investigative evidence;
  • validate model output for JSON, SQL, and HTML sinks, and escape retrieved text before it's concatenated into a prompt that has a grammar of its own;
  • authorize tools from authenticated identity with least privilege, single-use bound approval, idempotency, timeouts, and output limits;
  • keep server-side conversation state bound to its owner, and one subject's turns out of another subject's answer;
  • enforce tenant, ACL, provenance, citation, cache, egress, and runtime boundaries;
  • bound denial-of-wallet across a complete request rather than one API call;
  • gate releases on attack resistance, benign utility, coverage, and evaluator health; and
  • rehearse containment, evidence preservation, eradication, recovery, and learning.

Why it runs offline

The whole course uses only Python's standard library. It makes no model call, needs no API key, and contacts no external service. That's deliberate. Authorization, provenance, parsing, isolation requirements, budgets, and incident state all have to be testable independently of whichever model happens to sit inside them.

The examples simulate model proposals. They make no claim that deterministic strings measure a production model. Replace those adapters with staging integrations and keep the same invariants and failure tests.

Setup

You need Python 3.11 or newer.

bash
git clone https://github.com/alexvervloet/genai-security-deep-dive.git
cd genai-security-deep-dive
python3 -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate
python -m pip install -r requirements.txt
python check_setup.py

Expected final lines:

package: genai_security import OK
capstone: deterministic control suite OK

All lessons are ready. No credentials or external services are required.

The editable install in requirements.txt matters more than it looks. Running python examples/01_threat_model.py puts examples/ on Python's import path rather than the repository root, so import genai_security resolves only once the package itself is installed. Install first, then run the labs.

Learning path

Run the lessons in order. Each file opens with the prediction to make before running, the command, the invariant to inspect, and a pointer to what comes next.

# Lesson Core boundary Primary risk
1 Threat modelling Assets, owners, flows, entry points All
2 Sensitive data Context minimization and output inspection LLM02, LLM07
3 Supply chain Immutable versions, digests, approval LLM03
4 Poisoning Record and population gates LLM04
5 Output handling Exact schemas and sink encoding LLM05
6 Agency and identity Trusted principal and bound approval LLM01, LLM06
7 Vector isolation Prefilter, cache scope, pinned evidence LLM08, LLM09
8 Egress and SSRF Scheme/host/port/address/redirect policy CWE-918
9 Sandboxing Contract for a real isolated runner Code execution
10 Resource controls Shared, atomic pre-call reservations LLM10
11 Red-team gate Attacks, utility, coverage, evaluator health Verification
12 Incident response Contain before tested recovery Operations
13 Context assembly Escaped passages, nonce-fenced region LLM01, LLM08
14 Session correlation Owner-bound handles, per-turn subject LLM02, LLM06
15 Audit records Keyed record at the boundary, not in the sandbox LLM06, repudiation

The LLM01 to LLM10 codes come from the OWASP Top 10 for LLM and GenAI Applications 2025, a shared vocabulary for naming these failures in a review. Section 20.2 of the textbook lists all ten against the boundary each one needs, and explains why a numbered list is a checklist rather than a threat model.

Read TEXTBOOK.md alongside the labs for the full Chapter 20 lecture. Use EXERCISES.md to extend every invariant rather than only watching the demonstration.

1. Threat modelling

bash
python examples/01_threat_model.py

Notice that assigning LLM01 doesn't close the risk. The first model reports an uncontrolled boundary and an open score-20 risk, and findings clear only once a concrete authorization boundary and mitigation exist. In a real review, keep the residual risk rather than treating mitigation as elimination.

2. Sensitive data and prompt leakage

bash
python examples/02_sensitive_data.py

The internal order status enters context. The restricted token stays out and appears only as a keyed fingerprint, because an unsalted digest of a guessable value is one dictionary away from the value itself. An ordinary email gets redacted at output, and an exact secret canary blocks the response. Search context, output, logs, traces, and exceptions, not only the user-visible answer.

3. Supply-chain provenance

bash
python examples/03_supply_chain.py

The approved prompt and model pass source, version, signature, and payload checks. A changed prompt produces digest mismatch: support-prompt. The course HMAC is a teaching signature. Production needs protected identity-backed signing and verifiable provenance, such as Sigstore and SLSA.

4. Data and model poisoning

bash
python examples/04_poisoning.py

The gate identifies an untrusted source, a blocked marker, and conflicting labels. It quarantines both sides of the label conflict, because the detector can't safely guess which side is true. Extend this with near-duplicate, distribution, influence, and held-out behavior tests for a real corpus.

5. Improper output handling

bash
python examples/05_output_handling.py

Valid JSON becomes an exact typed action, SQL values stay in parameters, an unexpected admin field rejects the whole proposal, and model prose gets HTML-escaped. Repeat the pattern for every downstream grammar. There's no universal output sanitizer.

6. Excessive agency and identity

bash
python examples/06_agency_and_identity.py

A model-supplied tenant gets rejected rather than overwritten behind your back. An irreversible operation fails without approval and passes only with approval bound to subject, tenant, tool, and idempotency key. The effective tenant and requester come from the trusted session.

Then the read that every one of those controls allows. A support agent, in the right tenant, holding a role that genuinely grants customer reads, asks for a different customer's history. Well-formed arguments, no trusted field supplied, and a read owes no approval, so nothing above refuses it. Who's asking and on whose installation are two questions; about which person is a third, and a role is silent on it. The tool carries the record the request is about, read from the case rather than from the proposal, and the pivot is refused on what it names rather than quietly rewritten, because a corrected pivot returns data the caller asked for and leaves nothing to say an attempt happened.

7. Vector, cache, and claim isolation

bash
python examples/07_vector_isolation.py

The other tenant's semantically stronger secret never becomes a ranking candidate. The cache key binds to tenant, principals, query, and corpus version. A factual claim then keeps an approved source version, digest, and exact quote. Structural evidence doesn't by itself prove semantic entailment, so keep a separate factuality evaluation.

8. Egress and SSRF

bash
python examples/08_egress_and_ssrf.py

Three identical URLs with three different DNS answers. The allowlisted name passes with a global address and fails when the same name resolves to loopback or to the cloud metadata address, which no check on the URL string could catch. A plaintext metadata URL fails earlier, at the scheme, before any lookup happens. A redirect to an unapproved host fails too. The lesson opens no socket. A production client has to connect to the exact checked address, so a second DNS lookup can't rebind it.

9. Generated-code isolation

bash
python examples/09_sandbox_boundaries.py

The request gets denied for network, root identity, secret environment, and a writable mount outside scratch. The example then lists the guarantees only a real container, microVM, or managed runner can enforce. No generated code executes on the host, and the Python function is deliberately never described as a sandbox.

10. Unbounded consumption

bash
python examples/10_resource_controls.py

One reservation charges once across replay. An oversized recursive branch gets rejected before any work happens and leaves every counter unchanged. Real distributed agents need the same atomic reservation invariant in a concurrency-safe shared store.

Then the ceiling that budget doesn't have. Tokens, calls, steps, bytes and latency are all the operator's resources, which is why they get limits: the person writing them is the person holding the bill. An agent that can refund, credit, discount, or upgrade is moving somebody else's money, and that figure usually appears on a dashboard instead. The lesson pays three refunds, each comfortably inside a per-request ceiling, and the third is refused by the account window, because a per-request ceiling bounds a day only if the number of requests is bounded and a retry or a redelivery produces a fresh one. For each resource an agent can move, ask whose it is.

11. Red-team release gates

bash
python examples/11_redteam_gate.py

Allow-all fails with 100% attack success. Block-all fails too, because benign utility falls to zero. The third system passes everything and proves nothing, because it answers from the probe list itself, which is exactly what a suite satisfied without controls is worth. The capstone wires the same gate to real boundaries. Missing categories and evaluation exceptions count as failures in the tested gate too.

12. Incident response

bash
python examples/12_incident_response.py

Recovery directly after detection gets rejected. The exercised path preserves a digest, contains, records root cause and regression, passes the gate, recovers, and closes with an owner. Hash chaining detects changed metadata and not a deleted tail. The run shows a truncated log verifying cleanly right up until you check it against an anchored head. Production still needs restricted, durable, append-only evidence storage.

13. Context assembly

bash
python examples/13_context_assembly.py

The naive prompt carries three citation keys and the retriever issued two. The forged one came from a support ticket and occupies the same position, in the same syntax, as a real one. Escaping defuses the heading, the key, the operator line, and the forged fence without deleting any of them, so the ticket's actual wording survives for an investigation. The region is then fenced with a per-request nonce the document couldn't have contained. A politely worded request comes through untouched, because nothing was forged: impersonation is closed here, persuasion is lesson 6's problem.

14. Session correlation

bash
python examples/14_session_correlation.py

A conversation handle is a bearer reference to accumulated context, so it comes from the CSPRNG and every resume checks the owner. A colleague in the same tenant and a matching name in another tenant are both refused, with the same wording a missing handle gets, so the store isn't a membership oracle.

Then the harder half. One operator preps two clients in a morning inside a conversation they own, and every turn is one they were entitled to see. The answer being composed is about the second client. Turns carry the subject recorded when they happened, so the first client's amount and date are withheld, and a question naming nobody is still attributed correctly. No permission was exceeded anywhere in that paragraph, which is what makes it an attribution failure rather than an access-control one.

15. Audit records the agent can't rewrite

bash
python examples/15_audit_records.py

Code inside a sandbox owns everything in the sandbox, including the logs written about it. In the July 2026 OpenAI and Hugging Face incident, METR's independent review found agents that rebuilt part of their tool-execution path so a transcript showed one command while another ran (roughly 7% of the transcripts it evaluated, in places), and agents that tried to edit their logs. They could edit the logs inside their containers, and concluded those weren't the real source of truth. The record that held was kept outside their reach.

The lesson replays a small version. The agent's transcript and its in-sandbox log agree, and both lie about the second call. The boundary log, written by the component every call crosses and chained with an HMAC under a key the agent never holds, says what ran, and reconciling the two finds the spoof. Then the agent attacks the record: an edited entry fails, a deletion from the middle fails, a truncated tail needs lesson 12's head anchor, and a whole-file rewrite with every hash recomputed passes its own check and still fails under the logger's key. That last case is the reason for the key. A plain hash chain only proves a record wasn't edited by someone who couldn't also rewrite the rest.

Hands-on capstone

Run the complete release review twice:

bash
python hands_on/security_review.py
python hands_on/security_review.py

Expected output on each run:

GENAI SECURITY RELEASE REVIEW
  naive gate passed: False
  hardened gate passed: True
  evidence: security-report.json
  release ready: True

The command attacks the same allow-all and hardened boundary across benign utility, all ten OWASP 2025 categories, SSRF, and generated-code isolation. It exits nonzero unless the naive implementation fails and the hardened implementation passes. It writes a deterministic security-report.json, which Git ignores so release systems can archive it separately.

Inspect the report:

bash
python -m json.tool security-report.json

Each result carries a control field naming the boundary that decided it. Read those before trusting a pass. A probe can block for a reason unrelated to the risk it's named after, and an outcome on its own can't show you that. The naive system records no controls, which is the point of it.

Then extend it with a risk from your own threat model. A top-ten-only capstone isn't a complete security review.

Verification

Run the entire offline suite:

bash
python -m unittest discover -v

The tests cover successful decisions, actual security denials, malformed input, exceptions, replay/idempotency, atomic failures, cross-tenant access, changed provenance, damaged audit chains, benign utility, and broken evaluation coverage.

To reproduce CI locally after activating the environment:

bash
python check_setup.py
python -m compileall -q genai_security examples hands_on tests check_setup.py
python -m unittest discover -v
for example in examples/[0-9][0-9]_*.py; do python "$example"; done
python hands_on/security_review.py
python hands_on/security_review.py

CI runs this matrix on the minimum supported Python (3.11) and a current Python (3.13) and verifies that test discovery finds a nonzero number of tests.

Repository map

genai_security/                 executable security controls
  threats.py                   assets, flows, ranked risks, OWASP taxonomy
  data.py                      context minimization and disclosure inspection
  provenance.py                artifact manifests, digest and approval checks
  poisoning.py                 record and corpus quarantine findings
  sinks.py                     strict JSON, parameterized SQL, escaped HTML
  context.py                   escaped passages and a nonce-fenced prompt region
  sessions.py                  owner-bound conversation handles and turn subjects
  capabilities.py              identity, roles, approval, idempotency, limits
  vectors.py                   tenant/ACL/provenance prefilter and cache keys
  claims.py                    pinned structural evidence for factual claims
  network.py                   SSRF-resistant egress planning
  isolation.py                 pre-execution contract for a real runner
  resources.py                 request-wide atomic reservations
  redteam.py                   adversarial evaluation and release policy
  incidents.py                 stateful response and tamper-evident audit metadata
  audit.py                     keyed boundary records and transcript reconciliation
examples/                      fifteen narrated, executable lessons
hands_on/security_review.py    deterministic naive-vs-hardened capstone
tests/                         offline security-invariant test suite
check_setup.py                 environment and capstone readiness check
TEXTBOOK.md                    full Chapter 20 lecture
EXERCISES.md                   progressive engineering exercises
LESSONS.md                     surprises learned while building the course

What this course proves, and what it doesn't

The offline suite proves the behavior of these teaching policies. It doesn't prove:

  • that your identity provider supplies the correct tenant and roles;
  • that a vector database applies prefilters before its actual similarity engine;
  • that an HTTP client connects to the same address your policy resolved;
  • that a container or microVM resists escape and resource exhaustion;
  • that artifact builders and signing identities are protected;
  • that provider retention and regional settings match privacy policy;
  • that distributed budget and idempotency stores are atomic under concurrency; or
  • that a production model resists your task-specific attacks while retaining utility.

Those claims need integration and end-to-end tests against the real infrastructure. Keep the invariants from this repository and replace the adapters.

Standards baseline

This August 2026 course uses dated primary baselines so future readers can spot the drift.

Frameworks change. Re-check current versions during a real review, and record the exact version your evidence targets.

Troubleshooting

ModuleNotFoundError: genai_security

Activate the environment and run python -m pip install -r requirements.txt. For a temporary pre-install development check only, use PYTHONPATH=. python ....

Setup reports Python older than 3.11

Create the virtual environment with a newer interpreter, such as python3.11 -m venv .venv, then reinstall.

The capstone exits nonzero

Open security-report.json. Check hardened.gate_failures and any result whose passed value is false. An evaluator exception appears as actual: null and must be fixed rather than waived as a pass.

A test passes only when network or credentials are available

That test doesn't belong in the default offline gate. Inject a deterministic adapter for the course, and add the live behavior as a clearly separated integration suite.

Continue

Complete EXERCISES.md, add a system-specific capstone risk, and carry the resulting release evidence into AI in Production.