av / dives
source about me

AI Engineering: Deep Dives

A series of hands-on, build-it-from-scratch courses for learning to build with LLMs. Each one is a standalone repository you walk through. Every concept is a small runnable Python script, every section ends with something to run, and in most repos the first runnable thing needs no API key. No frameworks, no magic, just enough code to see how each piece works.

They share one house style. Provider-agnostic where it makes sense (OpenAI or Claude, often a local model too), offline-first examples, a real capstone, and an EXERCISES.md full of predict-then-run prompts. Assumes only basic Python.

The core pathdo these in order
01

OpenAI API

You send a list of messages. You get back a message. Everything else is detail on that request.

02

Claude API

The same idea, the Anthropic way: content blocks, tool use, and extended thinking.

03

Prompt Engineering

Shape what the model does with how you ask: zero/few-shot, chain-of-thought, roles, structure.

04

RAG

A model can only answer from what is in its context window. RAG is the discipline of putting the right text there.

05

Evals

If you cannot measure it, you cannot improve it: turn quality into a number you can rerun.

06

Agents

An agent is a loop: the model picks a tool, you run it, you feed the result back, until it is done.

07

Prompt Injection & Guardrails

Treat everything the model reads and writes as untrusted: contain the blast radius.

08

Production

The model call is one line. Production is the dozen lines around it that make it safe, cheap, observable, and reliable.

Bonus divesstandalone, each notes where it slots in
BONUS

Agent Harnesses

Once you have hand-written the loop, most agent work is building on a harness: hooks, permission policies, sandboxing, subagents, and headless runs.

Slots in after Agents (6)
BONUS

Context Engineering

The model only knows what is in its context window, so manage it: conversation memory, compaction, long-term recall, and what to drop when it will not all fit.

Slots in after Agents (6), pairs with RAG (4)
BONUS

Multimodal

A multimodal model takes more than text: images and audio. Put the right modality in the right slot, and mind the token cost.

Slots in after the API dives (1-2), pairs with RAG (4)
BONUS

Realtime Voice

Conversational voice is a low-latency, full-duplex loop: stream audio both ways, handle interruption, and choose a pipeline vs a speech-to-speech model.

Slots in after Multimodal, the API dives (1-2)
BONUS

Fine-tuning

Fine-tuning changes how a model behaves, not what it knows: teach behavior by example, then prove it beat your baseline.

Slots in after RAG (4) + Evals (5)
BONUS

MCP

The Model Context Protocol: hand an LLM tools, data, and prompts from a separate process. Write the server once, any client can use it.

Slots in after Agents (6)
BONUS

Local Models

An open-weight model on your machine speaks the same OpenAI API, so local is mostly an ops choice: privacy, cost, control.

Slots in after the API dives (1-2), pairs with Fine-tuning
BONUS

Observability

A prototype is judged once; a production system is judged continuously. Watch quality as a trend: drift, silent regressions, and alerting that does not cry wolf.

Slots in after Production (8), pairs with Evals (5)
BONUS

Architecture

Every other dive teaches a component; this one teaches where the boundaries between them go. Where conversation state lives, what a queue buys, what streaming costs your guardrails, and where the tenant boundary goes, each decision measured rather than asserted.

Slots in after Production (8), pairs with Observability
BONUS

Professional Tools

Volume 2: rebuild each from-scratch primitive with the tool professionals actually reach for, and measure both on the same eval.

Slots in after everything (you need the primitives first)
BONUS

AI Data Engineering

A retrieval index is a disposable, derived view of source truth: version documents, carry permissions into every chunk, make deletes stick, and prove the corpus can be rebuilt.

Slots in after RAG (4), before Production (8)
BONUS

GenAI Security

Treat the model as an untrusted principal, not a security boundary: authorize every effect in code, verify the supply chain, isolate data and execution, bound what one request can spend, and make attacks block the release.

Slots in after Prompt Injection (7), before Production (8)
BONUS

Inference Platform Engineering

A self-hosted model becomes a service only when memory and queue scheduling turn finite GPUs into measured latency, throughput, reliability, and cost.

Slots in after Local Models (15); Production (8); Architecture (21)
BONUS

Testing & Delivery

A release is an evidence pipeline. Requirements defined independently decide whether the reproducibility, compatibility, security, rollout, and recovery evidence is good enough to promote.

Slots in after Evals (5) + Production (8); pairs with GenAI Security (20)
BONUS

ML Foundations

A model is a chain of numeric contracts. Trace shapes, logits, loss, gradients, masked attention, sampling, calibration, quantization, and retained memory through runnable NumPy and PyTorch code.

Slots in after the API dives (1, 2); before Fine-tuning (13), Local Models (15), and Inference Platforms (22)
BONUS

Structured Data + AI

A generated query is a hypothesis, and only the database settles it. Score by executing, expect the join that multiplies a total, recognize an undefined metric as a specification gap, and put the boundary in a role rather than a prompt.

Slots in after Evals (5); pairs with GenAI Security (20)
Capstoneeverything at once
END

Capstone: askrepo

One codebase Q&A tool built across eight eval-gated stages, from a first retrieval pass to a hardened, observable app.

Companionoutside the sequence, for readers whose AI work ships in TypeScript
TS

AI in TypeScript

The same ideas in TypeScript. Your types stop at the network boundary, so everything the model says is unknown until you check it at runtime.

Referencethe shared docs