A series of hands-on, build-it-from-scratch courses for learning to build with LLMs. Each one is a standalone repository you walk through. Every concept is a small runnable Python script, every section ends with something to run, and in most repos the first runnable thing needs no API key. No frameworks, no magic, just enough code to see how each piece works.
They share one house style. Provider-agnostic where it makes sense (OpenAI or Claude,
often a local model too), offline-first examples, a real capstone, and an
EXERCISES.md full of predict-then-run prompts. Assumes only basic Python.
You send a list of messages. You get back a message. Everything else is detail on that request.
02The same idea, the Anthropic way: content blocks, tool use, and extended thinking.
03Shape what the model does with how you ask: zero/few-shot, chain-of-thought, roles, structure.
04A model can only answer from what is in its context window. RAG is the discipline of putting the right text there.
05If you cannot measure it, you cannot improve it: turn quality into a number you can rerun.
06An agent is a loop: the model picks a tool, you run it, you feed the result back, until it is done.
07Treat everything the model reads and writes as untrusted: contain the blast radius.
08The model call is one line. Production is the dozen lines around it that make it safe, cheap, observable, and reliable.
Once you have hand-written the loop, most agent work is building on a harness: hooks, permission policies, sandboxing, subagents, and headless runs.
Slots in after Agents (6) BONUSThe model only knows what is in its context window, so manage it: conversation memory, compaction, long-term recall, and what to drop when it will not all fit.
Slots in after Agents (6), pairs with RAG (4) BONUSA multimodal model takes more than text: images and audio. Put the right modality in the right slot, and mind the token cost.
Slots in after the API dives (1-2), pairs with RAG (4) BONUSConversational voice is a low-latency, full-duplex loop: stream audio both ways, handle interruption, and choose a pipeline vs a speech-to-speech model.
Slots in after Multimodal, the API dives (1-2) BONUSFine-tuning changes how a model behaves, not what it knows: teach behavior by example, then prove it beat your baseline.
Slots in after RAG (4) + Evals (5) BONUSThe Model Context Protocol: hand an LLM tools, data, and prompts from a separate process. Write the server once, any client can use it.
Slots in after Agents (6) BONUSAn open-weight model on your machine speaks the same OpenAI API, so local is mostly an ops choice: privacy, cost, control.
Slots in after the API dives (1-2), pairs with Fine-tuning BONUSA prototype is judged once; a production system is judged continuously. Watch quality as a trend: drift, silent regressions, and alerting that does not cry wolf.
Slots in after Production (8), pairs with Evals (5) BONUSEvery other dive teaches a component; this one teaches where the boundaries between them go. Where conversation state lives, what a queue buys, what streaming costs your guardrails, and where the tenant boundary goes, each decision measured rather than asserted.
Slots in after Production (8), pairs with Observability BONUSVolume 2: rebuild each from-scratch primitive with the tool professionals actually reach for, and measure both on the same eval.
Slots in after everything (you need the primitives first) BONUSA retrieval index is a disposable, derived view of source truth: version documents, carry permissions into every chunk, make deletes stick, and prove the corpus can be rebuilt.
Slots in after RAG (4), before Production (8) BONUSTreat the model as an untrusted principal, not a security boundary: authorize every effect in code, verify the supply chain, isolate data and execution, bound what one request can spend, and make attacks block the release.
Slots in after Prompt Injection (7), before Production (8) BONUSA self-hosted model becomes a service only when memory and queue scheduling turn finite GPUs into measured latency, throughput, reliability, and cost.
Slots in after Local Models (15); Production (8); Architecture (21) BONUSA release is an evidence pipeline. Requirements defined independently decide whether the reproducibility, compatibility, security, rollout, and recovery evidence is good enough to promote.
Slots in after Evals (5) + Production (8); pairs with GenAI Security (20) BONUSA model is a chain of numeric contracts. Trace shapes, logits, loss, gradients, masked attention, sampling, calibration, quantization, and retained memory through runnable NumPy and PyTorch code.
Slots in after the API dives (1, 2); before Fine-tuning (13), Local Models (15), and Inference Platforms (22) BONUSA generated query is a hypothesis, and only the database settles it. Score by executing, expect the join that multiplies a total, recognize an undefined metric as a specification gap, and put the boundary in a role rather than a prompt.
Slots in after Evals (5); pairs with GenAI Security (20)