Core path - 3 of 8
Prompt Engineering: A Guided Deep Dive
A hands-on playground for learning prompt engineering from the basics zero/few-shot, chain-of-thought, roles, structured output, through to optimized prompts for real use cases, and finally to the discipline that separates a guess from a result: measuring that a prompt change actually helped. Every concept is a small, runnable Python script you read, run, and tweak. No framework magic just enough code to see how each idea changes what the model does.
This is the third of eight core repos in the series. The first two teach the API call (OpenAI, Claude); this one teaches you to get more out of that same call by asking better.
Like its siblings, it's meant to be walked through, not just read. Each script prints a before/after so you can see the effect, and EXERCISES.md has a predict-then-run prompt for each lesson.
0. The one big idea
The model is fixed; the prompt is the program. You don't touch the weights you change what you ask and how, and that is most of the quality you'll ever get.
Everything below is a variation on that. Zero/few-shot is how many examples you show; chain-of-thought is giving room to think; roles and system prompts are who the model is and what the rules are; structured output is the exact shape of the answer. None of it changes the model; it changes the request. And the last step, the capstone, is the habit that makes it real: measure the change instead of trusting that the new prompt "reads better."
1. Setup (5 minutes)
# 1. Create an isolated Python environment
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 2. Install dependencies
pip install -r requirements.txt
# 3. Choose your provider (set PROVIDER in .env); your key loads separately
cp .env.example .env
# Your API key does NOT go in .env. Store it in your OS keychain and run
# lessons with `secrun`: 2-minute setup in ../SECRETS.md.
# 4. Confirm everything is wired up (makes no API call, costs nothing)
secrun python check_setup.py # secrun injects your key so the check can see it
Prompt engineering is provider-agnostic, so this repo is too. Pick whichever stack
you set up in the sibling repos with PROVIDER in .env:
PROVIDER |
Chat model | Key needed |
|---|---|---|
openai (default) |
OpenAI gpt-4o-mini |
OPENAI_API_KEY |
claude |
Claude claude-haiku-4-5 |
ANTHROPIC_API_KEY |
The only file that knows which provider you picked is
common/providers.py; every lesson is pure prompting. Because
the OpenAI stack uses the OpenAI SDK, it also reaches any OpenAI-compatible local
server (Ollama, LM Studio, llama.cpp, vLLM): keep PROVIDER=openai and set
OPENAI_BASE_URL to the local endpoint. So the same lessons run against hosted
OpenAI, hosted Claude, or a model on your laptop, for free.
No offline mode here. Unlike most sibling repos, every lesson makes a small, real call; the whole point is to watch a prompt change the output. The calls are cheap (a fraction of a cent each), or free against a local model.
2. The fundamentals: the core techniques
Each file is a self-contained before/after. Run them in order.
secrun python fundamentals/01_zero_shot.py
| File | Technique | One-line idea |
|---|---|---|
| 01_zero_shot.py | Zero-shot | Ask clearly, constrain the output, no examples. |
| 02_few_shot.py | Few-shot | Teach format & conventions with 2–5 examples. |
| 03_chain_of_thought.py | Chain-of-thought | Let the model reason step by step before answering. |
| 04_role_prompting.py | Role / persona | Assign expertise + audience to steer tone & depth. |
| 05_system_prompts.py | System prompts | Set durable behavior: role, rules, format, fallback. |
| 06_structured_output.py | Structured output | Force machine-readable JSON (json=True / a schema). |
| 07_delimiters_and_context.py | Delimiters & grounding | Separate instructions from data; answer only from context. |
| 08_prompt_chaining.py | Prompt chaining | Decompose into a pipeline; generate→critique→revise. |
| 09_self_consistency.py | Self-consistency | Sample N times and majority-vote for accuracy. |
| 10_parameters.py | Decoding params | temperature, top_p, max_tokens, seed, stop. |
| 11_react.py | ReAct | Interleave Thought→Action→Observation; the pattern under "agents". |
| 12_reflexion.py | Reflexion | Attempt→verify→reflect→retry against a real check, not vibes. |
| 13_meta_prompting.py | Meta-prompting | Use the model to rewrite a weak prompt into a strong one. |
| 14_reasoning_models.py | Reasoning models | Drop the "think step by step" scaffolding; give goal + constraints. |
3. The use cases: optimizing prompts for real tasks
Each shows a naïve prompt vs an optimized prompt for a real job, and explains why every change helps.
secrun python examples/03_code_review.py
| File | Use case | Techniques combined |
|---|---|---|
| 01_customer_support_reply.py | Support email responder | system prompt · policy constraints · tone · fallback |
| 02_data_extraction.py | Unstructured text → typed JSON | schema · normalization · null policy · json=True |
| 03_code_review.py | Security-aware code review | persona · rubric · severity · fixed format |
| 04_summarization.py | Audience-targeted TL;DR | audience · length limits · focus · grounding |
| 05_text_to_sql.py | Natural language → SQL | schema grounding · dialect · safety · few-shot |
| 06_classification.py | Ticket routing / classification | closed label set · 'other' escape hatch · confidence · few-shot edges |
4. The mental model (cheat sheet)
A reliable prompt usually answers these questions for the model:
- Role: who are you? (persona / expertise)
- Task: what exactly should you do?
- Context: what information do you have? (clearly delimited)
- Constraints: what are the hard rules / what must you not do?
- Format: exactly how should the output be structured?
- Examples: what does a good answer look like? (if needed)
- Fallback: what to do when you can't comply or don't know?
General heuristics:
- Be specific. Vagueness is the #1 cause of bad output.
- Show, don't just tell. Examples lock in format and conventions cheaply.
- Give room to think on hard problems; hide the reasoning if the end user doesn't need it.
- Constrain the output when code will parse it, and still parse defensively.
- Match temperature to the task:
0for extraction/classification/code, higher for creative work. - Iterate. Prompt engineering is empirical: change one thing, observe, repeat.
5. The capstone: optimize.py
Everything points here. Every lesson argues a tuned prompt beats a naive one; the capstone stops arguing and measures it: it runs both prompts over a small labeled set, scores each, and tells you which won and by how much. That's the whole discipline in one tool, and the bridge to the Evals deep dive.
# Compare the built-in naive vs tuned prompt on a sentiment task:
secrun python hands_on/optimize.py
# A different built-in task (support-ticket priority), showing the misses:
secrun python hands_on/optimize.py --task priority --show-misses
# Bring your own: two prompt files + a JSONL of {"text","expected"} rows:
secrun python hands_on/optimize.py --prompt-a naive.txt --prompt-b tuned.txt --data cases.jsonl
Read hands_on/optimize.py: evaluate() is the whole loop
(run a prompt over the cases, score each, average) and compare() prints the
verdict. Suggested exercise: add two hard cases (a sarcastic review, a
backhanded compliment) and rerun, then watch which prompt cracks. The first time it
tells you your "better" prompt was actually worse, prompt engineering has clicked.
Notes & costs
- Running a script makes real API calls. Against a hosted API this costs a fraction of a cent each; against a local model it's free.
09_self_consistency.py,10_parameters.py, and the capstone make several calls each by design (they sample or score multiple times).- Never commit your
.env; it's already in.gitignore.
Where to go next
You've learned to shape a single call. The series builds outward from here:
- Ground it in your data: when the model needs facts it doesn't have, retrieve the right text and put it in the context. → RAG
- Measure it at scale: the capstone is a tiny eval; the real discipline (judges, metrics, significance, CI gates) is its own dive. → Evals
- Let it act: ReAct (lesson 11) by hand is the seed of an agent loop with real tools. → Agents
- Harden it: delimiters (lesson 07) are the first, weakest injection defense; the real defense-in-depth is its own dive. → Prompt Injection & Guardrails
- Reasoning models: lesson 14 is the start; prompting o-series / extended thinking well is a growing skill.
From teaching code to production
Every lesson here optimizes one prompt in isolation. Production is about operating prompts like the code they are:
| This repo's teaching shortcut | In production |
|---|---|
| The prompt is a string literal in the script | A versioned prompt behind config, promoted only past an eval gate |
| You eyeball the before/after | The capstone's compare, run as a CI gate that blocks a quality regression |
chat() is called bare |
The call wrapped in retries + backoff, a budget, and a response cache |
| You trust the model's output shape | Schema validation (structured()) + guardrails on every request |
| One prompt, one model, by hand | A prompt registry with staged rollouts, A/B tested on live traffic |
These shortcuts are right for learning and wrong for production. All of those concerns: observability, cost, reliability, caching, guardrails, prompt versioning, and eval gates) are built from scratch and wired into one running app in Production (#8 in the series). It runs offline on a mock provider, so you can see the whole ops machinery with no key and no cost.
File map
check_setup.py ← run first: verifies Python, packages, provider, key
README.md ← this guide
EXERCISES.md ← predict-then-run prompts, one per lesson
common/ ← shared plumbing (read providers.py!)
providers.py ← the ONLY provider-specific file: chat / chat_stream / structured
display.py ← tiny terminal helpers (header, rule)
fundamentals/ ← the core techniques (run in order)
01_zero_shot.py ... 14_reasoning_models.py
examples/ ← naïve-vs-optimized prompts for 6 real use cases
01_customer_support_reply.py ... 06_classification.py
hands_on/
optimize.py ← capstone: A/B-compare two prompts on a labeled set
Troubleshooting
Run secrun python check_setup.py first; it catches most problems. Then, by symptom:
| What you see | What it means / the fix |
|---|---|
PROVIDER=... needs ... in the environment |
Set PROVIDER in .env, then load the key from your keychain by running under secrun. See SECRETS.md. |
ModuleNotFoundError (openai / anthropic / rich) |
Dependencies aren't installed or the venv isn't active. source .venv/bin/activate then pip install -r requirements.txt. |
AuthenticationError / 401 |
The key is present but wrong; check it matches the PROVIDER you set. |
| A JSON lesson prints prose, not JSON | A weaker (often local) model ignored the format. json=True / structured() help; the lessons also parse defensively. |
| Running against a local model and it's flaky on JSON/ReAct | Small models follow strict formats less reliably, so try a more capable one (qwen2.5, llama3.1) or the hosted stack. |
SyntaxError / odd type errors on startup |
You're likely on Python 3.9 or older; this repo needs 3.10+. check_setup.py confirms your version. |
Still stuck? Every file is small and self-contained. Open it, read the docstring at the top, and run it directly.
The series
This is one of sixteen standalone, hands-on deep dives into building with LLM APIs: eight core, plus eight bonus dives. Each one stands on its own, with its own setup, examples, and capstone, and they all share the same house style: provider-agnostic, built from scratch (no frameworks), offline-first examples, and a real capstone. Do them in any order; this sequence builds naturally:
- OpenAI API: the API from zero
- Claude API: the same ideas, the Anthropic way
- Prompt Engineering: shape model behavior with better prompts (zero/few-shot, chain-of-thought, roles)
- RAG: answer questions over your own documents
- Evals: measure whether a change actually helps
- Agents: give a model tools and a loop so it can act
- Prompt Injection & Guardrails: attack and defend all of the above
- Production: operate one app end to end: observability, cost, reliability, caching, guardrails, prompt versioning, eval gates
Bonus dives, standalone and slotting in where they're most useful:
- Context Engineering: manage what's in the window: memory, compaction, assembly
- Multimodal: images & audio, not just text
- Fine-tuning: teach a model new behavior by example
- MCP: serve tools, data & prompts to any LLM over a standard protocol
- Local Models: run open-weight models on your own machine
- Agent Harnesses: build on the loop: hooks, permissions, sandboxing, subagents
- Realtime Voice: low-latency speech-to-speech agents
- Observability: watch a running app over time: drift, quality, alerting, the flywheel
You are here: #3, Prompt Engineering.