Bonus dive
Fine-tuning: A Guided Deep Dive
A hands-on playground for learning fine-tuning, meaning teaching a model new behavior from examples, from the ground up. You'll build a training set from scratch, validate it, run a fine-tune job, and then do the step most people skip: prove the fine-tuned model actually beat the base model on data it never saw. No framework magic, just enough code to see how each step works.
Here's what makes this repo work. It runs completely offline on a mock provider, with no
API key. Real fine-tuning costs money and takes minutes to hours, which is a terrible way
to learn the shape of it. So the default PROVIDER=mock ships a tiny deterministic
"model" and a simulated fine-tune lifecycle, upload then job then poll then use, that runs
in-process in under a second for $0.
The real OpenAI path is closing. OpenAI is winding down self-serve fine-tuning. Orgs that had never fine-tuned lost the ability to start new jobs on 2026-05-07, orgs with no recent fine-tuned inference on 2026-07-02, and the remainder go on 2027-01-06. Inference on already-tuned models keeps working until the base model retires. If you're new to this,
--realwill returntraining_not_availablewhatever base model you name.That doesn't make the repo obsolete, and nothing here was removed. The mechanics transfer: build a dataset, validate it, run the lifecycle, prove the tuned model beat the baseline. The method is identical whether the trainer is OpenAI, a cloud GPU, or MLX on your laptop. What changed is where you run it. Section 9 on distillation and Section 10 on open-weight fine-tuning with LoRA and PEFT are where that now happens.
This repo is standalone and teaches everything it needs on its own. It's the hands-on version of the RAG deep dive's "RAG, fine-tuning, or something else?" section, and Section 7 borrows the win-rate method from the Evals deep dive. Its code depends on neither.
Like its siblings, walk through it. Each section ends with something to run, and every section runs offline and free on the mock. EXERCISES.md has a predict-then-run prompt for each one.
0. The one big idea
Fine-tuning teaches behavior well and facts badly. You teach a default behavior with examples, and then you have to prove it beat your baseline.
That's the whole repo. RAG and long context change what's in the context window, which is knowledge. Fine-tuning changes how the model responds by default, meaning format, tone, or one narrow skill, taught only by showing it input and output examples. So the training set is the product, and most of the work is building and validating it. And because "it feels better" is worth nothing, the step that makes fine-tuning real is the last one. Measure the tuned model against the base model on a held-out set, and ship only if it wins. Hold onto that and none of this feels complicated.
The honest version of "facts badly" is worth a sentence, because the slogan you'll hear elsewhere is "fine-tuning can't teach knowledge" and that isn't true. Training does write into the weights, and enough of it does install knowledge; continued pretraining on a large domain corpus is how domain-specific base models get made. What's true is that the few thousand examples of a hosted fine-tune are far too little repetition to land facts reliably, and that even when a fact does land you can't cite it, update it, or tell whether it's still there. So the rule survives its own correction: use retrieval for knowledge, because facts in weights are unreliable, unattributable, and stale the moment the world moves, not because training is incapable of storing them.
1. Setup (5 minutes)
# 1. Create an isolated Python environment
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 2. Install dependencies (the default mock stack needs only python-dotenv)
pip install -r requirements.txt
# 3. Copy the env file: the default runs keyless (no API key needed)
cp .env.example .env
# (Real provider instead of the mock? Its key goes in your OS keychain,
# not .env: see ../docs/SECRETS.md, then run scripts as `secrun python ...`.)
# 4. Confirm everything is wired up (makes no API call, costs nothing)
python check_setup.py
The default stack is offline and free, and the whole learning arc runs on a mock provider with no key. Pick a real provider only when you want to run an actual, paid fine-tune.
PROVIDER |
What runs | Key needed |
|---|---|---|
mock (default) |
Deterministic in-process model + a simulated fine-tune lifecycle. The entire repo, free. | none |
openai |
Real chat. The fine-tuning API is being retired (see the note above); --real now fails with training_not_available for most accounts. See Sections 9-10 for what replaces it. Still the default chat/teacher stack. |
OPENAI_API_KEY |
claude |
Chat only. Anthropic fine-tuning is a limited/enterprise program, not self-serve. Usable as a base/teacher model. | ANTHROPIC_API_KEY |
You can complete every section for $0. The mock simulates uploading, training, polling, and serving a fine-tuned model. That's now the only way most readers can run the lifecycle end to end, which makes it the main path rather than a convenience.
2. When to fine-tune, against prompt, few-shot, and RAG
python examples/01_when_to_finetune.py
The most valuable fine-tuning skill is knowing when not to. Fine-tuning is the slow, expensive, provider-specific option, and reaching for it first is the most common and costly mistake. One rule resolves most cases.
- RAG and long context change what's in the context → reach for them when you need facts that change or must be cited.
- Fine-tuning changes how the model behaves by default → reach for it when you need the same format, tone, or skill every time, or lower cost and latency on a fixed, high-volume task.
- Tools and agents change what the model can do → reach for them when it has to act or fetch live data.
The example walks a handful of real scenarios and lands each one on the right tool. The headline rule of thumb, don't fine-tune first, is here for a reason. A better prompt or a few examples solves most of what people think needs training.
3. The dataset is the product
python examples/02_dataset_format.py
A fine-tune learns the behavior its examples demonstrate, so the examples are the product. Hosted fine-tuning uses JSON Lines: one conversation per line, each ending in the assistant turn you want the model to learn to produce.
{"messages": [
{"role": "system", "content": "You are Acme's support triage assistant. Reply EXACTLY as 'category: <...> | reply: <...>'."},
{"role": "user", "content": "i forgot my password"},
{"role": "assistant", "content": "category: account | reply: Use Settings > Security > Reset password."}
]}
The example builds one from scratch, loads the hand-made
datasets/support_train.jsonl, and splits it into train
and validation. The running example throughout the repo is a support-triage assistant
taught to answer in one rigid house format, which is exactly the kind of narrow, repeated
behavior fine-tuning is good at.
4. Validate your data before you pay to train on it
python examples/03_validate_data.py
A fine-tune job is slow, and on a real provider it costs money. The cheapest way to waste neither is to check the dataset first. This runs every offline check on the hand-made set, covering schema, duplicates, class balance, and a token and cost estimate, then runs them again on a deliberately broken set so you can watch the checks catch real problems. Garbage in, garbage model. Most bad fine-tunes are bad datasets that nobody validated.
5. Run a fine-tune job, from upload to create to poll to use
python examples/04_run_finetune.py
Every hosted fine-tune follows the same lifecycle:
- upload the training file (
files.create) - create a job from it (
fine_tuning.jobs.create), this is what trains - poll the job until it's done (
fine_tuning.jobs.retrieve) - use the new model id
The example runs this two ways with nearly identical code, which is the point.
- The default,
PROVIDER=mock, simulates the whole lifecycle in-process, deterministically, in under a second, for $0. - The opt-in real run, with
PROVIDER=openaiand--real, uploads your file and starts a genuine job. It prints a cost warning and asks for confirmation first, and it can take a while.
6. Use the fine-tuned model
python examples/05_use_model.py
Once a job succeeds, the provider hosts your model under a new id, something like
ft:gpt-4o-mini-2024-07-18:.... Using it is a normal chat call with that id and no
special API. The example asks the base and the fine-tuned model the same questions side
by side. The base handles the one or two categories it happens to know, then rambles and
ignores the house format on the rest. The tuned model snaps every one into the trained
category: ... | reply: ... shape. That behavior change, taught only by examples, is the
whole idea.
7. Did it actually help?
python examples/06_did_it_help.py
This is where the repo pays off. A fine-tune you think is better is worth nothing. The
only thing worth shipping is one you can prove beat your baseline on data the training
never saw. This points the
Evals deep dive's method at one
decision, base against fine-tuned on the held-out
datasets/support_eval.jsonl, with two numbers.
- Accuracy. The percentage of held-out examples where the predicted category matches gold.
- Win-rate. Pairwise, the way
evals/07_pairwise.pydoes it. Show a judge both answers and tally which is better. Here the judge is an offline format rubric so it runs free; in production you'd use an LLM-as-judge.
The baseline is the exact model the tune started from, read out of the fine-tuned id
(ft:<base>:...), not whatever the series' chat default happens to be. A
gpt-4o-mini fine-tune measured against gpt-6-luna would mix a model change into the
tuning effect, and you couldn't tell which one moved the number.
If the tuned model doesn't beat the baseline, the honest move is to not ship it and go back to the dataset.
8. Hyperparameters and reading the loss curve
python examples/07_hyperparameters.py
You don't need many knobs and the defaults are usually fine, but three of them matter.
- n_epochs is how many times the model sees the whole dataset. Too few and it hasn't learned. Too many and it overfits, which you spot when validation loss stops dropping, or rises, while training loss keeps falling.
- learning_rate_multiplier is how big each step is. Higher is faster and can overshoot. Lower is steadier and slower.
- batch_size is examples per update, mostly a speed and stability tradeoff. Leave it on auto unless you have a reason.
The example renders the simulated loss curve and shows what overfitting looks like, so the abstract knobs turn into a picture.
9. Distillation: train a small model on a strong model's outputs
python examples/08_distillation.py
The most common production shape of fine-tuning isn't hand-labeling. It's distillation.
Take a big, expensive, smart model, the teacher, that already does your task well, run it
over a pile of inputs, and use its answers as training data for a small, cheap, fast
model, the student. The labels write themselves, which is what makes a set of hundreds or
thousands of examples cheap to build. The example builds a distillation dataset. On the
mock, the teacher is the mock in the house format; on a real provider you'd point it
at gpt-4o or claude. Then it validates the result, proving it's a normal training
file you can feed straight into Section 5.
10. Beyond hosted: open-weight fine-tuning with LoRA and PEFT
python examples/09_open_weights_lora.py
Everything so far was hosted fine-tuning: hand a provider a JSONL file, and they train
and host the result. The other world is open-weight fine-tuning. Download a model whose
weights are public (Llama, Mistral, Qwen, Gemma) and train it yourself on your own GPU, or
a rented one. This section is conceptual, because running open weights needs a GPU and a
different stack (PyTorch plus Hugging Face transformers and peft) and deserves its own
deep dive. It explains the two ideas you need, full fine-tuning against LoRA and PEFT, and
shows that your dataset is the same asset either way. The file you built in Sections 3 and
4 is exactly what an open-weight trainer consumes. To actually run open weights locally,
see the
Local Models deep dive.
11. Preference tuning, learning from comparisons (DPO and RLHF)
python examples/10_preference_tuning.py
Every example so far taught by demonstration: show the one right answer and imitate, which
is SFT. But some goals have no single right answer. "Be warmer", "be more concise", "refuse
this more firmly". You can't write the correct reply, and you can say which of two replies
is better. Preference tuning learns from exactly that, pairs of
{prompt, chosen, rejected}. RLHF trains a reward model from the rankings. DPO, the modern
shortcut, trains directly on the pairs. This section is conceptual. It shows the data
shape, where the pairs come from (the Production dive's thumbs up/down loop is the cheapest
source), and how it gets run (Hugging Face trl's DPOTrainer, usually as a LoRA). The
discipline is unchanged. Still gate on a held-out eval before shipping.
12. Reinforcement fine-tuning (RFT): learning from a grader
SFT learns from demonstrations, meaning one right answer to imitate. Preference tuning learns from comparisons, where A is better than B. Reinforcement fine-tuning learns from a grader. The model generates an answer, a scoring function rates it, and training pushes the model toward higher-scoring answers. There's no labeled target and no pair, just a way to score an attempt. This section is conceptual with no runnable example, because graders are expensive and fiddly and this is the most complex rung here.
The whole game is the grader, a function score(prompt, answer) -> number. It can be a
hard programmatic check (do the unit tests pass? does the JSON validate against the
schema? does the math answer match?) or a model-as-judge scoring against a rubric. This is
the "reinforcement learning from verifiable rewards" that trains modern reasoning models.
When correctness is checkable but you can't write down the one right output, a grader
beats labeled data.
Reach for RFT when success is easy to verify and hard to demonstrate, meaning there are many correct programs, proofs, or plans, so you can check one but not enumerate them. Or when writing thousands of gold answers or preference pairs costs more than writing one scoring function. Or when you're optimizing a multi-step behavior where only the outcome is gradeable. Stick with SFT when you can cheaply demonstrate the target, and with preference tuning when quality is a matter of taste a judge can rank but not score objectively.
There's a catch beyond cost. A model optimizing a score will hack a weak grader, passing the letter of the check while missing the point. That's the eval-gaming failure the Evals dive warns about, now inside the training loop. So the grader needs the same scrutiny as an eval, and the shipping discipline is unchanged. Gate on a held-out set the grader never saw.
The capstone: finetune_run.py
Everything assembled into one command that does the real workflow:
validate → tune → eval-gate vs. baseline → ship ONLY if it wins
# Offline, free, the full arc on the mock:
python hands_on/finetune_run.py
# Point at a different training file (e.g. the distilled set from Section 9):
python hands_on/finetune_run.py --train datasets/support_distilled.jsonl
# Require the tuned model to clear a minimum win-rate to "ship":
python hands_on/finetune_run.py --min-winrate 0.6
# The real, PAID path (opt-in, confirmed):
PROVIDER=openai secrun python hands_on/finetune_run.py --real
The gate is the discipline the whole repo is about. A fine-tune ships only when it
provably beats the base model on a held-out set. If it doesn't, the gate says so and exits
non-zero, the same shape as a CI eval gate. Read
hands_on/finetune_run.py. It's the library, validate plus
mock_tuner plus evaluate, wired to a CLI.
Should I fine-tune at all? (the decision)
Fine-tuning is rarely the first thing to reach for. Match the problem to the tool before you train anything.
| Your problem | Reach for | Why |
|---|---|---|
| The model needs facts that change, or must cite sources | RAG / long context | You're changing what it knows, not how it behaves |
| You need a consistent format, tone, or narrow skill, every time | Fine-tuning | You're teaching behavior, and that's what training adjusts |
| A few examples in the prompt already get it right | Few-shot prompting | Cheaper, instant, no training loop, so try this first |
| It must act or fetch live data | Tools / agents | Capability, not behavior or knowledge |
| Lower latency/cost on a fixed, high-volume task | Distill + fine-tune a smaller model | Push known-good behavior into a cheaper model |
Two rules of thumb. Don't fine-tune first. It's the slow, expensive, provider-specific option, and a better prompt or RAG solves most of what looks like a training problem. And never fine-tune on vibes. The only way to know it helped is to measure it against a baseline, as Section 7 does. They also complement each other rather than competing. A common production shape is fine-tune for behavior plus RAG for knowledge, in the same app.
Where to go next
You've taught a model a behavior and proved it stuck. What comes next is more of the same idea, with more control.
- Preference tuning (DPO and RLHF). Covered conceptually in §11. Train on comparisons, where A is better than B, as well as on demonstrations, to shape subtler behavior.
- Reinforcement fine-tuning (RFT). Covered conceptually in §12. Train against a grader, either a verifiable check or a rubric judge, when success is checkable but not easily demonstrated. That's how reasoning models are trained.
- Open-weight LoRA in practice. Actually run Section 10 on a GPU with
transformers,peft, andtrl. Pairs with the Local Models deep dive. - Bigger, cleaner datasets. The real lever is almost always more and better data rather than more epochs. Active learning means mining the cases your model gets wrong.
- Continuous fine-tuning. Re-tune on production traffic as it drifts, gated by evals each time.
- Function-calling and structured-output fine-tunes. Teach a small model to emit reliable tool calls or JSON for a fixed schema.
From teaching code to production
The teaching shortcuts that make this repo free and fast are exactly what you'd replace once a fine-tuned model sits on a live request path.
| This repo's teaching shortcut | In production |
|---|---|
| Simulated tune on the mock provider | A real job on real hardware, tracked by id, with the trained model pinned in config |
| Tiny hand-made dataset (dozens of rows) | Hundreds–thousands of cleaned, deduplicated, deliberately-balanced examples |
| Offline format rubric stands in for a judge | A real LLM-as-judge (debiased, both orderings) and/or human review |
| Eval gate runs once, by hand | The gate runs in CI on every candidate model; nothing ships unless it clears the bar |
| The base model is fixed | A model registry + the option to re-tune as the base model and your traffic change |
| Cost is estimated, never spent | A training-cost budget and per-call cost tracking once the tuned model serves traffic |
The general ops machinery (observability, cost, reliability, caching, guardrails, prompt versioning, eval gates) gets built from scratch and wired into one running app in Production, #8 in the series, which also runs offline on a mock provider.
File map
check_setup.py ← run first: Python, packages, provider, key
README.md ← this guide
EXERCISES.md ← predict-then-run prompts, one per section
finetune/ ← the from-scratch library (read it!)
providers.py ← chat for mock / openai / claude (one interface)
mock_tuner.py ← the offline simulated fine-tune: upload→job→poll→use
dataset.py ← load / split the chat-JSONL training data
databuild.py ← build training examples (incl. the distillation set)
validate.py ← offline checks: schema, dupes, balance, token+cost
evaluate.py ← base vs. tuned: accuracy + pairwise win-rate
datasets/
support_train.jsonl ← the hand-made training set (the running example)
support_eval.jsonl ← a HELD-OUT eval set (none of it is in training)
support_distilled.jsonl ← built by examples/08 (git-ignored; regenerate it)
hands_on/
finetune_run.py ← capstone: validate → tune → eval-gate → ship-if-wins
examples/
01_when_to_finetune.py ← when NOT to fine-tune (offline)
02_dataset_format.py ← the chat JSONL format; build + split (offline)
03_validate_data.py ← catch a bad dataset before training (offline)
04_run_finetune.py ← upload→create→poll→use (mock; --real for OpenAI)
05_use_model.py ← base vs. tuned, side by side (offline on mock)
06_did_it_help.py ← prove it beat the baseline on held-out data (offline)
07_hyperparameters.py ← epochs, LR, batch size + the loss curve (offline)
08_distillation.py ← build a training set from a teacher model (offline)
09_open_weights_lora.py ← LoRA/PEFT & open weights, explained (offline)
10_preference_tuning.py ← DPO/RLHF: learning from {chosen, rejected} pairs (offline)
Troubleshooting
Run python check_setup.py first; it catches most problems. Then, by symptom:
| What you see | What it means / the fix |
|---|---|
PROVIDER=openai needs OPENAI_API_KEY |
You picked a real provider. Either add the key, or set PROVIDER=mock to stay offline and free. |
A real fine-tune won't start without --real |
Working as intended; the paid path is opt-in. Add --real (and confirm the cost prompt) only when you mean it. |
PROVIDER=claude can't run a fine-tune |
Working as intended; Anthropic fine-tuning isn't self-serve. Use mock to learn, or openai for a real job. |
| Validation reports duplicates / imbalance | That's the check doing its job. Fix the dataset before training; that's far cheaper than a wasted run. |
| The tuned model didn't beat the baseline | The honest outcome sometimes. Don't ship it; improve the dataset (more, cleaner, better-balanced examples) and re-measure. |
ModuleNotFoundError (openai / anthropic) |
Only needed for real providers. On the default mock stack you need only python-dotenv. |
SyntaxError / odd type errors on startup |
You're likely on Python 3.10 or older; this repo needs 3.11+. check_setup.py confirms your version. |
Still stuck? Every file is small and self-contained. Open it, read the docstring at the top, and run it. finetune/mock_tuner.py is the whole "fine-tune lifecycle" in one readable file.
The series
This is one of the standalone, hands-on deep dives into building with LLM APIs. Eight core dives, plus the bonus ones listed below. Each one stands on its own, with its own setup, examples, and capstone, and they all share one house style. Provider-agnostic where it makes sense, built from scratch with no frameworks, offline-first examples, and a real capstone at the end. Do them in any order. This sequence builds naturally.
- OpenAI API: the API from zero
- Claude API: the same ideas, the Anthropic way
- Prompt Engineering: shape model behavior with better prompts
- RAG: answer questions over your own documents
- Evals: measure whether a change actually helps
- Agents: give a model tools and a loop so it can act
- Prompt Injection & Guardrails: attack and defend all of the above
- Production: operate one app end to end
Bonus dives, standalone and slotting in where they're most useful:
- Context Engineering: manage what's in the window, with memory, compaction, and assembly
- AI Data Engineering: the corpus behind the index, with versions, lineage, ACLs, and deletes
- Multimodal: images and audio as well as text
- Fine-tuning: teach a model new behavior by example
- MCP: serve tools, data, and prompts to any LLM over a standard protocol
- Local Models: run open-weight models on your own machine
- Agent Harnesses: build on the loop, adding hooks, permissions, sandboxing, and subagents
- Realtime Voice: low-latency speech-to-speech agents
- Observability: watch a running app over time, covering drift, quality, alerting, and the feedback loop
- Architecture: the seams between the components, each decision measured rather than asserted
- GenAI Security: treat the model as an untrusted principal, and put identity, supply chain, isolation, budgets, and release gates around it
- Inference Platform Engineering: turn finite GPU memory and a request queue into latency, throughput, and a fleet size you can defend
- Testing & Delivery: decide whether a build is fit to promote, using evidence, gates, staged rollout, and rollback
- Professional Tools: rebuild each hand-written piece with the tool professionals reach for, and measure both
And the whole series lands in one codebase in the capstone: a codebase Q&A tool built step by step, one tag per dive.
Fine-tuning is a bonus dive in the series. It slots most naturally after RAG (#4), whose "RAG, fine-tuning, or something else?" decision this repo makes hands-on, and leans on Evals (#5) to prove a tune actually helped.