Bonus dive
Context Engineering: A Guided Deep Dive
A hands-on playground for the skill the other dives keep bumping into: managing what's in the context window. A model only knows what you put in front of it right now, so as conversations get long, documents pile up, and agents loop, the real work becomes deciding what to keep, what to drop, what to summarize, and in what order. You'll build a token budgeter, three kinds of conversation memory, a persistent long-term store, and a context assembler from scratch, then watch a chat that remembers under a fixed budget.
Here's what makes this repo work. It runs completely offline on a mock provider, with no API key. The mock is a deterministic "model" that answers recall questions only from facts actually present in the window. So when a fact falls off a sliding window the mock genuinely forgets it, and when compaction or long-term memory keeps it, the mock genuinely remembers. The whole thesis is visible, offline, for $0. Flip one env var and the same code runs against a real OpenAI or Claude model.
This repo is standalone, and it's also the missing half of prompt engineering. If Prompt Engineering is how you ask, this is what the model can see when you ask. It extends the "memory is the message list you resend" idea from the API and Agents dives, and its long-term memory is the RAG pattern pointed at a conversation. Its code depends on none of them.
Like its siblings, walk through it. Each section ends with something to run, and every section runs offline and free on the mock. EXERCISES.md has a predict-then-run prompt for each one.
0. The one big idea
The model only knows what's in its context window right now. Context engineering is deciding what goes in, in what order, and what to drop when it won't all fit.
That's the whole repo. Memory isn't a model feature. It's a policy you implement. Resend the message list, but the list has a budget, so you choose what survives. Compaction summarizes what won't fit. Long-term memory stores what should outlive the conversation. Assembly decides which competing sources make the cut and where they sit. Every section is a variation on that one sentence. Hold onto it and none of this feels complicated.
1. Setup (5 minutes)
# 1. Create an isolated Python environment
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 2. Install dependencies (the default mock stack needs only python-dotenv)
pip install -r requirements.txt
# 3. Copy the env file: the default runs keyless (no API key needed)
cp .env.example .env
# (Real provider instead of the mock? Its key goes in your OS keychain,
# not .env: see ../docs/SECRETS.md, then run scripts as `secrun python ...`.)
# 4. Confirm everything is wired up (makes no API call, costs nothing)
python check_setup.py
That's it, and no key is required. The default PROVIDER=mock is a deterministic
in-process "model". Pick your stack with PROVIDER in .env.
PROVIDER |
What runs the model | Keys needed | Cost |
|---|---|---|---|
mock (default) |
a deterministic offline "model" | none | $0 |
openai |
OpenAI gpt-6-luna |
OPENAI_API_KEY |
tiny |
claude |
Claude claude-haiku-4-5 |
ANTHROPIC_API_KEY |
tiny |
The only file that knows which one you picked is context/providers.py. Everything else is provider-neutral logic about what goes in the window.
Why a mock is the right call here. The subject is the context window rather than the model. The mock answers recall questions only from what's actually in the messages you pass it, so forgetting and remembering are observable offline and deterministic. Real models behave the same way, just less predictably.
2. The window is a budget
python examples/01_token_budget.py # offline, no model needed
Everything starts with arithmetic. The context window is a fixed number of tokens, and a conversation only grows. context/tokens.py estimates tokens with a simple 4-characters-per-token heuristic, no tokenizer and no key, and this example grows a chat turn by turn until it overflows a deliberately tiny 2,000-token window. That overflow is the problem the rest of the repo solves.
3. The sliding window, and what it forgets
python examples/02_sliding_window.py
The simplest fix for overflow is to keep the system prompt plus the most recent turns that
fit, and let the oldest scroll off. WindowMemory does exactly that. It's bounded, cheap,
and genuinely forgetful. Dana introduces herself, the chat runs long, and when you ask her
name it's gone. That's the failure the "simple trim" in the other dives has and never
mentions. For the model, the conversation is the window.
4. Compaction: summarize what falls off
python examples/03_compaction.py
Same budget as §3, but instead of deleting old turns, SummaryMemory folds them into a
running summary and keeps the recent turns verbatim. The exact words are gone. The facts
survive. Run the same long conversation and this time the model recalls both the name and
the thing it was asked to remember. This is the most important technique in the repo, and
it's what real assistants do when a long chat remembers. You trade exact wording, plus
one summarization call, for durable memory under a fixed budget.
5. Long-term memory: remembering across sessions
python examples/04_long_term_memory.py
Compaction keeps a fact alive within a conversation. Close the session and it's gone. Long-term memory writes durable facts to a store outside the window and retrieves the relevant few back in when a new turn needs them. RAG pointed at the conversation. context/longterm.py persists facts to JSON and retrieves by overlap; swap in real embeddings from the RAG dive for production. The example runs two sessions against a fresh, empty window, and session two answers correctly only because it recalled a fact session one stored.
6. Order matters, so don't bury the lede
python examples/05_ordering.py # offline
Fitting the right text is half the job. Where you put it is the other half. Models attend
most reliably to the start and the end of a long context and can miss what's buried in the
middle, the "lost in the middle" effect. order_for_recall places the highest-priority
sections at the edges and the filler in the middle. Same tokens, better recall.
7. Assembling the window under a budget
python examples/06_assemble_budget.py # offline
A real request's context competes for space between the system prompt, tools, retrieved
docs, long-term memory, and recent turns, and they rarely all fit. assemble()
prioritizes, then packs. Keep the highest-priority sections that fit and drop the rest, so
a marketing blog can't crowd out the user's actual question. Assembling the window on
purpose is the difference between a focused request and a bloated one.
8. More context isn't better (context rot)
python examples/07_context_rot.py
A big window is a budget rather than a goal. Padding it just in case dilutes the signal, invites the model to latch onto an irrelevant passage, and bills you for every wasted token on every turn. That's context rot. The example answers the same question with a lean context and with a bloated one where a plausible distractor names a different person, and the bloated window costs about 10× the tokens and returns the wrong name. On the mock that flip is deterministic, since it naively takes the last "my name is ..." it sees, a crude stand-in for a real model's attention wandering to a distractor. Add a key and a harder question to watch a subtler version on a real model. Relevance beats volume, and the cheapest, fastest, most accurate token is the one you didn't send.
9. Pruning an agent's observations
python examples/08_pruning_observations.py # offline
Agents bloat context fastest. Every step appends a tool call and its full result, and tool results are huge. Ten steps in, the window is mostly stale observations the agent already used. Observation pruning keeps the reasoning and the most recent results verbatim and stubs the old ones. Here that gives a 75% smaller window with the agent's train of thought intact. It's compaction, long-term memory, and assembly applied to an agent's context.
10. The hidden cost, where compaction breaks your prompt cache
python examples/09_caching_vs_compaction.py # offline
Every technique so far shrinks the window. This one shows the bill they can raise instead. Providers cache the prompt prefix, and the rule is unforgiving. Any change anywhere in the prefix invalidates everything after it. An append-only history never touches its prefix, so every prior token is a cheap cache read at about 0.1× and only the new turn is written at about 1.25×. Compaction rewrites the prefix. It changes the system prompt, which is the summary, and drops old turns, so the next request is a cache miss that pays full write price on the whole context. The example bills the same conversation both ways in context/cost.py and finds compaction costing about 1.5× as much here. Fewer tokens, bigger bill. That's the honest tradeoff this series insists on: cheaper context and cheaper bill are different axes. There's a crossover, though. On very long chats the unbounded append-only window finally loses, so when you have to compact, do it rarely and in bulk, paying the cache miss once instead of every turn.
On the newest Claude models, the rewrite can also be an error
python examples/11_append_only_contract.py # offline simulation
secrun python examples/11_append_only_contract.py --real # the real check, about a cent
Claude Fable 5.1, Opus 5.5, and Sonnet 5.5 sign each thinking block with the exact prefix
it was produced after: the system prompt, the tools, and every earlier message. Send the
block back after editing any of those and it's invalid. Accounts created on or after
2026-08-31 get a 400 by default ("the block is bound to a different conversation"). Older
accounts get no error and the model still reads the block, so code can pass on your key
and fail on your users'. The example measures it live: an append-only request is
accepted, the same request with one sentence added to the system prompt is rejected, and
prefix_mismatch_behavior: "drop_block" makes it go through by silently dropping the
reasoning.
The edits that break thinking are the edits that restart the cache, so the rule is the
one this section already argued for on cost: freeze system and tools for the session
and only append to messages. To change instructions mid-session, append a
role: "system" message rather than editing system. When the history has to shrink,
summarize the whole session into one new user message and send no earlier turns. That
leaves no old thinking to check. The sliding window and the summary-in-the-system-prompt
compaction from sections 3 and 4 both fail the check, and so would this dive's chat.py
on those models. It runs on Haiku 4.5, which doesn't do this check.
11. When the API does it for you
secrun python examples/10_server_side_compaction.py
Everything above is hand-rolled on purpose, because you can't reason about a tradeoff you've never implemented. But two of these jobs now exist as server-side features on the Anthropic API, and knowing which is which saves you writing them twice.
| Feature | What it does | Maps to | Model |
|---|---|---|---|
Compaction (compact_20260112) |
Summarizes earlier turns near a threshold (150K default) | §4 | Sonnet/Opus 4.6+; Haiku 4.5 is a 400 |
Context editing (clear_tool_uses_20250919) |
Clears old tool results outright, no summary | §9 | works on Haiku 4.5 |
Summarize against clear is the whole decision. A summary costs tokens to produce and keeps
a lossy trace. Clearing costs nothing and keeps nothing. For a chat transcript you usually
want the summary, because the user will refer back to it. For the raw output of a grep
forty agent steps ago, clearing beats paying a model to write a paragraph about it.
Two things to carry away. First, a trap. With compaction on you have to append
response.content, the whole block list, to your history rather than the extracted text,
because the API returns a compaction block that carries the state. Code that keeps only
.text works fine right up until the first real compaction, and then loses it with no
warning. Second, what hasn't changed. Server-side compaction is still a prefix rewrite,
so §10's cache arithmetic applies exactly as before. Moving the work to the server makes
it easier to maintain, not free to run.
12. The capstone: chat.py
Everything assembled into a chat you'd actually use. It stays inside a token budget no matter how long you talk, through compaction, and it remembers durable facts across sessions, through long-term memory. Each turn, the system prompt is your persona, the running summary, and the long-term facts relevant to what you just asked.
# State some facts (offline on the mock: no key, no cost):
python hands_on/chat.py "Hi, my name is Dana. Remember our launch is Friday."
# A BRAND-NEW run, and it still knows, because the fact was persisted:
python hands_on/chat.py "When is my launch?"
# Interactive REPL ('/context' to watch the window, '/memory' to list stored facts):
python hands_on/chat.py
# See exactly what's sent each turn; shrink the budget to watch compaction kick in:
python hands_on/chat.py --show-context --budget 200
# Wipe the long-term store:
python hands_on/chat.py --forget
Read hands_on/chat.py. respond() is the whole turn: recall relevant
facts, assemble the system prompt, generate, persist any new durable facts. The library
does the work and the capstone wires it to a CLI. Suggested exercise: chat for a dozen
turns with --show-context --budget 200 and watch the compaction count climb while the
window stays under budget. Then quit, run it again, and notice it still greets you by
name.
Where to go next
You've built memory from scratch. What comes next is more of the same idea, at more scale and more rigor.
- Real embeddings for long-term memory. Swap the keyword overlap for the vector store from the RAG dive, so recall works by meaning rather than shared words.
- Smarter compaction. Summarize hierarchically, or keep structured state, a running JSON of facts and decisions, alongside the prose summary.
- Memory as a tool. Let an agent decide what to remember and recall by calling
remember()andrecall()tools, instead of doing it on every turn. - Eviction and freshness. Expire stale facts, resolve contradictions ("I moved to the Team plan"), and rank memory by recency as well as relevance.
- Measure it. Score whether compaction preserved the facts that mattered, using the Evals dive. A memory bug is a silent quality regression.
- Multi-agent context isolation. Give sub-agents their own focused windows so one agent's clutter never pollutes another's.
From teaching code to production
The teaching shortcuts here are exactly what you'd harden once a memory layer sits on a live path.
| This repo's teaching shortcut | In production |
|---|---|
| Token count is a ~4-chars/token estimate | The real tokenizer / the API's usage, and budgets enforced per request |
| Long-term memory is a local JSON file | A real vector DB with embeddings, per-user namespaces, and access control |
| Compaction summarizes inline, every overflow | Summarization wrapped in retries + a cost budget (it's an extra model call) |
| Recall is keyword overlap | Embedding similarity + reranking, and a relevance threshold so junk isn't injected |
| Facts are trusted and never expire | Eviction, contradiction resolution, and provenance on stored memory |
| A summary might drop a needed fact with nothing to show for it | An eval gate on memory quality, so a compaction regression fails the build |
| Compaction/pruning run on every overflow (§10) | Cache-aware memory: keep the prefix stable, compact rarely and in bulk, and watch cache_read_input_tokens vs cache_creation_input_tokens so shrinking the window doesn't grow the bill |
| Stored memory is trusted text | Guardrails: memory is untrusted input and a classic indirect-injection vector |
The general ops machinery (observability, cost, reliability, caching, guardrails, prompt versioning, eval gates) gets built from scratch and wired into one running app in Production, #8 in the series, which also runs offline on a mock provider.
File map
check_setup.py ← run first: Python, packages, provider
README.md ← this guide
EXERCISES.md ← predict-then-run prompts, one per section
context/ ← the from-scratch library (read it!)
providers.py ← the ONLY provider file: mock (default) + openai + claude
tokens.py ← estimate the token budget (offline heuristic)
cost.py ← a tiny prompt-cache cost model (read vs write vs miss)
memory.py ← Full / Window / Summary (compaction) conversation memory
longterm.py ← persistent cross-session memory with retrieval
assemble.py ← fit & order sections under a token budget
hands_on/
chat.py ← capstone: a chat that compacts + remembers across runs
examples/
01_token_budget.py ← the window is a budget (offline)
02_sliding_window.py ← keep recent, drop old, and what it forgets
03_compaction.py ← summarize what falls off instead of deleting it
04_long_term_memory.py ← remembering across sessions (RAG over the conversation)
05_ordering.py ← lost in the middle: put what matters at the edges (offline)
06_assemble_budget.py ← prioritize, then pack the window (offline)
07_context_rot.py ← more context is not better
08_pruning_observations.py← trim stale tool results in an agent loop (offline)
09_caching_vs_compaction.py← compaction blows the prompt cache: fewer tokens, bigger bill (offline)
10_server_side_compaction.py ← the API's own compaction & context editing
11_append_only_contract.py ← history edits that 400 on the newest Claude models (offline + --real)
(.ctx_memory.json is created by the capstone's long-term memory and is git-ignored.)
Troubleshooting
Run python check_setup.py first; it catches most problems. Then, by symptom:
| What you see | What it means / the fix |
|---|---|
ModuleNotFoundError: dotenv |
Dependencies aren't installed or the venv isn't active. source .venv/bin/activate then pip install -r requirements.txt. |
PROVIDER=... needs ... in the environment |
You switched to a real provider without a key. Load it from your keychain with secrun (see SECRETS.md), or go back to PROVIDER=mock. |
| The capstone "remembers" things from a previous run | That's long-term memory working; it persists to .ctx_memory.json. Run python hands_on/chat.py --forget to wipe it. |
| On a real provider, recall is fuzzier than the mock | The mock is deterministic; real models paraphrase and occasionally miss. That's why §8 (don't overload) and the Evals dive matter. |
| The summary dropped a fact I needed | Compaction is lossy; that's the tradeoff. Keep more recent turns verbatim (keep_recent) or store the fact in long-term memory. |
SyntaxError / odd type errors on startup |
You're likely on Python 3.10 or older; this repo needs 3.11+. check_setup.py confirms your version. |
Still stuck? Every file is small and self-contained. Open it, read the docstring at the top, and run it directly. context/memory.py is the heart of the dive.
The series
This is one of the standalone, hands-on deep dives into building with LLM APIs: eight core, plus the bonus dives listed below. Each one stands on its own, with its own setup, examples, and capstone, and they all share the same house style: provider-agnostic, built from scratch (no frameworks), offline-first examples, and a real capstone. Do them in any order; this sequence builds naturally:
- OpenAI API: the API from zero
- Claude API: the same ideas, the Anthropic way
- Prompt Engineering: shape model behavior with better prompts
- RAG: answer questions over your own documents
- Evals: measure whether a change actually helps
- Agents: give a model tools and a loop so it can act
- Prompt Injection & Guardrails: attack and defend all of the above
- Production: operate one app end to end
Bonus dives, standalone and slotting in where they're most useful:
- Context Engineering: manage what's in the window, with memory, compaction, and assembly
- AI Data Engineering: the corpus behind the index, with versions, lineage, ACLs, and deletes
- Multimodal: images and audio as well as text
- Fine-tuning: teach a model new behavior by example
- MCP: serve tools, data, and prompts to any LLM over a standard protocol
- Local Models: run open-weight models on your own machine
- Agent Harnesses: build on the loop, adding hooks, permissions, sandboxing, and subagents
- Realtime Voice: low-latency speech-to-speech agents
- Observability: watch a running app over time, covering drift, quality, alerting, and the feedback loop
- Architecture: the seams between the components, each decision measured rather than asserted
- GenAI Security: treat the model as an untrusted principal, and put identity, supply chain, isolation, budgets, and release gates around it
- Inference Platform Engineering: turn finite GPU memory and a request queue into latency, throughput, and a fleet size you can defend
- Testing & Delivery: decide whether a build is fit to promote, using evidence, gates, staged rollout, and rollback
- Professional Tools: rebuild each hand-written piece with the tool professionals reach for, and measure both
And the whole series lands in one codebase in the capstone: a codebase Q&A tool built step by step, one tag per dive.
Context Engineering is a bonus dive in the series. It slots most naturally after Agents (#6), whose stateless "resend the message list" memory it extends into compaction and long-term recall, and pairs with RAG (#4), which its long-term memory reuses.