av / dives /Context Engineering
source about me

Bonus dive

Context Engineering: A Guided Deep Dive

A hands-on playground for the skill the other dives keep bumping into: managing what's in the context window. A model only knows what you put in front of it right now, so as conversations get long, documents pile up, and agents loop, the real work becomes deciding what to keep, what to drop, what to summarize, and in what order. You'll build a token budgeter, three kinds of conversation memory, a persistent long-term store, and a context assembler from scratch, then watch a chat that remembers under a fixed budget.

Here's what makes this repo work. It runs completely offline on a mock provider, with no API key. The mock is a deterministic "model" that answers recall questions only from facts actually present in the window. So when a fact falls off a sliding window the mock genuinely forgets it, and when compaction or long-term memory keeps it, the mock genuinely remembers. The whole thesis is visible, offline, for $0. Flip one env var and the same code runs against a real OpenAI or Claude model.

This repo is standalone, and it's also the missing half of prompt engineering. If Prompt Engineering is how you ask, this is what the model can see when you ask. It extends the "memory is the message list you resend" idea from the API and Agents dives, and its long-term memory is the RAG pattern pointed at a conversation. Its code depends on none of them.

Like its siblings, walk through it. Each section ends with something to run, and every section runs offline and free on the mock. EXERCISES.md has a predict-then-run prompt for each one.


0. The one big idea

The model only knows what's in its context window right now. Context engineering is deciding what goes in, in what order, and what to drop when it won't all fit.

That's the whole repo. Memory isn't a model feature. It's a policy you implement. Resend the message list, but the list has a budget, so you choose what survives. Compaction summarizes what won't fit. Long-term memory stores what should outlive the conversation. Assembly decides which competing sources make the cut and where they sit. Every section is a variation on that one sentence. Hold onto it and none of this feels complicated.


1. Setup (5 minutes)

bash
# 1. Create an isolated Python environment
python3 -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate

# 2. Install dependencies (the default mock stack needs only python-dotenv)
pip install -r requirements.txt

# 3. Copy the env file: the default runs keyless (no API key needed)
cp .env.example .env
#    (Real provider instead of the mock? Its key goes in your OS keychain,
#     not .env: see ../docs/SECRETS.md, then run scripts as `secrun python ...`.)

# 4. Confirm everything is wired up (makes no API call, costs nothing)
python check_setup.py

That's it, and no key is required. The default PROVIDER=mock is a deterministic in-process "model". Pick your stack with PROVIDER in .env.

PROVIDER What runs the model Keys needed Cost
mock (default) a deterministic offline "model" none $0
openai OpenAI gpt-6-luna OPENAI_API_KEY tiny
claude Claude claude-haiku-4-5 ANTHROPIC_API_KEY tiny

The only file that knows which one you picked is context/providers.py. Everything else is provider-neutral logic about what goes in the window.

Why a mock is the right call here. The subject is the context window rather than the model. The mock answers recall questions only from what's actually in the messages you pass it, so forgetting and remembering are observable offline and deterministic. Real models behave the same way, just less predictably.


2. The window is a budget

bash
python examples/01_token_budget.py        # offline, no model needed

Everything starts with arithmetic. The context window is a fixed number of tokens, and a conversation only grows. context/tokens.py estimates tokens with a simple 4-characters-per-token heuristic, no tokenizer and no key, and this example grows a chat turn by turn until it overflows a deliberately tiny 2,000-token window. That overflow is the problem the rest of the repo solves.


3. The sliding window, and what it forgets

bash
python examples/02_sliding_window.py

The simplest fix for overflow is to keep the system prompt plus the most recent turns that fit, and let the oldest scroll off. WindowMemory does exactly that. It's bounded, cheap, and genuinely forgetful. Dana introduces herself, the chat runs long, and when you ask her name it's gone. That's the failure the "simple trim" in the other dives has and never mentions. For the model, the conversation is the window.


4. Compaction: summarize what falls off

bash
python examples/03_compaction.py

Same budget as §3, but instead of deleting old turns, SummaryMemory folds them into a running summary and keeps the recent turns verbatim. The exact words are gone. The facts survive. Run the same long conversation and this time the model recalls both the name and the thing it was asked to remember. This is the most important technique in the repo, and it's what real assistants do when a long chat remembers. You trade exact wording, plus one summarization call, for durable memory under a fixed budget.


5. Long-term memory: remembering across sessions

bash
python examples/04_long_term_memory.py

Compaction keeps a fact alive within a conversation. Close the session and it's gone. Long-term memory writes durable facts to a store outside the window and retrieves the relevant few back in when a new turn needs them. RAG pointed at the conversation. context/longterm.py persists facts to JSON and retrieves by overlap; swap in real embeddings from the RAG dive for production. The example runs two sessions against a fresh, empty window, and session two answers correctly only because it recalled a fact session one stored.


6. Order matters, so don't bury the lede

bash
python examples/05_ordering.py        # offline

Fitting the right text is half the job. Where you put it is the other half. Models attend most reliably to the start and the end of a long context and can miss what's buried in the middle, the "lost in the middle" effect. order_for_recall places the highest-priority sections at the edges and the filler in the middle. Same tokens, better recall.


7. Assembling the window under a budget

bash
python examples/06_assemble_budget.py        # offline

A real request's context competes for space between the system prompt, tools, retrieved docs, long-term memory, and recent turns, and they rarely all fit. assemble() prioritizes, then packs. Keep the highest-priority sections that fit and drop the rest, so a marketing blog can't crowd out the user's actual question. Assembling the window on purpose is the difference between a focused request and a bloated one.


8. More context isn't better (context rot)

bash
python examples/07_context_rot.py

A big window is a budget rather than a goal. Padding it just in case dilutes the signal, invites the model to latch onto an irrelevant passage, and bills you for every wasted token on every turn. That's context rot. The example answers the same question with a lean context and with a bloated one where a plausible distractor names a different person, and the bloated window costs about 10× the tokens and returns the wrong name. On the mock that flip is deterministic, since it naively takes the last "my name is ..." it sees, a crude stand-in for a real model's attention wandering to a distractor. Add a key and a harder question to watch a subtler version on a real model. Relevance beats volume, and the cheapest, fastest, most accurate token is the one you didn't send.


9. Pruning an agent's observations

bash
python examples/08_pruning_observations.py        # offline

Agents bloat context fastest. Every step appends a tool call and its full result, and tool results are huge. Ten steps in, the window is mostly stale observations the agent already used. Observation pruning keeps the reasoning and the most recent results verbatim and stubs the old ones. Here that gives a 75% smaller window with the agent's train of thought intact. It's compaction, long-term memory, and assembly applied to an agent's context.


10. The hidden cost, where compaction breaks your prompt cache

bash
python examples/09_caching_vs_compaction.py        # offline

Every technique so far shrinks the window. This one shows the bill they can raise instead. Providers cache the prompt prefix, and the rule is unforgiving. Any change anywhere in the prefix invalidates everything after it. An append-only history never touches its prefix, so every prior token is a cheap cache read at about 0.1× and only the new turn is written at about 1.25×. Compaction rewrites the prefix. It changes the system prompt, which is the summary, and drops old turns, so the next request is a cache miss that pays full write price on the whole context. The example bills the same conversation both ways in context/cost.py and finds compaction costing about 1.5× as much here. Fewer tokens, bigger bill. That's the honest tradeoff this series insists on: cheaper context and cheaper bill are different axes. There's a crossover, though. On very long chats the unbounded append-only window finally loses, so when you have to compact, do it rarely and in bulk, paying the cache miss once instead of every turn.

On the newest Claude models, the rewrite can also be an error

bash
python examples/11_append_only_contract.py           # offline simulation
secrun python examples/11_append_only_contract.py --real   # the real check, about a cent

Claude Fable 5.1, Opus 5.5, and Sonnet 5.5 sign each thinking block with the exact prefix it was produced after: the system prompt, the tools, and every earlier message. Send the block back after editing any of those and it's invalid. Accounts created on or after 2026-08-31 get a 400 by default ("the block is bound to a different conversation"). Older accounts get no error and the model still reads the block, so code can pass on your key and fail on your users'. The example measures it live: an append-only request is accepted, the same request with one sentence added to the system prompt is rejected, and prefix_mismatch_behavior: "drop_block" makes it go through by silently dropping the reasoning.

The edits that break thinking are the edits that restart the cache, so the rule is the one this section already argued for on cost: freeze system and tools for the session and only append to messages. To change instructions mid-session, append a role: "system" message rather than editing system. When the history has to shrink, summarize the whole session into one new user message and send no earlier turns. That leaves no old thinking to check. The sliding window and the summary-in-the-system-prompt compaction from sections 3 and 4 both fail the check, and so would this dive's chat.py on those models. It runs on Haiku 4.5, which doesn't do this check.


11. When the API does it for you

bash
secrun python examples/10_server_side_compaction.py

Everything above is hand-rolled on purpose, because you can't reason about a tradeoff you've never implemented. But two of these jobs now exist as server-side features on the Anthropic API, and knowing which is which saves you writing them twice.

Feature What it does Maps to Model
Compaction (compact_20260112) Summarizes earlier turns near a threshold (150K default) §4 Sonnet/Opus 4.6+; Haiku 4.5 is a 400
Context editing (clear_tool_uses_20250919) Clears old tool results outright, no summary §9 works on Haiku 4.5

Summarize against clear is the whole decision. A summary costs tokens to produce and keeps a lossy trace. Clearing costs nothing and keeps nothing. For a chat transcript you usually want the summary, because the user will refer back to it. For the raw output of a grep forty agent steps ago, clearing beats paying a model to write a paragraph about it.

Two things to carry away. First, a trap. With compaction on you have to append response.content, the whole block list, to your history rather than the extracted text, because the API returns a compaction block that carries the state. Code that keeps only .text works fine right up until the first real compaction, and then loses it with no warning. Second, what hasn't changed. Server-side compaction is still a prefix rewrite, so §10's cache arithmetic applies exactly as before. Moving the work to the server makes it easier to maintain, not free to run.


12. The capstone: chat.py

Everything assembled into a chat you'd actually use. It stays inside a token budget no matter how long you talk, through compaction, and it remembers durable facts across sessions, through long-term memory. Each turn, the system prompt is your persona, the running summary, and the long-term facts relevant to what you just asked.

bash
# State some facts (offline on the mock: no key, no cost):
python hands_on/chat.py "Hi, my name is Dana. Remember our launch is Friday."

# A BRAND-NEW run, and it still knows, because the fact was persisted:
python hands_on/chat.py "When is my launch?"

# Interactive REPL ('/context' to watch the window, '/memory' to list stored facts):
python hands_on/chat.py

# See exactly what's sent each turn; shrink the budget to watch compaction kick in:
python hands_on/chat.py --show-context --budget 200

# Wipe the long-term store:
python hands_on/chat.py --forget

Read hands_on/chat.py. respond() is the whole turn: recall relevant facts, assemble the system prompt, generate, persist any new durable facts. The library does the work and the capstone wires it to a CLI. Suggested exercise: chat for a dozen turns with --show-context --budget 200 and watch the compaction count climb while the window stays under budget. Then quit, run it again, and notice it still greets you by name.


Where to go next

You've built memory from scratch. What comes next is more of the same idea, at more scale and more rigor.

  • Real embeddings for long-term memory. Swap the keyword overlap for the vector store from the RAG dive, so recall works by meaning rather than shared words.
  • Smarter compaction. Summarize hierarchically, or keep structured state, a running JSON of facts and decisions, alongside the prose summary.
  • Memory as a tool. Let an agent decide what to remember and recall by calling remember() and recall() tools, instead of doing it on every turn.
  • Eviction and freshness. Expire stale facts, resolve contradictions ("I moved to the Team plan"), and rank memory by recency as well as relevance.
  • Measure it. Score whether compaction preserved the facts that mattered, using the Evals dive. A memory bug is a silent quality regression.
  • Multi-agent context isolation. Give sub-agents their own focused windows so one agent's clutter never pollutes another's.

From teaching code to production

The teaching shortcuts here are exactly what you'd harden once a memory layer sits on a live path.

This repo's teaching shortcut In production
Token count is a ~4-chars/token estimate The real tokenizer / the API's usage, and budgets enforced per request
Long-term memory is a local JSON file A real vector DB with embeddings, per-user namespaces, and access control
Compaction summarizes inline, every overflow Summarization wrapped in retries + a cost budget (it's an extra model call)
Recall is keyword overlap Embedding similarity + reranking, and a relevance threshold so junk isn't injected
Facts are trusted and never expire Eviction, contradiction resolution, and provenance on stored memory
A summary might drop a needed fact with nothing to show for it An eval gate on memory quality, so a compaction regression fails the build
Compaction/pruning run on every overflow (§10) Cache-aware memory: keep the prefix stable, compact rarely and in bulk, and watch cache_read_input_tokens vs cache_creation_input_tokens so shrinking the window doesn't grow the bill
Stored memory is trusted text Guardrails: memory is untrusted input and a classic indirect-injection vector

The general ops machinery (observability, cost, reliability, caching, guardrails, prompt versioning, eval gates) gets built from scratch and wired into one running app in Production, #8 in the series, which also runs offline on a mock provider.


File map

check_setup.py              ← run first: Python, packages, provider
README.md                   ← this guide
EXERCISES.md                ← predict-then-run prompts, one per section
context/                    ← the from-scratch library (read it!)
  providers.py              ← the ONLY provider file: mock (default) + openai + claude
  tokens.py                 ← estimate the token budget (offline heuristic)
  cost.py                   ← a tiny prompt-cache cost model (read vs write vs miss)
  memory.py                 ← Full / Window / Summary (compaction) conversation memory
  longterm.py               ← persistent cross-session memory with retrieval
  assemble.py               ← fit & order sections under a token budget
hands_on/
  chat.py                   ← capstone: a chat that compacts + remembers across runs
examples/
  01_token_budget.py        ← the window is a budget (offline)
  02_sliding_window.py      ← keep recent, drop old, and what it forgets
  03_compaction.py          ← summarize what falls off instead of deleting it
  04_long_term_memory.py    ← remembering across sessions (RAG over the conversation)
  05_ordering.py            ← lost in the middle: put what matters at the edges (offline)
  06_assemble_budget.py     ← prioritize, then pack the window (offline)
  07_context_rot.py         ← more context is not better
  08_pruning_observations.py← trim stale tool results in an agent loop (offline)
  09_caching_vs_compaction.py← compaction blows the prompt cache: fewer tokens, bigger bill (offline)
  10_server_side_compaction.py ← the API's own compaction & context editing
  11_append_only_contract.py   ← history edits that 400 on the newest Claude models (offline + --real)

(.ctx_memory.json is created by the capstone's long-term memory and is git-ignored.)


Troubleshooting

Run python check_setup.py first; it catches most problems. Then, by symptom:

What you see What it means / the fix
ModuleNotFoundError: dotenv Dependencies aren't installed or the venv isn't active. source .venv/bin/activate then pip install -r requirements.txt.
PROVIDER=... needs ... in the environment You switched to a real provider without a key. Load it from your keychain with secrun (see SECRETS.md), or go back to PROVIDER=mock.
The capstone "remembers" things from a previous run That's long-term memory working; it persists to .ctx_memory.json. Run python hands_on/chat.py --forget to wipe it.
On a real provider, recall is fuzzier than the mock The mock is deterministic; real models paraphrase and occasionally miss. That's why §8 (don't overload) and the Evals dive matter.
The summary dropped a fact I needed Compaction is lossy; that's the tradeoff. Keep more recent turns verbatim (keep_recent) or store the fact in long-term memory.
SyntaxError / odd type errors on startup You're likely on Python 3.10 or older; this repo needs 3.11+. check_setup.py confirms your version.

Still stuck? Every file is small and self-contained. Open it, read the docstring at the top, and run it directly. context/memory.py is the heart of the dive.


The series

This is one of the standalone, hands-on deep dives into building with LLM APIs: eight core, plus the bonus dives listed below. Each one stands on its own, with its own setup, examples, and capstone, and they all share the same house style: provider-agnostic, built from scratch (no frameworks), offline-first examples, and a real capstone. Do them in any order; this sequence builds naturally:

  1. OpenAI API: the API from zero
  2. Claude API: the same ideas, the Anthropic way
  3. Prompt Engineering: shape model behavior with better prompts
  4. RAG: answer questions over your own documents
  5. Evals: measure whether a change actually helps
  6. Agents: give a model tools and a loop so it can act
  7. Prompt Injection & Guardrails: attack and defend all of the above
  8. Production: operate one app end to end

Bonus dives, standalone and slotting in where they're most useful:

  • Context Engineering: manage what's in the window, with memory, compaction, and assembly
  • AI Data Engineering: the corpus behind the index, with versions, lineage, ACLs, and deletes
  • Multimodal: images and audio as well as text
  • Fine-tuning: teach a model new behavior by example
  • MCP: serve tools, data, and prompts to any LLM over a standard protocol
  • Local Models: run open-weight models on your own machine
  • Agent Harnesses: build on the loop, adding hooks, permissions, sandboxing, and subagents
  • Realtime Voice: low-latency speech-to-speech agents
  • Observability: watch a running app over time, covering drift, quality, alerting, and the feedback loop
  • Architecture: the seams between the components, each decision measured rather than asserted
  • GenAI Security: treat the model as an untrusted principal, and put identity, supply chain, isolation, budgets, and release gates around it
  • Inference Platform Engineering: turn finite GPU memory and a request queue into latency, throughput, and a fleet size you can defend
  • Testing & Delivery: decide whether a build is fit to promote, using evidence, gates, staged rollout, and rollback
  • Professional Tools: rebuild each hand-written piece with the tool professionals reach for, and measure both

And the whole series lands in one codebase in the capstone: a codebase Q&A tool built step by step, one tag per dive.

Context Engineering is a bonus dive in the series. It slots most naturally after Agents (#6), whose stateless "resend the message list" memory it extends into compaction and long-term recall, and pairs with RAG (#4), which its long-term memory reuses.