Bonus dive
MCP (Model Context Protocol): A Guided Deep Dive
A hands-on playground for learning the Model Context Protocol from the ground up. It's
the open standard for handing an LLM tools, data, and prompts from a separate process.
You'll build MCP servers, write a client that talks to them, and finally let a model drive
those tools over the protocol, understanding every moving part along the way. The three
primitives (tools, resources, prompts), self-describing JSON-RPC requests, stdio against
HTTP transports, MRTR, cacheable discovery, wiring a server into a real host, and the
security model. It targets MCP 2026-07-28 and the official Python SDK v2.
Here's what makes this repo click. Most of it runs offline and free. A server and a client talk to each other with no model involved, so Sections 2 through 7 (your first server, the client, resources, prompts, a multi-tool server) need no API key at all. You only need a provider for Section 8 and the capstone, where an LLM host chooses tools.
This repo is standalone and teaches everything it needs on its own. It goes far deeper than the "Bonus: MCP" section of the Agents deep dive, where Section 8 here is that agent loop with tools served over MCP. Its security section builds on the Prompt Injection deep dive. Its code depends on neither.
Like its siblings, walk through it. Each section ends with something to run, and the first six run offline and free. EXERCISES.md has a predict-then-run prompt for each section.
0. The one big idea
MCP is a standard way to hand an LLM tools, data, and prompt templates from a separate process. Write the server once, and any MCP-speaking client or host can discover and use it.
That's the whole repo. Before MCP, every app re-implemented its own tools and glued them to its own model in its own way. MCP makes the connector standard. A server exposes capabilities. A client, sitting inside a host like Claude Desktop, an IDE, or the capstone here, connects and uses them over plain JSON-RPC. The model never knows or cares where a tool came from. To it, a tool is a name, a description, and a schema. Everything below, from resources and prompts to HTTP transport and security, is a small addition to that one idea. Hold onto it and none of this feels complicated.
1. Setup (5 minutes)
# 1. Create an isolated Python environment
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 2. Install dependencies (the official MCP SDK + a provider SDK)
pip install -r requirements.txt
# 3. Copy the env file: offline sections need no key; §8 + capstone do
cp .env.example .env
# (Real provider instead of the mock? Its key goes in your OS keychain,
# not .env: see ../docs/SECRETS.md, then run scripts as `secrun python ...`.)
# 4. Confirm everything is wired up (makes no API call, costs nothing)
secrun python check_setup.py
The mcp SDK and Python 3.11+ are required for everything. A PROVIDER and its key are
required only for the LLM-in-the-loop sections, Section 8 and the capstone.
PROVIDER |
Used for | Key needed |
|---|---|---|
| (none) | Sections 2-7: server to client with no model. Fully offline. | none |
openai (default) |
The host loop (§8 + capstone): OpenAI chat + function calling. | OPENAI_API_KEY |
claude |
The host loop (§8 + capstone): Claude messages + tool use. | ANTHROPIC_API_KEY |
MCP-first means free-first. The protocol is the subject, and the protocol doesn't need a model. You can learn the entire mechanism, covering servers, the three primitives, transports, and even the security model, without spending a cent. The LLM shows up only at the end, to use what you built.
2. The protocol in one page
python examples/01_protocol.py
MCP is a small, boring idea, and the vocabulary is worth getting straight before you launch anything.
- Host. The app the user interacts with, whether that's Claude Desktop, an IDE, or the capstone here. It contains one or more clients.
- Client. A connector inside the host that holds one connection to one server and speaks the protocol.
- Server. A separate program that exposes capabilities. It contains no model. It answers requests.
They talk over JSON-RPC 2.0, plain JSON request and response, across a transport: stdio
for a local subprocess, or streamable HTTP/SSE for a network service. And a server
exposes exactly three primitives, which the rest of the repo walks through one at a
time.
MCP 2026-07-28 has no initialize handshake and no protocol session. Each
request carries the protocol version, client identity, and capabilities it needs;
server/discover is optional. See PROTOCOL_2026.md for the
production migration and compatibility notes.
| Primitive | What it is | Who's in control |
|---|---|---|
| Tool | A function the model can call to act | Model-controlled |
| Resource | Read-only data exposed by URI (like a GET) | App-controlled |
| Prompt | A reusable, parameterized prompt template | User-controlled |
3. Your first server, and the raw client
python examples/02_first_server_and_client.py
Section 2 showed the JSON messages. Here you send real ones, using the official SDK's
client API with no wrapper, so you see the actual ceremony exactly as the SDK docs describe
it. The server, servers/calculator.py, is a dozen lines: an
MCPServer instance with one @mcp.tool() function. The client spawns it as a subprocess
over stdio, the high-level Client selects modern MCP automatically, then it runs
list_tools() and call_tool(...). This is the only example that uses the raw API
directly. After this we use a small MCPClient wrapper so the protocol stays in focus
instead of the async boilerplate.
4. A client that lists and calls a tool
python examples/03_client_calls_tool.py
This is the free-first runnable to really sit with. Same idea as §3, where a client lists
and calls a tool, but through the small MCPClient wrapper, so the
steps stand out: connect, list, call. It proves the core claim of the whole repo. You write
a server once, and any MCP-speaking client can discover and use its tools, with no LLM
anywhere. Everything later is a small addition to exactly this.
5. Resources, read-only data for the model
python examples/04_resources.py
A tool is something the model calls to act. A resource is read-only data the
server publishes by URI, closer to a GET endpoint than a function call. The distinction is
about control. Your application decides to read a resource and put its contents into the
model's context. The model doesn't invoke it. The notes server exposes
a static resource, notes://all, and a templated one, notes://note/{key}. The example
lists and reads them, still with no LLM.
6. Prompts, reusable templates served by MCP
python examples/05_prompts.py
The third primitive. A prompt is a parameterized template the server owns and the user
picks, like the slash-commands in a chat app (/summarize). Why serve prompts over a
protocol instead of hard-coding them in the host? Because the server is the expert on its
own data. The team that runs the notes server can ship a good "summarize my notes" prompt
with the correct field names and the right tone, and improve it server-side without every
host re-implementing it. The host lists what's available and offers it to the user.
7. A real, many-tool server
python examples/06_multi_tool_server.py
So far, one tool at a time. Real servers expose a handful of related tools, and the client
discovers them all the same way. This connects to servers/toolbox.py,
which has calculator, search_notes, word_count, and save_note plus resources and a prompt,
and exercises several over one connection. Two things to notice. You didn't change the
client to get new tools; the server grew and tools/list returns more. And save_note has
a side effect, since it writes a file, which is exactly the kind of tool you gate behind
approval once a model is driving, as in §8 and §11.
8. Put an LLM in the loop
secrun python examples/07_llm_calls_mcp_tools.py # needs a key
The first example that costs money. Everything before it was offline. Now a model drives
the MCP tools. The host lists the server's tools, describes them to the model, and when the
model asks to call one, the host runs it over the protocol and feeds the result back. This
is the agent loop from the Agents deep dive with one change: the tools live in a separate
process behind MCP. And the model has no idea the tools came from an MCP server. To it they're
names, descriptions, and schemas. That invisibility is the reason MCP exists. The loop
lives in host/loop.py, and it carries over the agent-dive safety logic: a
max_steps ceiling, approval for side-effecting tools, and in-band error results so a
failing tool doesn't crash the host.
9. Transports, stdio against HTTP
# terminal 1: start the HTTP server and leave it running
python servers/calculator_http.py
# terminal 2:
python examples/08_http_transport.py
The stdio examples launched the server themselves as a subprocess. An HTTP server is
different. It's already running somewhere and you connect to it by URL. Same tools, same
tools/list and tools/call, with a network transport underneath. Rule of thumb: stdio
for local tools that ship with the host as a subprocess on your machine, streamable HTTP
for a shared service that several hosts connect to over the network.
Self-contained, routable requests
Run the example and inspect its raw HTTP call. Modern MCP requires
MCP-Protocol-Version, Mcp-Method, and, for a named primitive, Mcp-Name.
Gateways and rate limiters can route on those headers without parsing JSON. The
request's _meta carries client identity/capabilities. The response has no
Mcp-Session-Id: any replica can serve the next request without sticky routing
or a shared protocol-session store.
This doesn't ban application state. Make state explicit instead. A tool can mint a workflow or job handle and require the model to pass it back later. The handle is visible in the tool contract instead of hidden in transport state.
stateless_http=True still exists in SDK v2, but only changes how the server
supports pre-2026 legacy clients. It isn't the switch for modern traffic; modern
MCP is already self-contained.
Multi-round trips and cacheable catalogs
python examples/10_multi_round_trip.py # typed input_required loop
python examples/11_cacheable_catalogs.py # ttlMs/cacheScope and a visible hit
Removing the session also removes the back-channel used by old server-initiated
elicitation, sampling, and roots requests. Multi Round-Trip Requests (MRTR)
replace it: the server returns resultType: "input_required"; the client obtains
typed answers and retries the original call with inputResponses and opaque
requestState. The SDK's Resolve(...) dependency and high-level Client drive
that loop in example 10.
Discovery results now carry ttlMs and cacheScope. Clients can reuse
fresh tool/resource/prompt catalogs; private data stays in one authorization
partition, while genuinely identical public data may be shared. Example 11
makes the cache hit visible. The full migration matrix, including Tasks,
extensions, authorization hardening, and deprecated features, is in
PROTOCOL_2026.md.
10. Security: MCP + prompt injection
python examples/09_security.py # offline, no key
MCP is a trust decision. When your host connects to a server, that server's tool descriptions and resource contents flow straight into your model's context, and the model's tool calls get executed by your host. A server you didn't write is untrusted input, exactly like a web page in the Prompt Injection deep dive. This connects to servers/sneaky.py, a deliberately hostile server, and shows two attacks, a malicious tool description that tries to hijack the model and a tool result that smuggles instructions, along with the defenses: least privilege, human approval for side-effecting tools, and treating every server's text as untrusted. No LLM needed. You can see the malice in the raw data, which is the whole point.
11. Wiring your server into a real host
Here's what a standard protocol buys you. A server you wrote here works in real hosts
unchanged. To use servers/toolbox.py in Claude Desktop, add it to the
MCP config in claude_desktop_config.json:
{
"mcpServers": {
"toolbox": {
"command": "/absolute/path/to/.venv/bin/python",
"args": ["/absolute/path/to/mcp-deep-dive/servers/toolbox.py"]
}
}
}
In Claude Code, register it from the CLI:
claude mcp add toolbox -- /absolute/path/to/.venv/bin/python servers/toolbox.py
Restart the host and your tools, resources, and prompts appear, the same ones the client in
§4 saw. You can point a host at existing third-party servers (filesystem, GitHub,
databases) the same way. The mcp CLI, installed via mcp[cli], can inspect or run a
server during development with mcp dev servers/toolbox.py.
The capstone: assistant.py
Everything assembled into one runnable command. A multi-turn chat assistant whose every capability comes from an MCP server. The model holds no tools of its own. It discovers them over the protocol and calls them over the protocol. Swap the server and the assistant gains new powers without you touching the capstone.
# interactive chat (Ctrl-D or "quit" to exit):
secrun python hands_on/assistant.py
# one-shot question, then exit:
secrun python hands_on/assistant.py "What does the Plus plan cost for a year?"
# point at a different MCP server:
secrun python hands_on/assistant.py --server servers/notes.py
# auto-approve side-effecting tools (skip the prompt before save_note):
secrun python hands_on/assistant.py --yes
Read hands_on/assistant.py. It's the client (MCPClient), the
host loop (run_host), and a human-approval callback wired to a CLI. The whole repo in one
file. Suggested exercise: write your own small MCPServer with one tool you'd
actually use, and point the capstone at it with --server. When the assistant calls your
tool with no other change, MCP has clicked.
Where to go next
You've built servers, a client, and a host. What comes next is more of the same idea, at more scale.
- Tasks and extensions. Durable, pollable work and optional capabilities, without expanding the protocol core.
- MRTR resolvers. Typed elicitation or model assistance returned as
input_required, never pushed down a hidden back-channel. - OAuth and remote servers. Authenticating to hosted MCP servers you don't run.
- Real third-party servers. Wire the official filesystem, GitHub, or Postgres servers into Claude Desktop and feel the write-once-use-anywhere payoff.
- Streaming results and long-running tools. Progress notifications over the protocol.
- Building a host UI. The capstone is a REPL. A real host renders tools, resources, and prompt slash-commands as UI.
From teaching code to production
The teaching shortcuts here are exactly what you'd harden once an MCP host sits on a live path.
| This repo's teaching shortcut | In production |
|---|---|
| Connect to any server script | Vet and pin servers; treat unknown servers as untrusted code |
Approval is a terminal y/N prompt |
A real authorization layer with policy, audit, and per-tool scopes |
| Tool/resource text is read as-is | Guardrails on everything a server returns; it's untrusted input (§10) |
| stdio subprocess on your machine | Auth'd HTTP servers with TLS, rate limits, and least-privilege creds |
| The host loop prints a trace | Observability: structured traces of every tool call, with cost |
| One server, hard-coded | A registry of approved servers, versioned and health-checked |
The general ops machinery (observability, cost, reliability, caching, guardrails, prompt versioning, eval gates) gets built from scratch and wired into one running app in Production, #8 in the series, which runs offline on a mock provider.
File map
check_setup.py ← run first: Python, the mcp SDK, provider, key
README.md ← this guide
EXERCISES.md ← predict-then-run prompts, one per section
PROTOCOL_2026.md ← migration, deployment, auth, extensions, deprecations
servers/ ← MCP servers (the capability side)
calculator.py ← the minimal one-tool server (stdio)
calculator_http.py ← the same, over modern streamable HTTP (Section 9)
notes.py ← all THREE primitives: tools, resources, prompts
toolbox.py ← a realistic multi-tool server (used by the capstone)
sneaky.py ← a deliberately HOSTILE server (Section 10)
client/ ← the client side
mcp_client.py ← MCPClient: a small readable wrapper over the SDK
host/ ← the LLM side
providers.py ← neutral tool schema + run_turn for openai / claude
loop.py ← run_host: the agent loop, tools served over MCP
hands_on/
assistant.py ← capstone: an MCP-powered, multi-turn chat assistant
examples/
01_protocol.py ← the protocol & vocabulary in one page (offline)
02_first_server_and_client.py ← the raw SDK client, once (offline)
03_client_calls_tool.py ← connect → list → call, via MCPClient (offline)
04_resources.py ← read-only data by URI (offline)
05_prompts.py ← server-owned prompt templates (offline)
06_multi_tool_server.py ← many tools over one connection (offline)
07_llm_calls_mcp_tools.py ← a model drives the tools (needs a key)
08_http_transport.py ← connect to a server over HTTP (offline)
09_security.py ← a hostile server; attacks & defenses (offline)
10_multi_round_trip.py ← typed input_required / MRTR loop (offline)
11_cacheable_catalogs.py ← ttlMs/cacheScope + client cache hit (offline)
(workspace/ is created by the notes/toolbox servers' save_note tool and is
git-ignored.)
Troubleshooting
Run secrun python check_setup.py first; it catches most problems. Then, by symptom:
| What you see | What it means / the fix |
|---|---|
ModuleNotFoundError: mcp |
The SDK isn't installed. pip install -r requirements.txt (it pulls mcp[cli]). |
ModuleNotFoundError: mcp.server.fastmcp |
You're on the 1.x SDK, or following a 1.x tutorial. This repo targets 2.x, where the server class moved: mcp.server.fastmcp.FastMCP became mcp.server.mcpserver.MCPServer. pip install -r requirements.txt pins the right major version. |
Code calls ClientSession.initialize() |
That's a legacy protocol path. Use the high-level Client; its default auto mode selects 2026-07-28 and falls back only for an old server. |
Tool result attributes are missing (isError, inputSchema) |
2.x renamed response fields to snake_case: result.is_error, tool.input_schema. The JSON on the wire is unchanged and still camelCase, which is why examples/01_protocol.py still shows inputSchema. |
| A server example just hangs | A stdio server talks over stdin/stdout, so don't run servers/*.py directly expecting output; run the example (or the capstone), which launches the server for you. |
08_http_transport.py can't connect |
The HTTP server isn't up. Start python servers/calculator_http.py in another terminal first (it stays running on :8000). |
TypeError: MCPServer.__init__() got an unexpected keyword argument 'host' (or port, stateless_http) |
Another 1.x/2.x split. Transport options go to run(). Note that stateless_http only affects legacy clients; modern MCP is session-free already. |
PROVIDER=... needs ... in the environment |
Only Section 8 and the capstone need a key; the protocol, transport, MRTR, cache, and security examples are offline. Load the key from your keychain with secrun (see SECRETS.md), or stick to the offline examples. |
Import errors from host / client / servers |
Run from the repo root (python examples/03_...py), not from inside a subfolder; the examples add the repo root to sys.path. |
| Claude Desktop doesn't see my server | Use absolute paths to the venv's python and the script in the config, then fully restart the app. mcp dev servers/toolbox.py helps debug locally. |
SyntaxError / odd type errors on startup |
You're likely on Python 3.10 or older; this repo needs 3.11+. check_setup.py confirms your version. |
Still stuck? Every file is small and self-contained. Open it, read the docstring at the top, and run the matching example. client/mcp_client.py and host/loop.py are the whole story.
The series
This is one of the standalone, hands-on deep dives into building with LLM APIs. Eight core dives, plus the bonus ones listed below. Each one stands on its own, with its own setup, examples, and capstone, and they all share one house style. Provider-agnostic where it makes sense, built from scratch with no frameworks, offline-first examples, and a real capstone at the end. Do them in any order. This sequence builds naturally.
- OpenAI API: the API from zero
- Claude API: the same ideas, the Anthropic way
- Prompt Engineering: shape model behavior with better prompts
- RAG: answer questions over your own documents
- Evals: measure whether a change actually helps
- Agents: give a model tools and a loop so it can act
- Prompt Injection & Guardrails: attack and defend all of the above
- Production: operate one app end to end
Bonus dives, standalone and slotting in where they're most useful:
- Context Engineering: manage what's in the window, with memory, compaction, and assembly
- AI Data Engineering: the corpus behind the index, with versions, lineage, ACLs, and deletes
- Multimodal: images and audio as well as text
- Fine-tuning: teach a model new behavior by example
- MCP: serve tools, data, and prompts to any LLM over a standard protocol
- Local Models: run open-weight models on your own machine
- Agent Harnesses: build on the loop, adding hooks, permissions, sandboxing, and subagents
- Realtime Voice: low-latency speech-to-speech agents
- Observability: watch a running app over time, covering drift, quality, alerting, and the feedback loop
- Architecture: the seams between the components, each decision measured rather than asserted
- GenAI Security: treat the model as an untrusted principal, and put identity, supply chain, isolation, budgets, and release gates around it
- Inference Platform Engineering: turn finite GPU memory and a request queue into latency, throughput, and a fleet size you can defend
- Testing & Delivery: decide whether a build is fit to promote, using evidence, gates, staged rollout, and rollback
- Professional Tools: rebuild each hand-written piece with the tool professionals reach for, and measure both
And the whole series lands in one codebase in the capstone: a codebase Q&A tool built step by step, one tag per dive.
MCP is a bonus dive in the series. It slots most naturally right after Agents (#6), since Section 8 here is that dive's loop with tools served over MCP, and its security section (§10) builds on Prompt Injection & Guardrails (#7).