Core path - 1 of 8
OpenAI API: A Guided Deep Dive
A hands-on playground for learning the OpenAI API from zero. You'll build a real CLI
tool that answers questions about your code. The main path teaches Chat Completions,
roles, sampling controls, token counting, and cost. A focused track under responses/
then teaches OpenAI's Responses API without turning it into a separate course.
Walk through this repo rather than reading it. Each section ends with something to run. Do the running. That's where the learning is. And once a section clicks, EXERCISES.md has a quick predict-then-run prompt for it. Committing to an answer before you run is what makes it stick.
0. The one big idea
The OpenAI API is simpler than it looks.
You send a list of messages. You get back a message.
That's it. Everything else, the roles, the knobs, the token math, is detail on top of that one request and response. Hold onto that and nothing below will feel complicated.
1. Setup (5 minutes)
# 1. Create an isolated Python environment
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 2. Install dependencies
pip install -r requirements.txt
# 3. Set up your API key (it does NOT go in .env)
cp .env.example .env # optional; holds no secrets
# Store your key in your OS keychain and run scripts with `secrun`: 2-minute
# setup in ../docs/SECRETS.md. Get a key: https://platform.openai.com/api-keys
# 4. Confirm everything is wired up correctly (makes no API call, costs nothing)
secrun python check_setup.py
check_setup.py is your first stop if anything goes wrong. It checks your Python
version, your installed packages, and your key, and tells you exactly what to fix. Green
across the board means you're ready for Section 2.
You can learn a lot before spending a cent. The token-counting and cost-estimation parts (Sections 5 and 6) run entirely offline. Skip ahead to them if you don't have a key yet.
2. Your first request
secrun python examples/01_basic_chat.py
Open examples/01_basic_chat.py and read it. It's tiny. The shape of every call you'll ever make is right there:
response = client.chat.completions.create(
model="gpt-6-luna",
reasoning_effort="none",
messages=[{"role": "user", "content": "In one sentence, what is an API?"}],
)
print(response.choices[0].message.content)
Five things to internalize:
| Thing | What it is |
|---|---|
model |
Which model answers. gpt-6-luna is the cheap, fast default. |
reasoning_effort |
How long the model thinks before it answers. Luna defaults to "medium", and at anything above "none" it rejects temperature and function tools here. The examples send "none"; Reasoning models turns it back up. |
messages |
A list of messages, your half of the conversation. |
response.choices[0].message.content |
The model's reply text. |
response.usage |
Exactly how many tokens you were billed for. |
3. The three roles
A conversation is a transcript, and every line carries a role.
systemholds standing instructions. Persona, rules, tone. You set it once and it steers everything, which makes it the strongest lever you have.useris what the human says.assistantis what the model said. You re-send past assistant messages to give the model memory. The API itself is stateless, so the conversation exists only in the list you send each time.
secrun python examples/02_roles.py
Experiment: open examples/02_roles.py, change the system
message to "You are a grumpy pirate.", and rerun. Same question, completely different
voice. That's the system role doing its job.
4. The knobs that shape a response
These four parameters control how the model answers. Each one has its own runnable example.
temperature, or how bold the word choices are
0.0 is focused and repeatable, 0.7 is balanced (the default), 1.5+ goes wild. For
code and facts, go low. For brainstorming, go high.
secrun python examples/03_temperature.py
max_completion_tokens, a hard cap on the answer's length
This caps output tokens and not input. When the budget runs out the model gets cut off,
possibly mid-sentence. Watch finish_reason. "length" means truncated and "stop"
means it finished on its own.
secrun python examples/04_max_tokens.py
top_p, or how many options the model may consider
Nucleus sampling. 0.1 allows only the most obvious tokens, 1.0 allows everything and
is the default. Tune temperature or top_p, never both. They interact in confusing ways.
secrun python examples/05_top_p.py
stop, to halt generation at a marker
Up to four strings, each of which ends generation the moment it would appear. The stop text itself isn't included. Good for cutting lists short or stopping at a delimiter.
secrun python examples/06_stop_sequences.py
Quick reference:
| Knob | Range | Raise it to... | Default |
|---|---|---|---|
temperature |
0.0–2.0 | get more variety/creativity | 1.0 |
top_p |
0.0–1.0 | widen the pool of candidate words | 1.0 |
max_completion_tokens |
≥1 | allow a longer answer | model max |
stop |
up to 4 strings | end at a specific marker | none |
5. Tokens: what you actually pay for
Models don't read characters or words. They read tokens, which are chunks of text, often word fragments. Rough rule: 1 token is about 4 English characters, roughly three quarters of a word. Rough isn't good enough for budgeting, so we count exactly with tiktoken, locally, with no API call.
python utils/tokens.py # see a sentence broken into tokens
Counting matters for three reasons.
- Cost. You're billed per token, as the next section covers.
- Limits. Every model has a maximum context window covering input and output. Overflow it and the request fails.
- Intuition. Watching the count change as you edit a prompt teaches you how the model reads your text.
See utils/tokens.py for count_tokens() (a raw string) and
count_message_tokens() (a full chat list, including the small per-message
overhead the API adds).
6. Cost estimation
OpenAI charges separately for input and output tokens, and output usually costs several
times more. utils/pricing.py holds a small price table and an
estimate_cost() helper.
python examples/07_token_counting.py # tokens -> dollars, across models, offline
The sample output shows the same request costing wildly different amounts depending on the model, which is why choosing the right model is part of prompt engineering.
Prices change. The table in
utils/pricing.pyis a snapshot, so always confirm at https://platform.openai.com/docs/pricing.
7. The first capstone: ask.py
Everything above comes together in one tool. Ask a question about a code file, see the token count and estimated cost before you spend anything, get the answer, then see the actual usage and cost after.
# See the cost first: no API call, no key needed:
secrun python hands_on/ask.py snippets/buggy.py "Is there a bug here?" --dry-run
# For real (needs your key; run under secrun):
secrun python hands_on/ask.py snippets/buggy.py "Is there a bug here?"
# Now turn the knobs you just learned:
secrun python hands_on/ask.py snippets/buggy.py "Rewrite this cleanly" --temperature 0
secrun python hands_on/ask.py snippets/buggy.py "List the issues" --max-tokens 200 --stop "4."
secrun python hands_on/ask.py snippets/buggy.py "Explain this" --model gpt-4o
Run secrun python hands_on/ask.py --help to see every knob explained inline. Read the
source in hands_on/ask.py. It's commented as a tutorial, especially
build_messages(), which assembles the request, and the usage and cost reporting at the
end.
Suggested exercise: point hands_on/ask.py at your own code, try the same question
at --temperature 0 vs --temperature 1.2, and watch both the answers and the
cost change.
8. Beyond the basics
With the core down, here are the most useful next capabilities. Each one is a runnable example in the same numbered style, and every one is still a variation on "send messages, get a message."
Streaming, to get the answer as it's typed
With stream=True the response arrives in small chunks as it's generated, so the
user sees text appear immediately instead of waiting for the whole thing.
secrun python examples/08_streaming.py
Structured outputs, to make the model return real JSON
Force the reply to be valid JSON, or even match an exact schema you define
(response_format with "strict": True). The end of fragile "please reply in
JSON" prompting.
secrun python examples/09_structured_outputs.py
Function and tool calling, to let the model use your code
You describe functions. The model decides when to call one and with what arguments. You run it and feed the result back. This is how a model gets to actually do things, like query a database or hit an API.
secrun python examples/10_function_calling.py
Embeddings, turning text into vectors for search and similarity
A different endpoint, client.embeddings.create, converts text into numbers that
capture meaning. It's what semantic search and retrieval (RAG) are built on. The
example ranks sentences by similarity to a query, including ones that share no words
with it at all.
secrun python examples/11_embeddings.py
Multi-turn conversations, because the API has no memory
Each request is stateless and the model remembers nothing. A chatbot that "remembers" is
you re-sending the whole messages list every turn, appending each new user and
assistant message. The example is a tiny REPL that grows that list.
secrun python examples/12_conversation.py
Error handling and retries, for surviving the real world
The network blips, you hit a rate limit, a model name has a typo. The SDK already
retries transient failures (429/5xx/connection) with backoff; your job is to tune
timeout/max_retries and catch the typed exceptions so "fix your request"
errors are handled differently from "try again later" ones.
secrun python examples/13_error_handling.py
Pydantic validation, for typed and validated responses
Instead of a hand-written JSON Schema plus json.loads into an untyped dict, define
your shape as a Pydantic model and pass it as response_format. The SDK sends the
schema, constrains the model, and hands back a validated instance with typed attributes,
enforced constraints, and editor autocomplete.
secrun python examples/14_pydantic_validation.py
Formatting output as Markdown, tables, and code blocks
Models answer in Markdown, and dumped raw to a terminal that is a mess of literal
**asterisks**. The rich library renders Markdown, syntax-highlighted code, and real
tables in the terminal. That's the difference between output you skim and output you
squint at.
secrun python examples/15_rich_output.py
Server-Sent Events (SSE), the protocol under streaming
Every streaming AI response travels over SSE, a plain HTTP response that stays open and
drips data: <json> lines until the server is done. The SDK hides the parsing, but you
need the wire format the moment you build a backend that forwards tokens to a browser. This example shows raw events, per-token
timing, and partial response accumulation.
secrun python examples/16_sse.py
Local models, the same client with a different base_url
Nothing here is tied to OpenAI's servers. Local runtimes like Ollama and llama.cpp
expose an OpenAI-compatible endpoint, so the same openai SDK talks to a model on your
own machine. Change base_url, pass any non-empty api_key, and everything else works
unchanged: roles, knobs, streaming, usage. You get privacy, no per-token bill, and
offline use, in exchange for running the server yourself. The example prints a useful
message instead of crashing if no local server is up.
secrun python examples/17_local_serving.py # needs a local runtime; prints how to start one
Vision, sending images alongside text
Multimodal models accept images in the same message. The user content becomes a list
of parts, text plus image_url, where the image is either a URL or a local file inlined
as a base64 data: URI. Images are billed as tokens, scaled by pixel size. Set detail
yourself: on luna the default ("auto") doesn't shrink large images, so a 3000x3000
photo costs about 3.5 times what "high" charges. The example reads a public sample
image, or your own local file.
secrun python examples/18_vision.py # or: secrun python examples/18_vision.py my_image.png
Reasoning models, which think first and answer second
Reasoning models generate hidden reasoning tokens before answering, which makes them far
better on math, logic, and coding. You drop temperature and steer with
reasoning_effort instead. usage reports the hidden thinking you still pay for.
Reasoning is no longer a separate family: the o-series is being switched off (o1, o1-pro,
and o4-mini on 2026-10-23), and the dial now lives on the mainline tiers. The example runs
on gpt-6-luna, the same model as every other lesson, with the dial at "high" instead
of "none".
secrun python examples/19_reasoning.py
The Batch API, half price for non-urgent work
For work that isn't interactive, like classifying 10k reviews or summarizing a backlog, upload a JSONL of requests and get results within 24 hours at 50% off. The example builds a tiny batch, submits it, and shows how to poll for and fetch results.
secrun python examples/20_batch_api.py
Prompt caching, so you don't re-pay for a repeated prefix
On OpenAI this happens automatically. A prompt's long, identical prefix gets cached and
re-billed at a discount on later calls. The one rule is structural. Put the constant
part first, meaning the big system prompt, the tool catalog, the document, and the
variable question last. The example shows cached_tokens kicking in.
secrun python examples/21_prompt_caching.py
Async and concurrency, for many requests at once
A single call is mostly idle waiting on the network, so independent prompts should
run concurrently. AsyncOpenAI + asyncio.gather + a Semaphore (bounded
concurrency) finishes a batch in roughly the time of the slowest call, while staying
under your rate limit. The example times sequential vs. concurrent.
secrun python examples/22_async_concurrency.py
Moderation, a free safety filter
A dedicated free classifier that flags hateful, violent, sexual, and self-harm content with per-category scores. Moderate user input on the way in and model output on the way out, and refuse or redact whatever it flags.
secrun python examples/23_moderation.py
Logprobs, or how confident was the model?
With logprobs=True / top_logprobs=k the API returns the probability of each
chosen token and the alternatives it weighed. Turn that into a 0-1 confidence, for
calibrated classification, flagging shaky answers, or debugging.
secrun python examples/24_logprobs.py
Seed and reproducibility, pinning down a random model
A fixed seed makes the same inputs reproduce the same output, which you want for
tests, caching, and reproducible evals, even with real randomness in play. The example
runs at temperature=0.9 on purpose so you can watch the seed do the work. The same
seed twice gives identical output, no seed twice gives different output. At
temperature=0 the model is already deterministic, so the seed would have nothing
visible to do; in production you'd combine both for the strongest guarantee. Either
way this is best-effort rather than a guarantee, so watch system_fingerprint, which
signals a backend change that can break determinism.
secrun python examples/25_seed_determinism.py
Responses API mini-track
Everything above uses chat.completions.create. OpenAI recommends
responses.create for new OpenAI projects, while Chat Completions remains supported.
The request changes from messages to input and instructions. The response changes
more: output is a list of typed items, not a list of text choices. A response may
contain messages, reasoning items, custom tool calls, or hosted tool calls. Use
output_text when you only need its combined text. Inspect output when item type or
tool metadata matters.
These eight are terser than the numbered examples above, and deliberately so: they assume you already have the reflexes from Sections 2 to 8 and spend their words on what's different rather than re-teaching what a role or a token is. If a script here feels dense, the matching example above is the gentler version of the same idea.
Work through these in order:
| Example | What it makes observable |
|---|---|
| 01 request and items | The request shape, response metadata, usage, and typed output items. |
| 02 conversation state | An ID replaces transcript upload, but earlier context is still billed. Instructions must be repeated. |
| 03 streaming events | Text deltas share the stream with lifecycle and item events. |
| 04 custom tool loop | The model proposes a call. Your application validates and executes it. |
| 05 hosted web search | OpenAI executes a required hosted tool and returns call evidence and sources. |
| 06 background responses | A response ID supports later retrieval, bounded polling, and cancellation. |
| 07 structured outputs | response_format becomes text.format. A schema constrains shape, not completion. |
| 08 conversation object | A durable object collects items across jobs, and refuses to combine with a response chain. |
secrun python responses/01_request_and_items.py
secrun python responses/02_conversation_state.py
secrun python responses/03_streaming_events.py
secrun python responses/04_custom_tool_loop.py
secrun python responses/05_hosted_web_search.py
secrun python responses/06_background_responses.py start
secrun python responses/07_structured_outputs.py
secrun python responses/08_conversation_object.py
Two state mechanisms deserve separate names. previous_response_id chains one response
to the next. A Conversation is a durable object that can collect items across sessions,
jobs, or devices. The two can't be supplied on the same request. Neither is a token
discount. The model still receives the usable context, and OpenAI bills those earlier
input tokens again. Response objects are stored for 30 days by default. Conversation
objects and their items don't use that 30-day expiry, so a Conversation is storage
you own and eventually have to delete. Example 08 builds one, reads its items back,
provokes the 400 you get for sending both mechanisms, and cleans up after itself.
Read the official
conversation-state guide
before choosing either mechanism for user data.
Hosted tools remove your client-side execution loop. They also move execution and data handling into OpenAI's service. Choose the tool on the server, narrow its scope, inspect the returned call items, and treat web results as untrusted input. For functions in your own process, keep the allowlist and argument checks shown in example 04.
Background mode solves connection lifetime, not job ownership. Store the response ID,
poll only while its status is queued or in_progress, stop on every other status, and
put a deadline around the polling loop. Example 06 also makes the retention tradeoff
plain: store=False still requires temporary storage while asynchronous work runs. See
the official background-mode guide
for the current retention rules.
There's still a portability cost. /v1/chat/completions is implemented by Ollama, LM
Studio, vLLM, LiteLLM, and many hosted providers. That's why
example 17 can switch to a local model by changing
base_url. The Responses API is OpenAI's endpoint. Use it when its item model, hosted
tools, or state handling saves real application work. Use Chat Completions when provider
portability matters.
Official references: Responses API, streaming, function calling, and web search.
9. The second capstone: extract.py
Where ask.py returns prose, extract.py returns data. Point it at messy free-form
text and it pulls out a clean, typed, validated structure, then shows it as a Markdown
summary and a real table. This is where examples 14 (Pydantic) and 15 (rich) finally do
something useful on a realistic task.
# See tokens + cost first: no API call:
secrun python hands_on/extract.py snippets/meeting_notes.txt --dry-run
# Extract action items (owner, due date, inferred priority) into a table:
secrun python hands_on/extract.py snippets/meeting_notes.txt
# Want the raw validated JSON instead? (e.g. to pipe into another tool)
secrun python hands_on/extract.py snippets/meeting_notes.txt --json
Read the source in hands_on/extract.py. The Extraction and
ActionItem Pydantic models are the schema the model must follow, and render() is the
rich table. Suggested exercise: point it at your own meeting notes or an email, or
change the models to extract something else entirely, like contacts or invoice line
items. The prompt barely changes.
10. The third capstone: streaming_server.py
Where ask.py and extract.py are CLI tools, streaming_server.py is a web service. A
FastAPI backend that streams AI responses to a browser over SSE, and it shows three
production concerns: token-by-token forwarding, client disconnect detection, and error
recovery with retries.
# Start the server (auto-reloads on file saves):
uvicorn hands_on.streaming_server:app --reload
# Then open: http://localhost:8000
Open the browser's Network tab and click the /stream request to see the raw
text/event-stream response. Close the tab mid-stream to watch the server log "client
disconnected" and stop the AI call. Read the source in
hands_on/streaming_server.py. The three-phase generator
_stream_tokens is the pattern every production streaming endpoint follows.
Suggested exercise: point the server at gpt-4o and ask it something long,
then close the browser tab mid-response. Notice in the server logs that generation
stops immediately, with no wasted tokens.
11. The fourth capstone: rag.py
The embeddings example in Section 8 ranked sentences by similarity. rag.py puts that
to work. It answers questions over a small knowledge base by retrieving the most
relevant facts and pasting them into the prompt, which is the smallest thing you could
still call retrieval-augmented generation (RAG). No vector database, no framework.
Just the embeddings and chat calls you already know, wired together from scratch.
Hold onto one idea. A model can only answer from what's in its context window, and RAG decides what to put there.
# Answer the built-in demo question from the knowledge base:
secrun python hands_on/rag.py
# Ask your own:
secrun python hands_on/rag.py "Can I get a refund?"
# The contrast that makes the point: the same question with NO retrieved context:
secrun python hands_on/rag.py "How long are deleted notes kept?" --no-rag
# See exactly what gets retrieved and what prompt gets sent:
secrun python hands_on/rag.py "What plans are there?" -k 5 --show-prompt
The knowledge base describes a made-up app, so the model can't fall back on training
and a correct answer can only come from retrieval. Run it with --no-rag and watch the
model guess or refuse. That contrast is the lesson. The embeddings call and the chat
call use the same OPENAI_API_KEY.
Read the source in hands_on/rag.py. retrieve() is the whole embed,
score, rank loop, and build_user_message() is the entire augment step. RAG is mostly
good string assembly. Suggested exercise: add a fact to
KNOWLEDGE_BASE, then ask a question only that fact can answer.
Where to go next
You've now covered the essentials, the common extensions, and four capstone projects. Further on:
- RAG at scale. You just built a minimal version in
rag.py, which re-embeds a handful of facts on every run. Real systems embed once into a vector database, chunk long documents, and add reranking and evaluation. Enough moving parts to deserve a deep dive of its own. - The context window. What happens as conversations get long, and smarter ways to manage history than the simple trim in example 12, like summarizing old turns and sliding windows. That's a whole dive: Context Engineering.
- Vision and audio. Passing images to multimodal models, and speech-to-text.
- Streaming and tools together. The pattern most production assistants use, built hands-on in the Agents dive, where streaming happens inside the tool loop.
Every one of these builds on the request, context, output, and usage concepts you met at the start. The wire format changes between Chat Completions and Responses. Those four concerns don't.
Troubleshooting
Hit a snag? Run secrun python check_setup.py first. It catches most problems. The
rest, sorted by the error you see:
| What you see | What it means / the fix |
|---|---|
ModuleNotFoundError: No module named 'openai' |
Dependencies aren't installed (or your venv isn't active). Run source .venv/bin/activate then pip install -r requirements.txt. |
Set OPENAI_API_KEY ... on every script |
No key found. Store it in your keychain and run the script under secrun. See SECRETS.md. (The offline token/cost parts in Sections 5–6 still run without a key.) |
AuthenticationError / 401 |
The key is present but wrong: expired, revoked, or a typo. Make a fresh one at the API keys page. |
RateLimitError / 429 |
Too many requests, or you're out of credit. Wait a moment, or check your billing/usage in the dashboard. |
NotFoundError / 404 about the model |
A model name was mistyped or your account can't access it. The examples use widely-available IDs; if you changed one, check it against the models list. |
SyntaxError or odd type errors on startup |
You're likely on Python 3.10 or older. This repo needs 3.11+; check_setup.py will confirm your version. |
| It "hangs" with no output | Some examples stream, others wait for the full reply before printing. Give it a few seconds; for streaming examples you'll see text appear word by word. |
Still stuck? Every example is small and self-contained. Open the file, read the docstring at the top, and run it directly. The error message almost always points at the line.
From teaching code to production
Every example here takes shortcuts that are perfect for learning and wrong for a real deployment. Here's the map from each shortcut to what production uses.
| This repo's teaching shortcut | In production |
|---|---|
The answer goes to print() |
One structured trace per request (id, timing, tokens) you can search after the fact |
estimate_cost() just prints a number |
An enforced budget that refuses the call before it overspends |
A bare client.chat.completions.create(...) |
The call wrapped in retries + backoff and a circuit breaker for 429s/503s/timeouts |
| Every call hits the API | A response cache so repeat questions cost nothing |
| Model id and system prompt are string literals in the script | Versioned prompts/models behind config, promoted only past an eval gate |
| You trust whatever the model returns | Input/output guardrails on the request path |
All seven concerns (observability, cost, reliability, caching, guardrails, prompt versioning, and eval gates) get built from scratch and wired into one running app in Production, which is #8 in the series. It runs offline on a mock provider, so you can see the whole ops machinery with no key and no cost.
File map
check_setup.py ← run first: verifies Python, packages, and your key
EXERCISES.md ← active-recall prompts, one per README section
hands_on/
ask.py ← capstone CLI: ask a question about a code file
extract.py ← capstone CLI: extract validated data from free text
streaming_server.py ← capstone server: stream AI responses over SSE
rag.py ← capstone CLI: answer questions over a knowledge base (RAG)
static/index.html ← browser UI for the streaming server
utils/
tokens.py ← tiktoken-based token counting
pricing.py ← price table + cost estimation
snippets/buggy.py ← a sample file to ask questions about
snippets/meeting_notes.txt ← sample free-form text for extract.py
examples/
01_basic_chat.py ← the minimal request
02_roles.py ← system / user / assistant
03_temperature.py ← randomness
04_max_tokens.py ← length cap + finish_reason
05_top_p.py ← nucleus sampling
06_stop_sequences.py ← halting generation
07_token_counting.py ← tokens & cost, fully offline
08_streaming.py ← stream the answer as it's generated
09_structured_outputs.py ← guaranteed JSON / schema-conformant output
10_function_calling.py ← let the model call your functions
11_embeddings.py ← vectors & semantic similarity
12_conversation.py ← multi-turn chat & the stateless API
13_error_handling.py ← timeouts, retries & typed exceptions
14_pydantic_validation.py ← typed, validated responses via Pydantic
15_rich_output.py ← Markdown, tables & code blocks in the terminal
16_sse.py ← SSE protocol: raw events, timing, partial accumulation
17_local_serving.py ← same client, local model via base_url (Ollama/llama.cpp)
18_vision.py ← send an image (URL or local base64) alongside text
19_reasoning.py ← reasoning models: reasoning_effort, hidden tokens
20_batch_api.py ← submit many requests at 50% off, results within 24h
21_prompt_caching.py ← automatic prefix caching; structure prompts to hit it
22_async_concurrency.py ← AsyncOpenAI + asyncio.gather + a Semaphore (throughput)
23_moderation.py ← the free safety classifier (flags + per-category scores)
24_logprobs.py ← token probabilities -> confidence & calibrated classification
25_seed_determinism.py ← seed pins down randomness (best-effort reproducibility)
responses/
01_request_and_items.py ← request shape, output items, metadata & usage
02_conversation_state.py ← response chains, instructions & billing
03_streaming_events.py ← typed stream events, text deltas & terminal status
04_custom_tool_loop.py ← validate and execute a local function request
05_hosted_web_search.py ← require hosted search and inspect its sources
06_background_responses.py ← start, retrieve, poll & cancel asynchronous work
07_structured_outputs.py ← text.format, parse, and what truncation does to both
08_conversation_object.py ← durable server-side items vs a response chain
Footnote: quieting Pylance/type-checker noise
Two patterns trip the type checker repeatedly with the OpenAI SDK. Pre-empt them and new files stay clean:
- Assigning
messages/toolsto a variable (rather than passing the literal straight intocreate()) makes Pylance infer a too-narrow type likelist[dict[str, str]]. Annotate with the SDK's own param types; you also get key autocomplete:pythonfrom openai.types.chat import ChatCompletionMessageParam, ChatCompletionToolParam messages: list[ChatCompletionMessageParam] = [...] tools: list[ChatCompletionToolParam] = [...] .contentis typedstr | None(it'sNonewhen the model returns only tool calls).print()and f-strings accept it, butjson.loads()and+concatenation don't, so guard with... or ""/... or "{}".
The repo's .vscode/settings.json also sets
python.analysis.typeCheckingMode to basic, which keeps the useful checks
(undefined names, bad attrs/args) while dropping the strict dict-vs-TypedDict
complaints.
The series
This is one of the standalone, hands-on deep dives into building with LLM APIs. Eight core dives, plus the bonus ones listed below. Each one stands on its own, with its own setup, examples, and capstone, and they all share one house style. Provider-agnostic, built from scratch with no frameworks, offline-first examples, and a real capstone at the end. Do them in any order. This sequence builds naturally.
- OpenAI API: the API from zero
- Claude API: the same ideas, the Anthropic way
- Prompt Engineering: shape model behavior with better prompts, using zero-shot and few-shot, chain-of-thought, and roles
- RAG: answer questions over your own documents
- Evals: measure whether a change actually helps
- Agents: give a model tools and a loop so it can act
- Prompt Injection & Guardrails: attack and defend all of the above
- Production: operate one app end to end, across observability, cost, reliability, caching, guardrails, prompt versioning, and eval gates
Bonus dives, standalone and slotting in where they're most useful:
- Context Engineering: manage what's in the window, with memory, compaction, and assembly
- AI Data Engineering: the corpus behind the index, with versions, lineage, ACLs, and deletes
- Multimodal: images and audio as well as text
- Fine-tuning: teach a model new behavior by example
- MCP: serve tools, data, and prompts to any LLM over a standard protocol
- Local Models: run open-weight models on your own machine
- Agent Harnesses: build on the loop, adding hooks, permissions, sandboxing, and subagents
- Realtime Voice: low-latency speech-to-speech agents
- Observability: watch a running app over time, covering drift, quality, alerting, and the feedback loop
- Architecture: the seams between the components, each decision measured rather than asserted
- GenAI Security: treat the model as an untrusted principal, and put identity, supply chain, isolation, budgets, and release gates around it
- Inference Platform Engineering: turn finite GPU memory and a request queue into latency, throughput, and a fleet size you can defend
- Testing & Delivery: decide whether a build is fit to promote, using evidence, gates, staged rollout, and rollback
- Professional Tools: rebuild each hand-written piece with the tool professionals reach for, and measure both
And the whole series lands in one codebase in the capstone: a codebase Q&A tool built step by step, one tag per dive.
You're here: #1, OpenAI API.