Chapter 5 · Python for AI Work
Advanced Python for AI Apps
- Page 18 of 23
- 6 min read
You can now call a model. Real AI applications need a few more tools: type hints to keep larger code correct, validation because a model's output cannot be trusted blindly, async to make many API calls at once, streaming so users see answers as they are written, and logging to know what happened in production.
Type hints
Type hints write down what type each value should be. Python itself does not enforce them when the program runs, but your editor and checkers like mypy or pyright use them to catch mistakes before you run anything — and they document the code for the next reader:
from dataclasses import dataclass
@dataclass
class Document:
title: str
text: str
tags: list[str]
def search(docs: list[Document], word: str, limit: int = 3) -> list[str]:
"""Titles of documents containing `word`, at most `limit` of them."""
hits = [d.title for d in docs if word.lower() in d.text.lower()]
return hits[:limit]
def first_tag(doc: Document) -> str | None:
return doc.tags[0] if doc.tags else None
docs = [
Document("Intro", "Python is a language for AI.", ["basics"]),
Document("Pricing", "Plans start at 10 dollars.", []),
]
print(search(docs, "python"))
print(first_tag(docs[1]))['Intro']
Nonelist[str]is a list of strings;dict[str, int]maps strings to ints.str | Nonemeans "a string, or None" — use it for anything that may be missing.-> list[str]after the brackets describes what the function returns.
Never trust model output: validate it
Ask a model for JSON and you will usually get valid JSON with the fields you asked for. Usually is not always. Treat its reply like user input: parse it, check every field, and reject anything unexpected:
import json
from dataclasses import dataclass
@dataclass
class ReviewAnalysis:
sentiment: str
score: float
topics: list[str]
ALLOWED = {"positive", "negative", "neutral"}
def parse_analysis(raw: str) -> ReviewAnalysis:
"""Turn a model's JSON reply into a checked Python object, or raise ValueError."""
data = json.loads(raw)
if data.get("sentiment") not in ALLOWED:
raise ValueError(f"unexpected sentiment: {data.get('sentiment')!r}")
score = float(data["score"])
if not 0 <= score <= 1:
raise ValueError(f"score out of range: {score}")
return ReviewAnalysis(data["sentiment"], score, list(data.get("topics", [])))
good = '{"sentiment": "negative", "score": 0.12, "topics": ["delivery", "packaging"]}'
print(parse_analysis(good))
for bad in ['{"sentiment": "angry", "score": 0.3}', '{"sentiment": "positive", "score": 7}', 'Sure! Here is the JSON']:
try:
parse_analysis(bad)
except (ValueError, KeyError) as error:
print("Rejected:", type(error).__name__, "-", error)ReviewAnalysis(sentiment='negative', score=0.12, topics=['delivery', 'packaging'])
Rejected: ValueError - unexpected sentiment: 'angry'
Rejected: ValueError - score out of range: 7.0
Rejected: JSONDecodeError - Expecting value: line 1 column 1 (char 0)One except catches the broken JSON too, because json.JSONDecodeError is a kind of ValueError.
Providers also offer structured outputs (you pass a JSON schema and the model is constrained to follow it), and libraries such as Pydantic turn this checking into a class definition. They reduce failures but do not remove your responsibility to check values that matter — a well-formed answer can still be wrong.
async: many requests at once
An API call spends almost all its time waiting for the network. Waiting for three calls one after another takes three times as long as waiting for all three at the same time. async/await lets one program wait on many things together. Here the network is simulated with asyncio.sleep:
import asyncio
import time
async def fake_api_call(question: str) -> str:
await asyncio.sleep(1) # stands in for waiting on the network
return f"answer to {question!r}"
async def main():
questions = ["What is NumPy?", "What is pandas?", "What is an API?"]
start = time.perf_counter()
for q in questions: # one after another
await fake_api_call(q)
print(f"One by one: {time.perf_counter() - start:.0f} s")
start = time.perf_counter()
answers = await asyncio.gather(*(fake_api_call(q) for q in questions)) # all at once
print(f"Together: {time.perf_counter() - start:.0f} s")
print(answers[0])
asyncio.run(main())One by one: 3 s
Together: 1 s
answer to 'What is NumPy?'async defdefines a coroutine, a function that can pause while it waits.awaitpauses until the result is ready, letting other work run meanwhile.asyncio.gather()runs several coroutines together and collects their results in order.
With a real SDK it looks the same — use the async client and limit how many requests run at once so you do not hit rate limits:
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI()
limit = asyncio.Semaphore(5) # at most 5 requests in flight at once
async def summarise(text: str) -> str:
async with limit:
response = await client.responses.create(
model="gpt-5.5", input=f"Summarise in five words: {text}"
)
return response.output_text
async def main():
reviews = ["The delivery was late but the product is great.",
"Terrible battery life, would not buy again.",
"Exactly as described, very happy with it."]
summaries = await asyncio.gather(*(summarise(r) for r in reviews))
for s in summaries:
print("-", s)
asyncio.run(main())Example output:
- Late delivery, great product.
- Poor battery, would not rebuy.
- Exactly as described, satisfied.Streaming: show the answer as it is written
A long answer can take many seconds to finish. With stream=True the API sends small pieces as soon as they are generated, and you print each one immediately — this is how chat apps "type". It is the generator idea from page 8:
from openai import OpenAI
client = OpenAI()
stream = client.responses.create(
model="gpt-5.5",
input="Give me one tip for learning Python.",
stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True) # show each piece as it arrives
print()Example output (appearing a few words at a time):
Write a little code every day, and read the error messages carefully.Logging instead of print
print is fine while learning. In a real service, use the logging module: every message gets a level (DEBUG, INFO, WARNING, ERROR), you choose which levels to show without changing code, and logs can go to files or monitoring tools:
import logging
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(name)s: %(message)s")
log = logging.getLogger("review-bot")
def classify(text: str) -> str:
log.debug("classifying %d characters", len(text)) # hidden at INFO level
label = "negative" if "bad" in text.lower() else "positive"
log.info("classified as %s", label)
if len(text) < 5:
log.warning("very short input: %r", text)
return label
classify("Bad")
classify("A wonderful phone")INFO review-bot: classified as negative
WARNING review-bot: very short input: 'Bad'
INFO review-bot: classified as positiveLog what helps you debug — model name, token counts, timings, errors — but never log API keys or users' private data.
Try it yourself
- Add type hints to three functions you wrote earlier in this tutorial.
- Extend
parse_analysissotopicsmust be a list of at most 5 strings. - Change the async example to process 10 questions with at most 3 running at once, and measure the time.