Chapter 3 · Writing Real Programs
Comprehensions, Iterators, Generators and Decorators
- Page 8 of 23
- 7 min read
You have written loops that build a new list item by item. That pattern is so common that Python has a short, readable form for it: the comprehension. You will see comprehensions in almost every piece of data or AI code, so it is worth learning to read and write them fluently. Later on this page you will also see what a for loop really does underneath (iterators), and how to wrap extra behaviour around any function (decorators).
List comprehensions
The form is [expression for item in collection] — "make a list of this for each item":
prices = [120, 450, 80, 990]
# The loop way
with_vat = []
for p in prices:
with_vat.append(p * 1.15)
# The comprehension way: same result, one line
with_vat = [p * 1.15 for p in prices]
print([round(p) for p in with_vat])[138, 518, 92, 1138]Filtering with if
Add if condition at the end to keep only some items. Cleaning text data like this is a daily task before sending it to a model:
reviews = ["Great!", "", "Terrible support", " ", "Loved it"]
# keep only non-empty reviews, cleaned
clean = [r.strip() for r in reviews if r.strip()]
print(clean)
lengths = [len(r) for r in clean]
print(lengths)['Great!', 'Terrible support', 'Loved it']
[6, 16, 8]Read it aloud: "r.strip() for each r in reviews, if r.strip() is not empty". Remember from page 4 that empty text counts as False.
Dictionary and set comprehensions
Curly braces with a key: value pair make a dictionary; curly braces with a single expression make a set:
words = ["python", "ai", "data", "ai", "python", "ai"]
counts = {w: words.count(w) for w in set(words)}
print(dict(sorted(counts.items())))
prices_usd = {"basic": 10, "pro": 25}
prices_bdt = {plan: usd * 122 for plan, usd in prices_usd.items()}
print(prices_bdt)
lengths = {len(w) for w in words} # a set comprehension
print(sorted(lengths)){'ai': 3, 'data': 1, 'python': 2}
{'basic': 1220, 'pro': 3050}
[2, 4, 6]Iterators: what a for loop really does
A list, a string, a dict, a file, a range — anything you can loop over is an iterable. When for starts, it asks the iterable for an iterator with iter(), then calls next() on it again and again. When there is nothing left, the iterator raises StopIteration and the loop ends quietly. You can do the same steps by hand:
words = ["tokens", "embeddings", "context"]
it = iter(words) # ask the list for an iterator
print(next(it))
print(next(it))
print(next(it))
try:
next(it)
except StopIteration:
print("No more items")tokens
embeddings
context
No more itemsThe difference matters in practice. An iterable can give you a fresh iterator every time; an iterator only goes forward, once:
it = iter([1, 2, 3])
print(list(it))
print(list(it)) # the iterator is used up
numbers = [1, 2, 3]
print(list(numbers))
print(list(numbers)) # a list can be looped over again and again[1, 2, 3]
[]
[1, 2, 3]
[1, 2, 3]This is why reading a file or a streamed model answer a second time gives you nothing: both are iterators, and the first loop used them up.
Your own iterator
A class becomes an iterator when it has two special methods: __iter__, which returns the iterator (here the object itself), and __next__, which returns the next value or raises StopIteration:
class Countdown:
def __init__(self, start):
self.current = start
def __iter__(self):
return self
def __next__(self):
if self.current <= 0:
raise StopIteration
value = self.current
self.current -= 1
return value
for n in Countdown(3):
print(n, end=" ")
print()
print(list(Countdown(5)))
def countdown(start): # the same thing as a generator
while start > 0:
yield start
start -= 1
print(list(countdown(5)))3 2 1
[5, 4, 3, 2, 1]
[5, 4, 3, 2, 1]Look at the last three lines: a function with yield does exactly the same job in five lines instead of twelve. A generator is the easy way to write an iterator — Python writes __iter__ and __next__ for you. That is what the next two sections are about.
Generator expressions: lazy and memory-friendly
Write a comprehension with round brackets — or directly inside a function call — and Python does not build the whole list. It produces one value at a time, as it is needed. For a million numbers that saves a lot of memory. any() and all() pair perfectly with them:
numbers = range(1, 1_000_001)
total = sum(n * n for n in numbers) # no list is ever built
print(total)
scores = [0.4, 0.9, 0.7]
print(any(s > 0.8 for s in scores)) # is at least one above 0.8?
print(all(s > 0.3 for s in scores)) # are all above 0.3?333333833333500000
True
TrueWriting your own generator: yield
A function that uses yield instead of return becomes a generator: each yield hands out one value and pauses until the next one is asked for. A classic AI use is batching — sending documents to an API a few at a time instead of all at once:
def chunks(items, size):
"""Yield successive pieces of `items`, each at most `size` long."""
for start in range(0, len(items), size):
yield items[start:start + size]
documents = [f"doc{i}" for i in range(1, 8)]
for batch in chunks(documents, 3):
print("Sending batch:", batch)Sending batch: ['doc1', 'doc2', 'doc3']
Sending batch: ['doc4', 'doc5', 'doc6']
Sending batch: ['doc7']The same idea powers streaming: when ChatGPT types its answer word by word, your code receives the pieces from a generator-like stream. You will use one on page 18.
Decorators: wrapping a function in extra behaviour
In Python a function is a value like any other: you can store it in a variable, pass it to another function and return it from one. A decorator uses that. It is a function that takes a function and returns a new one that does something extra — timing it, logging it, retrying it, checking a permission — and then calls the original.
import time
from functools import wraps
def timed(func):
"""Print how long each call to `func` takes."""
@wraps(func)
def wrapper(*args, **kwargs):
start = time.perf_counter()
result = func(*args, **kwargs)
elapsed = time.perf_counter() - start
print(f"{func.__name__} took {elapsed:.1f}s")
return result
return wrapper
@timed
def slow_answer(question):
time.sleep(0.5) # pretend this is a model call
return f"Answer to: {question}"
print(slow_answer("What is a token?"))
print(slow_answer.__name__)slow_answer took 0.5s
Answer to: What is a token?
slow_answer@timedabovedef slow_answeris shorthand forslow_answer = timed(slow_answer). From then on, callingslow_answerreally callswrapper.wrapper(*args, **kwargs)accepts any arguments and passes them on unchanged (page 7), so one decorator works on any function.@wraps(func)copies the original's name and docstring onto the wrapper. Without it,slow_answer.__name__would saywrapper, which makes error messages and logs confusing.
A decorator with settings
To give a decorator options, add one more layer: a function that takes the options and returns the decorator. Here is a retry decorator — exactly what you want around a network call that sometimes fails:
from functools import wraps
def retry(times):
"""Call the function again, up to `times` attempts, on ConnectionError."""
def decorator(func):
@wraps(func)
def wrapper(*args, **kwargs):
for attempt in range(1, times + 1):
try:
return func(*args, **kwargs)
except ConnectionError as error:
print(f"attempt {attempt} failed: {error}")
raise ConnectionError(f"gave up after {times} attempts")
return wrapper
return decorator
calls = 0
@retry(times=3)
def flaky_api():
global calls
calls += 1
if calls < 3:
raise ConnectionError("server busy")
return "ok"
print(flaky_api())attempt 1 failed: server busy
attempt 2 failed: server busy
okYou will meet decorators everywhere in AI code, usually written by someone else: @dataclass and @property on page 12, @app.get("/") in web frameworks such as FastAPI, @tool in agent libraries that turn your function into something a model can call, and caching decorators like @functools.cache. Now you know what they do: they wrap your function.
When not to use a comprehension
Comprehensions are for building a collection. If the logic needs several steps, nested conditions or print calls, a normal loop is clearer. Readable code beats clever code.
Try it yourself
- From a list of file names, make a list of only the ones ending in
.pdf. - Turn
["Apple", "banana", "Cherry"]into a dict mapping each word to its length. - Write a generator
numbered(lines)that yields"1. first line","2. second line"and so on. - Write a decorator
loggedthat prints the function name and its arguments before every call.