Chapter 5 · Python for AI Work
Calling an LLM from Python
- Page 17 of 23
- 5 min read
This is the moment the whole tutorial has been building towards: calling a large language model from your own Python code. With the official SDKs it takes about ten lines. The hard part is not the call — it is doing it safely, reliably and without a surprise bill.
1. Get an API key, and keep it secret
Create a key in your provider's dashboard (OpenAI or Anthropic). Usage is billed per token, so set a monthly spending limit in the dashboard straight away. A key is like a password to your account:
- Never write it in your code, and never paste it into a chat or a screenshot.
- Keep it in an environment variable. In development, a
.envfile loaded bypython-dotenvis the usual way:
# .env — keep this file OUT of git (add ".env" to .gitignore)
OPENAI_API_KEY=sk-...your key...
ANTHROPIC_API_KEY=sk-ant-...your key...Add .env to .gitignore before your first commit. Keys pushed to GitHub are found by bots within minutes. If one ever leaks, delete it in the dashboard immediately and create a new one.
2. Install the SDK
python -m pip install openai python-dotenv (or anthropic for Claude). An SDK (software development kit) wraps the HTTP calls from the previous page in friendly Python classes and handles retries, timeouts and errors for you.
3. Your first call
OpenAI's current API is the Responses API. instructions sets the model's behaviour; input is the user's message:
from dotenv import load_dotenv
from openai import OpenAI
load_dotenv() # reads .env into the environment
client = OpenAI() # finds OPENAI_API_KEY by itself
response = client.responses.create(
model="gpt-5.5",
instructions="You are a friendly Python tutor. Answer in two sentences.",
input="What is a Python dictionary?",
)
print(response.output_text)
print("Tokens in:", response.usage.input_tokens, "| out:", response.usage.output_tokens)Example output (the model's exact words will differ every time):
A dictionary stores data as key-value pairs, like a real dictionary maps words to meanings. You look values up by key, for example person["name"].
Tokens in: 29 | out: 36The usage numbers are what you pay for: input tokens (everything you sent) and output tokens (everything it wrote). Log them. A loop that accidentally sends a whole document 1,000 times is an expensive bug.
The same with Claude
Every provider's SDK follows the same shape — a client, a create call, the text in the result. Anthropic's Messages API requires max_tokens, the upper limit on the reply length:
from dotenv import load_dotenv
from anthropic import Anthropic
load_dotenv()
client = Anthropic() # finds ANTHROPIC_API_KEY by itself
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=300,
system="You are a friendly Python tutor. Answer in two sentences.",
messages=[{"role": "user", "content": "What is a Python dictionary?"}],
)
print(message.content[0].text)Example output:
A Python dictionary is a collection of key-value pairs, written with curly braces. It lets you find a value quickly by its key instead of by position.Model names change often: new versions arrive and old ones are retired. Check the provider's models page and keep the name in one place (a constant or a setting) so you can change it easily.
4. Conversations: the model has no memory
An API call is stateless: the model does not remember your previous call. A chat "remembers" only because the program sends the whole history every time — the list of message dicts from page 5:
from openai import OpenAI
client = OpenAI()
history = []
def ask(question):
history.append({"role": "user", "content": question})
response = client.responses.create(model="gpt-5.5", input=history)
answer = response.output_text
history.append({"role": "assistant", "content": answer}) # the model remembers nothing by itself
return answer
print(ask("My name is Nadia. Suggest one Python project for a beginner."))
print(ask("What was my name?"))Example output:
Try a to-do list app that saves tasks to a JSON file, Nadia.
Your name is Nadia.This is also why long chats cost more and eventually hit the context window limit: every turn resends everything before it.
5. When things go wrong
The SDK raises specific exceptions, so you can apply page 11's lessons directly. The OpenAI SDK already retries connection errors, 408, 409, 429 and 5xx responses twice with exponential backoff; max_retries changes that number. Here the key was wrong:
import openai
from openai import OpenAI
client = OpenAI(timeout=30, max_retries=3) # retries 429 and 5xx with backoff
try:
response = client.responses.create(model="gpt-5.5", input="Hello")
print(response.output_text)
except openai.AuthenticationError:
print("Check your API key.") # 401: retrying will not help
except openai.RateLimitError:
print("Still rate limited after retries — slow down or check your quota.")
except openai.APIConnectionError:
print("Could not reach the API — check the network.")
except openai.APIStatusError as error:
print("API error", error.status_code)Check your API key.A checklist before you ship
- The key comes from the environment, never from the code.
- A timeout and retries are set; errors give the user a helpful message.
- Token usage is logged, and a spending limit is set in the dashboard.
- Users' personal data is not sent unless your provider's terms and your own privacy policy allow it.
- The model's output is treated as untrusted text: checked before it is shown, stored or acted upon.
Try it yourself
- Write
translate(text, language)that asks a model to translate text, and try it with English → Bangla. - Print the input and output token counts and estimate the cost using your provider's current price list.
- Turn the conversation example into a loop that keeps chatting until the user types
quit.