Chapter 3 · Writing Real Programs
Files, JSON and CSV
- Page 10 of 20
- 4 min read
Programs forget everything when they stop. To keep data — your notes, a dataset, the answers a model gave — you write it to a file. AI work is full of files: CSV spreadsheets of training data, JSON configuration, JSON Lines datasets for fine-tuning. This page covers all three.
Reading and writing text files
from pathlib import Path
notes = Path("notes.txt")
# Write (creates the file, or replaces it)
with open(notes, "w", encoding="utf-8") as f:
f.write("Python is fun\n")
f.write("পাইথন শেখা সহজ\n")
# Append a line
with open(notes, "a", encoding="utf-8") as f:
f.write("Files are easy\n")
# Read line by line
with open(notes, encoding="utf-8") as f:
for number, line in enumerate(f, start=1):
print(number, line.rstrip())
print(notes.exists(), notes.suffix, notes.stat().st_size, "bytes")1 Python is fun
2 পাইথন শেখা সহজ
3 Files are easy
True .txt 68 byteswith open(...) as f:opens the file and always closes it at the end of the block, even if an error happens. Always usewith.- The mode:
"r"read (the default),"w"write — replaces the whole file,"a"append to the end. encoding="utf-8": always write it. Without it, Windows may use an old encoding and Bangla text becomes garbage or crashes with aUnicodeDecodeError.pathlib.Pathrepresents a file path and works the same on Windows, macOS and Linux. Note how the Bangla line takes more bytes than its 14 characters — UTF-8 uses 3 bytes for each Bangla character.
JSON: the language of APIs
JSON (JavaScript Object Notation) is a text format for structured data. It looks almost exactly like Python dictionaries and lists, and every AI API sends and receives it. The json module converts both ways:
import json
settings = {
"model": "gpt-5.5",
"temperature": 0.2,
"languages": ["en", "bn"],
"greeting": "স্বাগতম",
}
# Python → JSON text in a file
with open("settings.json", "w", encoding="utf-8") as f:
json.dump(settings, f, ensure_ascii=False, indent=2)
# JSON text → Python
with open("settings.json", encoding="utf-8") as f:
loaded = json.load(f)
print(loaded["languages"][1], loaded["greeting"])
print(json.dumps({"ok": True, "items": None})) # to a string, not a filebn স্বাগতম
{"ok": true, "items": null}The file it wrote, settings.json:
{
"model": "gpt-5.5",
"temperature": 0.2,
"languages": [
"en",
"bn"
],
"greeting": "স্বাগতম"
}| Function | Does |
|---|---|
json.dump(obj, file) / json.load(file) | Python ↔ a JSON file |
json.dumps(obj) / json.loads(text) | Python ↔ a JSON string (the "s" is for string) |
Notice the small differences: JSON writes true, false and null where Python has True, False and None. ensure_ascii=False keeps Bangla readable in the file instead of turning it into \u09b8 codes, and indent=2 makes it pretty.
CSV: spreadsheets as text
A CSV (comma-separated values) file is a table in plain text — what you get when you export from Excel or Google Sheets. Here is reviews.csv:
product,stars,comment
Headphones,5,Great sound
Charger,2,Stopped working
Keyboard,4,Nice to type oncsv.DictReader turns each row into a dictionary, using the first line as the keys:
import csv
with open("reviews.csv", newline="", encoding="utf-8") as f:
rows = list(csv.DictReader(f))
print(len(rows), "rows")
print(rows[0])
positive = [r for r in rows if int(r["stars"]) >= 4]
print([r["product"] for r in positive])3 rows
{'product': 'Headphones', 'stars': '5', 'comment': 'Great sound'}
['Headphones', 'Keyboard']Everything read from a CSV is text — that is why we wrote int(r["stars"]). For bigger tables you will use pandas (page 14), which does these conversions for you.
JSON Lines: one record per line
AI datasets — fine-tuning data, evaluation sets, batch API jobs — usually use JSON Lines (.jsonl): one complete JSON object on each line. You can add a line without rewriting the file, and read a huge file one line at a time:
import json
examples = [
{"prompt": "2 + 2", "answer": "4"},
{"prompt": "Capital of Bangladesh", "answer": "Dhaka"},
]
with open("dataset.jsonl", "w", encoding="utf-8") as f:
for row in examples:
f.write(json.dumps(row, ensure_ascii=False) + "\n")
with open("dataset.jsonl", encoding="utf-8") as f:
for line in f:
row = json.loads(line)
print(row["prompt"], "->", row["answer"])2 + 2 -> 4
Capital of Bangladesh -> DhakaTry it yourself
- Write five of your favourite quotes to
quotes.txt, one per line, then read them back and print the longest. - Save a dictionary of your study plan (subject → hours) as JSON, load it again and print the total hours.
- From
reviews.csv, write a new filebad_reviews.jsonlcontaining only reviews with fewer than 3 stars.