Chapter 6 · Project and Next Steps
Project: An AI Review Analyser
- Page 21 of 23
- 5 min read
Time to put everything together. You will build a small but realistic AI tool: it reads customer reviews from a CSV file, works out the sentiment (positive, negative or neutral) and the topic of each one, and prints a summary per product. Companies run exactly this kind of pipeline on thousands of reviews and support tickets.
It uses nearly every page of this tutorial: files and CSV (p. 10), pandas (p. 14), functions and dataclasses (p. 7, 12), error handling (p. 11), JSON (p. 10), an LLM API (p. 17) and validation (p. 18).
The design
- Two classifiers with the same shape.
classify_offlineuses keyword rules: free, instant, deterministic — ideal for building and testing the rest of the program without spending money.classify_with_llmasks a model and is far more accurate. Both return anAnalysis, so the rest of the code does not care which one ran. - Validate the model's answer. Anything that is not the JSON we asked for is skipped with a message, not allowed to crash the run or pollute the results.
- Cache results. Every finished analysis is appended to
cache.jsonl. Run the program again and reviews already done are read from the cache instead of being sent (and paid for) again. If the run crashes halfway, nothing finished is lost. - A command-line interface.
argparseturns command-line words into options and writes a--helpmessage for you.
The data: reviews.csv
id,product,review
r1,Headphones,Great sound and fast delivery. Love them!
r2,Headphones,Stopped working after a week. I want a refund.
r3,Charger,Arrived late but works perfectly.
r4,Charger,Terrible. It broke on the first day.
r5,Keyboard,"Excellent keys, very happy with it."
r6,Keyboard,It is a keyboard. It types.
r7,Headphones,"Delivery was slow, sound is great."
r8,Charger,"Fast charging, perfect for my phone."The program: analyse_reviews.py
"""Analyse customer reviews with an LLM — or offline, with simple keyword rules."""
import argparse
import json
from dataclasses import asdict, dataclass
from pathlib import Path
import pandas as pd
MODEL = "gpt-5.5"
ALLOWED = {"positive", "negative", "neutral"}
CACHE = Path("cache.jsonl")
@dataclass
class Analysis:
sentiment: str
topic: str
def classify_offline(text: str) -> Analysis:
"""A rough keyword baseline: free, instant, and good for testing the pipeline."""
words = text.lower()
good = ["love", "great", "excellent", "perfect", "fast", "happy"]
bad = ["broke", "late", "terrible", "refund", "slow", "stopped"]
score = sum(w in words for w in good) - sum(w in words for w in bad)
sentiment = "positive" if score > 0 else "negative" if score < 0 else "neutral"
topic = "delivery" if any(w in words for w in ["deliver", "arrived", "late"]) else "product"
return Analysis(sentiment, topic)
def classify_with_llm(client, text: str) -> Analysis:
response = client.responses.create(
model=MODEL,
instructions=(
"Classify the customer review. Reply with JSON only, like "
'{"sentiment": "positive", "topic": "delivery"}. '
"sentiment is positive, negative or neutral; topic is one lower-case word."
),
input=text,
)
data = json.loads(response.output_text) # may raise ValueError
if data.get("sentiment") not in ALLOWED:
raise ValueError(f"unexpected sentiment {data.get('sentiment')!r}")
return Analysis(data["sentiment"], str(data.get("topic", "other")).lower())
def load_cache() -> dict[str, dict]:
if not CACHE.exists():
return {}
with open(CACHE, encoding="utf-8") as f:
return {row["id"]: row for row in map(json.loads, f)}
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("csv", help="a CSV file with id, product and review columns")
parser.add_argument("--offline", action="store_true", help="use keyword rules instead of an LLM")
args = parser.parse_args()
reviews = pd.read_csv(args.csv)
cache = load_cache()
client = None
if not args.offline:
from openai import OpenAI
client = OpenAI(timeout=30, max_retries=3)
results = []
with open(CACHE, "a", encoding="utf-8") as cache_file:
for row in reviews.itertuples():
if row.id in cache: # analysed before: no second API call
results.append(cache[row.id])
continue
try:
if args.offline:
analysis = classify_offline(row.review)
else:
analysis = classify_with_llm(client, row.review)
except (ValueError, KeyError) as error:
print(f"Skipping {row.id}: {error}")
continue
record = {"id": row.id, "product": row.product, **asdict(analysis)}
cache_file.write(json.dumps(record, ensure_ascii=False) + "\n")
results.append(record)
df = pd.DataFrame(results)
df.to_csv("results.csv", index=False)
print(f"Analysed {len(df)} of {len(reviews)} reviews\n")
print(pd.crosstab(df["product"], df["sentiment"]))
print("\nTopics:", df["topic"].value_counts().to_dict())
if __name__ == "__main__":
main()Run it
python analyse_reviews.py reviews.csv --offline # free: keyword rules
python analyse_reviews.py reviews.csv # real: calls the modelOffline mode (this is the real output):
Analysed 8 of 8 reviews
sentiment negative neutral positive
product
Charger 1 1 1
Headphones 1 1 1
Keyboard 0 1 1
Topics: {'product': 5, 'delivery': 3}With a model (example output — the model's judgements will vary a little; here it returned prose instead of JSON for r6, which the program caught and skipped):
Skipping r6: Expecting value: line 1 column 1 (char 0)
Analysed 7 of 8 reviews
sentiment negative neutral positive
product
Charger 1 0 2
Headphones 1 1 1
Keyboard 0 0 1
Topics: {'durability': 2, 'delivery': 2, 'sound': 1, 'typing': 1, 'charging': 1}Read the results critically
Compare the two runs. The keyword rules called r3 ("Arrived late but works perfectly") neutral: one good word and one bad word cancel out. The model understood that the customer is happy with the product despite the late delivery. Its topics — durability, charging, typing — also say far more than the rules' plain "product". Keyword rules cannot read meaning; models can — but they cost money, take time and occasionally return something unexpected. That trade-off is the daily reality of AI engineering, and a cheap baseline like classify_offline is also how you check that the expensive model is actually better.
Make it your own
- Add a
--limit Noption that analyses only the first N reviews, for cheap test runs. - Retry a review once when the model's JSON is invalid, adding "Reply with JSON only" to the input.
- Use
AsyncOpenAIand a semaphore (p. 18) to analyse 5 reviews at once. - Draw a bar chart of sentiment per product with Matplotlib and save it as
report.png. - Write ten reviews in Bangla and check whether the model handles them as well as English.