Chapter 5 · Probability
Conditional Probability and Bayes' Theorem
- Page 14 of 17
- 3 min read
Conditional probability is the probability of something given that you already know something else. It is written P(A | B): "the probability of A, given B". The probability that an email is spam is one thing; the probability that it is spam given that it contains the word "free" is another — and that second kind is what every classifier really computes.
Bayes' theorem
Bayes' theorem tells you how to update a belief when you see new evidence:
P(A | B) = P(B | A) × P(A) / P(B)P(A)— the prior: what you believed before the evidence (the base rate).P(B | A)— the likelihood: how likely the evidence is if A is true.P(A | B)— the posterior: your updated belief.
The surprising medical test
A disease affects 1% of people. A test catches 99% of real cases, and wrongly says "positive" for 5% of healthy people. You test positive. How likely is it that you have the disease? Most people guess about 95%. Here is the real answer — first with the formula, then by simulating a million people:
import numpy as np
p_disease = 0.01 # 1% of people have it (the base rate)
p_pos_if_disease = 0.99 # sensitivity: the test catches 99% of real cases
p_pos_if_healthy = 0.05 # 5% of healthy people still test positive
p_pos = p_pos_if_disease * p_disease + p_pos_if_healthy * (1 - p_disease)
p_disease_if_pos = p_pos_if_disease * p_disease / p_pos
print(f"P(disease | positive test) = {p_disease_if_pos:.1%}")
# Check by simulating one million people
rng = np.random.default_rng(2)
sick = rng.random(1_000_000) < p_disease
positive = np.where(sick, rng.random(1_000_000) < p_pos_if_disease,
rng.random(1_000_000) < p_pos_if_healthy)
print(f"simulated: {sick[positive].mean():.1%}")P(disease | positive test) = 16.7%
simulated: 16.5%Only about 17%. Picture 10,000 people: 100 are sick and 99 of them test positive. But 5% of the 9,900 healthy people — 495 — also test positive. Of the 594 positives, only 99 are really sick. When something is rare, even a good test produces more false alarms than true ones. Ignoring the base rate is called the base-rate fallacy, and it affects every AI system that looks for rare things: fraud, intrusions, defects, disease.
The spam filter
The first successful spam filters were built on Bayes' theorem. Count how often a word appears in spam and in normal email, then update the probability:
# Out of 1,000 emails: 200 spam, 800 not spam.
# "free" appears in 120 of the spam emails and in 40 of the others.
spam, ham = 200, 800
free_in_spam, free_in_ham = 120, 40
p_spam = spam / (spam + ham)
p_free_given_spam = free_in_spam / spam
p_free = (free_in_spam + free_in_ham) / (spam + ham)
print(f"P(spam) = {p_spam:.2f}")
print(f"P(spam | 'free') = {p_free_given_spam * p_spam / p_free:.2f}")P(spam) = 0.20
P(spam | 'free') = 0.75Before reading the email, 20% chance of spam. After seeing "free", 75%. A naive Bayes classifier does this for every word and combines the evidence. It is still a fast, strong baseline for text classification — scikit-learn has it as MultinomialNB.
Try it yourself
- In the medical example, change the disease rate to 10%. What is P(disease | positive) now?
- What if the false-positive rate drops from 5% to 1%?
- "win" appears in 50 spam and 10 normal emails. Compute P(spam | "win").