Chapter 6 · Measuring Models and Next Steps
Measuring Models: Precision, Recall and the Confusion Matrix
- Page 16 of 17
- 4 min read
"The model is 95% accurate" sounds good — until you learn that it never catches the one thing it was built to catch. This page gives you the numbers that tell the real story. They apply to any yes/no decision: spam filters, fraud alerts, content moderation, an LLM judging whether an answer is correct.
The confusion matrix
Every prediction falls into one of four boxes:
| Actually yes | Actually no | |
|---|---|---|
| Predicted yes | True positive (TP) | False positive (FP) — a false alarm |
| Predicted no | False negative (FN) — a miss | True negative (TN) |
import numpy as np
# 1 = spam, 0 = not spam, for 12 emails
actual = np.array([1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0])
predicted = np.array([1, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0])
tp = np.sum((predicted == 1) & (actual == 1)) # caught spam
fp = np.sum((predicted == 1) & (actual == 0)) # good email wrongly flagged
fn = np.sum((predicted == 0) & (actual == 1)) # spam that got through
tn = np.sum((predicted == 0) & (actual == 0)) # good email left alone
print(f"TP={tp} FP={fp} FN={fn} TN={tn}")
precision = tp / (tp + fp)
recall = tp / (tp + fn)
print(f"accuracy {(tp + tn) / len(actual):.2f}")
print(f"precision {precision:.2f}")
print(f"recall {recall:.2f}")
print(f"F1 {2 * precision * recall / (precision + recall):.2f}")TP=3 FP=1 FN=1 TN=7
accuracy 0.83
precision 0.75
recall 0.75
F1 0.75| Metric | Formula | Question it answers |
|---|---|---|
| Accuracy | (TP + TN) / all | What share of all predictions were right? |
| Precision | TP / (TP + FP) | When it says "yes", how often is it right? |
| Recall | TP / (TP + FN) | Of all the real "yes" cases, how many did it find? |
| F1 | 2 × P × R / (P + R) | One number that is high only when both are high |
Recall is also called sensitivity — the same idea as the medical test on page 14.
Why accuracy lies
When one class is rare, accuracy is almost meaningless. A "model" that never predicts fraud:
import numpy as np
rng = np.random.default_rng(4)
actual = (rng.random(10_000) < 0.01).astype(int) # 1% of transactions are fraud
lazy = np.zeros_like(actual) # a "model" that never says fraud
print(f"accuracy of always saying 'not fraud': {np.mean(lazy == actual):.1%}")
print(f"frauds it catches (recall): {np.sum((lazy == 1) & (actual == 1))} of {actual.sum()}")accuracy of always saying 'not fraud': 98.9%
frauds it catches (recall): 0 of 10998.9% accurate, and it catches none of the 109 frauds — recall is 0. On imbalanced data, always look at precision and recall for the rare class.
The trade-off: choosing a threshold
Most models output a probability, and you choose the threshold above which you say "yes". Moving it trades precision against recall:
import numpy as np
rng = np.random.default_rng(8)
actual = rng.random(2000) < 0.2 # 20% positives
# A model's probabilities: higher, on average, for the real positives
score = np.clip(np.where(actual, rng.normal(0.7, 0.15, 2000), rng.normal(0.35, 0.15, 2000)), 0, 1)
for threshold in [0.3, 0.5, 0.7]:
pred = score >= threshold
tp = np.sum(pred & actual)
precision = tp / pred.sum()
recall = tp / actual.sum()
print(f"threshold {threshold}: precision {precision:.2f}, recall {recall:.2f}")threshold 0.3: precision 0.27, recall 1.00
threshold 0.5: precision 0.56, recall 0.89
threshold 0.7: precision 0.93, recall 0.49A low threshold catches every positive (recall 1.00) but most alerts are false (precision 0.27). A high one is almost always right (0.93) but misses half. There is no correct threshold in general — it depends on the cost of each mistake:
- Missing is expensive (cancer screening, fraud, safety filters): favour recall.
- False alarms are expensive (blocking a real customer's payment, flagging a good email as spam): favour precision.
In practice use sklearn.metrics — confusion_matrix, precision_score, recall_score, classification_report — but now you know exactly what they compute.
Try it yourself
- Change one prediction in the confusion example so that precision becomes 1.0. What happens to recall?
- Find the threshold in the last example that gives the highest F1.
- For an LLM-based moderation filter on a children's app, would you favour precision or recall? Why?