Chapter 4 · Statistics: Reading Data Honestly
Describing Data: Mean, Median, Spread and Percentiles
- Page 9 of 17
- 3 min read
Before you train anything, you look at the data. Descriptive statistics sum up thousands of numbers in a few: where the middle is, and how spread out they are. These are also the numbers you report when you monitor an AI system in production.
The middle: mean or median?
The mean is the sum divided by the count. The median is the middle value once the data is sorted. They disagree when the data has outliers:
import numpy as np
# Response times of an API in milliseconds — one very slow call
latency = np.array([210, 190, 230, 205, 220, 198, 215, 2400, 207, 212])
print("mean: ", latency.mean())
print("median:", np.median(latency))
print("std: ", latency.std().round(1))
print("without the slow call: mean", np.delete(latency, 7).mean().round(1),
" std", np.delete(latency, 7).std().round(1))mean: 428.7
median: 211.0
std: 657.2
without the slow call: mean 209.7 std 11.1One slow call out of ten more than doubles the mean (428.7 ms), yet the median (211 ms) still describes a typical call. Use the median when a few extreme values could mislead you — incomes, house prices, response times. The mode is the most common value; it is mostly used for categories ("which answer did users pick most?").
The spread: variance and standard deviation
Two classes can have the same average and be completely different:
import numpy as np
a = np.array([48, 50, 50, 52]) # two classes with the same average mark
b = np.array([20, 40, 60, 80])
for name, marks in [("class A", a), ("class B", b)]:
deviations = marks - marks.mean()
variance = np.mean(deviations ** 2)
print(f"{name}: mean {marks.mean():.0f}, variance {variance:.0f}, std {np.sqrt(variance):.1f}")class A: mean 50, variance 2, std 1.4
class B: mean 50, variance 500, std 22.4- Subtract the mean from every value (the deviations).
- Square them, so negatives do not cancel positives.
- Average the squares: that is the variance.
- Take the square root to get back to the original units: the standard deviation (std, σ).
Class A's marks are all within a couple of marks of 50; class B's are spread across 60 marks. Same mean, very different stories. Always report a spread next to an average.
NumPy's
std()divides by n. pandas'.std()divides by n − 1 (the "sample" standard deviation), so the two give slightly different answers on small data. Both are correct; they answer slightly different questions.
Percentiles: p50, p95, p99
The p95 (95th percentile) is the value that 95% of the data is below. For an AI service, it says "95 out of 100 users waited less than this". Teams set targets on p95 and p99 rather than the mean, because the slowest users are the ones who complain:
import numpy as np
rng = np.random.default_rng(7)
latency = rng.lognormal(mean=5.3, sigma=0.4, size=10_000) # 10,000 simulated API calls
for p in [50, 95, 99]:
print(f"p{p}: {np.percentile(latency, p):6.0f} ms")
print(f"mean: {latency.mean():5.0f} ms")p50: 199 ms
p95: 384 ms
p99: 501 ms
mean: 216 msThe median (p50) is 199 ms, but 1 user in 100 waits more than half a second. The mean (216 ms) hides that completely.
Try it yourself
- Replace the 2,400 ms call with 240 ms. How much do the mean and median change?
- Compute the variance of
[2, 4, 4, 4, 5, 5, 7, 9]by hand, then check with NumPy (the answer is 4). - Load a CSV of your own with pandas and run
df.describe(). Which numbers on this page does it show?