Chapter 4 · Statistics: Reading Data Honestly
Distributions, the Normal Curve and Scaling Features
- Page 10 of 17
- 3 min read
A distribution describes which values are possible and how often each one turns up. Rolling a fair die gives a uniform distribution: every face equally likely. Heights, measurement errors and many other things that come from lots of small effects added together follow the normal distribution, the famous bell curve.
The normal distribution
A normal distribution is fully described by two numbers: its mean (the centre) and its standard deviation (the width). And it follows a reliable rule, which you can check by simulating 100,000 heights:
import numpy as np
rng = np.random.default_rng(42)
heights = rng.normal(loc=165, scale=7, size=100_000) # mean 165 cm, std 7 cm
mean, std = heights.mean(), heights.std()
print(f"mean {mean:.1f}, std {std:.2f}")
for k in [1, 2, 3]:
inside = np.mean(np.abs(heights - mean) < k * std)
print(f"within {k} std: {inside:.1%}")mean 165.0, std 7.03
within 1 std: 68.4%
within 2 std: 95.4%
within 3 std: 99.7%| Within… | Share of the data | For heights (mean 165, std 7) |
|---|---|---|
| 1 standard deviation | about 68% | 158–172 cm |
| 2 standard deviations | about 95% | 151–179 cm |
| 3 standard deviations | about 99.7% | 144–186 cm |
So something more than 3 standard deviations from the mean is very rare — which is a simple way to flag anomalies, such as a sudden spike in API errors or an unusual transaction.
Not everything is normal. Incomes, response times (page 9) and word frequencies have long right tails: most values small, a few huge. Always look at your data (a histogram, page 15 of Python for AI) before assuming a bell curve.
Z-scores and standardisation
A z-score says how many standard deviations a value is from the mean: z = (x − mean) / std. Turning a whole feature into z-scores is called standardisation. Afterwards every feature has mean 0 and standard deviation 1, whatever its original units:
import numpy as np
# Two features on very different scales
salary = np.array([25_000, 40_000, 55_000, 120_000]) # taka per month
age = np.array([22, 30, 41, 35]) # years
def standardise(x):
return (x - x.mean()) / x.std()
print(standardise(salary).round(2))
print(standardise(age).round(2))
print(standardise(salary).mean().round(2), standardise(salary).std().round(2))[-0.97 -0.55 -0.14 1.66]
[-1.44 -0.29 1.29 0.43]
0.0 1.0Salary (tens of thousands) and age (tens) are now on the same scale. Without this, distance-based models and gradient descent would treat salary as hundreds of times more important, purely because its numbers are bigger — the problem you saw with the flats on page 2. This is exactly what scikit-learn's StandardScaler does inside a pipeline.
Try it yourself
- What is the z-score of a height of 186 cm when the mean is 165 and the standard deviation is 7?
- Simulate 100,000 rolls of a die with
rng.integers(1, 7, ...). Is it normal? Check the share within 1 standard deviation. - Standardise the flats from page 2 and recompute the distances. Which flat is now closest to flat A?