Chapter 4 · Statistics: Reading Data Honestly
Samples, Averages and How Sure You Can Be
- Page 12 of 17
- 3 min read
You almost never see all the data. You test a model on 200 questions, not every question users will ever ask. You survey 1,000 people, not the whole country. That smaller set is a sample, and the key question is: how much can you trust what it tells you?
The law of large numbers
The more data in your sample, the closer its average gets to the true value. A fair coin lands heads half the time — but watch how long it takes to show:
import numpy as np
rng = np.random.default_rng(0)
flips = rng.integers(0, 2, 100_000) # 1 = heads, 0 = tails
for n in [10, 100, 1_000, 10_000, 100_000]:
print(f"after {n:>7,} flips: share of heads = {flips[:n].mean():.3f}")after 10 flips: share of heads = 0.400
after 100 flips: share of heads = 0.560
after 1,000 flips: share of heads = 0.537
after 10,000 flips: share of heads = 0.503
after 100,000 flips: share of heads = 0.500After 10 flips, 40% heads; after 100, 56%. Only with thousands does it settle near 50%. Small samples wander. Conclusions from 10 or 20 test examples are mostly noise.
Standard error: the uncertainty of an average
The standard error measures how much an average would change if you took a new sample. For an accuracy p measured on n examples:
standard error = √( p × (1 − p) / n )Because of the square root, four times as much data only halves the uncertainty. Roughly, the true value lies within two standard errors of what you measured, 95% of the time — that range is a 95% confidence interval.
A real question: is prompt B better?
You test two prompts on the same 200 questions. A gets 86% right, B gets 88%. Is B better? The bootstrap answers without any formula: resample the results many times, with replacement, and see how much the accuracy moves:
import numpy as np
rng = np.random.default_rng(11)
# Two prompts, each tested on the same 200 questions (1 = correct)
prompt_a = np.array([1] * 172 + [0] * 28) # 86% correct
prompt_b = np.array([1] * 176 + [0] * 24) # 88% correct
def bootstrap_interval(results, rounds=10_000):
"""Resample with replacement; the middle 95% of the means is the interval."""
means = [rng.choice(results, size=len(results)).mean() for _ in range(rounds)]
return np.percentile(means, [2.5, 97.5])
for name, r in [("A", prompt_a), ("B", prompt_b)]:
low, high = bootstrap_interval(r)
print(f"prompt {name}: {r.mean():.0%} (95% interval {low:.0%} to {high:.0%})")
se = np.sqrt(0.86 * 0.14 / 200)
print(f"standard error of A: {se:.3f}")prompt A: 86% (95% interval 81% to 90%)
prompt B: 88% (95% interval 84% to 92%)
standard error of A: 0.025The two intervals overlap heavily (81–90% and 84–92%). With 200 questions, a 2-point difference could easily be luck. To trust a difference this small you need more test questions, or a larger gap. This one check saves teams from shipping "improvements" that are only noise — and it is exactly how serious AI evaluation reports results.
Comparing two prompts on the same questions lets you do even better: count only the questions where they disagree. That is called a paired test, and it is what evaluation tools use.
Try it yourself
- Calculate the standard error for 86% accuracy on 2,000 questions. How wide is the interval now?
- Run the bootstrap for prompt B with 400 questions (352 correct). Do the intervals still overlap?
- Change the coin's seed several times. How far from 0.5 can 10 flips land?