Chapter 5 · Probability
Probability: The Basic Rules
- Page 13 of 17
- 4 min read
AI systems are uncertain by nature: a classifier says "90% likely spam", a language model picks words with probabilities, an agent's step succeeds most of the time. Probability is the language for that uncertainty. A probability is a number from 0 (impossible) to 1 (certain), often written as a percentage.
Three rules
| Rule | Formula | Example with a die |
|---|---|---|
| Not | P(not A) = 1 − P(A) | P(not six) = 1 − 1/6 = 5/6 |
| And (independent events) | P(A and B) = P(A) × P(B) | two sixes in a row = 1/6 × 1/6 = 1/36 |
| Or | P(A or B) = P(A) + P(B) − P(A and B) | even or above 4: 3/6 + 2/6 − 1/6 = 4/6 |
In the last rule you subtract the overlap (a 6 is both even and above 4) so it is not counted twice. Independent means one event does not change the chance of the other — true for two dice, not true for "it is cloudy" and "it rains".
Check it by simulation
When a probability is hard to work out, simulate it: repeat the experiment many times with random numbers and count. NumPy makes that a few lines:
import numpy as np
rng = np.random.default_rng(5)
rolls = rng.integers(1, 7, size=600_000) # a fair die, rolled 600,000 times
print("P(six) ≈", np.mean(rolls == 6).round(3), " exact", round(1 / 6, 3))
print("P(not six) ≈", np.mean(rolls != 6).round(3), " exact", round(5 / 6, 3))
print("P(even or > 4) ≈", np.mean((rolls % 2 == 0) | (rolls > 4)).round(3), " exact", round(4 / 6, 3))P(six) ≈ 0.167 exact 0.167
P(not six) ≈ 0.833 exact 0.833
P(even or > 4) ≈ 0.666 exact 0.667Why long agent chains fail
An AI agent that books a trip might search, compare, fill a form, pay and confirm. If every step must succeed, the "and" rule multiplies the probabilities:
# An agent finishes a task only if every step succeeds
for p_step in [0.99, 0.95, 0.90]:
row = [f"{p_step ** n:.0%}" for n in [1, 5, 10, 20]]
print(f"each step {p_step:.0%} reliable -> 1, 5, 10, 20 steps: {', '.join(row)}")each step 99% reliable -> 1, 5, 10, 20 steps: 99%, 95%, 90%, 82%
each step 95% reliable -> 1, 5, 10, 20 steps: 95%, 77%, 60%, 36%
each step 90% reliable -> 1, 5, 10, 20 steps: 90%, 59%, 35%, 12%Even with each step 95% reliable, a 20-step task succeeds only about a third of the time. This is one of the most important numbers in agent engineering. It is why good agents keep chains short, check their own results, retry failed steps and ask a human when unsure.
Expected value
The expected value is the average outcome over many repeats: multiply each outcome by its probability and add up. It is how you estimate the running cost of an AI feature:
# What does one support question cost on average?
outcomes = [
# (probability, cost in US cents)
(0.70, 0.2), # answered by a small, cheap model
(0.25, 1.5), # escalated to a large model
(0.05, 40.0), # handed to a human agent
]
expected = sum(p * cost for p, cost in outcomes)
print(f"expected cost: {expected:.3f} cents per question")
print(f"for 100,000 questions a month: ${expected * 100_000 / 100:,.0f}")expected cost: 2.515 cents per question
for 100,000 questions a month: $2,515The rare human hand-off (5% of questions) is 80% of the cost — 2 cents out of 2.5. Expected value shows where to optimise first.
Try it yourself
- What is the probability of at least one six in four rolls? (Hint: 1 − P(no six in four rolls).) Check by simulation.
- How reliable must each step be for a 10-step agent to succeed 90% of the time?
- If the human hand-off falls to 2%, what is the new expected cost?