Chapter 1 · Why Maths for AI
Why Maths for AI (and How Much You Need)
- Page 1 of 17
- 4 min read
You can call a language model without knowing any maths. But the moment you need to understand what is happening — why a search returns the wrong document, why a model gets worse when trained longer, whether 88% is really better than 86% — you need a little maths. Not a degree: a handful of ideas, each one used every day in AI work. This tutorial teaches them with NumPy, so every formula is also code you can run.
A whole model in four lines
Here is a tiny spam detector. It looks at three numbers describing an email and gives a probability that it is spam:
import numpy as np
# One email, described by three numbers (features):
# number of links, share of CAPITAL letters, number of "!" (all scaled 0–1)
x = np.array([0.9, 0.7, 0.4])
w = np.array([2.0, 1.5, 0.5]) # weights: how much each feature matters
b = -1.5 # bias: the starting point
z = x @ w + b # dot product (page 3)
p = 1 / (1 + np.exp(-z)) # sigmoid turns a score into a probability (page 8)
print(f"score z = {z:.2f}")
print(f"probability of spam = {p:.2f}")score z = 1.55
probability of spam = 0.82Every line is a piece of maths you will learn here. The email is a vector (page 2). Combining it with the weights is a dot product (page 3). Turning the score into a probability uses the sigmoid function (page 8). Finding good weights is done with gradient descent (page 7). Checking whether the detector is any good needs precision and recall (page 16). A large language model is this same pattern, repeated billions of times.
The four areas
| Area | What it does in AI | Chapter |
|---|---|---|
| Linear algebra | Represents data as vectors and matrices: embeddings, images, model weights | 2 |
| Calculus | Explains how a model learns: slopes, gradients, loss | 3 |
| Statistics | Describes data and tells you whether a result is real or luck | 4 |
| Probability | Handles uncertainty: predictions, Bayes, how an LLM picks the next word | 5 |
How much you need
For an AI engineer who builds applications, you need to understand the ideas and recognise them in code. You will rarely calculate a derivative by hand; you will often need to know what a gradient is, why features should be scaled, or why a test set must be kept apart. Researchers who design new models go much deeper — this tutorial is the foundation either way.
Reading the notation
Maths in papers and documentation uses a few symbols again and again. You do not need to memorise them now; come back to this table whenever you meet one.
| Symbol | Read it as | In NumPy |
|---|---|---|
x, 𝐱 | a number; a vector (often bold) | x = np.array([...]) |
xᵢ | the i-th element of x | x[i] |
Σ xᵢ | the sum of all the xᵢ | x.sum() |
x̄ or μ | the mean (average) | x.mean() |
σ | the standard deviation | x.std() |
‖x‖ | the length of a vector | np.linalg.norm(x) |
x · y, xᵀy | the dot product | x @ y |
W, Wᵀ | a matrix; its transpose | W, W.T |
ŷ | "y hat": a prediction | y_pred |
∂L/∂w | how the loss L changes when w changes | a gradient (page 7) |
P(A | B) | the probability of A, given that B happened | page 14 |
How to use this tutorial
- Run every example yourself — in a notebook (Python for AI, page 20) it is easy to change a number and see what happens.
- When a formula looks frightening, read the code next to it first. Code is just maths written out step by step.
- Take the three quizzes: after chapter 2, after chapter 4, and the final exam at the end.
Try it yourself
- Change the email in the first example to
[0.0, 0.1, 0.0]. What happens to the probability, and why? - Which weight makes the biggest difference? Change each one to 0 in turn and compare.