Chapter 3 · Calculus: How Models Learn
Functions, Slopes and Derivatives
- Page 6 of 17
- 3 min read
Calculus sounds hard, but the part AI needs is one idea: the derivative, which tells you how fast something changes. If you stand on a hill, the derivative is the steepness under your feet — and which way is downhill. Training a model is walking downhill on a "how wrong am I" landscape, so this idea is at the heart of machine learning.
A function and its slope
A function takes an input and gives an output: f(x) = x² turns 3 into 9. Its slope at a point is "rise over run": how much the output changes for a tiny change in the input. Make the step tiny, and you get the derivative. You can estimate it directly in Python:
def f(x):
return x ** 2
def slope(f, x, h=1e-6):
"""Rise over run across a tiny step h."""
return (f(x + h) - f(x)) / h
for x in [-2, 0, 1, 3]:
print(f"x = {x:2}: slope ≈ {slope(f, x):.3f} exact 2x = {2 * x}")x = -2: slope ≈ -4.000 exact 2x = -4
x = 0: slope ≈ 0.000 exact 2x = 0
x = 1: slope ≈ 2.000 exact 2x = 2
x = 3: slope ≈ 6.000 exact 2x = 6The estimate matches the exact rule: the derivative of x² is 2x. At x = 0 the slope is 0 — the bottom of the curve, where it is flat. For negative x the slope is negative: going right makes the output smaller. The sign of the derivative tells you which way is downhill.
The rules you will meet
| Function | Derivative | In words |
|---|---|---|
c (a constant) | 0 | a flat line does not change |
a·x | a | a straight line has the same slope everywhere |
xⁿ | n·xⁿ⁻¹ | the power rule: x² → 2x, x³ → 3x² |
eˣ | eˣ | grows exactly as fast as its own value |
ln x | 1/x | grows ever more slowly |
f(g(x)) | f′(g(x)) · g′(x) | the chain rule: multiply the slopes of each step |
The chain rule: slopes through a pipeline
A model is a chain of functions: a dot product, then an activation, then a loss. To know how a weight deep inside affects the final loss, you multiply the slopes of each link of the chain. That is the chain rule, and applying it backwards through a network, layer by layer, is called backpropagation. Libraries such as PyTorch do it for you automatically — "autograd" — but it is still the chain rule. Here the known derivative of the sigmoid (which comes from the chain rule) matches the numerical estimate:
import numpy as np
def sigmoid(z):
return 1 / (1 + np.exp(-z))
def slope(f, x, h=1e-6):
return (f(x + h) - f(x)) / h
z = 0.5
exact = sigmoid(z) * (1 - sigmoid(z)) # the known derivative of the sigmoid
print(round(slope(sigmoid, z), 6), round(exact, 6))0.235004 0.235004More than one input: partial derivatives
A real model has millions of weights. A partial derivative, written ∂L/∂w, is the slope with respect to one weight while all the others stay fixed. Put all of them in a list, and you have the gradient: a vector that points uphill in every direction at once. The next page uses it.
Try it yourself
- Use
slope()onf(x) = x³at x = 2. Does it match 3x² = 12? - Try
h = 1e-1andh = 1e-12. Why are both worse than1e-6? - Where is the slope of
f(x) = (x − 4)²zero? What is special about that point?