Chapter 3 · Calculus: How Models Learn
Gradient Descent: How Models Learn
- Page 7 of 17
- 4 min read
Every model you have heard of — from a spam filter to GPT — learns the same way. Start with random weights. Measure how wrong the model is with a loss number. Work out which direction makes the loss smaller (the derivative from the last page). Take a small step that way. Repeat. That is gradient descent.
new weight = old weight − learning rate × gradientThe minus sign is the key: the gradient points uphill, so you step the other way. The learning rate decides how big each step is.
One weight, three learning rates
Here the loss is (w − 3)², which is smallest at w = 3. Each run starts at w = 0 and takes 20 steps:
def loss(w):
return (w - 3) ** 2 # smallest at w = 3
def gradient(w):
return 2 * (w - 3) # its derivative
for learning_rate in [0.1, 0.01, 1.1]:
w = 0.0
for step in range(20):
w = w - learning_rate * gradient(w)
print(f"learning rate {learning_rate:<4}: w = {w:10.4f}, loss = {loss(w):.4f}")learning rate 0.1 : w = 2.9654, loss = 0.0012
learning rate 0.01: w = 0.9972, loss = 4.0113
learning rate 1.1 : w = -112.0128, loss = 13227.9441- 0.1 — just right: w is almost at 3 after 20 steps.
- 0.01 — too small: it moves in the right direction but has only reached 1.0. It would get there, very slowly.
- 1.1 — too big: every step overshoots the bottom, further each time, and the loss explodes. When training loss suddenly becomes huge or
nan, a too-large learning rate is the first suspect.
Training a real model from scratch
Now two weights. We fit a straight line, predicted mark = w × hours + b, to six students' data. The loss is the mean squared error (MSE); its partial derivatives tell us how to change w and b:
import numpy as np
# Study hours and exam marks for 6 students
hours = np.array([1, 2, 3, 4, 5, 6], dtype=float)
marks = np.array([52, 55, 61, 64, 70, 74], dtype=float)
w, b = 0.0, 0.0 # the line: predicted = w * hours + b
learning_rate = 0.02
for step in range(5001):
predicted = w * hours + b
error = predicted - marks
loss = np.mean(error ** 2) # mean squared error
grad_w = 2 * np.mean(error * hours) # ∂loss/∂w
grad_b = 2 * np.mean(error) # ∂loss/∂b
w -= learning_rate * grad_w
b -= learning_rate * grad_b
if step in (0, 100, 1000, 5000):
print(f"step {step:4}: w = {w:5.2f}, b = {b:5.2f}, loss = {loss:8.2f}")
print("NumPy's exact answer:", np.polyfit(hours, marks, 1).round(2))
print("prediction for 7 hours:", round(w * 7 + b, 1))step 0: w = 9.30, b = 2.51, loss = 3987.00
step 100: w = 9.36, b = 26.14, loss = 84.36
step 1000: w = 4.52, b = 46.84, loss = 0.45
step 5000: w = 4.51, b = 46.87, loss = 0.45
NumPy's exact answer: [ 4.51 46.87]
prediction for 7 hours: 78.5The loss falls from 3,987 to 0.45, and the weights settle on exactly the answer NumPy's exact solver gives: each extra hour of study is worth about 4.5 marks, starting from about 47. You have just trained a linear regression model — the same loop, scaled up to billions of weights and run on GPUs, trains a language model.
What changes in real training
- Stochastic gradient descent (SGD): instead of computing the gradient on all the data every step, use a small random mini-batch. Each step is noisier but far cheaper.
- Epoch: one pass over the whole dataset. Training usually runs for several.
- Optimisers such as Adam adapt the step size for each weight automatically. You still choose a starting learning rate.
- Local minima: a real loss landscape has many valleys. In practice, for large networks, gradient descent still finds weights that work well.
Try it yourself
- In the first example, find the largest learning rate that still converges. (Hint: try values between 0.9 and 1.0.)
- In the line fit, try
learning_rate = 0.05and then0.07. Why does one still work while the other ends innan? - Add a seventh student who studied 8 hours and got 90. How do w and b change?