Chapter 3 · Calculus: How Models Learn
exp, log, Sigmoid, Softmax and Loss Functions
- Page 8 of 17
- 4 min read
A model produces raw scores. To use them you need two more tools: functions that turn scores into probabilities, and loss functions that measure how wrong a prediction is. Both are built from two functions you may remember from school: exp and log.
exp and log
import numpy as np
print(np.exp([0, 1, 2]).round(3)) # e⁰, e¹, e²
print(np.log([1, np.e, 100]).round(3)) # log undoes exp
print(round(np.log(0.9 * 0.8), 4), round(np.log(0.9) + np.log(0.8), 4)) # log turns × into +[1. 2.718 7.389]
[0. 1. 4.605]
-0.3285 -0.3285exp(x)= eˣ, with e ≈ 2.718. It is always positive and grows very fast — useful for turning any score into a positive number.log(the natural log) undoesexp: log(e) = 1, log(1) = 0. It turns multiplication into addition. Multiplying many small probabilities gives numbers too tiny for a computer, so models add their logs instead.
Sigmoid: a score into a probability
The sigmoid squashes any number into the range 0 to 1, so it can be read as a probability of "yes". A score of 0 means 50/50; large positive scores approach 1:
sigmoid(z) = 1 / (1 + e⁻ᶻ)Softmax: scores for many classes
With more than two choices, softmax turns a list of scores (called logits) into probabilities that are all positive and add up to 1. A bigger score gets a bigger share. This is the last step of every classifier — and of every language model choosing its next word (page 15).
import numpy as np
def sigmoid(z):
return 1 / (1 + np.exp(-z))
def softmax(z):
e = np.exp(z - np.max(z)) # subtracting the max avoids overflow; the result is the same
return e / e.sum()
print(sigmoid(np.array([-4, -1, 0, 1, 4])).round(3))
scores = np.array([2.0, 1.0, 0.1]) # raw scores ("logits") for three classes
probs = softmax(scores)
print(probs.round(3), probs.sum().round(3))[0.018 0.269 0.5 0.731 0.982]
[0.659 0.242 0.099] 1.0Loss functions: how wrong is the model?
Training needs one number to make smaller. The two you will see most often:
- Mean squared error (MSE) for predicting numbers (prices, marks): the average of the squared differences. Squaring makes big mistakes count much more than small ones.
- Cross-entropy (also called log loss) for predicting classes: minus the log of the probability the model gave to the right answer.
import numpy as np
# Mean squared error: for predicting numbers
actual = np.array([3.0, 5.0, 2.0])
predicted = np.array([2.5, 5.5, 4.0])
print("MSE:", np.mean((predicted - actual) ** 2))
# Cross-entropy (log loss): for predicting a class with a probability
# The true answer is class 0. How much loss for each prediction?
for p_correct in [0.9, 0.6, 0.1, 0.01]:
print(f"model gave the right class p = {p_correct:<4}: loss = {-np.log(p_correct):.2f}")MSE: 1.5
model gave the right class p = 0.9 : loss = 0.11
model gave the right class p = 0.6 : loss = 0.51
model gave the right class p = 0.1 : loss = 2.30
model gave the right class p = 0.01: loss = 4.61Look at how cross-entropy behaves. Confident and right (0.9): almost no loss. Unsure (0.6): some loss. Confident and wrong — only 0.01 for the right class — costs 4.61, forty times more than being right. That is exactly what you want a model to learn: do not be confidently wrong. Language models are trained with this same loss, on the probability they gave to the real next token.
Try it yourself
- What is
sigmoid(0), and why does that make sense? - Add 10 to every score before softmax. Do the probabilities change? Why?
- What is the cross-entropy loss when the model gives the right class a probability of 1.0?