Chapter 5 · Probability
Probability Inside a Language Model
- Page 15 of 17
- 3 min read
Everything on the last few pages comes together inside a language model. At every step it reads the text so far and produces a score for every token in its vocabulary — tens of thousands of them. Softmax (page 8) turns the scores into probabilities. Then it samples one token, adds it to the text and repeats. The settings you pass to an API — temperature, top-p — change only this last step.
Temperature
Imagine the text so far is "The capital of Bangladesh is". These are made-up scores (logits) for five candidate words. Temperature divides the logits before softmax:
import numpy as np
words = ["Dhaka", "Chittagong", "Sylhet", "Paris", "banana"]
logits = np.array([6.0, 4.5, 4.0, 2.0, -1.0]) # the model's raw scores for the next word
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
for temperature in [0.5, 1.0, 2.0]:
probs = softmax(logits / temperature)
print(f"T={temperature}: " + " ".join(f"{w} {p:.2f}" for w, p in zip(words, probs)))T=0.5: Dhaka 0.94 Chittagong 0.05 Sylhet 0.02 Paris 0.00 banana 0.00
T=1.0: Dhaka 0.73 Chittagong 0.16 Sylhet 0.10 Paris 0.01 banana 0.00
T=2.0: Dhaka 0.50 Chittagong 0.24 Sylhet 0.18 Paris 0.07 banana 0.02- Low temperature (0.5) sharpens the distribution: "Dhaka" takes 94%. Output becomes focused and repeatable — good for extraction, classification and code.
- Temperature 1 uses the model's own probabilities.
- High temperature (2.0) flattens it: even "Paris" and "banana" get a real chance. Output becomes more varied and creative — and more likely to be wrong.
- Temperature 0 means always pick the top token (called greedy decoding). Even then, answers are not always perfectly identical, for technical reasons inside the servers.
Sampling, top-k and top-p
Sampling means drawing a token at random according to those probabilities. Over 1,000 draws at temperature 1, the counts follow the probabilities closely:
import numpy as np
from collections import Counter
words = ["Dhaka", "Chittagong", "Sylhet", "Paris", "banana"]
logits = np.array([6.0, 4.5, 4.0, 2.0, -1.0])
probs = np.exp(logits - logits.max())
probs /= probs.sum()
rng = np.random.default_rng(9)
picks = rng.choice(words, size=1000, p=probs) # sampling: what the model "says" 1,000 times
print(Counter(picks.tolist()).most_common())
top_k = 2 # top-k: keep only the k most likely words
keep = np.argsort(probs)[::-1][:top_k]
print("top-2 candidates:", [words[i] for i in keep])[('Dhaka', 721), ('Chittagong', 177), ('Sylhet', 90), ('Paris', 11), ('banana', 1)]
top-2 candidates: ['Dhaka', 'Chittagong']One draw in a thousand said "banana". Over a long answer, those rare bad picks add up, so two settings cut off the unlikely tail before sampling:
- Top-k keeps only the k most likely tokens.
- Top-p (nucleus sampling) keeps the smallest set of top tokens whose probabilities add up to p — for example 0.9. When the model is sure, that is one or two tokens; when it is unsure, more.
What this explains
- Why the same prompt gives different answers: it is sampling, by design.
- Why models sound confident when wrong: the probabilities are about which text is likely, not which fact is true (Level 0, topic 1).
- Log-probabilities: some APIs return the log of each token's probability. A very low value marks a word the model was unsure about — a useful signal for spotting possible mistakes.
- Training is cross-entropy (page 8) on the real next token, averaged over trillions of tokens.
Try it yourself
- Run the temperature example with T = 0.1 and T = 5. Describe what happens to "Dhaka".
- Implement top-p: sort the probabilities, use
np.cumsum, and keep tokens until the total passes 0.9. Which words remain? - Sample 1,000 times at T = 2. How often does "banana" appear now?