Chapter 1 · What AI Really Is
How Machines Learn from Data
- Page 2 of 8
- 2 min read
"The model learned it" sounds mysterious. It is not. Learning, for a machine, means adjusting numbers until its answers on known examples become good — then using those numbers on new inputs.
The three ingredients
- Data — many examples. For a house-price model: size, location, price. For a language model: an enormous amount of text.
- A model — a mathematical function with adjustable numbers called parameters (or weights). Large language models have billions of them.
- A training procedure — repeat: make a prediction, measure how wrong it was (the loss), nudge the parameters to be a little less wrong.
Training versus inference
| Training | Inference | |
|---|---|---|
| What happens | Parameters are adjusted from data | Parameters are fixed; the model produces an output |
| When | Before release — weeks or months on many GPUs | Every time you send a message |
| Does it learn from you? | Your chat does not change the model's parameters in the moment. (Whether providers later use conversations for training depends on their policy and your settings.) |
A tiny learning loop you can read
# Learn y = w * x from examples where the true w is 3.
data = [(1, 3), (2, 6), (3, 9), (4, 12)]
w = 0.0 # the single "parameter", starting wrong
learning_rate = 0.01
for step in range(200):
for x, y in data:
prediction = w * x
error = prediction - y # how wrong we are
w -= learning_rate * error * x # nudge w to be less wrong
print(round(w, 3)) # ~3.0 — learned from the examples, never written by hand
Real models do exactly this idea at enormous scale: billions of parameters instead of one, and text instead of four number pairs.
Three common kinds of learning
- Supervised — examples come with the correct answer (a label): "this email is spam".
- Unsupervised — no labels; the model finds structure, such as groups of similar customers.
- Self-supervised — the answer is hidden inside the data itself. Language models learn this way: hide the next word, ask the model to predict it. No human labelling needed, which is why they can learn from so much text.
Key takeaways
- Learning = adjusting parameters to reduce error on examples.
- Training happens before release; inference is what happens when you use the model.
- Language models learn self-supervised, by predicting hidden next words.