Chapter 2 · Inside a Large Language Model
From Prompt to Answer
- Page 5 of 8
- 2 min read
You type a question and words stream back. Here is what happens in between — and how a raw text predictor became a helpful assistant.
How an assistant is made
- Pretraining — the model learns to predict the next token over a huge amount of text. The result knows a lot but behaves like a document completer, not an assistant.
- Instruction tuning — further training on examples of instructions and good responses teaches it to follow requests.
- Learning from human feedback — people (and other models) compare answers; the model is tuned toward the ones judged helpful, honest and safe. This step is often called RLHF.
Generating an answer, one token at a time
At inference the model runs a loop. This sketch shows the idea (it is illustrative, not a real library):
tokens = tokenize(conversation) # your prompt, as token ids
while True:
probabilities = model(tokens) # a probability for every possible next token
next_token = sample(probabilities, temperature=0.7)
if next_token == END_OF_TURN:
break
tokens.append(next_token) # the answer so far becomes input
show(detokenize([next_token])) # why replies appear word by word
Every new token is chosen using everything before it — your prompt and what the model has already written. That is why an early mistake can carry through the rest of an answer.
Temperature: steady or creative
The model gives a probability to every candidate token. Temperature controls how it picks:
- Low (near 0) — almost always the most likely token. More predictable and repeatable; good for extraction, classification, code.
- Higher — less likely tokens get picked more often. More varied; good for brainstorming and creative writing.
Low temperature makes answers consistent, not correct: a model can be consistently wrong.
Key takeaways
- Pretraining gives knowledge; instruction tuning and human feedback make an assistant.
- Answers are generated token by token, each one depending on all the text before it.
- Temperature trades predictability for variety — it does not control accuracy.