Chapter 2 · Inside a Large Language Model
What Is a Large Language Model?
- Page 3 of 8
- 2 min read
A large language model (LLM) is a deep learning model trained on a very large amount of text to do one deceptively simple job: given some text, predict what comes next.
Autocomplete, at an enormous scale
Your phone keyboard suggests the next word. An LLM does the same thing — but it has learned from a vast share of the written web, books and code, and it has billions of parameters to store the patterns it found: grammar, facts, styles of reasoning, how code is structured.
When you ask "What is the capital of Bangladesh?", the model is not looking it up. It continues the text with the answer that its training made most likely: Dhaka. Most of the time the most likely continuation is also the correct one — which is why LLMs are so useful, and also why they are sometimes confidently wrong (chapter 3).
What "large" means
- Large data — trillions of tokens of text.
- Large models — billions of parameters.
- Large compute — training runs on thousands of specialised chips.
The architecture: the Transformer
Almost every modern LLM is built on the Transformer, introduced in the 2017 paper Attention Is All You Need. Its key idea is attention: when processing a word, the model weighs how relevant every other word in the text is to it. In "The bank of the river was muddy", attention lets "bank" draw meaning from "river", so the model treats it as a riverbank, not a money bank.
LLMs you have heard of
ChatGPT (OpenAI), Claude (Anthropic), Gemini (Google), Llama (Meta) and many others are all LLMs, trained and tuned differently. The ideas in this tutorial apply to all of them.
Key takeaways
- An LLM predicts the next piece of text, repeatedly.
- It produces the most likely continuation, which is usually — not always — the correct one.
- The Transformer and its attention mechanism are what make modern LLMs work.