Chapter 2 · Inside a Large Language Model
Tokens: How an LLM Reads Text
- Page 4 of 8
- 2 min read
An LLM never sees letters or words directly. Text is first cut into tokens — small pieces that may be a whole word, part of a word, a space plus a word, or a punctuation mark — and each token becomes a number the model can compute with.
See it for yourself
OpenAI publishes its tokenizer as the Python library tiktoken:
# pip install tiktoken
import tiktoken
enc = tiktoken.get_encoding("o200k_base") # the tokenizer used by GPT-4o
for text in ["Artificial intelligence is changing how we work.",
"unbelievable"]:
ids = enc.encode(text)
print(len(ids), [enc.decode([i]) for i in ids])
# 8 ['Artificial', ' intelligence', ' is', ' changing', ' how', ' we', ' work', '.']
# 3 ['un', 'bel', 'ievable']
Common words are one token; a less common word like "unbelievable" is split into pieces. You can also try the free OpenAI Tokenizer page in your browser.
A handy rule of thumb (English)
OpenAI's guide to tokens gives these estimates for English: 1 token ≈ 4 characters ≈ ¾ of a word, so 100 tokens ≈ 75 words. For exact counts, use the tokenizer.
Bangla uses more tokens
We measured the same sentence in both languages with the same tokenizer:
| Sentence | Characters | Tokens (o200k_base) | Tokens (older cl100k_base) |
|---|---|---|---|
| Artificial intelligence is changing how we work. | 48 | 8 | 9 |
| কৃত্রিম বুদ্ধিমত্তা আমাদের কাজের ধরন বদলে দিচ্ছে। | 49 | 20 | 59 |
Tokenizers learn their pieces from training text, which contains far more English than Bangla, so Bangla is split into smaller pieces. Newer tokenizers are much better (20 instead of 59), but the gap is still real.
Why tokens matter to you
- Cost — APIs charge per token, for both your input and the model's output.
- Limits — every model can only handle a certain number of tokens at once (its context window, page 7).
- Oddities — because the model sees pieces, not letters, tasks like counting letters in a word can trip it up.
Key takeaways
- Models read tokens, not words: whole words, word pieces, spaces and punctuation.
- English: about 4 characters or ¾ of a word per token.
- The same message in Bangla usually costs more tokens.