Chapter 2 · Vectors and Matrices
Dot Product and Cosine Similarity
- Page 3 of 17
- 3 min read
The dot product is the single most used operation in AI. Every layer of a neural network is built from dot products, and every semantic search ranks documents with one. It is simple: multiply matching elements, then add them up.
import numpy as np
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
print(a * b) # element by element
print(np.sum(a * b)) # add those up: the dot product
print(a @ b) # the same thing, written with @
print(np.dot(a, b)) # and again[ 4 10 18]
32
32
321×4 + 2×5 + 3×6 = 32. In NumPy, write a @ b.
What the dot product tells you
The dot product is large and positive when two vectors point the same way, zero when they are at right angles (they have nothing in common), and negative when they point in opposite directions. Its size also grows with the vectors' lengths. To compare direction only, divide by both lengths. The result is the cosine similarity, always between −1 and 1:
cosine similarity(a, b) = (a · b) / (‖a‖ × ‖b‖)import numpy as np
pairs = {
"same direction": (np.array([1, 1]), np.array([3, 3])),
"at right angles": (np.array([1, 0]), np.array([0, 5])),
"opposite": (np.array([1, 2]), np.array([-2, -4])),
}
for name, (a, b) in pairs.items():
cos = a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
angle = np.degrees(np.arccos(np.clip(cos, -1, 1)))
print(f"{name:16} cosine {cos:+.2f} angle {angle:.0f}°")same direction cosine +1.00 angle 0°
at right angles cosine +0.00 angle 90°
opposite cosine -1.00 angle 180°Notice [1, 1] and [3, 3] have cosine 1 even though one is three times longer. For text that is exactly what you want: a long document and a short one about the same topic should count as similar.
Semantic search in ten lines
This is the core of every RAG system and every "search by meaning" feature. Each document and the question are turned into embeddings; the documents are ranked by cosine similarity to the question:
import numpy as np
def cosine_similarity(a, b):
return (a @ b) / (np.linalg.norm(a) * np.linalg.norm(b))
# Made-up 4-number embeddings; real ones have 768 to 3,072 numbers
docs = {
"How do I get a refund?": np.array([0.90, 0.10, 0.05, 0.20]),
"Return an item for my money back": np.array([0.80, 0.20, 0.10, 0.25]),
"When does my parcel arrive?": np.array([0.10, 0.90, 0.15, 0.10]),
"Reset my password": np.array([0.05, 0.10, 0.95, 0.05]),
}
question = np.array([0.85, 0.15, 0.05, 0.22]) # "Can I get my money back?"
ranked = sorted(docs, key=lambda d: cosine_similarity(question, docs[d]), reverse=True)
for d in ranked:
print(f"{cosine_similarity(question, docs[d]):.3f} {d}")0.998 How do I get a refund?
0.995 Return an item for my money back
0.303 When does my parcel arrive?
0.136 Reset my password"Return an item for my money back" ranks near the top although it shares almost no words with "Can I get my money back?" — because their embeddings point the same way. That is the difference between searching by meaning and searching by keywords.
- Many embedding APIs return vectors already normalised to length 1. Then cosine similarity is simply the dot product, which is faster:
question @ doc. - For millions of documents, vector databases use clever indexes to find the top matches without comparing against every one (Python for AI, page 19, on why
O(n)matters).
Try it yourself
- Compute the dot product of
[2, 0, 1]and[1, 3, 2]by hand, then with@. - Normalise every document vector and the question, then rank with plain
@. Is the order the same? - Find two vectors with a cosine similarity of exactly 0.