Chapter 2 · Vectors and Matrices
Matrix Multiplication and Shapes
- Page 5 of 17
- 3 min read
Multiplying two matrices is the work that GPUs were built to do fast, and it is where most of a model's computing time goes. You will almost never do it by hand, but you must know the one rule that decides whether it works.
The shape rule
(m × n) @ (n × p) → (m × p)
inner sizes must match: n and nimport numpy as np
A = np.ones((2, 3)) # 2 × 3
B = np.ones((3, 4)) # 3 × 4
print((A @ B).shape) # inner sizes match (3 and 3): the result is 2 × 4
try:
B @ A # 3 × 4 times 2 × 3: inner sizes 4 and 2 differ
except ValueError as error:
print("ValueError:", error)(2, 4)
ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0, with gufunc signature (n?,k),(k,m?)->(n?,m?) (size 2 is different from 4)The error message is long, but the end says it all: size 2 is different from 4. When you meet it, print the .shape of both sides, write them next to each other, and check the inner sizes. Often the fix is a transpose (.T).
Each element of the result is a dot product: row i of the left matrix with column j of the right one. Unlike ordinary numbers, the order matters — A @ B is usually not the same as B @ A, and often one of them is not even allowed.
A whole batch in one step
Models do not process one example at a time. They stack many inputs into a matrix, one per row, and push the whole batch through a layer with a single matrix multiplication:
import numpy as np
W = np.array([[0.2, 0.8, -0.5],
[1.0, -0.3, 0.4]]) # the same 2 × 3 layer as before
b = np.array([0.1, -0.2])
X = np.array([[0.5, -1.0, 2.0], # a batch: 4 inputs, one per row
[1.0, 0.0, 0.0],
[0.0, 1.0, 1.0],
[2.0, 2.0, 2.0]])
Z = X @ W.T + b # (4 × 3) @ (3 × 2) = 4 × 2
print(Z.shape)
print(Z)(4, 2)
[[-1.6 1.4]
[ 0.3 0.8]
[ 0.4 -0.1]
[ 1.1 2. ]]The first row matches the single-input result from the previous page. Four inputs, one line of code, and on a GPU it costs about the same as one. Batching is why training and serving models is affordable.
An embedding lookup is a matrix product
A language model keeps one embedding row per token in a big matrix. Picking a token's row can be written as a one-hot vector (all zeros, one 1) times that matrix:
import numpy as np
vocab = ["the", "cat", "sat", "mat"]
E = np.array([[0.1, 0.3], # one row (embedding) per word
[0.9, 0.2],
[0.4, 0.8],
[0.7, 0.1]])
one_hot = np.array([0, 1, 0, 0]) # "cat"
print(one_hot @ E) # the matrix product picks out row 1
print(E[vocab.index("cat")]) # which is why libraries just index the row[0.9 0.2]
[0.9 0.2]Mathematically it is a product; in practice libraries simply read the row, which is much faster. When you read in a paper that "the input is multiplied by the embedding matrix", now you know it means "look up each token's vector".
Try it yourself
- What shape is
(5 × 2) @ (2 × 7)? Is(2 × 7) @ (5 × 2)allowed? - Make
A2 × 3 andB2 × 3 and multiplyA @ B.T. What shape comes out? - Look up the embedding for "mat" with a one-hot vector.