Chapter 2 · Vectors and Matrices
Matrices: Tables That Transform
- Page 4 of 17
- 3 min read
A matrix is a table of numbers: rows and columns. A dataset is a matrix (one row per example, one column per feature). A greyscale image is a matrix of pixels. And every layer of a neural network stores its knowledge in a matrix of weights.
import numpy as np
# 3 students × 4 tests: one row per student, one column per test
scores = np.array([
[72, 85, 90, 64],
[88, 79, 95, 70],
[55, 60, 58, 81],
])
print(scores.shape) # (rows, columns)
print(scores[1]) # one row: the second student
print(scores[:, 2]) # one column: the third test
print(scores.T.shape) # transpose: rows become columns(3, 4)
[88 79 95 70]
[90 95 58]
(4, 3)- The shape
(3, 4)means 3 rows and 4 columns. Always check shapes first when something goes wrong — most bugs in numeric code are shape bugs. scores[1]is a row;scores[:, 2]means "every row, column 2".- The transpose
scores.Tflips the table so rows become columns:(3, 4)becomes(4, 3).
Matrix times vector: many dot products at once
Multiplying a matrix by a vector takes the dot product of each row with the vector. The result has one number per row. This is exactly what one layer of a neural network does: each row of W is one "neuron", which weighs all the inputs and adds its own bias.
z = W x + bimport numpy as np
x = np.array([0.5, -1.0, 2.0]) # input: 3 features
W = np.array([ # weights: 2 outputs × 3 inputs
[0.2, 0.8, -0.5],
[1.0, -0.3, 0.4],
])
b = np.array([0.1, -0.2]) # one bias per output
z = W @ x + b # each output is a dot product plus a bias
print(z)
print(np.maximum(z, 0)) # ReLU: negative values become 0[-1.6 1.4]
[0. 1.4]First output: 0.2×0.5 + 0.8×(−1.0) + (−0.5)×2.0 + 0.1 = −1.6. Then an activation function such as ReLU (keep positives, turn negatives into 0) adds the bend that lets stacked layers learn curved patterns rather than only straight lines. A deep network is many of these steps, one after another.
A matrix as a transformation
There is a second way to see W x: the matrix moves the vector — rotates, stretches or squashes it into a new space. The layer above takes a point in 3-dimensional space and puts it in 2-dimensional space. Deep networks keep transforming their input this way until the answer becomes easy to read off. The 3Blue1Brown videos in the course show this beautifully.
Try it yourself
- Print the average of each test with
scores.mean(axis=0)and of each student withaxis=1. - Calculate the second output of the layer by hand and check it is 1.4.
- Change
xso that both outputs are negative. What does ReLU give?