Chapter 4 · Python for Data
NumPy: Fast Maths on Arrays
- Page 13 of 23
- 5 min read
Machine learning is maths on large tables of numbers: an image is a grid of pixel values, a sentence becomes a list of numbers called an embedding, a model is millions of numbers called weights. Doing that maths with Python loops would be far too slow. NumPy — "Numerical Python" — gives you the array: a fast block of numbers you work on all at once. pandas, scikit-learn, PyTorch and nearly every other data library are built on its ideas.
Install it with python -m pip install numpy. By convention it is always imported as np.
Arrays
import numpy as np
scores = np.array([72, 85, 90, 64, 78])
print(scores)
print(scores.shape, scores.dtype)
print(scores + 5) # added to every element
print(scores * 1.1)
print(scores.mean(), scores.max(), scores.std().round(2))[72 85 90 64 78]
(5,) int64
[77 90 95 69 83]
[79.2 93.5 99. 70.4 85.8]
77.8 90 9.22Every array has a shape (its size in each direction) and a dtype (the type of all its elements — one type for the whole array, which is part of why it is fast). Operations apply to every element at once; this is called vectorisation. No loop is written, and the work runs in optimised C code — typically tens to hundreds of times faster than a Python loop over a list.
Arrays are not lists
import numpy as np
prices = [100, 200, 300]
print(prices * 2) # a list repeats itself
print(np.array(prices) * 2) # an array multiplies each element[100, 200, 300, 100, 200, 300]
[200 400 600]Two dimensions: rows and columns
A 2-D array is a table. Index it with [row, column]; a colon : means "all":
import numpy as np
# 3 students × 4 tests
marks = np.array([
[70, 82, 91, 65],
[88, 79, 94, 90],
[55, 61, 72, 58],
])
print(marks.shape) # (rows, columns)
print(marks[1]) # row 1: the second student
print(marks[:, 2]) # column 2: every student's third test
print(marks[0, 3]) # row 0, column 3
print(marks.mean(axis=1)) # average per student (across columns)
print(marks.mean(axis=0)) # average per test (down the rows)(3, 4)
[88 79 94 90]
[91 94 72]
65
[77. 87.75 61.5 ]
[71. 74. 85.66666667 71. ]The axis argument confuses everyone at first. Think of it as "the direction that gets collapsed": axis=0 squashes the rows together (one result per column), axis=1 squashes the columns (one result per row).
Boolean masks: filtering without loops
A comparison on an array gives an array of True/False — a mask. Indexing with a mask keeps only the matching elements. This is how you keep a model's confident predictions:
import numpy as np
confidences = np.array([0.91, 0.42, 0.77, 0.30, 0.88])
labels = np.array(["cat", "dog", "cat", "bird", "dog"])
sure = confidences > 0.75 # an array of True/False
print(sure)
print(labels[sure]) # keep only where True
print(sure.sum(), "confident predictions")
print(np.where(confidences > 0.75, labels, "unsure"))[ True False True False True]
['cat' 'cat' 'dog']
3 confident predictions
['cat' 'unsure' 'cat' 'unsure' 'dog']Broadcasting and creating arrays
When shapes differ, NumPy broadcasts the smaller array across the bigger one, as long as the sizes line up from the right. Here a row of two bonuses is applied to every student. You will also often create arrays rather than type them:
import numpy as np
marks = np.array([[70, 82], [88, 79], [55, 61]])
bonus = np.array([5, 0]) # +5 on the first test only
print(marks + bonus) # the row is "stretched" over every student
rng = np.random.default_rng(42) # a seeded random generator
noise = rng.normal(0, 1, size=(2, 3)).round(2)
print(noise)
print(np.zeros((2, 3)), np.arange(0, 10, 3), np.linspace(0, 1, 5))[[75 82]
[93 79]
[60 61]]
[[ 0.3 -1.04 0.75]
[ 0.94 -1.95 -1.3 ]]
[[0. 0. 0.]
[0. 0. 0.]] [0 3 6 9] [0. 0.25 0.5 0.75 1. ]np.random.default_rng(42) is the modern way to get random numbers; the seed makes them repeatable, exactly as on page 9.
Why this matters for AI: similarity of meaning
AI systems turn text into embedding vectors, placed so that texts with similar meanings point in similar directions. Cosine similarity measures how closely two vectors point the same way: 1.0 means the same direction, near 0 means unrelated. It is the heart of semantic search and RAG (retrieval-augmented generation). With NumPy it is one line — @ is the dot product:
import numpy as np
# Pretend embeddings: real ones have hundreds or thousands of numbers.
sentences = ["How do I reset my password?", "I forgot my login password", "What is the price of the pro plan?"]
vectors = np.array([
[0.90, 0.10, 0.05],
[0.85, 0.20, 0.10],
[0.05, 0.10, 0.95],
])
def cosine_similarity(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
query = vectors[0]
for sentence, vector in zip(sentences, vectors):
print(f"{cosine_similarity(query, vector):.3f} {sentence}")1.000 How do I reset my password?
0.991 I forgot my login password
0.118 What is the price of the pro plan?The two password questions score close to 1 even though they share few words; the pricing question scores low. Replace these three made-up vectors with real embeddings from an API and this is a working search engine for your documents.
Try it yourself
- Make an array of 7 daily temperatures and print the days (indexes) above the average.
- Create a 3×3 array of the numbers 1–9 with
np.arange(1, 10).reshape(3, 3)and print its column sums. - Add a fourth sentence and vector to the similarity example, and find the most similar sentence to it with
argmax().