Chapter 6 · Project and Next Steps
From Python to Machine Learning: Your First Model
- Page 22 of 23
- 8 min read
On the last page you classified reviews in two ways: with keyword rules you wrote yourself, and with a language model. Machine learning is the third way, and the idea behind both of the others' modern forms: instead of writing the rules, you show the computer many examples with answers, and an algorithm finds the rules for you. Large language models are the same idea at enormous scale. This page trains your first models with scikit-learn, the standard Python library for classic machine learning, using everything you learned about NumPy and pandas.
The words you need
| Word | Meaning | In the example below |
|---|---|---|
Features (X) | The inputs the model looks at | Four measurements of a flower |
Label / target (y) | The answer it should learn to predict | Which of three species it is |
Training (fit) | Letting the algorithm find patterns in examples | 120 flowers with answers |
Prediction (predict) | Using what it learned on new inputs | 30 flowers it has never seen |
| Classification / regression | Predicting a category / a number | Species is a category |
When the examples come with answers, it is supervised learning — the most common kind. Finding groups in data without answers is unsupervised learning.
Step 1: look at the data
Install it with python -m pip install scikit-learn (a notebook from the previous page is a good place for this whole page). scikit-learn comes with small practice datasets. The classic one measures 150 iris flowers:
from sklearn.datasets import load_iris
iris = load_iris(as_frame=True)
df = iris.frame # a pandas DataFrame
print(df.shape)
print(df.head(3).to_string())
print(iris.target_names.tolist())
print(df["target"].value_counts().sort_index().tolist())(150, 5)
sepal length (cm) sepal width (cm) petal length (cm) petal width (cm) target
0 5.1 3.5 1.4 0.2 0
1 4.9 3.0 1.4 0.2 0
2 4.7 3.2 1.3 0.2 0
['setosa', 'versicolor', 'virginica']
[50, 50, 50]150 rows, four measurement columns and a target column: 0, 1 or 2 for the three species, 50 of each. Balanced classes like this make accuracy easy to read.
Step 2: split, train, predict, measure
import pandas as pd
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
iris = load_iris(as_frame=True)
df = iris.frame
X = df.drop(columns="target") # features: the four measurements
y = df["target"] # label: 0, 1 or 2
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
print(X_train.shape, X_test.shape)
model = DecisionTreeClassifier(max_depth=3, random_state=42)
model.fit(X_train, y_train) # learn from the training set
predictions = model.predict(X_test) # predict flowers it has never seen
print(predictions[:10])
print(y_test.to_numpy()[:10])
print(f"accuracy: {accuracy_score(y_test, predictions):.2f}")
new_flower = pd.DataFrame([[5.0, 3.4, 1.5, 0.2]], columns=X.columns)
print(iris.target_names[model.predict(new_flower)[0]])(120, 4) (30, 4)
[0 2 1 1 0 1 0 0 2 1]
[0 2 1 1 0 1 0 0 2 1]
accuracy: 0.97
setosatrain_test_splithides 20% of the rows — 30 flowers — before training.random_state=42makes the shuffle repeatable (page 9);stratify=ykeeps the three species in the same proportions in both parts.fittrains the model on the 120 training rows only. A decision tree learns a series of yes/no questions such as "is the petal shorter than 2.5 cm?".predictanswers for the 30 hidden flowers. The first ten predictions match the true labels exactly.- Accuracy is the share of correct answers: 0.97 means 29 of 30.
Every scikit-learn model works the same way: create it, fit(X_train, y_train), then predict(X) or score(X, y). Learn that pattern once and you can use hundreds of algorithms.
Step 3: why the test set matters — overfitting
Why hide data at all? Because a model can do perfectly on data it has seen and still fail on new data. Watch how the tree's depth — how many questions it may ask — changes the two scores:
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X, y = load_iris(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
for depth in [1, 3, None]:
model = DecisionTreeClassifier(max_depth=depth, random_state=42).fit(X_train, y_train)
print(f"max_depth={depth}: train {model.score(X_train, y_train):.2f}, test {model.score(X_test, y_test):.2f}")max_depth=1: train 0.67, test 0.67
max_depth=3: train 0.98, test 0.97
max_depth=None: train 1.00, test 0.93- Depth 1 is too simple: 0.67 on both. That is underfitting.
- No depth limit scores a perfect 1.00 on the training data — it has memorised every flower — yet it does worse on the test set than depth 3. That is overfitting, the same thing you saw in the training-curve chart on page 15.
- Depth 3 is the sweet spot here: nearly as good on new data as on old.
The rule that follows: never judge a model on the data it was trained on. Only the test score tells you how it will do in real use.
Step 4: compare models with pipelines
Many algorithms work better when every feature is on the same scale. A pipeline chains a preprocessing step and a model into one object that still has fit and predict, so the scaling is learned from the training data only:
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
X, y = load_iris(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
models = {
"decision tree": DecisionTreeClassifier(max_depth=3, random_state=42),
"logistic regression": make_pipeline(StandardScaler(), LogisticRegression()),
"nearest neighbours": make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=5)),
}
for name, model in models.items():
model.fit(X_train, y_train)
print(f"{name:20} {model.score(X_test, y_test):.2f}")decision tree 0.97
logistic regression 0.93
nearest neighbours 0.93With 30 test flowers, one mistake is a difference of 0.03, so these three are really about equally good. On a dataset this small, do not read much into small gaps — a lesson that applies just as much when you compare prompts or LLMs on a handful of examples.
Step 5: a first text classifier
Models need numbers, so text must be turned into numbers first. TfidfVectorizer does that by counting words (and down-weighting words that appear everywhere). Here it is with the review problem from the last page:
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
reviews = [
"great sound and battery", "love it, works perfectly", "excellent quality, fast delivery",
"very happy with this keyboard", "amazing value for money", "good product, would buy again",
"broke after two days", "terrible support, never again", "stopped charging, very poor",
"arrived late and damaged", "waste of money", "bad quality, keys stopped working",
]
labels = ["positive"] * 6 + ["negative"] * 6
model = make_pipeline(TfidfVectorizer(), LogisticRegression())
model.fit(reviews, labels)
new_reviews = ["the battery is great", "poor quality, it broke", "fast delivery, happy"]
for text, label in zip(new_reviews, model.predict(new_reviews)):
print(f"{label:9} <- {text}")positive <- the battery is great
negative <- poor quality, it broke
positive <- fast delivery, happyTwelve examples are far too few for real use — the model only knows the words it has seen — but the shape is exactly what a production text classifier looks like. Embeddings from a language model (page 13) are a much richer way of turning text into numbers, and you can feed them to these same scikit-learn models.
The machine-learning workflow
data → clean & explore → features (X) and label (y) → train/test split
→ fit on train → evaluate on test → improve → use it on new dataThree mistakes to avoid from day one:
- Testing on the training data. It always looks great and tells you nothing.
- Leakage: letting information from the test set, or from the answer itself, slip into the features — for example scaling with statistics from the whole dataset before splitting. Pipelines prevent the common cases.
- Trusting accuracy on unbalanced data. If 95% of emails are not spam, a model that always says "not spam" scores 0.95 and is useless. Look at the classes separately.
Where this goes next
To understand why these models work — vectors, distances, probability, averages and spread — the next Level 0 topic is Maths and Statistics essentials. After Level 0, the ML & Data path goes deep into these models, and the AI Engineer path uses them alongside LLMs.
Try it yourself
- Change
test_sizeto 0.3 andrandom_stateto 7. Do the scores change? Why does that suggest you should not trust one split too much? - Load
load_winefromsklearn.datasetsand repeat steps 1 to 4 on it. - Add ten of your own reviews (in English or Bangla) to the text classifier and test it on five new ones.