Chapter 6 · Project and Next Steps
Jupyter Notebooks and Project Structure
- Page 19 of 22
- 5 min read
Data scientists and AI engineers work in two places. They explore in a notebook — load some data, try a prompt, draw a chart, change one line and run it again. Then they build in ordinary .py files organised as a project, which can be tested, reviewed and run on a server. This page covers both, and how to move from one to the other.
What a notebook is
A Jupyter notebook (a .ipynb file) is a document made of cells. A code cell holds Python; when you run it, its output — text, a table, a chart — appears right underneath and is saved in the file. A Markdown cell holds formatted notes, so the reasoning sits next to the code. Behind the notebook runs a kernel: one Python process that keeps every variable in memory between cells.
# Google Colab: nothing to install — open colab.research.google.com
# On your computer, inside the project's virtual environment (page 2):
python -m pip install jupyterlab
jupyter labVS Code opens .ipynb files too: choose your .venv as the kernel in the top-right corner. Google Colab (page 2) is the same idea in the browser, with free GPUs.
Working with cells
# Cell 1
import pandas as pd
reviews = pd.read_csv("../data/reviews.csv")
# Cell 2
reviews.shape # the last line of a cell is shown automatically
# Cell 3
reviews["rating"].value_counts()| Key / command | What it does |
|---|---|
Shift + Enter | Run the cell and move to the next one |
Esc then M / Y | Make the cell Markdown / code |
Esc then A / B | Add a cell above / below |
%pip install pandas | Install into the kernel's own environment (safer than !pip) |
%%time as the first line | Show how long the cell took |
| Kernel → Restart and Run All | Start fresh and run every cell from the top |
The hidden-state trap
Cells run in the order you run them, not the order they appear. The number beside a cell, like [7], tells you when it last ran. Suppose you run a cell that sets limit = 10, later change it to limit = 5 and run it again, then delete it. Every other cell still sees limit = 5 — a value that is no longer written anywhere in the notebook. Send that notebook to a colleague and it fails or, worse, gives different results.
- Before you share a notebook or trust its results, use Restart and Run All. If it runs from top to bottom, it is honest.
- Keep cells in the order they should run, and keep each one short.
- Never paste an API key into a cell: notebooks get shared and uploaded with their contents. Load keys from
.envas on page 17, or use Colab's Secrets panel.
Notebook or .py file?
| Use a notebook for | Use .py files for |
|---|---|
| Looking at new data for the first time | Code you will run again and again |
| Trying prompts and comparing model answers | Anything that runs on a server or on a schedule |
| Charts and a written report of findings | Functions and classes other code imports |
| Teaching and step-by-step demonstrations | Anything that needs tests and code review |
The usual path: explore in a notebook, and once a piece of code works, move it into a module and import it back into the notebook. These two lines make the notebook pick up your edits to the module without restarting:
%load_ext autoreload
%autoreload 2
from my_ai_tools import clean, estimate_tokensA real project layout
Here is how the review analyser you are about to build on the next page would look as a proper project:
review-analyser/
├── .venv/ virtual environment (never committed)
├── .env secret keys (never committed)
├── .gitignore
├── README.md
├── requirements.txt
├── data/
│ └── reviews.csv
├── notebooks/
│ └── 01-explore-reviews.ipynb
├── review_analyser/ your package (page 9)
│ ├── __init__.py
│ ├── classify.py
│ └── report.py
├── tests/
│ └── test_classify.py
└── main.py the entry point you run- One package (
review_analyser/) holds the real code, split into modules by job.main.pyonly reads arguments and calls it. data/holds input files,notebooks/holds exploration, numbered so their order is clear.requirements.txt(page 9) lets anyone recreate the environment.README.mdsays what the project does and how to run it..gitignorekeeps things that must never reach Git out of it:
.venv/
.env
__pycache__/
.ipynb_checkpoints/
*.logTests: let Python check your code
A test is a small function that runs your code and checks the answer with assert. The pytest tool finds every file called test_*.py in tests/ and runs every function starting with test_. Here are two tests for the package from page 9:
# tests/test_text.py
from my_ai_tools import clean, estimate_tokens
def test_clean_collapses_spaces():
assert clean(" a b ") == "a b"
def test_estimate_tokens_is_never_zero():
assert estimate_tokens("") == 1python -m pip install pytest
python -m pytest -q.. [100%]
2 passed in 0.01sTwo dots, two passes. Change clean() by mistake and the run turns red and tells you exactly which check failed. For AI code this is how you keep the parts that do not involve the model — parsing, cleaning, validation, cost limits — correct while you change everything around them.
Try it yourself
- Open a notebook (Colab or JupyterLab), load a CSV with pandas and draw one chart from page 15.
- Run a cell twice with a changed value, delete it, then use Restart and Run All and see what breaks.
- Create the project layout above for a small idea of your own, with a
.gitignore, a package and one passing test.