Module 09: Convolutions
A convolution is a tiled matmul with structured data reuse. That reuse is what lets vision kernels hit near-peak FLOPS on hardware that would stall at fully-connected equivalents. The inductive bias of shared weights is the pedagogical story. The memory access pattern is the systems story.
ARCHITECTURE TIER | Difficulty: ●●●○ | Time: 6-8 hours | Prerequisites: 01-08
Prerequisites: Modules 01-08 means you should have:
- Built the complete training pipeline (Modules 01-08)
- Implemented DataLoader for batch processing (Module 05)
- Comfort with parameter initialization, forward/backward passes, and optimization
If you can train an MLP on TinyDigits using your training loop and DataLoader, you’re ready.
Overview
You have just finished the foundations tier. You can train an MLP, run optimization, move batches through a working pipeline. Welcome to the Architecture Tier, where the question changes from can the network learn? to what should the network’s structure assume about the data?
Convolution is the answer for images. A photograph has structure that a generic MLP throws away: neighboring pixels matter together, the same edge can appear anywhere in the frame, and shifting a cat one pixel to the left should not require relearning what a cat looks like. A convolutional layer bakes those assumptions in — locality, weight sharing, and translation equivariance — and that structural prior is the reason CNNs dominated computer vision for a decade.
In this module you implement Conv2d as explicit nested loops and read MaxPool2d and AvgPool2d, which follow the same pattern. No vectorization tricks, no GPU kernel calls. The loops are the point: they expose where every multiply-accumulate goes and why a single forward pass on a \(224{\times}224\) batch costs billions of operations. Once you have felt that cost in code, the optimizations real frameworks pile on top — im2col, Winograd, cuDNN — stop feeling like magic.
Commands
# first time
tito module start 09
# later sessions
tito module resume 09
# when your tests pass
tito module complete 09Your notebook is modules/09_convolutions/convolutions.ipynb.
Learning objectives
- Implement Conv2d with explicit 7-nested loops revealing O(B×C×H×W×K²×C_in) computational complexity
- Master spatial dimension calculations with stride, padding, and kernel size interactions
- Understand receptive fields, parameter sharing, and translation equivariance in CNNs
- Analyze memory vs computation trade-offs: 2x2 pooling cuts the spatial area 4x (each dimension 2x) while keeping the strongest features
- Connect your implementations to production CNN architectures like ResNet and VGG
What you’ll build
The pattern you’ll enable:
# Building a CNN block
conv = Conv2d(3, 64, kernel_size=3, padding=1)
pool = MaxPool2d(kernel_size=2, stride=2)
x = Tensor(image_batch) # (32, 3, 224, 224)
features = pool(ReLU()(conv(x))) # (32, 64, 112, 112)What you’re not building yet
To keep this module focused, you will not implement:
- Dilated convolutions (PyTorch supports this with
dilationparameter) - Grouped convolutions (that’s for efficient architectures like MobileNet)
- Depthwise separable convolutions (advanced optimization technique)
- Transposed convolutions for upsampling (used in GANs and segmentation)
- Optimized implementations (cuDNN uses Winograd algorithm and FFT convolution)
You are building the foundational spatial operations. Advanced convolution variants and GPU optimizations come later.
What you write
The notebook arrives with the surrounding code already written and explained. You write 2 functions, each marked # YOUR CODE HERE and followed by a test cell:
Conv2d._convolve_loops- The core convolution: sliding window dot products over the input.
Conv2d.forward- Forward pass through Conv2d layer.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.core.spatial; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (14). Each prints a ✅ line when it passes.
- Conv2d Output Shape Computation
- Conv2d Padding
- Conv2d Convolution Loops
- Conv2d Forward (Composition)
- MaxPool2d Output Shape
- MaxPool2d Loops
- AvgPool2d Output Shape
- AvgPool2d Loops
- BatchNorm2d._validate_input
- BatchNorm2d._get_stats
- BatchNorm2d
- BatchNorm2d Gradients
- Pooling Operations
- SimpleCNN Integration
Integration tests after export (20).
tests/09_convolutions/test_09_convolutions_progressive.py
Completing this module unlocks Milestone 04, CNN Revolution (1998) (tito milestone run 04).
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
Expected: ... Got: [[[[5. 6.] [8. 9.]]]]-
From
test_unit_conv2d_convolve_loops. Each window’s sum was overwritten instead of accumulated. Useconv_sum += input_val * weight_valinside the kernel loops. ValueError: Conv2d expected 4D input (batch, channels, height, width), got 3D: (32, 8, 8)-
From calling the layer. A batch of grayscale images has no channel axis yet. Reshape to
(32, 1, 8, 8); the error message offers both this and the single-image reshape. ValueError: Conv2d expected 1 input channels, got 3-
From calling the layer. The input’s channel axis does not match the
in_channelsthe layer was built with. Check the layout is(batch, channels, height, width).
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Convolutions: Sliding Windows and the im2col Trick in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
Convolution is the structural prior for grids of pixels. The next data type — text — is not a grid. It is a discrete sequence with no fixed alphabet, no fixed length, and no notion of intensity. Before any model can read English you have to answer a more basic question: how do you turn a string into numbers a network can multiply against? In Module 10 you implement character and BPE tokenizers and meet the first real architectural decision of the language stack: vocabulary.
Next: Module 10: Tokenization
How later modules use this one
| Module | What it does | Your spatial ops in action |
|---|---|---|
| Milestone 04: CNN | LeNet-style CNN on TinyDigits | Stack your Conv2d, ReLU, and MaxPool2d for classification |
| Module 17: Acceleration | Optimize convolution | Replace loops with im2col and vectorized GEMM |