Module 09: Convolutions

A convolution is a tiled matmul with structured data reuse. That reuse is what lets vision kernels hit near-peak FLOPS on hardware that would stall at fully-connected equivalents. The inductive bias of shared weights is the pedagogical story. The memory access pattern is the systems story.

NoteModule Info

ARCHITECTURE TIER | Difficulty: ●●●○ | Time: 6-8 hours | Prerequisites: 01-08

Prerequisites: Modules 01-08 means you should have:

  • Built the complete training pipeline (Modules 01-08)
  • Implemented DataLoader for batch processing (Module 05)
  • Comfort with parameter initialization, forward/backward passes, and optimization

If you can train an MLP on TinyDigits using your training loop and DataLoader, you’re ready.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

You have just finished the foundations tier. You can train an MLP, run optimization, move batches through a working pipeline. Welcome to the Architecture Tier, where the question changes from can the network learn? to what should the network’s structure assume about the data?

Convolution is the answer for images. A photograph has structure that a generic MLP throws away: neighboring pixels matter together, the same edge can appear anywhere in the frame, and shifting a cat one pixel to the left should not require relearning what a cat looks like. A convolutional layer bakes those assumptions in — locality, weight sharing, and translation equivariance — and that structural prior is the reason CNNs dominated computer vision for a decade.

In this module you implement Conv2d as explicit nested loops and read MaxPool2d and AvgPool2d, which follow the same pattern. No vectorization tricks, no GPU kernel calls. The loops are the point: they expose where every multiply-accumulate goes and why a single forward pass on a \(224{\times}224\) batch costs billions of operations. Once you have felt that cost in code, the optimizations real frameworks pile on top — im2col, Winograd, cuDNN — stop feeling like magic.

Commands

# first time
tito module start 09

# later sessions
tito module resume 09

# when your tests pass
tito module complete 09

Your notebook is modules/09_convolutions/convolutions.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement Conv2d with explicit 7-nested loops revealing O(B×C×H×W×K²×C_in) computational complexity
  • Master spatial dimension calculations with stride, padding, and kernel size interactions
  • Understand receptive fields, parameter sharing, and translation equivariance in CNNs
  • Analyze memory vs computation trade-offs: 2x2 pooling cuts the spatial area 4x (each dimension 2x) while keeping the strongest features
  • Connect your implementations to production CNN architectures like ResNet and VGG

What you’ll build

Figure 1: A concrete convolution. A 4 by 4 input containing 1 through 16 is convolved with a 3 by 3 all-ones kernel, stride 1 and no padding, producing the 2 by 2 output 54, 63, 90, 99. The bias is zero here. TinyTorch computes Conv2d with explicit sliding-window loops, each filter spanning all input channels, and backward accumulates input, weight, and bias gradients.

The pattern you’ll enable:

# Building a CNN block
conv = Conv2d(3, 64, kernel_size=3, padding=1)
pool = MaxPool2d(kernel_size=2, stride=2)

x = Tensor(image_batch)  # (32, 3, 224, 224)
features = pool(ReLU()(conv(x)))  # (32, 64, 112, 112)

What you’re not building yet

To keep this module focused, you will not implement:

  • Dilated convolutions (PyTorch supports this with dilation parameter)
  • Grouped convolutions (that’s for efficient architectures like MobileNet)
  • Depthwise separable convolutions (advanced optimization technique)
  • Transposed convolutions for upsampling (used in GANs and segmentation)
  • Optimized implementations (cuDNN uses Winograd algorithm and FFT convolution)

You are building the foundational spatial operations. Advanced convolution variants and GPU optimizations come later.

What you write

The notebook arrives with the surrounding code already written and explained. You write 2 functions, each marked # YOUR CODE HERE and followed by a test cell:

Conv2d._convolve_loops
The core convolution: sliding window dot products over the input.
Conv2d.forward
Forward pass through Conv2d layer.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.spatial;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (14). Each prints a ✅ line when it passes.

  • Conv2d Output Shape Computation
  • Conv2d Padding
  • Conv2d Convolution Loops
  • Conv2d Forward (Composition)
  • MaxPool2d Output Shape
  • MaxPool2d Loops
  • AvgPool2d Output Shape
  • AvgPool2d Loops
  • BatchNorm2d._validate_input
  • BatchNorm2d._get_stats
  • BatchNorm2d
  • BatchNorm2d Gradients
  • Pooling Operations
  • SimpleCNN Integration

Integration tests after export (20).

  • tests/09_convolutions/test_09_convolutions_progressive.py

Completing this module unlocks Milestone 04, CNN Revolution (1998) (tito milestone run 04).

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

Expected: ... Got: [[[[5. 6.] [8. 9.]]]]
From test_unit_conv2d_convolve_loops. Each window’s sum was overwritten instead of accumulated. Use conv_sum += input_val * weight_val inside the kernel loops.
ValueError: Conv2d expected 4D input (batch, channels, height, width), got 3D: (32, 8, 8)
From calling the layer. A batch of grayscale images has no channel axis yet. Reshape to (32, 1, 8, 8); the error message offers both this and the single-image reshape.
ValueError: Conv2d expected 1 input channels, got 3
From calling the layer. The input’s channel axis does not match the in_channels the layer was built with. Check the layout is (batch, channels, height, width).

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Convolutions: Sliding Windows and the im2col Trick in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

NoteUp next: Module 10, Tokenization

Convolution is the structural prior for grids of pixels. The next data type — text — is not a grid. It is a discrete sequence with no fixed alphabet, no fixed length, and no notion of intensity. Before any model can read English you have to answer a more basic question: how do you turn a string into numbers a network can multiply against? In Module 10 you implement character and BPE tokenizers and meet the first real architectural decision of the language stack: vocabulary.

Next: Module 10: Tokenization

How later modules use this one

Table 1: How spatial ops feed into later milestone and optimization modules.
Module What it does Your spatial ops in action
Milestone 04: CNN LeNet-style CNN on TinyDigits Stack your Conv2d, ReLU, and MaxPool2d for classification
Module 17: Acceleration Optimize convolution Replace loops with im2col and vectorized GEMM
Back to top