Milestone 04: The CNN Revolution (1998)

NoteMilestone Info

Architecture Milestone | Difficulty: ●●●○ | Time: 15–30 min (CIFAR-10 option: about 1.5 hours) | Prerequisites: Modules 01–07 and 09

TipWhat You’ll Learn
  • What weight sharing buys on small images (3× fewer parameters) and why the accuracy case needs larger ones
  • How weight sharing enables translation invariance
  • The hierarchical feature learning that powers all computer vision

Overview

This is the first Architecture Milestone. The Foundation Milestones proved your training loop learns; the next two prove your architectures do the work they were invented for. Here you pick up the Conv2d and MaxPool2d layers you just built in Module 09 and first validate them on TinyDigits, then optionally scale to natural images.

1998. Yann LeCun deploys LeNet-5: a convolutional neural network that reads handwritten zip codes for the US Postal Service and dollar amounts on bank checks for NCR. It is not a research demo. It is production software, sorting mail and clearing checks at industrial scale — the first commercial success of deep learning.

The breakthrough is structural. Images are not bags of pixels; nearby pixels matter more than distant ones, and the same edge detector works whether you put it in the corner or the center. Exploit those two facts — local connectivity and weight sharing — and a convolutional layer’s parameter count depends on its filter size, not on the size of the image.

You are about to reproduce those same principles using your own Conv2d and MaxPool2d from Module 09. The default milestone runs offline on TinyDigits; the CIFAR-10 script is the optional scale-up path when you want the natural-image benchmark.

What You’ll Recreate

CNNs that exploit image structure:

  1. TinyDigits — compare your CNN with Milestone 03’s MLP on the same 8×8 images
  2. CIFAR-10 — scale to natural color images (32×32 RGB)
Images --> Conv2d(1->8, k=3) --> ReLU --> MaxPool2d(2,2) --> Flatten(72) --> Linear(72->10) --> Classes

Prerequisites

Table 1 lists the modules you need to have completed before starting.

Table 1: Prerequisite modules for the CNN milestone.
Module Component What It Provides
01–07 Foundation Tensors through your optimizer (the script runs its own training loop, so Module 08 is not required)
09 Convolutions Your Conv2d + MaxPool2d

Running the Milestone

Before running, ensure you have completed Modules 01–07 and 09. You can check your progress:

tito module status
tito milestone run 04            # TinyDigits, the default: 86-87% after 50 epochs, about 2 minutes
tito milestone run 04 --part 2   # optional scale-up to CIFAR-10 (downloads the dataset; not needed to complete 04)

Expected Results

Table 2 records the accuracy and runtime you should expect to see.

Table 2: Expected accuracy for the CNN milestone on TinyDigits and CIFAR-10. The CIFAR-10 row is a single run of 02_lecun_cifar10.py --quick-test. The flag only subsets the data after the full ~170 MB download; if CIFAR-10 is not on disk and is not downloaded, the script exits with status 2 (--test-only checks the architecture without any download). Without the flag the script trains on up to 100 batches of 32 images per epoch and scores the first 2,000 test images (20 batches of 100), which would take several hours, because Module 09’s convolution runs as Python loops. Module 17’s im2col convolution does the same arithmetic thousands of times faster.
Script Dataset Architecture Accuracy vs MLP
01 (TinyDigits) 1K train, 8×8 Simple CNN, 810 parameters 86–87% MLP matches or beats it at equal epochs; CNN has 3× fewer parameters
02 (CIFAR-10), --quick-test 1,000 train and 500 test images cut from the full CIFAR-10 download (~170 MB), 32×32 RGB Deeper CNN, 3 epochs 37% in one measured run (chance is 10%) Optional scale-up; about 1.5 hours of training on an Apple M5 Max

Pass Criteria

Part 1 (TinyDigits) is the required part. Before training, it checks YOUR CrossEntropyLoss on one batch against a NumPy computation and stops if the two disagree, because training can still converge (the gradient comes from Module 06’s backward) while every printed loss is wrong. After training, it passes only when all three of these hold, and exits with status 1 otherwise:

  • Test accuracy is at least 75%.
  • YOUR Conv2d delivered a non-zero gradient to its filters during training.
  • The filters moved at least 25% of their initial norm from their random start.

Accuracy alone cannot tell whether the convolution learned: with the filter gradients zeroed, the Linear head still reaches about 81% on random filters. A correct run moves the filters about 76%.

Part 2 (CIFAR-10) is an optional extension, recorded separately. It passes when every conv and linear weight moved from its initial value and test accuracy is at least 25% (20% with --quick-test; chance is 10%). These floors are not yet calibrated on the full dataset. --test-only is a smoke check on synthetic data that trains nothing and records nothing.

The Aha Moment: Structure Matches Reality

An MLP sees an image as 3,072 unrelated numbers. It does not know that pixel (0,0) is next to pixel (0,1). It learns brittle correlations like “if pixel 1,234 is bright and pixel 2,891 is dark…” — patterns tied to absolute positions, which fall apart the moment the cat shifts a few pixels to the left.

A CNN bakes spatial structure into the architecture itself:

  1. Local connectivity — each neuron only looks at a small neighborhood (3×3 or 5×5). Edges, corners, and textures are local patterns; the network does not need a global view to detect them.
  2. Weight sharing — one filter scans the entire image. “Cat in the top-left” and “cat in the bottom-right” trigger the same feature detector, so the network learns the concept once instead of 1,024 times.
  3. Translation invariance — pooling makes the output insensitive to small shifts. The network learns that a cat is present, not where the pixels happened to land.

On 8×8 TinyDigits that structure buys parameters, not accuracy. Trained for the same number of epochs, Milestone 03’s 2,410-parameter MLP is as accurate as your 810-parameter CNN, and both approach 90 percent after 100 epochs. With only 64 pixels, a dense layer can afford to connect everything.

The balance shifts as images grow. A dense layer’s weights scale with the number of input values; a convolution’s do not. On CIFAR-10’s 32×32 color images, a 32-unit dense first layer needs 3,072 × 32 = 98,304 weights, while eight 3×3 filters over three channels need 216. That is the constraint LeNet was designed around.

Expect the CNN to train far more slowly than the MLP, about 87 times longer per epoch on the laptop we measured. Module 09’s convolution runs in explicit Python loops, which is exactly the overhead its im2col discussion shows how to remove.

The default milestone validates your implementations on TinyDigits, and that part alone completes Milestone 04. The optional Part 2 scales them: 50,000 natural color images, 32×32×3 = 3,072 dimensions per image, 10 categories (airplanes, cars, birds, cats, ships…). This is the hard problem.

Your DataLoader streams batches from disk. Your Conv2d layers extract features hierarchically — first layer finds edges, second finds textures, third finds object parts. Your MaxPool2d shrinks the spatial map while preserving what matters.

When the optional CIFAR-10 run prints its test accuracy, sit with it for a second, even if it is well short of LeNet’s era-defining numbers. The quick run reached 37% on ten classes where guessing gets 10%. You did not download a pretrained model. Every tensor op, every gradient, every parameter update is code you wrote — running on your laptop, training a model that would have made headlines twenty years ago. That is systems engineering.

Your Code Powers This

Table 3 names the TinyTorch components that power this milestone.

Table 3: TinyTorch components that power the CNN milestone.
Component Your Module What It Does
Tensor Module 01 Stores images and feature maps
Conv2d Module 09 Your convolutional layers
MaxPool2d Module 09 Your pooling layers
ReLU Module 02 Your activation functions
Linear Module 03 Your classifier head
CrossEntropyLoss Module 04 Your loss computation
DataLoader Module 05 Your batching pipeline
backward() Module 06 Your autograd engine

Historical Context

LeNet-5 was deployed for zip code recognition at the US Postal Service and check-amount reading at NCR — the first neural networks to ship in production at meaningful scale.

CIFAR-10 (2009) became the standard pre-ImageNet benchmark. Reaching 70%+ on it was the signal the field was ready for the next jump in scale.

The 2012 “ImageNet moment” — AlexNet — applied the same CNN principles to 1.2 million images on GPUs. The blueprint was already in LeCun’s 1998 paper. The hardware just had to catch up.

Systems Insights

  • Memory: 810 parameters on TinyDigits, 3,240 bytes of float32; weight sharing keeps the count fixed as images grow
  • Compute: 3,312 multiply-accumulates per image against the MLP’s 2,368; weight reuse trades stored weights for arithmetic
  • Architecture: Hierarchical feature learning (edges → textures → objects)

Read why

The companion book’s chapter Synthesis II: Convolution, Attention, and Generated Text covers Milestones 04 and 05, including the matched-budget comparison with the MLP. Read it in TinyTorch: From Tensors to Transformers (PDF) after you have run the milestone.

What’s Next

CNNs are the right inductive bias for grid-structured data, but most of the world’s interesting signals are sequential — text, audio, time series. Milestone 05 introduces Transformers, the architecture that ate sequence modeling first and, eventually, vision itself.

Further Reading

Back to top