Historical Milestones
Proof-of-Mastery Demonstrations | 7 Milestones | Prerequisites: Vary by milestone
Milestones are runnable recreations of historical ML breakthroughs that use YOUR TinyTorch implementations. Each one validates that the components you built across the modules can reproduce results that once made headlines.
Overview
You’ve spent the modules building a working ML framework: tensors, autograd, layers, optimizers, attention. The milestones answer the only question that matters: does it actually run the experiments that defined the field?
You’ll find out by rebuilding history. Each milestone reproduces a landmark result (Rosenblatt’s Perceptron, the XOR crisis, backpropagation, convolutional networks, transformers, MLPerf generative serving, custom kernels) using your code. When the Perceptron computes its first predictions, it’s running your tensor, layer, and activation stack. When attention processes a sequence, it’s running your multi-head attention on top of your transformer block. When a CNN recognizes TinyDigits, those are your convolutional layers extracting the features. When TinyGPT writes Shakespearean verse, it’s running your end-to-end generative language stack. And when your tiled matrix multiply and im2col convolution match NumPy on ragged shapes, you time them against compiled C++ SIMD, Apple Metal, and OpenAI Triton kernels that apply the same ideas below the Python runtime.
That makes these milestones evidence, to yourself and to anyone reading your repository, that the framework you built can run the experiments that defined the field.
The Journey
Table 1 traces the historical milestone timeline and the modules each one requires.
| Year | Milestone | The Landmark Experiment | Required Modules |
|---|---|---|---|
| 1958 | Perceptron | First neural network forward pass | 01-03 |
| 1969 | XOR Crisis | Experience the AI Winter trigger | 01-03 |
| 1986 | MLP Revival | Backprop solves XOR + digit recognition | 01-07 |
| 1998 | CNN Revolution | TinyDigits CNN, with optional CIFAR-10 scale-up | 01-07, 09 |
| 2017 | Transformer | Train TinyGPT on Shakespeare, TinyCopilot, and concept Q&A | 01-08, 10-13 |
| 2018 | MLPerf to Generative Serving | INT8 quantization, KV-cache serving, and Pareto frontier | 01-04, 06, 07, 09, 11-19 |
| 2024 | Custom Kernels | Verify your Module 17 kernels, then time them against bundled C++ SIMD, Metal, and Triton kernels | 01, 06, 09, 14, 17 |
Why Milestones Transform Learning
You’ll feel the historical struggle. When no setting of your single-layer perceptron’s weights gets past 75% on XOR, however many you try, you’ll understand in your bones why Minsky’s proof stalled neural-network research for a decade. The AI Winter wasn’t abstract skepticism; it was researchers watching their perceptrons fail in exactly the way yours just did.
You’ll experience the breakthrough. Then you add one hidden layer. Same data, same training loop. Suddenly: 100% accuracy. Loss collapses to zero. You didn’t just read about how depth unlocks non-linear representations; you watched your two-layer network solve what your one-layer network couldn’t. That’s lived experience, not summary.
You’ll build something real. By Milestone 04 you’re done with toy demos. You’re training a LeNet-style CNN on TinyDigits, extracting spatial features with your convolutional layers, and optionally scaling the same code path to CIFAR-10 with a network you wrote line by line, on a framework you wrote module by module.
How to Use Milestones
tito module status
tito milestone run 01Run milestones through tito rather than calling the script with python: tito milestone run checks the prerequisite modules and records each part’s result, and a direct run does neither. A milestone counts as complete only when every required part has passed (Parts 1 and 2 for Milestones 03, 05, and 06), so --part N alone records progress without completing a multi-part milestone. Completions of Milestones 03 through 06 recorded before this rule have to be earned again.
Each tinytorch/milestones/NN_yyyy_name/ folder contains:
README.md: full historical context and instructions- Python scripts: runnable demonstrations for each milestone part
Learning Philosophy
Module teaches: HOW to build the component
Milestone proves: WHAT you can build with it
Modules give you the parts. Milestones force the parts to do real work: the same work that, in each case, moved the field forward.
What’s Next?
Start at the beginning. Run tito milestone run 01 and watch a single-layer network, built on your tensor, Linear, and Sigmoid implementations, compute real predictions before any training infrastructure exists. From there the path is chronological: each milestone exposes the next constraint, then later milestones introduce the ideas that broke through it.
Build the future by understanding the past.