Milestone 02: The XOR Crisis (1969)
Foundation Milestone | Difficulty: ●●○○ | Time: 30–45 min | Prerequisites: Modules 01–03 (Part 1) / Modules 01–04, 06, and 07 (Part 2)
- Why single-layer networks have fundamental mathematical limits
- How hidden layers enable non-linear decision boundaries
- Why “deep” learning is called DEEP
Overview
It’s 1969. Neural networks are the hottest thing in AI. Funding is pouring in. Then Marvin Minsky and Seymour Papert publish a 308-page mathematical proof that destroys everything: perceptrons cannot solve XOR. Not “struggle with” — CANNOT. Mathematically impossible.
Funding evaporates overnight. Research labs shut down. The field dies for 17 years — the infamous AI Winter.
You’re about to live that crisis. You’ll watch your own perceptron — built from your own modules — fail on four points despite trying many weight configurations. Accuracy never gets past 75%. Then you’ll add one hidden layer and watch the impossible collapse into the trivial.
What You’ll Recreate
Two demonstrations of perceptron limitations and the multi-layer solution:
- The Crisis — watch a perceptron fail on XOR due to linear boundary limits
- The Solution — add a hidden layer and solve the “impossible” problem
Crisis: Input --> Linear --> Output (FAILS)
Solution: Input --> Linear --> ReLU --> Linear --> Output (100%!)
The XOR Problem
Inputs Output
x1 x2 XOR
0 0 --> 0 (same)
0 1 --> 1 (different)
1 0 --> 1 (different)
1 1 --> 0 (same)
Plot those four points. The two zeros sit on one diagonal, the two ones on the other. No straight line separates them — and a single-layer perceptron can only draw straight lines. No amount of training fixes that. It’s geometry, not optimization.
Prerequisites
Table 1 lists the modules you need to have completed before starting.
| Module | Component | What It Provides | Required For |
|---|---|---|---|
| 01 | Tensor | YOUR data structure | Parts 1 & 2 |
| 02 | Activations | YOUR sigmoid/ReLU | Parts 1 & 2 |
| 03 | Layers | YOUR Linear layers | Parts 1 & 2 |
| 04 | Losses | YOUR loss functions | Part 2 (Solved) |
| 06 | Autograd | YOUR automatic differentiation | Part 2 (Solved) |
| 07 | Optimizers | YOUR SGD optimizer | Part 2 (Solved) |
Running the Milestone
The crisis needs Modules 01–03. The multi-layer solution needs Modules 01–04, 06, and 07, and tito runs it as the first part of Milestone 03. Running that part alone records it but does not complete Milestone 03, which also needs Part 2 (TinyDigits); tito milestone run 03 runs both:
tito module statustito milestone run 02 # the crisis: a single layer tops out at 75%
tito milestone run 03 --part 1 # the solution: one hidden layer reaches 100% (records Part 1 only)Expected Results
Table 2 records the accuracy and runtime you should expect to see.
| Script | Layers | Loss | Accuracy | What It Shows |
|---|---|---|---|---|
| 01 (Single Layer) | 1 | ~0.69 / N/A | \(\le\) 75% | Cannot learn XOR; a single layer that scores 100% fails the milestone, since only broken code can do that |
| 02 (Multi-Layer, run as Milestone 03 part 1) | 2 | 0.7468 before training, 0.0060 after 500 epochs | 100% by epoch 100 | Hidden layers solve it. Before training, YOUR BinaryCrossEntropyLoss must match a NumPy computation on the untrained outputs. Passes when all four truth-table rows are right, the hidden layer received a non-zero gradient, and its weights moved at least 0.25× their initial norm |
The Aha Moment: Depth Changes Everything
The numbers in the table are the aftermath. Live, the experiment feels different.
Script 01 tries weight after weight, first hand-picked and then random. Every one tops out at 75%. Did you break something?
You check the code. The script already did: every forward pass matched the same arithmetic computed in NumPy. Your Linear layer works. Your sigmoid works. But no line gets all four points right.
Then it lands: it’s not broken. It’s impossible. This is what Minsky proved. This is why funding died. Your code is slamming into the same mathematical wall that nearly ended AI research — every component working perfectly, all of it useless against XOR’s geometry.
Then you run script 02. Add one hidden layer. Loss drops steadily: 0.75… 0.18… 0.03… 0.01… 0.006. Accuracy: 100%.
Depth enables non-linear decision boundaries. The hidden layer learns to bend the input space until XOR becomes linearly separable. A single layer can only draw straight lines. Stack two, and you can draw any shape you need.
Same code. Same training loop. Same four points. The impossible is now trivial — and you’ve earned the right to call this deep learning.
Your Code Powers This
Table 3 names the TinyTorch components that power this milestone.
| Component | Your Module | What It Does |
|---|---|---|
Tensor |
Module 01 | Stores inputs and weights |
ReLU |
Module 02 | YOUR activation for hidden layer |
Linear |
Module 03 | YOUR fully-connected layers |
BinaryCrossEntropyLoss |
Module 04 | YOUR loss computation |
backward() |
Module 06 | YOUR autograd engine |
SGD |
Module 07 | YOUR optimizer |
Systems Insights
- Memory: O(n²) with hidden layers (vs O(n) for perceptron)
- Compute: O(n²) operations
- Breakthrough: Hidden representations unlock non-linear problems
Historical Context
Minsky and Papert’s proof was mathematically airtight — and read as a verdict on the whole research program. Multi-layer networks were known, but no one had a practical way to train them. That gap took 17 years to close: Rumelhart, Hinton, and Williams published backpropagation through hidden layers in 1986, and the field exhaled.
The lesson is uncomfortable. A correct theorem, applied to the wrong abstraction, set an entire field back nearly two decades.
Read why
The companion book’s chapter Synthesis I: From Perceptrons to Rumelhart’s MLP covers Milestones 01 through 03 together. Read it in TinyTorch: From Tensors to Transformers (PDF) after you have run the milestone.
What’s Next
XOR is a toy: four points, two dimensions, a problem you can solve in your head. The real question is whether the same trick — stack a hidden layer, let it learn its own representation — survives contact with image data. Milestone 03 points the same training stack at TinyDigits and finds out.
Further Reading
- The Crisis: Minsky, M., & Papert, S. (1969). “Perceptrons: An Introduction to Computational Geometry”
- The Solution: Rumelhart, Hinton, Williams (1986). “Learning representations by back-propagating errors”
- Wikipedia: AI Winter