Milestone 02: The XOR Crisis (1969)

NoteMilestone Info

Foundation Milestone | Difficulty: ●●○○ | Time: 30–45 min | Prerequisites: Modules 01–03 (Part 1) / Modules 01–04, 06, and 07 (Part 2)

TipWhat You’ll Learn
  • Why single-layer networks have fundamental mathematical limits
  • How hidden layers enable non-linear decision boundaries
  • Why “deep” learning is called DEEP

Overview

It’s 1969. Neural networks are the hottest thing in AI. Funding is pouring in. Then Marvin Minsky and Seymour Papert publish a 308-page mathematical proof that destroys everything: perceptrons cannot solve XOR. Not “struggle with” — CANNOT. Mathematically impossible.

Funding evaporates overnight. Research labs shut down. The field dies for 17 years — the infamous AI Winter.

You’re about to live that crisis. You’ll watch your own perceptron — built from your own modules — fail on four points despite trying many weight configurations. Accuracy never gets past 75%. Then you’ll add one hidden layer and watch the impossible collapse into the trivial.

What You’ll Recreate

Two demonstrations of perceptron limitations and the multi-layer solution:

  1. The Crisis — watch a perceptron fail on XOR due to linear boundary limits
  2. The Solution — add a hidden layer and solve the “impossible” problem
Crisis:   Input --> Linear --> Output (FAILS)
Solution: Input --> Linear --> ReLU --> Linear --> Output (100%!)

The XOR Problem

Inputs    Output
x1  x2    XOR
0   0  -->  0   (same)
0   1  -->  1   (different)
1   0  -->  1   (different)
1   1  -->  0   (same)

Plot those four points. The two zeros sit on one diagonal, the two ones on the other. No straight line separates them — and a single-layer perceptron can only draw straight lines. No amount of training fixes that. It’s geometry, not optimization.

Figure 1: The XOR crisis and non-linear space folding. Left: Diagonally paired classes cannot be separated by any single straight line (\(w_1 x_1 + w_2 x_2 + b = 0\)), capping single-layer accuracy at 75%. Right: Adding a hidden layer with non-linear activation (ReLU) transforms the coordinate space (\([h_1, h_2]\)), folding the input so that a simple linear decision boundary achieves 100% accuracy.

Prerequisites

Table 1 lists the modules you need to have completed before starting.

Table 1: Prerequisite modules for the XOR milestone.
Module Component What It Provides Required For
01 Tensor YOUR data structure Parts 1 & 2
02 Activations YOUR sigmoid/ReLU Parts 1 & 2
03 Layers YOUR Linear layers Parts 1 & 2
04 Losses YOUR loss functions Part 2 (Solved)
06 Autograd YOUR automatic differentiation Part 2 (Solved)
07 Optimizers YOUR SGD optimizer Part 2 (Solved)

Running the Milestone

The crisis needs Modules 01–03. The multi-layer solution needs Modules 01–04, 06, and 07, and tito runs it as the first part of Milestone 03. Running that part alone records it but does not complete Milestone 03, which also needs Part 2 (TinyDigits); tito milestone run 03 runs both:

tito module status
tito milestone run 02            # the crisis: a single layer tops out at 75%
tito milestone run 03 --part 1   # the solution: one hidden layer reaches 100% (records Part 1 only)

Expected Results

Table 2 records the accuracy and runtime you should expect to see.

Table 2: Expected loss and accuracy for the XOR milestone scripts.
Script Layers Loss Accuracy What It Shows
01 (Single Layer) 1 ~0.69 / N/A \(\le\) 75% Cannot learn XOR; a single layer that scores 100% fails the milestone, since only broken code can do that
02 (Multi-Layer, run as Milestone 03 part 1) 2 0.7468 before training, 0.0060 after 500 epochs 100% by epoch 100 Hidden layers solve it. Before training, YOUR BinaryCrossEntropyLoss must match a NumPy computation on the untrained outputs. Passes when all four truth-table rows are right, the hidden layer received a non-zero gradient, and its weights moved at least 0.25× their initial norm

The Aha Moment: Depth Changes Everything

The numbers in the table are the aftermath. Live, the experiment feels different.

Script 01 tries weight after weight, first hand-picked and then random. Every one tops out at 75%. Did you break something?

You check the code. The script already did: every forward pass matched the same arithmetic computed in NumPy. Your Linear layer works. Your sigmoid works. But no line gets all four points right.

Then it lands: it’s not broken. It’s impossible. This is what Minsky proved. This is why funding died. Your code is slamming into the same mathematical wall that nearly ended AI research — every component working perfectly, all of it useless against XOR’s geometry.

Then you run script 02. Add one hidden layer. Loss drops steadily: 0.75… 0.18… 0.03… 0.01… 0.006. Accuracy: 100%.

Depth enables non-linear decision boundaries. The hidden layer learns to bend the input space until XOR becomes linearly separable. A single layer can only draw straight lines. Stack two, and you can draw any shape you need.

Same code. Same training loop. Same four points. The impossible is now trivial — and you’ve earned the right to call this deep learning.

Your Code Powers This

Table 3 names the TinyTorch components that power this milestone.

Table 3: TinyTorch components that power the XOR milestone.
Component Your Module What It Does
Tensor Module 01 Stores inputs and weights
ReLU Module 02 YOUR activation for hidden layer
Linear Module 03 YOUR fully-connected layers
BinaryCrossEntropyLoss Module 04 YOUR loss computation
backward() Module 06 YOUR autograd engine
SGD Module 07 YOUR optimizer

Systems Insights

  • Memory: O(n²) with hidden layers (vs O(n) for perceptron)
  • Compute: O(n²) operations
  • Breakthrough: Hidden representations unlock non-linear problems

Historical Context

Minsky and Papert’s proof was mathematically airtight — and read as a verdict on the whole research program. Multi-layer networks were known, but no one had a practical way to train them. That gap took 17 years to close: Rumelhart, Hinton, and Williams published backpropagation through hidden layers in 1986, and the field exhaled.

The lesson is uncomfortable. A correct theorem, applied to the wrong abstraction, set an entire field back nearly two decades.

Read why

The companion book’s chapter Synthesis I: From Perceptrons to Rumelhart’s MLP covers Milestones 01 through 03 together. Read it in TinyTorch: From Tensors to Transformers (PDF) after you have run the milestone.

What’s Next

XOR is a toy: four points, two dimensions, a problem you can solve in your head. The real question is whether the same trick — stack a hidden layer, let it learn its own representation — survives contact with image data. Milestone 03 points the same training stack at TinyDigits and finds out.

Further Reading

Back to top