Module 03: Layers

A layer is a function with weights. Three stacked is an MLP; ninety-six is GPT-3. The interface discipline, every layer exposing the same forward() and parameters(), is what lets the optimizer find your weights, autograd walk your graph, and a future compiler fuse your kernels.

NoteModule Info

FOUNDATION TIER | Difficulty: ●●○○ | Time: 5-7 hours | Prerequisites: 01, 02

Prerequisites: Modules 01 and 02 means you have built:

  • Tensor class with arithmetic, broadcasting, matrix multiplication, and shape manipulation
  • Activation functions (ReLU, Sigmoid, Tanh, Softmax) for introducing non-linearity
  • Understanding of element-wise operations and reductions

If you can multiply tensors, apply activations, and reason about shape transformations, you’re ready.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

A layer is a function with weights. Stack three of them and you have a multi-layer perceptron; stack ninety-six and you have GPT-3. The discipline that makes such composition possible — every layer exposing the same forward() and parameters() interface — is what you build in this module.

You’ll work with four pieces: a Layer base class that fixes the interface, a Linear layer that applies the learned transformation y = xW + b, a Dropout layer that prevents overfitting by randomly zeroing activations, and a Sequential container that chains layers into networks. Together they’re the smallest set of abstractions that lets you write model = Sequential(Linear(784, 256), ReLU(), Linear(256, 10)) and have it just work.

The payoff is parameters(). Once every layer hands its weights to a single list, optimizers (Module 07) and autograd (Module 06) can update them without caring what’s inside. PyTorch’s nn.Module follows the exact same contract — you’re building the real abstraction, not a toy version of it.

Commands

# first time
tito module start 03

# later sessions
tito module resume 03

# when your tests pass
tito module complete 03

Your notebook is modules/03_layers/layers.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement the Linear and Dropout forward passes, and read how the supplied constructors handle weight initialization and parameter management for gradient-based training
  • Master the mathematical operation y = xW + b and understand how parameter counts scale with layer dimensions
  • Understand memory usage patterns (parameter memory vs activation memory) and computational complexity of matrix operations
  • Connect your implementation to production PyTorch patterns, including nn.Linear, nn.Dropout, and parameter tracking

What you’ll build

Figure 1: A shared layer interface. Linear and Dropout inherit forward() and parameters() from Layer, while Sequential provides the same interface to chain them into networks.

The pattern you’ll enable:

# Building a multi-layer network
layer1 = Linear(784, 256)
activation = ReLU()
dropout = Dropout(0.5)
layer2 = Linear(256, 10)

# Manual composition for explicit data flow
x = layer1(x)
x = activation(x)
x = dropout(x)  # training mode by default; call dropout.eval() (or model.eval()) for inference
output = layer2(x)

What you’re not building yet

To keep this module focused, you will not implement:

  • Automatic gradient computation (autograd is a later module)
  • Parameter optimization (optimizers are a later module)
  • Hundreds of layer types (PyTorch has Conv2d, LSTM, Attention - you’ll build Linear and Dropout)

You are building the core building blocks. Training loops and optimizers come later.

What you write

The notebook arrives with the surrounding code already written and explained. You write 2 functions, each marked # YOUR CODE HERE and followed by a test cell:

Linear.forward
Forward pass through linear layer.
Dropout.forward
Forward pass through dropout layer.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.layers;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (6). Each prints a ✅ line when it passes.

  • Linear Layer
  • Linear Edge Cases
  • Linear Parameter Collection
  • Dropout Decision Logic
  • Dropout Mask Generation
  • Dropout Layer

Integration tests after export (13).

  • tests/03_layers/test_03_layers_progressive.py

Completing this module unlocks Milestone 01, Perceptron (1958) (tito milestone run 01) and Milestone 02, XOR Crisis (1969) (tito milestone run 02).

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

ValueError: Matrix multiplication shape mismatch: (32, 784) @ (256, 784)
From test_unit_linear_layer. Linear.forward transposed the weight. TinyTorch stores it as (in_features, out_features), so the forward pass is x.matmul(self.weight) with no transpose.
Inference should pass through unchanged
From test_unit_dropout_layer. Dropout ran with training=False. Return the input unchanged when self._should_apply_dropout(training) is false.
ValueError: setting an array element with a sequence.
From test_unit_dropout_layer. The mask was applied to x.data and wrapped in a new Tensor. Multiply the tensor itself, x * mask, so the operation stays on the path Module 06 will differentiate.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Layers: Parameters, State, and Initialization in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

NoteUp next: Module 04, Losses

Your layers can produce predictions, but you have no way to say how wrong a prediction is. Module 04 introduces loss functions — MSELoss for regression, CrossEntropyLoss for classification — that turn a prediction and a target into a single scalar. That scalar becomes the signal autograd (Module 06) backpropagates, which optimizers (Module 07) then apply to the very parameters() you collected here.

Next: Module 04: Losses

How later modules use this one

Table 1: How the Layer abstractions feed into later training modules.
Module What It Does Your Layers In Action
04: Losses Quantify prediction error loss = CrossEntropyLoss()(model(x), y)
06: Autograd Backpropagate through the stack loss.backward() fills layer.weight.grad
07: Optimizers Update parameters from gradients optimizer.step() consumes model.parameters()
Back to top