Module 03: Layers
A layer is a function with weights. Three stacked is an MLP; ninety-six is GPT-3. The interface discipline, every layer exposing the same forward() and parameters(), is what lets the optimizer find your weights, autograd walk your graph, and a future compiler fuse your kernels.
FOUNDATION TIER | Difficulty: ●●○○ | Time: 5-7 hours | Prerequisites: 01, 02
Prerequisites: Modules 01 and 02 means you have built:
- Tensor class with arithmetic, broadcasting, matrix multiplication, and shape manipulation
- Activation functions (ReLU, Sigmoid, Tanh, Softmax) for introducing non-linearity
- Understanding of element-wise operations and reductions
If you can multiply tensors, apply activations, and reason about shape transformations, you’re ready.
Overview
A layer is a function with weights. Stack three of them and you have a multi-layer perceptron; stack ninety-six and you have GPT-3. The discipline that makes such composition possible — every layer exposing the same forward() and parameters() interface — is what you build in this module.
You’ll work with four pieces: a Layer base class that fixes the interface, a Linear layer that applies the learned transformation y = xW + b, a Dropout layer that prevents overfitting by randomly zeroing activations, and a Sequential container that chains layers into networks. Together they’re the smallest set of abstractions that lets you write model = Sequential(Linear(784, 256), ReLU(), Linear(256, 10)) and have it just work.
The payoff is parameters(). Once every layer hands its weights to a single list, optimizers (Module 07) and autograd (Module 06) can update them without caring what’s inside. PyTorch’s nn.Module follows the exact same contract — you’re building the real abstraction, not a toy version of it.
Commands
# first time
tito module start 03
# later sessions
tito module resume 03
# when your tests pass
tito module complete 03Your notebook is modules/03_layers/layers.ipynb.
Learning objectives
- Implement the Linear and Dropout forward passes, and read how the supplied constructors handle weight initialization and parameter management for gradient-based training
- Master the mathematical operation
y = xW + band understand how parameter counts scale with layer dimensions - Understand memory usage patterns (parameter memory vs activation memory) and computational complexity of matrix operations
- Connect your implementation to production PyTorch patterns, including
nn.Linear,nn.Dropout, and parameter tracking
What you’ll build
forward() and parameters() from Layer, while Sequential provides the same interface to chain them into networks.
The pattern you’ll enable:
# Building a multi-layer network
layer1 = Linear(784, 256)
activation = ReLU()
dropout = Dropout(0.5)
layer2 = Linear(256, 10)
# Manual composition for explicit data flow
x = layer1(x)
x = activation(x)
x = dropout(x) # training mode by default; call dropout.eval() (or model.eval()) for inference
output = layer2(x)What you’re not building yet
To keep this module focused, you will not implement:
- Automatic gradient computation (autograd is a later module)
- Parameter optimization (optimizers are a later module)
- Hundreds of layer types (PyTorch has Conv2d, LSTM, Attention - you’ll build Linear and Dropout)
You are building the core building blocks. Training loops and optimizers come later.
What you write
The notebook arrives with the surrounding code already written and explained. You write 2 functions, each marked # YOUR CODE HERE and followed by a test cell:
Linear.forward- Forward pass through linear layer.
Dropout.forward- Forward pass through dropout layer.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.core.layers; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (6). Each prints a ✅ line when it passes.
- Linear Layer
- Linear Edge Cases
- Linear Parameter Collection
- Dropout Decision Logic
- Dropout Mask Generation
- Dropout Layer
Integration tests after export (13).
tests/03_layers/test_03_layers_progressive.py
Completing this module unlocks Milestone 01, Perceptron (1958) (tito milestone run 01) and Milestone 02, XOR Crisis (1969) (tito milestone run 02).
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
ValueError: Matrix multiplication shape mismatch: (32, 784) @ (256, 784)-
From
test_unit_linear_layer.Linear.forwardtransposed the weight. TinyTorch stores it as(in_features, out_features), so the forward pass isx.matmul(self.weight)with no transpose. Inference should pass through unchanged-
From
test_unit_dropout_layer. Dropout ran withtraining=False. Return the input unchanged whenself._should_apply_dropout(training)is false. ValueError: setting an array element with a sequence.-
From
test_unit_dropout_layer. The mask was applied tox.dataand wrapped in a newTensor. Multiply the tensor itself,x * mask, so the operation stays on the path Module 06 will differentiate.
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Layers: Parameters, State, and Initialization in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
Your layers can produce predictions, but you have no way to say how wrong a prediction is. Module 04 introduces loss functions — MSELoss for regression, CrossEntropyLoss for classification — that turn a prediction and a target into a single scalar. That scalar becomes the signal autograd (Module 06) backpropagates, which optimizers (Module 07) then apply to the very parameters() you collected here.
Next: Module 04: Losses
How later modules use this one
| Module | What It Does | Your Layers In Action |
|---|---|---|
| 04: Losses | Quantify prediction error | loss = CrossEntropyLoss()(model(x), y) |
| 06: Autograd | Backpropagate through the stack | loss.backward() fills layer.weight.grad |
| 07: Optimizers | Update parameters from gradients | optimizer.step() consumes model.parameters() |