Module 02: Activations
Activations are the first operator in your framework with almost no arithmetic intensity: every element is read once, used for a handful of operations, and never reused. That makes them pure memory-bandwidth workloads. Stacking them also breaks the linear collapse that would otherwise flatten a hundred layers into one matrix multiply, so both the math and the memory-wall lesson start here.
FOUNDATION TIER | Difficulty: ●○○○ | Time: 3-5 hours | Prerequisites: 01
Prerequisites: Module 01 means you need:
- Completed Tensor implementation with element-wise operations
- Understanding of tensor shapes and broadcasting
- Familiarity with NumPy mathematical functions
If you can create a Tensor and perform element-wise arithmetic (x + y, x * 2), you’re ready.
Overview
A neural network without activation functions isn’t a neural network — it’s a single matrix multiplication wearing a costume. Stack a hundred linear layers, and the composition collapses to one: \(W_2(W_1 x) = (W_2 W_1)x\). Depth buys you nothing until you break the linearity.
Activations are how you break it. ReLU zeros out negatives. Sigmoid squashes any real number into \((0, 1)\). Softmax turns raw scores into a probability distribution. Each is just a few lines of math, but together they’re what lets a network learn to distinguish a cat from a dog instead of computing one giant linear regression.
You’ll work with five of them, writing ReLU, Sigmoid, and Softmax yourself and reading the supplied Tanh and GELU. Along the way you’ll meet the chapter’s load-bearing insight — why every production softmax subtracts the max before exponentiating — and the dead-neuron problem that explains why ReLU’s apparent simplicity hides a real failure mode.
Commands
# first time
tito module start 02
# later sessions
tito module resume 02
# when your tests pass
tito module complete 02Your notebook is modules/02_activations/activations.ipynb.
Learning objectives
- Implement three activation functions (Sigmoid, ReLU, Softmax) with the numerical-stability tricks production frameworks use, and read the supplied Tanh and GELU alongside them
- Explain why nonlinearity turns a stack of matrix multiplies into a function approximator
- Quantify the compute cost of each activation and decide when the extra accuracy of GELU is worth the extra exponentials
- Map your implementations onto the corresponding
torch.nn.functionalcalls so PyTorch stops feeling like a black box
What you’ll build
The pattern you’ll enable:
# Transforming tensors through nonlinear functions
relu = ReLU()
activated = relu(x) # Zeros negatives, keeps positives
softmax = Softmax()
probabilities = softmax(logits) # Converts to probability distribution (sums to 1)What you’re not building yet
To keep this module focused, you will not implement:
- Gradient computation (
backward()methods are attached in Module 06 via autograd) - Learnable parameters (activations are fixed mathematical transformations with
parameters() -> []) - Advanced variants (LeakyReLU, ELU, Swish — PyTorch ships dozens; you’ll build the core five that most architectures use)
- GPU acceleration (your NumPy implementation runs on CPU)
The forward pass is enough to build intuition. Gradients show up in Module 06.
What you write
The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:
SigmoidFunction.forward- Apply sigmoid activation element-wise.
ReLUFunction.forward- Apply ReLU activation element-wise.
SoftmaxFunction.forward- Apply softmax activation along specified dimension.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.core.activations; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (5). Each prints a ✅ line when it passes.
- Sigmoid
- ReLU
- Tanh
- GELU
- Softmax
Integration tests after export (12).
tests/02_activations/test_02_activations_progressive.py
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
sigmoid overflowed on extreme inputs (overflow encountered in exp)-
From
test_unit_sigmoid. The direct formula1 / (1 + np.exp(-x))overflows for large negativex. Computez = np.exp(-np.abs(x)), whose exponent is never positive, and choose1 / (1 + z)orz / (1 + z)by the sign ofx. Softmax should handle large numbers-
From
test_unit_softmax. Exponentiating the raw scores overflows. Subtract each row’s maximum beforenp.exp; the result is unchanged and every exponent is at most zero. Each row should sum to 1-
From
test_unit_softmax. The normalizing sum ran over the whole array. Sum alongself.dimwithkeepdims=True, so each row is divided by its own total.
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Activation Functions: Non-Linearity and Numerical Stability in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
You now have nonlinearity. You still don’t have anything to be nonlinear about. ReLU on a raw input vector accomplishes nothing — the network needs a learnable transformation between activations, something with weights and biases that gradient descent can shape.
That’s the next module.
The question Module 03 answers: what is Linear(x), exactly, and how does it compose with the activations you just built into the canonical Linear → activation → Linear → activation → ... stack that defines an MLP?
You’ll implement the Linear layer — weight matrix, bias vector, forward pass — and then chain it with your ReLU and Softmax to build the first thing in this course that deserves to be called a neural network.
Next: Module 03: Layers
How later modules use this one
| Module | What It Does | Your Activations In Action |
|---|---|---|
| 03: Layers | Neural network building blocks | Linear(x) followed by ReLU()(output) |
| 04: Losses | Training objectives | Softmax probabilities feed into cross-entropy loss |
| 06: Autograd | Automatic gradients | ReLUFunction.backward computes activation gradients |