Module 02: Activations

Activations are the first operator in your framework with almost no arithmetic intensity: every element is read once, used for a handful of operations, and never reused. That makes them pure memory-bandwidth workloads. Stacking them also breaks the linear collapse that would otherwise flatten a hundred layers into one matrix multiply, so both the math and the memory-wall lesson start here.

NoteModule Info

FOUNDATION TIER | Difficulty: ●○○○ | Time: 3-5 hours | Prerequisites: 01

Prerequisites: Module 01 means you need:

  • Completed Tensor implementation with element-wise operations
  • Understanding of tensor shapes and broadcasting
  • Familiarity with NumPy mathematical functions

If you can create a Tensor and perform element-wise arithmetic (x + y, x * 2), you’re ready.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

A neural network without activation functions isn’t a neural network — it’s a single matrix multiplication wearing a costume. Stack a hundred linear layers, and the composition collapses to one: \(W_2(W_1 x) = (W_2 W_1)x\). Depth buys you nothing until you break the linearity.

Activations are how you break it. ReLU zeros out negatives. Sigmoid squashes any real number into \((0, 1)\). Softmax turns raw scores into a probability distribution. Each is just a few lines of math, but together they’re what lets a network learn to distinguish a cat from a dog instead of computing one giant linear regression.

You’ll work with five of them, writing ReLU, Sigmoid, and Softmax yourself and reading the supplied Tanh and GELU. Along the way you’ll meet the chapter’s load-bearing insight — why every production softmax subtracts the max before exponentiating — and the dead-neuron problem that explains why ReLU’s apparent simplicity hides a real failure mode.

Commands

# first time
tito module start 02

# later sessions
tito module resume 02

# when your tests pass
tito module complete 02

Your notebook is modules/02_activations/activations.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement three activation functions (Sigmoid, ReLU, Softmax) with the numerical-stability tricks production frameworks use, and read the supplied Tanh and GELU alongside them
  • Explain why nonlinearity turns a stack of matrix multiplies into a function approximator
  • Quantify the compute cost of each activation and decide when the extra accuracy of GELU is worth the extra exponentials
  • Map your implementations onto the corresponding torch.nn.functional calls so PyTorch stops feeling like a black box

What you’ll build

Figure 1: Independent activation results for the same input. Elementwise functions preserve shape, while Softmax normalizes across the vector; GELU uses TinyTorch’s sigmoid approximation.

The pattern you’ll enable:

# Transforming tensors through nonlinear functions
relu = ReLU()
activated = relu(x)  # Zeros negatives, keeps positives

softmax = Softmax()
probabilities = softmax(logits)  # Converts to probability distribution (sums to 1)

What you’re not building yet

To keep this module focused, you will not implement:

  • Gradient computation (backward() methods are attached in Module 06 via autograd)
  • Learnable parameters (activations are fixed mathematical transformations with parameters() -> [])
  • Advanced variants (LeakyReLU, ELU, Swish — PyTorch ships dozens; you’ll build the core five that most architectures use)
  • GPU acceleration (your NumPy implementation runs on CPU)

The forward pass is enough to build intuition. Gradients show up in Module 06.

What you write

The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:

SigmoidFunction.forward
Apply sigmoid activation element-wise.
ReLUFunction.forward
Apply ReLU activation element-wise.
SoftmaxFunction.forward
Apply softmax activation along specified dimension.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.activations;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (5). Each prints a ✅ line when it passes.

  • Sigmoid
  • ReLU
  • Tanh
  • GELU
  • Softmax

Integration tests after export (12).

  • tests/02_activations/test_02_activations_progressive.py

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

sigmoid overflowed on extreme inputs (overflow encountered in exp)
From test_unit_sigmoid. The direct formula 1 / (1 + np.exp(-x)) overflows for large negative x. Compute z = np.exp(-np.abs(x)), whose exponent is never positive, and choose 1 / (1 + z) or z / (1 + z) by the sign of x.
Softmax should handle large numbers
From test_unit_softmax. Exponentiating the raw scores overflows. Subtract each row’s maximum before np.exp; the result is unchanged and every exponent is at most zero.
Each row should sum to 1
From test_unit_softmax. The normalizing sum ran over the whole array. Sum along self.dim with keepdims=True, so each row is divided by its own total.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Activation Functions: Non-Linearity and Numerical Stability in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

You now have nonlinearity. You still don’t have anything to be nonlinear about. ReLU on a raw input vector accomplishes nothing — the network needs a learnable transformation between activations, something with weights and biases that gradient descent can shape.

That’s the next module.

NoteUp next: Module 03, Layers

The question Module 03 answers: what is Linear(x), exactly, and how does it compose with the activations you just built into the canonical Linear → activation → Linear → activation → ... stack that defines an MLP?

You’ll implement the Linear layer — weight matrix, bias vector, forward pass — and then chain it with your ReLU and Softmax to build the first thing in this course that deserves to be called a neural network.

Next: Module 03: Layers

How later modules use this one

Table 1: How activations feed into subsequent TinyTorch modules.
Module What It Does Your Activations In Action
03: Layers Neural network building blocks Linear(x) followed by ReLU()(output)
04: Losses Training objectives Softmax probabilities feed into cross-entropy loss
06: Autograd Automatic gradients ReLUFunction.backward computes activation gradients
Back to top