Module 06: Autograd

Gradient computation is a major source of memory growth during training. Autograd’s choices about what to cache during the forward pass and when to release it decide whether a model fits in memory. The runtime you build here is the memory manager for the rest of the system.

NoteModule Info

FOUNDATION TIER | Difficulty: ●●●○ | Time: 6-8 hours | Prerequisites: 01-05

You need to be fluent with everything from Modules 01–05:

  • Tensor operations (matmul, broadcasting, reductions)
  • Activation functions (the source of non-linearity)
  • Neural network layers (what gradients will flow through)
  • Loss functions (the scalar gradients flow back from)
  • DataLoader for batched iteration

If you can hand-compute a forward pass through a small network and explain why we minimize loss, you’re ready.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

A neural network learns by nudging every parameter in the direction that lowers the loss. To find that direction you need a gradient — one number per parameter. A modern model has billions of parameters, so deriving those gradients by hand is not just tedious, it is impossible. Every framework you have ever used — PyTorch, TensorFlow, JAX — solves this with the same trick: automatic differentiation.

In this module you build reverse-mode autograd from scratch. The forward pass records each operation into a small graph; loss.backward() walks that graph in reverse, applying the chain rule one operation at a time. When you finish, calling loss.backward() on your tensors does the same thing it does in PyTorch — and you will know exactly why.

This is the conceptually hardest module in the Foundation tier. It is also the one that unlocks everything that follows: optimizers, training loops, and any model that learns from data.

Commands

# first time
tito module start 06

# later sessions
tito module resume 06

# when your tests pass
tito module complete 06

Your notebook is modules/06_autograd/autograd.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Trace the supplied Function base class that enables gradient computation for all operations
  • Follow how computation graphs track dependencies between tensors during forward pass
  • Master the chain rule by implementing the backward passes for matrix multiplication, ReLU, and mean squared error
  • Understand memory trade-offs between storing intermediate values and recomputing forward passes
  • Connect your autograd implementation to PyTorch’s design patterns and production optimizations

What you’ll build

Figure 1: A reverse-mode trace. For x = 3, the forward graph computes y = x times x = 9 and L = y + x = 12. Backward combines the direct contribution of 1 with two multiply contributions of 3 to give dx = 7.

The pattern you’ll enable:

# Automatic gradient computation
x = Tensor([2.0], requires_grad=True)
y = x * 3 + 1  # y = 3x + 1
y.backward()   # Computes dy/dx = 3 automatically
print(x.grad)  # [3.0]

What you’re not building yet

To keep this module focused, you will not implement:

  • Higher-order derivatives (gradients of gradients)—PyTorch supports this with create_graph=True
  • Static graph compilation (tracing a graph once and optimizing it ahead of time, as torch.compile and XLA do)
  • GPU kernel fusion—PyTorch’s JIT compiler optimizes backward pass operations
  • Checkpointing for memory efficiency—that’s an advanced optimization technique

You are building the core gradient engine. Advanced optimizations come in production frameworks.

What you write

The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:

MatMul.backward
Gradient computation for matrix multiplication.
ReLUFunction.backward
Gradient computation for ReLU activation.
MSEFunction.backward
Gradient computation for Mean Squared Error Loss.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.autograd;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (10). Each prints a ✅ line when it passes.

  • Broadcast Gradient Reduction
  • Arithmetic Backward Passes
  • Shape and Reduction Backward Passes
  • Broadcasting in Gradients
  • Activation Backward Passes
  • Stable Softmax Helper
  • Loss Backward Passes
  • One-Hot Encoding Helper
  • Tensor Autograd Enhancement
  • Gradients Through a Reused Tensor

Integration tests after export (17).

  • tests/06_autograd/test_06_autograd_progressive.py

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

MatMul.backward grad_a: [[12. 14.] [12. 14.]]
From test_unit_arithmetic_backward. The gradient for a used b as is. For C = A @ B, grad_A = grad_C @ B.T; transpose the last two axes of b (np.swapaxes(b, -2, -1)).
ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0
From test_unit_tensor_autograd. The same missing transpose, on non-square shapes, where NumPy cannot even multiply the arrays.
MSE.backward failed: [-0.125 -0.05 0.05 0.22499999] vs [-0.25 -0.09999999 0.1 0.45]
From test_unit_loss_backward. The gradient is off by a constant factor. For the mean of \((p - t)^2\) it is \(2(p - t)/N\); a missing 2 halves it and a missing \(N\) inflates it.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Automatic Differentiation: Recording the Graph and Walking It Backward in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

You can now compute a gradient for every parameter in any network you build. That gradient tells you which way is downhill — but it does not tell you how big a step to take, or how to dampen oscillations, or how to adapt the step size per-parameter. That is the optimizer’s job.

NoteUp next: Module 07, Optimizers

You’ll implement SGD, momentum, and Adam: the rules that turn the param.grad tensors produced by backward() into actual parameter updates. With autograd plus an optimizer, you have the entire machinery a training loop needs.

Next: Module 07: Optimizers

How later modules use this one

Table 1: How autograd feeds into subsequent optimizer and training modules.
Module What It Does Your Autograd In Action
07: Optimizers Update parameters using gradients optimizer.step() uses param.grad computed by backward()
08: Training Complete training loops loss.backward() → optimizer.step() → repeat
12: Attention Multi-head self-attention Gradients flow through Q, K, V projections automatically
Back to top