Module 07: Optimizers
A gradient is a direction, not a step. The optimizer picks a step size, a history window, and a per-parameter scaling rule. Its state, momentum buffers and running second moments, often costs more memory than the model itself. Choose well and a model converges in an afternoon. Choose poorly and VRAM runs out on epoch one.
FOUNDATION TIER | Difficulty: ●●○○ | Time: 3-5 hours | Prerequisites: 01-06
Prerequisites: Modules 01-06 means you need:
- Tensor operations and parameter storage
- DataLoader for efficient batch processing
- Understanding of forward/backward passes (autograd)
- Why gradients point toward higher loss
If you understand how loss.backward() computes gradients and why we need to update parameters to minimize loss, you’re ready.
Overview
You have gradients. Now what? An optimizer is the rule that turns a gradient into a parameter update — the difference between a model that converges in an afternoon and one that diverges on the first batch. Picture optimization as hiking in fog: you can feel the slope under your feet but cannot see the valley. Each optimizer is a different strategy for choosing your next step.
You’ll meet three, writing the first two yourself: SGD with momentum (the foundation), Adam with adaptive per-parameter learning rates (the modern workhorse), and AdamW with decoupled weight decay (the default for transformers). They differ in memory cost, convergence speed, and how forgiving they are when you guess the learning rate wrong — and that last point matters more than most practitioners admit.
Commands
# first time
tito module start 07
# later sessions
tito module resume 07
# when your tests pass
tito module complete 07Your notebook is modules/07_optimizers/optimizers.ipynb.
Learning objectives
- Implement SGD with momentum to reduce oscillations and accelerate convergence in narrow valleys
- Master Adam’s adaptive learning rate mechanism with first and second moment estimation
- Understand memory trade-offs (plain SGD: 2x, SGD with momentum: 3x, Adam/AdamW: 4x the parameter bytes, counting weights, gradients, and optimizer state) and computational complexity per step
- Connect optimizer state management to checkpointing and distributed training considerations
What you’ll build
The pattern you’ll enable:
# Training loop with optimizer
optimizer = Adam(model.parameters(), lr=0.001)
loss.backward() # Compute gradients (Module 06)
optimizer.step() # Update parameters using gradients
optimizer.zero_grad() # Clear gradients for next iterationWhat you’re not building yet
To keep this module focused, you will not implement:
- Learning rate schedules (that’s Module 08: Training)
- Gradient clipping (that’s Module 08: Training)
- Second-order optimizers like L-BFGS (rarely used in deep learning due to memory cost)
- Distributed optimizer sharding (production frameworks use techniques like ZeRO)
You are building the core optimization algorithms. Advanced training techniques come in Module 08.
What you write
The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:
SGD.step- Perform SGD update step with momentum.
Adam._update_moments- Update first and second moment estimates with bias correction.
Adam.step- Perform Adam update step by composing helpers.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.core.optimizers; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (8). Each prints a ✅ line when it passes.
- Gradient Extraction
- Base Optimizer
- SGD Optimizer
- Adam Moment Updates
- Adam Optimizer
- AdamW Moment Updates
- AdamW Optimizer
- Adam Checkpoint State
Integration tests after export (10).
tests/07_optimizers/test_07_optimizers_progressive.py
Completing this module unlocks Milestone 03, MLP Revival (1986) (tito milestone run 03).
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
assert np.allclose(param2.data, expected_second, rtol=1e-5)-
From
test_unit_sgd_optimizer. The second SGD step is wrong because the momentum buffer did not accumulate. Update it asmomentum * buffer + grad, then step with the buffer. m_hat should equal grad at step 1, got [0.01 0.02]-
From
test_unit_adam_update_moments. No bias correction. Dividemby1 - beta1**tandvby1 - beta2**t. m_hat should equal grad at step 1, got [inf inf]-
From
test_unit_adam_update_moments. The correction ran before the update count was incremented, sot = 0and1 - beta1**0is zero. Increment the count first.
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Optimizers: Momentum, AdamW, and Decoupled Weight Decay in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
You have an optimizer that takes one step. Module 08 wraps it in the loop that takes a million: epochs, validation, checkpointing, gradient clipping, and learning rate schedules. The question it answers is the one this chapter raised but did not solve — how do you actually drive a model to convergence without babysitting it?
Next: Module 08: Training
How later modules use this one
| Module | What It Does | Your Optimizers In Action |
|---|---|---|
| 08: Training | Complete training loops | for epoch in range(10): loss.backward(); optimizer.step() |
| 09: Convolutions | Convolutional networks | AdamW optimizes millions of CNN parameters efficiently |
| 13: Transformers | Attention mechanisms | Large models require careful optimizer selection |