Module 07: Optimizers

A gradient is a direction, not a step. The optimizer picks a step size, a history window, and a per-parameter scaling rule. Its state, momentum buffers and running second moments, often costs more memory than the model itself. Choose well and a model converges in an afternoon. Choose poorly and VRAM runs out on epoch one.

NoteModule Info

FOUNDATION TIER | Difficulty: ●●○○ | Time: 3-5 hours | Prerequisites: 01-06

Prerequisites: Modules 01-06 means you need:

  • Tensor operations and parameter storage
  • DataLoader for efficient batch processing
  • Understanding of forward/backward passes (autograd)
  • Why gradients point toward higher loss

If you understand how loss.backward() computes gradients and why we need to update parameters to minimize loss, you’re ready.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

You have gradients. Now what? An optimizer is the rule that turns a gradient into a parameter update — the difference between a model that converges in an afternoon and one that diverges on the first batch. Picture optimization as hiking in fog: you can feel the slope under your feet but cannot see the valley. Each optimizer is a different strategy for choosing your next step.

You’ll meet three, writing the first two yourself: SGD with momentum (the foundation), Adam with adaptive per-parameter learning rates (the modern workhorse), and AdamW with decoupled weight decay (the default for transformers). They differ in memory cost, convergence speed, and how forgiving they are when you guess the learning rate wrong — and that last point matters more than most practitioners admit.

Commands

# first time
tito module start 07

# later sessions
tito module resume 07

# when your tests pass
tito module complete 07

Your notebook is modules/07_optimizers/optimizers.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement SGD with momentum to reduce oscillations and accelerate convergence in narrow valleys
  • Master Adam’s adaptive learning rate mechanism with first and second moment estimation
  • Understand memory trade-offs (plain SGD: 2x, SGD with momentum: 3x, Adam/AdamW: 4x the parameter bytes, counting weights, gradients, and optimizer state) and computational complexity per step
  • Connect optimizer state management to checkpointing and distributed training considerations

What you’ll build

Figure 1: Optimizer state and weight decay. SGD uses coupled weight decay and optional momentum. Adam forms bias-corrected moments from the gradient plus coupled decay. AdamW forms moments from the gradient and applies weight decay separately.

The pattern you’ll enable:

# Training loop with optimizer
optimizer = Adam(model.parameters(), lr=0.001)
loss.backward()  # Compute gradients (Module 06)
optimizer.step()  # Update parameters using gradients
optimizer.zero_grad()  # Clear gradients for next iteration

What you’re not building yet

To keep this module focused, you will not implement:

  • Learning rate schedules (that’s Module 08: Training)
  • Gradient clipping (that’s Module 08: Training)
  • Second-order optimizers like L-BFGS (rarely used in deep learning due to memory cost)
  • Distributed optimizer sharding (production frameworks use techniques like ZeRO)

You are building the core optimization algorithms. Advanced training techniques come in Module 08.

What you write

The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:

SGD.step
Perform SGD update step with momentum.
Adam._update_moments
Update first and second moment estimates with bias correction.
Adam.step
Perform Adam update step by composing helpers.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.optimizers;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (8). Each prints a ✅ line when it passes.

  • Gradient Extraction
  • Base Optimizer
  • SGD Optimizer
  • Adam Moment Updates
  • Adam Optimizer
  • AdamW Moment Updates
  • AdamW Optimizer
  • Adam Checkpoint State

Integration tests after export (10).

  • tests/07_optimizers/test_07_optimizers_progressive.py

Completing this module unlocks Milestone 03, MLP Revival (1986) (tito milestone run 03).

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

assert np.allclose(param2.data, expected_second, rtol=1e-5)
From test_unit_sgd_optimizer. The second SGD step is wrong because the momentum buffer did not accumulate. Update it as momentum * buffer + grad, then step with the buffer.
m_hat should equal grad at step 1, got [0.01 0.02]
From test_unit_adam_update_moments. No bias correction. Divide m by 1 - beta1**t and v by 1 - beta2**t.
m_hat should equal grad at step 1, got [inf inf]
From test_unit_adam_update_moments. The correction ran before the update count was incremented, so t = 0 and 1 - beta1**0 is zero. Increment the count first.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Optimizers: Momentum, AdamW, and Decoupled Weight Decay in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

NoteUp next: Module 08, Training

You have an optimizer that takes one step. Module 08 wraps it in the loop that takes a million: epochs, validation, checkpointing, gradient clipping, and learning rate schedules. The question it answers is the one this chapter raised but did not solve — how do you actually drive a model to convergence without babysitting it?

Next: Module 08: Training

How later modules use this one

Table 1: How optimizers feed into subsequent training modules.
Module What It Does Your Optimizers In Action
08: Training Complete training loops for epoch in range(10): loss.backward(); optimizer.step()
09: Convolutions Convolutional networks AdamW optimizes millions of CNN parameters efficiently
13: Transformers Attention mechanisms Large models require careful optimizer selection
Back to top