Module 04: Losses
A loss function reduces a full output tensor to a single scalar. That reduction is the only signal the optimizer ever sees. Build the reduction wrong, or apply it to the wrong axis, and every gradient flowing backward through autograd optimizes the wrong objective.
FOUNDATION TIER | Difficulty: ●●○○ | Time: 4-6 hours | Prerequisites: 01, 02, 03
Prerequisites: Modules 01-03 means you should understand:
- Tensor operations and broadcasting (Module 01)
- Activation functions and their role in neural networks (Module 02)
- Layers and how they transform data (Module 03)
If you can build a simple neural network that takes input and produces output, you’re ready to learn how to measure its quality.
Overview
A neural network without a loss function is a guess machine. The loss is the single scalar that tells optimization which direction to move — turning a forward pass into a learning step. Get it wrong and training stalls, diverges, or silently optimizes the wrong objective.
In this module you’ll work with three losses that cover most supervised learning, writing the first two yourself: Mean Squared Error for regression, CrossEntropy for multi-class classification, and Binary Cross-Entropy for multi-label or binary decisions. Along the way you’ll implement the log-sum-exp trick — the one numerical safeguard that separates a softmax that trains from one that returns nan on the first batch with large logits.
By the end, you’ll know not just how to compute each loss, but why the choice of loss reshapes what your model learns, and where naive implementations break at production scale.
Commands
# first time
tito module start 04
# later sessions
tito module resume 04
# when your tests pass
tito module complete 04Your notebook is modules/04_losses/losses.ipynb.
Learning objectives
- Implement MSELoss for regression and CrossEntropyLoss for multi-class classification, then read the supplied BinaryCrossEntropyLoss for binary decisions
- Master the log-sum-exp trick for numerically stable softmax computation
- Understand computational complexity (O(B×C) for cross-entropy with large vocabularies) and memory trade-offs
- Analyze loss function behavior across different prediction patterns and confidence levels
- Connect your implementation to production PyTorch patterns and engineering decisions at scale
What you’ll build
The pattern you’ll enable:
# Measuring prediction quality
loss = criterion(predictions, targets) # Scalar feedback signal for learningWhat you’re not building yet
To keep this module focused, you will not implement:
- Class weights (handling imbalanced datasets)
- Label smoothing (regularization technique)
- Advanced loss functions (Focal Loss, Triplet Loss, Contrastive Loss)
- GPU acceleration (your NumPy implementation runs on CPU)
These three losses cover most supervised training you will do in this course.
What you write
The notebook arrives with the surrounding code already written and explained. You write 2 functions, each marked # YOUR CODE HERE and followed by a test cell:
MSEFunction.forward- Compute mean squared error between predictions and targets.
CrossEntropyFunction.forward- Compute cross-entropy loss between logits and target class indices.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.core.losses; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (4). Each prints a ✅ line when it passes.
- Log-Softmax
- MSE Loss
- Cross-Entropy Loss
- Binary Cross-Entropy Loss
Integration tests after export (12).
tests/04_losses/test_04_losses_progressive.py
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
Expected 0.18000000000000002, got 0.5400000214576721-
From
test_unit_mse_loss. The squared errors were summed. Mean squared error divides by the number of elements: usenp.mean. Uniform predictions should have loss ≈ log(3) = 1.099, got -1.099-
From
test_unit_cross_entropy_loss. The sign is missing. Cross-entropy is the negative mean of the selected log-probabilities. Loss should not be NaN with large logits-
From
test_unit_cross_entropy_loss. The log-probabilities came fromnp.log(np.exp(...) / ...), which overflows. UseLogSoftmax, which subtracts the row maximum first.
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Losses: Cross-Entropy and Numerical Stability in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
You now have a feedback signal. The next problem is feeding it: a single sample at a time is too slow, an entire dataset at once won’t fit in memory, and unshuffled data trains a different model than shuffled data. Module 05 builds the DataLoader — batching, shuffling, and iteration — so the loss you just wrote can be averaged over B samples per step instead of one.
Next: Module 05: DataLoader
How later modules use this one
| Module | What It Does | Your Loss In Action |
|---|---|---|
| 05: DataLoader | Batching + shuffling | Feeds (inputs, targets) batches whose predictions your loss scores every step |
| 06: Autograd | Automatic differentiation | loss.backward() traces the loss back into parameter gradients |
| 07: Optimizers | Parameter updates | optimizer.step() consumes those gradients to shrink the loss |
| 08: Training | Complete training loop | loss = criterion(outputs, targets) becomes the heartbeat of every epoch |