Module 04: Losses

A loss function reduces a full output tensor to a single scalar. That reduction is the only signal the optimizer ever sees. Build the reduction wrong, or apply it to the wrong axis, and every gradient flowing backward through autograd optimizes the wrong objective.

NoteModule Info

FOUNDATION TIER | Difficulty: ●●○○ | Time: 4-6 hours | Prerequisites: 01, 02, 03

Prerequisites: Modules 01-03 means you should understand:

  • Tensor operations and broadcasting (Module 01)
  • Activation functions and their role in neural networks (Module 02)
  • Layers and how they transform data (Module 03)

If you can build a simple neural network that takes input and produces output, you’re ready to learn how to measure its quality.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

A neural network without a loss function is a guess machine. The loss is the single scalar that tells optimization which direction to move — turning a forward pass into a learning step. Get it wrong and training stalls, diverges, or silently optimizes the wrong objective.

In this module you’ll work with three losses that cover most supervised learning, writing the first two yourself: Mean Squared Error for regression, CrossEntropy for multi-class classification, and Binary Cross-Entropy for multi-label or binary decisions. Along the way you’ll implement the log-sum-exp trick — the one numerical safeguard that separates a softmax that trains from one that returns nan on the first batch with large logits.

By the end, you’ll know not just how to compute each loss, but why the choice of loss reshapes what your model learns, and where naive implementations break at production scale.

Commands

# first time
tito module start 04

# later sessions
tito module resume 04

# when your tests pass
tito module complete 04

Your notebook is modules/04_losses/losses.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement MSELoss for regression and CrossEntropyLoss for multi-class classification, then read the supplied BinaryCrossEntropyLoss for binary decisions
  • Master the log-sum-exp trick for numerically stable softmax computation
  • Understand computational complexity (O(B×C) for cross-entropy with large vocabularies) and memory trade-offs
  • Analyze loss function behavior across different prediction patterns and confidence levels
  • Connect your implementation to production PyTorch patterns and engineering decisions at scale

What you’ll build

Figure 1: Stable cross-entropy on large logits. For logits [1000, 1002, 1001] and target class 1, direct float32 exponentiation overflows. Shifting by the maximum gives [-2, 0, -1], and log-sum-exp yields a finite scalar loss of approximately 0.4076. Cross-entropy negates the target log-probability; decimals are rounded.

The pattern you’ll enable:

# Measuring prediction quality
loss = criterion(predictions, targets)  # Scalar feedback signal for learning

What you’re not building yet

To keep this module focused, you will not implement:

  • Class weights (handling imbalanced datasets)
  • Label smoothing (regularization technique)
  • Advanced loss functions (Focal Loss, Triplet Loss, Contrastive Loss)
  • GPU acceleration (your NumPy implementation runs on CPU)

These three losses cover most supervised training you will do in this course.

What you write

The notebook arrives with the surrounding code already written and explained. You write 2 functions, each marked # YOUR CODE HERE and followed by a test cell:

MSEFunction.forward
Compute mean squared error between predictions and targets.
CrossEntropyFunction.forward
Compute cross-entropy loss between logits and target class indices.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.losses;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (4). Each prints a ✅ line when it passes.

  • Log-Softmax
  • MSE Loss
  • Cross-Entropy Loss
  • Binary Cross-Entropy Loss

Integration tests after export (12).

  • tests/04_losses/test_04_losses_progressive.py

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

Expected 0.18000000000000002, got 0.5400000214576721
From test_unit_mse_loss. The squared errors were summed. Mean squared error divides by the number of elements: use np.mean.
Uniform predictions should have loss ≈ log(3) = 1.099, got -1.099
From test_unit_cross_entropy_loss. The sign is missing. Cross-entropy is the negative mean of the selected log-probabilities.
Loss should not be NaN with large logits
From test_unit_cross_entropy_loss. The log-probabilities came from np.log(np.exp(...) / ...), which overflows. Use LogSoftmax, which subtracts the row maximum first.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Losses: Cross-Entropy and Numerical Stability in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

NoteUp next: Module 05, DataLoader

You now have a feedback signal. The next problem is feeding it: a single sample at a time is too slow, an entire dataset at once won’t fit in memory, and unshuffled data trains a different model than shuffled data. Module 05 builds the DataLoader — batching, shuffling, and iteration — so the loss you just wrote can be averaged over B samples per step instead of one.

Next: Module 05: DataLoader

How later modules use this one

Table 1: How losses feed into subsequent training modules.
Module What It Does Your Loss In Action
05: DataLoader Batching + shuffling Feeds (inputs, targets) batches whose predictions your loss scores every step
06: Autograd Automatic differentiation loss.backward() traces the loss back into parameter gradients
07: Optimizers Parameter updates optimizer.step() consumes those gradients to shrink the loss
08: Training Complete training loop loss = criterion(outputs, targets) becomes the heartbeat of every epoch
Back to top