Module 05: DataLoader
Training is I/O-bound before it is compute-bound. DataLoader is where production systems overlap disk reads, preprocessing, and accelerator compute so the GPU never sits idle. Your batch, shuffle, and collate logic is the synchronous core that those systems parallelize.
FOUNDATION TIER | Difficulty: ●●○○ | Time: 3-5 hours | Prerequisites: 01-04
Prerequisites: Modules 01-04 means you should be comfortable with tensors, activations, layers, and losses. This module introduces data loading infrastructure that will be used by autograd, optimizers, and training loops in the following modules.
Overview
A naive training loop reaches into a 50,000-image dataset, picks one sample, computes a gradient, and repeats. It works. It also wastes the GPU and gets the math wrong: gradients computed on a sorted sequence of samples are not the gradients you intended. Every framework solves this with the same abstraction — a DataLoader sitting between storage and computation, turning raw samples into shuffled, contiguous batches.
In this module you build that abstraction. A Dataset says how to find sample i. A DataLoader decides how many samples to group, in what order, and when to load them. The result is a single iterator that works identically on 1,000 tensors in RAM or 100 GB of JPEGs on disk — and that you will reuse, unchanged, in every later module that trains a model.
Commands
# first time
tito module start 05
# later sessions
tito module resume 05
# when your tests pass
tito module complete 05Your notebook is modules/05_dataloader/dataloader.ipynb.
Learning objectives
- Trace the supplied Dataset abstraction and TensorDataset to see how in-memory data storage is exposed to the loader
- Implement the DataLoader iteration method, the batching and shuffling loop that makes iteration memory efficient
- Master the Python iterator protocol for streaming data without loading entire datasets
- Analyze throughput bottlenecks and memory scaling characteristics with different batch sizes
- Connect your implementation to PyTorch data loading patterns used in production ML systems
What you’ll build
The pattern you’ll enable:
# Transform individual samples into training-ready batches
dataset = TensorDataset(features, labels)
loader = DataLoader(dataset, batch_size=32, shuffle=True)
for batch_features, batch_labels in loader:
# batch_features shape: (32, 784)
# batch_labels shape: (32,)
train_step(batch_features, batch_labels)What you’re not building yet
To keep this module focused, you will not implement:
- Multi-process data loading (PyTorch uses
num_workersfor parallel loading) - Automatic dataset downloads (you’ll use pre-downloaded data or write custom loaders)
- Prefetching mechanisms (loading next batch while GPU processes current batch)
- Custom collation functions for variable-length sequences
You are building the batching foundation. Parallel loading is left to production frameworks.
What you write
The notebook arrives with the surrounding code already written and explained. You write one function, each marked # YOUR CODE HERE and followed by a test cell:
DataLoader.__iter__- Return iterator over batches.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.core.dataloader; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (8). Each prints a ✅ line when it passes.
- Dataset Abstract Base Class
- TensorDataset
- _pad_image
- _random_crop_region
- Data Augmentation Transforms
- DataLoader
- DataLoader Deterministic Shuffling
- Training Workflow
Integration tests after export (12).
tests/05_dataloader/test_05_dataloader_progressive.py
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
Expected 3 batches, got 2-
From
test_unit_dataloader. The final partial batch was dropped. Step through the indices withrange(0, len(indices), self.batch_size)and yield the short last batch too. AttributeError: 'tuple' object has no attribute 'data'-
From
test_unit_dataloader.__iter__yielded the list of samples. Pass it throughself._collate_batch(batch), which stacks it into one features tensor and one labels tensor. Different seeds should produce different shuffles-
From
test_unit_dataloader_deterministic.__iter__never shuffled. Whenself.shuffleis true, callrandom.shuffle(indices)at the start of every pass.
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Dataloaders: Batching and Memory Management in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
You can now move data through a model. You cannot yet learn from it — every batch leaves the network as a loss number with no gradient attached. Module 06 fixes that by building automatic differentiation: every tensor remembers the operations that produced it, and loss.backward() walks the resulting graph to assign a gradient to every parameter the loader’s batch touched.
DataLoader and autograd compose directly: the iterator you just built becomes the input edge of every computation graph in the rest of the book.
Next: Module 06: Autograd
How later modules use this one
| Module | What it does | Your DataLoader in action |
|---|---|---|
| 06: Autograd | Reverse-mode differentiation | Each batch tensor becomes a leaf of the computation graph |
| 08: Training | End-to-end training loops | for batch in loader: is the outer loop of every example |
| 09: Convolutions | Convolutional layers | The same iterator now feeds 4-D image batches to CNNs |