Module 05: DataLoader

Training is I/O-bound before it is compute-bound. DataLoader is where production systems overlap disk reads, preprocessing, and accelerator compute so the GPU never sits idle. Your batch, shuffle, and collate logic is the synchronous core that those systems parallelize.

NoteModule Info

FOUNDATION TIER | Difficulty: ●●○○ | Time: 3-5 hours | Prerequisites: 01-04

Prerequisites: Modules 01-04 means you should be comfortable with tensors, activations, layers, and losses. This module introduces data loading infrastructure that will be used by autograd, optimizers, and training loops in the following modules.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

A naive training loop reaches into a 50,000-image dataset, picks one sample, computes a gradient, and repeats. It works. It also wastes the GPU and gets the math wrong: gradients computed on a sorted sequence of samples are not the gradients you intended. Every framework solves this with the same abstraction — a DataLoader sitting between storage and computation, turning raw samples into shuffled, contiguous batches.

In this module you build that abstraction. A Dataset says how to find sample i. A DataLoader decides how many samples to group, in what order, and when to load them. The result is a single iterator that works identically on 1,000 tensors in RAM or 100 GB of JPEGs on disk — and that you will reuse, unchanged, in every later module that trains a model.

Commands

# first time
tito module start 05

# later sessions
tito module resume 05

# when your tests pass
tito module complete 05

Your notebook is modules/05_dataloader/dataloader.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Trace the supplied Dataset abstraction and TensorDataset to see how in-memory data storage is exposed to the loader
  • Implement the DataLoader iteration method, the batching and shuffling loop that makes iteration memory efficient
  • Master the Python iterator protocol for streaming data without loading entire datasets
  • Analyze throughput bottlenecks and memory scaling characteristics with different batch sizes
  • Connect your implementation to PyTorch data loading patterns used in production ML systems

What you’ll build

Figure 1: TinyTorch Data Pipeline: From raw dataset storage to training-ready batches.

The pattern you’ll enable:

# Transform individual samples into training-ready batches
dataset = TensorDataset(features, labels)
loader = DataLoader(dataset, batch_size=32, shuffle=True)

for batch_features, batch_labels in loader:
    # batch_features shape: (32, 784)
    # batch_labels shape: (32,)
    train_step(batch_features, batch_labels)

What you’re not building yet

To keep this module focused, you will not implement:

  • Multi-process data loading (PyTorch uses num_workers for parallel loading)
  • Automatic dataset downloads (you’ll use pre-downloaded data or write custom loaders)
  • Prefetching mechanisms (loading next batch while GPU processes current batch)
  • Custom collation functions for variable-length sequences

You are building the batching foundation. Parallel loading is left to production frameworks.

What you write

The notebook arrives with the surrounding code already written and explained. You write one function, each marked # YOUR CODE HERE and followed by a test cell:

DataLoader.__iter__
Return iterator over batches.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.dataloader;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (8). Each prints a ✅ line when it passes.

  • Dataset Abstract Base Class
  • TensorDataset
  • _pad_image
  • _random_crop_region
  • Data Augmentation Transforms
  • DataLoader
  • DataLoader Deterministic Shuffling
  • Training Workflow

Integration tests after export (12).

  • tests/05_dataloader/test_05_dataloader_progressive.py

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

Expected 3 batches, got 2
From test_unit_dataloader. The final partial batch was dropped. Step through the indices with range(0, len(indices), self.batch_size) and yield the short last batch too.
AttributeError: 'tuple' object has no attribute 'data'
From test_unit_dataloader. __iter__ yielded the list of samples. Pass it through self._collate_batch(batch), which stacks it into one features tensor and one labels tensor.
Different seeds should produce different shuffles
From test_unit_dataloader_deterministic. __iter__ never shuffled. When self.shuffle is true, call random.shuffle(indices) at the start of every pass.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Dataloaders: Batching and Memory Management in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

NoteUp next: Module 06, Autograd

You can now move data through a model. You cannot yet learn from it — every batch leaves the network as a loss number with no gradient attached. Module 06 fixes that by building automatic differentiation: every tensor remembers the operations that produced it, and loss.backward() walks the resulting graph to assign a gradient to every parameter the loader’s batch touched.

DataLoader and autograd compose directly: the iterator you just built becomes the input edge of every computation graph in the rest of the book.

Next: Module 06: Autograd

How later modules use this one

Table 1: How the DataLoader feeds into subsequent training modules.
Module What it does Your DataLoader in action
06: Autograd Reverse-mode differentiation Each batch tensor becomes a leaf of the computation graph
08: Training End-to-end training loops for batch in loader: is the outer loop of every example
09: Convolutions Convolutional layers The same iterator now feeds 4-D image batches to CNNs
Back to top