Datasets Guide
TinyTorch follows an offline-first data philosophy. Four curated micro-datasets ship directly inside the repository as about 475 KB of data, enabling instantaneous repository cloning, zero-network classroom labs, and 5-second CPU training cycles. When scaling beyond toy regimes, standard vision and language benchmarks download on demand from their original hosts and are cached locally.
Design Philosophy
TinyTorch adopts a strict offline-first, dual-tier dataset architecture:
1. Bundled Micro-Datasets (TinyVerse)
About 475 KB of data shipped directly in the Git repository. Zero network dependencies, instant CPU execution in 5 to 30 seconds, 100% offline verification for classroom labs, and transparent debugging where every tensor activation fits inside a terminal window.
2. Scalable Benchmarks (Hugging Face Hub)
Industry-standard benchmarks downloaded on demand and cached locally in milestones/datasets/. Maintained under the official harvard-edge organization on Hugging Face with interactive confirmation before download and full Gebru et al. datasheets.
Core Principle: Following Andrej Karpathy’s minimal sample methodology: compact, curated datasets for learning core mechanisms, and standard benchmarks for industrial scaling.
Dataset Catalog
The following table summarizes the TinyTorch dataset suite across bundled offline packages and scalable on-demand benchmarks.
| Dataset | Type | Payload | Format | Target Milestones | Core Task |
|---|---|---|---|---|---|
| TinyDigits | Bundled | 310 KB | Python pickle (.pkl) |
Milestones 03, 04, 06 | 8x8 handwritten digit classification |
| TinyPy | Bundled | 73 KB | Python source (.txt) |
Milestone 05 Part 3 | TinyCopilot code generation and AST compiler validation |
| TinyTalks | Bundled | 63 KB | Plain text dialog (.txt) |
Milestone 05 Part 4 | Conversational Q&A and memorization versus generalization |
| TinyShakespeare | Bundled | 21 KB | Plain text verse (.txt) |
Milestone 05 Part 1 | Character-level language modeling and autoregressive sampling |
| CIFAR-10 | On-Demand | ~170 MB | Binary archive (.tar.gz) |
Milestone 04 Part 2 | Natural 32x32 RGB image classification and convolution scaling |
| MNIST | On-Demand | ~11 MB | Gzip binary (.gz) |
Optional Scale-Up | Full-resolution 28x28 handwritten digit benchmark |
The bundled data totals about 475 KB, keeping repository cloning instantaneous on any network.
Bundled Micro-Datasets
These four datasets ship directly inside the TinyTorch repository under the datasets/ directory. They require no internet access, no external downloads, and no API keys.
TinyDigits
Handwritten Digit Recognition · 8x8 Grayscale · Bundled
- Repository location:
datasets/tinydigits/ - Payload size: ~310 KB (
train.pkl: 258 KB,test.pkl: 52 KB) - Sample count: 1,000 training images, 200 test images (10 balanced classes, digits 0 to 9)
- Data shape and dtype:
(N, 8, 8)or flattened(N, 64), NumPyfloat32normalized to[0.0, 1.0], labelsint64 - Target milestones: Milestone 03 (MLP Revival), Milestone 04 (CNN Revolution), Milestone 06 (MLPerf Serving)
Systems and Learning Role
Why 8x8 grayscale instead of full-sized 28x28 MNIST?
- Rapid CPU training: The Milestone 03 MLP reaches over 80 percent in about a second, and the Milestone 04 CNN reaches about 85 percent in about two minutes, on a laptop CPU (both milestones pass at 75 percent). Students iterate on mathematical logic and loss functions without waiting for long training runs.
- Transparent debugging: An 8x8 grid contains exactly 64 values. Spatial activation maps and convolutional filter weights print entirely inside a standard terminal window without line wraps or ellipsis truncation.
- Architectural equivalence: Even at 8x8 resolution, the data exercises the exact same 2D convolution (
Conv2d), max pooling (MaxPool2d), stride indexing, tape autograd, and softmax cross-entropy gradients as industrial vision models.
Loading and Usage
from milestones.data_manager import DatasetManager
# Load via DatasetManager helper
dm = DatasetManager()
(x_train, y_train), (x_test, y_test) = dm.get_tinydigits()
print(f"Train images: {x_train.shape}, dtype: {x_train.dtype}")
print(f"Test images: {x_test.shape}, dtype: {x_test.dtype}")
# Train images: (1000, 8, 8), dtype: float32
# Test images: (200, 8, 8), dtype: float32Alternatively, load the raw pickle files directly:
import pickle
from pathlib import Path
data_dir = Path("datasets/tinydigits")
with open(data_dir / "train.pkl", "rb") as f:
train_dict = pickle.load(f)
x_train, y_train = train_dict["images"], train_dict["labels"]TinyPy
TinyCopilot Python Code Generation · Bundled
- Repository location:
datasets/tinypy/ - Payload size: ~73 KB (
tinypy_sample.txt) - Sample count: 132 top-level Python functions and 24 classes
- Data format: UTF-8 plain text Python source code
- Target milestones: Milestone 05 Part 3 (TinyCopilot:
03_tinycopilot.py)
Systems and Learning Role
Why train a language model on Python code rather than English prose?
- Rigid syntactic structure: Code enforces strict indentation scoping, colon delimiters, bracket balancing, and variable namespace rules. A character-level model must learn these formal structural constraints.
- The AST compiler gate: Natural language text quality is subjective and difficult to evaluate automatically. In contrast, generated Python code can be passed directly to Python’s built-in Abstract Syntax Tree parser (
ast.parse). This creates an objective compiler gate that scores whether generated completions are syntactically executable.
Loading and Usage
from milestones.data_manager import DatasetManager
# Load bundled TinyPy source corpus
dm = DatasetManager()
code_corpus = dm.get_tinypy(sample_only=True)
print(f"Corpus size: {len(code_corpus):,} characters")
print(code_corpus[:180])TinyTalks
Conversational Dialogue and the Overfitting Detective · Bundled
- Repository location:
datasets/tinytalks/ - Payload size: ~63 KB combined (
tinytalks_v1.txt: 18 KB,tinytalks_tinytorch.txt: 13.7 KB, and five train/val/test split files: 31.7 KB) - Sample count: 301 general Q&A pairs plus 81 TinyTorch concept pairs
- Data format: Plain text formatted as
Q: <question>\nA: <answer>\n\n - Target milestones: Milestone 05 Part 4 (
04_tinygpt_chat.py)
Systems and Learning Role
TinyTalks supports two distinct learning objectives:
- Conversational turn-taking: Formatted Q&A pairs teach a word-level transformer how to recognize prompt prefixes and generate coherent responses in an interactive terminal chat.
- The Overfitting Detective experiment: A 218,000-parameter TinyGPT model contains far more capacity than the 13 KB concept dataset has characters. Students train on the training split (
tinytorch_train.txt) and evaluate on the held-out test split (tinytorch_test.txt). The milestone finds the epoch with the lowest held-out loss and reports overfitting from the measured curve: on quick runs held-out loss is lowest at epoch 1 and rises while training loss keeps falling.
Loading and Usage
from milestones.data_manager import DatasetManager
dm = DatasetManager()
# Load train and test splits for the Overfitting Detective
train_qa = dm.get_tinytalks(sample_only=True, topic="tinytorch", split="train")
test_qa = dm.get_tinytalks(sample_only=True, topic="tinytorch", split="test")
print(f"Train split: {len(train_qa):,} characters")
print(f"Test split: {len(test_qa):,} characters")TinyShakespeare
Autoregressive Language Modeling · Bundled
- Repository location:
datasets/tinyshakespeare/ - Payload size: ~21 KB bundled sample (
tinyshakespeare_sample.txt); ~1.1 MB full dataset on demand - Sample count: ~21,000 characters in sample; ~1.1 million characters in full text
- Data format: Plain text UTF-8 character stream
- Target milestones: Milestone 05 Part 1 (
01_tinygpt_shakespeare.py)
Systems and Learning Role
- Classic autoregressive modeling: Serves as the classic character-level language modeling benchmark popularized by Andrej Karpathy.
- Generative sampling dynamics: Provides rich Elizabethan vocabulary and dramatic cadence to test temperature-scaled sampling, top-k filtering, and next-token probability distributions without network downloads.
Loading and Usage
from milestones.data_manager import DatasetManager
dm = DatasetManager()
# Instant offline sample (21 KB)
sample_text = dm.get_tinyshakespeare(sample_only=True)
print(f"Offline sample length: {len(sample_text):,} characters")
# Full dataset (prompts before downloading 1.1 MB)
# full_text = dm.get_tinyshakespeare(sample_only=False)Scalable Benchmarks
When scaling implementations to industrial workloads, TinyTorch provides automated access to standard benchmarks. These datasets are downloaded on demand and cached locally in milestones/datasets/. CIFAR-10 asks before its ~170 MB download (set TINYTORCH_AUTO_DOWNLOAD=1 to skip the prompt); MNIST downloads without asking.
CIFAR-10
Natural Image Classification · 32x32 RGB · On-Demand
- Storage path:
milestones/datasets/cifar-10/ - Download size: ~170 MB compressed archive (~180 MB extracted)
- Sample count: 50,000 training images, 10,000 test images (10 balanced natural object classes)
- Data format and shape:
(N, 3, 32, 32)NumPy array, 3-channel RGB, uint8 or normalized float32 - Target milestones: Milestone 04 Part 2 (
02_lecun_cifar10.py)
Systems and Learning Role
- Multi-channel convolution: Scales student
Conv2dlayers from single-channel 8x8 grayscale to 3-channel color images with 32x32 spatial dimensions. - Exposing compute bottlenecks: Training on 50,000 color images exposes the runtime cost of nested Python loops in naive convolution. This provides empirical motivation for the im2col transformation in Module 17 and the custom kernels of Milestone 07.
Loading and Usage
from milestones.data_manager import DatasetManager
# Interactive confirmation prompt before download
dm = DatasetManager()
(x_train, y_train), (x_test, y_test) = dm.get_cifar10()
print(f"CIFAR-10 train shape: {x_train.shape}, test shape: {x_test.shape}")
# CIFAR-10 train shape: (50000, 3, 32, 32), test shape: (10000, 3, 32, 32)MNIST
Full-Scale Handwritten Digits · 28x28 Grayscale · On-Demand
- Storage path:
milestones/datasets/mnist/ - Download size: ~11 MB compressed (~55 MB extracted)
- Sample count: 60,000 training images, 10,000 test images (10 balanced classes, digits 0 to 9)
- Data format and shape:
(N, 28, 28)NumPy float32 normalized to[0.0, 1.0], labelsint64 - Target milestones: Optional scaling benchmark across Milestone 03 and Milestone 04
Systems and Learning Role
- Historical parity: Reproduces Yann LeCun’s 1998 LeNet-5 benchmark at authentic 28x28 resolution.
- Resolution scaling: Proves that dense layers scale quadratically with input resolution (28x28 = 784 inputs vs 8x8 = 64 inputs), whereas convolutional layer parameter counts remain fixed by kernel size (3x3 = 9 weights per filter).
Loading and Usage
from milestones.data_manager import DatasetManager
dm = DatasetManager()
(x_train, y_train), (x_test, y_test) = dm.get_mnist()
print(f"MNIST train shape: {x_train.shape}, test shape: {x_test.shape}")
# MNIST train shape: (60000, 28, 28), test shape: (10000, 28, 28)Connecting Datasets to Systems Milestones
The TinyVerse suite supplies the input data across every milestone in the curriculum:
| Tier | Milestones | Primary Datasets | Systems and Algorithmic Role |
|---|---|---|---|
| Foundation | 01, 02, 03 | Synthetic XOR, TinyDigits | Validates linear separation limits, non-linear activation learning, and dense backpropagation on CPU in under 5 seconds. |
| Architecture | 04, 05 | TinyDigits, CIFAR-10, TinyShakespeare, TinyPy, TinyTalks | Demonstrates spatial parameter sharing in convolutions and next-token prediction in causal transformers across code, dialogue, and verse. |
| Optimization | 06 | TinyDigits, TinyGPT | Serves as the benchmark suite for INT8 post-training quantization, structured pruning, and KV-cache memoization along the MLPerf Pareto frontier. |
| Hardware | 07 | Seeded random matrices | Fixed-size GEMM and bias+GELU inputs cross the foreign function interface into C++ SIMD, Apple Metal MPS, and OpenAI Triton kernels. |
Hugging Face Hub Integration
The complete suite of TinyVerse datasets is maintained under the official Harvard Edge organization on the Hugging Face Hub:
- Organization:
harvard-edge - Collection:
tinytorch - Dataset repositories:
Standardized Documentation and Provenance
Each dataset repository contains a standardized Datasheet for Datasets following the Gebru et al. specification. These datasheets document data motivation, composition, collection methodology, preprocessing pipelines, and recommended usage boundaries.
Loading via the Hugging Face Datasets Library
All TinyVerse datasets can be loaded directly using the standard Hugging Face datasets Python package:
from datasets import load_dataset
# Load TinyDigits vision benchmark
tinydigits = load_dataset("harvard-edge/tinydigits")
print(tinydigits)
# Load TinyPy code completion corpus
tinypy = load_dataset("harvard-edge/tinypy")
print(tinypy["train"][0])
# Load TinyTalks conversational Q&A
tinytalks = load_dataset("harvard-edge/tinytalks")
print(tinytalks["train"][0])