Architecture Tier (Modules 09-13)
Build modern neural architectures: from computer vision to language models.
What You’ll Learn
The Architecture tier teaches you how to build the neural network architectures that power modern AI. You’ll implement CNNs for computer vision and transformers for language understanding, building on the foundational training infrastructure from the previous tier.
By the end of this tier, you’ll understand:
- Why convolutional layers are essential for computer vision
- How attention mechanisms enable transformers to understand sequences
- What embeddings do to represent discrete tokens as continuous vectors
- How modern architectures compose these components into powerful systems
Module Progression
Why This Order?
The Architecture tier branches into two parallel tracks (Vision and Language) because these domains have fundamentally different data structures and operations. But both follow the same principle: build components in the order they compose.
Vision Track: Spatial Processing (09)
Convolutions (09) stands alone because CNNs have a relatively simple pipeline:
- Images come in, convolutions extract spatial features, pooling reduces dimensions
- One module gives you everything needed for computer vision
Language Track: Sequential Processing (10-13)
Tokenization (10) → Embeddings (11) → Attention (12) → Transformers (13)
Language requires more infrastructure, and the order is non-negotiable:
- Tokenization converts text to integers: you cannot process raw strings
- Embeddings convert integers to vectors: attention needs continuous representations
- Attention computes context-aware representations: the core transformer operation
- Transformers compose attention with MLPs and normalization: the complete architecture
Each step transforms the data representation:
"hello" → [72, 101, 108, 108, 111] → [[0.1, 0.3, ...], [...]] → attention → output
text token IDs embeddings transformer
Why Not Merge Them?
Vision and language students have different goals. A computer vision engineer building image classifiers doesn’t need tokenization; an NLP engineer building chatbots doesn’t need convolutions. Parallel tracks let students focus on their domain while building on shared foundations.
Module Details
09. Convolutions - Convolutional Neural Networks
What it is: Conv2d (convolutional layers) and pooling operations for processing images.
Why it matters: CNNs revolutionized computer vision by exploiting spatial structure. Understanding convolutions, kernels, and pooling is essential for image processing and beyond.
What you’ll build: Conv2d, MaxPool2d, and related operations with proper gradient computation.
Systems focus: Spatial operations, memory layout (channels), computational intensity
Historical impact: This module enables Milestone 04 (1998 CNN Revolution) : a LeNet-style CNN trained on handwritten digits with your implementations.
10. Tokenization - From Text to Numbers
What it is: Converting text into integer sequences that neural networks can process.
Why it matters: Neural networks operate on numbers, not text. Tokenization is the bridge between human language and machine learning: understanding vocabulary, encoding, and decoding is fundamental.
What you’ll build: Character-level and subword tokenizers with vocabulary management and encoding/decoding.
Systems focus: Vocabulary management, encoding schemes, out-of-vocabulary handling
11. Embeddings - Learning Representations
What it is: Learned mappings from discrete tokens (words, characters) to continuous vectors.
Why it matters: Embeddings transform sparse, discrete representations into dense, semantic vectors. Understanding embeddings is crucial for NLP, recommendation systems, and any domain with categorical data.
What you’ll build: Embedding layers with proper initialization and gradient computation.
Systems focus: Lookup tables, gradient backpropagation through indices, initialization
12. Attention - Context-Aware Representations
What it is: Self-attention mechanisms that let each token attend to all other tokens in a sequence.
Why it matters: Attention is the breakthrough that enabled modern LLMs. It allows models to capture long-range dependencies and contextual relationships that RNNs struggled with.
What you’ll build: Scaled dot-product attention, multi-head attention, and causal masking for autoregressive generation.
Systems focus: O(n²) memory/compute, masking strategies, numerical stability
13. Transformers - The Modern Architecture
What it is: Complete transformer architecture combining embeddings, attention, and feedforward layers.
Why it matters: Transformers power GPT, BERT, and virtually all modern LLMs. Understanding their architecture (positional encodings, layer normalization, residual connections) is essential for AI engineering.
What you’ll build: A complete decoder-only transformer (GPT-style) for autoregressive text generation.
Systems focus: Layer composition, residual connections, generation loop
Historical impact: This module enables Milestone 05 (2017 Transformer Era) - solving structured sequence tasks, reversal and copying, with YOUR attention implementation.
What You Can Build After This Tier
After completing the Architecture tier, you’ll be able to:
- Milestone 04 (1998): Train a LeNet-style CNN on TinyDigits (86–87% in our runs), with an optional CIFAR-10 scale-up
- Milestone 05 (2017): Implement transformers that solve structured sequence tasks
- Train on real datasets (TinyDigits, with an optional CIFAR-10 path)
- Understand why modern architectures (ResNets, Vision Transformers, LLMs) work
Prerequisites
Required:
- Foundation Tier (Modules 01-08) completed
- Understanding of tensors, data loaders, autograd, and training loops
- Basic understanding of images (height, width, channels)
- Basic understanding of text/language concepts
Helpful but not required:
- Computer vision concepts (convolution, feature maps)
- NLP concepts (tokens, vocabulary, sequence modeling)
Time Commitment
Per module: 3 to 10 hours; each module page gives its own estimate
Total tier: 26-36 hours, the sum of the per-module estimates for modules 09-13
Recommended pace: 1 module per week (2 modules/week for intensive study)
Learning Approach
Each module follows the Build → Use → Reflect cycle with real datasets:
- Build: Implement the architecture component (Conv2d, attention, transformers)
- Use: Train on real data (TinyDigits, with an optional CIFAR-10 path)
- Reflect: Analyze systems trade-offs (memory vs accuracy, speed vs quality)
Key Achievements
Milestone 04: CNN Revolution (1998)
After Module 09, you’ll recreate Yann LeCun’s breakthrough:
tito milestone run 04 # a LeNet-style CNN on TinyDigits; --part 2 scales to CIFAR-10What makes this special: You’re not just importing torch.nn.Conv2d: you built the entire convolutional architecture from scratch.
Milestone 05: Transformer Era (2017)
After Module 13, you’ll implement the attention revolution:
tito milestone run 05 # TinyGPT on Shakespeare, then reversal, copy, and mixed sequence tasksWhat makes this special: Structured sequence tasks stress-test YOUR attention and transformer stack: the same core pattern behind GPT-style models.
Two Parallel Tracks
The Architecture tier splits into two parallel paths. tito completes modules in numeric order, so Module 09 comes before 10–13 even if language is your focus:
Vision Track (Module 09):
- Convolutions (Conv2d + Pooling)
- Enables computer vision applications
- Culminates in CNN milestone
Language Track (Modules 10-13):
- Tokenization → Embeddings → Attention → Transformers
- Enables natural language processing
- Culminates in Transformer milestone
Order: Complete both tracks in order (09→10→11→12→13); tito module start stays locked until every earlier module is complete.
Next Steps
Ready to build modern architectures?
# Start the Architecture tier with vision
tito module start 09Or explore other tiers:
- Foundation Tier (Modules 01-08): Mathematical foundations
- Optimization Tier (Modules 14-19): Production-ready performance
- Torch Olympics (Module 20): Compete in ML systems challenges