Module 11: Embeddings
Embeddings are the largest single tensor in most language models (GPT-3’s 2.3 GB) but every forward pass only touches a handful of rows. That makes them a row-sparse scatter/gather problem, not a matmul, and it puts them squarely in the memory-bandwidth regime of the roofline. This module builds the lookup table, the positional encodings, and the scatter-add backward pass that every transformer input layer depends on.
ARCHITECTURE TIER | Difficulty: ●●○○ | Time: 3-5 hours | Prerequisites: 01-08, 10
Prerequisites: Modules 01-08 and 10 means you should understand:
- Tensor operations (shape manipulation, matrix operations, broadcasting)
- Training fundamentals (forward/backward, optimization)
- Tokenization (converting text to token IDs, vocabularies)
If you can explain how a tokenizer converts “hello” to token IDs and how to multiply matrices, you’re ready.
Overview
Neural networks operate on vectors. Language is made of tokens. Embeddings are how the two meet: a learnable lookup table that turns each integer token ID into a dense vector, so the rest of the network can do calculus on it.
Your tokenizer from Module 10 produces IDs like [42, 7, 15]. By the end of this module you have built the layer that turns those IDs into geometry — the same layer sitting at the input of every transformer from BERT to GPT-4 — and the positional encodings that tell the network where in the sequence each token lives.
Commands
# first time
tito module start 11
# later sessions
tito module resume 11
# when your tests pass
tito module complete 11Your notebook is modules/11_embeddings/embeddings.ipynb.
Learning objectives
- Implement the embedding forward lookup that converts token IDs to dense vectors, and the backward pass that scatters gradients back into the table
- Master positional encoding strategies including learned and sinusoidal approaches
- Understand memory scaling for embedding tables and the trade-offs between vocabulary size and embedding dimension
- Connect your implementation to production transformer architectures used in GPT and BERT
What you’ll build
The pattern you’ll enable:
# Converting tokens to position-aware dense vectors
embed_layer = EmbeddingLayer(vocab_size=50000, embed_dim=512)
tokens = Tensor([[1, 42, 7]]) # Token IDs from tokenizer
embeddings = embed_layer(tokens) # (1, 3, 512) dense vectors ready for attentionWhat you’re not building yet
To keep this module focused, you will not implement:
- Attention mechanisms (that’s Module 12: Attention)
- Full transformer architectures (that’s Module 13: Transformers)
- Word2Vec or GloVe pretrained embeddings (you’re building learnable embeddings)
- Subword embedding composition (PyTorch handles this at the tokenization level)
You are building the foundation for sequence models. Context-aware representations come next.
What you write
The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:
EmbeddingFunction.backward- Compute gradient for embedding lookup.
Embedding.forward- Forward pass: lookup embeddings for given indices.
_compute_sinusoidal_table- Compute the raw sinusoidal positional encoding table as a numpy array.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.core.embeddings; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (9). Each prints a ✅ line when it passes.
- Embedding.__init__
- Embedding.forward
- Embedding gradients
- PositionalEncoding.__init__
- PositionalEncoding.forward
- Sinusoidal Table Computation
- Sinusoidal Embeddings
- EmbeddingLayer Initialization
- Complete Embedding System
Integration tests after export (16).
tests/11_embeddings/test_11_embeddings_progressive.py
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
Row 0 is used twice so its gradient should be [2, 2], got [1. 1.]-
From
test_unit_embedding_backward.grad_weight[indices] += gradcounts a repeated ID once. Usenp.add.at(grad_weight, indices, grad), which accumulates every occurrence. Even dims at pos 0 should be sin(0)=0-
From
test_unit_sinusoidal_table. Sine and cosine are swapped. Even columns takenp.sin, odd columnsnp.cos. Lower dims should oscillate faster-
From
test_unit_sinusoidal_table. The frequency exponent has the wrong sign. The frequency term isexp(-log(10000) * 2i / d), which shrinks as the dimension index grows.
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Embeddings and Positional Encoding: Turning Token IDs into Vectors in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
You now have static geometry: every token ID maps to a fixed vector in space, with positional information layered on top. But the embedding for “bank” is the same whether the next word is “river” or “loan” — the representation has no context. That is the question Module 12 answers.
You’ll build scaled dot-product attention: the mechanism that lets every embedding look at every other embedding in the sequence and re-weight itself accordingly. Your static (B, S, D) embedding tensor becomes a context-aware (B, S, D) tensor where each position has been shaped by the rest of the sentence.
Next: Module 12: Attention
How later modules use this one
| Module | What It Does | Your Embeddings In Action |
|---|---|---|
| 12: Attention | Context-aware representations | attention(embed_layer(tokens)) produces query, key, value |
| 13: Transformers | Decoder-only GPT stack | GPT.forward(tokens) embeds tokens and positions before the blocks |