Module 11: Embeddings

Embeddings are the largest single tensor in most language models (GPT-3’s 2.3 GB) but every forward pass only touches a handful of rows. That makes them a row-sparse scatter/gather problem, not a matmul, and it puts them squarely in the memory-bandwidth regime of the roofline. This module builds the lookup table, the positional encodings, and the scatter-add backward pass that every transformer input layer depends on.

NoteModule Info

ARCHITECTURE TIER | Difficulty: ●●○○ | Time: 3-5 hours | Prerequisites: 01-08, 10

Prerequisites: Modules 01-08 and 10 means you should understand:

  • Tensor operations (shape manipulation, matrix operations, broadcasting)
  • Training fundamentals (forward/backward, optimization)
  • Tokenization (converting text to token IDs, vocabularies)

If you can explain how a tokenizer converts “hello” to token IDs and how to multiply matrices, you’re ready.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

Neural networks operate on vectors. Language is made of tokens. Embeddings are how the two meet: a learnable lookup table that turns each integer token ID into a dense vector, so the rest of the network can do calculus on it.

Your tokenizer from Module 10 produces IDs like [42, 7, 15]. By the end of this module you have built the layer that turns those IDs into geometry — the same layer sitting at the input of every transformer from BERT to GPT-4 — and the positional encodings that tell the network where in the sequence each token lives.

Commands

# first time
tito module start 11

# later sessions
tito module resume 11

# when your tests pass
tito module complete 11

Your notebook is modules/11_embeddings/embeddings.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement the embedding forward lookup that converts token IDs to dense vectors, and the backward pass that scatters gradients back into the table
  • Master positional encoding strategies including learned and sinusoidal approaches
  • Understand memory scaling for embedding tables and the trade-offs between vocabulary size and embedding dimension
  • Connect your implementation to production transformer architectures used in GPT and BERT

What you’ll build

Figure 1: Token lookup and position addition. Token IDs 2, 0, 2 gather rows W2, W0, W2. Position vectors broadcast over the batch and are added to token vectors; repeated IDs accumulate gradients into the same table row. Positions may be learned or sinusoidal, within the configured position-table length, and optional token scaling precedes the addition.

The pattern you’ll enable:

# Converting tokens to position-aware dense vectors
embed_layer = EmbeddingLayer(vocab_size=50000, embed_dim=512)
tokens = Tensor([[1, 42, 7]])  # Token IDs from tokenizer
embeddings = embed_layer(tokens)  # (1, 3, 512) dense vectors ready for attention

What you’re not building yet

To keep this module focused, you will not implement:

  • Attention mechanisms (that’s Module 12: Attention)
  • Full transformer architectures (that’s Module 13: Transformers)
  • Word2Vec or GloVe pretrained embeddings (you’re building learnable embeddings)
  • Subword embedding composition (PyTorch handles this at the tokenization level)

You are building the foundation for sequence models. Context-aware representations come next.

What you write

The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:

EmbeddingFunction.backward
Compute gradient for embedding lookup.
Embedding.forward
Forward pass: lookup embeddings for given indices.
_compute_sinusoidal_table
Compute the raw sinusoidal positional encoding table as a numpy array.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.embeddings;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (9). Each prints a ✅ line when it passes.

  • Embedding.__init__
  • Embedding.forward
  • Embedding gradients
  • PositionalEncoding.__init__
  • PositionalEncoding.forward
  • Sinusoidal Table Computation
  • Sinusoidal Embeddings
  • EmbeddingLayer Initialization
  • Complete Embedding System

Integration tests after export (16).

  • tests/11_embeddings/test_11_embeddings_progressive.py

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

Row 0 is used twice so its gradient should be [2, 2], got [1. 1.]
From test_unit_embedding_backward. grad_weight[indices] += grad counts a repeated ID once. Use np.add.at(grad_weight, indices, grad), which accumulates every occurrence.
Even dims at pos 0 should be sin(0)=0
From test_unit_sinusoidal_table. Sine and cosine are swapped. Even columns take np.sin, odd columns np.cos.
Lower dims should oscillate faster
From test_unit_sinusoidal_table. The frequency exponent has the wrong sign. The frequency term is exp(-log(10000) * 2i / d), which shrinks as the dimension index grows.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Embeddings and Positional Encoding: Turning Token IDs into Vectors in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

You now have static geometry: every token ID maps to a fixed vector in space, with positional information layered on top. But the embedding for “bank” is the same whether the next word is “river” or “loan” — the representation has no context. That is the question Module 12 answers.

NoteUp next: Module 12, Attention

You’ll build scaled dot-product attention: the mechanism that lets every embedding look at every other embedding in the sequence and re-weight itself accordingly. Your static (B, S, D) embedding tensor becomes a context-aware (B, S, D) tensor where each position has been shaped by the rest of the sentence.

Next: Module 12: Attention

How later modules use this one

Table 1: How embeddings feed into attention and transformer modules.
Module What It Does Your Embeddings In Action
12: Attention Context-aware representations attention(embed_layer(tokens)) produces query, key, value
13: Transformers Decoder-only GPT stack GPT.forward(tokens) embeds tokens and positions before the blocks
Back to top