Tensor Tetris
Pack tensors into HBM before you OOM.
How to play
Tetris, but with tensor blocks. Each falling block represents one of the four canonical training-memory citizens, sized roughly to its real footprint:
- Parameters (purple, 3 cells) — the model weights themselves. Spawn once, anchor the bottom of HBM, never freed during training.
- Activations (blue, 6 cells) — the largest and most frequent: forward-pass outputs that backward needs. Two variants spawn often.
- Gradients (red, 1 cell) — small but plentiful: one per parameter, written during backward, consumed by
step(). - Optimizer states (orange, 6 cells) — Adam keeps an fp32 master + first moment + second moment per parameter, so ~6× weight bytes. Big and persistent.
Pack them into the HBM box without the stack overflowing.
- Every 3 placements triggers a backward event: the newest activations are consumed (LIFO — the order autograd actually uses).
- Every 6 placements triggers a step() event: gradients are used and cleared.
- Parameters and optimizer states are never freed.
Important: blocks do not disappear because rows line up. The Tetris packing is the visual metaphor; the freeing rule comes from the training loop. Memory only disappears when backward or step() fires.
The Systems Concept
ML practitioners spend a lot of time tuning batch size, gradient accumulation, and recomputation strategies because GPU memory is the binding constraint at scale. The single most important fact: activations dominate the training memory bill, not weights — which is why selective recomputation, ZeRO sharding, and smaller batches all exist. Korthikanti et al. 2022 is the canonical reference; PyTorch’s autograd allocator manages this every step.
A note on simplifications: real frameworks free on graph traversal and optimizer.zero_grad(set_to_none=True), not on a placement counter — the rhythm in this game is for your eyes, not PyTorch’s.
Part of MLSysBook Playground. Found a bug? Report an issue.