Module 01: Tensor

A tensor is an array of numbers with a shape, and every other part of TinyTorch is built on it. In this module you write the Tensor class: how it stores its numbers, and how two tensors are added and multiplied. Why its memory layout matters for speed comes later, in the book and in the Optimization tier.

NoteModule Info

FOUNDATION TIER | Difficulty: ●○○○ | Time: 4-6 hours | Prerequisites: None

Prerequisites: None means exactly that. This module assumes:

  • Basic Python (lists, classes, methods)
  • Basic math (matrix multiplication from linear algebra)
  • No machine learning background required

If you can multiply two matrices by hand and write a Python class, you’re ready.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

Every neural network you have ever used — image classifiers, language models, self-driving perception stacks — is, at runtime, a sequence of operations on one data structure: the tensor. Get the tensor right and the rest of the framework practically writes itself. Get it wrong and every layer above leaks confusion.

In this module you build that data structure from scratch. By the end, your Tensor supports arithmetic, broadcasting, matrix multiplication, and shape manipulation — the same surface area you would call on torch.Tensor, backed by NumPy instead of CUDA.

Commands

# first time
tito module start 01

# later sessions
tito module resume 01

# when your tests pass
tito module complete 01

Your notebook is modules/01_tensor/tensor.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement the Tensor constructor and the add and matmul operations, then trace how the shape-manipulation and reduction operations build on them
  • Master broadcasting semantics that enable efficient computation without data copying
  • Understand computational complexity (O(n³) for matmul) and memory trade-offs (views vs copies)
  • Connect your implementation to production PyTorch patterns and design decisions

What you’ll build

A Tensor carries its values, its shape, and the gradient state later modules fill in. Every operation runs through the same dispatch path (Figure 1).

Figure 1: Tensor values, metadata, and operations. A Tensor holds an independent NumPy float32 array, shape metadata, and gradient state. Operations dispatch through Function.apply and wrap results in copied storage; Module 06 adds graph recording and backward propagation. Transpose swaps axes, but wrapping the NumPy result copies its values.

What you write

The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:

Tensor.__init__
Create a new tensor from data.
Add.forward
Add two arrays element-wise with broadcasting support.
MatMul.forward
Multiply two matrices.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.core.tensor;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (6). Each prints a ✅ line when it passes.

  • Tensor Creation
  • Arithmetic Operations
  • Validate Matmul Shapes
  • Matrix Multiplication
  • Shape Manipulation
  • Reduction Operations

Integration tests after export (10).

  • tests/01_tensor/test_01_tensor_progressive.py

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

AttributeError: 'Tensor' object has no attribute 'shape'
From every unit test. __init__ stored the array but not its metadata. Set self.shape, self.size, and self.dtype from self.data in the constructor; every later operation reads them.
assert scalar.dtype == np.float32
From test_unit_tensor_creation. __init__ kept NumPy’s default dtype, so integer input stays integer. Convert with np.array(data, dtype=np.float32).
assert np.array_equal(result.data, expected)
From test_unit_matrix_multiplication. MatMul.forward is not computing row-times-column. Entry (i, j) is the dot product of row i of a with column j of b, that is a[i, :] and b[:, j]; an elementwise a * b or b[j, :] fails here.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Tensors: Strided Memory and Multi-Dimensional Arrays in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

You now have the universal carrier of every value the rest of the framework will ever compute. A Tensor holds the data, knows its shape, and supports the arithmetic that linear algebra and eventually backpropagation depend on.

The next module is Module 02: Activations. It answers a single question: given a tensor of pre-activations, how do you apply a non-linearity element-wise — and why is that the one operation that lets a deep network represent anything beyond a glorified linear regression? You will implement ReLU, Sigmoid, Tanh, and Softmax on top of the Tensor you just built, treating each one as a pure function from tensor to tensor. Everything works because __add__, __mul__, broadcasting, and shape preservation are already in place. That is the dividend of getting Module 01 right.

Next: Module 02: Activations

Back to top