Module 15: Quantization

Quantization maps floating-point values onto a limited set of integer codes. Packed INT8 storage uses one byte per code instead of four for FP32. TinyTorch simulates these codes in float32 Tensor arrays and performs float32 arithmetic; the storage sizes it reports describe the packed representation, not the NumPy arrays actually allocated.

NoteModule Info

OPTIMIZATION TIER | Difficulty: ●●●○ | Time: 4-6 hours | Prerequisites: 01-14

Prerequisites: Modules 01-14 means you should have:

  • Built the complete foundation (Tensor through Training)
  • Implemented profiling tools to measure memory usage
  • Understanding of neural network parameters and forward passes
  • Familiarity with memory calculations and optimization trade-offs

If you can profile a model’s memory usage and explain the cost of FP32 storage, you’re ready.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

Packed INT8 weights can reduce weight storage relative to FP32, with an accuracy trade-off that must be measured. In this module you build quantize/dequantize functions, calibration, QuantizedLinear, and model conversion. Codes remain in float32 arrays, so the implementation teaches quantization without promising a smaller checkpoint or faster integer execution.

The math you implement is the same math TensorFlow Lite, PyTorch Mobile, and ONNX Runtime use to fit models on phones, IoT boards, and edge hardware without ever touching the cloud.

Commands

# first time
tito module start 15

# later sessions
tito module resume 15

# when your tests pass
tito module complete 15

Your notebook is modules/15_quantization/quantization.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement asymmetric INT8 quantization: scale, zero-point, and the quantize/dequantize round trip, distinguishing modeled packed bytes from actual storage.
  • Build calibration that fits scale and zero-point to a real activation distribution from sample inputs.
  • Reason about quantization error: where it comes from, how it bounds (±scale/2), and why neural networks tolerate it.
  • Connect your implementation to TensorFlow Lite, PyTorch Mobile, and ONNX Runtime — same math, different kernels.
  • Quantify the memory–accuracy trade-off across model sizes and quantization choices.

What you’ll build

Figure 1: Integer codes and execution storage. Floating-point values map to affine integer codes and reconstruct through scale and zero point. Packed storage models one byte per code, but TinyTorch Tensor stores codes as float32 and QuantizedLinear uses float32 arithmetic.

The pattern you’ll enable:

# Simulate quantization and model packed storage
quantize_model(model, calibration_data=sample_inputs)
# Dense arrays remain float32; measure output error and accuracy

What you’re not building yet

To keep this module focused, you will not implement:

  • Per-channel quantization (PyTorch supports this for finer-grained precision)
  • Mixed precision strategies (keeping sensitive layers in FP16/FP32)
  • Quantization-aware training
  • INT8 GEMM kernels (production uses hardware instructions like AVX-512 VNNI)

You are building per-tensor asymmetric INT8 quantization. Packed serialization and integer kernels require additional implementations.

What you write

The notebook arrives with the surrounding code already written and explained. You write 7 functions, each marked # YOUR CODE HERE and followed by a test cell:

quantize_int8
Quantize FP32 tensor to INT8 using asymmetric (min-max) quantization.
dequantize_int8
Dequantize INT8 tensor back to FP32.
QuantizedLinear.calibrate
Calibrate input quantization parameters using sample data.
QuantizedLinear.forward
Forward pass with quantized computation.
QuantizedLinear.memory_usage
Model packed INT8 bytes, including metadata; not actual NumPy storage.
_measure_layer_bytes
Measure parameter count and byte usage for a single layer.
analyze_model_sizes
Compare memory usage between original and quantized models.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.perf.quantization;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (8). Each prints a ✅ line when it passes.

  • INT8 Quantization
  • INT8 Dequantization
  • QuantizedLinear
  • Collect Layer Inputs
  • Quantize Single Layer
  • Model Quantization
  • Measure Layer Bytes
  • Model Size Analysis

Integration tests after export (20).

  • tests/15_quantization/test_constant_roundtrip.py
  • tests/15_quantization/test_quantization_composition.py
  • tests/15_quantization/test_quantization_integration.py
  • tests/15_quantization/test_quantizer_core.py

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

Worst-element quantization error 0.166667 exceeds the INT8 bound scale/2 = 0.009804
From test_unit_quantize_int8. The range was taken from the data alone. Extend it to include 0.0 before computing the scale, so the zero point fits in the int8 range instead of being clipped.
ZeroDivisionError: float division by zero
From test_unit_quantize_int8. A constant tensor has no range, so the scale is zero. Handle max == min separately before dividing.
Worst-element round-trip error 1.741569 exceeds the INT8 bound scale/2 = 0.009216.
From test_unit_dequantize_int8. A zero-point sign is flipped. Quantize with x / S + Z and dequantize with (q - Z) * S.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter INT8 Quantization: Compressing Floats into 8-Bit Integers in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

You have mapped weights onto a quantization grid. Module 16 changes which weights are active. Its pruning functions zero individual weights or whole channels while preserving dense shapes. Sparse storage or structural compaction is needed before zeros translate into smaller allocations or fewer dense operations.

Next: Module 16: Compression

How later modules use this one

Table 1: How quantization stacks with compression, acceleration, and capstone modules.
Module What it adds The stack so far
16: Compression Pruning removes redundant weights Compare rounded and masked candidates independently
17: Acceleration Compare dense execution strategies Benchmark vectorized and tiled matrix multiplication
20: Capstone Benchmark an optimized candidate baseline → prune/quantize → measure → submission.json
Back to top