Module 15: Quantization
Quantization maps floating-point values onto a limited set of integer codes. Packed INT8 storage uses one byte per code instead of four for FP32. TinyTorch simulates these codes in float32 Tensor arrays and performs float32 arithmetic; the storage sizes it reports describe the packed representation, not the NumPy arrays actually allocated.
OPTIMIZATION TIER | Difficulty: ●●●○ | Time: 4-6 hours | Prerequisites: 01-14
Prerequisites: Modules 01-14 means you should have:
- Built the complete foundation (Tensor through Training)
- Implemented profiling tools to measure memory usage
- Understanding of neural network parameters and forward passes
- Familiarity with memory calculations and optimization trade-offs
If you can profile a model’s memory usage and explain the cost of FP32 storage, you’re ready.
Overview
Packed INT8 weights can reduce weight storage relative to FP32, with an accuracy trade-off that must be measured. In this module you build quantize/dequantize functions, calibration, QuantizedLinear, and model conversion. Codes remain in float32 arrays, so the implementation teaches quantization without promising a smaller checkpoint or faster integer execution.
The math you implement is the same math TensorFlow Lite, PyTorch Mobile, and ONNX Runtime use to fit models on phones, IoT boards, and edge hardware without ever touching the cloud.
Commands
# first time
tito module start 15
# later sessions
tito module resume 15
# when your tests pass
tito module complete 15Your notebook is modules/15_quantization/quantization.ipynb.
Learning objectives
- Implement asymmetric INT8 quantization: scale, zero-point, and the quantize/dequantize round trip, distinguishing modeled packed bytes from actual storage.
- Build calibration that fits scale and zero-point to a real activation distribution from sample inputs.
- Reason about quantization error: where it comes from, how it bounds (±scale/2), and why neural networks tolerate it.
- Connect your implementation to TensorFlow Lite, PyTorch Mobile, and ONNX Runtime — same math, different kernels.
- Quantify the memory–accuracy trade-off across model sizes and quantization choices.
What you’ll build
The pattern you’ll enable:
# Simulate quantization and model packed storage
quantize_model(model, calibration_data=sample_inputs)
# Dense arrays remain float32; measure output error and accuracyWhat you’re not building yet
To keep this module focused, you will not implement:
- Per-channel quantization (PyTorch supports this for finer-grained precision)
- Mixed precision strategies (keeping sensitive layers in FP16/FP32)
- Quantization-aware training
- INT8 GEMM kernels (production uses hardware instructions like AVX-512 VNNI)
You are building per-tensor asymmetric INT8 quantization. Packed serialization and integer kernels require additional implementations.
What you write
The notebook arrives with the surrounding code already written and explained. You write 7 functions, each marked # YOUR CODE HERE and followed by a test cell:
quantize_int8- Quantize FP32 tensor to INT8 using asymmetric (min-max) quantization.
dequantize_int8- Dequantize INT8 tensor back to FP32.
QuantizedLinear.calibrate- Calibrate input quantization parameters using sample data.
QuantizedLinear.forward- Forward pass with quantized computation.
QuantizedLinear.memory_usage- Model packed INT8 bytes, including metadata; not actual NumPy storage.
_measure_layer_bytes- Measure parameter count and byte usage for a single layer.
analyze_model_sizes- Compare memory usage between original and quantized models.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.perf.quantization; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (8). Each prints a ✅ line when it passes.
- INT8 Quantization
- INT8 Dequantization
- QuantizedLinear
- Collect Layer Inputs
- Quantize Single Layer
- Model Quantization
- Measure Layer Bytes
- Model Size Analysis
Integration tests after export (20).
tests/15_quantization/test_constant_roundtrip.pytests/15_quantization/test_quantization_composition.pytests/15_quantization/test_quantization_integration.pytests/15_quantization/test_quantizer_core.py
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
Worst-element quantization error 0.166667 exceeds the INT8 bound scale/2 = 0.009804-
From
test_unit_quantize_int8. The range was taken from the data alone. Extend it to include 0.0 before computing the scale, so the zero point fits in the int8 range instead of being clipped. ZeroDivisionError: float division by zero-
From
test_unit_quantize_int8. A constant tensor has no range, so the scale is zero. Handlemax == minseparately before dividing. Worst-element round-trip error 1.741569 exceeds the INT8 bound scale/2 = 0.009216.-
From
test_unit_dequantize_int8. A zero-point sign is flipped. Quantize withx / S + Zand dequantize with(q - Z) * S.
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter INT8 Quantization: Compressing Floats into 8-Bit Integers in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
You have mapped weights onto a quantization grid. Module 16 changes which weights are active. Its pruning functions zero individual weights or whole channels while preserving dense shapes. Sparse storage or structural compaction is needed before zeros translate into smaller allocations or fewer dense operations.
Next: Module 16: Compression
How later modules use this one
| Module | What it adds | The stack so far |
|---|---|---|
| 16: Compression | Pruning removes redundant weights | Compare rounded and masked candidates independently |
| 17: Acceleration | Compare dense execution strategies | Benchmark vectorized and tiled matrix multiplication |
| 20: Capstone | Benchmark an optimized candidate | baseline → prune/quantize → measure → submission.json |