Module 14: Profiling
You cannot optimize what you have not measured. Profiling is the systems skill that turns “I think this is slow” into “this layer spends 62% of its time waiting on HBM reads at 5% of peak GFLOP/s.” Before you reach for quantization, kernel fusion, or KV caching, you need numbers: parameter count, FLOPs per forward pass, peak activation memory, median latency, and an arithmetic-intensity reading on the roofline. This module builds those instruments end-to-end so every subsequent optimization module has ground truth to point at.
OPTIMIZATION TIER | Difficulty: ●●○○ | Time: 3-5 hours | Prerequisites: 01-13
Prerequisites: Modules 01-13 means you should have:
- Built the complete ML stack (Modules 01-08)
- Implemented CNN architectures (Module 09) or Transformers (Modules 10-13)
- Models to profile and optimize
Why these prerequisites: You’ll profile models built in Modules 01-13. Understanding the implementations helps you interpret profiling results — for example, why attention is memory-bound.
Overview
You have built a working ML framework. Now you have to make it fast. The Optimization Tier starts here, and it starts with a rule that almost every engineer breaks at least once: measure before you optimize. Guess at the bottleneck and you will spend a week speeding up code that was never on the critical path.
This module gives you the instruments. You’ll build a profiler that counts parameters, estimates FLOPs, tracks memory, and measures latency with enough statistical rigor that the numbers actually mean something. By the end you can answer the questions every optimization decision rests on: Is this model compute-bound or memory-bound? Which layer dominates? Where will quantization or caching pay off — and where will it waste your time?
Every later module in this tier — quantization, compression, acceleration, KV-caching — depends on the data this profiler produces. Build the instrument first. Then optimize.
Commands
# first time
tito module start 14
# later sessions
tito module resume 14
# when your tests pass
tito module complete 14Your notebook is modules/14_profiling/profiling.ipynb.
Learning objectives
- Implement the linear-layer FLOP count, the arithmetic-intensity calculation, and the bottleneck classifier that the supplied Profiler calls when it reports parameters, FLOPs, memory, and latency
- Analyze performance characteristics to identify compute-bound vs memory-bound workloads
- Master statistical measurement techniques with warmup runs and outlier handling
- Connect profiling insights to optimization opportunities in quantization, compression, and caching
What you’ll build
The pattern you’ll enable:
# Comprehensive model analysis for optimization decisions
profiler = Profiler()
profile = profiler.profile_forward_pass(model, input_data)
print(f"Bottleneck: {profile['bottleneck']}") # "memory" or "compute"What you’re not building yet
To keep this module focused, you will not implement:
- GPU profiling (we measure CPU performance with NumPy)
- Distributed profiling (that’s for multi-GPU setups)
- CUDA kernel profilers (PyTorch uses
torch.profilerfor GPU analysis) - Layer-by-layer visualization dashboards (TensorBoard provides this)
You are building the measurement foundation. Visualization and GPU profiling come with production frameworks.
What you write
The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:
_count_linear_flops- Count FLOPs for a Linear layer forward pass.
arithmetic_intensity- Compute arithmetic intensity and place it against a machine’s ridge point.
_analyze_bottleneck- Illustrate a heuristic memory/compute classification.
How you know it works
tito module complete stops at the first step that fails:
- the unit tests inside your notebook run;
- your code is exported into
tinytorch.perf.profiling; - the integration tests run against that exported package, together with the modules before it;
- the module is recorded as done, and
tito module statusshows it.
Unit tests in your notebook (17). Each prints a ✅ line when it passes.
- _count_layer_parameters
- _count_conv_flops
- _count_linear_flops
- arithmetic_intensity
- _analyze_bottleneck
- _calculate_memory_efficiency
- _compute_derived_metrics
- _estimate_backward_costs
- _estimate_optimizer_memory
- Helper Functions
- Parameter Counting
- _count_sequential_flops
- FLOP Counting
- _calculate_parameter_memory
- Memory Measurement
- Latency Measurement
- Advanced Profiling Functions
Integration tests after export (18).
tests/14_profiling/test_14_profiling_progressive.py
When it fails
A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.
Expected 16384, got 8192-
From
test_unit_count_linear_flops. Each multiply-accumulate counts as two floating-point operations. Multiply by 2. High bandwidth should be memory-bound-
From
test_unit_analyze_bottleneck. The comparison is inverted. The heuristic labels a workload memory-bound whenmemory_bandwidth_mbs > gflops_per_second * 100.
Finished? Read why
The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Profiling: The Roofline Model in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.
What’s next
You now have the instrument. Module 15 picks up the first real optimization it enables: quantization.
You’ll map FP32 weights onto INT8 codes (a modeled 4\(\times\) reduction in packed weight storage; TinyTorch keeps the codes in float32) and use this profiler to answer the question that decides whether quantization is worth applying: which layers tolerate reduced precision, and which ones break? Profile first, quantize second, profile again to verify. You’re about to see why this loop is the foundation of every production deployment.
Next: Module 15: Quantization
How later modules use this one
| Module | What It Does | Your Profiler In Action |
|---|---|---|
| 15: Quantization | Reduce precision to INT8 | profile_layer() identifies quantization candidates |
| 16: Compression | Prune and compress weights | count_parameters() measures the compression ratio |
| 17: Acceleration | Vectorize computations | measure_latency() validates the speedup |
| 19: Benchmarking | Compare across systems | profile_forward_pass() produces the comparable numbers |