Module 14: Profiling

You cannot optimize what you have not measured. Profiling is the systems skill that turns “I think this is slow” into “this layer spends 62% of its time waiting on HBM reads at 5% of peak GFLOP/s.” Before you reach for quantization, kernel fusion, or KV caching, you need numbers: parameter count, FLOPs per forward pass, peak activation memory, median latency, and an arithmetic-intensity reading on the roofline. This module builds those instruments end-to-end so every subsequent optimization module has ground truth to point at.

NoteModule Info

OPTIMIZATION TIER | Difficulty: ●●○○ | Time: 3-5 hours | Prerequisites: 01-13

Prerequisites: Modules 01-13 means you should have:

  • Built the complete ML stack (Modules 01-08)
  • Implemented CNN architectures (Module 09) or Transformers (Modules 10-13)
  • Models to profile and optimize

Why these prerequisites: You’ll profile models built in Modules 01-13. Understanding the implementations helps you interpret profiling results — for example, why attention is memory-bound.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

You have built a working ML framework. Now you have to make it fast. The Optimization Tier starts here, and it starts with a rule that almost every engineer breaks at least once: measure before you optimize. Guess at the bottleneck and you will spend a week speeding up code that was never on the critical path.

This module gives you the instruments. You’ll build a profiler that counts parameters, estimates FLOPs, tracks memory, and measures latency with enough statistical rigor that the numbers actually mean something. By the end you can answer the questions every optimization decision rests on: Is this model compute-bound or memory-bound? Which layer dominates? Where will quantization or caching pay off — and where will it waste your time?

Every later module in this tier — quantization, compression, acceleration, KV-caching — depends on the data this profiler produces. Build the instrument first. Then optimize.

Commands

# first time
tito module start 14

# later sessions
tito module resume 14

# when your tests pass
tito module complete 14

Your notebook is modules/14_profiling/profiling.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement the linear-layer FLOP count, the arithmetic-intensity calculation, and the bottleneck classifier that the supplied Profiler calls when it reports parameters, FLOPs, memory, and latency
  • Analyze performance characteristics to identify compute-bound vs memory-bound workloads
  • Master statistical measurement techniques with warmup runs and outlier handling
  • Connect profiling insights to optimization opportunities in quantization, compression, and caching

What you’ll build

Figure 1: An illustrative roofline bound. A peak of 1000 GFLOP/s and bandwidth of 200 GB/s give a ridge at 5 FLOP per byte. At intensity 2, attainable performance is bounded by 400 GFLOP/s; measurements may be lower.

The pattern you’ll enable:

# Comprehensive model analysis for optimization decisions
profiler = Profiler()
profile = profiler.profile_forward_pass(model, input_data)
print(f"Bottleneck: {profile['bottleneck']}")  # "memory" or "compute"

What you’re not building yet

To keep this module focused, you will not implement:

  • GPU profiling (we measure CPU performance with NumPy)
  • Distributed profiling (that’s for multi-GPU setups)
  • CUDA kernel profilers (PyTorch uses torch.profiler for GPU analysis)
  • Layer-by-layer visualization dashboards (TensorBoard provides this)

You are building the measurement foundation. Visualization and GPU profiling come with production frameworks.

What you write

The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:

_count_linear_flops
Count FLOPs for a Linear layer forward pass.
arithmetic_intensity
Compute arithmetic intensity and place it against a machine’s ridge point.
_analyze_bottleneck
Illustrate a heuristic memory/compute classification.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.perf.profiling;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (17). Each prints a ✅ line when it passes.

  • _count_layer_parameters
  • _count_conv_flops
  • _count_linear_flops
  • arithmetic_intensity
  • _analyze_bottleneck
  • _calculate_memory_efficiency
  • _compute_derived_metrics
  • _estimate_backward_costs
  • _estimate_optimizer_memory
  • Helper Functions
  • Parameter Counting
  • _count_sequential_flops
  • FLOP Counting
  • _calculate_parameter_memory
  • Memory Measurement
  • Latency Measurement
  • Advanced Profiling Functions

Integration tests after export (18).

  • tests/14_profiling/test_14_profiling_progressive.py

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

Expected 16384, got 8192
From test_unit_count_linear_flops. Each multiply-accumulate counts as two floating-point operations. Multiply by 2.
High bandwidth should be memory-bound
From test_unit_analyze_bottleneck. The comparison is inverted. The heuristic labels a workload memory-bound when memory_bandwidth_mbs > gflops_per_second * 100.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Profiling: The Roofline Model in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

You now have the instrument. Module 15 picks up the first real optimization it enables: quantization.

NoteUp next: Module 15, Quantization

You’ll map FP32 weights onto INT8 codes (a modeled 4\(\times\) reduction in packed weight storage; TinyTorch keeps the codes in float32) and use this profiler to answer the question that decides whether quantization is worth applying: which layers tolerate reduced precision, and which ones break? Profile first, quantize second, profile again to verify. You’re about to see why this loop is the foundation of every production deployment.

Next: Module 15: Quantization

How later modules use this one

Table 1: How the profiler feeds into optimization-tier modules.
Module What It Does Your Profiler In Action
15: Quantization Reduce precision to INT8 profile_layer() identifies quantization candidates
16: Compression Prune and compress weights count_parameters() measures the compression ratio
17: Acceleration Vectorize computations measure_latency() validates the speedup
19: Benchmarking Compare across systems profile_forward_pass() produces the comparable numbers
Back to top