Module 19: Benchmarking

A benchmark is not a number; it is a claim about throughput, latency, and hardware utilization under a specific workload. This module builds the measurement discipline that separates “my model is 2x faster” from “my model is 2x faster on batch=1, input-length=128, FP16, on an A100, with warmup discarded and 95% confidence intervals reported.” Every performance claim in an ML paper or an MLPerf submission lives or dies in the harness you write here.

NoteModule Info

OPTIMIZATION TIER | Difficulty: ●●●○ | Time: 5-7 hours | Prerequisites: 01-18

This module assumes familiarity with the complete TinyTorch stack (Modules 01-13), profiling (Module 14), and optimization techniques (Modules 15-18). You should understand how to build, profile, and optimize models before tackling systematic benchmarking and statistical comparison of optimizations.

Audio overview (AI-generated)
Open in Binder → Runs in your browser with nothing to install; the session is discarded when you leave, so download your notebook to keep it.
Lecture slides AI-generated · opens an in-page viewer
🔥 Slide Deck · AI-generated
1 / -
Loading slides...

Overview

“My model is 3x faster!” Faster than what? Measured how? At what input size, on what hardware, after how many warmup runs? Most “speedup” claims dissolve under those four questions — and yours will too if you don’t measure with care.

Modules 14-18 gave you optimizations. This module gives you the one tool that tells you which of them actually worked: a benchmarking harness that controls for noise, runs proper warmup, and produces confidence intervals you can defend. It follows the same discipline MLPerf enforces (fixed workloads, warmup, repeated samples), built from pieces you write yourself.

By the end, you will have the evaluation framework that drives the Torch Olympics capstone — and the discipline to never publish a number you cannot justify.

Commands

# first time
tito module start 19

# later sessions
tito module resume 19

# when your tests pass
tito module complete 19

Your notebook is modules/19_benchmarking/benchmarking.ipynb.

Learning objectives

TipBy completing this module, you will:
  • Implement the precise timer and the latency measurement loop at the core of the harness, with its warmup protocol, confidence intervals, and variance control
  • Quantify trade-offs between accuracy, latency, and memory using Pareto frontiers
  • Diagnose measurement noise — coefficient of variation, outliers, cold-start effects — before reporting numbers
  • Connect optimizations from Modules 14-18 into a single comparison workflow for the Torch Olympics capstone

What you’ll build

Figure 1: A repeatable benchmark. Fix inputs, model mode, and environment; warm up, retain repeated latency samples, and report their distribution. Measure quality and storage with clearly labeled assumptions.

The pattern you’ll enable:

# Compare baseline vs optimized model with statistical rigor
benchmark = Benchmark([baseline_model, optimized_model], datasets=[test_dataset])
latency_results = benchmark.run_latency_benchmark()
# e.g. model_0_latency_ms: 12.3000 ± 0.8000 (n=10), model_1_latency_ms: 4.1000 ± 0.3000 (n=10)
# (mean ± std, illustrative numbers; each result also carries a 95% CI in ci_lower/ci_upper)

What you’re not building yet

To keep this module focused, you will not implement:

  • Hardware-specific benchmarks (GPU profiling requires CUDA, covered in production frameworks)
  • Energy measurement (requires specialized hardware like power meters)
  • Distributed benchmarking (multi-node coordination is beyond scope)
  • Automated hyperparameter tuning for optimization

You are building the statistical foundation for fair comparison. The hardware-specific extensions are mechanical once the methodology is right.

What you write

The notebook arrives with the surrounding code already written and explained. You write 3 functions, each marked # YOUR CODE HERE and followed by a test cell:

precise_timer
High-precision timing context manager for benchmarking.
Benchmark.run_latency_benchmark
Benchmark model inference latency using Profiler.
pareto_frontier
Return the names of the non-dominated points, in input order.

How you know it works

tito module complete stops at the first step that fails:

  1. the unit tests inside your notebook run;
  2. your code is exported into tinytorch.perf.benchmarking;
  3. the integration tests run against that exported package, together with the modules before it;
  4. the module is recorded as done, and tito module status shows it.

Unit tests in your notebook (28). Each prints a ✅ line when it passes.

  • BenchmarkResult
  • Precise Timer
  • Benchmark.__init__
  • Benchmark.run_latency_benchmark
  • _simulated_accuracy
  • Benchmark.run_accuracy_benchmark
  • Benchmark.run_memory_benchmark
  • Benchmark (Full Class Integration)
  • BenchmarkSuite.__init__
  • BenchmarkSuite._estimate_energy_efficiency
  • BenchmarkSuite.run_full_benchmark
  • BenchmarkSuite.plot_results
  • pareto_frontier
  • BenchmarkSuite._format_results_summary
  • BenchmarkSuite._format_recommendations
  • BenchmarkSuite (Full Class Integration)
  • MLPerf.__init__
  • MLPerf._run_latency_test
  • _extract_pred_array
  • MLPerf._run_accuracy_test
  • MLPerf.run_standard_benchmark
  • MLPerf._compile_report_data
  • MLPerf._format_compliance_summary
  • MLPerf (Full Class Integration)
  • _collect_base_metrics
  • _calculate_improvements
  • _generate_recommendations
  • analyze_optimization_techniques (Full Integration)

Integration tests after export (23).

  • tests/19_benchmarking/test_benchmark_contracts.py
  • tests/19_benchmarking/test_benchmark_core.py
  • tests/19_benchmarking/test_benchmarking_integration.py

Completing this module unlocks Milestone 06, MLPerf to Generative Serving (2018) (tito milestone run 06).

When it fails

A bare NotImplementedError with no message means a cell reached a function you have not written yet: the notebook ships each one as # YOUR CODE HERE followed by raise NotImplementedError(). The messages below are ones this module actually prints when an implementation is present but wrong.

Expected ~0.01s, got 0.0s
From test_unit_precise_timer. The timer never recorded the elapsed time. Set timer.elapsed = time.perf_counter() - timer.start_time in a finally block after the yield, so it runs when the with block exits.

Finished? Read why

The reasoning behind this module (why it is built this way, what it costs, and how production frameworks differ) is the chapter Benchmarking: Measuring Reliable Speedups in the companion book, TinyTorch: From Tensors to Transformers (PDF). The book prints complete reference implementations, so read it after you finish the module, not while you are working on it.

What’s next

You now have a tool that separates real speedups from noise. Before the Capstone, the next chapter puts the whole Optimization Tier through a historical end-to-end run. The Optimization Milestone, MLPerf to Generative Serving (2018), chains your Profiler (Module 14), Quantization (15), Compression (16), Acceleration (17), and KV-Cache (18) into a single measure → optimize → validate loop, the same discipline MLPerf forced onto an industry that previously measured whatever made its hardware look fastest. You measure what your compression and caching actually bought, on your own framework, with every number traceable to code you wrote.

After that, Module 20 — the Capstone: Torch Olympics — is where the benchmarking harness earns its keep. You will combine the optimizations from Modules 14-18 (quantization, pruning, fusion, caching), benchmark them locally, record the results in a schema-validated submission.json, and defend every number against the same statistical scrutiny you just built.

The capstone has no “easy” mode. Hold every entry to the harness from this chapter: overlapping confidence intervals are not a win, and a mean measured without warmup is not a result. The discipline you practiced here is the price of admission.

NoteUp next: Milestone 06 (MLPerf), then Module 20, Torch Olympics

First: the MLPerf milestone runs your Profiler → Quantization → Compression → Acceleration → KV-Cache pipeline on real models, the same measure-optimize-validate loop production teams use. Then Module 20 combines everything from Modules 01-19 in the Torch Olympics: stack optimizations, benchmark them honestly on your own machine, and save the results for your fastest, smallest, or most accurate candidate.

Next: Milestone 06: MLPerf, then Module 20: Torch Olympics

How later modules use this one

Table 1: How benchmarking powers each capstone competition event.
Competition Event Metric Optimized Your Benchmark In Action
Latency Sprint Minimize inference time benchmark.run_latency_benchmark() measures each candidate’s latency
Memory Challenge Minimize model size benchmark.run_memory_benchmark() tracks footprint
Accuracy Contest Maximize accuracy under constraints benchmark.run_accuracy_benchmark() measures each candidate’s accuracy
All-Around Balanced Pareto frontier pareto_frontier(points, lower_is_better) finds the non-dominated trade-offs
Back to top