Interactive Labs
What happens to accuracy?
Lab 01 is one instance of a pattern that repeats 34 times. The rest of this page is about the pattern.
Each lab is a structured confrontation with a quantitative reality that surprises. The pedagogical design rests on a simple observation: a student who predicts wrong and then discovers why has learned more than a student who reads a correct answer. The prediction lock is what makes that possible. You cannot passively watch the simulator; you have to commit first.
Every part within every lab follows the same rhythm:
Stakeholder Scenario: A fictional but realistic message from a CTO, VP of Engineering, or ML lead frames a real-world problem. These are not toy examples. They are the decisions engineers make every day.
Prediction Lock: Before seeing any data, you must commit a structured prediction (multiple choice or numeric estimate). The simulator is locked until you predict. This forces you to surface your assumptions.
Interactive Instruments: Sliders, toggles, and charts powered by the mlsysim physics engine let you explore the design space. Every number traces to a specific textbook claim, with no magic constants.
Prediction Reveal: The lab shows you what you predicted versus what actually happened, with specific numbers: “You predicted 2\(\times\). Actual: 50\(\times\). You were off by 25\(\times\).” This gap is the learning moment.
Math Peek: A collapsible accordion reveals the governing equation. You can always see the physics behind the simulator.
Briefing ~2 min Learning objectives, prerequisites, core question
Part A ~12 min Calibration --- correct a wrong prior with data
Part B ~12 min Deepening --- quantify the mechanism behind Part A
Part C ~12 min Cross-context --- same system, different hardware
Part D ~12 min Design challenge --- make a decision with trade-offs
Synthesis ~5 min Key takeaways, connections, self-assessment
At least one part includes a failure state: push a slider too far and the system crashes (OOM, SLA violation, thermal throttle). These failures are reversible and instructive: the point is to find the boundary, not to punish.
Your predictions and design decisions persist across labs in the Design Ledger, a browser-based save system. Lab 08’s training memory budget builds on Lab 05’s activation analysis, which builds on Lab 01’s magnitude calibration. The capstone labs (Lab 16 in each volume) synthesize your full Design Ledger into a portfolio.
How do these labs work? A 5-minute walkthrough of the predict-discover-explain ritual every lab follows.
If a model fails for three different physical reasons on three hardware targets, how do you diagnose which axis to fix?
If you double compute power, why doesn't latency halve?
Why does discovering a deployment constraint late cost 16× more than finding it early?
When is moving compute to data cheaper than moving data to compute?
ReLU and Sigmoid produce similar accuracy, so why does activation memory dictate cache fit?
Why does self-attention scale quadratically O(N²) while convolutions scale linearly?
Why does compiled execution run 17× faster than eager mode without changing a single weight?
Why does a 7B parameter model need 112 GB of memory before storing a single activation?
When does curating data produce more accuracy per dollar than adding more raw data?
Can you compress a model 4× without losing accuracy? Where is the cliff?
Is your workload compute-bound or memory-bound, and why does the answer change everything?
Amdahl's Law says 5% sequential code limits speedup to 20× regardless of parallelism. Is that right?
Your server looks healthy at 50% utilization, so why is it on fire at 80%?
Your model shipped Monday. By Friday it lost 3 accuracy points while your dashboard is green. Why?
If you add 10× more GPUs, do you get 10× more throughput?
Why does Model FLOPs Utilization rarely exceed 50% on real datacenter hardware?
When does topology, latency, bandwidth, or bisection become the real system limit?
Can your distributed storage hierarchy feed your GPUs fast enough, or are they starving?
Data, tensor, or pipeline parallelism, and which fits your model and your cluster?
Which collective algorithm fits your topology, message size, and residual risk?
At 10,000 GPUs, what is the probability of zero failures in 24 hours?
FIFO scheduling wastes 40% of your cluster. Can gang scheduling and DRF do better?
How does FlashAttention SRAM tiling eliminate memory bandwidth bottlenecks?
How does PagedAttention virtual memory eliminate KV-cache fragmentation at scale?
When does moving inference to the edge save energy vs. trigger thermal throttling?
How do grey failures and stragglers silently amplify tail latency across a fleet?
Differential privacy adds noise. How much accuracy do you lose for how much privacy?
Adversarial training costs 8× compute. When is it worth it?
Moving your training from Iowa to Quebec cuts carbon 10×. Why?
Where does software authority end and physical delegation begin?
Can a learned policy on Qualcomm Linux close a physical loop with MCU permission?
Can the microcontroller governor reject plausible but stale or dangerous intent?
What evidence supports a deployment release claim under real-world disturbance?
Already running in your browser — nothing to install. Power users who want offline access or want to hack the simulations can optionally grab the package:
python3 -m pip install -r labs/requirements.txt
python3 -m pip install -e mlsysim
cd labs
marimo run vol1/lab_01_ml_intro.py
Comprehensive theory across the full ML systems stack.
Beamer decks and teaching materials for every chapter.
Interactive Marimo notebooks that measure the book's claims.
Hands-on embedded ML deployment on real devices.
Course maps, syllabi, and adoption resources for teaching.
Build your own ML framework from scratch, module by module.
The analytical modeling engine behind the book's quantitative figures.