MLSysBook Playground
Fast, playable systems challenges across all four volumes.
Pick a volume, play a round, then read the systems takeaway below the game. Try another strategy and see what changes.
Try a first run
Volume I · MemoryTensor Tetris
Pack a training step before you run out of memory.Play →
Volume II · ReliabilityCheckpoint Roulette
Decide when saving progress is worth the pause.Play →
Volume III · AgentsContext Cache
Keep the next action's evidence in a tiny working set.Play →
Volume IV · Physical AISafety Gate
Guide a rover while keeping unsafe moves out.Play →
Browse by volume
Volume I: Foundations
Six quick challenges about data, memory, model size, and hardware limits.
Tensor Tetris
Pack training memory before you OOM.
Parameters anchor at the bottom; activations, gradients, and optimizer states fall in their real proportions. Activations dominate — and that’s why you’ll OOM.
Pulse Prune
Shrink a network without breaking it.
A neural network on a canvas. You’ve got 45 seconds to remove most of the weights without crashing accuracy. The faint ones are safe; the bright ones are doing real work.
Quantization Sharp Shot
Compress a model before the target blurs.
Shoot a downrange target with per-layer precision dials as your sight. Lower precision saves bits but blurs, jitters, or drifts the target depending on which layer you compressed.
Batch Size Balancer
Push throughput to the edge of OOM.
A tug-of-war against time. Increase batch size to process more images per second, but watch the memory gauge—if it spikes into the red, you OOM and lose the run.
Data Loader Dash
The CPU preparing data before the GPU starves.
Data blocks slide toward a glowing GPU. Type or tap the matching letter in the processing zone to prepare each block. Miss too many and the GPU starves.
Roofline Rider
Ride from the memory limit to the compute limit.
Ride a glowing kernel from the sloped memory limit to the flat compute limit. Steer toward the cyan roof, dodge stalls, and reach the finish line.
Volume II: Scaling
Eight games about training and serving when one machine is not enough.
Gradient Lander
Balance batch size and learning rate to converge safely.
Pilot your model down the loss landscape to the global minimum. “Thrust” increases your Batch Size—giving you stable descent, but rapidly burning through your limited VRAM budget. Tilt to steer your Learning Rate. Hit the minimum too fast, or run out of VRAM, and your model diverges.
Pipeline Pacer
Keep the GPUs fed without bubbling.
Schedule micro-batches across four pipeline stages. If you send too many forwards, memory fills up; if you pause too long, later stages sit idle (pipeline bubble). Balance forward and backward scheduling to maximize throughput.
MoE Router
Route tokens to the right experts before they expire.
Colored tokens descend toward four experts. Press 1–4 or tap an expert to route the oldest token to its matching destination. A wrong route or a missed token ends the run.
Checkpoint Roulette
Fault tolerance and checkpointing at scale.
Hold Spacebar to train, then release to save your progress. Each checkpoint pauses training while it writes; a node failure during the write restores the previous save. Reach 100% before the one-minute timer runs out.
All-Reduce Rhythm
Keep the gradients flowing in a perfect ring.
GPUs are arranged in a ring. Follow the displayed GPU order and tap each one on the beat to send a gradient chunk. Miss a slot and your combo resets; the round ends after 30 seconds.
Topology Tycoon
Build the fabric, avoid the bottlenecks.
You have 8 GPUs and 20 seconds to build their network topology. Click between two nodes to toggle their connection: None, InfiniBand, or NVLink. Once the run phase begins, packets will flood the network. Maximize your delivered GB/s!
KV Cache Packer
Fit the prompts before memory fragments.
A Tetris-like grid of requests. Place contiguous blocks to serve them. As they finish at different times, your memory fragments. Unlock Paged Mode to shatter requests into 1x1 blocks and defrag your cache!
Cluster Commander
Schedule jobs without fragmenting the fleet.
Schedule 1x1, 2x2, and 4x4 jobs on an 8x8 GPU cluster. Don’t let scattered small jobs fragment the cluster, or a massive 4x4 pre-training run will block the entire queue and kill your utilization!
Volume III: Agentic Systems
Keep a digital agent’s state useful and its tool actions recoverable.
Context Cache
Keep the next action’s evidence inside a limited working set.
Pick the facts a tool call will need while a five-slot context budget forces you to evict the rest. Four short missions reveal whether the next step has enough evidence.
Tool Trail
What should an agent do when a tool call fails?
An agent’s tool call failed. Retry it, undo a prior effect, or stop and inspect. A careless retry can duplicate a real-world change.
Volume IV: Physical AI
Dispatch physical actions before their deadlines and keep unsafe proposals from crossing the boundary.
Latency Line
A correct action can still arrive too late.
Choose speed and quality at each stage of a sense–plan–act cycle. Watch the latency strip and meet the deadline before you dispatch.
Safety Gate
Keep the rover inside the safe set.
A learned policy proposes moves; you decide what crosses into actuation. Allow safe steps, project unsafe ones, or brake before the tick budget runs out.
🚧 Early development · iterating fast
MLSysBook Playground is part of mlsysbook.ai. The games share design vocabulary with the textbook — compute blue, data green, error red, MIT accent — so a figure in the book and a game in your browser feel like the same world.













