Cluster Commander
Schedule jobs without fragmenting the fleet.
How to play
You have an 8x8 GPU cluster. Jobs appear in the queue with varying shapes (1x1 inference, 2x2 fine-tuning, 4x4 pre-training). Click a valid empty space in the grid to schedule the next job. If the grid becomes too fragmented to fit a large job, it will block the entire queue!
The Systems Concept
Scheduling diverse jobs on a shared GPU cluster (like Slurm or Kubernetes) often leads to fleet fragmentation. When many small jobs scatter across the cluster, a large contiguous job (e.g., a massive pre-training run requiring 16 interconnected GPUs) might be unable to start, causing cluster utilization to plummet. Modern schedulers use defragmentation and backfilling to mitigate this.
Part of MLSysBook Playground. Found a bug? Report an issue.