Gradient Lander

Balance batch size and learning rate to converge safely.

Press Enter to launch. Use Up to thrust and Left or Right to steer toward the green minimum.
↑ thrust (batch size) · ← → steer learning rate · R retry VRAM 100%, descent speed 0

How to play

The game starts paused so you can read the controls. Press Enter (or Space, or ↑) to launch when you’re ready.

  • ↑ hold to thrust — increases batch size, stabilizes descent, burns VRAM.
  • ← → tilt — steer the learning-rate trajectory across the loss landscape.
  • Watch the dashed line under your ship — that’s your altitude over the ground beneath. Watch the small blue circle ahead of your ship — that’s where you’ll land if you do nothing.
  • Touch down on the green pad (the global minimum) at low speed and near-vertical angle. The green pad pulses so you can spot it instantly.
  • Hit the orange pads → local minimum (suboptimal). Hit anything else → divergence. Run out of VRAM → OOM.
  • Click ↺ TRY AGAIN in the canvas, or press R, to retry. The terrain is the same for everyone playing today.
  • On a phone or tablet: the canvas is divided into three vertical zones — tap-and-hold the left third to steer left, the right third to steer right, the center to thrust (and to launch from the READY screen). The game also respects your OS Reduce Motion preference.

The Systems Concept

Large batches make optimization feel smoother because gradient noise falls — but the memory bill rises. Lander turns that tradeoff into a landing problem: stability only helps if you still have enough VRAM left and the learning-rate trajectory actually reaches the basin you wanted.

Each failure mode in the game maps to a real one in training:

  • Diverged (LR too high) — over-aggressive learning rate overshoots the minimum.
  • Local minimum — optimizer settles in a sub-optimal basin without enough exploration.
  • Missed the basin — landed in a flat or saddle region, no clear gradient signal.
  • Off course — weights drift out of the parameter space the model can handle.
  • OOM — VRAM exhausted; the run dies the way real training processes do.

A “soft landing” on the global minimum is the only outcome where every constraint held simultaneously. That is the dream of large-batch training.


Part of the MLSysBook Playground — try another game. Found a bug? Report an issue.