KV Cache Packer

Fit the prompts before memory fragments.

Click a row to place each request. Watch gaps appear, then use paged mode to recover capacity.
score 0 Click to place · Phase 2: Spacebar to defrag · R retry

How to play

  1. In Phase 1, click a row with enough contiguous space for the incoming request bar.
  2. Watch requests complete at different times. The holes they leave behind create fragmentation.
  3. In Phase 2, paged mode splits requests into small blocks. Press Space to defrag and recover capacity.

The Systems Concept

In LLM inference, serving requests requires storing key-value tensors for every active sequence. Contiguous allocation wastes capacity when prompts and generations finish at different times. PagedAttention-style block allocation turns the KV cache into a paging problem, reducing external fragmentation and allowing larger effective batches.


Part of MLSysBook Playground. Found a bug? Report an issue.