KV Cache Packer
Fit the prompts before memory fragments.
How to play
- In Phase 1, click a row with enough contiguous space for the incoming request bar.
- Watch requests complete at different times. The holes they leave behind create fragmentation.
- In Phase 2, paged mode splits requests into small blocks. Press Space to defrag and recover capacity.
The Systems Concept
In LLM inference, serving requests requires storing key-value tensors for every active sequence. Contiguous allocation wastes capacity when prompts and generations finish at different times. PagedAttention-style block allocation turns the KV cache into a paging problem, reducing external fragmentation and allowing larger effective batches.
Part of MLSysBook Playground. Found a bug? Report an issue.