solvers.ContinuousBatchingModel
solvers.ContinuousBatchingModel()Compares static KV reservation with PagedAttention under continuous batching.
Decode is memory-bound, so the number of requests sharing a step is the throughput lever, and the KV-cache allocator decides how many fit. A static allocator reserves a contiguous slot of max_seq_len tokens for every request; PagedAttention hands out fixed-size blocks on demand. This model sizes both from the same KV budget and a request-length distribution, then compares memory-bound decode throughput at each allocator’s concurrency.
Kwon et al. (2023) profile contiguous pre-allocation (their Orca baselines) and find that only 20.4% to 38.2% of KV-cache memory holds actual token states (Fig. 2). The rest is reserved slots for future tokens, internal fragmentation from over-provisioning for the maximum sequence length, and external fragmentation from the memory allocator (Sec. 3.1, Fig. 3). vLLM limits each request’s waste to one block (Sec. 4.2) and reaches 96.3% token-state usage (Fig. 2).
Scope and assumptions:
- The static baseline is max-length reservation, like Kwon’s “Orca (Max)”. It is modeled with no external fragmentation because every slot has the same size, although Kwon et al. still attribute 8.9% of Orca (Max) KV memory to external fragmentation and other overhead. Exact-size and power-of-two contiguous allocators, which trade reservation waste for external fragmentation, are not modeled.
- Both allocators are sized with every admitted request at its full length (peak occupancy). A request mid-generation holds fewer paged blocks, so paged capacity is conservative, while static reserves
max_seq_lenfor the whole request lifetime either way. - Request lengths are exponential, cut off at
max_seq_len, with meanmean_request_tokens(calc_capped_exponential_scale).
Literature Source: 1. Kwon et al. (2023), “Efficient Memory Management for Large Language Model Serving with PagedAttention,” SOSP ’23. 2. Yu et al. (2022), “ORCA: A Distributed Serving System for Transformer-Based Generative Models.”
Methods
| Name | Description |
|---|---|
| solve | Compare static max-length KV reservation with paged allocation. |
solve
solvers.ContinuousBatchingModel.solve(
model,
hardware,
max_seq_len,
mean_request_tokens=None,
max_batch_size=1,
page_size=16,
precision='fp16',
efficiency=0.5,
)Compare static max-length KV reservation with paged allocation.