Read the abstract, introduction and design section of "Efficient Memory Management for Large Language Model Serving with PagedAttention" (Kwon et al., SOSP 2023) — or the vLLM blog post if you cannot access the paper.
- Map the paper's concepts to this chapter's: what plays the role of pages, frames, page tables, fragmentation, copy-on-write and swapping?
- Numerical example. A GPU has 24 GB free for the KV cache. Each token of a request needs 0.5 MB of KV cache. Requests have a maximum length of 2,048 tokens, but their actual lengths are 300 tokens on average. Compute how many concurrent requests fit when (a) memory is reserved contiguously for the maximum length, and (b) memory is allocated in blocks of 16 tokens on demand (assume on average half a block is wasted per request). What is the improvement factor?
- The paper uses copy-on-write for requests that share a prompt (e.g., parallel sampling). Explain how, and what is saved.
- Discuss one limitation or open issue, and relate it to edge deployment (smaller GPUs, many users).