October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

Learn LLM Serving as a Memory and Scheduling Problem

LLM serving depends on fitting growing KV caches into accelerator memory while scheduling prompt and decode work at each step. Here’s how paging, chunked prefill, and benchmark context fit together.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

LLM serving is a coordination problem: the system must fit each active request’s growing key-value (KV) cache into accelerator memory, then decide which requests get compute at each model step. Memory limits how much work can run concurrently; scheduling determines how that work affects throughput and response latency.

Why serving needs both memory and scheduling

During autoregressive inference, a model generates output one token at a time. It reuses key and value tensors from earlier tokens rather than recomputing the entire context on every step. The serving system therefore retains a KV cache for each active sequence, and that cache grows as the sequence grows.

As an Amazon Associate I earn from qualifying purchases.

Requests differ: one may have a long prompt and a short answer, another a short prompt and a long answer. Their cache requirements change over time, so the server cannot treat memory as a fixed amount per request. If caches are fragmented or duplicated unnecessarily, usable capacity falls and fewer requests fit in a batch. The PagedAttention paper describes these constraints and its paging-based approach to KV-cache management.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

At each iteration, the scheduler has to make two related decisions: which requests can remain active given available cache and other resources, and which eligible requests should participate in the next forward pass. TensorRT-LLM’s PyTorch scheduler documentation describes these as separate capacity and microbatch scheduling stages.

How KV-cache allocation affects batch capacity

A serving batch can contain multiple requests, but each active sequence consumes cache space. As sequences grow, they consume more of the memory that could otherwise support additional requests. Capacity is therefore not just a question of how much accelerator memory is installed; it also depends on how efficiently the serving system allocates and reuses that memory.

PagedAttention applies an operating-system paging idea to the KV cache. Rather than require each sequence’s cache to occupy one contiguous physical region, it maps cache contents into fixed-size blocks that can be allocated dynamically and shared where appropriate. The paper presents this design as a way to reduce wasted cache memory and support cache sharing. Its authors describe the motivation as “an attention algorithm inspired by the classical virtual memory and paging techniques in operating systems.”

Paging is one design choice, not a guarantee of a particular batch size or latency. Kernel implementation, block management, sharing opportunities, model configuration, and workload all affect results. vAttention explores another approach: it preserves contiguous virtual address space for the KV cache while mapping physical memory on demand through CUDA virtual-memory mechanisms. Its authors describe the goal as mitigating physical-memory fragmentation while retaining virtual contiguity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why prefill and decode need different scheduling

Prompt prefill

Prefill processes the input prompt, often handling many prompt tokens in a forward pass. A long prefill can use substantial compute and take time that would otherwise serve requests already generating output.

Token decode

Decode produces output incrementally, typically one token per iteration for each sequence. Because users are waiting for those tokens, decode work is often latency-sensitive. Combining a large prefill with ongoing decode can make iteration times uneven and delay output.

Chunked prefill

Sarathi-Serve addresses this scheduling tension by dividing prompt prefill into chunks. Its paper describes stall-free schedules intended to let new requests contribute prompt work without pausing ongoing decode work. The relevant trade-off is how much prefill to schedule alongside decode: chunk size and the prompt-to-generation mix influence both utilization and latency.

How the main design approaches differ

Approach Core idea Questions to evaluate
PagedAttention / vLLM Fixed-size KV blocks and block mapping support dynamic allocation and sharing. How much cache capacity and sharing does the workload gain? What are the kernel and block-management costs? How do throughput and latency compare under matched conditions?
Sarathi-Serve Chunked prefills and stall-free schedules balance prompt processing with ongoing decode. What chunk size and prompt/decode mix are used? What are the latency target, hardware, parallelism, and serving capacity?
TensorRT-LLM scheduler Capacity selection and microbatch selection are distinct stages at each step. How does the admission policy handle KV-cache capacity, batch formation, and paused requests for this workload?
vAttention Contiguous virtual address space is paired with on-demand physical-memory allocation. Are the required kernels compatible? What allocation granularity, runtime overhead, portability, and measured throughput apply?

These are different system designs documented in papers or software documentation, not a directly comparable product ranking. A meaningful comparison holds the model, accelerator, input and output lengths, concurrency, latency objective, and implementation version as constant as possible. TensorRT-LLM’s scheduler guide is on its main branch, so its behavior may change; pin a software version when relying on operational details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to read serving-capacity claims

Published capacity gains describe particular experiments, not universal advantages. Sarathi-Serve’s authors reported 2.6× higher serving capacity for Mistral-7B on one A100 and up to 3.7× for Yi-34B on two A100 GPUs, compared with vLLM. They also reported up to a 5.6× end-to-end serving-capacity gain for Falcon-180B using pipeline parallelism. Those figures belong to the paper’s stated setups.

The vAttention authors reported up to 1.23× serving throughput compared with PagedAttention-based FlashAttention and FlashInfer kernels in their evaluation. That result is specific to the paper’s workloads, hardware, and implementation; it is not a general engine-to-engine guarantee.

The vAttention paper also gives per-token KV-memory examples of 64 KB for Yi-6B, 128 KB for Llama-3-8B, and 240 KB for Yi-34B. These figures apply to the models and configurations used in that paper, not every deployment of those model families.

Do not combine these figures into a leaderboard: the papers use different models, hardware, baselines, and methods. For any benchmark, check the exact model, accelerator count, parallelism, sequence lengths, batch or concurrency, latency target, and software implementation before applying the result to a deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What serving configuration can and cannot tell you

The vLLM stable CLI reference documents controls for KV-cache sizing and dtype, optional CPU KV-cache offloading, a scheduler admission watermark, and asynchronous scheduling. These options expose useful trade-offs, but the documentation does not establish one best setting for every workload. Defaults and feature availability can vary by release, so use the documentation for the specific version and hardware you deploy.

Operationally, changing a cache or scheduling control can affect which requests fit, how much work is admitted, or when requests are processed. Measure the result against the workload’s own prompt and output lengths, concurrency, and latency goals rather than assuming that a setting that improves one metric will improve all of them.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.