LLM serving is a coordination problem: the system must fit each active request’s growing key-value (KV) cache into accelerator memory, then decide which requests get compute at each model step. Memory limits how much work can run concurrently; scheduling determines how that work affects throughput and response latency.
Why serving needs both memory and scheduling
During autoregressive inference, a model generates output one token at a time. It reuses key and value tensors from earlier tokens rather than recomputing the entire context on every step. The serving system therefore retains a KV cache for each active sequence, and that cache grows as the sequence grows.
As an Amazon Associate I earn from qualifying purchases.
Requests differ: one may have a long prompt and a short answer, another a short prompt and a long answer. Their cache requirements change over time, so the server cannot treat memory as a fixed amount per request. If caches are fragmented or duplicated unnecessarily, usable capacity falls and fewer requests fit in a batch. The PagedAttention paper describes these constraints and its paging-based approach to KV-cache management.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallAt each iteration, the scheduler has to make two related decisions: which requests can remain active given available cache and other resources, and which eligible requests should participate in the next forward pass. TensorRT-LLM’s PyTorch scheduler documentation describes these as separate capacity and microbatch scheduling stages.
#1 Best Overall
How KV-cache allocation affects batch capacity
A serving batch can contain multiple requests, but each active sequence consumes cache space. As sequences grow, they consume more of the memory that could otherwise support additional requests. Capacity is therefore not just a question of how much accelerator memory is installed; it also depends on how efficiently the serving system allocates and reuses that memory.
PagedAttention applies an operating-system paging idea to the KV cache. Rather than require each sequence’s cache to occupy one contiguous physical region, it maps cache contents into fixed-size blocks that can be allocated dynamically and shared where appropriate. The paper presents this design as a way to reduce wasted cache memory and support cache sharing. Its authors describe the motivation as “an attention algorithm inspired by the classical virtual memory and paging techniques in operating systems.”
Paging is one design choice, not a guarantee of a particular batch size or latency. Kernel implementation, block management, sharing opportunities, model configuration, and workload all affect results. vAttention explores another approach: it preserves contiguous virtual address space for the KV cache while mapping physical memory on demand through CUDA virtual-memory mechanisms. Its authors describe the goal as mitigating physical-memory fragmentation while retaining virtual contiguity.
Why prefill and decode need different scheduling
Prompt prefill
Prefill processes the input prompt, often handling many prompt tokens in a forward pass. A long prefill can use substantial compute and take time that would otherwise serve requests already generating output.
Rank #3
Token decode
Decode produces output incrementally, typically one token per iteration for each sequence. Because users are waiting for those tokens, decode work is often latency-sensitive. Combining a large prefill with ongoing decode can make iteration times uneven and delay output.
Chunked prefill
Sarathi-Serve addresses this scheduling tension by dividing prompt prefill into chunks. Its paper describes stall-free schedules intended to let new requests contribute prompt work without pausing ongoing decode work. The relevant trade-off is how much prefill to schedule alongside decode: chunk size and the prompt-to-generation mix influence both utilization and latency.
How the main design approaches differ
| Approach | Core idea | Questions to evaluate |
|---|---|---|
| PagedAttention / vLLM | Fixed-size KV blocks and block mapping support dynamic allocation and sharing. | How much cache capacity and sharing does the workload gain? What are the kernel and block-management costs? How do throughput and latency compare under matched conditions? |
| Sarathi-Serve | Chunked prefills and stall-free schedules balance prompt processing with ongoing decode. | What chunk size and prompt/decode mix are used? What are the latency target, hardware, parallelism, and serving capacity? |
| TensorRT-LLM scheduler | Capacity selection and microbatch selection are distinct stages at each step. | How does the admission policy handle KV-cache capacity, batch formation, and paused requests for this workload? |
| vAttention | Contiguous virtual address space is paired with on-demand physical-memory allocation. | Are the required kernels compatible? What allocation granularity, runtime overhead, portability, and measured throughput apply? |
These are different system designs documented in papers or software documentation, not a directly comparable product ranking. A meaningful comparison holds the model, accelerator, input and output lengths, concurrency, latency objective, and implementation version as constant as possible. TensorRT-LLM’s scheduler guide is on its main branch, so its behavior may change; pin a software version when relying on operational details.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsHow to read serving-capacity claims
Published capacity gains describe particular experiments, not universal advantages. Sarathi-Serve’s authors reported 2.6× higher serving capacity for Mistral-7B on one A100 and up to 3.7× for Yi-34B on two A100 GPUs, compared with vLLM. They also reported up to a 5.6× end-to-end serving-capacity gain for Falcon-180B using pipeline parallelism. Those figures belong to the paper’s stated setups.
The vAttention authors reported up to 1.23× serving throughput compared with PagedAttention-based FlashAttention and FlashInfer kernels in their evaluation. That result is specific to the paper’s workloads, hardware, and implementation; it is not a general engine-to-engine guarantee.
The vAttention paper also gives per-token KV-memory examples of 64 KB for Yi-6B, 128 KB for Llama-3-8B, and 240 KB for Yi-34B. These figures apply to the models and configurations used in that paper, not every deployment of those model families.
Do not combine these figures into a leaderboard: the papers use different models, hardware, baselines, and methods. For any benchmark, check the exact model, accelerator count, parallelism, sequence lengths, batch or concurrency, latency target, and software implementation before applying the result to a deployment.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →What serving configuration can and cannot tell you
The vLLM stable CLI reference documents controls for KV-cache sizing and dtype, optional CPU KV-cache offloading, a scheduler admission watermark, and asynchronous scheduling. These options expose useful trade-offs, but the documentation does not establish one best setting for every workload. Defaults and feature availability can vary by release, so use the documentation for the specific version and hardware you deploy.
Operationally, changing a cache or scheduling control can affect which requests fit, how much work is admitted, or when requests are processed. Measure the result against the workload’s own prompt and output lengths, concurrency, and latency goals rather than assuming that a setting that improves one metric will improve all of them.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




