Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Red Hat announced the open-source llm-d project on May 20, 2025, to help coordinate large language model inference across Kubernetes clusters. It is not a model server that replaces vLLM: it adds inference-aware routing and scheduling above serving engines, with features such as KV-cache-aware request placement and prefill/decode disaggregation. The project is now a CNCF Sandbox project, while some Red Hat commercial deployment paths remain Technology Preview. Red Hat’s launch announcement and the project repository describe its evolution.
What Red Hat launched
At Red Hat Summit on May 20, 2025, Red Hat introduced llm-d as an open-source community and project for distributed generative AI inference. The goal is to make Kubernetes better suited to operating fleets of model servers, especially when performance depends on more than simply adding replicas and distributing requests.
Red Hat named CoreWeave, Google Cloud, IBM Research and NVIDIA among the project’s founding contributors. Its launch announcement also listed AMD, Cisco, Hugging Face, Intel, Lambda, Mistral AI, UC Berkeley’s Sky Computing Lab and the University of Chicago’s LMCache Lab as contributors or partners. That roster signals a broad set of interests in the project; it is not evidence that every organization has adopted llm-d in production.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsSince launch, llm-d has joined the Cloud Native Computing Foundation as a Sandbox project. The CNCF designation places it in a community-governed open-source ecosystem, but Sandbox status is not a certification of production readiness or a guarantee of support.
#1 Best Overall
What llm-d is—and what it is not
llm-d is a Kubernetes-native distributed inference stack. It coordinates model-serving instances and can make routing and scheduling decisions with inference-specific information. vLLM remains a model-serving engine; llm-d works above engines such as vLLM, and current project materials also discuss SGLang and other integrations. vLLM’s llm-d integration documentation makes that relationship explicit.
A simplified view of the layers looks like this, though real deployments vary:
Application requests
↓
Inference Gateway / Gateway API extensions
↓
llm-d routing and scheduling
↓
KServe integration (for documented deployment paths)
↓
vLLM or another supported model-serving engine
↓
Accelerators and Kubernetes infrastructure
Kubernetes supplies the underlying orchestration and resource management. KServe can provide serving abstractions such as the LLMInferenceService resource. The gateway and llm-d components contribute inference-aware request placement. None of these layers is a foundation model, chatbot or general-purpose Kubernetes distribution.
Recommended Free Tools
Why ordinary load balancing can be insufficient
LLM requests have state and stages that a conventional stateless web request does not. A model server may have useful attention state—often called the KV cache—from a previous request or repeated prompt prefix. If a generic round-robin balancer sends the next related request to a different worker, that worker may need to recompute the same prompt tokens. The result can be wasted accelerator work, more cache fragmentation and longer time to first token.
Inference also has two workloads with different resource profiles. Prefill processes the input prompt; decode generates output tokens, one step at a time. A system optimized for prompt processing may not be the best arrangement for sustained token generation. Consequently, performance depends on which worker receives a request, what it already has cached, the prompt and output sizes, and the available compute and network resources—not only on the number of replicas.
How llm-d tries to improve distributed inference
KV-cache-aware routing
Rather than treating workers as interchangeable, inference-aware routing can consider where useful KV-cache state resides and how busy workers are. Reusing cached state can avoid repeated computation and improve latency or throughput when prompts share prefixes or conversations revisit context.
It is not a universal win. The benefit depends on prompt repetition, cache capacity, memory pressure, traffic patterns and the cost of moving or rebuilding state. Short, unrelated requests may offer little cache reuse, and aggressive locality can conflict with load balancing if one worker becomes a hotspot.
Prefill/decode disaggregation
llm-d supports separating prompt processing and token generation into different worker pools. This lets operators provision and tune resources for each phase rather than asking the same pool to handle both. It can help workloads with a pronounced imbalance between long prompts and output generation, but it adds network transfers, scheduling decisions and failure modes. Network bandwidth and latency become more consequential, so the design must be evaluated against the actual workload.
Latency-aware scheduling and service objectives
The project describes routing and scheduling that account for load and latency, including predicted-latency scheduling and service-level-objective-related request handling. Such mechanisms aim to make request placement more informed than round-robin distribution. Their real-world impact depends on accurate signals, the traffic mix and the SLOs operators configure.
KV-cache offloading
The launch announcement also described moving KV-cache pressure beyond scarce GPU memory, using CPU memory or network storage and technologies such as LMCache. Offloading can expand effective cache capacity, but it is not free capacity: operators must account for memory bandwidth, network traffic, storage behavior and cache invalidation.
Rank #3
Multiple engines and accelerator environments
llm-d’s stated direction is portability across model servers, clouds and accelerator types. Current project and Red Hat materials discuss a range of accelerator environments, but compatibility is not a blanket promise. A working configuration depends on the model architecture, serving engine, kernels, quantization, hardware topology, transport libraries and maturity of the specific integration.
How llm-d differs from related tools
| Technology | Primary role |
|---|---|
| vLLM | Runs and serves models efficiently on supported hardware. |
| llm-d | Coordinates multiple model-serving instances and adds inference-aware routing and distributed scheduling. |
| Kubernetes | Orchestrates containers, nodes, accelerators, networking and scaling. |
| KServe | Provides model-serving APIs, abstractions and deployment integrations, including llm-d paths. |
| Inference Gateway / Gateway API extensions | Routes requests with inference-specific context rather than relying only on generic load balancing. |
| NVIDIA Dynamo | An alternative integrated inference stack focused on high-scale serving, particularly relevant to NVIDIA-standardized environments. |
The core distinction is that llm-d complements vLLM rather than replacing it. A single vLLM server on one GPU may not need another orchestration layer. llm-d is more relevant when model size, concurrency, prompt length, availability requirements or traffic volume make a multi-worker or multi-node system worthwhile.
KServe and llm-d can be used together: KServe supplies an abstraction for managing serving deployments, while llm-d provides distributed inference behavior. NVIDIA Dynamo is an alternative architecture, not a component required by llm-d. Comparisons between these projects should be tested against a specific workload; descriptions in the llm-d proposal reflect the project’s perspective, not neutral third-party benchmarking.
Project status and commercial deployment status are different
As of August 2026, the llm-d repository lists v0.7, released in May 2026. The project’s release notes describe a stabilized optimized baseline, kustomize-first guides, expanded nightly CI across OpenShift, GKE and CoreWeave, and generally available predicted-latency scheduling. They also identify an experimental batch gateway and document capabilities developed in earlier releases, including hierarchical KV offloading, cache-aware LoRA routing, active-active high availability and scale-to-zero autoscaling. These are project-reported release details, not independent proof that every feature is appropriate for every production workload.
“Production-ready” therefore needs a deployment-specific answer. The open-source project is designed for production-scale inference and publishes deployment guides and releases. Red Hat’s documented distributed-inference path on selected managed Kubernetes services, however, is marked Technology Preview and is not covered by production SLAs. Preview status matters to buyers who need a contractual support commitment. Check the current support terms and validated configurations for the exact product, release and cloud before relying on them.
Free tools Windows power users keep installed
One-click scans. No signup required.
Red Hat has incorporated llm-d into its commercial Red Hat AI Inference stack. In May 2026, the company described validated deployment blueprints for CoreWeave Kubernetes Service and Azure Kubernetes Service. This commercial offering is distinct from the community project and from broader products such as OpenShift AI. Open-source availability does not mean that every configuration is supported by Red Hat, nor does it remove the costs of GPU capacity, cloud services, subscriptions or operations.
Deployment paths and a Red Hat preview example
Depending on the version and environment, documented paths include direct Kubernetes deployment, vLLM integration, KServe’s LLMInferenceService, OpenShift AI, and Red Hat AI Inference on selected managed Kubernetes services. The setup and support status differ among those paths; a command for one should not be treated as a universal llm-d installer.
For the Red Hat AI Inference managed-Kubernetes Technology Preview documented in June 2026, prerequisites include Kubernetes 1.33 or later, Helm 3.17 or later with OCI support, GPU nodes, authentication for registry.redhat.io and quay.io, and Red Hat AI Inference Server early-access credentials. Red Hat’s guide installs supporting components including KServe, cert-manager, Istio and LeaderWorkerSet. The following Azure example is specific to that guide and requires the stated access and prerequisites:
helm registry login registry.redhat.io
helm upgrade rhaii oci://quay.io/rhoai/rhai-on-xks-chart
--install
--create-namespace
--namespace rhaii
--set azure.enabled=true
--set-file imagePullSecret.dockerConfigJson=~/pull-secret.json
For the guide’s CoreWeave option, the provider settings are:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11--set azure.enabled=false --set coreweave.enabled=true
The same guide estimates that deployment takes about 5–10 minutes, but that is an indicative documentation estimate, not a guarantee. Its example workload uses the serving.kserve.io/v1alpha1 API and an LLMInferenceService to deploy two replicas of Qwen3 8B, requesting one NVIDIA GPU per pod. It is an example configuration, not a general hardware recommendation. Consult the Red Hat deployment guide for the complete, version-specific procedure and caveats.
Best Value
How to evaluate llm-d without overreading benchmarks
Red Hat has reported production results attributed to Red Hat and Tesla engineers using Llama 3.1 70B: 3× higher output throughput and a 2× reduction in time to first token with intelligent routing. The available claim should not be read as a general llm-d guarantee; results depend on hardware, quantization, context lengths, concurrency, cache reuse, network and the round-robin baseline.
The CNCF announcement also describes a project benchmark using Qwen3-32B, eight vLLM pods and 16 NVIDIA H100 GPUs, reporting near-zero time to first token and approximately 120,000 tokens per second under its test conditions. Treat that as a cited project benchmark, not an independently established baseline for other models or deployments.
A useful proof of concept compares llm-d with a simpler serving setup under controlled conditions:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →- Establish a baseline. Record results from the existing server or ordinary routing before changing architecture.
- Hold variables constant. Use the same model, quantization, hardware, prompt set, output limits and concurrency in each comparison.
- Measure more than tokens per second. Track time to first token (TTFT), inter-token latency, output throughput, GPU utilization, cache hit rate, error rate and cost per output token.
- Test representative traffic. Include short and long prompts, repeated prefixes, retrieval-augmented generation (RAG), and multi-turn or agentic conversations.
- Compare deployment shapes. Measure one-node, multi-node and prefill/decode-disaggregated configurations where relevant.
- Inject failures and spikes. Test worker and node loss during prefill and decode, cache eviction, node replacement, autoscaler reaction time, cold model loads, traffic surges, cancellation and retry behavior, and long-running workflows.
- Include operating costs. Count accelerator rental, CPU and memory nodes, networking, cache storage, Kubernetes and observability, model loading and scaling overhead, engineering and on-call work, and any applicable subscription or support costs.
Results should answer a practical question: does inference-aware coordination improve the latency, throughput, reliability or cost that matters to this service enough to justify additional operational complexity?
When llm-d is a plausible fit
Consider an evaluation if your organization already operates Kubernetes or OpenShift, serves traffic across multiple model replicas or GPU nodes, and has measurable pressure on latency, throughput or accelerator utilization. Long prompts, repeated prefixes, RAG and agentic traffic can make cache locality more important. A team seeking a more portable stack across infrastructure providers may also find the project relevant, subject to verifying its specific model and hardware combination.
It is less compelling for a low-volume workload running on one GPU or one vLLM instance, for a team without Kubernetes expertise, or when a managed model API already meets privacy, latency and cost requirements. A simpler server can be cheaper and easier to operate. “Distributed” is not automatically better.
Alternatives to consider
- vLLM alone: A simpler choice for one server or a modest deployment when an efficient model engine is the main need.
- KServe with vLLM: A modular approach for teams standardizing model-serving APIs and managing multiple deployments, without necessarily adopting every llm-d optimization.
- NVIDIA Dynamo: An alternative to evaluate for organizations committed to NVIDIA’s integrated inference ecosystem.
- AIBrix: A fast-moving, research-oriented alternative to assess for governance, support, release maturity and production commitments.
- Managed model APIs: A better fit when reducing infrastructure work and accelerating application delivery matter more than control over model weights, hardware or data locality.
Likewise, choosing a supported Red Hat path is a commercial platform decision, not a requirement for using the community project. Red Hat AI Inference may suit enterprises seeking validated integrations and vendor support; OpenShift AI is broader for teams that also need development, training, serving and monitoring capabilities. For either, verify current support scope and pricing directly—no universal public price follows from the project’s open-source status.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The practical takeaway
Red Hat’s llm-d project addresses a real limitation of generic load balancing: LLM inference has cache state, distinct prompt and generation phases, and latency goals that can make request placement important. Its combination of Kubernetes orchestration, inference-aware routing and model-server integrations could help at scale, but the gains depend on workload and configuration. Start with a measured comparison against a simpler vLLM deployment, and distinguish the CNCF community project from Red Hat’s product support and preview status.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

