.
All right. Welcome, everyone, to yet another Inference Talk. I hope you have had a good conference so far. In this session, I'm sure people who have been in the room must have heard these terms many times by now, so we're going to do a little bit more deep dive into the challenges of LLM deployments for agentic workloads. In this session, we'll focus specifically on KV cache-aware routing and PD disaggregation. Also, when you look at public inference benchmark results, you are typically looking at very steady-state, isolated, highly sanitized numbers.
What those benchmarks actually don't show you is the chaotic reality of multi-turn interactions and massive context fluctuations, which are very typical of agentic workloads. So we'll also try to pull the curtain back on some of those complexities. By way of introduction, my name is Ashish Kamra. I'm a senior manager of performance engineering at Red Hat. And with me... Hi, I'm Yu Chen. I'm the product manager at Red Hat Inference, working closely with vLLM and LLMD core maintainers, also a contributor myself. So here's the agenda for the next 20 minutes or so.
Yu Chen will start with an analysis of inference behavior in the agentic era and some of the core characteristics and challenges. Next, you will walk us through the KV cache utilization and management strategies. I will break down the mechanics of prefill decode disaggregation and walk you through some results. And then Yu Chen will, again, bring it all back together with our ongoing case study on our favorite open coding model, GLM 5.2.
Just a couple of sources from our side. If you are more interested in learning more about open source inference, we have a free course on deeplearning.ai by Cedric and with Andrew Ng. The other is a series of blogs on the Red Hat Developer Portal on distributed inference concepts, troubleshooting, and deployment patterns.
For those who may not be aware, Red Hat is better known as the Linux company for Enterprise Linux and the Kubernetes company for OpenShift, but more recently we are also a major player in open source AI inference, with us being the top contributor in vLLM, LLMD, and the KServe projects, and also having incubated guide LLM for benchmarking, LLM compressor for model quantization, and speculators for speculative decoding models. We also bring all of that together in an optimized model hub on Hugging Face under the Red Hat AI org.
And we are also building the platform for the next wave of agentic inference workloads, and with that I will hand over to Yu-Chen to walk you through more of it. So we are currently at this inflection point, moving from the era of classic LLM inference to the agentic era. When we look at real-world agentic workloads such as SWE-bench and also traces from real-world cloud code sessions, they fundamentally break many assumptions we made with classic LLM serving. As you heard many times in previous sessions, for example, multi-turns are the new standard. We found from a few turns all the way to 3,000 turns.
Also, because agents frequently reuse the system prompt and the tool definitions, we usually see super high cache hit rates, oftentimes exceeding 90%. Another thing is the input-output ratio is massive, oftentimes over a 100:1 ratio and even higher in many cases. On top of that, the context management is incredibly complex due to this high variance, because we can't just simply take the average, and oftentimes we need to look at distributions and the P90 numbers, especially when you do capacity planning. We also observe really interesting patterns like sub-agent fan-out, which further complicates scheduling.
To help communities study these patterns, we collaborated with Google and also IBM, our parent company, to add a trace replay tool in Inference Perf. You heard from earlier sessions from Ashok and Jason. So feel free to check it out, and the link is here. Next slide. Transitioning from the characteristics we just saw, for agentic workloads, we're no longer chasing raw throughput in a steady state. We often need to optimize, for example, interactive latency, and there are very highly volatile and client-driven contexts because users and clients define the prompt structure.
So this introduces several critical challenges. First of all, KV cache management becomes super volatile because the context is client-determined, as I said. Oftentimes we face frequent evictions and rewrites. Secondly, we also need to tune the engine, like vLLM, with upper-layer scheduling and routing. It needs that coordination, such as prefix-aware routing, especially when latency becomes a primary scheduling metric rather than a secondary or afterthought. Thirdly, we also need to rethink our metrics. For example, we need to measure cache throughput separately. Why? Because on the right, it's really clear that the economic stakes are very high.
This is the Anthropic API pricing you also heard from earlier sessions. There's a 10x cost difference between cached and non-cached tokens. A 10x difference on your token balance sheet is a pretty serious impact on your business. So let's look at how the KV cache is both utilized and managed in LLMD. LLMD router has these really flexible endpoint picker plugins, we call them EPP, that can route the request to the optimal pods based on the KV cache locality and also the load criteria. The EPP continues to probe each pod via pod metrics to score each pod, for example, by running and waiting requests, and then the KV cache utilization, and also prefix cache availability.
So we can schedule requests to the optimal pod with the lowest load and also highest possibility of a cache hit.
Going down to the KV cache management layer, you also heard from the earlier session right before this. For agentic sessions, when you have hot, warm, and cold cache, our current effort focuses on, for example, more offloading tiers like NVMe, SSD, and also file system cache, along with KV-centric stores like Mooncake, and also implementing smarter and session-aware eviction policies, such as priority and also session pinning, to ensure this really important context persists exactly when and where it's needed. So I'm going to play this video really quick. It's a short demo. I'm going to stand here to look at it. Okay. So this is an example of KV cache-aware routing.
As you see, when we send the very first request and it populates the KV cache, it takes roughly three seconds. When we actually look at where the KV cache is going, there's no KV cache hit because it's the very first turn. Then when we have the second turn, the request actually reuses the KV cache because, as you see, the system prompt is the same, and this time it takes about one second. When you actually look at the pod address, it's exactly the same because we defined the KV cache. Now going to the third turn, a new request with a different system prompt, it takes about three seconds.
As you see, right now we don't find any KV cache hit because you can tell it's a different pod address. Then if you just change the user prompt and keep the same system prompt, on the next turn, it reuses the KV cache, and this time it takes roughly about one second. Yeah. It's a pretty intuitive demo. And I'll turn it to Ashish to talk about the next slide. But before that, what problem does it solve? Oftentimes, the prefix-aware routing, KV cache routing, helps you solve the TTFT problem. Of course, it will improve your throughput. But oftentimes for agentic workloads, it's not just the TTFT. Your throughput is about your inter-token latency. How do we solve that?
So prefill decode disaggregation is a really powerful technique, but there are times that it works and times it doesn't work. I'll turn it to Ashish to give you a preview of the PD disaggregation. So before we dive into PD, let's just look at what LLMD is. LLMD is a high-performance Kubernetes-native, and actually now works on non-Kubernetes environments as well, distributed LLM inference framework hosted under the CNCF umbrella. LLMD provides a unified intelligent control plane designed specifically for the agentic era of inference workloads. Yu Chen already talked about the router and the EPP at the top of the slide.
The other aspects are workload APIs, such as leader-worker set and disaggregated set, that orchestrate complex multi-node model execution. And then autoscalers that monitor capacity bounds and real-time traffic mixes to independently scale up and scale down your pods, depending on the system load. So now let's look at prefill decode disaggregation in detail. Okay, so why does PD exist in the first place? One of the most powerful patterns implemented by LLMD is prefill decode disaggregation, and you must have heard from some of the previous talks as well.
What happens is, in a non-PD situation, in aggregated serving, one pod is responsible for optimizing both your time-to-first-token and your inter-token latencies. But in PD, prefill and decode become independently scalable inference pods. To understand why we actually need this, we have to look at the physics of LLM execution. Co-locating both prefill and decode tasks on the same GPU creates something called phase interference. The prefill phase is the phase that creates the KV caches for your initial prompt. It wants high compute, it's highly bursty, utilizes GPUs at high FLOPS, and thrives on large batch parallelism to process the prompts and build the initial KV cache.
The decode phase, on the other hand, is generating one token at a time, and it's more memory-bandwidth hungry. It's highly latency-sensitive and requires high KV cache residency. In a traditional aggregated pod, if there's a sudden influx of a long prefill prompt, it will completely stall the ongoing decode token generation process, causing massive problems and jitter in user streaming latency. So how does PD actually work in practice in LLMD? LLMD uses, okay, let's start with step one.
An incoming request hits the gateway router, which dynamically evaluates cluster states using something known as the endpoint picker, Yu Chen talked about, and schedules the request to use PD disaggregation, selecting the optimal prefill and decode workers. The router then coordinates the transaction directly with the designated prefill worker. The prefill worker processes the prompt, constructs the initial KV cache of the prompt, and outputs the standard KV transfer metadata. The target decode worker actually pulls the computed KV caches across the network fabric, utilizing the KV transfer metadata that the prefill pod had generated.
Okay, so with that, yes, that's kind of how PD is implemented in practice in LLMD. Next, I would like to show you some experimental results on where PD actually shines. In this graph, you can see that in the standard aggregated deployment, which is the top red line, the P99 ITL hovers roughly around 900 milliseconds, and you can see some fluctuations up and down. But the bottom blue line is the P99 inter-token latency on a PD deployment, and you can see that it's drastically, almost nine times better, at 100 milliseconds, and it's also much smoother than the aggregated serving. This is some of our own internal results at Red Hat.
For a GPT-OSS 120b model, 16 H100s, the aggregated config is four replicas, tensor parallelism four, and the disaggregated is two prefill, two decode, all with tensor parallelism four. It's a highly multi-turn workload with a 10,000-token prefix and 128 tokens for every turn. This is a great chart. You can see at the bottom-most line is a standard aggregated config that is doing the default Kubernetes scheduling, and it's aggregated. So that's kind of our baseline. Then the middle blue line is still aggregated, but with the LLMD KV cache-aware routing, and you can almost see the gains just based on the routing.
The red line is actually the PD, the prefill decode config, with two prefill and two decode workers. You can actually see that it's very similar to the aggregated config at the lower concurrency regimes, and very similar at the higher concurrency regimes. But it's actually the middle part of the concurrency regime where PD actually shines. These are some of the classic Pareto curves that we see when you actually do PD and aggregated side by side. These results are again from the GPT-OSS 120b model, 64 H100s.
Aggregated is eight replicas, TP8, and disaggregated is three prefill, five decode, again TP8, and a prefill-heavy workload with 5,000 average input sequence length and 500 output sequence length. You can actually see the blue line is the PD curve, and the red line is the aggregated curve, and the PD curve kind of dominates the aggregated curve across the entire interactivity spectrum. Okay, but I don't want to leave you guys with the idea that PD is the answer to everything and it's a magic bullet, because it's essentially a phase-separation tradeoff and not a magic bullet. So we created this matrix to help you decide when PD might be good for you.
If you're managing long context with high ISL-OSL ratios, and if you have a large model that you're serving that can apply rich model parallelism techniques, you're facing that middle concurrency regime that I showed you in the previous graphs, and the very important part is that if you want strict ITL streaming requirements, like you want the token generation to be much smoother, then you want to consider PD. But we also saw that it requires transfer of KV caches from your prefill workers to your decode workers. So you must possess an advanced high-speed network fabric like RDMA or RoCE to support that KV cache transfer.
If you do not have such requirements, short to moderate context, any model size, low concurrency regimes, or if you have strict TTFT requirements because you can actually tune them on an aggregated serving, and the biggest point is if you don't have the network fabric to support those KV cache transfers, then you might actually just want to stick with aggregated. So here is my key takeaway from all of this. Architecting this complex platform requires balancing a lot of knobs in a highly multidimensional design space, all of which is supported in LLMD, as we saw.
The scheduler must constantly evaluate SLO targets, queue depths, KV cache locality metrics, PD ratios, and network topologies to be able to route the request to the optimal pod. While in the PD design space, you need dynamic PD rate matching to adapt to PD ratios, because you can start with a static PD ratio, but it needs to evolve with the autoscaler as the traffic changes. And you need autoscaling to scale PD pools independently and constantly tweak model parallelism techniques like tensor parallelism and data parallelism to meet your SLOs.
So I think with these, I will hand it over to Yu Chen to anchor some of the concepts that we showed in the real-world case study of serving the GLM 5.2 model, which is still ongoing as we speak. Yeah, still ongoing. You probably have seen tons of impressive numbers of GLM 5.2 on B200. When we talk to our customers, they usually don't have the luxury of B200. They have a lot of H200. So we have to figure out how to put all the knobs together and make GLM 5.2 work really well for a cluster of H200. So let's anchor all the concepts together.
We went through, for example, the KV cache routing, PD disaggregation, we kind of call them a well-lit path in LLMD, and also we combine with different parallelism strategies so we can independently scale prefill paths because agentic workloads are super lumpy, heavy prefill. In this case, we designed the prefill pool using up to three workers, optimized for high throughput with DeepEP. Then for the decode pool, we use one dedicated worker and that's optimized for low latency. So we use NIXL for efficient KV transfer between the pools. Also, with each worker, we have the leader-worker stack group with TP1, DP8, and also EP8, expert parallelism 8.
So the architecture is highly modular because you can actually scale the throughput by simply adding prefill workers without reconfiguring the decode pool. This highlights how LLMD effectively manages the complexity of combining TP and DP and EP at scale.
We also found some interesting fun facts a couple of days ago. BF16 KV cache actually is faster than using FP8 KV cache for longer prefills. This is also something we continue to explore and find more interesting patterns. But more importantly, we want to just show the result really quick. For this dataset, an agentic workload dataset, the ISL-OSL ratio is pretty high, a 5:1 ratio. Prefill is really the constraint, you can tell. So for this dataset, we have 4x faster TTFT and also 60% more requests. And this is continuously a work in progress. The next step is we need to also put in the upper-layer scheduler, lower TTFT, and also add more prefill replicas.
I know we are running out of time really quickly. So for this quick summary, for the fundamental shift for agentic workloads, we're continuing to have this LLMD, the agentic North Star, with session graph orchestration, program-aware scheduling, state reuse lifecycle, and also the agentic benchmark we're working on. You can find them in LLMD, upstream LLMD. Also, feel free to join the SIG group and contribute. And this is the very last slide. Distributed inference is not a challenge every single company can solve alone.
We're proud to be building this future in the open alongside our incredible ecosystem collaborators, CoreWeave, Google, IBM, NVIDIA, a growing list of launch partners, and industry adopters. So if you're passionate about the future of open source inference, we invite you to join us. We do have a booth downstairs. Feel free to stop by and ask any questions. And thank you so much for your time.
Thank you. but it's essentially a phase separation trade-off and not a magic bullet. So we created this matrix to help you decide when PD might be good for you. So if you're managing long context with high ISL-OSL ratios, and if you have a large model that you're serving that you can apply rich model parallelism techniques, you're facing that middle concurrency regime that I showed you in the previous graphs, and the very important part is that if you want strict ITL streaming requirements, like you want the token generation to be much more smooth, then you want to consider PD.
But we also saw that it requires transfer of KV caches from your pre-fill workers to your decode workers. So you must possess an advanced high-speed network fabric like RDMA or ROCKEY to support that KV cache transfer. And if you do not have such requirements, short moderate context, any model size, low concurrency regimes, or if you have strict TTFT requirements because you can actually tune them on an aggregate serving. And the biggest point is if you don't have the network fabric to support those KV cache transfers, so you might actually just want to stick with aggregated. So here is my key takeaway from all of this.
So architecting this complex platform requires balancing a lot of knobs and a highly multidimensional design space, all of which is supported in LLMD, as we saw. The scheduler must support or constantly evaluate SLO targets, Q-Depts, KV cache locality metrics, PD ratios, and network topologies to be able to route the request to the optimal FOD. While in the PD design space, you need dynamic PD rate matching to adapt to PD ratios, because you can start with a static PD ratio, but it needs to evolve with the autoscaler as the traffic changes. And you need autoscaling to scale PD pools independently and constantly tweaking model parallelism techniques
like tensor parallelism, data parallelism to meet your SLOs. So I think with these, I will hand it over to Yuchen to anchor some of the concepts that we showed the real-world case study of solving the GLM 5.2 model, which is still ongoing as we speak. Yeah, still ongoing. You probably have seen tons of impressive numbers of GLM 5.2, B200. When we talk to our customers and they usually don't have the luxury of B200, they have a lot of H200. So we have to figure out how to put all the knobs together and make GLM 5.2 work really well for cluster of H200. So let's anchor all the concepts together.
We went through, for example, the KV cache routing, PD desegregation, we kind of call them a wild-lit path in LMD, and also we combine with different parallelism strategies so we can independently scale pre-fuel paths because for agentic workload is super lump, you know, like heavy pre-fuel. So in this case, we designed the pre-fuel pool using up to three workers, optimized for high throughput with deep EP. And then for deco pool, we use one dedicated worker and that's optimized for low latency. So we use Nixo for efficient KV transfer between the pools. And also with each worker, we have the leader worker stack group with TP1, DP8, and also EP8, expert parallelism 8.
So the architecture is highly modular because you can actually scale the throughput by simply adding a pre-fuel workers without reconfiguring and the deco pool. So this highlights how LMD effectively manage the complexity of combining like TP and DP and EP at scale. And also we found some interesting fun fact actually a couple days ago.
BIF16KV cache actually is faster than using like FPKV cache for longer pre-fuel. This is also like we continue to explore and found like more interesting patterns. But more importantly, we want to kind of just show the result really quick. So for this data set, agentic workload data set, the ISO ratio is pretty high for the 5 to 1 ratio. Pre-fuel is really the constraint you can tell. So for this data set, we have 4X passers TTFT and also 60% more requests. And this is continuous like a work in progress. So the next step is we need to also put in the upper layer scheduler, lower TTFT, and also adding more pre-field replicas.
So I know we are running out of time really quick. So for this quick, we, the fundamental shift for agentic workload, we're continuing to have this LMD, the agentic North Star with session graph orchestration, program award scheduling, state reuse, lifecycle, and also the agentic benchmark we're working on. So you can find them in LMD, upstream LMD. And also, you know, feel free to join the seed group and contribute. And this is the very last slide. So distributed inference is not challenge every single company can solve a lot. We're proud to be building this future in the open alongside our incredible ecosystem collaborators,
CoreWave, Google, IBM, NVIDIA, growing list of launch partners and industry adopters. So if you're passionate about the future of open source inference, we invite you to join us. We do have a booth downstairs. Feel free to stop by, ask any questions. And thank you so much for your time.
Thank you.