Open Reader

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat

completed 21:47 Aug 27, 2026 Watch on YouTube

Current Status

completed

Video ID

YXowceUKYJI

RAG / Chat

Enabled
KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat
Description

Agentic sessions in Red Hat's traces run from a few turns to 3,000, cache hit rates routinely clear 90%, and input to output token ratios often pass 100 to 1. A public inference benchmark shows none of that, because it reports steady state numbers from one sanitized run. Yuchen Fama and Ashish Kamra spend the talk on the two levers that matter once the client rather than the server controls the cache lifecycle, and a live demo makes the first one concrete. An opening request takes about 3 seconds, the next turn reuses the cache on the same pod and takes about 1, and a fresh system prompt lands on a different pod and pays the full 3 again. With a 10x gap between cached and uncached token costs, routing is the cheaper lever to reach for before adding GPUs. The second lever splits compute bound prefill from memory bound decode, so a long incoming prompt cannot stall token generation midstream. Across 16 H100s serving gpt-oss, P99 inter token latency falls from roughly 900 milliseconds to about 100, and the curve gets visibly smoother. The useful part is where they draw the boundary. Disaggregation wins in the middle concurrency band, roughly ties at both ends, and needs an RDMA or RoCE fabric to move cache between workers at all. Without one, stay aggregated. The closing case study runs GLM 5.2 on the H200s customers actually have instead of B200s, at three prefill workers to one decode, for 4x faster time to first token and 60% more requests. Speaker info: - https://www.linkedin.com/in/yuchen-fama - https://www.linkedin.com/in/ashishkamra/ - https://github.com/llm-d/llm-d Timestamps: 0:00 - What public inference benchmarks leave out 1:28 - Red Hat's inference stack, and the agenda 3:29 - Agentic traces: 3,000 turns, 90% cache hits, 100 to 1 ratios 5:12 - Volatile cache, and the 10x cached token gap 6:28 - How llm-d routes: endpoint picker, offload tiers, eviction 7:46 - Demo: cache hits, pod affinity, 3 seconds versus 1 9:28 - What llm-d is, and why prefill and dec

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agentic LLM serving requires treating KV-cache locality, prefill/decode separation, and dynamic scheduling as a coordinated control-plane problem rather than optimizing steady-state model throughput alone.
  • Why it matters: For multi-turn agents with long, volatile contexts, cache hits can materially reduce both latency and token cost, while prefill/decode (P/D) disaggregation can protect streaming quality—but only under the right workload, hardware, and network conditions.
  • Best use: Use this as an architecture and evaluation guide for Kubernetes-based agent inference: first implement cache-aware routing, then selectively test independently scaled prefill/decode pools against real trace-replayed workloads and SLOs.

Executive Summary

Red Hat argues that conventional inference benchmarks obscure the operating reality of agentic workloads: conversations can span from a few turns to 3,000 turns, input-to-output ratios can exceed 100:1, contexts fluctuate substantially, and sub-agents can fan out unpredictably. These workloads repeatedly reuse system prompts and tool definitions, producing cache-hit rates often above 90%, so routing a request to a worker holding its KV cache becomes a primary latency and cost lever rather than a minor optimization.

Their LLMD control plane addresses this with endpoint-picker plugins that continuously score pods using running and queued requests, KV-cache utilization, and prefix-cache availability. In their demo, first turns with a new system prompt took about three seconds, while subsequent turns routed to the same cache-bearing pod took roughly one second. The stated economic rationale is Anthropic-style pricing where uncached input tokens can cost 10 times cached tokens.

P/D disaggregation separates compute-heavy, bursty prefill from memory-bandwidth-bound, latency-sensitive decode. This prevents long prompt prefills from interrupting active token streams on the same GPU. Red Hat reports a P99 inter-token-latency result of roughly 100 ms for P/D versus roughly 900 ms for aggregated serving in one GPT-OSS 120B, 16-H100 experiment, and shows that cache-aware routing alone improves an aggregated baseline.

The speakers are careful not to position P/D as universal. It is most compelling for long-context, prefill-heavy workloads at middle concurrency with strict streaming-latency requirements and RDMA/RoCE-class network fabric for KV transfer. For short contexts, low concurrency, strict time-to-first-token needs, or weaker networking, aggregated serving may be preferable. The operational conclusion is that the scheduler, autoscaler, cache lifecycle, P/D worker ratio, model parallelism, and topology must adapt together as traffic changes.

Key Takeaways

  • Claim: Agentic workloads invalidate steady-state inference assumptions because they are multi-turn, context-variable, and heavily dependent on prompt-state reuse. | Evidence: The speakers cite workloads ranging from a few turns to 3,000 turns, cache-hit rates often exceeding 90% from repeated system prompts and tool definitions, input-output ratios over 100:1, and sub-agent fan-out. They recommend studying distributions and P90 values rather than averages for capacity planning. | Implication: Capacity and routing decisions for Ken's agent systems should be trace-driven and session-aware; standard isolated throughput tests will underrepresent cache churn, tail latency, and prefill bursts. | Caveat: These figures are presented as observations from SWE-bench and cloud coding-session traces, not as universal workload guarantees.
  • Claim: KV-cache-aware, prefix-aware routing is an immediate high-value optimization for multi-turn agent sessions because it reduces time to first token by preserving cache locality. | Evidence: LLMD's endpoint picker plugins score pods using running and waiting requests, KV-cache utilization, and prefix-cache availability. In the demo, a cold first request took roughly three seconds; a subsequent request with the same system prompt was routed to the same pod and took about one second. A different system prompt went to another pod and returned to about three seconds. | Implication: Before introducing more complex inference topology, Ken should prioritize session affinity/prefix routing and instrument cache hit rate, cache residency, queue depth, and TTFT by request class. | Caveat: Routing requires sufficiently current per-pod cache and load telemetry; cache locality must be balanced against queue depth rather than treated as the sole routing criterion.
  • Claim: Cache reuse has direct unit-economics significance, not merely performance value. | Evidence: The presentation references Anthropic API pricing with a stated 10x cost difference between cached and non-cached tokens. | Implication: Ken should treat reusable instruction, tool-schema, and long-lived agent context design as a cost-control problem: avoid gratuitous prompt rewrites and measure uncached-prefill token exposure. | Caveat: The cited price differential is an external API pricing example, not a claim that every self-hosted deployment realizes a literal 10x infrastructure saving.
  • Claim: P/D disaggregation can substantially improve streaming smoothness by isolating prefill bursts from decode work. | Evidence: Prefill is described as compute-intensive, bursty, and suited to large batches, whereas decode is memory-bandwidth-bound and latency-sensitive. In a Red Hat GPT-OSS 120B test on 16 H100s, with four TP4 aggregated replicas versus two TP4 prefill and two TP4 decode workers, reported P99 inter-token latency was roughly 900 ms aggregated versus roughly 100 ms with P/D. | Implication: If Ken's user experience depends on stable streamed tokens during long-context agent work, P/D is a credible architecture to benchmark, with P99 ITL—not aggregate throughput—as the principal success metric. | Caveat: This is an internal result on a specific highly multi-turn workload with a 10,000-token prefix and 128 output tokens per turn; it should not be generalized without replaying representative traffic.
  • Claim: P/D is a selective phase-separation trade-off, not a default serving architecture. | Evidence: The speakers recommend P/D for long contexts, high input-to-output ratios, large models capable of rich parallelism, middle-concurrency regimes, and strict ITL requirements. They recommend retaining aggregated serving for short-to-moderate context, low concurrency, strict TTFT optimization, or environments without high-speed fabric. | Implication: Ken should gate P/D adoption on a decision matrix: workload shape, concurrency band, desired TTFT versus ITL SLO, and measured KV-transfer performance on the actual cluster fabric. | Caveat: P/D requires transferring KV cache between workers, making advanced networking such as RDMA or RoCE a stated prerequisite; otherwise the transfer overhead can erase the benefit.
  • Claim: The operational problem is dynamic coordination of routing, independently autoscaled prefill/decode capacity, and parallelism—not a static P/D ratio. | Evidence: LLMD exposes workload APIs including leader-worker sets and disaggregated sets, while its autoscalers use capacity bounds and real-time traffic mixes to scale pods independently. The presenters state that static P/D ratios need dynamic rate matching as traffic shifts, alongside adjustment of tensor parallelism and data parallelism. | Implication: Ken should avoid hard-coding a fixed prefill/decode fleet split; scaling policy must respond to observed prefill pressure, decode queues, cache locality, and SLO violations. | Caveat: The talk presents the control-plane design and intended capabilities more than a detailed production validation of all dynamic policies.
  • Claim: A GLM 5.2 H200 deployment case study suggests modular P/D capacity can improve prefill-constrained agent workloads, while low-level cache-format choices remain empirical. | Evidence: For an agentic dataset with a 5:1 input-to-output ratio, the presenters report 4x faster TTFT and 60% more requests using up to three throughput-oriented prefill workers and one latency-oriented decode worker. Each worker used TP1, DP8, and EP8; DeepEP was used for the prefill pool and NIXL for KV transfer. They also observed BF16 KV cache outperforming FP8 KV cache for longer prefills. | Implication: Do not assume FP8 KV cache is automatically superior, or that benchmark results on B200 transfer to H200 clusters. Validate precision, parallelism, and pool sizing jointly on Ken's hardware and workload mix. | Caveat: The GLM 5.2 work is explicitly ongoing, with further upper-layer scheduling and additional prefill replicas planned; the reported gains lack full baseline and test-condition detail.

Detailed Brief

LLMD's control-plane framing

  • Claims: LLMD is presented as a Kubernetes-native, now also non-Kubernetes-capable, distributed LLM inference framework under the CNCF umbrella.; Its design goes beyond request routing: workload APIs orchestrate complex multi-node model execution, while autoscaling is intended to react separately to capacity bounds and live traffic mix.; Red Hat's longer-term 'agentic North Star' includes session-graph orchestration, program-aware scheduling, state-reuse lifecycle management, and agentic benchmarking.
  • Evidence: The named workload abstractions are leader-worker set and disaggregated set.; The endpoint picker is described as a pluggable component that probes pod metrics and selects workers based on load and cache-locality signals.; The team references a trace-replay contribution to Inference Perf made with Google and IBM to help study agentic traffic patterns.
  • Caveats: The future-state features are directionally described and should not be assumed to be mature, available, or production-ready without checking the upstream LLMD release and documentation.; The presentation is partly a product and ecosystem talk; its benchmark claims need independent replication and workload-specific validation.
  • Implications: The useful abstraction for an agent platform is a state-aware inference control plane, not simply a load balancer in front of interchangeable GPU replicas.; Trace replay should be part of performance qualification before changing topology, since synthetic fixed-context benchmarks will miss the scheduling problem being solved.

Performance-shape lessons from the experiments

  • Claims: KV-cache-aware routing improves an aggregated serving baseline even without P/D.; P/D's strongest advantage may occur in a middle-concurrency range rather than at either low or high extremes.; In a separate 64-H100 GPT-OSS 120B test, the P/D Pareto curve reportedly dominated aggregated serving across the displayed interactivity spectrum.
  • Evidence: The first experiment used a 10,000-token prefix and 128 output tokens per turn; the second used average 5,000-token inputs and 500-token outputs.; For the 64-H100 test, the aggregated configuration was eight TP8 replicas; the P/D configuration was three TP8 prefill workers and five TP8 decode workers.
  • Caveats: The transcript does not provide arrival rates, exact SLO definitions, transfer-network specifications, model settings, or confidence intervals for the charts.; The statement that P/D is similar to aggregated at lower and higher concurrency in one test reinforces that the performance envelope is workload-dependent.
  • Implications: Benchmarking should produce Pareto curves across concurrency and both TTFT and ITL, rather than declaring a topology superior from one throughput point.; Cache-aware routing is likely the lower-complexity baseline to establish before attributing gains to P/D.

Notable Concepts & Terms

  • KV cache-aware / prefix-aware routing: Routes a request to a worker likely to retain its prompt prefix KV cache, improving cache-hit probability and reducing prefill work and TTFT.
  • LLMD endpoint picker plugin (EPP): LLMD's pluggable pod-selection mechanism, using metrics such as queued/running requests, KV-cache utilization, and prefix availability.
  • Prefill/decode (P/D) disaggregation: Separates prompt processing and autoregressive token generation into independently scaled worker pools to prevent prefill bursts from disrupting decode latency.
  • Phase interference: The contention caused when compute-heavy prefills and memory-bandwidth-sensitive decode work share GPU resources in aggregated serving.
  • ITL / P99 ITL: Inter-token latency, especially tail inter-token latency; the speaker treats it as the critical streaming-quality metric that P/D can protect.
  • Dynamic P/D rate matching: Continuously adapting the relative prefill and decode capacity rather than retaining a static worker split as traffic composition changes.
  • RDMA / RoCE: High-speed networking technologies identified as important for moving KV cache between prefill and decode workers without nullifying P/D benefits.
  • TP, DP, and EP: Tensor, data, and expert parallelism; the case study combines these dimensions to serve a large MoE-style workload while scaling P/D pools modularly.

Operator Notes / Why Ken Should Care

  • Collect and replay representative agent traces segmented by turn count, prefix reuse, input/output ratio, concurrency, sub-agent fan-out, TTFT, and P99 ITL; do not use average request shape as the primary planning input.
  • Establish a cache-locality routing baseline and dashboard before testing P/D: track prefix-cache hit rate, same-session worker affinity, cache evictions, queue depth, uncached prefill tokens, TTFT, and P99 ITL.
  • Run an A/B topology experiment across low, middle, and high concurrency with aggregated serving, aggregated plus cache-aware routing, and P/D; compare both user streaming behavior and total request capacity.
  • Require a network-fabric readiness test for P/D, measuring end-to-end KV-transfer overhead under load. Avoid P/D on commodity networking unless measured transfer costs still preserve the desired ITL benefit.
  • Treat KV-cache precision as a benchmark variable rather than a presumed optimization; test BF16 versus FP8 cache for long-prefill workloads on the target GPU generation.
  • Design autoscaling around separate prefill and decode pressure signals, with safeguards against static pool ratios drifting out of balance as workload composition changes.

Source/Metadata

  • Title: KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat
  • Transcript words: 3984
  • Duration seconds: 1307
  • Timestamp note: No usable timestamps or chapters were provided. The transcript repeats the final P/D decision matrix and GLM 5.2 case-study segment.

Transcript

3119 words en Processed in 118.4s

. All right. Welcome, everyone, to yet another Inference Talk. I hope you have had a good conference so far. In this session, I'm sure people who have been in the room must have heard these terms many times by now, so we're going to do a little bit more deep dive into the challenges of LLM deployments for agentic workloads. In this session, we'll focus specifically on KV cache-aware routing and PD disaggregation. Also, when you look at public inference benchmark results, you are typically looking at very steady-state, isolated, highly sanitized numbers. What those benchmarks actually don't show you is the chaotic reality of multi-turn interactions and massive context fluctuations, which are very typical of agentic workloads. So we'll also try to pull the curtain back on some of those complexities. By way of introduction, my name is Ashish Kamra. I'm a senior manager of performance engineering at Red Hat. And with me... Hi, I'm Yu Chen. I'm the product manager at Red Hat Inference, working closely with vLLM and LLMD core maintainers, also a contributor myself. So here's the agenda for the next 20 minutes or so. Yu Chen will start with an analysis of inference behavior in the agentic era and some of the core characteristics and challenges. Next, you will walk us through the KV cache utilization and management strategies. I will break down the mechanics of prefill decode disaggregation and walk you through some results. And then Yu Chen will, again, bring it all back together with our ongoing case study on our favorite open coding model, GLM 5.2. Just a couple of sources from our side. If you are more interested in learning more about open source inference, we have a free course on deeplearning.ai by Cedric and with Andrew Ng. The other is a series of blogs on the Red Hat Developer Portal on distributed inference concepts, troubleshooting, and deployment patterns. For those who may not be aware, Red Hat is better known as the Linux company for Enterprise Linux and the Kubernetes company for OpenShift, but more recently we are also a major player in open source AI inference, with us being the top contributor in vLLM, LLMD, and the KServe projects, and also having incubated guide LLM for benchmarking, LLM compressor for model quantization, and speculators for speculative decoding models. We also bring all of that together in an optimized model hub on Hugging Face under the Red Hat AI org. And we are also building the platform for the next wave of agentic inference workloads, and with that I will hand over to Yu-Chen to walk you through more of it. So we are currently at this inflection point, moving from the era of classic LLM inference to the agentic era. When we look at real-world agentic workloads such as SWE-bench and also traces from real-world cloud code sessions, they fundamentally break many assumptions we made with classic LLM serving. As you heard many times in previous sessions, for example, multi-turns are the new standard. We found from a few turns all the way to 3,000 turns. Also, because agents frequently reuse the system prompt and the tool definitions, we usually see super high cache hit rates, oftentimes exceeding 90%. Another thing is the input-output ratio is massive, oftentimes over a 100:1 ratio and even higher in many cases. On top of that, the context management is incredibly complex due to this high variance, because we can't just simply take the average, and oftentimes we need to look at distributions and the P90 numbers, especially when you do capacity planning. We also observe really interesting patterns like sub-agent fan-out, which further complicates scheduling. To help communities study these patterns, we collaborated with Google and also IBM, our parent company, to add a trace replay tool in Inference Perf. You heard from earlier sessions from Ashok and Jason. So feel free to check it out, and the link is here. Next slide. Transitioning from the characteristics we just saw, for agentic workloads, we're no longer chasing raw throughput in a steady state. We often need to optimize, for example, interactive latency, and there are very highly volatile and client-driven contexts because users and clients define the prompt structure. So this introduces several critical challenges. First of all, KV cache management becomes super volatile because the context is client-determined, as I said. Oftentimes we face frequent evictions and rewrites. Secondly, we also need to tune the engine, like vLLM, with upper-layer scheduling and routing. It needs that coordination, such as prefix-aware routing, especially when latency becomes a primary scheduling metric rather than a secondary or afterthought. Thirdly, we also need to rethink our metrics. For example, we need to measure cache throughput separately. Why? Because on the right, it's really clear that the economic stakes are very high. This is the Anthropic API pricing you also heard from earlier sessions. There's a 10x cost difference between cached and non-cached tokens. A 10x difference on your token balance sheet is a pretty serious impact on your business. So let's look at how the KV cache is both utilized and managed in LLMD. LLMD router has these really flexible endpoint picker plugins, we call them EPP, that can route the request to the optimal pods based on the KV cache locality and also the load criteria. The EPP continues to probe each pod via pod metrics to score each pod, for example, by running and waiting requests, and then the KV cache utilization, and also prefix cache availability. So we can schedule requests to the optimal pod with the lowest load and also highest possibility of a cache hit. Going down to the KV cache management layer, you also heard from the earlier session right before this. For agentic sessions, when you have hot, warm, and cold cache, our current effort focuses on, for example, more offloading tiers like NVMe, SSD, and also file system cache, along with KV-centric stores like Mooncake, and also implementing smarter and session-aware eviction policies, such as priority and also session pinning, to ensure this really important context persists exactly when and where it's needed. So I'm going to play this video really quick. It's a short demo. I'm going to stand here to look at it. Okay. So this is an example of KV cache-aware routing. As you see, when we send the very first request and it populates the KV cache, it takes roughly three seconds. When we actually look at where the KV cache is going, there's no KV cache hit because it's the very first turn. Then when we have the second turn, the request actually reuses the KV cache because, as you see, the system prompt is the same, and this time it takes about one second. When you actually look at the pod address, it's exactly the same because we defined the KV cache. Now going to the third turn, a new request with a different system prompt, it takes about three seconds. As you see, right now we don't find any KV cache hit because you can tell it's a different pod address. Then if you just change the user prompt and keep the same system prompt, on the next turn, it reuses the KV cache, and this time it takes roughly about one second. Yeah. It's a pretty intuitive demo. And I'll turn it to Ashish to talk about the next slide. But before that, what problem does it solve? Oftentimes, the prefix-aware routing, KV cache routing, helps you solve the TTFT problem. Of course, it will improve your throughput. But oftentimes for agentic workloads, it's not just the TTFT. Your throughput is about your inter-token latency. How do we solve that? So prefill decode disaggregation is a really powerful technique, but there are times that it works and times it doesn't work. I'll turn it to Ashish to give you a preview of the PD disaggregation. So before we dive into PD, let's just look at what LLMD is. LLMD is a high-performance Kubernetes-native, and actually now works on non-Kubernetes environments as well, distributed LLM inference framework hosted under the CNCF umbrella. LLMD provides a unified intelligent control plane designed specifically for the agentic era of inference workloads. Yu Chen already talked about the router and the EPP at the top of the slide. The other aspects are workload APIs, such as leader-worker set and disaggregated set, that orchestrate complex multi-node model execution. And then autoscalers that monitor capacity bounds and real-time traffic mixes to independently scale up and scale down your pods, depending on the system load. So now let's look at prefill decode disaggregation in detail. Okay, so why does PD exist in the first place? One of the most powerful patterns implemented by LLMD is prefill decode disaggregation, and you must have heard from some of the previous talks as well. What happens is, in a non-PD situation, in aggregated serving, one pod is responsible for optimizing both your time-to-first-token and your inter-token latencies. But in PD, prefill and decode become independently scalable inference pods. To understand why we actually need this, we have to look at the physics of LLM execution. Co-locating both prefill and decode tasks on the same GPU creates something called phase interference. The prefill phase is the phase that creates the KV caches for your initial prompt. It wants high compute, it's highly bursty, utilizes GPUs at high FLOPS, and thrives on large batch parallelism to process the prompts and build the initial KV cache. The decode phase, on the other hand, is generating one token at a time, and it's more memory-bandwidth hungry. It's highly latency-sensitive and requires high KV cache residency. In a traditional aggregated pod, if there's a sudden influx of a long prefill prompt, it will completely stall the ongoing decode token generation process, causing massive problems and jitter in user streaming latency. So how does PD actually work in practice in LLMD? LLMD uses, okay, let's start with step one. An incoming request hits the gateway router, which dynamically evaluates cluster states using something known as the endpoint picker, Yu Chen talked about, and schedules the request to use PD disaggregation, selecting the optimal prefill and decode workers. The router then coordinates the transaction directly with the designated prefill worker. The prefill worker processes the prompt, constructs the initial KV cache of the prompt, and outputs the standard KV transfer metadata. The target decode worker actually pulls the computed KV caches across the network fabric, utilizing the KV transfer metadata that the prefill pod had generated. Okay, so with that, yes, that's kind of how PD is implemented in practice in LLMD. Next, I would like to show you some experimental results on where PD actually shines. In this graph, you can see that in the standard aggregated deployment, which is the top red line, the P99 ITL hovers roughly around 900 milliseconds, and you can see some fluctuations up and down. But the bottom blue line is the P99 inter-token latency on a PD deployment, and you can see that it's drastically, almost nine times better, at 100 milliseconds, and it's also much smoother than the aggregated serving. This is some of our own internal results at Red Hat. For a GPT-OSS 120b model, 16 H100s, the aggregated config is four replicas, tensor parallelism four, and the disaggregated is two prefill, two decode, all with tensor parallelism four. It's a highly multi-turn workload with a 10,000-token prefix and 128 tokens for every turn. This is a great chart. You can see at the bottom-most line is a standard aggregated config that is doing the default Kubernetes scheduling, and it's aggregated. So that's kind of our baseline. Then the middle blue line is still aggregated, but with the LLMD KV cache-aware routing, and you can almost see the gains just based on the routing. The red line is actually the PD, the prefill decode config, with two prefill and two decode workers. You can actually see that it's very similar to the aggregated config at the lower concurrency regimes, and very similar at the higher concurrency regimes. But it's actually the middle part of the concurrency regime where PD actually shines. These are some of the classic Pareto curves that we see when you actually do PD and aggregated side by side. These results are again from the GPT-OSS 120b model, 64 H100s. Aggregated is eight replicas, TP8, and disaggregated is three prefill, five decode, again TP8, and a prefill-heavy workload with 5,000 average input sequence length and 500 output sequence length. You can actually see the blue line is the PD curve, and the red line is the aggregated curve, and the PD curve kind of dominates the aggregated curve across the entire interactivity spectrum. Okay, but I don't want to leave you guys with the idea that PD is the answer to everything and it's a magic bullet, because it's essentially a phase-separation tradeoff and not a magic bullet. So we created this matrix to help you decide when PD might be good for you. If you're managing long context with high ISL-OSL ratios, and if you have a large model that you're serving that can apply rich model parallelism techniques, you're facing that middle concurrency regime that I showed you in the previous graphs, and the very important part is that if you want strict ITL streaming requirements, like you want the token generation to be much smoother, then you want to consider PD. But we also saw that it requires transfer of KV caches from your prefill workers to your decode workers. So you must possess an advanced high-speed network fabric like RDMA or RoCE to support that KV cache transfer. If you do not have such requirements, short to moderate context, any model size, low concurrency regimes, or if you have strict TTFT requirements because you can actually tune them on an aggregated serving, and the biggest point is if you don't have the network fabric to support those KV cache transfers, then you might actually just want to stick with aggregated. So here is my key takeaway from all of this. Architecting this complex platform requires balancing a lot of knobs in a highly multidimensional design space, all of which is supported in LLMD, as we saw. The scheduler must constantly evaluate SLO targets, queue depths, KV cache locality metrics, PD ratios, and network topologies to be able to route the request to the optimal pod. While in the PD design space, you need dynamic PD rate matching to adapt to PD ratios, because you can start with a static PD ratio, but it needs to evolve with the autoscaler as the traffic changes. And you need autoscaling to scale PD pools independently and constantly tweak model parallelism techniques like tensor parallelism and data parallelism to meet your SLOs. So I think with these, I will hand it over to Yu Chen to anchor some of the concepts that we showed in the real-world case study of serving the GLM 5.2 model, which is still ongoing as we speak. Yeah, still ongoing. You probably have seen tons of impressive numbers of GLM 5.2 on B200. When we talk to our customers, they usually don't have the luxury of B200. They have a lot of H200. So we have to figure out how to put all the knobs together and make GLM 5.2 work really well for a cluster of H200. So let's anchor all the concepts together. We went through, for example, the KV cache routing, PD disaggregation, we kind of call them a well-lit path in LLMD, and also we combine with different parallelism strategies so we can independently scale prefill paths because agentic workloads are super lumpy, heavy prefill. In this case, we designed the prefill pool using up to three workers, optimized for high throughput with DeepEP. Then for the decode pool, we use one dedicated worker and that's optimized for low latency. So we use NIXL for efficient KV transfer between the pools. Also, with each worker, we have the leader-worker stack group with TP1, DP8, and also EP8, expert parallelism 8. So the architecture is highly modular because you can actually scale the throughput by simply adding prefill workers without reconfiguring the decode pool. This highlights how LLMD effectively manages the complexity of combining TP and DP and EP at scale. We also found some interesting fun facts a couple of days ago. BF16 KV cache actually is faster than using FP8 KV cache for longer prefills. This is also something we continue to explore and find more interesting patterns. But more importantly, we want to just show the result really quick. For this dataset, an agentic workload dataset, the ISL-OSL ratio is pretty high, a 5:1 ratio. Prefill is really the constraint, you can tell. So for this dataset, we have 4x faster TTFT and also 60% more requests. And this is continuously a work in progress. The next step is we need to also put in the upper-layer scheduler, lower TTFT, and also add more prefill replicas. I know we are running out of time really quickly. So for this quick summary, for the fundamental shift for agentic workloads, we're continuing to have this LLMD, the agentic North Star, with session graph orchestration, program-aware scheduling, state reuse lifecycle, and also the agentic benchmark we're working on. You can find them in LLMD, upstream LLMD. Also, feel free to join the SIG group and contribute. And this is the very last slide. Distributed inference is not a challenge every single company can solve alone. We're proud to be building this future in the open alongside our incredible ecosystem collaborators, CoreWeave, Google, IBM, NVIDIA, a growing list of launch partners, and industry adopters. So if you're passionate about the future of open source inference, we invite you to join us. We do have a booth downstairs. Feel free to stop by and ask any questions. And thank you so much for your time. Thank you. but it's essentially a phase separation trade-off and not a magic bullet. So we created this matrix to help you decide when PD might be good for you. So if you're managing long context with high ISL-OSL ratios, and if you have a large model that you're serving that you can apply rich model parallelism techniques, you're facing that middle concurrency regime that I showed you in the previous graphs, and the very important part is that if you want strict ITL streaming requirements, like you want the token generation to be much more smooth, then you want to consider PD. But we also saw that it requires transfer of KV caches from your pre-fill workers to your decode workers. So you must possess an advanced high-speed network fabric like RDMA or ROCKEY to support that KV cache transfer. And if you do not have such requirements, short moderate context, any model size, low concurrency regimes, or if you have strict TTFT requirements because you can actually tune them on an aggregate serving. And the biggest point is if you don't have the network fabric to support those KV cache transfers, so you might actually just want to stick with aggregated. So here is my key takeaway from all of this. So architecting this complex platform requires balancing a lot of knobs and a highly multidimensional design space, all of which is supported in LLMD, as we saw. The scheduler must support or constantly evaluate SLO targets, Q-Depts, KV cache locality metrics, PD ratios, and network topologies to be able to route the request to the optimal FOD. While in the PD design space, you need dynamic PD rate matching to adapt to PD ratios, because you can start with a static PD ratio, but it needs to evolve with the autoscaler as the traffic changes. And you need autoscaling to scale PD pools independently and constantly tweaking model parallelism techniques like tensor parallelism, data parallelism to meet your SLOs. So I think with these, I will hand it over to Yuchen to anchor some of the concepts that we showed the real-world case study of solving the GLM 5.2 model, which is still ongoing as we speak. Yeah, still ongoing. You probably have seen tons of impressive numbers of GLM 5.2, B200. When we talk to our customers and they usually don't have the luxury of B200, they have a lot of H200. So we have to figure out how to put all the knobs together and make GLM 5.2 work really well for cluster of H200. So let's anchor all the concepts together. We went through, for example, the KV cache routing, PD desegregation, we kind of call them a wild-lit path in LMD, and also we combine with different parallelism strategies so we can independently scale pre-fuel paths because for agentic workload is super lump, you know, like heavy pre-fuel. So in this case, we designed the pre-fuel pool using up to three workers, optimized for high throughput with deep EP. And then for deco pool, we use one dedicated worker and that's optimized for low latency. So we use Nixo for efficient KV transfer between the pools. And also with each worker, we have the leader worker stack group with TP1, DP8, and also EP8, expert parallelism 8. So the architecture is highly modular because you can actually scale the throughput by simply adding a pre-fuel workers without reconfiguring and the deco pool. So this highlights how LMD effectively manage the complexity of combining like TP and DP and EP at scale. And also we found some interesting fun fact actually a couple days ago. BIF16KV cache actually is faster than using like FPKV cache for longer pre-fuel. This is also like we continue to explore and found like more interesting patterns. But more importantly, we want to kind of just show the result really quick. So for this data set, agentic workload data set, the ISO ratio is pretty high for the 5 to 1 ratio. Pre-fuel is really the constraint you can tell. So for this data set, we have 4X passers TTFT and also 60% more requests. And this is continuous like a work in progress. So the next step is we need to also put in the upper layer scheduler, lower TTFT, and also adding more pre-field replicas. So I know we are running out of time really quick. So for this quick, we, the fundamental shift for agentic workload, we're continuing to have this LMD, the agentic North Star with session graph orchestration, program award scheduling, state reuse, lifecycle, and also the agentic benchmark we're working on. So you can find them in LMD, upstream LMD. And also, you know, feel free to join the seed group and contribute. And this is the very last slide. So distributed inference is not challenge every single company can solve a lot. We're proud to be building this future in the open alongside our incredible ecosystem collaborators, CoreWave, Google, IBM, NVIDIA, growing list of launch partners and industry adopters. So if you're passionate about the future of open source inference, we invite you to join us. We do have a booth downstairs. Feel free to stop by, ask any questions. And thank you so much for your time. Thank you.