Open Reader

Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI

completed 18:12 Sep 19, 2026 Watch on YouTube

Current Status

completed

Video ID

sOB3HSiG8vo

RAG / Chat

Enabled
Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI
Description

The routing weights inside OpenAI's inference load balancer used to come out of a feedback loop. Engines reported signals, a controller smoothed them into a score, compared it to the fleet average, and nudged each weight up or down. A proportional controller, Lu Zhang notes, with real virtues: many signals folded into one decision, and constrained engines balanced themselves. It also produced behavior nobody could explain well. Ask why one engine got a higher weight and there was no clean answer. Tune one property and another moved. Worst was the oscillation: shift traffic off a hot engine, it cools, the controller reads cool as spare capacity and sends the traffic back, and the bouncing wrecks the KV cache locality routing was meant to protect. Qianru Lao walks through what replaced it: a control plane with a global view of every CPU cluster and GPU engine, and a data plane in each cluster that answers the one synchronous question, which engine serves this request, from a cached snapshot of routing weights. Signals still flow, into an optimizer rather than a loop. Its goal is to minimize expected end to end latency across all traffic, counting network distance and engine side queueing, under hard constraints that every request is routed and no engine exceeds capacity. Her example of why nearest is not enough: one region sends 120 requests a second at an engine that serves 100, while an engine two regions away sits at 40 of 80, so the farther engine wins once you count the wait. Zhang closes with the protections: outlier penalties, retry budgets that tighten as utilization climbs to prevent retry storms, and load shedding as last resort. Speaker info: - https://linkedin.com/in/qianru-lao - https://openai.com - https://www.linkedin.com/in/luzhang1/ Timestamps: 0:00 - From engine signal feedback loops to explicit policy 2:44 - What makes inference routing different 3:41 - Early days: weighted consistent hashing 4:23 - Weights from a proportional controller 6:57 - O

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Production LLM routing should move from opaque reactive feedback-loop weighting to an explicit global optimization policy, executed locally and protected by fast guardrails.
  • Why it matters: The talk provides a reusable control-plane/data-plane pattern for routing across heterogeneous, geographically distributed GPU capacity while balancing network latency, queueing/engine latency, cache locality, and failure resilience.
  • Best use: Use it as an architectural reference for an inference gateway or agent/model-routing control plane, especially the split between globally computed policy, locally cached execution, and overload protections.

Executive Summary

OpenAI describes the evolution of its inference load balancer (ILB), which sits between CPU-based frontend clusters and GPU-based inference-engine clusters. Its engine-selection problem is not conventional load balancing: routing must account for model and geography constraints, heterogeneous GPU capacity, live health and latency signals, network distance, and KV-cache locality for conversational follow-ups.

The initial approach used weighted consistent hashing with engine weights periodically adjusted from smoothed performance signals relative to fleet averages. This was adaptive and reduced manual intervention, but it was difficult to explain or tune. In particular, coupled feedback could oscillate traffic between engines as their apparent load changed, undermining KV-cache reuse and making heterogeneous hardware fleets harder to manage predictably.

The replacement is an explicit control-plane/data-plane design. A global control plane combines demand, network overhead, engine health and capacity, and offline latency regressions for TTFT and TBOT to compute per-frontend routing weights that minimize expected end-to-end latency under capacity constraints. Frontend data planes pull and cache these policy snapshots, make synchronous routing decisions locally, and apply live local guardrails without waiting for the control plane.

The operational lesson is that optimization alone is insufficient near production limits. The system penalizes outlier engines, uses dynamic retry budgets to prevent retry storms under high utilization, and proactively sheds load as a final mechanism for graceful degradation when demand exceeds available inference capacity.

Key Takeaways

  • Claim: LLM inference routing must optimize more than request distribution because engine selection affects both latency and expensive GPU efficiency. | Evidence: The ILB considers TTFT, TBOT/token throughput, health and utilization signals, geographic locality, and KV-cache reuse; returning a conversation follow-up to an engine with cached context avoids recomputation and reduces latency. | Implication: A routing layer for agent or model-serving systems should treat session affinity and cache locality as first-class inputs rather than relying on generic load-balancer algorithms.
  • Claim: A periodic feedback loop that converts aggregate engine performance into routing weights is adaptive but too opaque and unstable for fine-grained production control. | Evidence: OpenAI's earlier design filtered infeasible engines and used weighted consistent hashing; a controller smoothed engine signals, compared performance against the fleet average, and raised or lowered weights. Traffic shifts could cool an engine, cause the controller to send traffic back, and create oscillation that disrupted KV-cache utilization. | Implication: Favor explicit objectives, constraints, and inspectable policy outputs over a single blended health score when routing decisions must be debuggable and safely tunable. | Caveat: The speakers say the feedback-loop design did self-balance constrained versus less-constrained traffic and reduced the need for manual intervention; the problem is not feedback itself, but allowing a coupled, implicit loop to be the primary policy.
  • Claim: The scalable pattern is a globally informed control plane paired with a fast, independent data plane. | Evidence: The data plane's engine selector reads asynchronously refreshed local candidate-engine and routing-weight state, while the control plane continuously computes and publishes future snapshots. Only request routing is synchronous; engine-signal collection and policy distribution are asynchronous. | Implication: Do not put a central optimizer on the request critical path. Distribute versioned or cached routing policy to local gateways so centralized optimization failure or latency does not block inference requests.
  • Claim: Nearest-engine routing can be slower end-to-end than cross-region routing when local GPU capacity is saturated. | Evidence: In the example, cluster B has 120 RPS of demand while nearby engine B can serve 100 RPS; engine C has 40 RPS of spare capacity. Sending B's excess 20 RPS to C adds network latency but avoids waiting behind an overloaded local engine. | Implication: Route against predicted total latency—network plus queueing/engine service time—not geographic distance or static regional affinity alone. | Caveat: This conclusion depends on measured network latency and realistic engine-side latency-versus-load profiles; distance alone is not enough to select the destination.
  • Claim: The routing optimizer should explicitly allocate traffic fractions subject to demand and capacity constraints. | Evidence: Its inputs are demand from each CPU cluster, latency to each engine, available capacity and health, and offline regressions of capacity, TTFT, and TBOT as load increases. It outputs the fraction of each frontend's traffic sent to each GPU engine, minimizing expected end-to-end latency while routing all demand, staying within effective capacity, and keeping weights non-negative. | Implication: For a production implementation, make latency curves and capacity estimates explicit operational artifacts, then validate optimizer recommendations against stale telemetry, forecast error, and abrupt failures. | Caveat: The talk does not specify the optimization method, update cadence, cache-affinity formulation, or how uncertainty and sudden demand changes are modeled.
  • Claim: Production inference routing needs explicit overload protections because normal reliability mechanisms can amplify incidents. | Evidence: Outlier engines receive reduced routing weight; retries are bounded by dynamic caps or budgets because extra retries near saturation can create a failure-feedback loop ('retry storm'); load shedding is used when demand exceeds capacity to degrade gracefully. | Implication: Tie retry allowance to real-time utilization or saturation state, define an explicit shedding policy, and ensure operators can remove degraded capacity before it contaminates fleet-wide routing decisions. | Caveat: Load shedding is presented as a last resort, not a substitute for capacity planning or accurate admission control.

Detailed Brief

Three operational paths and ownership boundaries

  • Claims: The request path, engine-signal path, and routing-weight path have different latency and reliability requirements.; Both control and data planes consume engine telemetry, but for different purposes: global policy computation versus immediate local protection.
  • Evidence: The request path runs from frontend CPU cluster to its local data-plane selector and then to the chosen GPU engine.; Signals named include ready-replica count, engine health, TTFT, and TBOT.; The control plane publishes weights and each data plane pulls them into a local cache.
  • Caveats: The transcript provides no staleness bounds, rollout/rollback mechanism for policy snapshots, or consistency guarantees across frontend clusters.
  • Implications: Policy-distribution observability should distinguish stale global policy from a local engine-health event.; The control-plane contract should be narrow and auditable: candidate eligibility and traffic-allocation weights, with local safety overrides.

What the talk does not substantiate

  • Claims: Although the opening promises a concrete case study on reducing global network overhead, the supplied transcript contains no measured before/after result or detailed rollout evidence.
  • Evidence: The only numerical example is illustrative: 90 RPS into 100-RPS local capacity, 120 RPS into 100-RPS local capacity, and 40 RPS spare on an 80-RPS remote engine.
  • Caveats: The design should be treated as a high-level architecture pattern rather than evidence for a specific expected latency reduction, cost saving, or optimizer performance level.
  • Implications: Require internal benchmarks before adopting equivalent routing logic: end-to-end tail latency, cache-hit impact, cross-region traffic cost, optimizer error, and behavior during capacity loss.

Notable Concepts & Terms

  • Inference Load Balancer (ILB): The layer between frontend CPU clusters and GPU inference engines that chooses an engine for each inference request.
  • TTFT: Time to first token; a core user-visible latency signal used to model engine performance.
  • TBOT: Time between output tokens, also described as token throughput; captures generation responsiveness after the first token.
  • KV cache locality: Keeping follow-up turns on an engine that already holds useful conversation context, avoiding recomputation and improving latency and efficiency.
  • Weighted consistent hashing: The prior engine-selection method, using per-engine weights derived from feedback signals while preserving some routing affinity.
  • Control plane / data plane: The control plane computes globally optimized routing policy asynchronously; the data plane makes low-latency local decisions from cached policy and live guardrails.
  • Effective capacity: The usable engine capacity enforced as an optimizer constraint, rather than a nominal hardware throughput figure.
  • Retry storm: A positive feedback failure mode in which retries add load to an already saturated system, producing more failures and still more retries.

Operator Notes / Why Ken Should Care

  • Define routing policy as an explicit optimization problem with named objectives and hard constraints; avoid a single unexplained composite health score.
  • Keep the centralized policy solver off the request path and give every gateway a locally cached policy plus a degraded-mode default.
  • Instrument and test cache-affinity loss as a routing-regression metric alongside TTFT, TBOT, utilization, and end-to-end latency.
  • Implement utilization-sensitive retry budgets before scaling retries during incidents; test for retry-driven saturation in load drills.
  • Establish measurable admission-control/load-shedding rules and a capacity-removal workflow for outlier GPU nodes or clusters.
  • Before investing in a global optimizer, benchmark whether cross-region dispatch improves tail latency after including network delay, queueing, cache loss, and traffic cost.

Source/Metadata

  • Title: Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI
  • Transcript words: 3995
  • Duration seconds: 1092
  • Timestamp note: No timestamps or chapters were provided; the latter portion of the transcript substantially repeats earlier content.

Transcript

2443 words en Processed in 93.0s

Hi everyone, thanks for joining our talk. I'm Lu and this is my colleague Chen Ru. So today we're going to talk about routing LLM inference in production. Specifically how our system evolved from routing based on feedback loops driven by engine signals to a more explicit and predictable policy, which is still informed by engine signals. However, the way we use it is different. So for the agenda today, we're going to begin by introducing the inference load balancer, what it is, what it does, and how it has evolved. And then Chen Ru will walk us through the newer control plane and data plane driven architecture, what are the responsibilities of each, and followed by a concrete case study of how we reduce the global network overhead. And in the end, I return to discuss the protection mechanisms that help keep the system stable under production level stress. So to begin with, what is the inference load balancer? And where does it sit? So this is a very high level diagram of the system we are talking about. On the left hand side are the front end clusters. Those are the CPU clusters that act as gateways into our system. And they receive user requests, then prepare them into the inference request that can be processed by the inference engines. And on the right hand side are the engine clusters, which are usually GPU clusters, and each hosting multiple inference engines. So that's why they got the name of engine clusters. And as you may have already heard, nowadays the GPUs are popular and expensive. So sitting in the middle, it is the ILB or inference load balancer. It actually runs on the front end clusters, but it's also a bridge into our inference stack. It has two main responsibilities: select an engine and approximate the request. For this talk, we are going to focus on the engine selection parts. So in some ways, ILB resembles a very traditional load balancer, because a request usually targets a model, and a model is backed by multiple engines. They may live on different clusters in different regions, or even across continents, because that gives us good resiliency towards localized degradation or cluster failures. However, the inference stack, or the uniqueness of inference, introduces a lot of nuances. Like it has to consider a bunch of signals reported in real time, like the well-known time-to-first token, TTFT, time between output tokens, also known as token throughput, or time between tokens, and other healthiness and utilization signals. Besides, there's an important concept of a KV cache, which is also well-known. But for example, when the conversation already has a lot of useful context cached in one engine, sending the follow-up terms of the same conversation back to the same engine, we are avoiding recomputation, improve efficiency, and reduce latency. So the combination of performance, reliability, locality, cache awareness is what makes it such an interesting problem. So how we attempted in the problem, let's take a look at the early days. And to be honest, early days in this industry sounds a lot more historic than it really is. And the routing process at that time began with a filter of each request may not be served by all the engines because of constraints such as capabilities, or geo-restrictions due to compute or data residency. And among the remaining engines, ILB used a weighted consistent hashing to select the best destination engine for a request for a certain user. Then the important question becomes, where do the weights come from? So they were generated by a periodic feedback loop. The inference engines, as mentioned earlier, report all kinds of the signals we care about. And the controller will periodically smooth out those signals and compute performance score. The performance score then will be compared against the fleet average. Then the weight will be adjusted for each engine, if the weight goes up, if the performance is better, or it goes down when the performance is worse than the fleet average. And this generated weight will impact the routing, and then it's basically a control loop. Conceptually, it's very similar to the PID controller. And no, this PID controller will not help you cure a Linux process, but instead it's a classic control theory technique that continuously steers the system towards its desired state. And we just borrowed this important concept, the proportional part of it, and applied it into our load balancer. So it has a lot of nice properties. For example, it could combine the useful signals we care about into the single routing decision. And because it adapts to the observed performance, as what we mentioned earlier, there's a lot of constraints, and those constraints might have some engines busier because they can serve more requests, more kinds of requests than the remaining. But those busier signals will be fit into the next loop, and resulting in the less constrained request that can go to more of those kinds of engines. So basically, they're self-balanced out. And to some extent, this just means we don't need to do a lot of manual intervention, and it just works. However, that kind of adaptability comes with big trade-offs. Because of the same reason that it combines so many signals, it's also very hard to reason about a particular routing decision, or why some engine gets a higher weight than we expect. And every time we want to fine-tune towards some aspect, it's almost impossible to not impact something else. And the load is not always very evenly distributed, because sometimes a model is served by engines on different GPU SKUs, and they have different characteristics. Then the problem becomes a lot trickier. And the feedback loop sometimes creates bad oscillations, because when you shift an engine away some traffic, the engine turns a bit cooler, and this signal gets fit to the controller. The controller now thinks, hey, this engine can take a lot more traffic. Then some traffic is going to be shifted back and forth between a few engines, and disrupting the KV cache utilization. So all those limitations motivated us to rethink about the architecture, and see if we have new ways to address the problem. So I'm going to hand over to Chen Ru to deep dive into the new architecture we tried out. Yeah, thank you, Lu. So I'm going to talk about the architecture of the load balancer, and how do we reduce the overall overhead with our routing algorithm. So the load balancer answers one question: for each request from a CPU cluster, which engine should serve it? A most naive baseline might be round robin, which sends requests across engines evenly. But if you think a little bit more, that doesn't make sense. Because engines are not homogeneous, they can have different hardware and capacity, different health, and also different distance from CPU cluster. Also, round robin could break cache locality. Related requests that could reuse the same engine cache might be sent to different engines. A probably better solution might be for each CPU cluster, it choose the best engine from its own local view. But that's not enough either. Think about one extreme case. Multiple CPU clusters route traffic to the same engines independently, which could overload that engine while leaving other engines underutilized. So what we need is a globally optimized solution. A control plane that has a global view for all the CPU clusters and GPU engines, and could compute a globally optimized routing answer. And the data plane can make a routing decision quickly based on the answer pulled from the control plane. Now, let's look inside the control plane and data plane. In the data plane, there is an engine selector, which selects engine for each request. It reads the local routing state, which includes the candidate engines and the routing weights for each candidate engine. Both of them are refreshed asynchronously in the background. So we don't need to ask the control plane before we make a routing decision for each request. Also, the data plane collects real-time engine signal, such as number of ready replicas, engine health, etc., to service fast local guardrails. In the control plane, the data loader combines those live engine signals and network overhead. And with offline regressions of capacity, TTFT and TBOT, the optimizer could turn those data into routing weights. And the control plane will publish the routing weights for each data plane to pull. In this way, no request needs to wait on the data plane. The control plane continuously computes the next globally optimized routing weights snapshot, while the data plane makes a routing decision based on the latest snapshot already installed locally. In summary, there are three important paths through the system. The first path is the inference request path. The request arrives to the CPU cluster and the data plane inside the CPU cluster will select engine for that request based on the local routing state and forward the request to the selected engines. The second path is the engine signal path. The system continuously collects real-time engine signal, such as TTFT, TBOT, number of ready replicas, and engine health, etc. Both planes need those real-time engine signals. The control plane needs them to compute a globally optimized routing weights, while the data plane needs them to service fast local guardrails. And the third path is the routing weights path. The control plane computes and publishes the routing weights, and the data plane pulls the updates to its local cache. So only the first path is synchronous, but it's fast and only local inside the data plane of the CPU cluster. The other two loops are asynchronous loops, and they are to improve future routing decisions. So that's pretty much of the architecture part, but that still leaves one question: how do we compute those routing weights? But before answering that question, let's answer another question first. Why not just send a request to the nearest engine? That's because the traffic demand and GPU capacity are not geographically balanced. For example, in region one, CPU cluster A sends 90 RPS, and the nearby engine A can serve 100 RPS. So in this case, nearest engine only is fine. While in region two, CPU cluster B sends 120 RPS, and the nearby engine B could only serve 100 RPS. So in this case, if we insist on keeping everything local, the extra 20 RPS needs to wait on an overloaded engine B. While in region three, we are only using 40 RPS of an 80 RPS engine C. That still leaves 40 RPS spare. So if we send the extra 20 RPS from cluster B to engine C, that will add network distance. So in this case, a further engine C is faster end to end. That's why we need something better than the nearest only routing. Now let's open the black box of the optimizer. The optimizer accepts four types of input. The request from each CPU cluster, the network latency to each engine, the available engine capacity and health, and also the TTFT, TBOT latency profiles that tell us how the engine side latency changes as the load increases. And with those inputs, the optimizer turns the input to the output routing weights. The routing weights say for each CPU cluster, what fraction of its traffic should go to each GPU engine. And the optimization goal is straightforward: it's to minimize the expected end-to-end latency across all routed traffic. The important part is that the end-to-end latency includes both the network distance and the engine side latency. That means a nearby engine might be attractive when it still has room to serve traffic. While a further engine might be better if all the nearby engines are close to full. And the optimizer also needs to respect several hard constraints. First, it needs to route all the traffic demand. Second, it needs to ensure all the engines stay within the effective capacity. Third, it needs to keep the routing weights non-negative. And that's pretty much my part. And Lu will continue to talk about the protection mechanisms in the system. Thanks, Chen Ru. So, as AI engineers, we all know that production in many cases is not behaving in the most ideal case. So, clusters can fail, GPUs or individual nodes can degrade, and networking can just get all kinds of mysterious issues. So, how do we keep our production system healthy as much as possible under heavy load? The first thing we have is penalties. Basically, when an engine is an outlier, we detect the abnormality and try to reduce the routing weight to that engine. In that way, we give it a chance to either recover by themselves if there's some transient issue, or we can have a human intervene to rotate it out or replace the faulty hardware. And secondly, the retries, which is a very common technique used to mitigate problems. However, in some cases, it actually could make them even worse. Like when the system is very close to a tipping point or very heavily utilized. Retries send more load. And this more load causes more failures and causes more retries, which is the infamous retry storm. So, we implemented caps or budgets to constrain retries into an acceptable region. And this actually needs to be dynamic because in the happy time or in the normal time, we can tolerate a lot more retries than when the system is heavily utilized. And finally, we have load shedding, which is our last resort when the production capacity couldn't meet the increasing amount of inference demands. So, we instead will try to have the system fail gracefully. We basically proactively load shed a portion of the traffic to have the system degrade gracefully. So, that pretty much concludes our talk today. And thanks for joining us. Both of us will be around in our booth area this afternoon. So, if you have further questions, feel free to walk to the area and chat with us. Thank you. that can go to more of those kind of engines. So basically, they're self-balanced out. And to some extent, this just means we don't need to do a lot of manual intervention, and it just works. However, that kind of adaptability comes with big trade-offs. Because of the same reason that it combines so many signals, it's also very hard to reason about a particular routing decision, or like why search engine gets a higher weight than we expect. And every time we want to fine-tune towards some aspect, it's almost impossible to not impact something else. And the load is not always very evenly distributed, because sometimes a model is served by engines on different GPU skills, and they have different characteristics. Then the problem becomes a lot more trickier. And the feedback loop sometimes creates bad oscillations, because when you shift an engine away some traffic, the engine turns a bit cooler, and this signal gets fit to the controller. The controller now thinks, hey, this engine can take a lot more traffic. Then some traffic is going to be shifted back and forth between a few engines, and disrupting the Kiwi cache utilization. So all those limitations motivated us to rethink about the architecture, and see if we have new ways to address the problem. So I'm going to hand over to Qian Ru to deep dive into the new architecture we tried out. Yeah, thank you, Lu. So I'm going to talk about the architecture of the load balancer, and how do we reduce the overall overhead with our routing algorithm. So the load balancer answers one question. For each request from a CPU cluster, which engine should serve it? Well, most naive baseline might be round robin, which send requests across engine evenly. But if you think a little bit more, that doesn't make sense. Because engines are not homogeneous, they can have different hardware and capacity, different health, and also different distance from CPU cluster. Also, round robin could break cache locality. Related requests that could reuse the same engine cache might be sent to different engines. A probably better solution might be for each CPU cluster, it choose the best engine from its own local view. But that's not enough either. Think about one extreme case. Multiple CPU cluster route traffic to the same engines independently, which could overload that engine while leaving other engines underutilized. So what we need is a globally optimized solution. A control plan that has a global view for all the CPU cluster and GPU engines, and could compute a globally optimized routing answer. And the data plane can make a routing decision quickly based on the answer pulled from the control plane. Now, let's look inside the control plane and data plane. In the data plane, there is an engine selector, which selects engine for each request. It reads the local routing state, which includes the candidate engines and the routing weights for each candidate engine. Both of them are refreshed asynchronously in the background. So we don't need to ask the control plane before we make a routing decision for each request. Also, the data plane collects real-time engine signal, such as number of ready replica, engine house, etc., to service fast local guardrails. In the control plane, the data loader combines those live engine signals and never overhead. And with offline regressions of capacity, TDFT and TBOT, the optimizer could turn those data into routing weights. And the control plane will publish the routing way for each data plane to pull. the control plane is a control. In this way, no request need to wait on the data plane. The control plane continuously computes the next globally optimized routing way snapshot, while the data plane makes a routing decision based on the latest snapshot already installed locally. In summary, there are three important paths through the system. The first path is the inference request path. The request arrives to the CPU cluster and the data plane inside the CPU cluster will select engine for that request based on the local routing state and forward the request to the selected engines. The second path is the engine signal path. The system continuously collects real-time engine signal, such as TTFT, TBOT, number of ready replica, and engine house, etc. Both planes need those real-time engine signals. The control plane need them to compute a globally optimized routing way, while the data plane need them to service fast local guardrails. And the third path is the routing way path. The control plane computes and publishes the routing way, and the data plane for the updates to its local cache. So only the first path is synchronous, but it's fast and only local inside the data plane of the CPU cluster. The other two loops are asynchronous loop, and they are to improve future routing decision. So that's pretty much of the architecture part, but that still leaves one question. How do we compute those routing weights? But before answering that question, let's answer another question first. Why not just send a request to the nearest engine? That's because the traffic demand and GPU capacity are not geographically balanced. For example, in region one, CPU cluster A sends 90 RPS, and the nearby engine A can serve 100 RPS. So in this case, nearest release only is fine. While in region two, CPU cluster B sends 120 RPS, and the nearby engine B could only serve 100 RPS. So in this case, if we insist on keeping everything local, the extra 20 RPS needs to wait on an overloaded engine B. While in region three, we are only using 40 RPS of an 80 RPS engine C. That still leaves 40 RPS spare. So if we send the extra 20 RPS from cluster B to engine C, that will add network distance. So in this case, a further engine C is a faster end to end. That's why we need something better than the nearest only routing. Now let's open the black box of the optimizer. Let's open the optimizer. The optimizer accepts four types of input. The request from each CPU cluster, the network latency to each engine, the available engine capacity and health, and also the TTFT, TBOT latency profiles. That tell us how the engine side latency change as the low increases. And with those inputs, the optimizer, turn the input to the output routing weights. The routing weights say for each CPU cluster, what fraction of its traffic should go to each GPU engine. And the optimization goal is straightforward. It's to minimize the expected end-to-end latency across all routed traffic. The important part is that the end-to-end latency includes both the network distance and the engine side latency. That means a nearby engine might be attractive when it still has room to serve traffic. While a further engine might be better if all the nearby engines are close to full. And the optimizer also need to respect several hard constraints. First, it needs to route all the traffic demand. Second, it needs to ensure all the engines stay within the effective capacity. Third, it needs to keep the routing weights non-negative. Third, it needs to ensure that the control plane is the routing weights and the optimizer. And the data plane pulls them and uses them to make a globally optimized routing decision. And that's pretty much my part. And Lu will continue to talk about the protection mechanisms in the system. Lu Wang Huang Thanks, Chenru. Lu Wang Huang So, as AI engineers, we all kind of know that production in many cases are not behaving in the most ideal case. So, clusters can fill, GPUs or individual nodes can degrade, and networking can just get to all kind of mysterious issues. So, how do we keep our production system healthy as much as possible under the heavy load? Lu Wang Huang The first thing we have is penalties. Basically, when an engine is an outlier, we detect the abnormally and try to reduce the routing weight to that engine. In that way, we give it a chance to either recover by themselves if there's some transient issue, or we can have a human intervene to rotate it out or replace the faulty hardware. Lu Wang Huang And secondly, the retries, which is a very common technique used to mitigate problems. However, during some cases, it actually could make them even worse. Like when the system is very close to like a tip over or very heavily utilized. Retries, we are sending more load. And this more load, we are causing more failures and causing more retries, which is the infamous retries storm. So, we implemented caps or budgets to constantly retries into an acceptable region. And this is actually even need to be dynamic because in the happy time or in the normal time, we can tolerate a lot more retries than when the system are heavily utilized. And finally, we have the load shedding, which is our last result when the production capacity couldn't meet the increasing amount of inference demands. So, we instead will try to have all the system fill. We basically proactively load shed a portion of the traffic to have the system degrade gracefully. So, that pretty much concludes our talk today. And thanks for joining us. Both of us will be around in our booth area this afternoon. So, if you have further questions, feel free to walk to the area and chat with us. Thank you. Thank you.