Good morning, everyone.
Welcome to the first Influence Talk on the last day of AI Engineering World Fair. My name is Nishan Gupta, and I'm joined today by my co-speaker Naman Ahuja. We work on building the efficiency, training, and inference infrastructure at Meta. Today, we're going to be talking about how do you operate distributed inference systems at scale. As we all know, inference is no longer just a research artifact boiled to a product. It's a foundational hyperscale infrastructure workload, which is growing at a tremendous rate. The inference traffic already outpaces the largest microservices in the world, and the rate of growth is fastest of any workload we have ever seen.
Let's rewind back to around 2008 and try to compare the AI era with the cloud era.
In around 2008, the cloud started as virtual machine offerings. The interesting engineering was virtualization. Then, over time, the value moved up to the stack, to the schedulers, Kubernetes, Mesos, then to service meshes, then to autoscalers, and then to various platforms that were built on top of it. The orchestration layer is what actually captured the value and the complexity. AI is on the exact same trajectory, but just compressed into the last few years instead of a decade. We started with simple models running on GPUs. Then we saw the evolution of model-serving frameworks like VLLM,
Torchserve, Triton, and now we are watching the orchestration layer emerging in real time, addressing the complex challenges of routing, KV cache management, prefill decode disaggregation, and multi-model multiplexing. In this next phase of AI era, it's not just about the best models or the kernels or the optimizations. It's about the whole ecosystem. It's about the control plane and the orchestration. And this is what we're going to be focusing in this race talk. So let's talk a little bit about agentic demand explosion.
In the classical web-serving before the AI inference workloads kicked off, the capacity scaled roughly linearly with users, depending upon the workload type. Double the users, mostly double the QPS, double the infrastructure fleet, if there are no optimizations, and capacity planning was mostly a spreadsheet exercise. In this new agentic serving, capacity is scaling with number of users times number of calls per user times number of tokens, which varies depending on the model, the optimizations, the hardware SKU you have. A chatbot can have one model call per turn, which is evolving now with 10 to 20 for co-pilots,
and 50 for research agents, and now it's thousands of such calls for autonomous workloads, with no human in the loop. The key takeaway is that you cannot plan capacity for the agents the same way we did for microservices. We need to think about elasticity and implement workload-aware scheduling and admission control. Now let's try to dive a little bit deep into the differences, pros and cons, differences between the traditional microservice serving versus the modern inference serving across these key dimensions. The request shape. Microservices assume short, uniform requests, where the LLM requests we are seeing in our workloads,
they can vary from 50 tokens to 100,000 tokens, with vastly different compute profiles between prefill and decode stage. For batching, classical stacks for microservices was doing batching, batching mostly at the load-balance layer, if at all. However, LLM serving requires continuous in-flight batching, otherwise the throughput collapses by an order of magnitude or more. State. Most of the classical microservices were stateless, when we're not talking about storage layer. However, LLM serving does require a huge per-request state, the KV Cache, which is very expensive to build, and even more expensive to throw away. Scaling units.
When we think about traditional microservices, we could run them in cheap CPUs, in pods. However, for modern inferencing, we require to run on GPUs, which are 100 times more expensive, which are 10 times slower to acquire, and we cannot over-provision them casually. Otherwise, it will lead to a huge wastage. Failure mode. When you think about traditional microservices, like most of us have built over the last couple of years, we could, even if a host crashed or a pod crashed, we could restart it. We could rebuild the state if need be. However, for the model inferencing, it takes a huge amount of time to go from cold to hot startup. And if a GPU is mid-decode,
it can drop thousands of in-flight tokens, and it can lead to a queue buildup. The takeaway is that the bottleneck is not just the model, it's the orchestration itself. Now, as we can see in these hidden decisions behind a prompt, when we go to any agentic application, it requires a bunch of steps which are behind the scenes. We have to authenticate. We have to choose a model depending on the request type. We have to select the region where it goes. We have to do admission control. We have to do caching lookup. We have to run it on the GPU. We have to do batching. And there's a bunch of other steps involved. And as we can see, out of all these steps,
only one step requires the model, which is the prefill decode, if it's disaggregated inference or just if it's not disaggregated inference. Other steps require infrastructure. The intelligence might lie in the model, but the economics, the reliability, and the user experience are all in the infrastructure. And this is why a lot of platform teams across a lot of companies are having much more impact on the product quality and the success, much more than before. Most of us in this room have a deep expertise across one or two or three layers.
We might own kernels or kernel optimizations. We might own routing or the product itself, or we might be operating the GPU infrastructure or the cluster itself. But very few of us have operated the whole stack or seen, thought about it end to end. As you can see in these layers, these layers are not new. They have been around for 20 years or more. What is new is the combination and the coupling between them. A decision at the routing layer can change the cache hit rate at the model layer, which can change the batch composition, which can change the GPU utilization, which can change the autoscaling decision because of the change in GPS simplization.
So everything is entangled. When we have any regressions in our inference workloads, it's not just about understanding what happened at the caching layer or admission control. We need to think about the stack top to bottom. And when we see any bottleneck, it's very important to understand at which layer is that bottleneck so that we can invest properly. Now, diving a little bit deep into how a prompt works.
When we have a prompt for any application, be it if you want to generate an image or if you have a research task or if you have complex multi-agent orchestration, more or less, it involves a bunch of these steps. The prompt goes to the gateway. Then it goes to the router, after which it does the cache lookup if the request was already seen before. Then it goes to the schedulers, which decides on which GPU cluster it should run, and on which hardware. It can be on NVIDIA or AMD or your in-house silicon chip, which then goes to the appropriate serving runtime, VLLM, MIG-Lang, or whatever we are working on. And then we stream the response back to the user
according to the SLO profiles of time-to-first token and time between each token, and while making sure the throughput is what the user desired. Now, as we can see, this inference behaves like a distributed transaction. Each arrow in this diagram is a network hop. Every one of these hops can retry. It can time out. It can fall back. It can even fail. And each of these hops will have SLOs, and it is streaming back to the user. So if it fails, the partial philis mantix are much harder to deal with than for a regular RPC call. Think about what happens if we have already streamed 200 tokens back to the user, and suddenly a GPU host is preempted
due to a scheduled or a planned or unplanned maintenance event. We cannot just retry. We have to think about it holistically. This is why reliability, we cannot build reliability at the edge. It has to be a property of the control plane, because the control plane is the one which sees the whole workflow. Now, let's talk about schedulers and some of the optimizations and how we can think about it. For traditional microservices, we used to think about bin packing of traditional microservices across three, four dimensions. Could be across the virtualization, CPU memory, or could be across four domains, depending upon if you're using AWS,
if you're using your own in-house cloud providers. But for inference, the scheduler needs to be aware across at least seven axes. When we schedule a particular request, it has to be aware of the GPU type. There can be n number of heterogeneous hardwares in your cluster, H100 versus E100 versus B200 with different network topologies. It has to be aware of the HBM headroom, KV cache state. The model weights, whether they're already loaded, whether they're cold or we have to, whether they've already warmed up, we have to cold start it. It has to be aware of the tenant priority. and some of the optimizations and how we can think about it.
So for the traditional microservices, we used to think about pin packing of traditional microservices across three, four dimensions. Could be across the socialization, CPU memory, or could be across four domains, depending upon if you're using AWS, if you're using your own in-house cloud providers. But for inference, the scheduler needs to be aware across at least seven axes. When we schedule a particular request, it has to be aware of the GPU type. There can be n number of heterogeneous hardwares in your cluster, H100 versus E100 versus B200 with different network topologies. It has to be aware of the HBM headroom, KV cache state.
The model weights, whether they're already loaded, whether they're cold or we have to, whether they've already warmed up, we have to cold start it. It has to be aware of the tenant priority. There can be n number of tenants running on that multi-tenant cluster with different SLO profiles. We have to also be aware of the workflow context. Are we in the third step of reasoning that has already spent X dollar? Or are we at the initial stages and we can terminate the workflow if we are over-provisioned? We have to also think about latency budget, depending upon the type of agentic application we are building. So this brings us to the idea that we have to make sure
that we implement agentic-aware scheduling. We have to make sure that we have placed the work which will finish in the fastest and the cheapest time, as opposed to just placing it on a random GPU. A concrete example might be that the scheduler needs to be aware that request R is at step three of a five-step workflow, and step one and two has already spent X plus Y dollar. So if a step three fails, the whole workflow will be terminated and we have wasted all that compute resources. That's why workflow-aware orchestration is very, very important, because it will change the admission decisions, the priority, and how we retry. Now let's talk about optimizations.
I will not go too deep into a lot of optimizations. There's a lot of research that has already been done outside, but I would like to share a framework, which at least I like to use when it comes to it, and we can place it into four quadrants. First, can we avoid the work? Meaning, can we skip it entirely through caching, through techniques like prefix caching, response caching, semantic caching? The second, can we share the work? Can multiple requests share compute through batching?
Think continuous batching, prefilled decode, chunk prefilled, speculative decoding. Third, can we move away the work? Can we send it somewhere to a cheaper model or closer to the user? Through mostly routing, can we route it to a smaller model or a cheaper region or some other techniques? And lastly, can we delay the work? Can we wait for a better moment through admission control and queuing, which requires us to understand the priority classes of these requests and implement deadline-aware scheduling. Now this framework is very powerful because it transfers to various stacks. You might be using vLLM or SG-Lang or TensorRT,
but every technique fits into one of these quadrants. So whenever we think about any optimization to our model, we have to apply it.
We have to do a comparison and contrast with the previous techniques and see how all these stack with each other. Now whenever we think about scale, it's not just important to think about the performance of the model. We have to think about the cost, the economics as well. This is where it's important to understand what metric we are trying to optimize. Because the cost is not just the cost of the GPU or the model. It is the cost of all these parameters, retries, storage, failures, network, and of course the operational cost of development and all that stuff. What is important is to understand what is the key performance indicator
for your product, which will add value to the users. So it's not important to optimize just cost per token or cost per request. We have to optimize cost per successful task because this is what actually users care about. And if we're able to optimize that, the cost for the overall product decreases and the users are much more happier. Now let's talk about reliability, a little bit about reliability and what it means to prevent cascading failures. Now the failure story is never a GPU is preempted or a GPU has died.
The interesting story is the feedback loop that follows. So a GPU can degrade, the latency can rise, the client retries, the queue depth increases, the healthy GPUs will saturate, which will follow more retries, much more full regional failures. This is the classic cascading failures, but with a twist for agentic application, the KV cache. We cannot just casually restart or reroute to a different cluster. A cold pool has to warm up before it can absorb traffic, during which the hot pool has to take on all those requests. This is why it's important to design the loop breakers very deliberately. The circuit breakers at the routing layer, the admission control,
which is rather than just queuing, the load shedding tied to queue depth, not just CPU or memory utilization. And we have to also think about retry budgets, because if we don't think about all these things, the cost can scale much, much more quickly. Now I'll hand it over to my co-speaker, and I want to talk about the remaining talk. So once inference reaches production scale, it starts looking much more like a distributed system. We are no longer just calling a model. It's more like a classic distributed system problem. So in distributed system, we talk about queues, scheduling, auto-scaling, fault isolation. These are some of the dimensions.
Inference has all these problems, but there are new constraints now. Instead of CPU memory alone, we have CPU, HBM, KVCache, and cost per successful task. So the operating question becomes, how does the platform know what to do next? And that's where observability comes into play. It's not just about dashboards. It's about how to provide input signal to the control loop. Telemetry, fields, analysis, analysis drive decisions, issues and changes of scheduling and routing, and finally we just rinse and repeat. I'll just give an overview of what are some important metrics.
First one is time to first token, which tells me about how much time does it really take to get the first response. Then we have utilization ratio, which tells me whether memory or compute is the bottleneck. We have success per dollar that tells us whether the platform is actually delivering and working efficiently. And finally, we have end-to-end request latency that tells us what's the time being spent across the full request path. There's a trade-off between latency, cost, and throughput. You cannot just get all of them. It's pretty analogous to cap theorem. If I increase the batch size, I improve the throughput and cost efficiency, but I may hurt tail latency.
If I use speculative decoding, I may improve latency, but there are some extra compute. Ultimately, we increase the cost per token. And finally, I can use a simple model, smaller model. I can reduce the latency and cost, but the response will be of low quality. Ultimately, I'll do failure analysis, do retries, which brings the cost back up. So every serving decision moves systems somewhere in this triangle. And our job is to find the perfect setting. It's just an optimization problem now. This is where the industry is heading right now. Inference needs its own control plane. Everything we discussed, routing, batching, caching, scheduling, reliability,
they cannot be a separate knob now. They're converging into a logical layer. Let's call it inference control plane. We used to manage VMs before in distributed system. We have auto-scaling schedulers. And we had Kubernetes, which turned these into control plane.
Inference is going through the same transition now. Models are becoming resources. GPU, KVCache, token, latency, cost are now scheduled around. The control plane decides which models serve which request and how it is batched. So whether we build this layer internally, use open source, or from a vendor, the key design is to assume layer will exist. Now let's discuss what are some of the operating lessons we have experienced in A-infra, and how they are applicable here. The first lesson is that infrastructure bottlenecks usually show up before model bottlenecks. In production, many failures can come, but they can be just about scheduling and routing, or capacity breakdowns.
So these are not related to inference. It's about infrastructure problems. Then we have elasticity. We need elasticity in the system. We can have more GPUs, but this will not really solve the problem. We are just hiding the problem. Then our solution to schedule the decision, overpowered efficiency. Models are becoming resources. GPU, KVCache, token, latency, cost are now scheduled around. The control plane decides which models serve which request and how it is batched. So whether we build this layer internally, use open source, or from a vendor, the key design is to assume the layer will exist.
Now let's discuss what are some of the operating lessons we have experienced in A-infra and how they are applicable here. The first lesson is that infrastructure bottlenecks usually show up before model bottlenecks. In production, many failures can come, but they can be just about scheduling and routing or capacity breakdowns. So these are not related to inference. It's about infrastructure problems. Then we have elasticity. We need elasticity in the system. We can have more GPUs, but this will not really solve the problem. We are just hiding the problem. Then our solution to schedule the decision overpowered efficiency.
The same fleet can deliver very optimally, depending on how you're scheduling it or how you're batching it. Then we have control loops, which beats the manual process. The platform has to sense, detect, and automatically adapt to the system. So main takeaways: do not optimize for tokens. Optimize for successful tasks. So this is a broader shift I want to leave you with. The first phase of AI infrastructure was about better models. We invested a lot of time in improving our models, making the models smarter, and coming with better benchmarks. The current phase right now is faster inference, lower latency, better batching, and better GPU utilization.
But the next phase is about orchestration. That means GPU, memory, cache, and everything—these are just resources, and they need to be scheduled and controlled. The teams that understand this early on will build infrastructure for the future.