AI Engineer

Operating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta

1930 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: At production scale, LLM inference is fundamentally a distributed-systems orchestration problem: the durable advantage comes from an inference control plane that jointly manages routing, caching, scheduling, batching, admission, reliability, and cost.
  • Why it matters: Agentic workloads multiply calls, tokens, state, and failure exposure far beyond conventional microservices, so capacity planning and per-model optimization alone will produce poor economics and cascading reliability failures.
  • Best use: Use this as an architecture and operating-model primer for designing an agent/inference control plane, defining the right metrics, and prioritizing reliability mechanisms before scaling GPU capacity.

Executive Summary

Meta's Nishant Gupta and Naman Ahuja argue that AI infrastructure is repeating the cloud transition from raw compute to orchestration, but on a compressed timeline. Early AI value accrued to models, kernels, and serving runtimes such as vLLM, TorchServe, and Triton; the next source of complexity and value is a control plane that makes coordinated decisions across model routing, KV-cache placement, prefill/decode execution, batching, and hardware allocation.

Their central operational claim is that agentic demand cannot be planned like a conventional service. Instead of scaling roughly with users and QPS, capacity depends on users multiplied by calls per user multiplied by highly variable token volume. A chatbot may make one model call per turn, copilots may make 10–20, research agents about 50, and autonomous workflows potentially thousands. This requires workload-aware scheduling, elasticity, and admission control rather than spreadsheet-based fleet planning.

The speakers frame each inference request as a distributed transaction with streaming semantics. A prompt traverses authentication, routing, cache lookup, scheduling, hardware selection, a serving runtime, and response streaming; every hop can fail, retry, or time out. Because the request carries expensive KV-cache state and may already have streamed output, reliability cannot be bolted onto an edge service. The control plane must make workflow-aware retry, priority, routing, and load-shedding decisions end to end.

The practical standard they advocate is to optimize for cost per successful task, not cost per token or request. This requires observability that feeds automated control loops, explicitly balancing latency, throughput, quality, and cost. Their message to infrastructure teams is that adding GPUs often masks scheduling failures; a well-designed orchestration layer can extract materially more useful work from the same fleet.

Key Takeaways

  • Claim: Inference infrastructure is moving toward a distinct control plane, analogous to how cloud infrastructure evolved from VMs into schedulers, Kubernetes, service meshes, and autoscalers. | Evidence: The speakers trace cloud's value shift from virtualization to orchestration, then map AI's compressed evolution from GPU-hosted models to serving frameworks including vLLM, TorchServe, and Triton, followed by emerging systems for routing, KV-cache management, prefill/decode disaggregation, and multi-model multiplexing. | Implication: Treat serving runtimes and individual model endpoints as data-plane components; invest in a policy/control layer that can coordinate cross-cutting resource and request decisions.
  • Claim: Agentic systems require a new capacity model because demand grows as users × calls per user × tokens, rather than roughly linearly with user count or QPS. | Evidence: The talk contrasts a one-call chatbot with copilots making 10–20 calls per turn, research agents making roughly 50, and autonomous workflows making thousands of calls without a human in the loop; token lengths can range from 50 to 100,000. | Implication: Capacity systems should forecast and gate work by workflow type, token demand, deadline, and priority, rather than treating all requests as interchangeable QPS. | Caveat: Actual demand depends on model choice, token patterns, hardware SKU, and serving optimizations, so fixed call-count assumptions are insufficient.
  • Claim: LLM serving differs materially from classic microservices because it has variable request shapes, continuous batching needs, costly per-request state, expensive GPUs, and disruptive cold starts. | Evidence: The speakers note that throughput can collapse by an order of magnitude or more without continuous in-flight batching; KV cache is expensive to construct and discard; GPUs are described as roughly 100 times more expensive and 10 times slower to acquire than CPU capacity; a mid-decode GPU loss can drop thousands of in-flight tokens. | Implication: Do not directly port stateless pod autoscaling and restart assumptions into inference. Scheduling, failure recovery, and scaling must preserve or account for model weights, cache state, and active decoding work.
  • Claim: Inference reliability must be designed as a control-plane property because an LLM request is a multi-hop, streaming distributed transaction with difficult partial-failure semantics. | Evidence: A request may pass through a gateway, router, cache, scheduler, heterogeneous GPU cluster, serving runtime, and streaming layer. If a host is preempted after 200 tokens have already been sent, a normal retry cannot simply recreate the user-visible interaction. | Implication: Build end-to-end failure policies—rather than isolated service retries—that understand response state, cache locality, workflow value, retry budgets, and available warm capacity. | Caveat: The transcript does not prescribe a universal recovery protocol for partially streamed outputs; the appropriate behavior remains product- and workflow-dependent.
  • Claim: Scheduling should be agent- and workflow-aware, not merely a bin-packing decision over compute and memory. | Evidence: The proposed scheduler considers at least GPU type and topology, HBM headroom, KV-cache state, whether model weights are warm, tenant priority and SLOs, workflow context, and latency budget. Their example is a request at step three of a five-step workflow, where steps one and two have already incurred cost that should influence priority and retry decisions. | Implication: Represent workflow progress and sunk compute in request metadata so the scheduler can prioritize work whose failure would invalidate high-value prior steps, while shedding lower-value work sooner.
  • Claim: Optimization should be evaluated through four levers—avoid, share, move, or delay work—and measured by cost per successful task rather than token-level efficiency alone. | Evidence: Avoid includes prefix, response, and semantic caching; share includes continuous batching, chunked prefill, prefill/decode techniques, and speculative decoding; move includes routing to a smaller model, cheaper region, or closer location; delay includes admission control, queues, priorities, and deadline-aware scheduling. The speakers warn that a smaller model may lower direct cost but reduce quality and trigger retries, restoring total cost. | Implication: Evaluate routing and serving changes against successful end-to-end task completion, including retries and quality failures, instead of local metrics such as cost per token. | Caveat: Latency, throughput, cost, and output quality cannot all be optimized simultaneously: larger batches improve throughput but can worsen tail latency, while speculative decoding can reduce latency at additional compute cost.
  • Claim: Adding GPU capacity is not a substitute for elasticity and automated control loops; infrastructure bottlenecks often appear before model bottlenecks. | Evidence: The speakers identify routing, scheduling, and capacity breakdowns as common production failure sources, and describe a cascade where a degraded GPU raises latency, clients retry, queues deepen, healthy GPUs saturate, and the region can fail. They recommend circuit breakers, queue-depth-based load shedding, admission control, and retry budgets. | Implication: Prioritize telemetry-driven automation that senses degradation and changes routing, scheduling, and admission decisions before expanding fleet size or manually tuning the system.

Detailed Brief

Control-loop observability and operating metrics

  • Claims: Observability is not principally a dashboarding function; it supplies the signals used by the control loop to alter scheduling and routing decisions.; The platform should optimize within an explicit latency, cost, throughput, and quality trade-off space rather than pursue a single utilization metric.
  • Evidence: The speakers describe a closed loop of telemetry, analysis, decisions, scheduling/routing changes, and repetition.; They name time to first token as a responsiveness measure, utilization ratio as an indicator of whether compute or memory is constraining the system, end-to-end request latency as a full-path measure, and success per dollar as an efficiency metric.; Their trade-off examples are that larger batches improve throughput and cost efficiency while harming tail latency, and smaller models reduce latency and direct cost but may degrade quality enough to create failure-analysis and retry expense.
  • Caveats: The talk provides a useful metrics taxonomy but does not specify target thresholds, alert policies, or how to weight user experience against cost for different product classes.
  • Implications: Separate leading indicators for queue pressure, cache/memory pressure, warm-model availability, and retry amplification from lagging business measures such as successful tasks per dollar.; Use product-specific service classes to define acceptable trade-offs rather than applying one global latency or utilization target to interactive and autonomous workloads.

Failure containment in stateful inference

  • Claims: The dangerous event is not the original GPU or host failure but the feedback loop it initiates across retries, queues, and healthy capacity.; Warm capacity is operationally significant because a cold pool cannot immediately replace a failed or overloaded hot pool when weights and caches must be established.
  • Evidence: The proposed failure sequence is GPU degradation, increased latency, client retries, rising queue depth, saturation of healthy GPUs, more retries, and potential regional failure.; Recommended loop breakers are routing-layer circuit breakers, admission control rather than unrestricted queuing, queue-depth-based load shedding, and explicit retry budgets.
  • Caveats: The transcript emphasizes queue depth over CPU or memory utilization for load shedding, but a production policy still needs to account for tenant priority, workflow value, and user-facing streaming semantics.
  • Implications: Make retry traffic a first-class capacity consumer in load tests and incident simulations.; Maintain enough validated warm failover capacity for critical workloads, rather than assuming autoscaling can safely absorb abrupt failures.

Notable Concepts & Terms

  • Inference control plane: The logical orchestration layer that jointly governs model selection, routing, batching, caching, scheduling, reliability, latency, and cost rather than exposing each as an isolated tuning knob.
  • KV cache: Per-request model state accumulated during generation; its cost and locality make rerouting, restart, and failure recovery much harder than for stateless RPCs.
  • Prefill/decode disaggregation: Separating prompt processing from token generation so each phase can be scheduled and optimized independently, reflecting their different compute and latency profiles.
  • Continuous in-flight batching: Dynamically batching active LLM requests during serving; the speakers characterize it as necessary to avoid an order-of-magnitude-or-greater throughput collapse.
  • Workflow-aware orchestration: Scheduling and retry decisions that consider an agent's step within a multi-step workflow and the compute already spent, not only the characteristics of the current model call.
  • Cost per successful task: The preferred economic metric because it incorporates model quality, failures, retries, networking, storage, and operations rather than rewarding cheap but unsuccessful token generation.
  • Success per dollar: An operational efficiency metric intended to show whether the platform is converting infrastructure spend into completed useful work.
  • Retry budget: A deliberate limit on retries used to prevent a degraded component from generating traffic amplification and cascading fleet or regional failures.

Operator Notes / Why Ken Should Care

  • Define a request/workflow metadata contract that carries tenant tier, deadline, model-quality requirement, token estimate, workflow step, accumulated spend, and retry allowance into routing and scheduling.
  • Audit whether current autoscaling and load shedding use queue depth, warm-model availability, KV-cache/HBM pressure, and retry rate—not only CPU/GPU utilization.
  • Create a scorecard centered on successful task completion: time to first token, end-to-end latency, success per dollar, retry amplification, cache hit rate, and quality-adjusted task success.
  • Run failure drills for mid-stream GPU loss and cold-pool failover; explicitly test partial-response handling, retry policy, queue buildup, and protection of high-value in-progress workflows.
  • Evaluate any inference gateway, runtime, or vendor control-plane strategy by its ability to make coordinated cache-, workflow-, hardware-, and priority-aware decisions, rather than by token throughput alone.

Source/Metadata

  • Title: Operating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta
  • Transcript words: 3458
  • Duration seconds: 1190
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.
Full transcript 3252 words · 15 min read
0:00

Good morning, everyone.

0:12

Welcome to the first Influence Talk on the last day of AI Engineering World Fair. My name is Nishan Gupta, and I'm joined today by my co-speaker Naman Ahuja. We work on building the efficiency, training, and inference infrastructure at Meta. Today, we're going to be talking about how do you operate distributed inference systems at scale. As we all know, inference is no longer just a research artifact boiled to a product. It's a foundational hyperscale infrastructure workload, which is growing at a tremendous rate. The inference traffic already outpaces the largest microservices in the world, and the rate of growth is fastest of any workload we have ever seen.

0:47

Let's rewind back to around 2008 and try to compare the AI era with the cloud era.

0:54

In around 2008, the cloud started as virtual machine offerings. The interesting engineering was virtualization. Then, over time, the value moved up to the stack, to the schedulers, Kubernetes, Mesos, then to service meshes, then to autoscalers, and then to various platforms that were built on top of it. The orchestration layer is what actually captured the value and the complexity. AI is on the exact same trajectory, but just compressed into the last few years instead of a decade. We started with simple models running on GPUs. Then we saw the evolution of model-serving frameworks like VLLM,

1:33

Torchserve, Triton, and now we are watching the orchestration layer emerging in real time, addressing the complex challenges of routing, KV cache management, prefill decode disaggregation, and multi-model multiplexing. In this next phase of AI era, it's not just about the best models or the kernels or the optimizations. It's about the whole ecosystem. It's about the control plane and the orchestration. And this is what we're going to be focusing in this race talk. So let's talk a little bit about agentic demand explosion.

2:05

In the classical web-serving before the AI inference workloads kicked off, the capacity scaled roughly linearly with users, depending upon the workload type. Double the users, mostly double the QPS, double the infrastructure fleet, if there are no optimizations, and capacity planning was mostly a spreadsheet exercise. In this new agentic serving, capacity is scaling with number of users times number of calls per user times number of tokens, which varies depending on the model, the optimizations, the hardware SKU you have. A chatbot can have one model call per turn, which is evolving now with 10 to 20 for co-pilots,

2:40

and 50 for research agents, and now it's thousands of such calls for autonomous workloads, with no human in the loop. The key takeaway is that you cannot plan capacity for the agents the same way we did for microservices. We need to think about elasticity and implement workload-aware scheduling and admission control. Now let's try to dive a little bit deep into the differences, pros and cons, differences between the traditional microservice serving versus the modern inference serving across these key dimensions. The request shape. Microservices assume short, uniform requests, where the LLM requests we are seeing in our workloads,

3:23

they can vary from 50 tokens to 100,000 tokens, with vastly different compute profiles between prefill and decode stage. For batching, classical stacks for microservices was doing batching, batching mostly at the load-balance layer, if at all. However, LLM serving requires continuous in-flight batching, otherwise the throughput collapses by an order of magnitude or more. State. Most of the classical microservices were stateless, when we're not talking about storage layer. However, LLM serving does require a huge per-request state, the KV Cache, which is very expensive to build, and even more expensive to throw away. Scaling units.

4:03

When we think about traditional microservices, we could run them in cheap CPUs, in pods. However, for modern inferencing, we require to run on GPUs, which are 100 times more expensive, which are 10 times slower to acquire, and we cannot over-provision them casually. Otherwise, it will lead to a huge wastage. Failure mode. When you think about traditional microservices, like most of us have built over the last couple of years, we could, even if a host crashed or a pod crashed, we could restart it. We could rebuild the state if need be. However, for the model inferencing, it takes a huge amount of time to go from cold to hot startup. And if a GPU is mid-decode,

4:41

it can drop thousands of in-flight tokens, and it can lead to a queue buildup. The takeaway is that the bottleneck is not just the model, it's the orchestration itself. Now, as we can see in these hidden decisions behind a prompt, when we go to any agentic application, it requires a bunch of steps which are behind the scenes. We have to authenticate. We have to choose a model depending on the request type. We have to select the region where it goes. We have to do admission control. We have to do caching lookup. We have to run it on the GPU. We have to do batching. And there's a bunch of other steps involved. And as we can see, out of all these steps,

5:19

only one step requires the model, which is the prefill decode, if it's disaggregated inference or just if it's not disaggregated inference. Other steps require infrastructure. The intelligence might lie in the model, but the economics, the reliability, and the user experience are all in the infrastructure. And this is why a lot of platform teams across a lot of companies are having much more impact on the product quality and the success, much more than before. Most of us in this room have a deep expertise across one or two or three layers.

5:50

We might own kernels or kernel optimizations. We might own routing or the product itself, or we might be operating the GPU infrastructure or the cluster itself. But very few of us have operated the whole stack or seen, thought about it end to end. As you can see in these layers, these layers are not new. They have been around for 20 years or more. What is new is the combination and the coupling between them. A decision at the routing layer can change the cache hit rate at the model layer, which can change the batch composition, which can change the GPU utilization, which can change the autoscaling decision because of the change in GPS simplization.

6:29

So everything is entangled. When we have any regressions in our inference workloads, it's not just about understanding what happened at the caching layer or admission control. We need to think about the stack top to bottom. And when we see any bottleneck, it's very important to understand at which layer is that bottleneck so that we can invest properly. Now, diving a little bit deep into how a prompt works.

6:53

When we have a prompt for any application, be it if you want to generate an image or if you have a research task or if you have complex multi-agent orchestration, more or less, it involves a bunch of these steps. The prompt goes to the gateway. Then it goes to the router, after which it does the cache lookup if the request was already seen before. Then it goes to the schedulers, which decides on which GPU cluster it should run, and on which hardware. It can be on NVIDIA or AMD or your in-house silicon chip, which then goes to the appropriate serving runtime, VLLM, MIG-Lang, or whatever we are working on. And then we stream the response back to the user

7:30

according to the SLO profiles of time-to-first token and time between each token, and while making sure the throughput is what the user desired. Now, as we can see, this inference behaves like a distributed transaction. Each arrow in this diagram is a network hop. Every one of these hops can retry. It can time out. It can fall back. It can even fail. And each of these hops will have SLOs, and it is streaming back to the user. So if it fails, the partial philis mantix are much harder to deal with than for a regular RPC call. Think about what happens if we have already streamed 200 tokens back to the user, and suddenly a GPU host is preempted

8:08

due to a scheduled or a planned or unplanned maintenance event. We cannot just retry. We have to think about it holistically. This is why reliability, we cannot build reliability at the edge. It has to be a property of the control plane, because the control plane is the one which sees the whole workflow. Now, let's talk about schedulers and some of the optimizations and how we can think about it. For traditional microservices, we used to think about bin packing of traditional microservices across three, four dimensions. Could be across the virtualization, CPU memory, or could be across four domains, depending upon if you're using AWS,

8:45

if you're using your own in-house cloud providers. But for inference, the scheduler needs to be aware across at least seven axes. When we schedule a particular request, it has to be aware of the GPU type. There can be n number of heterogeneous hardwares in your cluster, H100 versus E100 versus B200 with different network topologies. It has to be aware of the HBM headroom, KV cache state. The model weights, whether they're already loaded, whether they're cold or we have to, whether they've already warmed up, we have to cold start it. It has to be aware of the tenant priority. and some of the optimizations and how we can think about it.

9:21

So for the traditional microservices, we used to think about pin packing of traditional microservices across three, four dimensions. Could be across the socialization, CPU memory, or could be across four domains, depending upon if you're using AWS, if you're using your own in-house cloud providers. But for inference, the scheduler needs to be aware across at least seven axes. When we schedule a particular request, it has to be aware of the GPU type. There can be n number of heterogeneous hardwares in your cluster, H100 versus E100 versus B200 with different network topologies. It has to be aware of the HBM headroom, KV cache state.

9:55

The model weights, whether they're already loaded, whether they're cold or we have to, whether they've already warmed up, we have to cold start it. It has to be aware of the tenant priority. There can be n number of tenants running on that multi-tenant cluster with different SLO profiles. We have to also be aware of the workflow context. Are we in the third step of reasoning that has already spent X dollar? Or are we at the initial stages and we can terminate the workflow if we are over-provisioned? We have to also think about latency budget, depending upon the type of agentic application we are building. So this brings us to the idea that we have to make sure

10:35

that we implement agentic-aware scheduling. We have to make sure that we have placed the work which will finish in the fastest and the cheapest time, as opposed to just placing it on a random GPU. A concrete example might be that the scheduler needs to be aware that request R is at step three of a five-step workflow, and step one and two has already spent X plus Y dollar. So if a step three fails, the whole workflow will be terminated and we have wasted all that compute resources. That's why workflow-aware orchestration is very, very important, because it will change the admission decisions, the priority, and how we retry. Now let's talk about optimizations.

11:13

I will not go too deep into a lot of optimizations. There's a lot of research that has already been done outside, but I would like to share a framework, which at least I like to use when it comes to it, and we can place it into four quadrants. First, can we avoid the work? Meaning, can we skip it entirely through caching, through techniques like prefix caching, response caching, semantic caching? The second, can we share the work? Can multiple requests share compute through batching?

11:47

Think continuous batching, prefilled decode, chunk prefilled, speculative decoding. Third, can we move away the work? Can we send it somewhere to a cheaper model or closer to the user? Through mostly routing, can we route it to a smaller model or a cheaper region or some other techniques? And lastly, can we delay the work? Can we wait for a better moment through admission control and queuing, which requires us to understand the priority classes of these requests and implement deadline-aware scheduling. Now this framework is very powerful because it transfers to various stacks. You might be using vLLM or SG-Lang or TensorRT,

12:26

but every technique fits into one of these quadrants. So whenever we think about any optimization to our model, we have to apply it.

12:36

We have to do a comparison and contrast with the previous techniques and see how all these stack with each other. Now whenever we think about scale, it's not just important to think about the performance of the model. We have to think about the cost, the economics as well. This is where it's important to understand what metric we are trying to optimize. Because the cost is not just the cost of the GPU or the model. It is the cost of all these parameters, retries, storage, failures, network, and of course the operational cost of development and all that stuff. What is important is to understand what is the key performance indicator

13:13

for your product, which will add value to the users. So it's not important to optimize just cost per token or cost per request. We have to optimize cost per successful task because this is what actually users care about. And if we're able to optimize that, the cost for the overall product decreases and the users are much more happier. Now let's talk about reliability, a little bit about reliability and what it means to prevent cascading failures. Now the failure story is never a GPU is preempted or a GPU has died.

13:43

The interesting story is the feedback loop that follows. So a GPU can degrade, the latency can rise, the client retries, the queue depth increases, the healthy GPUs will saturate, which will follow more retries, much more full regional failures. This is the classic cascading failures, but with a twist for agentic application, the KV cache. We cannot just casually restart or reroute to a different cluster. A cold pool has to warm up before it can absorb traffic, during which the hot pool has to take on all those requests. This is why it's important to design the loop breakers very deliberately. The circuit breakers at the routing layer, the admission control,

14:33

which is rather than just queuing, the load shedding tied to queue depth, not just CPU or memory utilization. And we have to also think about retry budgets, because if we don't think about all these things, the cost can scale much, much more quickly. Now I'll hand it over to my co-speaker, and I want to talk about the remaining talk. So once inference reaches production scale, it starts looking much more like a distributed system. We are no longer just calling a model. It's more like a classic distributed system problem. So in distributed system, we talk about queues, scheduling, auto-scaling, fault isolation. These are some of the dimensions.

14:55

Inference has all these problems, but there are new constraints now. Instead of CPU memory alone, we have CPU, HBM, KVCache, and cost per successful task. So the operating question becomes, how does the platform know what to do next? And that's where observability comes into play. It's not just about dashboards. It's about how to provide input signal to the control loop. Telemetry, fields, analysis, analysis drive decisions, issues and changes of scheduling and routing, and finally we just rinse and repeat. I'll just give an overview of what are some important metrics.

15:26

First one is time to first token, which tells me about how much time does it really take to get the first response. Then we have utilization ratio, which tells me whether memory or compute is the bottleneck. We have success per dollar that tells us whether the platform is actually delivering and working efficiently. And finally, we have end-to-end request latency that tells us what's the time being spent across the full request path. There's a trade-off between latency, cost, and throughput. You cannot just get all of them. It's pretty analogous to cap theorem. If I increase the batch size, I improve the throughput and cost efficiency, but I may hurt tail latency.

15:47

If I use speculative decoding, I may improve latency, but there are some extra compute. Ultimately, we increase the cost per token. And finally, I can use a simple model, smaller model. I can reduce the latency and cost, but the response will be of low quality. Ultimately, I'll do failure analysis, do retries, which brings the cost back up. So every serving decision moves systems somewhere in this triangle. And our job is to find the perfect setting. It's just an optimization problem now. This is where the industry is heading right now. Inference needs its own control plane. Everything we discussed, routing, batching, caching, scheduling, reliability,

16:41

they cannot be a separate knob now. They're converging into a logical layer. Let's call it inference control plane. We used to manage VMs before in distributed system. We have auto-scaling schedulers. And we had Kubernetes, which turned these into control plane.

17:03

Inference is going through the same transition now. Models are becoming resources. GPU, KVCache, token, latency, cost are now scheduled around. The control plane decides which models serve which request and how it is batched. So whether we build this layer internally, use open source, or from a vendor, the key design is to assume layer will exist. Now let's discuss what are some of the operating lessons we have experienced in A-infra, and how they are applicable here. The first lesson is that infrastructure bottlenecks usually show up before model bottlenecks. In production, many failures can come, but they can be just about scheduling and routing, or capacity breakdowns.

17:32

So these are not related to inference. It's about infrastructure problems. Then we have elasticity. We need elasticity in the system. We can have more GPUs, but this will not really solve the problem. We are just hiding the problem. Then our solution to schedule the decision, overpowered efficiency. Models are becoming resources. GPU, KVCache, token, latency, cost are now scheduled around. The control plane decides which models serve which request and how it is batched. So whether we build this layer internally, use open source, or from a vendor, the key design is to assume the layer will exist.

18:12

Now let's discuss what are some of the operating lessons we have experienced in A-infra and how they are applicable here. The first lesson is that infrastructure bottlenecks usually show up before model bottlenecks. In production, many failures can come, but they can be just about scheduling and routing or capacity breakdowns. So these are not related to inference. It's about infrastructure problems. Then we have elasticity. We need elasticity in the system. We can have more GPUs, but this will not really solve the problem. We are just hiding the problem. Then our solution to schedule the decision overpowered efficiency.

18:48

The same fleet can deliver very optimally, depending on how you're scheduling it or how you're batching it. Then we have control loops, which beats the manual process. The platform has to sense, detect, and automatically adapt to the system. So main takeaways: do not optimize for tokens. Optimize for successful tasks. So this is a broader shift I want to leave you with. The first phase of AI infrastructure was about better models. We invested a lot of time in improving our models, making the models smarter, and coming with better benchmarks. The current phase right now is faster inference, lower latency, better batching, and better GPU utilization.

19:25

But the next phase is about orchestration. That means GPU, memory, cache, and everything—these are just resources, and they need to be scheduled and controlled. The teams that understand this early on will build infrastructure for the future.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note