AI Engineer

Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave

1814 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: CoreWeave argues that a durable inference platform must support both serverless and dedicated consumption while dynamically optimizing across workload shapes, heterogeneous GPUs, and KV-cache locality to deliver price-performance.
  • Why it matters: The talk identifies practical control-plane and routing choices that become important when agentic workloads create repeated, long input contexts and when real-time, streaming, and batch jobs must share expensive GPU capacity.
  • Best use: Use it as a concise infrastructure-design reference for inference routing, cache strategy, capacity segmentation, and scheduling—not as a deep implementation tutorial.

Executive Summary

Sitanshu Gupta frames “vertical mobility” as serving the full range of inference—from small MVP deployments to extremely large models—without rebuilding the platform for each customer or workload. CoreWeave’s proposed answer is one platform with two primary commercial and operational modes: serverless, billed per token, and dedicated capacity, billed per GPU-hour. Serverless additionally offers provisioned throughput so customers can retain API simplicity while receiving reserved capacity and SLA protection from multi-tenant noisy-neighbor effects.

The technical center of the presentation is workload-aware routing and KV-cache management. Gupta says agentic and chat requests generally have long inputs and short outputs, while agentic loops have much tighter inter-turn latency. Because roughly 80–90% of input context may repeat across agentic requests for some customers, cache-aware routing can avoid recomputing expensive prefill work. CoreWeave prioritizes KV-cache locality first and least-loaded capacity second, including across heterogeneous GPU capacity, zones, and regions.

The platform is intended to expose multiple inference engines—vLLM, SGLang, and TensorRT-LLM—and support differing GPU generations, deployment patterns, parallelization strategies, and optional prefill/decode disaggregation. Gupta stresses that disaggregation is a conditional optimization rather than a universal default, because its added cost and complexity are not justified for every workload shape.

CoreWeave also treats capacity as a scheduling problem: real-time workloads occupy reserved capacity during active periods, while customers can schedule that same dedicated capacity down and use it for loose-SLA batch jobs overnight. Additional performance levers include NVFP4 quantization, speculative decoding, and asynchronously trained customer-specific speculative models. The closing benchmark references support the positioning, but the presentation offers little methodology or quantified economic impact beyond qualitative claims.

Key Takeaways

  • Claim: Inference infrastructure should offer distinct serverless and dedicated modes rather than forcing all customers into one deployment and billing model. | Evidence: CoreWeave describes serverless as API/UI access with per-token billing and provider-managed hardware and orchestration; dedicated inference uses isolated customer capacity, customer-controlled deployment choices, and per-GPU-hour billing. | Implication: For an AI product or agent system, select serverless for simplicity and variable demand; move to dedicated capacity when isolation, predictable hardware, deployment control, or sustained utilization justifies reservation. | Caveat: Dedicated capacity gives customers more control over engines, deployment topology, and isolation, but the customer also bears more responsibility for deployment and model-performance decisions.
  • Claim: Provisioned throughput is a middle ground between ordinary multi-tenant serverless and dedicated clusters. | Evidence: Gupta says serverless customers can provide a known traffic profile so CoreWeave carves out backing capacity, preserves throughput and SLAs, and still bills by tokens rather than exposing hardware management. | Implication: If Ken operates latency- or throughput-sensitive agents but does not want to manage GPUs, a provisioned-serverless product tier is a valuable procurement and platform pattern to seek or build. | Caveat: The benefit depends on being able to forecast and communicate a sufficiently reliable traffic profile.
  • Claim: KV-cache-aware routing is the highest-leverage inference optimization for agentic workloads with recurrent context. | Evidence: The speaker estimates that 80–90% of input context can be the same across different agentic requests for some customers, while prefill is compute-bound and expensive; the router therefore selects cache locality before falling back to the least-loaded resource. | Implication: Agent architecture should maximize reusable prompt prefixes, tool context, and conversation state; infrastructure routing should preserve cache affinity instead of treating each request as stateless load-balancing traffic. | Caveat: The 80–90% figure is presented as customer- and workload-dependent, not as a universal agentic-workload benchmark.
  • Claim: Workload classification should drive latency, scheduling, and cache policy because chat, agents, voice/video, and batch have materially different operating constraints. | Evidence: Agentic and chat workloads are described as long-input/short-output real-time traffic, but agents have fast multi-turn loops; voice and video are steady streaming and highly latency-sensitive; batch jobs may tolerate seconds, minutes, or even 10–12 hours. | Implication: Do not apply one SLA, autoscaling policy, queue, or cache-eviction strategy to all model calls; separate interactive agent loops, streaming experiences, and deferred back-office processing. | Caveat: Gupta characterizes agentic and chat needs as more throughput-oriented than latency-oriented overall, which may not hold for user-facing or tool-execution-critical agent experiences.
  • Claim: Prefill/decode disaggregation is a selectable optimization, not a baseline architecture requirement. | Evidence: CoreWeave exposes the ability to split prefill and decode hardware or keep them combined, and Gupta explicitly says disaggregation is not cheap for every use case. | Implication: Treat prefill/decode separation as a benchmark-driven design decision tied to context length, output length, concurrency, and cache-hit behavior—not as an automatic scaling pattern. | Caveat: The presentation does not specify the workload thresholds, networking requirements, or measured gains that determine when disaggregation wins.
  • Claim: Idle dedicated inference capacity can be monetized or utilized better by scheduling batch jobs around real-time demand. | Evidence: CoreWeave describes a dedicated customer using capacity for real-time workloads during U.S. daytime, then scaling that workload down via API on a schedule and opening the same capacity for batch processing overnight. | Implication: Ken should design non-urgent evaluation, enrichment, indexing, and data-processing workloads to be interruptible and schedulable into off-peak inference capacity. | Caveat: This approach requires predictable demand windows and sufficiently loose batch SLAs; it is less useful when real-time traffic is globally distributed or highly volatile.
  • Claim: Performance gains compound across quantization, speculative decoding, cache design, engine selection, and parallelization rather than coming from a single serving engine. | Evidence: The talk names NVFP4 quantization, speculative decoding, customer-specific speculative-model training, parallelization choices, and vLLM/SGLang/TensorRT-LLM as platform levers; CoreWeave claims recent leaderboard strength for Kimi 2.6/2.7 in Artificial Analysis and near-Fireworks speed on GLM traffic in OpenRouter. | Implication: Evaluate inference vendors and internal serving stacks on workload-specific end-to-end price-performance, not engine branding or a single public leaderboard. | Caveat: The benchmark claims lack configurations, absolute throughput/latency figures, cost data, and controlled methodology; Gupta himself distinguishes synthetic benchmark workloads from OpenRouter’s production traffic.

Detailed Brief

Control plane, isolation, and data handling

  • Claims: Both serverless and dedicated requests traverse a control plane responsible for authentication, authorization, rate limiting, usage tracking, billing support, observability, and SLA monitoring.; Dedicated customers use private gateways with isolation intended to prevent other tenants from accessing their capacity.; CoreWeave says it tracks token usage for serverless billing without retaining the token contents under a zero-data-retention policy.
  • Evidence: The speaker explicitly separates the request gateway/control plane from the inference engines and GPU hardware layer.; The stated engine layer includes vLLM, SGLang, and TensorRT-LLM, while the hardware layer spans different GPU generations.
  • Caveats: Zero data retention is stated at a high level; the transcript does not define operational details such as logging exclusions, cache retention duration, customer controls, or audit mechanisms.; The talk does not describe how routing, isolation, and observability are implemented across regions or failure domains.
  • Implications: For enterprise agent deployments, data retention, request logging, cache persistence, and tenant-isolation semantics should be evaluated separately from headline model performance.; A usable inference control plane needs product-grade identity, quotas, billing, and observability alongside model-serving capability.

Cache persistence and tailored speculative decoding

  • Claims: For chat workloads with long pauses between turns, retaining all KV cache in GPU HBM is inefficient, but full eviction makes the next turn slower because prefill must be repeated.; CoreWeave offloads cached prefills to high-bandwidth storage and reloads them into HBM when the corresponding conversation returns.; Customers can provide datasets for asynchronous training and deployment of speculative models intended to improve acceptance length and output throughput.
  • Evidence: Gupta references LMCache and Mooncake as external examples in the cache-offloading space.; The proposed speculative-decoding workflow is asynchronous: collect data, train speculative models, then deploy them to the customer environment.
  • Caveats: No cache-hit rates, reload penalties, storage architecture, quality effects from quantization, or speculative-decoding acceptance-rate results are provided.; Customer-specific speculative training introduces a data-governance and deployment-validation surface that the talk does not address.
  • Implications: Conversation-memory design has an infrastructure cost and latency dimension: persist enough reusable state to avoid repeat prefills, but tier it rather than pinning everything in GPU memory.; Custom speculative models are a potential optimization only after measuring stable request distributions and validating quality, safety, and rollback procedures.

Notable Concepts & Terms

  • Vertical mobility: CoreWeave’s framing for a single inference platform that can serve workloads ranging from small deployments to very large-model inference without architectural replacement.
  • Provisioned throughput: A serverless tier where capacity is carved out for a customer’s expected traffic profile, reducing noisy-neighbor risk while retaining token-based pricing.
  • KV cache: Stored model attention state from prior tokens; maximizing its reuse avoids expensive prefill recomputation, especially for repeated agent context.
  • Prefill/decode disaggregation: Separating prompt processing from token generation onto different serving resources; potentially useful but not cost-effective for every workload.
  • KV-cache locality routing: Sending a request to infrastructure that already holds its relevant cache state, which CoreWeave prioritizes before least-loaded routing.
  • NVFP4 quantization: A low-precision inference optimization cited as one of CoreWeave’s major performance levers.
  • Speculative decoding: Using a smaller draft/speculative model to propose tokens that a target model verifies, with custom training intended to improve acceptance and output throughput.
  • LMCache and Mooncake: External systems cited as examples of offloading KV-cache state to high-bandwidth storage for later reuse.

Operator Notes / Why Ken Should Care

  • Instrument agent traffic to measure repeated-prefix ratio, cache-hit rate, prefill cost share, inter-turn delay, and tail latency before selecting a serving topology.
  • Create separate service classes and queues for interactive agent/tool loops, streaming voice/video, and interruptible batch work; assign explicit latency and interruption policies to each.
  • For predictable demand, compare provisioned serverless commitments against dedicated GPU reservations using real utilization, burstiness, and SLA costs—not just token price.
  • Preserve stable system prompts, tool schemas, retrieval scaffolding, and conversation prefixes where possible so cache-aware routing can produce real savings.
  • Run workload-specific experiments before adopting prefill/decode disaggregation, NVFP4, or speculative decoding; require quality, latency, throughput, and unit-cost results plus rollback criteria.
  • If using cache offload or customer-trained speculative models, define retention boundaries, tenant isolation, data consent, validation, and observability requirements up front.

Source/Metadata

  • Title: Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
  • Transcript words: 4771
  • Duration seconds: 922
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript; the transcript substantially repeats the presentation.
Full transcript 2370 words · 21 min read
0:12

Good afternoon everybody. I'm Sitan Shu from Corvi. I'm going to be talking about vertical mobility. It's a fancy topic, the title that we came up with, but going to be talking about the inference platform that we have at Corvi that we are building to serve small to big models and various different types of workloads. A quick intro about me, I joined Corvi just about four months back, leading all of inference over there and before this I was managing everything at AWS Annapurna Labs for training and before that inference and training at Sambanova. So quite a bit of experience in this particular space. The way I'll be taking you through is explaining to you the consumption models that we have and from that how we have derived what the platform should look like so that we do not need to keep changing the platform and we keep making enhancements in the platform that we have for serving inference and how and why performance plays such an important role over there. I think a little bit of this might be common with the previous topic that was discussed over here. Consumption models. So we have at large two biggest consumption models. One is serverless, which is where the customers can come in, consumers can come in, do not need to worry about managing the hardware themselves, do not need to worry about managing the clusters, orchestration, anything at all. There's API, there's UI, you come in, you pay per token and you get your model served. Biggest thing over here is that the type of models that we serve in the catalog, that is the breadth of models that the customer will be able to go through. I'll talk about dedicated and then I'll come back to serverless because there is one thing unique on the serverless side. Dedicated inference service that we provide is more for customers who want to know exactly what hardware they are going to be using and running on. But the model deployment also depends on them. So they use our service, they use our orchestration layers, but the model deployment depends on them. Model performance also depends on them as long as we provide in the platform the capability and the knobs to serve those features. Coming back to serverless, one of the interesting pieces over here is typically serverless models are noisy neighbor problems where if let's say everyone is banging on the exact same model, then you might be timing out quite a bit depending on how much capacity I have behind it. So another feature that we have on the serverless side is what we are calling a provisioned throughput. So as a customer, if you know your traffic profile and if you can let us know about that, we can carve it out specifically for you behind the scenes. You still do not need to worry about what hardware it is exactly running on as long as your throughput, your SLAs are maintained. So that is another one on the serverless side and that is still charged by the token, but you know that you are not running into the noisy neighbor problem over there. Let me take a quick stab at a few different types of workloads, workload shapes that we have, that we are seeing and the ratio between these is continuously changing though agentic is really high up there. Agentic and chat are very similar, super high on the input sequence lengths, very low on the output sequence lengths typically, but the biggest difference between agentic and chat being the fact that the multi-turns in agentic are super low latency versus in chats because when you get the response of the user, you have to read the answer and then you respond to it. So there are differences over there and that big difference ultimately converts into something related to the KV cache management. But these two are both real-time and another real-time workload is your voice and videos, which are steady streaming and super latency sensitive. On the agentic and chat side, largely the requirements are from throughput point of view, not so much from latency, but real-time voice and videos are absolutely latency sensitive. Coming to batch, batch is where the SLAs are super loose, they run into seconds and minutes, sometimes for some customers actually even in hours, they're like I'll just throw, give me 10 to 12 hours of workload capability and I'll throw whatever I can, process it whenever you can. These batch workloads, the way they come into the picture over here in deciding, sorry, being the requirement for some of the design choices that we make in the stack, imagine these four different types of workload shapes. In the time dimension, you have to play the game of Tetris on how you can fit it in to utilize the underlying infrastructure the most. I'll give a high level on how our stack is shaped right now and I'll walk you through a bit of a request flow over here. So for both serverless and dedicated, if you look at the right hand side of the screen, you'll see that on the platform side, you'll go to the control plane to have your authorizations, your rate limitings and your usage being tracked, etc. so that you can be built accordingly and observability so that we can make sure that we are not violating the SLAs that have been signed. Underlying on the platform, I've shown at our super high level that we have these different inference engines, VLLMs, SGLangs and TensorRTLM, but there's quite a bit of detail over here that I'll touch up on. And underlying that, what I'm trying to show over here in green is various different pieces of hardware. So it's not that, so the platform needs to be capable enough of sharing, of having the workload getting distributed across various different generations of these GPUs, specifically in media GPUs that we use. So let's take a few examples over here. Let's say the request originates from the client side through apps or notebooks, any of those, or through the agents. It hits the gateway. Once it hits the gateway, then as I mentioned on the control plane, goes through authentication, etc., but then comes either the serverless or dedicated. So in the case of serverless, it will be paper tokens, so the token usage would be monitored over here. Not the exact tokens, but the token usage, because we maintain ZDR, zero data retention policies. It is, depending on the multi-tenancy or the provisioned, if it is provisioned, then we know underlying for the router, it needs to go in and target the explicit deployments for the provisioned throughput customers. For the multi-tenant customers, there are separate deployments. Router over here, specifically, the router is very important, since the router is responsible for making KV cache-aware routing choices. Why is it important? Because as I mentioned when we were discussing the workload profiles, the agentic use cases are typically super heavy on the input sequence lengths and bulk of the input sequence length, about 80 to 90 percent, depending on which company it is, depending on the customers, 80 to 90 percent of it is the same for various different requests. So there is no point in going in and re-computing the pre-fill or redoing the pre-fill for that. Pre-fill is super compute bound, very expensive, that's why as much as you can hit the cache, more you can save, which is why if you look at the token pricing anywhere, there's a specific price for input tokens and there's a way cheaper price for the cache input tokens. So caching becomes really important over here. Underlying, how you want to split the hardware is totally dependent on the choice in the platform and we provide the capability to do either. Either do a pre-fill decode disaggregation if the use case desires it or do not do it because pre-fill decode disaggregation is not cheap for every type of use case. Let's take another request flow over here. Let's see if when it was a dedicated customer, then what will happen. A dedicated customer, again, will go through the gateways that have been set up for them with proper isolations. Billing is not based on tokens. Billing is based on usage of per GPU per hour. It's a private gateway so that there is no noisy neighbor problem. No one else can get in. Same router logic over here so that if there are requests which are very similar, then it hits the cache most. Depending on the deployment that the customer makes on their dedicated GPUs, they can decide if they want to do pre-fill decode disaggregation or not. They can decide which engine to use. VLM or SG-LANG or TENSOR RTLM. Given the bulk of capacity that the customer has reserved, they can decide if they want to have just one deployment with the ability to scale through the whole cluster or they want to have multiple different models, multiple different deployments. One thing that I do want to mention about the router over here is the fact that heterogeneous capacity across different zones and regions is supported. It is quite a bit of a hard problem to load balance across that so the priority order that we typically take is first KV cache locality and then the least loaded fallback. That's that. Another request flow that I want to go over here which might be a little hard to see from the diagram is I want to take the batch workflow. For the batch workflow what we would actually do is the underlying capacity that the customer has, let's say the same dedicated inference customer, during US daytime they're running their real-time workloads and from evening to night they want to run batch workloads, the same capacity after time can be scheduled to run the batch workloads. So we provide the capability in the API to tell when to scale up and when to scale down and as per schedule if we can if they tell us that we can have to scale down we will scale down and open it up for batch processing for the night. I think I've spoken quite a bit about optimizations on the KV cache side but I do want to repeat a little bit because this is one of the most interesting pieces. If we can hit on the cache more you can avoid the cost of pre-fill which is the most expensive piece over here. Reusing the KV cache across multiple different turns in your agentic workloads between turns also there is a lot of similar pre-fill that comes in in the input sequence length. Think about the chat workloads which is where offloading KV cache also becomes extremely important because with cache with the chat workloads we have a lot of latency between different between multiple turns that we as users put in but if we completely evict whatever we had in our particular conversation then the next time we ask a question in the same chat it's going to take a little bit longer. So instead of actually completely evicting and redoing the pre-fill again what the techniques being used are maybe using we are using our own but externally we know about LM cache and Mooncakes the pre-fill again what we do is we will offload the KV cache to a high bandwidth storage so that we can store a lot of these pre-fills such that whenever the accompanying request comes for that particular conversation it can be loaded in right away into the HBM. On the performance lever I would I just want to mention a few performance levers that we've discussed the period disag that is one but quantization and specular decoding are others and how to carefully choose the parallelization degrees and the strategies that is actually very important. Two of the biggest levers that we have been working with are quantization to NVFE4 and spec deck. We do provide capability where if the customer has their data set and they want us to train speculators for their data sets for better acceptance lens which will ultimately make the output throughput significantly higher. We do have that as well so but that happens async we gather data async we train the speculators async and then we deploy the speculators into the customer deployments if that's what they wanted. You see three screenshots over here I have posted them from the last one month one month's worth of work that some of us in my team have done. You can see we came quickly on top of the leaderboard on Kimi 2.6, 2.7 and those are from artificial analysis and going back to the session before this can we trust that that's why for GLM I have the results from open router. So artificial analysis when they run benchmarks they're running very specific workloads. Open router is actual user traffic and you can see on the open router side weights and biases so the branding is different but weights and biases basically curve we bought weights and biases about a year back. You can see the speed over here that we have from our deployment is pretty close to what Fireworks is providing as Fireworks fast but underlying techniques that we are using is what I want to emphasize the most over here for performance optimization. That becomes critical because ultimately what you want to serve to the customer what we want to serve to the customer is price performance benefit. Quick recap, single platform is what I've been trying to emphasize is what I'm trying to show two different consumption models serverless and dedicated for customers and within serverless I describe two different consumption models as well pay as you go and provision throughput if you care about that and ultimately compounding the gains through performance optimizations in the stack. That's all thank you folks.

0:19

mobility. It's a fancy topic, the title that we came up with, but basically going to be talking about the inference platform that we have at Corvi that we are building to serve small to big models and various different types of workloads. A quick intro about me, I joined Corvi just about four months back, leading all of inference over there and before this I was managing everything at AWS Annapurna Labs for training and before that inference and training at Sambanova. So quite a bit of experience in this particular space. The way I'll be taking you through is explaining to you

0:59

the consumption models that we have and from that how we have derived what the platform should look like so that we do not need to keep changing the platform and we keep making enhancements in the platform that we have for serving inference and how and why performance plays such an important role over there. I think a little bit of this might be common with the previous topic that was discussed over here. Consumption models. So we have at large two biggest consumption models. One is a serverless, which is where the customers can come in, consumers can come in, do not need to worry about managing the

1:35

hardware themselves, do not need to worry about managing the clusters, orchestration, anything at all. There's API, there's UI, you come in, you pay per token and you get your model served. Biggest thing over here is that the type of models that we serve in the catalog, that is the breadth of models that the customer will be able to go through. I'll talk about dedicated and then I'll come back to serverless because there is one thing unique on the serverless side. Dedicated inference service that we provide is more for customers who want to know exactly what hardware they are going to be using and running on.

2:12

But the model deployment also depends on them. So they use our service, they use our orchestration layers, but the model deployment depends on them. Model performance also depends on them as long as we provide in the platform the capability and the knobs to serve those features. Coming back to serverless, one of the interesting pieces over here is typically serverless models are noisy neighbor problems where if let's say everyone is banging on the exact same model, then you might be timing out quite a bit depending on how much capacity I have behind it. So another feature that we have on the serverless side is what we are calling

2:49

a provisioned throughput. So as a customer, if you know your traffic profile and if you can let us know about that, we can carve it out specifically for you behind the scenes. You still do not need to worry about what hardware it is exactly running on as long as your throughput, your SLAs are maintained. So that is another one on the serverless side and that is still charged by the token, but you know that you are not running into the noisy neighbor problem over there. Let me take a quick stab at a few different types of workloads, workload shapes that we have, that we are seeing and the ratio between

3:29

these is like continuously changing though agentic is like really high up there. Agentic and chat kind of very similar, super high on the input sequence lengths, very low on the output sequence lengths typically, but the biggest difference between agentic and chat being the fact that the multi-turns in agentic are super low latency versus in chats because when you get the response of the user, you have to read the answer and then you respond to it. So there are differences over there and that big difference ultimately converts into something related to the KV cache management. But these two are both real-time

4:07

and another real-time workload is your voice and videos, which are steady streaming and super latency sensitive. On the agentic and chat side, largely the requirements are from throughput point of view, not so much from latency, but real-time voice and videos are absolutely latency sensitive. Coming to batch, batch is where the SLAs are like super loose, they run into like seconds and minutes, sometimes for some customers actually even in hours, they're like I'll just throw, give me 10 to 12 hours of workload capability and I'll throw whatever I can, process it whenever you can. These batch workloads, the way they come into the picture

4:51

over here in deciding, sorry, being the requirement for some of the design choices that we make in the stack, imagine these four different types of workload shapes. In the time dimension, you have to play the game of petris on how you can fit it in to utilize the underlying infrastructure the most.

5:12

I'll give a high level on how our stack is shaped right now and I'll walk you through a bit of a request flow over here. So for both serverless and dedicated, if you look at the right hand side of the screen, you'll see that on the platform side, you'll go to the control plane to have your authorizations, your rate limitings and your usage being tracked, etc. so that you can be built accordingly and observability so that we can make sure that we are not violating the SLAs that have been signed. Underlying on the platform, I've shown at our super high level that we have these different

5:52

inference engines, VLLMs, SGLangs and TensorRDLL, but there's quite a bit of detail over here that I'll touch up on. And underlying that, what I'm trying to show over here in green is various different pieces of hardware. So it's not that, so the platform needs to be capable enough of sharing, of having the workload getting distributed across various different generations of these GPUs, specifically in media GPUs that we use. So let's take a few examples over here. Let's say the request originates from the client side through apps or notebooks, any of those, or through the agents. It hits the gateway. Once it hits the gateway, then like I mentioned on the

6:40

control plane, goes through authentication, etc., etc., etc., but then comes either the serverless or dedicated. So in the case of serverless, it will be paper tokens, so the token usage would be monitored over here. Not the exact tokens, but just the token usage, because we maintain ZDR, zero data retention policies. It is, depending on the multi-tenancy or the provisioned, if it is provisioned, then we know underlying for the router, it needs to go in and target the explicit deployments for the provisioned throughput customers. For the multi-tenant customers, there are separate deployments.

7:17

Router over here, specifically, the router is very important, since the router is responsible for making KV cache-aware routing choices. Why is it important? Because like I mentioned when we were discussing the workload profiles, the agentic use cases are typically super heavy on the input sequence lengths and bulk of the input sequence length, about 80 to 90 percent, depending on which company it is, depending on the customers, 80 to 90 percent of it is the same for various different requests. So there is no point in going in and re-computing the pre-fill or redoing the pre-fill for that.

7:57

Pre-fill is super compute bound, very expensive, that's why as much as you can hit the cache, more you can save, which is why if you look at the token pricing anywhere, there's a specific price for input tokens and there's a way cheaper price for the cache input tokens. So caching becomes like really important over here. Underlying, how you want to split the hardware is totally dependent on the choice in the platform and we provide the capability to do either. Either do a pre-fill decode disaggregation if the use case desires it or do not do it because pre-fill decode disaggregation is not cheap for every

8:39

type of use case. Let's take another request flow over here. Let's see if when it was a dedicated customer, then what will happen. A dedicated customer, again, will go through the gateways that have been set up for them with proper isolations. Billing is not based on tokens. Billing is based on usage of per GPU per hour. It's a private gateway so that there is no noisy neighbor problem. No one else can get in. Same router logic over here so that if there are requests which are very similar, then it hits the cache most. Depending on the deployment that the customer makes on their dedicated GPUs, they can decide if they

9:25

want to do pre-fill decode disaggregation or not. They can decide which engine to use. VLM or SG-LANG or TENSOR RTLM. Given the bulk of capacity that the customer has reserved, they can decide if they want to have just one deployment with the ability to scale through the whole cluster or they want to have multiple different models, multiple different deployments. One thing that I do want to mention about the router over here is the fact that heterogeneous capacity across different zones and regions is supported. It is a little it's quite a bit of a hard problem to load balance across that so the priority order that we

10:07

typically take is first KV cache locality and then the least loaded fallback.

10:16

That's that. Another request flow that I want to go over here which might be a little hard to see from the diagram is I want to take the batch workflow. For the batch workflow what we would actually do is the underlying capacity that the customer has, let's say the same dedicated inference customer, during US daytime they're running their real-time workloads and from evening to night they want to run batch workloads, the same capacity after time can be scheduled to run the batch workloads. So we provide the capability in the API to tell when to scale up and when to scale down and as per

10:52

schedule if we can if they tell us that we can have to scale down we will scale down and open it up for batch processing for the night.

11:04

I think I've spoken quite a bit about optimizations on the KV cache side but I do want to repeat a little bit because this is one of the most interesting pieces. If we can hit on the cache more you can you avoid the cost of pre-fill which is the most expensive piece over here. Reusing the KV cache across multiple different turns in your agentic workloads between turns also there is a lot of similar pre-fill that comes in in the input sequence length. Think about the chat workloads which is where offloading KV cache also becomes extremely important because with cache with the chat workloads we have a lot of latency between

11:48

different between multiple turns that we as users put in but if we completely evict whatever we had in our particular conversation then the next time we ask a question in the same chat it's going to take a little bit longer. So instead of actually completely evicting and redoing the pre-fill again what the the techniques being used are maybe using we are using our own but externally we know about LM cache and Mooncakes the pre-fill again what we do is we will offload the KV cache to a high bandwidth storage so that we can store a lot of these pre-fills such that whenever the accompanying request comes for that particular

12:29

conversation it can be loaded in right away into the HBM.

12:37

On the performance liver I would I just want to mention a few performance livers that we've discussed the period disag that is one but quantization and specular decoding are others and how to carefully choose the parallelization degrees and the strategies that is actually very important. Two of the biggest livers that we have been working with are quantization to NVFE4 and spec deck. We do provide capability where if the customer has their data set and they want us to train speculators for their data sets for better acceptance lens which will ultimately make the output throughput significantly higher. We do have that as well so but that happens async we

13:19

we gather data async we train the speculators async and then we deploy the speculators into the customer deployments if that's what they wanted. You see three screenshots over here I have posted them from the last one month one month's worth of work that some of us in my team have done. You can see we came quickly on top of the leaderboard on Kimi 2.6, 2.7 and those are those are from artificial analysis and going back to the session before this can we trust that that's why for GLM I have the results from open router. So artificial analysis when they run benchmarks they're running very specific workloads.

13:58

Open router is actual user traffic and you can see on the open router side weights and biases so the branding is different but weights and biases basically curve we bought weights and biases about a year back. You can see the speed over here that we have from our deployment is pretty close to what Fireworks is providing as Fireworks fast right but underlying techniques that we are using is what I want to emphasize the most over here for performance optimization. That becomes critical because ultimately what you want to serve to the customer what we want to serve to the customer is price performance benefit.

14:36

Quick recap, single platform is what I've been trying to emphasize is what I'm trying to show two different consumption models serverless and dedicated for customers and within serverless I describe two different consumption models as well pay as you go and provision throughput if you care about that and ultimately compounding the gains through performance optimizations in the stack. That's all thank you folks.

15:19

I you

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note