Good afternoon everybody. I'm Sitan Shu from Corvi. I'm going to be talking about vertical mobility. It's a fancy topic, the title that we came up with, but going to be talking about the inference platform that we have at Corvi that we are building to serve small to big models and various different types of workloads. A quick intro about me, I joined Corvi just about four months back, leading all of inference over there and before this I was managing everything at AWS Annapurna Labs for training and before that inference and training at Sambanova. So quite a bit of experience in this particular space. The way I'll be taking you through is explaining to you the consumption models that we have and from that how we have derived what the platform should look like so that we do not need to keep changing the platform and we keep making enhancements in the platform that we have for serving inference and how and why performance plays such an important role over there. I think a little bit of this might be common with the previous topic that was discussed over here. Consumption models. So we have at large two biggest consumption models. One is serverless, which is where the customers can come in, consumers can come in, do not need to worry about managing the hardware themselves, do not need to worry about managing the clusters, orchestration, anything at all. There's API, there's UI, you come in, you pay per token and you get your model served. Biggest thing over here is that the type of models that we serve in the catalog, that is the breadth of models that the customer will be able to go through. I'll talk about dedicated and then I'll come back to serverless because there is one thing unique on the serverless side. Dedicated inference service that we provide is more for customers who want to know exactly what hardware they are going to be using and running on. But the model deployment also depends on them. So they use our service, they use our orchestration layers, but the model deployment depends on them. Model performance also depends on them as long as we provide in the platform the capability and the knobs to serve those features. Coming back to serverless, one of the interesting pieces over here is typically serverless models are noisy neighbor problems where if let's say everyone is banging on the exact same model, then you might be timing out quite a bit depending on how much capacity I have behind it. So another feature that we have on the serverless side is what we are calling a provisioned throughput. So as a customer, if you know your traffic profile and if you can let us know about that, we can carve it out specifically for you behind the scenes. You still do not need to worry about what hardware it is exactly running on as long as your throughput, your SLAs are maintained. So that is another one on the serverless side and that is still charged by the token, but you know that you are not running into the noisy neighbor problem over there. Let me take a quick stab at a few different types of workloads, workload shapes that we have, that we are seeing and the ratio between these is continuously changing though agentic is really high up there. Agentic and chat are very similar, super high on the input sequence lengths, very low on the output sequence lengths typically, but the biggest difference between agentic and chat being the fact that the multi-turns in agentic are super low latency versus in chats because when you get the response of the user, you have to read the answer and then you respond to it. So there are differences over there and that big difference ultimately converts into something related to the KV cache management. But these two are both real-time and another real-time workload is your voice and videos, which are steady streaming and super latency sensitive. On the agentic and chat side, largely the requirements are from throughput point of view, not so much from latency, but real-time voice and videos are absolutely latency sensitive. Coming to batch, batch is where the SLAs are super loose, they run into seconds and minutes, sometimes for some customers actually even in hours, they're like I'll just throw, give me 10 to 12 hours of workload capability and I'll throw whatever I can, process it whenever you can. These batch workloads, the way they come into the picture over here in deciding, sorry, being the requirement for some of the design choices that we make in the stack, imagine these four different types of workload shapes. In the time dimension, you have to play the game of Tetris on how you can fit it in to utilize the underlying infrastructure the most. I'll give a high level on how our stack is shaped right now and I'll walk you through a bit of a request flow over here. So for both serverless and dedicated, if you look at the right hand side of the screen, you'll see that on the platform side, you'll go to the control plane to have your authorizations, your rate limitings and your usage being tracked, etc. so that you can be built accordingly and observability so that we can make sure that we are not violating the SLAs that have been signed. Underlying on the platform, I've shown at our super high level that we have these different inference engines, VLLMs, SGLangs and TensorRTLM, but there's quite a bit of detail over here that I'll touch up on. And underlying that, what I'm trying to show over here in green is various different pieces of hardware. So it's not that, so the platform needs to be capable enough of sharing, of having the workload getting distributed across various different generations of these GPUs, specifically in media GPUs that we use. So let's take a few examples over here. Let's say the request originates from the client side through apps or notebooks, any of those, or through the agents. It hits the gateway. Once it hits the gateway, then as I mentioned on the control plane, goes through authentication, etc., but then comes either the serverless or dedicated. So in the case of serverless, it will be paper tokens, so the token usage would be monitored over here. Not the exact tokens, but the token usage, because we maintain ZDR, zero data retention policies. It is, depending on the multi-tenancy or the provisioned, if it is provisioned, then we know underlying for the router, it needs to go in and target the explicit deployments for the provisioned throughput customers. For the multi-tenant customers, there are separate deployments. Router over here, specifically, the router is very important, since the router is responsible for making KV cache-aware routing choices. Why is it important? Because as I mentioned when we were discussing the workload profiles, the agentic use cases are typically super heavy on the input sequence lengths and bulk of the input sequence length, about 80 to 90 percent, depending on which company it is, depending on the customers, 80 to 90 percent of it is the same for various different requests. So there is no point in going in and re-computing the pre-fill or redoing the pre-fill for that. Pre-fill is super compute bound, very expensive, that's why as much as you can hit the cache, more you can save, which is why if you look at the token pricing anywhere, there's a specific price for input tokens and there's a way cheaper price for the cache input tokens. So caching becomes really important over here. Underlying, how you want to split the hardware is totally dependent on the choice in the platform and we provide the capability to do either. Either do a pre-fill decode disaggregation if the use case desires it or do not do it because pre-fill decode disaggregation is not cheap for every type of use case. Let's take another request flow over here. Let's see if when it was a dedicated customer, then what will happen. A dedicated customer, again, will go through the gateways that have been set up for them with proper isolations. Billing is not based on tokens. Billing is based on usage of per GPU per hour. It's a private gateway so that there is no noisy neighbor problem. No one else can get in. Same router logic over here so that if there are requests which are very similar, then it hits the cache most. Depending on the deployment that the customer makes on their dedicated GPUs, they can decide if they want to do pre-fill decode disaggregation or not. They can decide which engine to use. VLM or SG-LANG or TENSOR RTLM. Given the bulk of capacity that the customer has reserved, they can decide if they want to have just one deployment with the ability to scale through the whole cluster or they want to have multiple different models, multiple different deployments. One thing that I do want to mention about the router over here is the fact that heterogeneous capacity across different zones and regions is supported. It is quite a bit of a hard problem to load balance across that so the priority order that we typically take is first KV cache locality and then the least loaded fallback. That's that. Another request flow that I want to go over here which might be a little hard to see from the diagram is I want to take the batch workflow. For the batch workflow what we would actually do is the underlying capacity that the customer has, let's say the same dedicated inference customer, during US daytime they're running their real-time workloads and from evening to night they want to run batch workloads, the same capacity after time can be scheduled to run the batch workloads. So we provide the capability in the API to tell when to scale up and when to scale down and as per schedule if we can if they tell us that we can have to scale down we will scale down and open it up for batch processing for the night. I think I've spoken quite a bit about optimizations on the KV cache side but I do want to repeat a little bit because this is one of the most interesting pieces. If we can hit on the cache more you can avoid the cost of pre-fill which is the most expensive piece over here. Reusing the KV cache across multiple different turns in your agentic workloads between turns also there is a lot of similar pre-fill that comes in in the input sequence length. Think about the chat workloads which is where offloading KV cache also becomes extremely important because with cache with the chat workloads we have a lot of latency between different between multiple turns that we as users put in but if we completely evict whatever we had in our particular conversation then the next time we ask a question in the same chat it's going to take a little bit longer. So instead of actually completely evicting and redoing the pre-fill again what the techniques being used are maybe using we are using our own but externally we know about LM cache and Mooncakes the pre-fill again what we do is we will offload the KV cache to a high bandwidth storage so that we can store a lot of these pre-fills such that whenever the accompanying request comes for that particular conversation it can be loaded in right away into the HBM. On the performance lever I would I just want to mention a few performance levers that we've discussed the period disag that is one but quantization and specular decoding are others and how to carefully choose the parallelization degrees and the strategies that is actually very important. Two of the biggest levers that we have been working with are quantization to NVFE4 and spec deck. We do provide capability where if the customer has their data set and they want us to train speculators for their data sets for better acceptance lens which will ultimately make the output throughput significantly higher. We do have that as well so but that happens async we gather data async we train the speculators async and then we deploy the speculators into the customer deployments if that's what they wanted. You see three screenshots over here I have posted them from the last one month one month's worth of work that some of us in my team have done. You can see we came quickly on top of the leaderboard on Kimi 2.6, 2.7 and those are from artificial analysis and going back to the session before this can we trust that that's why for GLM I have the results from open router. So artificial analysis when they run benchmarks they're running very specific workloads. Open router is actual user traffic and you can see on the open router side weights and biases so the branding is different but weights and biases basically curve we bought weights and biases about a year back. You can see the speed over here that we have from our deployment is pretty close to what Fireworks is providing as Fireworks fast but underlying techniques that we are using is what I want to emphasize the most over here for performance optimization. That becomes critical because ultimately what you want to serve to the customer what we want to serve to the customer is price performance benefit. Quick recap, single platform is what I've been trying to emphasize is what I'm trying to show two different consumption models serverless and dedicated for customers and within serverless I describe two different consumption models as well pay as you go and provision throughput if you care about that and ultimately compounding the gains through performance optimizations in the stack. That's all thank you folks.
mobility. It's a fancy topic, the title that we came up with, but basically going to be talking about the inference platform that we have at Corvi that we are building to serve small to big models and various different types of workloads. A quick intro about me, I joined Corvi just about four months back, leading all of inference over there and before this I was managing everything at AWS Annapurna Labs for training and before that inference and training at Sambanova. So quite a bit of experience in this particular space. The way I'll be taking you through is explaining to you
the consumption models that we have and from that how we have derived what the platform should look like so that we do not need to keep changing the platform and we keep making enhancements in the platform that we have for serving inference and how and why performance plays such an important role over there. I think a little bit of this might be common with the previous topic that was discussed over here. Consumption models. So we have at large two biggest consumption models. One is a serverless, which is where the customers can come in, consumers can come in, do not need to worry about managing the
hardware themselves, do not need to worry about managing the clusters, orchestration, anything at all. There's API, there's UI, you come in, you pay per token and you get your model served. Biggest thing over here is that the type of models that we serve in the catalog, that is the breadth of models that the customer will be able to go through. I'll talk about dedicated and then I'll come back to serverless because there is one thing unique on the serverless side. Dedicated inference service that we provide is more for customers who want to know exactly what hardware they are going to be using and running on.
But the model deployment also depends on them. So they use our service, they use our orchestration layers, but the model deployment depends on them. Model performance also depends on them as long as we provide in the platform the capability and the knobs to serve those features. Coming back to serverless, one of the interesting pieces over here is typically serverless models are noisy neighbor problems where if let's say everyone is banging on the exact same model, then you might be timing out quite a bit depending on how much capacity I have behind it. So another feature that we have on the serverless side is what we are calling
a provisioned throughput. So as a customer, if you know your traffic profile and if you can let us know about that, we can carve it out specifically for you behind the scenes. You still do not need to worry about what hardware it is exactly running on as long as your throughput, your SLAs are maintained. So that is another one on the serverless side and that is still charged by the token, but you know that you are not running into the noisy neighbor problem over there. Let me take a quick stab at a few different types of workloads, workload shapes that we have, that we are seeing and the ratio between
these is like continuously changing though agentic is like really high up there. Agentic and chat kind of very similar, super high on the input sequence lengths, very low on the output sequence lengths typically, but the biggest difference between agentic and chat being the fact that the multi-turns in agentic are super low latency versus in chats because when you get the response of the user, you have to read the answer and then you respond to it. So there are differences over there and that big difference ultimately converts into something related to the KV cache management. But these two are both real-time
and another real-time workload is your voice and videos, which are steady streaming and super latency sensitive. On the agentic and chat side, largely the requirements are from throughput point of view, not so much from latency, but real-time voice and videos are absolutely latency sensitive. Coming to batch, batch is where the SLAs are like super loose, they run into like seconds and minutes, sometimes for some customers actually even in hours, they're like I'll just throw, give me 10 to 12 hours of workload capability and I'll throw whatever I can, process it whenever you can. These batch workloads, the way they come into the picture
over here in deciding, sorry, being the requirement for some of the design choices that we make in the stack, imagine these four different types of workload shapes. In the time dimension, you have to play the game of petris on how you can fit it in to utilize the underlying infrastructure the most.
I'll give a high level on how our stack is shaped right now and I'll walk you through a bit of a request flow over here. So for both serverless and dedicated, if you look at the right hand side of the screen, you'll see that on the platform side, you'll go to the control plane to have your authorizations, your rate limitings and your usage being tracked, etc. so that you can be built accordingly and observability so that we can make sure that we are not violating the SLAs that have been signed. Underlying on the platform, I've shown at our super high level that we have these different
inference engines, VLLMs, SGLangs and TensorRDLL, but there's quite a bit of detail over here that I'll touch up on. And underlying that, what I'm trying to show over here in green is various different pieces of hardware. So it's not that, so the platform needs to be capable enough of sharing, of having the workload getting distributed across various different generations of these GPUs, specifically in media GPUs that we use. So let's take a few examples over here. Let's say the request originates from the client side through apps or notebooks, any of those, or through the agents. It hits the gateway. Once it hits the gateway, then like I mentioned on the
control plane, goes through authentication, etc., etc., etc., but then comes either the serverless or dedicated. So in the case of serverless, it will be paper tokens, so the token usage would be monitored over here. Not the exact tokens, but just the token usage, because we maintain ZDR, zero data retention policies. It is, depending on the multi-tenancy or the provisioned, if it is provisioned, then we know underlying for the router, it needs to go in and target the explicit deployments for the provisioned throughput customers. For the multi-tenant customers, there are separate deployments.
Router over here, specifically, the router is very important, since the router is responsible for making KV cache-aware routing choices. Why is it important? Because like I mentioned when we were discussing the workload profiles, the agentic use cases are typically super heavy on the input sequence lengths and bulk of the input sequence length, about 80 to 90 percent, depending on which company it is, depending on the customers, 80 to 90 percent of it is the same for various different requests. So there is no point in going in and re-computing the pre-fill or redoing the pre-fill for that.
Pre-fill is super compute bound, very expensive, that's why as much as you can hit the cache, more you can save, which is why if you look at the token pricing anywhere, there's a specific price for input tokens and there's a way cheaper price for the cache input tokens. So caching becomes like really important over here. Underlying, how you want to split the hardware is totally dependent on the choice in the platform and we provide the capability to do either. Either do a pre-fill decode disaggregation if the use case desires it or do not do it because pre-fill decode disaggregation is not cheap for every
type of use case. Let's take another request flow over here. Let's see if when it was a dedicated customer, then what will happen. A dedicated customer, again, will go through the gateways that have been set up for them with proper isolations. Billing is not based on tokens. Billing is based on usage of per GPU per hour. It's a private gateway so that there is no noisy neighbor problem. No one else can get in. Same router logic over here so that if there are requests which are very similar, then it hits the cache most. Depending on the deployment that the customer makes on their dedicated GPUs, they can decide if they
want to do pre-fill decode disaggregation or not. They can decide which engine to use. VLM or SG-LANG or TENSOR RTLM. Given the bulk of capacity that the customer has reserved, they can decide if they want to have just one deployment with the ability to scale through the whole cluster or they want to have multiple different models, multiple different deployments. One thing that I do want to mention about the router over here is the fact that heterogeneous capacity across different zones and regions is supported. It is a little it's quite a bit of a hard problem to load balance across that so the priority order that we
typically take is first KV cache locality and then the least loaded fallback.
That's that. Another request flow that I want to go over here which might be a little hard to see from the diagram is I want to take the batch workflow. For the batch workflow what we would actually do is the underlying capacity that the customer has, let's say the same dedicated inference customer, during US daytime they're running their real-time workloads and from evening to night they want to run batch workloads, the same capacity after time can be scheduled to run the batch workloads. So we provide the capability in the API to tell when to scale up and when to scale down and as per
schedule if we can if they tell us that we can have to scale down we will scale down and open it up for batch processing for the night.
I think I've spoken quite a bit about optimizations on the KV cache side but I do want to repeat a little bit because this is one of the most interesting pieces. If we can hit on the cache more you can you avoid the cost of pre-fill which is the most expensive piece over here. Reusing the KV cache across multiple different turns in your agentic workloads between turns also there is a lot of similar pre-fill that comes in in the input sequence length. Think about the chat workloads which is where offloading KV cache also becomes extremely important because with cache with the chat workloads we have a lot of latency between
different between multiple turns that we as users put in but if we completely evict whatever we had in our particular conversation then the next time we ask a question in the same chat it's going to take a little bit longer. So instead of actually completely evicting and redoing the pre-fill again what the the techniques being used are maybe using we are using our own but externally we know about LM cache and Mooncakes the pre-fill again what we do is we will offload the KV cache to a high bandwidth storage so that we can store a lot of these pre-fills such that whenever the accompanying request comes for that particular
conversation it can be loaded in right away into the HBM.
On the performance liver I would I just want to mention a few performance livers that we've discussed the period disag that is one but quantization and specular decoding are others and how to carefully choose the parallelization degrees and the strategies that is actually very important. Two of the biggest livers that we have been working with are quantization to NVFE4 and spec deck. We do provide capability where if the customer has their data set and they want us to train speculators for their data sets for better acceptance lens which will ultimately make the output throughput significantly higher. We do have that as well so but that happens async we
we gather data async we train the speculators async and then we deploy the speculators into the customer deployments if that's what they wanted. You see three screenshots over here I have posted them from the last one month one month's worth of work that some of us in my team have done. You can see we came quickly on top of the leaderboard on Kimi 2.6, 2.7 and those are those are from artificial analysis and going back to the session before this can we trust that that's why for GLM I have the results from open router. So artificial analysis when they run benchmarks they're running very specific workloads.
Open router is actual user traffic and you can see on the open router side weights and biases so the branding is different but weights and biases basically curve we bought weights and biases about a year back. You can see the speed over here that we have from our deployment is pretty close to what Fireworks is providing as Fireworks fast right but underlying techniques that we are using is what I want to emphasize the most over here for performance optimization. That becomes critical because ultimately what you want to serve to the customer what we want to serve to the customer is price performance benefit.
Quick recap, single platform is what I've been trying to emphasize is what I'm trying to show two different consumption models serverless and dedicated for customers and within serverless I describe two different consumption models as well pay as you go and provision throughput if you care about that and ultimately compounding the gains through performance optimizations in the stack. That's all thank you folks.
I you