[SPEAKER_01] Audrey, I am from RunPod.
SPEAKER_01
This is an intro to RunPod. Can I just get a quick hand to see how many people have already heard of RunPod or maybe even used RunPod before? Okay, newbies for everybody. Great. So, RunPod, we are a cloud AI infrastructure company. So, we have the hardware, we have the GPUs, and we make it easy for developers to deploy models. And that can be your own private model. It can be an open source model from Hugging Face. It doesn't matter to us. You bring your code and we'll bring the rest. Just really quickly, what problems does RunPod solve? Why are we even here today? Infrastructure can be hard, managing it.
SPEAKER_01
I think about back in the day before we had AWS, Google Cloud, when everybody would have to have on-prem servers and manage those, maintain those. That is something that we don't want to have to do as developers. Those are things that we happily have moved away from and given off to dev ops, and now it's even more abstracted for us.
SPEAKER_01
GPU access is slow and opaque. So, I don't know if anybody has tried to buy a GPU recently. We're in a global supply crunch. It's a bit like in COVID when everybody went to the store and bought all the toilet paper because we didn't know how long they would need to be at home for. We're a little bit in that right now, but we expect the market will recover as customers, companies, people figure out a little bit better, get a little bit better at estimating what kind of compute they need. And then last, builder primary focus should be building. So, again, we want to build apps.
SPEAKER_01
We as software developers, we bring the value through the applications that we build, not for managing the infrastructure. And I think RunPod has a pretty unique story. These are our founders, Zen and Pardeep. So, they had a couple of GPU rigs in their basement in 2022, failed crypto mining, and then so they were thinking, what are we going to do with our GPUs now? So, they prototyped what is now the foundations of RunPod. They posted on Reddit and said, hey, anyone want to use these GPUs for free? Just give us feedback on it. And that is literally how our company has started, and we have been revenue generating ever since.
SPEAKER_01
And the reason why I want to tell this story is not because it's bootstrappy, but because the origin story of RunPod has always started with builders and getting feedback from the community, and that is still true today. So, I won't promise that we'll be perfect, but we are definitely very engaged with our users on Reddit, on Discord, so we're always trying to stay engaged with you all. Just at a glance, to give you an idea of RunPod, we have over 500,000 developers on our platform, 30 plus data centers across the world, including Europe and the EU, and we've just passed a significant revenue milestone for us, 120 million in annual recurring revenue.
SPEAKER_01
These are just a few of our customers. You might be surprised to see some of the AI Cloud native companies on here too, but they come to us for the same reasons that most of our customers come to us. It's because they need flexible and reliable GPU infrastructure. This is a really high-level overview of different ways you can build on RunPod. So, I would say at our core, pods, it's our sandbox virtual environment. We spin up a container for you, allocate GPUs to it, and we manage the rest. So, you just bring your Docker files, you bring your code. Serverless, it's our auto-scaling product.
SPEAKER_01
So, when you're thinking more about bursty workloads or batch workloads, serverless is really great because instead of being always on like a container is, serverless, your workers spin down, and when they're idle, you don't pay for anything. Clusters, if you're doing some heavy-duty training, there's a place for you as well on RunPod. Multi-node clusters with high-speed networking. And then the hub, which I'll switch to in a second, it's our central repository for AI repos. These are already pre-configured, pre-vetted. We have a couple of examples of listings by RunPod for popular models, but also our community contributes to them as well.
SPEAKER_01
So, they're repos that you can fork, you can watch, and then you can star and deploy on RunPod. So today we're going to be talking mostly about serverless.
SPEAKER_01
So serverless is best for real-time inference. I talked about the auto-scaling that comes with it. Why teams use it is mostly because they don't need to preempt and figure out how much compute they need ahead of time. You can set you can configure the number of max workers that you want to scale up to.
SPEAKER_01
You can set limits for caps, for spending caps, and you can also configure workers that are always on. So they already have your models downloaded and they can respond to requests immediately. For a lot of teams, serverless is the fastest way if you want to start deploying a production-ready API. And now I'm going to switch over and just show you really quick how easy it is to get started and deploy something. Okay. Where are we? Okay. So right now I'm going to do everything via the console so that it's nice and pretty for you guys to see. But we also have CLI support. We have skills to help work with RunPod.
SPEAKER_01
Everything that's ready for your agent so you don't have to read our documents. But since we're all humans here today, I'm going to show you via the console. We'll start in the hub, which is if you're just trying to explore and see what's out there, what is something that you can get up and running right now, the hub is a great place to start. So as I mentioned, these are already vetted open source listings for AI repos. And I am going to pick the LLM, and I'll just open the underlying repository as well so you guys can see what- It is literally just a GitHub repo. So it tells you how to get set up for it. We can see there's already the Docker file here.
SPEAKER_01
It's already pre-configured for you. It's got some defaults for you, depending on the listing. You can pass in different environmental variables to configure it how you wish. But I'm just going to go ahead and click deploy. And I have a model that I wanted. Let me see. I was going to just pick Gwen.
SPEAKER_01
Works well.
[SPEAKER_01] This is going to download it from Hugging Face and just expand the advanced options and look for the max model length.
SPEAKER_01
And I'm going to bump this up for the context window and leave everything else as the defaults.
SPEAKER_01
But there are settings for max loras. All of these configuration options get passed as flags to the VLM serve. And I'm going to spin it up as an endpoint here. So this might take a minute or two since this is the very first time I just created it. We've got to initialize my workers. Let's check out. So the default configuration here is it's going to deploy on some H100s and A100s are the backup here. I have my pricing. This is a fraction of a cent per second. As I mentioned before, this is only going to be charged for while the worker is actually running and handling a request.
SPEAKER_01
Max workers is where I can bump this up if I want to have my workload scale up to 15 workers at a time. And I can set some active workers once that I want always to be on that I don't want the container to ever spin down. And I can save that. Okay, so how does one interact with the serverless endpoint? This is just an API HTTP endpoint right here. We provisioned this endpoint for you.
SPEAKER_01
You can send requests to this. Your customers can send requests to this.
SPEAKER_01
If I just hit run and I'm going to add a few, let's, what should we ask the LLM today? Does anyone have a suggestion? Okay. I'm American. So how did Big Ben get its name? I don't know. Well, these requests are queued.
SPEAKER_01
Let me check on our workers.
SPEAKER_01
Okay. We have a handful that are initializing. This is the containers being created. That's the model being downloaded. Getting ready and the ones that are running, they've already finished. These are probably going to be the ones who are going to pick up those requests that we just added. I've got telemetry about it's blank right now, but the number of requests, execution time, delay time. So you have observability into how your endpoints are operating. And, let's see.
SPEAKER_01
Okay.
SPEAKER_01
It's already done. Got a request back in. It sat in the queue for about 41 seconds. That's going to be a little bit longer than all of the subsequent requests because of some of the cold start time that I talked about, like downloading the model, initializing the first container. But execution time, only about one and a half seconds. So, that was probably less than five minutes to get started and get something deployed on serverless from a hub listing. Does anyone have any questions?
SPEAKER_01
This is a very short and sweet intro. We have another session later today at four o'clock. And that one is going to be focused on our Python flash SDK. And that one is going to be completely via the terminal.
SPEAKER_01
And I'm going to walk you through how I can spin up and deploy my code as a remote function onto a GPU, and deploy in the end and make it a production ready endpoint here. And I'm going to go ahead and get a little bit more as well. Okay. But that's all I got for today. So, thanks. Thanks for coming. That's the model being downloaded. Um, getting ready and the ones that are running, they've already finished. These are probably gonna be the ones who are gonna pick up those requests that we just added. I've got telemetry about, um, it's blank right now, but the number of requests, execution time, delay time. So you have observability into how your endpoints are operating.
SPEAKER_01
And, let's see. Okay. It's already done.
SPEAKER_01
Got a request back in. It sat in the queue for about 41 seconds. Um, that's going to be a little bit longer than all of the subsequent requests because of some of the cold start time that I talked about, like downloading the model, um, initializing the first container. But, um, execution time, only about one and a half seconds. So, yeah, that was probably less than five minutes to get started and get something deployed, um, on serverless from a hub listing.
SPEAKER_01
Does anyone have any questions? This is, this is a very short and sweet intro. Um, we have another session later today at four o'clock. Um, and that one is gonna be focused on our Python flash, um, SDK. And that one is going to be completely via the terminal. Um, and I'm gonna walk you through how I can spin up, uh, and deploy my code on my code as a remote, remote function onto a GPU, um, and, uh, deploy in the end and make it like a production ready endpoint here. And I'm gonna go ahead and get a little bit more as well.
SPEAKER_01
Okay. But that's all I got for today. So, yeah, thanks. Thanks for coming. .