Under 5 minutes to a deployed LLM endpoint — Audry Hsu, RunPod
Description
Two failed crypto mining rigs in a basement in 2022. The founders posted on Reddit offering the GPUs for free in exchange for feedback. That is the origin of RunPod, now at $120 million in annual recurring revenue with 500,000 developers on the platform. The demo runs in under five minutes: pick a model from the Hub, configure a context window, deploy a serverless endpoint on H100s. First request queues for 41 seconds on cold start while the container initializes and the model downloads. Every request after that executes in about 1.5 seconds. You pay only while a worker is handling a request. Speaker info: - https://www.linkedin.com/in/audry-hsu/
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Skim
- Core thesis: RunPod provides rapid GPU deployment for LLM endpoints via a hub of pre-configured repos, serverless auto-scaling, and a pay-per-second model targeting developers who want infrastructure abstracted away.
- Why it matters: Ken can deploy production LLM APIs in under 5 minutes without managing GPU infrastructure, paying only for active inference time, which streamlines agent/AI ops prototyping and GTM testing.
- Best use: Reference for quick LLM endpoint deployment tactics, pricing model (fraction of cent per second), and understanding RunPod's serverless vs. pods vs. clusters product tiers.
Executive Summary
Audrey Hsu from RunPod delivers a live demo showing deployment of an LLM endpoint (Qwen model from Hugging Face) in under 5 minutes using RunPod's serverless product. The company positions itself as a cloud AI infrastructure provider solving three pain points: difficulty managing on-prem GPU infrastructure, slow and opaque GPU access due to global supply constraints, and developer distraction from core application building. RunPod's origin story—two founders with failed crypto mining rigs in 2022 who posted on Reddit offering free GPU access for feedback—anchors their builder-first community engagement strategy. The platform now serves 500,000+ developers, operates 30+ data centers globally, and reports $120M ARR.
The demo walks through RunPod's Hub (a repository of vetted, pre-configured AI repos), showing how to deploy a Qwen LLM via console by selecting a listing, adjusting environment variables (max model length for context window), and spinning up workers on H100/A100 GPUs. The serverless product charges only when workers handle requests (fraction of a cent per second), supports auto-scaling to configurable max workers, and offers 'active workers' that remain always-on to eliminate cold starts. Initial request took ~41 seconds (including model download and container initialization), while execution time was 1.5 seconds. Subsequent requests would be faster due to warm containers.
RunPod's product suite includes Pods (sandbox virtual environments with persistent containers), Serverless (auto-scaling for bursty/batch workloads), Clusters (multi-node with high-speed networking for training), and the Hub (community-contributed repos). Audrey emphasizes serverless is best for real-time inference and production-ready APIs where teams don't want to pre-estimate compute needs. The platform provides telemetry (request count, execution time, delay time), spending caps, and CLI/SDK support. A follow-up session at 4pm was teased covering Python Flash SDK for deploying code as remote functions on GPUs via terminal.
Key Takeaways
- Claim: RunPod can deploy a production-ready LLM endpoint in under 5 minutes using pre-configured Hub listings. | Evidence: Live demo deployed Qwen model from Hugging Face by selecting a Hub listing, adjusting max model length, and spinning up H100/A100 workers; first request returned in ~41 seconds (including cold start), execution time 1.5 seconds. | Caveat: Initial cold start included model download and container initialization time (~41 seconds queue time); subsequent requests would be faster. No discussion of model size limits, token throughput benchmarks, or failure handling. | Implication: Ken can rapidly prototype LLM APIs for agent systems or content workflows without GPU procurement or infrastructure setup, paying only for active inference time. | Timestamp: 05:30
- Claim: RunPod's serverless pricing is 'fraction of a cent per second' and charges only when workers actively handle requests, not idle time. | Evidence: Audrey showed pricing in console during demo; serverless workers spin down when idle, eliminating costs during downtime. Platform supports configurable max workers and 'active workers' that remain always-on. | Caveat: No specific pricing examples given (e.g., cost per 1M tokens, cost per H100 worker-second). 'Active workers' that stay always-on would incur continuous charges even if idle. | Implication: Ken should evaluate serverless for bursty agent workloads (e.g., batch content generation, on-demand inference) where idle time savings offset cold start delays; compare to always-on pods for latency-critical use cases. | Timestamp: 04:15
- Claim: RunPod's Hub contains vetted, pre-configured AI repos (Docker files, environment variables, vLLM serve flags) contributed by RunPod and the community. | Evidence: Demo showed selecting an LLM listing that was 'literally just a GitHub repo' with defaults for vLLM serve, environment variables for max_loras and max_model_length, and one-click deploy to serverless endpoint. | Caveat: No information on vetting process rigor, versioning, security audits, or what happens if a community-contributed repo has bugs or supply chain issues. Unclear how often repos are updated for new model releases. | Implication: Ken can fork and customize Hub repos for rapid deployment but should audit dependencies and security before production use, especially for community-contributed listings. | Timestamp: 03:45
- Claim: RunPod was founded in 2022 by two founders with failed crypto mining rigs who posted on Reddit offering free GPU access for feedback, achieving revenue from day one. | Evidence: Audrey stated founders Zen and Pardeep had GPU rigs in basement, posted on Reddit, got user feedback, and have been 'revenue generating ever since'; now $120M ARR and 500,000+ developers. | Caveat: No financial details on burn rate, profitability, or unit economics. Community-first origin story may not reflect current go-to-market or support capacity at scale. | Implication: Ken should expect strong community engagement (Reddit, Discord) but verify SLA/support tiers for production workloads; company may prioritize developer-friendly features over enterprise reliability guarantees. | Timestamp: 01:30
- Claim: GPU supply crunch is compared to COVID toilet paper hoarding, but market expected to recover as customers improve compute estimation. | Evidence: Audrey stated 'we're in a global supply crunch... a bit like in COVID when everybody went to the store and bought all the toilet paper' and expects recovery as companies 'get better at estimating what kind of compute they need.' | Caveat: No timeline given for market recovery. No discussion of how RunPod sources GPUs (owned vs. community-sourced), whether they face allocation constraints, or how they prioritize customers during shortages. | Implication: Ken should ask RunPod about GPU availability guarantees and failover to backup GPU types (demo showed H100s with A100 backup) before committing production workloads; supply risk may affect scaling plans. | Timestamp: 01:15
- Claim: RunPod supports CLI, SDKs, and 'skills' for agents to interact with the platform without reading documentation. | Evidence: Audrey mentioned 'we have CLI support, we have skills to help work with RunPod, everything that's ready for your agent so you don't have to read our documents' but used console for demo clarity. | Caveat: No details on what 'skills' means (OpenAI function calling schemas? MCP servers? custom API wrappers?), SDK language support beyond Python Flash, or agent integration examples. | Implication: Ken should investigate RunPod's agent-friendly tooling for agentic workflows that auto-provision GPU endpoints, potentially enabling AI-native GTM where agents deploy their own infrastructure. | Timestamp: 03:20
Detailed Brief
RunPod's Product Positioning and Pain Points Addressed
- Claims: RunPod is a 'cloud AI infrastructure company' that provides GPUs and simplifies model deployment (private or open source models).; Three pain points: infrastructure management difficulty (like pre-AWS on-prem servers), slow/opaque GPU access due to supply crunch, and developer distraction from core building.; Platform serves 500,000+ developers, operates 30+ data centers (including EU), and reports $120M ARR.
- Evidence: Audrey compared GPU infrastructure management to pre-cloud on-prem server days, now abstracted by DevOps and cloud providers.; GPU supply crunch likened to COVID toilet paper hoarding; recovery expected as customers improve compute estimation.; Customers include AI-native companies who come 'for the same reasons'—flexible, reliable GPU infrastructure.; Origin story: founders Zen and Pardeep had failed crypto mining rigs in 2022, posted on Reddit for free GPU access and feedback, revenue generating since inception.
- Caveats: No details on gross margins, unit economics, or profitability despite $120M ARR claim.; No specifics on GPU sourcing model (owned hardware vs. community/partner GPUs), which affects supply resilience.; Community-first origin may not reflect current enterprise SLA/support maturity.
- Implications: Ken should evaluate RunPod for rapid prototyping and bursty workloads where cost efficiency and speed to deploy outweigh enterprise support needs.; Strong community engagement (Reddit, Discord) may provide fast troubleshooting but verify production SLA terms.; GPU supply risk exists; ask about allocation guarantees and failover GPU types before production commitments.
Serverless Product Deep Dive and Live Demo
- Claims: Serverless is best for real-time inference with auto-scaling; pay only when workers handle requests, not idle time.; Pricing: 'fraction of a cent per second'; supports configurable max workers, spending caps, and 'active workers' that stay always-on.; Demo deployed Qwen LLM from Hugging Face Hub listing in under 5 minutes; first request ~41 seconds (cold start), execution 1.5 seconds.
- Evidence: Demo selected Hub listing (pre-configured GitHub repo with Docker file, vLLM serve flags), adjusted max_model_length for context window, deployed on H100/A100 GPUs.; Workers initialized (container created, model downloaded), then processed queued requests.; Telemetry shown: request count, execution time, delay time for observability.; HTTP endpoint provisioned for API requests; can be hit by customers or internal systems.
- Caveats: No specific pricing numbers given (e.g., cost per 1M tokens, cost per H100 worker-second).; Initial cold start delay (41 seconds) would impact latency-sensitive use cases; 'active workers' mitigate this but incur continuous charges.; No discussion of max concurrent requests per worker, token throughput benchmarks, or failure/retry behavior.; Demo showed single model deployment; multi-model or LoRA support not demonstrated despite 'max loras' config option visible.
- Implications: Ken should use serverless for batch/bursty agent workloads where idle time savings justify cold start delays; use 'active workers' or pods for latency-critical inference.; Request RunPod pricing calculator or benchmarks for Ken's expected token volumes to compare vs. OpenAI/Anthropic APIs or self-hosted alternatives.; Telemetry/observability shown is basic; Ken may need external monitoring (Datadog, Prometheus) for production-grade observability.
Hub, Pods, Clusters, and Product Ecosystem
- Claims: Hub is a central repository of vetted, pre-configured AI repos (RunPod and community contributed) that can be forked, watched, starred, and deployed.; Pods are 'sandbox virtual environments'—RunPod spins up containers, allocates GPUs, manages infrastructure; user brings Docker files and code.; Clusters offer multi-node setups with high-speed networking for heavy-duty training.; CLI, SDKs, and 'skills' available for agent interaction without reading docs.
- Evidence: Demo showed Hub listing as a GitHub repo with README, Docker file, and environment variable defaults.; Audrey mentioned 'you can fork, you can watch, and then you can star and deploy on RunPod.'; Pods described as always-on containers vs. serverless which spins down when idle.; Follow-up session at 4pm to cover Python Flash SDK for deploying code as remote functions on GPUs via terminal.
- Caveats: No details on Hub vetting process, security audits, or versioning for community repos.; Pods vs. serverless cost comparison not provided; unclear when to choose each.; Clusters mentioned briefly; no pricing, minimum node count, or interconnect specs (InfiniBand, NVLink?).; 'Skills' for agents not explained—unclear if OpenAI function schemas, MCP servers, or custom tooling.
- Implications: Ken should audit Hub repos before production use, especially community-contributed ones, for supply chain security and dependency risks.; Evaluate Pods for always-on, low-latency inference vs. serverless for cost-optimized bursty workloads; request pricing comparison from RunPod.; Investigate 'skills' and SDK capabilities for agentic workflows that auto-provision GPU infrastructure—could enable AI-native GTM automation.; Clusters may be relevant for Ken if fine-tuning large models; follow up on training-specific features and networking specs.
Notable Concepts & Terms
- RunPod Hub: Central repository of pre-configured, vetted AI repos (Docker files, vLLM serve configs, environment variables) that can be one-click deployed to RunPod serverless or pods; community-contributed and RunPod-maintained.
- Serverless (RunPod product): Auto-scaling GPU inference product that charges only when workers actively handle requests (fraction of cent/second); supports max workers, spending caps, and 'active workers' that stay always-on to eliminate cold starts.
- Active Workers: Configurable number of serverless workers that remain always-on with models pre-loaded, eliminating cold start delays but incurring continuous charges even when idle.
- Pods (RunPod product): Sandbox virtual environments—RunPod spins up containers, allocates GPUs, manages infrastructure; always-on (vs. serverless auto-scaling); user brings Docker files and code.
- Cold Start Time: Initial delay when first serverless worker spins up, including container creation and model download; demo showed ~41 seconds for Qwen model; subsequent requests faster due to warm containers.
- Python Flash SDK: RunPod SDK (covered in follow-up session at 4pm) for deploying code as 'remote functions' on GPUs via terminal, creating production-ready endpoints programmatically.
- vLLM Serve: Underlying inference server used in Hub listings; environment variables in RunPod Hub repos pass as flags to vLLM serve for configuration (max_loras, max_model_length, etc.).
Operator Notes / Why Ken Should Care
- Ken can deploy LLM endpoints in <5 minutes using RunPod Hub pre-configured repos, paying fraction of cent/second for active inference—ideal for rapid agent/AI ops prototyping without GPU procurement.
- Serverless product suits bursty/batch workloads (content generation, on-demand inference) where idle time savings offset cold start delays; 'active workers' mitigate latency but incur continuous costs.
- RunPod's 'skills' for agents (unspecified tooling) could enable agentic workflows that auto-provision GPU infrastructure—investigate for AI-native GTM automation where agents deploy their own endpoints.
- Hub repos are GitHub-based; Ken should audit community-contributed listings for supply chain security before production use, especially dependencies and versioning.
- GPU supply risk acknowledged (COVID toilet paper analogy); ask RunPod about allocation guarantees and failover GPU types (demo showed H100 with A100 backup) before production commitments.
- Strong community engagement (Reddit, Discord) from founders' origin story may mean fast troubleshooting but verify enterprise SLA/support tiers for production-grade reliability.
- Pricing transparency lacking—no specific cost per 1M tokens or per H100 worker-second given; request calculator or benchmarks to compare vs. OpenAI/Anthropic APIs or self-hosted alternatives.
- Python Flash SDK session at 4pm may reveal programmatic deployment tactics for Ken's agentic workflows—worth attending or reviewing transcript for code-first patterns.
Watch Map
- 00:00: Intro: Audrey from RunPod, audience poll on RunPod familiarity (mostly new users).
- 00:45: RunPod positioning: cloud AI infrastructure for deploying models (private or open source).
- 01:15: Pain points: infrastructure management, GPU supply crunch (COVID toilet paper analogy), developer distraction from building.
- 01:30: Origin story: founders' failed crypto mining rigs, Reddit post for free GPU feedback, revenue from day one.
- 02:00: RunPod stats: 500K+ developers, 30+ data centers, $120M ARR; customer examples including AI-native companies.
- 02:30: Product overview: Pods (sandbox containers), Serverless (auto-scaling), Clusters (multi-node training), Hub (repo of AI listings).
- 03:20: Serverless deep dive: best for real-time inference, auto-scaling, pay-per-second, max workers, spending caps, active workers.
- 03:45: Live demo begins: navigating Hub, selecting LLM listing (GitHub repo with Docker file, vLLM serve configs).
- 04:15: Deploy Qwen model: adjust max_model_length, configure H100/A100 GPUs, pricing shown (fraction of cent/second).
- 05:00: Workers initializing: container creation, model download, telemetry (requests, execution time, delay time).
- 05:30: First request result: ~41 seconds queue time (cold start), 1.5 seconds execution time; subsequent requests faster.
- 06:15: Q&A invitation; mention of 4pm follow-up session on Python Flash SDK for deploying code as remote functions.
- 06:45: Closing: thanks for attending.
Source/Metadata
- Title: Under 5 minutes to a deployed LLM endpoint — Audry Hsu, RunPod
- Transcript words: 1987
- Duration seconds: 806
- Timestamp note: Timestamps estimated from transcript flow and 806-second video duration; no explicit chapter markers in transcript.
Transcript
[SPEAKER_01] Audrey, I am from RunPod. This is an intro to RunPod. Can I just get a quick hand to see how many people have already heard of RunPod or maybe even used RunPod before? Okay, newbies for everybody. Great. So, RunPod, we are a cloud AI infrastructure company. So, we have the hardware, we have the GPUs, and we make it easy for developers to deploy models. And that can be your own private model. It can be an open source model from Hugging Face. It doesn't matter to us. You bring your code and we'll bring the rest. Just really quickly, what problems does RunPod solve? Why are we even here today? Infrastructure can be hard, managing it. I think about back in the day before we had AWS, Google Cloud, when everybody would have to have on-prem servers and manage those, maintain those. That is something that we don't want to have to do as developers. Those are things that we happily have moved away from and given off to dev ops, and now it's even more abstracted for us. GPU access is slow and opaque. So, I don't know if anybody has tried to buy a GPU recently. We're in a global supply crunch. It's a bit like in COVID when everybody went to the store and bought all the toilet paper because we didn't know how long they would need to be at home for. We're a little bit in that right now, but we expect the market will recover as customers, companies, people figure out a little bit better, get a little bit better at estimating what kind of compute they need. And then last, builder primary focus should be building. So, again, we want to build apps. We as software developers, we bring the value through the applications that we build, not for managing the infrastructure. And I think RunPod has a pretty unique story. These are our founders, Zen and Pardeep. So, they had a couple of GPU rigs in their basement in 2022, failed crypto mining, and then so they were thinking, what are we going to do with our GPUs now? So, they prototyped what is now the foundations of RunPod. They posted on Reddit and said, hey, anyone want to use these GPUs for free? Just give us feedback on it. And that is literally how our company has started, and we have been revenue generating ever since. And the reason why I want to tell this story is not because it's bootstrappy, but because the origin story of RunPod has always started with builders and getting feedback from the community, and that is still true today. So, I won't promise that we'll be perfect, but we are definitely very engaged with our users on Reddit, on Discord, so we're always trying to stay engaged with you all. Just at a glance, to give you an idea of RunPod, we have over 500,000 developers on our platform, 30 plus data centers across the world, including Europe and the EU, and we've just passed a significant revenue milestone for us, 120 million in annual recurring revenue. These are just a few of our customers. You might be surprised to see some of the AI Cloud native companies on here too, but they come to us for the same reasons that most of our customers come to us. It's because they need flexible and reliable GPU infrastructure. This is a really high-level overview of different ways you can build on RunPod. So, I would say at our core, pods, it's our sandbox virtual environment. We spin up a container for you, allocate GPUs to it, and we manage the rest. So, you just bring your Docker files, you bring your code. Serverless, it's our auto-scaling product. So, when you're thinking more about bursty workloads or batch workloads, serverless is really great because instead of being always on like a container is, serverless, your workers spin down, and when they're idle, you don't pay for anything. Clusters, if you're doing some heavy-duty training, there's a place for you as well on RunPod. Multi-node clusters with high-speed networking. And then the hub, which I'll switch to in a second, it's our central repository for AI repos. These are already pre-configured, pre-vetted. We have a couple of examples of listings by RunPod for popular models, but also our community contributes to them as well. So, they're repos that you can fork, you can watch, and then you can star and deploy on RunPod. So today we're going to be talking mostly about serverless. So serverless is best for real-time inference. I talked about the auto-scaling that comes with it. Why teams use it is mostly because they don't need to preempt and figure out how much compute they need ahead of time. You can set you can configure the number of max workers that you want to scale up to. You can set limits for caps, for spending caps, and you can also configure workers that are always on. So they already have your models downloaded and they can respond to requests immediately. For a lot of teams, serverless is the fastest way if you want to start deploying a production-ready API. And now I'm going to switch over and just show you really quick how easy it is to get started and deploy something. Okay. Where are we? Okay. So right now I'm going to do everything via the console so that it's nice and pretty for you guys to see. But we also have CLI support. We have skills to help work with RunPod. Everything that's ready for your agent so you don't have to read our documents. But since we're all humans here today, I'm going to show you via the console. We'll start in the hub, which is if you're just trying to explore and see what's out there, what is something that you can get up and running right now, the hub is a great place to start. So as I mentioned, these are already vetted open source listings for AI repos. And I am going to pick the LLM, and I'll just open the underlying repository as well so you guys can see what- It is literally just a GitHub repo. So it tells you how to get set up for it. We can see there's already the Docker file here. It's already pre-configured for you. It's got some defaults for you, depending on the listing. You can pass in different environmental variables to configure it how you wish. But I'm just going to go ahead and click deploy. And I have a model that I wanted. Let me see. I was going to just pick Gwen. Works well. [SPEAKER_01] This is going to download it from Hugging Face and just expand the advanced options and look for the max model length. And I'm going to bump this up for the context window and leave everything else as the defaults. But there are settings for max loras. All of these configuration options get passed as flags to the VLM serve. And I'm going to spin it up as an endpoint here. So this might take a minute or two since this is the very first time I just created it. We've got to initialize my workers. Let's check out. So the default configuration here is it's going to deploy on some H100s and A100s are the backup here. I have my pricing. This is a fraction of a cent per second. As I mentioned before, this is only going to be charged for while the worker is actually running and handling a request. Max workers is where I can bump this up if I want to have my workload scale up to 15 workers at a time. And I can set some active workers once that I want always to be on that I don't want the container to ever spin down. And I can save that. Okay, so how does one interact with the serverless endpoint? This is just an API HTTP endpoint right here. We provisioned this endpoint for you. You can send requests to this. Your customers can send requests to this. If I just hit run and I'm going to add a few, let's, what should we ask the LLM today? Does anyone have a suggestion? Okay. I'm American. So how did Big Ben get its name? I don't know. Well, these requests are queued. Let me check on our workers. Okay. We have a handful that are initializing. This is the containers being created. That's the model being downloaded. Getting ready and the ones that are running, they've already finished. These are probably going to be the ones who are going to pick up those requests that we just added. I've got telemetry about it's blank right now, but the number of requests, execution time, delay time. So you have observability into how your endpoints are operating. And, let's see. Okay. It's already done. Got a request back in. It sat in the queue for about 41 seconds. That's going to be a little bit longer than all of the subsequent requests because of some of the cold start time that I talked about, like downloading the model, initializing the first container. But execution time, only about one and a half seconds. So, that was probably less than five minutes to get started and get something deployed on serverless from a hub listing. Does anyone have any questions? This is a very short and sweet intro. We have another session later today at four o'clock. And that one is going to be focused on our Python flash SDK. And that one is going to be completely via the terminal. And I'm going to walk you through how I can spin up and deploy my code as a remote function onto a GPU, and deploy in the end and make it a production ready endpoint here. And I'm going to go ahead and get a little bit more as well. Okay. But that's all I got for today. So, thanks. Thanks for coming. That's the model being downloaded. Um, getting ready and the ones that are running, they've already finished. These are probably gonna be the ones who are gonna pick up those requests that we just added. I've got telemetry about, um, it's blank right now, but the number of requests, execution time, delay time. So you have observability into how your endpoints are operating. And, let's see. Okay. It's already done. Got a request back in. It sat in the queue for about 41 seconds. Um, that's going to be a little bit longer than all of the subsequent requests because of some of the cold start time that I talked about, like downloading the model, um, initializing the first container. But, um, execution time, only about one and a half seconds. So, yeah, that was probably less than five minutes to get started and get something deployed, um, on serverless from a hub listing. Does anyone have any questions? This is, this is a very short and sweet intro. Um, we have another session later today at four o'clock. Um, and that one is gonna be focused on our Python flash, um, SDK. And that one is going to be completely via the terminal. Um, and I'm gonna walk you through how I can spin up, uh, and deploy my code on my code as a remote, remote function onto a GPU, um, and, uh, deploy in the end and make it like a production ready endpoint here. And I'm gonna go ahead and get a little bit more as well. Okay. But that's all I got for today. So, yeah, thanks. Thanks for coming. .