GPU Cloud Deployment Without Leaving Your IDE — Audry Hsu, RunPod
Description
The iteration cycle before Flash: commit, push, build a Docker image, pull it from the registry, load it onto a server, allocate a GPU, then find out if it works. Audrey Hsu demos what replacing that with a single decorator looks like — add `@flash.endpoint` to an async Python function and it deploys to GPU cloud from your IDE, with hot reload so a model swap is one line of code rather than a container rebuild. The second demo chains three models: Qwen 3 generates image prompts, DreamShaper renders them, Nano Banana 2 composes the results into a single photo. H100 pricing is $0.00116 per second, charged only while a worker is handling a request. RunPod's recommendation: start with pods while experimenting, switch to serverless when you need hundreds of workers autoscaling across data centers. Speaker info: - https://www.linkedin.com/in/audry-hsu/
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: RunPod Flash eliminates Docker/Git iteration cycles by letting developers deploy GPU workloads directly from their IDE using a Python decorator, charging only for actual GPU compute time at $0.00116/sec for H100s
- Why it matters: Removes infrastructure friction from AI development—no more commit-push-build-deploy loops during iteration, enabling sub-minute changes to production GPU workloads
- Best use: Reference for GPU deployment tooling, developer experience patterns, and serverless GPU pricing models; valuable for understanding how to reduce time-to-iteration in AI workflows
Executive Summary
RunPod is a 2022-founded GPU cloud provider that started when two crypto miners with spare GPUs posted on Reddit offering free compute for feedback. They're now at $120M ARR with 500+ developers across 30+ data centers in 10 countries, serving both AI-native startups and enterprises. Their core value proposition: flexible GPU infrastructure without configuration overhead (CUDA alignment, PyTorch versioning, GPU SKU testing).
The session demonstrates Flash, RunPod's Python SDK that solves the 'iteration tax' problem. Traditionally, testing changes to GPU workloads requires: commit → push to GitHub → build Docker image → pull from registry → allocate GPU → test. Flash replaces this with a single decorator (@flash.Endpoint) that deploys functions directly to cloud GPUs from local dev environments with hot reload. Code outside the decorated function runs locally; code inside runs on RunPod GPUs. Changes trigger immediate repackaging and deployment.
Live demo shows three use cases: (1) basic Stable Diffusion XL Turbo inference via local endpoint, (2) model swap to DreamShaper without rebuild cycle, (3) multi-model pipeline chaining Qwen3 prompt engineering → DreamShaper generation → NanoBanana2 photo composition. The pipeline example illustrates orchestration value—not just single model calls but complex workflows across multiple endpoints. Pricing is pay-per-second ($0.00116/sec for H100), charged only during active request processing with auto-scaling workers.
Key differentiators: serverless auto-scaling (vs. reserved Pods), sub-second billing granularity, and elimination of idle costs. Recommendation hierarchy: Pods for experimentation (1-2 GPUs, predictable load), serverless for production scale (hundreds of workers, variable load, multi-datacenter distribution). The tool targets the gap between local prototyping and production deployment where infrastructure friction kills velocity.
Key Takeaways
- Claim: RunPod reached $120M ARR starting from founders posting spare crypto mining GPUs on Reddit in 2022 | Evidence: Company began when founders Zen and Pardeep offered free GPUs for feedback after failed crypto venture; built in public, revenue-generating from day one; now serves 500+ developers across 30+ data centers in 10 countries including Oxford University for LLM training | Caveat: Session doesn't detail customer concentration, churn, or how much revenue comes from enterprises vs. individual developers | Implication: Demonstrates product-led growth model for infrastructure—community feedback loop and usage-based pricing can scale faster than traditional enterprise sales in GPU cloud market | Timestamp: timestamp unavailable
- Claim: Flash decorator eliminates commit-push-build-deploy iteration cycles, enabling GPU deployment directly from IDE with hot reload | Evidence: Add @flash.Endpoint decorator to async Python function; everything inside decorator runs on cloud GPU, everything outside runs locally; demonstrated live model swap from Stable Diffusion XL Turbo to DreamShaper without Docker rebuild—just comment/uncomment code and rerun | Caveat: Requires async Python functions; session doesn't address debugging experience, error handling, or limitations for stateful workloads or complex dependencies | Implication: Reduces iteration time from minutes (full CI/CD cycle) to seconds; particularly valuable for prompt engineering, model comparison, and hyperparameter tuning phases where you need 10-100 fast iterations | Timestamp: timestamp unavailable
- Claim: RunPod serverless charges $0.00116/second for H100 GPUs with per-second billing only during active request processing | Evidence: Pricing shown in console during demo; workers auto-scale based on requests (3 workers spun up for 3-image generation request); uptime shown in dashboard is what gets billed; no charges for idle time when no requests queued | Caveat: Serverless has pricing premium over Pods (reserved instances); presenter recommends Pods for experimentation with 1-2 GPUs, serverless for production scale with hundreds of workers | Implication: Economics favor bursty/variable workloads over continuous training—best for inference APIs, batch jobs with unpredictable timing, or dev/test environments where you want zero idle cost | Timestamp: timestamp unavailable
- Claim: Flash enables multi-model pipeline orchestration as easily as single model calls, demonstrated with Qwen3 → DreamShaper → NanoBanana2 chain | Evidence: Live demo: sent prompt to Qwen3 (public endpoint) for prompt engineering → passed result to DreamShaper (Flash endpoint on H100) for image generation → sent to NanoBanana2 (Google model) for photo composition; all orchestrated from local Python script with no infrastructure code | Caveat: Demo doesn't show error handling, retry logic, or latency considerations when chaining models across different endpoints and providers | Implication: Real value isn't just GPU access—it's making complex AI workflows composable without ops overhead; enables rapid experimentation with model combinations, A/B testing different generators, or building agentic systems that chain reasoning/generation/refinement | Timestamp: timestamp unavailable
- Claim: RunPod offers three deployment patterns—Pods (reserved VMs), Serverless (auto-scaling), Clusters (multi-node training)—plus Hub (pre-vetted open source repos) | Evidence: Pods: on-demand rental, pay-by-second, persistent VM with reserved GPU; Serverless: auto-scaling workers, no idle cost; Clusters: multi-node for training; Hub: one-click deploy for Comfy UI, Stable Diffusion, vLLM; presenter uses serverless for demo but explains Pods for experimentation | Caveat: Session doesn't detail cluster setup complexity, networking between nodes, or migration path from Pods to serverless or vice versa | Implication: Choose deployment model by workload predictability: Pods for prototyping and steady-state inference, serverless for production APIs with variable load, clusters for large-scale training; Hub useful for quick evaluation of standard tools | Timestamp: timestamp unavailable
Detailed Brief
RunPod Origin Story and Market Position
- Claims: Founded 2022 by Zen and Pardeep after failed crypto mining venture left them with spare GPUs; Posted on Reddit offering free GPUs for feedback as initial go-to-market; Revenue-generating from day one, now $120M ARR serving 500+ developers; Customers include AI-native companies and large enterprises across 30+ data centers in 10 countries (France, Romania, Iceland, Asia Pacific); Example customer: Oxford University using RunPod for LLM training
- Evidence: Presenter Audrey works at RunPod, asked audience about usage and got confirmation from Oxford student named Eunice using for LLM training; Built in public with community from inception—unusual for infrastructure company to be revenue-positive immediately; 'Punching above our weight class' comment suggests competing with larger cloud providers despite startup size
- Caveats: No detail on customer concentration, retention, or what portion of revenue is enterprise vs. individual developers; Unclear how much of $120M ARR is from serverless vs. Pods vs. clusters; No competitive positioning against AWS SageMaker, GCP Vertex, Modal, Replicate, or other GPU clouds
- Implications: Reddit-driven PLG model validated for infrastructure—developer community can bootstrap cloud business; Usage-based pricing (pay-by-second) aligns with developer preference over reserved commitments; Geographic distribution (Europe, Asia Pacific) suggests focus on data sovereignty and latency beyond US-centric clouds; Oxford use case for LLM training implies competitiveness on H100 availability and pricing vs. academic alternatives
Flash SDK: IDE-Native GPU Deployment
- Claims: Traditional iteration cycle: commit → push to GitHub → build Docker → pull from registry → allocate GPU → test; Flash eliminates this cycle by deploying functions directly from local dev environment using Python decorator; Code inside @flash.Endpoint decorator runs on cloud GPU; code outside runs locally; Hot module reload: any file change triggers immediate repackaging and deployment without manual rebuild; Spins up local FastAPI dev server for testing endpoints before production
- Evidence: Live demo showed
flash run image_generation.pystarting local dev server; Demonstrated model swap from Stable Diffusion XL Turbo to DreamShaper by commenting/uncommenting code—no Docker rebuild, just rerun script; Endpoint decorator includes parameters: name, gpu_type (Ada 80 Pros = H100 variant), max_workers=5, active_workers=1, timeout; Demo showed real-time log output: 'sees the request, started the job, queued it' as inference ran - Caveats: Requires async Python functions—no mention of support for other languages or synchronous code; No discussion of debugging experience, logging, or how to inspect GPU-side errors; Unclear how large model weights are handled—does Flash cache them or download on every deployment?; Demo had minor bugs (forgot to pass prompt as CLI flag initially), suggesting rough edges in UX; No mention of environment dependency management or how to handle custom CUDA libraries
- Implications: Targets the 'iteration tax' that makes GPU development feel slower than CPU development—brings GPU workflows closer to local Python scripting; Competitive with Modal, Banana, Replicate's deployment UX but focused on iteration speed over production scale; Hot reload particularly valuable for prompt engineering, model comparison, hyperparameter tuning where you need 10-100 fast iterations; Local dev server pattern familiar to web developers—lowers barrier for non-ML engineers to work with GPU workloads; Open question: does this scale to teams? No mention of version control, collaboration, or shared endpoints
Multi-Model Pipeline Orchestration
- Claims: Flash enables chaining multiple models without infrastructure code; Demo pipeline: Qwen3 (prompt engineering) → DreamShaper (image generation) → NanoBanana2 (photo composition); Each model can be on different endpoint (public vs. Flash vs. Google), orchestrated from local Python; Qwen3 generated detailed prompt from simple input: 'thoughtful expressions, weathered faces, soft focus on background clouds, muted urban palette with grays and deep blues, overcast lighting'; Final output composited faces from reference photos onto generated scene
- Evidence: Live demo sent request 'two men with glasses walking in London on a cloudy day, close up of faces' to Qwen3; Qwen3 output fed to DreamShaper running on Flash endpoint (H100 worker); DreamShaper result sent to NanoBanana2 (described as 'premium Google model good at composing photos'); Demo generated 3 variations in parallel, spinning up 3 workers to handle requests concurrently; Final images showed improved prompt quality from Qwen3 vs. raw user input
- Caveats: No error handling, retry logic, or timeout handling shown in pipeline code; Unclear how latency compounds across three model calls—no timing metrics shown; Demo was non-production (using founder faces)—no discussion of how to handle failures in production pipelines; NanoBanana2 described as 'premium Google model' but no pricing or API details provided
- Implications: Real value isn't just GPU rental—it's making complex workflows composable without Kubernetes/Airflow/ops overhead; Enables rapid experimentation with model combinations: A/B test different generators, chain reasoning models with image models, build agentic loops; Pattern extends to RAG pipelines: embed → search → rerank → generate, or agent systems: plan → act → observe → reflect; Prompt engineering offload to Qwen3 demonstrates pragmatic use of small LLMs to improve specialized model outputs; Photo composition use case suggests applicability to content generation pipelines for marketing, e-commerce product shots, personalized avatars
Pricing Model and Deployment Options
- Claims: Serverless charges $0.00116/second for H100 GPUs, billed only during active request processing; Serverless has pricing premium over Pods due to auto-scaling overhead; Recommendation: Pods for experimentation (1-2 GPUs), serverless for production (hundreds of workers); Workers auto-scale based on request volume: demo spun 3 workers for 3-image request; Console shows uptime per worker—that's what gets billed; no charges when workers scale to zero; Four deployment patterns: Pods (reserved VMs), Serverless (auto-scaling), Clusters (multi-node training), Hub (pre-vetted repos)
- Evidence: Presenter showed console during demo: 5-6 workers provisioned (max_workers=5 in decorator), 3 running for 3 requests; Uptime column in console shows active billing time per worker; H100 pricing: $0.00116/sec = ~$4.18/hour = ~$100/day for continuous usage; Comparison: 'If you're experimenting, start with very low worker count or start with Pods... serverless for when you need hundreds of workers'; Hub mentioned as one-click deploy for Comfy UI, Stable Diffusion, vLLM—pre-vetted open source repos
- Caveats: No detail on cold start times when scaling from zero; Unclear what 'serverless premium' over Pods actually costs—no specific pricing comparison; No discussion of network egress, storage, or other charges beyond GPU compute time; Console UI shown but not explained—unclear what monitoring/observability tools are available; No mention of reserved capacity, volume discounts, or enterprise pricing
- Implications: Economics strongly favor bursty/variable workloads: pay $0 when idle vs. continuous Pod rental; At $0.00116/sec, 1-minute inference request costs ~$0.07—competitive for API workloads but expensive for continuous training; Auto-scaling model assumes stateless inference; unclear how to handle stateful workloads or model preloading; Recommendation hierarchy makes sense: Pods for learning/prototyping, serverless for production scale, clusters for training; Hub useful for quick evaluation but likely templates code—question is how customizable and whether it uses Flash under the hood; Pricing transparency (shown live in console) builds trust vs. AWS-style surprise bills
Notable Concepts & Terms
- Flash (RunPod Flash SDK): Python SDK that deploys GPU functions directly from IDE using decorators; eliminates Docker build/push cycles by packaging code and pushing to cloud GPUs with hot reload on file changes
- @flash.Endpoint decorator: Python decorator that marks async functions to run on cloud GPUs; includes parameters for gpu_type, max_workers, active_workers, timeout; everything inside decorator runs remote, everything outside runs local
- Iteration tax: Time/friction cost of traditional GPU deployment cycle (commit → push → build Docker → deploy → test); Flash's core value prop is eliminating this to enable rapid iteration
- Active workers vs. max workers: Active workers = always-on GPU instances (pay continuously); max workers = ceiling for auto-scaling; serverless scales between active and max based on request load
- Pods vs. Serverless vs. Clusters: RunPod deployment patterns: Pods = reserved VMs with persistent GPU; Serverless = auto-scaling workers with pay-per-use; Clusters = multi-node for distributed training
- DreamShaper: Fine-tuned Stable Diffusion 1.5 model optimized for art/illustrative styles; used in demo as higher-quality alternative to SD XL Turbo for controlled image generation
- NanoBanana2: Google model described as 'premium' for photo composition; used in demo to composite reference faces onto generated scenes; unclear if this is internal Google model or third-party
Operator Notes / Why Ken Should Care
- For agent systems: Flash's decorator pattern and multi-model orchestration directly applicable to agentic workflows where you chain reasoning → action → observation loops across different models without infrastructure code; hot reload enables rapid iteration on agent prompts and tool definitions
- For AI ops: Pay-per-second serverless model with auto-scaling eliminates idle cost problem for variable workloads; useful for batch inference jobs, development/staging environments, or API backends with unpredictable traffic; contrast with reserved instances for steady-state production
- For content/business: Multi-model pipeline (prompt engineering → generation → composition) demonstrates pragmatic pattern for content generation at scale; applicable to marketing assets, e-commerce product visualization, personalized avatars; Qwen3 for prompt improvement is transferable technique
- For investing: $120M ARR from 2022 launch suggests strong product-market fit in GPU infrastructure space; Reddit-driven PLG model and immediate revenue generation rare for infrastructure; competitive positioning unclear but 'punching above weight class' implies taking share from AWS/GCP despite smaller scale
- For GTM: Community-driven growth (Reddit post → feedback loop → public building) more effective than enterprise sales for developer tools; usage-based pricing and sub-second billing align with developer expectations; Flash targets iteration speed pain point that resonates with ML engineers frustrated by DevOps overhead
- For workflow: Flash represents shift toward 'infrastructure as code annotation'—@decorator replaces YAML configs and CI/CD pipelines; local dev server pattern makes GPU workloads feel like normal web development; consider for any workflow where iteration speed bottlenecked by deployment friction
Watch Map
- timestamp unavailable: Timestamps unavailable; session structure: intro (RunPod overview, origin story, customer examples) → Flash concept (decorator pattern, hot reload) → live demo 1 (basic Stable Diffusion inference) → live demo 2 (model swap DreamShaper) → live demo 3 (multi-model pipeline Qwen3 → DreamShaper → NanoBanana2) → pricing discussion (serverless vs. Pods, console walkthrough) → Q&A on pricing/use cases
Source/Metadata
- Title: GPU Cloud Deployment Without Leaving Your IDE — Audry Hsu, RunPod
- Transcript words: 3244
- Duration seconds: 1218
- Timestamp note: Timestamps not present in transcript; session duration ~20 minutes based on 1218 seconds; live coding demo with real-time inference and iterative model swaps
Transcript
[SPEAKER_00] Hey, everyone. I'm Audrey. I work at RunPod. Were any of you in my earlier session? OK, good. Because then I'm going to say this intro is the same, but what I'm going to show is a little bit different. Has anyone heard of RunPod or used RunPod before? You have. Do you mind if I ask you how you've used us or heard about us before? [SPEAKER_01] We have some RunPod. So I use this for LLM training. LLM training? OK, at your university. [SPEAKER_00] And where do you go to uni? Oxford. [SPEAKER_00] Oxford. I did a study abroad there one summer. It's awesome there. OK, love it. [SPEAKER_00] Thank you. OK, and your name is? Eunice. Eunice. OK, so Eunice might know a little bit about this already, but I'll talk you guys through a little intro. What do we do? We're an AI cloud infrastructure company, and our mission is to build the foundational platform for developers to scale their AI workloads. What that means is we bring the hardware, we bring the GPUs and the compute. We make it easy for you guys to bring your code, bring your models, and deploy as quickly as possible. We don't want you spending time configuring infrastructure and thinking about things like scaling. Why does RunPod exist? A lot of teams that we've talked to are all wrestling with the same thing: infrastructure. They're spending more time with the infrastructure than they are with the models. Things like CUDA version alignment, what versions of PyTorch run well together, which new GPU skews have been tested, and figuring out the bugs there. A lot of those are things that we try to take that configuration problem away from you guys so you could just focus on training your model or building your apps. And a little bit of backstory about our company. This is Zen and Pardeep, our two founders. They started RunPod in 2022. They had a failed crypto mining venture, so they had a bunch of spare GPUs in their basement. They built a prototype of what is the foundation of RunPod today. And they just posted on Reddit and said, does anyone want some free GPUs in exchange for feedback? And that is literally how our company started. And ever since then, we've been building in public with the community. We've been revenue generating from the very beginning, which is very rare. And even today, we have around 500 developers on our platform. We're in 30-plus data centers across 10 countries. In Europe, that includes France, Romania, Iceland, if that's part of Europe, Asia Pacific. And we recently hit a pretty big milestone of 120 million in annual recurring revenue. So we're going to look at some of the customers that we have. You might be surprised seeing that some of these are AI native companies and some large enterprises as well. The bottom line of what they have in common is that they need flexible and reliable GPU infrastructure. I would definitely say we're punching above our weight class. Really quickly, there are different ways to build on RunPod, depending on what you're trying to do. If you need a more persistent VM environment, then Pods is a great use case. You can rent a pod on demand, pay by the second, and once you're done, you can tear it all down and start again. Pods are if you need reserved GPU. As long as your pod is running, the GPU is yours, and no one can take it away from you. Serverless, if you're ready to deploy something and you care more about scaling, so your workloads are more variable in terms of frequency and load, we help you auto-scale your workers and scale them back down when you don't have any requests happening, so you don't pay for any idle time. Clusters are a great use case for training, multi-node. And then Hub is also a place where you can deploy already open source AI repos that have already been pre-vetted by us for popular models like Comfy UI, Stable Diffusion, VLLM. And that's one way if you're just exploring to click around and get started really quickly. I'm going to talk about serverless today and the product that we just... I'm going to switch my displays again here so we can mirror my screen. One of the things that is a huge pain for developers is if they're still in the iteration or the development phase. Normally, when you are working on, let's say, some code around your inference model and you're still testing things out, you have to make a commit, push it to GitHub, build your Docker image, pull it down from the container registry, and then load it onto a server, allocate a GPU to it, and then you get to test it and see if it's working as you expect. And then you do that all over again until you're ready. The problem that Flash is trying to solve here—and Flash is our Python SDK—is that we want to eliminate all of that iteration cycle so that you can deploy your function on a GPU right from your local development environment. I'll zoom in here really quick. This is all you need to know about Flash in one little paragraph. You have a regular async Python function, you add our Flash endpoint decorator, and it's going to deploy and package everything inside your function onto a GPU cloud. Everything around it, your main function, any helper functions that you have, those all run on your local development environment. But if you need GPU compute, that can run on the cloud. And we have hot mod file reload. If you change anything in your application anywhere, then it gets repackaged and pushed up immediately, and you can test and iterate super quickly. And I'm just going to show an example of this. OK. So I have a function here, generate_image. And it's going to deploy and package everything inside your function onto a GPU cloud. Everything around it, your main function, any helper functions that you have, those all run on your local development environment. But if you need GPU compute, that can run on the cloud. And you can, we have hot mod of file reload. So if you change anything in your application anywhere, then it gets repackaged and pushed up immediately, and you can test and iterate super quickly. And I'm just going to show an example of this. Okay. So I have a function here, generate image. Simply, I'm loading PyTorch. I'm loading a pre-trained stable diffusion model, stable diffusion XL turbo. Really great for fast generation of images. And I'm going to save the image down, and that's going to return it base64 encoded. So I can run this right now here. I've already installed all my dependencies. I already have a Flash project going. I'm going to Flash Run Image Generation Pi. And what I'm going to actually do is I have a little Flash Run. So Flash Run spins up a local development server here. It's just a fast API server. And I can send my request here to this endpoint. And I'm going to do that really quickly. Just get to my project. And this is just a little helper script that's going to send a post request to it and then decode that image so that you guys can actually get to see what it looks like once it's generated. And it was image generation async. There we go. And let's pass a prompt to it. Can I get help from the audience? What do we want to generate today? Literally anything. Anything random. [SPEAKER_01] Cats flying in the sky. Okay. Cats flying in the sky. What does the sky look like? What time of day is it? [SPEAKER_03] Cloudy. [SPEAKER_03] In London. [SPEAKER_03] Yeah. Flying in. Flying on a cloudy day in the sky somewhere in London. And I passed it correctly. He's looking at it so closely. He's helping me debug live. I love it. Thank you. I passed URL and that's true. I must have. Is it my HTTP? Okay. There we go. Okay. Going back to the local dev server. It sees the request. It started the job. It's queued it. And we're just going to wait for a second to see if it finishes. And so while that's happening, let me bring your attention back to... I'll make it bigger for you guys. The endpoint decorator. So this is where all the magic happens. I have passed a name for my endpoint. I specify a GPU family. So the ADA 80 Pros. These are different variations of NVIDIA H100 cards. I can specify my max number of workers to be five. So I can have at max five of them running at once. I just put one active worker. So this is one that's always going to be running and always on. And that's definitely a dragon. And it didn't take my prompt probably because I... [SPEAKER_00] Did I not pass it as a... I didn't pass it as a flag. Prompt. There we go. Okay. [SPEAKER_00] Now it's definitely generating cats flying. [SPEAKER_00] Okay. [SPEAKER_00] Back to the endpoint decorator. And then there's other different configurations for timeout, which is how long a worker is idle. Here we go. Okay. This looks terrible, guys. They are cats. They're abstract cats. And I'm not from London, but maybe someone can tell me if this looks like a London chimney. Maybe. Okay. So I don't like what just happened. So what we're going to do instead is we're going to switch out our model. So I'm just going to comment out this code here. And then down here... Let's swap in DreamShaper, which is a fine-tuned model based off of Stable Diffusion 1.5. So this one is... While Stable Diffusion XL Pro is more optimized for just quick generation, I think this one is going to generate a better quality image for us. And it's specifically better for more art and illustrative styles. So we've changed some of the parameters in it. It's going to have a few more inference steps to it. We're going to set that to 25. Height and width 10 by 24. That's fine. And let's just send the same request again. And let's see what happens. It's different. So again, what made this really fast is instead of making a code change, committing it, rebuilding my Docker, uploading it somewhere, and then allocating GPU infrastructure, all of this is happening right here from my IDE, and I never have to leave. This is good, right, guys? We like this one. Okay. So one last thing that I'm just going to show you guys to round things out is I think where using a developer tool like Flash makes a big difference is when you're trying to not just make one single call to one model. It's all about all the orchestration code around it, right? So I have here a pipeline that I've pre-prepared. And what it's going to do is instead of me generating and writing out every prompt, it's going to send a request to Gwen that's already hosted on a public endpoint. And Gwen3 is going to generate all the prompts for me. And then after that, it's going to send that to our Dream Shaper running on our endpoint. And then after that, there's one more pipeline that it goes through. It's going to send the request to Nano Banana 2, which is a premium Google model that's really good at composing photos together. And I'm hoping that I can compose some cool pictures of our founders and I can send them to them after this demo is done. Okay. Now let's run the whole pipeline here. Check. Okay. Prompt. And Gwen3 is going to generate all the prompts for me. And then after that, it's going to send that to our Dream Shaper running on our endpoint. And then after that, there's one more pipeline that it goes through. It's going to send the request to Nano Banana 2, which is a premium Google model that's really good at composing photos together. And I'm hoping that I can compose some cool pictures of our founders and I can send them to them after this demo is done. Okay. Now let's run the whole pipeline here. Check. Okay. Prompt. Two men walking in London on a... It is cloudy today. Close up of their faces. Any other requests for these two men? How do they look? Are they doing something? Glasses. Glasses? Yeah. Okay. Two men with glasses. Close up of their faces. Okay. And let's generate... Let's generate three of those and let's compose it together. [SPEAKER_02] So how does it work in terms of pricing? Pricing? Sure. So every request that we send to it, you're only charged for how long that request is running. So let's see. I'm going to... I said this whole session was going to be only in the terminal, but I'm going to go back into the console just to show you what's running. Let's see. This is the endpoint that we created from the terminal. Here are the workers. I think we said five workers. So we have about five or six here that are provisioned. Three of them are running because I asked for three photos. And so this is uptime. This is what you're being charged for. [SPEAKER_00] And let me see. The cost of an H100 right now is .00116 cents per second. [SPEAKER_02] Is it the same as for pods or is it... Pricing is a little bit different for serverless versus pods because pods, you don't get any of the scaling with it. So there's a little bit of a premium for serverless. So what we usually recommend is if you're still experimenting, then either start with a very low worker count or start with pods, right? Because when you're experimenting, you might only need limited number of GPUs. One GPU at a time, two GPUs at a time. Serverless for when you need hundreds of workers running on hundreds of GPUs and you want them distributed for a better availability across different data centers. Okay, guys, here's our final presentation. So this was our original prompt. [SPEAKER_00] Two men with glasses walking in London on a cloudy day. Close up of their faces. So on the left, this is what Dream Shaper generated based off of the prompt engineering that Quen3 did for us. So it's a lot better of a prompt than what I sent in. It has a lot better cues about notes on like, I can read it out to you since I know it's hard to see. Thoughtful expressions and weathered faces. Soft focus on background clouds. Muted urban palette with grays and deep blues. Overcast lighting. And then on the right is the final composed photo. [SPEAKER_00] This is a very handsome picture of Pardeep. And this is somehow a very old photo of Zen. And I'm just going to scroll down to show you that this is the reference photo that I sent it. But overall, I just wanted to show you guys this is how you can get started really quickly. You can start in your local development environment. You can use open sourced models. You can bring your own model, private model. And this was really fun to do. [SPEAKER_00] So thank you guys for hanging out with me. This is the endpoint that we created from the terminal. Here are the workers. I think we said like five workers. So we have about five or six here that are provisioned. Three of them are running because I asked for three photos. And so this is uptime. This is what you're being charged for. And let me see. The cost of an H100 right now is .00116 cents per second. Is it the same as for pods or is it... Pricing is a little bit different for serverless versus pods because pods, you don't get any of the scaling with it. So there's a little bit of a premium for serverless. So what we usually recommend is if you're still experimenting, then either start with a very low worker count or start with pods, right? Because when you're experimenting, you might only need limited number of GPUs. One GPU at a time, two GPUs at a time. Serverless for when you need hundreds of workers running on hundreds of GPUs and you want them distributed for a better availability across different data centers. Okay, guys, here's our final presentation. So this was our original prompt. Two men with glasses walking in London on a cloudy day. Close up of their faces. So on the left, this is what Dream Shaper generated based off of the prompt engineering that Quen3 did for us. So it's a lot better of a prompt than what I sent in. It has a lot better cues about notes on like, I can read it out to you since I know it's hard to see. Thoughtful expressions and weathered faces. Soft focus on background clouds. Muted urban palette with grays and deep blues. Overcast lighting. And then on the right is the final composed photo. This is a very handsome picture of Pardeep. And this is somehow a very old photo of Zen. And I'm just going to scroll down to show you that this is the reference photo that I sent it. But overall, overall, yeah, I just wanted to show you guys like this is how you can get started really quickly. You can start in your local development environment. You can use open sourced models. You can use bring your own model, private model. And this was really fun to do. So thank you guys for hanging out with me. So thank you guys for just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to