Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai
Description
GPU utilization is a lie. It read 100% straight through pretraining while the cluster was nowhere near well used, so Gabriel Jorge Menezes tracks tensor core utilization instead, and watched it climb as training resolution stepped from 128 pixels up to 1024. That is one of several numbers he argues you cannot train at this scale without. InfiniBand counters are exported by nothing off the shelf, and most of their failures turned out to be cross node communication, so they built that collection themselves. Any GPU running hotter than 78 degrees gets pulled rather than debugged, because one warm card throttles and destabilizes the entire run. This is the infrastructure half of Krea 2, the model trained from scratch on thousands of GPUs. Crashes scaled with the cluster and often failed silently, with communication timing out while every dashboard stayed green, and the practical answer was to stop treating each one as a mystery. Let it crash, and the same nodes running the same code will frequently go 24 hours on the next attempt. What made that survivable was checkpointing aggressively against a filesystem quick enough to write a terabyte in under 30 seconds. Production and training then share one cluster, with training holding priority and inference evicted to outside providers through a fake Kubernetes node, migrated back gradually rather than all at once so the site never drops. Speaker info: - https://www.linkedin.com/in/gabriel-jorge-menezes/ - https://gab-menezes.github.io/ Timestamps: 0:00 - Krea 2, trained from scratch, and two open checkpoints 3:26 - Crashes at scale, and the silent ones 4:18 - Metrics are everything, starting with temperature 5:58 - GPU utilization is a lie, use tensor cores 6:48 - InfiniBand and NVLink metrics you have to build yourself 8:29 - Checkpointing hard against a fast filesystem 9:21 - Gang scheduling, and training outranking production 11:01 - Flipping inference out through a fake node 14:23 - Taints that stop you wasting GPUs 1
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Krea trained and serves its Krea 2 image model efficiently by treating large-scale GPU infrastructure as a reliability and scheduling problem: instrument the hardware deeply, checkpoint aggressively, and dynamically shift inference off an internal cluster whenever training needs the GPUs.
- Why it matters: This is a concrete architecture for maximizing utilization of a shared GPU fleet without forcing researchers to manage capacity or causing customer-facing inference downtime.
- Best use: Use it as a practical reference for GPU-cluster observability, fault-tolerant distributed training, Kubernetes gang scheduling, and a control-plane pattern for bursting inference to external capacity.
Executive Summary
Gabriel Jorge Menezes explains the infrastructure behind Krea 2, a from-scratch, open-source image-generation model with both a raw pretrained checkpoint and a post-trained Turbo version that generates images in under a second. Krea's goal was not merely photorealism but a model that gives creatives more unusual, exploratory, and stylistically diverse output. The infrastructure challenge was therefore both training a large diffusion transformer on thousands of GPUs and serving production inference from the same overall GPU estate.
His central operational lesson is that distributed training failures are normal, frequently silent, and worsen nonlinearly as GPU count rises. Krea's response was not to eliminate every fault before progressing, but to make failures diagnosable and inexpensive: collect hardware and fabric telemetry beyond default NVIDIA exports, proactively remove suspect equipment, and checkpoint frequently enough that crashes do not materially delay training.
The most actionable technical claim is that conventional GPU utilization is a misleading health metric. Krea instead watches Tensor Core utilization, GPU temperature, InfiniBand throughput/wait/error metrics, and NVLink errors. Cross-node InfiniBand communication caused most failures, while NVLink telemetry helped isolate node-local faults that did not show up in ordinary GPU health signals.
For cluster economics, Krea uses one Kubernetes-based cluster for both training and production. Training gets higher priority and can preempt inference, while a Virtual Kubelet-based abstraction sends displaced inference workloads to another cluster or external GPU provider. Taints, Prometheus-driven automation, and a descheduler then move workloads back gradually when local GPUs return, avoiding both idle GPU spend and a disruptive mass eviction of production traffic.
Key Takeaways
- Claim: At large GPU counts, training infrastructure should be designed around expected intermittent failure rather than assuming stable long-running jobs. | Evidence: Krea's small experiments ran for days, but larger 128-, 256-, and 512-GPU runs crashed much more often, at times lasting less than eight hours; identical machines, code, and data could repeatedly fail after an hour and later run successfully for 12-24 hours. | Implication: For any serious distributed training program, reliable recovery and failure visibility are more valuable than spending excessive time trying to explain every transient crash before continuing. | Caveat: The speaker references a Meta paper on expected failure rates but says Krea observed the same broad pattern rather than the same absolute numbers.
- Claim: GPU utilization is not an adequate measure of whether expensive accelerators are doing useful work. | Evidence: Krea saw nominal GPU utilization at 100% during pretraining but calls it misleading because it only reflects whether the GPU is busy, not whether Tensor Cores are being efficiently used; Tensor Core utilization rose as image resolution increased from 128 through 1024 pixels. | Implication: Capacity dashboards should distinguish device busy time from effective compute utilization, otherwise teams can misdiagnose low training throughput as a lack of GPU capacity.
- Claim: Fabric-level telemetry—especially InfiniBand metrics—is essential for diagnosing multi-node training failures. | Evidence: Krea identified cross-node communication as the source of most failures and built custom collection for InfiniBand throughput, message wait time, error types, packet counts, and related exported counters because NVIDIA's standard DCGM telemetry did not provide all of it. It also collected NVLink errors, which are not fully exported by NVIDIA, to identify problematic single-node communication paths. | Implication: A multi-node GPU platform without fabric error and latency telemetry has a major observability gap; node replacement decisions should be based on communication-path evidence, not only GPU process-level health. | Caveat: InfiniBand was materially more important for Krea than NVLink because its primary issues were cross-node; NVLink was chiefly useful for isolating individual bad machines.
- Claim: Aggressive checkpointing is the practical mechanism that converts unstable large-scale training from a schedule-killing problem into a recoverable one. | Evidence: Krea abandoned Ceph after reliability problems and moved to Weka, reporting roughly 1.2 TB/s read throughput and nearly 1 TB/s writes; this allowed checkpoints every 20-30 minutes, with roughly 1 TB written in under 30 seconds without choking training. | Implication: When distributed-job uptime is poor, checkpoint cadence and checkpoint write performance should be treated as first-class training economics, not as a storage implementation detail. | Caveat: The recommendation to use a paid storage system is explicitly conditional on having the budget, but the speaker's point is that storage trustworthiness matters more than avoiding storage cost.
- Claim: A shared GPU cluster can prioritize training over inference without user-visible production disruption if inference has a transparent external escape path. | Evidence: Krea gives admitted training pods high Kubernetes priority, allowing them to evict inference from internal GPUs. A Virtual Kubelet presents a controllable fake Kubernetes node, receives the displaced pod specification, selects an external cluster or GPU provider, translates the specification, deploys it, and reconciles its state back to Kubernetes. | Implication: A control plane can make internal GPU capacity function as a training-first resource while preserving inference availability through portable workload placement and external bursting. | Caveat: This pattern requires integration work for each provider and a routing algorithm to choose among them; it is not an off-the-shelf autoscaling configuration.
- Claim: Kubernetes should be allowed to perform standard reconciliation rather than embedding bespoke recovery logic in the external-provider adapter. | Evidence: In Krea's Virtual Kubelet implementation, when a workload fails on either side, the adapter marks the pod failed; Kubernetes detects that state and creates a replacement pod through its normal mechanisms. | Implication: Keep provider adapters narrow: report desired and observed state accurately, then rely on the orchestrator's native restart and reconciliation semantics rather than duplicating them.
- Claim: Gradual workload migration is required when returning inference from burst capacity to local GPUs, because immediate eviction can take production down. | Evidence: Krea uses Prometheus-derived signals to add or remove node taints based on internal GPU availability. When GPUs become available again, a descheduler slowly moves pods back because they no longer tolerate the taint. The speaker rejects a Kubernetes NoExecute taint because it would evict everything simultaneously. | Implication: Repatriation needs a controlled drain/migration policy, and any scheduler relying on static capacity declarations needs automation or operational safeguards for node churn. | Caveat: Krea's Kueue resource accounting is manually specified, and its fluid cluster—with nodes entering maintenance or disappearing—can cause the declared resource total to drift and break gang scheduling.
Detailed Brief
Training scheduler design: research-facing simplicity, cluster-wide allocation
- Claims: Krea aims to remove GPU-capacity management from the researcher workflow: researchers submit work into a queue, and the infrastructure decides when and where it can run.; The system uses Kueue for gang scheduling and two layers of priority: workload priority within Kueue and ordinary Kubernetes priority.; Gang scheduling matters because large distributed jobs require their full set of resources together; partial allocation is not useful for these training runs.
- Evidence: The speaker describes Kueue as an open-source project that provides queueing and gang-scheduling behavior for training workloads.; Krea's production inference pods are lower priority, while accepted training pods are high priority and therefore schedule once admitted.; He notes that Kubernetes 1.35 reportedly includes out-of-the-box gang scheduling similar to Kueue, though Krea had not yet evaluated it.
- Caveats: Kueue's queue resource quantities were manually configured in Krea's setup; changing live capacity can cause the configured values to become stale and interfere with gang scheduling.; The speaker does not provide an automated solution for synchronizing Kueue resource definitions with real cluster capacity.
- Implications: Separate the research API from physical capacity allocation so experimentation does not require each researcher to reason about individual GPU availability.; Evaluate native Kubernetes gang scheduling before committing to a separate scheduling layer, but validate feature maturity and operational behavior in the target Kubernetes version.
Hardware triage and practical observability rules
- Claims: Krea favors rapid isolation and replacement of suspect hardware over prolonged diagnosis during active training campaigns.; Simple metrics can be highly effective if they are directly tied to known training failure modes.
- Evidence: Krea removes GPUs above 78°C rather than attempting to repair or reason through intermittent thermal throttling during a run.; A single warmer GPU can throttle, slow the collective job, and create instability or seemingly unrelated failures.; NVLink error spikes can reveal a faulty node even when its GPUs otherwise appear healthy.
- Caveats: The 78°C threshold is a Krea operating heuristic, not a universal hardware specification or a substitute for understanding a provider's thermal design.
- Implications: Create explicit quarantine/replacement thresholds for leased or managed GPU hardware so on-call teams can act quickly.; Instrument at the physical-device, intra-node-link, and cross-node-fabric layers rather than treating the GPU as a single opaque unit.
Notable Concepts & Terms
- Krea 2 / K2: Krea's from-scratch diffusion image model, released with both a raw pretrained checkpoint for downstream post-training and a fast post-trained Turbo version.
- Diffusion Transformer (DiT): The model category Krea worked on; its research effort included transferring techniques from LLM research into diffusion transformers.
- Tensor Core utilization: Krea's preferred proxy for useful GPU compute efficiency, contrasted with generic GPU utilization that can show a busy device without effective compute.
- InfiniBand: The high-speed interconnect between GPU nodes; Krea considers its detailed error, wait, and packet telemetry essential because most large-run failures were cross-node.
- NVLink: The intra-node GPU interconnect; less central than InfiniBand for Krea but useful for detecting faults localized to a machine.
- Kueue: An open-source Kubernetes queueing and gang-scheduling project Krea uses to prioritize and admit distributed training jobs.
- Virtual Kubelet: An open-source Kubernetes extension used by Krea to represent external capacity as a virtual node and dispatch inference pods to external clusters or GPU providers.
- Taints, tolerations, and descheduler: The Kubernetes primitives Krea combines to control whether workloads use local versus external capacity and to migrate them back gradually without a production-wide eviction.
Operator Notes / Why Ken Should Care
- Audit GPU observability for the distinction between generic utilization and Tensor Core efficiency; add fabric wait/error/packet metrics before scaling distributed jobs.
- Set an explicit checkpoint recovery objective: acceptable lost work per crash, target checkpoint interval, maximum checkpoint duration, and storage reliability requirements.
- If sharing GPU capacity between training and serving, prototype a training-first priority policy with an inference burst path before relying on preemption in production.
- Avoid NoExecute-style bulk eviction for inference repatriation; require gradual drain behavior, health gates, and rollback protection.
- If evaluating Kueue or a similar scheduler, automate capacity synchronization or monitor for drift caused by maintenance and node churn.
Source/Metadata
- Title: Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai
- Transcript words: 3216
- Duration seconds: 1015
- Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript.
Transcript
. Hello, everyone. My name is Gabriel. I work at Korea, and I will be talking about the infrastructure that allows us to train K2 and also how we serve it. So what is K2? K2 is our pre-trained from-scratch model. We just released it less than a month ago. And the whole idea behind training this model was because we were bored of AI images. They are quite saltless. They have no spice. And the whole idea was we want to give creatives tools to explore out of distribution, extremely interesting images, do composition, and actually give tools to creatives. And that was the whole idea of the model. So the model was trained from scratch, no base checkpoints, nothing, everything done in-house. And also, as I said, built for exploration. And right now, this is what you can get out of K2: very different styles, pixel art, photo real, and some silly stuff. And the whole idea of the model has been, let's explore this medium. Create2 is open source. Right now, you can go play with it. There are two checkpoints. We also serve it in production. There's a raw checkpoint, which is pre-trained so people can post-train and do whatever they wish to do with it. And there's also the post-trained version, which is the turbo one, which is very, very fast. You can get an image in less than a second. And this is the type of image you can get in less than a second. On Huggy Face, GitHub, go play with it. Also, you can go to production on kreya.ai and go play with it. So let's talk about how we train this model. As I said, it's going to be how we train and how we serve. First, the model was trained from scratch on thousands of GPUs. We have a big cluster, one main cluster with a lot of GPUs, all InfiniBand-connected. And you put those GPUs to work and it's trained. But I wish it were that simple, but it's not. So at the beginning, we did a bunch of small ablations on a small number of GPUs to see how things go, right? So you want to test some hypotheses, and you use a small number of GPUs and let it train for a little bit. Oh, this works, it doesn't work. Let's scale. And as we were training this model, the whole idea was to bridge the gap between LLM research and diffusion transformers. So my researchers ported a lot of research from LLMs into DITs. And the whole architecture of the model was meant to be extremely, extremely simple. And so it is very, very dumb, but very effective. And so let's start talking about numbers. Incredibly, maybe still an issue on our part, maybe our cluster was very interesting as we were scaling. When you do small experiments, experiments would run for days and even less than we would like to, but they would still run fine. And as we started scaling, getting more and more GPUs, like 128, 256, 512, whatever number, and you scale and scale, things are crashing more. That's expected, right? There's more surface area for things to break, and things are going to go wrong. And a lot of the time, things would go wrong in silent ways. I know, NCCL timeout just crashes, and the metrics are all good. And it is extremely annoying stuff. At the beginning, we were paranoid, swap nodes, change nodes, whatever, whatever, whatever. And we learned that sometimes you just let it crash. It crashes, it runs for an hour, it crashes, it runs for an hour, it crashes, and then it runs again on the same setup, machine, same code, same data for 12 hours, 16 hours, 24 hours. There is this paper from Meta that gives you a rough estimate of how many failures you should expect. We kind of saw the same pattern, but not the same numbers. Our runs would last way, way less than this. So it was extremely annoying, and you can imagine doing large KOP training on runs that last less than eight hours. It is a problem, right? You want to keep those GPUs fed, and if things are crashing, you are not making progress and are losing time, and the model is going to be delayed. So for us, extremely important, and at least from the infra side, was to get metrics. Metrics are everything. That's how I can support my researchers. That's how I have visibility in the system. And if you're doing large KOP training, I highly, highly recommend for you to invest heavily in metrics. Don't go blind, because you're going to go crazy. So I'm going to share some of the metrics that were important for us, and they're quite silly, but extremely effective. First one, GPU temperature. GPUs are very, very annoying. If you have a single GPU that is a bit warmer than the others, it's going to start throttling and slow down, and the training is going to be unstable and have weird problems. So for us, it was, if there are any GPUs above 78 degrees, you remove them. Don't think about it. Don't try to fix it. Don't try to be smart. Just remove the GPU. You're going to save time, and just ask your provider, this GPU is hot. Please replace it. And there are two metrics that, at the beginning, we did not fully understand. And as we are getting more and more used to GPUs and how they work, there is GPU utilization, which is a lie. Don't trust this. This is dumb. This tells you, oh, the GPU is doing work. And this is the amount of time the GPU is doing work, but not good work. Not how efficiently the GPU is working. So, as you can see during our pre-training, the GPU is at 100%. This is not true. We are not fully utilizing the GPU. This is 100% a lie. What we would use as a proxy was TensorFlow utilization. This is actually how much of the TensorFlow cores you're using and how effective they are being. And it was also very interesting, as we're doing pre-training, and then you go from pre-training to mid-training to post-training, you start scaling the resolution of the images. So, for example, in pre-training, we do 128, 256, 512, 1024 pixels as they scale. And you could see the TensorFlow core utilization go up as we would scale on these resolutions, because now we're doing more work on images. And also very, very interesting were InfiniBand and NVLink metrics. These, by default, are not exported by the NVIDIA metrics, the DCGM stuff. Some NVLink stuff, yes, but no InfiniBand. So, if you don't have InfiniBand metrics, go get them. I'm telling you right now, if you're doing large-scale training with a bunch of GPUs talking to each other between machines, and you have no InfiniBand metrics, you are doing something wrong. This was probably the most important stuff for us, because most of our failures were related to cross-node communication. So, for example, here, you just have throughput. But on our refining dashboard, we have a bunch of stuff, from wait, like when a message is sent on the fabric, how much time the message is waiting, or the number of errors and different types of errors and number of packets, and all of the things that InfiniBand exports, we collect. We had to build custom stuff to get this. It was not hard. You can figure it out. It was very, very easy. Same thing for NVLink. NVIDIA exports some stuff about NVLink. But, for example, NVLink errors, NVIDIA doesn't export this. So, you can collect this. And, as I said, InfiniBand was extremely important. NVLink was a little bit less. In some cases, these helped us catch some problems, especially because NVLink is single nodes, right? The communication inside a node. So, sometimes a single node would have a weird failure where the GPU seemed to be fine, but some weird error happens. And then you can look at NVLink errors and you see all errors happening. And then replace that machine. So, go get these metrics. and all of the things that infinite bands exported, we collect. We had to build custom stuff to get this. It was not hard. You can figure it out. It was very, very easy. Same thing for NVLink. NVidia exports some stuff about NVLink. But, for example, NVLink errors. NVidia doesn't export this. So, you can collect this. And, as I said, infinite band was extremely important. NVLink was a little bit less. In some cases, these helped us catch some problems, especially because NVLink is single nodes, right? The communication inside a node. So, sometimes a single node would have a weird failure where the GPU seemed to be fine, but some weird error happens. And then you can look at NVLink errors, and you see all errors happening. And then replace that machine. So, go get these metrics. They're extremely important. And without this, we would not be able to train at all. Also, as I said, our trains would crash constantly. And a hecky way to do it to fix the problem is just checkpoints. Use and abuse the file system that you have. At the beginning, we used Ceph. Ceph did not work well. It was very annoying. It broke. We did not trust the data. So, I recommend if you have the money, go with something paid because you can trust your data. You can see numbers. This is our Weka cluster. We can do 1.2 terabytes a second of reads, almost a terabyte of writes. And the file system would not choke on the training. So, we could checkpoint every 30 minutes, 20 minutes, produce a terabyte of data in less than 30 seconds. So, this would not delay trainings. That was probably one of the most important things that we did to recoup the loss. Just checkpoint. Don't think about it. And how we serve. This goes in connection on how the trainings are launched. Because at the beginning, I don't want my researchers to think about GPUs. I just want them to launch stuff. And this goes into a queue. And if we have GPUs, we have GPUs. If we don't have GPUs, we don't have GPUs. So, queue is an open source project. You can look it up. It does game scheduling. Game scheduling is extremely important for trainings in general. And this gives us a semantic of two tiers of priority where you have a workload priority. And these, you can say, this training is more important than this one. So, it skips in the queue, in front of the queue. And then, after this, we have the normal Kubernetes priority for use to Kubernetes. And the way the system works, the training pods, they always have high priority for everything. So, once they are admitted, they're going to schedule. If there is inference running on those machines, the inference gets kicked out. And you would say, oh, this is bad. Production is going to go down. No. You can build on top of that to make production not go down, which is very, very cool. The only one of the problems with queue, which is annoying, you can automate that. We have not. It's just that you specify the queues. The queues have amount of resources: CPU, GPUs, memory, whatever. But this is manually specified. And at least our cluster is quite fluid. Nodes phasing in and out of existence. They go into maintenance, whatever. You lose a few nodes here and there. This number gets out of sync. And sometimes this would break gang scheduling. So, FYI, this is a bit annoying. You're going to face this if you use queue. But, yeah, very good project. Kubernetes 135. We have not had the chance to play with it. Has gang scheduling out of the box. Something very similar to queue. So, maybe you can use Kubernetes 135. And, as I said, this is the system that we built that allowed us to train using the whole cluster. As I said, we have one big cluster that runs production and trainings. So, I don't want my researchers to think about GPUs. And I don't want to make the trainings and research be delayed because production is running, right? Production is lower priority. The site still needs to work. People still need to use the website. But the GPUs, the value that we get of the GPUs doing training is higher than we get out of production. So, the whole system works by default where there is this magical system that I'm going to explain in a bit that allows us to flip traffic between clusters magically. And not just clusters: external providers, GPU rentals, whatever. And you can see the green, the dark green is inference running in cluster. And then someone launches a train. And then suddenly it starts flipping to the other cluster. And then training is done or whatever happens, it flips back. So, we stop wasting money. And this is seamless. No one needs to think about it. The whole system handles itself. And you get this very nice pattern of I'm going to use all the GPUs in my cluster for trainings. Production is going to run somewhere else. I don't need to think about it. My users in production, they're not going to feel anything. Research is going to be happy. And we can get values out of the GPUs. So, how does this work? There's this very nice project called Virtual Kubelet, also open source. You can build on top of it. It is a very nice code base. And this works by creating a fake machine in Kubernetes. Kubernetes has nodes. This creates a fake machine that is up to you to control how it works. So, Kubernetes does normal scheduling, as you would expect. Things would go into these nodes. For example, here, all of the GPUs in the cluster are used, right? So, this pod goes into the Virtual Kubelet node. And in there, you can do whatever. This is the system that we built. There is, you receive the pod spec and then you find a provider. This is up to you. Let's say you have a deal with some provider that gives you nice prices. You integrate into here. We built some nice interfaces to not leak things. So, we just implement a provider and there's an algo that decides which one it goes to. You translate the spec of the pod into the provider stuff and you deploy. And then you have something that reconciles between both sides. And it was extremely, extremely nice. If you guys know about Kubernetes, Kubernetes has the horizontal pod out of the scaler which scales the number of replicas. Number of replicas. Could you stop being annoying? Thank you, sir. I appreciate it. There you go. Let's go back. Now we go back. Back, back. There you go. Kubernetes has the HPA, and the HPA scales the pods. And so, if something fails, it is very interesting. You don't need to handle the fail. The only thing you need to handle is, oh, something failed. You mark the pod as failed. Kubernetes is going to detect that something has failed and create a new one. You don't need to try to save the world. Let Kubernetes handle it for you, which is an extremely nice way of handling stuff. If something breaks on your side, something breaks on the other side, just mark as failed. Let Kubernetes handle it. Create a new one. And things keep working. Very, very nice way to handle stuff. And also very interesting, let's say you have GPUs on your cluster available, right? You don't want to waste money. This would be very, very bad. So, the system works using taints, Kubernetes taints. They allow and disallow things to run. Pods have tolerations for the taints. And when we have GPUs in the cluster, if you look back, there is the taint system at the bottom. This taint system, it is what would by itself decide if we have GPUs or not GPUs in the cluster. And this adds a taint into the node when we have a lot of GPUs. So, a lot of GPUs in the cluster. We taint the node. Nothing can schedule on it. So, we stop wasting GPUs. The pods, they go into the GPUs in the cluster. We don't waste money. And then imagine someone launches a train, right? This train is going to take all the GPUs in the cluster. It's going to hog all of the GPUs. No GPUs in the cluster. The system detects this, removes the taint, new pods schedule there. Very, very nice. You also don't think about it. And for us, for example, we use just some Prometheus metrics. That's how we do it. It's very simple. But it works very, very, very well. You let the system run by itself. And this adds a taint into the node when we have a lot of GPUs. So, a lot of GPUs in the cluster. We taint the node. Nothing can schedule on it. So, we stop wasting GPUs. The pods, they go into the GPUs in the cluster. We don't waste money. And then imagine someone launches a train, right? This train is going to take all the GPUs in the cluster. It's going to hog all of the GPUs. No GPUs in the cluster. The system detects this, removes the taint, new pods scheduled there. Very, very nice. You also don't think about it. And for us, for example, we use just some Prometheus metrics. That's how we do it. It's very simple. But it works very, very, very well. You let the system run by itself. Someone's going to launch stuff. You're going to, the train is going to kick out the pods. It's going to schedule. It's going to take the GPUs. The system is going to detect that, remove the taint. Pod scheduled there. Very nice. Someone finished the train. Now we have pods running on the other side. You're wasting money. How do we fix this? Same thing. You run something else that detects the system and removes things back. In this case, a descheduler, once the taint is added back, GPUs available, we add the taint. The descheduler says, oh, these pods, they don't tolerate the taint. I'm going to migrate them back. And you can ask, oh, why don't you use a no execute taint? No execute in Kubernetes would kick everything out at the same time. So, the moment you put the taint, everything would be kicked out, and that's bad. Production would go down. So, this system slowly migrates the pods back so production doesn't go down and we don't waste money. It is a very self-healing system.