.
Hello, everyone. My name is Gabriel. I work at Korea, and I will be talking about the infrastructure that allows us to train K2 and also how we serve it. So what is K2? K2 is our pre-trained from-scratch model. We just released it less than a month ago. And the whole idea behind training this model was because we were bored of AI images. They are quite saltless. They have no spice. And the whole idea was we want to give creatives tools to explore out of distribution, extremely interesting images, do composition, and actually give tools to creatives. And that was the whole idea of the model.
So the model was trained from scratch, no base checkpoints, nothing, everything done in-house. And also, as I said, built for exploration. And right now, this is what you can get out of K2: very different styles, pixel art, photo real, and some silly stuff. And the whole idea of the model has been, let's explore this medium. Create2 is open source. Right now, you can go play with it. There are two checkpoints. We also serve it in production. There's a raw checkpoint, which is pre-trained so people can post-train and do whatever they wish to do with it. And there's also the post-trained version, which is the turbo one, which is very, very fast.
You can get an image in less than a second. And this is the type of image you can get in less than a second. On Huggy Face, GitHub, go play with it. Also, you can go to production on kreya.ai and go play with it. So let's talk about how we train this model. As I said, it's going to be how we train and how we serve. First, the model was trained from scratch on thousands of GPUs. We have a big cluster, one main cluster with a lot of GPUs, all InfiniBand-connected. And you put those GPUs to work and it's trained. But I wish it were that simple, but it's not. So at the beginning, we did a bunch of small ablations on a small number of GPUs to see how things go, right?
So you want to test some hypotheses, and you use a small number of GPUs and let it train for a little bit. Oh, this works, it doesn't work. Let's scale. And as we were training this model, the whole idea was to bridge the gap between LLM research and diffusion transformers. So my researchers ported a lot of research from LLMs into DITs. And the whole architecture of the model was meant to be extremely, extremely simple. And so it is very, very dumb, but very effective. And so let's start talking about numbers. Incredibly, maybe still an issue on our part, maybe our cluster was very interesting as we were scaling.
When you do small experiments, experiments would run for days and even less than we would like to, but they would still run fine. And as we started scaling, getting more and more GPUs, like 128, 256, 512, whatever number, and you scale and scale, things are crashing more. That's expected, right? There's more surface area for things to break, and things are going to go wrong. And a lot of the time, things would go wrong in silent ways. I know, NCCL timeout just crashes, and the metrics are all good. And it is extremely annoying stuff. At the beginning, we were paranoid, swap nodes, change nodes, whatever, whatever, whatever.
And we learned that sometimes you just let it crash. It crashes, it runs for an hour, it crashes, it runs for an hour, it crashes, and then it runs again on the same setup, machine, same code, same data for 12 hours, 16 hours, 24 hours. There is this paper from Meta that gives you a rough estimate of how many failures you should expect. We kind of saw the same pattern, but not the same numbers. Our runs would last way, way less than this. So it was extremely annoying, and you can imagine doing large KOP training on runs that last less than eight hours. It is a problem, right?
You want to keep those GPUs fed, and if things are crashing, you are not making progress and are losing time, and the model is going to be delayed. So for us, extremely important, and at least from the infra side, was to get metrics. Metrics are everything. That's how I can support my researchers. That's how I have visibility in the system. And if you're doing large KOP training, I highly, highly recommend for you to invest heavily in metrics. Don't go blind, because you're going to go crazy. So I'm going to share some of the metrics that were important for us, and they're quite silly, but extremely effective. First one, GPU temperature. GPUs are very, very annoying.
If you have a single GPU that is a bit warmer than the others, it's going to start throttling and slow down, and the training is going to be unstable and have weird problems. So for us, it was, if there are any GPUs above 78 degrees, you remove them. Don't think about it. Don't try to fix it. Don't try to be smart. Just remove the GPU. You're going to save time, and just ask your provider, this GPU is hot. Please replace it. And there are two metrics that, at the beginning, we did not fully understand. And as we are getting more and more used to GPUs and how they work, there is GPU utilization, which is a lie. Don't trust this. This is dumb.
This tells you, oh, the GPU is doing work. And this is the amount of time the GPU is doing work, but not good work. Not how efficiently the GPU is working. So, as you can see during our pre-training, the GPU is at 100%. This is not true. We are not fully utilizing the GPU. This is 100% a lie. What we would use as a proxy was TensorFlow utilization. This is actually how much of the TensorFlow cores you're using and how effective they are being. And it was also very interesting, as we're doing pre-training, and then you go from pre-training to mid-training to post-training, you start scaling the resolution of the images.
So, for example, in pre-training, we do 128, 256, 512, 1024 pixels as they scale. And you could see the TensorFlow core utilization go up as we would scale on these resolutions, because now we're doing more work on images. And also very, very interesting were InfiniBand and NVLink metrics. These, by default, are not exported by the NVIDIA metrics, the DCGM stuff. Some NVLink stuff, yes, but no InfiniBand. So, if you don't have InfiniBand metrics, go get them. I'm telling you right now, if you're doing large-scale training with a bunch of GPUs talking to each other between machines, and you have no InfiniBand metrics, you are doing something wrong.
This was probably the most important stuff for us, because most of our failures were related to cross-node communication. So, for example, here, you just have throughput. But on our refining dashboard, we have a bunch of stuff, from wait, like when a message is sent on the fabric, how much time the message is waiting, or the number of errors and different types of errors and number of packets, and all of the things that InfiniBand exports, we collect. We had to build custom stuff to get this. It was not hard. You can figure it out. It was very, very easy. Same thing for NVLink. NVIDIA exports some stuff about NVLink.
But, for example, NVLink errors, NVIDIA doesn't export this. So, you can collect this. And, as I said, InfiniBand was extremely important. NVLink was a little bit less. In some cases, these helped us catch some problems, especially because NVLink is single nodes, right? The communication inside a node. So, sometimes a single node would have a weird failure where the GPU seemed to be fine, but some weird error happens. And then you can look at NVLink errors and you see all errors happening. And then replace that machine. So, go get these metrics. and all of the things that infinite bands exported, we collect. We had to build custom stuff to get this. It was not hard.
You can figure it out. It was very, very easy. Same thing for NVLink. NVidia exports some stuff about NVLink. But, for example, NVLink errors. NVidia doesn't export this. So, you can collect this. And, as I said, infinite band was extremely important. NVLink was a little bit less. In some cases, these helped us catch some problems, especially because NVLink is single nodes, right? The communication inside a node. So, sometimes a single node would have a weird failure where the GPU seemed to be fine, but some weird error happens. And then you can look at NVLink errors, and you see all errors happening. And then replace that machine. So, go get these metrics.
They're extremely important. And without this, we would not be able to train at all. Also, as I said, our trains would crash constantly. And a hecky way to do it to fix the problem is just checkpoints. Use and abuse the file system that you have. At the beginning, we used Ceph. Ceph did not work well. It was very annoying. It broke. We did not trust the data. So, I recommend if you have the money, go with something paid because you can trust your data. You can see numbers. This is our Weka cluster. We can do 1.2 terabytes a second of reads, almost a terabyte of writes. And the file system would not choke on the training.
So, we could checkpoint every 30 minutes, 20 minutes, produce a terabyte of data in less than 30 seconds. So, this would not delay trainings. That was probably one of the most important things that we did to recoup the loss. Just checkpoint. Don't think about it. And how we serve. This goes in connection on how the trainings are launched. Because at the beginning, I don't want my researchers to think about GPUs. I just want them to launch stuff. And this goes into a queue. And if we have GPUs, we have GPUs. If we don't have GPUs, we don't have GPUs. So, queue is an open source project. You can look it up. It does game scheduling.
Game scheduling is extremely important for trainings in general. And this gives us a semantic of two tiers of priority where you have a workload priority. And these, you can say, this training is more important than this one. So, it skips in the queue, in front of the queue. And then, after this, we have the normal Kubernetes priority for use to Kubernetes. And the way the system works, the training pods, they always have high priority for everything. So, once they are admitted, they're going to schedule. If there is inference running on those machines, the inference gets kicked out. And you would say, oh, this is bad. Production is going to go down. No.
You can build on top of that to make production not go down, which is very, very cool. The only one of the problems with queue, which is annoying, you can automate that. We have not. It's just that you specify the queues. The queues have amount of resources: CPU, GPUs, memory, whatever. But this is manually specified. And at least our cluster is quite fluid. Nodes phasing in and out of existence. They go into maintenance, whatever. You lose a few nodes here and there. This number gets out of sync. And sometimes this would break gang scheduling. So, FYI, this is a bit annoying. You're going to face this if you use queue. But, yeah, very good project. Kubernetes 135.
We have not had the chance to play with it. Has gang scheduling out of the box. Something very similar to queue. So, maybe you can use Kubernetes 135. And, as I said, this is the system that we built that allowed us to train using the whole cluster. As I said, we have one big cluster that runs production and trainings. So, I don't want my researchers to think about GPUs. And I don't want to make the trainings and research be delayed because production is running, right? Production is lower priority. The site still needs to work. People still need to use the website. But the GPUs, the value that we get of the GPUs doing training is higher than we get out of production.
So, the whole system works by default where there is this magical system that I'm going to explain in a bit that allows us to flip traffic between clusters magically. And not just clusters: external providers, GPU rentals, whatever. And you can see the green, the dark green is inference running in cluster. And then someone launches a train. And then suddenly it starts flipping to the other cluster. And then training is done or whatever happens, it flips back. So, we stop wasting money. And this is seamless. No one needs to think about it. The whole system handles itself. And you get this very nice pattern of I'm going to use all the GPUs in my cluster for trainings.
Production is going to run somewhere else. I don't need to think about it. My users in production, they're not going to feel anything. Research is going to be happy. And we can get values out of the GPUs. So, how does this work? There's this very nice project called Virtual Kubelet, also open source. You can build on top of it. It is a very nice code base. And this works by creating a fake machine in Kubernetes. Kubernetes has nodes. This creates a fake machine that is up to you to control how it works. So, Kubernetes does normal scheduling, as you would expect. Things would go into these nodes. For example, here, all of the GPUs in the cluster are used, right?
So, this pod goes into the Virtual Kubelet node. And in there, you can do whatever. This is the system that we built. There is, you receive the pod spec and then you find a provider. This is up to you. Let's say you have a deal with some provider that gives you nice prices. You integrate into here. We built some nice interfaces to not leak things. So, we just implement a provider and there's an algo that decides which one it goes to. You translate the spec of the pod into the provider stuff and you deploy. And then you have something that reconciles between both sides. And it was extremely, extremely nice.
If you guys know about Kubernetes, Kubernetes has the horizontal pod out of the scaler which scales the number of replicas. Number of replicas. Could you stop being annoying? Thank you, sir. I appreciate it. There you go. Let's go back. Now we go back. Back, back. There you go. Kubernetes has the HPA, and the HPA scales the pods. And so, if something fails, it is very interesting. You don't need to handle the fail. The only thing you need to handle is, oh, something failed. You mark the pod as failed. Kubernetes is going to detect that something has failed and create a new one. You don't need to try to save the world.
Let Kubernetes handle it for you, which is an extremely nice way of handling stuff. If something breaks on your side, something breaks on the other side, just mark as failed. Let Kubernetes handle it. Create a new one. And things keep working. Very, very nice way to handle stuff. And also very interesting, let's say you have GPUs on your cluster available, right? You don't want to waste money. This would be very, very bad. So, the system works using taints, Kubernetes taints. They allow and disallow things to run. Pods have tolerations for the taints. And when we have GPUs in the cluster, if you look back, there is the taint system at the bottom.
This taint system, it is what would by itself decide if we have GPUs or not GPUs in the cluster. And this adds a taint into the node when we have a lot of GPUs. So, a lot of GPUs in the cluster. We taint the node. Nothing can schedule on it. So, we stop wasting GPUs. The pods, they go into the GPUs in the cluster. We don't waste money. And then imagine someone launches a train, right? This train is going to take all the GPUs in the cluster. It's going to hog all of the GPUs. No GPUs in the cluster. The system detects this, removes the taint, new pods schedule there. Very, very nice. You also don't think about it.
And for us, for example, we use just some Prometheus metrics. That's how we do it. It's very simple. But it works very, very, very well. You let the system run by itself. And this adds a taint into the node when we have a lot of GPUs. So, a lot of GPUs in the cluster. We taint the node. Nothing can schedule on it. So, we stop wasting GPUs. The pods, they go into the GPUs in the cluster. We don't waste money. And then imagine someone launches a train, right? This train is going to take all the GPUs in the cluster. It's going to hog all of the GPUs. No GPUs in the cluster. The system detects this, removes the taint, new pods scheduled there. Very, very nice.
You also don't think about it. And for us, for example, we use just some Prometheus metrics. That's how we do it. It's very simple. But it works very, very, very well. You let the system run by itself. Someone's going to launch stuff. You're going to, the train is going to kick out the pods. It's going to schedule. It's going to take the GPUs. The system is going to detect that, remove the taint. Pod scheduled there. Very nice. Someone finished the train. Now we have pods running on the other side. You're wasting money. How do we fix this? Same thing. You run something else that detects the system and removes things back.
In this case, a descheduler, once the taint is added back, GPUs available, we add the taint. The descheduler says, oh, these pods, they don't tolerate the taint. I'm going to migrate them back. And you can ask, oh, why don't you use a no execute taint? No execute in Kubernetes would kick everything out at the same time. So, the moment you put the taint, everything would be kicked out, and that's bad. Production would go down. So, this system slowly migrates the pods back so production doesn't go down and we don't waste money. It is a very self-healing system.