Hello everyone. Thank you for coming. I'll start the talk now. So this talk is around a paper that I did, which is very simple. The proposition is very clear. It's two lines of algebra that make the RMSNorm layer in transformers cheaper, quicker, and improve it as a layer in the transformer architecture. Similar to how layer norm once used to be the standard and then it was substituted by RMSNorm. This follows along this way of thinking. And I got the chance to meet some people from the open source world, and I co-authored this paper together with Nils Graf, who was the creator of this. And the work follows from there. So this is presented on archive. You can have a look, read it, test it out. There is a repo as well. And the concept, the idea and the way of thinking, I think it's easiest to explain with flash attention. So in a similar way that flash attention waits until there is a multiplication and tries to limit the communications between memory so that the whole process is faster. This is a similar thought along those lines. And it does certain improvements that make the RMSNorm process much quicker and in effect improve the whole transformer. And one question is: why RMSNorm? Since that layer does almost none of the math. And that's true. So the share of the math portion, if you look at it, is quite small. However, the clock time or wall time, as they say, is quite big. And for example, in one decode step, so when inferences perform, the RMSNorm can be started 33 times. Of course, it depends on the model and so on. In the paper, you have the specific models and how this was tested. And the question is how this can be improved and how the wait for the matrix multiplication can be avoided. And the reason why this is slow is because the GPUs are not slow or bad at math, but they're bad at everything else around the actual math. So that means starting the work, the actual work. So for example, starting the process as it happens in some of the experiments 33 times, that takes a long time. And for example, fusing each normalization into the matrix multiplication can help avoid this. Also doing weight folding can help in moving data between memory. And that's a process that's also slow for GPUs. And also weighting. So for example, deferring the division that's done in the RMSNorm layer is also a way to avoid this weighting step. So basically what this paper does is it improves all these three aspects by doing a few algebraic tricks in the way RMSNorm is computed. That's it. And math-wise, these are the tricks. It's mainly around the first two propositions. One is weightless normalization. You can see that here. And deferred normalization. So that's the second one. And now in more newer architectures, there is a situation where RMS can appear twice. For example, in GEMMA 4 this happens. So cancelling the pre-normalization also works. And all of this is algebraically proven in the paper. And the first proposition is this, where the gain and the weight fold into one matrix. W, you can see here with an asterisk. And that is computed offline, similar to how in flash attention you compute some stuff on the side so that there is no communication between memory all the time. So this is one step that's done, this weight folding. And the other step is deferring the scalar divide of the matmul so that they can be done in parallel. So in a normal case, you would have to compute once, then weight, and compute again. In this case, the idea is to split this so that it can be parallelized. And the third one, which is a version of this, is that there is, if there are two, because this is scale invariant, one of them can be dropped and this still works. And this is applicable to newer models that can have this architecture and implementation. So in order to make this happen in real life, especially this proposition number two, so for this one, for example, it's easy. There is a repo called transformer tricks. You can just apply this to any model and it works. But in order to do this, there is some kernel work. So it's not as straightforward to do. So in order for me to do that, I was implementing this and I came out with this experiment once. So it looks okay in general, where it's like, okay, the prompt is the transformer architecture, revolutionized NLP because, and then there is some expected output. But in the output I got, I saw this repetition and one step lag, as you can see here, the word because appears again. And there was something happening with the GPU streams and I was trying to figure out what was happening. And I was getting this one step lag and outputs that were from the past in a way. And in debugging all of this, I realized that in the process of building something like this, so as I explained the proposition of deferring these operations. In CUDA, you can do two things. You can do tensor cores that do one part of the matrix multiplication and you can do CUDA cores that run stuff like element-wise operations, reductions, square roots, and so on. So the idea was to do this in parallel and get the benefit of what I was explaining in the paper to actually test out this concept. So this is how it was supposed to look. So there is, if you do things sequentially, there is this idle waiting time when the vector unit computes the RMS and scaling. And then there is a matrix multiplication. So the idea was, okay, with flash norm, which is the technique in the paper, you're supposed to do those both in parallel. So the matrix unit computes the matmul and the vector unit computes the RMS. So in that way you save time. However, you cannot just do this in Python. You have to go a bit lower. And I did that with CUDA code like this. And this looked in general okay at that time. However, I realized that I did something slightly wrong. And that thing was that the join in the end where you're supposed to join the two streams was implicit in my case. And when I tested this out, the unit test worked. The quality seemed similar like perplexity testing and so on because it's just similar generation. But over long generation I was able to see this problem. So I had no idea what this was. And the reason was that when I was doing this implicit join, basically one of the streams hadn't finished the work. So I got race conditions that read the past from the unfinished matrix multiplication. So the idea that I had to fix this was around the fact that I had to be explicit about the join and wait until one of the operations is finished so that I'm certain that when I join I'm not reading from the past. So that was the realization in this exploration of CUDA streams. So the post scale read an old buffer value. And how this is fixed is with this where basically you need to mark the end of the matrix multiplication. Then mark the end of the RMS. And then post scale wait for the first stream. And then wait for the second stream. And that fixed the bug and made the paper work and the model speak forwards instead of backwards. And that was the cool academic perspective. But I also wanted to try things, right? Deploy this, test it out, see how I can make it work in a more production setting. And you can also read the paper and see all the tests. Some are done, most are done around Llama models. But this works for other architectures as well. So what you can do for this specific paper is, for example, the weight folding that I explained, the proposition one. You can just do it with some code in the repo. That's like flash, you say flashify and it does that. However, with this second thing that I mentioned, you need to do a bit of kernel work if you want to do that. As I explained in my example. And these are some results that are based on Llama models. And there are different details that you can have a look at as well. Like what happens if you do only deferred normalization? What happens if you do a full fused kernel? So there are a lot of experiments going lower here to test all the propositions. And these have been our results in different levels of scrutiny and detail. But even the simple one with weight folding shows some improvement. And this also works with the day-to-day tools that you use in a model. So it's not like you have to reinvent the wheel or do things from scratch. So it works with Torch Compile because it's a new checkpoint and that's it. Flash attention does similar tricks at a different layer. And also it works with quantized models. So it's totally cool to actually apply this and you can get a model that has this cool new normalization layer. And where you can get these details and codes to actually run this is this transformer tricks repo. So it has different algebra tricks as I explained as well as this paper that I mentioned. And also there is the GitHub, not the GitHub, but the HuggingFace model repo where I've done this with some models. And you can have a HuggingFace link to the model and test it out. And what you also can do with this HuggingFace models is to deploy them in production. So when I was thinking about doing this, I realized that, okay, now that the science is done and there is a link to a HuggingFace model, Superlink's inference engine was a cool way to actually deploy any HuggingFace model. And we've done this at hackathons where people would bring a custom HuggingFace model or checkpoint that they have with their fine-tuned stuff. And you can test out, even if you have some version of this algebraic tricks that you want to improve a model and test your own research ideas. You can actually try that out and have a deployed version on a cluster of this model and not have to worry about this glue code around deploying models. So that's pretty cool. And the point is that if you have the full cluster open source and the model inference open source, you can actually test out this kind of more novel research ideas where if you want to do kernel manipulation or flash norm and things like that, it's much more difficult to do this at the rented endpoint where you don't own the inference. You want something that's portable and flexible to actually allow you to do this stuff, but it's also production ready enough so that you can test things out at scale. And you can, for example, use Skye to combine this with other models. Like as you can see in the top left, you can have this flashified models with different other models to do agentic tasks if you want and do that end-to-end bigger use case. And the way Skye works is this production cluster helps you deploy the models. So you can have a look at Skye's repo as well for more details on this. And also there is a smarter queuing mechanism that helps you, especially if you work with smaller models. Because when doing the flash norm stuff, I worked with smaller Llama models and also with small agents from HuggingFace. So having a way to deploy smaller models that can also work on the same GPU so that you don't have to spend your money on GPU costs, but actually switch models around, especially smaller models. It was quite useful. And you can also control the model configs through an API as well as the cluster, which is also pretty convenient without having an infra person supporting you in your open source research. So that's cool as well. And you own your cloud, which is useful if you want open weights, open models, open source. And there is also a catalog that Skye has of different models. Not just the ones I mentioned, but you can have a look. There's also re-ranking embedding models if you're building something along those lines. And with that, I'm finishing this story of my research journey where I co-authored this paper around the technique that improves the transformer, but also found a way to bring this to production and test it out and find a way to play around with this open source models. And feel free to contact me on LinkedIn. Maybe if you have any questions or contributions. A lot of this stuff that I've mentioned, like some of them are PRs on VLLM or on Hugging Face. You might find them all around. You can also check out the paper. That's the archive link that you have there. And you also have the Skye repo and my LinkedIn. So thank you very much for attending. And you can catch me for questions. We'll be here close by.
Hello everyone. Thank you for coming. And I'll start the talk now. So, this talk is around a paper that I did, which is very simple. The proposition is very clear. It's basically two lines of algebra that make the RMSNorm layer in transformers cheaper, quicker, and kind of improve it as like a layer in the transformer architecture. Similar to how layer norm once used to be the standard and then it was substituted by RMSNorm. This follows along this way of thinking. And I got the chance to kind of meet some people from the open source world, and I co-authored this paper together with Nils Graf, who was the kind of the creator of this. And the work follows from there.
So, this is presented on archive. You can have a look, read it, test it out. There is a repo as well. And the concept, the, let's say the idea and the way of thinking, I think it's easiest to explain with maybe flash attention. So, in a similar way of how flash attention kind of waits until there is a multiplication and tries to limit this communications between memory so that the whole process is faster. This is kind of a similar thought along those lines. And it does certain improvements that make the RMSNorm process much quicker and in effect improve the whole transformer. And one question is, okay, why RMSNorm? Since that layer does almost none of the math.
And that's true. So, the share of the kind of math portion, if you look at it, is quite small. However, the clock time or wall time, as they say, is quite big. And for example, in one decode step, so, right, when like inferences perform, the RMSNorm can be started like 33 times. Of course, it depends on the model and so on. In the paper, you have the specific models and how this was tested. And the question is how this can be improved and how this wait for the matrix multiplication can be kind of avoided. And the reason why this is slow is because the GPUs are not slow or bad at math, but they're bad at everything else around the actual math.
So, that means starting the work, the actual work. So, for example, starting the process as it happens in some of the experiments 33 times, that takes a long time. And for example, fusing each normalization into the matrix multiplication can help avoid this. Also doing weight folding can help in kind of moving data between memory. And that's a process that's also slow for GPUs. And also weighting. So, for example, deferring the division that's done in the RMSNorm layer is also a way to avoid this weighting step. So, basically what this paper does is it improves all these three aspects by doing a few algebraic tricks in the way RMSNorm is computed. That's it.
And math-wise, these are the tricks. It's mainly around the first two propositions. One is weightless normalization. You can see that here. And deferred normalization. So, that's the second one. And now in more newer architectures, there is a situation where RMS can appear twice. For example, in GEMA 4 this happens. So, cancelling the pre-normalization also works. And all of this is algebraically proven in the paper. And the first proposition is this, where kind of the gain and the weight fold into one matrix. W, you can see here with an asterisk.
And that is computed offline, similar to how maybe in flash attention you compute some stuff on the side so that there is no communication between memory all the time. So, this is one step that's kind of done, this weight folding. And the other step is deferring the scalar divide of the math-mull so that they can be done in parallel. So, in a normal case, you would have to compute once, then weight, and compute again. In this case, the idea is to kind of split this so that it can be parallelized.
And the third one, which is kind of a version of this, is that there is kind of, if there are two, because this is scale invariant, one of them can be dropped and this still works. And this is applicable to newer models that can have this architecture and implementation. So, in order to make this happen in real life, especially this proposition number two, so for this one, for example, it's easy. There is a repo called transformer tricks. You can just apply this to any model and it works. But in order to do this, there is some kernel work. So, it's not as straightforward to do.
So, in order for me to do that, I was implementing this and I came out with this experiment once. So, it looks okay in general, where it's like, okay, the prompt is the transformer architecture, revolutionized NLP because, and then there is some kind of expected output. But in the output I got, I saw this repetition and one step lag, as you can see here, the word because appears again. And there was something happening with the GPU streams and I was trying to figure out what was happening. And I was getting this one step lag and kind of outputs that were from the past in a way.
And in debugging all of this, I realized that in the process of building something like this, so as I explained the proposition to or deferring this to operations, In CUDA, you can do two things. You can do like tensor cores that do one part of the matrix multiplication and you can do CUDA cores that kind of run stuff like element-wise operations, reductions, square roots, and so on. So, the idea was to do this in parallel and get the benefit of what I was explaining in the paper to actually test out this concept. So, this is how it was supposed to look like.
So, there is, if you do things sequentially, there is this idle waiting time when the vector unit computes the RMS and scaling. And then there is a matrix multiplication. So, the idea was, okay, with flash norm, which is the technique in the paper, you're supposed to do those both in parallel. So, the matrix unit computes the matmul and the vector unit computes the RMS. So, in that way you save time. However, you cannot just do this in Python. You have to go a bit lower. And I did that with CUDA codes like this. And this looked in general okay at that time. However, I realized that I did something slightly wrong.
And that thing was that the join in the end where you're supposed to join the two streams was implicit in my case. And when I tested this out, the unit test worked. The quality seemed similar like perplexity testing and so on because it's just like similar generation. But over long generation I was able to see this problem. So, I had no idea what this was. And the reason was that when I was doing this implicit join, basically one of the streams hadn't finished the work. So, I got race conditions that kind of read the past from the unfinished matrix multiplication.
So, the idea that I had to fix this was around the fact that I had to be explicit about the join and wait until one of the operations is finished so that I'm certain that when I join I'm not reading from the past. So, that was the realization in this exploration of CUDA streams.
So, the post scale read like an old buffer value. And how this is fixed is with this where basically you need to mark the end of the matrix multiplication. Then mark the end of the RMS. And then post scale wait for the first stream. And then wait for the second stream. And that fixed the bug and made kind of the paper work and the model speak forwards instead of backwards. And that was the cool maybe academic perspective. But I also wanted to try things, right? Deploy this, test it out, see how I can make it work in maybe a more production setting. And you can also read the paper and see all the tests. Some are done, most are done around LAMA models.
But like this works for other architectures as well. So, what you can do for this specific paper is, for example, the weight folding that I explained, the proposition one. You can just do it with some code in the repo. That's like flash, you say flashify and it does that. However, with this second thing that I mentioned, you need to do a bit of kernel work if you want to do that. Like I explained in my example. And these are some results that are based on LAMA models. And there are different kind of details that you can have a look at as well. Like what happens if you do only deferred normalization? What happens if you do a full fused kernel?
So, there are a lot of experiments of going lower here to test all the prepositions. And these have been our results in different, let's say, levels of scrutiny and detail. But even the simple one with like weight folding shows some improvement. And this also works with like the day-to-day tools that you use in a model. So, it's not like you have to reinvent the wheel or, you know, do things from scratch. So, it works with Torch Compile because it's kind of like a new checkpoint and that's it. Flash attention does similar tricks at a different layer. And also, it works with quantized models.
So, it's totally cool to actually apply this and you can get a model that has this cool new normalization layer. And where you can get this details and codes to actually run this is this transformer tricks repo. So, it has different algebra tricks like I explained as well as this paper that I mentioned. And also, there is the GitHub, not the GitHub, but the HuggingFace model repo where I've done this with some models. And you can have a HuggingFace link to the model and test it out. And what you also can do with this HuggingFace models is to deploy them in production.
So, when I was thinking about doing this, I realized that, okay, now that let's say the science is done and there is a link to a HuggingFace model.
Superlink's inference engine was a cool way to actually deploy any HuggingFace model. And we've done this at Hackathons where people would bring like a custom HuggingFace model or checkpoint that they have with their fine-tuned stuff. And you can test out like even if you have some version of this algebraic tricks that you want to improve a model and test your own research ideas. You can actually try that out and have a deployed version on a cluster of this model and not have to worry about this glue code around deploying models. So, that's pretty cool.
And the point is that if you have the full cluster open source and the model inference open source, you can actually test out this kind of maybe more novel research ideas where if you want to do kernel manipulation or flash norm and things like that, it's much more difficult to do this at the rented end point where you don't own the inference. You want something that's portable and flexible to actually allow you to do this stuff, but it's also production ready enough so that you can test things out at scale. And you can, for example, use site to combine this with other models.
Like as you can see in the top left, you can have this flashified models with different other models to do agentic tasks if you want and kind of do that end-to-end bigger use case. And the way site works is this production cluster helps you deploy the models. So, you can have a look at size repo as well for more details on this. And also, there is a smarter queuing mechanism that helps you, especially if you work with smaller models. Because when doing the flash norm stuff, I worked with like smaller Lama models and also with small agents from HikingFace.
So, having a way to deploy smaller models that can also work on like the same GPU so that you don't have to spend your money on GPU costs, but actually kind of switch models around, especially smaller models. It was quite useful. And you can also control the model configs through an API as well as the cluster, which is also pretty convenient without having like an infra guy supporting you in your open source research. So, that's cool as well. And you own your cloud, which is useful if you want open weights, open models, open source. And there is also like a catalog that Sai has of different models. Not just the ones I mentioned, but you can have a look.
There's also re-ranking embedding models if you're building something along those lines. And with that, I'm kind of finishing this story of my research journey where I co-authored this paper around the technique that improves the transformer, but also found a way kind of to bring this to, let's say, production and test it out and find a way to play around with this open source models. And feel free to contact me on LinkedIn. Maybe if you have any questions or contributions. A lot of this stuff that I've mentioned, like some of them are PRs on like VLLM or on Hugging Face. You might find them all around. You can also see the, check out the paper.
That's the archive link that you have there. And you also have the Sai repo and my LinkedIn. So, thank you very much for attending.
And you can catch me for questions. We'll be here. Close by. .