Open Reader

Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards — Filip Makraduli

completed 17:18 Sep 19, 2026 Watch on YouTube

Current Status

completed

Video ID

c1hGBoWw20A

RAG / Chat

Enabled
Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards — Filip Makraduli
Description

An RMS norm layer does almost none of the arithmetic in a transformer, yet a single decode step can launch it around 33 times, and a GPU is fast at math and slow at everything else: starting work, moving data, waiting. That gap is what FlashNorm attacks. Filip Makraduli wrote the paper with Nils Graef, and the idea fits in two lines of algebra. Fold the norm's gain into the projection weights offline so one matrix absorbs both. Defer the scalar divide so the matrix unit and the vector unit run at once instead of one idling for the other. And in newer architectures that normalize twice in a row, drop one, because the operation is scale invariant and the second adds nothing. Together they buy a 33 to 35 percent speedup on the norm plus projection operation, and the folded checkpoint works with torch compile and quantized models. The deferral is where it got interesting, because you cannot do it from Python. He wrote the CUDA to run the matmul on tensor cores and the RMS reduction on CUDA cores in parallel. Unit tests passed, perplexity looked normal, and then over a long generation the model began repeating itself with a one step lag, outputs from the past. The join between the two streams was implicit. One stream had not finished, so the post scale read a stale buffer from an unfinished multiply. The fix was to mark the end of each stream explicitly and make the post scale wait on both. He closes on why the second half needed Superlinked's open inference engine to deploy a modified checkpoint, since you cannot do kernel surgery on a rented endpoint. Speaker info: - https://x.com/f_makraduli - https://www.linkedin.com/in/filipmakraduli/ - https://filipmakraduli.substack.com/ Timestamps: 0:00 - Two lines of algebra for RMS norm 2:07 - Why a layer with no math costs so much time 3:44 - Weight folding, deferred division, dropped pre norm 6:33 - The output that came from the past 7:25 - Tensor cores and CUDA cores in parallel 8:45 - An implicit join and a race conditio

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Skim
  • Core thesis: FlashNorm reformulates RMSNorm so its gain can be folded into adjacent weights offline and its normalization divide can be deferred or parallelized with matrix multiplication, reducing inference wall-time overhead without changing the model's intended computation.
  • Why it matters: RMSNorm is computationally small but repeatedly introduces kernel-launch, memory-traffic, and synchronization costs during decoding; this offers a potentially practical optimization path for self-hosted open-weight inference stacks.
  • Best use: Use the talk as a concise implementation and debugging case study before evaluating the paper/repository for a targeted benchmark on models and hardware you control.

Executive Summary

Makraduli presents FlashNorm, a set of algebraic RMSNorm transformations co-authored with Nils Graf. The central observation is that normalization is not expensive because of FLOPs, but because autoregressive inference repeatedly pays GPU overhead around small operations: kernel launches, memory movement, and dependencies. He cites an example where RMSNorm is invoked 33 times in one decode step.

The first transformation, weight folding (or weightless normalization), combines RMSNorm's learned gain with the following projection weight offline, producing a modified checkpoint. A second transformation defers the scalar normalization divide until after matrix multiplication, allowing matrix/tensor-core work and RMS/vector-core work to proceed concurrently. Where architectures contain two compatible RMSNorms, such as Gemma 4 according to the talk, scale invariance can permit cancellation of a pre-normalization.

The most useful portion is the implementation failure mode: an initial multi-stream CUDA kernel produced plausible unit-test and perplexity results yet generated long outputs with a one-token lag and repeated words. The cause was an implicit rather than explicit join between asynchronous CUDA streams; post-scaling read a stale output buffer before the matrix multiplication completed. Explicit completion events and waits on both streams corrected the race.

The talk argues the lightweight weight-folding transformation can be applied from the Transformer Tricks repository and remains compatible with torch.compile and quantized models, while deferred normalization requires custom kernel work. Its production section is largely a pitch for deploying modified Hugging Face checkpoints on an open, self-managed cluster, but the underlying operational lesson is sound: custom inference-level optimizations require ownership of the serving path rather than a fixed managed endpoint.

Key Takeaways

  • Claim: RMSNorm can be a meaningful decode-latency target despite representing little arithmetic, because its repeated orchestration and memory costs dominate its wall-clock contribution. | Evidence: The speaker says RMSNorm may start 33 times in a single decode step in tested configurations, and attributes its cost to kernel/work launch, memory movement, and waiting rather than GPU mathematical throughput. | Implication: For self-hosted inference, profile per-token latency at the kernel and synchronization level rather than ranking optimization candidates by FLOP share alone. | Caveat: The 33-call figure depends on model architecture and configuration; the transcript does not provide absolute latency gains or hardware-specific benchmark numbers.
  • Claim: FlashNorm's easiest optimization is offline weight folding: absorb RMSNorm's learned gain into an adjacent matrix weight so the inference path removes a separate scaling operation. | Evidence: The talk describes folding the gain and weight into a modified matrix W* computed offline, and says the Transformer Tricks repository can apply this transformation to a model as a new checkpoint. | Implication: This is the low-friction experiment: transform a candidate open-weight checkpoint, run correctness and throughput comparisons, and retain the original checkpoint for rollback. | Caveat: The speaker claims improvement but gives no numerical speedup in the transcript; implementation and validation should be benchmarked on the target model and serving engine.
  • Claim: Deferred normalization can expose parallelism by running matrix multiplication on tensor cores while RMS-related vector operations run concurrently on CUDA cores, then applying the scale after both complete. | Evidence: Rather than serially computing RMS/scaling and then matmul, the proposed kernel overlaps matmul with RMS computation; the speaker contrasts tensor cores for matmul against CUDA cores for elementwise operations, reductions, and square roots. | Implication: Treat deferred normalization as an inference-kernel engineering project with a higher integration and maintenance burden, not as a generic checkpoint conversion. | Caveat: Unlike offline folding, this approach requires lower-level fused/custom CUDA kernel work and cannot be obtained merely through Python-level changes.
  • Claim: Asynchronous CUDA stream bugs can evade standard quality checks while corrupting long-form generation. | Evidence: The faulty implementation passed unit tests and appeared comparable under perplexity testing, but longer generation showed repeated tokens and a one-step lag, including an extra occurrence of the prompt word "because." | Implication: Any custom parallel inference kernel needs long-horizon generation regression tests and deterministic stream-synchronization tests in addition to unit tests and aggregate perplexity. | Caveat: The failure described is specific to this two-stream implementation, but it illustrates a broader validation gap whenever output buffers are shared across asynchronous GPU work.
  • Claim: The observed backward/lagged generation was caused by a race condition: post-scaling consumed an old buffer value because the matrix-multiplication stream had not completed. | Evidence: Makraduli says the final join was implicit; the fix was to explicitly mark completion of both matmul and RMS work, then make post-scale wait on the first stream and the second stream. | Implication: In multi-stream kernels, define data ownership and event dependencies explicitly; never infer completion merely from kernel launch order or seemingly correct short outputs. | Caveat: Explicit synchronization can reduce or eliminate the anticipated overlap if dependencies are incorrectly designed, so performance must be measured after correctness is restored.
  • Claim: The optimization is positioned as operationally compatible with existing open-model workflows, including torch.compile and quantized models. | Evidence: The speaker says weight-folded output is simply a new checkpoint, reports compatibility with torch.compile and quantized models, and notes experiments were primarily on Llama models while asserting applicability beyond them. | Implication: The practical adoption path is strongest for controlled open-weight Llama-family deployments, with architecture-by-architecture validation before broad platform standardization. | Caveat: Most reported tests were around Llama models; compatibility claims for other architectures and quantization formats should be independently verified.

Detailed Brief

Algebraic scope and architecture conditions

  • Claims: FlashNorm comprises more than one optimization: weightless normalization via folding, deferred normalization via reordered scaling, and a cancellation opportunity when two compatible RMS normalization locations appear.; The speaker characterizes the latter as enabled by scale invariance and identifies Gemma 4 as an example of a newer architecture where RMSNorm can appear twice.
  • Evidence: The paper is described as providing algebraic proofs for the propositions.; The speaker analogizes the design philosophy to FlashAttention: precompute or reorder work to reduce recurrent memory communication and waiting.
  • Caveats: The transcript does not state the exact layer-placement requirements, numerical tolerances, or architecture coverage for the double-RMS cancellation.; The FlashAttention comparison is an analogy in optimization philosophy, not a claim that FlashNorm uses the same algorithm or produces comparable magnitude of gains.
  • Implications: Model graph structure determines which transformation is legal; conversion tooling should inspect architecture rather than apply a uniform rewrite.; The paper and code, rather than this presentation, are necessary sources for proof details and conversion eligibility.

Deployment and ownership of the inference stack

  • Claims: The speaker argues that custom model transformations and kernel manipulation are harder to evaluate on rented, fixed inference endpoints because the operator cannot modify the serving path.; He presents a self-managed, open-source cluster and inference engine as a way to deploy modified Hugging Face checkpoints, compose them with other models for agentic tasks, share GPUs among smaller models, and manage configuration through an API.
  • Evidence: The named deployment products/repos in the transcript are Superlink's inference engine and Skye, though their spelling is somewhat uncertain from the transcript.; The stated use case is testing custom or fine-tuned Hugging Face checkpoints without building all deployment glue code.
  • Caveats: This portion is promotional and provides no comparative operational metrics, reliability evidence, or cost data for the named infrastructure.; Owning the cluster increases flexibility but also transfers security, observability, capacity-planning, and operational responsibility.
  • Implications: If kernel-level differentiation is strategically important, inference deployment architecture should preserve control over model artifacts, serving runtime, and GPU scheduling.; For ordinary hosted-model usage, the deployment pitch has limited immediate value because the custom CUDA path is unavailable.

Notable Concepts & Terms

  • RMSNorm: A transformer normalization layer targeted because repeated small normalization operations can impose significant decode-time orchestration overhead.
  • FlashNorm: The paper's umbrella term for algebraic RMSNorm transformations intended to reduce memory traffic, synchronization, and serial work.
  • Weight folding / weightless normalization: An offline rewrite that absorbs RMSNorm's learned gain into a neighboring projection matrix, allowing a modified checkpoint to eliminate separate scaling.
  • Deferred normalization: A reordering that postpones scalar division so RMS-related work can overlap with matrix multiplication before a final scale is applied.
  • CUDA streams: Asynchronous GPU execution queues; the talk's key engineering lesson is that they require explicit event-based synchronization when one stream consumes another's output.
  • Stale buffer read / race condition: The specific failure in which post-scaling read a previous matmul result, creating repeated, one-token-lagged generation despite passing superficial evaluation.
  • Tensor cores versus CUDA cores: The hardware parallelism exploited by the proposed kernel: tensor cores handle matrix multiplication while CUDA cores handle reductions, square roots, and elementwise operations.
  • Transformer Tricks: The cited repository containing implementation tooling for the algebraic transformations, including the simpler checkpoint-level conversion.

Operator Notes / Why Ken Should Care

  • Run a controlled benchmark of the repository's checkpoint-level weight-folding path on one self-hosted Llama-family workload; measure per-token decode latency, throughput, VRAM, output equivalence, and quantized-model behavior.
  • If considering custom deferred-normalization kernels, require an event/dependency diagram, explicit CUDA event waits, buffer-lifetime assertions, and long-context generation tests before accepting performance results.
  • Add adversarial long-generation regression cases to any custom inference-runtime CI; perplexity and short completions are insufficient to catch stale-buffer decoding errors.
  • Decide whether control of the inference runtime is strategically valuable enough to justify self-hosted kernel experimentation; do not assume such optimizations transfer to managed API endpoints.
  • Review the underlying paper and Transformer Tricks code for exact architecture constraints and reported benchmarks before treating FlashNorm as a platform-level optimization.

Source/Metadata

  • Title: Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards — Filip Makraduli
  • Transcript words: 4673
  • Duration seconds: 1038
  • Timestamp note: No timestamps or chapters were present. The transcript contains a near-complete duplicate of the talk, so its effective unique content is roughly half the supplied word count.

Transcript

2271 words en Processed in 90.0s

Hello everyone. Thank you for coming. I'll start the talk now. So this talk is around a paper that I did, which is very simple. The proposition is very clear. It's two lines of algebra that make the RMSNorm layer in transformers cheaper, quicker, and improve it as a layer in the transformer architecture. Similar to how layer norm once used to be the standard and then it was substituted by RMSNorm. This follows along this way of thinking. And I got the chance to meet some people from the open source world, and I co-authored this paper together with Nils Graf, who was the creator of this. And the work follows from there. So this is presented on archive. You can have a look, read it, test it out. There is a repo as well. And the concept, the idea and the way of thinking, I think it's easiest to explain with flash attention. So in a similar way that flash attention waits until there is a multiplication and tries to limit the communications between memory so that the whole process is faster. This is a similar thought along those lines. And it does certain improvements that make the RMSNorm process much quicker and in effect improve the whole transformer. And one question is: why RMSNorm? Since that layer does almost none of the math. And that's true. So the share of the math portion, if you look at it, is quite small. However, the clock time or wall time, as they say, is quite big. And for example, in one decode step, so when inferences perform, the RMSNorm can be started 33 times. Of course, it depends on the model and so on. In the paper, you have the specific models and how this was tested. And the question is how this can be improved and how the wait for the matrix multiplication can be avoided. And the reason why this is slow is because the GPUs are not slow or bad at math, but they're bad at everything else around the actual math. So that means starting the work, the actual work. So for example, starting the process as it happens in some of the experiments 33 times, that takes a long time. And for example, fusing each normalization into the matrix multiplication can help avoid this. Also doing weight folding can help in moving data between memory. And that's a process that's also slow for GPUs. And also weighting. So for example, deferring the division that's done in the RMSNorm layer is also a way to avoid this weighting step. So basically what this paper does is it improves all these three aspects by doing a few algebraic tricks in the way RMSNorm is computed. That's it. And math-wise, these are the tricks. It's mainly around the first two propositions. One is weightless normalization. You can see that here. And deferred normalization. So that's the second one. And now in more newer architectures, there is a situation where RMS can appear twice. For example, in GEMMA 4 this happens. So cancelling the pre-normalization also works. And all of this is algebraically proven in the paper. And the first proposition is this, where the gain and the weight fold into one matrix. W, you can see here with an asterisk. And that is computed offline, similar to how in flash attention you compute some stuff on the side so that there is no communication between memory all the time. So this is one step that's done, this weight folding. And the other step is deferring the scalar divide of the matmul so that they can be done in parallel. So in a normal case, you would have to compute once, then weight, and compute again. In this case, the idea is to split this so that it can be parallelized. And the third one, which is a version of this, is that there is, if there are two, because this is scale invariant, one of them can be dropped and this still works. And this is applicable to newer models that can have this architecture and implementation. So in order to make this happen in real life, especially this proposition number two, so for this one, for example, it's easy. There is a repo called transformer tricks. You can just apply this to any model and it works. But in order to do this, there is some kernel work. So it's not as straightforward to do. So in order for me to do that, I was implementing this and I came out with this experiment once. So it looks okay in general, where it's like, okay, the prompt is the transformer architecture, revolutionized NLP because, and then there is some expected output. But in the output I got, I saw this repetition and one step lag, as you can see here, the word because appears again. And there was something happening with the GPU streams and I was trying to figure out what was happening. And I was getting this one step lag and outputs that were from the past in a way. And in debugging all of this, I realized that in the process of building something like this, so as I explained the proposition of deferring these operations. In CUDA, you can do two things. You can do tensor cores that do one part of the matrix multiplication and you can do CUDA cores that run stuff like element-wise operations, reductions, square roots, and so on. So the idea was to do this in parallel and get the benefit of what I was explaining in the paper to actually test out this concept. So this is how it was supposed to look. So there is, if you do things sequentially, there is this idle waiting time when the vector unit computes the RMS and scaling. And then there is a matrix multiplication. So the idea was, okay, with flash norm, which is the technique in the paper, you're supposed to do those both in parallel. So the matrix unit computes the matmul and the vector unit computes the RMS. So in that way you save time. However, you cannot just do this in Python. You have to go a bit lower. And I did that with CUDA code like this. And this looked in general okay at that time. However, I realized that I did something slightly wrong. And that thing was that the join in the end where you're supposed to join the two streams was implicit in my case. And when I tested this out, the unit test worked. The quality seemed similar like perplexity testing and so on because it's just similar generation. But over long generation I was able to see this problem. So I had no idea what this was. And the reason was that when I was doing this implicit join, basically one of the streams hadn't finished the work. So I got race conditions that read the past from the unfinished matrix multiplication. So the idea that I had to fix this was around the fact that I had to be explicit about the join and wait until one of the operations is finished so that I'm certain that when I join I'm not reading from the past. So that was the realization in this exploration of CUDA streams. So the post scale read an old buffer value. And how this is fixed is with this where basically you need to mark the end of the matrix multiplication. Then mark the end of the RMS. And then post scale wait for the first stream. And then wait for the second stream. And that fixed the bug and made the paper work and the model speak forwards instead of backwards. And that was the cool academic perspective. But I also wanted to try things, right? Deploy this, test it out, see how I can make it work in a more production setting. And you can also read the paper and see all the tests. Some are done, most are done around Llama models. But this works for other architectures as well. So what you can do for this specific paper is, for example, the weight folding that I explained, the proposition one. You can just do it with some code in the repo. That's like flash, you say flashify and it does that. However, with this second thing that I mentioned, you need to do a bit of kernel work if you want to do that. As I explained in my example. And these are some results that are based on Llama models. And there are different details that you can have a look at as well. Like what happens if you do only deferred normalization? What happens if you do a full fused kernel? So there are a lot of experiments going lower here to test all the propositions. And these have been our results in different levels of scrutiny and detail. But even the simple one with weight folding shows some improvement. And this also works with the day-to-day tools that you use in a model. So it's not like you have to reinvent the wheel or do things from scratch. So it works with Torch Compile because it's a new checkpoint and that's it. Flash attention does similar tricks at a different layer. And also it works with quantized models. So it's totally cool to actually apply this and you can get a model that has this cool new normalization layer. And where you can get these details and codes to actually run this is this transformer tricks repo. So it has different algebra tricks as I explained as well as this paper that I mentioned. And also there is the GitHub, not the GitHub, but the HuggingFace model repo where I've done this with some models. And you can have a HuggingFace link to the model and test it out. And what you also can do with this HuggingFace models is to deploy them in production. So when I was thinking about doing this, I realized that, okay, now that the science is done and there is a link to a HuggingFace model, Superlink's inference engine was a cool way to actually deploy any HuggingFace model. And we've done this at hackathons where people would bring a custom HuggingFace model or checkpoint that they have with their fine-tuned stuff. And you can test out, even if you have some version of this algebraic tricks that you want to improve a model and test your own research ideas. You can actually try that out and have a deployed version on a cluster of this model and not have to worry about this glue code around deploying models. So that's pretty cool. And the point is that if you have the full cluster open source and the model inference open source, you can actually test out this kind of more novel research ideas where if you want to do kernel manipulation or flash norm and things like that, it's much more difficult to do this at the rented endpoint where you don't own the inference. You want something that's portable and flexible to actually allow you to do this stuff, but it's also production ready enough so that you can test things out at scale. And you can, for example, use Skye to combine this with other models. Like as you can see in the top left, you can have this flashified models with different other models to do agentic tasks if you want and do that end-to-end bigger use case. And the way Skye works is this production cluster helps you deploy the models. So you can have a look at Skye's repo as well for more details on this. And also there is a smarter queuing mechanism that helps you, especially if you work with smaller models. Because when doing the flash norm stuff, I worked with smaller Llama models and also with small agents from HuggingFace. So having a way to deploy smaller models that can also work on the same GPU so that you don't have to spend your money on GPU costs, but actually switch models around, especially smaller models. It was quite useful. And you can also control the model configs through an API as well as the cluster, which is also pretty convenient without having an infra person supporting you in your open source research. So that's cool as well. And you own your cloud, which is useful if you want open weights, open models, open source. And there is also a catalog that Skye has of different models. Not just the ones I mentioned, but you can have a look. There's also re-ranking embedding models if you're building something along those lines. And with that, I'm finishing this story of my research journey where I co-authored this paper around the technique that improves the transformer, but also found a way to bring this to production and test it out and find a way to play around with this open source models. And feel free to contact me on LinkedIn. Maybe if you have any questions or contributions. A lot of this stuff that I've mentioned, like some of them are PRs on VLLM or on Hugging Face. You might find them all around. You can also check out the paper. That's the archive link that you have there. And you also have the Skye repo and my LinkedIn. So thank you very much for attending. And you can catch me for questions. We'll be here close by. Hello everyone. Thank you for coming. And I'll start the talk now. So, this talk is around a paper that I did, which is very simple. The proposition is very clear. It's basically two lines of algebra that make the RMSNorm layer in transformers cheaper, quicker, and kind of improve it as like a layer in the transformer architecture. Similar to how layer norm once used to be the standard and then it was substituted by RMSNorm. This follows along this way of thinking. And I got the chance to kind of meet some people from the open source world, and I co-authored this paper together with Nils Graf, who was the kind of the creator of this. And the work follows from there. So, this is presented on archive. You can have a look, read it, test it out. There is a repo as well. And the concept, the, let's say the idea and the way of thinking, I think it's easiest to explain with maybe flash attention. So, in a similar way of how flash attention kind of waits until there is a multiplication and tries to limit this communications between memory so that the whole process is faster. This is kind of a similar thought along those lines. And it does certain improvements that make the RMSNorm process much quicker and in effect improve the whole transformer. And one question is, okay, why RMSNorm? Since that layer does almost none of the math. And that's true. So, the share of the kind of math portion, if you look at it, is quite small. However, the clock time or wall time, as they say, is quite big. And for example, in one decode step, so, right, when like inferences perform, the RMSNorm can be started like 33 times. Of course, it depends on the model and so on. In the paper, you have the specific models and how this was tested. And the question is how this can be improved and how this wait for the matrix multiplication can be kind of avoided. And the reason why this is slow is because the GPUs are not slow or bad at math, but they're bad at everything else around the actual math. So, that means starting the work, the actual work. So, for example, starting the process as it happens in some of the experiments 33 times, that takes a long time. And for example, fusing each normalization into the matrix multiplication can help avoid this. Also doing weight folding can help in kind of moving data between memory. And that's a process that's also slow for GPUs. And also weighting. So, for example, deferring the division that's done in the RMSNorm layer is also a way to avoid this weighting step. So, basically what this paper does is it improves all these three aspects by doing a few algebraic tricks in the way RMSNorm is computed. That's it. And math-wise, these are the tricks. It's mainly around the first two propositions. One is weightless normalization. You can see that here. And deferred normalization. So, that's the second one. And now in more newer architectures, there is a situation where RMS can appear twice. For example, in GEMA 4 this happens. So, cancelling the pre-normalization also works. And all of this is algebraically proven in the paper. And the first proposition is this, where kind of the gain and the weight fold into one matrix. W, you can see here with an asterisk. And that is computed offline, similar to how maybe in flash attention you compute some stuff on the side so that there is no communication between memory all the time. So, this is one step that's kind of done, this weight folding. And the other step is deferring the scalar divide of the math-mull so that they can be done in parallel. So, in a normal case, you would have to compute once, then weight, and compute again. In this case, the idea is to kind of split this so that it can be parallelized. And the third one, which is kind of a version of this, is that there is kind of, if there are two, because this is scale invariant, one of them can be dropped and this still works. And this is applicable to newer models that can have this architecture and implementation. So, in order to make this happen in real life, especially this proposition number two, so for this one, for example, it's easy. There is a repo called transformer tricks. You can just apply this to any model and it works. But in order to do this, there is some kernel work. So, it's not as straightforward to do. So, in order for me to do that, I was implementing this and I came out with this experiment once. So, it looks okay in general, where it's like, okay, the prompt is the transformer architecture, revolutionized NLP because, and then there is some kind of expected output. But in the output I got, I saw this repetition and one step lag, as you can see here, the word because appears again. And there was something happening with the GPU streams and I was trying to figure out what was happening. And I was getting this one step lag and kind of outputs that were from the past in a way. And in debugging all of this, I realized that in the process of building something like this, so as I explained the proposition to or deferring this to operations, In CUDA, you can do two things. You can do like tensor cores that do one part of the matrix multiplication and you can do CUDA cores that kind of run stuff like element-wise operations, reductions, square roots, and so on. So, the idea was to do this in parallel and get the benefit of what I was explaining in the paper to actually test out this concept. So, this is how it was supposed to look like. So, there is, if you do things sequentially, there is this idle waiting time when the vector unit computes the RMS and scaling. And then there is a matrix multiplication. So, the idea was, okay, with flash norm, which is the technique in the paper, you're supposed to do those both in parallel. So, the matrix unit computes the matmul and the vector unit computes the RMS. So, in that way you save time. However, you cannot just do this in Python. You have to go a bit lower. And I did that with CUDA codes like this. And this looked in general okay at that time. However, I realized that I did something slightly wrong. And that thing was that the join in the end where you're supposed to join the two streams was implicit in my case. And when I tested this out, the unit test worked. The quality seemed similar like perplexity testing and so on because it's just like similar generation. But over long generation I was able to see this problem. So, I had no idea what this was. And the reason was that when I was doing this implicit join, basically one of the streams hadn't finished the work. So, I got race conditions that kind of read the past from the unfinished matrix multiplication. So, the idea that I had to fix this was around the fact that I had to be explicit about the join and wait until one of the operations is finished so that I'm certain that when I join I'm not reading from the past. So, that was the realization in this exploration of CUDA streams. So, the post scale read like an old buffer value. And how this is fixed is with this where basically you need to mark the end of the matrix multiplication. Then mark the end of the RMS. And then post scale wait for the first stream. And then wait for the second stream. And that fixed the bug and made kind of the paper work and the model speak forwards instead of backwards. And that was the cool maybe academic perspective. But I also wanted to try things, right? Deploy this, test it out, see how I can make it work in maybe a more production setting. And you can also read the paper and see all the tests. Some are done, most are done around LAMA models. But like this works for other architectures as well. So, what you can do for this specific paper is, for example, the weight folding that I explained, the proposition one. You can just do it with some code in the repo. That's like flash, you say flashify and it does that. However, with this second thing that I mentioned, you need to do a bit of kernel work if you want to do that. Like I explained in my example. And these are some results that are based on LAMA models. And there are different kind of details that you can have a look at as well. Like what happens if you do only deferred normalization? What happens if you do a full fused kernel? So, there are a lot of experiments of going lower here to test all the prepositions. And these have been our results in different, let's say, levels of scrutiny and detail. But even the simple one with like weight folding shows some improvement. And this also works with like the day-to-day tools that you use in a model. So, it's not like you have to reinvent the wheel or, you know, do things from scratch. So, it works with Torch Compile because it's kind of like a new checkpoint and that's it. Flash attention does similar tricks at a different layer. And also, it works with quantized models. So, it's totally cool to actually apply this and you can get a model that has this cool new normalization layer. And where you can get this details and codes to actually run this is this transformer tricks repo. So, it has different algebra tricks like I explained as well as this paper that I mentioned. And also, there is the GitHub, not the GitHub, but the HuggingFace model repo where I've done this with some models. And you can have a HuggingFace link to the model and test it out. And what you also can do with this HuggingFace models is to deploy them in production. So, when I was thinking about doing this, I realized that, okay, now that let's say the science is done and there is a link to a HuggingFace model. Superlink's inference engine was a cool way to actually deploy any HuggingFace model. And we've done this at Hackathons where people would bring like a custom HuggingFace model or checkpoint that they have with their fine-tuned stuff. And you can test out like even if you have some version of this algebraic tricks that you want to improve a model and test your own research ideas. You can actually try that out and have a deployed version on a cluster of this model and not have to worry about this glue code around deploying models. So, that's pretty cool. And the point is that if you have the full cluster open source and the model inference open source, you can actually test out this kind of maybe more novel research ideas where if you want to do kernel manipulation or flash norm and things like that, it's much more difficult to do this at the rented end point where you don't own the inference. You want something that's portable and flexible to actually allow you to do this stuff, but it's also production ready enough so that you can test things out at scale. And you can, for example, use site to combine this with other models. Like as you can see in the top left, you can have this flashified models with different other models to do agentic tasks if you want and kind of do that end-to-end bigger use case. And the way site works is this production cluster helps you deploy the models. So, you can have a look at size repo as well for more details on this. And also, there is a smarter queuing mechanism that helps you, especially if you work with smaller models. Because when doing the flash norm stuff, I worked with like smaller Lama models and also with small agents from HikingFace. So, having a way to deploy smaller models that can also work on like the same GPU so that you don't have to spend your money on GPU costs, but actually kind of switch models around, especially smaller models. It was quite useful. And you can also control the model configs through an API as well as the cluster, which is also pretty convenient without having like an infra guy supporting you in your open source research. So, that's cool as well. And you own your cloud, which is useful if you want open weights, open models, open source. And there is also like a catalog that Sai has of different models. Not just the ones I mentioned, but you can have a look. There's also re-ranking embedding models if you're building something along those lines. And with that, I'm kind of finishing this story of my research journey where I co-authored this paper around the technique that improves the transformer, but also found a way kind of to bring this to, let's say, production and test it out and find a way to play around with this open source models. And feel free to contact me on LinkedIn. Maybe if you have any questions or contributions. A lot of this stuff that I've mentioned, like some of them are PRs on like VLLM or on Hugging Face. You might find them all around. You can also see the, check out the paper. That's the archive link that you have there. And you also have the Sai repo and my LinkedIn. So, thank you very much for attending. And you can catch me for questions. We'll be here. Close by. .