[SPEAKER_00] Okay.
SPEAKER_00
Hope everyone is awake after lunch. And nice to meet you all. I'm Ziv. I'm in the AI Labs team in NVIDIA based out of Paris and working with different Frontier Model Builders across a lot of domains. Of course, Diffusion is one of them. And we'll hear about a couple of examples of work we do with them. We have only 20 minutes. So obviously it's a mix between going deep and going very high level. Each of these topics will probably be a full day or full conference to cover. So I'll try to cover everything I can within this timeframe, but feel free to reach out afterwards, either through LinkedIn or I'll stay here a few minutes afterwards.
SPEAKER_00
So without further ado, Diffusion models, I assume everyone here knows about it. Anyone who doesn't know what VideoGen, ImageGen, how they work on a high level, denoising? Perfect. Okay. The idea is that, of course, unlike autoregressive architectures, LLM, the idea is that you have a lot of iterations to denoise the image or the video, usually between 20 to 50 steps. And we see, I think in the last year, an influx of very good high quality models, both for image generation, whether it's Flux 2, video generation, LTX 2, 1, Google with the Nano Banana, and the later generations. And we do see a lot of more practical use cases for that. And the main challenge,
SPEAKER_00
once we have some interesting use cases, is how to make it actually usable, right? We know that it's cool to generate videos or to generate images, but now if we talk about a developer context or an enterprise context, this should be fast. Okay, we want it to be mature, we want it to be scalable. And these are usually the challenges that are hard to solve as this ecosystem is not as mature as the autoregressive LLM, VLM ecosystem. Okay? So we try to borrow a lot of the concepts we see work very well for LLM, and we gradually distill them, if I'll use this terminology, into the world of diffusion models. Okay? We'll cover a few of the topics here,
SPEAKER_00
but again, every day we see more and more research in this domain. And I expect this world to be even more mature in the next AI engineer. Use case enablement, real-time image, real-time video is obviously the holy grail. Okay? Imagine how many new use cases where it's a world models for robotics, for computer games, for content generation. It opens a lot of new avenues for companies and developers to use it. And the big challenge together is, of course, the latency. Okay? It takes a lot of time to get a first image, and then to obviously get high quality. Okay? If we talk about 1080p or 720p content out there. And to bridge this gap, I'll talk about three concepts.
SPEAKER_00
Of course, it's not covering all the ways you can optimize your video gen, image gen models, but I'll touch on quantization, caching, and distillation. It's not necessarily the order you'll deploy it yourself. Okay? Usually you'll start with distillation, then do some quantization, then some caching. But I started from the simple to the more complex. Okay? Simple is usually quantization. Okay? For those of you who tried it in LLM, concepts are quite similar. Okay? Then we'll talk about caching and distillation. When we talk about quantization, we have two approaches. Okay? We have two approaches to post-training quantization and quantization-aware training.
SPEAKER_00
And I'd say in many cases, of course, we do want to use the more simple approach like PTQ, but we know that at least to maintain the image quality, the video quality, it's a little bit more complex with the diffusion models. Okay? And we also know that these types of models are more attention-heavy, which means that the impact of doing quantization is not as impactful as the LLMs, VLMs. But it is still quite a low-hanging fruit when we talk about taking advantage of the more advanced features of Blackwell, for example, and more modern compute. In this example, just the work we did with Black Forest Labs on Flux2, you can see that using usually dynamic quantization,
SPEAKER_00
okay, we don't want to use static, which means that we compute all the range of all the different parameters up front, deploy it and use this static range for the quantization. In this case, we use dynamic approach, okay, which means that some of the range will be computed on the fly. Okay? Again, to make sure that the distribution is in line with the different data distribution that you'll probably want to use when running these models. It's something that you can either do it yourself, okay? We recently released a good example in our tiered TLM visual gen repository, open source, you can start using it and see how it goes.
SPEAKER_00
What we also try to do to help the community to adopt it is also to help our partners to do pre-quantized checkpoints. So you can just go to Hugging Face, load the quantized checkpoint and start using it, okay? If you don't need to fine-tune or to do some lower adapters afterwards, it's something that is quite handy, and you can already see the impact. Of course, when we talk about quantization, the impact is both on the memory, okay? It will require less memory, which means you can run it on lower-end GPUs, whether it's consumer GPUs or lower-end data center GPUs, but also something that will help you in the performance, okay?
SPEAKER_00
is also to help our partners to do pre-quantized checkpoints. So you can just go to Hugging Face, load the quantized checkpoint and start using it, okay? If you don't need to fine-tune or to do some LoRA adapters afterwards, it's quite handy, and you can already see the impact. Of course, when we talk about quantization, the impact is both on the memory, okay? It will require less memory, which means you can run it on lower-end GPUs, whether it's consumer GPUs or lower-end data center GPUs, but also something that will help you in the performance, okay? So this is one part of the toolkit. A whole world sitting behind it to make sure that it's effective.
SPEAKER_00
Just today I've seen one of the latest research coming from MetaLab about attention, FP4, which as I mentioned, attention is quite heavy for this kind of model, so we do try to follow up with the latest research and make sure it's accessible for you as a community. Okay, when we are talking about the second stage, caching, okay? KV cache is something that anyone that worked a little bit with LLMs, with autoregressive models, it's something that everyone talks about how to use it efficiently, how to offload it, et cetera. It's a whole world. With the characteristics of diffusion models, it's not the same way, right?
SPEAKER_00
We don't generate a token every time, so it's harder to use this kind of techniques when we talk about denoising steps or making sure that we use the computation we had before in the way we'll generate future images or future videos. There are some, tcache is one example. It's not a very strong example, but it's a good example to understand the concept, okay? While we are doing denoising steps, right? We talked about 20 to 50 steps. There are areas between the denoising steps that are pretty much the same, okay? So we don't necessarily need to recompute them.
SPEAKER_00
What tcache is doing is if there was a minimal change, a very small change between the denoising steps, it compares it and you understand, okay, now I don't need to recompute for the next denoising step, okay? So it's more general, okay? It does it for the entire pixel space or latent space, okay? More modern techniques of caching will do it in a more chunk-based way, okay?
SPEAKER_00
Imagine that I don't know, now we are in the classroom here, most of you audience are sitting, staring at the screen, so nothing much changes, but I try to be a little bit more dynamic, so I'm moving, which means that this chunk of the video doesn't necessarily need to be recomputed, you guys don't need to recompute, I do need to, okay? So we'll isolate just this chunk and recalculate that, okay? Of course, you can define the threshold, and this is something that actually makes a lot of impact. We provided here some good examples of the expected boosts you can get from using this. But make sure that you try it, of course, and you maintain the quality, okay?
SPEAKER_00
Caching is something that if you don't do it the right way, can have quite a significant impact on the quality of the image, okay? And as content creators, world models, et cetera, it's something you want to make sure that you maintain while you get the boost, okay? So that's caching. I encourage you to read more about different techniques. This is something that is already available in the TensorRT LLM Visual Gen I mentioned. Just a flag you enable, and you set up the threshold, and you can experiment with it, but also it's available in VLLM-OMNI, SG-Line Diffusion, and other serving libraries.
SPEAKER_00
Distillation, okay? And this goes to the fact that you don't necessarily need 50 steps. Distillation is something we've seen, I'd say probably the big bang for distillation was during the DeepSeek first release, how they managed to distill from a very big model to much smaller models and get acceptable quality, but with a much smaller model. In Diffusion, the goal is not to get to a smaller model, okay? You'll still have the same number of parameters. This is more about step distillation, okay? Training the model, the student model, to generate as good quality images or videos, but by using much fewer steps.
SPEAKER_00
Okay? Instead of 50 steps, going to four steps, eight steps, in some cases one shot, okay? And maintaining the quality, okay? And this is the big challenge. Imagine if you are able to reduce this significant number of steps, but maintain the quality, it's something that can give you 10x, 200x improvement in performance. And if you go back to real-time generation, this is something today, it's probably the only way that it can get us there in good quality, okay? There's the next one, I think there's some demo we did in the last GTC conference a couple of weeks ago in San Jose, with two different distillation techniques, and we got to real-time generation, okay?
SPEAKER_00
And this is something that everyone is looking for, all the AI labs, and I'm sure also the bigger players, because this means that we can actually get to streaming something that will open a lot of new use cases. Okay, so how do we get it? Okay, we are, when we talk about distillation, we always have a teacher model and a student model. Currently, we have two main approaches when we talk about distillation, okay? One is trajectory-based, which means we'll try to teach the student how to follow the trajectory of the denoising steps as the teacher is doing, okay? And the second is distribution-based, which means we'll only look at the output distribution, okay?
SPEAKER_00
We want the student to get to the same point at the end, but we'll let the student understand how to get there, okay? And not by following the exact trajectory, okay? The more common, I would say, better quality technique these days is distribution-based, and we also see a lot of ways that can be combined. These techniques can be combined. In the last fast video release, they actually managed to do a hybrid approach that maintained the quality, but also got to more stable training. The challenge, and why I kept it to the last, is that distillation usually is a post-training technique, okay?
SPEAKER_00
Which means that if you do want it to work with your data, it's something you'll need to use some data for that technique, and you want it to converge in a good way, right? Because otherwise, it will be garbage in, garbage out, okay? and we also see a lot of ways that can be combined. These techniques can be combined. In the last fast video release, they actually managed to do a hybrid approach that maintained the quality, but also got to a more stable training. The challenge, and why I kept it to the last, is that distillation usually is a post-training technique, okay?
SPEAKER_00
Which means that if you do want it to work with your data, it's something you'll need to use some data for that technique, and you want it to converge in a good way, right? Because otherwise, it will just garbage in, garbage out, okay? So it will require more compute. It will require more time. Also, more proficiency. Again, as it's an exploratory, still, or research-driven domain. There's a lot of different techniques out there, and we expect more to come, but we are starting to see more mature techniques coming, and some very good examples shown in the latest open source models. Of course, closed source model builders are also using this approach.
SPEAKER_00
So FastGen is something that came out of our Envy research group, okay? It's an open source repository. You can go, there's a lot of different techniques there, okay? It's not a distillation technique or method, but the idea is that because it's so complex when we talk about large models, okay? A lot of the new video diffusion models are 20, 30, 40 B parameters, and we expect it to get to hundreds of billions of parameters. It requires post-training, it requires scale, sharding, all across different GPUs.
SPEAKER_00
So to manage all of this, we came with FastGen as a way to structure this process for you and enable you to focus only on the quality and fine-tuning the exact recipe that you want to use. So you can see there's an optional training data here. If you're not using, well, you can always use open source data and it will work up to a point, right? And we're actually happy about the results there. But if you want it to work for your use case with very specific data distribution, then we recommend you to use your own data for the fine-chain.
SPEAKER_00
Some of the results quoted here, the speedup, it's actually something we got, not just the speedup doesn't come only in time, it's also in using much smaller, much less compute to get to real-time. Okay, we got at GTC, as I mentioned, we got to one GPU of Blackwell B200 to generate near real-time video, or real-time, again, depends on the quality of the output. So it does something that we highly recommend you to look into if you want to get to this point, okay? We do expect a lot of the other autoregressive techniques to come and gradually be relevant for the video generation and image generation.
SPEAKER_00
We also see a lot of new model builders working a transfusion or autoregressive diffusion approach, okay? So you use the diffusion to generate a frame, but then it generates frame after frame in an autoregressive manner. So again, we expect a lot more of these techniques to get into this domain, but it's still a lot of research driven. So make sure this is one very good example you can take a look at. And I think the best value about it is all of it is incremental, okay? You can use this plus this plus this. You don't necessarily need to decide I'm doing only distillation or only quantization or only context parallelism.
SPEAKER_00
There's a lot of different techniques out there and they're all incremental, okay? So you can start with quantization, as I mentioned, which is the easier approach. If it's good enough for you, stay there. If not, let's move to now multi-GPU. Maybe do some context parallelism, maybe add some caching techniques, okay? And then last and the most impactful, that's the distillation. And hope to see a lot of you trying it and getting into the real-time performance. Now try it yourself, okay? All of it are open source resources that you can use.
SPEAKER_00
We have added support for the open source models as well, whether it's the one family, flux two family, LTX two family, and other ongoing. So hopefully we'll be able to see you guys contributing to this and making video diffusion as good as we see with LLM-VLM. And I think I'm almost at time. So if there is maybe one, two questions, happy to try and answer. If not, we can let you one minute of breathing. Thank you. On average, what would you say are there requirements for you to find in this model? Because access to GB200s are not that easy right now. And in terms of data set, how big are the data sets that you've seen work well with some of these models?
SPEAKER_00
Okay, so the question was about the compute needed for that and then the data set needed for that, just for everyone to hear. What's good about distillation is that you don't need GB200, right? You can do it with hoppers. You can do it with H200, H100, B200, B300. So it's not necessarily that you need very big compute as you do for pre-training, but you still need to compute. Okay? So it's not something you just take your one instance and start doing it. Of course, it depends on the size of the model, right? If your model is small, you have video generation models that are very small, 2B, 4B parameters. So this requires obviously much less compute.
SPEAKER_00
On the data front, I think it's very important to make sure that one, you know how to evaluate. Okay? So you can understand what's different if I just use a general purpose data set versus your specific data requires for your use case. And in such cases, we have seen differences. So for the more general demos, we don't use any special data set and it works well. [SPEAKER_01] But again, if it's something that protein generation or something around that, it will require something more specific. [SPEAKER_01] I'm at time, I think. [SPEAKER_01] But yeah, until they kick me out. Anyone other question? Okay, we can afterwards, I think. Thanks, everyone.
SPEAKER_00
[SPEAKER_01] I'm at time, I think. [SPEAKER_01] But yeah, until they'll kick me out. Anyone other question? Okay, we can afterwards, I think.
SPEAKER_00
Thanks, everyone. Because access to GB200s are not that easy right now. And in terms of data set, how big are the data sets that you've seen work well with some of these models? Okay, so the question was about the compute needed for that and then the data set needed for that, just for everyone to hear. What's good about distillation is that you don't need GB200, right? You can do it with hoppers. You can do it with H200, H100, B200, you know, B300. So it's not necessarily that you need very big compute as you do for pre-training, but you still need to compute. Okay? So it's not something you just, you know, take your, I don't know, just one instance and start doing it.
SPEAKER_00
Of course, it depends on the size of the model, right? If your model is small, you have video generation models that are very small, 2B, 4B parameters. So this requires obviously much less compute. On the data front, I think it's very important to make sure that, one, you know how to evaluate. Okay? So you can understand what's different if I just use just a general purpose data set versus your specific data requires for your use case. And in such cases, we have seen differences. So for the more general demos, we don't use any special data set and it works well.
SPEAKER_01
But again, if it's something that, I don't know, protein generation or something around that, that it will require, you know, something more specific. I'm at time, I think. But yeah, until they'll kick me out.
SPEAKER_00
Anyone other question? Okay, we can afterwards, I think. Thanks, everyone. I think so? I think so? I think so? I think so? I think so? I think so? I think so? I think so?
I think so?