You Might Not Need 50 Diffusion Steps — Ziv Ilan, Nvidia
Description
At GTC a few weeks ago, Ziv Ilan's team at NVIDIA got a video diffusion model generating in close to real time on a single Blackwell B200. The trick wasn't a new architecture, it was stripping out most of the fifty step denoising process diffusion models default to, by combining quantization, caching, and step distillation: training a student model to match a teacher's output using four steps, eight steps, or in some cases just one. Ilan walks through each layer of that stack: dynamic quantization work done with Black Forest Labs on Flux 2, a caching method that skips recomputing latent chunks that barely change between denoising steps, and distillation approaches split into trajectory based training, where the student copies the teacher's exact path, and distribution based training, where it only has to land on the same output, now the more common and higher quality of the two. NVIDIA's open source FastGen repo packages the post training and GPU sharding work needed to apply all this at scale, and Ilan frames the gains as additive, quantization alone can be enough on its own, or you can stack it with caching and distillation to reach the ten to two hundred times speedup that real time generation needs. Speaker info: - https://www.linkedin.com/in/ziv-ilan-deci/
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Diffusion models (image/video) can be accelerated from 20-50 steps to real-time or near-real-time generation by incrementally applying quantization, caching, and distillation techniques, each validated in NVIDIA's work with Flux2, LTX2, and other frontier models.
- Why it matters: Real-time video/image generation unlocks entirely new use cases—world models for robotics, gaming, streaming content—but requires mature optimization stacks similar to the LLM ecosystem; Ziv shows that the diffusion stack is reaching that maturity now.
- Best use: Reference for anyone deploying diffusion models at scale; Ziv's FastGen repository, TensorRT LLM Visual Gen, and checkpoint quantization workflow are concrete paths Ken can adopt or recommend for agent/content systems.
Executive Summary
Ziv Ilan, NVIDIA AI Labs (Paris), presents three incremental optimization techniques—quantization, caching, and distillation—that together reduce diffusion model inference from the typical 20-50 denoising steps to real-time or single-GPU performance without sacrificing image or video quality. His team works with frontier model builders (Black Forest Labs, LTX, Google Nano) to bridge the gap between research novelty and production readiness. The talk is framed around making diffusion as usable as autoregressive LLMs: fast, scalable, mature enough for enterprise and developer contexts.
Quantization (post-training and dynamic) is the first low-hanging fruit: NVIDIA's Flux2 collaboration and pre-quantized Hugging Face checkpoints show memory and throughput wins on consumer/data-center GPUs. Caching (e.g., t-cache and chunk-based approaches) skips redundant computation between denoising steps; it works best when parts of the latent space remain stable (like a static classroom audience). Distillation—trajectory-based or distribution-based—trains a student model to achieve teacher quality in far fewer steps (4–8 instead of 50), yielding 10-200× speedups and enabling real-time generation on a single Blackwell B200 GPU.
Ziv emphasizes that distillation is post-training, requires compute (H100/H200/B200 suffice) and domain-specific data if your use case diverges from general-purpose distributions. He released FastGen, an open-source orchestration toolkit for large-scale distillation (20-40B+ parameter models), sharding, and evaluation. All three techniques stack: quantize first, add caching, then distill. NVIDIA provides pre-quantized checkpoints, VLLM-OMNI/SG-Line integration, and support for Flux2, LTX2, and other open models.
The Q&A clarifies that distillation does not require GB200s—Hopper (H100/H200) and smaller Blackwell GPUs work—but model size and data quality determine final compute needs. General demos run on open datasets; specialized domains (protein generation) benefit from custom data. Ziv's team demonstrated real-time video at GTC 2024 (March) using these techniques, positioning diffusion inference as production-ready for streaming and interactive use cases.
Key Takeaways
- Claim: Diffusion models (20-50 denoising steps) can achieve real-time generation by stacking quantization, caching, and distillation optimizations incrementally. | Evidence: NVIDIA demonstrated near-real-time video generation on a single Blackwell B200 GPU at GTC 2024 using FastGen distillation, dynamic quantization on Flux2, and chunk-based caching techniques. | Caveat: Real-time depends on output quality/resolution (720p vs 1080p) and whether the use case requires domain-specific data; general open datasets may not suffice for specialized applications like protein generation. | Implication: Ken can recommend or deploy incremental optimization stacks for video/image agents; the diffusion ecosystem now parallels LLM maturity for production. | Timestamp: timestamp unavailable
- Claim: Dynamic quantization (compute ranges on the fly) preserves quality better than static quantization and enables lower-end GPUs (consumer/data-center) to run diffusion models. | Evidence: NVIDIA's Flux2 collaboration with Black Forest Labs uses dynamic quantization; pre-quantized checkpoints are available on Hugging Face; TensorRT LLM Visual Gen repository is open-source with examples. | Caveat: Quantization is less impactful for attention-heavy diffusion models than for LLMs/VLMs, but remains a low-hanging fruit; static quantization can degrade quality when data distribution shifts. | Implication: For Ken's agent workflows, load pre-quantized checkpoints if fine-tuning/LoRA isn't needed; otherwise, apply dynamic PTQ via TensorRT LLM Visual Gen before distillation. | Timestamp: timestamp unavailable
- Claim: Caching (e.g., t-cache, chunk-based) skips redundant computation between denoising steps when latent-space changes fall below a threshold, delivering throughput gains. | Evidence: T-cache compares denoising steps and reuses computation if minimal change is detected; chunk-based methods isolate stable regions (static classroom audience) and recompute only dynamic regions (moving speaker). | Caveat: Aggressive thresholds can degrade image/video quality; users must experiment and validate; caching is already available in TensorRT LLM Visual Gen, VLLM-OMNI, and SG-Line Diffusion as a flag with tunable threshold. | Implication: Ken should enable caching after quantization, tune the threshold per use case (e.g., robotics world models need different stability than content gen), and monitor quality metrics. | Timestamp: timestamp unavailable
- Claim: Distillation (trajectory-based or distribution-based) trains a student model to match teacher quality in 4-8 steps instead of 50, yielding 10-200× speedups; distribution-based is currently higher-quality. | Evidence: NVIDIA's FastGen distills 20-40B+ parameter models (LTX2, Flux2 family); fast-video release uses a hybrid trajectory+distribution approach for stable training; DeepSeek popularized distillation for LLMs; diffusion applies it to step reduction, not parameter count. | Caveat: Distillation is a post-training technique requiring compute (H100/H200/B200 suffice, not GB200), domain-specific data for non-general use cases, and careful evaluation to avoid garbage-in-garbage-out; it is the most impactful but most complex optimization. | Implication: Ken should treat distillation as the final layer after quantization and caching; FastGen orchestrates sharding, parallelism, and evaluation for large models; expect ongoing research (autoregressive diffusion, transfusion) to refine this further. | Timestamp: timestamp unavailable
- Claim: Autoregressive techniques (KV cache, context parallelism) and transfusion (autoregressive+diffusion hybrids) are gradually entering diffusion, but the field is still research-driven. | Evidence: Ziv mentions transfusion models (diffusion per frame, autoregressive across frames), MetaLab's attention FP4 research, and NVIDIA's expectation that diffusion will adopt LLM-style parallelism and serving optimizations. | Caveat: Techniques are exploratory and not yet as mature as LLM stacks; each new model architecture may require custom tuning; the field is evolving daily. | Implication: Ken should monitor NVIDIA FastGen, TensorRT LLM Visual Gen, and research releases; the diffusion serving stack will converge toward LLM-style tooling (VLLM, SG-Line, etc.) over the next 6-12 months. | Timestamp: timestamp unavailable
Detailed Brief
Why Diffusion Optimization Matters Now
- Claims: Diffusion models (image/video) require 20-50 denoising steps, unlike autoregressive LLMs that generate tokens sequentially.; High-quality models (Flux2, LTX2, Google Nano) are mature, but latency/scalability remain blockers for enterprise/developer adoption.; Real-time generation unlocks new use cases: world models for robotics, computer games, streaming content creation.
- Evidence: NVIDIA AI Labs works with Flux2 (Black Forest Labs), LTX2, and Google's models.; GTC 2024 (March San Jose) demo: real-time video on one Blackwell B200 GPU.; Ziv states, 'This is probably the only way that can get us there in good quality.'
- Caveats: Diffusion ecosystem is less mature than LLM/VLM stacks; research is ongoing.; Quality (720p vs 1080p) and use-case specificity (general vs protein generation) affect feasibility.
- Implications: Ken should treat diffusion optimization as production-ready for streaming/interactive agents.; The field is converging toward LLM-style tooling (VLLM, SG-Line, TensorRT) for diffusion models.
Quantization: Low-Hanging Fruit for Memory and Throughput
- Claims: Dynamic quantization (compute ranges on the fly) outperforms static quantization (fixed ranges) for diffusion quality.; Quantization is less impactful for attention-heavy diffusion models than for LLMs but still valuable for lower-end GPUs and Blackwell FP4 ops.
- Evidence: Flux2 collaboration uses dynamic PTQ; pre-quantized checkpoints on Hugging Face.; TensorRT LLM Visual Gen repository provides open-source examples.; MetaLab's attention FP4 research (released today per Ziv) targets attention layers specifically.
- Caveats: Static quantization degrades quality when data distribution shifts.; Quantization alone may not suffice for real-time; it's the first incremental step.
- Implications: Ken: load pre-quantized checkpoints if no fine-tuning is needed; otherwise apply dynamic PTQ before caching/distillation.; Blackwell GPUs enable advanced FP4 ops; quantization is a prerequisite to exploit hardware.
Caching: Skip Redundant Computation Across Denoising Steps
- Claims: T-cache compares denoising steps and reuses computation when change is minimal.; Chunk-based caching isolates stable regions (e.g., static audience) and recomputes only dynamic regions (e.g., moving speaker).; Caching is already integrated in TensorRT LLM Visual Gen, VLLM-OMNI, and SG-Line Diffusion as a tunable threshold flag.
- Evidence: Ziv's classroom analogy: audience is static, speaker moves; only recompute the speaker's chunk.; Caching can degrade quality if threshold is too aggressive; users must validate.
- Caveats: Incorrect thresholds harm quality; requires per-use-case tuning.; Caching is effective for stable content (static scenes) but less so for high-motion video.
- Implications: Ken should enable caching after quantization, validate quality per use case, and tune threshold.; For world models (robotics), stable backgrounds and moving agents benefit from chunk-based caching.
Distillation: Step Reduction for Real-Time Generation
- Claims: Distillation trains a student model to match teacher quality in 4-8 steps instead of 50, yielding 10-200× speedups.; Trajectory-based distillation teaches the student to follow the teacher's denoising path; distribution-based teaches only the final output distribution.; Distribution-based and hybrid (trajectory+distribution) approaches currently deliver better quality and training stability.
- Evidence: NVIDIA FastGen orchestrates large-scale distillation (20-40B+ models) with sharding and parallelism.; Fast-video release (hybrid approach) maintained quality and improved training stability.; DeepSeek pioneered LLM distillation for parameter reduction; diffusion applies it to step reduction instead.
- Caveats: Distillation is post-training, requiring compute (H100/H200/B200, not GB200), domain-specific data, and evaluation proficiency.; General-purpose datasets work for demos; specialized domains (protein gen) need custom data.; Garbage-in-garbage-out: poor training data or convergence yields poor results.
- Implications: Ken should treat distillation as the final, most impactful optimization after quantization and caching.; FastGen provides the orchestration layer for large models; evaluate with domain metrics before deployment.; Real-time video (GTC demo) is achievable on one B200 GPU with proper distillation.
Future Convergence: Autoregressive Techniques and Transfusion
- Claims: Autoregressive LLM techniques (KV cache, context parallelism) are gradually entering diffusion.; Transfusion models (diffusion per frame, autoregressive across frames) combine both paradigms.; Research is ongoing; new models and techniques emerge daily.
- Evidence: Ziv mentions MetaLab's attention FP4, transfusion models, and NVIDIA's expectation of LLM-style tooling for diffusion.; FastGen already supports multi-GPU context parallelism and sharding for large models.
- Caveats: Techniques are still exploratory and not as mature as LLM stacks.; Each new architecture may require custom tuning; field is research-driven.
- Implications: Ken should monitor FastGen, TensorRT LLM Visual Gen, and autoregressive-diffusion research for agent/content systems.; Expect diffusion serving to converge with LLM tooling (VLLM, SG-Line) in 6-12 months.
Notable Concepts & Terms
- FastGen: NVIDIA's open-source orchestration toolkit for large-scale diffusion distillation (20-40B+ models), handling sharding, parallelism, and evaluation; supports Flux2, LTX2, and other models.
- Dynamic quantization (PTQ): Post-training quantization that computes parameter ranges on the fly (vs. static, which uses fixed ranges); preserves quality for diffusion models when data distribution shifts.
- T-cache / Chunk-based caching: Techniques to skip redundant computation between denoising steps by comparing latent-space changes (t-cache: global; chunk-based: regional) and reusing previous results.
- Trajectory-based vs. distribution-based distillation: Trajectory: teach student to follow teacher's denoising path step-by-step; distribution: teach only the final output distribution, letting student find its own path. Distribution-based is currently higher-quality.
- Transfusion models: Hybrid architectures using diffusion to generate individual frames and autoregressive generation across frames, combining both paradigms for video generation.
- TensorRT LLM Visual Gen: NVIDIA's open-source repository for diffusion model quantization, caching, and serving; integrates with VLLM-OMNI and SG-Line Diffusion; provides pre-quantized checkpoints on Hugging Face.
Operator Notes / Why Ken Should Care
- For agent systems: Real-time video generation (world models, robotics) is now feasible with FastGen distillation and Blackwell GPUs; treat quantization → caching → distillation as the standard optimization stack.
- For content/business: Pre-quantized checkpoints (Flux2, LTX2 on Hugging Face) enable immediate deployment on lower-end GPUs; streaming content and interactive generation are production-ready.
- For AI ops: NVIDIA's TensorRT LLM Visual Gen, FastGen, and VLLM-OMNI integration provide mature serving infrastructure similar to LLM stacks; context parallelism and KV-cache-like techniques are arriving.
- For investing/GTM: Diffusion model serving is reaching LLM-level maturity; real-time generation unlocks new markets (gaming, robotics, live content); monitor transfusion research and autoregressive-diffusion convergence.
- For workflow/tooling: Incremental optimization is key—quantize first (low effort), then cache (medium), then distill (high effort, highest payoff); validate quality per step with domain metrics.
Watch Map
- timestamp unavailable: Intro: Diffusion models overview, 20-50 denoising steps, maturity challenge vs LLMs
- timestamp unavailable: Use case enablement: Real-time generation, latency/quality tradeoffs, new applications
- timestamp unavailable: Quantization: Dynamic PTQ, Flux2 collaboration, pre-quantized checkpoints, TensorRT LLM Visual Gen, MetaLab attention FP4
- timestamp unavailable: Caching: T-cache, chunk-based methods, classroom analogy (static audience, moving speaker), threshold tuning, VLLM-OMNI/SG-Line integration
- timestamp unavailable: Distillation: Trajectory vs distribution, FastGen orchestration, 10-200× speedups, GTC 2024 real-time demo on B200, post-training data/compute needs
- timestamp unavailable: Future: Autoregressive techniques, transfusion models, ongoing research, incremental optimization stack summary
- timestamp unavailable: Q&A: Compute requirements (H100/H200/B200, not GB200), dataset needs (general vs specialized), evaluation importance
Source/Metadata
- Title: You Might Not Need 50 Diffusion Steps — Ziv Ilan, Nvidia
- Transcript words: 3483
- Duration seconds: 1126
- Timestamp note: Timestamps/chapters not present in transcript; watch_map reflects logical flow and topic order
Transcript
[SPEAKER_00] Okay. Hope everyone is awake after lunch. And nice to meet you all. I'm Ziv. I'm in the AI Labs team in NVIDIA based out of Paris and working with different Frontier Model Builders across a lot of domains. Of course, Diffusion is one of them. And we'll hear about a couple of examples of work we do with them. We have only 20 minutes. So obviously it's a mix between going deep and going very high level. Each of these topics will probably be a full day or full conference to cover. So I'll try to cover everything I can within this timeframe, but feel free to reach out afterwards, either through LinkedIn or I'll stay here a few minutes afterwards. So without further ado, Diffusion models, I assume everyone here knows about it. Anyone who doesn't know what VideoGen, ImageGen, how they work on a high level, denoising? Perfect. Okay. The idea is that, of course, unlike autoregressive architectures, LLM, the idea is that you have a lot of iterations to denoise the image or the video, usually between 20 to 50 steps. And we see, I think in the last year, an influx of very good high quality models, both for image generation, whether it's Flux 2, video generation, LTX 2, 1, Google with the Nano Banana, and the later generations. And we do see a lot of more practical use cases for that. And the main challenge, once we have some interesting use cases, is how to make it actually usable, right? We know that it's cool to generate videos or to generate images, but now if we talk about a developer context or an enterprise context, this should be fast. Okay, we want it to be mature, we want it to be scalable. And these are usually the challenges that are hard to solve as this ecosystem is not as mature as the autoregressive LLM, VLM ecosystem. Okay? So we try to borrow a lot of the concepts we see work very well for LLM, and we gradually distill them, if I'll use this terminology, into the world of diffusion models. Okay? We'll cover a few of the topics here, but again, every day we see more and more research in this domain. And I expect this world to be even more mature in the next AI engineer. Use case enablement, real-time image, real-time video is obviously the holy grail. Okay? Imagine how many new use cases where it's a world models for robotics, for computer games, for content generation. It opens a lot of new avenues for companies and developers to use it. And the big challenge together is, of course, the latency. Okay? It takes a lot of time to get a first image, and then to obviously get high quality. Okay? If we talk about 1080p or 720p content out there. And to bridge this gap, I'll talk about three concepts. Of course, it's not covering all the ways you can optimize your video gen, image gen models, but I'll touch on quantization, caching, and distillation. It's not necessarily the order you'll deploy it yourself. Okay? Usually you'll start with distillation, then do some quantization, then some caching. But I started from the simple to the more complex. Okay? Simple is usually quantization. Okay? For those of you who tried it in LLM, concepts are quite similar. Okay? Then we'll talk about caching and distillation. When we talk about quantization, we have two approaches. Okay? We have two approaches to post-training quantization and quantization-aware training. And I'd say in many cases, of course, we do want to use the more simple approach like PTQ, but we know that at least to maintain the image quality, the video quality, it's a little bit more complex with the diffusion models. Okay? And we also know that these types of models are more attention-heavy, which means that the impact of doing quantization is not as impactful as the LLMs, VLMs. But it is still quite a low-hanging fruit when we talk about taking advantage of the more advanced features of Blackwell, for example, and more modern compute. In this example, just the work we did with Black Forest Labs on Flux2, you can see that using usually dynamic quantization, okay, we don't want to use static, which means that we compute all the range of all the different parameters up front, deploy it and use this static range for the quantization. In this case, we use dynamic approach, okay, which means that some of the range will be computed on the fly. Okay? Again, to make sure that the distribution is in line with the different data distribution that you'll probably want to use when running these models. It's something that you can either do it yourself, okay? We recently released a good example in our tiered TLM visual gen repository, open source, you can start using it and see how it goes. What we also try to do to help the community to adopt it is also to help our partners to do pre-quantized checkpoints. So you can just go to Hugging Face, load the quantized checkpoint and start using it, okay? If you don't need to fine-tune or to do some lower adapters afterwards, it's something that is quite handy, and you can already see the impact. Of course, when we talk about quantization, the impact is both on the memory, okay? It will require less memory, which means you can run it on lower-end GPUs, whether it's consumer GPUs or lower-end data center GPUs, but also something that will help you in the performance, okay? is also to help our partners to do pre-quantized checkpoints. So you can just go to Hugging Face, load the quantized checkpoint and start using it, okay? If you don't need to fine-tune or to do some LoRA adapters afterwards, it's quite handy, and you can already see the impact. Of course, when we talk about quantization, the impact is both on the memory, okay? It will require less memory, which means you can run it on lower-end GPUs, whether it's consumer GPUs or lower-end data center GPUs, but also something that will help you in the performance, okay? So this is one part of the toolkit. A whole world sitting behind it to make sure that it's effective. Just today I've seen one of the latest research coming from MetaLab about attention, FP4, which as I mentioned, attention is quite heavy for this kind of model, so we do try to follow up with the latest research and make sure it's accessible for you as a community. Okay, when we are talking about the second stage, caching, okay? KV cache is something that anyone that worked a little bit with LLMs, with autoregressive models, it's something that everyone talks about how to use it efficiently, how to offload it, et cetera. It's a whole world. With the characteristics of diffusion models, it's not the same way, right? We don't generate a token every time, so it's harder to use this kind of techniques when we talk about denoising steps or making sure that we use the computation we had before in the way we'll generate future images or future videos. There are some, tcache is one example. It's not a very strong example, but it's a good example to understand the concept, okay? While we are doing denoising steps, right? We talked about 20 to 50 steps. There are areas between the denoising steps that are pretty much the same, okay? So we don't necessarily need to recompute them. What tcache is doing is if there was a minimal change, a very small change between the denoising steps, it compares it and you understand, okay, now I don't need to recompute for the next denoising step, okay? So it's more general, okay? It does it for the entire pixel space or latent space, okay? More modern techniques of caching will do it in a more chunk-based way, okay? Imagine that I don't know, now we are in the classroom here, most of you audience are sitting, staring at the screen, so nothing much changes, but I try to be a little bit more dynamic, so I'm moving, which means that this chunk of the video doesn't necessarily need to be recomputed, you guys don't need to recompute, I do need to, okay? So we'll isolate just this chunk and recalculate that, okay? Of course, you can define the threshold, and this is something that actually makes a lot of impact. We provided here some good examples of the expected boosts you can get from using this. But make sure that you try it, of course, and you maintain the quality, okay? Caching is something that if you don't do it the right way, can have quite a significant impact on the quality of the image, okay? And as content creators, world models, et cetera, it's something you want to make sure that you maintain while you get the boost, okay? So that's caching. I encourage you to read more about different techniques. This is something that is already available in the TensorRT LLM Visual Gen I mentioned. Just a flag you enable, and you set up the threshold, and you can experiment with it, but also it's available in VLLM-OMNI, SG-Line Diffusion, and other serving libraries. Distillation, okay? And this goes to the fact that you don't necessarily need 50 steps. Distillation is something we've seen, I'd say probably the big bang for distillation was during the DeepSeek first release, how they managed to distill from a very big model to much smaller models and get acceptable quality, but with a much smaller model. In Diffusion, the goal is not to get to a smaller model, okay? You'll still have the same number of parameters. This is more about step distillation, okay? Training the model, the student model, to generate as good quality images or videos, but by using much fewer steps. Okay? Instead of 50 steps, going to four steps, eight steps, in some cases one shot, okay? And maintaining the quality, okay? And this is the big challenge. Imagine if you are able to reduce this significant number of steps, but maintain the quality, it's something that can give you 10x, 200x improvement in performance. And if you go back to real-time generation, this is something today, it's probably the only way that it can get us there in good quality, okay? There's the next one, I think there's some demo we did in the last GTC conference a couple of weeks ago in San Jose, with two different distillation techniques, and we got to real-time generation, okay? And this is something that everyone is looking for, all the AI labs, and I'm sure also the bigger players, because this means that we can actually get to streaming something that will open a lot of new use cases. Okay, so how do we get it? Okay, we are, when we talk about distillation, we always have a teacher model and a student model. Currently, we have two main approaches when we talk about distillation, okay? One is trajectory-based, which means we'll try to teach the student how to follow the trajectory of the denoising steps as the teacher is doing, okay? And the second is distribution-based, which means we'll only look at the output distribution, okay? We want the student to get to the same point at the end, but we'll let the student understand how to get there, okay? And not by following the exact trajectory, okay? The more common, I would say, better quality technique these days is distribution-based, and we also see a lot of ways that can be combined. These techniques can be combined. In the last fast video release, they actually managed to do a hybrid approach that maintained the quality, but also got to more stable training. The challenge, and why I kept it to the last, is that distillation usually is a post-training technique, okay? Which means that if you do want it to work with your data, it's something you'll need to use some data for that technique, and you want it to converge in a good way, right? Because otherwise, it will be garbage in, garbage out, okay? and we also see a lot of ways that can be combined. These techniques can be combined. In the last fast video release, they actually managed to do a hybrid approach that maintained the quality, but also got to a more stable training. The challenge, and why I kept it to the last, is that distillation usually is a post-training technique, okay? Which means that if you do want it to work with your data, it's something you'll need to use some data for that technique, and you want it to converge in a good way, right? Because otherwise, it will just garbage in, garbage out, okay? So it will require more compute. It will require more time. Also, more proficiency. Again, as it's an exploratory, still, or research-driven domain. There's a lot of different techniques out there, and we expect more to come, but we are starting to see more mature techniques coming, and some very good examples shown in the latest open source models. Of course, closed source model builders are also using this approach. So FastGen is something that came out of our Envy research group, okay? It's an open source repository. You can go, there's a lot of different techniques there, okay? It's not a distillation technique or method, but the idea is that because it's so complex when we talk about large models, okay? A lot of the new video diffusion models are 20, 30, 40 B parameters, and we expect it to get to hundreds of billions of parameters. It requires post-training, it requires scale, sharding, all across different GPUs. So to manage all of this, we came with FastGen as a way to structure this process for you and enable you to focus only on the quality and fine-tuning the exact recipe that you want to use. So you can see there's an optional training data here. If you're not using, well, you can always use open source data and it will work up to a point, right? And we're actually happy about the results there. But if you want it to work for your use case with very specific data distribution, then we recommend you to use your own data for the fine-chain. Some of the results quoted here, the speedup, it's actually something we got, not just the speedup doesn't come only in time, it's also in using much smaller, much less compute to get to real-time. Okay, we got at GTC, as I mentioned, we got to one GPU of Blackwell B200 to generate near real-time video, or real-time, again, depends on the quality of the output. So it does something that we highly recommend you to look into if you want to get to this point, okay? We do expect a lot of the other autoregressive techniques to come and gradually be relevant for the video generation and image generation. We also see a lot of new model builders working a transfusion or autoregressive diffusion approach, okay? So you use the diffusion to generate a frame, but then it generates frame after frame in an autoregressive manner. So again, we expect a lot more of these techniques to get into this domain, but it's still a lot of research driven. So make sure this is one very good example you can take a look at. And I think the best value about it is all of it is incremental, okay? You can use this plus this plus this. You don't necessarily need to decide I'm doing only distillation or only quantization or only context parallelism. There's a lot of different techniques out there and they're all incremental, okay? So you can start with quantization, as I mentioned, which is the easier approach. If it's good enough for you, stay there. If not, let's move to now multi-GPU. Maybe do some context parallelism, maybe add some caching techniques, okay? And then last and the most impactful, that's the distillation. And hope to see a lot of you trying it and getting into the real-time performance. Now try it yourself, okay? All of it are open source resources that you can use. We have added support for the open source models as well, whether it's the one family, flux two family, LTX two family, and other ongoing. So hopefully we'll be able to see you guys contributing to this and making video diffusion as good as we see with LLM-VLM. And I think I'm almost at time. So if there is maybe one, two questions, happy to try and answer. If not, we can let you one minute of breathing. Thank you. On average, what would you say are there requirements for you to find in this model? Because access to GB200s are not that easy right now. And in terms of data set, how big are the data sets that you've seen work well with some of these models? Okay, so the question was about the compute needed for that and then the data set needed for that, just for everyone to hear. What's good about distillation is that you don't need GB200, right? You can do it with hoppers. You can do it with H200, H100, B200, B300. So it's not necessarily that you need very big compute as you do for pre-training, but you still need to compute. Okay? So it's not something you just take your one instance and start doing it. Of course, it depends on the size of the model, right? If your model is small, you have video generation models that are very small, 2B, 4B parameters. So this requires obviously much less compute. On the data front, I think it's very important to make sure that one, you know how to evaluate. Okay? So you can understand what's different if I just use a general purpose data set versus your specific data requires for your use case. And in such cases, we have seen differences. So for the more general demos, we don't use any special data set and it works well. [SPEAKER_01] But again, if it's something that protein generation or something around that, it will require something more specific. [SPEAKER_01] I'm at time, I think. [SPEAKER_01] But yeah, until they kick me out. Anyone other question? Okay, we can afterwards, I think. Thanks, everyone. [SPEAKER_01] I'm at time, I think. [SPEAKER_01] But yeah, until they'll kick me out. Anyone other question? Okay, we can afterwards, I think. Thanks, everyone. Because access to GB200s are not that easy right now. And in terms of data set, how big are the data sets that you've seen work well with some of these models? Okay, so the question was about the compute needed for that and then the data set needed for that, just for everyone to hear. What's good about distillation is that you don't need GB200, right? You can do it with hoppers. You can do it with H200, H100, B200, you know, B300. So it's not necessarily that you need very big compute as you do for pre-training, but you still need to compute. Okay? So it's not something you just, you know, take your, I don't know, just one instance and start doing it. Of course, it depends on the size of the model, right? If your model is small, you have video generation models that are very small, 2B, 4B parameters. So this requires obviously much less compute. On the data front, I think it's very important to make sure that, one, you know how to evaluate. Okay? So you can understand what's different if I just use just a general purpose data set versus your specific data requires for your use case. And in such cases, we have seen differences. So for the more general demos, we don't use any special data set and it works well. But again, if it's something that, I don't know, protein generation or something around that, that it will require, you know, something more specific. I'm at time, I think. But yeah, until they'll kick me out. Anyone other question? Okay, we can afterwards, I think. Thanks, everyone. I think so? I think so? I think so? I think so? I think so? I think so? I think so? I think so? I think so?