[SPEAKER_02] Some people are still filtering into the room.
SPEAKER_02
It's mostly intro stuff for the first couple of sites, so they won't miss anything. Okay, welcome everybody. My name is Brendan. I'm a research scientist at DeepMind. I'm talking today about text diffusion, which is a more forward-looking research area at DeepMind.
SPEAKER_02
So you're probably familiar with image and video diffusion, which is state-of-the-art for these modalities right now, where you take ground truth, say image, you add noise to it in training, and then you train a neural network to remove that noise gradually, and then at inference time you just initialize the picture with pure noise, and then you iteratively refine out the noise to recover back to whatever image or video or audio or whatever you're looking for.
SPEAKER_02
And the principle is essentially the same for text, for text diffusion, where you start with a clean sequence of tokens, like a clean sentence or something like that, and then you gradually add noise. You corrupt it somehow. There's lots of different ways to do that. You can do it in a continuous or discrete way, but let's just say discrete for now, which would in this case just mean adding random tokens or replacing tokens with other random tokens. And you do that for a bunch of different noise levels, and you train the neural network to try to fill in, to try to correct the mistakes basically in the text.
SPEAKER_02
And then at inference time, you initialize the sequence of tokens to just pure noise, like pure random discrete tokens from the vocabulary, and then you iteratively refine through that to fill in the information in the order that the neural network wants to do to recover back to say a clean sentence. And then in practice, so I showed you some GIFs here of what it looks like for images. You get very similar looking outputs for text where it starts off all noisy and then gradually fills in the text and you get clean, relatively clean outputs at the end.
SPEAKER_02
Okay, so the team I'm on, we had a research demo release one year ago now called Gemini Diffusion, which was a variant of a Gemini model with text diffusion instead of autoregressive next token generation. And that was a research preview that was open to about 100k people. And we're still keeping posted for new developments in that direction upcoming soon. And we did have some good numbers at the time, but again, it's a year ago, which is prehistoric times in this field.
SPEAKER_02
Our main comparator model was Gemini 2.0 Flashlight at the time, because that was the architecture we were branching from the text diffusion model. And we basically had very similar quality across the board there. Mostly a little bit of advantage in code, a little bit of disadvantage in some other areas, but relatively similar performance at much better latencies. But again, this is a year ago, so I wouldn't fixate too much on these numbers. Okay, so what's the difference between autoregressive generation and diffusion?
SPEAKER_02
So in the standard vanilla Gemini, Gemma, GPT, whatever, generation of text, you do this: you have some context that comes in and you want to generate some response to that. And the model does it one token at a time. So it generates the first token and then conditions on that generates the next token and so on. Whereas in diffusion, you have a context that will come in, whatever that is, and it'll initialize, like I mentioned, a long sequence of tokens, could be hundreds, could be thousands, could be shorter, depends on the model, to be random noise. And then it iteratively refines that canvas to remove the noise over the course of a few denoising steps.
SPEAKER_02
So rather than one token at a time, it does the entire block together, but over a couple of iterations. So it's not just one pass, it does multiple passes, but it gets to attend to the future tokens and so on. So it's a different way of generating text. So that obviously has some pros and cons. So the main pro that people really like, and probably is the biggest advantage that text diffusion models have, is that it's faster inference, it just generates faster tokens per second, because it makes much better use of the hardware, the TPU and the GPU. And I have some slides on that to explain why.
SPEAKER_02
But some other advantages are it can do bidirectional attention within this canvas of tokens. So autoregressive models can only attend to the past. They have causal attention within their transformer, whereas a text diffusion model is not restricted. They can attend to the future. And that has some interesting properties, like it can do self-corrected generation based on future tokens. So they could do some reasoning, see that I got the answer incorrect, and then go back and fix the reasoning and do it again. I have a demo of that.
SPEAKER_02
Because this process is iterative, it does a number of steps to respond. That means the model can actually do adaptive computation. It turns out that you can train the model to spend more time on harder problems and less time on easier problems. And the diffusion models in general, you can do things in-place editing, where you say, fix the last tokens and give me the prefix that corresponds to those tokens and stuff. But the main disadvantage it has, and the reason why it's not used everywhere right now, is lower throughput for large batches.
SPEAKER_02
So autoregressive models, they're slow, but you can have a big batch of queries together. And then if you push that through the neural network on the GPU, each individual user is slow, but you make good use of the TPU by doing that. And so you can serve a lot of queries. And so you keep your costs down. You can serve a lot. Whereas since text diffusion does multiple forward passes on the same data, it hits a compute threshold basically earlier. And even though it's lower latency for any one user, it tends to be lower throughput overall. So higher cost to serve.
SPEAKER_02
And right now, if you've played with Claude recently, you'll know that they have some throughput concerns. So people really care about throughput right now. And so no one's landing text diffusion into any of these big models, primarily because of that disadvantage. It's just too expensive to serve, even if it is much lower latency. Okay, so just leaning into why it doesn't have lower latency, in case you're not familiar with the architecture of how GPUs and TPUs run today. So in a GPU, there's a tensor core, which does these big matrix multiplies. It's very efficient, has a lot of flops or hops or whatever. And then the memory that sits on the TPU GPU, this HBM,
SPEAKER_02
So people really care about throughput right now. And so no one's landing text diffusion into any of these big models, primarily because of that disadvantage. It's just too expensive to serve, even if it is much lower latency. Okay, so just leaning into why it doesn't have lower latency, in case you're not familiar with the architecture of how GPUs and TPUs run today. So in a GPU, there's a tensor core, which does these big matrix multiplies. It's very efficient, has a lot of flops or hops or whatever. And then the memory that sits on the TPU GPU, this HBM, that's where the weights and the activations and everything are stored.
SPEAKER_02
And that has to transfer over from the memory all the weights and the activations and the KV cache into the tensor core in order to do the computation. And so it has to flow through this bandwidth channel. And that bandwidth channel is very tight. It turns out that both GPUs and TPUs have a lot of flops and not that much bandwidth. It's quite hard, it's expensive to put bandwidth onto these chips and it's easy to put flops. So because of that ratio, if you do more flops for each streaming amount of data you put through, the better. So when you're serving an autoregressive model, these chips are memory bound. They're basically bottlenecked by this bandwidth.
SPEAKER_02
So when you do autoregressive next token generation, for each token, you're doing one token at a time, let's say batch size one, you have to stream over the entire neural network and all the KV cache and everything to get one token and then you do it again for the next token and so on. Whereas for text diffusion, you're generating, say, 256 tokens, you still stream over everything. But if you can do that less times than the number of tokens, this iterative refinement process, then you'll get a speed up. So if you can do, say, 24 passes to generate 256 tokens, you'll be doing 10 times fewer memory transfers than an autoregressive model.
SPEAKER_02
And if you are truly memory bound, then you'll be 10 times faster, something like that. So that's the real reason, that's the hardware reason why text diffusion models are much lower latency than autoregressive models. Okay. So we had this Gemini diffusion demo, maybe some of you got access to it, last year. And that was able to hit something like 2,000 tokens a second, that was really good at the end pretty consistently, depending on the length of the query. Obviously, it depends. The longer sequence it's generating, the less it's pre-fill dominated. [SPEAKER_05] And so you can really lean into these very long sequences of very fast tokens.
SPEAKER_02
But if you're only generating one token, for instance, you'll be just dominated by the cost of the pre-fill. And the tokens per second number that was reported on this webpage was incorporated pre-fill and everything like that. So this was 2,000 tokens a second, as genuine raw tokens that you would receive in your web browser. Okay. So that was the whirlwind tour of text diffusion and its main advantage, which is latency. But I want to dig in a little bit into some of the other advantages that text diffusion has, which are not as talked about in the literature. But I think are pretty cool. And this is why I'm excited about it.
SPEAKER_02
So at Google I/O last year, they showed this demo for the text diffusion model, which is a really easy prompt. But lots of models actually make mistakes on it. So the prompt is, if you go to the next slide, what is the square root of 81 times 2 thirds squared plus, blah, blah, blah? And I think the answer is 39 to this problem. And so you pass that into the model and you ask the Gemini diffusion to respond to that. And after one forward pass, these are the tokens that have been generated. So one forward pass through the model, it's starting to respond. Now, it's doing this iterative refinement process. So one forward pass is not all it's going to do.
SPEAKER_02
But after one forward pass, it has this. So it has answer equals, and then it says 60. It's not correct, but that's what it's guessing for now. And then it starts to do the reasoning. So a solution, calculate the square root of 81, and so on. After two forward passes, it's changed 60 to 49. And it's gotten a little bit further into the reasoning. So it's gotten five steps into the reasoning. Two squared equals four. And some of the blue tokens are still going to change. And then after three forward passes, it's actually gotten all the way through the reasoning. So it gets the answer correct at the end. 36 plus 3 is equal to 39.
SPEAKER_02
And it's gone back and fixed the original response to say 39. So it had a mistake twice, 60 and 49. But once it finished the reasoning, it was able to return back and fix the mistake that it made at the start. Now it's going to do a couple more forward passes in order to fix some of these tokens that aren't quite right in the text. But overall, that's basically the structure of the output that it'll return. And this is a property that text diffusion models have. This ability to do bidirectional reasoning. So to not only see the past, but also see the future that it's going to utter, it's going to respond. And also to use that information to do self-correction.
SPEAKER_02
So it made a mistake, but it was able to have another forward pass. It was able to go through and fix that mistake. At the time, much bigger models than the one we were serving made a mistake for this problem. So both ChatGPT 4O, which was new at the time, and Gemini 2.5 Flash, which was brand new at the time, both made an error on this exact problem. So you give them the exact same prompt, and then they would say the answer is 39. The GPT 4O said 40, because that's the best guess it can do at that one token. It went through the reasoning, and then it did manage to figure out it was 39. It was able to go through and fix that mistake.
SPEAKER_02
At the time, much bigger models than the one we were serving made a mistake for this problem. So both ChatGPT4O, which was new at the time, and Gemini 2.5 Flash, which was brand new at the time, both made an error on this exact problem. So you give them the exact same prompt, and then they would say, remember the answer is 39. The GPT4O said 40, because that's the best guess it can do at that one token. It went through the reasoning, and then it did manage to figure out it was 39. It said, sorry, I made a mistake. It's 39, not 40. Gemini 2.5 Flash also made a mistake. It said 42, and then it actually just stuck to its guns and never changed it.
SPEAKER_02
It said 36 plus 3 is 42. So it incorporated the error into its reasoning later. And these are way bigger models than the Gemini diffusion model. So it really is a property, a flaw of autoregressive models that the text diffusion models don't have. And you can fix this with modern reasoning thinking models, but then you're just punting the problem into something else. But anyway, okay, so that's one advantage, which is bidirectional reasoning, self-correction. Another one is what I hinted at before, which is dynamic computation. So you can give the model more time at inference, more forward passes. You give it a bigger budget, and it can just do better.
SPEAKER_02
It's not exactly monotonic, but it is roughly monotonic that the quality across every eval basically just continues to go up. Because even if the solution is almost tightly clean and correct, it gets to look at it and see that it made a mistake and then fix it. So you get this nice curve where you always see, as the number of denoising steps, which is the forward passes, increases, overall the quality gets higher. And these are just six coding evals that we monitor internally.
SPEAKER_05
[SPEAKER_02] On top of that, a slightly different concept is the model can do adaptive computation, which is that you can allow the model, train it in a way, to determine itself when it is finished.
SPEAKER_02
And then for easy responses, it can use a little bit of compute, and for harder responses, it can take longer. So here's just three examples from the Gemini diffusion model, which is what are the first 100 digits of pi? This is actually 100 tokens. It looks like a short response. It's actually 100 tokens. And it only takes four steps to do that. Because the model, it's an easy response, because you've just memorized the 100 digits of pi. You just output it. Whereas an autoregressive model at the same time would have only done four tokens. So that's a very easy one. Slightly more challenging is to write a little bit of code.
SPEAKER_02
So that takes 18 forward passes to generate FizzBuzz. And then something more complicated is explaining quantum mechanics in a single paragraph. And that took 31 denoising steps. It just took its time, decided to spend longer on those ones. And so the model naturally gets to decide to determine when it's going to finish and return the response. And typically we see that harder evals take more time. So this is a year ago now. So the evals are old school. But on the harder end is GPQA diamond, which for the model size we were targeting was quite a hard eval. And that took a long time for it to respond to those ones.
SPEAKER_02
Whereas on the other end, MBPP, which is mostly basic Python programs, it was very easy to respond to. It took very little time for it to respond to these things. And this is entirely determined by the model itself. Just easier problems, easier prompts it could respond to quickly and harder ones it decided itself to spend more time reasoning. OK, so that's another property, which is the dynamic and adaptive computation. Lastly is the fast in-place editing. So diffusion models in general have this very nice property where you can take an image, cut something out of it, and give it a prompt, and it'll fill it in.
SPEAKER_02
And you can use the context that you haven't cut out to fill in the piece you've cut out correctly. And so you can use that for clever image editing. And the reason it can do that is because it's not autoregressive. There are autoregressive image generators, which go left to right, top to bottom, in raster order, generating pixels. But diffusion doesn't work like that. It'll just see the entire image and then start to denoise it. And because of that, because it gets to see every pixel, every pixel gets to see every pixel, it can fill in the missing information and do it in a way that's consistent with whatever prompt you're giving. So we can do something similar.
SPEAKER_02
So I have a couple of demos here. Can you see that? Yes. So this is just some code and you say, there's a bug in this code, can you fix it? And it'll just make the edit in the correct place. Like it won't, you can barely see that, but it's a little fix here of the indices. And you can say things like, can you add documentation? It'll go in. It's not just one by one generating all the tokens, it's doing a clever editing procedure to actually fill in the correct edits here. You can do that with more general text. Like you can take a story and then say, add a middle paragraph.
SPEAKER_02
And because it can see that the first and third paragraph, it can fill in the paragraph in a way that's consistent with the rest of the story. And this is just in-place editing. Okay. So those are some of the advantages. Don't have a lot of time. The biggest advantage, as I mentioned, is this low latency. And we really lean into that. And I just want to show you a couple of demos of some of the things that people internally have built to show what the advantage of low latency can give you. So it's not just the same thing faster. It can really unlock some new applications. So in your own work, you're all AI engineers.
SPEAKER_02
[SPEAKER_05] It'd be interesting to see when the next diffusion model comes out from our team, what the low latency could unlock. And what new applications can be built. So here are just some demos. So this is Wikipedia. Let me just pause it actually. This is Wikipedia where everything is generated on the fly, even the HTML. And I just want to show you a couple of demos of some of the things that people internally have built to show what the advantage of low latency can give you. So it's not just the same thing faster. It can really unlock some new applications. So in your own work, you're all AI engineers.
SPEAKER_02
[SPEAKER_05] It'd be interesting to see when the next diffusion model comes out from our team, what the low latency could unlock. And what new applications can be built. So here are just some demos. So this is Wikipedia. Let me just pause it actually. This is Wikipedia where everything is generated on the fly, even the HTML. So this is actually being generated by the model on the fly. So that it's a webpage with the HTML and the text and everything being generated on the fly. So it looks like regular Wikipedia. And when you click on it, the latency is low enough that it can just fill in the page as if it was a real Wikipedia page.
SPEAKER_02
So that's Wikipedia generated on the fly by just a very low latency model. We have a similar thing where we did it for Reddit. So now all the responses to your posts will be by bots. They weren't already. And it's generating fake comments. So the Gemini diffusion model was not an image generating model. So this demo links in the startup, state of the art image generator model at the time, which I think was Juno. So it was before nano banana. So it's the two of these models working together to fill in the webpage. So the image generation model is a little slower.
SPEAKER_02
But you can see that you can invent any Reddit you want sharks in this case, and it'll generate the page with the text. The images follow a minute later. And then you can interact with this website as if it was a real website with real users and so on. Just entirely, all the comments, all the images, all the HTML, everything is being generated entirely on the fly here. This is being generated by the model. I love this one. This is my favorite one. This is an operating system also being entirely generated on the fly. So every click here is generating the next page of the operating system.
SPEAKER_02
So it looks like a real operating system, but it's all being generated by the model on the fly, responding to every click. Every time you enter the readme, it generates the text. It also generates the web pages you can go back to the desktop and so on. I think this, yeah, okay. And then this is a demo from someone on Twitter who used the Gemini Diffusion API. Or sorry, not the API, the web page to do some vibe coding with his voice. So I really like this one. Create a to-do app. Add 10 random to-dos. Allow to-dos to have a completed state. Mark four random to-dos as completed. Allow me to sort to-dos by name and by state. All right, let's see if this works.
SPEAKER_02
Sort by name, sort by state. Testing, enter. It was added to the bottom. We'll try deleting a few. Let's add one. Everything's working. Please convert this to dark mode. And this was literally 15 seconds of work. Okay, so that was someone outside of our team, so he could say that. Yeah, vibe coding by voice. But in general, we think that low latency models can really unlock some new experiences for users and new products. [SPEAKER_00] And so we're excited to see what people will do when the next generation comes out. [SPEAKER_00] Okay, and on that note, thank you very much. [SPEAKER_00] Questions? [SPEAKER_00] Yeah? [SPEAKER_00] Yeah. [SPEAKER_00] Yeah, yeah.
SPEAKER_02
[SPEAKER_00] So, yeah, we use all the same data. [SPEAKER_00] Yeah. [SPEAKER_00] The algorithms have to change a bit, but we use all the same data.
SPEAKER_05
[SPEAKER_00] Yeah.
SPEAKER_02
[SPEAKER_00] And also maybe, can you distill these models? [SPEAKER_00] Are there any techniques for them? [SPEAKER_00] You can distill them, yeah. Yeah. I'm not sure if there are any published ones. Are there any published ones? I don't think so. Maybe there's some externally, but. Will we have the new future lecture about the. There's a, yeah, so we're going to release something soon. Yeah. Yeah, yeah. Yeah, so there's a, the bigger models tend to require less steps for the same output. So even if the model is getting bigger and the flops per forward pass are getting bigger, they tend to reduce the forward passes they need.
SPEAKER_02
So it's a situation where you have some sort of a diminishing cost of serving even the biggest models. [SPEAKER_05] [SPEAKER_05] So the next question, did you, how do you try to get it? I don't know. We haven't got to that yet. I'm a research scientist. Yeah. How do you define the size of the answer? Because for a few weeks, you expected I have this image and I have the output frame. [SPEAKER_05] How do you do that in coding or text? So there's a few different ways. The easiest way is to just fix some window length and then just iterate on that. [SPEAKER_05] So it's autoregressive, blockwise autoregressive. It's the standard way to do it.
SPEAKER_02
But you can have a head that will predict the length of the response and stuff like that. Yeah. Yeah. But you have to fix delta in order to. No, it still can generate unlimited text, but it's just, if you fix a window length, then [SPEAKER_06] it just does that window length autoregressively if it needs to generate many, many windows of text. Yeah? You might have already mentioned this, but when you give that example of the denoising steps being different ranges to different.
SPEAKER_02
[SPEAKER_04] Yeah.
SPEAKER_02
[SPEAKER_04] Could you ahead of time set a limit on the denoising steps that you apply for a problem so you can almost understand what your latency is going to be ahead of time?
SPEAKER_02
Mm-hm. Yeah, yeah.
SPEAKER_02
These are all with a limit, but they just finish earlier than the limit. [SPEAKER_06] It just does that window length autoregressively if it needs to generate many, many windows of text. Yeah? You might have already mentioned this, but when you give that example of the denoising steps being different ranges to different. [SPEAKER_04] Yeah. [SPEAKER_04] Could you ahead of time set a limit on the denoising steps that you apply for a problem so you can almost understand what your latency is going to be ahead of time? Mm-hm. Yeah, yeah.
SPEAKER_02
These are all with a limit, but they just finish earlier than the limit. Oh, right. Okay. Yeah. [SPEAKER_04] If you have multiple windows like you just said going through, can they then go back and attend to previous windows?
SPEAKER_00
[SPEAKER_02] Yeah. [SPEAKER_02] As the window goes, it's.
SPEAKER_00
[SPEAKER_02] Yeah. [SPEAKER_02] You could potentially, but for us we just set it in stone and continue. [SPEAKER_03] Yeah.
SPEAKER_00
[SPEAKER_03] Lots of versions here. [SPEAKER_03] The version is it can go a lot, but it just wouldn't make sense to make a hybrid. [SPEAKER_03] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] That's how it works. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah.
SPEAKER_02
The start to prepare with the autoregressive style and then go to the division mode or? [SPEAKER_01] Oh, no. [SPEAKER_01] So it's pre-fill is the same. [SPEAKER_01] It's just that you've got some context and you pre-fill. [SPEAKER_01] And then after that, the generation step is typically in blocks of some fixed size, like 512 or 1000 or 32 or whatever you want. And then that's autoregressive. Yeah.
[SPEAKER_02] I guess that you're doing the demo thing inside a bi-directional. [SPEAKER_02] So how do you, what is the form supposed to gain back to the open ID?
SPEAKER_02
[SPEAKER_05] Is it a latent institution instead of? So you can do that. The easiest way to do it is to just have a logits at the top. Just vanilla prediction head at the top. It's after the demo thing called that?
SPEAKER_05
[SPEAKER_02] Well, so this is all discrete diffusion, right? So it's always tokens in, tokens out.
SPEAKER_02
If I show you the... For every step. For every step, yeah. So if I go back here. It's always a discrete corruption process. And then you fill in a discrete token back in.
SPEAKER_05
[SPEAKER_02] So you're always in a discrete space.
SPEAKER_02
[SPEAKER_05] But you can do it in latent spaces, but people have done that. [SPEAKER_05] But most of the text diffusion models and literature is discrete diffusion today.
SPEAKER_05
Yeah.
Is that only a demo? It might be one day. Yeah. Yeah. [SPEAKER_05] What's your outlook like? [SPEAKER_02] Do you think one of the architectures will win in the future? [SPEAKER_02] Or do they have different use cases?
SPEAKER_02
I think for now they have different use cases. So if you think about what a low latency model provides, that's worse throughput. What's that trade off? On-device applications.
SPEAKER_04
[SPEAKER_02] So we are in a couple of on-device applications already. [SPEAKER_02] And within the alphabet ecosystem. [SPEAKER_02] So robotics, things like that.
SPEAKER_02
We want to run a model on the device itself. So your phone or a robot or whatever. And then you want it to be low latency. And you're not batching with thousands of other queries like Gemini being served on the server side. So you want the lowest latency model. Quality isn't really any difference, they're the same quality basically.
SPEAKER_04
[SPEAKER_05] So you may as well pick the low latency one. [SPEAKER_05] Because you don't have the throughput concerns.
SPEAKER_02
[SPEAKER_05] What do you mean in the future, can you get quality up to the bar with the current proxy models? [SPEAKER_05] Quality isn't the concern, it's the throughput for serving in a big batch setting. Yeah. Yeah.
SPEAKER_03
[SPEAKER_02] So currently you are getting more of the sense of quality of function models in the area. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah.
SPEAKER_03
[SPEAKER_02] Yeah.
SPEAKER_02
Yeah. Yeah. Yeah. [SPEAKER_05] You can do RL. Yeah. [SPEAKER_05] You just need to change the algorithm.
SPEAKER_01
[SPEAKER_05] Yeah. [SPEAKER_02] But you can still do it. [SPEAKER_05] Yeah. [SPEAKER_05] Yeah.
SPEAKER_02
Yeah. Yeah. Yeah. [SPEAKER_02] The problem with that is if you train it with noise it expects noise. You'd have to train it with this more model.
SPEAKER_05
[SPEAKER_03] As you like the output.
[SPEAKER_03] That just adds complexity. So we don't usually do that. Yeah? Is that a paper? [SPEAKER_02] A word you said, and if you want to come in with the idea of our team? [SPEAKER_02] Well, not from our team, but there's a bunch of literature out there, yeah. [SPEAKER_02] Okay.
SPEAKER_02
Yeah. Is that a spectrum? Yeah. You get the idea, I think. Yeah. Any other questions? There's a lot of questions.
[SPEAKER_02] I've gone through them all. [SPEAKER_02] That's good. [SPEAKER_02] Oh, one more. [SPEAKER_02] Okay. [SPEAKER_06] Just a random one. [SPEAKER_06] What happens if you ask it to generate? [SPEAKER_06] I don't know.
SPEAKER_05
[SPEAKER_02] I don't think we've tried to do that.
SPEAKER_02
Probably would work. We'd probably do something. Yeah. [SPEAKER_05] Cool. Okay. Thanks, everybody. [SPEAKER_06] I don't know. I don't think we've tried to do that. Probably would work. We'd probably do something. Yeah. [SPEAKER_05] Cool. [SPEAKER_02] Okay. Thanks, everybody. Oh, my God.
SPEAKER_05
[SPEAKER_02] Oh, my God. [SPEAKER_02] Oh, my God. [SPEAKER_02] Oh, my God. [SPEAKER_02] Oh, my God.
SPEAKER_02
Thank you. Yeah. Yeah. So currently you are getting more like the sense of quality of function models in the area. Yeah.
SPEAKER_02
Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. You can do RL. Yeah.
SPEAKER_05
You just need to change the algorithm. Yeah.
SPEAKER_02
But you can still do it.
SPEAKER_05
Yeah. Yeah.
SPEAKER_02
Yeah. Yeah.
SPEAKER_02
Yeah. Yeah.
SPEAKER_02
Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. The problem with that is if you train it with noise it expects noise. is you'd have to train it with this more model
as you like the output. That just adds complexity. So we don't usually do that. Yeah? Is that a paper? Like a sort of word you said, and if you want to come in with the idea of our team? Well, not from our team, but there's a bunch of literature out there, yeah. Okay. Yeah. Is that a spectrum? Yeah. You get the idea, I think. Yeah.
SPEAKER_02
Any other questions? There's a lot of questions. I've gone through them all. That's good. Oh, one more. Okay.
SPEAKER_06
Just a random one. What happens if you ask it to generate?
SPEAKER_06
I don't know.
SPEAKER_02
I don't think we've tried to do that.
SPEAKER_02
Probably would work. We'd probably do something. Yeah.
SPEAKER_05
Cool.
SPEAKER_02
Okay. Thanks, everybody.
SPEAKER_02
Oh, my God. Oh, my God. Oh, my God. Oh, my God. Oh, my God.
Oh, my God. Oh, my God. Thank you.