AI Engineer

SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind

2117 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Google DeepMind sees generative video as a complementary foundation model for space-time reasoning, with the near-term product opportunity in multimodal creation and editing and the longer-term opportunity in unified world-model and agent systems.
  • Why it matters: The panel offers unusually direct signals on where frontier generative-media capability is heading: joint audiovisual generation, natural-language video editing, multimodal model unification, evaluation bottlenecks, and the need to capture real production workflows rather than merely collect more content.
  • Best use: Use this as strategic input for media-agent product design, especially for building workflow traces, human-in-the-loop evaluation loops, controllable generation, and customer-feedback channels that improve models upstream.

Executive Summary

The panel combines product updates with a research and operating thesis. DeepMind announced Nano Banana 2 Lite as a fast, low-cost image-generation/editing model with roughly three-second latency, plus Gemini Omni Flash APIs for video generation and editing priced at the VO3 Fast level. The speakers position these not merely as novelty tools but as building blocks for short-form production, marketing assets, educational content, translation/localization, and natural-language video editing.

Their core technical argument is that video models should become a major complement to language models. Language remains valuable because it is a universal interface and a strong causal-style conditioning representation learned during pretraining, but it is lossy for visual, temporal, auditory, and aesthetic information. Video models can supply space-time and physical intuition; the expected direction is tighter integration of visual reasoning, video generation, video understanding, and text-based reasoning rather than treating generative video as isolated content creation.

The panel is candid that model quality is not solved by benchmark scores or generic preference optimization. Human raters, trusted creative users, live experiments, automated checks such as OCR, and model-based evaluators all play roles, but free-form editing and aesthetics remain hard to evaluate. Optimizing simple human preference can also produce undesirable defaults such as over-smoothed, oversaturated outputs or recurrent artifacts that creators notice but general evaluators miss.

The strongest operating lesson is about data and post-training. The scarce asset is not random web video; it is high-quality data plus traces of how professionals actually work: how an image becomes a campaign, then variants for different ad formats, through iteration and selection. The speakers frame field deployment engineers as a feedback system between customers and upstream model development, not simply a sales or implementation function.

Key Takeaways

  • Claim: Generative-media models are moving from one-off visual generation toward practical, low-latency creation and editing systems. | Evidence: Nicole Brichtova says Nano Banana 2 Lite is faster and cheaper than the original model, approaches larger-model quality, and operates at about three-second latency; Gemini Omni Flash APIs expose video generation and editing to developers and are priced at the VO3 Fast level. | Implication: For product workflows where iteration speed matters, architect around fast preview/edit loops with selective escalation to premium models instead of treating media generation as a slow, single-shot production step. | Caveat: The panel does not provide exact price figures, benchmark comparisons, API limits, or production reliability metrics.
  • Claim: The durable product surface for video generation is multimodal input-to-video and natural-language editing, not simply text-to-video prompting. | Evidence: Brichtova describes using image sets as storyboards and audio tracks as voice references, then producing video; she also highlights adding or removing objects, cleaning noisy video, marketing/ad creation, education, translation, text rendering, redubbing, and localization. | Implication: Media agents should accept and preserve rich source context—reference imagery, video, audio, brand assets, and desired edits—rather than collapsing the task into a text prompt. | Caveat: These are described as emerging use cases, and the panel expects many of the highest-value applications to be discovered through API users rather than Google first-party products.
  • Claim: Language is a powerful but insufficient intermediate representation; the expected architecture is joint text and video reasoning rather than language-only control. | Evidence: Shane Gu argues that language conditioning helps avoid spurious correlations because it encodes a causal description of what is being generated, while Dumitru Erhan says language alone is not sufficient and video is a required complementary foundation model if AI is to match human-like perception. | Implication: Do not assume a prompt-only agent interface is enough for visual-world tasks. Maintain multimodal state and use language as an explicit control layer on top of visual and temporal representations. | Caveat: The speakers regard the optimal intermediate representation as an open research question; code, continuous latent reasoning, and other representations remain under exploration.
  • Claim: DeepMind expects generative video to follow a trajectory similar to language models: early creative demos, better instruction following, reduced hallucinations, then increasingly reliable reasoning and test-time scaling. | Evidence: Gu compares current video models to early language models before robust instruction following and reasoning, and predicts stronger video instruction following and reduced inconsistencies could enable interleaving space-time simulation with text simulation for broader AGI tasks. | Implication: Treat current video models as valuable but fallible components for perception, simulation, and generation; build verification and constrained execution around them rather than granting them autonomous authority. | Caveat: This is a research outlook, not a stated product roadmap or a claim that current video models are reliable world simulators.
  • Claim: Joint audiovisual generation is architecturally preferable to generating video and audio separately and stitching them together. | Evidence: Erhan says VO3 jointly generated audio and video because both arise from one latent causal process: speech, lip movement, and environmental sound need to remain synchronized. He contrasts this with previous approaches that generated pixels then added lip-sync/audio layers. | Implication: For immersive or character-based content, evaluate native joint audiovisual models first; external lip-sync and audio post-processing may remain useful controls but should not be assumed to deliver coherent realism. | Caveat: The speakers acknowledge unresolved challenges in representing nonverbal audio qualities such as tone, acoustics, music, and other sensory properties that language describes imprecisely.
  • Claim: Generative-media evaluation remains fundamentally hybrid: automated checks can cover objective failures, but expert human judgment and real workflow feedback remain essential. | Evidence: The team uses OCR-like checks for text-rendering correctness, thousands of human evaluations, live experiments, side-by-side internal reviews, trusted testers, and feedback from daily creative users. They cite artifacts such as models adding wedding rings to hands and regressions such as blurred grass detail that a trusted tester surfaced. | Implication: Build layered evals: objective invariants, task-success tests, expert review, regression suites from real failures, and production telemetry. Avoid optimizing a media agent solely for generic aesthetic preference. | Caveat: Simple preference ratings are a weak objective: generated video may win side-by-side comparisons by looking sharper or more HDR-like while being less realistic or less useful for the task.
  • Claim: The highest-value training and evaluation data is workflow trajectory data from real users, and field deployment should feed that signal back into model development. | Evidence: The panel contrasts random web video with professional-quality data and asks for the full path from a product image to a video ad to channel-specific campaign assets. Gu defines post-training broadly as everything between pretraining and final user experience and argues FDEs should derive upstream product/model insights, not only deploy solutions. | Implication: Capture structured traces of user intent, references, iterations, selections, corrections, and final deployment outcomes. Position implementation/customer-success functions as an evaluation and product-learning loop, with appropriate consent and data governance. | Caveat: The speakers do not disclose concrete data partnerships, collection mechanisms, or terms for contributing workflow data.

Detailed Brief

Model portfolio strategy: unification is directional, specialization remains practical

  • Claims: Gemini Omni signals a long-term ambition for fully multimodal inputs and outputs, potentially including image generation and editing within the same family.; The panel does not expect every specialized model to disappear immediately because latency, cost, resolution, duration, training, and product constraints differ materially.; Whether modalities should share one checkpoint remains partly empirical: image and video have obvious transfer, joint audio-video has strong causal coupling, while transfer between coding, 3D, and video is less certain.
  • Evidence: Erhan contrasts a lightweight image model with a hypothetical 4K, 30-second video model as serving different engineering and user needs.; The speakers cite jointly generated audiovisual content as a case where combining modalities is clearly justified because lip movement and sound must arise coherently.
  • Caveats: Claims about a five-year unified model versus a six-month specialized portfolio are speculative framing from the panel, not a release commitment.; A universal model can be technically elegant while being inferior on cost, latency, or controllability for a specific product workflow.
  • Implications: Use a routing layer across specialized models today, but design shared multimodal asset/state schemas so products can adopt more unified models without rewriting workflows.; Make model selection an operational decision based on fidelity, turnaround time, editability, and cost rather than a branding decision based on which model is most general.

Control, aesthetics, and the limits of default generation

  • Claims: Prompting and reference-based control remain important even as models auto-prompt or produce stronger defaults, because users need a way to express trained aesthetic judgment.; Default aesthetics are consequential product decisions, not neutral behavior; teams implicitly choose color palettes, saturation, density, and composition preferences.; Professional creators detect subtle failures—eye gaze, micro-expressions, skin texture, physical scale, repeated patterns, and brand-specific shades—that broad user preference tests can miss.
  • Evidence: The team describes Nano Banana Pro producing overly dense infographics, internal tuning discussions over muted versus saturated palettes, and creator feedback that surfaced visually specific regressions.; Examples of real production constraints include preserving a rug pattern across custom sizes, correct earring-to-head scale for virtual try-on, and preserving a brand's exact rather than approximate visual language.
  • Caveats: The panel does not offer a robust formal solution for encoding brand language or expert aesthetic standards.; A generated asset that is preferred at a glance may still fail under professional inspection or be unusable in a production pipeline.
  • Implications: Give operators controls for references, style constraints, brand palettes, asset consistency, and iterative correction; do not rely on one global default.; Create domain-specific regression sets around failures that matter to paying users rather than using generic visual-quality leaderboards.

Notable Concepts & Terms

  • Nano Banana 2 Lite: DeepMind's fast, lower-cost image generation and editing model, presented as a replacement for many original Nano Banana use cases and an enabler of rapid ideation.
  • Gemini Omni Flash: Developer-facing API offering for video generation and editing, positioned around multimodal inputs and natural-language transformation of video.
  • Joint audiovisual generation: Generating pixels and audio together from a shared generative process so speech, lips, ambient sound, and scene events remain coherent.
  • World model: Used loosely in the discussion, but Gu anchors it to model-based reinforcement learning: a model that represents and predicts environment dynamics, with video offering space-time intuition.
  • Spurious correlation / causal conditioning: Gu's rationale for language conditioning: detailed language descriptions can constrain generative learning toward causal factors rather than accidental correlations in training data.
  • FDE: Field deployment engineer; the panel argues this role should close the loop from customer workflows and failures back to evaluation, post-training, and upstream modeling.
  • Post-training: Gu's expansive definition: all work between base pretraining and the final user experience, including customer harnesses, feedback, evaluation, and adaptation.
  • Workflow trajectory data: The sequence of intent, assets, iterations, choices, edits, and deployment outcomes that reveals how professionals actually perform a task; presented as more valuable than static output data alone.

Operator Notes / Why Ken Should Care

  • Instrument any generative-media workflow to retain consented traces of source assets, prompts, references, edits, candidate comparisons, accepted outputs, manual fixes, and downstream performance.
  • Establish a media-agent eval stack with hard checks for text, dimensions, brand constraints, and asset consistency; add expert review for aesthetic, physical, and workflow-specific failures.
  • Implement tiered model routing: use fast models for ideation and high-volume edits, then escalate only shortlisted assets to higher-fidelity generation or human review.
  • Treat references as first-class inputs: preserve image, video, audio, and brand artifacts in the agent state instead of translating them into prose-only prompts.
  • Create a customer-facing failure-intake loop modeled on an FDE function, with a taxonomy that turns recurring production failures into regression tests and product requirements.
  • Avoid setting success metrics around generic preference alone; measure whether outputs are deployable in the customer's actual workflow, including correct scale, consistency, localization, and brand adherence.

Source/Metadata

  • Title: SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind
  • Transcript words: 13161
  • Duration seconds: 3419
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript. The transcript also contains repeated passages and at least one apparent extraction artifact.
Full transcript 10045 words · 58 min read
0:00

And welcome back.

0:11

For those on the stream and those in person, we tend to take these longer sessions between all the main stage keynotes to reflect on things that are particularly important but don't have a significant launch moment. Today we're very lucky to have people working on Omni, Vio, Nano Banana, the world's best generative models, here with us. Demetrio, I first saw you when you were posting about your office. I think you're probably Google's number one office influencer, at least in San Francisco. I think you like to bike as well. You like to take photos of something. I'm not sure if you're on my bike here. Yeah. But also you work on video models. That's right.

0:30

Shane, I met you, I think, at a dinner. Yeah. And I remember you were trying to get me invested in one of the companies. I forget which one. No, forget about that.

0:46

But now you're working on Omni thinking, and just a bunch of other things. Gemini, Gemini RL. Yeah. Yeah. And Nicole, also the rest of the Gen Media models. Nano Banana and everything you just launched, actually, even this week. Yeah, we launched some APIs. Yeah, yeah, yeah. And I haven't tried to convince you to invest in anything, but maybe I should.

0:57

So I try not to be an investor. People just convince me anyway. I'm like, okay, well, I'm not that rich, but no. You can't not try to invest in some of these things. And for those of us who are not working at the Frontier Lab, this is the best, the closest they will ever get. So, actually, let's recap, since you're closest to it and we just did it. What was launched this week? What should people go try out?

0:59

Yeah. So, yesterday we had two launch moments. One of them, we launched Nano Banana 2 Lite, which is our fastest, cheapest image model in the Nano Banana model family. And it's better than the original Nano Banana. So, really, for most people, that model replaces what you used and loved the original Nano Banana for across generation and editing. And it gets really close to the frontier quality of the mainline bigger models.

1:01

So, that's really exciting. I think if you look at some of the demos or things that people have been trying, getting that three-second latency just unlocks a whole bunch of things that you can do with ideation and iteration. And it's just really fun. And the models are getting to a point where the quality is really good, where you can use it for ideation, but you can also use some of those outputs as ready production outputs. So, that's really exciting.

1:02

And then, second launch, we finally launched the Gemini Omni Flash APIs that we pre-announced at I.O., so thank you for waiting. And that is the first time that we're making the APIs available for developers. And it's really exciting video generation and editing. And we're pricing it the same as EO3 on Fast. So we're getting you really, really good quality for a really awesome price, hopefully. Yeah. I mean, that's incredible. I would actually really...

1:08

So, when you guys launched Omni for the first time, you also did a podcast with Logan, who couldn't be here today. And you added a sloth and ramen and all these things. I actually really want to do that to our videos. I just didn't have an API for it because, obviously, I have to automate the whole thing. So thank you for the API. That is my favorite use case. Everybody should do that.

1:13

I got a cat, which is probably the most boring of the animals. If you don't know what we're talking about, you should look it up. It's very funny. Fofur, who's on the team, did that. Fofur is the number one guy you should follow. You should follow Fofur. Get ideas on, okay, what can this thing do? Yes. Right? Yes. He's amazing at that. I've tried to get him for the last two years to come to AIE. He hasn't made it yet. He's actually come in person. He just didn't want to speak because he's anonymous. I know. I want to say his real name, but I can't say his real name.

1:25

No, no, no. We won't do that to him. But you should really follow him. He's amazing. He did all that work. I actually met him in the office when we did the podcast, I think, and I didn't realize it was him, so his badge doesn't say Fofur. Yeah. I know. He used to be part of Replicate, and Replicate had this joke where everyone was Deepfates. Deepfates is this mysterious character. Replicate is a very cool company, and Fofur was part of it. So, okay. One thing I wanted to get on there before I go into omniproper is we added cats, we added sloths. Very cool, very cute, very fun. What are the more workhorse use cases that are not just demos?

1:37

Yeah. So, obviously the hero capability of the model, or maybe there are two, one is the ability to take in anything as input and then get video on the other side. Obviously, in the future, and we've talked about this as a pre-announce, we want to get the other output modalities out as well. But basically what that means is, you can take a set of images that you have as maybe a storyboard, you can take an audio track as a reference of a voice that you want a character to speak, and then you can get a video on the other side.

1:38

So that just unlocks a whole bunch of things that you can do in short film production or shorts. We've launched on YouTube as well to help creators create content more easily. And then the other one is obviously video editing. That's another thing that we're really excited about that we're just making easier, because now you can use natural language to take a video, add something, remove something. Sloth is obviously a fun example.

1:41

But there are consumer use cases that we had in mind where you could take your beachification video that was too noisy and you want to clean up that noise. Maybe in the past you wouldn't have because you didn't have the tools or you didn't know what the tools were that you needed to go to. So that's one use case that you can go to. We've seen a lot of folks use it for marketing, ad campaign creation, and I'm excited to see more of those use cases as we launch the APIs. Because obviously we don't see all of it in the first-party products, but I'm really excited for people to start to explore that in the API.

1:43

So those are just some of the high-level things that have come up. People also use it to create education materials. Yes. And that's really exciting. I think we've all talked about being excited about the future of education, where everything can be customized to you and personalized to your knowledge level and the style that you prefer. And so this is just a step in that direction.

1:51

Yeah, I actually used Nano Banana yesterday. My parents are visiting, and there was a very fun use case. I bought some gadget on Amazon that they wanted, and the instructions to use it were only in English, and there were plenty of diagrams or whatever, and I took a picture of it and said, translate this into Romanian and keep everything else the same, right? So it was amazing, right? It looks identical, and it is perfectly translated, more or less, right? But it's using Gemini under the hood, obviously, to do the translation.

1:54

And so you can see this use case for video as well, right? The power of text rendering in Omni is quite next level. And you could think about plenty of use cases of both text rendering, translation, internalization, all sorts of things that would be genuinely useful to a lot of different people and broader access to either you could redub a video or whatever it is that you wanted to do. There are plenty of different things that you could think about doing. Yeah. Yes. And keep everything else the same, right? So it was amazing, right? It was just, yeah, it looks identical, and it has, it's perfectly translated, more or less, right?

2:07

But it's using Gemini under the hood, obviously, to do the translation. And so you can see this use case for video as well, right? The power of text rendering in Omni is quite next level. So, and you could think about plenty of use cases of both text rendering, translation, internalization, all sorts of things that would be actually genuinely useful to a lot of different people, and broader access to either you could redub a video or whatever it is that you wanted to do. There's plenty of different things that you could think about doing. Yeah. One of the most enlightening conversations I have on my podcast is with researchers at the frontier of these things.

2:33

I had one with Ethan from the XAI video team, the Grok video team, who was basically saying the next trend is actually not just single model. It's more of the video agents. And I don't know if that terminology resonates, obviously, very relevant for RL. But it was basically giving up on trying to do everything in effectively one pass. Do you feel that same way? Or is it still an open research question which way the trends are going? Yeah. So, what excited me most is really when the symbolic foundational models and this video foundational model can actually really work together.

2:56

And in a way, if you look at the beginning of the generative image generation, video generation, a lot of it started when the language model got good enough to provide very detailed captioning, from stable diffusion days or DOWE-2 days. So, basically, language is an extremely helpful representation. One is that it's universal. But the other more technical thing, my hypothesis, one very difficult thing about machine learning is this spurious coordination. So, you don't know if this feature, right, that's going to predict it, is actually a causal factor or not. There are two ways.

3:20

One is we can have really diverse data, training data, from every intervention of the causal graph. The other is you condition the causal information. And conditioning the language is like conditioning a causal information of the world. So... Which is a prompt or a concept or what? Yeah. Yeah, exactly. So, if you look at how are we going to describe this video, how are we going to describe this image, it's actually very close to how we describe this causality behind this, how this is generated. So, one is that can really allow for very rich generalization and then just a good model.

3:40

The other is, eight months ago, we put the evaluation paper called Video Models, Zero-Shot Learners and Reasoners. Yes. So, that was a confirmed paper. And then later on, actually, the Nano Banana team filled up with the Vision Banana paper that basically used Nano Banana to do. But essentially the idea is video model is extremely good for the model, sort of a foundation model for space and time kind of information. So, classic computer vision tasks, a lot of it could be zero-shotted. And when you say feeding something like a visual quiz, it can, there's definitely a lot to improve, it can solve.

3:54

And it can, robotics kind of scene, it has very good physical intuitions, world model. And I think that the key is really the mix of the visual reasoning and then the text reasoning all tied together. Obviously, what are doing it as unified model versus this agent calculation, I think that's more like it's going to be more incremental, how it's going to, I imagine everything's going to go into a single model eventually. Yeah. But right now there's a lot you can do if you basically take very good video understanding, image understanding, Gemini agentically with the nominee. And that's actually going to, yeah, our team is exploring a lot. Yeah. Yeah. Okay.

4:07

There's a lot in there. I think one question I am increasingly starting to wonder is, does it all trend towards one product for you guys, right? Now you have multiple models out. The naming of Omni does imply that eventually everything will go away and it just goes into Omni. Is that the plan? Is it? I don't know. I think maybe, I think eventually. I think there's different trade-offs, engineering, research, product trade-offs in, for the same reason, sorry, how's it called? Nanobanana light? I don't know what the product name is. Nanobanana too light? I don't know what the product name is. Nanobanana too light, yeah, right? It serves a particular niche, right?

4:36

And it probably doesn't necessarily fit immediately in the same model, literally checkpoint, as something that can do 4K, 30-second videos, right? They're probably not trainable in quite the same way, right? So, I don't know, it depends on how far into the future you look. Sure, in five years from now, will they all be the same model? Probably. But six months from now, we'll probably still have multiple different models doing different things because, pragmatically, the trade-offs are such that we should have multiple different kinds of models. Yeah. I think that's right.

4:59

And just on that note, we did call it Gemini Omni because we wanted to hint at the future where Gemini just becomes fully multimodal in and out, right? And so, it's definitely a move in that direction. I think we'll probably see a move in the direction where Omni also generates images and edits images and all those kinds of things. But Doom is right that I think on the way there, there's a bunch of really, really useful applications of some of these more specialized models. And so, we will probably continue to work on those as well because that serves a certain need at this point in time that may not exist a year from now.

5:12

There's also a research question about just how much transfer there is between different kinds of modalities, right? I think you may believe that there's some transfer between coding and video generation. And I think most people don't necessarily believe that. But you could try to think that there is something there. Or it could be a waste, right, to put them together, to try to learn both tasks at the same time, right? So, I think it's an interesting question to which extent image and video, obviously, there's some transfer. Not that different. There's value in learning to output video and audio at the same time because joint audio-visual, that's how it is.

5:41

And then there's other intersections of modalities that are not super obvious, right? Like 3D representation, coding, I don't know, maybe. Things like that, right? So, I think it's worth exploring the different corners there. And we are actively doing that with a focus towards what people actually want to do with this one. Yeah. One thing I feel surprised by, but also I feel like it's insufficiently answered, is what is the correct intermediate representation? So, captioning, right? XCI does captioning. Omni does captioning. And I understand how captioning works for images. And I understand that you can extend it into video and guide it across time.

6:19

It just feels very inefficient. There's got to be, I feel like there should be something better. Maybe it's code. And maybe we generate, and obviously I think a lot of FFmpeg and Matplot, what's the three blue, one brown one, Manim. A lot of video is generated through code. And maybe that's the optimal representation. Any hypothesis as to, is it better, or is just English all you need? Well, so, I'm in the Gemini and we do a lot of RL agent and, of course, coding. So, yeah, we're definitely exploring the coding representations. Yeah. That's a better way to represent. XCI does captioning. Omni does captioning. And I understand how captioning works for images.

7:04

And I understand that you can extend it into video and guide it across time. It just feels very inefficient. There's got to be, I feel like there should be something better. Maybe it's code. And maybe we generate, and obviously I think a lot of FFmpeg and Matplot, what's the Three Blue One Brown one, Manim. A lot of video is generated through code. And maybe that's the optimal representation. Any hypothesis as to, is it better or is English all you need? Well, so, I'm in the Gemini and we do a lot of RL agent and, of course, coding. So, yeah, we're definitely exploring the coding representations. Yeah. That's a better way to represent.

7:46

But what's your probability estimate on we just output binaries? We just, it's just ones and zeros. I guess maybe a similar discussion was, is the language the right representation? Right. So, one question, for example, professor, someone asked is, why does the chain of thought need to be in natural language? Yes. Can it just be any kind of continuous tokens, just any amount of additional computations? So, one is, obviously the test, adaptive compute is going to give better results. So, it's that. But what really made chain of thought, four years ago I wrote the larger model zero-shot reasoner and then self-improvement. So, I know from the very early days.

8:34

But the reason it works really well is right now the recipe that works is the pre-training that scales a lot and then learns intelligence. There are a lot of scaling RL, but those are still extremely compute center intensive to extract information. And you really want to rely the intelligence on that. So, basically, by tying the reasoning in natural language, you basically directly use the intelligence of the pre-training to it. So, if you remove that kind of constraint, then you're not. And these days, I feel a lot of advancements in the text, but also in this multimodal space, is really driven by this text as a great representation. Yeah. It's a good backbone. Yeah.

9:11

I think, to me, it's even simpler than that. It's text is how we communicate. So, I think fundamentally, if you're building product that humans will be interfacing with, that we will be using text somehow, if it's a text interface, right? Not everything. So, I think it's natural to default to that. Yeah. Obviously, there's a confused discussion, some RL maximalists, who are like, oh, we don't care about chain of thought. It's just additional compute. Sure. But I personally, yeah. RL maximalists. I wonder who qualifies in that description. David Silver. Ah. Okay. Yeah, I mean, they've just left to start their thing. Interesting. Okay.

10:16

So, I mean, I think I'm very interested in better representations. Because I think that's one of our themes that we're curating today at the World's Fair is world models. You mentioned the word world models, but it's not something that's super well defined. I think everyone's converging on some version of it that it's the ideal. Sure. Everything is a world model now. It's not that useful, right? So, I just gave a keynote at the iClear's world model workshop. Yeah. And then, essentially, I definitely encourage to check out the definition by Jitendra Malik. He's the OG computer vision professor. Yeah. He has a bit of word to say about world model.

11:11

But also, Ganschmidt-Huber is how he defined the word model from 2019. Like, I don't like 1990. When it was basically just that model base. For me, the world model is basically just the model in the model-based RL. And I feel that is sufficient to describe. But obviously, there are a lot of, Fei-Fei had a nice blog post about what she was about. Yeah, just broken down. But yeah. Yeah, I mean, I'll end this part of the conversation, but I do think that language to me relying on language as the narrow pipe through which everything goes through still is a lossy compression. No, no, no. But we're not seeing that, right?

11:41

We're basically seeing the video model and the language together. Yeah. So, I think the language alone is not sufficient. That's why we feel the video is a very complemented foundation model. Yeah. Right now, the view only many people in a few hours are generating pretty videos. But I think our vision, it's much more than that. It's a missing foundation model that's absolutely required if you want to make the AGI match the humans, not just the jacked one. Yeah. Okay. So, one other thing, you mentioned on the vision side, and I'm kind of curious how parallel, in terms of your research careers, this development is.

12:22

I think basically a lot of vision people have crossed over into world model people. A lot of vision people also become generative video and image people. And is it just as simple as reversing image to text and then now it's text to image? Is that, I mean, that effectively was the diffusion process. I just see the career paths of the people that I talk to and see, and I see this overall trend of research directions. And I just wanted you guys to reflect on that. I mean, I certainly went that way, right? I started a long time ago doing computer vision, object detection, recognition, things like that. I think that's just a simpler problem, right? Generation is just harder.

13:21

It's a different kind of mapping, right? You map from the inverse mapping is not as simple as just inverting the kind of, right? It's more ambiguous, right? To go from cat to image of a cat. And in some ways it's also a loop because your vision work creates the synthetic labels that then continues. Because, I mean, sure. I don't know. I try to validate my theories about how fields develop, how careers progress through this. I mean, for the better the understanding side gets, we have seen that the generation side also gets better, right? So there's... It's completely bootstrapping. Yeah. It's amazing. So there's definitely a there there to that thesis.

14:19

And I think, yeah, a lot of people have, I definitely worked with a lot of image understanding people who became image generation people. And then some of them have moved on to video because it's the next thing where you have so many more dimensions to work with. So yeah, I'm curious about you specifically because you're... Yeah. So I definitely recommend to start with understanding, recognition, because that's basically discriminator. And then that's going to lead to better generation. And that's where the bridge is basically reinforcement learning. So my journey is I initially worked on the algorithmic research in the gentle model against some generation.

14:46

It's completely bootstrapping. Yeah. It's amazing. So there's definitely a there, there to that thesis. And I think, yeah, I think a lot of people have, I definitely worked with a lot of image understanding people who became image generation people. And then some of them have moved on to video because it's the next thing where you have so many more dimensions to work with. So yeah, I'm curious about you specifically because you're... Yeah. So I definitely recommend to start with understanding, recognition, because that's discriminator. And then that's going to lead to a better generation.

15:24

And that's where the bridge is reinforced learning. So my journey is I initially worked on the algorithmic research in the gentle model against some generation. And then I worked on RL and robotics. And then six years ago, I was leading a moonshot on the dexterity. It was pretty early, but I see now everyone's doing it. Four years ago, I figured out that symbolic AGI is going to make sure it much faster than the physical AGI counterpart. So I decided to language models and then those things. And then recently worked with do me and then Omni team, I quite enjoy collaboration there.

15:57

What I quite enjoy, what I recommend definitely to the researcher is to definitely explore or at least get exposure to what the top people in each of the community are looking at, how they think about problems. So when I look at the video model, to me, it reminds me pretty early on of language model where very early language model was that creative demo, right? You try to write a story, novel and then in GPT-2 and then those days, LSTM days, right? And then, instruction tuning, you actually make it usable as a chatbot, but then at the chatbot stage, it still had so much hallucinations and the instruction following wasn't good enough. So it couldn't use for reasoning.

16:27

And when it got good enough in pre-training and post-training for reasoning, then this test time skating, the RL really took off. That's many of the best performing models. And right now I think the video model is, as we mentioned, it is a complementary foundational model and I can imagine it's going to follow a similar path. It's going to improve a lot in instruction following, a lot of this. It's going to improve a lot in reducing coders nations to the extent that you become a very reliable world model. So you can intermix the video, space-time simulation with the text simulation to solve arbitrary AGI problems.

16:59

Also, I think the difference still is between text models and image video models is that we haven't quite unified understanding and generation in multimedia, I'd say, yet? I think without going through the details, of course there's, it depends on at which level you're thinking about this, but generally there's not that many, as far as I know, models, SOTA frontier models, that are genuinely good at both understanding and generation of let's say videos, right? It's an interesting challenge. No, I'm not saying that we should do this, but I think it stands to reason that understanding and generation are two sides of the same coin.

17:12

So they should be in the same model in some ways. But we don't necessarily always do that. Yeah. You mentioned audio as well, right? Yeah. Is that as hard as video or qualitatively different? If so, in what way? What are the interesting directions? Three years ago, was people using, I guess, diffusion to do audio as in the diffusion approach. I don't know if you guys saw that. And I just think it's very interesting if a modality that we perceive, which is audio, is different than video, actually to machines is exactly the same. They see no difference. I think on a technical level there are some differences, but I think they're relatively minor.

17:34

I think from my perspective, audio came into my life when we shipped a VO3, which was, I believe, the first model that did a joint. Yes. With the slicing of the. Yeah. Yeah. Slanted gold bars or whatever. It was the first model that is joint audiovisual generation. Yes. I mean, there were other models that did kind of agentic hacking under the hood, but this one was truly generating everything at once. And the reason we did that is because we felt, and I think it was the right choice, we felt that it only makes sense to generate them at the same time because they're from a machine learning perspective, it is one latent causal generative process, right?

18:02

There's something that generates you speaking. It's not the pixels and then the audio somehow generated by some other process, lips have to move in sync with the audio, right? So I think that solved a lot of the issues that previous models had, or the way that people did video generation before, where it was like, okay, we generate pixels and then we're going to hack something on top of it that moves the lips with the audio that we generate. That was very bad. And so I think that was, to me, that's the, I mean, after VO3, people were like, what do you mean? There's no audio in your model. Yeah. That makes no sense. Once it's there, you have to have it.

18:24

So I think that was the right choice. I think that's the difference in doing it to one single generative model, I think was the right choice. One thing I want to also ask you guys' opinions. Well, one difference I find in the audio and then against the image and video is the audio information is less verbalized. Of course the TTS and stuff is trivial, right? But when you get her outside, how to describe music, how do you describe this person's tone, pitch, I feel the verbalization is insufficient. And the interesting thing is that you see that in two other things like taste, taste sense, and also smell. And then another interesting thing is the skin color.

18:48

So skin color, the language is pretty limited to describe the skin color. And the reason is that we're extremely sensitive to the small difference, perturbations on the skin color, because that basically shows us, is this person going to kill me? Or can I befriend this person? This kind of information. And then I feel the smell, taste, skin color, and sound kind of stuff. It's very tied into primitive survival kind of stuff.

19:02

And so our sensory system is so sensitive that it's intractable to, so for example, ask one, the wine taster and then professional, and then he basically said he used language from dating, describing a partner as a way to describe the taste, because there's no sufficient vocab to describe. So I'm curious, yeah, do you guys feel that? I think, well to some extent, I think the same is true for visual information, right? When you think about a certain style or a certain aesthetic, right? There are some people who just have a much more developed, whether it's palette or visual taste and aesthetic, right?

19:18

I think language just tends to be a bit of a limiting factor when you are trying to describe any of these things that we experience with sensory information. And to your point earlier, I think that is the reason why we are investing in world models and why we are pushing on perception and the generation side of things. Because it is such a large part of how we as humans navigate the world. It's a large part of how embodied AI navigates the world. So I'm curious, yeah, do you guys feel that? I think, to some extent, I think the same is true for visual information, right? When you think about a certain style or a certain aesthetic, right?

19:38

There are some people who just have a much more developed, whether it's palette or visual taste and aesthetic, right? I think language just tends to be a bit of a limiting factor when you are trying to describe any of these things that we experience with sensory information. And to your point earlier, I think that is the reason why we are investing in world models and why we are pushing on perception and the generation side of things. Because it is such a large part of how we as humans navigate the world. It's a large part of how embodied AI navigates the world. And I do think language does have a lot of, it's gotten us very far and it can probably get us really far.

20:02

But it feels limiting in a lot of these areas. And yeah, I don't really know how to describe sense and taste. But yeah, I'm curious. I don't know that I have thought that deeply about this yet. So, yeah, I don't have a good answer about audio. I don't know the limit because I'm thinking about, well, what is Omni bad at in terms of audio? But they're all solvable problems, I find. With more data or better data or whatever it is. So, I don't know that we have pushed the frontier so much that we have hit some sort of limits that are rooted in evolutionary limits imposed by humans. I don't know. He's feeling the limits of captioning, which is the thing I was thinking.

20:35

Yeah, yeah, yeah. There's a lot of information in the world and it connects to why we do world modeling. Mm-hm. You just need S refs. S ref 15476, and then that's what majority does, right? I guess maybe, I can't describe this vibe, but. Well, I think that that's the point of providing some of these references, right? Yeah. Because even just describing how someone talks and their tone and then prosody and all of these things. I think some of these terms, I didn't used to know what they mean, right? Now I know. Prosody, yes. Disfluencies. Exactly, there's an entire vocabulary that even if you're not steeped in a domain. Which is true for actually most human domains.

21:06

You don't even know what it means. And sometimes it's also a question of if we haven't focused on those things with the large language models, then they may also have gaps in those areas, right? And then we feel them on the other side with generation because we're fundamentally relying on the language model's understanding of the world to then be able to represent it. So I think, yeah, it all goes back to your question about language as intermediary. But yeah, I think, to me, some of these might just be focus areas and things that we haven't necessarily pushed on as much as we can, and as we will, we will discover what the actual ceiling is.

21:21

Yeah, as a podcaster, I think a lot about sound. Mm-hmm. And I would just offer a couple of things for discussion in case it triggers anything with you guys. I have three domains of rough audio, which is music, voice, SFX, is that rough? Okay, covers everything. And then also, even within voice, let's just focus on voice, forget the other two.

21:34

From another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another minute to another

21:38

the need for world models, because you need it even for audio, about, well, I'm further away from you, so I should sound a little bit softer or more diffuse. And the video models need to pick that up, because if they're going to do immersive video and audio, you need that.

21:42

I love that example of studio quality or not. In a way, we don't have enough language to really describe this kind of echoing or some kind of noise happening. We just don't have precise enough. And basically, the reason that I think it's quite important to have relatively information-rich captioning is that we rely on natural language as a representation. But if you don't have enough representation, that means the conditional language, the generation is very multi-modal. And if you, anything you can learn from the VAE, very old GmbA research, the idea is we really want to capture most of the stochasticity in the later representation. And then the X given the Z should be deterministic.

21:45

So, yeah. Yeah. Yeah. Well, I hope there's more progress there. And I'm sure you guys are doing... I even actually do facial expressions, right? And maybe this gets to your point about things that we're very sensitive to, right? I think you can tell a lot of AI content also just from people's facial expressions. Yes. Yes. And we try not to contribute to it, but, you know. Or skin textures, right? The things that make things look real in real life. I can tell from the way you're nodding or from the way your micro-expressions are changing how you're reacting to what I'm saying. We haven't quite crossed that chasm, I think. We're so much better than we were a year ago. Yeah.

22:11

But there's so much more hedrium in a lot of those things that we as humans are super sensitive to. And I think image arguably probably is there because there's a lot of images that I will see that really do look indistinguishable from reality. And I can't tell if they're generated or not. They're better than reality. Or... Well, that's a different... No, I think that one of the... Better than what I would take on my vacation as a photo, yes. One of the fun experiments that we did a while ago on the team is, can we generate videos that are better than real videos, right? So you just take the same caption from some video and then... And recycle it, yeah.

22:42

Just try to describe a real video and then generate the equivalent version with Omni and then do a human eval. How does it do? And then humans largely prefer AI-generated videos. Oh, really? Yeah. But... Because it's the RL process. That's the RL process working. It's however you want to rationalize it. It's not necessarily the RL process. I think it's just, I'm not saying this is a good result. I'm just saying we have optimized in a way that potentially triggers something in the human brain that, oh, it looks... A lot of the AI videos just look better. Yeah, yeah, yeah.

23:24

On inspection, on deeper inspection, they would not actually be more useful or whatever. But if you just say side by side, random YouTube video versus generated version of it, it will just look better. Because it's sharper, more HDR. The skin tone is better. Again, it's not more realistic. It doesn't solve your problem necessarily, but it looks better. I think it also depends on the sensitivity of the people. I was born and raised in Japan, and I think one thing I know is they're extremely, extremely sensitive about, you know, that's why architecture, know, triggers something in the human brain that, oh, it looks... A lot of the AI videos just look better. I'm not...

23:51

Yeah, yeah, yeah. On inspection, on deeper inspection, they would not actually be more useful or whatever. But if you just say side by side, random YouTube video versus generated version of it, you will just have a... It will just look better. Because it's more... It's sharper, more HDR. The skin tone is better. Again, it's not more realistic. It doesn't solve your problem necessarily, but it looks better. I think it also depends on the sensitivity of the people. I was born and raised in Japan, and I think one thing I know is they're extremely, extremely sensitive about, that's why architecture, food and stuff, they have... Yeah.

24:45

So I talked to a manga artist there, and he's disgusted by the inter-generation AI. And one thing he mentions is the eye gaze. Eye gaze, that slight difference makes him feel creepy about it, unnatural. If you're looking a little bit off. Yeah, it's just... Yeah, it looks too fake. Yeah. So I think it does depend on the sensitivity and... Yeah, yeah, yeah. All I'm saying is human preferences are not particularly a reliable barometer of what you should be optimizing for. If you just ask people, do you like this or not, you're not necessarily getting what you wanted.

25:16

Yeah, let me just add one thing, but four years ago, there was a debate that prompt engineering is gonna disappear. And some very powerful people say it's gonna disappear, but I said it shouldn't. Because prompt engineering, specifying that is the only way you can control the output. So, when you have control over the AI. And what allows you to prompt engineers is really that sensitivity. So, sure, maybe right now, the AI can do a lot of auto-prompting in that, and it's gonna generate something that's sufficient. But if it's like that, never be satisfied. Never be satisfied with the AI's generated content. Always fine-tune your sensitivity and always keep prompting.

25:46

What are the differences? I think, to that extent, there's also a big difference between the average human untrained eye, which I would put myself in that bucket. I have some aesthetic sensibilities, and I've done this long enough, but I have a preference. But your example of a manga artist, that's somebody who has honed a craft over possibly many decades. And anybody who does that, whether it's design, architecture, right, you just have a very different level of expertise, and you see things that the average human will not see. But Dumi's right.

26:04

When we look at, if you were to just poll ten people on the street, they would probably prefer the overly smooth, very saturated kind of content. Yeah, it's called the Instagram filter. It is, it is, yeah. And so there's also a little bit of a question of what does your default aesthetic look like if you don't specify? But then to Shane's point, one of the things we always try to get these models better at is instruction follow. So that when you want to get them to a different outcome, you should be able to, whether that's through language or whether that's through references, because language is sometimes too limiting.

26:18

And so these models continue to get better at it, but they so much at work. Do you feel pressure as a product director to set the default for the world? I mean, kind of. Maybe I should. I don't know. I haven't thought about this. But it's like someone has to have a default. The default has to exist. Actually, I would say we have thought about this. And I think one of the, so for example, actually, if you look at Nano Banana generations, we had an explosion of Nano Banana infographics when Nano Banana Pro came out. I tried it, yeah. Yeah, yeah, yeah. I think Nureb's papers were all, so many had infographics generated.

26:56

Ooh, can you run your watermarking on it and see how many? We probably could. Yeah, Synthetic. We haven't done that, but I saw my Twitter was, maybe this is just also the bias of my algorithm. But they were everywhere. And it was actually very painful because I think our default aesthetic was a little bit too, it was too cluttered. I think that the model was a bit of an overeager student that just learned, it was like, oh, I know all this information about this concept. Let me shove it into the same image. Japanese infographics, 5x that. Or maybe it was, but it just, and... Wait, wait, so same prompt, same content. If it's in Japanese, it's more? Density, density.

27:55

Oh, wow. Because that's the style in Japan. Yeah, some very bureaucrat. There's a famous word for it. No, but we do go through this process with Omni. We did it together, right? Where we had a bunch of, at the very end, okay, this is, we did some tuning and, okay, what kind of style do we prefer? Yes. Is it more muted, more saturated? We had a lot of saturation. Yeah, there were, I think Nicole just has PTSD, so has forgotten about it, but she was very much involved in this, of okay, which kind of color palette do we basically prefer, right? And it's not something that, you have to make a trade-off there. And it's not, because it ends up being us, right?

28:55

Actually, it is true. It ends up being the modeling teams, and you could ask the question legitimately of are we the best people to do that, or should we actually work with someone who has a really creative point of view, and is more of an art director, and has, and we kind of go back and forth on this whenever we... I mean, you have the trusted testers, I'm on the... We do, we have trusted testers who give us a lot of feedback, and we take that seriously. Very well organized, by the way. They have these weekly calls and stuff, it's amazing. Logan's team does a lot of that, so... Kudos to Logan. Kudos to Logan, who couldn't be here today.

29:40

And we have a lot of people, actually, internally at Google, like Fulfer, who give us a ton of... No, no, no, truly, who give us a ton of feedback on when we release new checkpoints, and sometimes it will be stuff that we don't see, right? We would be like, oh, yeah, this optimization seems okay, and then they would come back and say, what have you done? You completely ruined my grass, because now the detail is all blurry. I think he just noticed, not a super secret at this point, but that our model tends to put rings, wedding rings on hands. That's, yeah. Very strange. I had never noticed that, but he's like, here, I just saw it, and there's a Fulfer channel, basically.

29:59

Yeah. Where he posts, I was like, why is there a wedding ring in every hand? I'm like, that's strange. That sounds like very common reward hacking. Yeah, yeah, yeah, yeah. But something that we would not have noticed necessarily while developing this, right? You know, are we all at a factor? I don't know. You do have a lot of preference base, and then you may prefer that. It's fierce correlation, reward hacking. It can happen in many weird ways. I think he just noticed, not a super secret at this point, but that our model tends to put rings, wedding rings, on hands. That's, yeah.

30:44

Very strange. I had never noticed that, but he's like, here, I just saw it, and there's a Fulfer channel. Yeah. Where he posts, I was like, why is there wedding ring in every hand? I'm like, that's strange. That sounds like very common reward hacking. Yeah, yeah, yeah, yeah. So, something that we would not have noticed necessarily while developing this, right? Are we all at a factor? I don't know. You do have a lot of preference base, and then you may prefer that.

31:28

It's fierce correlation, reward hacking. It can happen in many weird ways. It does, it does. This is related to another topic that, again, I try to use these main stage things as introductions or ties in. We have an evals track. We have Character AI and YouTube talking about how they evaluate videos. How do you evaluate videos? Hmm. Apart from playing Fulfer. Not everyone has a Fulfer. But also, I think there needs to be something more quantitative. Well, I mean, you improve Gemini to improve the evaluation for the future. Yeah. That's the answer. No, no, that's definitely one way. It's actually very hard. It's very hard.

32:35

It's very hard to get auto-readers to evaluate things in a video, including especially things like aesthetics, right? Yeah. There are some things that are a little bit more objective, especially when we talk, let's say we talk about images and we look at infographics, text rendering. That's actually fine, right? Because you can OCR things out, and then you can look at, okay, this letter is messed up, and then the whole thing is actually useless because literally if a letter is off in rendered text, you just can't use that asset, right?

32:51

So those things are a little bit more auto-readable from what we found. We do rely a lot on humans looking at things. And so we do do a lot of human evals. We do a lot of human evals. We do a lot of human evals. We do a lot of human evals. And every time... Shane is like... And every time we have a new model, we want to do more things, and we want to jam in more capabilities, and then we have more evals that we have to run. And then at some point, you do get two models that are kind of close to each other, and then we literally make decisions based on looking at outputs side by side.

33:16

Sometimes in a room, I've been in rooms where there's ten of us, and we're just looking at videos side by side, and we're like, do you prefer this or do you prefer that? Oh, wow. Yeah. I mean, but it is genuinely very complicated. The more capabilities you add, even just the one capability, but it's almost AGI-complete capabilities, like video, video editing, right? Think about video editing as a... and editing with audio and... My editor will be very happy to hear this. Yeah. I mean, it's just the hardest problem in Gen Media.

33:43

I mean, I don't know if it's the hardest, but it's definitely there, right? In terms of complexity of evaluation, free-form video editing is... you can do anything... Yes. And... I spent a lot of money on that, and it's very hard. Please help me. Adding those... we don't have add a sloth eval, right? Well, now we should. Now we should. Yeah, yeah, yeah. But things like that, it's... it's... it's not that easy to track. I think I'm just surprised at the sample size that you have, right? To test the entire surface of your models, you still rely on audio magnitude of hundreds. No, no, no, no, no. So, yeah, well, we do a ton of human evals on thousands of things.

34:06

I think there's also an element of, we can talk about things like live experiments, right? Which is also where you get signal on... some of these more minute differences at much larger scale. Then there's auto-readers, which is definitely kind of a more... It's a very well-defined space, I think, for LLMs. Much more nascent for media models. And then sometimes you still do rely on human judgment. And we do rely on things like feedback from people who just have a very honed aesthetic, and people who just use these models in their workflows day to day, right?

34:14

Because we could also... you could have a model that is really well on some slice of human evals, but then it really breaks the workflow for somebody. And so this is why we do early access programs, and we try to get feedback, and then we try to incorporate it before we release something more broadly. I feel like Shane had a hot take based on his... He always does. I feel like I'm a good expression. Always. When we were talking about his... Every kind of human work should be gradually amortized. And then the interesting thing is the video understanding, especially against AI-generated video, like detecting AI stuff, is an extremely interesting vision task.

34:26

And then some of it is aesthetics or this kind of visual quality, but for some of the kind of cases, semantically it doesn't make sense. For example, you're taking a famous scene from a movie and trying to construct that. And then if you generate it, it can generate something there, but at some point, some of the semantic information doesn't make sense. It's actually inconsistent. So can the AI actually detect that? So when I evaluate the AI video, I was like, oh, I feel I am so smart. AI is still kind of behind. But we should make a lot of effort. I think video understanding is an extremely important intelligence task beyond just the pure aesthetics or the preference.

34:38

And, yeah, we should always try to amortize the human label. Yeah. What data do you need? A lot of people I talk to want to get in front of you, actually. They want to be nice about it. They have a lot of video data. They have gaming data. They have real-world video data. They have images. They have labelers. What do you want? Are you offering? I'm just like, this is your request for, okay, okay, we got... I'm sure you get a lot of pitches, right? You got a lot of people who want to talk to you. What's... I think, actually, it's the signal, it is probably sorting out signal from noise is the main problem.

34:58

So creating a nice API of, okay, if you actually do A, B, and C, we are interested in that. Loaded question there. So, I don't know that there's an easy, if you do... I think we do already have a lot of data. I think it's hard to talk about this. Yeah, you don't want to talk about it in public. I don't want to get you in trouble. Yeah. But I think. No, no, I just want to say it's hard to talk about this without trying to... I have to think about what I am revealing about our project and where we're going. Yeah. Generally high-quality data, I think, maybe let's just put it this way, right? It's not the secret. Embodied? I'm sorry? Embodied data? I mean, I think.

35:38

Yeah, sure. I mean, we have announced, I think, publicly, right, that we have some sort of robotics collaboration, right? Or because we have a robotics team at GDM, so they're always interested in things like that. Yeah. But I think. No, no, I just want to say it's hard to talk about this without trying to, without, I have to think about what I am revealing about our project and where we're going. Yeah. Generally high quality data, I think, maybe, maybe let's just put it this way, right? It's not the secret. Embodied? I'm sorry? Embodied data? I think.

36:14

Yeah, sure. We have announced, I think, publicly, right, that we have some robotics collaboration, right? So I think it's, or, but, because we have a robotics team at GDM, so they're always interested in things like that. I mean, for Omni specifically, I think we're just quite interested in high quality data, right? It's not necessarily, oh, random YouTube video, but some more professional shop, things like that, right? Things like that, those are things that we're always on the lookout for. And, yeah. And I think for, maybe this is easier to some extent to answer for some of the agentic work as well. Actual, what are the tasks that people are trying to do, right?

36:36

These things are actually difficult to manufacture if you're doing it yourself or if you're doing it with a vendor. What is the actual, if you're creating a marketing campaign, what does that look like, right? Do you start from, here's a picture of my new product and then I want to turn that into a video ad and I want to turn that into a bunch of assets that fit all these different ad formats that I need to push onto the various platforms to promote. And then you kind of go from this to that and what is that kind of trajectory of tasks that you're experiencing along the way?

36:46

That is really useful and that is actually difficult to get, right? Because we don't always have the right first party surface where people are actually doing some of these things. Or you might work with someone who's a vendor but they also don't have that product surface, right? A lot of this kind of information lives in the places where people are doing these tasks and so that's difficult to get. If anyone's figured that out, you should reach out to us. Every channel of thoughts, yeah. Every channel of thoughts. Every channel of thoughts. Yeah. And maybe the data that Chinese lab is using. Yes. Yeah, yeah, yeah. Yeah.

37:50

As a media person myself, right, there's so many podcasters and people in marketing departments and all these, they would be happy to be your data. Just put a BCI on my head. And talk to us. And watch the things. Because there's just an endless amount of work to do. There's so much work. And this is all, this needs to somewhat be a commodity. Obviously, you can be an art, an artism. You can be Hollywood for the really high quality stuff. But actually a lot of work is commodity and should be modelable. And we want you to do it. And we want the high quality, to Dumi's point, right? We do want the high quality. We want commodity, yeah, yes, yes. You want it on both sides.

38:38

I just like the folks. Thank you for the solicitation. We also, I also added a data quality track. I think that people want to understand what at AI, how to raise the bar. Right? And a lot of it is just educating the market and educating researchers and engineers and founders on this is where we're going. And I think a lot of this is slop. Stop doing that. Do this instead. And people will listen. Yeah. I don't know. To that extent, I mean... But I think to that point, there's a lot of, again, craft that goes into this, right? And there's a lot of process. Even to the marketing campaign example, you don't create that in five minutes, right?

39:27

You go through a process and you iterate and you pick something over something else because you liked it for whatever reason. Maybe the eye gaze was correct, right? We don't know these things, right? Yeah. Because none of us are marketing directors and the models don't know these things. I even say this for the natural language as well. I always say 99% of information is inside people. You can only extract it through active dialogue and befriending them. So most of the stuff on the internet is the outcome, the output of that. Yes. But what are all the trajectories? How did this person have this inspiration to write this paper? Mm-hm. What is the starting point?

40:20

What is the inspiration? What is the dialogue that sparked it? What is the dialogue that was created? Yeah. It's like when you write a novel, right? A novel speaks to you because usually there's some sort of personal connection that you feel to the story or the trajectory or the characters. If you read most of the stuff that's written by LLMs today, it falls into these default patterns and the language starts to feel really similar and all the descriptions sound really similar. You can quickly read it as, oh, this is not that interesting because I can't connect to it, right? And, again, that's human expertise.

40:48

One nice thing recently is that Google Cloud and Google DeepMind are starting to invest a lot more in the FDEs for the product engineers. And I also saw some recruiting for the creative GenMedia space as well. So I think those are really the effort because we feel what we can do with a lot of public data has limits. But really, partnering with that, we can provide better models and products and we can feedback. We have an FDE track here for the first time. Every lab is announcing it. It's crazy. Yeah. One thing I'm actually very keen on doing, and I push for this at Cognition as well, is to turn the FDEs not just into sales and solutions but also to evals workers.

41:20

FDEs is not the sales. FDEs is way, way bigger than that. How do you frame FDEs then? Because I do think about it as sales. You're, the more, you customize the solution. So I define post-training as anything between the pre-training and the final user experience. Anything. Anything is a post-training. And to me, when I first learned a lot about, I mean, FD kind of originally came from here and then that. So I guess the history is different. But, yeah, I think the key is really that the key is not only to work with them and ensure that they know how to use, but also to code, derive insights that can help both parties. They can put a lot of harness how they use the model.

42:04

We can improve very upstream. So how to get the customer feedback to the modeling, I feel, is more the role I want for the FDs. Yeah. And even if I start, just on that, if you want to talk to us, or at least me, I'm not going to offer up your time. But it's really helpful for us to actually talk to people who are using our models and understand where they're struggling. Because, again, they're just, it's the real world task that you're actually trying to use them for, right? I will talk to people who do interior design with some of our image models, and they will say, hey, I really want to take this pattern, but then I want to scale it across ten different ruck sizes.

42:50

We can improve very upstream. So how to get the customer feedback to the modeling, I feel, is more the role I want for the FDs. Yeah. And even if I start, just on that, if you want to talk to us, or at least me, I'm not going to offer up your time. But it's really helpful for us to actually talk to people who are using our models and understand where they're struggling. Because, again, it's the real-world task that you're actually trying to use them for, right? I will talk to people who do interior design with some of our image models, and they will say, hey, I really want to take this pattern, but then I want to scale it across ten different ruck sizes.

43:25

And sometimes I have a very custom ruck size, and then the model fails at replicating the pattern the same way. Or I want to do a try-on for these earrings, and then the earrings have a certain size, and then my head has a certain size. It has to make sense if you're actually trying to try things on, and the models fail at a bunch of these things that actually happen in the real world, right? And so that's useful for us, because for some of these things, we don't think about them, because we don't use the models for those tasks.

43:43

Or I think, to your point about ad campaigns or whatever, people have notions of brand languages or whatever, which is a bunch of images or PDFs saying things. It's a pretty ambiguous question as well. What is the IKEA brand language? Is it blue and yellow? That's not a very... But what shade of blue? Yeah, yeah, yeah. So there are, you know, and the brands are pretty specific, pretty, you know, they do care about the shade of blue. It shouldn't just be a random blue and a random yellow. That's not going to be IKEA, right? I'm just thinking about an example. But this is the kind of stuff that it's not necessarily part of our developing frontier models mandate.

44:41

But it's something that we do want to fundamentally build products that people will use to solve concrete tasks, not just research artifacts, right? So I think it's useful to understand what people do care about. Well, I'm sure a lot of people are very grateful for your work. And there's a lot more to do. You've made so much progress over the last even just a couple years of Nano Banana and Bio and Omni. And I don't know what else you got cooking, but we're very excited. This is one of those things where I was very disappointed when Sora shut down. And I think there needs to be more general exploration of generative models and not just coding. I think that is...

45:21

We obviously like this thing. We love coding. We love coding. And, yes. But thank you so much for your time. It's been a real pleasure. And I can't wait to see what this looks like next year. Thank you for having us. Great question. Thank you, everyone.

46:33

They, I mean, they want to be nice about it. They have a lot of video data. They have gaming data. They have real-world video data. They have images. They have labelers. What do you want? Are you, like, offering? I'm just like, this is your request for, like, okay, okay, we got, I'm sure you get a lot of pitches, right? You got a lot of people who want to talk to you. What's, like, I think, actually, it's the signal, it is probably, sorting out signal from noise is the main problem. So creating a nice API of, like, okay, if you actually do A, B, and C, we are interested in that.

47:10

Um, loaded question there. So, uh, I don't know that there's, like, an easy, like, you know, if you do. I think we do already have a lot of data. I think it's hard to talk about this, you know. Yeah, you don't want to talk about it in public. I don't want to get you in trouble. Yeah. But, like, I think. No, no, I just want to say it's, like, hard to talk about this in a sort of, you know, without trying to, without, I have to think about the, what I am revealing about our project and where we're going. Yeah. Um, generally high quality data, I think, maybe, maybe let's just put it this way, right? It's not the secret. Embodied? I'm sorry? Embodied data? I mean, I think.

47:44

Yeah, sure. I mean, we have sort of announced, I think, publicly, right, that we have some sort of robotics collaboration, right? Like, so I think it's, like, or, but, because we have a robotics team at GDM, so, you know, they're always interested in things like that. Um, I mean, for Omni specifically, I think we're just quite interested in just high quality data, right? Like, you know, it's not sort of, not necessarily, like, oh, random YouTube video, but, like, you know, some more professional shop, things like that, right? Like, things like that, those are, those are things that we're always on the lookout for. Like, uh, and, yeah.

48:18

And I think for, you know, maybe this is easier to some extent to answer for, like, some of the agentic work as well. Like, like, like, actual kind of, like, what are the tasks that people are trying to do, right? These things are actually kind of difficult to manufacture if you're doing it yourself or if you're, like, doing it with a vendor. Like, what is the actual, like, if you're creating a marketing campaign, like, what does that look like, right?

48:41

Like, do you start from, here's, like, a picture of my new product and then I want to turn that into a video ad and I want to turn that into a bunch of assets that, like, fit all these different ad formats that I need to push onto the various platforms to promote. And then, like, so you kind of go from this to that and, like, what is that kind of trajectory of tasks that you're, like, you know, experiencing along the way? Like, that is really useful and that is actually kind of difficult to get, right? Because, like, we don't always have the right first party surface where people are actually doing some of these things.

49:15

Or, like, you might work with someone who's a vendor but they don't, also don't have that product surface, right? Like, like, a lot of this kind of information lives in the places where people are doing these tasks and so that's kind of difficult to get. Like, if anyone's figured that out, you should reach out to us. Every channel of thoughts, yeah. Every channel of thoughts. Every channel of thoughts. Yeah. And maybe the data that Chinese lab is using. Yes. Yeah, yeah, yeah. You know, yeah. As a media person myself, right, like, there's so many podcasters and people in marketing departments and all these, like, they would be happy to be your data.

49:49

You know, like, you know, just, like, put a BCI on my head. And talk to us. And watch the things. Because, you know, there's just an endless amount of work to do. Like, there's so much work. And this is all, like, this needs to somewhat be a commodity. Like, obviously, you can be an art, like, an artism. Like, you can be Hollywood for, like, the really high quality stuff. But actually a lot of work is commodity and, like, should be modelable. And we want you to do it. And we want the high quality, like, to Dumi's point, right? Like, we do want, we want the high quality. We want commodity, yeah, yes, yes. You want it on both sides. I just like the folks.

50:24

Thank you for the solicitation. You know, we also, I also added a data quality track. I think that people want to understand, like, what at AI, like, how to raise the bar. Right? Like, and a lot of it is just educating the market and educating researchers and engineers and founders on, like, this is where we're going. And I think a lot of this is slop. Stop doing that. Do this instead. And, like, people will listen. Yeah. I don't know. To that extent, you know, I mean... But I think to that point, like, there's a lot of, again, just, like, craft that goes into this, right? And there's a lot of process.

51:01

Like, even to the marketing campaign example, you don't create that in, like, five minutes, right? You, like, go through a process and you iterate and you, like, pick something over something else because you liked it for whatever reason. Like, maybe the eye gaze was correct, right? Like, we just, we don't know these things, right? Yeah. Because none of us are marketing directors and, like, the models don't know these things. I even kind of say this for the natural, like, language as well. Like, I always kind of say 99% of information is inside people. You can only extract it through active dialogue and befriending them.

51:31

So, most of the stuff on the internet is, like, sort of the outcome, the output of that. Yes. But, you know, what are all the trajectories? You know, how did this person have this inspiration to write this paper? Mm-hm. What is the starting point? What is the inspiration? What is the dialogue that sparked it?

51:47

What is the dialogue that was created? Yeah. It's like when you write a novel, right? Like, a novel speaks to you because, like, usually there's some sort of, like, a personal connection that you feel to, like, the story or the trajectory or the characters. Like, if you read most of the stuff that's written by LLMs today, like, it's, you know, it falls into these, like, default patterns and, like, the language starts to feel really similar and all the descriptions sound really similar. You can kind of, like, quickly read it as, like, oh, this is not that interesting because, like, I can't connect to it, right? And, again, that's kind of, like, a human expertise.

52:26

One nice thing recently is that Google Cloud and Google DeepMind are kind of starting to invest a lot more in the FDEs for the product engineers. And I also kind of saw some recruiting for the creative, you know, GenMedia kind of space as well. So I think those are kind of really the effort because we kind of feel, you know, what we can kind of do with a lot of public data is limits. But really, you know, partnering with that, we can provide kind of better models and products and we can feedback. We have an FDE track here for the first time. Every lab is announcing it. It's crazy. Yeah. One thing I'm actually very keen on doing, and I push for this at Cognition as well,

53:01

is to turn the FDEs not just into sales and solutions but also to evals workers. FDEs is not the sales. FDEs is way, way bigger than that. How do you frame FDEs then? Because I do think about it as sales. Like, you're, you know, the more, like, you customize the solution. So I define post-training as anything between the pre-training and the final user experience. Anything. Anything is a post-training. And to me, when I first sort of, you know, learned a lot about, I mean, FD kind of, I guess, originally, you know, came from, like, here and then that. So I guess the kind of history is different.

53:37

But, yeah, I think the key is really that, you know, the key is, like, not only to kind of work with them and ensure that they kind of know how to use, but also to sort of code, like, derive kind of insights that can basically kind of help both parties. They can put, like, a lot of harness how they use the model. We can improve, like, very upstream. So how to get the customer feedback to the modeling, I feel, is the kind of more the role I kind of want for the FDs. Yeah. And even if I start, just on that, like, if you want to talk to us, or at least me, I'm not going to offer up your time.

54:11

But I, it's really helpful for us to actually talk to people who are using our models and, like, understand where they're struggling. Because, again, they're just, like, it's the real world task that you're actually trying to use them for, right? Like, I will talk to people who do kind of interior design with some of our image models, you know, and they will say, hey, like, I really want to take this pattern, but then I want to scale it across, like, ten different ruck sizes. And sometimes I have, like, a very custom ruck size, and then the model fails at, like, replicating the pattern the same way.

54:43

Or, you know, I want to do a try-on for these earrings, and then the earrings have a certain size, and then, like, my head has a certain size. Like, it has to make sense if you're actually trying to try things on, and, like, the models kind of fail at a bunch of these things that, like, actually happen in the real world, right? And so that's, like, useful for us, because for some of these things, like, we don't think about, because we don't, you know, we don't use the models for those tasks.

55:06

Or, like, you know, I think to your point about ad campaigns or whatever, like, people have, like, notions of brand languages or whatever, like, which is like a bunch of images or PDFs saying things. You know, it's a pretty kind of, you know, ambiguous question as well. What is the IKEA brand language, you know? Is it blue and yellow? I mean, that's not a very, like... But, like, what shade of blue, you know? Yeah, yeah, yeah. So there's, like, you know, and the brands are pretty specific, you know, pretty, you know, like, they do care about the shade of blue. It shouldn't just be a random blue and a random yellow. That's not going to be IKEA, right?

55:36

I'm just thinking about an example. But, like, this is the kind of stuff that, you know, it's not necessarily part of our, like, you know, developing frontier models kind of, you know, unnecessarily mandate. But it's something that we do want to fundamentally, like, build products that people will use to solve concrete tasks, not just research artifacts, right? So I think it's useful to understand what people do care about. Well, I'm sure a lot of people are very grateful for your work. And there's a lot more to do. You've made so much progress over the last, like, even just a couple years of, like, Nano Banana and Bio and Omni.

56:08

And I don't know what else you got cooking, but we're very excited. Like, this is one of those things where, like, I was very disappointed, you know, when Sora shut down. And I think, like, there needs to be more general exploration of, you know, generative models and not just, you know, coding. I think that is... We obviously like this thing. We love coding. We love coding. And, yes. But thank you so much for your time. It's been a real pleasure. And I can't wait to see what this looks like next year. Thank you for having us. Great question. Thank you, everyone.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note