Open Reader

Let's go Bananas with GenMedia — Guillaume Vernade, Google DeepMind

completed 1:17:14 May 18, 2026 Watch on YouTube

Current Status

completed

Video ID

BcWFc3H7Khg

RAG / Chat

Enabled
Let's go Bananas with GenMedia — Guillaume Vernade, Google DeepMind
Description

Guillaume Vernade from Google DeepMind takes a public domain book and runs it through the full gen media stack live. Gemini reads the whole text and writes image prompts for each character and chapter. Imagen generates the portraits. Veo animates them into video clips using those images as first frames. Lyria composes a different piece of music per chapter, with or without lyrics. The TTS model reads dialogue from the book using a trick that makes two voices sound like four distinct characters. The interesting layer underneath all of it is that Gemini acts as the prompt engineer for every other model, and it works well partly because the gen media models were trained on prompts written by Gemini. The workshop also covers the Lyria Realtime model, which generates music continuously and responds to new prompts mid-stream like a DJ, and a new interactions API that makes chained multi-turn calls cheaper by caching context server-side instead of resending the full book on every turn. Speaker info: - https://x.com/Giom_V - https://www.linkedin.com/in/guillaumevernade - https://github.com/Giom-V

Summary

Generated by claude-haiku-4-5-20251001

Summary: Let's Go Bananas with GenMedia — Guillaume Vernade, Google DeepMind

Main Topics

  • Google DeepMind's GenMedia Suite: Overview of image, video, and music generation models
  • World Models Vision: Building multimodal models that understand and generate multiple types of media
  • Developer Advocacy: Ensuring real-world usability of AI products
  • Practical Workshop: Live demonstration of creating illustrated content from a book using multiple GenMedia models
  • API Infrastructure: Differences between consumer apps, Vertex AI, and the Gemini Developer API
  • New Release Cadence: DeepMind shipping updates on average every 5 days across all products

Key Points

Products & Models

  • Nano Banana 2: Image generation with improved aspect ratios (520px to 4K), search grounding, and image grounding capabilities
  • Veo 3.1 Lite: Cheapest video generation model ($0.05/second, $0.40 per video) for iterative prompt development
  • Lyria: Music generation model with two variants:
  • Lyria Clip: 30-second music clips ($0.04)
  • Lyria Full Song: Up to 3-minute songs ($0.08)
  • Lyria Real Time: Live music generation that responds to prompts like a DJ, creating music in real-time with 2-second response times
  • Gemini Suite: Multimodal models (Gemini 1.5, 2.0, 3.x) now with full vision capabilities
  • Open Models: Gemma 4 (released last week) providing open-source alternatives

Technical Insights

Structured Outputs: Using JSON schemas to get consistent, predictable model outputs for seamless integration

File Upload API: Simplifies the complexity of Vertex AI by automatically handling file storage and access, allowing easy context windows for large documents

Chat Mode with Context: Maintaining conversation history enables consistency in image generation across multiple requests, useful for maintaining character styles and visual continuity

Auto-Retry Mechanisms: Essential for handling model overload, especially during peak US hours

Service Tier Pricing:

  • Standard: Normal pricing, normal queue
  • Flex: 50% discount but delayed responses (up to minutes)
  • Priority: 2x cost but guaranteed fast responses (like airport fast-track)

Interactions API (Preview): New stateful API that:

  • Returns interaction IDs for reusing context without re-uploading
  • Enables conversation forking (e.g., create lyrics, then split into image and music generation)
  • Automatic prompt caching for cost reduction
  • Session duration: ~2 days

Practical Applications Demonstrated

Book Illustration Workflow:

  • Upload full book text using File Upload API
  • Use Gemini to generate character descriptions and scene prompts
  • Generate images for characters with consistent styling
  • Create chapter-specific prompts using character references
  • Generate videos using images as first frames
  • Create musical accompaniment
  • Add voice narration with character differentiation

Voice Differentiation Trick: Use TTS with same voice but different textual style directions (e.g., "breathless and humble stutter" vs. "long poetic pauses") to create perception of multiple characters

Lyria Advanced Features:

  • Everything controlled via prompts (no separate parameters)
  • Timing controls: "first 30 seconds this, then switch to that"
  • Structural guidance: intro/verse/chorus/outro specifications
  • BPM and scale specifications
  • Multiple languages in single song
  • Lyrics with timestamp data for karaoke applications
  • Image-to-music generation using images as inspiration
  • Video game applications: Real-time music that responds to gameplay events

European Regional Limitations

  • Preview models only available on global endpoints (unlikely to change)
  • Goal is faster movement from preview to general availability (GA)
  • Rapid back-to-back releases (e.g., Gemini 3 → 3.1) reset preview counters
  • Acknowledged as "P0" priority to improve for European developers
  • Nano Banana 1 remains available with free tier for EU users

Notable Quotes

> "My job is to make sure that whenever we release things, you guys, the developers, have everything you need to work with our products."

> "A normal developer should be able to just swap the model name and it works."

> "This is the year where we are actually going to build agents" (compared to previous year of just talking about them)

> "Our vision of a world model is that it understands the world—ingesting as many modalities as possible (sound, videos, audio, sensors) and outputting in as many different modalities."

> "Everything is in the prompt" (regarding Lyria's parameter-free design)

Takeaways

  • Start Simple, Iterate: Use cheaper models (Veo 3.1 Lite) to iterate on prompts before upscaling to premium models
  • Leverage Gemini for Prompting: Use Gemini to generate prompts for GenMedia models—they're trained together and work synergistically
  • Use Structured Outputs: Define clear JSON schemas for consistent, predictable model outputs
  • Maintain Context Efficiently: Use chat mode or the new Interactions API to preserve consistency across multi-step generation workflows
  • Combine Modalities Creatively: Mix images, video, music, and voice to create rich multimedia experiences
  • Explore Lyria Real-Time: This underutilized model has potential applications in games, interactive experiences, and live content creation
  • Access Resources: Check the GitHub cookbook (goo.gle/cookbook-illustration) for quick starts and full-fledged examples
  • Cost Management: Monitor video generation costs; use priority tier only when necessary; consider Flex tier for non-time-sensitive tasks
  • Account for Model Limitations: Content filters may reject graphic content; Europe has data residency restrictions on preview models
  • Plan for Scale: When building production systems, improve upon demo approaches—async calls, better data persistence, optimized character referencing

Transcript

10143 words en Processed in 543.7s

[SPEAKER_02] Good morning, everyone. Thank you for being here so early in the morning, and to be able to be a few ones who pass security to actually be here. I'm Guillaume, and I will tell you about Gen Media in general. So this is my life as Nano Banana sees it. I joined Google six years ago. I used to be a video game producer before. I initially worked on Stadia, the streaming game company, that product that we killed, as so many other products. And I've been at DeepMind for two years now. And I worked as what we called a developer advocate. So if you're not familiar with what a developer advocate is, my job is to make sure that whenever we release things, or whatever we release, you guys, the developers, have everything you need to work with our products. So you need documentation. You need code samples. You need demos. Now, there are new things that you need as well, such as skills and prompt guides and things like that. So making sure that you can start right away and it works. And on the other side, when I'm talking internally, that's the reason why it's advocate. Because I'm advocating for the developers and trying to bring some common sense to the internal teams and making sure that what we release makes sense in the real world and it's not entirely developed for Google by Google. A very good example of that is the Imagine models. When we had Imagine and Nano Banana, each model has its own set of API. It doesn't make any sense. A normal developer should be able to just swap the model name and it works. And I've fought quite a long time for that. I never managed to win. But in the end, I think the Imagine brand doesn't exist anymore. So I win by default. But you see the kind of work I have to do on a daily basis. As I said, I'm mostly working on the Gen Media model. So media is everywhere. If you look at the world, your phone, and anything, media is everywhere. So you have images, videos, songs everywhere. So it's really cool in our world. And that was at the core of what DeepMind is building with our models. You're starting to hear with LeCain's new startup about world model. But that's exactly what we have been trying to build at Google since the beginning. And our vision of a world model is that it's a model that, as its name implies, understands the world. But meaning, it can ingest as many modalities as possible. So sound, videos, audio, sensors, whatever, all five senses. And then outputs or talks in as many different modalities as possible, so audio, text, and so on, and much more in the future. And we tend to have specific models. We have all image generation models. We have all video generation models. But deep down, the goal is really to have one model that encompasses all of that. It's just that for release purposes, it's easier to ship specific models and to always update the main model and have risks of breaking something else at the same time. But a quick story. When we released Gemini 1.0, it was two years ago. It feels like it was years ago, aeons ago. But the first Gemini model, Gemini 1.1, was meant to be multimodal. Because all of our models have always been multimodal. But since the testing was not finished, or whatever reason, I wasn't there yet, they removed the image understanding, the multimodal understanding inputs from the model. So the 1.1 was not multimodal. And then 1.5 came. And then this one was multimodal in. And I think that was the first one that was doing that. Which, once again, is crazy. It was a year and a half ago. And now you don't imagine working with models that are not multimodal. But still, when you were using that one, very often, you were giving it an image. And it was telling you, oh, I'm sorry. I'm just an LLM. I can't deal with images. And that was just because some of the training that was added at the end of 1.0 was, you know how to deal with image. But if you are asked, don't use it. Though some of that training was still remaining in 1.5 and coming up from time to time, so that was annoying until we switched to 2.0. So as I said, we don't only have the Gemini models in DeepMind. So we have all of the image, video, music generations. We have a couple of other very specific models. We have the robotics one that is multimodal, because vision is very important for robotics. We have a bunch of agents that we are shipping. I think this year is going to be the year of agents, a real one. Last year was everybody talking about agents. This year is the year where we are actually going to build agents. And we have the open models. And I forgot to update the slide, because it's not Gemma 3 anymore. It's Gemma 4 since last week. And then a bunch of research models, like alpha-evolve, alpha-genome, weathernext, and so on that are very specific for research purposes. Just on GenMedia, we ship things on average more than every month. Overall, if you just take all of DeepMind, we are shipping things every five days on average. Sometimes some weeks we're shipping two or three things. And if we had all of the small features, we are shipping multiple times per week. So that's why I and most of my colleagues are very busy. It's Gemma 4 since last week. And then a bunch of research models, like alpha-evolve, alpha-genome, weathernext, and so on that are very specific for research purposes. Just on GenMedia, we ship things on average more than every month. On the wall, if you just take all of DeepMind, we are shipping things every five days on average. Sometimes some weeks we're shipping two or three things. And if we had all of the small features, we are shipping multiple times per week. So that's why I and most of my colleagues are very busy, because we always have new things to document and talk about. So very quickly about the new things that we shipped recently. NanoBanana, we shipped NanoBanana 2, which has new spectra ratios, from 520 pixels to 4K. It has search grounding, as you all know. But it also has image grounding. So you can ask it to search for images on the web and use those images as grounding so that it knows. It has better knowledge about what things look like. It's very useful for architectural stuff, for example, for animals and things. There's a lot of rules internal that I was not able to all figure out. But for example, buildings, they have to be old enough. Otherwise, there are legal reasons we can't use images. But it's still very useful when it works. We have the VO models. You have heard about VO3, VO3.1. We just released last week VO3.1 Lite, which is our cheapest model. I think it's $0.05 per second, so $0.40 for one video, which is very cheap. And the idea is that you can iterate on the prompt that way, and then upscale afterwards. And we have Lyria, which is our music generation model that we released two weeks ago, with which you can create either 30-second clips or full songs of three minutes. I will show you demos of all of that afterwards. That's the point of the session, anyway. But there's also another Lyria model that people don't know about that is called Lyria real time. And that's actually my favorite model. And this one is you create music as well. But it's not a diffusion model like users, where you just give a prompt and you get something out of it. It's a predict model, which means it's a live model. So it creates music. And it continues to create music in real time until you stop it. And you can just send prompts. And it will, as a DJ, mix and swap to the new music in real time. And real time, two seconds. And that's quite fun to play with. If we have time at the end, I will show you a demo. But as I said, this is meant to be a workshop. So the goal is for you to play with the models and for me to show you codes instead of slides. So this is the content that we are going to use. That's this one as well, if you can read what I wrote. And if that doesn't work, just tell me. Are you all in? It doesn't work? Oops. Let me try to check that the link works. [SPEAKER_05] Yeah. No. [SPEAKER_04] I have the title. Cool. Are you all in? OK. So full disclosure before we start. This is using Gen Media models, which means they are all paid models. So running the notebook is going to cost you something $1. And you can just keep the video generation, because that's most of the cost of that $1. So the goal of that content I prepared is to illustrate a book, using the different Gen Media models. So what we are going to do is that we are going to take a book that is an open source one that I took on the Gutenberg online library. And we are going to use Gemini to come up with prompts, and then the Gen Media to create the content for the prompts so that we will have images of the characters, images of the scenes, videos of those things, and so on. And that example comes from what we call the cookbook. So we have this GitHub repo where we are posting examples of both quick start guides to explain how to use a new model, or how to use a new feature. And also examples, full-fledged examples, such as this one on how to go further and to mix different features into content. So if you are looking for ideas, that's a good place to check what you can do with the models. So introduction, billing. So let's start. So the first thing is that you need to install the SDK. OK, you need the latest one because of the music generation that was shipped last week, two weeks before. So you need the next one. So it's starting. And I should have done that. And you will need an API key if you don't have an API key from AI Studio. And you need the other side, you need that API key to be a paid one. If you don't have a paid API key, you can still use the image generation examples by using the Nano Banana 1 model, which has a free tier. So you can just swap the model we're going to use for this one. And why is that so slow? I think it's good. OK. So then I'm just loading the API key. And I'm creating the client, the GenAI client. And you can see there, I added these parts that we don't add in all of our examples. That is the auto-retry thing. Because if we are using Nano Banana 2 at the moment, especially in the evening when the US wakes up, the model can be overloaded. So that part says that it's going to be automatically retrying five times after two seconds. A bunch of imports. I think it's good. OK. So then I'm just loading the API key. And I'm creating the client, the GenAI client. And you can see there, I added these parts that we don't add in all of our examples. That is the auto-retry thing. Because if we are using Nano Banana 2 at the moment, especially in the evening when the US wakes up, the model can be overloaded. So that part says that it's going to be automatically retrying five times after two seconds. A bunch of imports. And then we are going to select all of the models that we are going to use. So in this case, 3.1 flash image preview, which is Nano Banana 2. [SPEAKER_02] Gemini 2.5 flash. Let's use 3 flash instead. Lyria clip to create 30-second music clips. And the TTS model, this one, the pro one. [SPEAKER_02] So I added this checkbox that should have been false by default, but I made a mistake yesterday evening. [SPEAKER_02] Just so that you should not be able to run the notebook by mistake if you don't want to pay. [SPEAKER_02] Then I'm just setting limits here, because a book can have a lot of chapters, lots of characters, and that can be quite long to generate. [SPEAKER_02] So that's why I set up limits on how many we want for the demo purposes and ghost purposes as well. So as I said, we are going to use an open source book. That's The Wind in the Willows from Kenneth Graham that I don't remember reading, but I think in the UK it's quite well known. So I'm just downloading it from the Gutenberg project. And here, I'm using client file upload, which is, you might know that we have two ways of using the Gemini models. We have the AI Studio Gemini API way, and you have Vertex. And the main difference between the two, and that's a good time to switch to this slide. But at Google, we like to create multiple products that are doing the same thing and confuse our users. That's our motto. So we're doing the same with AI. So we have a bunch of, but actually, when you think about it, it can make more sense. So on the left, we have what we call the consumers app. So it's a fully developed app that is easy to use for anybody who is not technical. So you can do plenty of things with Gemini, but you're lacking, as a developer, you might be frustrated because you're lacking control about which models, which features, which parameters are being used. And on the other end, we have Vertex AI, which is the exact opposite. It's meant for enterprise, so you have a lot of control. You can control on which data center it runs. You have control about your buckets, who has access to what, and so on. The only thing is that it comes with great powers, great responsibilities. So it can be a pain for people to start there. So that's why we have the Gemini developer API that is a middle ground for developers, where you can just create an API key and then start using the models right away, which comes with security risks as well. That's if your key leaks, anybody can just use it. And AI Studio is in the same vein that it's meant to be a place where you can test the model and play with them as easily as possible. And the good thing is that we have the same SDK between Vertex AI and developer API, so you can swap from one to another. So there's no wrong place to start playing with the models, because you can always change. And then we can skip that. And I can go back here, because actually what I was going to say is that we actually have a few differences between when you're using Vertex and when you're using the Gemini API, because the goal of the Gemini API is to hide all of the complexity from Vertex. And one of those complexities is creating buckets, creating ACL for the buckets, giving writes, and all of those things. So the Gemini API has that API that is called File Upload. And basically what it does is that you upload a file, and then it's easily accessible from the model. So I upload the file, and then I will use what we call chat mode. So basically what it does is that it changes requests, and it keeps the history of it, so that it's easier to keep all of the context. And in this case, it's going to be good, because we are going to feed the whole book to the model, thanks to the large context window. And that's why it's going to be useful in our case, because I'm going to run that while I talk. It's going to be useful because for image generation, it's always a good idea to know what has been generated previously, so it keeps the same styles and things like that. I'm also going to use structured outputs, so that we have a structure about what the model outputs, and we can talk the same language, which is going to be very simple. I want it to generate prompts, and I want the prompts to have a name, so that we know if it's a chapter, a character, or something, and the prompt in itself. So I'm just creating, initializing the chat clients with the response type JSON, the scheme that I'm providing. And something that we shipped yesterday, that is service tier priority. So don't do that yourself, I think you should remove that line. What it does is that it, we shipped that last week. So we have three service tiers. You have the normal one, you're paying the normal price, you're in the queue with everybody else. And we have another one that is called Flex, and that's basically, I don't care if that takes a long time, but I want to pay less. So you're going to have a 50% discount, but your request can be delayed and so on, up to a few minutes. And on the other hand, you have priority, where you are going to pay twice the price, but at the same time, you're guaranteed that it's going to be fast, because you will have the fast track, like in the airport or anywhere else. So yesterday, I did that, and I'm going to use that to be certain that it works well for me. [SPEAKER_02] And we have another one that is called Flex, and that's, I don't care if that takes a long time, but I want to pay less. So you're going to have a 50% discount, but your request can be delayed and so on, up to a few minutes. And on the other hand, you have priority, where you are going to pay twice the price, but at the same time, you're guaranteed that it's going to be fast, because you will have the fast track, in the airport or anywhere else. So yesterday, I did that, and I'm going to use that to be certain that it works well for me. But for you, you might want to save a few bucks and not add that here. [SPEAKER_04] How much more expensive? Twice. [SPEAKER_04] Twice, okay. Yeah. And I'm not sure it works with VO anyways, which is the most expensive ones of the model. And no, we have Liar yet, so. Okay, yeah, okay. So anyway, I'm starting with, I'm creating this chart, and what I'm going to do is to send it a first message that says, I'm feeding you the whole book. I don't need to do anything with the book yet, but you have it. It's in your context, and instructions will follow. And then we are going to define the style. Usually, I just said nothing, and then I let Gemini come up with its own style. But then I'm getting tired of having exactly the same style always, so let's write something. A colorful building blog style. Let's go with that. Let's see how it goes. So I'm just defining a style. And by the way, if you're using Collab, Collab has those things that they called magics, and that's nice to create those notebooks and to have the forms that people can fill, and that's just fitting into the code. Then some system instructions to direct the model, to the kind of images that we want. Because the problem I had at the beginning when I was working on those examples is whenever you ask a portrait image and the model knows it's about a book, it tends to create cover pages and to add a title. And I didn't want those styles at all to create ones with different panels, and I didn't want that either. So that's what the system instructions are about. And then we can start working, and I'm going to ask it to create prompts for each characters of the book. So can you describe the main characters, only the adults? Actually we could remove that, because it was just from the beginning of Nano Banana when you could not generate kids images in Europe, but it's not true anymore. So you can create image from nothing with kids, but you can't edit images with kids. That's a current limitation in Europe. Anyway, here's our prompts for each character. And then we can move to creating the images for each of them. And what I'm going to do is that I'm going to create another chat just for the images, because I don't want to mix the text and the images output. But I want it to be a chat so that it will have all of the history of the previous images it created. So I'm setting up with the responsibility image, the aspect ratio we want, the stem instructions we decided and the style, and priority as I said before. So here we go. So here's the mole. That's nice. Here's the water rod. And that's actually way faster now that I'm using priority, so this works. I could have done it in a better way and make all of the calls asynchronous. That wouldn't have worked with chat mode, but that would have been a way to make it faster. I think it's good enough, the toad. Way bigger than the car for some reason. Yeah. And then Mr. Badger. So this is mainly a demonstration of what you can do. Every time I run it, I'm thinking, oh, if I was to optimize it, there's plenty of better ways to do that. And the badger, and then the author will be at the end. One of the things I'm doing in the code, if I go up, is that I'm also saving all of the generated image from the character in an array. And I will explain afterwards why I'm doing that. And that's, once again, just appending them one after another. In real world, I would save them in a better way. And the author is here. So now we can move to the next phase, which is illustrating the book. So same thing, I'm going to ask the chat for each chapter. Give me a prompt to illustrate what's happening. It should be a single image, not a multi-tiled one. I'm trying to force it to describe the character again, even though I will have the images as references, because it's still better. So let's go with that. That should be quite fast. Well, same thing, the way it works with charts, it's keeping an history that is all the previous messages. And it's sending back the history to the model every time we make a call, which can be, in this case, since it's resending the book all the time to the model, that's the reason why it can be quite slow. We actually released new API a few months ago that are called the Interactions API. Let's run that while I talk. And the main difference between the new API and the old ones is that the new APIs are going to be stateful and stateless. And what changes is that every time you make a call, you get an Interactions ID. And you can reuse that Interactions ID in future calls, and it will recover all of the context directly from the server. So you don't need to re-upload the same context again and again at every turn of the conversation. And it's also making it easier to form the discussion. For example, you want to create a song with a cover image. You can create the lyrics with one model, and then you fork it. And on one end, you create these images, and the other end, you create the song. You had a question? [SPEAKER_09] How long do you start the session? That's a good question. I think it's two days, something like that. It's still in preview. But there's good chances that at I.O. we make it the default API. And it's also making it easier to form the discussion. For example, you want to create a song with a cover image. You can create the lyrics with one model, and then you fork it. And on one end, you create these images, and the other end, you create the song. You had a question? [SPEAKER_09] How long do you start the session? That's a good question. I think it's two days, something like that. It's still in preview. But there's good chances that at I.O. we make it the default API. But I'm not using it enough to know. One of the cool things it does as well is that since it knows that it's context that you're going to reuse, it's automatically caching it as well. So it makes it cheaper to run. But even though the normal API are also doing the same. So we have our chapter. So the third chapter is next to the river. So the moor and the characters are together. Then it's on the road with the toad. And then in the forest, in the snow. Oh, scary. So see. And what I did there is that I used the fact that we were using the history to trust the model to have all of the previous images of the character so that it would remember how they look like and how to create new images of them. But there are actually better ways to do that. So I tried another way, which is to create a new structured output that is the name of the chapter, the prompt, but also the list of the characters that are appearing in this chapter. And that's what I was doing if I wanted to do it at scale with more than a few characters. And I'm going to ask the model to give me a prompt for each chapter, but also to have, thanks to the list of characters, I will only give it as references the right images for the chapter. So I'm going to get the same thing here. And so in the first image, there should be the mole and the water rat. In the second one, Mr. Toad, mole, water rat, gray horse. And the third one, mole and water rat again. So I created a very dirty script that searched through the character images that we saved earlier, so that I can give it a list of characters. And it was going to give me the list of images to give to the model. Did I run it? Yeah. And I'm going to do the same thing and generate images for each chapter. But this time, instead of relying on chat mode, I'm going to use generate content, so the unary call. But I'm going to pass it the images that are from the characters that are specifically in this chapter image. So it should give slightly better context to the model about what to display and how to show them. So let's see if it works better. I think if. And as I said, if I wanted to do it at scale, I would have a lot of improvement I would do about that. And one of them would be that I think I would generate more than one image per character. I think I would have one portrait image and then one full body image. And then maybe from the side, from the back, and ask the model to tell me which exactly, how they are going to be displayed so that I can give the model exactly the reference we need for the generation. So here they are. And it's more or less the same, to be honest. Except that this time, I don't know why, it seems to be attacking them all for some reason. I don't think that's what the story is about. It's better. You might know better than me. We can check the prompt. What does the prompt say? Blah, blah, blah. No, it doesn't say anything about attacking the. No, it's rescuing him. Yeah. Yeah. That's how you can use the model. For those who came, the content I'm showing, you can open it there. And it's a colab that is about taking a book and creating images and videos to illustrate the book and the characters. So we went through creating prompts for each character and then generating images for each of those characters and then creating new prompts for each chapter and then creating images for the chapters using the reference that we have from the images. If you can't see what I wrote, it's goo.glee slash cookbook dash illustration. So, and then we're going to try to move to the next phase with videos now. So we are going to use Vio to generate videos of those images. So I'm going to swap to the bigger model because I can pay for it. You can skip to the, stay with the cheapest one if you don't want to spend too much. And I'm going to just do exactly, take the last chapter, take the last image and send it the same prompt and the reference image. So that's what I do here. So when I pass an image to Vio, that's, it's going to use it as the first frame for the video. And funnily enough, a lot of the training for video generation models is actually image generation because I think the most important part of generating a video is generating the first frame so that it knows where to start with and then what to do with it. So. So what's the best video model? [SPEAKER_04] Is it? [SPEAKER_04] It's the one that doesn't have light or fast. So. So fast is just a faster version of the same. [SPEAKER_04] Yeah, that's. Even more expensive, but faster. [SPEAKER_04] Yeah. So Vio 3.1 is the main Vio model and the other ones are smaller versions of it so that are running slightly faster and doing less generation turns. And yes, I said I was going to repeat the questions and I forgot the question was which one is faster, the better model among the three. So I guess the door just opened. I'm still very early. So up, we can see where it goes. Yeah, it saved. [SPEAKER_10] And there was sounds. I don't know if we can we have more sounds so we can see what. Quick, give me your hand. [SPEAKER_10] We must get out of this dreadful place. [SPEAKER_10] Thank you, water rat. of it so that are running slightly faster and doing less generation turns. And yes, I said I was going to repeat the questions and I forgot the questions was which one is faster, the better model among the three. So I guess the door just opened. I'm still very early. So up, we can see where it goes. Yeah, it saved. [SPEAKER_10] And there was sounds. I don't know if we can we have more sounds so we can see what. Quick, give me your hand. [SPEAKER_10] We must get out of this dreadful place. [SPEAKER_10] Thank you, water rat. [SPEAKER_10] I thought I was done for. [SPEAKER_10] That's not that bad except the wrong character is speaking. That's a problem with using the same prompt when you create the videos and the audio and the image because it doesn't have the extra content about what exactly is meant to be happening afterwards. [SPEAKER_02] So that's why I added another example that is we're going to do the same except we are going to add one more step that is asking Gemini to come up with a new prompt just for the video and to explain what's happening afterwards. So that's what I did. [SPEAKER_02] I'm going to animate this chapter image and can you create a prompt with Vio about what's happening in a few seconds after the initial image. And I'm passing it the last image so that it knows exactly where to start. And while it runs since we have yet again new people who came, if you want to follow on your laptop, that's a link to open the collab I'm showing. So goo.glee slash cookbook dash illustration. And what we are doing at the moment is illustrating the book that is called The Wheel of the Willow, something like that. And creating images and then videos to illustrate what's happening in the book. So it came up with a prompt that is in a colorful building block style. A water rat in his round, brown face and blue jersey lowers his silver pistols and offers a reassuring part to the mole's shoulder. The mole in his black velvet smoking suit exhales the puff of white plastic vapors in relief and as his pink snort switches, they turn and begin to walk together, blah, blah, blah, blah, blah, blah. What's interesting in the prompt is that I feel it got the style of the book. So the prompt is written in the same style, the same oldish style of speaking as well. And you can see that it's realized that it's not attacking the mole, it's just saving it and so on. And also what we can see from there is that you remember I said some system instruction at the beginning saying that whenever it describes characters, it should always try to describe what they look like and all. So it added how they are dressed all the time, which also helps with the character consistency. So let's see if this one is better. I think it's better except there's no discussion. So you can try with yourself. Just be careful, as I said earlier. Video generation can be expensive, so don't try it 100 times until you get it. It can be black. OK, so I see. Then we have this new Lyria model that we shipped two weeks ago. That is our new music generation model. And basically, as I said, it's a model to which you can give a prompt and it will come up with a song. And you have two different models. You have the clip model that is creating 30 second music. And I think it's for 4 cents per song. And you have the full song model that is creating up to three minutes songs. And it costs twice as much, so 8 cents per music. For the sake of the demonstration here, I'm using the clip model. So the fastest one. So we are going to use exactly the same trick as before. So we are going to have Gemini create the prompts for the music generation. So I'm asking it to create instrumental songs for each chapter. And to create them for Lyria. Keep the consistency between the chapters. But at the same time, highlighting what's specific in each chapter. So I don't want three times or four times the same song. I want different ones for each song. And by the way, something that I forgot to say is that, as you can imagine, we have multiple models. We have multiple teams inside of DeepMind. But actually, all of them are working together. And a big part of the training data for the Gen Media models is being made with the help of Gemini. So basically, all of our Gen Media models are trained with prompts that are written by Gemini. So that's also why Gemini is quite good at creating prompts for the Gen Media models. Because they have been trained to listen to him very well. So that's why these tricks of having Gemini write the prompts for you works quite well. And in any case, deep down, there's always a bit of rewriting of your prompts that are being done by the Gen Media models before it's actually starting to generate. Just because otherwise, when people are sending one-liners, the models won't do anything interesting with a one-liner. And so usually, the longer your prompt, the more interesting it's going to be. And the more likely it's going to be following what you're asking for. So I think I talked a lot because music generation is actually quite very fast. So we can see the different songs. I think that fits with a pastoral suit representing spring and flowing waters. And see, next one, the open road. Feels more adventurous, yes? And then the dark forest, let's see. And you can see the prompt here. It comes with which instruments to use, how to use them. What's interesting with the Lyria model is that it's actually, everything is managed in the prompt. You don't have parameters at all at the moment. So if you want the song to be a certain duration, you can just ask in the prompt. If you want the song to be using a certain scale, you can ask it in the prompt. If you want a certain BPM, you ask in the prompt as well. And the model is really good at understanding everything you ask in the prompt. And you can ask, it doesn't make sense for 30-second songs much. But if you're building longer songs, you can say, during the first 30 seconds, that's what I want the songs to be. And then you switch to something else after. Or you can say, this is the intro, this is the outro, this is the, always forget how to say that in English. But the part that is coming up multiple times in the song, so that it knows exactly how to chorus. Chorus. If you want a certain BPM, you ask in the prompt as well. And the model is really good at understanding everything you ask in the prompt. And you can ask, it doesn't make sense for 30-second songs much. But if you're building longer songs, you can say, during the first 30 seconds, that's what I want the songs to be. And then you switch to something else after. Or you can say, this is the intro, this is the outro, this is the, always forget how to say that in English. But the part that is coming up multiple times in the song, so that it knows exactly how to chorus. Chorus. It knows that this part needs to be repeated and that it needs to be the same thing and the same model. And if you want lyrics, you can either provide the lyrics or just let it invent lyrics. So let's say we are going to change that. We are going to create songs with chapter. Add lyrics to describe what's happening in the chapter. And let's see how it goes with lyrics this time. As I said, music generation is quite fast, especially the 30-second model. It takes a few seconds to generate things. The longer part here is actually sending the book again to the model and asking it for prompts. I think it didn't work. [SPEAKER_02] I think it didn't work. I think it is. [SPEAKER_08] I think it is. [SPEAKER_08] And let's see this one. [SPEAKER_08] Toad in his caravan of yellow and red. [SPEAKER_07] Dreamed of the dusty high roads ahead. [SPEAKER_07] And you can see that the theme of the song is quite the same. That's the price we have to pay for using Chat Mode. Because it remembers the prompt it did before. So we asked for new prompts. But it still has a memory of what it did before. So I think it anchored it to use the same kind of prompt and describe the scene the same way. But still, it's very funny to work with the lyrics. And as you can see in the prompt, it's basically just adding the lyrics in the prompt. And the model understands that these are the lyrics I need to play and add in the song. And you can also say that this specific part is said at exactly this moment in the song. This part is being said at another moment. And if you check the output of the Lyria model, you actually have the old lyrics with the times. So you can create a karaoke app or something like that using what you get out of the model. Let's try the third one. Into the wild wood where the shadows are deep [SPEAKER_08] And evil faces through the hollow trees peep [SPEAKER_08] Mole is in terror and lost in the snow [SPEAKER_08] Till Ratty arrives with his pistols aglow [SPEAKER_08] We could do a musical with that. So that's for music generation. And then we also have text generation models. And I'm going to show you something very fun. And I guess you all know about the old text to speech model because everybody loved the Notebook LM integration that can create a podcast. And that's actually great that you can select two different voices. So you have two characters talking with each other. But I'm going to show you a trick with which you can actually create something that is basically you can add more characters than actually two when you are creating discussions with the TTS model. So here's what I'm going to do. I'm going to ask the model to extract a specific dialogue from the book just because I didn't want to copy paste it. So I told it that it starts with small, neat ears and thick, silky ears and ends with his ears in the air. But then I asked it to write it as a play so that it's basically a transcript of what which character should be saying. And the trick is that I'm asking it to create a specific style of way of speaking for each character, even though it's going to be using the same voice, and to write the transcript that way. So when it's a narrator, I'm going to use one specific voice. And when it's all of the other characters, it's going to be the same voice for all of them. So narrator is saying something, and then character is saying something, and then write the style between parenthesis, and that's what is going to tell the model how to talk. And that's something that you can also use to say, oh, it is saying this part very whispering, and then this part has a lot of emotion in it, and you can play with the way the character talks. But I'm going to use that to ask the model to come up with very different ways from the same voice to talk when each character is talking. And that actually creates the feeling of actual different voices for each of them. And then I'm going to pass that to the TTS model and ask it to read it, basically. One of the tricks, and I lost 15 minutes because of that yesterday evening. So you cannot just give it the text to read. You always have to start with read this text or something like that. Otherwise, for some reason, it doesn't know that it needs to read the text that is given to it. And this is very complex. I think there's no simple way to set it up. But basically what I say is, speaker-narrator is using the voice Sulafat, and character is using Fenrir. And this is all text. So narrator talks, and the first character talk, and that is going to have long poetic pauses. And then the second one is breathless and humble stutter. And you can see that we can guess which character is which one because it's the same way of speaking that is reused for each of those lines. And it's still running. The problem with the TTS model is that it's a very good model, but it's not a very fast one. The reason for that is that... Small, neat ears and thick, silky hair [SPEAKER_06] It was the water rat. [SPEAKER_06] Then the two animals stood and regarded each other cautiously. [SPEAKER_06] Hello, mole. [SPEAKER_05] Hello, rat. [SPEAKER_06] Would you like to come over? [SPEAKER_05] Oh, it's all very well to...to talk. [SPEAKER_06] The rat said nothing, but stooped and unfastened a rope and hold on it, then lightly stepped into a little boat which the mole had not observed. [SPEAKER_06] The rat sculled smartly across and made fast. [SPEAKER_06] Then he held up his forepaw as the mole stepped gingerly down. [SPEAKER_06] Lean on that. [SPEAKER_01] Now then, step lively. [SPEAKER_01] The mole, to his surprise and rapture, found himself actually seated in the stern of a real boat. [SPEAKER_06] This has been a... [SPEAKER_06] Hello, mole. [SPEAKER_05] Hello, rat. [SPEAKER_06] Would you like to come over? [SPEAKER_05] Oh, it's all very well to...to talk. [SPEAKER_06] The rat said nothing, but stooped and unfastened a rope and hold on it, then lightly stepped into a little boat which the mole had not observed. [SPEAKER_06] The rat sculled smartly across and made fast. [SPEAKER_06] Then he held up his forepaw as the mole stepped gingerly down. [SPEAKER_06] Lean on that. [SPEAKER_01] Now then, step lively. [SPEAKER_01] The mole, to his surprise and rapture, found himself actually seated in the stern of a real boat. [SPEAKER_06] This has been a wonderful day. [SPEAKER_06] Do you know I've never, never been in a boat before in all my life? [SPEAKER_06] What? [SPEAKER_03] Never been in a... [SPEAKER_03] You never... [SPEAKER_03] Well, I... [SPEAKER_03] What have you been doing then? [SPEAKER_03] Is it... [SPEAKER_06] Is it so nice as all that? [SPEAKER_06] We can stop that. You can see, you couldn't guess that we are using the same voice for the two characters being directed in different ways. We're still in two different directions and you can use this trick to actually create multiple characters, multiple voices for those characters and make it seem seamless for users. As I said earlier, this is meant to be just a demonstration on how to do it if you were to do it at scale. That's exactly right. I would not do this kind of trick. I would actually create a full transcript with the actual names and then keep on the side a prompt for each character. And maybe sometimes you still want to not talk about them exactly the same way because sometimes they still need to be excited even though they talk very slow and so on. But still, that's just to show how good the TTS model is at creating different voices. And even though I asked it to force an accent, it didn't do it. But you can also play with this character has an Irish accent, this one is English, this one speaks with a German accent or whatever. [SPEAKER_02] And that's also a very easy way to create different feeling about the character by using the same voice. We're nearly at time for the questions. But just to finish, we use a very large context window of the model to fill it with a full book and fill it multiple times with a full book because we've tried putting it all in at the same time. [SPEAKER_02] But since it's a multimodal model, you can work with all things, not just text. [SPEAKER_02] So you can fill it as an audio book. [SPEAKER_02] You can fill it as a video as well. [SPEAKER_02] You can play with that and not get limited to text to speech and things. So this is another example with another book, which is The Adventures of Chatterer, The Red Squirrel. And we are basically going to do the same thing. I'm going to run all of it at the same time. And this time, I basically ask it to use a style that is futuristic science fiction, utopia, saturated neon lights. So it's going to be not the kind of squirrel you are expecting. Yeah. And while it runs, I think I said I was going to show you something. So we also have AI Studio, but in AI Studio we have a gallery with lots of example apps that we are building. And I wanted to show you it's going to be in GenMedia. As I told you, we have the Lyria model that is creating music songs. But we also have the... no, not this one. Let's go with it. This one's better. We also have the Lyria real time model I was talking about. And basically, you are asking it to create music that is post-punk, tunes and neo-sounds at the same time. But you can say, okay, I want more K-pop. And slightly more drums. And I don't know what post-punk is, so I don't want it. You can hear the music changing. And let's go with something more chill. So I think, as I said, that's my favourite model because I think it's underused. And there's plenty of things that I can imagine doing. I come from the video game industry. So one of the things I would have tried is can you recreate music in real time for the player depending on where they are. In which region are they? Are they in the forest? Are they jumping? Are they cooking? Are they fighting? How much HP do they have? And so the music could change in real time. And that's the kind of thing you can try. And for some reason the link is not there. But there's another very cool example our colleagues who are working on this model made. It's basically you are in space. And each planet is a prompt. And you can move through the planet. And depending on which planet you are close to, the music changes. So you can just move around the planets. And sometimes there are weird things happening. Because Christmas songs is just next to Viking metal. So the mix can be quite funny. So that's it for the presentation. I have some time for questions now. Thank you. Yeah. And for those who arrived too late, you can check the content afterwards. And as I said, we have this cookbook that is basically a GitHub repo where we are adding quick starts on how to use the models. Some tricks. And also examples of more complex things you can build when you are mixing different capabilities and models. I have a question. [SPEAKER_04] I don't know if you can answer it. [SPEAKER_04] Yeah. I think for questions, we need a mic. So that's it. I don't need to repeat them. One second. [SPEAKER_08] One second. [SPEAKER_04] So thank you first of all very much for this nice demo. Some tricks. And also examples of more complex things you can build when you are mixing different capabilities and models. I have a question. [SPEAKER_04] I don't know if you can answer it. [SPEAKER_04] Yeah. I think for questions, we need Mike. So that's it. I don't need to repeat them. One second. One second. One second. One second. One second. [SPEAKER_08] One second. [SPEAKER_08] One second. One second. [SPEAKER_08] One second. [SPEAKER_04] So thank you first of all very much for this nice demo. [SPEAKER_04] It was a lot of fun to follow along. [SPEAKER_04] I have a question. [SPEAKER_04] In our company, we are offering to all employees also some of the models. [SPEAKER_04] And I think we are still on Nano Banana 1 because we can only offer models hosted in Europe. [SPEAKER_04] And all the new models are still in preview. [SPEAKER_04] So we don't have access to it. [SPEAKER_04] Do you know if this will change? [SPEAKER_04] So the short answer is no. Okay. But I was expecting the question because it's a pain for everyone in Europe. As I said, my job is to bring the feedback from the developers and to try to make things change. So that's one of the fights I'm fighting at the moment. So that we have some ways to offer a better exit for better ways to use the model for people in Europe. Because in Europe we care about data privacy and data sovereignty and all of that. So I know it's a problem. So the core of the problem in a way is the rule that Google Cloud has that every preview model is only available on global endpoints. So that is unlikely to change. But what we are going to try to change is to release the model in global accessibility faster. [SPEAKER_02] The problem we have with Nano Banana 2 and the Pro and the Gemini 3.3 as well is that we release models too quickly back to back. [SPEAKER_02] And so instead of having Gemini 3 going GA, we release Gemini 3.1. [SPEAKER_02] And so we reset the preview counter. [SPEAKER_02] So we need to do something about that. [SPEAKER_02] But yeah, I know. I hear you. It's my P0 thing that I want to change. Thank you. And have you run the notebook at the same time? Yes. [SPEAKER_04] Which side, where do you go? Wow. Did someone else run it at the same time and then with a different style and that, oh, maybe a different book? No? I did the Frankenstein book. [SPEAKER_04] Oh. And choose the retro game. [SPEAKER_04] Sure. [SPEAKER_04] And it was quite interesting. [SPEAKER_04] Yeah. [SPEAKER_04] So the character looked very video game-like. [SPEAKER_04] All right. [SPEAKER_04] Oh, yeah. [SPEAKER_02] Yeah. The main difficulty with books like Frankenstein is that sometimes the model is not going to be allowing things that are too graphic. That could be happening in the book. So it can be a bit torn down or worse, it's not going to accept to show the image. Yeah. That's why I settled with kids' books for the example. That's easier. Except when I was not able to make images of children, which was also another limiting factor. So, yeah. I can show you other demos that we have related to GenMedia. In the meantime, if you have questions, just raise your hand and we can... We're going back here. See? That's the futuristic neon style version of Chatter the Squirrel. I can show you a bit more about Lyria because it's new. And you likely already know everything about Nano Banana. So as I said earlier, when you work with Lyria, you always get two outputs. If you set the modalities to be their audio and text, you will get two outputs. And the first output is going to be the lyrics, and the second output is going to be the music. And that's actually one of the few models where it's very interesting to use streaming. So when you do generate content here, you can use generatecontent_stream. And what it does is that you receive the first part first, and then the second part afterwards. So you get the lyrics first. So if you want to do something according to the lyrics, like creating an image or giving the song a title, then you can do it while the music is still generating, and you don't have to wait for the full output to be there. So you get the lyrics and you get the timing. So this sentence is going to be said at the beginning, and then after 4.8 sec, it's going to say something else and so on. So you can hear it. [SPEAKER_07] The air is still and cold up here. [SPEAKER_07] The mountain tops are sharp and clear. [SPEAKER_07] And then a streak of gentle gold. So yeah, and you can provide the same thing. And here it's only like you can see the last one. It's providing when it starts and when it ends, because it wants to have 1.2 seconds without things said at the end. You can also create music from images. So that's one of the things I forgot to do in today's demo. I should have also gave the images from the chapter, so that it would have been used as a reference to create the first image. So I use this picture of grocery lists for making a portofu, and then it will come up with a song about doing a portofu. What was the prompt again? An epic song with opera voices about these quests. See, it's becoming epic. The scrolls are dry, the ink is ancient, the list is long, the soul is patient. Beef shank and ribs, the holy prize, in the cellar where the shadow lies. [SPEAKER_08] The carrots and the turnips wait. [SPEAKER_08] And it's also nice that you can use multiple voices as well. [SPEAKER_09] I'm going to skip ahead a bit. [SPEAKER_02] The time is gone. [SPEAKER_02] The celery wilt, the celery wilt, the end is near. [SPEAKER_09] The final moment of the quest is here. [SPEAKER_09] The celery wilt. [SPEAKER_09] The celery wilt. [SPEAKER_03] The closed, the sauce. [SPEAKER_03] But then, I talked about the interactions API earlier, about being our new ways of using the API. So that's an example using those. [SPEAKER_02] So it basically was the same as the current API. [SPEAKER_02] So you give a model, you give an input, and you get response modalities. [SPEAKER_08] And it's also nice that you can use multiple voices as well. [SPEAKER_09] I'm going to skip ahead a bit. [SPEAKER_02] The time is gone. [SPEAKER_02] The celery wilt, the celery wilt, the end is near. [SPEAKER_09] The final moment of the quest is here. [SPEAKER_09] The celery wilt. [SPEAKER_09] The celery wilt. [SPEAKER_03] The closed, the sauce. But then, I talked about the interactions API earlier, about being our new ways of using the API. So that's an example using those. [SPEAKER_02] So it was the same as the current API. [SPEAKER_02] So you give a model, you give an input, and you get response modalities. [SPEAKER_02] Yeah, Philippe who worked on that is not there, so I can say it. [SPEAKER_02] I would not have renamed content to input because it's going to confuse everybody. [SPEAKER_02] But that's how it is. [SPEAKER_02] And then you get the output. And if we can check in the output, I think, no. Yeah, we don't see it here. But there's a, and then it works the same way, except you get this interactions ID that you can use to chain things. And then that's, yeah, that's the one with images and prompting. What you can do as well is you can use the BPM part. So you want a song that is very fast or very slow, so that's the big uses. And I even told you to use the Reset Accelerando Illusion. So it gives you the illusion that the music is getting faster and faster and faster. So if you want music for when you do your sports routine, that's how you do it. And as I said, you can give specific time. So the first 10 seconds are going to be fast acoustic guitar. And then it goes into piano for 20 seconds, for 10 more seconds. And then it's full band afterwards. Actually, it's not following it. OK, forgot what I just saw, Sean. For some reason. And then, but the easiest way is this. You can use the gear. You can use the structure. So that's how I want my intro. That's how I want my verse. That's how I want my outro. And that's 30 second song. So it's not that you want to have the chorus, but you can also add the chorus and the bridge and so on. What is that about? Ah, yes. The sunset. Yeah. I need to make a better, longer song, I think, for example. [SPEAKER_09] And but you can also chain all of that into everything together. So this is a full song where from the first two seconds I have an intro. I can tell exactly which scale to use, how intense it's going to be. And then it moves to another verse and so on and so on. And that's how you can get something very complex. So it starts very, very slow. And then if we move. There's the drums and that started and so on. So we should be in this part. See it's still laid back. And starting to be, yeah, to add grooves and so on. [SPEAKER_02] So if you want to create complex things, it's better to use the full song model because the short one is taking shortcuts to actually build something that is interesting in 30 seconds. And as I already showed, you can provide the lyrics. So it's creating a song about nano bananas. Yellow peel, a tiny sweet. The nano banana, a tropical treat. [SPEAKER_09] But wait, it hums, it starts to create. [SPEAKER_09] Switching into AI mode. [SPEAKER_09] But what's funny as well is that you can use it to create things that are not songs. [SPEAKER_09] So you can ask it to create music, but without very calm music in the background and just something reading a text or something. So this, and this one I'm using the reasoning capabilities of the model, and it's knowledge about what Shakespeare is doing. And so to create a text that is basically something that looks like Shakespearean. And I didn't really tell it to read. So that's why there's still music, but you can really steer it into not having background music at all and just say things. And it works also in different languages. So you can ask it to create songs in all languages that you want. Sometimes there are a few words that are not pronounced the right way. It's still getting better. And this one, I try to ask to have it to use two different languages in the same song. And since it's the same as TTS and live models that we have, it's really good at switching languages in the middle of the generation. And once again, I'm using the model knowledge of things because it's basically trying to explain how bubble sort is working in music. And instruments, you saw it. So yeah, there's plenty of very cool things to be able with the Lyria model. So give it a try. It's very easy to use and because everything is in the prompt. So, yeah. Do you have any other question? No? What do you want, Paige to show you afterwards? Huh? She has seven more minutes to prepare something new in our presentation. That's true. And also just here, for clarification, I'm not sure if this was announced to all of you. But there are some electrical issues in the building. [SPEAKER_00] So, they're you all are the lucky, valiant few that made it here earlier this morning. [SPEAKER_00] The most of the attendees were not allowed into the building. [SPEAKER_00] And so we'll do the session that's coming up at 1040, but we'll also be bringing back everybody in the afternoon who was going to be presenting in the morning to do a whistle stop tour of all of the Google D Mike things for the afternoon workshop. [SPEAKER_00] So if you would prefer to come back for the afternoon workshop, you can. [SPEAKER_00] It'll just be at 1 PM. But there are some electrical issues in the building. [SPEAKER_00] So you all are the lucky, valiant few that made it here earlier this morning. [SPEAKER_00] Most of the attendees were not allowed into the building. [SPEAKER_00] And so we'll do the session that's coming up at 1040, but we'll also be bringing back everybody in the afternoon who was going to be presenting in the morning to do a whistle stop tour of all of the Google D Mike things for the afternoon workshop. [SPEAKER_00] So if you would prefer to come back for the afternoon workshop, you can. [SPEAKER_00] It'll just be at 1 PM. [SPEAKER_00] So that was the example I was talking about. [SPEAKER_00] It's just, I don't know where Christmas songs are, but how can you drive us by your own? It just serves for Space DJ and it's available online. Can you move to put the songs higher? Yeah. She's on this, so it should be nice. [SPEAKER_02] So the molo is not meant to do voices, but it can do some vocalizations. [SPEAKER_02] It's not that bad at using the same voice and switching the style of the music. And you have an autopilot, so it just moves around and that's the music until it's done. Yeah, I saw a phone rock, I think Nashville songs. Yeah. Yeah. [SPEAKER_09] Yeah. [SPEAKER_09] Yeah. [SPEAKER_09] It's not moving fast enough. That's nice. [SPEAKER_02] Let's move to the other place. [SPEAKER_02] Australian E-pop. [SPEAKER_02] Turn table, oh, turn table. [SPEAKER_09] Australian E-pop, I guess. [SPEAKER_02] No, no Australian here. [SPEAKER_02] But that's a very cool model. [SPEAKER_02] And the only thing is that the session ends after 10 minutes so you don't run it at the time, but I can feel it should work. [SPEAKER_02] I'm just missing the surf button. [SPEAKER_02] Speed metal. [SPEAKER_02] Yeah! [SPEAKER_02] So I will stop with that. But give it a try, it's a really cool model to play with. [SPEAKER_02] I'm just missing the surf button. [SPEAKER_02] Okay. So that's it. I will still be around if you have other questions, things you want to discuss that you didn't want to be on camera. [SPEAKER_02] So thank you. Okay. But I was expecting the question because it's a pain for everyone in Europe. As I said, my job is to bring the feedback from the developers and to try to make things change. So that's one of the fights I'm fighting at the moment. So that we have some ways to offer a better exit for better ways to use the model for people in Europe. Because in Europe we care about data privacy and data sovereignty and all of that. So I know it's a problem. So the core of the problem in a way is the rule that Google Cloud that every preview model is only available on global endpoints. So that is unlikely to change. But what we are going to try to change is to release the model in global accessibility faster. The problem we have with Nano Banana 2 and the Pro and the Gemini 3.3 as well is that we release models too quickly back to back. And so instead of having Gemini 3 going GA, we release Gemini 3.1. And so we kind of reset the preview counter. So we need to do something about that. But yeah, I know. I hear you. It's kind of my P0 thing that I want to change. Thank you. And have you run the notebook at the same time? Yes. Which side, where do you go? Wow. Did someone else run it at the same time and then with a different style and that, oh, maybe a different book? No? I did the Frankenstein book. Oh. And choose like the retro game. Sure. And it was quite interesting. Yeah. So the character looked very video game-like. All right. Oh, yeah. Yeah. The main difficulty with books like Frankenstein is that sometimes the model is not going to be allowing things that are too, let's say, graphic. That could be happening in the book. So it can be a bit torn down or worse, it's not going to accept to show the image. Yeah. That's why I settled with kids' books for the example. That's easier. Except when I was not able to make images of children, which was also another limiting factor. So, yeah. I can show you other, I don't know what page afterwards is going to show. I can show you other cool demos that we have related to GenMedia. In the meantime, if you have questions, just raise your hand and we can... We're going back here. See? That's... Whatever. That's the futuristic neon style version of Chatter the Squirrel. I can show you a bit more about Lyria because it's new. And you likely already know everything about Nano Banana. So as I said earlier, when you work with Lyria, you always get two outputs. If you set the modalities to be their audio and text, you will get two outputs. And the first output is going to be the lyrics, and the second output is going to be the music. And that's actually one of the few models where it's very interesting to use streaming. So when you do generate content here, you can use generatecontent underscore stream. And what it does is that you receive the first part first, and then the second part afterwards. So you get the lyrics first. So if you want to do something according to the lyrics, like creating an image or giving the song a title, then you can do it while the music is still generating, and you don't have to wait for the full output to be there. So you get the lyrics and you get the timing. So this sentence is going to be said at the beginning, and then after 4.8 sec, it's going to say something else and so on. So you can hear it. The air is still and cold up here. The mountain tops are sharp and clear. And then a streak of gentle gold. So yeah, and you can provide the same thing. And here it's only like you can see the last one. It's providing when it starts and when it ends, because it wants to have like 1.2 seconds without things said at the end. You can also create music from images. So that's one of the things I forgot to do in today's demo. I should have also gave the images from the chapter, so that it would have been used as a reference to create the first image. So I use this picture of like grocery lists for making a portofu, and then it will come up with a song about doing a portofu. What was the prompt again? An epic song with opera voices about these quests. See, it's becoming epic. The scrolls are dry, the ink is ancient, the list is long, the soul is patient. Beef shank and ribs, the holy prize, in the cellar where the shadow lies. The carrots and the turnips wait. And it's also nice that you can use multiple voices as well. I'm going to skip ahead a bit. The time is gone. The celery wilt, the celery wilt, the end is near. The final moment of the quest is here. The celery wilt. The celery wilt. The closed, the sauce. But then, I talked about the interactions API earlier, about being our new ways of using the API. So that's an example using those. So it basically was the same as the current API. So you give a model, you give an input, and you get response modalities. Yeah, Philippe who worked on that is not there, so I can say it. I would not have renamed content to input because it's going to confuse everybody. But that's how it is. And then you get the output. And if we can check in the output, I think, no. Yeah, we don't see it here. But there's a, and then it works basically the same way, except you get this interactions ID that you can use to chain things. And then that's, yeah, that's the one with images and prompting. What you can do as well is you can use the BPM part. So you want a song that is very fast or very slow, so that's the big uses. And I even told you to use the Reset Accelerando Illusion. So it gives you the illusion that the music is getting faster and faster and faster. So if you want music for when you do your sports routine, that's how you do it. And as I said, you can give like specific time. So the first 10 seconds are going to be fast acoustic guitar. And then it goes into piano for 20 seconds, for 10 more seconds. And then it's full band afterwards. Actually, it's not following it. OK, forgot what I just saw, Sean. For some reason. And then, but the easiest way is like this. You can use the gear. You can use the structure. So that's how I want my intro. That's how I want my verse. That's how I want my outro. And that's 30 second song. So it's not that you want to have the chorus, but you can also add the chorus and the bridge and so on. So, um, What is that about? Ah, yes. The sunset. Yeah. I need to make a better, longer song, I think, for example. And, but you can also chain all of that into, like everything together. So this is a full song where from the first two seconds I have an intro. I can tell exactly which scale to use, how intense it's going to be. And then it moves to another verse and so on and so on. And that's, uh, how you can get something very, uh, complex. So it starts very, very slow. And then if we move. There's the drums, uh, and that started and so on. So we should be in this part. See it's still laid back. And starting to be, yeah, to add grooves and so on. So if you, if you want to create complex things, it's better to use the full song model because it's, uh, the, the short one is taking shortcuts to actually build something that is interesting in 30 seconds. And as, as I already showed, you can provide the lyrics. So it's creating a song about nano bananas. Yellow peel, a tiny sweet. The nano banana, a tropical treat. But wait, it hums, it starts to create. Switching into AI mode. Um, but, uh, and what's, what's funny as well is that you can use it to create things that are basically not, not, uh, songs. So you can, uh, you can ask it to create, uh, music, but without, with very calm, uh, music in the background and just some, something reading a text or, or, or something like that. So this, and this one I'm using the reasoning capabilities of the model to, and it's, it's knowledge about what Shakespeare is doing. And so to, to create a text that is basically, uh, something that looks like Shakespeareans. And I didn't really tell it to read. So that's why there's still music, but you can, you can really steer it into not having background music at all and, uh, and just say things. And you, it works also in different languages. So you can, uh, you can just ask it to, uh, to, uh, to create songs in all languages that you want. Sometimes there are a few words that are not pronounced the right way. It's still, uh, it's, it's still getting better. Uh, and this one, I try to ask, to have it to, uh, to use two different languages in the same song. And since it's the same as, uh, TTS and live models that we have, it's really good at switching, uh, switching, uh, switching the, the, the language in the middle of the, uh, the generation. And once again, I'm using the model knowledge of things because it's basically trying to explain how bubble sort is working, uh, in music. Um, and instruments, you saw it. So yeah, there's plenty of, uh, very cool things to, uh, to be able with, uh, with the Lyria model. So give, give it a try. It's, uh, it's, uh, it's very easy to use and, uh, because everything is in the prompt. So, yeah. Do you have any other question? No? What do you want, uh, Paige to show you afterwards? Uh, huh? She has, she has seven more minutes to prepare, uh, something new in our, in our presentation. That's true. And also just here, uh, for clarification, I'm not sure if this was announced to all of y'all. Um, but there are some electrical issues in the building. Um, so, uh, they're, uh, y'all are like the lucky, valiant few that made it here earlier this morning. Um, the, uh, most of the attendees were not allowed into the building. Um, and so, uh, the, uh, we'll do the session that's coming up at 1040, but we'll also be bringing back everybody in the afternoon who was going to be presenting in the morning to do kind of like a, a whistle stop tour of all of the Google D Mike things for the afternoon workshop. So if you would prefer to come back for the afternoon workshop, you can. It'll just be at 1 PM. So, and that was the example I was talking about. It's just, I don't know where Christmas songs are, but, uh, how can you drive us by your own? It just serves for Space DJ and it's available online. Can you move the, to put the songs, uh, higher? Yeah. She's on this, so it should be nice. So the molo is not meant to do voices, but it can do some kind of, like, vocalizations like that. Uh, it's not that bad at, uh, using the same voice and switching the style of the music. And, uh, and you have an autopilot, so it just moves around and that's the music, uh, until it's done. Yeah, I saw a phone rock, I think Nashville songs. Yeah. Yeah. Yeah. Yeah. Yeah, it's not moving fast enough. That's nice. Let's move to the other place. Australian E-pop. Turn table, oh, turn table. Australian E-pop, I guess. No, no Australian here. Um, but yeah, that's, uh, that's very cool model. And, uh, the only thing is that the session ends after 10 minutes so that you, so you don't run it, uh, at the time, but, uh, I can feel it should work. I'm just missing the surf button. Speed metal. Yeah! So I will stop with that. But yeah, give it a try, it's a really cool model to play with. Um, I'm just missing the surf button. Okay. So, yeah, I guess that's it. I will still be around if you have other questions, things you want to discuss that you didn't want to be on camera. So, yeah. Thank you. . .