SPEAKER_10
My name is Paige. I started doing machine learning a long time ago, around 2009-2010. I was primarily working with—though it feels like forever ago, I was just talking with a
SPEAKER_02
friend about this recently. Back in 2009-2010, it was wild that companies would even trust open source software to do business critical work. And so I was contributing to things like NumPy, SciPy, these microphones. Yeah. Sure. Cool, cool, cool. So NumPy, SciPy, Matplotlib, which is still just as excruciatingly painful to use. Scikit-learn, the early days of the scientific computing stack. And eventually I started working as an engineer. My background is geophysics and applied math for undergrad, [SPEAKER_07] and then computer science and carbonate geology for grad school. And so I started at Chevron
SPEAKER_07
[SPEAKER_02] doing work in subsurface geosciences, doing a lot of things like velocity modeling, drilling
SPEAKER_02
optimization, using very basic machine learning models. And also some large-scale compute. So if folks in the audience also have gray hair like me, you might remember Cloudera, which was one of the first companies that made some of the open source frameworks and tools available for consumption by Fortune 100 companies. So things like Spark. A lot of the Databricks team came from the Cloudera world. And so I was spinning up clusters of machines there. Eventually, TensorFlow got released in the open source world around 2015 towards the end of it. And I started contributing to that. The geosciences world is really big into GPUs. They were even into GPUs
SPEAKER_02
before the machine learning world. And so I had experience working with CUDA and all of the associated tools. And TensorFlow, when it was first released, only supported CPUs. I'm not sure how broadly that's known. But it was distributed deep learning across CPUs. And they needed somebody to help with getting GPUs to work with TensorFlow. That's also why there are three different code paths in the original TensorFlow 1 framework. Because they had to gut the back end and replace it for CPUs, for GPUs, and for TPUs, both for single node and for distributed computation across multiple nodes.
SPEAKER_02
So that's it. I owe my entire career to open source software and open source models. And I got hired at Google for specifically that reason. And then I left to go work at GitHub for about a year on VS Code, which is also open source. And early user experience testing for Copilot. And then came back to work on our large models. And so I was part of the original Palm 2, Gemini, and Gemma teams. Yep. So that is my journey. I'm not sure it's interesting. And you need to leave soon, you said? Because I need to have a chat with you. Oh, 4:15, because I have meetings. We have leadership team meetings from 4:30 to 5:30. I'll take over.
SPEAKER_02
Yep. And Guillaume will be here. And then also potentially our colleague Ian will come as well. Potentially, because I would not say something if I have an hour or an hour. I think there will be enough good stuff to talk about. Yeah. If anybody's curious, open source is a great way to work with teams or figure out if you would want to work with them. And then also make sure that your work is public so that there are other folks. I'd love to show you my project. Cool. Awesome. Guess how many stars it got? 700. Wow. Excellent. I'm catching up with you guys. 100. Right.
SPEAKER_02
Usually if you see everybody sprinting to do the same thing, that's a great indication that it's the wrong thing. Or a thing that eventually the model will have that capability to do. So one of my favorite examples of this is that when the models were first released, [SPEAKER_05] they had context windows of 8,000 tokens, 16,000 tokens, and so everybody was sprinting [SPEAKER_04] to build a vector database because they were thinking,
SPEAKER_05
[SPEAKER_04] oh, well, we have to work around this constraint
SPEAKER_02
[SPEAKER_04] that the models have this very small context window.
SPEAKER_04
And then obviously that's expanded over time. [SPEAKER_02] There's another great example of everybody sprinting [SPEAKER_02] to build fine tunes of models to support different languages, [SPEAKER_01] though now models support a variety of languages.
SPEAKER_02
If you've been listening along for some of the Google DeepLine sessions today,
SPEAKER_01
[SPEAKER_02] you've seen that in action,
SPEAKER_02
[SPEAKER_01] and then also many of our competitor models [SPEAKER_01] also support multiple languages. And so I think we also saw everybody sprinting
SPEAKER_01
[SPEAKER_02] to build an agent framework, when the reality is, [SPEAKER_02] I think that all of that will probably be absorbed
SPEAKER_02
into the model eventually, and also everybody sprinting [SPEAKER_04] towards things like building MCP servers, [SPEAKER_01] whereas now mostly people have moved away [SPEAKER_01] from MCP servers and are adopting skills,
SPEAKER_04
which are just fancy markdown files.
SPEAKER_01
[SPEAKER_04] And then obviously longer term, you could imagine, [SPEAKER_02] oh, hey, I just want a little listener watching everything
SPEAKER_04
[SPEAKER_02] that I do, and it will automatically create skills behind the scenes for me to use, and we're starting.
SPEAKER_02
[SPEAKER_04] Yep, yep. No, I was going to say, I do not necessarily agree,
SPEAKER_04
[SPEAKER_02] because you guys will always cater to the more generalized [SPEAKER_02] applications, but very specific ones,
SPEAKER_02
for instance, about the mics, which is where I'm working, Google is not going to work on that. I don't think that they will absorb those kind of functionalities. Yes, so counter example, our first implementation of Palm 2 and our first implementation of Gemini, we had to have a fine tune, one called MedLM and one called MedPalm, to support medical use cases. Now we see that all of the people who had previously needed to rely on those fine tunes are just using Gemini out of the box with either retrieval or with a custom prompt, because all of the data that we used for those fine tunes is just incorporated into Gemini itself. But you have the issue of reproducibility,
SPEAKER_02
you're not going to have the same results over time, you're going to have a lot of... So no large language model is deterministic. So that will be a problem regardless. But I do agree, there are—I think a lot of the magic is going to come from having a very opinionated view of use cases and being able to work directly with customers and solve their problems. So with that, and that's a perfect segue into the session today. is just incorporated into Gemini itself. But you have the issue of reproducibility. You're not going to have the same results every time. You're going to have a lot of variance. So no large language model is deterministic.
SPEAKER_02
So that will be a problem regardless. But I do agree there are things where I think a lot of the magic is going to come from having a very opinionated view of use cases and being able to work directly with customers and solve their problems. So with that, that's a perfect segue into the session today. So greetings, everyone. Thank you for being brave and for coming back. I know that there were a lot of people who were unable to join us earlier this morning for some of the sessions that we had. So just show of hands, how many folks came back this afternoon, were not here this morning? Cool, cool, cool. Excellent. So we have a treat for you today.
SPEAKER_02
For the sessions that we delivered earlier, we're going to do a recap of them. [SPEAKER_04] So we'll be walking through all of the examples.
SPEAKER_04
You won't walk away feeling like you've missed anything. And I'll be doing some demos within AI Studio and Anthropic. And then my colleague, Guillaume, and you can also see one of my agents that's doing its thing with computer use and browsing. I'm going to disconnect for a second so we don't accidentally see something that perhaps should not be shared. [SPEAKER_02] And then I'm going to go straight into Slideshow View.
SPEAKER_04
[SPEAKER_02] Awesome.
SPEAKER_02
So greetings, everyone. My name is Paige.
SPEAKER_02
So I am really excited to be here today. And I am especially excited to show you all of the things that we've been doing at Google DeepMind over the course of the last year, the last six months. [SPEAKER_04] It's been a wild ride. [SPEAKER_04] And never in my career have I been more excited to be a machine learning engineer working in this space. Just over the last month and a half, it feels a little bit of whiplash. We've released a ton of different things.
SPEAKER_04
[SPEAKER_02] So Gemini 3.1 Flash Live, which we'll take a look at in a second, which gives you the ability to have a real-time conversation with the model. [SPEAKER_02] Gemini 3.1 Pro and Flash Light, respectively, are the largest and very, very small but very capable models that are also very cost-effective and very performant. [SPEAKER_02] Augment Code, if you've heard of them, they're a company over in the Bay Area.
SPEAKER_02
They've just recently replatformed their entire agent infrastructure to default to Gemini 3.1 Pro, specifically because of performance and cost ratios. It can do a lot of really good work in a very small number of pennies. Also, Nano Banana 2 for image generation, image editing, including reverse image search, which we'll also take a look at in a second. And you'll be hearing about from my colleague Guillaume, who's our expert in generative media. Our Embeddings 2.0 model, which allows you to embed video, audio, images, code, and text in the same embedding space.
SPEAKER_02
So you can say, show me everything you have related to llamas, and you will see everything from stuffed llamas to pictures of llamas to videos of llamas to things of what does it sound like when a llama makes noises? Because I don't know, but the model somehow does. Lyria 3 for music generation, Genii 3 for world model building, our full stack runtime with AI Studio, which allows you to incorporate database, OAuth, custom API keys if you want to use other services and a whole bunch of other things, Gemma 4 for open models.
SPEAKER_02
We just released it last week under an Apache 2 license, which is really exciting if you care about open models, and then also VO 3.1 Light for video generation. So massive number of things across a broad spectrum of surfaces. And part of the reason for this is because Gemini is kind of unique in the industry in a couple of ways, one of which is that it's natively multimodal, so it can understand video and images and audio and text and code and all of the above all at once, but it can also output multiple modalities.
SPEAKER_02
So it can output text, it can output code, just like all of the other competitors on the market, but it can also output images, it can edit images, it can output images and text interleaved, and it can also output audio tokens. So pretty compelling use cases. But I think it's a lot more interesting to see it rather than to just have me talk about it. And for that, I am going to go into AI Studio real quick, and we're going to take a look at some of the things that Gemini can do. So first off, how many people have used AI Studio before? Excellent. I'm very glad that all of the DeepMinders have raised their hands.
SPEAKER_02
But for folks who have not, AI Studio is the best place to go to get access to DeepMind's models as soon as they're released. We have a playground feature, we also have a build feature, which we'll see in a second, which is very similar to 0.dev or Lovable. You can select different models here off to the right. So you can see if I click on the model name, we see some pills: everything from Gemini to live capabilities, image generation, video generation, audio generation and music, and then also our Gemma open model family, which we'll also take a look at in a second. You can select different models for the purposes of speed. I am selecting Gemini 3.1 Flashlight Preview.
SPEAKER_02
I am also on my personal instance of AI Studio, and this is an attempt to drive down costs. So Gemini 3.1 Flashlight is around 25 cents per million tokens analyzed, which is really good. Almost an order of magnitude lower than our Gemini 3.1 Pro model. And it can still analyze video, audio, etc. It's just a lot more lightweight, which means that you might not always get the same great capabilities, but it usually does a pretty good job. And you can do everything from analyzing images to analyzing video. So if I, as an example, was going to look up a dinosaur YouTube video, we already used this Rexy, the little T-Rex before. So I know it works.
SPEAKER_02
So I'm going to look at the same page. But I am going to find maybe this one, Carnadosaurus, which looks like a very, very long video, so around an hour long. We might chop it up a little bit. And it can still analyze video, audio, etc. It's just a lot more lightweight, which means that you might not always get the same great capabilities, but it usually does a pretty good job. And you can do everything from analyzing images to analyzing video.
SPEAKER_02
So if I, as an example, was going to look up a dinosaur YouTube video, we already used this Rexy, the little T-Rex before. So I know it works. So I'm going to look at the same page. But I am going to find maybe this one, Carnadosaurus, which is a very, very long video, so around an hour long. We might chop it up a little bit. But if you click this plus sign, you can see that you can add different files. You can either add files from Drive, everything from text files to PDFs. You can upload files directly, you can record audio live, add camera footage, add a link to a YouTube video, which we'll do right now.
SPEAKER_02
So I paste in the link to the YouTube video. I'll do a start time of 0 seconds and an end time of around 300. And it samples at around 1 frame per second. And then what you see is this is 30,900 tokens. I'm going to make sure to turn on Google Search grounding off to the right. And I'm going to say, please create a table with time stamps for all of the kinds of dinosaurs that you see in this video. Make sure to include a fun fact about each dinosaur. For all of the other dinosaur experts, armchair dinosaur experts in the room, you might have noticed that it said Carnadosaurus. I am skeptical that that's an actual dinosaur. So we'll see what happens.
SPEAKER_02
But what's going on behind the scenes is that the model is getting sent this YouTube video for inference. And it's not just the metadata associated with the YouTube video. It's also frame by frame the video itself. And question? Can you just put it in light? Light. Does anyone else have issues with seeing the settings? Or could we change the lighting? Oh, is this better? Oh, excellent. Awesome. So that's, thank you for the request. It's I can see it perfectly on my screen. Usually my eyes get a little irritated by light mode. But this is, as long as it's better.
SPEAKER_02
And so it took in the video. It defined the different dinosaurs that it sees. It says that the name means mediating bull. Mentioned triceratops. And then also mentioned that pteranodons are not necessarily dinosaurs. They're a group of flying reptiles. And then if I wanted to, if I wanted to get the code that was used to generate this experience and I wanted to replicate it in my own app, all I have to do is click get code. And it automatically configures the model. It configures the URL that I have inputted. The offset for the video. The prompt that I used. And it's in TypeScript, Python, or whatever your favorite language might be.
SPEAKER_02
So the TLDR is that if you can get it working in AI Studio, you can get it working as part of your app. All you have to do is click the get code button. So that is a Gemini 3.1 Flash for analyzing video. You can also use it to analyze images with a couple of other baked in tools. And if we look over to the right, we can see structured outputs, code execution. Things like function calling, even for custom functions. Also some things like URL context. And all of these are very special. But they're also just one-liners if you want to use them as part of your API.
SPEAKER_02
So as an example, if I turn on code execution, I'm going to see a Gemini 3.1 Flash selected. I'm going to go into compare mode. So I want to compare it with maybe the Gemini 3 Flash Preview. Also with code execution turned on. And I'm going to look for a picture of Lego bricks. This one. And copy the image. And right now we're in compare mode. So I'm comparing two different models. Both with the same tool turned on. And say something to the effect of draw bounding boxes around all of the green Lego bricks using Python. Make sure to display the image with bounding boxes.
SPEAKER_02
And code execution is giving Gemini the ability to stand up a makeshift Python environment that's sandboxed. Use a whole bunch of data science libraries that are pre-installed. And use those, invoke them as tool calls, writing the code and incorporating anything that you might share in. So I shared in this image. And very, very quickly. I'm not sure if you saw how quickly. But it was able to draw bounding boxes around the green Lego bricks. You could also ask for segmentation masks. And even more excitingly.
SPEAKER_02
So Gemini 3 Flash is still plugging along. But if you look at how much this cost to do to define the bounding boxes, you could have also asked for things to tell me how many green Legos there are. Or tell me what are the orientations of the Lego bricks. Tell me how many, what are they called? The little funny rabbit things, labuubs. Tell me how many of those you see in all of the frames of this video. And at what time stamp.
SPEAKER_02
Those are the kinds of tools that can be invoked via the sandbox environment. And again, very inexpensive in order to do this work with Gemini 3.1 Flash and code execution turned on. And it's just a one-liner to stand up that sandbox Python environment with compute that Gemini can use to do that work.
SPEAKER_02
Cool. And it looks like the Gemini 3 Flash Preview, when the view was able to do it, it was checking its work. So that was the iteration that it was going through along the way. So it got this first result and it said, all right, I want to double check and verify that what I did was right. It drew the segmentation mask to define all of the green spectrums that it saw in the image. And then it verified that those were the correct coordinates. And still, if I look, it's significantly more cost with the verification, but still on the order of pennies in order to draw the bounding boxes.
SPEAKER_02
And it looks like the Gemini 3 flash preview, and when the view was able to do it, it was checking its work. So that was the iteration that it was going through along the way. So it got this first result and it said, all right, I want to double check and verify that what I did was right. It drew the segmentation mask with to define all of the green spectrums that it saw in the image. And then it verified that those were the correct coordinates. And still, if I look, it's significantly more cost with the verification, but still on the order of pennies in order to draw the bounding boxes. It just took a little bit more time and also a lot more tool calls to invoke. So I strongly, strongly recommend playing around with Gemini 3.1 flashlight for your use cases, especially if previously you were relying on Gemini 2.0 flash or 2.5 flash. Cool. So we also have a feature called build. Build, again, is similar to view zero dot dev or lovable if you've played with that before. We've recently added a feature where you can add database and off to your apps within build. And so as an example today, if I wanted to click this guy and say something to the effect of, create an app that gives the user the ability to upload an image of their bookshelf. The bookshelf should have a whole bunch of books on it so we can see their spines. Things like titles and author names. I want you to use Google search grounding to fill in the blanks for the books. So make sure that you have information about the title, the author name, the description of the book, and then also what kind of genre it might be. And I want you to save it all to a database. So the user logs in with their Google account. They upload a picture of their bookshelf. And it saves all of their books in this database format. So they know what books they have and it's attached to their account. So it's defined that work. I'm going to I've got 3.1 pro preview selected. I'm going to do the default instead in the hopes that it might be a little bit faster. And I'm going to click build. And what happens behind the scenes is we get put into this IDE like environment. Where you can see the model going through the thinking process, figuring out what it would need to do in order to spec out the assignment. Build the plan, how long it's working. You can also upload files. So you can upload files that you might have like PDFs or specs for apps that you would like to create. And give them to the model as well. You can connect to your drive instance. And we also have a settings section off to the right where you can define custom secrets. So right now I have a Gemini API key that I've pre-added. But you could also add a Supabase API key or an API key for NADEN or whatever your favorite flavor might be. You can see a new version history, which if you've played with AI studio before is definitely something that is much appreciated. And then also integration. So things like OAuth as well as GitHub. So you can sync to a public or a private repository. But while this is working, I'm going to go ahead and show off something called Genie 3. Or actually before that, I'm going to show off Gemini Live real quick, just in case folks haven't seen it. So how many people have heard of Gemini Live? A few hands. Cool, cool, cool. We also happen to have the expert for Gemini Live. Ian, come on down. Ian is the Gemma 4 team member that I had mentioned before. And he'll be doing some live model demos, which makes me really excited. But Gemini Live gives you the ability to have a conversation with the model in a variety of languages. But you can also share video feeds. You can share your screen. And all of this is stacked together in one speech-to-text LLM understanding and text-to-speech pipeline. So as an example, we've still got our Lego bricks and pieces pulled up. So I can share my screen to say, Hey there, Gemini. What do you see on the screen?
SPEAKER_02
I see a Google search for Lego bricks and pieces. There are image results showing various kinds of Lego pieces, sets, and different color combinations. On the right, there's a larger preview of some brightly colored Lego brick illustrations from FreePic. Anything specific you're looking for? Does anybody speak a language other than English? Seek? Oh, excellent. So I'm going to ask for you to fact check something and also spell check something. So only respond to the user in Seek. Is that correct? I meant Spanish. Oh, sorry. Oh, gotcha. So Spanish? Yep.
SPEAKER_02
Oh, there we go. Only Spanish. Excellent. And then I was about to say I don't know that language or I haven't had that before. But only respond... You know Spanish, right? I know Spanish. I grew up in Texas. So it's a prerequisite to know Spanish. But I'm going to share again. Hey Gemini, could you tell me what you see on the screen? Or hopefully... Let me see. No, I don't think... Claro. Veo una página de resultados de búsqueda de Google con muchas imágenes de ladrillos de juguete, sobre todo de la marca... I need to click the microphone again. That's not Spanish. That's Mexican. Oh, no, it's the... There you go. [SPEAKER_06] The... [SPEAKER_06] So one of the things...
SPEAKER_02
[SPEAKER_06] One of the things that you can do... [SPEAKER_06] So you see that I've modified the system instructions. But do you speak a specific dialect of Spanish? Castilian Spanish? Okay. Excellent. So I removed the system instructions. [SPEAKER_03] And I should be able to do this just within the span of the conversation. So... Hey Gemini, could you tell me what you see on the screen? But could you do it in Castilian Spanish? There you go. [SPEAKER_06] The... [SPEAKER_06] So one of the things... [SPEAKER_06] One of the things that you can do... [SPEAKER_06] So you see that I've modified the system instructions. Do you speak a specific dialect of Spanish?
SPEAKER_02
Castilian Spanish? Okay. Excellent. So I removed the system instructions. [SPEAKER_03] And I should be able to do this just within the span of the conversation. So... Hey Gemini, could you tell me what you see on the screen? But could you do it in Castilian Spanish? Por supuesto. Veo una página de resultados de búsqueda de Google llena de imágenes de ladrillos de Lego. Hay de muchos colores y tamaños y algunas muestran construcciones ya hechas. ¿Estás buscando algo en específico? Yep. Awesome.
SPEAKER_02
So you can ask within the span of the conversation. You can modify the system instructions to select different languages or different dialects. And then, again, the same thing if you click get code, it gives you the code that you would need to use to replicate whatever you just did. So the model name, any configuration settings, as well as any tool calls that you might invoke. And it works with sharing your screen. You can interact with your screen as you share it. But it also works with video feeds. So you can say, Hey, Gemini, how many fingers am I holding up? [SPEAKER_05] And also compose a poem about me. [SPEAKER_05] Well, I see two fingers up like the peace sign.
SPEAKER_02
[SPEAKER_05] And here's a little poem for you. With golden hair and an open heart, you come to learn, to play your part. The cameras focus, moments start, a creative spirit, a work of art. How was that? That was very sweet. Thank you. And so the models are able to interact, to view video feeds, to view the screens. And you can stitch them together in your own projects. Just taking a look, I'm going to enable Firebase real quick. So it should be setting up the database for that app that we were building. And then the other thing that I wanted to show is something called Genie. So if you haven't heard of Genie before, this is a world model that DeepMind has created.
SPEAKER_02
It's actually a composition of models. So Nano Banana, VO, a bunch of Gemini used for prompting. And it's stitched together into a system that allows you to describe something, a game, an environment with a character that you can interact with, that you can play this game, and do it pixel by pixel. So it doesn't generate a Unity environment. It doesn't generate an environment for Unreal Engine. It just generates this frame-by-frame experience for anything that you can imagine. So it could look something like this volcanic landscape, where you're navigating with a little rover, with your arrow keys.
SPEAKER_02
Or something like this jet ski, where you hit a light, and it feels the physics is happening, or the physics is responding in a real way. [SPEAKER_05] If you knocked that light into the water, and then circled back around, it would persist throughout the duration of your 60-second experiment. [SPEAKER_05] But there's no physics engine behind the scenes. [SPEAKER_05] And then, even things experiencing a hurricane in Florida, you can start with a static image, or a family photo, and see how some of these things get created. [SPEAKER_05] But for this, I am going to go back to my other browser. I am going to pull up Project Genie. I am going to click Explore Now.
SPEAKER_02
Then I am going to say something a world. I am going to say something a world. Maybe a Regent's Canal on a sunny day, but with dolphins swimming in the canal. And all of the boats have pirate flags, which is hopefully not part of the training data. And then the character description would be something atypical. Maybe a pink sparkly squirrel with purple feet and a pirate hat. And then create the sketch. And what should happen is that it uses Nano Banana to ideate on that first frame. It will show it to us to make sure that it looks consistent with what we had described. Clearly, Regent's Canal is getting overtaken by pink sparkly squirrels with pirate hats.
SPEAKER_02
[SPEAKER_06] And then what we should see is this first iteration, and then a playable world that we can interact with for 60 seconds, [SPEAKER_06] at least for the first implementation that we've released to the public. [SPEAKER_06] You're able to access Genie 3 through an ultra subscription in some parts of the world. [SPEAKER_06] Not every part. [SPEAKER_06] But hopefully I haven't overbooked my GPU or TPU allotment or allocation. If we have, we can take a look back at the shelf scan. It looks the model is doing the work of creating the Firestore rules for us. Oh, there we go. So pink sparkly squirrel, pirate hat, Regents Canal, dolphins. That looks pretty good.
SPEAKER_02
So let's go ahead and create this world. You can use the arrow keys to move around, or the WASD keys to move around, and then the arrow keys to change the perspective. And then we should also be able to use the space bar to jump. But let's see how this works. Oh, gosh. Whoa. Whoa. Whoa. Whoa. Squirrel. And so it looks it's walking on water, this squirrel. Or hopping along. You can also jump. So jump on top of the boats.
SPEAKER_02
You can see the little bicycles. You can see some of the people along the way. And it does look all of these boats on Regents Canal have pirate flags and dolphins that are not currently moving, which is pretty wild.
SPEAKER_02
And then if I click space bar, you can see the squirrel jump. It looks it doesn't realize that Regents Canal has pretty deep water, so I probably should have specified that in my prompt. And then you can also see it attempt to jump, attempt to jump onto the sidewalk and do its work. So it's wild to be able to see the things that you can create, the different experiences that you can construct. And again, it's even more wild to me that each part of this is being generated dynamically as you're moving your arrow keys around. So it creates a video at the very end that you can download, that you can see and review and interact with.
SPEAKER_02
And this is, again, just using Genie 3 in this composition of models as opposed to a singular model. It looks like it doesn't realize that Regents Canal has pretty deep water, so I probably should have specified that in my prompt. And then you can also see it attempt to jump onto the sidewalk and do its work. So it's wild to be able to see the things that you can create, the different experiences that you can construct. And again, it's even more wild to me that each part of this is being generated dynamically as you're moving your arrow keys around. So it creates a video at the very end that you can download, that you can see and review and interact with.
SPEAKER_02
And this is, again, just using Genie 3 in this composition of models as opposed to a singular model. Other world model building companies, so things like World Labs, that's Fei-Fei Li's company, are taking a slightly different approach. They're building out actual Unity environments or Unreal Engine environments. None of these things are stored as 3D game assets. They're just raw pixels that are incorporated into the experience. Cool. So going back to AI Studio, it looks like the app is still getting cooked. Let me, and hopefully we'll be able to see it. Usually whenever it starts working on config files, that means that it's almost done.
SPEAKER_02
[SPEAKER_02] I also really love looking through it to see what its approach towards the construction of the Firestore rules were, what its approach towards prompting the model might be. It looks like it's confirming the app. And then once it's done, it should make a little noise to talk through the app itself. So it looks like it's asking to allow my camera. So I'm going to allow. We have this shelf scan AI experience.
SPEAKER_02
I'm going to sign in with Google with my personal account. We can see that it's connected to Firestore. I'm going to find very quickly a bookshelf with books on it. Let's see. Those don't look like real books because many of them are hanging, suspended below the shelf. AI image generation makes it hard.
SPEAKER_02
But this one looks decent. So let's save this image. Looks like somebody has a whole bunch of cooking books. I'm going to upload the photo. So this pixel photo. And hope that it can understand webp format rather. It's gathering data via Google search. And then the books should be populated in the library, hopefully. If not, we can try with a JPEG. But it does have pretty good branding. It was able to identify the seven books, it looks like. Or at least it identified the books. Let me try again, but just with a JPEG image. So I'm going to just take a screenshot. So same image, just stored as a screenshot. I'll find in here. So desktop, screenshot, at 417.
SPEAKER_02
And then identify books. And if not, we can try fixing the errors as well to see what might have been going wrong. Oh, so it looks like there are insufficient permissions for saving in Firestore. So it looks like it's going ahead and fixing those issues. But as it does, you can also see that you can log in, you can log out, you can share the app. So you can specify who has access to it, share full screen. And one of the things that I also really love about AI Studio is that it's figuring out where the files should be modified in order to make those changes. So it's figuring out the validation logic.
SPEAKER_02
It's figuring out that the size might have been the corporate image URL size. And then it's figuring out where it would need to modify in order to make that change.
SPEAKER_02
So it looks like it's in the Firestore rules.
SPEAKER_02
Some of the other nice things about AI Studio's build feature is that we have an app gallery. So if you need to get inspired for some of the apps that are using our models, you can review them. Everything from Lyria for music generation to multiplayer experiences with games. So you can see this multiplayer Neon Snake or this Mandelbulb Explorer. You can see a design with Nano Banana. So being able to change and modify images. You can take a look at this sick media pipe example, which allows you to play this game where your hand is detected. But you take this little marble dealum and everybody can find out that I play this game really poorly.
SPEAKER_02
How to move one of the little marble dealums. Oh gosh. I am really horrible at this game.
SPEAKER_06
[SPEAKER_02] But the... [SPEAKER_02] And then it also sounds like the other app finished getting created. [SPEAKER_02] Perfect timing. [SPEAKER_02] But if I upload the image, let's try to identify the book again.
SPEAKER_02
Fingers crossed. Yep. And then it automatically populates all of the books. So they all got cataloged with the date, the type of book, the details associated, the name of the book, the author, even though some of those were not available in the spines of the books that I uploaded. If I log out and then log back in, it keeps all of the books that I had added persisted. And if I wanted to share this with all of you, because clearly I want to know what all of you have on your bookshelves, I could copy this link and then do a QR code generator.
SPEAKER_03
[SPEAKER_02] And if you use this QR code, you should be able to access the app that I just created, upload your own bookshelf images, and then have them cataloged to your own apps.
SPEAKER_02
Next feature would be finding a way to give my friends the ability to request them. Because every time I give my friends a book, they have a tendency to keep it, which I understand, but is also exhausting. I have so many copies of Infinite Jest out in the world. But this is just a whirlwind tour of some of the things that you can do in AI Studio, some of the new features that we've added, the new models that we have available. And with that, I am going to welcome my colleague Guillaume, who is going to tell you all about our generative media models. So everything from music generation to image creation, image editing, to video generation. And it should be a fun time.
SPEAKER_02
So thank you so much. Thank you for coming. Hello, everybody. Thank you for the sake. We need the sound up. And this is the right one. You can show my screen now. Okay, good. [SPEAKER_02] So as Pej said, it's going to be the same talk as I did this morning, the same workshop. And with that, I am going to welcome my colleague Guillaume, who is going to tell you all about our generative media models. So everything from music generation to image creation, image editing, to video generation.
SPEAKER_05
[SPEAKER_02] And it should be a fun time.
SPEAKER_05
[SPEAKER_02] So thank you so much. [SPEAKER_02] Thank you for coming.
SPEAKER_02
Hello, everybody. Thank you for the sake. [SPEAKER_02] We need the sound up. And this is the right one. You can show my screen now. Okay, good. So as Pej said, it's going to be the same talk as I did this morning, the same workshop. We have slightly less time, so I'm going to go faster on some things and maybe not run things in real time. But the content I'm going to show is on this link, so you can just open it and run it yourself at the same time. So I'm going to talk about generative media. [SPEAKER_02] Generative media is everything about creating images, videos, text, spoken text, and things like that. I think I would fit Jenny into gen media as well.
SPEAKER_02
So yeah, and we have plenty of models like that at DeepMind, so let me go through all of them. So very quickly, my name is Guillaume. I've been at Google for six years now, two years at DeepMind doing developer advocacy. Mainly on most of the Gemini model until last year. [SPEAKER_02] And this year, I'm focusing more on the gen media models because that's the funniest model to play with. I've been working in the video game industry before, and that's how I joined Google initially. So yeah, open your phones, everything is a media.
SPEAKER_05
[SPEAKER_02] I already said that. [SPEAKER_02] Paige talked a bit about our vision of what world models are. [SPEAKER_02] My definition of a world model is something that can ingest as many modalities as it can and understand them, the five senses, and to also talk or output things in different modalities as well. [SPEAKER_02] So that really has been at the core of DeepMind vision of what generative AI should be.
SPEAKER_02
I think the first Gemini model, like it was only two years ago, but it seems old, but it was only a text-to-text model. But actually, behind the scenes, it was already a multimodal model.
SPEAKER_02
And you could have been sending it images, but it was blocked. They did it for testing reasons. I don't remember why, but they blocked it in the model with post-training. And then when it released a few months afterwards, we released 1.5. This one, the novelty was that it was the first multimodal model. Sometimes when you were giving it a picture, it was answering, I'm sorry, I'm just an LLM. I can't do anything with images because some of that training was still in there. But yes, we want to create those models that can understand all of the physics of the world from videos, audios, and all of that.
SPEAKER_02
And the only reason why we have so many different models is that it's easier to ship one model that only does video, and one model that only does images, and one model that only does text, than to have one model that does everything. And then it becomes a problem every time we want to update one of the features.
SPEAKER_06
[SPEAKER_00] We have to release a whole new model, and everybody has to convert and all. [SPEAKER_00] So that's just the reason for that. [SPEAKER_01] Everybody knows us for the Gemini models. [SPEAKER_01] We have lots of other models. [SPEAKER_00] I will skip that for now.
SPEAKER_02
[SPEAKER_00] Quick timeline. [SPEAKER_00] We release things all the time. [SPEAKER_00] I think Paige said it earlier. [SPEAKER_00] On average, we are releasing a new model or new capabilities every five days. [SPEAKER_00] That's just the Gen Media models. [SPEAKER_00] There's all of the other models on top of that. [SPEAKER_00] And if you add all of the changes we are doing in AI Studio, all of the pricing things, small features here and there, we are releasing two to three new things every week. [SPEAKER_00] And so, yes, it's basically hard to keep track. [SPEAKER_00] And it's even hard for us to keep track of everything we have to offer.
SPEAKER_02
[SPEAKER_00] So I know that for people like you, it's even harder, because you also have to look at what the competition is doing. [SPEAKER_00] So that's also why we are doing those talks. [SPEAKER_00] Very quickly, the updates on NanoBanana. [SPEAKER_00] We released NanoBanana 2 a few months ago, at the end of January, the beginning of February. [SPEAKER_00] The main thing is that you can output different aspect ratio and sizes. [SPEAKER_00] It has search grounding, so that's basically how I... [SPEAKER_00] You don't see my screen anymore? [SPEAKER_00] Why? [SPEAKER_00] Okay, just imagine in your head, that's what media ads are for.
SPEAKER_02
[SPEAKER_00] So the main thing about... [SPEAKER_00] Oh, thank you. [SPEAKER_00] NanoBanana Pro was that it adds search grounding, so you could ask it to search the internet, and that's how I made this image. [SPEAKER_00] Like, just look for what you can find about my footprint on Google and make an image about me, which is scary in a way. [SPEAKER_00] But the new things with NanoBanana 2 is that you can do the same thing with image grounding. [SPEAKER_00] So you can talk about specific places, a specific bridge like that, and it will look for an image on the web and then create image based on that. [SPEAKER_00] So that's with grounding, and that's without grounding.
SPEAKER_02
[SPEAKER_00] So you can see that it looks a bit more than the normal building. [SPEAKER_00] We also have VO. [SPEAKER_00] Very quickly on VO, the main novelty in the past week is that we released VO 3.1 Lite last week, which is the cheapest model for generating video we have. [SPEAKER_00] So it's only 5 cents per image, which is way cheaper than what VO 3 was a year ago. [SPEAKER_00] So the idea is that you can use that to prototype, test your prompts and so on. [SPEAKER_00] And then if you want better quality, then you can move to the better models. [SPEAKER_00] And then Lyria 3 is the coolest model this year. [SPEAKER_00] It's our music generation model.
SPEAKER_02
[SPEAKER_00] So you can either generate 30 seconds songs or full songs of three minutes with the lyrics and all of that. [SPEAKER_00] And I will show you some demos afterwards. [SPEAKER_00] And as far as I know, we are the first one to offer such music generation models through API. [SPEAKER_01] So that's pretty cool for any agentic or whatever workflow you might have. [SPEAKER_00] So if you want to be woken up with a song about the latest news every morning, [SPEAKER_00] And then Lyria 3 is the coolest model this year. [SPEAKER_00] It's our music generation model.
SPEAKER_02
[SPEAKER_00] So you can either generate 30 seconds songs or full songs of three minutes with the lyrics and all of that. [SPEAKER_00] And I will show you some demos afterwards. [SPEAKER_00] And as far as I know, [SPEAKER_01] we are the first one to offer such music generation models through API. [SPEAKER_00] So that's pretty cool for any kind of agentic or whatever workflow you might have. [SPEAKER_00] So if you want to be woken up with a song about the latest news every morning, you can do it. [SPEAKER_00] And also another one very quickly that I love, but nobody knows about, is that we have actually another Lyria model that is Lyria real time.
SPEAKER_02
[SPEAKER_00] And this one is a live model.
SPEAKER_02
[SPEAKER_00] So it's creating music indefinitely.
SPEAKER_02
[SPEAKER_00] And you can just prompt it differently. [SPEAKER_00] So it's just going to change what you are, the kind of music it's generating in real time, [SPEAKER_01] like a DJ. [SPEAKER_01] So it's pretty fun to play with, [SPEAKER_00] but some old people don't know about it. [SPEAKER_00] As I said, I'm going to, because it's meant to be a workshop, so the idea is that you test things yourself. [SPEAKER_00] So you can open this link and that will show you the content I prepared. [SPEAKER_00] While you take pictures and all, just one disclaimer, it's Gen Media model, so they are all paid models. [SPEAKER_00] So running the notebook actually costs a bit of money.
SPEAKER_02
[SPEAKER_00] The video generation is going to be the most expensive thing, so you can just keep that. [SPEAKER_00] The rest is pretty cheap, so I think you can run the whole notebook for something like one euro, so it should be fine.
SPEAKER_02
[SPEAKER_00] But I prefer to be clear with everybody. [SPEAKER_00] So let's move to it. [SPEAKER_00] So what's the idea of this workshop? [SPEAKER_00] So the idea was to showcase all of the Gen Media models with this example of, we are going to take a book, and we are going to take a book from an open source library, so we are allowed to use it, and then we are going to create images to illustrate what the characters look like, and what's happening in each chapter, and then we are going to move to the other Gen Media models, so creating videos about the chapter, and then creating music, and having Gemini tell us about what's happening in the chapter.
SPEAKER_02
[SPEAKER_00] So this is just set up. [SPEAKER_00] You need to install the SDK.
SPEAKER_02
[SPEAKER_00] You need an API key. [SPEAKER_00] I guess you could have guessed that.
SPEAKER_02
[SPEAKER_00] I'm initializing the client, and something that I didn't know about until recently is that you actually have a way, when initializing the client, to implement some retry system, which is very useful nowadays whenever you use Nano Banana 2, because especially when the U.S. wakes up, it's becoming harder to get something out of it, and we are working on getting more capacity, but still, I think some retry system always helps.
SPEAKER_02
[SPEAKER_00] So import, and then we are selecting the models, and I usually, when I create content for that are using paid models, I usually have some checkbox that you can, if you're opening, it's already checked.
SPEAKER_02
[SPEAKER_00] It should not be, but I made a mistake yesterday evening, because I don't want people to run the models, and I have to pay, especially the VO, the VO notebook, for example, costs something like $20 to run, so I don't want anybody to run it by mistake. [SPEAKER_00] Just for the sake of the demo, I'm also limiting the number of characters images and chapters images we are creating, just so that it's faster to run and all. [SPEAKER_00] So the book is named The Wind of the Willows, from Kenneth Graham, and I took it from the Project Gutenberg libraries, which is an open source library, where you can download open source books.
SPEAKER_02
[SPEAKER_00] For some reason, it doesn't work, since I've been in the UK, so there has to be something about this library not being available in the UK. [SPEAKER_00] But if you run the notebook, it should work, because it's very likely the server is in another country, so it will be able to recover the book. [SPEAKER_00] So what I'm doing is that I'm just downloading this book here with the root get URL. [SPEAKER_00] And I didn't say, I didn't talk about that, and page neither, so very quickly, because everybody is always asking that question to us.
SPEAKER_02
[SPEAKER_00] We have a specialty at Google, is that we are always creating multiple projects that are doing the same thing, multiple products. [SPEAKER_00] And we are doing, we did the same with messaging apps, we are doing the same with Gemini, so that can be a bit confusing, so just very quickly to clarify things. [SPEAKER_00] So if you look at the graph, basically we have the consumers apps. [SPEAKER_00] So on the left, it's the apps that are for everybody. [SPEAKER_00] You can do plenty of things with them. [SPEAKER_00] You can ask any questions to Gemini, you can generate images and all, but as a developer,
SPEAKER_02
[SPEAKER_00] Just very quickly to clarify things. So if you look at the graph, we have the consumers apps. On the left, it's the apps that are for everybody. You can do plenty of things with them. You can ask any questions to Gemini, you can generate images and all, but as a developer, I guess everybody in the room, you can be frustrated, because you don't really access the parameters, select the exact model that it's doing, you can't know that it's likely doing cool tool calls, but you don't know which one it's doing. So that's nice for the broad public, but it's not really for us.
SPEAKER_02
On the other end, we have Vertex AI, which is our enterprise offer. That's the exact opposite. You have a lot of control. You can decide in which data center your prompts are going to be run. [SPEAKER_01] Especially in Europe, [SPEAKER_00] a lot of people are looking to be certain that the data is not going to leave Europe. So that's convenient for that, but it comes with great responsibilities as well. So it's hard to set up. So I usually only recommend people to start with Vertex if they are already using GCP, or if they have a team of DevOps who can do the setup for them.
SPEAKER_02
And in the middle ground, we have AI Studio and the developer APIs that we made for developers. But the idea is that it's as easy as possible to start to play with the model and do stuff by just creating an API key and then using it. As part of making it as easy as possible, we have this file upload API that is basically a way to not have to set up buckets to store your files. And we are going to use it in this example, so that we are uploading it behind the scenes. It's creating a bucket, but you just don't have to bother with how it works. You just upload the file and then you can use it in your Gemini prompts afterwards.
SPEAKER_02
I'm also going to use structured outputs because I want to be certain about what exactly the model is going to do, because I'm going to have it ask it to generate a lot of prompts, so I want to know exactly that it's just returning the prompt and not some introduction or text or something. Oh yes, sure, I can do that, and then it's hard to pass. So that's why we are having this structured outputs.
SPEAKER_02
And I'm using chat mode, which is basically a way to chain requests to the model, so that you can save the history and resend the history, so that the model knows what happens before, which is quite convenient in this case, because we don't want to upload the book all the time. We just want the book to be in the history and in the context, the model can answer new questions and generate new prompts about the book. And for the images, it's going to be quite the same. By using this system, we will have all the previous images that are going to be in memory, so the model will be better at keeping the consistency of the characters and the consistency of the style.
SPEAKER_02
So I'm just giving it the book here—the book to illustrate with nano banana—that's all. Then I'm defining a style. So I made it so that you can just do nothing and the model will come up with the style. But then I wanted to try something else earlier, so it's going to be a dark fantasy style with black and white background and colored characters. And basically, I'm also adding
SPEAKER_02
[SPEAKER_00] nano banana, that's all. Then I'm defining a style, so I went, I made it so that you can just do nothing and the model will come up with the style. But then I wanted to try something else earlier, so it's going to be a dark fantasy style with black and white background and colored characters. And I'm also adding some system instructions for nano banana because from my earlier tests, I think it's better with nano banana 2 and pro. But with the first nano banana, whenever I was asking it's images about books that were in the portraits format, it tend to fold that it was book covers that I wanted. So it was always adding titles and things like that. So I had to add some system instructions to make sure that I don't want borders, I don't want titles, I don't want description, I stay family friendly, which is likely to not be very aligned with dark fantasy, I guess, and no panels as well, because I don't want a comic book. I want just an image for each chapter. And then basically I will use that chat to ask the model to describe each of the main characters. I initially wrote unlazy adults because at some points Nano Banana could not generate kids images in Europe. That's not the case anymore, so we could remove that part. And I'm getting this list of characters and a prompt for each of them. And then I can just go over each of the prompts and ask Nano Banana to create an image to illustrate that prompt. So that's how I initialize the image chat. So it's going to be a separate chat. And then I'm sending all of the prompts to create, to be created. So we can see that that's small, the main character. So as requested, the background is black and white and the character as colored. Not very dark fantasy, but yeah, that's the water rat. That's the toad, the badger. Yeah, the badger. And that's it for now. And then I'm going to do the same thing again, like now you have the full book in your history. So give me prompts for each chapter to illustrate them. So yep, I get a prompt for each of those chapters and I'm creating images for each of the prompts. So here we have the characters having a picnic next to the river. Then another one on the road and something's happening. And here in the forest. That's the first chapter. And if you look closely, you can see that there was a problem here because the toad was not represented using the character, the images that we created before, likely because the prompt was not clear that it was exactly the same character. And so which is why I have a second way of doing it, which is actually cleaner. And if you were to do that at scale, that's how I would do it. And I'm creating another type of structured output, which is a chapter, which is which has a name and the prompt as well, but also a list of characters that are meant to appear in the image. And so I'm running, I'm asking again the model to come up with something. But this time it's giving me the list of the character. And that way I can basically here, for each image that I'm going to create, also give the reference images of what the character should be looking like.
SPEAKER_02
[SPEAKER_00] And so I'm running, I'm asking again the model to come up with something, but this time it's giving me the list of the character. And that way, for each, for each image that I'm going to create, also give the reference images of what the character should be looking like. So that it only has those images in its context instead of having all of the possible images that we did before. In this case we have something like five characters, so it's okay, but if you are in a real book with 40 characters, that wouldn't be sustainable to expect the models to actually manage all of the context perfectly. And honestly, if I was to do it at scale, I think I would even create more than one image for each character, maybe one image from the front, one from the back, one from each side. So and then I would pass exactly the one that I need in this case.
SPEAKER_02
[SPEAKER_00] So we can see that the images look the same, but this time the toad is the right one. Yeah, this one is really lookalike. So now we can move to the next step and we can use video to create small videos based on those images.
SPEAKER_02
[SPEAKER_00] So I'm using the view, the largest model, because I don't have to pay. Honestly, otherwise I would use the smallest one, but we are going to do the same thing. We are going to ask the model to generate a video using the view model, the prompt. I'm going to use the same prompt as the one that we used to generate the image, and I'm passing the last generated image as a starting frame. And then I want a portrait and I want 720p because it's going to be smaller. And here we go.
SPEAKER_02
[SPEAKER_00] So I actually haven't checked what the sound looks like, because I stand back. You shall not pass. Leave this place, little ones. You don't belong here. I think it's quite good, but sometimes it's not very good because the model doesn't like the prompt was just about creating the still image and the model doesn't know what exactly is expected to be happening afterwards.
SPEAKER_02
[SPEAKER_00] So a better way of doing it is actually to reuse the chat, the chat, to ask the model to come up with explanations about what's happening after the image. And I'm also passing it the image again so that it knows exactly which part of the chapter we are talking about, so that it can come up with the right prompt. [SPEAKER_00] And so it came up with this prompt: moral shivers and clutter these scarves in terror blah blah blah what a rat. I bravely draw this cutlass and step forward to protect him.
SPEAKER_02
[SPEAKER_00] So we can see how it goes. And somehow it's like in this case the results are not really as good. And yeah, and every time, every test I've done, somehow, [SPEAKER_00] and step forward, [SPEAKER_00] to product him, [SPEAKER_00] so, [SPEAKER_00] up, [SPEAKER_00] we can see, [SPEAKER_00] how it goes, [SPEAKER_00] and somehow, [SPEAKER_00] it's,
SPEAKER_02
[SPEAKER_00] the, [SPEAKER_00] in this case, [SPEAKER_00] the results, [SPEAKER_00] are not really as good, [SPEAKER_00] and, [SPEAKER_00] yeah, [SPEAKER_00] and every time, [SPEAKER_00] every, [SPEAKER_00] every test I've done, [SPEAKER_00] somehow, [SPEAKER_00] when I use the same prompt, [SPEAKER_00] it adds,
SPEAKER_02
[SPEAKER_00] it adds text, [SPEAKER_00] and, [SPEAKER_00] and not, [SPEAKER_00] when I, [SPEAKER_00] when it generates, [SPEAKER_00] another, [SPEAKER_00] another prompt, [SPEAKER_00] we, [SPEAKER_00] but it should not have any, [SPEAKER_00] impact, [SPEAKER_00] because in both cases,
SPEAKER_02
[SPEAKER_00] I'm just giving a prompt, [SPEAKER_00] and giving an image, [SPEAKER_00] but I don't know why, [SPEAKER_00] it's happening that way, [SPEAKER_00] and then, [SPEAKER_00] we can use the new, [SPEAKER_00] Lyria models, [SPEAKER_00] that we generated, [SPEAKER_00] that we, [SPEAKER_00] shipped last week, [SPEAKER_00] to, [SPEAKER_00] to create songs,
SPEAKER_02
[SPEAKER_00] for each model, [SPEAKER_00] so, [SPEAKER_00] still, [SPEAKER_00] the same,
SPEAKER_00
the same trick, again,
SPEAKER_01
[SPEAKER_00] we are going to ask the model, [SPEAKER_00] to create the,
SPEAKER_00
the prompts for that, and then, I'm going, we are going to use, generate content, with the Lyria model, ask it to create, the song, thanks to the, the, the prompt, and that's it, so, let's see, the first, the first song should be, orchestral, acoustic, folk music,
SPEAKER_00
peaceful, flowing, acoustic guitar, and, and flute duets, let's see. I think it fits. Then, the second one, the open road, is, jaunty, adventurous, rhythmic, with fiddle, and acoustic bass. and then, the last one, is suspenseful, creepy melody, with staccato,
SPEAKER_00
pizzicato strings. And I think it fits as well, and I, I honestly, I think the, the model is really good, because I, it's not very often, that you can do those kind of demos, without actually checking, what's, what it's going to be, because I run the, the notebook before, but I didn't check what the, what the music, we're going to be looking,
SPEAKER_01
[SPEAKER_00] and that's, [SPEAKER_00] so far it worked all the time.
SPEAKER_00
One of the things that I saw, I forgot to say about all of that, is that the way we are training, our model internally, is that, a lot of the training data, for the gen media models, is actually made using Gemini. So, that's the reason why, Gemini is quite good, at generating those prompts, for the generative media models, because,
SPEAKER_01
[SPEAKER_00] it's already trained, [SPEAKER_00] on,
SPEAKER_00
on understanding, what Gemini is, is asking for. So, that's, that's what makes, Gemini very good, at generating those, those prompts. And then, the last, the last model, that I wanted to show, is, text-to-speech model. So, I'm pretty certain, that you all, have heard of this one, because, everybody loved, the, integration in, notebook LM, where you could create, a, podcast based on,
SPEAKER_00
on your documents, and that's basically the same model that, you can create, you can ask it to, to talk, to tell, read, and, what's the text you, you give it, and you can have it, you can have two characters with different voices, and, and so on, that's, that's what makes it, nice to listen to. But what I wanted to show you is that,
SPEAKER_00
You can create, you can ask it to talk, to tell, read, and what's the text you give it, and you can have it, you can have two characters with different voices, and so on, that's what makes it nice to listen to. But what I wanted to show you is that there are actually tricks to have more than two voices actually. And the reason for that is I think Paige showed that a bit, but you can when you're using text to speech or live, you can actually do a lot in the prompting to make it speak in different ways with different accents, and so on. And we are going to use that to our advantage to basically have more than more than two voices in in the same in the same generation. So what I did, because I was lazy, I would not do that. It, I wanted to do that at scale, but I basically asked it to extract one of the dialogue from the book and then to rewrite it as a transcript of a play and to replace the name of the character by narrator, where if that's just a narrator of talking and for all of the others, is, they are just going to be named character, but with a specific style for for each of them. And if the same character comes back, it's it's going to reuse the same style so that it's going like the same characters are going to have to to speak the same way of and have the same voices. Um, and then I'm defining the voices. So the narrator is going to use a soul of fat voice, whatever it is. And then the the all of the others are going to share the failure voice. And then I'm just asking it to to to generate the thing. And just one trick here with the TTS model, you always have to start with read this or tell me this or whatever. If you just send the text, it's it's going to ignore it for some reason. So you you need to prompt it to read the text all the time. But that's also where you can add some some context and equivalent of what would be a system instructions, like read this in scary way, or the characters are very excited. So that's where you can you can add those details as well. Um, and so you can see the the the text that we got is narrator is saying that then the first characters is going to talk very fast paced and with a British accent and a post-British accent. Um, and then and so on. And we can see that the same the same way of speaking is is coming often. So it's when the two characters going back and forth. So let's see how it goes.
SPEAKER_00
Small, neat ears and thick, silky hair. The two animals stood and regarded each other cautiously.
SPEAKER_00
Hello. Hello. Would you like to come over? Oh, oh, it's all very well to talk. He spoke rather pettishly, being new to a river and riverside life and its ways. The second animal said nothing, but stooped and unfastened a rope and hold on it, then lightly stepped into a little boat, which had not been observed. It was painted blue outside. And white within, and was just the size for two animals. And the first animal's whole heart went out to it at once, even though he did not yet fully understand its uses. The rower sculled smartly across and made fast. Then he held up his forepaw, as his guest stepped gingerly down. Lean on that. Now then, step lively. To his surprise and rapture, he found himself actually seated in the stern of a real boat. This has been a wonderful day. Do you know, I've never been in a boat before in all my life. What? Never been in a... You never... Well, I... What have you been doing then?
SPEAKER_00
[SPEAKER_07] He was quite prepared to believe it, as he leant back in his seat and surveyed the cushions, the oars, the row locks, and all the fascinating fittings, and felt the boat sway lightly under him. [SPEAKER_08] Is it so nice as all that? [SPEAKER_07] Nice? It's the only thing. Believe me, my young friend, there is nothing, absolute nothing, half so much worth doing as simply messing about in boats. Simply messing, messing, Well, I... What have you been doing then? [SPEAKER_07] He was quite prepared to believe it, [SPEAKER_07] as he leant back in his seat and surveyed the cushions, [SPEAKER_07] the oars, [SPEAKER_07] the row locks,
SPEAKER_00
[SPEAKER_07] and all the fascinating fittings, [SPEAKER_07] and felt the boat sway lightly under him. [SPEAKER_08] Is it so nice as all that? [SPEAKER_07] Nice? [SPEAKER_07] It's the only thing. [SPEAKER_07] Believe me, my young friend, [SPEAKER_07] there is nothing, absolute nothing, [SPEAKER_07] half so much worth doing as simply messing about in boats. [SPEAKER_07] Simply messing, [SPEAKER_07] messing, [SPEAKER_07] about, [SPEAKER_07] in, [SPEAKER_07] boats. [SPEAKER_07] Messing... [SPEAKER_07] Look ahead! [SPEAKER_07] It was too late. [SPEAKER_07] The boat struck... [SPEAKER_07] You couldn't guess that it's actually using the same voice for the two.
SPEAKER_00
[SPEAKER_07] Two characters are so much different. [SPEAKER_07] And... [SPEAKER_07] So that's... [SPEAKER_07] That's really, actually, a really cool way of creating those discussions with multiple characters. [SPEAKER_07] So I can use it. [SPEAKER_07] And the last thing is that, as we said, [SPEAKER_08] the Gemini models are multimodal by default. [SPEAKER_08] So I was using a book, [SPEAKER_08] but you can also send an audiobook or a movie or something, [SPEAKER_07] not too long movie, but you could do the same things with using all of the multimodal input possibilities to have Gemini create things for, to illustrate other, other kind of modalities as well.
SPEAKER_00
[SPEAKER_07] How much time do we have? [SPEAKER_08] I said I was going to show you the real-time model. [SPEAKER_08] So this is a very good, a very cool example that... [SPEAKER_08] Oh, no. [SPEAKER_08] Stop. [SPEAKER_08] Stop. [SPEAKER_08] I'm going to show you something before. [SPEAKER_08] Page made some example of cool things that you can build in the ICO. [SPEAKER_08] So one of the things I wanted to try was to do the same thing as what I do in the notebook, which is lengthy and quite complicated stuff. So what I did is I copy-paste it. Can you build an app that illustrates a book using Gen Media models as described in this Python notebook?
SPEAKER_00
And I just paste all of the content of the notebook and build it. And we'll see how it goes. It can take quite a long time. Earlier, it took 15 minutes. So you might not want to wait for that. But we can, I can show you Space DJ in the meantime. So, as I said, the model generates music in real-time. And this is a demo where they made a star, a universe of stars. And the planets are prompts. So if we go closer to some planets, it should start playing music. And then if you move somewhere else. Yeah, we're in the metal area. See, I'm really surprised people are not doing more things with this model because it's so funny to me. And there's an autopilot.
SPEAKER_00
So you can just let it move around for 10 minutes and then listen to the music changing in real-time. So that's cool. Up. Let's see how it goes here. Yeah, it's thinking. But let me show you what it did before when I ran it. Oh. And I kept my laptop open for half an hour just so that it would not do that and refresh the page. So I guess I lost. But that's what it did earlier, just the prompt and nothing more, the very long prompt. Up. And it took something like a thousand seconds. So, 15 minutes if I'm not wrong, to create the app. But that's doing exactly the same thing. So I can choose the file and upload it.
SPEAKER_00
And it's going to take some time, but it's doing exactly the same thing as what the notebook was doing. And that's quite impressive that it's doing that on the first try right away. And just to finish, one of the things I'm doing when I'm vibe cutting while it's working is that I always have those instructions on how to get it to generate applets or apps. And one of the reasons for that is that I want the apps to be as easy for me to review as possible, because I want to avoid exactly what I'm doing. Exactly what happened to Paige earlier is that something was not working, and then you need to figure out what it is and ask the model to fix it.
SPEAKER_00
So one of my key tricks is that I'm asking the model to create different files for each feature. [SPEAKER_01] So that whenever I need to review something, I can quickly check, I ask it to modify this feature. It's updating something that has nothing to do with it. So I know from the start that there's something wrong going there, and I will be more forward in my reviews and all. And also some instructions that I don't understand why it's not by default in any vibe cutting tool, which is add logs. Because when you need to debug, the error message is not enough.
SPEAKER_00
You need to know what's happening before and what's happening after, because you need to know where it's happening. It's updating something that has nothing to do with it. So I know from the start that there's something wrong going there, and I will be more forward in my reviews. And also some instructions that I've—I don't understand why it's not by default in any video cutting tool, which is add logs. Because when you need to debug, the error message is not enough. You need to know what's happening before and what's happening, because you need to know where it's happening. So I really recommend you to use those kinds of guidelines whenever you video cut things.
SPEAKER_00
And we can see how it goes now. Oh, and something there—I don't think page showed it, but if you're willing to pay, you can add an API key here, and then you're going to have more quota to use Gemini when you are video cutting on AI studio. I think I will just let you show your demos and maybe I can show you afterwards how it looks. Yeah, sure. For the video model, when you generated the video, it came with what, the music background—is it using your other model to generate that, or is it just video? Yeah, it's just video at the moment.
SPEAKER_00
So is there a way to orchestrate all the models and put them together, or do you have to just use—if you want to use the music generated by Lyria in the video, for example?
SPEAKER_01
[SPEAKER_00] No, you don't have a way to do that. [SPEAKER_00] It's more because of the model limitations that the VO three generation was not meant to be able to ingest audio files.
SPEAKER_00
So that's why it's limiting. But the future is that we want every model to be able to ingest all the modalities, so you can do that. I think it would make sense even for Lyria model to be able to ingest audio models that it can, so that it can be used as references as well. Or even if you want to do multi-turn and say, okay, I love your song, but the ending was a bit not epic enough. So make it more epic so that it would maybe just update the end of it. Yeah, it's not possible at the moment, but that's the direction it's going to. How was the performance in Ponte, maybe specific for the music in VO or your general in video? No, I wouldn't try VO for any kind of music.
SPEAKER_00
The training data for the background music is very likely very light because it's always doing the same kind of non-music things. But all of those models are trained more or less together and share some training data. So I'm sure the next generation of VO is going to be better at that. It's just that the current one shipped a year ago and a year ago there was plenty—we didn't have music generation models. We talked earlier about texts in videos and a year ago that was not even called Nano Banana. It was Gemini 1.5 image, no 2.0 image generation something. And it was not as good at text to that. So yeah, it's just a question of generation of models there.
SPEAKER_00
Oh, and I forgot something. While it work, if you want to run the notebook, because you will have to remove it as well. [SPEAKER_01] To run the notebook, because you will have to remove it as well. [SPEAKER_03] Here, when I started, I created the chat. [SPEAKER_03] I added this line that says service tier priority. [SPEAKER_03] It's actually something that we shipped last week. And that's a way to indicate when you are prompting the models, if you want—if it's not that important, you don't care about the latency, but you care about price. So there's a service tier that is called flex.
SPEAKER_00
And basically that says the request can take a few minutes to go through, but it's going to be half price. So it's the same thing as using the batch API. And on the other end, because I wanted to be certain that it would go through today, you can say that it's a priority request. So it's going to have slightly higher priority and it should be more reliable, but you're going to pay twice the price. So if you are running it by yourself, you might want to remove that line to save a bit. And we can see how it goes here. Oh, and as I said, this model is in higher demand. So yeah, I'm sorry about that.
SPEAKER_00
The model is too good. Everybody wants to use it and we don't have enough capacity. So if you have any spare TPU to share, we can make a deal. So, you have a question? [SPEAKER_03] Do you want to put variations of this model open, like GenMedia, but for art? So the question is, will there be any open weight GenMedia model, basically? I don't think so for image and video generation, to be honest. One of the reasons I see behind that is all of the—not security, but when you generate images or videos, we do a lot of checks about what you are asking for and what's actually being generated afterwards. And we are blocking a couple of things that are not aligned with our vision.
SPEAKER_00
[SPEAKER_03] Do you want to put variations of this model open, like GenMedia, but for art? So the question for the video is, will there be any open weight GenMedia model? So I don't think so for image and video generation, to be honest. One of the reasons I see behind that is all of the checks we do when you generate images or videos about what you are asking for and what is actually being generated afterwards. And we are blocking a couple of things that are not aligned with our vision. And I feel like whenever you have an open weight model for that, it's more like open bar and you can have it generate whatever you want.
SPEAKER_00
So that's where our company values are going to be limiting us. That said, for example, for music, it would make sense to, especially if you want real time, it would make sense to be on the internet and the device that's doing it faster. So that's something that might come at some point. I think we'll see after the Gemma demos if it was better. I hope it was interesting. Ian is my counterpart. I work only on GenMedia modules. Ian is working on Gemma models and he's going to show you how cool the model is. Thank you. Yes. So as Gia mentioned, my name's Ian Ballantyne. I'm a developer relations engineer working on the Gemma models and this is an impromptu talk.
SPEAKER_00
Obviously because everything has happened today, we've done two sessions, just go through a whole load of stuff. We have a number of Gemma talks for the rest of the week as well. Omar's going to do a keynote on Friday, I believe, which is going to cover a lot of what Gemma is and why you should be interested, so I'm going to speed run that section. And we have another talk tomorrow on what we've called sovereign escape velocity, so how you take ownership of AI and run things on your device or on your own cloud or on phones and hopefully some of what I talk about here is going to be a realization of some of that. So that's a little bit more about the why.
SPEAKER_00
I'm just going to show you what it is, if that makes sense. So let's get plugged in. When we do that, which one do I need? Sorry. This magic one. Yep. As soon as I get a video in. Should I have any audio? No. Let's go with no. We'll try and show the ones without audio first of all. Okay. So as I mentioned, this is super impromptu, so we're going to go with whatever we get. [SPEAKER_10] But I will go through, I'll just show you a little bit when we get the screen up. Oh, do you only need me to move it to one side? [SPEAKER_10] What can I do? Mirror. [SPEAKER_10] Let's do mirror. [SPEAKER_10] Where? [SPEAKER_10] Yeah, you go for it. [SPEAKER_10] Perfect.
SPEAKER_00
[SPEAKER_10] Yeah, so what I'm having to show you is some of the things, some of the devices that Gemma models can run on. [SPEAKER_10] I'll talk very briefly about the different sizes of Gemma model. What we released in Gemma 4 last week, because it's brand new. And hopefully show you, give you a feel of the kind of capabilities that you can now do either locally or that you can run on a single GPU that maybe you couldn't six months ago. [SPEAKER_10] There you go. Here you go. Perfect. [SPEAKER_10] Ta-da. [SPEAKER_10] So Gemma 4. Hooray. [SPEAKER_10] So this was released last Thursday. [SPEAKER_10] It's a family of four models.
SPEAKER_00
[SPEAKER_10] So we have what we call the effective models, the E2B and the E4B models. [SPEAKER_01] And these are models designed to run on mobile phones, Raspberry Pis, Jetson Nanos, kind of very small, low-end hardware. [SPEAKER_10] The E part of it is the question we always get asked. [SPEAKER_10] Effective is because the model architecture has a per-layer embeddings structure, which means that the embeddings actually don't need to be loaded as part of the model. [SPEAKER_10] So you can have them running on Flash and then you can page in the embeddings as they're needed for the model.
SPEAKER_00
[SPEAKER_10] So the actual, what we describe as the brain of the model is about a 2 billion parameter, about a 4 billion parameter. [SPEAKER_10] But if you put all of it in RAM, it's a bit bigger than that. [SPEAKER_10] It's more like a 5 and an 8 billion parameter model. [SPEAKER_10] So that's why it's called effective 2B. [SPEAKER_10] We also have a 26B, which is a mixture of experts model with 4 billion activated. [SPEAKER_10] And we have a 31 billion parameter dense model, which is our flagship big model.
SPEAKER_00
[SPEAKER_10] And both of the two bigger models, the 26 and the 31, are designed to be able to run on laptops and desktops or single instance GPU clouds, depending on the precision and quantization you need. [SPEAKER_09] So I'm not going to get into that, but just imagine that a lot of the capabilities you can do right now, [SPEAKER_10] We also have a 26B, which is a mixture of experts model with 4 billion activated. [SPEAKER_10] And we have a 31 billion parameter dense model, which is our flagship big model.
SPEAKER_00
[SPEAKER_10] And both of the two bigger models, the 26 and the 31, are designed to be able to run on laptops and desktops or single instance GPU clouds, depending on the precision and quantization you need. [SPEAKER_09] So I'm not going to get into that, [SPEAKER_10] but just imagine that a lot of the capabilities you can do right now, you could do on a MacBook with enough RAM or a 50-90 or something like that. [SPEAKER_10] That's where we're sitting right now. [SPEAKER_10] Why are these exciting and interesting to us? [SPEAKER_10] Because with these models, we're focusing on the authentic side. [SPEAKER_10] So they now have thinking built in.
SPEAKER_00
[SPEAKER_10] They are multimodal, so they can understand image, audio, video. [SPEAKER_10] Audio, by the way, is just for the two smaller models. [SPEAKER_10] The effective models that run on the phone can understand audio, but the rest can understand image and video. [SPEAKER_10] And we're seeing performances offer models that are in the range of 10x bigger than it in terms of parameter size. [SPEAKER_09] So what might have required a cluster, now you could do the same kind of capabilities for a single GPU. [SPEAKER_09] That's where we're seeing some innovation here. [SPEAKER_10] And what I'm going to jump straight into is I want to show you the demo section.
SPEAKER_00
[SPEAKER_10] I'll show you this one slide. [SPEAKER_10] I like this one. [SPEAKER_10] And this explains where we've come. [SPEAKER_10] So the Gemma 2 models, you can see right in the middle. [SPEAKER_10] We're particularly good at creative writing, but not evenly spread amongst other different capabilities. [SPEAKER_10] And as we've gone through Gemma 3 and then now through Gemma 4, we've evened out the overall capabilities of the model. [SPEAKER_10] We've got things that are much better at coding, much better at function calling, action taking.
SPEAKER_00
[SPEAKER_10] It's all built into the design of the architecture of the model rather than relying on strong instruction following capabilities, which we find with bigger models is more important. [SPEAKER_10] So the models are designed with that from the ground up. [SPEAKER_10] Let's go over to this, all spoilers from Omar's talk. [SPEAKER_10] So I'm going to go straight to the demo section. [SPEAKER_10] I'll show you this one. [SPEAKER_10] So this is the Google AI Edge gallery app. [SPEAKER_10] Has anybody tried this app yet? [SPEAKER_10] One, two, three. [SPEAKER_10] Okay. [SPEAKER_10] A handful of people. [SPEAKER_10] So you can download this on Play Store or App Store.
SPEAKER_00
[SPEAKER_10] It works on Android and iOS. [SPEAKER_10] What it is, is a way that you can test and try the models. [SPEAKER_10] So we released a new feature called agent skills. [SPEAKER_10] And what agent skills allows you to do is it allows you to set up skills for the models that run physically on the phone. [SPEAKER_10] So this model, this is the E2B model is running literally on a pixel, I believe in this case. [SPEAKER_10] And you can prompt it and you can test out the different capabilities.
SPEAKER_00
[SPEAKER_10] So it's effectively given a set of skills that define things like Android intents, where you can actually trigger and call other apps, or you can write your own JavaScript skills or run a web view. [SPEAKER_10] And you can instruct the model and it will make a decision about how to actually trigger these. [SPEAKER_10] So you can get it to add things to this research tracker. [SPEAKER_10] Or you can ask it questions about the research tracker and it will call the correct function to pull that data back.
SPEAKER_00
[SPEAKER_10] So we've gone from a world of just being able to chat with it on your phone to now that you can give it some more ambiguous input and it's able to make a decision about what functions to call. [SPEAKER_10] The next one, if it plays. [SPEAKER_10] Yeah, so in this example, we've got a number of different... [SPEAKER_10] I'll show you that one last, there we go. [SPEAKER_10] A number of different skills that you can try out. [SPEAKER_10] So everything from just loading locations in maps to using APIs and other services. [SPEAKER_10] So you can build these things yourself or you can just use some of the preloaded ones to try and understand what it's capable of.
SPEAKER_00
[SPEAKER_10] And this is our playground for doing so. [SPEAKER_10] And then, lastly, I wanted to show that this is the 2B model on the left-hand side. [SPEAKER_10] So you could also try things like vibe coding on device. [SPEAKER_10] The model's actually quite capable of doing this. [SPEAKER_10] You're not going to build a big architectural system, but it can do small web apps. [SPEAKER_10] It understands lots of different languages. [SPEAKER_10] You can write things in Python, HTML, TypeScript. [SPEAKER_10] It can validate some of the stuff. [SPEAKER_10] And you could build little apps that can run on the phone itself.
SPEAKER_00
[SPEAKER_10] So you got the whole feedback loop on the palm of your hand. [SPEAKER_10] So this one, they just generated a really simple calculator. [SPEAKER_10] And importantly, it does actually support divide by zero and gives an error correctly. [SPEAKER_10] So it's even able to reason about those things. [SPEAKER_10] And to add, I think at the moment, it's got with thinking turned off. [SPEAKER_10] But if you add thinking turned on, we see that you get a bump in terms of the quality of the output in terms of what the model's actually able to reason about. [SPEAKER_10] Because it adds that planning step before it actually does the execution. [SPEAKER_10] Okay.
SPEAKER_00
[SPEAKER_10] So next, I'm going to show you a demo running on the device. [SPEAKER_10] And importantly, it does actually support divide by zero and gives an error correctly. [SPEAKER_10] So it's even able to reason about those kinds of things. [SPEAKER_10] And to add, I think at the moment, it's got thinking turned off. [SPEAKER_10] But if you add thinking turned on, we see that you get a bump in terms of the quality of the output in terms of what the model's actually able to reason about. [SPEAKER_10] Because it adds that planning step before it actually does the execution. [SPEAKER_10] Okay. [SPEAKER_10] So next, I'm going to show you a demo running on the device.
SPEAKER_00
[SPEAKER_10] Who here has used or uses LM Studio or O-Llama or any of those tools? [SPEAKER_10] Okay. [SPEAKER_10] So maybe three or four people. [SPEAKER_10] So actually, this is useful. [SPEAKER_10] So one thing that's really important to us when we build these models is making it compatible with a lot of the tools in the ecosystem. [SPEAKER_10] So we partner with folks like Ollama, LM Studio, VLM, SGLang to make sure our models work really well with them and that they can be deployed efficiently. [SPEAKER_10] LM Studio is a tool for running local model instances. [SPEAKER_10] So in this case, I've just got the 26B model running on my device.
SPEAKER_00
[SPEAKER_10] And you can load it from here. [SPEAKER_10] And at the moment, this one is configured. [SPEAKER_10] This one will take about seven point 18 gigabytes. [SPEAKER_10] But if you add, for instance, enough memory to do the kind of context, you're looking at maybe 22 gigabytes of RAM required to do it. [SPEAKER_10] This is an M4 Mac. [SPEAKER_10] So obviously it has unified memory. [SPEAKER_10] So you can run that, but you would need, if you wanted full performance, you'd need to have a GPU with enough RAM to be able to run the whole model. [SPEAKER_10] And as I mentioned, the 26B model is a mixture of experts.
SPEAKER_00
[SPEAKER_10] So it actually only needs four billion activated parameters. [SPEAKER_10] So it's much quicker than the 31B model, but it is more intelligent than just a four billion parameter model would be. [SPEAKER_10] So one cool thing about this is that with these models, you can serve them on your local machine via a compatible endpoint. [SPEAKER_10] So an OpenAI compatible endpoint or an Anthropic compatible endpoint, and then you can use them with other apps. [SPEAKER_10] So what I'm going to do here is I'm just going to serve this model on port 1234, and then I can use other apps to call the chat completions API directly on this one.
SPEAKER_00
[SPEAKER_10] So the code to do this would be OpenAI chat completions, and then point it at this server. [SPEAKER_10] And it would be the equivalent of just using your local machine rather than using a cloud machine. [SPEAKER_10] So I'm going to run a terminal here. [SPEAKER_10] I'm going to do demo SVG space. [SPEAKER_10] So what I'm going to do is I'm going to create one orchestrator instance, and then I'm going to create 10 sub agents, and then we're going to generate SVGs for me. [SPEAKER_10] So all they're going to be told to do, actually let me make sure I turn thinking off to make sure it's super fast. [SPEAKER_10] Glad I remembered that. [SPEAKER_10] Right.
SPEAKER_00
[SPEAKER_10] So each of these terminals here is a separate agent, and you can see the orchestrator's this one right here. [SPEAKER_10] So it's farming out those decisions to all the different sub agents, and each of those sub agents has been given a different thing to draw. [SPEAKER_10] So you can see this would be a very good way for you to do some quick prototyping on your local machine. [SPEAKER_10] You don't need the internet to do this. [SPEAKER_10] And the throughput in the corner up there is how much combined token generation is being done by each individual sub agent.
SPEAKER_00
[SPEAKER_10] So when they finally put it all together, the orchestrator will just compile this into a page, and you get a lovely array of different SVGs that they've generated. [SPEAKER_10] So you can imagine this being any kind of agentic task that you want to be able to do on your machine.
SPEAKER_00
[SPEAKER_10] You know, sorting out files, implementing bits of code, doing subdividing research, analyzing data, that kind of stuff. [SPEAKER_10] And you can give them all very different jobs to do. [SPEAKER_10] So some of them have finished already. [SPEAKER_10] That's good. [SPEAKER_10] So something hopefully will pop up in a second. [SPEAKER_10] There you go. [SPEAKER_10] So there are my SVGs. [SPEAKER_10] So they're not too bad, are they? [SPEAKER_10] And as I mentioned, this one's animated. [SPEAKER_10] A spinning planet. [SPEAKER_10] So it's somehow figured out how to make that look as though it's animated, which is quite cool. [SPEAKER_10] Yeah.
SPEAKER_00
[SPEAKER_10] And this is thinking turned off. [SPEAKER_10] So if I had thinking turned on, it does a little bit more thought in the planning stage and you generally get better SVGs out. [SPEAKER_10] More time spent thinking about it results in better SVGs. [SPEAKER_10] So the next thing I wanted to show you is I'm going to jump over here to open code. [SPEAKER_10] So as I mentioned, you can use any application [SPEAKER_10] How to make that look as though it's animated, which is quite cool. [SPEAKER_10] Yeah. [SPEAKER_10] And this is the thinking turned off.
SPEAKER_00
[SPEAKER_10] So if I had thinking turned on, it does a little bit more thought in the planning stage and you generally get better SVGs out. [SPEAKER_10] More time spent thinking about it results in better SVGs. [SPEAKER_10] So the next thing I wanted to show you is I'm going to jump over here to open code. [SPEAKER_10] So as I mentioned, you can use any application that can use an open AI compatible interface. [SPEAKER_10] So if you've got a programming environment, you can just point it at that and you can try the model out to just see how well it performs.
SPEAKER_00
[SPEAKER_10] So in this particular example, all you do to configure this in open code is you just, I'm going to kill all my terminals here. [SPEAKER_10] Bye. [SPEAKER_10] Goodbye. [SPEAKER_10] Let's get rid of these. [SPEAKER_10] If you go to, there is a single file. [SPEAKER_10] If I go vim.config.json and I scroll down to where is it? [SPEAKER_10] This section here, this is all you need to specify. [SPEAKER_10] So you give the name of the provider, what the schema is, the models that you want to expose, and any additional parameters that you need for the endpoint. [SPEAKER_10] And then you just literally point it at a URL, which in this case is just my local machine.
SPEAKER_00
[SPEAKER_10] And that's all you need to do. [SPEAKER_10] And I have another example. [SPEAKER_10] I'm not going to run this right now, but we have a guide for how you deploy Gemma 4 on say something like Cloud Run or a cloud provider. [SPEAKER_10] So if you don't have enough RAM to run the quantization that you want, you can just throw it up to a single GPU and you get this, this single command here, this one right here, the beta cloud deploy. [SPEAKER_10] And this will put the model on an RTX Pro 6000. [SPEAKER_10] So you can just, if you just want to test it out, that's one way that you can test whatever model that you fine-tuned or run yourself.
SPEAKER_00
[SPEAKER_10] And we have another way to access it as well. [SPEAKER_10] If you go to ai.dev, that Paige and Guillermo have been showing you, we've also got the demo models are available here. [SPEAKER_10] So if you just want to prompt them, we've got the two big models and you can test them out there. [SPEAKER_10] And they support video and file uploads, grounding with search is all built into that too. [SPEAKER_10] So you can test to see how they perform in isolation as well. [SPEAKER_10] So really there's no reason whether you have the hardware or don't have the hardware for you to be able to give it a shot.
SPEAKER_00
[SPEAKER_10] And yeah, what I wanted to show you was that I had generated a spec for a little game. [SPEAKER_10] Where is it? [SPEAKER_10] Okay. [SPEAKER_10] So I made a game called Nebula Drift. [SPEAKER_10] So what I did is I used the 31 billion parameter model to give me a spec for a game. [SPEAKER_10] And I'm just going to give it to my local model to go and implement. [SPEAKER_10] So if I go implement spec and then I just give it the file, this one, and show it's the right model. [SPEAKER_10] And let's make sure we turn thinking back on so that it's going to actually reason about it too. [SPEAKER_10] There we go. [SPEAKER_10] And we'll run that.
SPEAKER_00
[SPEAKER_10] And so why is this interesting? [SPEAKER_10] Because we haven't changed anything about open code at all. [SPEAKER_10] We've just literally given it the model to try on the local machine. [SPEAKER_10] And in the open code spec, it's going to give it all the different tools that it needs to be able to run open code and to behave in that environment. [SPEAKER_10] So it should think about the task it needs to do and then it needs to look at the file system to either read files from the file system or write files to the file system. [SPEAKER_10] And if it goes wrong, I can, similar to any coding harness, I can just prompt it and say, oh, did you check this?
SPEAKER_00
[SPEAKER_10] Or can I feed the error back into you and have it reason through it? [SPEAKER_10] So this is just it following these instructions. [SPEAKER_10] So yeah, it's decided to make a directory. [SPEAKER_10] It's just going to use a shell directory from that. [SPEAKER_10] So if you've used any coding tools, it shouldn't be too unfamiliar, but maybe you've used it with another model. [SPEAKER_10] Maybe you haven't used it with a local model or an open model running on your hardware. [SPEAKER_10] So I just want to show you that this is possible. [SPEAKER_10] Oh, yeah, there you go. [SPEAKER_10] It's already created an index file. [SPEAKER_10] That was quick.
SPEAKER_00
[SPEAKER_10] It's going to create, if it reads the spec correctly, it should generate two more files. [SPEAKER_10] It should generate a JavaScript one and a CSS one. [SPEAKER_10] And then what we'll do is we'll just try to see what game it managed to make from the spec. [SPEAKER_10] And to give you an idea, it should be an infinite racing game where you've got to avoid some asteroids. [SPEAKER_10] So if it finishes off the code, there we go. [SPEAKER_10] Should we give that a shot? [SPEAKER_10] That looks like it's actually done something correctly. [SPEAKER_10] So if I go run that, where did it write it? [SPEAKER_10] What was it called?
SPEAKER_00
[SPEAKER_10] Nebula drift with an underscore, it decided. [SPEAKER_10] Oh, I think it's even trying to change the file path as I look at it right here. [SPEAKER_10] what game it managed to make from the spec. [SPEAKER_10] And to give you an idea, it should be like an infinite racing game where you've got to avoid some asteroids. [SPEAKER_10] So if it finishes off the code, there we go. [SPEAKER_10] Should we give that a shot? [SPEAKER_10] That looks like it's actually done something correctly. [SPEAKER_10] So if I go run that, where did it write it? [SPEAKER_10] What was it called? [SPEAKER_10] Nebula drift with an underscore, it decided.
SPEAKER_00
[SPEAKER_10] Oh, I think it's even trying to change the file path as I look at it right here. [SPEAKER_10] It's trying to make edits to it. [SPEAKER_10] Stop. [SPEAKER_10] Stop editing. [SPEAKER_10] Let's just try and run that. [SPEAKER_10] Okay. [SPEAKER_10] We have a game. [SPEAKER_10] Oh, it doesn't run. [SPEAKER_10] Okay. [SPEAKER_10] So this is interesting. [SPEAKER_10] So what we can do is if we just go... [SPEAKER_10] Let's try and run that again. [SPEAKER_10] Ah, so there's a syntax problem. [SPEAKER_10] So if I say... [SPEAKER_10] If I just literally copy this.
SPEAKER_00
[SPEAKER_10] So imagine in most harnesses, you'd have a feedback loop where you could feed it back straight into the coding, but I'm just going to do it manually and see if it can spot its own error. [SPEAKER_10] Stop. [SPEAKER_10] Stop. [SPEAKER_10] I found an error. [SPEAKER_10] Let's check that. [SPEAKER_10] Did you actually read that or not? [SPEAKER_10] Let's try again. [SPEAKER_10] Okay. [SPEAKER_10] So now it should... [SPEAKER_10] Now it should do some investigation. [SPEAKER_10] So probably reread the files, see if it can spot what the typo is or what the problem is, and then it should try and edit it or fix it.
SPEAKER_00
[SPEAKER_10] So actually, in OpenCode, they give you two tools, one which can edit individual lines and one that can just rewrite the file. [SPEAKER_10] So it can even make a decision about which one of those it needs to use to do it. [SPEAKER_10] So hopefully it's going to edit it correctly. [SPEAKER_10] Run sub-agents. [SPEAKER_10] The 31B is pretty good at that. [SPEAKER_10] The 26B depends on the task. [SPEAKER_10] Sometimes it understands what sub-agents to call, and sometimes you have to give it a bit more prompting.
SPEAKER_00
[SPEAKER_10] But similarly, with an environment like this, you can expose MCP servers, you can create skills, you can do all sorts of stuff to help guide the model. [SPEAKER_10] Shall we see whether that's fixed anything? [SPEAKER_10] There we go. [SPEAKER_10] I don't even see. [SPEAKER_10] Did that refresh? [SPEAKER_10] Nope. [SPEAKER_10] Not yet. [SPEAKER_10] I was trying to edit the game file. [SPEAKER_10] Must match exactly. [SPEAKER_10] So, okay, so it's trying to edit a part of the file. [SPEAKER_10] I could also just say, no, okay. [SPEAKER_10] I think it probably got unloaded. [SPEAKER_10] Let's put that back in. [SPEAKER_10] Let's load that back up.
SPEAKER_00
[SPEAKER_10] Try again, Gemma. [SPEAKER_10] Right. [SPEAKER_10] So if I tell it, just output the full file again for game.js. [SPEAKER_10] Where did it put it? [SPEAKER_10] It was in that one. [SPEAKER_10] So you can tell it specifically which file it should be reading and writing to. [SPEAKER_10] So hopefully it should listen to what I say and then regenerate that file. [SPEAKER_10] While we're waiting, does anybody have any other ideas for a game to make? [SPEAKER_10] Paige is obsessed with pink squirrels, but... [SPEAKER_10] So you could try something in that vein. [SPEAKER_10] Any ideas for a game? [SPEAKER_10] A game where you can build your own game into the game.
SPEAKER_00
[SPEAKER_10] Whoa! [SPEAKER_10] Let's try it. [SPEAKER_10] So I'm going to use a game. [SPEAKER_10] Yeah. [SPEAKER_10] Right.
SPEAKER_00
[SPEAKER_10] Let's go ask. [SPEAKER_10] We'll ask the big model. [SPEAKER_10] So if I go... [SPEAKER_10] Okay. [SPEAKER_10] Write a spec for a game where you can build your own game inside the game. [SPEAKER_10] And use this as a reference. [SPEAKER_10] So we're going to build a spec first. [SPEAKER_10] Let's just go here and we'll use... [SPEAKER_10] Yeah, we'll use the spec for the other one as a reference. [SPEAKER_10] Let's try and run that. [SPEAKER_10] While we see what's going on with our open code. [SPEAKER_10] This looks like it's editing the file right now. [SPEAKER_10] Yeah, you can see it's processing the tokens right there.
SPEAKER_00
[SPEAKER_10] And if I go to this one... [SPEAKER_10] OmniForge! [SPEAKER_10] So what it's... What I've done here is I've told it to follow a particular spec. [SPEAKER_10] So it keeps the same pattern. [SPEAKER_01] So, for instance, if I had a format that I needed it to do in, it's pretty good at following those instructions and keeping to that record. [SPEAKER_10] This may not be impressive if you've used any model in 2026 other than to know that it's something [SPEAKER_10] I've told it to follow a particular spec. So it keeps the same pattern.
SPEAKER_00
[SPEAKER_01] So for instance, if I had a format that I needed it to do in, it's pretty good at following those instructions and keeping to that record.
SPEAKER_00
[SPEAKER_10] This may not be impressive if you've used any model in 2026 other than to know that it's something that you can actually run at home. That's for me the interesting part of this. All the top models can do this, but yeah, up until recently, not that you could run yourself. So let's leave that doing its thing there. Has it fixed my file yet? It's made more edits. We can go and see what it changed. So it's starting... Oh, it added some more delta time. Oh, it's changing the input? Okay. Well, you can see what it's generating here. Like it's got these little player cases and all sorts of stuff. Should we see if it runs again? Well, we're waiting for that. Right. Are you going to run yet? Ah, what I could do is I could just say to it, let's say pressing the button doesn't start the game. Regenerate the files. And hopefully it's just going to write them back out. Let's do that. Okay, thinking about it. See how our spec is getting on. So let's just get 31b. Can you just implement this spec? This spec, but as a single index.html with the CSS and JavaScript included. See if there can do that. Okay. And there's one other thing I wanted to show you. We'll do that in parallel. AI.dev. One thing that the model is actually quite good at doing being multimodal is to understand the context of the input that it's given. So I could for instance take a screenshot of a website like if I go to Gemma for DeepMind, let's go to this one and if I just go like I just grab this and I could just go to the model implement this web page as a single index.html and if I just upload that file, that one there you go, and try and run this. Yep. How's the open code getting on? So it's written the game file hopefully with no mistakes, written the style file and hopefully written the index file, so let's give that another whirl. Right. There we go. So there's our game. It said I don't see any asteroids. Okay, we've got a star field and we've got a movable ship. So again, this is just what the model has come up with. You could just iterate on this, you could add to it. I could ask it for more asteroids, but I think we're done with this one. Let's see how the other ones are getting on. So this is the game about making a game. I'm really curious to see what it does for this because that's going to be pretty nuts. And then this one. This is generating. This is generating. This is generating the recreated web page. So this is all free to play around with. So just be aware there's usage limits in terms of shared ownership, but you can just use it to test stuff out pretty easy. That's why it's a little bit slower than the Gemini models because there's only
SPEAKER_00
[SPEAKER_10] The recreated web page. So this is all free to play around with, so just be aware there's usage limits in terms of shared ownership, but you can just use it to test stuff out pretty easy. That's why it's a little bit slower than the Gemini models because there's only one poor Gemini model serving everybody. Let's see, is this one done yet? Still going. I don't know, maybe it's made a really epic adventure. I think it might just be, you know, ship it as soon as this is done. What else I want to show you? Oh okay, while we're doing that, I want to show you this cool demo. So one of our team is really into their robotics, so they made a version. If you've ever used any of the robot simulation tools, so this is called Open Duck, which is just a little simulator of duck and you can ask it questions or you can talk to it via the E2B model, which is running in the browser. So it's you actually download the model and it runs in WebGPU straight in the browser page. So when it tells it to do these actions, Gemma's interpreting what it's asked the model to do. It looks at what the robot can do and instructs it to. If you look really carefully, you can see where it says "perform action" yeah, before it actually does the action, it tells you what command it's trying to trigger. So you can imagine this actually running embedded on the device or in this case is in the simulator to just prove the point. So that's quite a cool one. The other one I wanted to show you was the Android Studio team have integrated it into making Android apps. So you've got a little chat window at the side and you can build the apps from there. So again, they have a little agentic builder setup. So if you were just building phone apps, you can use the model. I think they're using the 26B again. Yep, they're using the 26B for that one. And then we also have it working with the ADK, so our agent development kit. Again, it exposes different functionality. It's got a thinking loop. It's got a feedback system. So you can give it more longer running tasks and it can make decisions about pulling information and stuff like that. So all these different environments that you can run it in. Shall we see how our final? Okay, this one is done. So this is going to be our web page. Oh, that's not bad. That looks pretty good. I don't imagine these buttons go anywhere, but for "what's new." Oh, they've, I thought I literally thought I had made a video that would have been a bit intense. But yeah, but you can see even just from the layout it's pretty close. It's kind of matched the font correctly. The understanding of the actual page is pretty good and this is just one-shotted. So point it.
SPEAKER_00
[SPEAKER_10] I thought I made a video that would have been a bit intense, but you can see even from the layout, it's pretty close. It matched the font correctly. The understanding of the actual page is pretty good, and this is just one-shotted, so point it at random web pages and see what it can build. Let's see whether this—I am now really interested in this game. Let's see whether this actually works. Oh my gosh, right, is this—oh, I've got a full editor, so I can build a little—oh, I like that. So a goal, can I run the game? I can put things in. Oh my gosh, please tell me—oh yes, the triangle moves. Oh, I love it. Wow. Okay, there you go. That's a good point to end on, I think. So yeah, GMO models, they run on your phone, they run on your laptops, they run on GPUs. Go try them out. Try the AIG gallery if you want to explore. I'll be around today and tomorrow if you have any questions about how—but yeah, really excited to share that with you and enjoy the rest of the conference.
SPEAKER_00
basically, we are going, to do the same thing, we are going, to ask the model, to generate a video, using the view model, the prompt, I'm going, to use the same prompt, as the one, that we used, to generate the image, and I'm, I'm passing, the last generated image, as, as a starting, frame, and then, I want a portrait, and I want 720p, because it's going, to be smaller, and here we go, so I, actually, haven't checked, what the sounds, looks like, because I, stand back, you shall not pass, leave this place, little ones, you don't belong here, I think it's quite good, but sometimes, it's not, not very good, because the model, doesn't like, the prompt, was just about,
SPEAKER_00
creating the still image, and the model, doesn't know, what exactly, is expected, to be happening, afterwards, so a better way, of doing it, is actually, to reuse the chat, the chat, to ask the model, to come up, with the, with the, with some, explanations, about what's happening, after the image, and I'm also, passing it, the image again, so that it knows,
SPEAKER_09
exactly, which part, of the chapter, we are talking about, so that it can come up,
SPEAKER_00
with the, with the right prompt, and so it, like, it came up, with this, this prompt, moral shivers, and clutter, these scarves, in terror, blah blah blah, what a rat, I bravely draw, this cutlass, and step forward, to product him, so, up, we can see, how it goes, and somehow, like, it's, like, the, like, in this case, the results, are not really as good, and, yeah, and every time, every, every test I've done, somehow, when I use the same prompt, it's, it adds, it adds text, and, and not, when I, when it generates, another, another prompt, we, but it should not have any, impact, because in both cases, I'm just giving a prompt, and giving an image, but I don't know why,
SPEAKER_00
it's happening like that, um, and then, we can use the new, Lyria models, that we generated, like, that we, shipped last week, uh, to, to create songs, for each model, so, still, the same, the same trick, again, we are going to ask the model, to create the, the prompts for that, uh, and then, I'm going, we are going to use, generate content, with the Lyria model, uh, ask it to create, uh, the song, uh, thanks to the, the, the prompt, and that's basically it, so, let's see, the first, uh, the first song should be, orchestral, acoustic, folk music, peaceful, flowing, acoustic guitar, and, and flute duets, let's see.
SPEAKER_00
I think it fits. Then, the second one, the open road, is, uh, jaunty, adventurous, rhythmic, with fiddle, and acoustic bass.
SPEAKER_00
and then, the last one, is suspenseful, creepy melody, with staccato, pizzicato strings.
SPEAKER_00
And I think it fits as well, and I, I honestly, I think the, the model is really good, because I, it's not very often, that you can do those kind of demos, without actually checking, what's, what it's going to be, because I run the, the notebook before, but I didn't check what the, what the music, we're going to be looking like, and that's, uh, so far it worked all the time. Um, one of the things that I saw, I forgot to say about all of that, is that the way we are training, our model internally, is that, a lot of the training data, for the gen media models, is actually made using Gemini. So, that's the reason why, Gemini is quite good, at generating those prompt,
SPEAKER_00
for the generative media models, because, it's already trained, on, uh, on understanding, what Gemini is, is asking for. So, that's, uh, that's what makes, uh, Gemini very good, at generating those, um, those prompts. And then, uh, the last, the last model, that I wanted to show, is, uh, text-to-speech model. So, I'm pretty certain, that you all, have heard of this one, because, uh, everybody loved, the, uh, integration in, notebook LM, where you could create, a, um, uh, uh, uh, uh, podcast based on, on your documents, and that's basically the same model that, uh, you can create, you can ask it to, to talk, uh, to tell, read, and, what's the text you, you give it,
SPEAKER_00
and you can have it, you can have two characters with different voices, and, and so on, that's, that's what makes it, uh, like, nice to listen to. But what I wanted to show you is that, there are actually tricks to, uh, to have more than two voices actually. Um, and, uh, and the reason for that is, uh, I think Paige, uh, showed that a bit, but you can, uh, when you're using text to speech or live, you can actually do a lot in the prompting to make it speak in, in different ways with different accents, with, uh, and so on. And we are going to use that to our advantage to basically have more than, more than two voices in, uh, in the same, um, in the same generation.
SPEAKER_00
So what I did, because I was lazy, I would not do that. It, I was, I wanted to do that at scale, but I basically asked it to extract one of the dialogue from the book and then to rewrite it as, um, as a transcript of, of a play and to replace the name of the character by narrator, where if that's just a narrator of talking and for all of the others, is, they are just going to be named character, but with a specific style for, for each of them. And if the same character comes back, it's, it's going to reuse the same style so that it's going like the same characters are going to have to, to speak the same way of, and have the same voices. Um,
SPEAKER_00
and then I'm defining the voices. So the narrator is going to use a soul of fat voice, whatever it is. And then the, the, all of the others are going to share the failure voice. Um, and then I'm just asking it to, uh, to, uh, to generate the thing. And just one trick here with the TTS model, you always have to start with read this or, uh, tell me this or whatever. If you just send the text, it's, it's going to ignore it for some reason. So, uh, you, you need to prompt it to read the text all the time. Um, but that's also where you can add some, uh, some context and like equivalent of what would be a system instructions, like read this in, uh, scary way, uh,
SPEAKER_00
or the characters are very excited. So that's where you can, you can add those details as well. Um, and so you can see the, the, the text that we got is narrator is saying that then the first characters is going to talk very fast paced and with a British accent and a post-British accent. Um, and then, and so on. And we can see that the same, uh, the same way of speaking is, is coming often. So it's when, uh, the two characters, uh, going back and forth. So let's see how it goes. Small, neat ears and thick, silky hair. The two animals stood and regarded each other cautiously. Hello. Hello. Would you like to come over? Oh, oh, it's all very well to talk.
SPEAKER_00
He spoke rather pettishly, being new to a river and riverside life and its ways. The second animal said nothing, but stooped and unfastened a rope and hold on it, then lightly stepped into a little boat, which had not been observed. It was painted blue outside. And white within, and was just the size for two animals. And the first animal's whole heart went out to it at once, even though he did not yet fully understand its uses. The rower sculled smartly across and made fast. Then he held up his forepaw, as his guest stepped gingerly down. Lean on that. Now then, step lively. To his surprise and rapture, he found himself actually seated in the stern of a real boat.
SPEAKER_00
This has been a wonderful day. Do you know, I've never been in a boat before in all my life. What? Never been in a... You never... Well, I... What have you been doing then?
SPEAKER_07
He was quite prepared to believe it, as he leant back in his seat and surveyed the cushions, the oars, the row locks, and all the fascinating fittings, and felt the boat sway lightly under him.
SPEAKER_08
Is it so nice as all that?
SPEAKER_07
Nice? It's the only thing. Believe me, my young friend, there is nothing, absolute nothing, half so much worth doing as simply messing about in boats. Simply messing, messing, about, in, boats. Messing... Look ahead! It was too late. The boat struck... Basically, you couldn't guess that it's actually using the same voice for the two. Two characters are so much different. And... So that's... That's really, actually, a really cool way of, like, creating those discussions with multiple characters. So I can use it. And the last thing is that, as we said,
SPEAKER_08
like, the Gemini models are multimodal by default. So I was... I was using a book, but you can also send an audio book or a movie or something, like,
SPEAKER_07
not too long movie, but you could do the same things with using all of the multimodal input possibilities to, to have Gemini create things for, to illustrate other, other kind of modalities as well. How much time do we have?
SPEAKER_08
I... I said I was going to show you the, yeah, real-time model. So this is a very good, a very cool example that... Oh, no. Stop. Stop. I'm going to show you something before. Uh... Page made some, uh, example of cool things that you can build in the ICO. So one of the things I wanted to try
SPEAKER_00
was basically to do the same thing as what I do in the, in the notebook, which is kind of lengthy and quite complicated stuff. So what I did is I basically copy-paste it. Can you build an app that, uh, that illustrates a book using Gen Media models as described in this Python notebook? And I just paste all of the content of the notebook and, and build it. And we'll see how it goes. Um, it can take quite a long time. Like earlier, it took 15 minutes. So we, you might not want to wait for that. But, uh, we can, I can show you Space DJ in the meantime. So, uh, as I said, the model generates, uh, music in real-time. And this is, uh, this is a, uh, a demo
SPEAKER_00
where they made a star, like, uh, universe of stars. And the, the planets are prompts. So if we go closer to some planets, it should start playing music.
SPEAKER_00
And then if you move somewhere else.
SPEAKER_00
Yeah, we're, uh, yeah, I guess in the metal, uh, area.
SPEAKER_00
See, you, uh, I'm really surprised, like, people are not doing more things with this model because it's, uh, it's so funny to me.
SPEAKER_00
And there's an autopilot. So you can just let it move around for, uh, for 10 minutes and then, uh, listen to the music changing in real-time. Um, so that's kind of cool. Up. Let's see how it goes here. Like, yeah, it's thinking. But let me show you what it, what it did before when I ran it. Oh. And I, I keep, I kept my laptop open for half an hour just so that it would not do that and refresh the page. So I guess I lost. Um, but that's basically what it did in, uh, in, uh, like earlier, just, uh, as you can see, just the prompt and nothing more the very long prompt.
SPEAKER_00
Up. Uh, and, uh, it took something like, like a thousand seconds. So, like, 15 minutes if I'm not wrong, um, to create the app. But that's basically doing exactly the same thing. So I can choose the file and upload it. And it's going to take some time, but basically, like, trust me, it's doing exactly the same thing, uh, as, uh, as what the notebook was doing. And that's really, like, quite impressive that it's doing that, uh, on the first try right away. Um, and just to finish, I, one of the thing I'm doing when I'm vibe cutting, uh, while it's working, is that I always have those, uh, instructions on how to get, uh, to get it to generate, um, applets or apps.
SPEAKER_00
Uh, and one of the reasons for that is that I want the apps to be as easy for me as to review as possible, because I want to avoid exactly what I'm doing. Exactly what happened to Paige earlier is that something was not working, and then you need to figure out what it is and ask the model to fix it. Um, so one of my key tricks is that I'm asking the model to create different files
SPEAKER_01
for each, uh, each feature. So that, uh, whenever I need to review something, I can quickly check, like,
SPEAKER_00
I ask it to modify this feature. It's updating something that has nothing to do with it. So I know from the start that there's something wrong going there, and I will be more forward in the, in my, in my reviews and all. And also some instructions that I've, I don't understand why it's not by default in any vibe cutting, uh, tool, which is add logs. Because when you need to, uh, when you need to debug, you, you just not, the error message is not enough. You need to know what's happening before and what's happening, like, because, because you know, you need to know where, where it's happening. So that's, uh, I really recommend you to, uh, to, to use those kinds of, uh,
SPEAKER_00
of guidelines whenever you, you vibe cutting things. And we can see how it goes now. Oh, and something there, I don't think page showed it, but if you are, uh, if you're willing to pay, you can add an API key here, and then you are going to have more, uh, more quota to use Gemini when, when you are vibe cutting on, on AI studio. Uh, I think I will just let you show your demos and maybe I can show you afterwards. how it looks, uh, after. Yeah, sure. Um, for the video model, when you generated the video, it came with, what, the music background, is it using your other model to generate that, or is it just video? Yeah, it's just video at the moment.
SPEAKER_00
So is there a way to, kind of, orchestrate all the models and put them together, or do you have to just use, uh, if you, if you want to use the music generated by Lyria in the video, in the video video, for example? Uh, no, you don't have a way to do that. Uh, it's, it's more because, it's more because of the model limitations that the VO of three generation was not meant to be able to, uh, to ingest audio files. So, uh, so that's why it's, it's limiting. Uh, but I guess, like, the future is that we want every model to be able to ingest all the modalities, uh, so that we, you can, you can do that. Uh, uh, same thing.
SPEAKER_00
I think, uh, it, it would make sense even for Lyria model to be able to, uh, ingest audio models that it can, so that it can be used as references as well. Or, uh, even if you want to do multi-turn and say, okay, I love, I love your song, but the, the ending was a bit, uh, uh, not, not epic enough. So make it, make it more epic so that it would maybe just, just update the, the end of it. So, yeah, it's not possible at the moment, but that's, that's the direction it's going, uh, to. How was the performance in Ponte, maybe specific for the music in VO or your general in video? No, I, I wouldn't try VO for, for any kind of music.
SPEAKER_00
Like I, the, the, the training data for, for the background music and like is very, is likely very light because it's always doing the same kind of non-music things. Um, so, uh, no, but once again, all of those models are trained more or less together, uh, and share some training data. So I'm sure the next generation of VO is going to be better at that. It's just that the current one, it, it shipped a year ago and a year ago, there was plenty, like we didn't have music generation models, uh, nano, like, uh, we talked earlier about texts in videos and, uh, and, uh, like a year ago, nano, nano, nano was not even called nano banana.
SPEAKER_00
It was Gemini 1.5 image, no 2.0 image, uh, generation something. Uh, and it was, uh, it was not as good at, uh, as, as text to that. So that's, uh, yeah, it's just, uh, just a question of, of generation of models, uh, there. Oh, and I forgot something to, uh, while it, while it work, um, the, if you want to run the,
SPEAKER_01
to run the notebook, uh, because you will have to remove it as well.
SPEAKER_03
Um, e, e, e, e, here, here, when I started, I created the chat. I, I added this line that says service tier priority. It's actually something that we shipped last week.
SPEAKER_00
Um, and that's a way for, uh, to indicate when you are prompting the models to, if you want, uh, if it's not that important, you don't, you don't care about the latency, but, uh, but you, but you care about price. So there's, uh, a service tier that is called flex. And basically that says, uh, the, the request can take a few minutes to go through, but it's going to be half price. So it's kind of, uh, the same thing as, uh, as using the batch API. And on the other end, because I wanted to be so that certain that it would go through today, you can, you can say that it's a priority request. So it's going to have slightly higher priority and that's, uh, it should not, uh,
SPEAKER_00
and it's, it's going to be more reliable, but you're going to pay twice the price. So, uh, if you are running it by yourself, you might want to remove that line to save a bit. Um, and we can see how it goes here. Oh, and see, as I said, like this model is in, uh, higher demand. So, yeah, that's, um, I'm, I'm sorry about that. It's, uh, it's, uh, the model is too, uh, too, uh, too good people. Everybody wants to use it and we don't have enough capacity. So if you have, if you have any spare TPU to, uh, to share, um, we can make a deal. So, you have a question?
SPEAKER_03
Do you want to, uh, put, like, variations of this model open, like, GenMedia, but for, uh, for art?
SPEAKER_00
Um, so the question, like, for the video, for, uh, the question is, uh, will, will there be any open weight GenMedia model, basically? Um, so I don't think so for image and video generation, uh, to be honest. Um, one, one of the reasons I see behind that is all of the, um, not like security, but like the, um, when, when you, when you generate images or videos, um, we do a lot of checks about what you are asking for and what's, uh, what is actually being generated afterwards. And we are blocking a couple of things that are not aligned with our visions.
SPEAKER_00
And I, I feel like whenever you have an open weight model for that, it's, it's more like open bar and you can, you can have it generate whatever you want. So that's, I think that's, uh, that's going to be where our, uh, our company values are going to be, uh, to be limiting us. Um, that said, for example, for music, it would make sense to, uh, like, especially if you want real time, it would make sense to be, uh, to be, uh, to be on the internet. And the device that it's, it's, it's doing it faster. So that's, that's something that might come at some points. Um, yeah, I think we'll see after the Gemma demos if, uh, if it was better. Um, yeah, hope it was interesting.
SPEAKER_00
And so, uh, Ian is, uh, basically why my counterpart, I work only on Genmedia Modules. Ian is working on Gemma models and he's going to, like, the biggest release of last week was Gemma 4 and he's going to show you how cool the model is. Thank you. Yes. So as Gia mentioned, my name's Ian Ballantyne. I'm a developer relations engineer working on the Gemma models and this is an impromptu talk. We, obviously because everything has happened today, we've done, Paige mentioned that we're going to do two sessions, just kind of go through a whole load of stuff. We have a number of Gemma talks for the rest of the week as well. Omar's going to do a keynote on Friday, I believe,
SPEAKER_00
which is going to cover a lot of, like, the what Gemma is and why you should be interested, so I'm going to, like, speed run that section. And we have another talk tomorrow on what we've called sovereign escape velocity, so, like, how you take ownership of AI and
SPEAKER_01
run things on your device or on your
SPEAKER_00
own cloud or on phones and hopefully some of what I talk about here is going to be, like, a realization of some of that. So that's a little bit more about, like, the why. I'm just going to show you the what it is, if that makes sense. So let's get plugged in. When we do that, which one do I need? Sorry. This magic one. Yep. As soon as I get a video in. Should I have any audio? Uh, no. Let's go with no. We'll try and show the ones without audio first of all. Uh, okay. So as I mentioned, this is super impromptu, so
SPEAKER_10
we're going to go with whatever we get. Uh, but I will go through, I'll just show you a little bit when we get the screen up. Oh, do you only need me to move it to one side? What can I do? Mirror. Let's do mirror. Uh, where? Yeah, you go for it. Perfect. Yeah, so what I'm having to show you a little bit is, uh, you know, some of the things that, uh, some of the devices that Gemma models can run on. Like, I'll talk very briefly about the different sizes of Gemma model. Uh, what we released in Gemma 4 last week, because it's brand spanking new. Um, and hopefully show you, give you kind of a feel of, like, the kind of capabilities that you can now do either
SPEAKER_10
locally or that you can run on, like, a single GPU that maybe you couldn't, like, six months ago. There you go. Here you go. Perfect. Ta-da. So Gemma 4. Hooray. Uh, so this was released last, uh, Thursday. Um, it's a family of four models. So we have what we call the effective models, the E2B and the E4B models.
SPEAKER_01
And these are models designed to run, like,
SPEAKER_10
on mobile phones, Raspberry Pis, Jets and Nanos. Kind of like very small, low-end hardware. Uh, the E part of it is the question we always get asked. The effective is, uh, because the model architecture has a per-layer embeddings structure,
SPEAKER_01
which means that the embeddings actually don't need
SPEAKER_10
to be loaded as part of the model. So you can have them running on Flash, um, and then you can page in, uh, the embeddings as they're needed for the model. So the actual, what we kind of describe as, like, the brain of the model is about a 2 billion parameter, about 4 billion parameter. But if you put all of it in RAM, it's a bit bigger than that. It's kind of more like a 5 and an 8 billion parameter model. Um, so that's why it's called effective 2B. Uh, we also have a 26B, which is a mixture of experts model with, uh, 4 billion activated. And we have a 31 billion parameter dense model, which is our, like, our flagship kind of big model.
SPEAKER_10
Um, and both of the two bigger models, the 26 and the 31, uh, designed to be able to run on, like, laptops and desktops or, like, single instance GPU clouds, depending on, like, the precision and quantization you need.
SPEAKER_09
So I'm not going to get into that,
SPEAKER_10
but just imagine that a lot of the capabilities you can do right now, you could do on, like, a MacBook with, you know, enough RAM or a, you know, a 50-90 or something like that. That's kind of where we're sitting right now. Um, why are these kind of exciting and interesting to us? Because with these models, we're focusing kind of on the authentic side. So they now have thinking built in. They are multimodal, so they can understand image, audio, video. Uh, audio, by the way, is just for the two smaller models. The effective models that run on the phone can understand audio, but the rest can understand image and video. And we're seeing, uh, performances offer models
SPEAKER_10
that are in the range of, like, 10x bigger than it in terms of parameter size.
SPEAKER_09
So what might have required a cluster, now you could do the same kind of capabilities for, uh, for a single GPU. So that's kind of really where we're seeing, um, some innovation here.
SPEAKER_10
Um, and what I'm going to jump straight into is I want to show you the demo section. Oh, I'll show you this one slide. I like this one. And this is quite, kind of explains a little bit where, where we've come. So the Gemma 2 models, you can see, like, right in the middle. We're kind of, you know, particularly good at creative writing, but, uh, uh, not so kind of evenly spread amongst, uh, amongst other different capabilities. And as we've kind of gone through Gemma 3 and then now through Gemma 4, we've kind of evened out the overall capabilities of the model. We've got things that are much better at coding, much better at function calling, action taking.
SPEAKER_10
It's all kind of built into the design of the architecture of the model rather than relying on, uh, like, strong, um, instruction following capabilities, which we find with bigger models is, is kind of more important. So the models are designed with that from the ground up. So, uh, let's go over to, uh, this is all spoilers from Omar's talk. So I'm going to go straight to the demo section. Here, I'll show you this one. So this is the, uh, Google AI Edge gallery app. Uh, has anybody tried this app yet? Uh, have, uh, one, two, three. Okay. Like a handful of people. So you can download this on Play Store or App Store. Um, it works on Android and iOS.
SPEAKER_10
And, uh, what it is, is it's, uh, it's a way that you can test and try the models. So we released a new feature called agent skills. And what agent skills allows you to do is it allows you to set up skills for the models that run physically on the phone. So this model, this is the E2B model is running literally on a pixel, I believe in this case. And you can prompt it and you can test out the different capabilities. So it's effectively given a set of, um, uh, skills that define things like Android intents, where you can actually trigger and call other apps, or you can write your own JavaScript skills or, like, run, like, a web view.
SPEAKER_10
And, uh, you can instruct the model and it will make a decision about how to actually trigger these. So you can get it to, for instance, add things to this, uh, research tracker. Or you can ask it questions about the research tracker and it will call the correct function to pull that data back. So we've gone from, like, a world of just being able to chat with it on your phone to now that you can give it some more ambiguous input and it's able to make a decision about, uh, what functions to call. Uh, the next one, if it plays. Yeah, so in this, uh, in this example, we've got a number of different... I'll show you that one last, there we go.
SPEAKER_10
Uh, a number of different skills that you can try out. So everything from just, like, you know, loading locations in maps to, uh, like, uh, uh, using APIs and other services. Um, so you can build these things yourself or you can just use some of the preloaded ones to try and understand what it's kind of capable of. And this is, like, our playground for kind of doing so. And then, lastly, I wanted to show that, uh, this is the 2B model on the left-hand side. So you can... you could also try things like vibe coding on device. The model's actually quite capable of doing this. You're not going to build, like, you know, a big architectural system,
SPEAKER_10
but it can do, you know, small, like, web apps. It understands lots of different languages. Uh, you can write things in Python, HTML, TypeScript. Uh, it can validate some of the stuff. And, uh, you know, you could build, like, little apps that can run on the... on the phone itself. So you kind of got the whole feedback loop just on... in the palm of your hand, basically. So this one, they just generated, like, a really simple calculator. And importantly, it does actually support divide by zero and gives an error correctly. So it's even able to reason about those kind of things. And to, uh, to kind of add, I think this, at the moment, it's got with thinking turned off.
SPEAKER_10
But if you add thinking turned on, we see that you get a bump in terms of the quality of the output in terms of what the model's actually able to reason about. Because it adds that planning step before it actually does the execution. Okay. So next, I'm going to show you a demo running on the device. Um, who here has used or uses LM Studio or, uh, O-Lama or any of those kind of tools? Okay. So, like, maybe, like, three or four people. Um, so actually, this is useful. So one thing that's really important to us when we build these models is making it compatible with a lot of the tools in the ecosystem. So we partner with, uh, folks like, Alama, LM Studio, VLM, SGLang
SPEAKER_10
to make sure our models kind of, uh, like, work really well with them and that they can be deployed efficiently. Um, LM Studio is kind of like, uh, a tool for running local model instance. So in this case, I've just got the 26B model running on my device. And you can, uh, you can load it from here. And at the moment, this one is configured. This one will take about, uh, at the moment it says seven, 18 gigabytes. But if you add, for instance, uh, enough memory to do the kind of context, you're looking at, maybe, like, 22 gigabytes, uh, of, uh, RAM required to do it. This is an M4 Mac. So obviously it has unified memory. So you can, you can run that, but you would need,
SPEAKER_10
if you wanted, like, full performance, you'd need to have a GPU with enough RAM to be able to run the whole model. And as I mentioned, the 26B model is a mixture of experts. So it, uh, it actually only needs, uh, four billion activated parameters. So it's much quicker than the 31B model, but it is more intelligent than just a four billion parameter model would be. So, uh, one cool thing about this is that with these models, you can serve them, uh, on your local machine, uh, via a compatible endpoint. So, uh, like an open AI compatible endpoint or an anthropic compatible endpoint, and then you can use them with other apps. So what I'm going to do here
SPEAKER_10
is I'm just going to serve this model, uh, on port one, two, three, four, and then I can use other apps to call, uh, the chat completions API directly on this one. So the code to do this would be like open AI, uh, open chat completions, and then point it at this server. And it would be the equivalent of just using your local machine rather than using a cloud machine. So I'm going to run a, uh, terminal here. I'm going to do demo SVG space. So what I'm going to do is I'm going to create, uh, one orchestrator instance, and then I'm going to create 10 sub agents, and then we're going to generate SVGs for me. So all they're going to be told to do,
SPEAKER_10
uh, actually let me make sure I turn thinking off to make sure it's super fast. Uh, glad I remembered that. Right. So each of these terminals here is a separate agent, and you can see the one, the orchestrator's this one right here. Uh, so it's farming out those decisions to all the different sub agents, and each of those sub agents has been given a different thing to draw. So you can see this would be like a very good way for you to do some kind of quick prototyping on your local machine. Uh, you don't need the internet to do this. And the throughput in the corner up there is how much combined token generation is being done by each individual, um, sub agent.
SPEAKER_10
So when they finally put it all together, the orchestrator will just compile this into a page, and you get like a lovely array of different SVGs that they've generated. So you can imagine this being any kind of agentic task that you want to be able to do on your machine. You know, sorting out files, implementing bits of code, doing like subdividing research, analyzing data, that kind of stuff like that. Um, and you can give them all very different jobs to do this. So some of them have finished already. That's good. So something hopefully will pop up in a second. There you go. So there are my SVGs. So, oh, they're not too bad, are they? And as I mentioned,
SPEAKER_10
oh, this one's animated. Ooh, a spinning planet. So it's somehow figured out how to make that look as though it's, uh, animated, which is quite cool. Yeah. And this is the thinking turned off. So if I had thinking turned on, it does a little bit more thought in the planning stage and you generally get better SVGs out. Like more time spent thinking about it results in better SVGs. Um, so the next thing I wanted to show you is, uh, I'm going to jump over here to open code. So as I mentioned, uh, you can use any application that can use an open AI compatible interface. So if you've got like a programming environment, you can just basically point it at that
SPEAKER_10
and you can try the model out to just see how well it performs. So in this particular example, the, all you do to configure this in open code is you just, I'm going to kill all my terminals here. Bye. Goodbye. Let's get rid of these. If you go to, there is a single file. If I go vim.config.json and I scroll down to, where is it? This section here, this is all you need to specify. So you give the name of the provider, what the schema is, the models that you want to expose, and any additional parameters that you need for the endpoint. And then you just literally point it at a URL, which in this case is just my local machine. And that's all you need to do.
SPEAKER_10
And I have another example. I'm not going to run this right now, but we have a guide for how you deploy Gemma 4 on say something like Cloud Run or like a cloud provider. So if you don't have enough RAM to run the quantization that you want, you can just throw it up to a single GPU and you get this, this single command here, this one right here, the beta cloud deploy. And this will put the model on an RTX Pro 6000. So you can just, if you just want to test it out, that's one way that you can test whatever model that you fine-tuned or run yourself. And we have another way to access it as well. If you go to ai.dev, that Paige and Guillermo have been showing you,
SPEAKER_10
we've also got the demo models are available here. So if you just want to prompt them, we've got the two big models and you can test them out there. And they support like, you know, video and file uploads, grounding with search is all kind of built into that too. So you can test to see how they perform in isolation as well. So really there's no reason whether you have the hardware or don't have the hardware for you to be able to give it a shot. And yeah, what I wanted to show you was that I had generated a spec for a little game. Where is it? Okay. So I made a game called Nebula Drift. So what I did is I used the 31 billion parameter model
SPEAKER_10
to give me just like a spec for a game. And I'm just going to give it to my local model to go and implement. So if I go implement spec and then I just give it the file, this one, and show it's the right model. And let's make sure we turn thinking back on so that it's going to actually reason about it too. There we go. And we'll run that. And so why is this interesting? Because we haven't changed anything about open code at all. We've just literally given it the model to try on the local machine. And in the open code spec, it's going to give it all the different tools that it needs to be able to run open code and to behave in that environment.
SPEAKER_10
So it should think about the task it needs to do and then it needs to look at the file system to either read files from the file system or write files to the file system. And if it goes wrong, I can, similar to any coding harness, I can just prompt it and say, oh, did you check this? Or can I feed the error back into you and have it kind of reason through it? So this is just it following these instructions. So yeah, it's decided to make a directory. It's just going to use a shell directory from that. So if you've used any coding tools, it shouldn't be too unfamiliar, but maybe you've used it with another model. Maybe you haven't used it with a local model
SPEAKER_10
or an open model like running on your hardware. So I just want to show you that this is kind of possible. Oh, yeah, there you go. It's already created an index file. That was quick. It's going to create, if it reads the spec correctly, it should generate two more files. It should generate a JavaScript one and a CSS one. And then what we'll do is we'll just try to see what game it managed to make from the spec. And to kind of give you an idea, it should be like an infinite racing game where you've got to avoid some asteroids. So if it finishes off the code, there we go. Should we give that a shot? That looks like it's actually done something correctly. So if I go run that,
SPEAKER_10
where did it write it? What was it called? Nebula drift with an underscore, it decided. Oh, I think it's even trying to change the file path as I look at it right here. It's trying to make edits to it. Stop. Stop editing. Let's just try and run that. Okay. We have a game.
SPEAKER_10
Oh, it doesn't run. Okay. So this is interesting. So what we can do is if we just go...
SPEAKER_10
Let's try and run that again. Ah, so there's a syntax problem. So if I say... If I just literally copy this. So imagine in most harnesses, you'd have a feedback loop where you could feed it back straight into the coding, but I'm just going to do it manually and see if it can spot its own error. I... Uh... Stop. Stop. I found... an error.
SPEAKER_10
Let's check that.
SPEAKER_10
Did you actually read that or not? Let's try again.
SPEAKER_10
Okay. So now it should... Now it should do some investigation. So probably reread the files, see if it can spot what the typo is or what the problem is, and then it should try and edit it or fix it. So actually, in OpenCode, they give you two tools, one which can edit like individual lines and one that can just like rewrite the file. So it can even make a decision about which one of those it needs to use to do it. So hopefully it's going to edit it correctly.
SPEAKER_10
Um... Uh... Run sub-agents. Uh... The 31B is pretty good at that. The 26B depends, like, on the task. Sometimes it understands, like, what sub-agents to call, and sometimes you have to give it a bit more prompting. But similarly, with an environment like this, you can expose MCP servers, you can create skills, you can do all sorts of stuff to kind of, like, help guide the model. Um... Shall we see whether that's fixed anything?
SPEAKER_10
There we go. I don't even see. Did that refresh? Nope. Not yet. I was trying to edit the game file.
SPEAKER_10
Must match exactly. So, okay, so it's trying to edit a part of the file. I could also just say, like, uh... No, okay. I think it probably got unloaded. Let's put that back in. Let's load that back up. Try again, Gemma. Right. So if I tell it, just output the full file again for game.js. Uh... Where did it put it? It was in that one. So you can tell it, like, specifically which file it should be reading and writing to. So hopefully it should listen to what I say and then regenerate that file.
SPEAKER_10
While we're waiting, does anybody have any other ideas for a game to make? Paige is obsessed with, like, pink squirrels, but, like... So you could try something in that vein. Any ideas for a game? A game where you can build your own game into the game. Whoa! Let's try it. So I'm going to use a game. Yeah. Right. Let's go ask. We'll ask the big model. So if I go... Uh... Okay. Write a spec spec for a game where you can build your own game inside the game. And... use this as a reference. So we're going to build a spec first. Uh... Let's just go here and we'll use... Yeah, we'll use... We'll use the spec for the other one as a reference. Let's try and run that.
SPEAKER_10
While we see what's going on with our open code. This looks like it's editing the file right now. Yeah, you can see it's processing the tokens right there. And if I go to this one... OmniForge! So... What it's... What I've done here is I've told it to follow a particular spec. So it keeps
SPEAKER_01
the same pattern. So like, for instance,
SPEAKER_10
if I had like a format that I needed it to do in, it's pretty good at following those instructions and kind of keeping to that record. This may not be impressive if you've used any model in 2026 other than to know that it's something that you can actually run at home. That's, for me, is the interesting part of this. Like the, you know, all the top models can do this. But yeah, up until recently, not that you could run yourself. So let's leave that doing its thing there. Has it fixed my file yet? It's made more edits. We can go and see what it changed. So it's starting... Oh, it added some more delta time. Oh, it's changing the input? Okay. Well, you can see
SPEAKER_10
what it's generating here. Like it's got like these little player cases and all sorts of stuff. Should we see if it runs again? Well, we're waiting for that. Right. Are you going to run yet? Ah, do you know what I could do is I could just say to it, let's say pressing the button doesn't start the game. Regenerate the files.
SPEAKER_10
And hopefully it's just going to write them back out. Let's do that. Okay, thinking about it. See how our spec is getting on. So let's just get 31b. Can you just implement this spec? This spec. But as a single index.html with the CSS and JavaScript included. See if there can do that.
SPEAKER_10
Okay. And there's one other thing I wanted to show you. We'll do that in parallel. AI.dev. One thing that the model is actually quite good at doing being multimodal is to understand the context of the input that it's given. So I could for instance take a screenshot of a website like if I go to Gemma for DeepMind let's go to this one and if I just go like I just grab this and I could just go to the model implement this web page as a single index.html and if I just upload that file that one there you go and try and run this yep how's the open code getting on? So it's written the
SPEAKER_10
game file hopefully with no mistakes written the style file and hopefully written the index file so let's give that another whirl right there we go so there's our game it said I don't see any asteroids okay we've got a star field and we've got a movable ship so again this is just what the model has come up with you could just iterate on this you could add to it I could ask it for more asteroids but I think we're done with this one let's see how the other ones are getting on so this is the game about making a game I'm really curious to see what it does for this because that's going to be pretty nuts and then this one this is generating this is generating this is generating
SPEAKER_10
the recreated web page so this is all free to play around with so just be aware there's like usage limits in terms of like shared ownership ship but you can just use it like to test stuff out pretty easy that's why it's a little bit slower than the Gemini models because it's there's only one poor Gemini model serving everybody let's see is this one done yet still going I don't know maybe it's made like a really epic adventure I think it might just be this might be you know ship it as soon as this is done what else I want to show you oh okay while we're doing that I want to show you this cool demo so one of our team is really into their robotics so they they made a version
SPEAKER_10
if you've ever used any of the robot simulation tools so this is called open duck which is just like a little simulator of duck and you can ask it questions or you can talk to it via a the E2B model which is running in the browser so it's you don't you actually download the model and it runs in web GPU straight in the browser page so when it tells it to do these actions Gemma's interpreting like what it's asked the model to do it looks at what the robot can do and instructs it to if you look really carefully you can see where it says perform action yeah before it actually does the action it tells you what command it's trying to trigger so you can imagine this
SPEAKER_10
actually running embedded on the device or in this case is in the simulator to just like kind of like prove the point so that's quite a cool one the other one I wanted to show you was the Android Studio team have integrated it into making Android apps so you've got like a little chat window at the side and you can build the apps from there so again they have like a little agentic builder setup so if you were just building phone apps you can use the model I think they're using the 26B again yep they're using the 26B for that one and then we also have it working with the ADK so our agent development kit again it exposes like different functionality it's got a thinking loop
SPEAKER_10
it's got like a feedback system so you can give it more longer running tasks and it can make decisions about like pulling information and stuff like that so all these different environments that you can kind of run it in shall we see how our final okay this one is done so this is going to be this will be our web page oh that's not bad that looks pretty good I don't imagine these buttons go anywhere but for what's new oh they've I thought I literally thought I had made a video that would have been a bit intense but yeah but you can see even just from like the layout it's pretty close it's kind of matched the font correctly the understanding of the actual page is pretty
SPEAKER_10
good and this is just kind of one-shotted so point it at random web pages and see what it can build and let's see whether this I am now really interested in this game let's see whether this actually works oh my gosh right is this oh I've got like a full editor so I can build like a little oh I like that so a goal can I run the game I can put things in oh my gosh please tell me oh yes the triangle moves oh I love it wow okay there you go that's a good point to end on I think so yeah GMO models you know they run on your phone they run on your laptops they run on GPUs go try them out try the AIG gallery if you want to explore I'll be around today and tomorrow if you have any
SPEAKER_10
questions about how but yeah really excited to share that with
SPEAKER_01
you and enjoy
SPEAKER_10
the rest of the conference yeah good good good good good good good good good good good good good good good