SPEAKER_01
Okay, awesome. Yeah, thank you everyone for joining the session. If you joined the previous session, we're switching slightly the topic and talking more about Gemini now. But I have two slides about Gemma as well, so to make Gus happy. I'm Patrick. I'm a member of the technical staff at Google DeepMind. I work on the Gemini API and AI Studio. And today I want to talk about any-to-any building native multimodal agents. So I want to talk about multimodal understanding, multimodal generation, real-time interactions, and then also build an example app together. So at the end of this session, you should be able to build this for yourself, a little Notebook LM clone.
SPEAKER_01
So what does any-to-any mean? These are all the capabilities you can do with the Gemini API. So there's a lot of use cases this enables. Because Gemini does not only understand text, right? It's natively multimodal, so you can also feed in code, image, audio, video, and then some more like URLs and also Google search. And then it can not only generate text, but now we're also able to generate images, speech generations, video generations, function calls, and of course code generation. So yes, this enables a lot of really cool stuff. But this slide is slightly giving the wrong impression because actually there are still different models. It's not one multimodal model yet.
SPEAKER_01
This is a bigger vision that we have at Gemini, at Google DeepMind, to bring more of the generation capabilities also into Gemini. But currently, it looks like this. This is an ugly slide, I know. But yes, we have the main Gemini model series right now, Gemini 3. And it's able to understand multiple modalities, but it only outputs text. And then we have different specialized native generation models, for example, Nano Banana for native image generation, and speech generation based on the main Gemini models.
SPEAKER_01
And I want to talk about this in a moment. And also, I mentioned Gemini here, correct me if I'm wrong, but it allows text, image, and video input, and the smaller models also audio inputs. So you can also build multimodal agents locally. So I want to focus on four things, four models, the multimodal understanding with Gemini, native image generation, native speech generation, and then also, if we still have time, a little bit about the live API. And then build something together, or at least give you the building blocks, how you can build this, a little Notebook LM clone. Who has used Notebook LM before?
SPEAKER_01
Almost everyone. Okay, so I don't think I have to explain it. But yes, you can feed in multiple different sources. And then the audio overview is pretty popular, where you can generate a podcast to explain topics for you. And then also, infographics are pretty cool. So we want to build the same thing. And we want to build this as an agent rather than a workflow. So this means that the agent should be able to decide what to create, rather than where we hard code the pipeline. Here, we are having a reasoning model that can decide what to create. And then it's hooked up via tool calls or function calls, and then calls the other specialized models.
SPEAKER_01
So this is roughly how the app or the agentic architecture looks like. We have the phase one for multimodal understanding. And then we have the phase two, this is where we have the agentic loop, where we use Gemini as the reasoning model.
SPEAKER_01
And it can then call different tools. And these will then generate different modalities for us. And then it acts in a loop and reasons if we need more assets, or if it's good enough. And then in the end, we get text, speech, and infographics as output. And yes, I want to do this as an example with some learning about attention is all you need paper. So we want to be able to feed in PDFs, images, videos, this can be a lecture, for example, or tutorial, and then voice memos. And ideally, we also want this cross model understanding, right, that we can draw information from all the different sources together and let the model make connections.
SPEAKER_01
And it's actually extremely easy with Gemini to achieve this. This is basically the code you need. Who has built with the Google Gen AI SDK before? Almost half of the room? Yes, basically, this is how you set it up. You get your API key for free. It's AI.studio. And then you install the SDK, we have it available in different languages. And then you can simply upload different files. Like here, we are uploading a PDF, video and an MP3 file. Or you can also for smaller files directly use it as inline data.
SPEAKER_01
And as a tip, on the right side, I mentioned Gemini API skills. So you don't have to know this code by heart. You can just hook up your agent with the Gemini skill, and then tell it to create this. And then it should know how to work with the Gemini models. Yes, this is basically everything we need. And then we call client models generate content. And here we're using Gemini 3 Flash. And then we can put everything together into the contents list. And tell it, for example, analyze all these resources, give it a little bit more information what these resources contain, and then it should generate a summary.
SPEAKER_01
And then a little bit of practical tips or nice to know is for understanding. You can also use it to transcribe audio actually, Flash and even the smallest ones, Flashlight is pretty good at transcribing audio if you just tell it in the prompt generate a transcript of this file. And then maybe nice to know is for audio, one minute of audio translates to 1920 tokens. And Gemini has a token limit of 1 million. So if you do the math, it translates to more than nine hours of content you can feed in audio content. For video, it's roughly one hour. But there are configurations you can tweak that give you more control and you can even feed in longer audios.
SPEAKER_01
Then you can tell it to look at different timestamps. So for example, only analyze from minute five to minute 15. And then yes, you can use the file API that easily lets you upload larger files, you can even pass in URLs, YouTube URLs directly. And what's also nice to know is you can combine it with context caching. This is built into the API. This is especially useful if you're loading longer files into Gemini and doing repeated queries because then it saves you 90% of the costs. So yes, this is multimodal understanding in a nutshell. So here doing a quick checkpoint, we're now able to use Gemini to understand all these different resources and generate a summary.
SPEAKER_01
Actually, the timer is not working. So I don't know how much longer I have. But I think we are still good. So yes, then the next phase is the multimodal generation part. So for this, we're using the agentic loop, where we use Gemini as the brain behind this. And we combine it with function calling, and we combine it with function calling and we combine it with function calling and I will show you how to do this in a moment. And then these function call the specialized native generation models. And then it can reason if these assets are enough or if we need more.
SPEAKER_01
And the way to do to use these specialized models is also basically the same code. Once you have the SDKs, you call again client models generate content. In this case, we're using Gemini 3.5 Flash image preview. It's not the nicest model, but this is actually Nano Banana 2, the more famous model. And then we tell it to create a picture or in this case, we can it's pretty good at creating infographics, which is pretty cool. Just give it in your prompt and create an infographic. And then it's creating these nice infographics for us. And similar for text to speech, there we have a text to speech model, which currently is still based on Gemini 2.5.
SPEAKER_01
And you can combine it with different configurations. You can also do two speaker audio files. So for example, this podcast style and here is a nice example if the sound works. Does it work?
SPEAKER_01
Neural network.
SPEAKER_01
Neural network architecture introduced in a 2017 Google paper called attention is all you need. Transformers are the revolutionary models. We have most powerful AI like transformers in only two minutes for you. Neural network.
SPEAKER_01
String, which is then used for the prompts, so the detailed description of how the image should look. And then you do the same for the audio generation function. And then if you set up your model call, client models generate content, you configure the tools. And then you also need to add this to your prompt. So this is a small example prompt, an agent prompt, where you tell Gemini, hey, here's the study we synthesized before from the different modalities. And now you're a research agent partner. Your job is to enhance the study guide with multiple materials. And then you do, you tell it the two functions.
SPEAKER_01
So decide which concepts are complex enough to need a visual diagram. And for this call generate image. And which sections would benefit from audio summary. And for this call generate speech. And yes, this is basically everything you need to set up the agentic function calling for the multimodal generation part. And I also quickly wanted to touch on why native generation matters. So we call this native image generation models, for example, because they are based on Gemini. So all the training or a lot of the training that goes into the main Gemini models are now also available in these models.
SPEAKER_01
And this allows a lot of really cool use cases, because these models understand the world. This on the left side is, for example, I found on Twitter. An example I really like from Nano Banana One, where you can draw errors on maps and just tell it, hey, create a picture of what you see here. And since Gemini understands the world here, it's able to correctly create a picture of the Golden Gate Bridge for you. And then on the right side, this is a nice example from the educational space. So you can use Nano Banana. In this case, it was Nano Banana Two to directly correct your math homework, for example, and create pictures with the corrections, because it understands math.
SPEAKER_01
It can even generate code on images for you. So lots of nice use cases. And for the audio models, these are multilingual, and they understand accents and tone. Hey up, lads and lasses. We're getting stuck into building these multimodal agents today. No faffing about. Let's just get them sorted, so we can all nip to the pub for a proper pint. Was it a good British accent? Yes. Is anyone speaking German here? One, two, three. I have a Bavarian accent, because you can actually also tell it to create different accents. Of course not every accent in the world, but still. So here's one with a Bavarian accent. Servus miteinander.
SPEAKER_01
Heute schauen wir uns an, wie wir diese multimodalen Agenten zusammenschrauben, gäh. Wenn das drum... Was it a good Bavarian accent? Am Ende gescheit leuft. Geh mir erst am... Servus miteinander. So yes, I would say this is pretty... Okay. Yes. And yes, again, so quick checkpoint. You now know how to do the understanding part and the generation part. And this is basically already the Notebook LM clone. And now I quickly wanted to mention now, or we also have a model for real time interaction with it via our, what we call the live API. And for this, we have a very new model, which is also based on Gemini, Gemini 3.1 Flash life. And we call this an audio to audio model.
SPEAKER_01
So native audio generation, it's only one architecture. [SPEAKER_02] Audio goes in and audio goes out. [SPEAKER_00] So you no longer have this cascaded pipeline with different models. [SPEAKER_00] And this allows a lot of really cool natural sounding interactions. [SPEAKER_00] I think I don't have time for a live demo, but you can try it at AI.studio live. And here's a quick video from one of our colleagues, Thor. Hey Gemini. How are you today? You kind of got to start it and start talking to it. And then you can also activate your camera. I'm doing grand. Thanks for asking. Just enjoying the chat, you know? And how are things with you? Can you see me?
SPEAKER_01
Well, it's plain as day. I see you there with your short hair and beard.
SPEAKER_02
[SPEAKER_01] Wearing a grand dark jacket over a blue shirt.
SPEAKER_00
[SPEAKER_01] But yes, try it out for yourself at AI.studio slash live. [SPEAKER_01] And I think that's almost it. [SPEAKER_01] Yes, this is again how you can do it in the code.
SPEAKER_01
But there again, we have a skill for it that you can configure.
SPEAKER_01
And this, yes, now we are at all the three checkpoints. And yes, the pattern is transferable to every other field. And maybe a few shout outs as well to some other models. I'm not sure if you've seen the keynote this morning. But we now have a multimodal embedding model where you can combine all the different modalities into one unified vector space, which allows applications like multimodal search. And then again, you can go local with Gemma 4 and again have this multimodal understanding. And Vio for an image for video with native audio. And yes, so that's it. Thank you. Thanks, everyone, and have fun building multimodal agents.
SPEAKER_01
Thanks, everyone, and have fun building multimodal agents. I see you there with your short hair and beard. Wearing a grand dark jacket over a blue shirt. But yeah, try it out for yourself at AI.studio slash live. And I think that's almost it. Yeah, this is again how you can do it in the code. But there again, we have a skill for it that you can configure. And this, yeah, now we are at all the three checkpoints. And yeah, the pattern is transferable to every other fields. And maybe a few shout outs as well to some other models. I'm not sure if you've seen the keynote this morning.
SPEAKER_01
But we now have a multimodal embedding model where you can combine all the different modalities into one unified vector space, which allows applications like multimodal search. And then again, you can go local with Gemma 4 and again have this multimodal understanding. And Vio for an image for video with native audio. And yeah, so that's it. Thank you. Thanks, everyone, thank you, and have fun building multimodal agents. Thanks, everyone, thank you, and have fun building multimodal agents.