SPEAKER_00
What's new in AI audio? I'm sorry it's a little bit misleading because the title leaves out the at Google DeepMind so we're just looking at what we've been working on at DeepMind. If we were to look at everything in AI audio we'd be spending a lot of time here but I'd love to show you what we're working on at DeepMind. This is me hi everyone I'm Thor I work on the developer experience at Google DeepMind working on the Gemini API and Google AI studio. Hello zusammen, herzlich willkommen my name is Thorsten. Bonjour, je m'appelle Thor, je suis très désolé mon français c'est très mauvais. Konnichiwa, orwa, rei jingda.
SPEAKER_00
Daxianhao, wos shu zuzai megwater dökoren, buha yi se wode tongwen hai buhao. Okay that was for the demo and now I just need to make sure last time I did this demo I recorded over it and then it was all gone that was very sad but we'll come back to that in a bit.
SPEAKER_00
Yeah what have we been up to at DeepMind? There's been a couple releases I actually joined the team in November literally the day before Gemini 3 was released so I joined and they told me tomorrow we're releasing Gemini 3 and I was like yay didn't do anything but it was great. Most recently on the open model side we released Gemma 4 I think literally last week and yeah pretty incredible some cool stuff you can do there multi modality as well baked into Gemma 4 so there's audio understanding in the Gemma 4 models and you can do that on device on edge devices as well so that is some very exciting stuff in terms of GenMedia and audio. You're probably very familiar with our image generation models video generation models obviously VO has audio generation in there as well so this is the progression there most recently with VO 3.1 light on the GenMedia model side and then on the audio models we recently launched Gemini 3.1 flash life which is our full duplex sound to sound real-time conversational model also multimodal so you can ingest real-time text voice vision which we'll look at in a bit.
SPEAKER_00
So on audio very very broad topic but the baseline of everything we do are the frontier Gemini models and Gemini 3 is incredibly good at understanding audio and that's not just transcribing it but really understanding all the nuances that are in there so that might be obviously speech but also the context of the speech the emotion your pacing anything that swings within the audio that's not just text.
SPEAKER_00
So on the audio understanding our goal is to build models that deeply comprehend, richly transcribe and robustly reason through audio seamlessly handling a large mix of different languages dialects accents and modalities and anywhere and always Gemini is really good at transcribing even people that are talking over each other which is pretty incredible. Seamlessly switching between different languages that were the demo we're looking at now.
SPEAKER_00
So Echo Script is Gemini three flash preview to analyze audio recordings and extract information out of it. It is built with google AI studio so you can find it in the gallery in AI studio you can try it out I can give you the slides later as well. So that was what I was trying to demo earlier so different from just a pure transcription model we can extract a lot of information out of the audio within one single request to the model or one single API request if we're using the API.
SPEAKER_00
So you can see here you know summary I introduced myself by name so we're actually able to label the section the speaker by name I forgot there was no hecklers in the room otherwise we would have picked that up as well and maybe we can see if we have what time later and we can do that but so we can see here we're extracting time stamps we're labeling the speaker identifying the speaker we're identifying the language and the emotion of it right and happy to introduce myself it's great. Now you can see this was in german normally it would classify my german as angry but here I guess I'm very happy to be with you all so I just told it to label the emotion label the language if it's a language that is not english give me an english translation as well right. In french neutral normally I would say sad you know french is just a bit more of a no I said you know I'm sorry my french is very bad didn't sound sad enough so neutral in this case.
SPEAKER_00
Okay this didn't work so my japanese I gotta practice that if anyone reads japanese so it should actually say hello my name is thor unfortunately bit of a miss there. Let's see if my mandarin was any better hello everyone I'm a german living in the united states sorry my chinese yeah that is correct does anyone read chinese in the room. No okay well we'll just trust that that is correct and so we can see here that this was one request to the model where I basically just told it to identify the distinct speakers if you have contacts label it by name label the speakers by name give me the accurate time stamps give me the language if the language is not english give me the translation identify the emotion out of happy sad angry neutral and then also provide a brief summary of the entire audio at the beginning so this was one api call to gemini 3 flash preview and we got all this information out we could I just gave it a response schema so structured outputs and I was able to just populate that into my ui to have the structure.
SPEAKER_00
So this audio understanding and the base research in the gemini 3 models that is what powers the speech generation as well as the real-time conversational generation so having that audio understanding is really great in terms of knowing what certain things sound like including different pacing different accents and scenarios like that.
SPEAKER_00
So the foundation of all our models is the gemini 3 foundational research and then we're building the dedicated audio models on top of that and so with speech generation it's a bit different you know if you've used other tts providers before you probably have a huge library of different voices that you filter by gender by accent by languages what have you but in gemini you have just I think it's like 30-aust sort of base voices and then what you do is you direct that voice to act in a certain way and again
SPEAKER_00
different pacing different accents and scenarios like that so the foundation of all our models is now the gemini 3 foundational research and then we're building dedicated audio models on top of that and so with speech generation it's a bit different if you've used other TTS providers before you probably have a huge library of different voices that you filter by gender by accent by languages what have you but so in Gemini you have just I think it's like 30 or so base voices and then what you do is you direct that voice to act in a certain way and again because we have that audio understanding we can basically modify the voice to act in a certain way to use a certain accent and so we can go from a small set of base voices to a very specific voice that we're looking for for our speech generation. Again there's a little application that you can try out it's in the Google AI Studio gallery as well it's called the voice library and so what we can do is giving the prompt structure that we just saw we're building the audio profile the scene we're setting the scene we're instructing sort this director's note so we're giving guidance for the performance just like how you would direct a human to act out a certain way and then some sample context and the transcript that we want so now what we can do is we would just set we want high-pitch Irish male and so basically I just use Gemini 3 flash here again to then construct our system prompt for the speech generation so we're saying here we're sending sort of our audio profile you know Finion here in the scene cozy crowded pub on the coast of county Clare you know delivered the lines with a strong authentic Irish accent and so now we hope the TPUs don't disappoint me there we go it failed but I didn't you know I prepared it so we can listen to it here.
SPEAKER_00
Oh you wouldn't believe the size of the thing until you saw it with your own two eyes I'm telling you it was a grand old mess so it was and we were all laughing fit to burst by the end of the night.
SPEAKER_00
So as you can see this was the base voice here is this one what kind of problem could we solve so you know that is a fairly standard American accent but so now by giving it that director's note we can then sort of give oh you wouldn't believe the size of the thing until you saw it with your own two eyes I'm telling you it was a grand old mess so it was and we were all laughing or you know similarly here we have Sapphire so this voice is here ready to build something awesome today again fairly standard American English accent here and now we could say you know give it a Singaporean sort of scene wow you must try this chicken rice la the chili is damn sure confirm plus job you will love it faster cue before the uncle close shop okay anyone spend time in Singapore that you know yeah that's you know that's something you'd hear in the hawker center so again you know that is underpinned by the audio understanding so the model really understands what these different scenarios sound like and then can modify the speech generation to be like that.
SPEAKER_00
Yes and then you know finally the native audio sound to sound multimodal real time so we just launched a couple weeks ago Gemini 3.1 Flash Live so it is a speech to speech real-time multimodal model you can ingest text audio video in real time through a WebSocket connection and then you get a real-time audio response back as well as the text transcript of that you know obviously benchmarks are you know especially in the audio space benchmarks you can't really trust them you know it's great you can see the reasoning the thinking so here the thinking and the reasoning and the intelligence is baked directly into the model so that's different from a cascading pipeline where you would actually go through text to then go through an LLM to get the intelligence here the intelligence is baked into the audio model so that's the difference there but obviously in real scenarios you know you can try this out in ai.studio slash live so the great thing with AI Studio is you can try it out you know without paying anything so this allows you to try the models without actually needing to put down the credit card so AI Studio is a really great way to do that and again you know we have the audio understanding baked into the model here so what we can do is we can give it some system instructions you know for example speaking a friendly Irish accent and then also we can ingest our camera here for example and then we can say hey can you see me well hello there I can see you loud and clear so I can what's on your mind today what do you think of my outfit ah look at you with your Gemini shirt it looks grand all together and the backwards hat gives you a fierce laid back vibe so it does you're looking sharp.
SPEAKER_00
[SPEAKER_03] Ah wonderful, canst du mir ein gedicht auf deutsch erzählen bitte. Ah a poem in German is it sure I can give that a go for you here's a little one.
SPEAKER_00
So obviously you need to adjust your system instructions to not speak in an Irish accent in every language. So it's pretty funny you can switch between the different languages there. Again, ai.studio.com slash live you can try it out. You could also ingest your screen. So you're basically just ingesting video frames in addition to the audio at a maximum frame rate of one frame per second at the moment. That's the maximum you can ingest. There are a couple examples. If you're a developer yourself, feel free to try these out. They are linked from the docs as well. These are just some examples of how you would get started, say in Python for a server to server connection or in JavaScript with.
SPEAKER_00
[SPEAKER_03] speak in an Irish accent in every language. So it's pretty funny you can [SPEAKER_03] switch between the different languages there. Again, so ai.studio.com slash life you can try it out. You could also ingest your screen. So you're ingesting video frames in addition to the audio at a maximum frame rate of one frame per second at the moment. That's the maximum you can ingest. There are a couple of examples. If you're a developer yourself, feel free to try these out. They are linked from the docs as well. These are just some examples of how you would get
SPEAKER_00
started, say in Python for a server to server connection or in JavaScript with a client to server connection. So you can find those here. Also recommend the Gemini skills. So we have published coding agent skills for all of the Gemini APIs, including the live API. So real-time audio, working with real-time audio can just be a bit more challenging. So using these agent skills and installing them in your coding agents can really help steer them in the right way and give you that result that you're looking for. And lastly, we have music, so also audio. I put that in here. We recently released Lyria 3.
SPEAKER_00
So it's a music generation model. But it can now actually generate music with lyrics. There are two separate models. There's a Lyria 3 clip, which is a 30-second jingle generation model. And then Lyria 3 Pro is the full-length song generation model. Bringing everything together, I don't know if you remember the good old days when you would actually call the radio station and request a song. So I've built this application called Life Jukebox. And we can now try it out where we actually give the real-time Gemini Life model a tool to then generate a song using Lyria. And so let's see if we can do that here.
SPEAKER_00
Right then, you're live on air with the Jukebox DJ. What an absolute legend. Who am I speaking to? And what kind of vibe are we going to cook up today? Is it a proper rave starter or something a bit more... Can we get maybe a German technoschlager about the UK startup scene? German technoschlager about the UK startup scene? Mate, that is revolutionary. I love it. Right, picture this. Pounding techno beat, a proper Hans Zimmer style. Yeah, that's great. Let's go. Let's do that. Hold on tight. We're almost there. Just to make sure it's a certified banger. Any specific buzzwords or stories from the UK startup world you want in the lyrics?
SPEAKER_00
And are we thinking manic energy or something a bit more melodic within that techno madness? Nah, just manic energy and surprise me on the lyrics. Alright, you're on. Get ready to have your eardrums blasted. Manic German technoschlager with a British startup twist. Cooking up a proper banger for you. Check this out. All right. I'll leave you with that. Thank you very much. If you don't speak German, sorry. But thanks so much. Appreciate you all. No worries. Oh yeah. And if you want the slides, I can just rewind. There are all the links in there. If that's helpful, you can just grab them there.
SPEAKER_00
Awesome. Thank you. And yeah, enjoy the rest of the conference. And big thank you as well to our friends in the back on the audio. You know, it wouldn't be possible without them. Cheers. Oh yeah. Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. Oh yeah. [SPEAKER_03] Oh yeah.
SPEAKER_00
[SPEAKER_03] Oh yeah. Oh yeah. [SPEAKER_03] Oh yeah. different pacing different accents and and scenarios like that so um you know the foundation of kind of all our models is now sort of the gemini 3 um foundational research and then we're building kind of the um dedicated audio models on on top of that and so with speech generation uh it's a bit different you know if you've used kind of other um tts providers before uh you probably have a huge library of you know different voices that you sort of you know you filter by gender by by you know accent by languages what have you um but so in in in gemini you have you know just i think it's like 30-aust sort of base
SPEAKER_00
voices and then what you do is you you kind of direct that voice to act in a certain way and again because we have that audio understanding uh we can we can basically modify the voice to you know act in a certain way to you know use a certain accent and so we can go from kind of a small set of of base voices to a very specific you know kind of voice that we're looking for for our speech generation uh again there's a a little application that you can try out uh it's in the uh google ai studio gallery as well uh it's called the voice library uh and so what we can do is you know um kind of giving the
SPEAKER_00
the the prompt structure that we just saw you know we're building sort of the the audio profile the scene we're setting the scene we're instructing sort of this director's note so we're giving guidance for the performance you know just like how you would um direct a human you know to to act out a certain way um and then some sample context and kind of the transcript um that we want so now what we can do is you know we would just set sort of uh we want you know high-pitch irish mail uh and so basically i just use gemini um three flash here again to then construct our um system prompt for the
SPEAKER_00
speech generation so we're saying here you know um we're sending sort of our audio profile you know finion here in the scene sort of cozy crowded pop uh in the coast of county clare um you know delivered the lines with a strong authentic uh irish accent and so now we hope the the tpus don't uh disappoint me there we go it failed but um i didn't you know i prepared it so we can we can listen to it here oh you wouldn't believe the size of the thing until you saw it with your own two eyes i'm telling you it was a grand old mess so it was and we were all laughing fit to burst by the end of the night
SPEAKER_00
so as as you can see you know this was uh the the bass voice here is uh this one what kind of problem
SPEAKER_02
could we solve so you know that is a fairly sort of stand standard you know american accent but so now by you know giving it that director's note we can then sort of give oh you wouldn't believe the size of the
SPEAKER_00
thing until you saw it with your own toys i'm telling you it was a grand old mess so it was and we were all laughing or you know similarly here we have um sapphire so this voice uh is here ready to build something awesome today again you know kind of fairly standard sort of american um
SPEAKER_02
english accent here and now we could say you know give it kind of a singaporean sort of scene wow you
SPEAKER_00
must try this chicken rice la the chili is damn sure confirm plus job you will love it faster cue before the uncle close shop okay anyone spend time in singapore that you know yeah that's uh you know that's something you'd hear in the hawker center so um again you know that is kind of underpinned by
SPEAKER_04
the audio understanding so um the model really understands what you know these these different
SPEAKER_00
scenarios sound like and then can modify uh the speech generation to to to be like that um yes and then you know finally sort of the the native audio uh sound to sound multimodal real time so we uh just launched a couple weeks ago gemini 3.1 flash life um so it is a speech to speech just kind a real-time multimodal model uh you can ingest uh text audio video in real time through a websocket connection and then you get a real-time audio response back as well as kind of the text transcript of that um you know obviously benchmarks are you know especially in the audio space benchmarks you
SPEAKER_00
know you can't really trust them um you know it's great you can see sort of the reasoning the thinking uh so so here the the the thinking and the reasoning and the intelligence is baked directly into the model so that's you know different from a cascading pipeline where you would actually go uh through text to then go through an llm to get the intelligence here you know the intelligence is baked into the audio model so that's kind of the difference there um but obviously in in real scenarios um you know you can try this out in uh ai.studio slash life so the great thing with ai studio is uh you can try it out um
SPEAKER_00
you know without paying anything so this is uh allows you to you know try kind of um the models without actually needing to put down the credit card so ai studio is a really great way to do that uh and again you know we have the audio understanding kind of baked into the model here so what we can do is we can give it some system instructions you know for example speaking a friendly irish accent and then also we can ingest kind of our you know camera here for example and then um we can say hey can you see me well hello there i can see you loud and clear so i can what's on your mind
SPEAKER_00
today what do you think of my outfit ah look at you with your gemini shirt it looks grand all together and the backwards hat gives you a fierce laid back vibe so it does you're looking sharp
SPEAKER_03
ah wonderful canst du mir ein gedicht auf deutsch erzählen bitte ah a poem in german is it sure i can give that a go for you here's a little one
SPEAKER_03
just to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring
SPEAKER_00
as well. So obviously you need to adjust your system instructions to not
SPEAKER_03
speak in an Irish accent in every language. So it's pretty funny you can switch between the different languages there. Again, so ai.studio.com slash life
SPEAKER_00
you can try it out. You could also ingest your screen. So you're basically just ingesting video frames in addition to the audio at a maximum frame rate of one frame per second at the moment. That's kind of the maximum you can ingest. There is a couple examples. If you're a developer yourself, feel free to try these out. They are linked from the docs as well. These are just some examples of how you would get started, say in Python for like a server to server connection or in JavaScript with kind of a client to server connection. So you can find those here. Also recommend the Gemini skills. So we have published coding agent skills for kind of all of the Gemini APIs,
SPEAKER_00
including the live API. So, you know, real-time audio, working with real-time audio can just be a bit more challenging. So using these agent skills and kind of installing them in your, you know, coding agents can really help steer them sort of in the right way and give you sort of that result that you're looking for. And lastly, we have, okay, we have a bit of time. So music, you know, also audio. So I put that in here. We recently released Lyria 3.
SPEAKER_00
So, yeah, it's a music generation model. But so it now actually can generate music with lyrics. There's two separate models. There's a Lyria 3 clip, which is a 30-second kind of jingle generation model. And then Lyria 3 Pro is the full-length song generation model. And so, you know, kind of bringing everything together, I don't know if you remember the good old days when you would actually call the radio station and wish for a song. And so I've kind of built this application called Life Jukebox. And so we can now try it out where we actually give the real-time Gemini Life model a tool to then generate a song using Lyria. And so let's see if we can do that here.
SPEAKER_00
Right then, you're live on air with the Jukebox DJ. What an absolute legend. Who am I speaking to? And what kind of vibe are we going to cook up today? Is it a proper rave starter or something a bit more... Can we get maybe a German technoschlager about the UK startup scene? German technoschlager about the UK startup scene? Mate, that is revolutionary. I love it. Right, picture this. Pounding techno beat a proper Hans Zimmer style. Yeah, that's great. Let's go. Let's do that. Hold on tight. We're almost there. Just to make sure it's a certified banger. Any specific buzzwords or stories from the UK startup world you want in the lyrics?
SPEAKER_00
And are we thinking manic energy or something a bit more melodic within that techno madness? Nah, just manic energy and surprise me on the lyrics.
SPEAKER_00
Alright, you're on. Get ready to have your eardrums blasted. Manic German technoschlager with a British startup twist. Cooking up a proper banger for you. Check this out.
SPEAKER_00
All right. I'll leave you with that. Thank you very much. If you don't speak German, sorry. But thanks so much. Appreciate you all. No worries.
SPEAKER_00
Oh yeah. And if you want the slides, I can just rewind. There's like all the links in there. If that's helpful, you can just grab them there. Awesome. Thank you. And yeah, enjoy the rest of the conference. And big thank you as well to our friends in the back on the audio. You know, it wouldn't be possible without them. Cheers. Oh yeah. Oh yeah.
SPEAKER_03
Oh yeah.
SPEAKER_03
Oh yeah.
SPEAKER_03
Oh yeah. Oh yeah. Oh yeah.
SPEAKER_00
Oh yeah.
SPEAKER_03
Oh yeah. Oh yeah.
SPEAKER_00
Oh yeah.
SPEAKER_03
Oh yeah.