Open Reader

From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind

completed 19:33 Jun 09, 2026 Watch on YouTube

Current Status

completed

Video ID

Bc6Ojl2XS1w

RAG / Chat

Enabled
From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind
Description

One API call to Gemini 3 Flash Preview: speaker labels by name, timestamps, emotion tags, language detection with English translation, and a full summary. That is the audio understanding layer that underlies everything else Thor Schaeff demos here, including speech generation directed by a "director's note" rather than picked from a catalogue, and Gemini 3.1 Flash Live, a sound to sound real time multimodal model with thinking baked in rather than cascaded through a separate LLM. The talk ends with Lyria 3, Google DeepMind's music generation model that can now produce full songs with lyrics. The live demo has the Gemini Live model calling Lyria via tool use on request to generate a German techno schlager about the UK startup scene, live on stage. Speaker info: - https://x.com/thorwebdev - https://www.linkedin.com/in/thorwebdev

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Google DeepMind has built a unified audio stack on Gemini 2 that integrates deep audio understanding (transcription + nuance), voice direction (30 base voices modifiable via natural language), and real-time multimodal conversation (Gemini 2.1 Flash Live) plus music generation (Lyria 3), all underpinned by native audio intelligence rather than cascaded TTS pipelines.
  • Why it matters: This is the first major demo showing how a single foundation model (Gemini 2) powers transcription, voice synthesis, real-time duplex conversation, and music generation with unified audio understanding—critical for anyone building conversational agents, content tools, or audio-first products.
  • Best use: Watch the live demos (especially EchoScript multi-language transcription, voice library accent direction, and Live Jukebox tool-calling music generation) to understand the actual developer experience, API shape, and system instruction patterns for Gemini audio APIs.

Executive Summary

Thor Schaeff, Developer Experience lead at Google DeepMind, walks through the company's newly launched audio stack built on Gemini 2 foundation models. The core thesis is that Gemini 2 natively understands audio—not just transcribes words but captures emotion, pacing, speaker identity, language switching, and overlapping speech—and this deep understanding powers all downstream audio tasks. He demonstrates EchoScript (Gemini 2 Flash Preview), which extracts speaker labels, timestamps, language identification, emotion, and translations from a multilingual recording in a single API call via structured outputs. The model correctly identified English, German, French, and Mandarin segments (though Japanese failed), labeled emotions, and provided translations—all without cascading services.

For speech generation, DeepMind uses ~30 base voices that are directed via natural language prompts (a 'director's note' approach) rather than selecting from a library of hundreds of pre-recorded voices. Because Gemini 2 understands audio deeply, a single base voice (e.g., standard American English) can be instructed to adopt a strong Irish accent, Singaporean Singlish cadence, or other performance styles. Thor demos this with the Voice Library tool in Google AI Studio, showing how the same base voice transforms when given scene context (e.g., 'cozy crowded pub on the coast of County Clare') and accent instructions. This approach collapses the traditional TTS voice catalog problem into a prompt engineering problem.

The centerpiece is Gemini 2.1 Flash Live, a real-time speech-to-speech multimodal model that ingests text, audio, and video (up to 1 fps) over WebSocket and returns audio responses plus text transcripts. Unlike cascaded pipelines (audio → STT → LLM → TTS), the intelligence and reasoning are baked directly into the audio model, reducing latency and preserving nuance. Thor demos live conversation with camera input (the model comments on his Gemini shirt and backward hat) and language switching (it recites a German poem in an Irish accent until corrected). The model is accessible for free experimentation at ai.google.dev/aistudio/live, removing credit card friction for developers. Gemini coding agent skills are available to help developers scaffold real-time audio integrations correctly.

Finally, Thor demos Lyria 3 (music generation with lyrics), available in two variants: Lyria 3 Clip (30-second jingles) and Lyria 3 Pro (full songs). He showcases 'Live Jukebox,' a demo app where Gemini 2.1 Flash Live acts as a radio DJ, takes user requests via voice, and uses function calling to trigger Lyria 3 to generate a custom song (e.g., 'German techno schlager about the UK startup scene'). This closes the loop: real-time conversational AI with tool use for generative audio. The entire stack—transcription, voice direction, real-time conversation, music generation—is unified under Gemini 2's native audio understanding, offering a single API surface and model family for multimodal audio applications.

Key Takeaways

  • Claim: Gemini 2 models natively understand audio beyond transcription—capturing emotion, pacing, language, speaker overlap, and context—which powers all downstream audio tasks. | Evidence: EchoScript demo (Gemini 2 Flash Preview) extracted speaker names, timestamps, language labels (English, German, French, Mandarin), emotion (happy, neutral, sad), and English translations from a multilingual recording in one API call using structured outputs. The model identified that Thor spoke in German ('herzlich willkommen') and labeled it 'happy,' correctly inferred speaker intent ('I'm sorry my French is very bad') and labeled it 'neutral,' and provided translations—all without separate STT/translation/NLU steps. | Caveat: Japanese transcription failed in the live demo (output was garbled), and Thor noted this publicly. The model's performance on low-resource languages or heavy accents is not quantified. Emotion classification is limited to four categories (happy, sad, angry, neutral), which may oversimplify nuance. | Implication: For Ken: If building voice agents or content analysis pipelines, Gemini 2's single-pass multimodal understanding eliminates the need for stitching together STT + LLM + sentiment analysis services. This reduces latency, API complexity, and cost, but you'll need to test edge cases (accents, code-switching, non-Latin scripts) for your domain. The structured outputs feature (response schema) is key for populating UIs or downstream workflows. | Timestamp: 03:15
  • Claim: Gemini voice synthesis uses ~30 base voices directed via natural language prompts (scene, accent, pacing) rather than a catalog of hundreds of pre-recorded voices. | Evidence: Thor demoed the Voice Library tool in Google AI Studio, showing base voice 'Finion' (standard American English) transformed into a strong Irish accent by providing scene context ('cozy crowded pub on the coast of County Clare') and director's note ('delivered the lines with a strong authentic Irish accent'). Output: 'Oh you wouldn't believe the size of the thing until you saw it with your own two eyes…' in convincing Irish cadence. Another example: base voice 'Sapphire' (American English) became Singaporean Singlish ('Wow you must try this chicken rice la the chili is damn shiok confirm plus chop you will love it faster queue before the uncle close shop') when given Singaporean scene context. | Caveat: Thor did not disclose how many base voices exist (~30 mentioned), whether users can upload custom reference audio, or how fine-grained control works (e.g., pitch, speed, timbre). The system prompt construction (audio profile, scene, director's note, transcript) is manual prompt engineering; no GUI sliders for prosody. Quality of accent/dialect fidelity was not objectively measured (only live audience reaction). | Implication: For Ken: This is a paradigm shift from traditional TTS (pick from 500 voices) to voice-as-prompt-engineering. You can generate context-specific performances (e.g., customer support bot with regional accent matching caller, podcast narration with emotional arcs) from a small set of base voices, reducing storage/licensing overhead. However, you'll need to invest in prompt templates and QA for consistency. This approach favors generative flexibility over deterministic repeatability—ideal for content creation, less so for compliance-sensitive applications where exact pronunciation/tone must be reproducible. | Timestamp: 07:45
  • Claim: Gemini 2.1 Flash Live is a real-time speech-to-speech multimodal model (audio + text + video up to 1 fps over WebSocket) with intelligence baked into the audio model, not a cascaded STT → LLM → TTS pipeline. | Evidence: Thor demoed live conversation at ai.google.dev/aistudio/live with camera input. The model responded in real time with Irish accent (per system instruction), commented on visual input ('your Gemini shirt looks grand all together and the backwards hat gives you a fierce laid-back vibe'), and switched languages on request (recited a German poem in Irish accent, then corrected itself when Thor noted the accent mismatch). Response latency appeared sub-second. Thor emphasized that reasoning/intelligence is in the audio model itself, not a separate LLM step, which reduces cascading latency. | Caveat: Video frame rate is capped at 1 fps, limiting real-time vision use cases (no object tracking, fast motion). No disclosed latency/jitter SLAs, cost per minute, or rate limits for the Live API. The model's ability to handle interruptions, turn-taking, or multi-party conversation was not demonstrated. System instruction adherence (e.g., 'speak in Irish accent') is not guaranteed across all languages (the German poem example showed it defaulted to Irish accent until corrected). | Implication: For Ken: This is critical for real-time voice agents (customer support, companions, sales bots). The unified audio intelligence means you avoid STT → LLM → TTS round-trips, which historically add 500ms+ latency. The multimodal capability (screen sharing at 1 fps) enables co-browsing, visual troubleshooting, or context-aware assistance. However, you'll need robust WebSocket infrastructure, prompt engineering for system instructions (voice, personality, task), and fallback handling for accent/language drift. Free experimentation in AI Studio (no credit card) lowers the barrier for prototyping, but production cost/scale unknowns remain. | Timestamp: 11:30
  • Claim: Lyria 3 generates full-length songs with lyrics (Lyria 3 Pro) and 30-second jingles (Lyria 3 Clip), and can be called as a tool by Gemini 2.1 Flash Live for conversational music generation. | Evidence: Thor demoed 'Live Jukebox,' a radio DJ agent built with Gemini 2.1 Flash Live. User requested 'German techno schlager about the UK startup scene with manic energy,' and the model used function calling to invoke Lyria 3, which generated and played a custom techno track with German lyrics live on stage. The interaction was conversational (model asked clarifying questions: 'any specific buzzwords or stories?'), and the final output was a full audio track delivered in ~20 seconds. | Caveat: Music quality was not objectively assessed (audience reaction suggested it was recognizable as techno but unclear if production-ready). No details on licensing, copyright handling, training data provenance, or whether generated music can be used commercially. Lyria 3 availability (API access, pricing, rate limits) was not disclosed—Thor only mentioned it was 'recently released.' The function calling pattern (how the model decides to invoke Lyria vs. continuing conversation) was not explained. | Implication: For Ken: This is the first demo of real-time conversational AI orchestrating generative audio tools. Use cases: on-demand podcast intros, dynamic ad jingles, personalized birthday songs, game soundtracks, content creation workflows. The key innovation is the model's ability to gather requirements conversationally, disambiguate, and trigger tool use—this is function calling for media generation, not just API retrieval. However, you'll need to clarify licensing/IP rights for production use, test music quality across genres, and design fallback UX if generation fails or takes too long. | Timestamp: 15:20
  • Claim: Google AI Studio offers free, no-credit-card access to experiment with Gemini models, including Live API and example galleries (EchoScript, Voice Library), and DeepMind publishes Gemini coding agent skills to scaffold real-time audio integrations. | Evidence: Thor repeatedly directed developers to ai.google.dev/aistudio and ai.google.dev/aistudio/live for free experimentation. He noted that AI Studio has a gallery with pre-built examples (EchoScript for transcription, Voice Library for TTS direction). He also mentioned that 'Gemini coding agent skills' are published and available in the docs, specifically calling out that real-time audio is 'a bit more challenging' and these skills help steer coding agents (e.g., Cursor, GitHub Copilot) to generate correct WebSocket/Live API code. | Caveat: No details on usage limits for free tier (requests per day, token caps, model access). The coding agent skills are not further described (e.g., which agents support them, what they contain, how they differ from docs). The gallery examples are demos, not production-ready templates—unclear how much boilerplate/error handling is needed for real apps. | Implication: For Ken: This is a major friction reducer for experimentation. You can prototype voice agents, transcription workflows, and multimodal conversations without upfront payment or onboarding. The coding agent skills are a strategic move—DeepMind is explicitly targeting developer tooling (Cursor, Copilot) to make Gemini audio APIs easier to scaffold, which accelerates adoption. For your agent systems, consider using these skills to generate starter code for Live API, then customize. The free tier is ideal for proof-of-concept before committing to GCP billing. | Timestamp: 13:00
  • Claim: Gemini 2's audio understanding enables seamless multilingual transcription with automatic language detection, translation, and speaker diarization in a single API call, even with overlapping speakers. | Evidence: Thor's EchoScript demo showed a recording where he spoke in English, German, French, Mandarin, and Japanese sequentially. Gemini 2 Flash Preview identified each language, provided English translations for non-English segments, labeled speakers by name (because context was available), assigned timestamps, and classified emotion—all returned as structured JSON via response schema in one request. Thor noted the model handles 'people talking over each other,' which is a known hard problem in traditional diarization. | Caveat: Japanese segment failed (garbled transcription). No quantitative accuracy metrics provided (WER, translation BLEU, diarization error rate). Unclear how the model performs with heavy accents, dialects, or low-resource languages. The demo used clear studio audio; performance on noisy/reverberant audio was not tested. | Implication: For Ken: If building global voice products (multilingual support, contact centers, podcasts, meetings), this one-shot multilingual understanding is a major simplification over traditional pipelines (detect language → STT per language → translate → diarize). You get all outputs in one call, which reduces latency and API orchestration complexity. However, test your specific language pairs and audio conditions—relying on a single model means a failure (like the Japanese example) can't be patched by swapping in a specialized service. The structured output feature is key for integrating into UIs or databases. | Timestamp: 03:15

Detailed Brief

Gemini 2 Audio Understanding: Foundation for All Audio Tasks

  • Claims: Gemini 2 models (especially 2 Flash Preview) deeply comprehend audio, not just transcribe it—capturing emotion, pacing, context, speaker overlap, language switching, and accents.; This native audio understanding is the foundation for all downstream audio capabilities: transcription, voice synthesis, real-time conversation, and music generation.; The goal is to build models that 'deeply comprehend, richly transcribe, and robustly reason through audio, seamlessly handling a large mix of different languages, dialects, accents, and modalities.'
  • Evidence: EchoScript demo: Single API call to Gemini 2 Flash Preview extracted speaker names, timestamps, language labels (English, German, French, Mandarin), emotion (happy, neutral), and English translations from a multilingual recording.; Structured outputs (response schema) allowed Thor to define the JSON shape upfront (speaker, language, emotion, translation, timestamp), and the model populated it directly.; Model identified overlapping speakers and language switches in real time during the demo.; Audio understanding extends beyond words: emotion ('happy' for German intro, 'neutral' for French apology), pacing, and context (e.g., recognizing 'I'm sorry my French is very bad' as self-deprecation, not sadness).
  • Caveats: Japanese transcription failed live (garbled output), showing language coverage gaps.; Emotion is limited to four categories (happy, sad, angry, neutral), which oversimplifies nuance.; No disclosed accuracy metrics (WER, translation BLEU, emotion F1, diarization error rate).; Demo used clean studio audio—performance on noisy, reverberant, or low-bitrate audio unknown.; Speaker diarization required context (Thor introduced himself by name)—unclear if model can label unknown speakers or distinguish similar voices.
  • Implications: For Ken: This is a single-API replacement for traditional pipelines (STT + NLU + sentiment + translation + diarization). Reduces latency, cost, and orchestration complexity.; Ideal for global voice products (multilingual support, international contact centers, podcast/meeting transcription).; Structured outputs make it trivial to populate databases, UIs, or downstream workflows—no more regex parsing of transcripts.; Risk: Single model dependency means no fallback if a language/accent fails (unlike swappable microservices). Test your specific languages/conditions before production.; The audio understanding as foundation strategy means improvements in Gemini 2 core models will lift all audio tasks (TTS, Live, Lyria) simultaneously.

Voice Synthesis via Direction: 30 Base Voices + Natural Language Prompts

  • Claims: Gemini voice synthesis uses ~30 base voices that are directed via natural language prompts (scene, accent, pacing, emotion) rather than selecting from a library of hundreds of pre-recorded voices.; Because Gemini 2 understands audio deeply, it can modify a single base voice to adopt accents, dialects, and performance styles on demand.; The prompt structure is: audio profile (base voice name, gender, pitch) + scene (context/setting) + director's note (performance instructions) + transcript (text to speak).
  • Evidence: Voice Library tool in Google AI Studio demonstrates this. Thor showed base voice 'Finion' (standard American English) transformed into strong Irish accent by setting scene ('cozy crowded pub on the coast of County Clare') and director's note ('delivered the lines with a strong authentic Irish accent').; Output: 'Oh you wouldn't believe the size of the thing until you saw it with your own two eyes I'm telling you it was a grand old mess so it was and we were all laughing fit to burst by the end of the night' in convincing Irish cadence.; Second example: base voice 'Sapphire' (American English) became Singaporean Singlish ('Wow you must try this chicken rice la the chili is damn shiok confirm plus chop you will love it faster queue before the uncle close shop') when given Singaporean scene context.; Thor used Gemini 2 Flash to construct the system prompt for speech generation (meta-prompting: using LLM to generate TTS prompt).
  • Caveats: Number of base voices is ~30 (not precise)—unclear if users can upload custom reference audio or train new base voices.; No GUI controls for prosody (pitch, speed, timbre)—all control is via natural language, which is flexible but non-deterministic.; Quality of accent/dialect fidelity was not measured (only live audience reaction).; Unclear how well the model handles uncommon accents, dialects, or voice characteristics (e.g., raspy voice, child voice, elderly voice).; No mention of voice cloning (using user's own voice as reference).
  • Implications: For Ken: This is a paradigm shift from TTS-as-catalog to TTS-as-prompt-engineering. You can generate context-specific performances (e.g., bot matches caller's regional accent, narrator shifts tone across story arcs) without managing hundreds of voice files.; Reduces licensing/storage overhead but increases prompt engineering overhead—you'll need templates and QA for consistency.; Ideal for content creation (podcasts, audiobooks, ads, games) where flexibility and variety matter. Less ideal for compliance-sensitive apps (e.g., medical, financial) where exact pronunciation and tone must be reproducible.; The meta-prompting pattern (using Gemini to generate TTS prompts) is interesting—could be systematized for dynamic voice persona generation.; Watch for voice cloning capability (not mentioned here)—if added, this could unlock personalized TTS without retraining.

Gemini 2.1 Flash Live: Real-Time Multimodal Speech-to-Speech

  • Claims: Gemini 2.1 Flash Live is a real-time speech-to-speech multimodal model that ingests text, audio, and video (up to 1 fps) over WebSocket and returns audio responses plus text transcripts.; Intelligence and reasoning are baked directly into the audio model, not a cascaded pipeline (STT → LLM → TTS).; Available for free experimentation at ai.google.dev/aistudio/live (no credit card required).; Supports system instructions (e.g., 'speak in a friendly Irish accent'), visual input (camera/screen sharing at 1 fps), and language switching.
  • Evidence: Live demo: Thor used ai.google.dev/aistudio/live with system instruction 'speak in a friendly Irish accent.' Model responded in real time: 'Well hello there I can see you loud and clear so I can what's on your mind today.'; Visual input: Model commented on Thor's appearance ('your Gemini shirt looks grand all together and the backwards hat gives you a fierce laid-back vibe').; Language switching: Thor requested a German poem, model recited it in Irish accent (system instruction bleed), then corrected itself when prompted.; Response latency appeared sub-second (audio played immediately after Thor stopped speaking).; Thor noted the model handles 'text, audio, video in real time' and outputs 'real-time audio response plus text transcript.'
  • Caveats: Video frame rate capped at 1 fps—limits real-time vision use cases (no fast motion tracking, no high-frequency visual context).; No disclosed latency/jitter SLAs, cost per minute, rate limits, or production availability details.; Interruption handling, turn-taking, and multi-party conversation not demonstrated.; System instruction adherence is imperfect (Irish accent bled into German poem until corrected).; Unclear how well it handles noisy audio, low bandwidth, or packet loss (demo used conference-quality mic/network).
  • Implications: For Ken: This is the real-time conversational AI primitive for voice agents. The unified audio intelligence (no STT/TTS round-trip) should deliver lower latency than cascaded systems (Whisper → GPT-4 → ElevenLabs = ~1s+ total).; Multimodal capability (1 fps video) enables co-browsing, visual troubleshooting, or context-aware assistance (e.g., customer support sees user's screen, repair bot sees equipment).; Free experimentation in AI Studio removes friction for prototyping—you can test voice personas, interruption handling, and system instructions without billing.; Production deployment requires WebSocket infrastructure, prompt engineering for system instructions (voice, personality, task), and fallback handling for network issues or model failures.; Watch for: cost per minute (Gemini Live pricing not disclosed), rate limits, and whether the free tier persists or becomes paid in production.; The 1 fps video limit is a key constraint—this is not for real-time AR/VR or fast-moving visual tasks. It's for context (show me your screen, look at this document), not tracking.; For agent systems: this is the first major model to natively unify audio reasoning and generation. You can build voice agents that reason over conversation history, visual context, and tool use without stitching together multiple APIs.

Lyria 3 Music Generation + Live Jukebox Demo

  • Claims: Lyria 3 generates full-length songs with lyrics (Lyria 3 Pro) and 30-second jingles (Lyria 3 Clip).; Gemini 2.1 Flash Live can use function calling to invoke Lyria 3 for conversational music generation.; Demo: 'Live Jukebox' app where user requests a song via voice, model asks clarifying questions, then generates and plays the track.
  • Evidence: Thor built 'Live Jukebox,' a radio DJ agent using Gemini 2.1 Flash Live with a tool to call Lyria 3.; User request: 'Can we get maybe a German techno schlager about the UK startup scene?'; Model response (in Irish accent): 'German techno schlager about the UK startup scene? Mate that is revolutionary I love it. Right picture this pounding techno beat a proper Hans Zimmer style.' Model asked follow-up: 'Any specific buzzwords or stories from the UK startup world you want in the lyrics? And are we thinking manic energy or something a bit more melodic within that techno madness?'; User: 'Nah just manic energy and surprise me on the lyrics.'; Model: 'Alright you're on get ready to have your eardrums blasted manic German techno schlager with a British startup twist cooking up a proper banger for you check this out.'; Model then played a generated techno track with German lyrics (audience could hear it but transcript does not capture audio quality).
  • Caveats: Music quality not objectively assessed—unclear if production-ready or demo-quality.; No details on Lyria 3 API access, pricing, rate limits, or availability (Thor only said 'recently released').; No mention of licensing, copyright handling, training data provenance, or commercial use rights for generated music.; Function calling pattern (how model decides to invoke Lyria vs. continuing conversation) not explained.; Latency from request to playback was ~20 seconds (unclear if this includes generation time or just API round-trip).
  • Implications: For Ken: This is the first public demo of real-time conversational AI orchestrating generative audio tools. Use cases: on-demand podcast intros, dynamic ad jingles, personalized birthday songs, game soundtracks, social media content.; The key innovation is conversational disambiguation—model gathers requirements, asks clarifying questions, then triggers tool use. This is function calling for media generation, not just API retrieval.; For agent systems: you can build voice agents that not only answer questions but create artifacts (songs, jingles, soundscapes) on demand. This extends the agent action space from information retrieval to content creation.; Production blockers: need to clarify licensing/IP rights (can users commercially use generated music?), test music quality across genres, and design fallback UX if generation fails or takes too long.; Watch for: API access details (will Lyria 3 be available via Gemini API or separate endpoint?), pricing (per generation? per second of audio?), and whether generated music includes watermarking or attribution.; The 'Live Jukebox' pattern (conversational agent + generative tool) is generalizable—imagine 'Live Designer' (generate images), 'Live Animator' (generate video), or 'Live Coder' (generate code) using the same conversational + function-calling architecture.

Developer Experience: AI Studio, Coding Agent Skills, and Example Galleries

  • Claims: Google AI Studio offers free, no-credit-card access to experiment with Gemini models, including the Live API.; AI Studio has a gallery of example apps: EchoScript (transcription), Voice Library (TTS direction), and links to Live API at ai.google.dev/aistudio/live.; DeepMind publishes 'Gemini coding agent skills' to help coding agents (Cursor, GitHub Copilot) scaffold real-time audio integrations correctly.; Example starter code is available in Python (server-to-server) and JavaScript (client-to-server) for the Live API.
  • Evidence: Thor repeatedly directed audience to 'ai.google.dev/aistudio' and 'ai.google.dev/aistudio/live' for free experimentation.; EchoScript and Voice Library are in the AI Studio gallery and can be tried without payment.; Thor noted: 'AI Studio is a really great way to try the models without actually needing to put down the credit card.'; Gemini coding agent skills: 'We have published coding agent skills for all of the Gemini APIs including the live API. Real-time audio working with real-time audio can just be a bit more challenging so using these agent skills and installing them in your coding agents can really help steer them in the right way and give you that result that you're looking for.'; Example code linked from docs for Python (server-to-server WebSocket) and JavaScript (client-to-server WebSocket).
  • Caveats: No details on free tier limits (requests per day, token caps, model access).; Coding agent skills not further described (which agents support them, what they contain, how they differ from docs).; Gallery examples are demos, not production templates—unclear how much boilerplate/error handling is needed for real apps.; No mention of whether AI Studio supports team collaboration, version control, or deployment pipelines.
  • Implications: For Ken: This is a major friction reducer for experimentation. You can prototype voice agents, transcription workflows, and multimodal conversations without upfront payment or onboarding.; The coding agent skills are strategic—DeepMind is explicitly targeting developer tooling (Cursor, Copilot) to make Gemini APIs easier to scaffold. This accelerates adoption by reducing the 'blank page' problem for real-time audio (WebSocket setup, error handling, audio streaming).; For your agent systems: install the Gemini skills in Cursor/Copilot and generate starter code for Live API, then customize. This is faster than reading docs and writing boilerplate from scratch.; The free tier is ideal for proof-of-concept before committing to GCP billing. Use it to validate latency, quality, and system instruction patterns before production.; Watch for: whether the free tier persists (or becomes paid after launch window), whether AI Studio evolves into a full IDE (currently it's a playground), and whether the coding agent skills are open-source (could be forked/customized).

Notable Concepts & Terms

  • EchoScript: Gemini 2 Flash Preview-powered tool in AI Studio gallery that extracts speaker labels, timestamps, language, emotion, and translations from audio in a single API call via structured outputs. Example of native audio understanding beyond transcription.
  • Voice Library: AI Studio gallery tool for Gemini voice synthesis. Uses ~30 base voices directed via natural language prompts (scene, accent, director's note) rather than a catalog of pre-recorded voices. Demonstrates voice-as-prompt-engineering paradigm.
  • Director's Note: Natural language instruction in Gemini voice synthesis prompt structure that guides performance (accent, pacing, emotion), analogous to directing a human voice actor. Example: 'delivered the lines with a strong authentic Irish accent.'
  • Gemini 2.1 Flash Live: Real-time speech-to-speech multimodal model (audio + text + video at 1 fps over WebSocket) with intelligence baked into the audio model (not cascaded STT/LLM/TTS). Available at ai.google.dev/aistudio/live for free experimentation.
  • Lyria 3: Google DeepMind's music generation model with two variants: Lyria 3 Clip (30-second jingles) and Lyria 3 Pro (full-length songs with lyrics). Can be invoked via function calling by Gemini Live for conversational music creation.
  • Live Jukebox: Demo app by Thor Schaeff where Gemini 2.1 Flash Live acts as a radio DJ, takes voice requests, asks clarifying questions, and uses function calling to invoke Lyria 3 to generate custom songs. Example of conversational agent + generative tool orchestration.
  • Gemini Coding Agent Skills: Published prompts/templates for coding agents (Cursor, GitHub Copilot) to scaffold Gemini API integrations correctly, especially for real-time audio (Live API). Designed to reduce boilerplate and steer agents toward correct WebSocket patterns.
  • Structured Outputs / Response Schema: Gemini API feature allowing developers to define JSON schema upfront; model populates it directly (no post-processing). Used in EchoScript to return speaker, language, emotion, timestamp, translation fields in a single response.
  • Native Audio Intelligence: Gemini 2's core capability: audio understanding (transcription + emotion + context + pacing + speaker overlap) is built into the foundation model, not stitched together from separate services. Powers all downstream audio tasks (TTS, Live, Lyria).
  • Cascaded Pipeline: Traditional real-time voice architecture: audio → STT → LLM → TTS, where each step adds latency. Gemini 2.1 Flash Live avoids this by baking intelligence directly into the audio model (speech-to-speech).

Operator Notes / Why Ken Should Care

  • For agent systems: Gemini 2.1 Flash Live is the first major real-time voice model with native multimodal reasoning (audio + video + text) and tool use (function calling for Lyria 3). This enables voice agents that reason over conversation history, visual context, and create artifacts (songs, images, code) on demand—extending the action space from information retrieval to content generation.
  • For AI ops: The unified audio stack (one model family for transcription, TTS, real-time conversation, music) simplifies infrastructure (single API surface, fewer services to orchestrate) but creates single-point-of-failure risk. Test fallback strategies if Gemini has an outage or quality regression.
  • For content/business: The voice-as-prompt-engineering paradigm (30 base voices + natural language direction) unlocks dynamic voice personas for podcasts, audiobooks, ads, and games without managing hundreds of voice files. This is a cost/licensing win but requires investment in prompt templates and QA.
  • For investing: Google is betting that native audio understanding (vs. cascaded pipelines) will be the moat for real-time voice AI. The free experimentation tier (AI Studio) and coding agent skills are strategic plays to drive developer adoption and lock-in before OpenAI/Anthropic ship competitive real-time voice APIs.
  • For GTM: The 'Live Jukebox' demo (conversational agent + generative tool) is a template for new product categories: Live Designer (generate images via voice), Live Animator (generate video), Live Coder (generate code). The conversational disambiguation (model asks clarifying questions) is key to UX—users don't need to know the tool's parameters upfront.
  • For workflow: The structured outputs feature (response schema) is critical for integrating Gemini audio into existing systems (CRMs, databases, UIs). EchoScript's one-shot multilingual transcription + translation + emotion + diarization replaces 4-5 separate API calls, reducing orchestration complexity and latency.

Watch Map

  • 00:00: Intro: Title clarification ('at Google DeepMind'), Thor introduces himself in 5 languages (demo teaser)
  • 01:30: Recent releases: Gemini 3, Gemma 4 (on-device audio understanding), VO 3.1 (video + audio), Gemini 3.1 Flash Live (real-time multimodal)
  • 03:15: EchoScript demo: Multilingual transcription (English, German, French, Mandarin, Japanese) with speaker labels, timestamps, emotion, translation in one API call. Japanese transcription failed.
  • 07:45: Voice Library demo: ~30 base voices + natural language direction. Irish accent ('Finion' voice + 'cozy pub in County Clare') and Singaporean Singlish ('Sapphire' voice + hawker center scene) examples.
  • 11:30: Gemini 2.1 Flash Live demo: Real-time speech-to-speech with video (1 fps), system instructions (Irish accent), language switching (German poem in Irish accent). Available at ai.google.dev/aistudio/live.
  • 13:00: Developer resources: AI Studio (free experimentation), Gemini coding agent skills, example code (Python, JavaScript) for Live API WebSocket integration.
  • 15:20: Live Jukebox demo: Conversational music generation. User requests 'German techno schlager about UK startup scene,' model asks clarifying questions, invokes Lyria 3 via function calling, generates and plays custom track.
  • 18:30: Closing: Slide rewind for links, thank you to audience and audio team.

Source/Metadata

  • Title: From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind
  • Transcript words: 4876
  • Duration seconds: 1173
  • Timestamp note: Timestamps manually reconstructed from transcript structure and demo sequence; not provided in original transcript.

Transcript

2848 words en Processed in 326.6s

What's new in AI audio? I'm sorry it's a little bit misleading because the title leaves out the at Google DeepMind so we're just looking at what we've been working on at DeepMind. If we were to look at everything in AI audio we'd be spending a lot of time here but I'd love to show you what we're working on at DeepMind. This is me hi everyone I'm Thor I work on the developer experience at Google DeepMind working on the Gemini API and Google AI studio. Hello zusammen, herzlich willkommen my name is Thorsten. Bonjour, je m'appelle Thor, je suis très désolé mon français c'est très mauvais. Konnichiwa, orwa, rei jingda. Daxianhao, wos shu zuzai megwater dökoren, buha yi se wode tongwen hai buhao. Okay that was for the demo and now I just need to make sure last time I did this demo I recorded over it and then it was all gone that was very sad but we'll come back to that in a bit. Yeah what have we been up to at DeepMind? There's been a couple releases I actually joined the team in November literally the day before Gemini 3 was released so I joined and they told me tomorrow we're releasing Gemini 3 and I was like yay didn't do anything but it was great. Most recently on the open model side we released Gemma 4 I think literally last week and yeah pretty incredible some cool stuff you can do there multi modality as well baked into Gemma 4 so there's audio understanding in the Gemma 4 models and you can do that on device on edge devices as well so that is some very exciting stuff in terms of GenMedia and audio. You're probably very familiar with our image generation models video generation models obviously VO has audio generation in there as well so this is the progression there most recently with VO 3.1 light on the GenMedia model side and then on the audio models we recently launched Gemini 3.1 flash life which is our full duplex sound to sound real-time conversational model also multimodal so you can ingest real-time text voice vision which we'll look at in a bit. So on audio very very broad topic but the baseline of everything we do are the frontier Gemini models and Gemini 3 is incredibly good at understanding audio and that's not just transcribing it but really understanding all the nuances that are in there so that might be obviously speech but also the context of the speech the emotion your pacing anything that swings within the audio that's not just text. So on the audio understanding our goal is to build models that deeply comprehend, richly transcribe and robustly reason through audio seamlessly handling a large mix of different languages dialects accents and modalities and anywhere and always Gemini is really good at transcribing even people that are talking over each other which is pretty incredible. Seamlessly switching between different languages that were the demo we're looking at now. So Echo Script is Gemini three flash preview to analyze audio recordings and extract information out of it. It is built with google AI studio so you can find it in the gallery in AI studio you can try it out I can give you the slides later as well. So that was what I was trying to demo earlier so different from just a pure transcription model we can extract a lot of information out of the audio within one single request to the model or one single API request if we're using the API. So you can see here you know summary I introduced myself by name so we're actually able to label the section the speaker by name I forgot there was no hecklers in the room otherwise we would have picked that up as well and maybe we can see if we have what time later and we can do that but so we can see here we're extracting time stamps we're labeling the speaker identifying the speaker we're identifying the language and the emotion of it right and happy to introduce myself it's great. Now you can see this was in german normally it would classify my german as angry but here I guess I'm very happy to be with you all so I just told it to label the emotion label the language if it's a language that is not english give me an english translation as well right. In french neutral normally I would say sad you know french is just a bit more of a no I said you know I'm sorry my french is very bad didn't sound sad enough so neutral in this case. Okay this didn't work so my japanese I gotta practice that if anyone reads japanese so it should actually say hello my name is thor unfortunately bit of a miss there. Let's see if my mandarin was any better hello everyone I'm a german living in the united states sorry my chinese yeah that is correct does anyone read chinese in the room. No okay well we'll just trust that that is correct and so we can see here that this was one request to the model where I basically just told it to identify the distinct speakers if you have contacts label it by name label the speakers by name give me the accurate time stamps give me the language if the language is not english give me the translation identify the emotion out of happy sad angry neutral and then also provide a brief summary of the entire audio at the beginning so this was one api call to gemini 3 flash preview and we got all this information out we could I just gave it a response schema so structured outputs and I was able to just populate that into my ui to have the structure. So this audio understanding and the base research in the gemini 3 models that is what powers the speech generation as well as the real-time conversational generation so having that audio understanding is really great in terms of knowing what certain things sound like including different pacing different accents and scenarios like that. So the foundation of all our models is the gemini 3 foundational research and then we're building the dedicated audio models on top of that and so with speech generation it's a bit different you know if you've used other tts providers before you probably have a huge library of different voices that you filter by gender by accent by languages what have you but in gemini you have just I think it's like 30-aust sort of base voices and then what you do is you direct that voice to act in a certain way and again different pacing different accents and scenarios like that so the foundation of all our models is now the gemini 3 foundational research and then we're building dedicated audio models on top of that and so with speech generation it's a bit different if you've used other TTS providers before you probably have a huge library of different voices that you filter by gender by accent by languages what have you but so in Gemini you have just I think it's like 30 or so base voices and then what you do is you direct that voice to act in a certain way and again because we have that audio understanding we can basically modify the voice to act in a certain way to use a certain accent and so we can go from a small set of base voices to a very specific voice that we're looking for for our speech generation. Again there's a little application that you can try out it's in the Google AI Studio gallery as well it's called the voice library and so what we can do is giving the prompt structure that we just saw we're building the audio profile the scene we're setting the scene we're instructing sort this director's note so we're giving guidance for the performance just like how you would direct a human to act out a certain way and then some sample context and the transcript that we want so now what we can do is we would just set we want high-pitch Irish male and so basically I just use Gemini 3 flash here again to then construct our system prompt for the speech generation so we're saying here we're sending sort of our audio profile you know Finion here in the scene cozy crowded pub on the coast of county Clare you know delivered the lines with a strong authentic Irish accent and so now we hope the TPUs don't disappoint me there we go it failed but I didn't you know I prepared it so we can listen to it here. Oh you wouldn't believe the size of the thing until you saw it with your own two eyes I'm telling you it was a grand old mess so it was and we were all laughing fit to burst by the end of the night. So as you can see this was the base voice here is this one what kind of problem could we solve so you know that is a fairly standard American accent but so now by giving it that director's note we can then sort of give oh you wouldn't believe the size of the thing until you saw it with your own two eyes I'm telling you it was a grand old mess so it was and we were all laughing or you know similarly here we have Sapphire so this voice is here ready to build something awesome today again fairly standard American English accent here and now we could say you know give it a Singaporean sort of scene wow you must try this chicken rice la the chili is damn sure confirm plus job you will love it faster cue before the uncle close shop okay anyone spend time in Singapore that you know yeah that's you know that's something you'd hear in the hawker center so again you know that is underpinned by the audio understanding so the model really understands what these different scenarios sound like and then can modify the speech generation to be like that. Yes and then you know finally the native audio sound to sound multimodal real time so we just launched a couple weeks ago Gemini 3.1 Flash Live so it is a speech to speech real-time multimodal model you can ingest text audio video in real time through a WebSocket connection and then you get a real-time audio response back as well as the text transcript of that you know obviously benchmarks are you know especially in the audio space benchmarks you can't really trust them you know it's great you can see the reasoning the thinking so here the thinking and the reasoning and the intelligence is baked directly into the model so that's different from a cascading pipeline where you would actually go through text to then go through an LLM to get the intelligence here the intelligence is baked into the audio model so that's the difference there but obviously in real scenarios you know you can try this out in ai.studio slash live so the great thing with AI Studio is you can try it out you know without paying anything so this allows you to try the models without actually needing to put down the credit card so AI Studio is a really great way to do that and again you know we have the audio understanding baked into the model here so what we can do is we can give it some system instructions you know for example speaking a friendly Irish accent and then also we can ingest our camera here for example and then we can say hey can you see me well hello there I can see you loud and clear so I can what's on your mind today what do you think of my outfit ah look at you with your Gemini shirt it looks grand all together and the backwards hat gives you a fierce laid back vibe so it does you're looking sharp. [SPEAKER_03] Ah wonderful, canst du mir ein gedicht auf deutsch erzählen bitte. Ah a poem in German is it sure I can give that a go for you here's a little one. So obviously you need to adjust your system instructions to not speak in an Irish accent in every language. So it's pretty funny you can switch between the different languages there. Again, ai.studio.com slash live you can try it out. You could also ingest your screen. So you're basically just ingesting video frames in addition to the audio at a maximum frame rate of one frame per second at the moment. That's the maximum you can ingest. There are a couple examples. If you're a developer yourself, feel free to try these out. They are linked from the docs as well. These are just some examples of how you would get started, say in Python for a server to server connection or in JavaScript with. [SPEAKER_03] speak in an Irish accent in every language. So it's pretty funny you can [SPEAKER_03] switch between the different languages there. Again, so ai.studio.com slash life you can try it out. You could also ingest your screen. So you're ingesting video frames in addition to the audio at a maximum frame rate of one frame per second at the moment. That's the maximum you can ingest. There are a couple of examples. If you're a developer yourself, feel free to try these out. They are linked from the docs as well. These are just some examples of how you would get started, say in Python for a server to server connection or in JavaScript with a client to server connection. So you can find those here. Also recommend the Gemini skills. So we have published coding agent skills for all of the Gemini APIs, including the live API. So real-time audio, working with real-time audio can just be a bit more challenging. So using these agent skills and installing them in your coding agents can really help steer them in the right way and give you that result that you're looking for. And lastly, we have music, so also audio. I put that in here. We recently released Lyria 3. So it's a music generation model. But it can now actually generate music with lyrics. There are two separate models. There's a Lyria 3 clip, which is a 30-second jingle generation model. And then Lyria 3 Pro is the full-length song generation model. Bringing everything together, I don't know if you remember the good old days when you would actually call the radio station and request a song. So I've built this application called Life Jukebox. And we can now try it out where we actually give the real-time Gemini Life model a tool to then generate a song using Lyria. And so let's see if we can do that here. Right then, you're live on air with the Jukebox DJ. What an absolute legend. Who am I speaking to? And what kind of vibe are we going to cook up today? Is it a proper rave starter or something a bit more... Can we get maybe a German technoschlager about the UK startup scene? German technoschlager about the UK startup scene? Mate, that is revolutionary. I love it. Right, picture this. Pounding techno beat, a proper Hans Zimmer style. Yeah, that's great. Let's go. Let's do that. Hold on tight. We're almost there. Just to make sure it's a certified banger. Any specific buzzwords or stories from the UK startup world you want in the lyrics? And are we thinking manic energy or something a bit more melodic within that techno madness? Nah, just manic energy and surprise me on the lyrics. Alright, you're on. Get ready to have your eardrums blasted. Manic German technoschlager with a British startup twist. Cooking up a proper banger for you. Check this out. All right. I'll leave you with that. Thank you very much. If you don't speak German, sorry. But thanks so much. Appreciate you all. No worries. Oh yeah. And if you want the slides, I can just rewind. There are all the links in there. If that's helpful, you can just grab them there. Awesome. Thank you. And yeah, enjoy the rest of the conference. And big thank you as well to our friends in the back on the audio. You know, it wouldn't be possible without them. Cheers. Oh yeah. Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. Oh yeah. [SPEAKER_03] Oh yeah. [SPEAKER_03] Oh yeah. Oh yeah. [SPEAKER_03] Oh yeah. different pacing different accents and and scenarios like that so um you know the foundation of kind of all our models is now sort of the gemini 3 um foundational research and then we're building kind of the um dedicated audio models on on top of that and so with speech generation uh it's a bit different you know if you've used kind of other um tts providers before uh you probably have a huge library of you know different voices that you sort of you know you filter by gender by by you know accent by languages what have you um but so in in in gemini you have you know just i think it's like 30-aust sort of base voices and then what you do is you you kind of direct that voice to act in a certain way and again because we have that audio understanding uh we can we can basically modify the voice to you know act in a certain way to you know use a certain accent and so we can go from kind of a small set of of base voices to a very specific you know kind of voice that we're looking for for our speech generation uh again there's a a little application that you can try out uh it's in the uh google ai studio gallery as well uh it's called the voice library uh and so what we can do is you know um kind of giving the the the prompt structure that we just saw you know we're building sort of the the audio profile the scene we're setting the scene we're instructing sort of this director's note so we're giving guidance for the performance you know just like how you would um direct a human you know to to act out a certain way um and then some sample context and kind of the transcript um that we want so now what we can do is you know we would just set sort of uh we want you know high-pitch irish mail uh and so basically i just use gemini um three flash here again to then construct our um system prompt for the speech generation so we're saying here you know um we're sending sort of our audio profile you know finion here in the scene sort of cozy crowded pop uh in the coast of county clare um you know delivered the lines with a strong authentic uh irish accent and so now we hope the the tpus don't uh disappoint me there we go it failed but um i didn't you know i prepared it so we can we can listen to it here oh you wouldn't believe the size of the thing until you saw it with your own two eyes i'm telling you it was a grand old mess so it was and we were all laughing fit to burst by the end of the night so as as you can see you know this was uh the the bass voice here is uh this one what kind of problem could we solve so you know that is a fairly sort of stand standard you know american accent but so now by you know giving it that director's note we can then sort of give oh you wouldn't believe the size of the thing until you saw it with your own toys i'm telling you it was a grand old mess so it was and we were all laughing or you know similarly here we have um sapphire so this voice uh is here ready to build something awesome today again you know kind of fairly standard sort of american um english accent here and now we could say you know give it kind of a singaporean sort of scene wow you must try this chicken rice la the chili is damn sure confirm plus job you will love it faster cue before the uncle close shop okay anyone spend time in singapore that you know yeah that's uh you know that's something you'd hear in the hawker center so um again you know that is kind of underpinned by the audio understanding so um the model really understands what you know these these different scenarios sound like and then can modify uh the speech generation to to to be like that um yes and then you know finally sort of the the native audio uh sound to sound multimodal real time so we uh just launched a couple weeks ago gemini 3.1 flash life um so it is a speech to speech just kind a real-time multimodal model uh you can ingest uh text audio video in real time through a websocket connection and then you get a real-time audio response back as well as kind of the text transcript of that um you know obviously benchmarks are you know especially in the audio space benchmarks you know you can't really trust them um you know it's great you can see sort of the reasoning the thinking uh so so here the the the thinking and the reasoning and the intelligence is baked directly into the model so that's you know different from a cascading pipeline where you would actually go uh through text to then go through an llm to get the intelligence here you know the intelligence is baked into the audio model so that's kind of the difference there um but obviously in in real scenarios um you know you can try this out in uh ai.studio slash life so the great thing with ai studio is uh you can try it out um you know without paying anything so this is uh allows you to you know try kind of um the models without actually needing to put down the credit card so ai studio is a really great way to do that uh and again you know we have the audio understanding kind of baked into the model here so what we can do is we can give it some system instructions you know for example speaking a friendly irish accent and then also we can ingest kind of our you know camera here for example and then um we can say hey can you see me well hello there i can see you loud and clear so i can what's on your mind today what do you think of my outfit ah look at you with your gemini shirt it looks grand all together and the backwards hat gives you a fierce laid back vibe so it does you're looking sharp ah wonderful canst du mir ein gedicht auf deutsch erzählen bitte ah a poem in german is it sure i can give that a go for you here's a little one just to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring to bring as well. So obviously you need to adjust your system instructions to not speak in an Irish accent in every language. So it's pretty funny you can switch between the different languages there. Again, so ai.studio.com slash life you can try it out. You could also ingest your screen. So you're basically just ingesting video frames in addition to the audio at a maximum frame rate of one frame per second at the moment. That's kind of the maximum you can ingest. There is a couple examples. If you're a developer yourself, feel free to try these out. They are linked from the docs as well. These are just some examples of how you would get started, say in Python for like a server to server connection or in JavaScript with kind of a client to server connection. So you can find those here. Also recommend the Gemini skills. So we have published coding agent skills for kind of all of the Gemini APIs, including the live API. So, you know, real-time audio, working with real-time audio can just be a bit more challenging. So using these agent skills and kind of installing them in your, you know, coding agents can really help steer them sort of in the right way and give you sort of that result that you're looking for. And lastly, we have, okay, we have a bit of time. So music, you know, also audio. So I put that in here. We recently released Lyria 3. So, yeah, it's a music generation model. But so it now actually can generate music with lyrics. There's two separate models. There's a Lyria 3 clip, which is a 30-second kind of jingle generation model. And then Lyria 3 Pro is the full-length song generation model. And so, you know, kind of bringing everything together, I don't know if you remember the good old days when you would actually call the radio station and wish for a song. And so I've kind of built this application called Life Jukebox. And so we can now try it out where we actually give the real-time Gemini Life model a tool to then generate a song using Lyria. And so let's see if we can do that here. Right then, you're live on air with the Jukebox DJ. What an absolute legend. Who am I speaking to? And what kind of vibe are we going to cook up today? Is it a proper rave starter or something a bit more... Can we get maybe a German technoschlager about the UK startup scene? German technoschlager about the UK startup scene? Mate, that is revolutionary. I love it. Right, picture this. Pounding techno beat a proper Hans Zimmer style. Yeah, that's great. Let's go. Let's do that. Hold on tight. We're almost there. Just to make sure it's a certified banger. Any specific buzzwords or stories from the UK startup world you want in the lyrics? And are we thinking manic energy or something a bit more melodic within that techno madness? Nah, just manic energy and surprise me on the lyrics. Alright, you're on. Get ready to have your eardrums blasted. Manic German technoschlager with a British startup twist. Cooking up a proper banger for you. Check this out. All right. I'll leave you with that. Thank you very much. If you don't speak German, sorry. But thanks so much. Appreciate you all. No worries. Oh yeah. And if you want the slides, I can just rewind. There's like all the links in there. If that's helpful, you can just grab them there. Awesome. Thank you. And yeah, enjoy the rest of the conference. And big thank you as well to our friends in the back on the audio. You know, it wouldn't be possible without them. Cheers. Oh yeah. Oh yeah. Oh yeah. Oh yeah. Oh yeah. Oh yeah. Oh yeah. Oh yeah. Oh yeah. Oh yeah. Oh yeah. Oh yeah.