TLMs: Tiny LLMs and Agents on Edge Devices with LiteRT-LM — Cormac Brick, Google
Description
Tiny LLMs are making on-device agents much more practical. In this workshop, Cormac Brick walks through how LiteRT-LM brings language models to edge devices, with a focus on Gemma, agent skills, and the real engineering tradeoffs behind running LLM workflows on phones and other constrained hardware. The session covers performance across edge devices, on-device function calling, fine-tuning and deployment, platform support across Android and iOS, and the memory, safety, and UX constraints that shape edge-native AI systems. If you're building local agents or want a practical look at where edge LLMs are headed, this is a useful hands-on overview. Speaker info: - https://www.linkedin.com/in/cbrick/ Timestamps (0:00:00) Intro: AI on the Edge, Small Language Models, and Gemma (0:04:51) Enabling App Development: MediaPipe, LiteRT, and System Services (0:09:09) Small Language Models: Performance, Reach, and Fine-tuning (0:11:30) Gemma 4: Sizes (E2B and E4B) and AI Core Roadmap (0:16:10) Gemma on Edge Runtime: Performance Benchmarks (0:18:34) Agent Skills: Google AI Gallery, Mood Tracker, and Wikipedia Lookup (0:23:38) Skill Architecture: Efficiency, Progressive Disclosure, and Tool Loading (0:27:34) Reliability: Constrained Decoding and Tool Usage (0:29:18) Community and Custom Skills (0:31:30) Skill Development Deep Dive: Orchestrator and Registry (0:33:30) Rapid Skill Prototyping: Using Gemini CLI and ADB (0:38:35) Open Source: AI Edge Gallery and Community Engagement (0:41:00) Deploying Tiny Models (sub-1B parameters) In-App (0:47:44) Third-Party Models: Fast VLM and Hardware Acceleration (0:50:17) Model Examples: Function Gemma, Mobile Actions, and Embedding Gemma (0:55:41) AI Edge Eloquent: Transcription and Text Polishing (0:59:07) Modularity Playbook: ASR and Text Polishing Engines (1:01:23) Synthetic Data Workflows for Tiny Models (1:06:36) Web Support and Fine-tuning Documentation (1:08:20) Summary and Key Takeaways (1:12:49) Q&A: Multi-skill Execution, Context
Summary
Generated by claude-haiku-4-5-20251001TLMs: Tiny LLMs and Agents on Edge Devices with LiteRT-LM
Main Topics
- Edge AI Evolution: Google's decade-long journey bringing AI models to edge devices, from running GoogleNet on a USB accelerator (2016) to modern mobile deployment
- Two Deployment Patterns: System-level Gen AI (2-5B models built into OS) vs. In-app Gen AI (tiny models <500M parameters)
- Agent Skills on Device: New paradigm enabling function calling and tool use on mobile/edge platforms
- Tiny LLM Fine-tuning: Workflow for deploying specialized models across diverse hardware platforms
- Real Production Apps: Practical implementation examples including the Eloquent transcription app
Key Points
Edge AI Infrastructure
- LiteRT Stack: Unified framework supporting CPU, GPU, NPU across Android, iOS, macOS, Linux, Windows, Web, and IoT devices
- Gemma 4 Models: New small (E2B: 2B parameters, E4B: 4B parameters) models optimized for edge with built-in function calling and thinking capabilities
- Performance Metrics:
- E2B on Android GPU: thousands of tokens/second
- Raspberry Pi: ~133 tokens/second (suitable for image analysis)
- Qualcomm IoT platform with NPU: compelling performance gains
Two Deployment Strategies
System Gen AI (2-5B models)
- Built into OS (AI Core on Android, Apple Intelligence)
- Customized via prompting or skills
- Foundation models preloaded on device
- Wide device compatibility
In-App Gen AI (Tiny Models <500M)
- Custom-trained for specific tasks
- Loaded with app/webpage
- Wider reach across device types
- Requires fine-tuning for production reliability
Agent Skills Architecture
- Progressive Disclosure: One-line skill descriptions loaded in context; full details loaded only when model decides to use them
- Token Efficiency: Critical for lightweight models with limited reasoning capacity
- Dynamic Tool Access: Skills can integrate new tools, extend I/O patterns, and access domain-specific knowledge
- Constraint Decoding: Stronger constraints applied specifically to tool calls improve reliability, especially for small models
- JavaScript-Based Skills: Low-code skill development enabling 80+ skills created by teams
Tiny Model Workflow
`
Transformers → LiteRT-Torch → LiteRT-LM File → Cross-platform Deployment
(with quantization & optimization)
`
Fine-tuning Impact: 20-40 point improvement in eval metrics (e.g., 40% → 86% on function calling with 270M parameter Function Gemma)
Real Application Example: Eloquent
Features:
- Offline transcription with speech cleaning
- Removes filler words and interjections (um, ah, etc.)
- Biasing dictionary for technical terms and proper names
- Two-model approach: ASR engine + text polishing LLM
- Built on fine-tuned Gemma 3.2 derivatives using synthetic data
Workflow:
- Generate synthetic training data using larger cloud model
- Fine-tune tiny Gemma model (Gemma 3.2 7M lineage)
- Combine with quantization for efficient deployment
- Ship compelling feature on wide device range
Notable Quotes
> "Tiny LLMs are where it's at for deploying things widely"
> "If you want to do your own narrow task and ship something to lots and lots of devices, fine-tuning is the workflow of choice"
> "One file can work in CPU and GPU across a wide set of devices" (except NPU, which requires special compilation)
> "For really tiny models on device, this is a very helpful tool that helps us have stronger guardrails for the model"
> "There's no hidden secret sauce in what we've shown"
> "The safety risk profile is narrower with tiny models because they have narrower functional scope"
Takeaways
Model Size Guidelines
- 2-5B Models: Best for system-level Gen AI on mobile OS
- 500M-1B Models: Viable for general-purpose in-app tasks with fine-tuning
- <500M Models: Require task-specific fine-tuning for production reliability
Deployment Best Practices
- Use Skills for Extensibility: Progressive disclosure pattern reduces context window requirements while enabling dynamic capabilities
- Fine-tune Strategically:
- For medium models (2-4B): Use skills and prompting
- For tiny models (<1B): Fine-tune for specific tasks
- For embedded systems: Consider LoRA adapters (8-100MB each) for hot-swapping
- Leverage Synthetic Data: Generate training data from larger cloud models to enable effective small model training
- Consider Hardware:
- Mobile phones: System Gen AI trend emerging
- Laptops: 32GB+ RAM enables larger models
- IoT/Robotics: Multi-LoRA approach optimal
Future Directions
- Expansion of tiny Gemma models with fine-tuning workflows
- macOS version of AI Edge Gallery (coming soon)
- Broader NPU support across MediaTek, Intel, Qualcomm platforms
- Community skill ecosystem development
Open Source Resources
- AI Edge Gallery: iOS/Android app for prototyping (open source)
- LiteRT-LM: Python, C++, Java, Swift APIs
- LiteRT-Torch: Fine-tuning and optimization toolkit
- GitHub: Collab notebooks for fine-tuning workflows, community skill repository
Transcript
Hey, my name is Cormac Brick. I work on Google AI Edge, which is a way of bringing models to the edge. This is something we use internally for our own products. It's also something we make available as open source products. As part of this, we also work really close with the Gemma team because they publish a lot of models that are targeted to edge devices as well. I've been focusing on edge AI for the last 10 years at this point. Started off in 2016 with running a Google Net on a USB key hardware accelerator plugged into a Raspberry Pi at NeurIPS 2016. Then joined Intel, worked on and led the architecture for the NPU that goes into all of their laptops these days. Three years ago, I moved to Google to work as tech lead on edge AI here. It's been great fun. It's been a crazy few years. I'm happy to share where we're up to and also give an overview of where is mobile AI up to today. What's the state of the art? What are the different patterns we see for model deployment? The two things I want to focus on in this talk are tiny LLMs and agent skills. Tiny LLMs are very small models. Agent skills are now possible on device. Larger models are needed to make those work fluently on device. Both are exciting new directions that have only recently been made possible. As recently as last week when we launched Gemma 4 with the DeepMind team, we supported them with an Android and iOS app. That opened new possibilities for things we can do on mobile. We'll deep dive into that as well as looking at what types of things we can do with tiny models. [SPEAKER_03] OK. [SPEAKER_03] So that's the intro. So first, give a review of what AI and the edge are, looking at what we look at as small language models and tiny language models. Then taking a look at Gemma models, just because that's pretty topical, and the types of performance we see for Gemma models on various types of edge devices, because that's what our team does. Then looking at agent skills, which we built on top of the latest generation of Gemma models that can run on both Android and iOS, as well as many other platforms. In the second half, we're going to look at tiny models and how you can fine tune and deploy a tiny model to edge devices. Then finally, I have an example of a real app that our team built, not just an example app, using tiny LLMs based on Gemma technology as well. There's a lot of benefits to running on the edge. There's latency or UX improvements for some really sensitive in-the-loop things like live voice translation, for example. That's something our team shipped on Pixel last year, where you can do a live voice translation. It's very challenging to do that with the required latency using a cloud service. Being able to do that offline and on the edge is key to meet latency or user experience requirements. Privacy is definitely a thing as well. We do some work with AI within messaging apps, for example, and people like their messages to stay on their phone, encrypted, fully private. That's a good example of where we're seeing LLMs getting deployed to assist in those types of use cases where privacy is really important. [SPEAKER_07] Ooh. Offline use is obvious. And then savings. We see this being increasingly relevant for laptop users, where you see a trend towards folks experimenting with small language models to do some types of tasks they may do with a desktop agentic workflow. Some folks are interested in exploring savings as well. This is what our team does. We enable people to build apps. We do MediaPipe, LiteRT LLM, which is an LLM runtime that works on mobile and edge. And then LiteRT, which is a standard inference framework. This was known a long time ago as TensorFlow Lite. We've used that over the years in things like Photos on Blur. MediaPipe is used in Photos, YouTube Shorts. If you've used any of the funny effects in YouTube Shorts, they're things that our team has built along with the YouTube team using MediaPipe and various deep learning tracking models under the hood. [SPEAKER_00] That's also built on the same MediaPipe and LiteRT technology. [SPEAKER_00] We run this on all Android phones today. We use this as part of system services or third-party apps building on a version of the stack that ships on Android. Our stack also runs far beyond Android. Even with it being available as a system service in Android, it's also available to take the same TFLite file and ship that to iOS, Mac OS, Linux, Windows, Web, or IoT devices from that same file. That's why deployment is helpful. One caveat: that same file deploys on CPU and GPU. For NPU, we need to do some special compilation and you end up with a special NPU file. But certainly for the Gemma models we published last week, But our stack also runs far beyond Android. And so we also run it being available as a system service in Android. It's also, you can take the same TFLite file and ship that to iOS, Mac OS, Linux, Windows, Web, or IoT devices just from that same file. And so, yeah, that's why deployment is helpful. One caveat there, that same file deploys on CPU and GPU. For NPU, we need to do some special compilation and you end up with a special NPU file. But certainly for the GEMMA models we published last week, this is true that one file can work in CPU and GPU across a wide set of devices. So then, some of the things we're seeing for LLMs on device is privacy-centric stuff, voice agents and local agents with tool calling as a very popular workflow. And then, within our stack, we, Lider TLM is the thing that we'll look at a bit later that helps us run tiny LLMs on device, cross-platform, and it's fast because we support all of the different hardware accelerators. So this is an important concept—we see two trends happening today. One is system-level Gen AI. So for very large models, the way these tend to turn up on mobile phones isn't that when you launch a single app, it will download a 4 billion parameter model just to help you find a good restaurant in whatever app you're looking in, right? Instead, the trend is to build larger models into the OS. So we call this system-level Gen AI. These models tend to be in the 2 to 5 billion parameter range, right? It depends on the OS and whatever. But this is one choice that we see both the Android team, who we work closely with, and the Apple Intelligence team. We see similar choices being made by both mobile OS vendors, where there's a central model built in. And that is, on Android, called AI Core, and there's things like summarization APIs and an increasing set of APIs available for developers, including a prompt API, where you can use that API. And that's available on premium Android devices and premium Apple devices as well. Obviously, Apple has their own Apple Intelligence thing. But as a trend, this is worth noting. This is really relevant, right? So if you want to leverage a built-in model, this is a great way to go. So then the other trend we see is in-app Gen AI, which is where the tiny LLMs, or TLMs as we're calling them in this presentation, are more relevant, right? So system Gen AI, generally, you customize that via prompting or via skills, as we're going to see later. And it's a foundation model preloaded on device. [SPEAKER_00] In-app Gen AI, generally, these are custom to tasks. They're loaded with the app or with the web page. We also deploy some of these on the web. And generally, they're targeted at wider reach. So this may work not just on premium devices, but all devices, because apps' reach is really important for a lot of the application developer teams that we work with. And these, surprisingly, you can get LLMs to perform well—if you fine-tune for a single task, we've seen really strong performance on simple tasks like summarization, transcription, or voice-to-action type things. We can get really reliable performance from models in the 100 to 500 million parameter range, depending on the complexity of the task, right? So a good example here is we launched Function Gemma in December. And that was a 270 million parameter model that was dedicated for function calling. And we're able to show that doing voice-to-function calling for 10 different functions that were relevant to Android developers. And our internal evals with an internal eval set reached over 85 to 90 percent reliability, just using that very small model, which was widely deployable to iOS and Android devices. And that's actually something you can play with in a sample app that we have that we'll look at a bit later. So that's an example of a tiny LLM model. But we're seeing more interest from application developer teams now into fine-tuning models to deploy as in-app Gen AI. So the difference here is on the left, you customize via prompting or skills, customization on the right, we would encourage people to do some degree of fine-tuning and then make a tiny LLM work. In practice, certainly below 500 million parameters, that's true. Maybe for 500 million parameters and above, we can do more general purpose tasks with a model, but for the really tiny models, certainly less than 500, in our experience, you need to fine-tune to get production-level reliability. So now I'm going to talk about Gemma 4. And Gemma 4, the models that were launched last week, fall into that system Gen AI candidate category, right? So we can do lots of powerful things with Gemma 4, as you'll see. First, we're going to talk about sizes. So there's two small sizes, which are E2B and E4B. E2B is called that because it has, it only needs about 2 billion parameters to be present in RAM to run the model. Because one of the limiting factors we see with these models is how much RAM you need to have in a device in order to run it reliably. E4B has been optimized to run it with 4 billion parameters on device. There are more parameters used by the model, but the other parameters that the model uses are, DeepMind have talked on this later in the week, so you can go there for more detail. But the TLDR is the other parameters in the model are used for per layer embeddings. So these have, E2B is called that because it has, it only needs about 2 billion parameters to be present in RAM to run the model. Because one of the limiting factors we see with these models is how much RAM you need to have in a device in order to run it reliably. E4B has been optimized to run it with 4 billion parameters on device. There are more parameters used by the model, but the other parameters that the model uses are DeepMind have talked on this later in the week, so you can go there for more detail. But the TLDR is, the other parameters in the model are used for per layer embeddings. So in our runtime, we don't actually need to load all of those parameters. So we need to maintain them, the 2B and 4B, they need to stay resident in RAM. The other ones, we typically memory map them. And we only need to load one line of the embedding table of the per layer embedding table and the embedding table once in the auto-aggressive loop. And we actually only need to load a few hundred bytes or thousands of bytes of that in order to do the next token inference. So as you go through inference, you don't ever end up requiring to load the whole PLE table into memory. And then, depending on the OS, it'll do a reasonably good job of evicting memory used by older PLE tokens. So that's why we have this idea of effective. So the smaller models run on, you can see on the right, it runs on a variety of platforms. The E2B and E4B models. These models are on the AI core roadmap. So these models that are available now for experimentation, at the appropriate point in time in the future, the Android team will integrate this into AI core, and they'll be available more broadly on a wider side of devices. And there's another talk next week, or there's a talk here this week from Ali from the AI core team, who will probably share more details about exactly what that roadmap looks like. And also Omar is going to do it from the GDM team, is going to do a deeper dive on all of the Gemma world. The bottom two models, just for reference, these are also relevant for the edge. They're not really the focus of my talk today, because these are relevant for folks to run on laptops. The sizes have been optimized to run really well on consumer-grade laptops, albeit ones that have maybe 32 gigs of RAM. But yes, that's what these do. So I'm going to focus less on these other two models, even though they're very powerful and useful. So we can see, certainly the E4B model, both the E2B and E4B model have excellent performance on a wide range of things. Yes, on knowledge and reasoning. But one of the big step ups relative to the last generation, from my perspective as a user of the models, is they've built in function calling, which is excellent, and they also have built in thinking. So that combination of thinking plus function calling is what unlocks our ability to now do skills on device. So you can just, as you'll see in a second, we can describe a skill, give it to the model, and the model can just pick it up and use it. So that allows us to use that pattern that's proven very popular in the last few months to bring that pattern to mobile, which allows for new types of mobile experiences. Also for the E2B and E4B, these are multimodal. So they support audio, image, and text. The larger models support just image and text. Also the other change from a deployment point of view is this is the first time the Gemma models are released as just a stock standard Apache 2.0 license, which means they're more usable by more people. Again, you'll hear more about this in Omar's talk, just that I'd mention it here. So that's Gemma in general. Now, to deep dive Gemma on our runtime and our platform. So yes, that same picture earlier, right? We have a single LIDAR TLM file, which is a LIDAR TL file with the things like the tokenizer and the other things we need in one package to run a model. That single model runs across all of these classes of devices across mobile, desktop, and embedded. And we're pretty excited to do more in the embedded space. There's a lot, particularly with image input, there's a lot of scope for new IoT use cases. Okay, bit of an eye chart now, but just to dig into performance. So I would say this is a snapshot as of today. This is something we are continuing to work on, both ourselves and our team, and also with various partners that we have, right across like Intel, with the Raspberry Pi team and the Qualcomm team. We're continuing to optimize all of these numbers. But we can see we can do, certainly on, for the 2 billion parameter model, we can get really compelling performance on a high-end Android phone can do thousands of tokens per second on the GPU. Also thousands of tokens per second on MacBook, but the bottom two rows then are on a Raspberry Pi. We can get about 133 tokens per second, which is sufficient to do simple image analysis use cases with reasonable latency. And also the bottom one is us running on a hardware accelerator. This is a Qualcomm IoT slash robotics development platform that they have available. And there we can see pretty compelling performance as well, because we've gone and used the NPU, which gives a lot better performance. And then, you can imagine, we have corresponding performance on the E4B model, which works on a wide set of devices as well with proportionally less pre-fill and decode performance, right, given the size and the numbers of parameters we need to fetch. But broadly they're available on lots of platforms. [SPEAKER_03] Okay, so now what can we do with those models? Lots of things, right? So one of the things I wanted to talk more about, because it's net new, rather than just show you our image analysis or audio transcription or audio translation, right? There's a lot of things we've been able to do with models from my perspective for a while that these, that the Gemma 4 models are, much better than their predecessors at. But the thing that's newest from my perspective is agent skills on device. So this is, we have an app. I don't know if you guys have seen it. It's available on both iOS and Android. I call it Google AI Edge Gallery. Lots of things, right? So one of the things I wanted to talk more about, because it's net new, rather than just show you our image analysis or audio transcription or audio translation, right? There's a lot of things we've been able to do with models from my perspective for a while that these Gemma 4 models are much better than their predecessors at. But the thing that's newest from my perspective is agent skills on device. So we have an app. I don't know if you guys have seen it. It's available on both iOS and Android. I call it Google AI Edge Gallery. And this allows you to do lots of different things, like basic AI chat stuff or ask questions of an image or do transcription or translation use cases starting from audio or audio to function calling type use cases. But the one second from the left is agent skills. I'm going to do a deeper dive on today. Okay, hopefully this will play. Oh, okay. That is very annoying. Sorry. I will fix this before we post slides. [SPEAKER_07] Let's see if this will. [SPEAKER_03] Okay. [SPEAKER_03] Oh, okay. Who knows it. Will sound work? Morning Gemma. [SPEAKER_06] Let's log a new mood journal entry when it's a school night. [SPEAKER_06] Okay. So this is playing with sound not working. [SPEAKER_06] I got eight hours of sleep. [SPEAKER_06] And I'm looking forward to hanging out with Amy today. So this is a journal skill app we're looking at here. Or a mood tracker where you can see the chat. Where it logs mood and sleep and then. [SPEAKER_06] Perfect. [SPEAKER_06] Analyze the trend in my mood over the last seven days. Okay. So what's happening here is with the mood tracker app, it logs your moods or observations to a diary. Awesome. And then the LLM is able to go back in and summarize the contents. I'm busy making breakfast. Can you check my calendar for the day and show me a bullet list. What's interesting here is this is just, okay. This is just, I want to repeat this more. Okay. So what's interesting here though is the way this is done, which we're going to look in more detail. This isn't a custom fine tune. It is just us giving a particular skill with a little bit of JavaScript that the model can call to the model. And then through just a free text interface, the model is able to use and pick up that skill. So it's also able to call a, there's two skills actually happening here. One is the mood tracker skill and the other one is the map skill. And now it's using a query Wikipedia skill here to look for latest information from the Oscars. So versus having a model that's just pinned in time, it now becomes really easy to extend the model with skills such as map lookup, something like interactive journal that you can both add, subtract entries to from voice as well as query. And also adding more modern knowledge or relevant knowledge. So we're showing Wikipedia in this case. Oh, this is another fun skill that somebody in the team also developed, was the mood music skill, which calls a web service to compose music based on a single image and plays it. So yeah, it's able to compose some lo-fi music to go with a person's breakfast use case. But it's a new paradigm in how we're able to extend the models. And to do that in a pretty low code way, as you'll also see in a bit when we dig into how this works under the hood. Okay, so examples here and I'm not going to play all of these videos. So I'm going to try to do some of these videos, maybe to save time, because I think you've seen many of them. So one is we can augment the knowledge base. There's one pattern that we find interesting. We can produce rich interactive content like flashcards for visualizations. And yeah. Where we can, instead of if you ask to summarize something in three bullet points, you can have a JavaScript skill to show those bullet points as a card instead. And so you can just say, summarize and show a card, summarize this. And if it thinks that a card display will be more helpful for the user, that will get used. Or music sentences. Okay. So I'm going to skip forward a little bit from the demos to how we build them. Okay. So what's actually happening is the way we've built the skills is they're efficient. So the instructions can be loaded on demand. So this uses a principle that you may have seen elsewhere of progressive disclosure of conditional depth. So instead of in an MCP workflow where you need to describe everything about all of the functions that you need, the way we've structured the skills is there's a one line description. The one line descriptions of the skills is what the agent sees. And then if it thinks that sounds interesting, then it asks for more, it asks to load the skill. It is we teach it a skill to load a skill. And then it goes in and loads the skill and it finds out all of the details about how to use it, what function calls it can use to do that skill. So that pattern is particularly important for token efficiency and frankly, reliability on edge models. Because if we had to load all of the details for all of the skills into the edge model, that would be a lot of context for the model to reason over. And in a lighter weight model, that will hurt performance ultimately, because the lighter weight models are really great to be able to run on device. But in terms of reasoning over very long context windows, if you can, what function calls it can use to do that skill. So this is that pattern is particularly important for token efficiency and frankly, reliability on edge models. Because if we had to load all of the details for all of the skills into the edge model, that would be a lot of context for the model to reason over. And in a lighter weight model, that will hurt performance ultimately, right, because the lighter weight models are really great to be able to run on device. But in terms of reasoning over very, very long context windows, if you can have a more condensed context window, that will up your batting average in terms of quality metrics you're looking at to ship a particular app. The second part is tool access. It helps us integrate new tools dynamically. So things like there's a set of input tools, which is how do you get more information, which could be like Wikipedia or looking up a weather service or something like this. There's a set of things that are helpful to present new outputs to the user, like showing something on maps or showing something via cards. So you can have skills both to extend the input and extend the output patterns that a model can do. And also helps us to bring in domain specific knowledge bases. So that skill that's calling Wikipedia, you could easily imagine that calling some customer CRM internally or asking data from a local rag system as well. So our, within our structure of skills, we have skill.md and then there's optionally scripts or assets where the skill.md is the metadata that we always process. The example here is showing extracting text and tables for PDF and we trigger on PDF for extraction. And then the instructions are only loaded when the model thinks that it requires that skill. And this pattern is particularly important. Within both iOS and Android systems. So dig a bit deeper. We have within our system prompt our own system prompts that we ship with the app. Then there's a set of skill descriptions which get added to the system prompt. And then when the user asks for something, the model decides to trigger the skill by reading its metadata. It then calls the load skill app that we have under the hood. And the tool response from that function call is then the contents of the skill.md file, which is then in the context window. So it now knows about these functions. And then it calls the run. So then that also contains some JavaScript. And then we call the run JavaScript tool, which runs the JavaScript that we picked up from the skill file. And then we call a response on that as well. One other thing to mention as part of this workflow is one of the things we did when we were optimizing Gemma for reliability is within the runtime, we have constraint decoding that applies. But it's tuned to only apply to the output when we're generating a tool call. And we can also constrain it just to the particular tool that you're supposed to be calling. So in this system, instead of just having generic JSON constraints, we know that there's a finite set of tools that the model is supposed to be able to use. So we can therefore have stronger constraint decoding. This also helps us have a more reliable system overall by using stronger constraint decoding. We find that's helpful for the 2 billion parameter model. As models get more capable, we find the margin you get from this strict constraint decoding is less essential when you're running a very large model, like if you were running a 10 billion or a very large model. But for really small models on device, this is a very helpful tool that helps us have stronger guardrails for the model so that we can up the quality so we have something that is useful in production. So then we support, within the app, we support you can toggle the skills you want to use. So even the skill descriptions, you can decide how many skill descriptions you want to have live at any given point in time. So this is loading a piano playing skill, virtual piano, which will do its things. You can actually tap the keys and play sounds. This is more JavaScript stuff. But then we can also load custom skills. So you can write a skill yourself and load it from a URL, which is fun if you want to experiment with prototyping a skill with Gemma for an app idea you have. The skills can have an API key as well, a secret key. So if you need to use a web service, that's something that you can prompt the user to put in. This is showing the mood music example. We have on our GitHub page a GitHub discussion where users are posting skills that they have written themselves. And the skills in the community that we like, we are able to pull up and have as featured skills in the app. So ones that people develop, we can have a way of showcasing useful community skills to the wider community as well. So a secret key. So if you need to use a web service, that's something that you can prompt the user to put in. This is showing the mood music example again. But we have a GitHub discussion on our GitHub page where users are posting skills that they have written themselves. And the skills in the community that we like, we are able to pull up and have as featured skills in the app. So if something is developed, ones that people develop, we can have a way of showcasing useful community skills to the wider community as well. So this is at third party scale. I think this is going to show us. Okay, this is adding an animal intro scale. This is a kid's thing. [SPEAKER_00] It's a lot. To the app. Yeah. But the key point here is there's a really low barrier to extend the model, right? And to extend the model in a way that is relevant to downstream app users. So this is very easy. And we're going to go deeper now and see just how easy that is. [SPEAKER_04] If the clicker is, it's a good thing. Okay, so skill architecture, when we go one step deeper, this is relevant. So we have our own orchestrator. At the end of the day, it has a skill registry and we call the load skill skill in order to load the skills. And within skills, there's a JavaScript skill. There's native intents. So at least within the Android system, we're able to call Android system intents or native intents. And you can just call those in your skill. So things like if you wanted to turn on and off the Wi-Fi or something. So intents that are exposed to all Android that are available in JavaScript, you can certainly use. Then there's the roleplay, the skill.md, which is the persona and scenario data. [SPEAKER_00] And then there's the specific skill and resources. [SPEAKER_00] And like I was saying before, this can include JavaScript that runs entirely locally, which is a fully offline experience. [SPEAKER_00] Or in the music composer one, we actually called a web API and that needed an API key that the user was prompted for. And that also works within a gallery. So then, yeah. [SPEAKER_00] Then the predefined tools we have under the hood, just to give a mental model of how it works, we kind of load skill, which helps us load skills, run JavaScript, or run intent. [SPEAKER_00] And just the orchestrator, using these three skills is able to make the overall system work. [SPEAKER_00] So this is another video. [SPEAKER_05] Now let's test our effective tool. [SPEAKER_00] This is from our Gemini product manager showing Olivier. [SPEAKER_05] Hey Gemini, can you find a French restaurant in San Francisco? [SPEAKER_05] Please reply to me in English. This is a restaurant roulette skill. So yeah. [SPEAKER_00] We actually have about 80 of these skills. [SPEAKER_00] It was so easy to develop skills using both anti-gravity or Gemini CLI, we actually had the team develop 80 of these skills. [SPEAKER_00] So we had lots of things to do with lots of options in terms of what to show. And it was a lot of fun for folks to develop these. [SPEAKER_00] So here we can go in and just look at the structure itself of restaurant roulette. [SPEAKER_00] Yeah. [SPEAKER_00] This is the under the hood part. So this requires secret true. [SPEAKER_00] Okay. [SPEAKER_00] We wanted to get an API key and it searches for 10 restaurants. And then returns location and cuisine. Okay. And this is the index.js file, which is where we can have a simple web view then rendered to have the roulette wheel. And you can obviously go and code all of this yourself, right? And there's full source code here for the examples. Cloud code, yeah, I never got that actually published. But it works really well in cloud code as well. And we have both source code for the example and a skill spec for getting started. But there's full instructions on GitHub if you want to try writing your own skill. But this was the pattern we actually used most. It was using skills to write skills. So using something like Gemini CLI or cloud code or anti-gravity. This was our favorite pattern where we just say, hey, I want to write a skill for AI Edge Gallery. This example works in Gemini CLI and we just say, hey, this is the documentation. Here are some examples of some skills that you can go and read. Then this is the skill I want to build, which was an offline archiver. Oh yeah, the idea is when you based on things, pictures you took around London, for example, it would be able to give you information like, there's that cool Churchill statue I kind of walked past. It could then go and research a bit about Churchill for you. And then on your flight back, you could go back into this skill and it would have a bunch of information about the stuff that you saw that you could read on the flight home, for example. I think that was the inspiration for this one. Yeah, so it fetches Wikipedia content and then stores it locally and then has an index. Also here with CLI, we can also ask it, because we have an ADB skill in Gemini CLI, you can also ask it to test the skill itself by saying, hey, you have access to a phone connected via ADB. And then Gemini CLI uses its Android ADB skill to go in and test that the app actually works and that the skill is doing what it says in the app. And we can just add it, ask it then, you can put this in a prompt and ask it to iterate and it'll do some basic validation itself and return it. Yeah, so I think it's going to be, if this plays, this will be an example of doing that. I could also have done this. Yeah, so this is us using Gemini CLI just to do that. And this works really robustly. I think of the maybe 80 skills our team did and certainly more than half were vibe coded skills, at least initially, which really reduces the barrier for folks to extend an LNM to do new things in a way that's useful for their audience. Yeah, yeah, go for it. Yeah, so this, I think it's going to be, if this plays, this will be an example of doing that. I could also have done this. Yeah, so this is us using Gemini CLI just to do that. And this works really robustly. I think of the maybe 80 skills our team did and certainly maybe more than half were vibe coded skills, at least initially, which then really reduces the barrier for folks to extend an LNM to do new things in a way that's useful for their audience. Yeah, yeah, go for it. [SPEAKER_03] Are all examples here running on the smallest model in the app or is it on device? The, we, all of the examples here are defaulting to the, to the good question. So the question for the recording is, are all of the examples running on the smallest model? In this case, the examples are running on the 4b model, right? The 2b model, the skills will work with the 2b model as well. And you can try that. But your mileage may vary, simpler skills, fewer skills, that sort of thing. But all of the examples you've seen are running on the 4b model. [SPEAKER_04] Are you trying to do the Chrome CDP? Try and doing a version of skills in Chrome. [SPEAKER_04] Yeah, so using the Chrome CDP to, with the skill, to actually use the local model to work in your browser. But we haven't. No. Yeah. It's an interesting thing to try. [SPEAKER_03] Okay. And this then is a quick shout out to this link. So a few things about Gallery that I should also mention are, quick check in time. A few things about Gallery that I should also mention are, one is Gallery is an open source project. And the open source project itself builds on top of the Lydor TLM tooling that you saw earlier as part of the intro. So the app itself, as well as being something that's fun to use so you can prototype app ideas or to see, wow, is this possible in the model on a phone? Or how fast would this be if I ran it on a phone? Or is this skill actually viable to do? As well as being able to do model prototyping, you also have full access to the source code here. And it also builds on top of the same infrastructure, the open source Lydor T and Lydor TLM acceleration infrastructure under the hood. So if you see something you like in Gallery, you can deploy the Gemma model some other way as well. But if you want to get the same experience or the same speed, you can then just take the underlying open source APIs and pull a model from our Hugging Face page that we'll look at more in a little bit and just run that directly as well. So as part of that, we also have a set of discussions on the Gallery that shows some of the community skills that have been uploaded from things from Cat Entertainment to Blackjack, right? So these skills, you could also, you can, people have posted these skills. So if you develop a skill, feel free to post it here. And it'll make it more discoverable by other people in the community. And then the ones here that seem compelling or maybe useful to a lot of people, we can add to our third party preferred skills list in the app so that it'll make it easier for folks to discover. Yeah. So this is something we saw a reasonable amount of action over the weekend on. And it's been something that has only been up since last Thursday. Okay. Okay. Quick time check. Okay. So next up, I wanted to switch gears. So everything we saw with skills really applies to, from your question, applies to the 2B and 4B model, which mostly on mobile phones will be destined for system gen AIs, where that'll really turn up in production. Certainly for IoT or for desktop or for edge applications, you could just load that model yourself and run these skill use cases in production on that Qualcomm IoT platform we saw earlier, for example. But to deploy models in app today and ship them in production apps, we're seeing more people are using smaller models to do that. So we would call tiny models, which are less than 1 billion parameter models in our view. These are the sort of things that we work with teams to deploy LLMs within their app. And these are the types of things we're seeing. So I want to briefly take a look at what that workflow looks like. So LIDAR TLM, this is the engine that powers Gallery, right? And itself is an open source project that has C++ and Java APIs. There's Swift APIs coming soon. It also has a Python API as of last week, which is relevant for IoT developers who work in Python. And that takes an LLM file and has the required components to fully run an auto-aggressive loop and expose that for your easy to use APIs. Yeah, so it's cross-platform C++ APIs. And where relevant, the same API can be used with hardware acceleration, the Qualcomm example. We also, for other smaller models in the past, we've published, we don't have this yet published for Gemma 4, but in the past, we've also published models that work on MediaTek Silicon as well and Intel Silicon as well. So over the course of the year, and as there are more and more smaller models available, we'll see broader NPU support. So then the workflow is, you know, you start from transformers. We use a package called LIDAR T-Torch that can do, that has some hydrogenated optimizations that will optimize for LIDAR T and also has quantization built into the workflow. So then that gives you a LIDAR TLM file and you can then deploy that. You can either prototype it with the Gallery app or just deploy it directly to your production candidate use case using LIDAR TLM and whichever platform you want to work with. Yeah, this is another view of that flow. So for really advanced use cases, we have LIDAR T-Torch Generative API. So if you actually want to write your own, if you want to write your own tiny model from scratch and train it from scratch, we support that workflow as well as just fine tuning stock standard workflows. And there's an API called the Torch Generative API that allows you to basically, it's got building blocks for LLMs that supports many colon LMs that are in native PyTorch that are in a way that give really good performance when you run them on device. And so that's the flow. Wow, this is really in the weeds now. Yeah, this is another view of that flow. So for really advanced use cases, we have LIDAR T-Torch Generative API. So if you actually want to write your own, if you want to write your own tiny model from scratch and train it from scratch, we support that workflow as well as just fine tuning stock standard workflows. And there's an API called the Torch Generative API that allows you to—it's got building blocks for LLMs that supports many LLMs that are in native PyTorch in a way that give really good performance when you run them on device. And so that's the flow. Wow, this is really in the weeds now. So this is our stack to deploy on NPUs. The takeaways from this slide are maybe two things. One is under the hood, we actually invest a lot of work in optimization libraries. So we've X and Npac, which is a CPU optimization library, and ML Drift, which is an optimization library for GPUs. So we have teams who work on these to ensure that both of these have excellent performance. Lots of Google's 1P apps rely on both of these libraries. So we're very motivated to ensure they have really great performance and work in the widest possible set of devices. And these are used what we call the JIT workflow, where we produce a single artifact called a LightHort TLM file or a LightHort T file that can work across CPU and GPU and get deployed to lots of types of devices. For NPU, we need something a little more specialized. For NPU, we need to call a vendor compiler plugin upfront, and that uses an ahead of time compiler workflow. So you need to produce an artifact that's particular to a particular NPU. And then within our runtime, we call a dispatch to a particular API that's allowed to dispatch work to the device driver of the NPU. But both of these are available through a consistent API. So even though the path JIT versus AOT affects the build workflow, the actual app development workflow is very similar across NPU or CPU or GPU. Export and inference is really simple. We have LightHort Torch to just export. I know I've talked a lot about Gemma models because it's the Gemma week for us. So we're very excited about Gemma at the moment. It's worth calling out. We do support third-party models as well. So there's a QEN 0.6b parameter model is what you're seeing there. And those QEN models also work in the app, right, which is what you're going to see on the right-hand side. And the app also supports just loading any LightHort TLM file and running it and getting benchmark stats, which we'll hopefully see in the video in a second. So you can run it, say, on GPU. And then, dun-dun-dun. Dun-dun-dun-dun. The latest pixel that you're simulating here? [SPEAKER_01] Wow. Actually, that's a great question. To be honest, I don't know. We do a lot of testing on Pixel and a lot of testing on S25 as well, right? So I assume it's one of those two devices, right? But I don't know definitively. And we actually didn't show it there because recently, in the last release, we've added an optional icon underneath each chat. And if you click on it, it shows you the pre-fill decode stats for the model. So you can, if you want, find any model we have on Hugging Face and load that in the app and then run to see the benchmark stats for this model on a particular phone if you want to. So, yeah. So you can, and also in the command line, if you run on desktop, you can also just use LIDA or TLM run if you want to do some desktop prototyping to understand how models work on desktop, you can also do that. And then, yeah. So I guess this is another, yeah, this is another third party model. This is fast VLM. This is a model from Apple. It's actually really a really nice VLM model that's only 500 million parameters. And literally this is running with hardware acceleration on Qualcomm, which is why it's running so fast. And this is running, yeah, this is running on an S25, I think, because it's Qualcomm Silicon. And this is literally, we just have it running in the loop saying, describe the scene, describe the scene with video input. It runs, yeah, it runs really fast. So this is a good example of what's possible with a model that's feasible to deploy on device. We haven't in this case done 4-bit quantization, but I assume had we been motivated enough, we could have done that. And we would then have something that only requires maybe 250 or 260 megabytes in your app to give this type of experience. So, yeah, they're certainly feasible. And this is earlier when I was saying about models. This is an example of a general purpose 500 million parameter model. It's been trained on a fairly narrow set of things, which are scene descriptions or image to description. So it's a narrow-ish use case, but still general purpose also, right? So this is a good example of the larger tiny models. The smaller ones, the ones we have with Gemma 3 270M, for example, those models typically require fine tuning to do a particular task. But, yeah, this is a good example. This is why I included it. This is a good example of a general purpose tiny model that is very useful. There's also, it's worth noting, I don't think I've included in the talk. Other examples of general purpose tiny models include very small transcription models, or some narrow pairwise translation models out there at the moment as well. Many of which we support, many of which are on our Hugging Face page as well, if you want to check them out. [SPEAKER_00] There's a bunch of things available on Hugging Face. So, yeah, next slide. This is just showing a handful of models that I wanted to talk about. So Function Gemma is one that we published with a partnership with GDM last year. This is Function Gemma that was a general purpose model that you can further fine tune for function calling. And there are collab notebooks out there if you want to look it up in terms of how to format data sets and how to fine tune Function Gemma for function calling. The next two ones above that are mobile actions and tiny gardens. Mobile actions is the example I mentioned earlier that has the 10 different mobile actions that we fine tuned ourselves, So, next slide. This is just showing a handful of models that I wanted to talk about. So Function Gemma is one that we published with a partnership with GDM last year. This is Function Gemma that was a general purpose model that you can further fine tune for function calling. And there are collab notebooks out there if you want to look it up in terms of how to format data sets and how to fine tune Function Gemma for function calling. The next two ones above that are mobile actions and tiny gardens. Mobile actions is the example I mentioned earlier that has the 10 different mobile actions that we fine tuned ourselves, where we achieved, I think it was 86 or 70% reliability on those. Tiny garden is another example. This is a game that we built inside the gallery app that you can play with as well. It's another example of a voice to function calling fine tune model that you can use. [SPEAKER_04] What is the IT suffix? [SPEAKER_04] IT is instruction fine tune. [SPEAKER_00] So in models, when GDM publish models, they sometimes have PT suffixes and sometimes have IT suffixes. PT is when we publish a model just after pre-training before fine tuning. [SPEAKER_00] That's helpful for expert users because you can do all of your own instruction fine tuning. [SPEAKER_00] So you can fully control the model's personality and you're not trying to unlearn some fine tuning we did. In this case, the way we teach the model how to do function calling is actually via instruction fine tuning itself. So when we publish a model for further fine tuning for function calling, it already has a function calling personality. And that's what we wanted to publish. So in this case, this model is for further fine tuning, but it's an IT model. More typically, if you see for larger models, certainly for the Gemma 3 family, we had both IT and PT checkpoints. We had both IT and PT checkpoints for those, depending on if you had a large volume of data and you want to do full fine tuning yourself. There's a separate conversation on sovereign AI use cases that some of the deep mind people are going to do later in the week. And I imagine that will be a case where you start with our pre-trained checkpoint and add a huge corpus of data, sovereign data or enterprise data, and you can fully fine tune a 27 billion parameter model to do stuff. Another example here is embedding Gemma. I know that's technically not an LLM, but embedding Gemma is a text embedding model that we published with the Gemma team last September. And it does text embeddings for RAG type use cases. But it's also a very strong embedding model that only takes 300 million parameters. It's another example of a high utility model. It's not technically an LLM, even if it is a transformer inside. But it's another example of a tiny model that's very useful for on-device use cases. I've covered this already. We have both AOT. AOT compilation is our workflow for our own devices. So for hardware acceleration, on-device is best for distributing small models to lots of platforms. And I covered this as well. LIDAR TLM itself is built on LIDAR TLM, and it supports all of these types of models. And you can find non-LLM models also on Hugging Face. This is relevant because typically, when you're building a more complex app, you need some things around an LLM, like a voice activity detection model or a denoising model. So these models are also available using just the LIDAR T runtime. LIDAR TLM is the runtime that has the full autoregressive loop. It builds in LIDAR T. And there's a set of LLM models that are available to use with LIDAR TLM, but there's lots of supporting non-autoregressive models that are available to use with the LIDAR T API that are relevant to deploy, that are typically used to deploy a full app. And this was the earlier thing I was saying about for advanced usage, you can fully customize models. So if you wanted to write your own LLM or moonshine or phi or quen variant, you can using the codens directory. I'm probably on track to finish a little early. So that's it in terms of the tiny models, examples, and workflows. There's a lot, and there will be more. So there's a lot we can do now with very tiny models. 500M class are available for some standard features for smaller LLMs. We've had a lot of success fine tuning them for apps. And what I'm going to show next is an example of an app that we've built using tiny LLMs. So this is an app that's available on iOS only for some reason, called AI Edge Eloquent. This is a transcription model. But instead, you may have noticed that as a speaker, I say lots of ums and ahs, right? So if you got the transcript of this presentation, it wouldn't be a great transcript. There would be lots of interjections. Eloquent is built for that type of transcription user story where it does dictation, but then as a separate automatic polish step where it can remove all of the interjections and filler words. So if you want to dictate a message for use later, it's able to clean up idioms of speech. One of the other things it has as well is a biasing list. So one of the things we found is if we're talking about LLMs, you would say, oh, have you trained a LoRA for this or got it? And any standard transcription service would, when you say LoRA, translate that to the name L-A-U-R-A, which is typically what happens. Whereas you can give it a list of keywords or technical terms as well. And then the model will bias to those words, right? Because that's how real people spell. So the key features here are A, it runs entirely offline. B, it's got its own biasing dictionary. [SPEAKER_00] And then C, it cleans up the text. If you see in the middle, it has this polished step, right? And overall, this gives a pretty neat offline, no cost solution, right? Because it goes back to that cost motivation at the start as well. And then the model will bias to those words, right? Because that's how real people spell. So you can, the key features here is A, it runs entirely offline. B, it's got its own biasing dictionary, if you want to call it that. [SPEAKER_00] And then C, it cleans up the text. If you see in the middle, it has this polished step, right? [SPEAKER_00] And overall, this gives a pretty neat offline, no cost, right? Because it goes back to that cost motivation at the start as well, because there are some paid services that do this. But it can give you really clean, cleaned up text. Is that last step done by changing, setting special tokens in the RL, or is it done last, overlaying a map or something? [SPEAKER_04] Which is this? [SPEAKER_00] The last step, a lower example? [SPEAKER_04] Is it done in the second model, or is it done in the model? It is, we're going to give me two slides, I'll answer your question. But yeah, good question, right? The question for the recording was, is the polishing done inside the main model or outside, and how is biasing applied? We're going to answer it in two slides, hopefully. Yeah, personalization, this is the personalization. So you can add things like Laura, the example we always use, because it's near and dear to our hearts. So it allows you to, if you want to connect it to your Google account, it can import stuff from Gmail or something, and look for unusual words, and then add them to their biasing list, or you can add them in yourself. The things people usually put in here are names, because models will get uncommon names wrong frequently, right? As well as technical terms, right? So Gianning and Cyril, these are two of the people on the team who developed the app. Yeah, no surprise. That's the example we have. Okay, so this is where we, the text polishing engine is what we're calling in this stage. So we have two steps here. Important to note that both of these, firstly I'll describe, then I'll talk about models. So microphone impact goes into a speech recognition engine that delivers an unfiltered transcription. Then in the bottom half, we have the personalization flow where you get a set of uncommon or unique words that then goes into relative terms. Both of these go into a text polishing engine, which is then a dedicated mini LLM just for text polishing, right? We could probably have built one LLM to do all of this, right? But it's actually one of the realities of mobile development is sometimes with tiny models as well. There's this modularity story. So that same transcription engine in your app, you may have some other use case for that and you may not want to pay. You can recoup the cost of those weights by using them in multiple places, right? So this is the pattern we see emerging as we build more of these types of apps, that modularity playbook. So in theory, those two models could be stitched together in practice. It's a more pragmatic choice to have separate models. This is more depth as well. [SPEAKER_04] Yeah. And you can also inspect what's happening in the middle. It's easier to debug, right? We see the same a little bit with voice to function calling as well. But that's another story for another day. So, okay. Then the other thing to point out, the ASR engine and text generation, these have a lot of little, I don't know if you recognize the Gemma logo in the corner. These aren't officially published Gemma models, but these are, we've basically done the same workflow that we're advising to other people to do. We have taken the Gemma models, the smaller Gemma model, right? This would be a derivative from the Gemma 3.2.7.tm lineage, is the best way I would describe it, right? And we've basically taken that model and done a fine tuning just to have a transcription engine, and then just to have a text polishing engine. We've fine tuned them. Generally, what that workflow looks like, to give you a bit more insight, is we'll use a workflow of using a much stronger LLM in the cloud to generate lots and lots of synthetic data that corresponds to the type of thing we want. And then once you have a few, low digit millions or tens of millions, depending on how ambitious you want to be, of synthetic data, you put that into a fine tuning workflow and you fine tune the base tiny model you're working with to get a derived model. And that same workflow we're showing here, you can then use to, we've used that internally to ship a stronger note-taking app, right? But that same workflow is, the reason I'm showing it here is a real life example of how we're using Gemma-derived tiny models in order to build new production apps. And we're using this flow for lots of other use cases internally as well, supporting Google 1P products and lots of other ways that those products will probably talk about themselves in time, right? But smaller Gemma models are really powerful for this type of use case. And we're seeing good mileage coming from this. Yeah? [SPEAKER_01] I'm going to combine it with keyboard because this is a very useful thing. [SPEAKER_01] Yeah. Yeah. [SPEAKER_01] Yeah. So this is what the text polishing engine does, as part of that. The text polishing engine was trained, instruction fine-tuned, to have a system prompt of, these are your special words. Please correct anything that sounds like these words to these words. And then also remove interjections or lack of clarity or even things like I forgot to say this or scratch that. Things like this. Probably a scale, huh? But for the tiny models, the playbook we're generally seeing is really useful is synthetic data generation with a larger model. And then pick an off-the-shelf, Gemma 3 to 70m, right, or similar, and run with that with fine-tuning. And then you can, combined with quantization, ship a pretty compelling narrow feature to a very wide set of users powered by an LLM that works on lots of devices. [SPEAKER_01] Do you have a GitHub report, some workflow how to fine-tune it? [SPEAKER_01] Because not just fine-tune it, right? [SPEAKER_01] You had, warm up the training and then start? Things like this. Probably a scale, huh? But for the tiny models, the playbook we're generally seeing is really useful is synthetic data generation with a larger model. And then pick an off-the-shelf Gemma 3 to 70m, right, or similar, and run with that with fine-tuning. And then you can, combined with quantization, you can then ship a pretty compelling narrow feature to a very wide set of users powered by an LLM that works on lots of devices. Do you have a GitHub report, some workflow how to fine-tune it? [SPEAKER_01] Because not just fine-tune it, right? [SPEAKER_01] You had to warm up the training and then start? [SPEAKER_01] Yeah, we do. With the Gemma 3 to 70m publication, we do have a collab notebook. And with Gemma 3 to 70m and Function Gemma, both of these, when we ship those models, have collab notebooks that show how to do fully fine-tune models, how to do full fine-tuning. Yeah. Okay, so regarding the fine-tuning, what's your experience? [SPEAKER_03] Let's take the function calling example. [SPEAKER_03] Yeah. [SPEAKER_03] Have you said you had 80% chance of hitting the 10 functions? [SPEAKER_03] That was eight hours. We finished that. That is probably after fine-tuning. [SPEAKER_03] Yeah. And could you say some numbers before? [SPEAKER_03] What's the amount? [SPEAKER_03] Oh, wow. Oh, wow. Is it 10%? [SPEAKER_03] Is it 50%? 40-something percent to 86%, right? Within that 86% as well, right? Because I don't want to... There was one... Oh, wow. We had 10 functions, and maybe two of those functions dragged our average down a lot. There was eight simple functions that were over 90, 93% type thing. So on very simple functions, we had really, really high reliability. And then, yeah, I forget the actual details, but there was two that brought the average down to more 86%, which is where we finished. There's also a blog post you can read about that for more detail. Yeah. But we'll see this in... Yeah. We see this with smaller models all the time. It's a... Our experience is on a given eval. Fine-tuning is between 20 and 40 points on the eval. So it's a really, really significant win for tiny models, right? When we're talking 200 million parameters, right? Or 270, right? We published with 3M. Yeah. Fine-tuning is essential for most things, right? Unless you have a model that is already published to do a narrow task, right? You'll find transcription models out there that are in that size that work really well at one task and don't require further fine-tuning. But that's because they've been scoped to the narrow task from the outset, right? But if you want to do your own narrow task, yeah, right? And ship something to lots and lots of devices, then yeah, fine-tuning is at least for now, it's the workflow of choice. Yeah? It may change. Will we get a colab that can use the web to do? [SPEAKER_04] And how it is mobile is possible. [SPEAKER_04] That's a good question. So we have a colab to get as far as the LIDAR-T file, right? Our support for LIDAR-T LM on web is work in progress, right? Yeah, yeah. I'll just say that. Please look at our GitHub. Please look there for latest status, right? You'll see it. [SPEAKER_00] Yeah. So will we get a fine-tuning manual for Gemma 4 by any chance? [SPEAKER_01] There are... So Gemma 4, what was announced last week was small models and medium-sized models, right? In the past for Gemma 3, we also published tiny models, right? But this is they were the first models we shipped for Gemma 4 were last week, right? So there will be more Gemma models in future, I imagine, right? There are some fine-tuning already available for the larger models, right? But for the tiny models we've published at this point in time, right? And if anybody's reading or listening to this talk in a few months' time, please search the web for the latest information. But as of right now, the tiny models we have published are Gemma 3 tiny models. And they do have for Function Gemma and for Gemma 3.2.7EM, there are fine-tuning workflows available, right? For Gemma 4, as available last week, there are some workflows. There are fine-tuning recipes in Vertex and elsewhere, I believe, for those. But check out the other Gemma talks later in the week and you'll see and hear more. Oh, yeah. So the Eloquent thing is actually available on iOS, if anybody wants to give that a go. Yeah. Not in Europe. Not in Europe? Yeah, I could not download it. [SPEAKER_01] Oh, okay. Yeah. [SPEAKER_01] You can find it on web browser but not on the E-mail app store. [SPEAKER_01] Oh, wow. [SPEAKER_01] So that needs to be enabled, I guess. [SPEAKER_01] That's really helpful feedback. I'll pass that along. Okay. Yeah, and then wrap up. Yeah, the key takeaways. System Gen.AI, medium-sized models in our, or small models in our parlance, right? We'll be turning up in a mobile device near you. Those same models are excellent for use in embedded systems, embedded platforms. At least for now, given the memory we have in mobile phones, which doesn't look to be getting larger anytime soon, given the cost. Tiny models are where it's at for deploying things widely, right? We hope to make those easier and easier. We hope to have stronger and stronger models available through the partnership with JDM. And we want to also make the fine tuning workflows as easy as we can, right? To enable, to make this more accessible. We'll be turning up in a mobile device near you. Those same models are excellent for use in embedded systems, embedded platforms. At least for now, given the memory we have in mobile phones, which doesn't look to be getting larger anytime soon, given the cost. Tiny models are where it's at for deploying things widely, right? We hope to make those easier and easier. We hope to have stronger and stronger models available through the partnership with JDM. And we want to also make the fine tuning workflows as easy as we can, right? To enable, to make this more accessible. But yeah, happy to share what we've been doing in both of these fronts. So I think we've nine minutes if anybody else wants to ask any questions. Happy to, yeah? Is it challenging to handle safety on the edge models? Because you don't have the ability, you don't have a hosted server, you have all of this checking, because they have to shell a lot of the training. [SPEAKER_04] Yeah. And you don't want to restrict things too much. [SPEAKER_04] It's a bit of a lot of access. Yeah, great question. So a question is about safety on edge models. So firstly, the Gemma team spend an awful lot of time on this. And I would defer you to them for all the questions on safety for the models that were published last week. But they do, it's really top of mind, they spend an awful lot of time on safety for those models. Additionally, for system Gen AI within AI core, for example, right? What actually ships there, and you'll probably hear more of this from other people, is within system Gen AI when that actually ships as part of the OS, not as a role model that we published last week. The system vendor will typically have some sort of input and output safety checker on the model, right? As an aftermarket addition, because there's particular things for their product they want and they don't want, right? So that's another layer. For smaller models then, for really tiny models, the safety there is really important. But generally, the way, at least for us, the way we've deployed tiny models, is within a very narrow API or task. Like you can imagine with Eloquent, the app you saw, the risk profile there, technically that's more like a regenerative app than a generative app, if you know what I mean. It's not going to fully make things up on the fly. So we try, yeah. So there's a, you can look at the scope of what the model is trying to do and the API surface and the functional surface. [SPEAKER_00] And typically, tiny models have, to make a tiny model work, it generally has narrower functional scope, which allows for a more, you still need to do safety, but it's a narrower problem you need to solve when you're looking at it, right? Is what I would say. [SPEAKER_00] Yeah. Is there a way, a place where I can look for how to deploy small, a little bit bigger Gemma models on 5090? [SPEAKER_01] Because I have a 5090, or deploy it on there, and then drive out. [SPEAKER_01] Is there a way I can look for? [SPEAKER_01] Oh, so the, okay. The, hmm. So on 5090. So I believe, I'm not sure if we have, I know we don't have specific documentation for that, but our tool flow does support NVIDIA GPUs. NVIDIA is also a Gemma partner, right? They have supported some of the Gemma launches in the past. So it may be, if you look on NVIDIA's own web page, you may see some of that through their TensorRT LM. I know they've supported some of the Gemma launches in the past. But we don't, I can't point you to specific documentation of the fast answer. But there are the two places I would check. Yeah? Okay. [SPEAKER_03] Yeah. [SPEAKER_03] Yeah, sorry. [SPEAKER_02] Yeah, sorry. [SPEAKER_02] So the examples that you showed for agent skills, they're obviously one skill execution, based on the intent, you run the skill that's a good skill. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] Would you change the architecture for multi-skill execution? [SPEAKER_02] Or is that difficult with smaller models? It's a question of, so we do actually support multi-skill execution, right? It's something, so within the app, if you download the app, you can actually define, you can define the skills you want loaded, right? And you can just toggle them on or off, right? And then you can easily say, if you're very specific in your prompt, which is, wow. If you're very specific in your prompt, which is, you know, look up this topic in Wikipedia, summarize the three bullets, and then display as flashcards. So if you're really specific, that's going to work, right? [SPEAKER_00] Frankly, we've only had this model a couple of weeks, so we're still putting models on the clock and seeing what are the boundaries of how much can we do skill stacking and skill chaining, right? So we're literally still in the mode of putting models on the clock there. So most of the examples we were publishing were single-skill, but even in the diarization app, right, if you recall there, Alice, the person doing that, she asked for, she asked to, oh, summarize my mood, or what time am I meeting up with Amy? And then, oh, Wikipedia is going to come up, ask for blah, blah, blah. So within that example, she was able to show within a single conversation, individual turns using individual skills. Yeah, but skill stacking within an individual prompt, I believe we've seen that work where you're pretty explicit, but also we're still learning the boundaries of what's possible with these classes of models. Sorry, a follow-up question. [SPEAKER_02] Yeah. So the decision to identify with skills like what it's wrong, and given that it's a smart model, do you do instruction tuning from a larger model to build an introduction? [SPEAKER_02] So it's really good at? No. This is, the model was, well, the model was changed to be really good at agentic workflows and really good at function calling, right? It wasn't trained, nothing in the Gemma model was trained specifically for our skill pattern, right? That just came afterwards, when we got the model and we started playing well, it was, wow, this works, and then, well, can we do this, right? So it was more that workflow. [SPEAKER_02] So it's really good at? No. This is the model was, well, the model was changed to be really good at agentic workflows and really good at function calling, right? It wasn't trained, nothing in the Gemma model was trained specifically for our skill pattern, right? That just came afterwards, when we got the model and we started playing with it, it was like, wow, this works, and then, well, can we do this, right? So it was more that kind of workflow. [SPEAKER_00] So there's no specific training for our skill structure, right? [SPEAKER_00] But the team spent lots of time doing general purpose thinking and function calling, right? The GDM team did lots of great work there to give us a general purpose model that is really strong, but there's nothing special for our app, right? [SPEAKER_00] So if you had a slightly different take on skills, and maybe there's a better skill architecture than the one we've shown you today, right? It's entirely possible. Yeah. You could expect to be pretty successful, right? There's no hidden secret sauce in what we've shown. Yeah. I'll take class. Sorry, I think you're maybe next. Yeah. [SPEAKER_07] I'm building an agentic app. [SPEAKER_07] Yeah. [SPEAKER_07] I've been using all the kind of model providers, and I tried to switch to the general model. [SPEAKER_07] I've got a few challenges. [SPEAKER_07] The first big challenge is context. [SPEAKER_07] You know, particularly once you start doing the agentic loop, the context goes up to 100k quite easily. [SPEAKER_07] Yeah. [SPEAKER_07] Can you talk a little bit about context window on the model? [SPEAKER_07] Yeah. Ooh, E2B. [SPEAKER_00] Okay. [SPEAKER_00] Yeah. [SPEAKER_00] I would defer you to the Gemma team for official guidance. [SPEAKER_00] The medium sized models have a context window of 128k, and the smaller models I would actually need to double check, right? I think they said 32k. 32k. [SPEAKER_00] Yeah. That's why I was wondering if it was 32k. [SPEAKER_00] So our implementation in Gallery, we default to 8k or 12k or something just for performance reasons, right? [SPEAKER_00] But the models do support up to 32k, and the other models support up to 128k. With 32k, you use a lot of memory. [SPEAKER_07] Wow. [SPEAKER_00] I wonder if we have stats for that in the model card. It's for the E2B and the E4B model. The memory footprint for a larger context is not as bad as you think, right? There's actually been a lot of optimization. The team optimized that metric for those models because it was targeted for edge use cases. So the amount of KV cache that's required for each input token, that was something that was optimized. So the models behave pretty well on that front. I don't have a number off the top of my head of bytes per input token to give you, but it's good for its model class is what I would say. Yeah. And now the second question. [SPEAKER_07] Do you have the iOS version of the edge? [SPEAKER_07] Yeah. AI Edge Gallery works in both iOS and Android. In terms of being open source, because I think the edge version for. [SPEAKER_07] macOS, that's a good point. I will put that in the coming soon bucket, right? Certainly one of the items on our to-do list because we published the iOS app for the first time in January. Whereas the Android app has been available since last summer, right? But we do have the intention to have a what you see is what you get experience for developers, right? So you can use the app, have fun, experiment with the models, then also get the source code and see how it's built, et cetera, right? So yeah, that's certainly our intention. Yeah? [SPEAKER_03] Is there a— [SPEAKER_03] Probably last question because we're at— Is there a trade off between tuning individual models for specific tasks and the actual amount of memory on device they consume? [SPEAKER_07] So it's like you've got your, one type one would be like 3.6 gigabytes or something. Yeah. Does it not add up over time when you're chaining the models together? So to clarify: For which model is your question? The— I've just downloaded whatever the E2B— Okay. [SPEAKER_03] So yeah. For E2B, we would recommend customization via skills or via prompting, not via fine tuning, right? For the smaller— [SPEAKER_00] So for the small models that are published, we recommend customization through skills and prompting, right? [SPEAKER_00] For tiny models, we would recommend customization through fine tuning. If you were deploying a smaller model, right? On the other path that is available to you for the small models is LoRa fine tuning, right? I know Apple supports that in their foundation model framework. You can check with the AI core speaker if that's on their roadmap, right? There's an AI core speaker who has an AMA coming up. You can check with them about their roadmap for this. But certainly if you were deploying that on an embedded system, right? If somebody asked me about deploying the 2B on a robotics platform, I would be like, absolutely. You should fine tune LoRas for each of your things, for each of your tasks. And our runtime supports loading the model and hot swapping LoRas. [SPEAKER_00] So you don't even need to load and unload the model to load and unload LoRa adapters. It's built for that particular use case for robotics or IoT platforms. And those LoRas are maybe, yeah, it depends on the radix you choose, but they're much smaller. You know, maybe 16 to 100 megabytes, or even smaller, like 8 to 100 megabytes in that kind of range, depending on the radix you use. Yeah. Thanks so much. And then our runtime supports loading the model and hot swapping LoRa's. So you don't even need to load and unload the model to load and unload LoRa adapters. It's built for that particular use case for robotics or IoT platforms. And those LoRa's then are maybe—yeah, it depends on the radix you choose—but they're much smaller. It's maybe 16 to 100 megabytes, or actually even smaller, like 8 to 100 megabytes in that range, depending on the radix you use. Yeah. Thanks so much. Yeah. [SPEAKER_03] Cool. All right. That's a wrap. I'm going to let everybody get lunch. Yeah. Thank you. Thank you. Not in Europe? Yeah, I could not download it. Oh, okay. Yeah. You can find it on web browser but not on the E-mail app store. Oh, wow. So that needs to be enabled, I guess. That's really helpful feedback. I'll pass that along. Okay. Yeah, and then wrap up. Yeah, the kind of key takeaways. System Gen.AI, medium-sized models in our, or small models in our parlance, right? We'll be kind of turning up in a mobile device near you. Those same models are excellent for use in embedded systems, embedded platforms. At least for now, given the memory we have in mobile phones, which doesn't look to be getting larger anytime soon, given the cost. Tiny models are kind of where it's at for deploying things widely, right? We hope to make those easier and easier. Like, we hope to have stronger and stronger models available through the partnership with JDM. And we want to also make the fine tuning workflows as easy as we can, right? To enable, to make this kind of more accessible. But yeah, happy just to kind of share what we've been doing in both of these fronts. So I think we've nine minutes if anybody else wants to ask any questions. Happy to, yeah? Is it challenging to handle safety on the edge models? Because you don't have the ability, you don't have like a hosted server, you have all of this checking, because they have to shell a lot of the training. Yeah. And you don't want to restrict things too much. It's a bit of a lot of access. Yeah, like great question. So a question is about safety on edge models. So firstly, the Gemma team spend an awful lot of time on this. And I would defer you to them for all the questions on safety for the models that were published last week. But they do, like it's really top of mind, like they spend an awful lot of time on safety for those models. Additionally, for like for system Gen AI within AI core, for example, right? What actually ships there, and you'll probably hear more of this from other people, is within system Gen AI when that actually ships as part of the OS, not as a role model that we published last week. The system vendor will typically have some sort of input and output safety checker on the model, right? As a kind of aftermarket addition, because there's particular things for their product they want and they don't want, right? So that's another layer. For smaller models then, for like really tiny models, the like safety there is really important. But generally, the way, like at least for us, the way we've deployed tiny models, is to within a very narrow kind of API or task. Like you can imagine with Eloquent, like the app you saw, that kind of the risk profile there, like technically that's more like of a regenerative app than a generative app, if you know what I mean. It's not going to fully make things up on the fly. So we try, like, yeah. So there's a, you know, you can kind of look at the scope of what the model is trying to do and the API surface and the functional surface. And typically, tiny models have, to make a tiny model work, it generally has narrower functional scope, which allows for a more kind of like, you still need to do safety, but it's a narrower problem you need to solve when you're kind of looking at it, right? Is what I would say. Yeah. Is there a way, like, place where I can look for how to deploy small, like, a little bit bigger Jemma models on like 5090? Because I have a 5090, or deploy it on there, and then drive out. Is there a way I can look for? Oh, so the, okay. The, hmm. So on 5090. So I believe, I'm not sure if we have, like, so I know we don't have specific documentation for that, but our tool flow does support NVIDIA GPUs. NVIDIA is also, like, a Jemma partner, right? They, they have supported some of the Jemma launches in the past. So it may be, it may be, if you look on NVIDIA's own web page, you may see some of that through their TensorRT LM. I know they've supported some of the Jemma launches in the past. But we don't, like, I can't point you to specific documentation of the fast answer. But there are the two places I would check. Yeah? Okay. Yeah. Yeah, sorry. Yeah, sorry. So the examples that you showed for, like, agent skills, they're obviously, like, one skill execution, like, based on the intent, you run the, you run the, the skill that's a good skill. Yeah. Yeah. Would you sort of, like, change the architecture for, like, multi-skill execution? Or is that difficult with smaller models? It's a question of, so we do actually support multi-skill execution, right? It's something, so within the app, if you download the app, you can actually define, you can define the skills you want loaded, right? And you can just toggle them on or off, right? And then you can easily say, like, if you're very, very specific in your prompt, which is, like, wow. Like, if you're very specific in your prompt, which is, like, you know, look up this topic in Wikipedia, summarize the three bullets, and then display as flashcards. So if you're really, really specific, that's going to work, right? Frankly, like, we've only had this model a couple of weeks, so we're still putting models on the clock and seeing what are the boundaries of how, like, how much can we do skill stacking and skill chaining, right? So we're, yeah, we're literally still in the mood of putting models on the clock there. So most of the examples we were publishing were single-skill, but even in the diarization app, right, if you recall there, like, Alice, the person doing that, she asked for, she asked to, oh, summarize my mood, or what time am I meeting up with Amy? And then, oh, Wikipedia is going to come up, ask for blah, blah, blah. So, like, within, in that example, she was able to kind of show within a single conversation, individual turns using individual skills. Yeah, but, like, skill stacking within an individual prompt, I believe we've seen that work where you're pretty explicit, but also we're still, like, frankly, we're still learning, right, the boundaries of, you know, like this, yeah, we're still learning the boundaries of what's possible with these classes of models. Sorry, a follow-up question. Yeah. So it's, like, the decision to identify with skills like what it's wrong, and given that it's a smart model, do you, like, do instruction tuning from a larger model to build an introduction? So it's, like, really good at? No. This is, like, the model was, well, the model was changed to be really good at agentic workflows and really good at function calling, right? It wasn't trained, like, nothing in the Jemma model was trained specifically for our skill pattern, right? That just kind of came afterwards, like, when we got the model and we started playing well, it was like, wow, this works, and then, well, can we do this, right? So it was more that kind of workflow. So there's no specific training for our skill structure, right? But, you know, the team spent lots of time doing general purpose thinking and function calling, right? Like, the GDM team did lots of great work there to give us some general purpose model is really strong, but there's nothing special for our app, right? So if you had a slightly different take on skills, and maybe there's a better skill architecture than the one we've shown you today, right? It's entirely possible. Yeah. You could expect to be pretty successful, right? There's no, like, hidden secret sauce in what we've shown. Yeah. I'll take class. Sorry, I think you're maybe next. Yeah. I'm building, like, an agentic app. Yeah. I've been using all the kind of model providers, and I tried to switch to the general model. I've got a few chance. The first big challenge is context. How, you know, particularly once you start doing the agentic loop, you know, the context go up to, like, 100k quite easily. Yeah. Can you talk a little bit about context window on the, on the, on the, on the. Yeah. Ooh, E2B. Okay. Yeah. I would, I would defer you to the Gemma team for official guidance. Like, the, the medium sized models have a context window of 128k, and the smaller models I would actually need to double check, right? I think they said 32k. 32k. Yeah. That's why I was wondering if it was 32k. So, like, our implementation, like, in Gallery, we default to, like, 8k or 12k or something just for performance reasons, right? But the models do support up to, like, E2B and 4B support up to then 32k, and the other models support up to 128k. With 32k, you use a lot of memory. Wow. I wonder if we have stats for that in the model card. It's for the E2B and the E4B model. The, the memory footprint for a larger context is, it's, it's not as bad as you think, right? There's actually put a lot of optimize it. Like, the team optimized that metric for those models because it was targeted for edge use cases. So the amount of, kind of, KV cache that's required for each input token, that was something that was optimized. So the models behave pretty well on that front. I don't have, I don't have a number off the top of my head of, like, bytes per input token to give you, but it's, it's, it's, it's good for its model class is what I would say. Yeah. And now the second question. Do you have the iOS version of the edge? Yeah. The, uh, AI Edge Gallery works in both iOS and Android. In terms of being open source, because I think the edge version for. Ooh, I, I, I, macOS, uh, that's a good point. I will put that in the coming soon bucket, right? Uh, certainly one of the items on our to-do list because we published, yeah, the iOS app, we only published for the first time in January. Um, whereas the Android app has been available since last summer, right? Uh, but we do, like, our intention is to have, um, what you see is what you get, uh, experience for developers, right? So you can use the app, have fun, experiment with the models, then also get the source code and see how it's built, et cetera, right? Um, so yeah, that's certainly our intention. Yeah? Is there a- Probably last question because we're at- Is there a trade off between, uh, Uh, like, by tuning individual models for specific tasks and then, like, actual amount of memory on device they consume? So it's like, you know, you've got your, one type one would be like 3.6 gigawatts or something. Yeah. Does it not, like, add up over time when you're, like, chaining the models together? Ooh, so to clarify, so. For which, for which model is your question? The- I've just downloaded whatever the E2B- Okay. So, yeah. So for E2B, we would kind of recommend customization via skills or via prompting, not via fine tuning, right? Um, for the smaller- So for the small models that are published, we recommend customization through skills and prompting, right? Um, for tiny models, um, we would recommend, uh, we would recommend customization through fine tuning. If you were deploying a smaller model, right? I don't, like, on, on the other path that is available to you for the medium, for the small models is lower fine tuning, right? So I know Apple supports that in their foundation model framework. Uh, you can check with the AI core speaker if that's on their roadmap, right? There's an AI core speaker who has an AMA, uh, coming up. You can check with him about their roadmap for this. But certainly if you were deploying that on an embedded system, right? Like, uh, if somebody asked me about deploying the 2B on an, uh, on a robotics platform, I would be like, absolutely. You should fine tune LoRa's for each of your things, uh, for each of your tasks. And then, like, our runtime supports loading the model and, like, hot swapping LoRa's. So you don't even need to kind of load and unload the model to load and unload LoRa adapters, uh, that it's built for that particular use case for, like, robotics or IoT platforms. Um, and that's, and those LoRa's then are, like, maybe, yeah, it depends on the radix you choose, but they're much smaller. It's, like, you know, maybe 16 to 100 megabytes, or actually, even smaller, like, 8 to 100 megabytes in that kind of range, depending on the radix you use. Yeah. Thanks so much. Yeah. Cool. All right. That's a wrap. I'm going to let everybody get lunch. Yeah. Thank you. Thank you.