AI Engineer

From 46% to 90%: Fine-Tuning Tiny LLMs for On-Device Agents — Cormac Brick, Google

651 summary words 3 min summary Watch video

Start with the signal

3 min read

Summary

From 46% to 90%: Fine-Tuning Tiny LLMs for On-Device Agents

Main Topics

  • AI Edge Stack & On-Device AI: Google's infrastructure for running LLMs locally on devices
  • Tiny LLMs (TLMs): Language models under 1 billion parameters optimized for mobile deployment
  • Agent Skills: New capability to build function-calling agents on Gemma 4
  • Fine-Tuning for Production: Using synthetic data to achieve robust function calling
  • Real-World Applications: Transcription and function-calling use cases

Key Points

AI Edge Infrastructure

  • Lighter TLM: A runtime for deploying language models as a single file package
  • MediaPipe & Light RT: Cross-framework runtime supporting CPU, GPU, and NPU
  • Scale: Light RT runtime supports 2.7+ billion Android devices with millions of daily invocations
  • Platform Support: Works across Android, iOS, web, and embedded platforms

System-Level vs. Custom AI Approaches

  • System Gen AI (pre-installed): Gemma 4 Nano via AI Core—highly optimized, no app size increase
  • Custom Deployment: Use Lighter TLM for more specific, customized tasks with full control
  • Trade-off: System AI is easier but less flexible; custom models require more work but enable specialized functionality

Agent Skills on Device

  • Built using simple prompt engineering on top of Gemma 4
  • Model is aware of available skills without loading all function details upfront
  • Load Skill Tool Call: Selectively loads only needed skills during runtime
  • Skills can include JavaScript for UI rendering
  • Community has already created 80+ example skills since launch (one week after announcement)

Fine-Tuning Tiny Models: The 46% to 90% Success Story

  • Function Gemma: 270 million parameter model for robust function calling
  • Challenge: Out-of-box performance was only 46% for a 7-function calendar/email task
  • Solution: Fine-tune using synthetically generated datasets
  • Results: Achieved 90%+ accuracy for 8 of 10 functions, 80%+ for remaining 2
  • Deployment: Successfully runs on legacy devices (Pixel 7) at ~2,000 tokens/sec prefill, 140 decode

Performance & Efficiency

  • Visual Language Models: 500M parameter models running optimized on Qualcomm NPUs
  • Synthetic Data Generation: Using Gemini to create training datasets for specialized tasks
  • Hardware Acceleration: NPU optimization enables real-time performance on older devices

Real-World Application: Eloquent Transcription

  • Built using two tiny Gemma 3-based models (few hundred million parameters each)
  • ASR Engine: Speech recognition component
  • Text Polishing Engine: Removes filler words and improves accuracy
  • Supports personalization (custom keywords, technical terms, names)
  • Chains models together for compelling offline transcription service

Notable Quotes

> "Tiny models are here. Once you go down to 200 or 100 million parameters, for that model to work, it needs to have a very narrow and focused task."

> "You can get really robust and reliable function calling using this fine-tuning workflow. It's more work than just prompting a larger model, but it does allow you to ship something robust in your app at scale."

> "System gen AI, enough gen AI. That's the overall message."

> "We can use skills to write skills."

Takeaways

  • Fine-Tuning is Essential for Production: Moving from 46% to 90%+ accuracy requires synthetic data fine-tuning—not just prompt engineering
  • Tiny Models Are Production-Ready: With proper fine-tuning, sub-300M parameter models can handle specialized tasks reliably on-device
  • Skills Architecture Works at Scale: Gemma 4 can effectively manage 8+ skills in conversation; multi-skill calls in single interactions still being optimized
  • Hybrid Approach is Optimal:
  • Use system AI (Gemma 4 Nano via AI Core) when available for simplicity
  • Deploy custom fine-tuned tiny models for specialized, customized tasks
  • Tooling Ecosystem Enables Success:
  • Function Gemma fine-tuning lab (Hugging Face Space)
  • Open-source Google AI Edge Gallery app (Android, iOS)
  • Lighter TLM for deployment across platforms
  • Community-Driven Development: 80+ skills created by community in first week shows strong adoption potential
  • Hardware Matters: NPU optimization critical for real-time performance; even legacy hardware can achieve practical inference speeds
Full transcript 3641 words · 29 min read
0:15

SPEAKER_01

So while we wait for it to come up because I know we're short of time I'm going to talk about agents on device. So I know whoever asked the question about skills and AI core, we have an answer to that. We've built a simple skill harness on top of AI core that you can build skills on. Be able to show that also going to talk about tiny LLMs, which are what we would call LLMs, they're smaller than a billion parameters that are small enough to build into your app if you want to have more customization or you want to do something that isn't already available for you in AI core. So that's the gist.

0:21

SPEAKER_01

So a quick overview of AI edge, how we think about small language models, tiny LLMs and system gen AI, then we're going to take a quick look at agent skills, which is something we can build on top of system gen AI or the new models that are coming down the pipe. And then we're going to take a quick look at tiny models. So that's that. Okay, so yeah, cool, yeah, okay, yeah, feel free. Okay, so AI SLMs and TLMs. Okay, so I think Ali already covered this. We know it's great to do things on device: latency, privacy, offline use, reliability or savings depending on things. These are all motivations to do things locally. By way of intro, I didn't really do this. I'm a software engineering kind of tech lead working on the Google AI edge stack. So that's we have Media Pipe, which is a asset some people may be familiar with. We've Lighter TLM, which is a LM harness that you can integrate with your app where you download the model and ship the model with your app. And then we also have Light RT, which is a runtime that supports both Lighter TLM and Media Pipe. It's a cross-framework runtime for running models, and all of that can run on CPU, GPU, or NPU depending on the platform and depending what's best. And you as a developer get to choose.

0:25

SPEAKER_01

Yeah, it's already trusted at scale. The Lighter runtime, there's a version of that built into Android OS. Lots of Android apps already use it. So it supports over 2.7 billion devices with lots and lots of daily invocations and lots of Android apps leverage this. But also works far beyond Android as well. So we support all of these platforms. And for example, Gemma is available on many of these platforms. Our team is giving another talk tomorrow so you can hear more about Gemma performance on all of these types of platforms and how we're able to do really useful things with the latest Gemma for models.

0:31

SPEAKER_01

But then building on Ali's and Florina's talk, this is the key idea: we have system-level gen AI, which is something that will be pre-installed in the system. So there's Gemma 9 Nano via AI core. This is an example of the summarization API. Apple also has something going on with their intelligence on iOS that I probably know a lot less about. But as a concept, right, as an app developer when you go to build a mobile app, this is one choice: there will often be some form of intelligence built into the system that you can leverage, which is highly optimized, as Ali and Florina covered, and that's available for use with your app.

0:37

SPEAKER_01

Then so this is typically like small language models. For Nano, it is the Gemma 4 E2b and E4b, which are the base models for what we ship there. That's really capable, highly optimized, preloaded on device. If you can use it, it's great. Your app doesn't get any bigger. And if it meets your use case needs, it's a great place to start. If you want more, if you have a more specific task that you want to do that's highly customized or something really boutique, you can use an app gen API. So that's what the Lighter TLM runtime can be loaded with your app or even your web page, right. And this offers a higher degree of customization and reach. Definitely more work. But you kind of have access to smaller models that can run on lots of devices and full customization. So it's clearly a lot more work, but it's the other option that's available.

0:41

SPEAKER_01

Okay, so the rest of the talk, 50 minutes, going to cover two key ideas. One is: how do you do skills on device? Because this is something new that we can do with Gemma 4. It came out last week. We have a few examples of that. This is one key idea. The other idea I want to cover is: for tiny models, what can you actually do with those types of models today? Because we've actually made a lot of progress in the last six to twelve months. So I want to share what's the state of the art with tiny LMs and if you want to use one in your app, how'd you go about that.

0:46

SPEAKER_01

Okay, so there's a lot on the screen. This is an app that our team developed that works on both iOS and Android for running LLMs locally. And here we show both really tiny LLMs so you can see what they can do. But also because Gemma 4 just came out, we're also using this to showcase how Gemma 4 can work on Android and iOS as well. And this actually builds on AI core. When AI core is available on the device, it will use AI core to provide the Gemma model for the app.

0:55

SPEAKER_01

Skills is the thing I want to go into deeply today. But there's a bunch of other things in the app: you can do AI chat, you can ask image, you can do audio scribe, and there's lots of example models. The app also supports third-party models like Qwen or Fire if you just want to load a model, get a feel for how it performs on device. This app is also open source on Android and it's built using Ladder TLM. So it's both a neat way for you to try things out, but also if you're keen, you can dive into the code and see how it all hangs together. And as an example for Ladder TLM, all right. But we're going to dive into the skills because this is a topic du jour. Okay, I'm not going to play this video because I don't have enough time, but yeah, this app is available on Android, iOS, and code is available on GitHub as well.

1:00

SPEAKER_01

Okay, okay, and the app is called Google AI Edge Gallery. So this is the video we will watch because it's shorter and meets my time budget. [SPEAKER_00] Sorry. Hey Gemma, could you find a French restaurant in San Francisco? Please reply to me in English. This uses a restaurant roulette skill, and we'll see how that's built in a moment. Select one winner right. So that's an example of something neat that you can build like with a simple agent harness on top of Gemma 4. That's really just a few lines, pretty easy to do with a few lines of code or a few lines of prompt coding. We'll see in a minute.

1:28

SPEAKER_01

Okay, okay, I got lost a little bit. Okay, all right, sorry, back to where we're supposed to be. So what's actually happening under the hood? So like I was saying, this is built on a prompt. And here you can provide, we have our own system prompt in our app. Then we also put the skill descriptions into the prompt so the model is aware of the Select one winner, right? So that's an example of something neat that you can build with a simple agent harness on top of Gemma 4 that's really just a few lines of code. Pretty easy to do with a few lines of the right vibe coding promises we'll see in a minute.

1:42

SPEAKER_01

Okay, okay, here I don't know how that. Yeah, okay, okay. All right, sorry, back to where we're supposed to be. So what's actually happening under the hood? So, as I was saying, this is built on, and this is built just using a prompt, right? And here you can provide—we have our own system prompt in our app, then we also put the skill descriptions into the prompt, so the model is aware of the types of skills it can use. But it doesn't have to see all of the functions and details of the skill—that's only loaded on demand. And we actually have a load skill tool call built into the model that then selectively, so if you say, "Hey, can you show the select location of the Google office?" it will then know, "Wow, I should use the map scale," that was the skill for math navigation. The tool responds, and then it uses the show JS tool to show you the location in the app as well. So one of the things that's neat about being in an app is you can put simple JavaScript into the skill that we then call as part of the skill. So this is how—I don't have the corresponding demo for this, but this would kind of pop up a nice JavaScript UI of Google Maps to show you in the app right there, similar to the restaurant relish, which was a custom JavaScript to do the rendering to do the relish real piece.

1:48

SPEAKER_01

Okay, so you can create your own skill as well. The app supports this. Sorry, well, yeah, instructions on GitHub. I don't know if I can pass this page too fast. But also, if I'd create your own skill, there's full instructions there if you want to hand write it out. This works really well though. So we can use skills to write skills. So we have Gemini CLI or code. Our team have done about 80 skills and had a lot of fun with this. So this is an example of a prompt which works really reliably. And Gemini CLI, we actually have an ADB skill as well that our team uses a lot, so you can even debug and test by saying, "Hey, you have access to a device via ADB," and you can also ask to test that. So this type of thing actually works really well, and it's fun. And you can then create a skill, and then in the app, there is a dot dot dot button and you can go to load your own skill from a URL if you publish it. You're accustomed to your own get up. It's really easy to do from within the app. You can then also let us know in our discussion on GitHub that you've created a skill and then other people can check out your skill and use that as well. These are some things this has only been out since last Thursday, but these are some example skills that the community have built, so feel free to do it and tag it up here.

1:54

SPEAKER_01

Okay, so that's skills. In the last 10 minutes, we are going to spend on TLM's, or probably more ideally maybe five or six minutes, so this time for questions.

2:00

SPEAKER_01

Okay, so later TLM. This is the runtime that we have that we use for running models that runs models in light or TLM format, which is a single file that packages everything we need to know about the model in order to be able to run it. It's open source. It's fast and it works on multiple platforms. And there is a Swift API and a JavaScript API coming soon. At the moment, if you go to the GitHub, you can see the C++ and Java version, and when we publish the Swift version, we will also publish and we will open source the iOS app at that point in time. So if you go to gallery for the moment, you can only see the code for Android, but hopefully in the next few weeks, where we can get the Swift work finished, have a really good API, and then we'll be able to open source that as well.

2:05

SPEAKER_01

So yeah, and it supports Gemma 4 as well on all of these devices. Also supports loads of other models, but understandably Gemma 4 is our favorite.

2:11

SPEAKER_01

So then to deploy a tiny model, what do you do? So typically starting with Transformers, you then have a package called light or torch that can help you export the model. And then library TLM. There's actually a reference version of that that you can use on your desktop as well if you want to try out a model. And you can either try it out on desktop or you can load it into the gallery and see it perform there. And then you can deploy the light or TLM. It's worth noting for smaller models, you will either pick a fixed function model like a visual language model or a transcription model or something like this. So there are some pre-built models available on our Transformers page that you can use. But something we also see that's really common is people fine-tuning models because certainly once you go down to like 200 or 100 million parameters, for that model to work, it needs to have a very narrow and focused task. And we've had a lot of success deploying those models internally and in a different app that you're going to see in a minute by doing fine-tuning using synthetic data.

2:18

SPEAKER_01

So this is what the export and inference flow looks like. So on the left-hand side of choice exporting a Quen point six model and then running that on desktop with lighter TLM run, so you can just see how that behaves using a GPU for example. The right-hand side is showing a different example, which is Apple's fast BLM, and this is a visual language model. It's only 500 million parameters and this is optimized and running on a Qualcomm NPU, and that's also available through our stack NPU optimization. So this is an example of that happening at engine. This is running really quick because it's using hardware acceleration, and this particular model is 500 million parameters by way of example.

2:24

SPEAKER_01

Another example that we've spent a bunch of time with the Deep Mine team on was publishing Function Gemma, something we published last December. This was based on Gemma three technology. This is only 270 million parameters, but it's robust function calling when fine-tuned. And this is then small and it's really fast even on legacy devices. So if you go all the way back to a Pixel Seven, this still computes almost 2,000 tokens per second prefill and 140 decode, so it's really useful for lots of use cases. It's really useful for lots of simple use cases. You can do text to function calling or voice to function calling using this size model. And there is a whole YouTube video called Function Gemma if you want to find out lots more details about how to do this. We also have a Function Gemma fine-tuning lab, so if you search Function Gemma fine-tuning lab, I don't have it here.

2:31

SPEAKER_01

And this is then it's small and it's really fast even on legacy devices. So if you go all the way back to a Pixel Seven, this still compresses almost 2000 tokens per second prefill and 140 decode. So it's really useful for lots of simple use cases. You can do text to function calling or voice to function calling using this size model.

2:38

SPEAKER_01

There is a whole YouTube video called function Gemma. If you want to find out more details about how to do this, we also have a function Gemma fine-tuning lab. So if you search function Gemma fine-tuning lab, this is available as a Hugging Face space. So you can import, you can define functions, upload your own data, and see how that fine-tunes function Gemma. This is recommended for really robust function calling. We have an example in the app called app intense where it'll do things you saw previously, like add calendar or add email.

2:45

SPEAKER_01

So when we took function Gemma out of the box, our success rate was 46% or something like that. Then we put it through this fine-tuning flow where we're like, hey, we have these seven functions. And instead of providing that via a system prompt, which is what you would do if you're using a larger model or if you're on a device with AI core, you instead need to synthetically create a dataset. That's typically the workflow we use. Flash synthetically creates a dataset. Upload it to this type of tool, or we obviously have our own internal tools. That then got that 46% to over 90% for eight of the ten functions we were trying, and two of the functions were in the 80s.

2:50

SPEAKER_01

So you can get really robust and reliable function calling using this fine-tuning workflow. It's more work than just prompting a larger model, but it does allow you to ship something robust in your app at scale.

2:55

SPEAKER_01

Yeah, so tiny models are here. Okay, I'm going to pull stuff for questions. We have another app called eloquent, which is a transcription service. But what's more interesting than the app was how we built it. It also supports things like personalization, so it does transcription with your own favorite keywords. So if you use a lot of tech jargon or a lot of people's names, transcription services don't always get that correct. This is only available on iOS and not available in Europe. But this will be increasingly available soon.

3:05

SPEAKER_01

But the more interesting thing for this conversation is under the hood. This is something we built using tiny LLMs ourselves. This uses an ASR engine that we built based on Gemma 3 technology, and then also something we call a text polishing engine that we've also built with Gemma 3 technology. Both of these models are only a few hundred million parameters. But chained together, they create a really compelling offline transcription service that is able to leverage your personal addictive break. And the polishing also removes ums and that stuff, which is also a common gripe with offline transcription apps.

3:10

SPEAKER_01

So this does work in production once you put in the effort to fine-tune a model, and you can create pretty compelling things. It's not available on iOS in Europe yet, but it will be available soon. Takeaways: system gen AI, enough gen AI. That's the overall message. I'm happy to take questions. I have three minutes. [SPEAKER_02] So talking about skills, so you personally and your team, how many skills can you start to do with this tiny model before performance starts to degrade?

3:22

SPEAKER_01

Yeah, we are still putting models in the clock. We've literally been playing with the model for about two to three weeks now. It's been in public for about one week. Within a single conversation, we can provide certainly for the four billion parameter model—if you, by default, we enable about eight skills, and it's able to choose between the eight skills reasonably well. Within a conversation, you're able to say, like, find me out a fact on using a Wikipedia search, then wow, show me where that is on Google Maps. So if you have a conversation that uses multiple skills, that works really robustly.

3:28

SPEAKER_01

The thing we're still working on that's harder is, through a single interaction with the app, for the app to know to call multiple skills as part of a single answer. That works sometimes, and that's something we're still figuring out how to make more robust. But yeah, our agent harness thing is really simple, so I'm sure we'll figure that out. We're still discovering the limits of how far we can push the model. [SPEAKER_03] Yeah.

3:42

SPEAKER_01

So the network TLM file format for stock LLMs is effectively a replacement for .taskfile. That's a transition we made last year. .taskfiles are still useful for things like a task file creates more things than just an LLM model. There is a face mesh task, and obviously that has a lot of other code as well. But for LLMs, we want something dedicated and simpler that people can use with open developer tooling. Yeah, so it bundles things like the tokenizer, but it's just the LLM model. [SPEAKER_03] What about CPU? How is the performance on CPU?

4:00

SPEAKER_01

On TPU? CPU. Yeah, so there is another talk tomorrow from some of my colleagues, including Wei, who's here in the second row. That has lots of performance data on Gemma and various models on CPU. You can also check out our model card in the meantime if you search Gemma Laiherty LLM model card. We keep it running. Yeah, so it bundles things like the tokenizer, but it's just the LLM model. [SPEAKER_03] What about CPU? How is the performance on CPU?

4:20

SPEAKER_01

On TPU? CPU. CPU. Yeah, so there is another talk tomorrow from some of my colleagues, including Wei, who's here in the second row. And that has lots of performance data on Gemma and various models on... Yeah. You can also check out our model card in the meantime if you search Gemma Laiherty LLM model card. We keep it running and update that whenever we have new performance numbers on new platforms. So there's a lot of comprehensive data there as well. All right. I'm at zero seconds and it's flashing at me. Yeah. Yeah. 2:30. Sorry, Shinten's here. Sorry, Shinten's here. but it's the other option that's available okay so rest of the talk 50 minutes going to cover two key

4:41

SPEAKER_01

ideas one is hey how do you do skills on device because this is something new that we can do with Gemma 4 it came out last week we have a few examples of that this is one key idea the other idea I want to cover is hey for tiny models what can you actually do with those types of models today because we've actually made a lot of progress in this in the last six to twelve months so I kind of just want to share what's the state of the art with tiny LMS and if you want to use one in your app how'd you go about that okay so this is wow there's a lot on the screen this is an app that our team who've developed that works on

5:14

SPEAKER_01

both iOS and Android for running LLMs locally and here we show both really tiny LLMs so you can see what they can do but also because Gemma 4 just came out we're also using this to showcase what how Gemma 4 can work on Android and iOS as well and this actually builds on AI core when AI core is available on the device it will use AI core to kind of provide the Gemma model for the app so skills is the thing I want to kind of go into deeply today but there's a bunch of other things in the app like you can do AI chat you can ask image you can do audio scribe and there's lots of example models and the app also

5:53

SPEAKER_01

supports 3p models like kind of quen or fire these types of models if you just want to load a model get a feel for how it performs on device and this app is also open source in Android and it's built using ladder TLM so it's both a neat way for you to try things out but also if you're keen you can kind of dive into the code and see hey how does it all kind of hang together and as an example for ladder TLM all right but we're going to dive into the skills because this is kind of a topic du jour okay I'm not going to play this video because I don't have enough time but yeah this app is available Android iOS

6:31

SPEAKER_01

code available on GitHub as well okay okay and the app is called Google AI Edge Gallery so this is the video we will watch because it's shorter and meets my time budget and we don't could we get sounds

6:53

SPEAKER_00

sorry I'll go hey Gemma is-ce-que-tu-peu-trouver un restaurant français a San Francisco please reply to me in English

7:00

SPEAKER_01

this uses a restaurant roulette skill and we'll see how that's built in a moment select one winner right so that's an example of something neat that you can build like with a simple agent harness on top of Gemma 4 that's like really just a few line like pretty easy to do with a few lines of code or a few lines of the right vibe coding promises we'll see in a minute okay okay here I don't know how that yeah okay okay I kind of got lost a little bit okay all right sorry back back back to where we're supposed to be so what's actually happening under the hood so like I was saying this

7:44

SPEAKER_01

is built on like and this is built just using a prompt right and here you can provide we have our own system prompt in our app then we also put the skill descriptions into the prompt so the so the model is aware of the types of skills that can use but it doesn't have to see all of the functions and details of the skill that's only kind of loaded on demand and we actually have a loots a load skill tool call built into the model that then like selectively so if you say hey can you show the select location of the Google office it will then know wow I should use the map scale that was the skill for math navigation the true responds and then

8:25

SPEAKER_01

it uses the show js tool to show you the location in the app as well so one of the things that's neat about being in an app is you can put simple JavaScript into the skill that we then call as part of the skill So this is how like I don't have the corresponding demo for this but this would kind of pop up a nice Kind of like JavaScript UI of kind of Google Maps to kind of just show you in the app right there Similar to the restaurant relish that was a custom JavaScript to do the rendering to do the relish real piece Okay, so you can create your own skill as well. The app supports this Sorry, well, yeah instructions on github I don't know if I can pass this page too fast

9:09

SPEAKER_01

But also and I'd create your own skill there's full instructions there if you want to kind of hand write it out This works really well though so we can use skills to write skills. So we have Gemini CLI or code code like our team have done like 80 skills that had a lot of fun with this So this is an example of a prompt which works really reliably And Gemini CLI we actually have an ADB skill as well That we our team uses a lot so you can even debug and test by saying hey you have access to a device via ADB and you can also ask to test that So this type of thing actually works really really well And it's fun and you can then create a scale and then in the app

9:51

SPEAKER_01

There is a dot dot dot button and you can go to load your own scale from a URL if you kind of publish it You're accustomed to your own get up. It's kind of really easy to do from within the app You can then also let us know in our discussion on github and that you've created a skill and then other people can check out your scale And kind of use that as well. These are some things this has only been out like since last Thursday But these are some example skills that the community have built so feel free to do it and tag it up here Okay, that skills so 10 the last 10 minutes we are going to spend on TLM's

10:23

SPEAKER_01

Or probably more ideally maybe five or six minutes. So this time for questions Okay, so later TLM. This is the runtime that we have that We use for running models that runs models in light or TLM format Which is a single file that packages everything we need to know about the model in order to be able to run it It's open source. It's fast and it works on multiple platforms And there is a Swift API and a JavaScript API coming soon at the moment If you go to the github, you can see the C++ and Java version and when we publish the Swift version We will also publish and we will also open source the iOS app at that point in time

11:02

SPEAKER_01

So if you go to gallery for at the moment, you can only see the code for Android, but Hopefully in the next few weeks where we can get the switch work finished have a really good API and then we'll be able to Open source that as well So yeah, and it supports Gemma 4 as well on all of these devices Also supports loads of other models, but understandably Gemma 4 is our favorite So then to deploy a tiny model what do you do so typically starting transformers you then have a package called light or t torch That can help you export the model and then library TLM There's actually a reference version of that that you can use on your desktop as well

11:42

SPEAKER_01

If you want to try out a model and you can either try it out or desktop or you can load it into the gallery and see it perform there And then you can deploy the light or TLM It's worth noting for smaller models You will either pick a fixed function model like a visual language model or a transcription model or something like this So there are some pre-built models available on our on our on our transformers page that you can use but Something we also see that's really common is people fine-tuning models because Certainly once you go down to like 200 or 100 million parameters for that model to work. It needs to have a very narrow and focused task and

12:21

SPEAKER_01

We've had a lot of success deploying those models internally and in an app a different app that you're going to see in a minute By doing kind of fine-tuning using synthetic data So this is what the export and inference flow looks like so this is This is showing okay on the left hand side of choice exporting a quen point six model and then running that On desktop with lighter TLM run so you can just see how that behaves using a GPU for example the right hand side is showing a different example which is Apple's fast BLM and This is a visual language model. It's only 500 million parameters and this is optimized and running on like a

13:02

SPEAKER_01

This is running on the Qualcomm NPU and that's also available through our stack NPU optimization So this is an example of that happening at engine This is running really quick because it's using hardware acceleration and this is model is just that particular model is 500 million parameters by way of example Another example that we've spent a bunch of time with the deep mine team on was publishing function gemma something we published last December this was based on gemma three technology This is only 270 million parameters But it's robust function calling when fine-tuned type of there And this is then it's small and it's really fast even on legacy devices

13:43

SPEAKER_01

So if you go all the ways back to a pixel seven this still Compossess almost 2000 tokens per second prefill and 140 decode so it's really useful for lots of and It's really useful for lots of simple use cases like you can do text to function calling or voice to function calling using this size model and There is a whole YouTube video on that's called function Gemma If you want to find out lots more details about how to do is we also from Have a function Gemma fine-tuning lab, so if you Search function Gemma fine-tuning lab. I don't have here This is available as a hugging face space so you can kind of import you can define functions upload your own data and

14:26

SPEAKER_01

See how that kind of fine-tune function Gemma and this is kind of recommended for Really high for really robust function calling so we have an example in the app called Like app intense where it'll do like this the thing you saw previously of like add calendar or add email So when we took function Gemma out of the box our success rate and that was I think 46% or something like that then we put it through this fine-tuning flow where we're like hey We have these seven functions and instead of providing that via a system prompt Which is what you would do if you're using a larger model or if you're on a device with AI core for example but

15:07

SPEAKER_01

You instead need to kind of synthetically create a data set right is typically the workflow we use flash to synthetically create a data set Upload it to this type of tool or we obviously have our own internal tools, but that then got that 46% To over like it was over 90% for eight of the ten functions We were trying and two of the functions were a bit lower in the kind of 80s So you can get really robust and reliable function calling using this fine-tuning workflow. Yes It's a bit more work than just prompting a larger model and But it does allow you to kind of ship something robust in your app at scale Sorry going the wrong direction

15:45

SPEAKER_01

Yeah, so then pre-bolt tiny models are here Yeah, okay, I'm going to pull stuff for questions, okay, so we have I don't want to go into this in too much detail We've another app. I'll just speed run this for one minute. We also have another app called eloquent, which is a transcription service But what's more interesting than the app was just an example of like how we built that so it also supports things like Personalization so it does transcription with with your own favorite keywords So if you use a lot of like tech jargon or a lot of people's names transcription service don't always get that correct

16:22

SPEAKER_01

Sadly this is only available on iOS and not available in Europe. Yes, right So this will be increasingly available soon But the more interesting thing for the purpose of this conversation is under the hood This is something we built using tiny LLMs ourselves So this uses a ASR engine that we have built based on Gemma 3 technology And then also has like something we call like a text polishing engine that we've also built with Gemma 3 technology And both each of these models are only a few hundred million parameters But chained together they can create a really compelling

16:54

SPEAKER_01

Offline like offline transcription service that is able to leverage your personal addiction break, right? And also like the polishing also removes on the as and that sort of stuff, right? Which is also a common gripe with kind of kind of offline transcription apps But yeah, so not really available widely as will be available But for the purpose of this conversation, it's more just like a proof of life example So like this does work in production once you put in the effort to kind of fine-tune a model and you can create pretty compelling things Okay, so it's not available in iOS in Europe, so it's not too so yeah takeaways system gen AI enough gen AI

17:31

SPEAKER_01

That's the kind of the overall Yeah overall we kind of message and wrap up happy to take questions. I have three whole minutes I think person there was first

17:41

SPEAKER_02

So talking about skills, so you personally and your team how many skills can you start to do this tiny model before performance starts in 38?

17:52

SPEAKER_01

Yeah, we are still putting models in the clock there So we've literally been playing with the model for about two to three weeks now. It's been in public for about one week We we see like within a single convert so we can provide like certainly for the four billion parameter model Like if you by default we enable about eight skills and it's able to choose between the eight skills reasonably well, right? Within a conversation You're able to say hey like you know like um Like find me out of fact on using a wikipedia scale then oh wow show me where that is on google maps

18:27

SPEAKER_01

So if you have a conversation that uses scale scale scale that works really robustly the thing we're still working on that's harder is through a single like Interaction with the app for the app to know to call multiple skills as part of a single answer who that's and that works sometimes Right and we're still like that's something we're still figuring out how to make that more robust, right? Um, but yeah, like it's all in a Just our like our agent harness thing is really simple, so I'm sure we'll figure that out But we're still kind of discovering the limits of how far we can push the model

19:09

SPEAKER_03

Yeah Yeah Yeah Yeah Yeah Yeah Yeah Yeah Yeah Yeah Yeah Yeah Yeah Yeah Yeah

19:27

SPEAKER_01

So the network TLM file format for just for stock LLMs is effectively a replacement for .taskfile. That's a transition we made last year. .taskfiles are still useful for things like a task file creates more things than just an LLM model, right? So there is like a face mesh task. And obviously that is a lot of other code as well. But for LLMs, we want something dedicated, simpler, that people can use with open developer tooling.

19:54

SPEAKER_03

.

19:55

SPEAKER_01

Yeah, so it bundles things like the tokenizer, but it's just the LLM model. .

20:01

SPEAKER_03

What about CPU? How is the performance on CPU or...?

20:04

SPEAKER_01

On TPU? CPU. CPU. Yeah, so there is another talk tomorrow from some of my colleagues, including Wei, who's here in the second row. And that has lots of performance data on Gemma and various models on... Yeah. You can also check out our model card in the meantime if you search Gemma Laiherty LLM model card. We keep it kind of running... Like we update that whenever we have new performance numbers on new platforms. So there's a lot of comprehensive data there as well. Cool. All right. I'm at zero seconds and it's flashing at me. . Yeah. Yeah. 2.30. Sorry, Shinten's here. Sorry, Shinten's here. . . . . . . . . . .

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note