AI Engineer

Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google

2002 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: For broad consumer-edge, IoT, and entry-level robotics deployment, task-specific tiny models fine-tuned on synthetic data are often a better production choice than general-purpose small LLMs because they dramatically reduce RAM, latency, and cost while retaining high reliability for constrained tasks.
  • Why it matters: The talk provides a concrete architecture and training playbook for turning cloud-like agent capabilities—especially voice-to-function calling, dictation cleanup, summarization, and proofreading—into offline, low-cost on-device features.
  • Best use: Use it to inform a tiered model architecture: reserve larger local or cloud models for open-ended reasoning, while using fine-tuned tiny models as fast, private, single-purpose control and interaction components.

Executive Summary

Cormac Brick of Google's AI Edge team argues that the practical constraint on edge AI is not merely compute but DRAM cost. A 1–4B-parameter small LLM may be easy to prompt and capable enough for reasoning and agentic function calling, but its effective deployment footprint generally pushes a device toward 4GB+ RAM once weights, runtime, KV cache, OS, and concurrent workloads are included. That confines these models to higher-end phones, laptops, and more expensive robotics hardware.

The alternative is a tiny, task-specialized model. Rather than expect a sub-500M model to be generally intelligent from zero-shot prompting, the recommended playbook is to select a compact base model, synthesize a task-specific training dataset, and fine-tune it for a narrow input-to-output behavior. Google reports that roughly 10,000 to 10 million synthetic examples can yield high reliability, with task-specific tiny models matching or surpassing 2–4B models on the same narrowly defined task while running faster and on far more devices.

The strongest demonstrated use case is voice-to-function calling: speech recognition feeds a fine-tuned function-calling model that maps arbitrary user text into one of a defined set of device actions. In Google's mobile-actions example, a tiny model handling about 10 output functions reportedly exceeded 86% reliability. The broader implication is that tiny models can make voice the viable control plane for devices where graphical settings interfaces are poor, while avoiding cloud dependency.

The presentation is especially relevant as an implementation-oriented case for heterogeneous agent systems: use fixed-task models for ASR, vision, and embeddings; use small general models where the hardware budget permits; and deploy fine-tuned tiny models for recurring bounded workflows. The main remaining bottleneck is making synthetic-data generation and evaluation accessible enough that developers can cheaply build reliable task-specific models.

Key Takeaways

  • Claim: DRAM, not just accelerator throughput, is the decisive deployment constraint for edge LLMs, making model footprint a product and bill-of-materials decision. | Evidence: Brick says Raspberry Pi 3 16GB pricing has risen 2.5x since launch and notes some phone makers are reducing DRAM. A quantized 2B Gemma model has roughly 841MB of text-only weights, but runtime and KV cache can raise active needs to about 2GB; after OS and concurrent software, Google uses 4GB+ RAM as the practical deployment rule of thumb. | Implication: For mass-market devices or robots, choose the model size only after setting the RAM budget and latency target; otherwise a seemingly cheap model feature can force expensive hardware. | Caveat: The 4GB+ figure is a rule of thumb for this particular 2B deployment stack, not a universal requirement; memory use depends on context length, runtime, quantization, and other workloads.
  • Claim: Small 1–4B models are a strong option when hardware can support them because zero-shot prompting is often enough for useful reasoning, function calling, and agent skills. | Evidence: Google's 2B Gemma example uses mixed 2-bit, 4-bit, and 8-bit quantization at 2.9 bits per weight. It reaches about 7.6 decode tokens/sec on Raspberry Pi, potentially around 2x faster with MTP, about 24 tokens/sec on Jetson Orin Nano using Google's stack, and 31 tokens/sec decode on a Qualcomm IoT board. | Implication: Use this tier for richer local agent behavior on premium hardware, but do not assume it is a universal edge solution. | Caveat: These models remain too memory-intensive or slow for older laptops, lower-end browser targets, and devices where the LLM is only a background subsystem rather than the main feature.
  • Claim: A task-specific tiny model can equal or exceed the quality of a 2–4B general model on a bounded task after fine-tuning, while being materially faster and deployable on lower-RAM devices. | Evidence: Brick describes tiny-model deployments generally below 500M parameters and says Google has used fine-tuned compact Gemma models for summarization and proofreading. The reported training recipe is 10,000 to 10 million synthetically generated examples for high-reliability fine-tuning. | Implication: For repeated high-volume workflows, build narrow specialists rather than paying the latency, memory, and inference cost of a general model on every request. | Caveat: This is not a replacement for general reasoning: the claimed gains depend on tightly specifying the task, generating representative synthetic data, and validating reliability against real user inputs.
  • Claim: Voice-to-function calling is a particularly viable tiny-model agent pattern for IoT and robotics. | Evidence: Google's mobile-actions demonstration uses an ASR model followed by a fine-tuned function-calling model that knows roughly 10 output functions, such as scheduling a calendar event or toggling Wi-Fi, and reportedly calls them with more than 86% reliability from arbitrary free-text input. A 270M-class model is said to run at roughly 45 decode tokens/sec on Raspberry Pi versus mid-single-digit decode speed for the earlier 2B example. | Implication: Treat voice as a constrained intent-to-tool-routing layer, with a finite action schema and strong permissions, rather than as an unconstrained conversational agent. | Caveat: The reliability figure applies to a limited action vocabulary and the presented dataset/task setup; expanding action scope or handling safety-critical commands requires separate evaluation and authorization controls.
  • Claim: A production on-device pipeline can combine multiple specialized tiny models to replace a cloud-only subscription feature. | Evidence: Google's offline voice-dictation app uses a fine-tuned tiny Gemma-based ASR engine plus a separate text-policy model. Beyond transcription, it removes filler words and biases recognition toward a user's relevant names and terms; Brick says it is available to try on iOS. | Implication: Decompose product capabilities into sequential specialist models—recognition, cleanup/policy, personalization—rather than relying on one general multimodal model to perform all stages. | Caveat: The transcript does not provide benchmark accuracy, battery impact, or supported-language coverage for the app.
  • Claim: Hardware-aware model routing is necessary even within local robotics: a model can be functional on a low-end device without delivering an acceptable interaction experience. | Evidence: The OpenDuck Mini V2 example compares a Jetson Nano robot with a Raspberry Pi robot using voice and image input. Both can read signs and react, but Brick characterizes the Jetson interaction as genuinely real-time while the Raspberry Pi version works noticeably more slowly. On a Qualcomm NPU, he estimates a 2B vision model could process roughly three high-resolution image-token frames per second, given 1,120 tokens per high-resolution image and approximately 4,000 prefill tokens/sec. | Implication: Separate perception, reaction, and deliberation latency budgets; an edge agent that technically runs may still fail product expectations if its response cadence is too slow. | Caveat: The visual throughput estimate is model- and hardware-specific and is not a guarantee of end-to-end control-loop latency.

Detailed Brief

Deployment stack and available starting points

  • Claims: Google's AI Edge team develops LiteRT and MediaPipe, contributes edge AI technology to Google products, and works with the Gemma team on device execution.; The edge value proposition is fourfold: predictable low latency, on-device privacy, offline availability, and avoidance of cloud-token costs at large interaction volumes.; Fixed-function models are already a mature route for ASR, visual understanding, and embeddings, and can avoid the need to fine-tune a language model at all.
  • Evidence: AI Edge Gallery is an open-source iOS and Android application intended to let developers test on-device small-model performance and inspect an implementation using Google's open-source runtime.; Apple FastVLM is cited as an example of a 0.5B visual model running quickly on Android with hardware acceleration.; Chrome developer-preview summarization and proofreading APIs are cited as examples where tiny-model delivery broadens the set of users who can receive local AI features.
  • Caveats: The presentation is primarily a Google ecosystem and Gemma-oriented account; it does not offer an independent cross-vendor comparison of quality, tooling, or total deployment cost.; The talk identifies quantization as essential but does not provide a full accuracy-versus-quantization evaluation methodology.
  • Implications: Start by determining whether a fixed-task perception or embedding model solves the problem before introducing a generative model.; Open-source reference applications can accelerate prototyping, but production decisions still need device-specific memory, thermal, battery, and latency validation.

How the tiny-model workflow could evolve

  • Claims: The hard part of deploying tiny models is less model invocation than producing the specialized training data and fine-tuning the model for the intended behavior.; Brick identifies automated or agent-assisted synthetic-data generation as a key next step toward making robust voice-to-function systems broadly accessible.; Faster visual tiny models, including segmentation-capable systems, are identified as an important future direction.
  • Evidence: Google released a Mobile Actions dataset on Hugging Face for recreating the cited FunctionGemma fine-tuning demonstration.; FunctionGemma is described as having additional pre-training for function-calling patterns, while Gemma 3 is positioned as the general-purpose starting model.
  • Caveats: Synthetic data can efficiently cover defined action schemas, but it can also miss ambiguous real-world language, adversarial phrasing, long-tail device states, and policy-sensitive requests.; The talk does not specify a testing harness, error taxonomy, fallback behavior, or rollback mechanism for unreliable tool calls.
  • Implications: A reusable internal asset is not merely the fine-tuned model but the synthetic-data generator, action ontology, evaluator, and regression suite around it.; For tool-using edge agents, reliability should be measured per action, argument extraction, rejection behavior, and authorization outcome—not only as aggregate function-call accuracy.

Notable Concepts & Terms

  • Tiny models: Task-specialized, compact models generally described as below roughly 500M parameters in this talk; their value is lower RAM needs, lower latency, and wider device reach.
  • Small models: General-purpose models in the approximately 1–4B parameter class that can often work through zero-shot prompting but typically demand higher-end device memory budgets.
  • Quantization: Reducing the bits used to store model weights; Google cites mixed 2-, 4-, and 8-bit quantization as central to fitting a 2B model into an edge footprint.
  • KV cache: Inference memory used to retain prior-token attention state; it materially increases real runtime RAM beyond the compressed model-weight size.
  • MTP: A decoding-speed technique referenced as potentially doubling output speed on the Raspberry Pi example, depending on the task.
  • FunctionGemma: A Google model described as receiving extra pre-training for function-calling patterns, making it a starting point for intent-to-tool fine-tuning.
  • Mobile Actions: An open-sourced Hugging Face dataset associated with Google's mobile voice-to-function-calling demonstration.
  • Text policy engine: The second specialist model in Google's offline dictation architecture, used for cleanup such as removing filler words and applying personalization.

Operator Notes / Why Ken Should Care

  • Define a three-tier deployment policy: fixed-task perception/ASR/embedding models first, fine-tuned tiny LMs for bounded agent actions, and small/general or cloud models only for open-ended reasoning.
  • For every proposed edge-agent feature, set a hard device envelope before model selection: available RAM after OS reservation, peak KV-cache use, target tokens/sec, camera/audio pipeline load, thermal limits, and BOM impact.
  • Pilot a constrained voice-to-tool router with a small action schema, explicit per-tool authorization, a refusal/fallback path, and evaluation broken down by intent selection, argument accuracy, and unsafe-action rejection.
  • Invest in a synthetic-data and evaluation pipeline as reusable infrastructure; include real-user utterance collection and regression tests so synthetic coverage does not become the sole reliability signal.
  • Benchmark candidate task-specialist models on the lowest supported hardware rather than extrapolating from desktop or accelerator demos; reject experiences that are technically functional but fail interaction-latency requirements.

Source/Metadata

  • Title: Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google
  • Transcript words: 4601
  • Duration seconds: 1304
  • Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript; the closing portion is duplicated.
Full transcript 3501 words · 20 min read
0:12

Yeah, a bit of a change of speed from the last two talks, where we're looking at higher-end robots. If we want intelligence to get into lots and lots and lots of devices, and not just really expensive robots, we are going to need tiny models. And this talk is about what is the state of the art of tiny models at the moment, what are the things they're good at, and what are the things you can go start building today.

0:15

Okay, so firstly, a bit of background briefly on me and the team I work on. Then we're going to take a look at small models that you may be more familiar with, explore what they can do, what they can't do yet, and then see, hey, why do we need even smaller models? And then, just looking at the state of the art of tiny models today, and what you need to do to get them into a form where you can deploy them in production to do useful things. And lastly, we've got a couple of examples that we can look at from work from our team.

0:22

Okay, so me, I have worked in edge AI for a while. These days, I work as a tech lead on the AI Edge team at Google. Within the team, the types of things we do are we develop open source projects called Lider TLM, Lider T Media Pipe, and these make it easy to deploy AI to edge devices. We also do a lot of work delivering edge AI core technology to Google's own products, some of which would be via tiny models. And then we also work with the Gemmin team to ensure their models work well on and run well on lots of devices. And then we have a significant focus on small and tiny models, because that's what's useful for a lot of mobile phone applications, or if we want to be able to ship a model in browser, that also has to be really, really small. And generally, our playbook is we develop things for first party, for in-house use first. And then if we can figure out a way to share that via an open source package or make those tools available to the wider world, we do so. And that helps lots of other people build similar types of things using open source technology.

0:28

Okay, so why do edge AI? This is probably, as opposed to just doing everything in the cloud, obvious, but I'll go through it anyway. There's latency, you have fast, consistent speed. Privacy, data stays on the device. Offline use, it's reliably available. So that feature that you rely on in your mobile device will still work even when you don't have reception. That can be very helpful. And then savings, especially these days, if the alternative is to call even a faster model on the cloud, that will come at a cost. Particularly if you're shipping an app or a mobile phone app or something in browser, where the user interaction is at very, very large scale. Then even though those tokens are relatively cheap, you're multiplying it by a large number and it'll add up quickly.

0:33

So then, the main challenges of deploying AI on the edge is the leftmost one is new, which is DRAM cost. And it's a really significant constraint, and you'll even see some mobile phone manufacturers are putting less DRAM into their devices this year than previously. You'll also see that since launch, the cost of a Raspberry Pi 3 16 gigabytes has gone up by a factor of 2.5x. So DRAM cost is really, really significant.

0:40

That then casts a shadow over the rest of this talk, right? Where, in order to be able to get AI applications running on the edge, we need to really think a lot about quantization. And we also really need to think about what is the smallest possible model we can use for a given task. Other challenges are, yeah, there's a wider pool of target devices. And yet another challenge is, it's fair to say that a lot of the research hours that go into LLNs these days are into the much larger models and MOE techniques and these types of stuff. And the lower end of the LLN spectrum is a lot less studied. So, yeah, these are challenges of deploying the edge.

0:44

Okay, so small models. And when I say small, I would mean typically maybe one to two or one to four billion parameters. You may find that these are built into the OS. There's a version of a small model that ships in Android high-end phones today with AI Core. There's a version that ships with Apple with Apple Intelligence. Some app vendors will ship models this size in their app. We certainly work with some app vendors that do this. And for IoT and robotics, you would typically require maybe four to eight gigs of DRAM in order to be able to ship this grade of model, which then has an implied cost on the device, right? So it then restricts these models to things like laptops, mobile phones, or higher-end electronics, and puts it out of reach of maybe a lot of lower-tier web browsers or the wider IoT and consumer robotics market.

0:50

Yeah, and for smaller models, developing smaller models, and we'll look in a while, we do a lot of work to minimize footprints with quantization. And the playbook here is mostly prompting, right? If you want to deliver a particular feature using a smaller model, you can just use zero-soft prompting and get pretty good performance. Also, use lower adapters. And it's somewhat robust at doing things like function calling and agent skills.

0:56

Okay, so a really quick example is our, you know, working with the Gemma team, our favorite go-to example is always Gemma for these kinds of things. So we can see that the E2B model is pretty capable in terms of reasoning. It's certainly on par with a Gemma 3 much larger model from 12 months ago. And so we now have much smaller models that are pretty capable at reasoning. And we get pretty decent answers just with zero-shot prompting for a given task.

1:01

We've also done lots and lots of work to optimize the memory footprint of that 2 billion parameter model as much as we possibly can. And so it uses a mix of 2-bit, 4-bit, and 8-bit quantization, getting it down to 2.9 bits per weight if you look at the actual weights we need to hold in memory. We do other tricks like per layer embeddings. I won't go into all of the detail here. But the end result is we can, you know, you need maybe one, like here it's 841 megabytes for a text-only model in memory just for the weights. And then by the time you add in the runtime and a KV cache footprint, you might be up to requiring 2 gigs of active RAM to be able to run this model. Then you account for an OS and the fact that there are other things going on, that's where we get the 4 gig plus rule of thumb for deploying this on a device.

1:06

Then in terms of speed, this is using our runtime. This is just a list of devices that we run on. For the purpose of this talk, we're going to look more closely at the last three rows of the table, which is if we take that 2 billion parameter model and run it on a Raspberry Pi, that will give about 7.6 tokens per second decode. This is without MTP. If you turn on MTP, that will get maybe 2x faster depending on the task. If you go to a higher, more capable device like a Jetson or in Nano, we can get up to maybe 24 tokens per second decode, or maybe even faster if you used NVIDIA's own tool chain. This is with our tool chain. We also have done work to port this to a Qualcomm IoT board, which is pretty popular among higher-end robotics and IoT applications. There, you can see you can get about almost 4,000 tokens per second pre-fill, 31 tokens per second decode. That's useful for lots of almost real-time applications on an NPU because with GEMMA 4 models, one medium resolution image is 500 tokens. A high-resolution image is 1120 tokens. So you could get 3 frames per second of high-resolution tokens going through this model and have pretty decent decode speed as well.

1:17

So there's lots of compelling applications you can build with this type of small model if you're willing to have more expensive hardware and have a more expensive DRAM line on your bill of materials for the device you're building. Yeah, just like our tool chain, we also work with other models in the community that are of similar size, and they each have their strengths as well, right? So these are some of the other models that we support here. That's useful for lots of almost real-time applications on an NPU because with GEMMA 4 models, one medium-resolution image is 500 tokens. A high-resolution image is 1120 tokens.

1:43

So you could get 3 frames per second of high-resolution tokens going through this model and have pretty decent decode speed as well. So there's lots of compelling applications you can build with this type of, with a small model. If you're market or if you're willing to have more expensive hardware and have a more expensive DRAM line on your bill of materials for the device you're building. Yeah, just our tool chain, we also work with other models in the community that are of similar size, and they each have their strengths as well. Right? So these are some of the other models that we support here.

2:20

Really briefly, I won't go into this in too much detail, but we also, if I can get this to play.

2:27

We also have an app that you can use on both iOS and Android. So if you want to take one of these small models to see how fast it works on a phone, you can just go straight ahead and do that. So it's available on AI Edge Gallery. Also, all of the... Oh, I'm getting that buzzing. The app is also fully open source. So if you want to see how to build something similar using one of these models or to see how this is using the open source runtime that runs the models, you can see all of that. So this is a great way of just getting started and trying small models if this is what you want to do. Okay.

3:42

This is another example, which I'm not going to play, but you should definitely check it out. This is an example showing the open source OpenDuck Mini V2 robot. This is one Xavier, one of the engineers in DeepMind, built this. It's a hobby project. Really, really fun. So go check out this YouTube video. What you'll see is he has two robots. One uses the Jetson Nano. One uses the Raspberry Pi. And you'll see that the robot is able to... It's able to read signs and react to things and nod its head. It's also able to take both voice and image input. Yeah, and what you'll see is the Jetson Nano one performs, has really good real-time interaction.

4:52

The one based on Raspberry Pi, it works, but it's a lot slower, right? So for some examples, for some types of interaction, even the best models we have today are maybe not meeting user interaction requirements.

5:03

But yeah, this is a really fun video, so definitely check it out. So yes, then small models, while they're great, right? If your product can afford to use one of these, they're really easy to use, because you just need to zero-shot prompt in order to get it to work. Gemma Team has done great work in having low-footprint, high-capability models that are ready to use, and they're optimized to run on all of those devices you saw earlier.

5:41

And if you can live within those constraints, then great, right? Your journey would stop here, and you would build a feature you would want, right? For lots and lots of other things that we do in our work, and other people that we talk to, we're still at a point where small models are too big, because they can't reach older laptops, or more consumer edge devices. The user interaction needs to be more responsive. We also have the reality, and we do have this a lot of times, where the model you want to run isn't the main feature in the application. It's one tiny thing in a corner that needs to run while everything else in the system is running.

6:25

So we also need a smaller model for system health as a common pattern. So then enter tiny models, right? So these are typically as small as 50 billion parameters. We've deployed models that small to maybe 500 million parameters. They're easier to ship in native applications. They would run on the types of things you see on the right-hand side, and would require maybe less than two gigs of RAM, or even less than that. And they can also be made to run really, really fast. But the playbook to deploying here is a little more complicated.

7:27

So sometimes there are off-the-shelf models that will do what you want, and we'll look at those in the next slide. Or else, if that doesn't work, you're going to be left in a world of fine-tuning a model to achieve a given outcome, which works very, very well. So fixed-task models, there are a bunch of things around ASR, vision, and embedding models. And if you have something, yeah, so ASR, vision, and embedding, these are all stock features, and they work really, really well. This is an example of Apple FastVLM, which is a 0.5 billion parameter model running on an Android device using hardware acceleration. And you can see it runs really, really fast.

8:24

So if you needed to add a little bit of visual intelligence to an edge device or an IoT device, this class of model is an excellent option to get that first level of visual awareness. Or for ASR, yeah, there are some strong models listed here as well. And then lastly, embedding models are great at, yeah, this is just a text embedding model, which is really good at processing and matching text, which can be relevant in some cases. Okay. But then, the next scenario is you want to fine-tune a model. So here you can start with the models I'm citing here are Google-developed models.

9:40

So there are some starting at 270 million parameters and Gemma 3 and Function Gemma. Gemma 3 is a general-purpose model. Function Gemma is one that has extra pre-training for function-calling patterns. So here, the performance, if you remember earlier on the Raspberry Pi, our performance was at mid-single-digit tokens per second decode. So here, that jumps up to 45 tokens per second because we need to read less from memory each time. And we can fine-tune this to do pretty compelling things. So on the right-hand side, this is running what we call a mobile actions model. So this is text in and function calling out.

10:17

This model knows about 10 different output functions and can call them at over 86% reliability from a given arbitrary text input. And this is for doing common things on a mobile device like schedule a calendar or turn on and off Wi-Fi or things like this. And it can take arbitrary free-text input and convert that to appropriate function calling. And for this demo, we've taken another ASR model and put it in front of that, which gives voice to function calling as a feature. And voice to function calling is pretty key for lots of IoT and edge devices because smaller devices tend to require settings menus.

10:33

And that user interface can be really, really challenging for lots of people. So yeah, being able to just talk to something to ask for a given outcome. This is a pretty key capability, and we can do that reasonably reliably using a fine-tuned small model. So the playbook is generally, then, you pick a base model. You check the performance, if the performance and memory footprint are within the range that you want. And then the harder part is the playbook we've found works really, really well as we synthetically generate data to fine-tune that model.

10:49

Depending on the model, there's a data set we've open sourced here called Mobile Actions that's available on Hugging Face that corresponds to this. If you want to recreate that same demo yourself and fine-tune Function Gemma from scratch. But we've generally found that in the range of 10,000 to 10 million samples of synthetically generated data will be sufficient to fine-tune a smaller model to a really, really high degree of reliability. And so for other tasks we've done, things like summarization or proofreading. So something which you could do with a two or four billion parameter model reasonably reliably.

11:21

If you're willing to put the time and energy into creating a synthetic data set and fine-tuning a model, you can achieve similar, the same or greater quality with a model that is much, much smaller, will work on a much wider set of devices, and will be much, much more responsive. So yeah, and that's the type of outcome we're seeing now with just fine-tuning a model for a single task. And it's really, yeah, we found this is a really good playbook for deploying at very wide scale. So here's another example, this is one example in production where we have, this is an app that we've developed for voice dictation without subscription.

11:59

All of the voice dictation happens locally on device. And as well as just doing dictation, it also does, it also does, well, it cleans up ums and ahs, right? If you see on the right-hand side, it's able to clean up text. you can achieve a similar, the same or greater quality with a model that is much, much smaller, will work on a much wider set of devices, and will be much, much more responsive. So, yeah, and that's the type of outcome we're seeing now with just fine-tuning a model for a single task. And it's really, yeah, we found this is a really good playbook for deploying at very wide scale.

12:28

So here's another example in, this is one example in production where we have, this is an app that we've developed for voice dictation without subscription. All of the voice dictation happens locally on device. And as well as just doing dictation, it also does, it also does, well, it cleans up ums and ahs, right? If you see on the right-hand side, it's able to clean up text. It's also able to do biasing towards words and names that are relevant to you personally. So personalization. The left-hand side shows how we built that application. So there's an ASR engine and a text policy engine. And both of these are fine-tuned versions of tiny Gemma models.

13:10

And this allows us to take something that would have been a server-only feature of, where you require a subscription to do highly accurate voice dictation and have an app that's just able to do that completely offline with very, very good quality. So this is something you can try on iOS if you want to give this a go today. But, yeah, and the backbone of this app is two fine-tuned small Gemma-based models in the low single digits, hundreds of parameters, million parameters. We are also worth noting is there is also features in developer preview in Chrome, for example. That summarization and proofread APIs are built-in APIs in Chrome.

13:37

And delivering those features via tiny models allows the Chrome team to ship them to a much wider set of Chrome users than would otherwise be possible. Yeah, so that's, we've probably got to have one minute for questions. Some key takeaways. It's on the last slide, if I can get there. Yeah, so the takeaways from consumer devices and entry-level robotics is small LLMs are easy to use. And especially on NPUs, they're very, very fast. Tiny models will enable reach for a much, much larger pool of devices. And voice-to-function calling can now be built to be robust using tiny models. It just requires investing in an appropriate synthetic data set with enough samples.

14:23

And then you can fine-tune a model to get really good outcomes. Cool. So happy to take one or two questions, or if anybody has one. Yeah? Sorry, I'm going to plug this out. Yeah, sorry. Sorry, say again? Broader ambitions of where tiny models can go? Wow. I think generalizing voice-to-function calling is one key goal. Making that very easy for lots of people, because I think that's a key use case. That if we can figure, if we can figure out how to make, have an agent generate the synthetic data for you, right?

15:30

It's certainly possible to make that journey much easier than it is today and make it available to a lot more people. Yeah. And certainly the visual input as well, that takes a little bit of time at the moment. There's certainly scope to have faster models there that can do a wider set of things, segmentation and other things that would enable other use cases. Awesome. Yeah, due to the time, we probably don't have a Q&A session for today. Yeah, but Cormac will stay after the session, maybe? And you can ask more questions about the time. I'll stay after the session, or you can come grab me downstairs at the DeepMind booth at 4 o'clock. I'll be there 4 to 5, okay?

16:36

Depending on the model, like, there's a data set we've open sourced here called mobile actions that's available on hugging face that corresponds to this. If you want to kind of recreate that same demo yourself and fine tune function demo from scratch. But we've generally found that in the range of 10,000 to 10 million samples of synthetically generated data will be sufficient to fine tune a smaller model to a really, really high degree of reliability. And so for other tasks we've done, like, things like summarization or proofreading. So something which you could do with a two or four billion parameter model reasonably reliably.

17:18

If you're willing to put the time and energy into creating a synthetic data set and fine tuning a model, you can achieve a similar, like, the same or greater quality with a model that is much, much smaller, will work on a much wider set of devices, and will be much, much more responsive. So, yeah, and that's the type of outcome we're seeing now with just fine tuning a model for a single task. And it's really, like, yeah, we found this is a really good playbook for deploying at, like, very wide scale. So here's another example in, this is one example in production where we have, this is an app that we've developed for voice dictation without subscription.

17:58

All of the voice dictation happens locally on device. And as well as just doing dictation, it also does, it also does, well, it kind of cleans up ums and ahs, right? If you see on the right-hand side, it's able to clean up text. It's also able to do biasing towards kind of words and names that are kind of relevant to you personally. So kind of personalization. The left-hand side kind of shows how we built that application. So there's an ASR engine and a text policy engine. And both of these are fine-tuned versions of tiny Gemma models. And this allows us to take something that would have been a kind of, like, server-only feature of, you know,

18:36

where you require a subscription to do highly accurate voice dictation and have an app that's just able to do that completely offline with very, very good quality. So this is something you can try on iOS if you want to give this a go today. But, yeah, and it just, the backbone of this app is kind of two fine-tuned small Gemma-based models in the low single digits, hundreds of parameters, million parameters. We are also worth noting is there is also kind of features in developer preview in Chrome, for example. That kind of summarization and proofread APIs are a feature as built-in APIs in Chrome.

19:13

And delivering those features via tiny models allows the Chrome team to ship them to a much wider set of Chrome users than would otherwise be possible. Yeah, so that's, we've probably got to have, like, one minute for questions. Some kind of key takeaways. It's on the last slide, if I can get there. Yeah, so the takeaways from consumer devices and entry-level robotics is small LLMs are easy to use. And especially on NPUs, they're very, very fast. Tiny models will enable reach for much, much larger pool of devices. And voice-to-function calling can now be built to be robust using tiny models.

19:51

It just requires kind of investing in an appropriate synthetic data set with enough samples. And then you can fine-tune a model to get really good outcomes. Cool. So happy to take one or two questions, or if anybody has one. Yeah?

20:09

Sorry, I'm going to plug this out. Yeah, sorry. Sorry, say again?

20:26

Like broader ambitions of where tiny models can go? Wow. I think kind of generalizing voice-to-function calling is one key goal. Like making that very easy for lots of people, because I think that's a key use case. That if we can figure, like, if we can figure out how to make, you know, have, like, an agent generate the synthetic data for you, right? Like, it's certainly possible to make that journey much easier than it is today and make it available to a lot more people. Yeah. And certainly the visual input as well, that takes a little bit of time at the moment.

21:04

There's certainly scope to have faster models there that can do a wider set of things, like kind of segmentation and other things that would enable other use cases. Awesome. Yeah, due to the time, we probably don't have a Q&A session for today. Yeah, but Cormac will stay after the session, maybe? And you can ask for more questions about the time. I'll stay after the session, or you can come grab me downstairs at the DeepMind booth at 4 o'clock. I'll be there 4 to 5, okay?

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note