SPEAKER_00
Hi everyone, can you hear me? Yes, you can hear me. Hi, sorry, one minute late. I'll try to do my best to finish earlier so my friend can do a pretty cool demo for you. I'm Gus, this is Ian, we are from Google DeepMind and I'm specifically working on the Gemma product. Do any of you know what Gemma models are? Okay, perfect, perfect. Thank you very much. So today we're going to talk a little bit about ownership and open models and, well, you know who we are, but the idea is last Thursday we released our new family of models, Gemma 4, and I'm going to talk a little bit about them. There's going to be more information tomorrow in the keynote by Omar and there's another talk by Cassidy also tomorrow that she'll go into even more details. We are going to tell a little bit of the story, but the story is a little bit bigger. We'll try our best here.
SPEAKER_00
So why does it matter? If you ask me, I work for Google, of course, if you ask me which is the best model for you to try, the easiest one I will answer for you, Gemini. Gemini is the best model we have, pretty strong, multimodal, can do all kinds of things. But then there is more to this story than just having the strongest model possible. In some situations, you want to own the model. You want to be able to run on your own hardware. You want to customize it. You want to be able to send your proprietary data that cannot leave your infrastructure. So there are many situations where even the best proprietary model will not be able to help you directly. That's when you might need an open model. That's where Gemma comes in. So when you think, why does Google have two family of models? Because they complement each other. So Gemini is the most intelligent one, can do a lot of cool stuff, but it's hosted in Google servers. You need the API to access. If you need more control and access, you need an open model. That's why we have Gemma. That's why, and we are very proud that the quality is very, very strong. We're going to go into some details later. But the idea is you would be able to do a lot of cool stuff with it.
SPEAKER_00
Among the launches, we released four sizes. Two are targeted to mobile or IoT or smaller devices. It's an E2B and an E4B. These names are a little bit weird. We are the only ones that use this name. And the E stands for effective. And the idea here is the model uses as much as 2B, what a 2B model would use as memory, but it's larger than that. The 2B is around 5B parameters. But then you say, oh, but where is this other 3B in memory? The fun fact is that they are not really parameters from the transformers. They are mapping tokens. So you can leave them in other memory. So what you really need on your GPU memory is the 2 billion or the 4 billion. Why do we do that? So that you can run these models on a phone, on a Pixel phone or any phone you have there. You can run these models and they're very strong. The E2B and E4B, both of them have text, vision and audio input, and they do only text output. They can do thinking. They can do coding, function calling, all these kinds of cool things. These all run on your phone right now. You could download it right now, right? We also have two other models which are the larger ones. We have a 26B and a 31B. The 26 is a mixture of experts, which means that it's as if we had many other models working together where each one of those are like a 4B model. Why does it matter? Because it has 26 billion parameters, but it needs a space of a 4 billion parameter to do the work. And this makes it accessible to way more hardware, to way more people, and it's still pretty strong. But our strongest model is the 31B Dense, which is 31B LLaMA parameters model. And this is really, really strong. If we look into our ELO score on LM Arena, you can see that both our models are fourth and seventh as the leads on open source models, open models. And if you compare them to maybe the top 20, 30, all of them are at least twice, three times larger than our models. In some cases, 20 times larger. So we are talking about a disproportionate amount of intelligence per size. So our 31B model is the one I use very regularly. It can do basically anything from coding, energetic, everything, multilingual, all of that. So I strongly recommend you try those. They are so strong that they are, both of them are really good to use on your, as a cloud deployed model. They can run on your desktop, but if you use on your server as your endpoint to do your work, they are pretty good. And you ask, oh, is this the most intelligent model? No, it isn't. I'm very biased and I love them, but I know the capabilities. But the question is, do you need the most intelligent model on the planet to summarize your email, to do some manual tests, to help you code, to do some agent capabilities that are searching and interacting with docs? Probably not. That's why these models are so strong, because they're cheaper. They're very strong, but they're cheaper to run. They require way less hardware. A 31B running one GPU. The competitors need 200 gigabytes of memory, which would be maybe four or five GPUs. So you can see that the price here is really, really different. One easy place for you to try these models is on AI Studio, where you can try Gemma models, Gemini, all the other ones, but Gemma are there. Both 26 and 31B, you can try right now. They're free. You can play with it. And they can do some cool stuff, which is vision plus thinking plus code execution all at the same time, right? I'll try to post something about this later, but the idea is you can play with the models pretty easily. And right there, not now, let's finish the talk and then you'll play with it.
SPEAKER_00
And as I was saying, the intelligence per parameter that these models bring is pretty good. It's very, very strong. And if we use the ELO score for LM Arena, because it's a benchmark that's a person's preference, right? We can look into academic benchmarks. They are very, very strong. But how the model responds to your queries, that's very important, right? That's how your customers will see, how you will see and interact with it. So this is why this is so important. And why does all this matter? One of the reasons that we care so much is because you want the user to have ownership. And more than that, we are enabling sovereignty. And sovereignty means in terms of you own the model and you can adapt your use cases and you are not susceptible to,
SPEAKER_00
very strong. And if we use the ELO score for LA Marina, because it's a benchmark that's a person's preference. Right? We can look into academic benchmarks. They are very, very strong. But how the model responds to your queries, that's very important. Right? That's how your customers will see, how you will see and interact with it. So this is why this is so important. And why does all this matter? One of the reasons that we care so much is because you want the user to have ownership. And more than that, we are enabling sovereignty. And sovereignty means in terms of you own the model and you can adapt your use cases and you are not susceptible to,
SPEAKER_00
I don't know, loss of service or for some kind of someone saying, no, no, you cannot use this model anymore. It's all available to you. And one of the changes we made last year until Gemma 3 and others, we had our specific license, a Gemma license, which is pretty good, commercial friendly and all. But there's a problem. If you have a custom license, I don't know if you have any lawyers here. If I tell you, oh, we have this custom license, your lawyers will look at me with that face that I hate you guys. And then they will spend like 18 months doing procurement process to understand the
SPEAKER_00
license and trying to change. And that never works. So it's pretty hard for sovereign institutions to adopt this kind of thing. That's why we moved to a past 2.0 for Gemma 4 and going forward. And that makes, thank you, and that makes our life, your life much easier to convince your legal department, let's say like that, that look, we own this model we can use. So this is pretty important and it enables many, many sovereignty institutions to use our models. We have some examples. For example, Ukraine used Gemma to, in parts of their services. We have one version of the Gemma model that was
SPEAKER_00
fine-tuned for Bulgarian. It was their LLM for the country. That was based on Gemma 2. We are working to make sure they use Gemma 4 now. We also have a Brazilian version that is based on Gemma 3, was fine-tuned for Portuguese. And the challenge of these models today is that they, if you want to fine-tune Gemma 2, a specific language, it's becoming very hard to do that. And the problem is hard because not the tooling or anything, it's because the model is pretty strong on those languages already. So any gains you try to have, you might not get there. So you might spend a lot of time to get 1%. And then
SPEAKER_00
maybe, I don't know if it's the best use of your time. So this is good and bad at the same time, because, but it's good that you can automatically use in many languages, you can try right now. And if you're going to the LLM Arena for languages, in many languages are top two, three, and look, it's a 31B model. It's very, very small, right? So this is pretty good. That being said, I will let my colleague continue and show some demos. Thank you, Gus. So one thing that I think is really important about these models is that when you think about using open models, you think about using proprietary models,
SPEAKER_00
we're moving, we're seeing a shift now to more agentic capabilities and the kind of tasks that we're trying to do. And with that comes a cost in tokens and token generation. So one benefit of taking ownership of the models is your ability to control or in cases where you have sunken hardware cost to be able to iterate on top of that. This graph on the right hand side is from the state of AI report that OpenRouter did. And it shows the bits more for you on this diagram, but have a look at that link. It shows the different types of tasks that people are doing through OpenRouter at the moment.
SPEAKER_00
And you'll see the one that's about here, this one here is programming is right in the middle. And this is among some of the highest tasks in terms of token generation, both input and output combined. So the more we have agents work and do these kinds of tasks for us that have very high token generation costs, that's when you start to get more benefit from being able to take control of that in itself. So if, for instance, you have a laptop that is capable of doing a particular task that you need to be doing, like processing a document or analyzing some data or doing some research or in the cases
SPEAKER_00
Gus talked about doing some coding that's suitable for that, then you have a GPU that you can take advantage to do some of that stuff. Now, similar to what Gus said about, we don't necessarily still have frontier models for doing the best possible things. I wouldn't get this model to do a full systems architecture and redesign of your application, right? It's not for that. But what it is very good at doing is following very specific instructions about doing things like refactoring, analyzing, generating code in small modular bits. And you can offload a chunk of work in that style to these kinds of models to be able to do that, whether it's on a single GPU or in your own
SPEAKER_00
personal hardware. And the way that we think about this is like a set of thresholds. Like, when do we get to the point where these models are capable of doing the task, but then they also fit on the right hardware, that they also can do it with the right amount of latency, depending on the, if it's a task for a user, needs to happen in a couple seconds. If it's a task where you're doing things like batch processing, you maybe have slightly different thresholds for what needs to be done. And then also what the cost of actually doing that is. So if you have a sunken cost in terms of
SPEAKER_00
infrastructure that you already own, or that you're prepared to outlay, and then operating on that, or whether you're leasing GPU time or something else. So these are going to be very specific to the tasks that you're trying to achieve. But what you can do with open models is you can think very carefully about what, which of these tasks can I fully offload or can I fully own, compared to relying just on using the best possible models to do that in the cloud. And an example, so Gus talked about the different types of hardware that can run these things now. I'm just going to run
SPEAKER_00
this little demo in the side at the moment. So we now have models that will work directly on mobile and edge devices. This example here was built by Cormac's team is a set of agent skills that the model is running on a phone. So I'm going to mute the microphone for that. So you can talk to the model, you can show it images, you can show it the world around you, and you can prompt it and chat to it. And what this one is showing is that it can look through a set of skills that it has about things on the phone. So either it can take actions on the device itself, so trigger other applications,
SPEAKER_00
like trigger calendar apps, trigger maps apps, or you can define your own skill sets. And what's different now with the Gemma 4 models than we saw for the previous generation is that it's edge devices. This example here was built by Cormac's team is a set of agent skills that the model is running on a phone. So I'm going to mute the microphone for that. So you can talk to the model, you can show it images, you can show it the world around you, and you can prompt it and chat to it. And what this one is showing is that it can look through a set of skills that it has about things on the phone. So either it can take actions on the device itself, so trigger other applications,
SPEAKER_00
like trigger calendar apps, trigger maps apps, or you can define your own skill sets. And what's different now with the Gemma 4 models than we saw for the previous generation is that it's able to reason about what actions it needs to take and reliably make those function calls defined. So what this app will allow you to do is it acts as a playground. So this is Google AI Edge Gallery, and you can find it on iOS and Android, and you can experiment to see what the models of this size are actually able to do. So I think this is the 2 billion parameter model, but there's also the 4 billion parameter model depending on the size of your hardware.
SPEAKER_00
And when we get to desktops and single GPUs, as Gus mentioned, that's where you can use the 26 and the 31B models, again, on your local hardware, and I'll show you how to do that in a minute. But the key point here is that whereas we're not paying for these agents or models within tokens, we're actually paying for them in terms of energy cost, if we think about it. Because now you're thinking about utilization of GPUs, you're thinking about utilization of NPUs on the hardware itself. When are you going to do these tasks? Does the user need to get a response right now when you're taking a picture of something, or is it something that you
SPEAKER_00
can process offline as a background task when they plug their phone in at night? So what I'm trying to say here is that the thresholds and how you think about the usage of these models shifts when you come to on-device or ownership, because you think more about how they're being executed and why they're being executed. Yeah, perfect. And similarly, on the enterprise side, if you don't have a piece of hardware that can run the 31 billion parameter model, you can now be thinking about scaling that down. So maybe if you wanted to use a 300 plus billion parameter model before, you might have
SPEAKER_00
needed multiple GPUs. Now you can think about using a single H100 or A100, or even in some cases like an L4. And then the costs obviously related to that also go down. So again, it's a calculation that you'll have to do depending on your use cases. But there are ways that you could scale, for instance, running one of these models to serve a small team or to serve a company, depending on what you're trying to do. And the final point is that you also have the fine tuning component too, which is that because these models can be customized, you can deploy your own version of it. So for
SPEAKER_00
instance, we have a variant of Gemma models called MedGemma, which is specialized for medical use cases. So if you wanted to have something that would operate on private data that you can control yourself, you can now feasibly deploy this to one or maybe two GPUs to run that for a whole hospital, for instance. So these are worth considering for the enterprise case. I'm going to jump straight to demos now. I've shown you some demos on the phone. I'm going to show you a quick demo here. Quick show of hands, who's ever used a tool called LM Studio? Okay, just under half people. So LM Studio is a way that you can play around with local models.
SPEAKER_00
And I have here, this is the 26B model. So this is our faster of the two larger models with four billion activated parameters. And I'm at the moment, including the context, probably about 26 gigabytes in RAM. And this is an M4 Mac. So I've got unified memory. I've got up to about 48 gigabytes. So I can run it on this machine. And I'm just going to run this terminal right here. Let's give that a go. Oops. Pre-showing my demo. Let's try that again. Okay, so I'm just going to run a little process where I'm going to do some quick translation on my device. So what it's going to do is I've got an orchestrator on this side here,
SPEAKER_00
which is going to hopefully kick off my agent in a minute. Let's make sure we are loaded. Let's see what LM Studio is doing. Yeah, it's just processing at the moment. And then it's going to farm out this translation to all of these different windows. And each one of them represents a different sub-agent. So this is running on my device. And it's going to basically execute all these translations in one go. So I've given it the Gemma 4 announcement. And I just want to translate to all these different languages. So you'll see in a second, it should hopefully send it over there. Three, two, one. And hopefully we should be generating translations in a second. There we go.
SPEAKER_00
So you can imagine doing any kind of generic task on your local machine. You could have it processing files. You could have it doing additional analysis. And hopefully what you'll see in a minute is it will be able to compile all these back. And then it will generate me a quick web page. And then you can see the results of your translation. There you go. So there's the multilinguality of the model there as well. Thank you. Right. So in the interest of time, I just want to say that the main next step for exploring and trying out these models is as simple as this code on the right-hand side. You can take any OpenAI
SPEAKER_00
compatible interface that you've got, and you can point it at a service like OLLAMA or LM Studio. And you can just pick out the Gemma model. And that's all you need to change code-wise to at least try it out. So the first thing we recommend you do is to drop it into existing workflows that you have to then see what the model can handle. What is it working well at? What would it need tuning for? What is out of its depth in terms of the complexity of the task? Next is to bolster your evaluation suites because benchmarks are great for just saying what general capabilities are. But the reality is that how good the model is
SPEAKER_00
depends on how well it does on your task and not anybody else's task. The other thing I mentioned very briefly is thinking about how you actually serve these models in the end. So if you need to run your own GPU and you need to host it, yes, you're in control of uptime and downtime, but then there's maintenance costs and so on. So you have to consider that as one of the factors, the ongoing costs as well, as well as any upfront capex costs if you buy infrastructure or hardware to do that too.
SPEAKER_00
for just saying what general capabilities are. But the reality is that how good the model is depends on how well it does on your task and not anybody else's task. The other thing I mentioned very briefly is thinking about how you actually serve these models in the end. So if you need to run your own GPU and you need to host it, yes, you're in control of uptime and downtime, but then there's maintenance costs and stuff like that. So you have to consider that as one of the factors, like the ongoing costs as well, as well as any upfront capex costs if you buy infrastructure or hardware to do that too. On mobile devices, for instance, you have to think about if I'm going to offload stuff to a phone, what am I supporting? What accelerators do they have? What size RAM do they have? So the conversation becomes a little bit more complex, but then there's a whole heap of things you can unlock like working offline or working on users' private data that never leaves their device. And finally, if you want to scale this up to enterprise levels, you have to think again about the kind of infrastructure that you're running on and what the ongoing costs are of that as well. But it does unlock that. So with that, the summary is that you can use these models in pretty much any way you can think about, experiment what kind of tasks are possible with it, use some of the benchmarks to give you an indication of what's feasible. But really, we want to hear your feedback and how you get on with these and how you fine tune them and what kind of things you run into. And we want to help you on that journey as well. So with that, thank you very much.
SPEAKER_00
[SPEAKER_01] Thank you. [SPEAKER_01] Thank you. you'll have to do depending on your use cases. But there's ways that you could scale, for instance, running one of these models to serve, you know, a small team or to serve a company, depending on what you're trying to do. And the final point is that you also have the fine tuning component too, which is that because these models can be customized, you can deploy your own version of it. So for instance, we have a variant of Gemma models called MedGemma, which is specialized for medical use cases. So if you wanted to have something that would operate on private data that you can control yourself,
SPEAKER_00
you can now feasibly deploy this to like one or maybe two GPUs to run that for, I don't know, like a whole hospital, for instance. So these are kind of worth considering for the enterprise case. I'm going to jump straight to demos now. I've shown you some demos on the phone. I'm going to show you a quick demo here. Quick show of hands, who's ever used a tool called LM Studio? Okay, just under half people. So LM Studio is a way that you can play around with local models. And I have here, I have, this is the 26B model. So this is our faster of the two larger models with four billion activated parameters. And I'm at the moment, including the context is probably about 26
SPEAKER_00
gigabytes in RAM. And this is an M4 Mac. So I've got unified memory. I've got up to about 48 gigabytes. So I can run it on this machine. And I'm just going to run this terminal right here. Let's give that a go. Oops. Pre-showing my demo. Let's try that again. Okay, so I'm just going to run a little process where I'm going to do some quick, a trick, quick translation on my device. So what it's going to do is I've got an orchestrator on this side here, which is going to hopefully kick off my agent in a minute. Let's make sure we are loaded. Let's see what LM Studio is doing. Yeah, it's just processing at the moment. And then it's going to farm out
SPEAKER_00
this translation to all of these different windows. And each one of them represents a different sub-agent. So this is running on my device. And it's going to basically execute all these translations in one go. So I've given it like the Gemma 4 announcement. And I just want to translate to all these different languages. So you'll see in a second, it should hopefully send it over there. Three, two, one. And hopefully we should be generating translations in a second. There we go. So you can imagine doing any kind of a genetic task on your local machine. You could have it like
SPEAKER_00
processing files. You could have it doing additional analysis. And hopefully what you'll see in a minute is it will be able to compile all these back. And then it will generate me a quick web page. And then you can see the results of your translation. There you go. So there's the multilinguality of the model there as well. Thank you.
SPEAKER_00
Right. So in the interest of time, I just want to say that the main next step for exploring and trying out these models is as simple as this code on the right-hand side. You can take any open AI compatible interface that you've got, and you can point it at a service like OLAMA or LM Studio. And you can just pick out the Gemma model. And that's all you need to change code-wise to at least try it out. So the first thing we recommend you do is to drop it into existing workflows that you have to then see what the model can handle. Like what is it working well at? What would it need tuning for? What is kind of out of its depth in terms of like the complexity of the task?
SPEAKER_00
Next is to kind of bolster your evaluation suites because, you know, benchmarks are great and everything for just saying what general capabilities are. But the reality is that how good the model is depends on how well it does on your task and not anybody else's task. The other thing I mentioned very briefly is thinking about how you actually serve these models in the end. So if you need to run your own GPU and you need to host it, yes, you're in control of like uptime and downtime, but then there's like maintenance costs and stuff like that. So you have to be, you have to consider that as like one of the factors, like the ongoing costs as well,
SPEAKER_00
as well as any upfront capex costs if you buy infrastructure or hardware to do that too. On mobile devices, for instance, you have to think about like, if I'm going to offload stuff to a phone, like what am I supporting? What accelerators do they have? What size RAM do they have? So the conversation becomes a little bit more complex, but then there's a whole heap of things you can unlock like working offline or working on users' private data that never leaves their device. And finally, if you want to scale this up to enterprise levels, you have to think then again about like the kind of infrastructure that you're running on and
SPEAKER_01
what the ongoing costs are of that as well. But it does kind of unlock that. So with that, the summary is that you can use these models in pretty much any way you can think about, experiment what kind of tasks are possible with it, use some of the benchmarks to kind of give you an indication of like what's feasible. But really, we want to hear your feedback and how you get on with these and how you fine tune them and what kind of things you run into. And we want to help you on that journey as well. So with that, thank you very much. Thank you. Thank you.