AI Engineer

Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind

4147 summary words 18 min summary Watch video

Start with the signal

18 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Gemma 4 open models (2B–31B) achieve frontier-class performance per parameter, enabling full ownership and sovereignty while running on local hardware from phones to single GPUs, licensed Apache 2.0, and ranking 4th–7th on LM Arena despite being 2–20× smaller than competitors.
  • Why it matters: This talk demonstrates how open models now cross critical capability/cost/hardware thresholds for agentic workflows, on-device AI, and enterprise sovereignty use cases—especially important for anyone evaluating infrastructure costs, data privacy, or multi-agent systems where proprietary API costs scale linearly with token generation.
  • Best use: Ken should watch this to understand the strategic rationale for open vs proprietary models in agent systems, evaluate Gemma 4 for operator workflows (coding assistants, batch processing, on-device agents), and internalize the licensing/sovereignty arguments for legal/procurement contexts.

Executive Summary

Google DeepMind released Gemma 4 (Apache 2.0) in four sizes: 2B/4B for mobile/edge (both multimodal input, text output, with thinking/function-calling), and 26B MoE + 31B dense for servers/desktops. The 31B model ranks 4th on LM Arena among open models and 7th overall, despite being 2–20× smaller than competitors. The 26B MoE activates only 4B parameters per forward pass, making it run on 26GB RAM—accessible on single consumer GPUs or unified-memory laptops. Both larger models outperform GPT-3.5-class benchmarks and support vision, thinking, code execution, and function calling.

The strategic framing is ownership vs intelligence ceiling: Gemini is Google's frontier proprietary model, but Gemma enables sovereignty (no API dependency, no data exfiltration, custom fine-tuning, offline operation). The licensing shift from custom Gemma license to Apache 2.0 eliminates 18-month legal procurement cycles for governments and enterprises. Real deployments include Ukraine's public services, a Bulgarian national LLM (Gemma 2 fine-tune), and Brazilian Portuguese tuning. MedGemma (medical fine-tune) can serve a hospital on 1–2 GPUs with private patient data never leaving infrastructure.

The cost/control argument centers on token economics in agentic systems: OpenRouter data shows programming tasks generate among the highest token volumes (input + output). For batch processing, research, refactoring, or modular code generation, a local 31B model on a single GPU costs ~1/5 the infrastructure of 200GB+ competitors and zero marginal API cost once deployed. The tradeoff is capability ceiling—Gemma is not for full systems architecture redesign, but excels at instruction-following, summarization, testing, and agentic tool use within defined skill sets.

Live demos: (1) iOS/Android Google AI Edge Gallery runs 2B/4B models on-device with function calling and multi-agent skill routing (calendar, maps, custom APIs) powered by reasoning chains. (2) Desktop demo using LM Studio on M4 Mac (48GB unified memory) runs 26B MoE at ~26GB RAM to orchestrate parallel translation agents across 8+ languages, compiling results into a webpage. Integration is trivial—swap OpenAI SDK endpoint to OLLAMA/LM Studio with model name change. Gus and Ian emphasize evaluating tasks on your data, not benchmarks, and considering sunken hardware costs (capex), energy costs (on-device), and maintenance/uptime tradeoffs vs API SLAs.

Key Takeaways

  • Claim: Gemma 31B ranks 4th among open models on LM Arena ELO (human preference benchmark), despite being 2–20× smaller than top-20 competitors. | Evidence: Gus states 'both our models are fourth and seventh as the leads on open source models... all of them are at least twice, three times larger... in some cases, 20 times larger.' The 26B MoE is 7th. ELO score chosen over academic benchmarks because 'it's a person's preference... how the model responds to your queries, that's very important.' | Caveat: Not the most intelligent model available—Gus explicitly says 'I'm very biased and I love them, but I know the capabilities.' Unsuitable for complex systems architecture or tasks requiring frontier reasoning. Academic benchmarks are 'very strong' but not detailed in talk. | Implication: For Ken's agent systems: Gemma models are cost-effective for instruction-following, summarization, testing, coding assistance, and multi-agent orchestration where 'good enough' >> 'best possible' when scaled across hundreds of API calls. Evaluate by dropping into existing workflows and measuring task-specific performance. | Timestamp: 04:30
  • Claim: Gemma 26B (MoE) activates only 4B parameters per forward pass but has 26B total, enabling single-GPU deployment where competitors need 4–5 GPUs (200GB memory). | Evidence: Gus: 'The 26 is a mixture of experts... it needs a space of a 4 billion parameter to do the work.' Ian runs 26B on M4 Mac with 48GB unified memory at ~26GB RAM usage in LM Studio demo. Gus: competitors 'need 200 gigabytes of memory, which would be maybe four or five GPUs.' | Caveat: Assumes batch size 1 and quantization is acceptable. Production serving at scale (high concurrency) will still require infrastructure planning. No discussion of throughput vs latency tradeoffs or KV cache memory growth with longer contexts. | Implication: For Ken's infrastructure planning: Cost per GPU-hour drops dramatically—1 H100/A100 vs 4–5. For startups or teams with sunken hardware costs, this unlocks local deployment without cloud API bills. Evaluate energy costs ($/kWh × GPU TDP × uptime) vs API costs ($/token × volume). For multi-agent systems with high token generation (e.g., code generation, research), TCO can be 5–10× lower if traffic justifies dedicated hardware. | Timestamp: 06:15
  • Claim: Apache 2.0 licensing (vs prior custom Gemma license) eliminates 18-month legal procurement cycles for government/enterprise adoption. | Evidence: Gus: 'If you have a custom license... your lawyers will look at me with that face that I hate you guys. And then they will spend like 18 months doing procurement process... That's why we moved to Apache 2.0 for Gemma 4.' Examples: Ukraine public services, Bulgarian national LLM, Brazilian Portuguese model. MedGemma for hospitals with private patient data. | Caveat: Apache 2.0 still has indemnification/patent clauses that some organizations scrutinize. No mention of export controls (ITAR/EAR) or whether Google provides enterprise support/SLAs beyond community support. Fine-tuning for specific languages is now harder because base model multilingual performance is already strong ('you might spend a lot of time to get 1%'). | Implication: For Ken's B2B/enterprise positioning: Apache 2.0 is a critical unlock for selling into regulated industries (healthcare, finance, government). Sovereignty argument = data never leaves infrastructure + model ownership + no vendor lock-in. For multilingual products, base model may suffice without fine-tuning cost/effort. Legal/procurement teams can approve in weeks vs months, accelerating GTM cycles. | Timestamp: 08:45
  • Claim: On-device models (2B/4B Gemma) run on phones with text/vision/audio input, text output, thinking, function calling, and multi-agent skill routing—paying in energy cost, not token cost. | Evidence: Ian demo: Google AI Edge Gallery app on iOS/Android runs 2B model with function calling to calendar, maps, custom skills. Gus: '2B is around 5B parameters... you can leave [token embeddings] in other memory... on your GPU memory is the 2 billion.' Ian: 'We're not paying for these agents or models within tokens, we're actually paying for them in terms of energy cost... utilization of NPUs... when you plug your phone in at night.' | Caveat: No latency numbers provided. Background task scheduling mentioned (offline processing when charging) suggests real-time user-facing tasks may have constraints. Model output is text-only (no multimodal generation). Skill definition/integration complexity not discussed. Device compatibility (RAM, NPU support) varies. | Implication: For Ken's on-device AI strategy: Enables offline-first apps, user data privacy (never sent to cloud), and zero marginal cost per inference after deployment. Ideal for batch background tasks (email summarization, photo analysis, local search) executed during charging. For multi-agent systems, can decompose tasks into on-device + cloud hybrid (private data local, complex reasoning cloud). Threshold question: does task need real-time response or can it wait for optimal device state (charging, idle). | Timestamp: 11:20
  • Claim: Programming/coding tasks generate among the highest token volumes on OpenRouter; agentic workflows amplify API costs, making local ownership economically advantageous. | Evidence: Ian: 'Programming is right in the middle... among some of the highest tasks in terms of token generation, both input and output combined.' State of AI report graph shown. Ian: 'More we have agents work... high token generation costs, that's when you start to get more benefit from being able to take control.' Suitable for 'refactoring, analyzing, generating code in small modular bits.' | Caveat: Not suitable for 'full systems architecture and redesign'—Ian explicitly warns against using for complex architectural tasks. No quantitative ROI calculation provided (e.g., break-even point for GPU capex vs API spend). Maintenance/uptime responsibility shifts to user. Evaluation burden increases (need task-specific evals, not just benchmarks). | Implication: For Ken's coding agent economics: If building multi-agent coding systems (e.g., test generation, refactoring, documentation), API costs scale linearly with agent parallelism and iteration depth. A local 31B on 1 GPU may cost $1–2/hr GPU rental vs $0.10–1.00 per 1M tokens at cloud pricing; break-even at ~100k–1M tokens/hr depending on pricing tier. For internal tools with predictable load, capex a GPU ($10k–30k) amortized over 2 years = $0.57–1.71/hr vs variable API costs. Key: measure actual token volume in pilot workflows to model TCO accurately. | Timestamp: 13:00
  • Claim: Integration requires only changing OpenAI SDK endpoint to OLLAMA/LM Studio and model name—trivial code change to trial Gemma in existing workflows. | Evidence: Ian slide: 'You can take any OpenAI compatible interface... point it at a service like OLLAMA or LM Studio. And you can just pick out the Gemma model. And that's all you need to change code-wise to at least try it out.' Recommendation: 'Drop it into existing workflows... see what the model can handle.' | Caveat: Assumes OpenAI-compatible function calling/tool use schemas are honored (not all open models do). No discussion of tokenizer differences, context window limits, or prompt engineering adjustments. Serving infrastructure (OLLAMA/LM Studio vs vLLM/TGI for production) and optimization (quantization, batching) not covered. | Implication: For Ken's experimentation velocity: Can A/B test Gemma vs GPT-4/Claude in existing agent pipelines in <1 hour. Key workflow: (1) Swap endpoint, (2) Run evals on representative tasks, (3) Identify where Gemma suffices vs where frontier models are necessary, (4) Route tasks accordingly (hybrid architecture). For cost optimization: route 70–80% of 'good enough' tasks to local Gemma, reserve expensive API calls for complex reasoning/planning. This is path to 5–10× cost reduction in high-volume agentic systems. | Timestamp: 16:45

Detailed Brief

Gemma 4 Model Family: Sizes, Capabilities, and Positioning

  • Claims: Four models released: 2B and 4B for mobile/edge (multimodal input, text output); 26B MoE and 31B dense for servers/desktops.; E2B/E4B are 5B/9B total parameters but only 2B/4B 'effective' parameters in GPU memory (rest are token embeddings in separate memory).; 31B is 'really, really strong... can do basically anything from coding, energetic, everything, multilingual.'; Both larger models (26B, 31B) support vision + thinking + code execution simultaneously and are free to try on AI Studio.
  • Evidence: Gus: 'The E stands for effective... the model uses as much as 2B... but it's larger than that. The 2B is around 5B parameters... they are mapping tokens... you can leave them in other memory.'; LM Arena ranking: 4th (31B) and 7th (26B) among open models; competitors are 2–20× larger.; MoE architecture: '26 billion parameters, but it needs a space of a 4 billion parameter to do the work.'; Real-world deployments: Ukraine public services, Bulgarian national LLM (Gemma 2), Brazilian Portuguese (Gemma 3), MedGemma for hospitals.
  • Caveats: Not frontier intelligence—Gus: 'Is this the most intelligent model? No, it isn't.' Unsuitable for tasks like full systems architecture.; Multimodal output not mentioned; appears to be text-only generation.; Fine-tuning for specific languages now harder because base multilingual performance is so strong gains are marginal (1%).; No latency/throughput numbers provided; no discussion of quantization impact on quality.
  • Implications: For Ken's model selection: 31B is the 'daily driver' for coding, summarization, analysis where frontier intelligence is overkill. 26B MoE trades slight quality for 4× faster inference and lower memory. 2B/4B unlock on-device use cases (offline, privacy, zero marginal cost).; For cost modeling: 31B on 1 GPU vs competitors on 4–5 GPUs = 4–5× hardware cost reduction. API cost avoidance scales with token volume; break-even depends on traffic.; For sovereignty/compliance: Apache 2.0 + local deployment + MedGemma example = strong positioning for regulated industries (healthcare, finance, government). Can pitch 'your data never leaves your VPC/on-prem.'; For multilingual products: Base model likely sufficient for major languages without fine-tuning investment. Check LM Arena leaderboard for language-specific rankings (top 2–3 in many languages).

Ownership Economics: Token Costs, Energy Costs, and Agentic Workflows

  • Claims: Agentic workflows with high token generation (programming, research, multi-step reasoning) make API costs scale rapidly.; OpenRouter data shows programming tasks among highest for combined input + output token volume.; On-device models shift cost model from $/token to energy cost ($/kWh × device power × uptime) or sunken hardware capex.; Ownership allows control over latency, uptime, data residency, and customization (fine-tuning) impossible with proprietary APIs.
  • Evidence: Ian: 'Programming is right in the middle... among some of the highest tasks in terms of token generation.' State of AI report graph shown.; Ian: 'The more we have agents work... high token generation costs, that's when you start to get more benefit from being able to take control.'; Desktop demo: 26B MoE running on M4 Mac (48GB RAM) orchestrates parallel translation agents across 8+ languages locally.; On-device: 'We're not paying for these agents... within tokens... paying for them in terms of energy cost... utilization of NPUs... background task when they plug their phone in at night.'
  • Caveats: No quantitative break-even analysis provided (e.g., tokens/month where GPU rental < API costs).; Ownership incurs maintenance, uptime management, and infrastructure expertise costs not present with managed APIs.; Evaluation burden increases—must build task-specific evals; benchmarks are insufficient ('how good the model is depends on how well it does on your task').; On-device assumes tasks can tolerate background execution or variable latency; not all can.
  • Implications: For Ken's agent cost modeling: Measure actual token volume in pilot (e.g., multi-agent coding system). If >100k–1M tokens/hr sustained, local GPU likely breaks even vs API in weeks–months. If bursty/low-volume, API may remain cheaper due to zero capex.; For on-device strategy: Classify tasks by latency tolerance (real-time vs background) and data sensitivity (private vs non-sensitive). Private + background = ideal for on-device Gemma (email summarization, photo analysis). Real-time + complex = cloud API. Hybrid architecture optimizes cost + UX.; For enterprise sales: Ownership argument resonates with orgs that have sunken GPU costs (existing infra) or data residency mandates. Pitch: 'You already own the hardware; Gemma makes it useful for AI without per-token billing.'; For evaluation: Ian emphasizes 'drop it into existing workflows... see what the model can handle... bolster your evaluation suites.' Ken should build task-specific evals (code correctness, summary quality, function-call accuracy) and measure pass@k, latency, cost per task.

Sovereignty, Licensing, and Enterprise Adoption

  • Claims: Apache 2.0 licensing eliminates 18-month legal procurement cycles vs prior custom Gemma license.; Sovereignty = ownership + data residency + no vendor lock-in + offline operation + custom fine-tuning.; Real deployments: Ukraine (public services), Bulgaria (national LLM), Brazil (Portuguese fine-tune), MedGemma (hospital patient data on 1–2 GPUs).; Fine-tuning for languages now harder because base model is so strong; marginal gains (~1%) may not justify effort.
  • Evidence: Gus: 'If you have a custom license... lawyers will spend like 18 months doing procurement... That's why we moved to Apache 2.0.'; Gus: 'Ukraine used Gemma... Bulgarian LLM... based on Gemma 2... Brazilian version... based on Gemma 3.'; MedGemma: 'Specialized for medical use cases... operate on private data... deploy this to one or maybe two GPUs to run that for a whole hospital.'; Language fine-tuning: 'Model is pretty strong on those languages already... you might spend a lot of time to get 1%.'
  • Caveats: Apache 2.0 still requires legal review (indemnification, patents); not zero-friction, just faster than custom license.; No mention of Google enterprise support, SLAs, or deployment assistance beyond community/docs.; Fine-tuning cost/complexity not detailed; assumes user has ML engineering capability or partners.; Export controls, security certifications (FedRAMP, SOC 2), or compliance docs not discussed.
  • Implications: For Ken's B2B GTM: Apache 2.0 is a critical procurement unlock. Lead with 'industry-standard open-source license' in RFPs. Sovereignty pitch = data stays on-prem/in-VPC + no API dependency + model is yours forever + fine-tune to your domain.; For regulated industries: MedGemma example shows healthcare is addressable. Finance/legal similarly want data residency. Ken should build case studies/proof-of-concepts in these verticals (e.g., legal doc summarization on 31B, medical triage on MedGemma).; For multilingual products: If targeting major languages (Spanish, Portuguese, French, German, Chinese, etc.), base Gemma 4 likely sufficient without fine-tuning. Check LM Arena language leaderboard to confirm. Save fine-tuning budget for domain-specific terminology (legal, medical, financial jargon) vs general language fluency.; For partnerships: Google may co-market sovereign AI deployments (Ukraine, Bulgaria examples). Ken could pitch collaborative case studies to accelerate enterprise adoption in new regions/verticals.

Practical Deployment: Integration, Serving, and Evaluation

  • Claims: Integration is trivial: swap OpenAI SDK endpoint to OLLAMA/LM Studio, change model name—works with existing function-calling code.; Recommended workflow: (1) Drop into existing pipelines, (2) Measure task-specific performance, (3) Identify where Gemma suffices vs where frontier models needed.; Serving options: OLLAMA/LM Studio for local experimentation; vLLM/TGI/Ray for production. Mobile uses Google AI Edge Gallery app (iOS/Android).; Evaluation must be task-specific, not benchmark-driven: 'How good the model is depends on how well it does on your task and not anybody else's task.'
  • Evidence: Ian slide: 'Take any OpenAI compatible interface... point it at OLLAMA or LM Studio... pick out the Gemma model... that's all you need to change.'; Demo: LM Studio on M4 Mac runs 26B MoE orchestrating parallel translation agents, compiling to webpage. Runs at ~26GB RAM with 48GB unified memory.; Mobile demo: Google AI Edge Gallery shows 2B model doing function calling (calendar, maps, custom skills) with reasoning chains.; Ian: 'First thing we recommend... drop it into existing workflows... see what the model can handle... bolster your evaluation suites.'
  • Caveats: Serving complexity understated: production deployment (vLLM/TGI) requires expertise in quantization, batching, load balancing, monitoring not covered in talk.; No discussion of tokenizer compatibility, context window limits, or prompt engineering differences vs GPT-4/Claude.; Evaluation toolkit/framework not provided; assumes user builds custom evals.; Mobile demo on Google AI Edge Gallery is a playground, not production SDK—integration path for real apps not detailed.
  • Implications: For Ken's experimentation: Can trial Gemma in <1 hour by swapping OpenAI client to localhost:11434 (OLLAMA) or LM Studio endpoint. Run existing agent workflows, measure latency/quality/cost. Use to identify 'good enough' tasks (summarization, testing, refactoring) vs 'need frontier' tasks (architecture, complex reasoning).; For production deployment: Budget ML engineering time for vLLM/TGI setup, quantization testing (INT8/INT4), latency optimization (batch size, KV cache tuning), monitoring (Prometheus/Grafana). Not plug-and-play at scale; requires infra expertise.; For evaluation: Build task-specific evals (unit tests for code generation, ROUGE/BERTScore for summaries, function-call accuracy for agents). Measure pass@k (k=1,5,10), latency p50/p95/p99, cost per task. Compare Gemma vs GPT-4/Claude on Ken's tasks, not MMLU/HumanEval.; For mobile: Google AI Edge Gallery is proof-of-concept; real integration requires TensorFlow Lite or MediaPipe SDK. Evaluate latency/battery impact in real app before committing. Consider hybrid: simple tasks on-device, complex tasks cloud (Gemini API).

Notable Concepts & Terms

  • Mixture of Experts (MoE): 26B Gemma has 26B total parameters but only 4B activate per forward pass—reduces memory/compute while maintaining capacity. Enables single-GPU deployment where dense models need 4–5 GPUs.
  • Effective parameters (E2B/E4B): 2B/4B 'effective' means only transformer parameters count toward GPU memory; token embeddings stored separately. Total model is 5B/9B but runs in 2B/4B memory footprint, enabling phone deployment.
  • Sovereignty (AI context): User owns model + data never leaves infrastructure + no vendor lock-in + offline operation + fine-tuning control. Critical for governments, healthcare, finance. Enabled by Apache 2.0 + local deployment.
  • LM Arena / ELO ranking: Human preference benchmark (vs academic benchmarks like MMLU). Gemma 31B is 4th, 26B is 7th among open models. Preferred metric because 'how the model responds... that's how your customers will see it.'
  • Function calling / tool use: Model outputs structured JSON to invoke external APIs/tools (calendar, maps, search). Gemma 2B/4B support this on-device; enables multi-agent skill routing without cloud dependency.
  • Thinking (chain-of-thought reasoning): Model generates intermediate reasoning steps before final answer. Gemma supports this natively; improves accuracy on complex tasks. Shown in demos (function-call reasoning chains).
  • MedGemma: Medical-domain fine-tune of Gemma for healthcare use cases. Example: hospital deploys on 1–2 GPUs to process private patient data without sending to cloud. Demonstrates domain-specific fine-tuning value.
  • OLLAMA / LM Studio: Local model serving tools. OLLAMA is CLI/API server for Mac/Linux/Windows. LM Studio is GUI app. Both support OpenAI-compatible API, enabling drop-in replacement for experimentation.
  • Google AI Edge Gallery: iOS/Android app for experimenting with on-device Gemma models (2B/4B). Supports function calling, multimodal input, custom skill definition. Proof-of-concept for offline AI agents on phones.
  • Unified memory (Apple Silicon): M-series Macs share RAM between CPU/GPU, enabling large models (26B at 26GB) to run without discrete GPU. Ian's M4 Mac demo shows 48GB unified memory running 26B MoE.

Operator Notes / Why Ken Should Care

  • For agent systems: Gemma's function-calling + reasoning enables local multi-agent orchestration (demo: parallel translation agents). Key insight: route 'good enough' tasks (summarization, testing, refactoring) to local Gemma to cut API costs 5–10×; reserve GPT-4/Claude for complex reasoning. Integration is trivial (swap OpenAI endpoint), so A/B testing is <1hr.
  • For AI ops / infrastructure: 31B on 1 GPU vs competitors on 4–5 GPUs = 4–5× capex/opex reduction. MoE (26B) trades slight quality for 4× memory efficiency. Key decision: measure token volume in pilot workflows to calculate break-even (GPU rental vs API costs). If >100k–1M tokens/hr sustained, local deployment likely wins in weeks–months. If bursty, API may remain cheaper.
  • For content / business: On-device models (2B/4B) enable offline-first apps + user data privacy (zero cloud exfiltration) + zero marginal inference cost. Ideal for batch background tasks (email summarization, photo analysis, local search) during device charging. Hybrid architecture: private data on-device, complex reasoning cloud. Evaluate latency tolerance (real-time vs background) when designing UX.
  • For investing / GTM: Apache 2.0 licensing is critical unlock for enterprise sales (18-month procurement → weeks). Sovereignty pitch resonates with regulated industries (healthcare, finance, government) and data-sensitive orgs. MedGemma case study shows healthcare is addressable; Ken should build PoCs in finance/legal. Multilingual base model quality reduces fine-tuning need/cost for international GTM.
  • For workflow automation: Gemma excels at instruction-following in 'small modular bits' (Ian's words)—refactoring, test generation, documentation, code review. Not suitable for full systems architecture. Workflow: decompose complex tasks into Gemma-solvable subtasks (summaries, transformations, validations) and frontier-model checkpoints (planning, complex reasoning). Measure pass@k on task-specific evals, not benchmarks.
  • Strategic insight: The 'escape velocity' in the title refers to crossing capability/cost/hardware thresholds where open models become economically and technically viable for production. Gemma 4 crosses these for: (1) on-device agents (2B/4B on phones), (2) single-GPU server deployment (26B/31B), (3) enterprise sovereignty (Apache 2.0), (4) agentic token economics (local >> API for high-volume). Ken should evaluate which thresholds his use cases cross and model TCO accordingly.

Watch Map

  • 00:00: Introduction: Gus Martins & Ian Ballantyne, Google DeepMind. Context-setting: why open models (Gemma) complement proprietary (Gemini).
  • 02:15: Gemma 4 model family overview: 2B/4B (mobile/edge, multimodal input), 26B MoE, 31B dense (server/desktop). Effective parameters explained.
  • 04:30: LM Arena rankings: 31B is 4th, 26B is 7th among open models; 2–20× smaller than competitors. Human preference benchmark rationale.
  • 06:15: Performance per parameter: 31B on 1 GPU vs competitors on 4–5 GPUs (200GB memory). AI Studio demo mentioned (vision + thinking + code execution).
  • 08:45: Sovereignty & licensing: Apache 2.0 eliminates 18-month legal cycles. Examples: Ukraine, Bulgaria, Brazil, MedGemma. Fine-tuning now harder due to strong base multilingual performance.
  • 11:20: Ian takes over: Ownership economics. OpenRouter data on programming tasks = high token volume. Agentic workflows amplify API costs. Cost model shifts: tokens → energy/hardware.
  • 13:00: Thresholds framework: capability, hardware fit, latency, cost. When to own vs API? Sunken hardware costs vs marginal API costs. Task-specific evaluation critical.
  • 14:30: Mobile demo: Google AI Edge Gallery app (iOS/Android). 2B model with function calling (calendar, maps, custom skills), reasoning chains, multimodal input. Energy cost model explained.
  • 16:00: Desktop demo setup: LM Studio on M4 Mac (48GB unified memory), 26B MoE at ~26GB RAM. Parallel translation agents demo begins.
  • 16:45: Integration slide: OpenAI SDK → OLLAMA/LM Studio endpoint swap. Recommendation: drop into existing workflows, measure task performance, build task-specific evals.
  • 17:30: Translation demo executes: 8+ languages generated in parallel, compiled to webpage. Demonstrates local multi-agent orchestration.
  • 18:30: Serving considerations: capex vs opex, maintenance, uptime, infrastructure complexity (mobile: accelerators, RAM; enterprise: GPU hosting). Summary: experiment on tasks, evaluate, scale.
  • 19:45: Q&A / closing. More details in tomorrow's keynote (Omar) and Cassidy's talk.

Source/Metadata

  • Title: Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind
  • Transcript words: 5179
  • Duration seconds: 1252
  • Timestamp note: Timestamps estimated from ~20-minute duration (1252 seconds = 20:52). Transcript includes some repetition/rewinding; timestamps approximate key sections. No official chapter markers in transcript.
Full transcript 3820 words · 23 min read
0:15

SPEAKER_00

Hi everyone, can you hear me? Yes, you can hear me. Hi, sorry, one minute late. I'll try to do my best to finish earlier so my friend can do a pretty cool demo for you. I'm Gus, this is Ian, we are from Google DeepMind and I'm specifically working on the Gemma product. Do any of you know what Gemma models are? Okay, perfect, perfect. Thank you very much. So today we're going to talk a little bit about ownership and open models and, well, you know who we are, but the idea is last Thursday we released our new family of models, Gemma 4, and I'm going to talk a little bit about them. There's going to be more information tomorrow in the keynote by Omar and there's another talk by Cassidy also tomorrow that she'll go into even more details. We are going to tell a little bit of the story, but the story is a little bit bigger. We'll try our best here.

0:21

SPEAKER_00

So why does it matter? If you ask me, I work for Google, of course, if you ask me which is the best model for you to try, the easiest one I will answer for you, Gemini. Gemini is the best model we have, pretty strong, multimodal, can do all kinds of things. But then there is more to this story than just having the strongest model possible. In some situations, you want to own the model. You want to be able to run on your own hardware. You want to customize it. You want to be able to send your proprietary data that cannot leave your infrastructure. So there are many situations where even the best proprietary model will not be able to help you directly. That's when you might need an open model. That's where Gemma comes in. So when you think, why does Google have two family of models? Because they complement each other. So Gemini is the most intelligent one, can do a lot of cool stuff, but it's hosted in Google servers. You need the API to access. If you need more control and access, you need an open model. That's why we have Gemma. That's why, and we are very proud that the quality is very, very strong. We're going to go into some details later. But the idea is you would be able to do a lot of cool stuff with it.

0:26

SPEAKER_00

Among the launches, we released four sizes. Two are targeted to mobile or IoT or smaller devices. It's an E2B and an E4B. These names are a little bit weird. We are the only ones that use this name. And the E stands for effective. And the idea here is the model uses as much as 2B, what a 2B model would use as memory, but it's larger than that. The 2B is around 5B parameters. But then you say, oh, but where is this other 3B in memory? The fun fact is that they are not really parameters from the transformers. They are mapping tokens. So you can leave them in other memory. So what you really need on your GPU memory is the 2 billion or the 4 billion. Why do we do that? So that you can run these models on a phone, on a Pixel phone or any phone you have there. You can run these models and they're very strong. The E2B and E4B, both of them have text, vision and audio input, and they do only text output. They can do thinking. They can do coding, function calling, all these kinds of cool things. These all run on your phone right now. You could download it right now, right? We also have two other models which are the larger ones. We have a 26B and a 31B. The 26 is a mixture of experts, which means that it's as if we had many other models working together where each one of those are like a 4B model. Why does it matter? Because it has 26 billion parameters, but it needs a space of a 4 billion parameter to do the work. And this makes it accessible to way more hardware, to way more people, and it's still pretty strong. But our strongest model is the 31B Dense, which is 31B LLaMA parameters model. And this is really, really strong. If we look into our ELO score on LM Arena, you can see that both our models are fourth and seventh as the leads on open source models, open models. And if you compare them to maybe the top 20, 30, all of them are at least twice, three times larger than our models. In some cases, 20 times larger. So we are talking about a disproportionate amount of intelligence per size. So our 31B model is the one I use very regularly. It can do basically anything from coding, energetic, everything, multilingual, all of that. So I strongly recommend you try those. They are so strong that they are, both of them are really good to use on your, as a cloud deployed model. They can run on your desktop, but if you use on your server as your endpoint to do your work, they are pretty good. And you ask, oh, is this the most intelligent model? No, it isn't. I'm very biased and I love them, but I know the capabilities. But the question is, do you need the most intelligent model on the planet to summarize your email, to do some manual tests, to help you code, to do some agent capabilities that are searching and interacting with docs? Probably not. That's why these models are so strong, because they're cheaper. They're very strong, but they're cheaper to run. They require way less hardware. A 31B running one GPU. The competitors need 200 gigabytes of memory, which would be maybe four or five GPUs. So you can see that the price here is really, really different. One easy place for you to try these models is on AI Studio, where you can try Gemma models, Gemini, all the other ones, but Gemma are there. Both 26 and 31B, you can try right now. They're free. You can play with it. And they can do some cool stuff, which is vision plus thinking plus code execution all at the same time, right? I'll try to post something about this later, but the idea is you can play with the models pretty easily. And right there, not now, let's finish the talk and then you'll play with it.

0:32

SPEAKER_00

And as I was saying, the intelligence per parameter that these models bring is pretty good. It's very, very strong. And if we use the ELO score for LM Arena, because it's a benchmark that's a person's preference, right? We can look into academic benchmarks. They are very, very strong. But how the model responds to your queries, that's very important, right? That's how your customers will see, how you will see and interact with it. So this is why this is so important. And why does all this matter? One of the reasons that we care so much is because you want the user to have ownership. And more than that, we are enabling sovereignty. And sovereignty means in terms of you own the model and you can adapt your use cases and you are not susceptible to,

0:40

SPEAKER_00

very strong. And if we use the ELO score for LA Marina, because it's a benchmark that's a person's preference. Right? We can look into academic benchmarks. They are very, very strong. But how the model responds to your queries, that's very important. Right? That's how your customers will see, how you will see and interact with it. So this is why this is so important. And why does all this matter? One of the reasons that we care so much is because you want the user to have ownership. And more than that, we are enabling sovereignty. And sovereignty means in terms of you own the model and you can adapt your use cases and you are not susceptible to,

1:19

SPEAKER_00

I don't know, loss of service or for some kind of someone saying, no, no, you cannot use this model anymore. It's all available to you. And one of the changes we made last year until Gemma 3 and others, we had our specific license, a Gemma license, which is pretty good, commercial friendly and all. But there's a problem. If you have a custom license, I don't know if you have any lawyers here. If I tell you, oh, we have this custom license, your lawyers will look at me with that face that I hate you guys. And then they will spend like 18 months doing procurement process to understand the

1:59

SPEAKER_00

license and trying to change. And that never works. So it's pretty hard for sovereign institutions to adopt this kind of thing. That's why we moved to a past 2.0 for Gemma 4 and going forward. And that makes, thank you, and that makes our life, your life much easier to convince your legal department, let's say like that, that look, we own this model we can use. So this is pretty important and it enables many, many sovereignty institutions to use our models. We have some examples. For example, Ukraine used Gemma to, in parts of their services. We have one version of the Gemma model that was

2:39

SPEAKER_00

fine-tuned for Bulgarian. It was their LLM for the country. That was based on Gemma 2. We are working to make sure they use Gemma 4 now. We also have a Brazilian version that is based on Gemma 3, was fine-tuned for Portuguese. And the challenge of these models today is that they, if you want to fine-tune Gemma 2, a specific language, it's becoming very hard to do that. And the problem is hard because not the tooling or anything, it's because the model is pretty strong on those languages already. So any gains you try to have, you might not get there. So you might spend a lot of time to get 1%. And then

3:24

SPEAKER_00

maybe, I don't know if it's the best use of your time. So this is good and bad at the same time, because, but it's good that you can automatically use in many languages, you can try right now. And if you're going to the LLM Arena for languages, in many languages are top two, three, and look, it's a 31B model. It's very, very small, right? So this is pretty good. That being said, I will let my colleague continue and show some demos. Thank you, Gus. So one thing that I think is really important about these models is that when you think about using open models, you think about using proprietary models,

4:06

SPEAKER_00

we're moving, we're seeing a shift now to more agentic capabilities and the kind of tasks that we're trying to do. And with that comes a cost in tokens and token generation. So one benefit of taking ownership of the models is your ability to control or in cases where you have sunken hardware cost to be able to iterate on top of that. This graph on the right hand side is from the state of AI report that OpenRouter did. And it shows the bits more for you on this diagram, but have a look at that link. It shows the different types of tasks that people are doing through OpenRouter at the moment.

4:46

SPEAKER_00

And you'll see the one that's about here, this one here is programming is right in the middle. And this is among some of the highest tasks in terms of token generation, both input and output combined. So the more we have agents work and do these kinds of tasks for us that have very high token generation costs, that's when you start to get more benefit from being able to take control of that in itself. So if, for instance, you have a laptop that is capable of doing a particular task that you need to be doing, like processing a document or analyzing some data or doing some research or in the cases

5:23

SPEAKER_00

Gus talked about doing some coding that's suitable for that, then you have a GPU that you can take advantage to do some of that stuff. Now, similar to what Gus said about, we don't necessarily still have frontier models for doing the best possible things. I wouldn't get this model to do a full systems architecture and redesign of your application, right? It's not for that. But what it is very good at doing is following very specific instructions about doing things like refactoring, analyzing, generating code in small modular bits. And you can offload a chunk of work in that style to these kinds of models to be able to do that, whether it's on a single GPU or in your own

6:08

SPEAKER_00

personal hardware. And the way that we think about this is like a set of thresholds. Like, when do we get to the point where these models are capable of doing the task, but then they also fit on the right hardware, that they also can do it with the right amount of latency, depending on the, if it's a task for a user, needs to happen in a couple seconds. If it's a task where you're doing things like batch processing, you maybe have slightly different thresholds for what needs to be done. And then also what the cost of actually doing that is. So if you have a sunken cost in terms of

6:44

SPEAKER_00

infrastructure that you already own, or that you're prepared to outlay, and then operating on that, or whether you're leasing GPU time or something else. So these are going to be very specific to the tasks that you're trying to achieve. But what you can do with open models is you can think very carefully about what, which of these tasks can I fully offload or can I fully own, compared to relying just on using the best possible models to do that in the cloud. And an example, so Gus talked about the different types of hardware that can run these things now. I'm just going to run

7:24

SPEAKER_00

this little demo in the side at the moment. So we now have models that will work directly on mobile and edge devices. This example here was built by Cormac's team is a set of agent skills that the model is running on a phone. So I'm going to mute the microphone for that. So you can talk to the model, you can show it images, you can show it the world around you, and you can prompt it and chat to it. And what this one is showing is that it can look through a set of skills that it has about things on the phone. So either it can take actions on the device itself, so trigger other applications,

7:57

SPEAKER_00

like trigger calendar apps, trigger maps apps, or you can define your own skill sets. And what's different now with the Gemma 4 models than we saw for the previous generation is that it's edge devices. This example here was built by Cormac's team is a set of agent skills that the model is running on a phone. So I'm going to mute the microphone for that. So you can talk to the model, you can show it images, you can show it the world around you, and you can prompt it and chat to it. And what this one is showing is that it can look through a set of skills that it has about things on the phone. So either it can take actions on the device itself, so trigger other applications,

8:39

SPEAKER_00

like trigger calendar apps, trigger maps apps, or you can define your own skill sets. And what's different now with the Gemma 4 models than we saw for the previous generation is that it's able to reason about what actions it needs to take and reliably make those function calls defined. So what this app will allow you to do is it acts as a playground. So this is Google AI Edge Gallery, and you can find it on iOS and Android, and you can experiment to see what the models of this size are actually able to do. So I think this is the 2 billion parameter model, but there's also the 4 billion parameter model depending on the size of your hardware.

9:20

SPEAKER_00

And when we get to desktops and single GPUs, as Gus mentioned, that's where you can use the 26 and the 31B models, again, on your local hardware, and I'll show you how to do that in a minute. But the key point here is that whereas we're not paying for these agents or models within tokens, we're actually paying for them in terms of energy cost, if we think about it. Because now you're thinking about utilization of GPUs, you're thinking about utilization of NPUs on the hardware itself. When are you going to do these tasks? Does the user need to get a response right now when you're taking a picture of something, or is it something that you

10:03

SPEAKER_00

can process offline as a background task when they plug their phone in at night? So what I'm trying to say here is that the thresholds and how you think about the usage of these models shifts when you come to on-device or ownership, because you think more about how they're being executed and why they're being executed. Yeah, perfect. And similarly, on the enterprise side, if you don't have a piece of hardware that can run the 31 billion parameter model, you can now be thinking about scaling that down. So maybe if you wanted to use a 300 plus billion parameter model before, you might have

10:39

SPEAKER_00

needed multiple GPUs. Now you can think about using a single H100 or A100, or even in some cases like an L4. And then the costs obviously related to that also go down. So again, it's a calculation that you'll have to do depending on your use cases. But there are ways that you could scale, for instance, running one of these models to serve a small team or to serve a company, depending on what you're trying to do. And the final point is that you also have the fine tuning component too, which is that because these models can be customized, you can deploy your own version of it. So for

11:12

SPEAKER_00

instance, we have a variant of Gemma models called MedGemma, which is specialized for medical use cases. So if you wanted to have something that would operate on private data that you can control yourself, you can now feasibly deploy this to one or maybe two GPUs to run that for a whole hospital, for instance. So these are worth considering for the enterprise case. I'm going to jump straight to demos now. I've shown you some demos on the phone. I'm going to show you a quick demo here. Quick show of hands, who's ever used a tool called LM Studio? Okay, just under half people. So LM Studio is a way that you can play around with local models.

11:53

SPEAKER_00

And I have here, this is the 26B model. So this is our faster of the two larger models with four billion activated parameters. And I'm at the moment, including the context, probably about 26 gigabytes in RAM. And this is an M4 Mac. So I've got unified memory. I've got up to about 48 gigabytes. So I can run it on this machine. And I'm just going to run this terminal right here. Let's give that a go. Oops. Pre-showing my demo. Let's try that again. Okay, so I'm just going to run a little process where I'm going to do some quick translation on my device. So what it's going to do is I've got an orchestrator on this side here,

12:30

SPEAKER_00

which is going to hopefully kick off my agent in a minute. Let's make sure we are loaded. Let's see what LM Studio is doing. Yeah, it's just processing at the moment. And then it's going to farm out this translation to all of these different windows. And each one of them represents a different sub-agent. So this is running on my device. And it's going to basically execute all these translations in one go. So I've given it the Gemma 4 announcement. And I just want to translate to all these different languages. So you'll see in a second, it should hopefully send it over there. Three, two, one. And hopefully we should be generating translations in a second. There we go.

13:14

SPEAKER_00

So you can imagine doing any kind of generic task on your local machine. You could have it processing files. You could have it doing additional analysis. And hopefully what you'll see in a minute is it will be able to compile all these back. And then it will generate me a quick web page. And then you can see the results of your translation. There you go. So there's the multilinguality of the model there as well. Thank you. Right. So in the interest of time, I just want to say that the main next step for exploring and trying out these models is as simple as this code on the right-hand side. You can take any OpenAI

13:51

SPEAKER_00

compatible interface that you've got, and you can point it at a service like OLLAMA or LM Studio. And you can just pick out the Gemma model. And that's all you need to change code-wise to at least try it out. So the first thing we recommend you do is to drop it into existing workflows that you have to then see what the model can handle. What is it working well at? What would it need tuning for? What is out of its depth in terms of the complexity of the task? Next is to bolster your evaluation suites because benchmarks are great for just saying what general capabilities are. But the reality is that how good the model is

14:27

SPEAKER_00

depends on how well it does on your task and not anybody else's task. The other thing I mentioned very briefly is thinking about how you actually serve these models in the end. So if you need to run your own GPU and you need to host it, yes, you're in control of uptime and downtime, but then there's maintenance costs and so on. So you have to consider that as one of the factors, the ongoing costs as well, as well as any upfront capex costs if you buy infrastructure or hardware to do that too.

15:00

SPEAKER_00

for just saying what general capabilities are. But the reality is that how good the model is depends on how well it does on your task and not anybody else's task. The other thing I mentioned very briefly is thinking about how you actually serve these models in the end. So if you need to run your own GPU and you need to host it, yes, you're in control of uptime and downtime, but then there's maintenance costs and stuff like that. So you have to consider that as one of the factors, like the ongoing costs as well, as well as any upfront capex costs if you buy infrastructure or hardware to do that too. On mobile devices, for instance, you have to think about if I'm going to offload stuff to a phone, what am I supporting? What accelerators do they have? What size RAM do they have? So the conversation becomes a little bit more complex, but then there's a whole heap of things you can unlock like working offline or working on users' private data that never leaves their device. And finally, if you want to scale this up to enterprise levels, you have to think again about the kind of infrastructure that you're running on and what the ongoing costs are of that as well. But it does unlock that. So with that, the summary is that you can use these models in pretty much any way you can think about, experiment what kind of tasks are possible with it, use some of the benchmarks to give you an indication of what's feasible. But really, we want to hear your feedback and how you get on with these and how you fine tune them and what kind of things you run into. And we want to help you on that journey as well. So with that, thank you very much.

15:06

SPEAKER_00

[SPEAKER_01] Thank you. [SPEAKER_01] Thank you. you'll have to do depending on your use cases. But there's ways that you could scale, for instance, running one of these models to serve, you know, a small team or to serve a company, depending on what you're trying to do. And the final point is that you also have the fine tuning component too, which is that because these models can be customized, you can deploy your own version of it. So for instance, we have a variant of Gemma models called MedGemma, which is specialized for medical use cases. So if you wanted to have something that would operate on private data that you can control yourself,

15:50

SPEAKER_00

you can now feasibly deploy this to like one or maybe two GPUs to run that for, I don't know, like a whole hospital, for instance. So these are kind of worth considering for the enterprise case. I'm going to jump straight to demos now. I've shown you some demos on the phone. I'm going to show you a quick demo here. Quick show of hands, who's ever used a tool called LM Studio? Okay, just under half people. So LM Studio is a way that you can play around with local models. And I have here, I have, this is the 26B model. So this is our faster of the two larger models with four billion activated parameters. And I'm at the moment, including the context is probably about 26

16:35

SPEAKER_00

gigabytes in RAM. And this is an M4 Mac. So I've got unified memory. I've got up to about 48 gigabytes. So I can run it on this machine. And I'm just going to run this terminal right here. Let's give that a go. Oops. Pre-showing my demo. Let's try that again. Okay, so I'm just going to run a little process where I'm going to do some quick, a trick, quick translation on my device. So what it's going to do is I've got an orchestrator on this side here, which is going to hopefully kick off my agent in a minute. Let's make sure we are loaded. Let's see what LM Studio is doing. Yeah, it's just processing at the moment. And then it's going to farm out

17:20

SPEAKER_00

this translation to all of these different windows. And each one of them represents a different sub-agent. So this is running on my device. And it's going to basically execute all these translations in one go. So I've given it like the Gemma 4 announcement. And I just want to translate to all these different languages. So you'll see in a second, it should hopefully send it over there. Three, two, one. And hopefully we should be generating translations in a second. There we go. So you can imagine doing any kind of a genetic task on your local machine. You could have it like

17:55

SPEAKER_00

processing files. You could have it doing additional analysis. And hopefully what you'll see in a minute is it will be able to compile all these back. And then it will generate me a quick web page. And then you can see the results of your translation. There you go. So there's the multilinguality of the model there as well. Thank you.

18:20

SPEAKER_00

Right. So in the interest of time, I just want to say that the main next step for exploring and trying out these models is as simple as this code on the right-hand side. You can take any open AI compatible interface that you've got, and you can point it at a service like OLAMA or LM Studio. And you can just pick out the Gemma model. And that's all you need to change code-wise to at least try it out. So the first thing we recommend you do is to drop it into existing workflows that you have to then see what the model can handle. Like what is it working well at? What would it need tuning for? What is kind of out of its depth in terms of like the complexity of the task?

19:01

SPEAKER_00

Next is to kind of bolster your evaluation suites because, you know, benchmarks are great and everything for just saying what general capabilities are. But the reality is that how good the model is depends on how well it does on your task and not anybody else's task. The other thing I mentioned very briefly is thinking about how you actually serve these models in the end. So if you need to run your own GPU and you need to host it, yes, you're in control of like uptime and downtime, but then there's like maintenance costs and stuff like that. So you have to be, you have to consider that as like one of the factors, like the ongoing costs as well,

19:33

SPEAKER_00

as well as any upfront capex costs if you buy infrastructure or hardware to do that too. On mobile devices, for instance, you have to think about like, if I'm going to offload stuff to a phone, like what am I supporting? What accelerators do they have? What size RAM do they have? So the conversation becomes a little bit more complex, but then there's a whole heap of things you can unlock like working offline or working on users' private data that never leaves their device. And finally, if you want to scale this up to enterprise levels, you have to think then again about like the kind of infrastructure that you're running on and

20:05

SPEAKER_01

what the ongoing costs are of that as well. But it does kind of unlock that. So with that, the summary is that you can use these models in pretty much any way you can think about, experiment what kind of tasks are possible with it, use some of the benchmarks to kind of give you an indication of like what's feasible. But really, we want to hear your feedback and how you get on with these and how you fine tune them and what kind of things you run into. And we want to help you on that journey as well. So with that, thank you very much. Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note