How Google DeepMind Runs Agents at Scale — KP Sawhney & Ian Ballantyne, Google DeepMind
Description
Google DeepMind employees have worse token quotas than paying customers. That is not a mistake. KP Sawhney explains: customers get priority, and if an internal team spikes usage on a cluster someone monitoring 24/7 will just call and ask them to stop. This panel covers how DeepMind thinks about agents at scale from the inside: managing quota across thousands of power users, a Darwinian skills library where only the strongest skills survive as engineers contribute en masse, and where the Deep Research pipeline is going next. KP's current focus is replacing the pipeline's giant context blobs with a shared file system so each research component can collaborate the way human researchers would, and produce supporting artifacts that the current architecture cannot. Speaker info: - https://linkedin.com/in/ianballantyne - https://github.com/irbg - https://linkedin.com/in/kyle-sawhney - https://github.com/KPSawhney Timestamps: 0:00 Introduction of KP Sawhney and Ian Ballantyne 0:57 Demo of the anti-gravity agentic platform 4:32 Discussion on KP's work with the deep research agent 5:46 Using skills libraries and managing agent sprawl 7:52 Addressing scalability and token quotas at Google 9:44 Audience Q&A: Managing per-user agent behaviors 13:12 Audience Q&A: Observability and agent trajectory stores 14:46 Audience Q&A: Future of deep research pipelines 16:10 Audience Q&A: Handling multi-agent systems 17:45 Audience Q&A: Perspectives on skills vs. MCP 19:12 Audience Q&A: Evaluating agentic workflows 20:25 Audience Q&A: Handling model limits and quota management 22:48 Audience Q&A: Automated code review processes
Summary
Generated by claude-sonnet-4-530-second take
This is a live demo/panel with two DeepMind engineers (Ian Ballantine, DevRel; KP Sawhney, AI platform team) showing how Google runs agents internally using a tool called "Anti-Gravity"—a VSCode-like IDE with a built-in multi-agent orchestration platform. The core thesis: DeepMind is moving from passing huge context blobs between agents to having agents collaborate via shared file systems like digital coworkers. They're deeply focused on token efficiency, quota management at Google scale, skills-based architectures (preferring skills over MCP), and custom observability/eval infrastructure. The talk is light on architectural depth but reveals real operational constraints: power users hit quota walls, agents are expensive to run, and DeepMind is building custom web UIs to trace agent trajectories across massive internal monorepos.
Key takes
- Anti-Gravity is DeepMind's internal agentic IDE: Multi-agent manager with integrated browser control, DOM inspection, planning systems, and human-in-the-loop feedback. Agents can spawn tasks, capture videos/screenshots, and run in parallel on different project tracks. DeepMind uses this for coding at monorepo scale.
- Deep Research is being refactored to use shared workspaces: Current implementation passes "huge blobs of text" through the pipeline, consuming context and costing too much. They're redesigning it so agents collaborate via a shared file system (like human researchers would), enabling artifacts like infographics and reducing token overhead.
- Token consumption and quota management are the primary scaling bottleneck: DeepMind brute-forces this with per-user/per-team quotas. Power users get cut off manually. They're exploring auto-fallback from Pro → Flash → local models (like Gemma) when quota runs out, to avoid workflow interruptions. Subscription models (à la Anthropic blocking OpenClaw) don't work for token-hungry agents.
- Skills > MCP (controversial internal take): KP strongly prefers "skills + guard-railed CLI interactions" over Model Context Protocol. Skills scale well in large orgs because domain experts contribute them, and all users/agents inherit that expertise for free. MCP's auth layer is good, but skills have won internally at DeepMind.
- Custom observability and eval infrastructure: They built a custom web app to drill down into agent hierarchies, inspect raw predict requests, and trace where looping/failure happened. They also have an "agent trajectory store" for coding tasks. Eval remains hard—onus is on skill authors to write tests, but they're experimenting with agents generating evals (meta but real).
- Auto-review models for PRs at Google scale: Fine-tuned per-language models trained on style guides + prior good code. Product teams add custom prompts. Agents are now submitting PRs, compounding the review load. Tools like Jules (GitHub web interface) help, but the "trillions of lines" agents generate are creating a new bottleneck.
Useful details
- Anti-Gravity features: Agent manager, chat panel, planning system, browser automation, DOM inspection, video/screenshot capture, human-in-the-loop editing, scratchpad notes visible during execution.
- DeepMind has a "gigantic monorepo" and a "huge library of skills" with Darwinian selection (only best survive).
- They use mock TPUs to test agentic flows without burning TPU hours during eval.
- Internal limits: Googlers have worse quota limits than customers (customers prioritized). Ian had to click 100x to continue during the demo.
- Specific models mentioned: Gemini Flash (fallback), Gemma (local/free), Pro, Ultra. Flash recommended for poor connectivity.
- One audience member is building memory-as-a-harness with ontology graphs, handling ~3 million context graphs/month across multiple hardware instances.
- Pricing concern: Agentic systems break subscription models (e.g., Anthropic blocking open Claw). Future pricing likely usage-based.
- Google has 24/7 monitoring teams that manually tell engineers to shut down runaway jobs on clusters.
- Announcement tease: "Exciting things you'll learn more about at IO in a few months" (no details given).
Caveats / counterpoints
- Very little technical architecture detail: No specifics on how Anti-Gravity orchestrates sub-agents, how many levels of hierarchy exist, or how agent-to-agent communication works under the hood. Ian admits he doesn't know the exact agent nesting depth.
- Demo was flaky: Connectivity issues, model errors, agent didn't successfully play the game or fully complete the task on screen. The live demo didn't validate the tool's reliability.
- MCP vs. skills debate is one engineer's opinion: KP says MCP "may be a flash in the pan" but acknowledges DeepMind supports both. This isn't company policy, just his preference.
- Eval is unsolved: "Evaling this stuff is really hard." Mechanical setup is complex, dataset creation is manual, and skill authors are responsible for tests. Agent-generated evals are "a little bit meta" (i.e., not production-ready).
- No clear answer on interruption/looping detection: They have observability to diagnose where looping started, but no mention of auto-recovery or guardrails to prevent it mid-run.
- Vague on public availability: Ian says "we support what the community uses," but no timeline or commitment on releasing Anti-Gravity or Deep Research's new architecture publicly.
Ken relevance
High relevance for agent ops and GTM strategy:
- Token economics matter more than you think: If DeepMind's power users are hitting quota walls and the company has to brute-force limits, this is a real cost/pricing moat for agent platforms. Ken should model token burn rates for any agent product and plan fallback tiers (premium → mid → local) before users hit walls mid-workflow.
- Skills-based architectures scale better than tool protocols in orgs: If you're building agent systems for enterprises, prioritize curated skill libraries with expert contributions over open-ended MCP integrations. DeepMind's "Darwinian" selection model (best skills survive) is a forcing function Ken could replicate with usage analytics + quality gates.
- Shared file system > context passing: Deep Research's redesign (agents collaborating via shared workspace instead of passing blobs) is a pattern Ken should adopt for multi-agent pipelines. Reduces token overhead, enables richer artifacts, and mimics human team collaboration.
- Observability is a build-it-yourself problem: No off-the-shelf solution at DeepMind scale. If Ken's agent systems get complex, budget for custom trajectory tracing and per-step inspection UIs early.
- Auto-fallback models are table-stakes UX: Users expect agents to keep working when quota runs out, not get a "limit reached" error. Ken's agent products should auto-switch models (or notify + pause) to avoid breaking user flow.
Investing angle: Agent infrastructure companies solving quota management, trajectory observability, or skills marketplaces are solving real DeepMind-scale pain. Token burn rate transparency will be a wedge for enterprise agent platforms.
Watch verdict
Skim. Useful for understanding DeepMind's operational constraints (quota walls, token costs, skills > MCP preference, shared workspace pattern for Deep Research) but light on technical architecture. The demo was flakey and didn't land. Skip if you need code-level implementation details; watch if you're designing agent systems for scale or enterprise GTM and want to hear what bottlenecks actually matter at Google.
Transcript
[SPEAKER_00] Hello, everybody. So my name is Ian Ballantine. I'm a developer relations engineer at Google DeepMind. [SPEAKER_03] Hi, folks. I'm KP Sorney, software engineer in DeepMind's AI platform team. [SPEAKER_00] And we're going to do an agentic panel today to talk about how DeepMind thinks about agentic software and how we build our own stacks. [SPEAKER_00] We're going to start very briefly by just showing a quick demo. And then we'll go into a discussion about some of the stuff that KP works on. And hopefully, you can ask lots of questions and find out all the things you want to know about how DeepMind and Google think about agents. And yeah, let us know your thoughts as well. So quick show of hands. Who's actually used antigravity before as a tool? OK. So four or five people. I just want to show you one quick thing, because one of the things that we have in antigravity is a lot of people know that it's a Visual Studio style interface, but they don't actually know that it has a whole agent manager and agent manager framework behind it. So you can run and spawn multiple agents working on different projects. And it's integrated into the IDE, but it's also an agentic platform in itself. So you can do things. I've got this project here. And you can have this chat panel on the side. And I can just say, build an example of this spec. And I can give it a particular file. Let's do the spec file. Let's just implement that one. And I'll send that off to Flash. Oh, experience errors right now. Let's restart. Should we try that? I am connected to the Wi-Fi. I'm pretty sure I am. Yes. And if you're not connected to the Wi-Fi, you should use the Gemma models instead. That's my plug. So let's start that up again. And hopefully I can do the same thing. Build. Build the spec. Marked here. There you go. Let's see whether we get it this time. OK, there you go. So, again, the model will just go away thinking on the side there. And hopefully what it should do is it should have a look at the spec. It should analyze it and it should use a bunch of its own internal tools to decide how to do that. It's actually found that there are already existing files that I have built. It might even tell me that it's already done it, which is interesting. But the tool itself has built-in to-dos. It has a planning system. And what it's just done here is it's thrown up a browser so it can actually check how the applications run. And this is controllable by anti-gravity itself. So if I just do that, if I give it two seconds, it should take control over the browser and it should actually try and run. So what it's probably doing now is it's probably looking at the implementation already to see what's actually already done. And it's going to then analyze it. Let's just have a quick look, see what it's actually done there. Let's try that again. It'll pick up. So it should have snag. So it wants to try and analyze it. Oh, one thing that it can do is it can actually inspect the DOM. So it can look at the actual web page itself as part of that and feed that back. And when it finally finishes doing the task, it will give you a report at the end as to what it was able to achieve and what it implemented. So you can then review. It can also capture, for instance, a screenshot or a video. So if it's an interaction, if you say add this feature to my web page, it will go through and then actually try and run it. I don't know how good it is at playing games, but we're about to find out. It's either going to try and play the game or let's have a look. See what it's doing at the moment. It hasn't loaded it yet. Let's go back. Okay, we've got an implementation plan here. So this is what it thinks it should change about the file. So what you can do is you can just go in and edit a line and say, actually, I want this different behavior or that's not what I meant at all. So it gives you that human in the loop feedback. And then when you're done, you just say proceed. And then it will go away and do that. While we're waiting for that to happen, do you want to say a little bit about what you work on and how this relates to your data? What's the day job? [SPEAKER_03] Yeah, sure. [SPEAKER_03] So one of the things I worked on a few months ago was the deep research agent, which is now available via the interactions API. [SPEAKER_03] And that's been great. [SPEAKER_03] But as we continue to iterate on that, my focus has turned now to making best use of this anti-gravity harness internally. [SPEAKER_03] And that applies to scaling up for all of the coding we're doing. [SPEAKER_03] And we have a gigantic monorepo. [SPEAKER_03] So it's pretty complex. [SPEAKER_03] But now starting to think about how we generalize this to a variety of other use cases. [SPEAKER_03] So potentially deep research itself. Rather than passing around huge blobs of text from the searches that have been done, why not have the different parts of that pipeline collaborate in a shared file system? [SPEAKER_03] And so, that's really been the focus for me. Really tightening up this harness and making it excellent at not just coding, but a variety of other tasks, too. And do you have any interesting use cases for how people within Google or DeepMind are using it at the moment? [SPEAKER_00] What kind of things are they doing with agents in DeepMind? Yeah. So there's a huge amount that's been going on. We have quite a few exciting things that you'll probably learn more about in a few months at IO, which I can't go into too much detail about right now. But at least internally, there's been a huge amount of focus on building up a huge library of skills that enable folks to do their job better. And skills are great. But in an organization as large as Google, there's a risk of skills really sprawling out of control. [SPEAKER_00] And do you have any interesting use cases for how people within Google or DeepMind are using it at the moment? [SPEAKER_00] Like what kind of things are they doing with agents in DeepMind? Yeah. So there's a huge amount that's been going on. We have quite a few exciting things that you'll probably learn more about in a few months at IO, which I can't go into too much detail about right now. But at least internally, there's been a huge amount of focus on building up a huge library of skills that enable folks to do their job better. And skills are great. But in an organization as large as Google, there's a risk of skills really sprawling out of control. [SPEAKER_03] And so that's a big area of focus for us right now is improving those skills, making sure that only the best ones survive, almost a Darwinian nature. [SPEAKER_03] But it really is helping folks to deliver good code at a much faster pace, which is obviously awesome. Awesome. Thank you. So you can see this little blue bar around the edge at the moment. That's the anti gravity taking control of the game. It seems to figure out how to start the game. I don't think it knows what the controls are. So maybe I end up looking those up. [SPEAKER_00] But yeah, this is so this is what it edited. [SPEAKER_00] The file that was there before was actually generated by a different model. [SPEAKER_00] So I can tell you it's not even the same game. It's completely rewritten it from scratch based on the spec we gave it. [SPEAKER_00] And then what you would get at the end. [SPEAKER_00] I can see that it's actually looking at the DOM, looking for any errors, trying to figure out how to actually use it. [SPEAKER_00] You should get a video at the end here. Oh, yeah, this is a scratch pad. This is like its notes that it's writing as it goes through and does the task. So you can actually get a bit of a trace as to what behaviors it's trying to figure out. Again, you can go in and you can review these things. If you don't like what they're doing, you can interrupt it. So in terms of the workflow, this is pretty common to a lot of different agent harnesses at the moment. But this is how we think about using things like Gemini models as well so that we can use it for our own development. And I will close that off now. So one big question on my mind related to agents. How do we do things at Google scale? So if we think, research, deep research is a feature within Gemini app, but then also for everyone at Google to use it, what kind of challenges come along with that scaling of agents? Yeah. So the thing that's top of mind for us at the moment is how token hungry this stuff can be. And just managing the quota on a per user or per team basis is really quite important. And so there's a lot of work we're doing around making that more efficient and lower cost. And I think Ian and I were chatting before. And I think what's going to be really interesting for folks like yourselves is mixing and matching between models like Gemma, which are effectively free from a quota perspective, using whatever GPUs or TPUs you have, and then using the more advanced models for specific components of the agentic system. And evaluation is also a big thing we're focusing on at the moment, particularly with these really complicated workflows. How do you actually evaluate that it was successful? How do you minimize the cost of that? So looking into things like mock TPUs so that you can test the harness itself and the agentic flow, but not necessarily using up a ton of TPU hours. Yes, because I'm sure people are aware of this, but definitely limited in that capacity at the moment within the world trying to get enough compute to do a lot of this stuff. [SPEAKER_03] Just quick show of hands from the room. Who's at the moment either building their own agent architecture or harness at the moment? [SPEAKER_03] Okay, fantastic. So just quick question, what kind of scale are you looking at? Who's your customer? Anybody? [SPEAKER_01] We're building memory as a harness, intralayer and essentially an ontology graphic context that developers can use. [SPEAKER_01] We open source to different hardware users. Nice. Nice, nice. And how many users do you expect to be able to scale that to? [SPEAKER_01] We know like five or ten generations of context graphs, around three million a month. [SPEAKER_00] Wow. [SPEAKER_00] That's some scale. [SPEAKER_01] You know, one user can create thousands. [SPEAKER_00] Yep. [SPEAKER_00] That's the challenge too. [SPEAKER_00] I'm sure that's the challenge for us too. [SPEAKER_00] How do you stop one user from, well actually, okay, I'm going to turn that into a question. [SPEAKER_00] How do you stop one user from taking down a whole system? [SPEAKER_01] We could buy spawning multiple instances and multiple. Because I'm sure that's, you know. The more we get better at doing these tasks and we've got power users spinning stuff up and they've got their team of a hundred people working for them, a hundred agents. How do we manage the per user behaviors, if you see what I mean? Yeah, no, it's a great point. And like I said earlier, honestly, right now it's brute force with the quota. So we have some real power users at DeepMind and ultimately it gets to a point where it's okay, you've got to just stop right now. But in general, I think that raises an interesting point about how this stuff is going to be priced in the future as well. You saw Anthropic blocking the open claw stuff because these agentic systems are so token hungry and the subscription model doesn't really work for that. And so, yeah, I think that's really top of mind for me as well right now, how to mitigate that. [SPEAKER_00] Yeah. [SPEAKER_00] I always used to joke when I joined Google that you've got all these resources available in data centers to use for different projects. [SPEAKER_00] How do you know when too much is too much? [SPEAKER_00] And one of my colleagues once told me, he just said, oh, they'll tell you. [SPEAKER_00] And I'm like, who's they? [SPEAKER_00] And sure enough, there's people monitoring these things twenty four seven, looking at spikes and graphs and all our, I'm sorry, team and they do just reach out to you and say, can you just stop this job running on this one cluster, please? [SPEAKER_00] So yeah, I think that's also an interesting one. [SPEAKER_00] I always used to joke when I joined Google that you've got all these resources available in data centers to use for different projects. How do you know when too much is too much? And one of my colleagues once told me, he just said, oh, they'll tell you. And I'm like, who's they? And yeah, sure enough, there's people monitoring these things 24 seven, looking at spikes and graphs and all our, I'm sorry, team and they do just reach out to you and say, can you just stop this job running on this one cluster, please? So yeah, I think that's also an interesting one. Any questions for the audience? [SPEAKER_03] That's a fantastic question. So we built a custom web app for that essentially. And essentially, there's one agent backend system that is used for a lot of stuff at Google. And essentially, anytime a user issues a query to an agent host on that system, it then automatically appears in this UI, where you can drill down at various levels of hierarchy, each of the pieces of the system. If needed, you can drill all the way down to the raw predict request made to the model and so forth. And so that's been useful. And we also have a concept of an agent trajectory store as well. [SPEAKER_00] And that's more focused around the coding piece where obviously you can have a huge number of steps going on there. And it can be really important to diagnose at what exact point looping started happening or the model went off the rails. So, yeah, it's all custom internally for now. But I'd be interested to hear what you folks use for observability as well. Any other questions? Yep. [SPEAKER_00] Yeah, I mean, that's a fantastic question. Obviously, I can't go into too much detail on release plans and so forth. But what I can tell you is, yeah, that's something we're actively exploring. We hope that it will make it faster, cheaper, and hopefully better results if we can effectively orchestrate that deep research using the same harness. [SPEAKER_03] And because right now, without going into too much detail, there's a huge amount of context that's passed all the way through that deep research system, which gets quite expensive and consumes the context. And so we're really thinking about, okay, how do we make each element of this system more like a collaborator as part of a workspace, which is how it would work if humans were researching something deeply, right? And I think that opens up a lot of nice potential for things like infographics, additional supporting artifacts and documents. So, yeah, it's definitely an area of focus for me. [SPEAKER_03] Yeah. How many levels of stuff do you run within extra base? That's a great question. I, good question. I don't know, is the short answer to that. The, I think we, they're not, they're a bit more opaque in how they're presented. The way you think about it is, you have multiple simultaneous ones working on different tracks, but it's not we don't, yeah, the short answer is, it's not as obvious as to which agents are actually working on a particular task for that. [SPEAKER_00] It's not a massively parallel system in that sense. It's more you can give them different trains of operation among a particular project, but you can tell them to work on specific things or you can have jobs that overlap a little bit, but it's not, yeah, I don't have a huge amount of detail on the specifics of how the sub agents work. [SPEAKER_03] I do think that's going to be the future though—how do we make agent to agent communication efficient? And then also how do we give us as the human the ability to really shape that and almost act like a supervisor on a digital assembly line, you know? So yeah, watch this space, I guess. Question? You were next. [SPEAKER_03] There has been a recent debate, of course, about skills and skills, lives, communities moving so fast about it, we haven't found that—where do you see the gravity, 15 lines, the combination of skills with CLI, self-improvement? What's your take on that? [SPEAKER_03] Yeah, for me, I really like skills and they've been working very well for me. Perhaps this is controversial, but I did always think that MCP may be a little bit of a flash in the pan. I like it from the auth perspective, I think that's very powerful. But for me, a combination of skills and guard railed CLI interactions has worked really well. [SPEAKER_00] And it speeds up my job so much. I've got a skill to your point, debugging raw logs and it can do most of that from the CLI. And in a business of our size, the great thing is, we have these skills contributed by folks who are absolute experts in that particular area. And then I and the agent get that knowledge for free. So I'm definitely team skills, if that helps. [SPEAKER_00] I mean, we support both of them. And I think that's the intention going forward. Again, it's what the community uses. We want to make sure that they work with the harness, work with the models. [SPEAKER_03] So I think, yeah, whatever you guys keep using will probably still be supported. It's probably the way to think about it. Yeah. [SPEAKER_00] At the back, should we go over that? [SPEAKER_03] Yeah, no, that's a fantastic point. [SPEAKER_02] I think evaling this stuff is really hard. Even just the mechanical nature of spinning up all of these sandboxed environments set up in the way needed to evaluate a particular problem set. [SPEAKER_03] I think the trickiest part is coming up with new data sets. There's a lot of good open source ones that are good for benchmarking externally. But you're right for specific skills. The onus is almost on the author of the skill itself to come up with some form of test in that. But people are also experimenting with agents designing that as well. So it's a little bit meta, but yeah, a lot of work to do in that space. [SPEAKER_03] I think we have a question over here. Yeah. [SPEAKER_03] When I'm using it, I've got to see the browser testing I love. [SPEAKER_00] Mm. Where it works. Mm. [SPEAKER_03] But you're right for specific skills. [SPEAKER_03] The onus is almost on the author of the skill itself to come up with some form of test in that. [SPEAKER_03] But people are also experimenting with the agents designing that as well. [SPEAKER_03] So it's a little bit meta, but yeah, a lot of work to do in that space. I think we have a question over here. [SPEAKER_03] Yeah. [SPEAKER_03] When I'm using it, I've got to see the browser testing I love. Mm. [SPEAKER_00] Where it works. [SPEAKER_00] Mm. I'm really struggling with the living, so you also, anything from my side, how do I know when I need the limit? Oh, you mean in terms of the model and the model usage? [SPEAKER_00] Yeah. Yeah, yeah. Yeah, that's a good question. [SPEAKER_03] It's funny, actually, because we have worse limits than you do, because obviously we prioritize customers and not ourselves. [SPEAKER_03] So the fact I was clicking a hundred times to keep going is because it recognizes I'm a Googler and that's my fault. [SPEAKER_03] You know, it's a good question. [SPEAKER_03] I think there's two parts to this. [SPEAKER_03] I think that there will always be limits, especially within different tiers. I don't have Ultra, for instance. Sorry, my corporate account does, but my personal one doesn't. So I have to change my behavior. But I think we're going to end up in a world where it's commonplace for you to run out of credit or run out of capacity on one model or something and then move to something else, but do it seamlessly under the harness. So you can give us preference rather than having you've hit your limit on tokens for the pro model. So it will automatically put you onto Flash, or you've reached your limit on everything you have in your subscriptions, so you use a local model. But to your point, it will not interrupt the workflow that you're doing, or the notification you get is not the completion of the task. It's we've run out of quota. Like, what if you were off doing something else where you'd sent it off doing a job and you come back to find that it spent the last hour not doing anything because you hit a limit. So I think when I spoke to him last about this, it was very much more about trying to make it clearer when you get to that point and making sure that you can. [SPEAKER_00] So, yeah. [SPEAKER_00] We have time for one more question. [SPEAKER_00] Do you want to have a go there? [SPEAKER_00] I have a question. I guess in terms of how you all do PRs, we have thousands of PRs today and maybe closing the development loop somehow. I wondered if you maybe thought about fine-tuning some PR models that would go into the task, Google comments, and comments and perhaps then better review that code and reduce the load? [SPEAKER_00] Yeah. [SPEAKER_00] So you're right. [SPEAKER_00] And I really like the fact that we have some really good stuff in place for that already. So the way it works is, on a per language basis, we have a specific auto-review model that has been fine-tuned on all of our style guides and all the rest of it and previous good examples of code. But then also on a product basis, folks will come up with their own specific instructions and prompts and so forth to make sure that the other reviewers get a good signal for how good the code is. And just yesterday, actually, I sent a PR for review and I didn't even have to trigger the auto-review thing. I just got an agent that someone had spun up commenting on the PR with quite a good suggestion. So, yeah, I think it's important. I think to the point earlier about us being supervisors of a digital assembly line, it's how do we get help with that piece of it as well? [SPEAKER_03] Then we can all go sit on the beach. [SPEAKER_03] Yeah. [SPEAKER_03] Yeah, to your point, you can imagine the scale of all the Google engineers submitting 100,000 lines of code now being done by agents submitting even more code with more review. [SPEAKER_03] They've built a lot of infrastructure for us. But then also we have tools like Jules, for instance, which are, if you've ever played with that, you've got a web interface where you can go and do that on your own PRs and GitHub. So, and you get review components of that. So, yeah, I think this is an area that's going to be with the ballooning of, I think there was a comment yesterday about the amount, the trillions of lines that GitHub is getting at the moment generated by agents and that process. So as much as we hate our own boring work, I'm sure the agents hate their boring work too. So we've got to figure out a way to do that. But yeah. [SPEAKER_00] I think we're out of time for questions, but we will be around to chat afterwards if you want to come and join us or if you want to head down to the DeepMind booth later, we'll be around there too. [SPEAKER_00] So thank you very much for listening. [SPEAKER_00] Thank you all. The way, the way you kind of think about it is like, you have multiple simultaneous ones working on different tracks, but it's not, we don't kind of, yeah, the short answer is, it's not as obvious as to which agents are actually working on a particular task for that. It's not kind of like a massively parallel system in that sense. It's more like you can give them different, like trains of operation among a particular project, but you can tell them to kind of work on specific things or you can have jobs that kind of overlap a little bit, but it's not kind of, yeah, I don't have a huge amount of detail on the, like the specifics, how the sub agents work. I do think that's going to be the future though, is how do we make agent to agent communication efficient? And then also how do we give us as the human, the ability to really shape that and almost act like a supervisor on a digital assembly line, you know? So yeah, watch this space, I guess. Question? You were next. There has been a recent debate, of course, you know, skills and skills, lives, communities going so fast about it, we haven't found that, where do you see the kind of gravity, 15 lines, the combination of skills, with CLI, self-improvement, what's your take on that? Yeah, for me, I really like skills and they've been working very, very well for me. Perhaps this is controversial, but I did always think that MCP may be a little bit of a flash in the pan. I like it from the auth perspective, I think that's very powerful. But for me, a combination of skills and guard railed CLI interactions has worked really well. And it speeds up my job so much, you know, I've got a skill for, to your point, debugging raw logs and, you know, it can do most of that from the CLI. And in a business of our size, the great thing is, we have these skills contributed by folks who are absolute experts in that particular area. And then I kind of, I and the agent get that knowledge for free, you know. So I'm definitely team skills, if that helps. I mean, we support both of them. And I think that's the intention going forward. Again, it's like what the community uses. Like, you know, we want to make sure that they work with the harness, work with the models. So I think, yeah, whatever you guys keep using will probably still be supported. It's probably the way to think about it. Yeah. At the back, should we go over that? Yeah, no, that's, that's a fantastic sort of point. I think, um, evaling this stuff is really hard. Um, even just the, the mechanical nature of spinning up all of these sandboxed environments set up in, in, in the way needed to, to evaluate, evaluate a particular, um, problem set. Um, I think that the trickiest part is coming up with new data sets, you know, um, there's a lot of good, um, you know, open source ones that are good for benchmarking externally. But you're right for specific skills. The onus is almost on like the, the, the, the author of the skill itself to, to come up with, um, some form of test in that. But people are also experimenting with the agents designing that as well. Um, so it's a little bit meta, but, um, yeah, a lot of work to do in that space. I think we have a question over here. Yeah. When I'm using it, like, I've got to see, like, the browser testing I love. Mm. Where it works. Mm. I'm really struggling, uh, with the living, uh, so, you also, like, like, the, or, like, the, or, like, the, or, like, the, or, like, the, or, like, the, or, like, anything from my side, like, how do I know when I need the limit? Oh, you mean in terms of, for, like, the model and the model usage? Yeah. Yeah, yeah. Yeah, that's a good question. Uh, it's funny, actually, cause we have, we have worse limits than you do. because obviously we prioritize customers and not ourselves. So the fact I was like clicking like a hundred times to keep going is because it recognizes I'm a Googler and that's my fault. You know, it's a good question. I think there's kind of two parts to this. I think that there will always be limits, especially like within different tiers. I don't have Ultra, for instance. Sorry, my corporate account does, but my personal one doesn't. So I have to change my behavior. But I think we're going to end up in a world, I mean, we have hopefully somewhere around here is Kevin from the anti-gravity team. We could probably talk specifically on that. But I think we're going to be in a pattern whereby it's commonplace for you to run out of credit or run out of capacity on like one model or something and then move to something else, but do it seamlessly under the harness. So you can give us preference rather than having like, you know, you've hit your limit on tokens for the pro model. So it will automatically put you onto Flash or you've reached your limit on everything you have in your subscriptions. So you use a local model or, but to your point, like that it will not interrupt the workflow that you're doing or, you know, the notification you get is not the completion of the task. It's like, oh, we've run out of quota. Sorry. Like, you know, what if you were off doing something else where you'd sent it off doing a job and you come back to find that it spent the last hour not doing anything because you hit a limit. So I think the way when I spoke to him last about this, it was very much more about like trying to make it clearer when you get to that point and making sure that you can. So, yeah. We have time for one more question. Do you want to have a go there? I have a question. I guess in terms of how guys do you PR the tool as we have like thousands of PR today and maybe closing the development loop somehow. I wondered if you maybe thought about fine-tuning some like PR models that would go into the task, Google comments, and comments and perhaps then better review that code and reduce the load? Yeah. So you're right. And I really like the fact that we have some really good stuff in place for that already. So the way it works is, you know, on a per language basis, we have a specific like auto-review model that has been fine-tuned on all of our style guides and all the rest of it and previous like good examples of code. But then also on a PA or product basis, folks will come up with their own like specific SIs and prompts and so forth to make sure that, you know, the other reviewers get a good signal for how good the code is. And just yesterday, actually, I sent a PR for review and I didn't even have to trigger the auto-review thing. I just got an agent that someone had spun up commenting on the PR with quite a good suggestion. So, yeah, I think it's important. I think to the point earlier about us being supervisors of a digital assembly line, it's like how do we get help with that piece of it as well? Then we can all go sit on the beach. Yeah. Yeah, to your point, like you can imagine like the scale of all the Google engineers submitting 100,000 lines of code now being done by agents submitting even more code with more review. They've built a lot of kind of infrastructure for us. But then also we have like tools like Jules, for instance, like which are, if you've ever played with that, you've got like a web interface where you can go and do that on your own PRs and GitHub. So like, and you get review components of that. So, yeah, I think this is an area that's going to be with the ballooning of, I think there was a comment yesterday about like the amount of the trillions of lines that GitHub is getting at the moment generated by agents and like that process. So as much as we hate our own boring work, I'm sure the agents hate their boring work too. So we've got to figure out a way to do that. But yeah. I think we're out of time for questions, but we will be around to chat afterwards if you want to come and join us or if you want to head down to the DeepMind booth later, we'll be around there too. So thank you very much for listening. Thank you all.