AI Engineer

Get Out of the Model's Way — Kevin Hou, Google Antigravity

1857 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent products should replace fixed, human-designed workflows with primitives that let increasingly capable models dynamically orchestrate specialist subagents, react to external events, and generate the interface needed for the task.
  • Why it matters: This is a concrete control-plane thesis for agent systems: the durable product advantage is not a static IDE or workflow UI, but an environment in which a capable lead model can delegate, select models, use tools safely, and present inspectable outputs.
  • Best use: Use it as a product-architecture and roadmap lens for OpenClaw or other agentic systems, especially around dynamic delegation, event-driven execution, and generated operational interfaces.

Executive Summary

Kevin Hou frames Google Antigravity as an "agent-first" coding product built around a simple principle: as model intelligence improves, the product should expose that capability rather than preserve familiar but limiting interaction patterns. He argues that prior transitions—from deterministic autocomplete to agents, from chat sidebars to autonomous execution, and from terminal fear to permissioned agent access—show that product teams sometimes need to remove entrenched UX in order to let models perform higher-leverage work.

Antigravity 2.0 operationalizes that view by decoupling the IDE from an agent manager. Hou positions the IDE as a debugger-like layer beneath the main abstraction: useful for inspection and intervention, but not necessarily the default place where work is managed. The primary product becomes a mission-control environment for parallel agents, with a lead agent that decomposes a user goal, creates specialized subagents, assigns models and environments, and aggregates results.

The talk identifies three proposed "2026 primitives": dynamic subagents, sidecars, and generative UI. Dynamic subagents are created and configured by a lead agent for specific roles and can run in parallel or in isolated environments. Sidecars are long-lived, event-listening processes for triggers such as schedules, webhooks, SMS, and GitHub PRs. Generative UI allows the model to render the most useful interactive view—Kanban, timeline, charts, filters, tables—at runtime instead of relying on fixed application screens.

The most useful evidence is two internal/product examples rather than a generalized benchmark. Antigravity reportedly automated 90% of an internal rollout-comparison workflow by generating hypotheses and parallel investigations over eval deltas, then presenting an interactive analysis UI. A more extreme demonstration built an OS kernel capable of running Doom using 93 subagents over 12 hours, 15,000 requests, and 2 billion tokens for under $1,000. These examples support the architectural direction, but they do not establish routine reliability, quality control, or favorable economics for ordinary production work.

Key Takeaways

  • Claim: The product interface should scale with model intelligence rather than lock users into interaction patterns that were necessary for weaker models. | Evidence: Hou contrasts 2022-era deterministic developer tooling—embeddings, rules files, AST parsing, autocomplete, and chat sidebars—with the later rise of agents, MCPs, custom tools, permissions, skills, hooks, artifacts, and multi-agent management. | Implication: Ken should evaluate product primitives by whether a substantially better future model can exploit them without a wholesale redesign, rather than by whether they optimize only today's manual workflow. | Caveat: Hou explicitly says these product bets are not right 100% of the time; the argument is a directional design philosophy, not proof that every familiar interface should be removed.
  • Claim: Agent orchestration should become the primary abstraction, while the IDE becomes an optional lower-level inspection and intervention surface. | Evidence: Antigravity 2.0 separates the agent manager from the IDE; Hou analogizes the IDE's role to a debugger relative to an IDE—valuable when needed, but not always the main workspace. He predicts agent teams, swarms, or software factories will become the dominant operating model. | Implication: Design the control plane so users can set objectives, observe execution, and intervene at meaningful checkpoints without requiring them to micromanage files, prompts, or individual agent steps. | Caveat: This is a product prediction from the Antigravity team, not evidence that users broadly prefer a standalone orchestration layer.
  • Claim: A capable lead agent should dynamically assemble a team rather than rely on a fixed catalog of preconfigured agents. | Evidence: In Antigravity's /teamwork mode, the lead agent clarifies the task, determines arbitrary team size, dynamically configures and seeds subagents for roles such as frontend, backend, infrastructure, QA, and design, and may choose a different model for a subagent than for itself. | Implication: For complex workflows, favor a planner-orchestrator that can create task-specific workers, select an appropriate model or execution environment, and consolidate results over a rigid multi-agent template. | Caveat: Hou acknowledges substantial remaining headroom in multi-agent collaboration and task decomposition, so dynamic delegation should not be treated as solved.
  • Claim: Large agent teams can tackle long-running, high-complexity builds, but their cost and duration make them a selective capability rather than a default workflow. | Evidence: Antigravity's hero run reportedly created an OS kernel from scratch that ran Doom using 93 subagents over 12 hours, 15,000 requests, and 2 billion tokens, at under $1,000. Photo-editor and messaging-app examples each reportedly required hundreds of subagents and nearly half a day. | Implication: Treat broad parallelism as an escalation mode for bounded but expensive projects; impose budget caps, timeout policies, artifact-based review gates, and measurable success criteria before allowing large autonomous swarms. | Caveat: These are showcase runs reported by the product team; the transcript offers no independent quality evaluation, rework rate, success distribution, or comparison with human teams.
  • Claim: The immediate high-value application of agent teams may be automating research and eval-analysis workflows, not only code generation. | Evidence: For side-by-side rollout analysis, an internal Antigravity workflow reportedly automated 90% of a previously Jupyter-heavy process: it calculated deltas, proposed 100 hypotheses for the differences, spawned parallel hypothesis-specific subagents, and returned a consolidated report with interactive filtering and segmentation. | Implication: Prioritize agentic automation where work already has repeatable inputs, comparable outputs, and a reviewable analytical artifact—such as experiment diagnosis, incident investigation, pipeline triage, or competitive research. | Caveat: The system depends on loaded skills and access to Google's monorepo and internal data context; performance may not transfer directly to organizations with weaker data hygiene or fragmented tool access.
  • Claim: Event-driven sidecars and runtime-generated interfaces are proposed as core primitives for moving agents from request-response tools into persistent operational systems. | Evidence: Hou describes sidecars as long-lived plug-in processes that listen for outside events including SMS, webhooks, cron schedules, and GitHub PRs; Antigravity already uses them for scheduled tasks. He says Gemini Flash runs at nearly 900 tokens/second and can render inline, task-specific Kanbans, timelines, charts, tables, and other interactive UI rather than relying on templates or HTML files. | Implication: Separate durable event/listening infrastructure from ephemeral agent runs, and require generated UIs to be inspectable, permission-aware, and disposable views over trusted underlying data and actions. | Caveat: The sidecar specification was described as planned for release later in the summer, and the claim that human-written specialized UIs are "dead" is a hypothesis, not a demonstrated universal conclusion.

Detailed Brief

The product bets are rooted in earlier interface transitions

  • Claims: Hou uses prior controversy to argue that user attachment to a familiar interface is not sufficient reason to preserve it.; The safer route to greater autonomy is to improve the surrounding primitives as models improve, rather than simply granting unconstrained access.
  • Evidence: He cites early concerns that giving agents terminal access could cause catastrophic deletion or other damage, followed by adoption once permission systems and model judgment improved.; He cites backlash after Windsurf removed its chat sidebar in favor of an agent-only experience, arguing that multi-step research and execution later became the prevailing interaction pattern.
  • Caveats: The historical cases are selectively presented and do not identify adoption losses, failure rates, or contexts where chat-first and manual interfaces remain preferable.
  • Implications: Maintain an escape hatch and progressive autonomy model during interface transitions, but do not make legacy UX the architectural center of an agent-native product.

What the three primitives imply for system boundaries

  • Claims: Subagents are not merely parallel prompts; they are dynamically specified workers with specialized roles and potentially separate models and security environments.; Generative UI is intended as an adaptive presentation layer, not simply a code-generation novelty.
  • Evidence: Hou says subagents can operate in sandboxes or remote-execution systems and are configured, prompted, and seeded by the main agent.; The eval-analysis example uses generated controls for selecting, filtering, segmenting, and slicing results after the investigation completes.
  • Caveats: The talk leaves unspecified the governance model for agent-to-agent permissions, data isolation, generated-UI integrity, audit logs, and approval enforcement.
  • Implications: The necessary architecture is a governed orchestration layer: agent identity, scoped credentials, budget and concurrency controls, sandbox boundaries, durable artifacts, and a human review surface should be first-class rather than added after delegation works.

Notable Concepts & Terms

  • Scaling with intelligence: Hou's core product principle: each new generation of model capability should translate into a visibly more capable user experience, rather than being constrained by old workflow assumptions.
  • Agent manager: The standalone mission-control layer in Antigravity 2.0 for managing projects and parallel agents, separated from the IDE.
  • Agent teams / swarms / software factories: Alternate labels for a lead agent dynamically coordinating specialized workers to complete complex work in parallel.
  • Dynamic subagent: A task-specific worker created, prompted, configured, model-routed, and potentially environment-routed by a lead agent at runtime.
  • Sidecar: A long-lived plug-in process that listens for external events and establishes triggers, enabling scheduled and event-driven agent workflows.
  • Generative UI: Runtime-generated, inline interactive interfaces such as dashboards, timelines, charts, and filters that are tailored to the task rather than prebuilt screens.
  • Skills file: A loaded source of domain/workflow knowledge that lets an agent operate effectively against a specific codebase or organizational process.
  • Pareto curve of intelligence, speed, and cost: Hou's claim that Gemini 3.5 Flash improves the tradeoff among capability, latency, and cost enough to make larger-scale orchestration more viable.

Operator Notes / Why Ken Should Care

  • Define a reusable lead-agent contract for complex runs: task clarification, decomposition, delegated-worker creation, model selection, consolidation, and explicit completion criteria.
  • Add hard controls before enabling high-concurrency swarms: per-run budget ceilings, token/request quotas, maximum duration, concurrency limits, sandbox selection, and mandatory human approval for consequential actions.
  • Prototype one event-driven sidecar workflow around a concrete trigger such as a GitHub PR, scheduled KPI anomaly, or inbound message; measure reliability and operational burden before generalizing the pattern.
  • Choose one analysis workflow with structured artifacts and human review—such as eval deltas or incident postmortems—and test hypothesis fan-out plus evidence-linked synthesis rather than deploying multi-agent orchestration first on open-ended execution.
  • Keep generated UI as a view layer over durable logs, artifacts, and source data; do not make ephemeral agent-rendered screens the sole audit or decision record.
  • Monitor the planned Antigravity sidecar protocol release for interoperability patterns, but avoid coupling core architecture to an unreleased specification.

Source/Metadata

  • Title: Get Out of the Model's Way — Kevin Hou, Google Antigravity
  • Transcript words: 6409
  • Duration seconds: 1141
  • Timestamp note: No timestamps or chapters were present. The transcript contains substantial repeated passages and trailing transcription noise.
Full transcript 3357 words · 28 min read
0:28

Transcription by CastingWords All right, hello everyone. My name is Kevin. I'm going to be talking about anti-gravity.

0:36

So are there any World Cup fans out there? Woo! Imagine you are coaching Argentina, and you're in the 89th minute, and you have Messi on your team. What play are you running? It's called give Messi the ball and get the heck out of the way. LLMs aren't just role players anymore. They can be your star player if you build the right product around them. And to let your star player cook, you have to get out of the model's way. You might want to get the slide. Are the slides up? Oh, they are. Great.

0:44

So anti-gravity is Google's agentic coding product for technical and non-technical users. We launched back in November of 2025 and have been accelerating devs both within Google and externally ever since. My name is Kevin Howe, and I lead part of the engineering team on anti-gravity.

0:50

So let's talk a little bit more about what anti-gravity is. We have and always will be unapologetically agent first. So we debuted the anti-gravity IDE last year with a brand new agent manager concept, and it was a platform to manage and orchestrate many agents. Since then, we've actually extracted our agent and launched our own anti-gravity CLI, and last month at Google I.O., we had the pleasure of launching anti-gravity 2.0. In the theme of getting the model out of the way, we actually decoupled the IDE from the agent manager, so now you have two separate applications, and now you can use the agent manager in a standalone app.

0:58

And since pictures are worth a thousand words, here's a screenshot of anti-gravity 2.0 in action. As you can see, not only is it your own dedicated mission control for your agents and projects, you have sub-agents, you have all the new models, you have work trees, scheduled tasks, voice mode. There are so many things to unpack with the product. But I don't want to spend today telling you about the product. I want to tell you a little bit more about the behind the scenes, some of the principles that went into it, and notably, some of the things that led to its roadmap.

1:04

So as some of you, for the long-time AI Eng fans, this is actually my fifth time speaking at AI Eng, and I've been building developers' tools since 2022. And the one thing that has stood above all other lessons that I've talked about is the idea of scaling with intelligence. This means that as the model gets better, so should your product. And the frontier edge of whatever model you are serving should be apparent inside of your user's product experience.

1:11

So let's get into more concrete examples of what this means. So for those of you that follow me on X or hear me generally for the last four years, you'll know that I've been working on a number of these transformations year over year over year. In 2022, I was working on autocomplete and chat sidebars. This was based on embeddings, rules files, AST syntax tree parsing. Everything inside of that app is deterministic, because that's all that the model could really handle. And in 2024, when agents came onto the scene, it completely changed how developers were going to do work. With it came new primitives, like MCPs, custom tools, and permission systems.

1:15

And with 2025, we introduced Antigravity's agent manager with many other products following suit in that similar form factor, with users managing many agents at once in parallel. And this led to things like skills, hooks, artifacts, and a couple other primitives. And that sort of defined the 2025 era. So let's talk a little bit about 2026 and what those primitives might be. Before we answer this question, I want to take you back to some of these battle scars that are a little bit closer to home.

1:22

Scaling with intelligence really is not easy. It's really hard to take away something that users love and are familiar with to lead them down potentially, and that's a big keyword, a better path. We aren't right 100% of the time, but there are two that jumped to mind when I was putting together the slides for this talk.

1:30

The first one is giving AI a terminal. We all remember fears about son of Anton deleting your entire code base and doing catastrophic things to both your startup, your company, etc. But as models got better and people invested in primitives, such as permission systems, users ended up building faster, they ended up shipping more, and they did so safely. So we were able to overcome this, and as models got smarter, they were able to make better decisions about what they should and should not run in your terminal.

1:35

The second instance is this tweet, which is very representative of the yelling that I got when we removed chat from Windsurf. So a lot of users were yelling at our team because we took away something that was very dear to them, the chat sidebar, and replaced it with only an agent. Now, at the time, this is something that was familiar and rather difficult to swallow. But when we look back, models have advanced, multi-step research, agentic research, and execution became the new paradigm, and here we are today using and loving all these agentic products.

1:38

And so now I bring you to today's battle. What is going on today? So we decoupled the agent manager from the IDE, and with Antigravity 2.0, we split them into separate applications. We believe that the IDE is to the agent manager what the debugger was to the IDE. You don't always need a debugger, but it definitely is helpful to have it if you need to go a layer beneath and go one step deeper into that abstraction stack. And our prediction is that this idea of agent orchestration, you can call it agent teams, you can call it swarms, you can call it software factories, is the future, and we're willing to bet on that future.

1:43

So here are the primitives for what we're calling the agent teams 2026 era. These are things like subagents, generative UI, and sidecars. And we'll talk more concretely about what those things are and some examples of how they manifest inside of the product. But it's really important to first understand the why. What brought about these changes, and what model changes, what model properties actually led to the development of these new things? And as a product team, do you force the new era of primitives, or is it something that comes to you by using the model and experiencing the model? The answer is both, right?

1:53

And the privilege of being inside of Google DeepMind is that we do have that relationship between the product and the model.

1:59

So you remember the crux of antigravity 1.0 is to manage agents in parallel, to put the human in the driver's seat. And if you remember my last talk, I talked a lot more about this research product flywheel. And now, as promised, because of the antigravity product, Gemini has now learned a thing or two about how to manage a team of agents. There's still a lot of headroom to make multi-agent systems better, more collaborative, better at deconstructing tasks into smaller tasks, but we've got a really good head start with Gemini. And all the basics have been imbued to the model so that we can build a product like antigravity 2.0.

2:04

Gemini 3.5 Flash was launched back in April, and this brought to market a lot of those capabilities that we had been working on in the background with antigravity. And Flash now isn't just good at executing tasks, it's actually really good at leading teams. It's faster and cheaper, pushing the Pareto curve of what is intelligent versus the speed and the cost at which you run those things. And putting this all together, we were really excited to announce agent teams in public preview inside of antigravity. All you have to do is simply type the slash command, slash teamwork, and you'll see a new mode where you can enter and unleash a swarm of agents onto the task at hand.

2:18

So we'll talk a little bit about how this works. You as a user will specify your task. The more specific you are, the better, though the nature of these agentic communication styles is that if it needs something more, it can actually ask you for more. Until everything is basically clear, you'll work with that lead agent, and it will manage a team of arbitrary size to get that work done. And what I like to say, it's the Avengers, right? It'll take a bunch of specialized roles—front-end engineers, back-end engineers, infrastructure specialists, QA, design. The list goes on and on and on, and there are infinite possibilities for what each of those subagents could take on.

2:27

Each subagent is dynamically generated and can operate independently. And it can even actually select a different model from what the main agent is using, and this is done so by that main agent. Again, we are scaling with intelligence. And one of the coolest aspects of this is that it can use generative UI. With a model that is as fast as Flash, things can happen nearly instantaneously if you ask, hey, what is the status of my task? Show me a Kanban of what's going on. Or maybe you prefer something

2:36

and it will manage a team of arbitrary size to get that work done. And what I like to say, it's the Avengers, right? It'll take a bunch of specialized roles—front-end engineers, back-end engineers, infrastructure specialists, QA, design. The list goes on and on, and there are infinite possibilities for what each of those subagents could take on. Each subagent is dynamically generated and can operate independently. And it can even select a different model from what the main agent is using, and this is done by that main agent. Again, we are scaling with intelligence. And one of the coolest aspects of this is that it can use generative UI. With a model that is as fast as Flash, things can happen nearly instantaneously if you ask, "What is the status of my task? Show me a Kanban of what's going on." Or maybe you prefer something like the Chrome debugger tool. It can show you a timeline like that. And all these things are generated on the fly because it's able to generate UI on demand.

2:39

So some of the projects that the system has implemented—we've built a photo editor. You can actually edit raw photos directly inside your browser. We've also built a messaging app that might look familiar to those in the room. And each of these took hundreds of subagents and took almost half a day to run.

2:41

But to really put it through its paces, one of the hero runs that we did was actually building an entire OS kernel. This is something that we got to show off at Google I/O, but we built a complete OS kernel from scratch and actually played Doom on it. And my colleague Varun was able to demo this at Google I/O. We were super proud of this particular milestone because it really demonstrated that if you throw more intelligence, you throw more subagents at this sort of problem, a model like Gemini 3.5 Flash could do this in a way that was not only very powerful, but also scalable and mildly affordable. Obviously, we're not going to spend thousands of dollars to build an OS kernel every day, though it is possible.

2:48

And some of the stats out of this: it took 93 subagents over the course of 12 hours, made 15,000 requests, 2 billion tokens, and it was under $1,000, which was one of the really cool aspects of this project. And so as you can see with this particular example, subagent primitives are one of the defining parts about building a 2026 era of agent teams. So agent teams are just that first example, and I want to show you another example that our team uses internally that demonstrates some of these new primitives.

2:58

The second one is about automating research tasks. So we work inside of Gemini. We help make Gemini better at coding-related tasks, agentic-related tasks, and this is where the real magic starts happening with the product. We have an internal version of Antigravity that researchers, engineers, non-technical folks can use, and when they understand the primitives that Antigravity offers, it becomes a very powerful way to automate your own workflows.

3:03

So we'll take the example of side-by-side eval analysis. This is a very common workflow, not only at DeepMind, but generally in the industry. You essentially take multiple rollouts—one, two, three, four, et cetera—and you want to compare them. So you take a set of tasks, you do some rollouts, you get some results, and they'll be in two different tables. Now, you look at the control, you look at the experiment, and then you have to figure out not only what the difference was, but perhaps what are the reasons for those differences, and how can we actually iterate from there and make a better version for the next experiment.

3:05

Traditionally, this was a lot of Jupyter Notebook work, but when you start working with the new primitives in 2026, you end up with a much cleaner workflow. Researchers were able to automate 90% of this workflow by simply asking the agent about the evals in question using natural language. Then the agent that is now primed with skills and an understanding of Google's massive monorepo code base is able to crunch the numbers and get back to you with a delta.

3:09

Now, what's really cool here is instead of just taking that delta and handing it back to the user, it went the extra step. It spun up a research agent specialist that proposes 100 different hypotheses over why those deltas might occur. And then it uses subagents to spin up one subagent for each hypothesis and basically drill into that particular case in parallel, mapping back to a single response and then telling the researcher, "Here are some areas that I found. Now, here's a report that you can review."

3:15

And what's really cool is that it doesn't stop at just the report. It actually puts together generative UI for you to look through, interact, select drop-downs, filter, segment, slice, and actually interact richly with that data. And internally, we care a lot about this sort of workflow, improving the model, improving the product, and understanding the ways that users find success and failure internally at Google. So what used to be a very manual process now takes minutes.

3:24

So what used to be hand engineering—you'd have to build your own async pool of agents, you'd have to set up your judges, you'd have to tape together data pipelines. All of this now starts becoming grounded in these new primitives that we've established earlier in the slideshow. You have a subagent graph that is completely dynamic. The generative UI comes in at the end to richly convey the findings in a way that the user best understands or caters to their learning style. And all of these things can be regenerated and redone on the fly. All the user had to do was load up a skills file and ask away.

3:30

So with teamwork and this eval example, we start arriving at these 2026 primitives that I keep talking about. And these model characteristics really change the way that we have to think about the product and how we have to develop the product.

3:34

So the three examples that we've talked about—first, we have the dynamic subagent. And to provide a little more color here, basically no two subagents are the same. The main agent is the one that is orchestrating this entirely on its own. It's configuring and prompting and seeding these subagents on the fly. They can operate in parallel. They can operate in different types of secure environments, be it a sandbox, be it a remote execution system. And they can all take on an infinite number of specialized roles. So the scaling story here is quite obvious.

3:39

And from the last two examples, you can probably tell. As the model gets smarter, your team will become more specialized. It will become more collaborative. And ultimately, that means it will be capable of getting more complex work done for you.

3:43

And now the second is this new concept. We've alluded to it slightly in the past, but it's called sidecars. This is a new plug-in protocol that we're bringing to Antigravity. A sidecar process is essentially a sidecar process. The naming reflects what's going on under the hood. But it's a long-lived utility. And it's responsible for listening. It allows the model to listen to the outside world and set up its own triggers for things that might happen. For example, this could be SMS messages. This could be web hooks, cron jobs, hooking it up to GitHub PRs. The list goes on. But this is a generic plug-in primitive. Antigravity already uses sidecars for things that are time-based. This is where the scheduled task cron concept comes from. But under the hood, this is all this new sidecar primitive.

3:45

So we'll be releasing the spec for this so that you all can build on top of this new primitive later this summer. But there are some really cool ways that people internally have been using this sort of concept. And the third and final primitive is generative UI. So we hypothesize that human-written specialized UIs are dead. Gemini Flash on Antigravity clocks in at almost 900 tokens a second. This is 10x faster than a lot of the other frontier model experiences. And in a matter of seconds, you're able to go from whatever you're thinking inside your head into a prompt, into a use case that is designed and embedded inside your conversation view perfectly.

3:52

And rather than rely on templates or HTML files, Antigravity can render your generated UI in line. So you can do things like play Doom, but this also extends to things like bar charts, graphs, tables—anything that you would want to interact with and inspect a little further than just a markdown file or just a conversation.

3:58

And generative UI in many ways reminds me of the quote that the late Steve Jobs said when unveiling the iPhone. He justifies the removal of the keyboard and says, "They all have these keyboards. They are there whether you need them or not. And they all have these control buttons that are fixed in plastic and are the same for every application." In an analogous way, we built our product to dynamically scale with the needs of the agent. We skipped the heavy infrastructure and mechanical UIs.

4:02

inside of your head into a prompt, into a use case that is designed and embedded inside of your conversation view perfectly. And rather than rely on templates or HTML files, Antigravity can render your generated UI in line. So you can do things like this and play Doom, but this also extends to things like bar charts, graphs, tables, anything that you would want to interact with and maybe inspect a little bit further than just a markdown file or just a conversation.

4:05

And generative UI in many ways reminds me of the quote that the late Steve Jobs said when unveiling the iPhone. He justifies the removal of the keyboard and says, they all have these keyboards. They are there whether you need them or not. And they all have these control buttons that are fixed in plastic and are the same for every application. In an analogous way, we built our product to dynamically scale with the needs of the agent. We skipped the heavy infrastructure and mechanical UIs in favor of sidecars and generative UI. And that creates a product experience that is not fixed in plastic.

4:15

So subagents, sidecar triggers, and generative UI are the latest primitives that are powering Antigravity. We've tried our best to stay out of the way and let the model cook. And if you're building a product around an agent, you should consider what are the primitives that are in my product and how might they scale with the model's intelligence. We are all familiar with shipping features is now quite easy with all of these new tools. And it's about deciding what features to actually add so that the model, so that your product can scale with the next release of the next model, which will inevitably be faster, better, and cheaper.

4:23

And so with the right primitives, you as a builder or you as a product owner, you might be surprised at what the models can do. And in classic fashion, I'm going to keep using this slide until we've actually conquered the TPU crunch. So you can find me on Twitter. You can DM me for feedback. We're always looking for new ideas on how to build the latest and greatest. Thank you for watching. Thank you, Swix and Ben for having me. It's always a joy to be here. And I'll be at the Antigravity booth if you want to talk further. If you want to get to know the product a bit more, some team members will be there. So thank you so much for your time. Excited to meet you all.

4:34

to lead them down potentially, and that's a big keyword, a better path. We aren't right 100% of the time, but there are two that jumped to mind when I was putting together the slides for this talk. The first one is giving AI a terminal. We all remember fears about son of Anton deleting your entire code base and doing catastrophic things to both your startup, your company, etc., etc. But as models got better and people invested in primitives, such as permission systems, users ended up building faster, they ended up shipping more, and they did so safely. So we were able to overcome this, and as models got smarter,

5:11

they were able to make better decisions about what they should and should not run in your terminal. The second instance is this tweet, which is very representative of sort of the yelling that I got, when we removed chat from Windsurf. So a lot of users were yelling at our team because we took away something that was very dear to them, the chat sidebar, and replaced it with only an agent. Now, at the time, this is something that was familiar and rather difficult to swallow. But when we look back, models have advanced, multi-step research, agentic research, and execution became the new paradigm, and here we are today using and loving all these agentic products.

5:49

And so now I bring you to today's battle. What is going on today? So we decoupled the agent manager from the IDE, and with Antigravity 2.0, we split them into separate applications. We believe that the IDE is to the agent manager what the debugger was to the IDE. You don't always need a debugger, but it definitely is helpful to have it if you need to go a layer beneath and go one step deeper into that abstraction stack. And our prediction is that this idea of agent orchestration, you can call it agent teams, you can call it swarms, you can call it software factories, is the future, and we're willing to bet on that future.

6:26

So here are the primitives for what we're calling the agent teams 2026 era. These are things like subagents, generative UI, and sidecars. And we'll talk more concretely about what those things are and some examples of how they manifest inside of the product. But it's really important to first understand the why. What brought about these changes, and what model changes, what model properties actually led to the development of these new things? And as a product team, do you force the new era of primitives, or is it something that comes to you by using the model and experiencing the model? The answer is kind of both, right?

7:04

And the privilege of being inside of Google DeepMind is that we do have that relationship between the product and the model. So you remember the crux of antigravity 1.0 is to manage agents in parallel, to put the human in the driver's seat. And if you remember my last talk, I talked a lot more about this research product flywheel. And now, as promised, because of the antigravity product, Gemini has now learned a thing or two about how to manage a team of agents. There's still a lot of headroom to make multi-agent systems better, more collaborative, better at deconstructing tasks into smaller tasks, but we've got a really good head start with Gemini.

7:43

And all the basics have been imbued to the model so that we can build a product like antigravity 2.0. Gemini 3.5 Flash was launched back in April, and this brought to market a lot of those capabilities that we had been working on in the background with antigravity. And Flash now isn't just good at executing tasks, it's actually really good at leading teams. It's faster and cheaper, pushing the Pareto curve of what is intelligent versus the speed and the cost at which you run those things. And putting this all together, we were really excited to announce agent teams in public preview inside of antigravity.

8:20

All you have to do is simply type the slash command, slash teamwork, and you'll see a new mode where you can enter and unleash a swarm of agents onto the task at hand. So we'll talk a little bit about how this works. You as a user will specify your task. The more specific you are, the better, though the nature of these agentic communication styles is that if it needs something more, it can actually ask you for more. Until everything is basically clear, you'll work with that lead agent, and it will manage a team of arbitrary size to get that work done. And what I like to say, it's kind of like the Avengers, right? It'll take a bunch of specialized roles

8:56

and front-end engineers, back-end engineers, infrastructure specialists, QA design. The list goes on and on and on, and there are infinite possibilities for what each of those subagents could take on. Each subagent is dynamically generated and can operate independently. And it can even actually select a different model from what the main agent is using, and this is done so by that main agent. Again, we are scaling with intelligence. And one of the coolest aspects of this is that it can use generative UI. With a model that is as fast as Flash, things can happen nearly instantaneously if you ask, hey, what is the status of my task? Show me a Kanban of what's going on.

9:34

Or maybe, you know, you prefer something a little bit more like the Chrome debugger tool. It can show you a timeline like that. And all these things are generated on the fly because it's able to generate UI on demand. So some of the projects that the system has implemented, we've built a photo editor. You can actually edit raw photos directly inside of your browser. We've also built a messaging app that might look a little bit familiar to those in the room. And each of these took hundreds of subagents and took almost half a day to run. But to really put it through its paces, one of the hero runs that we did was actually building an entire OS kernel.

10:12

This is something that we got to show off at Google I.O., but we built a complete OS kernel from scratch and actually played Doom on it. And my colleague Varun was able to demo this at Google I.O. We were super proud of this particular milestone because it really demonstrated that if you throw more intelligence, you throw more subagents at this sort of problem, a model like Gemini 3.5 Flash could do this in a way that was not only very, very powerful, but also scalable and, you know, mildly affordable. Obviously, we're not going to spend thousands and thousands of dollars to build an OS kernel every day, though it is possible. And some of the stats out of this,

10:47

it took 93 subagents over the course of 12 hours, made 15,000 requests, 2 billion tokens, and it was under $1,000, which was one of the really cool aspects of this project. And so as you can see with this particular example, subagent primitives are one of the defining parts about building a 2026 era of agent teams. So agent teams are just that first example, and I want to show you another example that our team uses internally that sort of demonstrates some of these new primitives. The second one is about automating research tasks. So we work inside of Gemini. We help sort of make Gemini better at coding-related tasks, agentic-related tasks,

11:26

and this is where the real magic starts happening with the product. We have an internal version of anti-gravity that researchers, engineers, non-technical folks can use, and when they understand the primitives that anti-gravity offers, it becomes a very, very powerful way to automate your own workflows. So we'll take the example of side-by-side eval analysis. So this is a very common workflow, not only at DeepMind, but just generally in the industry. You essentially will take multiple rollouts, one, two, three, four, et cetera, and you want to compare them. So you'll take a set of tasks, you'll do some rollouts, you'll get some results,

11:59

and they'll essentially be in two different tables. Now, you'll look at the control, you'll look at the experiment, and then you'll have to figure out not only what the difference was, but perhaps what are the reasons for those differences, and how can we actually iterate from there and make a better version for the next experiment. Now, traditionally, this was a lot of Jupyter Notebook elbow grease, essentially, but when you start working with the new primitives in 2026, you end up with a lot cleaner of a workflow. So researchers were able to automate 90% of this workflow by simply asking the agent about the evals in question using natural language.

12:33

Then the agent that is now primed with skills and an understanding of Google's massive monorepo code base is able to crunch the numbers and get back to you with a delta. Now, what's really cool here is instead of just taking that delta and then handing it back to the user, it went the extra step. It spun up for a research agent specialist that proposes 100 different hypotheses over why those deltas might occur. And then it uses subagents to then spit up one subagent for each hypothesis and basically drills into that particular case in parallel, mapping back to a single response and then telling the researcher, hey, here are some areas that I found.

13:09

Now, here's a report that you can review. And what's really cool is that it doesn't stop at just the report. It actually puts together a generative UI for you to look through, interact, select drop-downs, filter, segment, slice, and actually interact richly with that data.

13:26

And internally, we care a lot about this sort of workflow, improving the model, improving the product, and understanding the ways that users find success and failure internally at Google. So what used to be a very manual process now takes minutes. So what used to be hand engineering, you'd have to build your own async pool of agents, you'd have to set up your judges, you'd have to tape together data pipelines. All of this now starts becoming grounded in these new primitives that we've established earlier in the slideshow. You have a subagent graph that is completely dynamic. The generative UI comes in at the end to richly convey the findings

13:59

in a way that the user best understands or maybe caters to their learning style. And all of these things can be regenerated and redone on the fly. All the user had to do was load up a skills file and ask away. So with teamwork and this eval example, we start arriving at these 2026 primitives that I keep talking about. And these model characteristics really change the way that we have to think about the product and how we have to develop the product. So the three examples that we've talked about, first, we have the dynamic subagent. And to provide a little bit more color here, basically no two subagents are the same. The main agent is the one

14:35

that is orchestrating this entirely on its own. It's configuring and prompting and seeding these subagents on the fly. They can operate in parallel. They can operate in different types of secure environments, be it a sandbox, be it a remote execution system. And they can all take on an infinite number of specialized roles. So the scaling story here is quite obvious. And from the last two examples, you can probably tell. As the model gets smarter, your team will become more specialized. It will become more collaborative. And ultimately, that means it will be capable of getting more complex work done for you. And now the second is this new concept.

15:10

We've alluded to it slightly in the past, but it's called sidecars. This is a new plug-in protocol that we're bringing to Antigravity. A sidecar process is essentially, it is a sidecar process. The naming sort of reflects what's going on under the hood. But it's a long-lived utility. And it's responsible for listening. It allows the model to listen to the outside world and set up its own triggers for things that might happen. For example, this could be SMS messages. This could be web hooks, cron jobs, hooking it up to GitHub PRs. The list goes on and on. But this is a generic plug-in primitive. Antigravity already uses sidecars for things that are time-based.

15:45

This is where the scheduled task cron concept comes from. But under the hood, this is all this new sidecar primitive. So we'll be releasing the spec for this so that you all can build on top of this new primitive later this summer. But there are some really, really cool ways that people internally have been using this sort of concept. And the third and final primitive is generative UI. So we hypothesize that human-written specialized UIs are kind of dead. Gemini Flash on Antigravity clocks in at almost 900 tokens a second. This is 10x faster than a lot of the other frontier model experiences. And in a matter of seconds, you're able to go from whatever you're thinking

16:22

inside of your head into a prompt, into a use case that is designed and embedded inside of your conversation view perfectly. And rather than rely on templates or even HTML files, Antigravity can render your generated UI in line. So you can do things like this and play Doom, but this also extends to things like bar charts, graphs, tables, anything that you would want to interact with and maybe inspect a little bit further than just a markdown file or just a conversation. And generative UI in many ways reminds me of the quote that the late Steve Jobs said when unveiling the iPhone. He justifies the removal of the keyboard and says, they all have these keyboards.

17:02

They are there whether you need them or not. And they all have these control buttons that are fixed in plastic and are the same for every application. In an analogous way, we built our product to dynamically scale with the needs of the agent. We skipped the heavy infrastructure and mechanical UIs in favor of sidecars and generative UI. And that creates a product experience that is not fixed in plastic. So subagents, sidecar triggers, and generative UI are the latest primitives that are powering Antigravity. We've tried our best to stay out of the way and let the model cook. And if you're building a product around an agent, you should consider what are the primitives

17:41

that are in my product and how might they scale with the model's intelligence. We all are familiar with shipping features is now quite easy with all of these new tools. And it's about deciding what features to actually add so that the model, so that your product can scale with the next release of the next model, which will inevitably be faster, better, and cheaper. And so with the right primitives, you as a builder or you as a product owner, you might be surprised at what the models can do. And in classic fashion, I'm going to keep using this slide until we've actually conquered the TPU crunch. So you can find me on Twitter. You can DM me for feedback.

18:16

We're always looking for new ideas on how to build the latest and greatest. Thank you for watching. Thank you, Swix and Ben for having me. It's always a joy to be here. And I'll be at the Antigravity booth if you want to talk further. If you want to get to know the product a bit more, some team members will be there. So thank you so much for your time. Excited to meet you all.

18:51

NextHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHeyHey you you

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note