Open Reader

Why (Senior) Engineers Struggle to Build AI Agents — Philipp Schmid, Google DeepMind

completed 10:39 May 30, 2026 Watch on YouTube

Current Status

completed

Video ID

3_gYbhABcAE

RAG / Chat

Enabled
Why (Senior) Engineers Struggle to Build AI Agents — Philipp Schmid, Google DeepMind
Description

A `deleteItem` endpoint is obvious to the developer who built it. An agent only sees the function schema and docstring. Philipp Schmid from Google DeepMind argues this is why senior engineers struggle most: they carry years of implicit context that agents do not, and design tools assuming it. He names four other shifts: text replaces structured state, errors are inputs not restart triggers (especially costly when an agent has been running for 15 minutes), evals replace unit tests because the right question is how often it works not whether a fixed input always produces a fixed output, and build to delete because you will rebuild the same agent with a better model anyway. Speaker info: - https://x.com/_philschmid - https://www.linkedin.com/in/philipp-schmid-a6a2bb196/ - https://github.com/philschmid

Summary

Generated by claude-sonnet-4-5

30-second take

Philipp Schmid (Google DeepMind, Gemini agents team) argues that senior engineers fail at agent development because they apply traditional software engineering mindsets—deterministic workflows, strict state management, unit tests—to fundamentally non-deterministic systems. His core thesis: building agents requires shifting from "traffic controller" (exact step control) to "dispatcher" (goal definition), accepting probabilistic success rates over binary pass/fail, and treating text/context as state instead of structured data. This matters because most agent failures stem from fighting the model's nature rather than designing for it.

Key takes

  • Text is the new state: Structured data models (booleans, enums) can't capture semantic meaning or dynamic context like "use Celsius except for cooking." Agents operate on unstructured text/context, requiring engineers to abandon rigid schema thinking and embrace semantic interfaces that let users approve a plan while adding constraints ("focus on US market, ignore California") in one natural input.
  • Hand over control, stop forcing workflows: Traditional customer support used intent classification → predefined churn flow, but agents should dynamically understand that a cancellation request might actually be solved by offering an alternative, changing the conversation entirely. Engineers struggle when they try to script exact step sequences instead of defining goals and trusting the model to route dynamically.
  • Errors are inputs, not restart triggers: When a 15-minute agent run fails mid-process, restarting from scratch wastes compute and loses context. Agents need error recovery built in—feed failures back to the model as context, let it work around issues, and keep forward progress rather than assuming cheap HTTP retries.
  • Move from unit tests to evals (success rate > binary pass): Agents are non-deterministic; same input ≠ same steps or output. Success requires measuring how often something works (e.g., 9/10 vs 1/10 reliability), using LLM-as-judge or human evals for subjective outcomes, and focusing on end results even if token consumption or step count varies per user.
  • APIs must be agent-ready with semantic docs: Internal APIs feel self-documenting to humans who built them ("delete_item" is obvious), but agents only see function schemas/docstrings without years of context. Engineers must write explicit, semantic tool definitions explaining what happens on failure, what IDs represent, etc.—not assume tribal knowledge.
  • Build to delete: Software is disposable now. Better models/agents will emerge constantly, so overengineering for permanence is wasteful. The mindset shift is rapid iteration and rebuilding over long-term architectural planning.

Useful details

  • Traffic controller vs dispatcher analogy: Traditional software = controlling streetlights, speed limits, exact roads (deterministic). Agents = telling a dispatcher "get me to London from Germany" and letting them choose train/plane/car dynamically.
  • Deep Research agent example: Returns a plan; old approach = binary accept/deny. New approach = approve and add constraints ("focus US, ignore California") in one semantic input instead of multi-step clarification loops.
  • Temperature preference example: Old user profile had a single Celsius/Fahrenheit flag. Agent systems need context-aware preferences ("Celsius normally, Fahrenheit for cooking") that can't be modeled as boolean fields.
  • Churn prevention scenario: Classification model detects cancellation intent → triggers predefined retention flow. Agent approach = understand semantic reason, offer dynamic alternatives, let conversation evolve beyond scripted paths.
  • Go error handling as model: Go treats errors and values equally in function returns; agents should treat tool failures as normal inputs to reason about, not exceptions to crash on.
  • Eval metrics over assertions: Instead of "assert output == expected," measure "this prompt succeeds 9/10 times" and determine acceptable reliability thresholds for production (subjective outcomes need LLM judges or expert review).
  • Docstring gap: delete_item(id) needs explicit docs like "ID is product UUID, returns 404 if not found, no undo" because agents lack developer context that makes endpoints "obvious."
  • Iterative loop: Define instructions → run → observe → adjust prompts/tools → rerun. Contrast with traditional: spec → code → test → deploy.

Caveats / counterpoints

  • Schmid provides no specific reliability thresholds (e.g., is 70% success production-ready? 90%?), leaving the "right balance" undefined.
  • No discussion of cost implications—agents consuming variable tokens and doing extra steps could blow budgets unpredictably.
  • Assumes trust in models is justified; doesn't address cases where models hallucinate or confidently do the wrong thing despite good prompting.
  • "Build to delete" mindset may conflict with enterprise needs for stable, auditable systems; no guidance on when permanence still matters.
  • Examples are all Google/DeepMind internal (Deep Research agent, Gemini API) without external validation or comparison to other frameworks.
  • Doesn't cover multi-agent coordination, security boundaries, or how to prevent agents from taking harmful actions when given control.

Ken relevance

High relevance for Ken's agent systems and AI ops work:

  • Agent GTM: The eval-over-tests and probabilistic success framing directly applies to selling agent reliability to enterprises. Ken can position evals/monitoring as the new testing layer clients must adopt.
  • Tool/API design: If Ken is building tools for agents (or advising portfolio companies doing so), the "semantic interface" requirement is critical—generic CRUD APIs won't cut it without rich docstrings and agent-first design.
  • Content/education angle: This mental model shift (dispatcher vs traffic controller, text as state) is a clear content hook for teaching engineers or investors why agent projects fail. Ken could package this as a framework.
  • Workflow implications: Ken's personal agent workflows should embrace error recovery (don't restart 15-min research from scratch) and dynamic context (not rigid branching logic).
  • Investing lens: When evaluating agent startups, look for teams building evals/observability, not just agent frameworks. The "build to delete" mindset means rapid iteration winners will dominate.

Watch verdict

Watch fully — Dense, opinionated practitioner advice from someone shipping agents at scale (Google DeepMind). The five mental model shifts are concrete enough to change how you design agent systems, not just abstract theory. Ten minutes well spent for anyone building or investing in agents.

Transcript

1727 words en Processed in 83.6s

Okay, cool. Awesome. Hi, everyone. My name is Philip. I work at DeepMind, everything related to agents on Gemini or Gemini API. So if you have some questions afterwards, some concerns, some bugs, some issue, please let me know. We're going to talk today, 10 minutes, about why engineers struggle to build agents. And I see this every day internally at Google, but also externally at Google. And I brought five examples on what's really different to how we built traditional software a few years ago and to now how we build agents. And if we, on a high level, compare them, right, when we wrote software, we created a spec, a PRD, wrote code, sometimes created tests to make sure our code works. We deployed it and then our user used it. And when building agents, things are a little bit different. We define instructions on what we want our agent to do. We run it, we observe what it does, we maybe adjust our prompts, maybe we adjust our tools, we run it again, and we have this iterative loop of how can we improve and make our agent way more reliable, which is very different to how we build software. And something I like to compare it to is traditional software is more like we acted as a traffic controller, right? We had control over the streetlights, over how fast you can go, which road you can use, basically how the car drives. And now with agents, we are more of a dispatcher. We tell the agent, hey, I want to go to London and I'm from Germany. I could use the train, I could fly, I could use my car and go under the water. And it's more about, okay, we define the goal on what we want the agent to do, but we don't define the exact step the agent needs to take to achieve that goal. And every one of you has probably seen in their coding agent that sometimes it does something very weird, but at the end it achieves the outcome. And that's what we want to do. So starting with the first example, text is our new state. Traditionally we had data structures and everything was mapped to Boolean or to flex we could check. So initially when we created, for example, a deep research agent, deep research agent returns a plan to you, okay, I'm going to research this and that. In traditional software we might have had an accept plan or deny plan, but we couldn't catch semantic meaning. And now what we have with LLMs is they can understand the semantic meaning. So for example, if I have a deep research request on doing some market research, I can approve the initial plan, but I can also at the same time provide additional information. So maybe I want to focus on the US market and ignore California. Maybe I want to provide something additional and not have this multiple steps, right? Traditionally, I would probably said decline and then it has a follow up. I might need it to provide more input, create a new plan and continue. And another good example is everything related to memory and personalization we do cannot really be mapped to data structures, right? The example I hear I have is like, I'm from Europe, so I mostly use Celsius, but what if I would like to use Fahrenheit for cooking, right? Previously, we might have had some flex on the user profile is Celsius or is Europe or use Fahrenheit, but I couldn't dynamically adjust based on the user preference based on what they provide. So really, it's all about text and context. It could be images, video, audio as well, but we no longer are really operating in those clear structured data concepts. The other thing is we should start handing over control and the trap or the example which we might have from previous customer support is like when a user reached out, hey, I want to cancel my subscription. I might have had a classification model which classified the intent, okay, the user wants to churn and then I had a predefined workflow of, okay, do you try to sell it? Do you cancel the subscription? But there was no dynamic option to react to it dynamically and maybe instead of going through the subscription cancel flow. What if your agent tries to understand the meaning and offer something in except to the subscription and the user changes their mind and now you have a whole different content and it's very hard to model all of those differences and uniqueness and all of those stateful workflows we had before. So we need to trust into the LLM or hand over control that we are no longer working in those purely deterministic environments. The third one is errors are just inputs. So if something in your agent flow fails, we need to treat it as a normal input, very similar to a user input. In Go, we already do this, right? A function call can be an error or can be a value and we treat them equally and we have to do this for agents very similarly. In the past HTTP requests were very cheap. When some search, some product search failed, you just rerun your request, you redid all of the work which was okay. But now if you have an agent which takes five minutes, 15 minutes and something in the flow breaks and you would start all over, you would need to spend a lot of compute again to do all of the previous steps and you also might lose the existing context. So we cannot just start over the whole process. We need to understand and treat errors differently, provide them back to the model, maybe have some other work around, some additional checks that we basically keep going forward in the flow and not start over from the beginning. The fourth example or step is we need to move from unit test to evals. So when building software before, right, we wrote integration tests, unit tests, smoke tests, and all kinds of different tests and we assume that when we provide input A for our code B, we will always get C as an output and that's no longer the case with agent. Agents are non-deterministic. We cannot always guarantee that the same input will lead to the same steps and the same result. So we need to move from unit tests to eval. We need to test how often something works because agents are only successful if they are really reliable, right? If you have a customer agent and the same prompt only works one out of ten times, it's nothing really you want to put in production and it becomes very flaky. So we need to test on evals on how many times it passes and compared to traditional software, results are very subjective, right? An outcome can be very different if you ask it to create a research report, if you ask it to create a customer feedback scenario, and we need more qualifying feedback. LLM as a judge or human expert for example is a good way and we always need to trace what the agent is doing, but we need to create on the output. Maybe the agent decides for one user it needs to do four more steps to do more research, then for the other user it consumes maybe a few more tokens, but at the end the outcome is really what we need to measure and want to measure in terms of success. And then the last part is agents evolve and APIs don't and if you have worked on the back end and if you have built an API, you might have seen a lot of methods API endpoints which feel very self-explaining to you like delete item feels very self-explaining if you are working on the product API. But an agent doesn't have the context and the background from all those years from you working on the API, so we need to build APIs or tools which are really agent ready, which are self-documenting with semantic interfaces. I would assume if you have a product microservice and you have delete item endpoint with an ID, you don't need to define a doc string what the ID is or what happens if something fails, but our agents only see the function schemas and the doc strings and the tool definitions. So on the first look they don't really see what the delete item method does. That's why we need to really adjust to hey, we need methods tools which are written for agents to be used and not assume long year developer expertise and people who have built the API. So to summarize everything we need to give trust but we also verify we should stop fighting the model. You should not try to force the model into this one specific workflow with step one do this, step two do the other thing. We need to preserve meaning. Everything is a context now. We no longer have those very well defined data structures for all of our applications. We need to design for recovery. Models are not perfect. Agents are not perfect, especially if we have longer running agents. There will be some very weird things happening. So you need to design for recovery. We need to evaluate and then don't only assert agents are not 100% reliable. We need to find the right balance between how many times our run need to be successful to provide it to the user. And last but not least, build to delete. The bit of lesson is what everyone of us is learning is software is disposable. We are going to rebuild many times the same things with better models, better agents. Things will change. And yes, it's also available on my blog. So if you want to look a bit deeper with some code examples. And if not, if you have any questions, feel free to reach out to me and perfect on time. Thanks. Thank you. bugs, some issue, please let me know. We're going to talk today, 10 minutes, about why engineers struggle to build agents. And I see this every day internally at Google, but also externally at Google. And I brought five examples on what's really different to how we built traditional software a few years ago and to now how we build agents. And if we, on a high level, compare them, right, when we wrote software, we created a spec, a PRD, wrote code, sometimes created tests to make sure our code works. We deployed it and then our user used it. And when building agents, things are a little bit different. We define instructions on what we want our agent to do. We run it, we observe what it does, we maybe adjust our prompts, maybe we adjust our tools, we run it again, and we have like this iterative loop of how can we improve and make our agent way more reliable, which is very different to how we build software. And like, something I like to compare it to is like traditional software is more like we acted as a traffic controller, right? We had control over the streetlights, over how fast you can go, which road you can use, basically how the car drives. And now with agents, we are more of a dispatcher. We tell the agent, hey, I want to go to London and I'm from like Germany. I could use the train, I could fly, I could use my car and go like under the water. And it's more about, okay, we define the goal on what we want the agent to do, but we don't define the exact step the agent needs to take to achieve that goal. And I mean, every one of you has probably seen in their coding agent that sometimes it does something very weird, but at the end it achieves the outcome. And that's what we want to do. So starting with the first example, text is our new state. I mean, traditionally we had data structures and everything was kind of mapped to Boolean or to like flex we could check. So initially when we created, for example, a deep research agent, deep research agent returns a plan to you, okay, I'm going to research this and that. In traditional software we might have had an accept plan or deny plan, but we couldn't catch semantic meaning. And now what we have with LLMs is they can understand the semantic meaning. So for example, if I have a deep research request on like doing some market research, I can approve the initial plan, but I can also on the same time provide additional information. So maybe I want to focus on like the US market and ignore California. Maybe I want to provide something additional and not have like this multiple steps, right? Traditionally, I would probably said decline and then it has a follow up. I might need it to provide more input, create a new plan and continue. And another good example is everything related to memory and personalization we do cannot really be mapped to data structures, right? The example I hear I have is like, I'm from Europe, so I mostly use Celsius, but what if I would like to use Fahrenheit for cooking, right? Previously, we might had some flex on like the user profile is Celsius or is Europe or use Fahrenheit, but I couldn't like dynamically adjust based on the user preference based on what they provide. So really, it's all about text and context. I mean, it could be images, video, audio as well, but we no longer are really operating in those clear structured data concepts. The other thing is, we should start handing over control and the trap or the example which we might have from like previous customer support is like when a user reached out, hey, I want to cancel my subscription. I might have had a classification model which kind of classified the intent, okay, the user wants to churn and then I had a predefined workflow of, okay, do you try to sell it? Do you cancel the subscription? But there was no like dynamic kind of option to react to it dynamically and maybe instead of like we're going through the subscription cancel flow. What if your agent like kind of tries to understand the meaning and like offer something in except to like the subscription and the user changes their mind and now you have like a whole different content and it's very hard to model all of those differences and uniqueness and to like all of those stateful workflows we had before. So, we need to like trust into the LLM or like hand over control that we are no longer working in those purely deterministic environments. The third one is errors are just inputs. So, if something in your agent flow fails, we need to treat it as a normal input as very similar to a user input. In Go, we already do this, right? A function call can be an error or can be a value and we treat them kind of equally and we have to do this for agents very similarly. In the past HTTP requests were very cheap. When some search, some product search failed, you just rerun your request, you redid all of the work which was okay. But now, if you have like an agent which takes five minutes, 15 minutes and something in the flow breaks and you would start all over, you would need to spend a lot of compute again to like do all of the previous steps and you also might lose the existing context. So, we cannot like just start over the whole process. We need to kind of understand and treat errors differently, provide them back to the model, maybe have some other work around, some additional checks that we basically keep going forward in the flow and not like starting over from the beginning. The fourth example or step is like we need to move from unit test to evals. So, when building software before, right, we wrote integration tests, unit tests, smoke tests, and all kinds of different tests and we assume that when we provide input A for our code B, we will always get C as an output and that's not longer the case with agent. Agents are non-deterministic. We cannot always guarantee that the same input will lead to the same steps and the same result. So, we need to move from unit tests to eval. We need to test how often something works because agents are only successful if they are really reliable, right? If you have a customer agent and the same prompt only works one out of ten times, it's nothing really you want to put in production and it becomes very flaky. So, we need to test on evals on how many times it passes and compared to traditional software, results are very subjective, right? An outcome can be very different if you ask it to create a research report, if you ask it to create a customer feedback kind of scenario, and we need more like qualifying feedback. LLM as a judge or human expert for example is a good way and we always need to trace what the agent is doing, but we need to create on the output. Maybe the agent decides for like one user it needs to do like four more steps to do more research, then for the other user it consumes maybe a few more tokens, but at the end the outcome is really what we need to measure and want to measure in terms of success. And then the last part is agents evolve and APIs don't and if you have worked on the back end and if you have built an API, you might have seen a lot of methods API endpoints which feel very self-explaining to you like delete item feels very self-explaining if you are working on like the product API. But an agent doesn't have the context and the background from all those years from you working on the API, so we need to build APIs or tools which are really agent ready, which are self-documenting with semantic interfaces. I would assume if you have like a product microservice and you have delete item endpoint with an ID, you don't need to like define a doc string what the ID is or what happens if something fails, but our agents only see like the function schemas and the doc strings and the tool definitions. So on the first look they don't really see what the delete item method does. That's why we need to really adjust to hey, we need methods tools which are written for agents to be used and not assume long year developer expertise and people who have built the API. So to summarize everything we need to give trust but we also verify we should stop fighting the model. You should not like try to force the model into this one specific workflow with step one do this, step two do the other thing. We need to preserve meaning. Everything is a context now. We no longer have those very well defined data structures for all of our applications. We need to design for recovery. Models are not perfect. Agents are not perfect, especially if we have longer running agents. There will be some very weird things happening. So you need to design for recovery. We need to evaluate and then don't only assert agents are not 100% reliable. We need to find the right balance between how many times our run need to be successful to provide it to the user. And last but not least, build to delete. The bit of lesson is what everyone of us is learning is like software is disposable. We are going to rebuild many, many times the same things with better models, better agents. Things will change. And yes, it's also available on my blog. So if you want to like look a bit deeper with some code examples. And if not, if you have any questions, feel free to reach out to me and perfect on time. Thanks. Thank you.