Open Reader

Don't Build Agents You Can't Answer For — Addy Osmani

completed 18:26 Jul 14, 2026 Watch on YouTube

Current Status

completed

Video ID

n97BCfyFIvw

RAG / Chat

Enabled
Don't Build Agents You Can't Answer For — Addy Osmani
Description

For his closing keynote, Addy Osmani explores the evolving role of software engineers in the age of AI agents. He argues that as coding tasks become increasingly automated, the true value of an engineer shifts from mere code production to accountability, judgment, and system ownership. https://addyosmani.com/ https://x.com/addyosmani/status/2074927530482835916 Timestamps 0:00 Introduction and the human side of engineering 1:46 Rebundling roles and ownership of systems 2:34 Harnesses, loop engineering, and software factories 3:34 The shift to answerability as an engineering requirement 4:26 Reviewing AI-assisted code and organizational bottlenecks 5:55 Redefining leverage through human judgment 6:15 Alpha, decay, and the role of "taste" 8:49 Defining the modern software engineer 9:50 Risks to avoid: cognitive debt and surrender 11:51 Orchestration tax and system design 12:39 Accountability as the foundation for scaling 13:16 Career math: credibility vs. capability 14:13 High agency and the decision-making ladder 15:13 Defining the boundary between agents and humans 16:13 Operational rule: explain it or don't ship it 17:20 Future outlook: unlocking latent demand

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Engineers must shift from task execution to evidence-based accountability; agents automate work but humans must own decisions, verify output, and answer for production choices.
  • Why it matters: Defines the engineering role in AI-assisted development where 96% distrust AI code but only 50% verify before commit, creating organizational risk at scale.
  • Best use: Strategic framework for Ken's agent system governance, team capability planning, and identifying what truly differentiates human judgment in automated workflows.

Executive Summary

Addy Osmani argues that the future engineer's value lies not in coding capability but in owning verdicts—deciding what's worth building, verifying evidence from agents, and accepting responsibility for production decisions. He frames this through three architectural layers: harness engineering (model + context/tools), loop engineering (iterative prompting/checking systems), and software factories (agents in inner loops, humans at high-leverage checkpoints). The core tension: AI-assisted code is becoming normal (Sonar 2025 data shows it's no longer marginal) but 96% of engineers don't fully trust it, yet only ~50% always verify before committing—creating a 'distrust without bandwidth' gap.

Osmani introduces two career concepts: 'alpha' (gap between human and model capability) and 'decay' (how fast that gap closes). Speed, recall, and even verification are decaying as capabilities; taste decays more slowly but still resets as models learn preferences. The scarce resource becomes 'judgment backed by evidence' and 'answerability'—who can explain intent, inspect evidence, accept risk, and improve systems when decisions fail. He warns against three failure modes: cognitive debt (losing understanding of systems you ship), cognitive surrender (blindly accepting AI responses—Wharton study found 73% accepted wrong AI answers with higher confidence), and orchestration tax (human attention doesn't parallelize even when agent count does).

The operational model separates inner loops (agents investigate/implement/test/report) from outer loops (humans decide/verify/approve/own). The boundary isn't 'human reviews AI output' but 'evidence meeting responsibility.' Key rule: 'Explain it or don't ship it'—someone must understand work well enough to defend it, like traditional OWNERS files in codebases. Osmani rejects the idea that engineering work will shrink; every time software got easier to write (high-level languages, frameworks, cloud, low-code), latent demand appeared and total work expanded. Agents will shift the bottleneck from 'can we build this?' to 'should this exist and can we answer for it?'

He positions 'high agency' not as doing everything yourself but as 'ownership with judgment attached'—knowing when to delegate, inspect, stop, and sign your name. The agency ladder ranges from flagging problems (low) to executing/diagnosing/proposing (mid) to discernment (high)—deciding which problems deserve investment at all. Clean code matters not just for humans but for agents: Sonar research found clean vs messy repos had similar pass rates, but clean code used fewer tokens and needed fewer revisits. The talk is a strategic reframing for teams adopting agent workflows at scale, especially relevant for governance, hiring, and defining what 'engineer' means when models can generate most implementation.

Key Takeaways

  • Claim: Quality produces evidence, but verdicts assign responsibility—answerability lets engineers stand behind production decisions. | Evidence: Osmani defines 'verdict' not as judgment but as accountability: 'Does something ship? Do we block/redirect/accept risk?' The engineer owns production outcomes even when agents generate the work. | Caveat: No explicit threshold given for what counts as sufficient evidence or how to measure answerability in practice. | Implication: Ken's agent systems need structured evidence trails (diffs, tests, logs, rationale, traces) so humans can verify and own decisions, not just review outputs. | Timestamp: timestamp unavailable
  • Claim: 96% of engineers don't fully trust AI code, but only ~50% always verify before committing—creating 'distrust without bandwidth.' | Evidence: Sonar 2025 survey data; Osmani notes the danger is 'distrust without bandwidth' where adoption moves faster than verification capacity or governance policy. | Caveat: No detail on survey methodology, sample size, or whether this applies equally across company types/maturity levels. | Implication: Safety comes from making verification cheaper/clearer/harder to skip; Ken should design systems where evidence is legible and verification is built into the workflow, not optional. | Timestamp: timestamp unavailable
  • Claim: Clean code helps the next agent, not just the next human—clean repos use fewer tokens and need fewer revisits. | Evidence: Sonar research found clean vs messy repos had roughly the same pass rates, but clean code cost fewer tokens and required fewer revisits. | Caveat: No definition of 'clean' vs 'messy' code provided; unclear if this holds across all model types or tasks. | Implication: Maintainability and code quality directly impact agent efficiency and cost; Ken should prioritize clean architecture/naming/structure even in agent-generated code to reduce future token spend and rework. | Timestamp: timestamp unavailable
  • Claim: Wharton study: When AI was wrong, 73% of people still thought they picked the correct answer and felt more confident. | Evidence: Direct citation of Wharton research showing 'borrowed confidence'—users adopt AI's wrong answers and experience increased certainty. | Caveat: No details on task type, model used, or participant background; may not generalize to all engineering contexts. | Implication: Cognitive surrender is a real risk; Ken's systems should force evidence inspection and human reasoning steps, not just yes/no approvals, to counter false confidence. | Timestamp: timestamp unavailable
  • Claim: 'Alpha' (human capability gap over models) decays faster than 'signature' (credibility/accountability); skills earn leverage, accountability turns leverage into trust. | Evidence: Osmani lists capabilities with decay rates: speed (decayed), recall (harnesses have memory), verification (moving into evals/static checks), taste (decays slowly but resets as models learn), judgment (slope not wall). Half-life of edge = one model release; half-life of signature = much longer. | Caveat: No empirical data on decay rates or timeline predictions; framework is conceptual. | Implication: Ken should invest in building reputation/track record for good judgment calls rather than protecting specific technical skills; the career bet is on being the person who can own outcomes, not the person who can code fastest. | Timestamp: timestamp unavailable
  • Claim: High agency = ownership with judgment attached—knowing when to delegate, inspect, stop, and sign your name; the agency ladder peaks at discernment (deciding what's worth solving). | Evidence: Agency ladder: flag problem (low) → execute/diagnose/propose/recommend/resolve → discernment (find problem, decide if worth investment). 'When agents make more paths possible, agency is deciding which paths deserve ownership, not chasing every path.' | Caveat: No operational playbook for how to build or assess discernment capability; assumes some problems are not worth solving but doesn't give criteria. | Implication: Ken should hire/develop for judgment and ownership, not just execution speed; the valuable skill is knowing what not to build or when to stop, especially as agent throughput increases. | Timestamp: timestamp unavailable
  • Claim: Operational rule: 'Explain it or don't ship it'—someone must understand work well enough to defend it, like OWNERS files in large codebases. | Evidence: Analogizes to enterprise OWNERS files where certain people are on the hook for subsystems; 'Model might write code, but question is can you explain changes, have evidence, understand risks.' | Caveat: No guidance on how much explanation is enough or how to enforce this rule in fast-moving teams. | Implication: Ken should implement ownership assignment and evidence requirements in agent workflows; every shipped change needs a named human who can defend the decision and explain the system. | Timestamp: timestamp unavailable

Detailed Brief

Evolution of Engineering Roles: From Tasks to Accountability

  • Claims: Engineer of the future owns evidence, understanding, and verdicts for automated work; Boris Churney taxonomy: roles rebundle around prototype/build/sweep/grow/maintain modes; question shifts from 'what's your title' to 'what part of system can you own'; Harness engineering = model + context/tools/file system/git; turns intelligence into something delegable; Loop engineering = systems that keep prompting/checking/remembering/deciding; agents become infrastructure; Software factories = agents in inner loop, evidence out, humans make production decisions
  • Evidence: Quality produces evidence; verdict assigns responsibility; answerability lets you stand behind verdict; Sonar 2026 survey: AI-assisted code no longer marginal, increasingly large role in codebases; Harness = context + tools, makes models delegable; Dex's talk covered software factories as pattern
  • Caveats: No empirical definition of 'answerability' or how to measure it; Churney taxonomy is descriptive, not prescriptive—doesn't say how to choose mode or who owns overlaps; Software factory framing assumes human judgment remains highest leverage, but doesn't address scenarios where agent judgment could be sufficient
  • Implications: Ken should design agent systems with evidence capture at every step—logs, traces, rationale—so humans can form verdicts; Hiring/role design should focus on system ownership capability, not just technical skill lists; Loop design and evidence design become core engineering work; build workflows where evidence is legible and responsibility is clear

The Distrust-Without-Bandwidth Gap and Verification Economics

  • Claims: 96% of engineers don't fully trust AI code, but only ~50% always verify before committing; Making generation cheaper doesn't automatically make review cheaper; Safety comes from making verification cheaper, clearer, harder to skip; Clean code helps next agent: clean repos = same pass rate but fewer tokens, fewer revisits; Governance can't catch up to adoption speed; hard questions arise: did model touch this file, what constraints, what evidence, what risk, who owned result?
  • Evidence: Sonar survey data on trust vs verification rates; Sonar research: clean vs messy repos, token/revisit cost difference; Adoption moving faster than policy-setting in organizations
  • Caveats: No details on what 'verification' means (manual review? tests? static analysis? all of above?); Survey sample/methodology not described; Clean code definition not provided; may vary by language/domain
  • Implications: Ken's systems should make evidence inspection fast and legible—dashboards, summaries, diff annotations—so verification doesn't become bottleneck; Invest in automated checks (evals, static analysis, model critique) to reduce human review load; Prioritize code quality/maintainability even in generated code; it pays off in agent efficiency and lower rework costs; Governance frameworks should define evidence requirements and ownership upfront, not retrofit after adoption

Three Failure Modes: Cognitive Debt, Surrender, Orchestration Tax

  • Claims: Cognitive debt = erosion of understanding/memory around problem-solving; gap between code existing and humans understanding it; Delegation depth matters: 30-second run = interaction, hour/day tasks = work stream; review can't be glance at end, must be control system; Cognitive surrender = blindly accepting AI responses; 'your answer is now my answer before I formed opinion'; Wharton study: 73% accepted wrong AI answers with higher confidence—failure mode is borrowed confidence; Orchestration tax = human cognitive bandwidth doesn't parallelize; every loop causes more decisions to route/merge/verify/integrate; Fix is not fewer agents but designing attention like a system: where you enter, what you acquire, what you reuse
  • Evidence: Cognitive debt manifests as build passes tests, PR merges, but team loses ability to explain system; Wharton study on borrowed confidence with wrong AI answers; Observation: people running hundreds/thousands of agents but attention doesn't scale
  • Caveats: No data on how common cognitive debt is or how to measure it quantitatively; Wharton study details missing (task type, participant background); No playbook for 'designing attention like a system'—concept is abstract
  • Implications: Ken should limit delegation depth for critical systems—long-horizon tasks need checkpoints, not just end review; Build evidence trails that preserve rationale/context so future humans can reconstruct decision logic; Force human reasoning steps before approval to counter borrowed confidence; don't just show output, show evidence + ask for judgment; Architect attention flows: where humans enter loop, what gets escalated, what gets auto-approved with evidence; make intentional choices about cognitive load

Alpha, Decay, and the Half-Life of Edges

  • Claims: Alpha = gap between human capability and model capability; Decay = clock on that gap; If special capability is what makes you valuable, frontier will eventually come for it; Speed (decayed), recall (harnesses have memory), verification (moving into evals/static checks), taste (decays slowly, resets as models learn preferences), judgment (slope not wall); Mitchell Hashimoto definition: taste = ability to make high-quality qualitative judgments where no objective metric exists yet; Taste is alpha too; not eternal moat, but matters when production gets cheaper; scarce skill = knowing which option deserves to exist; Best version of taste = making better calls, leaving examples team/system can learn from; Half-life of edge = one model release; half-life of signature (credibility/expertise) = much longer
  • Evidence: Paul Graham: 'When anyone can make anything, choosing what to make becomes important'; Mitchell Hashimoto: taste as qualitative judgment before metrics exist; Career math: skills can earn leverage, accountability turns leverage into trust
  • Caveats: No timeline predictions for decay rates; Taste definition still somewhat abstract—how to operationalize or teach it?; Assumes signature/credibility compound over time, but no data on how long that lasts or what breaks it
  • Implications: Ken should stop optimizing for speed/recall/raw capability (models catching up fast) and focus on building reputation for good judgment; Invest in documenting decisions, building examples, creating feedback loops where taste improves over time; Hire for judgment and ownership, not just technical skill; train teams to explain why decisions were made, not just what was built; Strategy = keep moving edges up a level; don't cling to any one capability, focus on what only humans can be answerable for

Inner Loop (Capability) vs Outer Loop (Agency) Operating Model

  • Claims: Agents run inner loop: investigate, implement, test, report; Humans run outer loop: decide, verify, approve, own; Inner loop = capability; outer loop = agency; Boundary is not 'human reviews AI output' but 'evidence meets responsibility'; Agent returns evidence (diffs, tests, logs, rationale, traces, trajectories, screenshots); engineering decides if work worth doing, verifies evidence, approves/redirects/owns production; Operational rule: 'Explain it or don't ship it'—someone must understand work well enough to defend it; Analogize to OWNERS files: who is accountable for that part of architecture/codebase?; Automation moves floor; engineering moves up a level; new work = loop design, evidence design, brownfield stewardship; Fewer keystrokes ≠ less engineering; more surface area needing taste/verification/ownership/care
  • Evidence: Historical pattern: every time software got easier (high-level languages, frameworks, cloud, low-code), latent demand appeared, total work expanded; Agents will move bottleneck from 'can we build this?' to 'should this exist and can we answer for it?'
  • Caveats: No concrete examples of evidence design or loop design practices; Doesn't address edge cases where human can't understand complex agent work—what then?; Historical analogy (more abstraction = more demand) may not hold if agents commoditize entire problem domains
  • Implications: Ken should structure agent workflows with explicit inner/outer loop separation; agents produce evidence packages, humans make go/no-go decisions; Evidence packages should be legible, not just raw outputs—summaries, risk assessments, test results, rationale; Define ownership model upfront: every system component/decision type has a named human accountable; Invest in evidence design (what to capture, how to present) and loop design (where humans enter, what gets escalated) as core engineering disciplines; Expect demand for software to grow, not shrink; prepare for more systems to manage, not fewer, even as agents automate implementation

Notable Concepts & Terms

  • Answerability: The ability to stand behind a verdict/decision with evidence and understanding; what lets engineers own production outcomes even when agents generate the work.
  • Verdict: Not 'judgment' in abstract sense but production decision: ship/block/redirect/accept risk. Assigns responsibility to a named human.
  • Alpha and Decay: Alpha = gap between human and model capability. Decay = how fast that gap closes. Framework for thinking about career durability as models improve.
  • Cognitive Debt: Erosion of understanding/memory around systems; gap between code existing and humans genuinely understanding it. Delegation depth risk.
  • Cognitive Surrender: Blindly accepting AI responses without forming independent opinion; 'borrowed confidence' where wrong answers feel more certain (Wharton study: 73% accepted wrong AI with higher confidence).
  • Orchestration Tax: Cost of managing multiple agents/loops in parallel; human attention doesn't parallelize, so every loop adds routing/merging/verification overhead.
  • Harness Engineering: Model + context/tools/file system/git/harness that turns intelligence into something delegable. First evolution from 'model as whole story.'
  • Loop Engineering: Designing systems that keep prompting, checking, remembering, deciding next steps; when agents become infrastructure, not just single-run prompts.
  • Software Factory: Architecture where agents run inner loop (investigate/implement/test/report), humans make production decisions at high-leverage checkpoints.
  • Inner Loop vs Outer Loop: Inner = agents investigate/implement/test/report (capability). Outer = humans decide/verify/approve/own (agency). Boundary is evidence meeting responsibility.
  • High Agency: Ownership with judgment attached—knowing when to delegate, inspect, stop, sign your name. Agency ladder peaks at discernment: deciding what's worth solving.
  • Taste (Mitchell Hashimoto definition): Ability to make high-quality qualitative judgments where no objective metric exists yet. Matters before benchmarks/market votes, but still decays as models learn preferences.

Operator Notes / Why Ken Should Care

  • This is a governance and capability framework for Ken's agent systems. Osmani is saying: agents will ship more than any team can review, so design for evidence legibility, not heroic manual review.
  • The distrust-without-bandwidth gap (96% skeptical, 50% verify) is an operational crisis waiting to happen. Ken should build systems where verification is automatic/cheap/required, not optional.
  • Clean code = fewer tokens + fewer revisits = lower cost for agents. Ken should enforce code quality even in generated code; it's not just aesthetic, it's economic.
  • Cognitive surrender (73% accept wrong AI with higher confidence) means yes/no approvals aren't enough. Ken's systems should force evidence inspection and reasoning steps to counter false confidence.
  • The inner/outer loop model is practical: agents produce evidence packages, humans make production decisions. Ken should structure workflows this way with explicit handoffs and evidence requirements.
  • Alpha/decay framing says: stop optimizing for speed/recall (models catching up), focus on judgment/accountability (signature, not skill). Ken should hire/develop for ownership capability.
  • Historical pattern: every abstraction layer (languages/frameworks/cloud/low-code) expanded demand, didn't shrink it. Expect more systems to manage, not fewer, as agents lower implementation cost.
  • Osmani's 'explain it or don't ship it' rule is actionable: every agent-generated change needs a named human who can defend the decision. Ken should implement OWNERS-style accountability in agent workflows.
  • Evidence design and loop design are new core disciplines. Ken should invest in how to capture/present evidence and where humans enter loops, not just in agent throughput.
  • For content/investing/GTM: answerability applies beyond code. If agent generates investment thesis, marketing copy, content brief, same rule applies—who can explain and defend the output? Ken should build evidence trails everywhere, not just in engineering.

Watch Map

  • timestamp unavailable: Introduction: human side before architecture; engineer of future owns evidence/understanding/verdicts
  • timestamp unavailable: Boris Churney taxonomy: roles rebundle around prototype/build/sweep/grow/maintain; harness/loop/factory evolution
  • timestamp unavailable: Distrust-without-bandwidth gap: 96% don't trust AI code, 50% verify; Sonar data on AI code becoming normal
  • timestamp unavailable: Three failure modes: cognitive debt, cognitive surrender (Wharton 73% study), orchestration tax
  • timestamp unavailable: Alpha and decay: speed/recall/verification decaying, taste decays slowly; half-life of edge vs signature
  • timestamp unavailable: Mitchell Hashimoto on taste; Paul Graham on choosing what to make; taste as alpha, not eternal moat
  • timestamp unavailable: High agency ladder: flag → execute → diagnose → propose → resolve → discernment (decide what's worth solving)
  • timestamp unavailable: Inner loop (agents investigate/implement/test/report) vs outer loop (humans decide/verify/approve/own)
  • timestamp unavailable: Operational rule: 'Explain it or don't ship it'; OWNERS file analogy; boundary is evidence + responsibility
  • timestamp unavailable: Closing: automation moves floor, engineering moves up; historical pattern = lower cost → latent demand appears; bottleneck shifts to 'should this exist and can we answer for it'

Source/Metadata

  • Title: Don't Build Agents You Can't Answer For — Addy Osmani
  • Transcript words: 3629
  • Duration seconds: 1106
  • Timestamp note: Timestamps unavailable in transcript; watch_map reflects logical flow/topics based on content sequence.

Transcript

3025 words en Processed in 146.3s

Howdy folks! So good afternoon or good whatever time it is when you're watching this on YouTube. I'm really excited to be here and today I want to talk to you about what it takes to keep the human in the loop where engineering is concerned. I really want to start with the human side before we talk about the architecture here. I think that the engineer of the future is going to be really defined by the person who is able to choose what is worth doing. They're going to own the evidence, they're going to own the understanding as well as the verdict around increasingly automated work that's being done by agents. Now when I use the term verdict, I don't mean that we're suddenly all going to be Judge Judy. We're not. But what I mean really is something just a little bit different. I mean we're going to be accountable for the production decisions. Does something shift? Do we block it? Do we redirect it or accept the risk? Quality is something that we all talk about a lot. But quality produces evidence. A verdict assigns responsibility. And answerability is really what lets us stand behind a verdict. And this of course is not the only way that our industry is starting to think about our roles evolving. Boris Churney recently put some useful language around what many teams are starting to feel. The old craft boundaries are getting blurry and roles are rebundling around the work itself. And the important question here becomes a lot less about what is your title and more what part of the system can you own? Now I like this taxonomy quite a lot. It's optimistic without being overly vague. So things like prototype, build, sweep, grow and maintain. And these are real engineering modes. Agents are going to help with all of them. But the scarce thing is not merely doing the task. It's going to be knowing which mode your product needs and what quality bar applies. And who owns the result at the end of the day. Now we've been talking about harnesses and loop engineering and software factories over the last couple of days. We can talk why this shift is happening. We've moved past the model as the whole story. Right? With harness engineering, the coding agent is the model plus the harness around it. Right? Your context, your tools, your file system, git. And the harness is what turns intelligence into something that you can delegate to. The next move was loop engineering. Where we weren't just prompting one run anymore. We were designing systems that kept prompting, checking and remembering and deciding what happened next. And that's really when agents started to feel like infrastructure. And once you start putting all of those things together, you get that software factory. Dex covered this well in his talk. But you have agents that are running inside that inner loop. And evidence that comes out. Humans still end up making the production decisions in this loop. And the wind really isn't moving us from it. The wind is moving human judgments the highest leverage checkpoint, I think. And this is why it starts to matter now. AI generated and AI assisted code is becoming normal code for a lot of us. One of Sonar's 2026 survey said that AI assisted code is no longer marginal. It's increasingly having a large role in our code bases. And once that happens, answerability stops being this philosophical world. It becomes an engineering requirement. And there's a quality point here as well. Right? We used to care about clean code. Code that people could read. But cleaner code is actually not just going to help the next human and the next person on your teams. It actually helps the next agent. Another one of Sonar's research studies found that clean and messy repos had roughly the same pass rates. But clean code actually used fewer tokens and cost fewer revisits. So there's a lot of benefit to maintainability that can fuel efficiency for your factories. Now, making generation cheaper does not automatically make review cheaper. Right? I think a lot of us are facing this moment. And we know that engineers are not naive. The Sonar numbers say that almost everybody is skeptical of AI code. Now, I love working in my software factory. I love building my engineering loops. But the problem is still capacity. If 96% of people don't fully trust that code, but only about half always verify before committing, we have this danger that we've got distrust without bandwidth. And so safety comes from making verification cheaper, clearer, and harder for people to skip. And if you zoom out from the individual reviewer to the organization, review and validation start becoming a bottleneck when governance isn't able to catch up. And adoption is already moving way faster than any company can go and set their policies. And this means that we have some hard questions we have to deal with. Like, did a model actually touch this file? And the hard questions are also like, what constraints guided that work? What evidence was produced? What risk was accepted? And who owned the result? Now, the agent can ship more than any of us can review. Right? So what are we still good for? I think it's a question that's on a lot of our minds. Right? And if Homer Simpson's experience automating computers can teach us anything, maybe this is our future. I don't think it is. But it's one direction things could take. Now, let's try that again. If change is where humans enter the loop. If generation scales faster than comprehension, the scarce resource becomes judgment that's backed by evidence. So the question is no longer how much can the agent do, but where does human judgment still create leverage? Now, I want to talk to you about two terms that I'm going to use for the career part of this talk. Alpha and decay. Alpha is the gap between what you can do today and what current models can do. That gap is a very real thing. And decay is the clock on that gap. If the thing that makes you special is a capability, the frontier is eventually going to come for it. Right? And there's a whole conversation around this. This is one of the reasons why taste keeps coming up. Paul Graham had a point here that I think is very right. When anyone can make anything, choosing what to make becomes very important. And I buy that. But I also think that we have to be very careful because taste can become a magic word for whatever part of the work we don't want to explain just yet. Mitchell Hashimoto gave us a more useful version of this definition. Taste is the ability to make high quality qualitative judgments where no objective metric exists yet. That matters because it puts taste before the benchmark and before the market has fully voted. When you try out a model and you see the kind of UX and the kind of experiences that it builds, you can often tell when you think it has taste or lacks taste or when there's a gap there that humans can fill. Now, this is also only useful if we can turn some of this concept around taste into critique, examples and better judgment over time. So, yes, taste matters when production gets cheaper. And if anyone can generate 10 options, the scarce skill is really knowing which option deserves to exist. But taste is not some eternal moat. It's alpha as well. Now, the people with taste are still going to matter. I personally think they're still going to matter for a long time. But the best version of that skill is not mystique. It's making better calls and leaving behind examples that your team and the system can learn from. Now, this is also only useful if we can turn some of this concept around taste into critique, examples and better judgment over time. So, yes, taste matters when production gets cheaper. And if anyone can generate 10 options, the scarce skill is really knowing which option deserves to exist. But taste is not some eternal moat. It's alpha as well. Now, the people with taste are still going to matter. I personally think they're still going to matter for a long time. But the best version of that skill is not mystique. It's making better calls and leaving behind examples that your team and the system can learn from. Now, let's apply the decay test. Well, we used to have speed that decayed. We used to have recall. Harnesses have memory. Verification is moving into harnesses, evals, static checks and model critique. Taste, I continue to think this is going to decay much more slowly. But it still resets as models learn from examples and preferences. Even judgment in some ways is a slope rather than a wall. So the strategy is not to cling to any one capability. It's for us to keep moving our edges up a level. So this is one of the reasons why what can the agent do is not the best strategic question anymore. The list of things that agents can't do just keeps shrinking. The better question for us is really what can only a human be answerable for? Not because any of us are magical in any way, but because some decisions actually require ownership. They require context, risk acceptance, and responsibility after that work ships. This is why the word engineer has to get just a little bit stricter. More people than ever can now make computers do things. And I think that's truly awesome. The total addressable market for builders has never been larger. And that's so cool. But it's a huge expansion of leverage. An engineer is not merely somebody who can code and get things to exist. An engineer can reason about systems. They think about constraints, they defend trade-offs, they can manage risk, and they're the person that can be reached out to when things start to break. So what are things that engineers should avoid if we want to stay effective and accountable in this moment? Well, the first thing to avoid really is cognitive debt. Now, cognitive debt is the erosion of your understanding and memory around how to solve problems. I think a lot of us start to feel this the more that we're using agents every single day. I know that I feel this a lot. And it's because we're deferring more and more to AI to solve our problems. For code, it's the gap between how much code exists in your repo and how much any human on your team genuinely understands. And this is why things like delegation depth end up mattering. You can have a build that passes your tests, a PR that you can merge, but your team can still end up losing its ability to actually explain the system that they are shipping to production. Now, a very real pressure is also how much we delegate. So agents can now stay inside the system long enough for the human to lose the thread. So a 30 second run can feel like an interaction, but an hour or a day scale tasks or something long horizon, that's a work stream. And when tasks can end up lasting that long, especially when you begin running many of them in parallel, review can't just be a glance at the end, it has to become a whole control system. The second thing to avoid is cognitive surrender. Now, this is when you blindly accept AI's responses. Delegation is important because delegation says, do the work, then show me enough evidence that I can judge it. I still make a judgment in that situation. Surrender is really saying, hey, your answer is now my answer before I have formed any opinions myself. Now, Wharton did a study that offers us a warning light here. When AI was wrong, 73% of people still thought that they picked the wrong answer and they felt more sure. So the failure mode is not using AI, but it's borrowed confidence. The third thing to avoid is orchestration tax. Now, if you've been in the Bay Area, you will see people who for better or worse are still walking around with their laptops open or are talking to you about cloud agents and we're increasingly trying to run more and more in parallel or telling each other that we're shipping with hundreds of agents or thousands of agents. More AI agents running does not mean that there is more of you available. Your cognitive bandwidth does not parallelize. So every loop that you create ends up causing more decisions to route, merge, verify and integrate. And the fix is not necessarily fewer agents, but it's about designing your attention like a system. Like where you enter, what you acquire, what you reuse. You just want to be very intentional about it. Now accountability can be a scary word for a lot of people. And I wouldn't be surprised if it made you want to go hide in the bushes and just tell your agent to deal with it. But accountability is not what remains after agents get good. It's what lets the rest of the whole system scale. If agents can do more work, if they can do it faster in parallel, better than what many of us could do, the scarce thing becomes the ability to explain intent, to inspect evidence, to accept risk and improve the system when the decision was wrong. Now here is the career math. The half-life of an edge might be one model release. Speed, recall, verification, even taste all move as the frontier moves. But the half-life of a signature, your credibility, your expertise is much longer. And by signature I really mean the name on the work. The person, the team, the institution, whoever stands behind what's actually shipped. So skills can earn leverage. Accountability can turn leverage into trust. And this is one of the lines that I want to draw pretty clearly. Agents can choose. They can route. They can merge. They can escalate. They can operate inside policy. And in many systems, they can. They should. But execution and responsibility are very different things. The agent can follow your runbook, but it can't inherit the consequences. When something fails, the question is, who understood the policy? Who accepted the risk and who owns the blast radius? High agency is something that a lot of us talk about these days as being this thing that we're looking for when we're hiring. High agency is actively taking ownership of your outcomes. So knowing when to delegate, when to inspect, when to stop, and when to put your name on the result. High agency in this world is not, I personally do everything. That version doesn't really scale. It's not just hustle theater. But it's ownership with judgment attached. This agency ladder tries to make that a little bit more concrete. At the bottom, you've got someone that flags a problem and leaves it for the system. Higher up, they execute, diagnose, propose, recommend, and resolve. And the rare top movement is discernment. You find a problem and you decide whether or not it's worth investing in. Maybe it's not. And maybe you move on. But when agents make more paths possible, agency is not chasing every single path. It's really just deciding which paths deserve your ownership and attention. So translate that into an operating model. Agents can run much more of the inner execution loop. They can investigate, implement, test, and report. I think that there's leverage in that. But that outer loop is still engineering. So deciding, verifying, approving, owning. That inner loop is capability. The outer loop is agency. And this is a boundary that I really care about. And the rare top movement is discernment. Maybe you find a problem and you decide whether or not it's worth investing in. Maybe it's not, and maybe you move on. But when agents make more paths possible, agency is not chasing every single path. It's deciding which paths deserve your ownership and attention. So translate that into an operating model. Agents can run much more of the inner execution loop. They can investigate, implement, test, and report. I think there's leverage in that. But that outer loop is still engineering. So deciding, verifying, approving, owning. That inner loop is capability. The outer loop is agency. And this is a boundary that I really care about. Your agent returns evidence. It returns diffs, tests, logs, rationale, traces, trajectories, screenshots. Whatever the work itself requires. But then the engineering really begins. We decide whether the work was worth doing. We verify whether the evidence is enough. And we approve or redirect or own what reaches production. It doesn't matter if you're someone that's just working with a small number of agents or whether you're working with thousands of agents. I still think these ideas apply. So the boundary is not human looks at AI output. The boundary is evidence and responsibility. So here's an operational rule: Explain it or don't ship it. And it's not because humans have to type every line or read every line. But because someone has to understand the work well enough to defend it. If you've ever worked in a large code base or an enterprise code base, some code bases have this concept of an owner's file, or certain sub directories where there are people who are on the hook for that part of the system. You can think about this in a very similar way. Who is accountable for that part of your architecture and your code base? Your model might write the code. And the question is really still whether you can explain those changes that the agent is shipping, whether you've got the evidence or you understand the risks. Now this is one of the things I want you to remember near the end. Automation moves the floor for all of us. Engineering continues to move up a level. And our new work might be loop design, evidence design, and brownfield stewardship. But fewer keystrokes doesn't mean less engineering over the next few years. It means there is more surface area that needs taste, verification, ownership, and ultimately care. I don't think I've ever been more excited about the future of this field. Every time that we've made it easier to write software, we've predicted that the world would need less of it. And in fact the opposite happened. Higher level languages happened. Frameworks, cloud, low code. The pattern always went the other way. And when you lower the cost, latent demand ends up appearing. Those ideas that people didn't think were feasible to build and get out there are suddenly unlocked. And agents are going to do the same thing for a lot of people. It's not going to remove engineering work. It's going to move the bottleneck from can we build this to should this exist and can we answer for it? So build the factories, keep the lights on, own the verdict. I hope this was useful. Thank you. I still very much think that these ideas apply. So the boundary is not human looks at AI output. The boundary is evidence and responsibility. So here's an operational rule. Explain it or don't ship it. And it's not because humans have to type every line or read every line. But because someone has to understand the work well enough to defend it. If you've ever worked in a large code base or an enterprise code base, some code bases have this concept of an owner's file. Or certain sub directories where there are people who are on the hook for that part of the system. You can think about this in a very similar way. Who is accountable for that part of your architecture and your code base? Your model might write the code. And the question is really still whether you can explain those changes that the agent is shipping. Whether you've got the evidence or you understand the risks. Now this is one of the things I want you to remember near the end. Automation moves the floor for all of us. Engineering continues to move up a level. And our new work might be loop design, evidence design, and brownfield stewardship. But fewer keystrokes doesn't mean less engineering over the next few years. It means that there is more surface area that needs taste, verification, ownership, and ultimately care. I don't think I've ever been more excited about the future of this field. Every time that we've made it easier to write software, we've predicted that the world would need less of it. And in fact the opposite happened. Higher level languages happened. Frameworks, cloud, low code. The pattern always went the other way. And when you lower the cost, latent demand ends up appearing. Those ideas that people didn't think were feasible to build and get out there are suddenly unlocked. And agents are going to do the same thing for a lot of people. It's not going to remove engineering work. It's going to move the bottleneck from can we build this to should this exist? And can we answer for it? So build the factories, keep the lights on, own the verdict. I hope this was useful. Thank you.