AI Engineer

Agents, codebases, and teams — Aditya Khandelwal, Amazon AGI Lab

1786 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent-enabled software delivery succeeds at team scale only when leadership treats the codebase, agent harness, feedback loops, and adoption process as a shared operating system rather than leaving individual engineers to optimize personal setups.
  • Why it matters: This is a concrete team-level playbook for moving from isolated developer leverage and unreliable AI-generated code to an agentic delivery system that can produce review-ready PRs with less babysitting, lower context waste, and stronger quality controls.
  • Best use: Use it to assess and design the shared harness, repository information architecture, automated quality loops, and change-management model for an AI-native engineering team.

Executive Summary

Khandelwal argues that the usual advice for coding agents—add an AGENTS.md/CLAUDE.md, install skills, and let engineers choose their own workflows—fails when a whole team shares a production codebase. The resulting pattern is uneven productivity: a few power users generate many PRs while less-confident engineers inherit review burden, lose trust after incidents, and retreat to manually supervising agents.

His central reframing is organizational: agent effectiveness is a leadership responsibility, not an individual-contributor productivity hack. A team needs per-repository harness engineering that supplies the right context at the right time, detects and removes inevitable low-quality output, and evolves continuously as models, tools, and the codebase change. The human side matters equally: mandates and token-maxing create skepticism, while real adoption requires winning over skeptics and allowing them to improve the shared system.

The operational example from his 10-person team is a shared workflow centered on one high-value “ship it” skill. It takes code from completion through PR creation, descriptions, review-comment handling, CI failures, and merge-related steps. They connected issues and boards to the repository, added CI/CD and agentic reviews, and ran a nightly “code gardener” to identify organizational drift. The goal is not zero oversight; it is reliable, asynchronous execution with closed-loop recovery.

The most reusable technical advice is progressive disclosure. Keep the initial instruction file as a thin index rather than a giant prompt, cap skill markdown at roughly 100 lines, embed runbook pointers in code comments, and organize the repository so agents can discover detailed context only after reaching relevant code. He recommends monitoring early-session context consumption: if an agent rapidly expands from an unavoidable 20–25K tokens to 40–50K before doing meaningful work, the repository’s retrieval and navigation design is likely failing.

Key Takeaways

  • Claim: Team-scale agent adoption should be owned as a leadership and organizational systems problem, not delegated to each engineer’s personal prompt and skill setup. | Evidence: Khandelwal says his team of 10 found that individual-repo setups break in shared production environments: some people can generate roughly 10 PRs per day while others ship one or two and become responsible for reviewing the AI-generated output. | Implication: Ken should standardize core agent workflows and repository interfaces at the team level; otherwise, agent leverage will compound for a small group while quality and review load concentrate on everyone else. | Caveat: He presents his team’s process as an example rather than a universal prescribed implementation.
  • Claim: Babysitting, excessive context consumption, long intervention-heavy sessions, and recurring slop are diagnostics of a weak harness or poorly structured codebase—not proof that the current model is inherently incapable. | Evidence: He flags 500K to 1M-token sessions that hit auto-compaction on non-complex work, repeated complaints that a model is “dumb today,” and constant human intervention as signs that context delivery and tooling changed or are misconfigured. | Implication: Instrument agent runs for context growth, handoffs, retries, interventions, and failure modes; treat these as operational signals for improving the environment. | Caveat: Model behavior can vary, but his argument is that teams should first investigate the harness and shared setup rather than blaming the model.
  • Claim: Progressive disclosure is the core repository-design pattern: give the agent a small navigational entry point and let it retrieve detailed instructions only when relevant. | Evidence: The team caps skill MD files at 100 lines, treats a skill as a folder rather than one large instruction document, uses the top-level agent file as a thin index, and places runbook/documentation cues in code comments so an agent that reaches a file can discover the appropriate deeper guidance. | Implication: Design agent instructions like an information-retrieval system: optimize the first prompt for routing, place local operational knowledge near the code, and avoid loading a whole repository’s policy into every task. | Caveat: The structure must be iterated against observed agent behavior; a static documentation hierarchy is not enough.
  • Claim: A single trusted end-to-end workflow can drive adoption better than many disconnected agent capabilities. | Evidence: Their highest-value shared skill, “ship it,” moves work from code completion to a review-ready PR by opening the PR, producing descriptions and merge comments, responding to review comments, and working through CI failures. It often runs for more than an hour, but users adopted it after seeing they no longer had to constantly supervise it. | Implication: Prioritize an autonomous, auditable completion loop around a valuable delivery boundary—such as task-to-PR or PR-to-merge—over expanding a catalog of shallow skills. | Caveat: Long-running agent work requires expectation-setting: elapsed time is acceptable only if users can confidently do other work while the agent progresses.
  • Claim: Because slop is inevitable, teams need closed-loop detection and repair rather than relying solely on initial prompts or human review. | Evidence: The team wired issues and boards into the repo, added CI/CD and agentic reviews, and operates a nightly “code gardener” that scans for code that is not organized according to the repository’s defined standards. | Implication: Build automated remediation and hygiene loops into the engineering control plane, while labeling prototypes and experiments so they do not silently lower production quality. | Caveat: What constitutes correct organization is codebase-specific, and generated experimental code should be explicitly separated from production-bound work rather than forced through every production standard.
  • Claim: Sustained adoption depends on treating agent rollout as a human change-management problem, especially by converting skeptics into contributors to the shared setup. | Evidence: He describes a cycle where mandates and poor outputs reduce confidence; his response is to let skeptical engineers edit the shared configuration, feed their complaints back into skills, and explicitly allocate some IC time to iteration that may not initially create feature PRs. | Implication: Create a visible ownership and feedback process for the shared agent environment, protect iteration time, and judge adoption by whether skeptical engineers help improve the system rather than merely comply with usage targets. | Caveat: The team must accept that the setup will never be permanently “finished”; model, tool, and codebase changes continually invalidate assumptions.

Detailed Brief

Failure modes the operating model must absorb

  • Claims: Connecting multiple agents to issue tracking without clear controls can create an unmanageable issue backlog.; Merge contention is a normal consequence of increased parallel agent activity and must be handled as a workflow-design problem.; Agent-generated experimental code should be governed differently from code intended to ship.
  • Evidence: The team accumulated approximately 400–500 issues within a couple of weeks because multiple agents were creating issues before being properly wired.; Khandelwal explicitly anticipates “merge hell” as agent throughput rises.; For prototypes that were not intended to ship, the team opted them out of the repository’s full rigorous standards.
  • Caveats: Removing production rigor from experiments is safe only when the boundary between prototype and shippable code is explicit and enforced.; Higher agent throughput can worsen coordination costs if task decomposition, ownership, and merge sequencing remain unchanged.
  • Implications: Add quotas, deduplication, ownership, and lifecycle rules before allowing agents to autonomously create issues.; Plan for integration capacity—branch strategy, task partitioning, CI capacity, and merge queues—not just code-generation capacity.

Practical measurement for progressive disclosure

  • Claims: The first observable test of a repository’s agent interface is whether the agent knows where to navigate after its initial prompt.; Context growth early in a task is a practical proxy for whether the agent is discovering relevant guidance efficiently.
  • Evidence: He asks teams to inspect whether the agent immediately greps broadly or navigates deliberately after the first prompt.; He estimates that 20–25K tokens may be consumed as baseline overhead, but reaching 40–50K tokens immediately suggests progressive disclosure is not functioning well.
  • Caveats: These figures are heuristics from the speaker, not universal thresholds; task complexity and the specific agent harness will affect normal context use.
  • Implications: Define repository-specific telemetry baselines and investigate abnormal context expansion before increasing model context windows or budgets.

Notable Concepts & Terms

  • Harness engineering: Per-codebase design of prompts, instructions, repository structure, tools, and feedback loops so agents receive relevant context and can act reliably.
  • Progressive disclosure: A context-management pattern in which agents start with a thin index and discover detailed guidance through relevant files, comments, and linked runbooks.
  • Thin index: A compact top-level AGENTS.md/CLAUDE.md-style file that routes the agent to more specific instructions instead of embedding all knowledge in the initial prompt.
  • Ship it skill: The team’s end-to-end autonomous workflow that turns finished code into a review-ready PR and handles downstream review and CI loops.
  • Code gardener: A nightly automated maintenance agent that scans the repository for organizational or hygiene drift and supports self-healing.
  • Closed-loop/self-healing system: An agentic development environment that detects failures or low-quality output and feeds corrections back into automated checks and shared workflows.
  • Token maxing: The rejected enterprise pattern of mandating heavy AI use or unconstrained token consumption without solving quality, cost, or workflow reliability.

Operator Notes / Why Ken Should Care

  • Audit one representative agent task end to end: initial context loaded, token growth before first meaningful action, tool calls, human interventions, CI outcomes, review iterations, and final merge status.
  • Replace any monolithic agent instruction document with a thin repository index plus local, discoverable runbook references; impose a concise size limit on skill-level instruction files.
  • Select one high-frequency delivery workflow and build an autonomous shared completion loop for it, with explicit handoff states, CI repair behavior, and review-comment handling.
  • Establish a weekly shared-harness review attended by both high-leverage users and skeptics; require that recurring complaints become tracked changes to instructions, tools, repository structure, or checks.
  • Add governance before enabling autonomous issue creation: deduplication, severity/ownership fields, rate limits, and a clear distinction between agent findings, experiments, and actionable engineering work.
  • Allocate explicit engineering capacity for harness iteration and repository hygiene rather than evaluating the work solely by immediate feature PR output.

Source/Metadata

  • Title: Agents, codebases, and teams — Aditya Khandelwal, Amazon AGI Lab
  • Transcript words: 3623
  • Duration seconds: 1016
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.
Full transcript 3334 words · 16 min read
0:12

Today I'm going to be talking about agents, code bases, and teams. Essentially, how do you get your team to actually ship together with agents? I think for the longest time, the one thing that's bugged me is there's so much content about how do you set up your own code base to work well with agents. What skills do you add? This skill's better, that setup's better. But it all seems to break the moment you actually try to use it with your team in your actual production setup. For individual repos, it makes sense. But the moment you actually try to use it with your own team setup, it tends to break.

0:47

I think over the past few months, I figured out how to make it work with a team of folks. I was leading a team of 10 people over the last few months, and I think we found a good solution. I want to share that with you guys. But before we get into that, I just want to recap. What's been the journey that we've been on?

1:06

Coding agents took off, and a few people got really, really good leverage. I think all of us were asking, is this AGI? Did we achieve it? And then companies took that and said, well, if one person can do so well, let's just get everyone, and let's mandate it, and token max. And that was clearly a galaxy brain moment. And then the inevitable happened. AI slop shipped, and there's a bunch of sev2s. I'm not going to name which companies. But essentially, you saw people retracting. They said, I don't think this is the best option here. And eventually, model prices climbed. We saw people figure out that tokens have to be paid for.

1:51

You just can't token max your way through life. And budgets got bolted on. And essentially, money is being lit on fire. And the money has to come from somewhere. So given this journey, I want to actually, this is the enterprise journey, right? And what does that do for a single developer? And I think this is an important framing, because it really talks about people as a part of a team, right? And I want to look at it from two axes. So there is the fear axis, where people lie on the spectrum, right? It's coming from, is it coming from my job? Am I going to be out of a job?

2:25

Or is it a really handy tool? And they're not that fearful, versus the confidence they have in how much they're executing it. So they can either use it a lot, or they can use it not that much, because they don't really know how to use it that well. Now, when we started, people said, oh, what is this? Is this the end? Am I needed? And fear was pretty high. Utilization was pretty low, because people didn't really know how to use it. And then when a few people got outsized leverage, you saw early adopters. People saw them. And people said, okay, well, it looks like I'm still needed if I figure out how to use this thing. So let me actually try using it, right?

3:04

And then we saw mandates and token maxing, and people got a little skeptical. Confidence stayed the same, but people tried to use it a lot more, right? And then we realized there's a bunch of slop shipping, there's sev2s, and it's like, I'm not really that scared because it just ships slop. I'm still going to be needed. And they don't even know how to use it that well, because now the confidence is cratered, right? And so you've got to figure out how to get people from wherever they are on the spectrum to where they're not fearful, and they're actually using it a whole lot more.

3:39

And this is the framing that I want everyone to keep in mind as they're actually trying to get a team to adopt good AI usage and good AI patterns, right? And so the question is, what does it take? Step one, create a Cloud MD. Step two, add some skills. Is that it? Did we solve it? I think we all know you guys are here because clearly life's not that simple. And stuff's messy, right? And I think a few people might ask, why doesn't this work? Isn't that what everyone does? And I want to just talk about a few things you might see that actually indicate that, yeah, this isn't working.

4:11

So the first thing is, if you're babysitting your agents, it's not the right setup, right? And you've got to realize that. If you're seeing people on your team babysitting their agents, something's wrong. One of the things that I heard a lot was, insert whatever latest model there is being really dumb today. The model didn't change, right? The harness may have changed underneath. But if it's really that susceptible to small changes in the harness, clearly, your own code base isn't set up well. It's silently burning context and money. You don't realize it.

4:34

You go, you blow through 500K context. You might go to 750K, a million, and hit auto compact, even though you're not doing a really complicated task. Clearly, something's wrong. If you have long-ass sessions, you're getting constant intervention, there's still something wrong. If you're getting a constant slop factory, you obviously know things are not good. And if you find yourself asking, how are these other companies shipping so fast? How are model companies releasing models at a month-and-a-half, two-month cadence? Clearly, they have something which we don't, right? And so, I guess everyone's thinking, how do we solve this correctly?

4:57

And so, I think the first thing to realize is we need to frame it correctly, right? It isn't really an IC's job. It's a job for leadership. It's a job for the company, right? Making engineers work well with their agents is truly the most impactful thing you could do as an organization, because that's going to enable your engineers to ship faster and with confidence and avoid a lot of incidents. If we live in this figure-it-out-for-yourself paradigm, people are going to get outsized productivity. Some people aren't. And the people who are generating 10 PRs a day are going to look like gods compared to people who are shipping one to two.

5:20

And the one-to-two-PR people are actually going to get left with the review burden. And that's actually a really, really bad thing. Because now, not only can they not ship, they're going to actually see bad code and then curse the agents, and hence not be able to get onto the let's-ship-10-PRs, right? And so, it's really important to do this. If it's a problem facing the team, there's a few things you can do, right? The most impactful things that you can do to set up your code base to make it work well require team buy-in. If you want to change the way your code base is organized, you can't do that as an IC, right?

5:47

And if it's treated as a leadership problem, then you can do things like this. So, the other thing this needs is harness engineering, right? Per code base. And I think there's a lot of content on this, so I just want to talk about a few principles. But I don't want to make this talk about that, because there are a lot of smart people. You're an AI engineer. This conference is all about people telling you how to best set up your code base to make things function well. So, I don't want to talk too much about this, but there's a few key principles here. Smart prompt injection is one of them.

6:17

You want to treat your entire code base as one way to that, so that you're able to smartly prompt and check the model with just the right context at just the right time, without you needing to do it. And that's the framing. You want to be able to say, okay, I've set it off on this task. It has a map of how to find the things it needs at the time it needs it. If it's looking at some code and that code has, let's say, some documentation, the documentation needs to live in the comments. Because there's a lot of smart people. You're an AI engineer. This conference is all about people telling you how to best set up your code base to make things function well.

6:38

So, I don't want to talk too much about this, but there's a few key principles here. Smart prompt injection is one of them. You want to treat your entire code base as one way to that, so that you're able to smartly prompt and check the model with just the right context at just the right time. Without you needing to do it. And that's the framing. You want to be able to say, okay, I've set it off on this task. It has a map of how to find the things it needs at the time it needs it. If it's looking at some code and that code has, let's say, some documentation, the documentation needs to live in the comments.

7:12

So, if it ever greps into that code, it reads the comment, goes to that file, finds all the information about it. That's just one example. The second is close the loop, right? You've got to make a self-healing system because slop is inevitable. There is going to be some slop that's going to seep in. But you need to have a pipeline and a way to close the loop to remove the slop, to detect it, and to be able to self-heal the system. And then you need to iterate continuously. And I can't emphasize this enough. You can't assume that you do this for a month and you're done. Things are going to change constantly underneath.

7:43

So, you need to keep this as one of the things that you have to do as an organization. And the third most important thing is, treat it like a human problem, guys. This isn't, it's not, oh, it's this tool, people will figure it out. Let's just mandate our way through life. That's just not going to work. So, treat it like a human problem. Fear is real. Human emotions are real. We should recognize it. So, enough gyan or, it's more like the Hindi way to say enough prof, I'm giving you sermons. But how do you really do this, right? These are our principles. What's the real playbook? So, here's what we did.

8:36

And here's, I'm not going to overemphasize that this is the exact way to do it. But this is roughly how we did it, and you can take from it what you choose. The first thing is, do the basics, right? You've got to do them right. Progressive disclosure, I can't emphasize this enough, is really, really powerful, right? Find your best ICs and find how they're making the code base work for them. Take those practices and pass them org-wide. People can't live in their own practices. And this is really hard for engineers to do. It's accepting that my setup isn't perfect. And engineers don't like to hear that.

9:03

But you've got to figure out a way to find those best practices and ship them across. Make sure that that's a shared setup. The second thing we did was, there's one high-value skill that we invested in. In our case, it was this thing called ship it. What it did was, the moment you're done with your code, it takes care of everything from code done to PR ready for review. Which means you've got to open a PR, figure out your opinions, handle all the comments, handle all the PR descriptions, the merge comments, everything, right? It handles CI failures. It runs through these loops. And what this meant was often the skill was running for over an hour.

9:32

And that scared people, but once they saw the value, they got invested, right? Because it's one skill which tells them, okay, this AI thing can actually work for me. I don't need to constantly babysit it. I can trust it. The third thing, and really important, is to close the loop, right? So we wired issues and boards into the repo. We added CI/CD, we added agentic reviews. We have a code gardener that actually goes back and looks through a whole bunch of things. Every night it'll run and look at the code and check if something is not organized correctly. What correct organization means will depend on your code base. Get people invested. And I can't emphasize this enough.

10:17

You have to win over the skeptics. It's really easy to say the skeptic is just someone who's scared. It's really hard to get them to buy in. But if you can get them to buy in, you know you're doing something right. You have to get them to be able to edit and play with the shared setup. Because that's the true way you know that they're actually invested, right? And this is where you've got to ensure you're iterating constantly. If people are, and this is the hardest thing for engineers, again, because you're basically saying, I'm never going to get to perfection in my setup. But you've got to be okay with that.

10:48

You have to do it, and you have to treat it like X percent of your IC time is probably going to be spent on iterating on this thing, which is not going to lead to meaningful PRs up front.

10:53

But it's useful and it's worth it. And I don't want to say this is perfect, right? We faced a ton of issues while doing this. And I'm just going to walk you through some of them. But it's an iteration loop. So you've got to treat it like a piece of feedback. So what are the problems we hit, right? There are too many issues. When we started, we blew up to four or five hundred issues, I think within a couple of weeks. Which is a crazy number for a repo. And then there's so many different agents all trying to create issues because they've not been wired correctly. There's a lack of agreement.

11:28

As soon as people saw, oh, this isn't working perfectly or the way I expected it, it's super easy for them to say, you know what, I'm just going to go back to babysitting my agent. You don't want that. You want to actually take their feedback and put it back into the skill and improve the skill. Agents are taking too long. This is actually one of those expectation-setting things. It's good if agents take too long. That means you can actually go off and do other things and you have confidence that they're doing the right thing. At the end of the day, the moment we hit this reasoning paradigm, the longer the agent thought, the better its output.

11:52

You can create a similar mindset for your entire code base and for your skills. There's going to be merge hell. And we just have to deal with it. We have to figure out a way to deal with this. There is going to be slop when you're going to write experiments. Treat it like its own thing, right? What we said was, okay, people are generating this code, but it's not relevant. It's not going to be shipped. It's a prototype. Treat it like one. Get it to opt out of all the rigorous other standards you've got across your code base. And realize people vary on the spectrum, right?

12:32

And depending on the day, depending on what they're going through, they're going to vary on the spectrum. You have to be able to talk to them and figure out, hey, okay, why are you facing this? If the model changed, the hardness changed again, you need to go revisit something. Figure that out. And I think the biggest, the easiest way to say this is, instead of saying the model is so dumb, we have to ask, how can I make it smarter? Or how can I edit, and not, this is where I've crossed out the my. It's not a personal setup. It's the shared setup that you have to invest in. And it's a mindset, right? You have to go full send. And I want to end with this.

13:06

I learned skiing a couple years back. And the hardest thing for me was, you actually have to commit to it. If you're pizza breaking, you're going to crash. No matter what. You have to commit to the speed in order to actually get and feel like, okay, that's how I can turn. And that's how I can truly ski. And so I'm going to leave you with this. Just be okay with failing. Or how can I edit, and not, this is where I've crossed out the my. It's not a personal setup. It's the shared setup that you have to invest in. And it's a mindset, right? You have to go full send. And I want to end with this. I learned skiing a couple years back.

13:57

And the hardest thing for me was, you actually have to commit to it. If you're pizza breaking, you're going to crash. No matter what. You have to commit to the speed in order to actually get and feel, okay, that's how I can turn. And that's how I can truly ski. And so I'm going to leave you with this. Just be okay with failing. You have to go full send and be okay with falling. It's fine. The point is to be able to recover from that. And that will allow you to truly feel the AGI. Yeah, well, that's me. And I'm happy to take any questions. Yeah.

14:40

So I'm going to repeat the question for the recording.

14:45

Strategies that you found best for progressive disclosure. So I think a couple things, right? The first thing is, even in your skill MD files, don't overload it. We've set a hard limit for 100 lines in your skill MD, because your skill is really a folder. So that's step one. Make sure, and I think I spoke about this during the talk, but when you have some code that requires a runbook, make sure the runbook is reflected in the comments. So that if somehow the code, if somehow the agent figures its way into wrapping into the code base and finds that file, it knows I need to go look at this for all the description of how this is relevant. Right? You have to organize.

15:27

And I think this is why I talk about harness engineering, because your entire code base can be set up to encourage progressive disclosure. Don't overload your Cloud MD or your agents MD file into one big thing. You want to make sure that it's a thin index that can point through the right files. And that's what the agent gets in its first prompt, because that's what gets loaded when it starts to work. So these are some really powerful strategies, and the way you know this is working is when you give it a prompt, when you give it the first prompt, see what it's doing. Is it gripping? Or does it know where to go? How much context is it burning immediately?

16:01

So is it like, I think 20, 25k tokens get taken anyway, but how much more is getting added? If you're coming to 40k, 50k, something's wrong. That's not really progressive disclosure. So you have to figure out these boundaries, and then based on this, it's an iteration cycle. All right, well, if there aren't any other questions, feel free to find me.

16:31

Happy to talk about harness engineering in general or anything else. But yeah, thank you for listening.

16:37

Happy to talk about discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering

16:49

you you

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note