Today I'm going to be talking about agents, code bases, and teams. Essentially, how do you get your team to actually ship together with agents? I think for the longest time, the one thing that's bugged me is there's so much content about how do you set up your own code base to work well with agents. What skills do you add? This skill's better, that setup's better. But it all seems to break the moment you actually try to use it with your team in your actual production setup. For individual repos, it makes sense. But the moment you actually try to use it with your own team setup, it tends to break.
I think over the past few months, I figured out how to make it work with a team of folks. I was leading a team of 10 people over the last few months, and I think we found a good solution. I want to share that with you guys. But before we get into that, I just want to recap. What's been the journey that we've been on?
Coding agents took off, and a few people got really, really good leverage. I think all of us were asking, is this AGI? Did we achieve it? And then companies took that and said, well, if one person can do so well, let's just get everyone, and let's mandate it, and token max. And that was clearly a galaxy brain moment. And then the inevitable happened. AI slop shipped, and there's a bunch of sev2s. I'm not going to name which companies. But essentially, you saw people retracting. They said, I don't think this is the best option here. And eventually, model prices climbed. We saw people figure out that tokens have to be paid for.
You just can't token max your way through life. And budgets got bolted on. And essentially, money is being lit on fire. And the money has to come from somewhere. So given this journey, I want to actually, this is the enterprise journey, right? And what does that do for a single developer? And I think this is an important framing, because it really talks about people as a part of a team, right? And I want to look at it from two axes. So there is the fear axis, where people lie on the spectrum, right? It's coming from, is it coming from my job? Am I going to be out of a job?
Or is it a really handy tool? And they're not that fearful, versus the confidence they have in how much they're executing it. So they can either use it a lot, or they can use it not that much, because they don't really know how to use it that well. Now, when we started, people said, oh, what is this? Is this the end? Am I needed? And fear was pretty high. Utilization was pretty low, because people didn't really know how to use it. And then when a few people got outsized leverage, you saw early adopters. People saw them. And people said, okay, well, it looks like I'm still needed if I figure out how to use this thing. So let me actually try using it, right?
And then we saw mandates and token maxing, and people got a little skeptical. Confidence stayed the same, but people tried to use it a lot more, right? And then we realized there's a bunch of slop shipping, there's sev2s, and it's like, I'm not really that scared because it just ships slop. I'm still going to be needed. And they don't even know how to use it that well, because now the confidence is cratered, right? And so you've got to figure out how to get people from wherever they are on the spectrum to where they're not fearful, and they're actually using it a whole lot more.
And this is the framing that I want everyone to keep in mind as they're actually trying to get a team to adopt good AI usage and good AI patterns, right? And so the question is, what does it take? Step one, create a Cloud MD. Step two, add some skills. Is that it? Did we solve it? I think we all know you guys are here because clearly life's not that simple. And stuff's messy, right? And I think a few people might ask, why doesn't this work? Isn't that what everyone does? And I want to just talk about a few things you might see that actually indicate that, yeah, this isn't working.
So the first thing is, if you're babysitting your agents, it's not the right setup, right? And you've got to realize that. If you're seeing people on your team babysitting their agents, something's wrong. One of the things that I heard a lot was, insert whatever latest model there is being really dumb today. The model didn't change, right? The harness may have changed underneath. But if it's really that susceptible to small changes in the harness, clearly, your own code base isn't set up well. It's silently burning context and money. You don't realize it.
You go, you blow through 500K context. You might go to 750K, a million, and hit auto compact, even though you're not doing a really complicated task. Clearly, something's wrong. If you have long-ass sessions, you're getting constant intervention, there's still something wrong. If you're getting a constant slop factory, you obviously know things are not good. And if you find yourself asking, how are these other companies shipping so fast? How are model companies releasing models at a month-and-a-half, two-month cadence? Clearly, they have something which we don't, right? And so, I guess everyone's thinking, how do we solve this correctly?
And so, I think the first thing to realize is we need to frame it correctly, right? It isn't really an IC's job. It's a job for leadership. It's a job for the company, right? Making engineers work well with their agents is truly the most impactful thing you could do as an organization, because that's going to enable your engineers to ship faster and with confidence and avoid a lot of incidents. If we live in this figure-it-out-for-yourself paradigm, people are going to get outsized productivity. Some people aren't. And the people who are generating 10 PRs a day are going to look like gods compared to people who are shipping one to two.
And the one-to-two-PR people are actually going to get left with the review burden. And that's actually a really, really bad thing. Because now, not only can they not ship, they're going to actually see bad code and then curse the agents, and hence not be able to get onto the let's-ship-10-PRs, right? And so, it's really important to do this. If it's a problem facing the team, there's a few things you can do, right? The most impactful things that you can do to set up your code base to make it work well require team buy-in. If you want to change the way your code base is organized, you can't do that as an IC, right?
And if it's treated as a leadership problem, then you can do things like this. So, the other thing this needs is harness engineering, right? Per code base. And I think there's a lot of content on this, so I just want to talk about a few principles. But I don't want to make this talk about that, because there are a lot of smart people. You're an AI engineer. This conference is all about people telling you how to best set up your code base to make things function well. So, I don't want to talk too much about this, but there's a few key principles here. Smart prompt injection is one of them.
You want to treat your entire code base as one way to that, so that you're able to smartly prompt and check the model with just the right context at just the right time, without you needing to do it. And that's the framing. You want to be able to say, okay, I've set it off on this task. It has a map of how to find the things it needs at the time it needs it. If it's looking at some code and that code has, let's say, some documentation, the documentation needs to live in the comments. Because there's a lot of smart people. You're an AI engineer. This conference is all about people telling you how to best set up your code base to make things function well.
So, I don't want to talk too much about this, but there's a few key principles here. Smart prompt injection is one of them. You want to treat your entire code base as one way to that, so that you're able to smartly prompt and check the model with just the right context at just the right time. Without you needing to do it. And that's the framing. You want to be able to say, okay, I've set it off on this task. It has a map of how to find the things it needs at the time it needs it. If it's looking at some code and that code has, let's say, some documentation, the documentation needs to live in the comments.
So, if it ever greps into that code, it reads the comment, goes to that file, finds all the information about it. That's just one example. The second is close the loop, right? You've got to make a self-healing system because slop is inevitable. There is going to be some slop that's going to seep in. But you need to have a pipeline and a way to close the loop to remove the slop, to detect it, and to be able to self-heal the system. And then you need to iterate continuously. And I can't emphasize this enough. You can't assume that you do this for a month and you're done. Things are going to change constantly underneath.
So, you need to keep this as one of the things that you have to do as an organization. And the third most important thing is, treat it like a human problem, guys. This isn't, it's not, oh, it's this tool, people will figure it out. Let's just mandate our way through life. That's just not going to work. So, treat it like a human problem. Fear is real. Human emotions are real. We should recognize it. So, enough gyan or, it's more like the Hindi way to say enough prof, I'm giving you sermons. But how do you really do this, right? These are our principles. What's the real playbook? So, here's what we did.
And here's, I'm not going to overemphasize that this is the exact way to do it. But this is roughly how we did it, and you can take from it what you choose. The first thing is, do the basics, right? You've got to do them right. Progressive disclosure, I can't emphasize this enough, is really, really powerful, right? Find your best ICs and find how they're making the code base work for them. Take those practices and pass them org-wide. People can't live in their own practices. And this is really hard for engineers to do. It's accepting that my setup isn't perfect. And engineers don't like to hear that.
But you've got to figure out a way to find those best practices and ship them across. Make sure that that's a shared setup. The second thing we did was, there's one high-value skill that we invested in. In our case, it was this thing called ship it. What it did was, the moment you're done with your code, it takes care of everything from code done to PR ready for review. Which means you've got to open a PR, figure out your opinions, handle all the comments, handle all the PR descriptions, the merge comments, everything, right? It handles CI failures. It runs through these loops. And what this meant was often the skill was running for over an hour.
And that scared people, but once they saw the value, they got invested, right? Because it's one skill which tells them, okay, this AI thing can actually work for me. I don't need to constantly babysit it. I can trust it. The third thing, and really important, is to close the loop, right? So we wired issues and boards into the repo. We added CI/CD, we added agentic reviews. We have a code gardener that actually goes back and looks through a whole bunch of things. Every night it'll run and look at the code and check if something is not organized correctly. What correct organization means will depend on your code base. Get people invested. And I can't emphasize this enough.
You have to win over the skeptics. It's really easy to say the skeptic is just someone who's scared. It's really hard to get them to buy in. But if you can get them to buy in, you know you're doing something right. You have to get them to be able to edit and play with the shared setup. Because that's the true way you know that they're actually invested, right? And this is where you've got to ensure you're iterating constantly. If people are, and this is the hardest thing for engineers, again, because you're basically saying, I'm never going to get to perfection in my setup. But you've got to be okay with that.
You have to do it, and you have to treat it like X percent of your IC time is probably going to be spent on iterating on this thing, which is not going to lead to meaningful PRs up front.
But it's useful and it's worth it. And I don't want to say this is perfect, right? We faced a ton of issues while doing this. And I'm just going to walk you through some of them. But it's an iteration loop. So you've got to treat it like a piece of feedback. So what are the problems we hit, right? There are too many issues. When we started, we blew up to four or five hundred issues, I think within a couple of weeks. Which is a crazy number for a repo. And then there's so many different agents all trying to create issues because they've not been wired correctly. There's a lack of agreement.
As soon as people saw, oh, this isn't working perfectly or the way I expected it, it's super easy for them to say, you know what, I'm just going to go back to babysitting my agent. You don't want that. You want to actually take their feedback and put it back into the skill and improve the skill. Agents are taking too long. This is actually one of those expectation-setting things. It's good if agents take too long. That means you can actually go off and do other things and you have confidence that they're doing the right thing. At the end of the day, the moment we hit this reasoning paradigm, the longer the agent thought, the better its output.
You can create a similar mindset for your entire code base and for your skills. There's going to be merge hell. And we just have to deal with it. We have to figure out a way to deal with this. There is going to be slop when you're going to write experiments. Treat it like its own thing, right? What we said was, okay, people are generating this code, but it's not relevant. It's not going to be shipped. It's a prototype. Treat it like one. Get it to opt out of all the rigorous other standards you've got across your code base. And realize people vary on the spectrum, right?
And depending on the day, depending on what they're going through, they're going to vary on the spectrum. You have to be able to talk to them and figure out, hey, okay, why are you facing this? If the model changed, the hardness changed again, you need to go revisit something. Figure that out. And I think the biggest, the easiest way to say this is, instead of saying the model is so dumb, we have to ask, how can I make it smarter? Or how can I edit, and not, this is where I've crossed out the my. It's not a personal setup. It's the shared setup that you have to invest in. And it's a mindset, right? You have to go full send. And I want to end with this.
I learned skiing a couple years back. And the hardest thing for me was, you actually have to commit to it. If you're pizza breaking, you're going to crash. No matter what. You have to commit to the speed in order to actually get and feel like, okay, that's how I can turn. And that's how I can truly ski. And so I'm going to leave you with this. Just be okay with failing. Or how can I edit, and not, this is where I've crossed out the my. It's not a personal setup. It's the shared setup that you have to invest in. And it's a mindset, right? You have to go full send. And I want to end with this. I learned skiing a couple years back.
And the hardest thing for me was, you actually have to commit to it. If you're pizza breaking, you're going to crash. No matter what. You have to commit to the speed in order to actually get and feel, okay, that's how I can turn. And that's how I can truly ski. And so I'm going to leave you with this. Just be okay with failing. You have to go full send and be okay with falling. It's fine. The point is to be able to recover from that. And that will allow you to truly feel the AGI. Yeah, well, that's me. And I'm happy to take any questions. Yeah.
So I'm going to repeat the question for the recording.
Strategies that you found best for progressive disclosure. So I think a couple things, right? The first thing is, even in your skill MD files, don't overload it. We've set a hard limit for 100 lines in your skill MD, because your skill is really a folder. So that's step one. Make sure, and I think I spoke about this during the talk, but when you have some code that requires a runbook, make sure the runbook is reflected in the comments. So that if somehow the code, if somehow the agent figures its way into wrapping into the code base and finds that file, it knows I need to go look at this for all the description of how this is relevant. Right? You have to organize.
And I think this is why I talk about harness engineering, because your entire code base can be set up to encourage progressive disclosure. Don't overload your Cloud MD or your agents MD file into one big thing. You want to make sure that it's a thin index that can point through the right files. And that's what the agent gets in its first prompt, because that's what gets loaded when it starts to work. So these are some really powerful strategies, and the way you know this is working is when you give it a prompt, when you give it the first prompt, see what it's doing. Is it gripping? Or does it know where to go? How much context is it burning immediately?
So is it like, I think 20, 25k tokens get taken anyway, but how much more is getting added? If you're coming to 40k, 50k, something's wrong. That's not really progressive disclosure. So you have to figure out these boundaries, and then based on this, it's an iteration cycle. All right, well, if there aren't any other questions, feel free to find me.
Happy to talk about harness engineering in general or anything else. But yeah, thank you for listening.
Happy to talk about discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering discovering
you you