Coding Agents Don't Scale Themselves. Neither Do Your Teams. — Patrick Debois, Tessl
Description
In 2009 people told Patrick Debois that continuous delivery was crazy. He hears the same thing now about the dark factory, and reads it the same way: not that the technology cannot work, but that the organization is not set up for it yet. His starting assumption is that harnesses and loops will commoditize, possibly into a service a frontier lab just sells you, so none of that will be anyone's differentiator. What actually changes is the team, the platform and the organization around them. Developers pushed back on the conductor framing, telling him they did not sign up to write better prompts. What brought them back was tooling. Once the team started building harnesses for the agent, a genuinely technical path reopened, and the loudest skeptics turned out to be the right people to hand context authoring to, precisely because they were angry about the output. The shift he pushes is to stop fixing the code the agent produced and improve the system instead. Retros stop being about the code and become about where the agent hit the same wall repeatedly. Planning splits into work scoped tightly enough to hand off and work that still needs a conversation. He tracks two numbers: how many human touches it takes to get the right result, which should fall, and how much of each fix is shared, because one improvement to a common harness lands for everybody rather than making a single person 10x. Above the team it is paved roads and a named owner, not a thousand flowers blooming. Speaker info: - https://x.com/patrickdebois - https://www.linkedin.com/in/patrickdebois/ - https://jedi.be Timestamps: 0:00 - It will not work here, and what that signals 2:46 - Developers who did not sign up for prompting 5:20 - Stop fixing the code, improve the system 7:04 - Retros, planning, and the downstream squeeze 9:38 - Two metrics: human touches and reuse 10:29 - The platform team's new problems 12:12 - Sprawl, paved roads, and making spend visible 14:43 - Enabling the organization without a
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Coding-agent advantage will come less from individual prompting skill than from turning context, harnesses, evaluations, guardrails, and operating practices into shared organizational infrastructure.
- Why it matters: This is directly relevant to building scalable agent systems: it frames the real bottlenecks as reusable context, platform ownership, risk controls, downstream workflow capacity, and organizational learning rather than raw model capability.
- Best use: Use it as an operating-model blueprint for moving from individual coding-agent adoption to an internal agent enablement platform and governance program.
Executive Summary
Patrick Debois argues that autonomous software delivery—or the “dark factory”—is primarily an organizational-readiness problem, not a question of whether the underlying agent technology will eventually work. Coding-agent harnesses, loops, and model capabilities will commoditize; the durable differentiator is how a company changes its engineering system, team rituals, platform services, and cross-functional workflows around them.
The key behavioral shift is from repairing individual agent outputs to improving the system that produces them. Teams should capture reusable context, encode expectations such as tests and documentation, and treat repeated agent failures as system defects surfaced in retrospectives. Planning should separate well-scoped work that agents can execute from ambiguous, conversational work that still requires human judgment.
Debois sees the multiplier moving from an individual developer to shared infrastructure. Teams can begin inside a repository, but organizations eventually need centrally maintained registries for skills, reusable context, harness components, evaluations, identity, security guardrails, and cost visibility. This should provide a limited catalog of tested “paved roads,” not an ungoverned sprawl of competing prompts and skills.
He also warns against simplistic adoption and ROI stories: licenses, hackathons, training, and productivity claims alone do not create a scalable system. Leaders should mandate team leads and platform/DX functions to own the transition, measure declining human intervention and rising reuse, optimize token spend through better context and fewer iterations, and calibrate autonomy by feature risk using provenance, verification, and situational awareness.
Key Takeaways
- Claim: The strategic advantage of coding agents will not be the agent itself but the organizational system built around it. | Evidence: Debois compares current skepticism about autonomous “dark factories” to the early reception of continuous delivery in 2009: “it will not work here” usually signals that the organization is not set up for it. He expects agent loops and harnesses to become commodity capabilities, potentially delivered by frontier labs. | Implication: Treat agent enablement as a systems and operating-model investment, not as a procurement decision or an individual developer productivity tool. | Caveat: He does not claim full autonomy is immediately appropriate; readiness and feature-level risk determine the permissible degree of autonomy.
- Claim: Engineering teams should stop fixing agent-produced code repeatedly and instead improve the context, harness, and feedback system that generated it. | Evidence: The speaker explicitly reframes the job as “build the thing that builds the thing,” using reusable context, harnesses, and loops. He says teams should require agents to test, update documentation, and follow the same engineering practices expected of good human engineers. | Implication: Convert recurring failure modes into codified instructions, tools, tests, guardrails, and evaluators rather than absorbing them through manual cleanup. | Caveat: This is not an endorsement of unreviewed vibe coding; conventional engineering practices remain necessary both to maintain the system and to improve future agent performance.
- Claim: Planning and retrospectives need to become agent-system management rituals, with humans focusing on ambiguity and agents taking sufficiently defined work. | Evidence: In advanced teams, retrospectives examine why an agent repeatedly hit a problem and how to fix the system. During planning, well-scoped tasks are routed directly to agents, while humans retain work that is underspecified or requires team conversation and decisions. | Implication: Build an explicit work-classification process: agent-ready tasks need clear contracts and constraints, while ambiguous product and architecture decisions should remain human-led until context is mature enough. | Caveat: The split is based on scope clarity, not a blanket rule that agents should own all implementation work.
- Claim: The scalable productivity multiplier is shared agent infrastructure, not turning isolated developers into “10x” operators. | Evidence: Debois proposes tracking two signals: the number of human touches needed for an agent to deliver the right outcome, which should decline; and the reuse of improvements, where a single fix to context or a harness benefits every developer using it. | Implication: Prioritize shared assets whose improvement compounds across users, and instrument intervention rate, iterations/turns, reuse, and associated cost rather than relying on self-reported productivity. | Caveat: Token spend and headline delivery-speed claims alone are insufficient measures because they obscure whether the system itself is becoming more effective.
- Claim: A platform-like central function must own governed reuse for coding agents, including skills, context, evaluations, guardrails, identities, and cost observability. | Evidence: He cites reusable authentication knowledge, common linters, and security tools as examples that belong in a shared registry. Without ownership, organizations get skill sprawl: duplicated or forked artifacts with no clear choice, maintenance model, testability, or security posture. | Implication: Assign a clear accountable owner—likely a blend of platform engineering and developer experience—and create a curated, versioned catalog rather than an open-ended repository of agent artifacts. | Caveat: A single universal standard is unlikely to be practical; Debois expects a catalog of roughly three or four maintained paved roads, while custom approaches remain possible but are funded by the teams choosing them.
- Claim: Organization-wide adoption requires explicit mandates and a broader workflow redesign; bottom-up experimentation alone will not scale. | Evidence: Debois characterizes hackathons, lunch-and-learns, Slack channels, champions programs, licenses, and generic education as standard transformation tactics that are inadequate by themselves. He notes that faster engineering output can overwhelm GTM, users, and requirements intake unless those downstream and upstream workflows are automated too. | Implication: Give team leads and the central platform/DX function formal responsibility and capacity to redesign the entire delivery flow, including requirement capture, release/GTM support, and customer-facing throughput constraints. | Caveat: The talk offers a directional operating model rather than a detailed change-management rollout plan or empirical comparison of adoption programs.
- Claim: “Dark factory” autonomy should be risk-tiered and supported by auditability, verification, and situational awareness rather than assumed as an all-or-nothing destination. | Evidence: Debois calls the near-term reality a “dim factory”: not all features should become autonomous. He recommends investing in provenance of code changes, verifiers to assess usefulness/correctness, and situational awareness when failures occur. | Implication: Define autonomy levels by change type and risk, then raise autonomy only where evidence, audit trails, and rollback/diagnostic mechanisms support it. | Caveat: He provides no specific risk taxonomy, acceptance thresholds, or verifier implementation design; each organization must decide its own autonomy boundary.
Detailed Brief
Developer identity, leadership, and hiring
- Claims: The move from writing code to prompting and specification work can create identity friction for engineers who did not choose a managerial role over agents.; Building tools for agents creates a new technical career path that can re-engage skeptical engineers.; Team leads should actively set the maturity sequence—from prompting to reusable context to harnesses and loops—rather than telling individuals to experiment independently.; Hiring should test AI leverage, engineering judgment, and collaborative behavior as distinct dimensions rather than relying on vague AI-related job titles.
- Evidence: Debois recommends directing quality skeptics to encode their knowledge into better context and harnesses, turning their objections into systematic improvements.; His suggested hiring process starts with an AI-permitted exercise, follows with a walkthrough of why the solution is sound, and separately assesses willingness to share rather than operate as a solo player.; He argues labels such as “AI engineer,” “agentic engineer,” “AI product engineer,” and “forward deployed engineer” do not reliably signal maturity.
- Caveats: A candidate may be strong in only one or two of AI leverage, engineering taste, and collaboration; this should guide mentoring rather than collapse into a simplistic junior/senior rating.; Small, highly capable teams still need redundancy, production support, bug handling, and junior development, limiting the practical extent of one- or two-person team models.
- Implications: Frame agent-system construction as serious engineering work to retain and mobilize technical staff who resist prompt-centric role definitions.; Build hiring rubrics around demonstrated agent use, explainable technical judgment, and contribution to shared systems.
Economics and continuous learning
- Claims: When vendor costs rise, the appropriate default is spend optimization, not indiscriminate restriction of agent usage.; Better model selection, user education, context, and harnesses can reduce cost by reducing unnecessary agent iterations.; The long-term operating goal evolves from continuous delivery to continuous learning: retain reliability while increasing the rate at which the system can change.
- Evidence: The platform function should expose both what an agent workflow costs and what it accomplishes, allowing teams to see the value of reducing turns or iterations.; Debois argues that reuse and measured improvement in agent effectiveness are more defensible leadership metrics than trying to prove a clean productivity comparison with versus without coding agents.
- Caveats: The speaker acknowledges that faster delivery and better quality are commonly promised but difficult to prove cleanly.; Cost reduction should not be confused with blanket model downgrades; it is a systems optimization problem.
- Implications: Create cost-and-quality telemetry at the workflow level so teams can optimize agent behavior rather than respond to budget pressure with arbitrary caps.; Evaluate progress by whether the organization can safely absorb more frequent change, not simply whether a single code-generation task completes faster.
Notable Concepts & Terms
- Dark factory / dim factory: The aspiration for autonomous software work; Debois uses “dim factory” to stress that near-term autonomy must be selectively applied according to risk.
- Build the thing that builds the thing: The central abstraction shift: improve reusable agent-producing systems rather than manually repairing each generated output.
- Context engineering: Moving beyond a one-off prompt toward tested, distributed, optimized context that reliably shapes agent behavior.
- Harnesses and loops: Programmatic scaffolding and iterative feedback mechanisms that let agents execute work under engineering constraints rather than through ad hoc prompting.
- Human touches: A proposed effectiveness metric: the number of human interventions needed to get the agent to a correct result should fall as context and harnesses improve.
- Paved roads: A small, maintained catalog of approved reusable patterns and tools that offers an easy, safe default while avoiding forced uniformity.
- Skill registry: A centralized, governed repository for reusable agent capabilities, context, and related components, with ownership, testing, modularity, and security controls.
- Provenance and verifiers: Audit and validation mechanisms for autonomous changes: identify who or what changed code and assess whether the output is useful or correct.
Operator Notes / Why Ken Should Care
- Establish an agent-enablement owner with explicit cross-functional remit spanning platform engineering, developer experience, security/identity, and engineering leadership.
- Inventory recurring agent failures and convert the highest-frequency ones into reusable context, deterministic tooling, tests, and evaluation cases; track intervention count before and after.
- Create an internal registry governance model now: artifact ownership, versioning, evaluation requirements, security review, deprecation rules, and a curated set of approved paved roads.
- Add agent-readiness classification to planning: define the minimum task specification, tests, and acceptance criteria required before work can be delegated autonomously.
- Instrument workflow-level telemetry for iterations/turns, human interventions, reuse, model choice, and cost; avoid using licenses issued or raw token spend as primary success measures.
- Set a risk-tiered autonomy policy for code changes, with provenance, verification, escalation, and rollback expectations per tier.
- Assess whether requirements intake, release/GTM processes, support, and user onboarding will become the next bottlenecks if engineering throughput increases.
Source/Metadata
- Title: Coding Agents Don't Scale Themselves. Neither Do Your Teams. — Patrick Debois, Tessl
- Transcript words: 6247
- Duration seconds: 1325
- Timestamp note: No timestamps or chapters were present. The transcript contains a substantial repeated segment near the end.
Transcript
Well, welcome. Last day, I guess that's what happens. I'm going to talk to you maybe not on the technical side, but more on the organizational side. So if you're here for any technology, you can still leave if you want to. So in 2009, a lot of people were telling me the idea of continuous delivery was crazy, and I feel we're in the same era or the same thing right now with a dark factory. It will not work here. That's what I keep hearing over and over again. But what they are actually signaling to me is we're not ready yet. So it's not the technology that can't make it work. It's not something they won't be able to do eventually, but they're just not set up for this. And there's been a lot of conference talks here about optimizing agents with loops and harnesses and all those pieces, and I think that's great, but eventually we'll get there, right? It's not that this is rocket science, and yes, we'll have to assemble this in a good way. But one day this will become commodity somewhere, maybe even going into one of the frontier labs that just offers this as a service, and we'll make this work. And that's not going to be the differentiator for your organization. So I'm starting from there up. Assume we're heading towards the dark factory, some kind of form of autonomous working within an organization. What I've seen for the people adopting this within our organization, including here where I work at TESOL, is it changes the dynamic of the way you collaborate around this. And for those familiar, there's Conway's law, the way you organize yourselves and the tools, there is a relationship in how they interact and work together on this. But today I'm not talking about how you become better with your agent, but it is about how this will change your team dynamics, your platform, and your organization. So that's what I'll take you through. Enabling the team. I assume most of you work in a team and that you're not a solopreneur working somewhere. So it works differently than just you with your cloud code, and a team working together around that with cloud or any of the coding agents there as well. The narrative that I heard a lot is, well, the developer eventually becomes more of a conductor and an orchestrator of agents. And then I think that's fair. That's been an evolution that we're on, on the path, or more like becoming the managers of the agent, dealing with the agents. Now what I've seen is that eventually a lot of developers told me, we didn't sign up for this. We didn't sign up for better prompting, writing better specs. We're engineers. We're technical. And that creates friction, like, is this the role that we really want to do? There was a thing that came around which maybe is more context engineering that put a first step around, hey, it's not just a prompt. We'll test the prompt. We'll evaluate the prompt. We'll distribute the prompt and optimize the prompt. So yes, there's a little bit of engineering, but still a lot of developers felt empty just working with a prompt and a specification as such. What I've seen is that when we started introducing harnesses and loops, and eventually more autonomous work within the whole organization, a new technical path opened. All of a sudden we were helping the agent with tooling, building tooling for the agent, and that reignited some of the developers who felt that it wasn't for them. Now all of a sudden they were like, yes, we can do this. We have that knowledge. We're somehow helping this even in a programmatic way. So I think that's interesting, that the identity where we say abstraction, abstraction, abstraction, technically all of a sudden the craft created some new location for more engineering stuff to go to. Now when I get the question, can we please help skeptical people, what do they do? I also say that these are really great people to engage in creating better context for the agent. Because you tell them, please improve, please put all your knowledge to improve the result of the agent. And the same with the harness. So if you have those more resistant people that complain maybe about the quality that things were produced by just the vanilla coding agent, use that anger, use that skepticism to make it better. And the big mentality shift, if I would advise a company right now for their developers, is stop fixing the code that the agent produced, but improve the system. I'm not the only one saying this at this event, but that is the difference. You improve the system. And I think it was Swix a couple of years ago who said it: stop building the thing, but build the thing that builds the thing. Right? So we're going on that abstraction where that is with context, with harness, with loops, and that is the change that a lot of people who are still very tightly in the loop, auto-completion, prompting, need to think about, elevating this to system thinking. So what we're really trying to do is minimize the human touches, but still with good engineering practices. And some of the narrative that comes up more often, in the beginning we're like, oh great, vibe coding, a prompt and it gives a result and we can keep going. Where we now see, we're instructing it through prompts, but we're also instructing this like please do it with tests, please update the documentation, please do this. All the things that we're saying to good engineers, we're now asking the agents to do. So if you still have people who are YOLOing their way into this, I think you should tell them no, stop doing this. Engineering practices still matter for you to maintain the system and also for the agent to keep getting better at this. What I started seeing in some of the more advanced teams is that their rituals of, hey, we're doing a planning and we're doing a retro in a team, they weren't about, hey, we had issues with the code, but we're seeing we had issues with the system. So on the retro part it's like, hey, the agent went over and over, hit this problem, can we fix the system? That's something you learn in the retro. And on the planning side, what I started seeing is that things that were sufficiently scoped were easy to pick up by agents because they were well defined, and what still was left for the humans were the things that weren't scoped out well. So there was a split in the planning where you said these things can straight go into agents, well defined, and the harness is getting better, and these are conversational things that we need to decide as a team. And what I find important is there's a certain cycle that developers go through. Yes, they learn first about prompting, they get better specs, context, harness, loop, also the industry is learning like that. But the lead of the team can say, well, stop prompting, make the context reusable. Now we got that, now we jump to the next. So part of the team lead is putting that pace, almost that constraint and the directive in the team, where it doesn't work if you just say go figure it out and do something on your own. And one of the impacts of that is that if you start producing more as a team, the people downstream, GTM, people like that, they have a hard time keeping up. Even users have a hard time keeping up. So you need to help them also with automation. So your harness doesn't stop at your coding. It also is extended to those people as well. And the same thing with gathering requirements. The input might not come fast enough for your team. So that's another piece that you need to tap into, that workflow as well. There's a lot of metrics that people are saying, like, hey, is your token spend and all that stuff. I started to believe in these two metrics to see how to be more productive. One is you start measuring how many human touches you still do to have the agent do the right thing. That's supposed to go down. The better your harness is, the better your context is, the better your guidelines are. And on the other hand, if you're going from solo to shared system, that becomes a multiplier. You fix something once, everybody gets the benefit. This is not the multiplier from the one person becoming the 10x person, but the one change that optimizes the agents has an impact on all the people. So that is the path that we're on. You can start that in a team, working together within your repo, sharing the context, working on harness. But what you want to do is scale this out. So you come into the realm of the platform people, right? Because they're the typical shared organization working on this. Now the platform people, they might not be paying close attention because they're infrastructure and cloud and working on an MTP gateway and stuff like that. But there are new things bubbling up there. They need to think about maybe skill registries or eval systems for context and guardrails specifically for coding agents and identities and stuff. So they maybe need a little bit of a hand growing into that role. And that central role, it's hard. You need an owner to drive that program. But is it the impact on all the people. So that is the pot that we're on. You can start that in a team, working together within your repo, sharing the context, working on harness. But what you want to do is you want to scale this out. So you come into the realm of the platform people, right? Because they're the typical shared organization working on this. Now the platform people, they might not be paying close attention because they're infrastructure and cloud and working on an MTP gateway and stuff like that. But there's new things bubbling up there. They need to think about maybe skill registries or eval systems for a context and guardrails specifically for coding agents and identities and stuff. So they need maybe a little bit of a hand growing to that role. And that central role, it's hard. You need an owner to drive that program. But is it the platform team? Is it developer experience team? They don't typically own any of those pieces of the infrastructure, and the other people don't really do the development. So there's somewhere a blend. But you need to make sure that there's an owner driving this centralized piece and not just within your team. Because you won't have paved roads. And that's how I see it. Reusable context across teams. Why are we all inventing how we do the authentication system? Right? This is a shared component. Let's put it in the registry. Why are you building all your harnesses? Well, if we're all using the same linters and the same security tools, that's a reusable component. So I think that will centralize similar to the paved paths for cloud into that platform registry of reuse. But if everybody can put stuff on the internet in the repo, it becomes a sprawl. And it becomes a thing like, well, he has a skill, he's maintaining it. That person also has a similar skill and forked it. Now what do I do? Which one do I pick? So there is a thing that you say, there's an owner for this area. And they also care about making it testable. They make sure that it's modular, that other people can extend the context, for example, or the harness that is security scan. So you build more centralized, and the fact that it's secured and maintained as something, instead of just something I share around in my organization. Now that consensus is hard. I'm not saying this is tabs versus spaces, but at times it feels like that. If you have two developer teams having to have consensus on the way they work, that requires a lot of communication and brokerage. So you probably don't end up with one thing but a catalog of three, four paved roads where they can pick off. And they can still do their own, but that's on the wrong budget. All right. The centralized pieces will be maintained, and that is supposed to be the easy way of adoption to go there. Now, if they do this blindly, we also want to make sure that they know what this costs. Because if we visualize the cost, they might be eager to do some optimization in there. And that is part of the platform team, making that visible. How much is this spending? How much is that helping? If I can reduce the number of iterations the agent has to run through, that is an optimization that I can run. But if I don't visualize that and just see the end result, then we don't know. Right? So that is part of the platform team helping people. And so what I'm arguing is that we should somewhere move from the solo developer to the team-shared context and pieces to a multiplayer system in your organization. And I think that's where the multiplication effect will happen. Right? Because you have this flywheel of improvements to go into multiple directions. Now, one layer higher, the VP engineer says, how do I enable the organization? Right? And that is, I can predict the story in your organization. A hackathon, a lunch and learn, let's share the successes, have a shared Slack channel, have a champions program. That's all generic transformation. It could have been agile that transformed like that. It could have been DevOps. It doesn't matter. And on the other side, we know that the strategy of just give licensing and educate people, do something, let a thousand flowers bloom, it doesn't work. So what I'm advocating is that on the organizational side, you give the team leads and the platform that mandate to start doing that work. And it's not a solo developer piece. Now, finding people that help you externally is a mess. Yes, we have all the titles, the new job titles, AI product engineer, forward deployed engineer, there was a whole talk on this, agentic engineer, AI engineer. It doesn't mean anything. You cannot judge the maturity of this because nobody's really that mature. But it's a signal when you put a job posting out there that people might, with the new intention, be looking there. But it's not a validation of the skills as such. Right? So that is challenging for people hiring people. Now, they come to the interview, and I heard stories about people using AI to reflect in their ears the response to the interview person and stuff like that. I think what I hear from most companies is they say, first step is we give them an exercise and we want them to really go nuts on AI to solve this. If they have help from AI, that's all good. That shows you how much they can leverage the AI to do this. Now, after they pass this, you do a walkthrough and you actually say, please explain to me what happened. Why is this a good idea? That's where you are testing the taste and the engineering skills on why they're doing this. First part AI, then engineering. And this third thing is how do you collaborate? Are you willing to share? Are you open, or are you a solo player? That's another signal that you tap into. Right? But that fits into that whole thing of making it shareable, making it reusable, making it engineering great within an organization. Those are the people that you look for. Not people who've studied ML or AI, not people who are experts per se at the coding. There's a blend on this. Now, you might not find a person who has all three, which is okay. But at least you know, hey, they're very savvy on this piece, but then for the other piece, they need mentoring and they need tutoring. But don't put all three pieces into one, saying they're junior or senior. They have different skills on there. Now, the VP engineering has to defend this, and they have to make the case, right? Well, we have X amount of licenses that we sold. We have faster delivery. Maybe they can promise, but hard to prove. We have quality that improved. Again, hard to say. But similar to what I said with the metrics of how effective are your agents, you can show how much turns and how much improvement that you're making on that journey. And same thing, how much there is reuse. So it's an easier way to show metrics than comparing productivity with and without agent decoding. That helps you in those discussions as well. And so when people say, the vendors are charging completely nuts, so we're going to limit the spend, you shouldn't say, let's limit all the spend. Your reflex should be, let's optimize the spend and help them reduce that in a good way. Whether that's as simple as saying, pick the right model, educate them on the model, but also on giving them better context and harnesses. Because that will make your cost go down there as well. The debate on smaller and bigger teams. Yes, it's nice to have one person who can do it all. That's the ultimate dream. They can do everything. Typically, they're paired with a complementary skill, maybe PM, design, and so on. Okay, then we need a backup if one of them is on holiday. So that amounts back to three. And then maybe somebody has to care about production and tickets coming in. It could be the same people if you're really productive. But you lose speed of features if you're still doing bugs. And that depends a little bit on your quality. And then there's the junior you want to get on the road as well to make sure they're still learning what good looks like in one of those three areas. So I think we're still limited in the way in an organization that we're not going to each team being a solo or one or two. Yes, a lot of experience, but I think that is the thing. Now we keep investing in education for that piece as well. So one of the final things is the dark factory, which is probably a dim factory. You have to see what risk you're willing to take for what features. So not all features will become autonomous. But you can invest more in auditing, like provenance, like who changed the code, verifiers that check whether that code was useful. And when it fails, you invest in situational awareness as well. And that depends a little bit on your quality. And then there's the junior you want to get on the road as well, to make sure they're still learning what good looks like in one of those three areas. So I think we're still limited in the way, in an organization, that we're not going to each team being a solo or one or two. Yes, a lot of experience, but I think that is the thing. Now we keep investing in education for that piece as well. So one of the final things is the dark factory, which is probably a dim factory. You have to see what risk you're willing to take for what features. So not all features will become autonomous. But you can invest more in auditing, provenance, who changed the code, verifiers, that kind of check whether that code was useful. And when it fails, you invest in situational awareness as well. So there's a whole spectrum from being a micromanager to being on autonomous approval that everything is correct. But you make a decision on what your risk level is. And I think your mode is capturing the knowledge, the knowledge you're putting now into skills in your context and maybe in your harness, the way you restrain this, your business context. And for me, that brings continuous delivery actually to continuous learning. And if you ask the question, how fast can we swap in, swap out something new? That's your reactive mode. And if you can improve that, ultimately, it's not about making the whole system more reliable, but can I keep it reliable while changing more of the system? I'm working on a website where I try to list some of the agent enablement patterns that I described. I couldn't list them all within this time. Tell me what you're missing. I'm trying to source social stories. So if you have a story of how things are going in your organization, please tell me. And I'm happy to put a link in there as well. And if you're interested in the slides, happy to share those. And I think if there's one takeaway, it's not the solo player that will win the game. It's at the different levels how we improve our organizations. Thank you very much for listening, and I hope it was useful. of years who said it like stop building the thing but build the thing that builds the thing. Right? So we going on that abstraction where that is with context, with harness, with loops and that is kind of the change that a lot of people who are still very tightly in the loop, auto-completion, prompting, that they kind of need to think about elevating this to the system thinking. So what we're really trying to do is minimize the human touches but still with good engineering practices. And some of the narrative that comes up more often in the beginning we're like oh great, vibe coding, a prompt and it gives a result and we can keep going. Where we now see well we're kind of instructing it through prompts but we're also instructing this like please do it with tests, please update the documentation, please do this. All the things that we're saying to good engineers we're now asking the agents to do. So if you still have people who kind of YOLOing their way into this I think you should tell them no stop doing this. Like engineering practices still matter for you to maintain the system and also for the agent to keep getting better at this. What I started seeing in some of the more advanced kind of teams is that their rituals of hey we're doing a planning and we're doing a retro in a team that they weren't about like hey we had issues with the code but we're seeing we had issues with the system. So on the retro part is like hey the agent went over and over hit this problem can we fix the system? That's something you learn in the retro. And on the planning side what I started seeing is that things that were sufficiently scoped enough were easy to pick up by agents because they were well defined and what still was left for the humans were the things that weren't scoped out well. So there were like a split in the planning where you said these things can straight go into agents well defined and the harness is getting better and this is conversational things that we need to decide as a team. And what I find important is you there's a certain kind of cycle that developers go to. Yes they learn first about prompting they get better specs, context, harness, loop, also the industries learning like that. But there is the lead of the team can say well stop prompting make the context reusable. Now we got that now we jump to the next. So part of the team lead is putting that pace at almost that constraint and the directive in the team where it is doesn't work where you just say go figure it out and do something on your own. And one of the impacts of that is that if you start producing as a team more the people downstream, GTM, people like that, they have a hard time keeping up. Even users have a hard time keeping up. So you need to help them also with automation. So your harness doesn't stop at your coding. It also is extended to those people as well. And the same thing with kind of requiring like gathering requirements. The input might not come fast enough for your team. So that's another kind of piece that you need to tap into that workflow as well. There's a lot of metrics that people are saying like hey is your like tokens spend and all that stuff. I started to believe in these two metrics to kind of see on how to be more productive. One is you start measuring how many human touches you still do to have the agent do the right thing. That's supposed to go down. The better your harness is, the better your context is, the better your guidelines are. And on the other hand, if you're going from solo to shared system, that becomes a multiplier. You fix something once, everybody gets the benefit. This is not the multiplier from the one person becoming the 10x person, but the one change that optimizes the agents has an impact on all the people. So that is kind of the pot that we're on. You can start that in a team, working together within your repo, sharing the context, working on harness. But what you basically want to do is you want to scale this out. So you come into the realm of the platform people, right? Because they're the typical shared organization working on this. Now the platform people, they might not be paying close attention because they're like infrastructure and cloud and working on an MTP gateway and stuff like that. But there's new things like bubbling up there. They need to think about like maybe skill registries or eval systems for a context and guardrails specifically for coding agents and identities and stuff. So they need maybe a little bit of a hand kind of growing to that role. And that kind of central role, it's hard. You need an owner to drive that program. But is it the platform team? Is it developer experience team? They don't typically own any of those pieces of the infrastructure and the other people don't really do the development. So there's somewhere a blend. But you need to kind of make sure that there's an owner driving this centralized piece and not just within your team. Because you won't have paved roads. And that's how I see it. Reusable context across teams. Why are we all inventing how we do the authentication system? Right? This is a shared component. Let's put it in the registry. Why are you building all your harnesses? Well, if we're all using the same linters and the same security tools, that's a reusable component. So I think that will centralize similar to the paved paths for cloud into that platform registry of reuse. But if everybody can put stuff like on the internet in the repo, it becomes a sprawl. And it becomes a thing like, well, he has a skill, he's maintaining it. That person also has a similar skill and forked it. Now what I do? Like, which one do I pick? So there is a kind of thing that you say, there's an owner for this area. And they also care about making it testable. They make sure that it's modular, that other people can extend kind of the context, for example, or the harness, that is security scan. So you build kind of more centralized and the fact that it's secured and kind of maintained as something, instead of just something I share around in my organization. Now that consensus is hard. I'm not saying this is tap versus spaces, but at times it feels like that. If you have two developer teams having to have consensus on how the way they work, that requires a lot of communication and brokerage. So you probably don't end up with one thing but a catalog of three, four paved roads where they can pick off. And they can still do their own, but that's on the wrong budget. All right. The centralized pieces will be maintained and that is supposed to be the easy way of adoption to go there. Now, if they do this blindly, we also want to make sure that they know what this costs. Because if we visualize the cost, they might be eager to do some optimization in there. And that kind of is part of the platform team is making that visible. How much is this spending? How much is that kind of helping? If I can reduce the number of iterations the agent has to run through, that is an optimization that I can run. But if I don't visualize that and just see the end result, then we don't know. Right? So that is part of the platform team helping people. And so what I'm arguing is that we should somewhere move from the solo developer to the team shared kind of context and pieces to a multiplayer system in your organization. And I think that's where the multiplication effect will happen. Right? Because you have this flywheel of improvements to go into multiple directions. Now, one layer higher, the VP engineer says, how do I enable the organization? Right? And that is, you know, I can predict the story in your organization. A hackathon, a lunch and learn, let's share the successes, have a shared Slack channel, have a champions program. That's all generic transformation. It could have been agile that transformed like that. It could have been DevOps. It doesn't matter. And on the other side, we know that the strategy of just, you know, give licensing and educate people, do something, let a thousand flowers bloom, it doesn't work. So what I'm advocating is that the kind of on the organizational is that you give the team leads and the platform that mandate to start doing that work. And it's not a solo developer piece. Now, finding people that help you externally is a mess. Yes, we have all the titles, the new job titles, AI product engineer, forward deployed engineer, you know, there was a whole talk on this, agentic engineer, AI engineer. It doesn't mean anything. You cannot judge what kind of the maturity of this because nobody's really that mature. But it's a signal when you put a job posting out there that people might with the new intention, they'll be looking there. But it's not a validation of the skills as such. Right? So that is challenging for people, kind of hiring people. Now, they come to the interview and I heard stories about people using AI to reflect in their ears, be response to the interview person and stuff like that. I think what I hear from most companies is they say, first step is we give them an exercise and we want them to really go nuts on AI to solve this. You know, if they have help from AI, that's all good. That shows you kind of like how much they can kind of leverage the AI to do this. Now, after they pass this, you do a walkthrough and you actually say, please explain me what happened. Why is this a good idea? That's where you are testing the taste and the engineering skills on why they're doing this. First part of AI, then engineering. And this third thing is how do you collaborate? Are you willing to share? Are you open or are you a solo player? That's another signal that you tap into. Right? But that fits into that whole thing of like making it shareable, making it reusable, making it engineering great within an organization. Those are the people that you look for. Not people who've studied ML or AI, not people who are like experts per se at the coding. There's a blend on this. Now, you might not find a person who has all three, which is okay. But at least you know like, hey, they're very savvy on this piece, but then for the other piece, they need mentoring and they need tutoring. But like, don't put all the three pieces into one kind of saying like they're junior or senior. They have like different skills on there. Now, the VP engineering has to defend this and they have to make the case, right? Well, we have X amount of licenses that we sold. We have faster delivery. Maybe they can promise, but hard to prove. We have quality that improved. Again, hard to say. But similar to what I said with the metrics of how effective are your agents, you can show that how much turns and how much improvement that you're making on that journey. And same thing, how much there is reuse. So it's an easier way to kind of show metrics than comparing productivity with and without agent decoding. That help you in kind of those discussions as well. And so when people say, the vendors are charging completely nuts, so we're going to limit the spend. You shouldn't say like, let's limit all the spends. Your reflex should be, let's optimize the spend and help them kind of reduce that in a good way. Whether that's as simple as saying, pick the right model, educate them on the model, but also on like giving them better context and harnesses. Because that will make your cost go down there as well. The debate on smaller and bigger teams. Yes, it's nice to have like one person who can do it all. That's the ultimate dream. They can do everything. Typically, they're paired with a complementary skill, maybe PM, design, and so on. Okay, then we need a backup if one of them is on holiday. So that amounts back to three. And then maybe somebody has to care about production and tickets coming in. Could be the same people if you're really productive. But yeah, you know, you lose speed of features if you're still doing bugs. And that depends a little bit on your quality. And then there's the junior you want to get on the road as well to kind of make sure they're still learning what good looks like in one of those three areas. So I think we're still limited in the way in an organization that we're not going to each team being a solo or one or two. Yes, a lot of experience, but I think that is the thing. Now we keep investing in actually education for that piece as well. So one of the final things is the dark factory, which is probably a dim factory. You have to see what risk you're willing to take for what features. So not all features will become autonomous. But you can invest more in auditing like provenance, like who changed the code, verifiers, that kind of check whether that code was useful. And when it fails, you invest in situational awareness as well. So there's a whole spectrum from being a micromanager to being on autonomous approval that everything kind of is correct. But you make a decision on what your risk level is. And I think your mode is capturing the knowledge. The knowledge you're putting now into skills in your context and maybe in your harness, the way you kind of restrain this, your business context. And for me, that kind of brings continuous delivery actually to continuous learning. And if you ask the question, how fast can we swap in, swap out something new? That's your reactive mode. And if you can improve that, ultimately, it's not about making the whole system more reliable, but can I keep it reliable while changing more of the system? I'm working on a website that kind of where I try to list some of the agent enablement patterns that I described. I couldn't list them all within this time. Tell me what you're missing. I'm trying to source social kind of stories. So if you have a story of how things are going in your organization, please tell me. And I'm happy to put on a link in there as well. And if you're interested in kind of the slides, happy to share those. And I think if there's one takeaway, it's not the solo player that will win the game. It's kind of like at the different levels, how we improve our organizations. Thank you very much for listening and I hope it was useful.