Open Reader

Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs

completed 18:07 Aug 08, 2026 Watch on YouTube

Current Status

completed

Video ID

7vn4WpqNpck

RAG / Chat

Enabled
Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs
Description

Wisedocs processes medical claims that arrive as PDFs over 10,000 pages long, some of them larger than video files, through a pipeline of ML models spread across ten repositories nobody enjoyed touching. Denys Linkov's team spent six months collapsing that into a monorepo, and this talk is an honest audit of whether they should have just waited for the models to get good enough to do it for them. The benchmark he keeps coming back to is a single refactor task. With o3 it took three hours of back and forth in Cursor and still shipped ten major mistakes. Rerun on newer models, Sonnet 4.6 needed one extra iteration and Opus 4.8 essentially got it in one pass, at roughly a fifth of the original effort. The counterweight is what happens when you hand a current model the whole job. GPT 5.5 extra high declared the refactor done in 10 minutes 22 seconds and wrote 2,000 lines, which turned out to be scaffolding with the actual models missing, something it admitted in its own output by noting it had not added the deployment or bootstrap command yet. That gap is why Linkov reads the METR task length curve at 80% or 90% success instead of the usual 50%. Launching an hour long agent run on coin flip odds mostly buys you a wasted hour and a broken attention span. His verdict is that doing the refactor beat deferring it, and the evidence is as much social as technical: commit velocity rose and never flattened, work that used to take months ships in under a week, and developers across the company now volunteer into the repo even outside their own area, which was never true of the ten it replaced. Speaker info: - https://x.com/denyslinkov - https://www.linkedin.com/in/denyslinkov/

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Wisedocs' six-month refactor from more than 10 legacy repositories into a monorepo was worth doing despite rapidly improving coding agents, because architectural cleanup and verification discipline unlocked faster delivery, lower costs, larger-file support, and broader developer participation.
  • Why it matters: For AI-enabled engineering organizations, stronger models reduce the cost of refactoring but do not remove the need for clean system boundaries, end-to-end testability, explicit specifications, and human-validated guardrails.
  • Best use: Use this as a practical case study for deciding whether to defer a major cleanup for better agents or refactor now while building an agent-ready engineering harness.

Executive Summary

Denys Linkov presents Wisedocs' response to a scaling crisis: its pipeline for processing complex medical-claims PDFs, sometimes exceeding 10,000 pages, was too slow to meet demand, difficult to update, and spread across more than 10 legacy repositories that engineers avoided touching. The company chose a six-month refactor, consolidating into a monorepo and rebuilding the pipeline while preserving business functionality.

The talk's central comparison is between old and current coding-agent capability. An earlier O3-assisted Temporal refactor took roughly three hours of interactive work and required correction of 10 major mistakes. Re-running the task with Sonnet 4.6 required one extra iteration, while Opus 4.8 reportedly solved it in one shot. The improvement was not only model quality: modern harnesses also use planning, subagents, shell commands, and verification steps that reduce manual intervention.

Linkov argues against treating headline agent benchmarks as proof that agents can reliably own long autonomous projects. He prefers evaluating success at 80%, and ideally 90-99%, rather than the common 50% threshold, because a one-hour agent run with a 50% completion rate can waste both compute and human attention. His own zero-shot test of GPT 5.5 extra high appeared successful in 10 minutes 22 seconds but delivered only 2,000 lines of scaffolding and omitted critical model/deployment work.

The refactor paid off operationally: Wisedocs reduced pipeline time and cost, supported larger files, and moved feature work from multiple months to under a week. The new repository also became a cultural and organizational asset, with nearly all developers able to contribute across schemas, APIs, and other layers. The key qualification is that this success depended on requirements checking, evolving AI-verification processes, human PR review, and deliberate end-to-end validation—not autonomous coding alone.

Key Takeaways

  • Claim: A major refactor can be justified now even if coding agents will make future refactors cheaper, when technical debt is already constraining customer delivery and operational scalability. | Evidence: Wisedocs was adding customers without sufficient throughput; its complex AI pipeline was hard to update and distributed across more than 10 repositories. After the refactor, it reduced pipeline time and costs, supported larger files, and delivered features that previously took multiple months in under a week. | Implication: Decide refactor timing by the current business cost of slow delivery, high operating cost, and inaccessible code—not solely by the expectation that next year's agents will be better. | Caveat: The speaker does not claim that every organization needs a full rewrite; he says teams can isolate portions of a codebase and pursue partial refactors when that better fits business value.
  • Claim: Modern coding agents materially improve refactor execution, but the gain comes from models plus the surrounding harness rather than a model alone. | Evidence: An earlier O3 implementation of a Temporal workflow took about three hours of Cursor back-and-forth and made 10 major mistakes. On the same task, Sonnet 4.6 succeeded with one additional iteration and Opus 4.8 one-shot it; Linkov also observed more planning, subagent use, shell commands, and verification in modern harnesses. | Implication: Invest in planning, tool access, sandboxing, tests, and verification workflows alongside model selection; improving the harness can convert raw model capability into lower-touch engineering throughput. | Caveat: The examples are one organization's task-level benchmarks rather than a general controlled comparison across codebases.
  • Claim: Agent reliability should be judged at high success thresholds, not the 50% task-completion figures commonly used in frontier-model capability charts. | Evidence: Linkov argues that an hour-long agent run with a 50% chance of completion risks wasting an hour of compute and attention. He cites METR-style task-duration curves and notes that success declines sharply around the four-hour mark for a referenced frontier-model evaluation, with meaningful failures even on much shorter tasks. | Implication: For unattended or high-cost workflows, set acceptance criteria around 80% or preferably 90-99% reliable completion, and retain checkpoints for tasks outside demonstrated reliability envelopes. | Caveat: The transcript does not provide the underlying benchmark methodology or identify the referenced model clearly beyond a likely transcription error.
  • Claim: A well-written spec is now a primary control surface for agentic development, because incomplete requirements or flawed plans cause expensive autonomous failure loops. | Evidence: Linkov asks the audience who has launched an agent only to realize that the prompt, plan, or requirements were incomplete. In Wisedocs' orchestration selection work, the team evaluated five open-source projects against 17 criteria, built proofs of concept with three people, and ultimately met 15 of the 17 requirements. | Implication: Treat research, requirements decomposition, and POC-based validation as first-class agent workflows; do not promote research output directly into architecture decisions without checking product reality and hidden assumptions. | Caveat: Deep-research outputs can create false confidence: Linkov warns that a polished 20-page report may claim features that do not actually exist in the evaluated product.
  • Claim: A monorepo improved Wisedocs' agent and developer operating environment, especially for end-to-end testing, deployment, and sandbox setup. | Evidence: The company consolidated its legacy repositories into a monorepo. Linkov says current models can navigate multiple repositories when placed under a higher-level directory, but full end-to-end testing, verification, deployment, and cloning/setup for an AI-factory sandbox remain harder across multiple repos. | Implication: If cross-repo coordination is preventing reliable agent execution, prioritize a unified development/test environment or equivalent integration layer before expecting autonomous agents to manage system-level changes. | Caveat: He does not argue that agents cannot work across repositories; the limitation is most acute for integrated verification and execution rather than repository navigation itself.
  • Claim: Passing a zero-shot agent run is not evidence that a refactor is complete; agents can satisfy superficial goals with scaffolding while omitting critical implementation. | Evidence: GPT 5.5 extra high completed a stated end-to-end refactor goal in 10 minutes 22 seconds and wrote about 2,000 lines, but inspection found it had created scaffolding rather than implemented the models; a required Ray Serve deployment/bootstrap command was absent. | Implication: Evaluate coding agents on executable acceptance criteria—implemented components, deployment behavior, integration tests, and production-relevant validation—not token count, elapsed time, or a self-reported completion state. | Caveat: Linkov predicts substantial refactors may become consistently one-shot capable within six months, but this is his forecast rather than demonstrated current capability.
  • Claim: Codebase cleanup has a compounding organizational benefit: it expands who can safely contribute and spreads better patterns beyond the refactored system. | Evidence: After the rebuild, nearly every developer at Wisedocs was contributing to the monorepo, including outside their previous specialty through schema and API changes. Developers actively requested to work in the cleaner codebase, and its patterns spread to other company repositories. | Implication: Measure refactor ROI beyond output velocity: contribution breadth, onboarding friction, shared context, and the portability of engineering standards can be durable returns.

Detailed Brief

How Wisedocs would redo architecture research with agents

  • Claims: The team believes its two-month orchestrator evaluation could now be completed roughly 90% faster using an agentic workflow.; Agent-assisted research should be decomposed rather than delegated as a single undifferentiated deep-research request.
  • Evidence: The original process assessed five open-source orchestrators against 17 internally defined criteria, documented findings in Confluence, and used a three-person team to build proofs of concept.; The proposed current workflow begins with deep research tied to the actual problem statement, assigns subagents to product-and-criterion combinations, then proceeds to POCs and evaluation.
  • Caveats: The claimed 90% reduction is a retrospective estimate, not a measured rerun.; Library capability claims require direct verification because features may exist only in beta or not at all.
  • Implications: Separate discovery from decision validation: automate evidence gathering and comparison, but make implementation POCs the decision gate.; Build reusable evaluation templates that encode product requirements, source checking, test scenarios, and rejection conditions for future tool selection.

Governance practices used during the refactor

  • Claims: The verification process matured during the project rather than being fully designed at the outset.; Human PR review served both as quality control and as a mechanism for building shared repository context.
  • Evidence: Plan mode was only emerging in Claude Code and was unavailable in Cursor when the work began, but Wisedocs incorporated it into its development lifecycle as tools evolved.; PR reviews were entirely human during the refactor, supplemented by local AI checks that asked tools to review code quality.
  • Caveats: The team expects review to become more autonomous over time, but the transcript offers no evidence that autonomous review was sufficient for this refactor.
  • Implications: Do not remove human review merely because coding generation improves; use the review process to distribute architectural understanding while the codebase is changing.; Version-control the engineering workflow itself—planning, local checks, acceptance tests, and review expectations—so agent adoption does not outrun governance.

Notable Concepts & Terms

  • Technical debt as financial debt: Linkov frames debt as complexity that compounds and can eventually exceed the ROI gained from rapidly shipping features or acquiring customers.
  • Monorepo: Wisedocs' consolidation of more than 10 repositories; it improved integrated development, sandbox setup, end-to-end validation, and contribution across teams.
  • Agentic workflow: A multi-stage workflow using research, decomposed subagents, tool calls, planning, implementation, and verification rather than a single model prompt.
  • AI psychosis: The speaker's term for over-trusting polished AI research outputs without validating whether their factual claims, such as product features, are real.
  • METR task-duration reliability curve: A benchmark framing model capability by the duration of tasks it can complete; Linkov stresses judging it at high reliability rather than 50% success.
  • Plan mode: An agent-development step adopted by Wisedocs to improve requirement and implementation planning before code generation.
  • Ray Serve deployment/bootstrap command: A missing operational component that exposed the gap between an agent's apparent refactor completion and an actually runnable deployment.

Operator Notes / Why Ken Should Care

  • Create a refactor decision scorecard that quantifies delivery delay, run-cost burden, scalability limits, contributor friction, and customer-impact risk against the cost of a cleanup.
  • Benchmark coding agents using repository-level acceptance suites: required components present, build/deployment bootstraps, end-to-end tests, performance checks, and reviewable diffs—not claimed task completion.
  • Set an autonomy policy by reliability tier: unattended execution only for task classes that demonstrate high historical pass rates; require plan approval and staged checkpoints for longer or integrated changes.
  • For multi-repository agent work, prioritize a reproducible unified sandbox and integrated test/deploy path before attempting broad autonomous refactors.
  • Turn tool and framework selection into a reusable agent-assisted pipeline with independently verified capability evidence and mandatory POCs for critical requirements.

Source/Metadata

  • Title: Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs
  • Transcript words: 4823
  • Duration seconds: 1087
  • Timestamp note: No timestamps or chapter markers were provided in the transcript; portions of the Q&A and conclusion are duplicated.

Transcript

3402 words en Processed in 101.0s

It's not just my AI pipeline that's on fire, but also my PowerPoint. So it's 2025, we're scaling as a business, and things are going poorly. We're adding too many customers, we're not getting the throughput we need, and we need to improve our underlying technology. And there's three main issues that we're facing. The first one is that we're too slow to meet customer demand. The second one is that this AI pipeline that we've built is too complicated to update. And the third one is that because it's a legacy code base, or actually more than 10 repos, nobody actually wants to touch the code. It's not a fun experience. So we made this decision to refactor over the course of six months. And the real question for this talk today was, is this the right move to do? So I'll spend this time answering this question, but let's start off with the use case. The company I work at, WiseDocs, processes complex medical claims, which are PDFs that are more than 10,000 pages in size. Some of these files are bigger than video files. So it's a pretty complex application. And because of this, it's actually non-trivial to scale the different parts. So we're going to talk about the pipeline today, which has a number of ML models. To divide this talk into a number of chapters, we'll start off with the first one, which is the concept of tech debt. So I think we all have this feeling universally, if we've been developers for a while, that we all write bad code. The question is, do we do this intentionally or not? And if I look back to some of the earliest code I used to write, it was bad. This was more than 15 years ago. I tried to print an image of this character from a video game, and I didn't understand that you can't system.out.println in Java to render something on the screen. So hopefully I've come further from that point in time, but there are these moments where we all know that we've written bad code before. Now, if we think about technical debt as financial debt, it compounds in mysterious and sometimes unexpected ways. But you should think about it in a rigorous format as well. For us to achieve some kind of ROI by taking on technical debt, such as building a feature or getting new customers, we want to make sure that the ROI makes sense. If we introduce additional complexity into our code base, we can very quickly outrun the ROI we've generated. Now, with AI engineering, you've probably seen a number of different stories that have come out to showcase the progress that's been made. These are two case studies from Anthropic, one from Spotify and the other from Stripe, talking about the immense progress that they've made, both in shipping velocity and also the ability to refactor code. So at this point in time, writing code or making changes is something that teams are doing faster and faster. Now, I'll pause here. Who here thinks that products have gotten better in the past 20 years? Technical products. Also raise your hand. I hope everybody, right? Phones are pretty cool. About five years. About the past year. Okay. So the challenge is that we're going faster and faster through the technology life cycle, but we've lost something. The product focus on customers, in some way, has degraded. The maintainability of the code and the reliability has degraded. You can see some of the uptimes here from two leading companies. I've blurred out their names, for it doesn't actually matter who they are, but we are below a 3.9 or even a 4.9 reliability. So even though we're shipping faster and faster, the code quality and the product quality have not necessarily gone up. So let's talk about the refactor that we did. So we started this refactor with actual code implementation in April and did some pre-work earlier. So I'll go through five different tasks that we did and share some of the findings that we had before and after, especially as new models have come out. So we spent around two months evaluating orchestrators for our AI pipeline. We looked at five open source projects, and we wanted to benchmark and see how effective they were for our use case. And we started this off before deep research came out as part of Google and OpenAI, so that web search capability to do a comprehensive analysis was still not there. Now, after we actually gathered these requirements, we built out proof of concepts with a team of three to make sure that we actually got the right results. Now, I'm pretty confident we could do this 90% faster now with the tooling that we have. Before, we would manually go through, use a little bit of AI, but put everything into a Confluence doc, and we'd evaluate across 17 different criteria that we came up with. Nowadays, we could build a much more agentic workflow to do that, starting off with deep research, making sure that we match that against the problem statements that we have, creating subagents for each of these criteria and products, and then finally building POCs and evaluating. So things have changed in the past year and a half, where we could actually go much, much faster. But we still have to maintain that same set of quality because it's very easy to undergo AI psychosis, where you look at a deep research report that's 20 pages long, and you say, wow, this looks good, and then those features don't actually exist in the product, and you set yourself back. Now, after we've done the initial orchestration research and model-serving research, we went into actually committing code. This is just an example of what happened when we were experimenting. I was doing some initial research with Temporal and committed some activities and workflow code to make sure that we can actually replicate what we have in the legacy code base. So then I did what we wanted to do over a number of iterations, and at the time gave it to O3 to actually try to implement this code. And it did it much faster than I would have. This refactor took three hours of back-and-forth chatting within Cursor, but it made ten major mistakes. So at the time when we were going through this refactor, agentic coding was getting better and better, but it still hadn't reached the point of where it is now. And it was still a very manual process where you had to intervene and actually guide the model and manually edit or delete code. Now, I re-ran these benchmarks on some modern models. So we have Sonnet 4.6 and Opus 4.8, and things were much faster. Sonnet 4.6, with one additional iteration, was able to solve the task, and with Opus, it was able to one-shot this problem. So models are getting significantly better along with harnesses. And the interesting part here as well is that the way that the models interacted has changed substantially as well. Before with O3, there weren't substantial tool calls on certain categories, and then as we moved into Sonnet 4.6 and Opus, we see now that in modern harnesses, we get sub-agents, we get some of those planned calls, we get different shell commands, and we get different verifications. And overall, this process, even though the model execution was a little bit more expensive, was a lot less manual, so we could actually accomplish a lot more. So if I was rebuilding the same task that I had for this refactor, it would take around one-fifth of the time to accomplish, which is pretty good progress. So I think all of us realize the scenario that models are substantially better now than they were before. Now, this is really important because it shapes the way we think about the software development lifecycle. We think about 2025 and the types of work that we were doing. We were making some small changes. We would give specific code snippets to models. We were just starting to get into this agentic framework of the type of work we can do. And now, if we provide a well-constructed spec to a model, it could generally execute it at a very high capability level. And we can see this both in anecdotal experiences as well as some of the thought leadership that has been coming out of the big labs. This image is one from Anthropic. Let me ask the group a question. Who here has kicked off an agent and realized that either the prompt, the plan, or the requirements were incomplete or missing? A lot of people, yeah? It's very frustrating, right? You're like, okay, I'm ready to go. It's 11 p.m. or 5 p.m. I'm going to set off an agent and then come back. And then you realize there's a critical flaw. Now, the reason I bring this up is it's very important to have a good mental model and understanding of how accurately models can accomplish tasks. Who here has seen this meter graph before? I think a decent number of people. So this is pretty common in actually mapping how much time models can complete tasks of certain categories for. So the idea being that as models get better and better, they can do longer-running tasks. Now, typically, this graph is shared with the 50% accuracy rate, but I think it's much better to actually look at the 80% accuracy rate or higher. And you can see there, you can still see a similar exponential trend, but we're no longer claiming that models can accomplish tasks that would take a human 18-plus hours. Now, I actually think it's much better I'm going to set off an agent and then come back. And then you realize there's a critical flaw. Now, the reason I bring this up is it's very important to have a good mental model and understanding how accurate models can be in accomplishing tasks. Who here has seen this meter graph before? I think a decent number of people. So this is pretty common on actually mapping how much time models can complete tasks of certain categories for. So the idea being that as models get better and better, they can do longer-running tasks. Now, typically, this graph is shared with the 50% accuracy rate, but I think it's much better to actually look at the 80% accuracy rate or higher. And you can see there, you can still see a similar exponential trend, but we're no longer claiming that models can accomplish tasks that would take a human 18-plus hours. Now, I actually think it's much better to measure the accuracy at 90 or 99% because this is where the mental model is most efficient. You construct a plan, you create a spec, you hand it off to an agent, and you're pretty sure that it'll get things done, right? You don't want to be creating a plan or a spec and then have a 50-50 chance of coming back and knowing that you wasted compute and your attention span. Now, if you're kicking off a process that is going to take an hour and it has a 50% chance of completing, there's a very high chance you just wasted that hour and you could have been doing something different. If we think about broader evaluation, Meter does have some more information about their frontier models. So this is one for Mythos preview that they did roughly a month ago. And you can see here that generally, the success rate starts to decline significantly at that four-hour mark. But even before then, at the 15-second mark or even before the 15-minute mark, there are certain tasks that Mythos, in all its glory, cannot complete effectively and consistently. So we're making rapid progress in the AI model space, but we're still not there where you can just kick off an agent and have something be completed reliably. So again, this is really important for your software engineering teams and for you as an IC to understand what is your mental model and how are you going to contribute to that. And I think what's really important is, I think you've been hearing this throughout this conference, that there are a number of different frameworks and primitives that you need to have implemented in order to have good agentic development. And this is no different from what we found. As we were continuing to mature as an organization and going through our refactor, these are the things that made sure that we can implement the solutions effectively and not waste our time just running doom loops with models. So let's go into chapter three. Let's talk about the refactor itself and some of the productivity gains that we saw. So the core idea is that we had these 10 repositories, we put them into a monorepo, and we wanted to build additional features on top of it. So this is the result. The previous repos had been around for more than six years, and you can see the progress that was being made. It was pretty slow. Part of it was because of the tech debt that was taken on. Other parts were because we didn't have AI coding tools. And you can see that within the first six months of the rebuild, when we got to parity that we had before, the steepness of that curve is immense. And it didn't slow down after we kept shipping. So after that dotted line in the middle there, we kept adding new and new features into the repository. And we shipped a lot faster, both in terms of the amount of code, even though that's not a great metric, but also the commit rate that we had among developers. And we actually saw that a lot more developers joined in the contributions. So this is a log graph on the commits that we had from the repository initially. And then we slowly onboarded more and more people. And we had fewer commits because it's much easier to commit code when you're just refactoring and replicating something. But we still kept up that velocity as we were adding product features towards the end. And now almost every developer within the company is committing to this new monorepo, even though it might not be their area of expertise, but they might need to make changes to schemas, API calls, and other parts of the stack. So let's go on to chapter four. Can a modern LLM zero-shot this problem? Can I say, hey, amazing LLM, go refactor this code base? So I ran this experiment with GPT 5.5 extra high, and I gave it this goal, giving some of the names of the repositories with the underlying models and other components, and it completed its goal in 10 minutes and 22 seconds. Now, it only wrote 2,000 lines of code, which was a little bit fishy, so I dug deeper. And it actually just implemented a bunch of scaffolding and didn't implement the models. So you can see here, I did not add a raise serve deployment or bootstrap command yet, right? So we're still not there where models can self-validate and just one-shot these kinds of problems, but we're getting close. I think in six months, we'll get to the point that we can complete pretty substantial refactors, as we saw in the Stripe example, consistently across the board. So we get to the core question. Was this refactor worthwhile? Should we have waited a year to do this refactor as models and harnesses continue to get better, or did it make sense to do it at the time? Now, I'll say the other side of the argument, right? Things are getting substantially better. Models are getting better. They can call tools better. We have a lot more infrastructure, like sandboxes and monitoring frameworks, in order for us to actually understand what's happening under the hood with these models. So taking on technical debt and refactoring later is getting exponentially easier as the days go by. Now, the problem is that a lot of times when you build a lot of code and you do this kind of development in an AI-native world, it starts looking like some of the legacy code we've seen in the past. There's a lot of code written. It's written with low performance or quality, and the broader problem is people don't actually understand what's happening there. So if you have some issues within the code base or you want to adjust based on customer requirements, it's actually much harder to do so. So you do have to make sure that there are appropriate guardrails, whether or not you do a full refactor or only a partial one. So if you ask me, was it worthwhile? I'd say yes. We had built out the patterns that we had earlier with the number of different repos in order to match customer requirements and demands. It took an amount of time, but we ultimately achieved the goals of the business. And then we came back and refactored, and we were able to accelerate. We were able to actually reduce the amount of time the pipeline took. We were able to reduce the costs. We could support larger files. And now we can ship features that would take multiple months in under a week. So the monorepo refactor, the cleanup, was worthwhile, and we have some of the productivity metrics we saw there. The other part is that beyond just shipping velocity, developers actually want to work in this code base. So everybody comes along and says, hey, can I work in this code base? It's much cleaner compared to the other ones. Can we actually contribute in a way that makes sense? And a lot of the patterns we have adopted here have spread to other repos within the company. Now, whether or not you refactor, the AI delivery system is a layered approach. You can isolate different parts of your code base to avoid a full refactor, but there are so many components that you need to keep in mind. And hopefully throughout this conference, you've heard more details about this. But I really encourage everybody to think about the business value of delivering a big refactor and the trade-offs of doing it now versus in the future. So models will continue to get better, but sometimes it's good to pause, build a monorepo, and forge ahead. So thank you, everybody. Happy to take any questions. Thank you. Yeah, so the question was, before, we had multiple repos, and did we move into a monorepo? Yes, we did that. One of the things we found now is that models are much better at navigating multiple repos. So if you put it into a higher-level folder, right, they could navigate the file directory. But for doing that end-to-end testing and verification and deployment, it's still much harder to do with multiple repos. And if you're building a sandbox environment to run a full AI factory, it also takes more time to clone repos and get everything set up. Yeah? You mentioned a bunch of questions. Yeah, so the question was, when we defined certain features and requirements, did we go back and check them and make changes, as well as the guardrails framework? We did. I think we got 15 out of 17 requirements right when we were going ahead with the refactor. And some of the processes that we added for the actual AI engineering verification evolved over time. So for example, when we started, plan mode was just barely coming into cloud code and didn't exist in cursor, is that models are much better at navigating multiple repos. So if you put it into a higher-level folder, right, they could navigate the file directory. But for doing that end-to-end testing and verification and deployment, it's still much harder to do with multiple repos. And if you're building a sandbox environment to run a full AI factory, it also takes more time to clone repos and get everything set up. Yeah? You mentioned a bunch of questions. Yeah, so the question was, when we define certain features and requirements, did we go back and check them and make changes, as well as the guardrails framework? We did. I think we got 15 out of 17 requirements right when we were going ahead with the refactor. And some of the processes that we added for the actual AI engineering verification, that evolved over time. So, for example, when we started, plan mode was just barely coming into cloud code and didn't exist in cursor, but we adopted it as part of our development lifecycle. Yeah? So our PR reviews were all human PR reviews during that refactor. We did some local checks where we ran skills to say, hey, review this code, make sure that it's good. And they're continuing to get more autonomous as time goes on. But at that point, PRs were a really good way for us to build context for that repo, as we only had a few developers working on it. And we wanted to make sure people understood what had gone into the refactor. Yep? What's the number one factor you think things about labs or what do you do? In terms of factors, I think that the complexity of the task you can give to a model is going to be different. And many more companies will have more scaffolding in terms of actually doing a refactor. So, for example, when I showed the lifecycle of doing the research, the POC work, validating the code quality, checking hidden assumptions, you thought an open source library had this feature, but it was actually in a beta, for example, I think that is going to be much, much faster on top of the standard refactoring of, hey, here's a file, rewrite it to match this set of requirements. All right. Great. Thank you, everybody. Have a great rest of the conference. Thank you. which was a little bit fishy, so I dug deeper. And it actually just implemented a bunch of scaffolding and didn't implement the models. So you can see here, I did not add a raise serve deployment or bootstrap command yet, right? So we're still not there where models can self-validate and just one-shot these kinds of problems, but we're getting close. I think in six months, we'll get to the point that we can complete pretty substantial refactors, as we saw in the Stripe example, consistently across the board. So we get to the core question. Was this refactor worthwhile? Should we have waited a year to do this refactor as models and harnesses continue to get better, or did it make sense to do it at the time? Now, I'll say the other side of the argument, right? Things are getting substantially better. Models are getting better. They can call tools better. We have a lot more infrastructure, like sandboxes and monitoring frameworks, in order for us to actually understand what's happening under the hood with these models. So taking on technical debt and refactoring later is getting exponentially easier as the days go by. Now, the problem is that a lot of times when you build a lot of code and you do this kind of development in an AI-native world, it starts looking like some of the legacy code we've seen in the past. There's a lot of code written. It's written with low performance or quality, and the broader problem is people don't actually understand what's happening there. So if you have some issues within the code base or you want to adjust based on customer requirements, it's actually much harder to do so. So you do have to make sure that there are appropriate guardrails, whether or not you do a full refactor or only a partial one. So if you ask me, was it worthwhile, I'd say yes. We had built out the patterns that we had earlier with the number of different repos in order to match customer requirements and demands. It took an amount of time, but we ultimately achieved the goals of the business. And then we came back and refactored, and we were able to accelerate. We were able to actually reduce the amount of time the pipeline took. We were able to reduce the costs. We could support larger files. And now we can ship features that would take multiple months in under a week. So the monorepo refactor, the cleanup was worthwhile, and we have some of the productivity metrics we saw there. The other part is that beyond just shipping velocity, developers actually want to work in this code base. So everybody comes along and says, hey, can I work in this code base? It's much cleaner compared to the other ones. Can we actually contribute in a way that makes sense? And a lot of the patterns we have adopted here have spread to other repos within the company. Now, whether or not you refactor, the AI delivery system is a layered approach. You can isolate different parts of your code base to avoid a full refactor, but there's so many components that you need to keep in mind. And hopefully throughout this conference, you've heard more details about this. But I really encourage everybody to think about the business value of delivering a big refactor and the trade-offs of doing it now versus in the future. So models will continue to get better, but sometimes it's good to pause, build a monorepo, and forge ahead. So thank you, everybody. Happy to take any questions. Thank you. Yeah, so the question was before we had multiple repos and did we move into monorepo? Yes, we did that. One of the things we found now is that models are much better at navigating multiple repos. So if you put it into a higher-level folder, right, they could navigate the file directory. But for doing that end-to-end testing and verification and deployment, it's still much harder to do with multiple repos. And if you're building a sandbox environment to run sort of a full AI factory, it also takes more time to clone repos and get everything set up. Yeah? You mentioned a bunch of questions. Yeah, so the question was when we define certain features and requirements, did we go back and check them and make changes, as well as sort of the guardrails framework? We did. I think we got 15 out of 17 requirements right when we were going ahead with the refactor. And some of the processes that we added for the actual AI engineering verification, that evolved over time. So for example, when we started, plan mode was just barely coming into cloud code and didn't exist in cursor, but we adopted it as part of our development lifecycle. Yeah? So our PR reviews were all human PR reviews during that refactor. We did some local checks where we ran skills to say, hey, review this code, make sure that it's good. And they're continuing to get more autonomous as time goes on. But at that point, PRs were a really good way for us to build context for that repo as we only had a few developers working on it. And we wanted to make sure people understood what had gone into the refactor. Yep? What's the number one factor you think things about labs or what do you do? In terms of factors, I think that the complexity of the task you can give to a model is going to be different. And many more companies will have more scaffolding in terms of actually doing a refactor. So for example, when I showed the lifecycle of doing the research, the POC work, validating the code quality, checking hidden assumptions, like you thought an open source library had this feature, but it was actually in a beta, for example, I think that is going to be much, much faster on top of sort of the standard refactoring of, hey, here's a file, rewrite it to match this set of requirements. All right. Great. Thank you, everybody. Have a great rest of the conference. Thank you.