Open Reader

Guide, Verify, Solve — Anirban Chatterjee, Sonar

completed 22:31 Aug 09, 2026 Watch on YouTube

Current Status

completed

Video ID

03l29gJXpCE

RAG / Chat

Enabled
Guide, Verify, Solve — Anirban Chatterjee, Sonar
Description

A Carnegie Mellon study sorted GitHub projects by whether an AI tool wrote the code, and found the productivity gain ran out after about three months while the static analysis warnings and the added complexity stayed. That residue is verification debt, and how much it costs scales with criticality: a short lived internal tool can live with the gap between the quality a model gives you and the quality the application needs, a large codebase with adversarial users cannot. The obvious backstop is human review, and a Wharton study suggests it leaks badly. Participants took the AI's advice 92.7% of the time when it was correct, and still followed it nearly 80% of the time when it had been instructed to lie confidently. Anirban Chatterjee's argument is that the check has to be zero trust and multi layered. Zero trust means assuming the code could have come from anywhere and verifying it by a different method than the one that wrote it, since a model grading its own output inherits its own blind spots. Multi layered means computational review running alongside reasoning based review, because no single technique catches syntax, data flow, architecture and control flow at once. Sonar's leaderboard makes the blind spots concrete: across their metrics one Claude model rates well on correctness and reliability while the other is the better choice when maintainability, security or lower complexity is what matters. The loop he proposes wraps generation on both sides, handing the agent architectural constraints and coding standards before it starts, running verification inside the inner loop so issues get fixed before they propagate into later loops, and giving the agent the tools to remediate what comes back instead of queueing it for a person who is already rubber stamping too much. Speaker info: - https://www.linkedin.com/in/anirbanc/ - https://www.sonarsource.com/the-coding-personalities-of-leading-llms/leaderboard/

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI coding delivers an initial productivity spike but creates persistent complexity and quality risks unless organizations embed zero-trust, multilayered verification into both agent loops and CI/CD.
  • Why it matters: For teams scaling agent-written code, verification—not code generation—is presented as the control plane that prevents verification debt, security exposure, rubber-stamped reviews, and unmanageable technical debt.
  • Best use: Use this as an architecture and governance briefing for designing a verified AI coding workflow; separate the reusable operating model from Sonar's product claims.

Executive Summary

Anirban Chatterjee argues that AI-assisted development has crossed from experimentation into engineering, where repeatability, scalability, safety, and trust become prerequisites for broader deployment. His central concern is “verification debt”: as application criticality rises, the gap widens between the baseline quality of AI-generated code and the quality required for production systems with large codebases, many users, adversarial behavior, security needs, and regulatory obligations.

The talk cites a Carnegie Mellon analysis of GitHub projects that used Cursor versus traditional tooling. It reportedly found a temporary productivity increase lasting roughly three months, while static-analysis warnings and code complexity remained elevated afterward. Chatterjee’s interpretation is that the resulting quality burden eventually slows developers down, offsetting the short-term generation advantage.

Human review is not treated as an adequate sole control. A Wharton study described in the talk found participants followed correct AI advice 92.7% of the time, but also followed incorrect, confidently delivered AI advice nearly 80% of the time. In high-volume agentic development, the speaker expects this to translate into reviewer overload and rubber-stamping, making independent automated verification necessary.

The proposed operating model is ACDC—Agent-Centric Development Cycle—with three stages: Guide, Verify, Solve. Teams first encode architecture, approved dependencies, coding standards, observability requirements, and quality thresholds; agents receive only task-relevant context; verification runs during generation to catch and remediate issues before they propagate; then formal PR-level review and quality gates run in CI/CD. Sonar positions SonarQube, the newly acquired GitAIr, SonarVortex, and a remediation agent as implementations of this model, so product efficacy claims should be independently validated.

Key Takeaways

  • Claim: The apparent productivity benefit from AI coding can be temporary because AI-generated code creates persistent complexity and static-analysis burdens that later reduce developer throughput. | Evidence: Chatterjee cites a Carnegie Mellon study that categorized GitHub projects by whether Cursor was used: productivity rose temporarily for about three months, while static-analysis warnings and code complexity stayed elevated beyond that period; SonarQube was used to collect the code-quality data. | Implication: Ken should evaluate AI coding programs on downstream remediation load, complexity, and security—not only task-completion speed or lines of code generated. | Caveat: The transcript provides the speaker's interpretation rather than methodological detail, effect sizes, or evidence that the findings generalize across models, organizations, and codebases.
  • Claim: Verification debt grows with application criticality: the quality tolerated in a prototype is materially below what is required for production software with broad and potentially adversarial use. | Evidence: The speaker contrasts small, short-lived internal experiments with large, frequently changing production codebases serving many users, including users actively trying to break the system. | Implication: Agent autonomy and approval requirements should be tiered by system criticality; a low-risk internal tool should not establish the control standard for customer-facing or regulated systems.
  • Claim: Human code review alone is an unreliable backstop for agent-generated code because people can over-trust confident AI output and face review-volume overload. | Evidence: A Wharton experiment described by Chatterjee found participants accepted AI advice 92.7% of the time when it was correct and nearly 80% of the time when the AI was intentionally wrong; he expects similar rubber-stamping in code review when multiple agents generate large volumes of code. | Implication: Use deterministic and independently auditable controls before human approval, and reserve human attention for exceptions, design decisions, high-risk changes, and failures of automated checks. | Caveat: The cited experiment was not specifically a software code-review study, so its transfer to engineering review is an informed extrapolation rather than direct evidence.
  • Claim: A sound AI-code verification regime should be zero-trust and multilayered rather than relying on the generating model or a single review technique. | Evidence: The speaker defines zero trust as applying the same comprehensive review regardless of whether code came from a human or AI, using a methodology distinct from generation and producing repeatable, explainable audit evidence. Multilayering combines computational analysis, LLM-driven review, and other methods across quality, security, and compliance. | Implication: Treat generation and verification as separate system roles, ideally with diverse methods and failure modes, rather than asking one model to certify its own output.
  • Claim: The proposed ACDC workflow improves agent reliability by putting verification inside the generation loop as well as at the CI/CD gate. | Evidence: ACDC consists of Guide, Verify, and Solve: provide constraints and task-relevant context before coding, run verification while the agent writes, let an agent remediate detected issues, then conduct broad PR-level automated review and block release until quality criteria pass. | Implication: For Ken's agent systems, design an inner-loop tool-feedback cycle that can detect and correct local errors before they contaminate subsequent agent steps, while retaining independent outer-loop release gates. | Caveat: The talk demonstrates the workflow conceptually and through vendor tooling, but does not quantify its impact on defect rates, latency, token consumption, or false-positive handling.
  • Claim: Context management is a first-order control for both agent effectiveness and token efficiency: agents need precise, task-relevant specifications rather than the entire codebase. | Evidence: Before agent execution, the speaker recommends encoding architecture, allowed and prohibited dependencies, coding patterns, syntax standards, and logging/observability practices. He warns that dumping an entire codebase into the context window causes exploration, thrashing, and token burn. | Implication: Build or select a context-retrieval layer that supplies scoped repository knowledge and organizational constraints at task time; governance should be machine-readable, not merely documented in wikis.
  • Claim: Model routing should account for different quality profiles, not only token cost or raw task-solving ability. | Evidence: Sonar's stated 4,000-task LLM leaderboard evaluates correctness, complexity, task success, maintainability, reliability, and security. In the example, Claude Sonnet 4.6 is described as strong on correctness, task solving, and reliability, whereas Claude Opus 4.6 may be preferable when maintainability, security, or lower complexity matter more. | Implication: Define routing policies around the actual risk profile of work—for example, use higher-capability models or stricter verification for security-sensitive and high-maintenance changes—rather than optimizing only for cost. | Caveat: This comparison comes from Sonar's proprietary evaluation framework and the transcript does not provide benchmark design, statistical results, or independent validation.

Detailed Brief

Formalize the specifications before agent execution

  • Claims: The speaker treats upfront specification as a prerequisite to safe agentic coding, not as an optional documentation exercise.; Quality thresholds should be explicitly encoded so release decisions can be applied consistently across teams and tools.
  • Evidence: The recommended specification set includes desired architecture, architectural constraints, acceptable coding patterns, dependency allow/deny lists, syntax standards, and logging, observability, and tracing practices.; The speaker says teams should define acceptable levels of security, quality, and maintainability, using default quality criteria or organization-specific criteria.
  • Caveats: The talk does not address how to resolve conflicts among constraints, how specifications evolve across repositories, or how to test whether encoded policy reflects actual business requirements.
  • Implications: Governance becomes operational when it is represented as executable policy that agents and CI systems can consume.; A shared rulebook reduces drift between projects, teams, and coding assistants.

Sonar product architecture and automation posture

  • Claims: Sonar positions its stack as spanning deterministic analysis, LLM-driven review, in-loop agent tooling, PR automation, and legacy-code remediation.; The recommended path to autonomous PR approval is progressive trust rather than enabling automatic merges by default.
  • Evidence: SonarQube is described as analyzing syntax, data flow, architecture, and control flow across languages.; GitAIr, acquired a few weeks before the talk, is described as providing AI code review and CI workflow automation, including issue detection, fix generation, approval, and optional automatic PR merging.; SonarVortex is described as giving coding agents context and real-time verification during inner-loop development; a remediation agent is positioned to work through backlogs and legacy technical debt.; The speaker says SonarQube has more than 7 million developers and analyzes close to 750 billion lines of code daily.
  • Caveats: These are vendor claims in a conference presentation; the transcript contains no implementation detail, pricing, benchmark comparisons, security model, or independent evidence for the claimed outcomes.; Automatic fixes and merges create a governance risk unless rollback, repository permissions, evaluation thresholds, and audit logging are designed before autonomy is expanded.
  • Implications: The relevant product-evaluation question is whether a platform can expose verifiable feedback to agents at both authoring time and release time, not simply whether it finds issues in pull requests.; Legacy remediation is a distinct agent use case that can be isolated from feature delivery, making it a potentially safer domain for controlled autonomy.

Notable Concepts & Terms

  • Verification debt: The widening gap between AI-generated code quality and the quality required for a production system; humans must spend effort closing it before release.
  • Zero-trust verification: Code is verified consistently regardless of whether it was written by a person or model, using methods independent from the code-generation process and producing auditable results.
  • Multilayered verification: A defense-in-depth approach combining computational/static analysis, LLM-driven review, and other review methods because no single technique catches every software failure mode.
  • ACDC (Agent-Centric Development Cycle): Sonar's Guide–Verify–Solve framework for agentic development: supply constraints, check generated code continuously, remediate findings, and iterate.
  • Inner agentic loop: The code-generation cycle in which an agent receives context, writes code, receives verification findings, and fixes issues before continuing.
  • Outer CI/CD loop: The formal pull-request, automated review, quality-gate, test, build, and deployment process that independently validates a completed change.
  • Quality gate: A release-blocking threshold across criteria such as security, quality, and maintainability; code cannot progress until it passes or is remediated.
  • Model quality routing: Choosing a coding model based on the relevant dimensions of quality—such as security, maintainability, complexity, and reliability—not just price or generation speed.

Operator Notes / Why Ken Should Care

  • Create a risk-tiered policy for agent-written changes: define which repositories, change types, and security classes require inner-loop verification, independent CI gates, and human sign-off.
  • Represent architecture, dependency restrictions, coding standards, and observability requirements as machine-consumable task context before allowing coding agents to modify repositories.
  • Instrument AI-development evaluation beyond velocity: track static-analysis findings per change, remediation time, complexity drift, escaped defects, security findings, reviewer overrides, and token cost.
  • Require separation between the code generator and the verification mechanism; test whether validators can be bypassed, influenced by generated output, or lack adequate audit trails.
  • Pilot autonomous remediation first on bounded technical-debt backlogs with rollback and diff-review controls before permitting agents to approve or merge feature PRs.
  • If evaluating Sonar, validate its claimed in-loop integrations, quality-gate behavior, false-positive burden, policy customization, auditability, and permission model in Ken's actual repositories.

Source/Metadata

  • Title: Guide, Verify, Solve — Anirban Chatterjee, Sonar
  • Transcript words: 4686
  • Duration seconds: 1351
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.

Transcript

4532 words en Processed in 133.1s

. All right. Thank you. That's very helpful. My name is Anir Ban Chatterjee. I do product marketing at Sonar. I'm really excited to be talking to this group today. It's actually my first time here at this conference, and so I've been having a blast. I want the rest of my team here meeting a whole bunch of AI engineers, as well as leaders, influencers, and founders. There's a lot going on in this space. I think this year there's really been a turning point from experimentation to engineering, and that warms my heart very deeply because I started my career many, many, many, many, many years ago as a software engineer, writing code for servers, if you can believe it. And I think there's a turning point that's happening right now where we're starting to add the capabilities that we need to add to these systems in order to make them repeatable, make them scalable, make them consistent, much in the way we were doing with cloud computing not too long ago in order to expand the access that IT technology gave to small businesses and other innovators. I think AI is going to do the same thing for software development going forward. But in order to do that, in order to get there, we need to start adding safety and trust to these systems so that they can be used more widely across a wide variety of use cases so that we can build new things and solve bigger problems. And how we get there is what we're going to talk about today. And for those of you who were in Tarek's keynote yesterday, he presented some of this data, and I'm going to talk about it a little bit deeper today. So there was a study that Carnegie Mellon did where they actually looked at projects that were posted on GitHub. And they were able to use the metadata to sort them into projects where there were just traditional tools being used and projects where an AI tool was used to write the code. And in this case, it was Cursor, although it could have been any AI tool. And what they found was interesting. They found that there was, in fact, a temporary spike in productivity, but it lasted about three months and then it went back down. And the reason for that, we think, is because there was also a persistent increase in static analysis warnings and code complexity. They were actually using SonarQube to collect the data on this. And they saw that there was a persistent increase in these types of issues that went beyond the three-month mark and persisted well into the future. And so it's these types of issues that end up actually slowing developers down even more. And this is what makes it a challenge to deliver high-quality code using AI tools. The reason for this is that there is a differing need for quality depending on the criticality of the application. If you're experimenting, if you're playing around, if you're just one person building things to see what's possible, it's an internal non-critical application with just a few users. Maybe it's just you. Maybe it's a small team. Maybe it's just a short-lived project that's not going to last very long. The gap between the quality that you're getting from the AI tool and the quality you need from the application is quite small. And so you can live with that gap. But as you move to higher levels of criticality, as you run into situations where you're supporting many, many users, it's a larger code base with many lines of code and many changes happening across that code base all the time. You have many, many users. Some of them could be adversarial users actively trying to break your software. And so, in those cases, the quality level you need is quite a bit higher than the quality level you're getting by default from these AI tools. And that's where this verification debt comes in. That's where you have to bring the humans in, bring your software engineers in, to try to close that gap and make sure that the quality level is brought up to an acceptable level before you ship that code into production. So why is this happening? Why is this gap actually occurring? We know these models are excellent. They're getting better and better all the time. I'm really excited to start playing with Fable now that that's out to see what levels of code we can get out of Fable going forward. But we do know that because of the technology, because of the way that these models are built, they will still make mistakes. They will still have quality issues. They are still somewhat error-prone. And if you let these errors go into production code, you could have a catastrophic effect on your organization. They're also missing context. They only know what you tell them. They don't know the broader things that are happening elsewhere in the code base. They don't know what's happening with your business. They don't know what happened in the meeting you had with somebody else two weeks ago that's going to influence the code you're writing today. They don't have all the context that you have as an engineer. And so they don't always know your objectives the way you do. And that is going to also cause gaps between what you need from the software and the way it's built. So we also know that models are diverse. No two models are the same, and they have diverse quality issues. And we actually wanted to explore this. And so we actually have a leaderboard that you can go to on our website right now. It's called the LLM leaderboard. And what we do is we take all of the major new models that come out and we evaluate them. We give them 4,000 or so coding tasks. And we evaluate them using all of the metrics that Sonar Cube uses to evaluate code. We look at their correctness, complexity, the rate at which they're solving the tasks we assign them, and then our classic things: maintainability, reliability, and security. And we're able to graph all of these models across these different axes and show you where models perform well and where they have room to improve. And what you're looking at on the screen right now is actually Cloud Opus 4.6 and Cloud Sonic 4.6. If you're a Cloud customer, you might be toggling between these two models to control your token burn rates. And you'll see that Cloud Sonic is actually quite good from a correctness standpoint, from solving tasks, and from a reliability standpoint. But if you're requiring higher levels of maintainability or higher levels of security, if you're trying to get a lower complexity out of your code, you might benefit from switching to Opus for tasks like that. And so we run these kinds of analyses across a lot of different models. And you're always able to go to our website to get the latest analyses that we run. I think we're actually doing the latest Cloud and OpenAI models pretty soon. But this kind of data is helpful, because it tells you where models are good and where models are not good. And it also serves to put some sign on the fact that you still need to be vigilant with these models. None of these models are ever going to be perfect. You're always going to have some kind of need for verification in the loop to make sure that the code that you're getting is the code you actually want to ship. Now, classically, that verification can be human verification. It can be you: your own eyes reading the code, your own intellect reviewing the code to make sure that it's successful. But we know, based on experience and now based on research, that human review can also be compromised. This is a study that was done earlier this year by Wharton. And they actually gave quite a lot of human participants tasks to complete. And they gave those human participants the use of an AI tool to complete those tasks. But unbeknownst to those participants, the AI was told to confidently lie to these participants some of the time. And what they found in the data is that while participants did follow the AI advice 92.7% of the time when the AI was correct, they unfortunately also listened to the AI nearly 80% of the time when the AI was wrong. This is almost surely happening in code review as well. Especially when there are higher amounts of code being written, when there are multiple agents writing code simultaneously, when you now have to bring all those pieces together into a single software application. The load is just too great. There's only so many hours in the day. And you still have to ship something. And so there's a lot of rubber-stamping that I'm sure is happening in all your organizations. It's happening everywhere. And so we need to backstop that somehow with an automated verification tool. What can we do about it? As Tarek was talking about yesterday, I think all of us, when we got involved with software, one of the things that we found most attractive about it is that code is quite, once you write code properly, it's going to run the same way every single time. And there's a certain level of comfort in that. There's a certain level of comfort in knowing that if I write this function the right way, it is going to work this way every single time. And there's a clarity that comes to that. And there's a certain confidence you get out of being able to build something that you know is going to work well for every user going forward. But we also know that code is written by humans. It's happening everywhere. And so, we need to backstop that somehow with an automated verification tool. What can we do about it, right? As Tarek was talking about yesterday, I think all of us, when we got involved with software, one of the things that we found most attractive about it is that code is quite, once you write code properly, it's going to run the same way every single time. And there's a certain level of comfort in that, right? There's a certain level of comfort in knowing that if I write this function the right way, it is going to work this way every single time. And there's a clarity that comes to that. And there's a certain confidence you get out of being able to build something that you know is going to work well for every user going forward. But we also know that code is written by humans. Humans have requirements. And those requirements and externalities have impact in how this code functions. And as you add more and more and more code to the application, they can interact in predictable ways. As you now allow users to use those applications, those users can do all kinds of things you didn't expect. And so, software is not provable in the same way that code is provable. Software can break in interesting and novel ways. And as you're using AI to write more and more software to solve bigger and bigger problems, you're going to run into these limitations more and more often. And so, having automatic verification as part of this process is an important part of the solution. It's going to help you control some of the risks that you're introducing by maybe releasing some of the control you have over the code that's actually being written. And so, we believe that verification is going to be a key enabler and a key unblocker for all of the amazing things that we're going to be able to achieve with AI-driven software development going forward. And so, when we say verification, what do we mean? We think there are two core elements to successful automated verification when it comes to AI coding. One is that it needs to be zero trust. What do we mean by that? Zero trust in this context means that the code could really have come from anywhere. It could still be written by a human. It could be written by an AI. As I just showed you a few slides ago, different AIs will write code in different ways. And you're not going to want to use that same AI to validate the code because you're going to want a diversity of tools being used to make sure that you're catching all the different issues that can happen. And so, no matter where the code is coming from, you want to have a similar comprehensive regime to verify that code that works the same no matter how that code was written. Right? It uses a different methodology to review the code than was used to write the code. It's completely audible, completely explainable, so you can prove that verification was run the same way every single time. And it's algorithmic and repeatable and consistent no matter how you run it. It also needs to be multilayered. You need to have multiple ways, multiple techniques being used, multiple approaches being used to review the code that is being generated. Right? Because you're never going to be able to find every single problem that can occur in software by just using one or two methods. You need to use computational review. You also need to use L1-driven view and everything else in between. You heard a little bit about agentic, we've been hearing a lot about agentic loops this week. And Tarek talked yesterday about our framework for agentic loops. We call it ACDC, or agent-centric development cycle. And there are three phases in the ACDC that we like to talk about. The easiest, by easiest I mean the fastest one to implement now, the one that many of you are probably already on a path implementing, is the verification step, which is top center. This is the most important piece that allows you to write code in these agentic loops in a way that is going to be easily shippable. It needs to be multilayered, it needs to be reasoning-based, and it needs to cut across quality issues, security issues, and compliance issues to make sure that you're shipping quality that you can stand behind. Before the verification step, there's a guidance step. And what guide allows you to do is provide guardrails and context and constraints to make sure that the agent has everything it needs up front to write better code the first time. And then finally, after verification, you need to solve the issues that come up. And that's where the solve state comes. That's where you can remediate any issues that are found in the code. Hopefully, you're allowing the agent to have the agency to do so itself by providing access to the tools it needs to find the issues and fix them itself, and then just repeat the loop. And so, these agentic loops with verification at the core is how you can get to shipping quality software using AI agents. And there are different reasons why you would want to do this. And when we talk to customers, and we've talked to a lot of customers about this, the driving functions that are forcing them to adopt verification across all of their AI coding processes are very similar, right? They want to make sure that AI code is verified consistently. They don't want to have different methods of verification applying to different projects or different teams. They want to have a standard rulebook that applies everywhere, no matter what tool is being used. They also want to make sure they're using their AI tools effectively, right? Some of that, a big part of this, is token efficiency or just efficiency in general, but also just making sure that the tools are being used for the things that are being designed to do in ways that we know they're good at doing, right? So we talk a lot about token efficiency, we talk a lot about making sure that the right models are used for the right projects, and so on. Finally, and third, catching issues from a security standpoint as early as possible in the development cycle. Shifting left on security issues has been very important for a number of years now, and now that AI is writing more and more code, catching security issues up front is extremely critical, especially now that we're in a world where CVEs are announced and then immediately exploited by bad actors, often the same day. And so, you need to make sure that your code is as hardened as possible from those types of issues creeping into production. And finally, maintaining compliance. Many of you, I'm sure, work in a regulated industry, and for those types of situations where you need to be able to prove that verification is run constantly and consistently across the board, maintaining an audit trail that allows you to prove that is extremely important. We have a number of solutions that help with that. SonarCube has been around for quite a while. There are probably quite a few of you that are already SonarCube users. It is a zero-trust, multilayered verification platform that works across syntax issues, data flow issues, architectural issues, control flow issues. And it works across basically any language you would be using. We have a lot of deep hooks that I'm going to take you through in a moment that allow agents to have first-party access to the SonarCube verification so they can more effectively write high-quality code. And we also just recently, and by recently, just a few weeks ago, acquired a company called Guitar, based right here in San Mateo. And they do AI code review. And more than that, they actually build and automate the full CI workflow so that you can not only find issues using our LLM approach, but you can also automatically block if those issues cause a quality issue that you would want to push forward. It can write fixes, and it can approve those fixes and merge those PRs completely automatically if you wanted it to. Now, we'd want to earn that trust. It doesn't happen that way by default. Usually, by default, it will just find the issues and show them to you and enter a dialogue with you so you can have those issues fixed. But as you use it more and more and gain confidence, you can turn on more and more features and completely automate the PR review workflow if you like using Guitar. So, the other big news from earlier this week, and Tarek alluded to this during his talk yesterday, is that we also announced a new agentic loop capability with a product called Sonar Vortex. And Sonar Vortex, I'm going to show another flow chart that shows what it does, but it is providing your agents with tools in the inner loop, in the agentic loop, to not only get constraints and guardrails up front to write better code, but also run verification as it's writing code in real time. So, I can find and fix the issues that are being created. And then, we also released the remediation agent. And the remediation agent allows you to tackle backlog issues and take down your tech debt at a scale that you might not have the bandwidth to do now with human developers. You can basically take your tech debt or your older issues, your legacy code, point them at remediation agent, and it can then improve your code base almost in the background while you focus on the innovation work at the front end that you're working on now. And those are both GAAs of this week. And Sonar Vortex, I'm going to show another flow chart that shows what it does, but it is providing your agents with tools in the inner loop, in the agentic loop, to not only get constraints and guardrails up front to write better code, but also run verification as it's writing code in real time. So, I can find and fix the issues that are being created. And then we also released the remediation agent. And the remediation agent allows you to tackle backlog issues and take down your tech debt at a scale that you might not have the bandwidth to do now with human developers. You can take your tech debt or your older issues, your legacy code, point them at remediation agent, and it can then improve your code base almost in the background while you focus on the innovation work at the front end that you're working on now. And those are both GAAs of this week. Now, there was a very detailed chart that was shown during the keynote yesterday that I'm going to break down for you and explain what the different pieces of this chart mean. This is how we see the ACDC applying not only to the inner agentic loops, but also the outer CI-CD loops that we're all working in to develop code. Before you start either of those loops, though, it is really important to have a sense of the specifications of what you actually are going to want to accomplish with the software that's being written. This is your architectural constraints. This is your desired architecture for the software. This is your coding standards and your coding patterns that are acceptable. These are the list of dependencies that you are and are not allowed to use. These are your coding standards and syntax standards that you obey in your organization. This could be your logging practices or your observability and tracing practices. All of that goes into your specs. And then you also need to define what your quality criteria are, right? And we have quality criteria that we ship with that you can use by default, or you can adjust them as it makes sense for your organization. But this is what are the levels of security, of quality, of maintainability that you're willing to accept in your code that you're pushing into production. You need to write that down and encode it. And now you're ready to start using LLM coding agents to write code. And whenever you're initiating a coding task with an agent, one of the first things that we can help with is providing context and constraints so that the agent starts from the ground floor with an understanding of the code base and the guardrails that are relevant to it in that moment. You have to manage the context window of the agent. You can't just throw your entire code base at the agent up front. It's going to spend a lot of time thrashing and exploring and burning tokens while it's doing it. We are able to efficiently provide just the context that it needs based on the work that it's being given so that it can get to work writing productive code very quickly. It then generates the source code, as you can see. And we also provide in-loop verification to the agent. As it's writing, it can call into us and get a list of issues that we are finding in real time in the code that's being written. And the great thing about that is those issues can then be fixed immediately by the agent so they don't propagate into future agentic loops that are going to run in order to fully build out the software project that you're doing. And this is being enabled by SonarVortex as of this week. Once all of the inner loops have run, you reach a point later on when you have to start entering the formal review and shipping process for the code. And this is the CI-CD process, right? And so there's a PR flow that gets initiated that I'm sure we're all familiar with. Guitar can live in that flow. SonarCube also lives in that flow in order to run a broad automated review of all of the code that is in the PR. And an automated verification that actually returns issues across quality, security, and maintainability. There's a superhuman review that is LLM-driven by Guitar. And there's a computational review that is run by SonarCube that actually assigns grades for all three of those things and won't allow the PR to go past into production unless it gets a passing grade across that criteria. So if there are issues that come up, you can actually use a fix agent to fix all the issues that are discovered there. And then once you're actually able to pass that quality gate, that is when you're able to proceed and test and build and deploy the application. The verification needs to run in both the inner agentic loop and also in the outer loop for CI-CD. I'm running low on time, so I'm going to hope this video completes. This is a video demo of the inner loop, the agentic loop. It's a demo of SonarVortex. And what you've just seen happen is we've given it a task, and in starting that task, this is actually Cursor that's running right now. Cursor called into our SonarVortex context tool to get some context up front to give it an understanding of the code that's working so it knows how to write the code. It is now writing the code, and once it's completed the initial write, it's going to call into our verification process to get a list of issues that it finds. And then if issues get provided, it will actually fix those issues immediately in the inner loop. This is all happening automatically through an integration that we have directly with Cursor. We have similar integrations with Cloud Code, with Codex, with Antigravity, with any major AI coding tool that you would have. You get the gist. It flagged an issue. It's going to fix it immediately, and it's going to run the analysis again. And then it will not proceed until it actually is able to get a passing grade from us on the verification pass. I'm going to move past this because I'm out of time now. So, if you take one thing away from this presentation, it's that we believe very strongly, we're very convicted about this, that a governance and verification regime engine is extremely critical to unlock the next level of success that we need to be able to get from AI coding tools so we can solve bigger and bigger problems. And we know, based on our data, that Sonar customers and Sonar users are able to get higher levels of success from AI coding tools. I'm not going to go through all of these now because I'm out of time, but you can stop by our big red booth downstairs and we'll be happy to talk to you about any of these. But the good news is I think many of you probably have access to some of this stuff already. SonarCube is one of the most widely adopted verification tools in existence today. We have over 7 million developers around the world using us, and we analyze close to 750 billion lines of code across our solutions every single day. And if there are a few of you in the room who care about Gartner at all, it's nice to know that we're a Gartner Market Magic Quadrant leader as well. So, final slide, key takeaways, right? What are the things that we want to walk away from this?