Can LLMs generate Enterprise Quality Code? — Prasenjit Sarkar, Sonar
Description
Sonar ran 4,444 Java programming assignments through 53 models and measured what actually came out. GPT-4o generated under 250,000 lines for those assignments. GPT 5.4 generated 1.2 million. Claude Sonnet 4.6 generated 627,000 with the highest security issue rate at 300 per million lines of code. Prasenjit Sarkar from Sonar walks through the full leaderboard: pass rate, cyclomatic complexity, bug density, and security issues per model. Their response is a three-stage framework called ACDC: guide, verify, solve. The verify stage runs SonarQube analysis in 1 to 5 seconds before a commit, against 1 to 5 minutes in CI. If issues slip through to the PR, a remediation agent creates one fix per issue, runs it through analysis and compilation to check for regressions, and only presents it if it passes. Speaker info: - https://www.linkedin.com/in/jit2600/
Summary
Generated by claude-sonnet-4-530-second take
Sonar's Prasenjit Sarkar argues that while LLMs achieve 80%+ pass rates on functional tests (SWE-bench), they produce enterprise-unready code with severe maintainability, security, and verbosity issues. Sonar evaluated 53+ models across 4,444 Java assignments and found that top models like GPT-4o generate 1.2M+ lines for problems solvable in <250k, with 300+ security issues per million LOC (Claude 3.5) and high cognitive complexity. They've built a public leaderboard and an "ACDC" framework (guide/verify/solve) to treat LLM code quality via real-time analysis agents and automated remediation that won't create regressions. This is a vendor pitch with concrete benchmarking data, useful for understanding hidden costs of AI-generated code at scale.
Key takes
- Pass rate ≠ production quality: Models hitting 84% on SWE-bench (Gemini 3.1 Pro High) still produce 614 bugs and 210 security issues per million LOC. Functional correctness metrics (HumanEval, MBPP) ignore maintainability, security, architectural discipline, and tech debt—the actual enterprise concerns.
- Newer models are more verbose, not better: GPT-4o/4o-mini generate 1.2M LOC for the same 4,444 assignments that earlier models solved in 250k. Claude 3.5 Sonnet emits 627k LOC with 300 security issues/M LOC. Trend is northward on bloat and complexity as models improve functional accuracy.
- Root causes are data and probabilistic limits: Training sets contain mixed-quality open-source code with embedded security flaws. LLMs are non-deterministic (same prompt → different code each run), lack company-specific context, and aren't explainable, making diagnosis and improvement hard.
- Sonar's ACDC framework treats quality at three stages: Guide (context augmentation + data cleaning via Sonar Sweep, private beta); Verify (SonarCube agentic analysis in 1-5 sec pre-commit via MCP, open beta); Solve (remediation agent auto-fixes issues in PRs or tech debt, re-analyzes to avoid regressions, open beta). Positions Sonar as quality gatekeeper for agentic workflows.
Useful details
- Leaderboard: sonar.com/leaderboard (open-sourced, 53+ models, updated continuously). Evaluates bugs/M LOC, security issues/M LOC, cyclomatic complexity, cognitive complexity (Sonar proprietary metric for human readability).
- Specific model data:
- Gemini 3.1 Pro High: 84.17% pass rate, 307k LOC, 234 cyclomatic complexity, 614 bugs/M LOC, 210 security issues/M LOC
- Claude 3.5 Sonnet: 627k LOC, 300 security issues/M LOC (highest risk)
- GPT-4o/4o-mini: 1.2M LOC (highest bloat)
- Complexity types: Cyclomatic = branch count (ifs, loops); Cognitive = Sonar's readability/maintainability metric for humans
- Pragmatic Engineer Survey (March 2026): 55% of developers regularly use AI agents (likely a typo; should be 2025 or 2024 given current date)
- Sonar products: SonarCube Enterprise for analysis; Sonar Sweep (data treatment, private beta); Context Augmentation (pushes codebase to LLM); Agentic Analysis (runtime pre-commit check via MCP); Remediation Agent (auto-fix with regression prevention). Supports 40+ languages/frameworks, all major DevOps/IDEs.
Caveats / counterpoints
- Vendor pitch with proprietary metrics: Cognitive complexity is Sonar's own measure; not industry-standard or independently validated. Leaderboard data could favor Sonar's tooling.
- Dataset scope limited to Java: Evaluation used only Java assignments; findings may not generalize to Python, TypeScript, Rust, etc., where LLM training distributions differ.
- No cost or latency analysis: Doesn't discuss developer time saved by bloated code vs. time spent reviewing/fixing. Also silent on Sonar tooling costs or performance overhead of real-time analysis.
- Trend claims lack longitudinal data: "Newer models going northward" on bloat/complexity isn't quantified with version-to-version deltas or statistical significance.
- Regression-free claim unverified: "We don't give code that creates regressions" is asserted but not independently tested or explained in detail (e.g., test coverage requirements, failure modes).
- Survey date likely incorrect: "March 2026" is either a typo or speculative; no source link provided.
Ken relevance
High relevance for AI ops and agent quality control. If Ken is shipping or evaluating code-gen agents (or helping clients do so), this data quantifies hidden costs: bloat, security debt, and maintainability tax that don't show up in pass-rate benchmarks. The ACDC framework maps directly to agent workflow hygiene—pre-commit analysis and auto-remediation could be integration opportunities or competitive moats for Ken's own agent stacks. Leaderboard is a tactical tool for model selection (e.g., avoid Claude 3.5 for security-critical code; GPT-4o if LOC bloat is a cost driver). Also relevant for investing thesis: SonarCube is positioning as infrastructure for agentic era, potentially a category winner if code quality becomes a bottleneck at scale. Worth tracking their beta products and leaderboard updates as market signals for enterprise AI maturity.
Watch verdict
Skim. The leaderboard data and ACDC framework are useful tactical takeaways, but the talk is a product demo with repetitive sections (transcript duplicates several segments). Read the summary, visit the leaderboard, and consider beta access to agentic analysis if Ken's shipping code-gen workflows. Full video adds little beyond slides.
Transcript
All right, okay, sorry guys for the little hiccup. Okay, so my name is Prasenjit Sarkar and today's session is all about whether our LLMs are generating code which is enterprise ready, right? So let's look at the first slide. In the first slide, we are talking about Adrik Arfati who two months back said that a lot of things have been changed in the software development area. Earlier we used to write code in IDE, now things have been changed, now it's all about agenting. So you are spinning up agents, you are just giving instructions in English, and English is now the new programming language, everybody's talking about that. And then you are letting it go, but humans are actually reviewing the code that is being generated by the agent. So a lot of things have been changed. Earlier we used to start opening up IDEs, fancy IDEs like starting from VS Code or JetBrains, to Cursor now, or Windsurf, or anti-gravity. And now we are moving towards the agentic coding platforms, which is Codex, or Cloud, or Devin, or Gemini CLI. And according to the Pragmatic Engineer Survey, which was done in March 2026, we have seen that 55% of the developers are now using regularly some of the AI agents, right? But the question is, do you trust the code that is being generated by these LLMs, right? Is it maintainable? Is it secure? Is it readable? And those are the questions that we are going to debunk right now in these sessions. So let's look at why, and let's look at by evaluating those models and what they are generating out of the box. So there are two aspects that we are seeing. One aspect is all the LLM leaders, all the LLM companies, they are saying that okay, my pass rate is 80 plus percentage, 84%, 83%, 82%. Those are EVAL coming from human EVAL, MBPP, SWB bench. Those are fine. Those are the functional correctness on the test cases, which is mostly known for. But what we are missing is the security aspects, the real world reliability aspect, the engineering architectural problems, the engineering discipline that you have. So code maintainability and the tech debt that is going to be generated by the LLMs itself. And then the context of our analysis. So these are things which are missing. Now, what Sonar has done, we have created an evaluation framework. That EVAL framework run through 4,444 plus distinct Java programming assignments. It's an open source data set. We took up the assignments and then we run through our models. Now, when you run through the models, we saw a huge amount of data that is coming out of the analysis. And that is something that we have put into the open world. So we did the analysis using the SonarCube Enterprise and we got critical insights to choose the right LLM. Now, let's look at the right LLMs or probably not. Let's look at the LLMs that we have evaluated. So here I'm showing you just about the five LLMs. And here you can see that Gemini 3.1 pro high, the pass rate is coming from the SWB bench. So you can see that 84.17%, but it is verbose. So those 4,444 Java assignments that I talked about, to solve that problem, we have seen that it is creating 307,000 line of code, right? Which is pretty concise. It's not that bad. We have seen the complexity, cyclomatic complexity is 234. It's really, really buggy as well, which is 614 bugs that we found out per million line of code. And obviously we have the security issues per million line of code, which is 210. So you see that although these models are generating the code, although these models are pretty much high in models from the foundation models, but you see that for example, Gemini 3 Pro is creating the highest, sorry, the Claude 3 4.6 is creating the highest risk. It's 300 security issues per million line of code that we have seen, right? It is also high-bloat. So for those issue, for those number of assignments, we see 627,000 line of code, which is being generated by Claude 3 4.6. And you'll be stunned if you look at the GPT 4o and GPT 4o Pro high model, you will see that 1.2 million line of code being generated for those 4,000 plus Java assignments. That's a huge amount of line of code that is creating, right? That's a high-bloat. Now, why is it happening? Well, we have seen the mixed-quality code. So the training sets that you see, the training sets actually have mixed-quality code coming from open source, coming from some other places. And that is actually creating the problem as well. Then the built-in security flaws. So the data sets that you are using to train the model that has inbuilt security flaws, and that we have seen where the model is picking up those insecure code examples, along with the good examples as well. So the data sets that you see in the data, and that is actually causing your models to produce the code, which fails or misbehaves in a different way. And, of course, the LLMs themselves, right? So LLMs are probabilistic, right? So obviously we know that the prompt that you are giving to one model today, tomorrow when you give the same prompt to the same model, it is not going to generate the same code. It is going to create a different amount of code, a different set of code, right? It does have the limited context, which is obviously it doesn't understand the company's data or company's code base or company's architecture. And obviously it is not explainable. So it is very hard to diagnose and improve when it is generating the code. So we created this leaderboard called sonar.com slash leaderboard. Here we have given all the data about all the different models that we have evaluated. So far we have 53 plus models and all the different versions. So you see Gemini 3 Pro High, Gemini 3 Pro. So different combination of the thinking aspect as well. So we evaluated 53 plus models and we open sourced all of the data for the people to see how the models are now behaving in a certain way. And obviously it is not explainable. So it is very hard to diagnose and improve when it is generating the code. So we created this leaderboard called sonat.com slash leaderboard. Here we have given all the data about all the different models that we have evaluated. So far we have 53 plus models and all the different versions. So you see Gemini 3 Pro High, Gemini 3 Pro. So different combinations of the thinking aspect as well. So we evaluated 53 plus models and we open sourced all of the data for people to see how the models are now behaving in a certain way. So as of now, you see the Gemini 3.1 Pro High, which was evaluated February 19th. And that has the highest pass rate, which is 84.17. Not that bad of the issued NCD as well. And the lines of code, cyclomatic complexity and cognitive complexity are also fine. It's not that bad. But this is the leaderboard that we have created where we are evaluating all the different models that are coming up continuously. And then we are evaluating that and uploading the data. So you can see not only that, when you go to each and every model inside, there are a lot more details that we have provided about what exactly they are doing so that you can take a concise decision about whether you are going to take this model or the other models according to your architecture. So you see this key inside, which is Gemini 3.1 Pro High at 84.17 correctness. That is the functional correctness I'm talking about. And that's an accuracy leader. But you have other models which are five models that we have given, which are crossing the 80 plus percentage of accuracy. And these are leaders that we have. So we talked about two different complexities. One is cognitive complexity. One is cyclomatic complexity. So cyclomatic complexity is how many branches do you have? How many ifs and elses? How many for loops? How many ifs and other loops? How many while loops do you have? And the cognitive complexity is a Sonar proprietary one, where we measure how difficult a code is for a human being to read and understand and maintain that code. So these are the two different complexities that we maintain. And if you look at the models and the data, you will see the amount of verbosity that we have seen. So the newer models that we are seeing coming up day by day, we are seeing the lines of code going northward. If you see the GPT 5.2 high, it has created actually a million lines of code for those 4,400 plus Java assignments. And if you see the earlier models like GPT 4.0, that's less than 250,000 lines of code. But the models which are going northward, the number of lines of code being written is too high. You have seen the models which are also going higher that also have higher complexity and higher cyclomatic and cognitive complexity. You also need to see that the number of total bugs per model is also going high. But what we have seen is that the models which are getting matured enough day by day are getting finer bugs, finer security issues rather than the old issues. So they're doing a good job in terms of running the reinforcement learning. And they're securing the problems that they have seen already. But in doing that, they're also creating some finer bugs that is very hard for a human being to detect. We have seen the total vulnerabilities per model is now decreasing. But the kind of vulnerabilities that we have seen is going in a different direction. So we have seen that it's generating code that doesn't meet your engineering standards. But what can we do? So in this slide, we talked about the agent-centric development cycle. We call it ACDC. So Sonar has that's a funny name. So it's called ACDC framework. So in the ACDC framework, we have three stages. We have guide stage, we have verify stage, and we have solve phase. So in this one, we have an inner loop and we have an outer loop. So in the guide phase, we have introduced two different products. One is called Sonar context augmentation and Sonar sweep, which is in private beta. So Sonar sweep is basically treating the data that you actually have and the data that you are actually using to train your model. So if you have the problematic data, that means your model is going to create problematic code. If I treat the data right there itself, the code which is going to be generated is going to be good enough. Context augmentation is going to push the entire code base into the LLM itself. Then we have the verify stage. Verify stage is SonarCube. We have various different ways of utilizing that. We have introduced SonarCube agent-tick analysis, which is in beta right now. Open beta, anybody can participate in that, which is actually taking your code in the runtime. What does it mean is that you're using Claude or Codex or Gemini CLR or whatever, which have an MCP inbuilt. And then you can say, generate this code. Now, before you commit that, before you push that into PR, you just need to analyze my code. So it is going to analyze this code way before your CI runs. The CI runs takes about one to five minutes. And then this analysis is going to take about one to five seconds. Within one to five seconds, the code which is being generated right now, before you even commit, it will analyze and it will tell you that these are the problems that I found. Fair enough. It will be pushed down to the agent and the agent is going to fix that problem right there before you even commit that. And then you commit, then you push that code back to the PR and the PR analysis is going to run. Now, the solve part is where we have introduced the SonarCube remediation agent. So let's say that even after all of that, if there are issues that have slipped through your verify stage and it has gone back to the PR stage. So you committed the code, you push a PR, and then you found out SonarCube found that there are issues and your quality gate fails. If that fails, then the remediation agent is right there. It's in open beta where you can just click and then say that I want to fix all of the issues that are there in the PR. Not only that, let's say that you have tech debt, right? A huge amount of tech debt. Now, the solve part is where we have introduced the SonarCube remediation agent. So let's say that even after all of that, if there are issues that have slipped through your verify stage and it has gone back to the PR stage, right? So you committed the code, you pushed a PR, and then you found out that SonarCube found issues that are there and your quality gate fails. If that fails, then the remediation agent is right there. It's in open beta where you can just click and say that I want to fix all of the issues that are there in the PR. Not only that, let's say that you have tech debt, right? A huge amount of tech debt. So you go to the SonarCube dashboard and you see all those tech debts that you have, you just click and select all of the issues that you want to select and fix and then say assign it to the agent. And we are going to create each PR per issue and we are going to fix that one, giving it back to the developers. Developers are going to review that. If they're happy, they're going to approve it and then merge it. The beautiful thing that we have built for the remediation agent is that the remediation agent is going to create the fix, run it through the analysis again, run it through the compilation part again, and see whether that is creating any issues or not. If there is an issue, it is going to discard it. We are not going to give you the code which is going to create a regression. We don't do that, right? So that's the kind of verify loop that we run through. So yeah, this is the entire product portfolio where we are providing guide and verify and then solve. We have all this, 40 plus programming languages and frameworks. All the DevOps are supported. IDEs are supported. We are partnering with this. We are in the marketplace as well. So yeah, that's how we are solving the issues that you are seeing where the LLMs are generating the code, but we are not trusting that, right? Yeah, so if you want some more info, we are in the expo booth. Come and visit us and maybe we can show you one or two demos as well for the product that we have built. Right? Thank you. Thank you. Developers are going to review that. If they find, if they're happy, they're going to approve it and then merge it. The beautiful thing that we have built for the remediation agent is that the remediation agent is going to create the fix, run it through the analysis again, run it through the compilation part again, and see whether that is creating any issues or not. If there is an issue, it is going to discard it. We are not going to give you the code which is going to create a regression. We don't do that, right? So that's the kind of a verify loop that we run through. So yeah, this is the entire, I would say the product, this is our product portfolio where we are providing from guide and to verify and then solve. We have all this, you know, 40 plus programming language and framework. All the DevOps are, you know, supported IDs. You know, we are partnering with this. We are in the marketplace as well. So yeah, that's, that's how we are solving the issues that you are seeing where the LLMs are generating the code, but we are not trusting that, right? Yeah, so if you want some more info, we are in the expo booth. Come and visit us and maybe we can show you one or two demos as well for the product that we have built. Right? Thank you. Thank you.