SPEAKER_00
All right, okay, sorry guys for the little hiccup. Okay, so my name is Prasenjit Sarkar and today's session is all about whether our LLMs are generating code which is enterprise ready, right? So let's look at the first slide. In the first slide, we are talking about Adrik Arfati who two months back said that a lot of things have been changed in the software development area. Earlier we used to write code in IDE, now things have been changed, now it's all about agenting. So you are spinning up agents, you are just giving instructions in English, and English is now the new programming language, everybody's talking about that.
SPEAKER_00
And then you are letting it go, but humans are actually reviewing the code that is being generated by the agent. So a lot of things have been changed. Earlier we used to start opening up IDEs, fancy IDEs like starting from VS Code or JetBrains, to Cursor now, or Windsurf, or anti-gravity. And now we are moving towards the agentic coding platforms, which is Codex, or Cloud, or Devin, or Gemini CLI. And according to the Pragmatic Engineer Survey, which was done in March 2026, we have seen that 55% of the developers are now using regularly some of the AI agents, right? But the question is, do you trust the code that is being generated by these LLMs, right?
SPEAKER_00
Is it maintainable? Is it secure? Is it readable? And those are the questions that we are going to debunk right now in these sessions. So let's look at why, and let's look at by evaluating those models and what they are generating out of the box. So there are two aspects that we are seeing. One aspect is all the LLM leaders, all the LLM companies, they are saying that okay, my pass rate is 80 plus percentage, 84%, 83%, 82%. Those are EVAL coming from human EVAL, MBPP, SWB bench. Those are fine. Those are the functional correctness on the test cases, which is mostly known for.
SPEAKER_00
But what we are missing is the security aspects, the real world reliability aspect, the engineering architectural problems, the engineering discipline that you have. So code maintainability and the tech debt that is going to be generated by the LLMs itself. And then the context of our analysis. So these are things which are missing. Now, what Sonar has done, we have created an evaluation framework. That EVAL framework run through 4,444 plus distinct Java programming assignments. It's an open source data set. We took up the assignments and then we run through our models. Now, when you run through the models, we saw a huge amount of data that is coming out of the analysis.
SPEAKER_00
And that is something that we have put into the open world. So we did the analysis using the SonarCube Enterprise and we got critical insights to choose the right LLM. Now, let's look at the right LLMs or probably not. Let's look at the LLMs that we have evaluated. So here I'm showing you just about the five LLMs. And here you can see that Gemini 3.1 pro high, the pass rate is coming from the SWB bench. So you can see that 84.17%, but it is verbose. So those 4,444 Java assignments that I talked about, to solve that problem, we have seen that it is creating 307,000 line of code, right? Which is pretty concise. It's not that bad.
SPEAKER_00
We have seen the complexity, cyclomatic complexity is 234. It's really, really buggy as well, which is 614 bugs that we found out per million line of code. And obviously we have the security issues per million line of code, which is 210. So you see that although these models are generating the code, although these models are pretty much high in models from the foundation models, but you see that for example, Gemini 3 Pro is creating the highest, sorry, the Claude 3 4.6 is creating the highest risk. It's 300 security issues per million line of code that we have seen, right? It is also high-bloat.
SPEAKER_00
So for those issue, for those number of assignments, we see 627,000 line of code, which is being generated by Claude 3 4.6. And you'll be stunned if you look at the GPT 4o and GPT 4o Pro high model, you will see that 1.2 million line of code being generated for those 4,000 plus Java assignments. That's a huge amount of line of code that is creating, right? That's a high-bloat. Now, why is it happening? Well, we have seen the mixed-quality code. So the training sets that you see, the training sets actually have mixed-quality code coming from open source, coming from some other places. And that is actually creating the problem as well. Then the built-in security flaws.
SPEAKER_00
So the data sets that you are using to train the model that has inbuilt security flaws, and that we have seen where the model is picking up those insecure code examples, along with the good examples as well. So the data sets that you see in the data, and that is actually causing your models to produce the code, which fails or misbehaves in a different way. And, of course, the LLMs themselves, right? So LLMs are probabilistic, right? So obviously we know that the prompt that you are giving to one model today, tomorrow when you give the same prompt to the same model, it is not going to generate the same code.
SPEAKER_00
It is going to create a different amount of code, a different set of code, right? It does have the limited context, which is obviously it doesn't understand the company's data or company's code base or company's architecture. And obviously it is not explainable. So it is very hard to diagnose and improve when it is generating the code. So we created this leaderboard called sonar.com slash leaderboard. Here we have given all the data about all the different models that we have evaluated. So far we have 53 plus models and all the different versions. So you see Gemini 3 Pro High, Gemini 3 Pro. So different combination of the thinking aspect as well.
SPEAKER_00
So we evaluated 53 plus models and we open sourced all of the data for the people to see how the models are now behaving in a certain way. And obviously it is not explainable. So it is very hard to diagnose and improve when it is generating the code. So we created this leaderboard called sonat.com slash leaderboard. Here we have given all the data about all the different models that we have evaluated. So far we have 53 plus models and all the different versions. So you see Gemini 3 Pro High, Gemini 3 Pro. So different combinations of the thinking aspect as well.
SPEAKER_00
So we evaluated 53 plus models and we open sourced all of the data for people to see how the models are now behaving in a certain way. So as of now, you see the Gemini 3.1 Pro High, which was evaluated February 19th. And that has the highest pass rate, which is 84.17. Not that bad of the issued NCD as well. And the lines of code, cyclomatic complexity and cognitive complexity are also fine. It's not that bad. But this is the leaderboard that we have created where we are evaluating all the different models that are coming up continuously. And then we are evaluating that and uploading the data.
SPEAKER_00
So you can see not only that, when you go to each and every model inside, there are a lot more details that we have provided about what exactly they are doing so that you can take a concise decision about whether you are going to take this model or the other models according to your architecture. So you see this key inside, which is Gemini 3.1 Pro High at 84.17 correctness. That is the functional correctness I'm talking about. And that's an accuracy leader. But you have other models which are five models that we have given, which are crossing the 80 plus percentage of accuracy. And these are leaders that we have. So we talked about two different complexities.
SPEAKER_00
One is cognitive complexity. One is cyclomatic complexity. So cyclomatic complexity is how many branches do you have? How many ifs and elses? How many for loops? How many ifs and other loops? How many while loops do you have? And the cognitive complexity is a Sonar proprietary one, where we measure how difficult a code is for a human being to read and understand and maintain that code. So these are the two different complexities that we maintain. And if you look at the models and the data, you will see the amount of verbosity that we have seen. So the newer models that we are seeing coming up day by day, we are seeing the lines of code going northward.
SPEAKER_00
If you see the GPT 5.2 high, it has created actually a million lines of code for those 4,400 plus Java assignments. And if you see the earlier models like GPT 4.0, that's less than 250,000 lines of code. But the models which are going northward, the number of lines of code being written is too high. You have seen the models which are also going higher that also have higher complexity and higher cyclomatic and cognitive complexity. You also need to see that the number of total bugs per model is also going high. But what we have seen is that the models which are getting matured enough day by day are getting finer bugs, finer security issues rather than the old issues.
SPEAKER_00
So they're doing a good job in terms of running the reinforcement learning. And they're securing the problems that they have seen already. But in doing that, they're also creating some finer bugs that is very hard for a human being to detect. We have seen the total vulnerabilities per model is now decreasing. But the kind of vulnerabilities that we have seen is going in a different direction. So we have seen that it's generating code that doesn't meet your engineering standards. But what can we do? So in this slide, we talked about the agent-centric development cycle. We call it ACDC. So Sonar has that's a funny name. So it's called ACDC framework.
SPEAKER_00
So in the ACDC framework, we have three stages. We have guide stage, we have verify stage, and we have solve phase. So in this one, we have an inner loop and we have an outer loop. So in the guide phase, we have introduced two different products. One is called Sonar context augmentation and Sonar sweep, which is in private beta. So Sonar sweep is basically treating the data that you actually have and the data that you are actually using to train your model. So if you have the problematic data, that means your model is going to create problematic code. If I treat the data right there itself, the code which is going to be generated is going to be good enough.
SPEAKER_00
Context augmentation is going to push the entire code base into the LLM itself. Then we have the verify stage. Verify stage is SonarCube. We have various different ways of utilizing that. We have introduced SonarCube agent-tick analysis, which is in beta right now. Open beta, anybody can participate in that, which is actually taking your code in the runtime. What does it mean is that you're using Claude or Codex or Gemini CLR or whatever, which have an MCP inbuilt. And then you can say, generate this code. Now, before you commit that, before you push that into PR, you just need to analyze my code. So it is going to analyze this code way before your CI runs.
SPEAKER_00
The CI runs takes about one to five minutes. And then this analysis is going to take about one to five seconds. Within one to five seconds, the code which is being generated right now, before you even commit, it will analyze and it will tell you that these are the problems that I found. Fair enough. It will be pushed down to the agent and the agent is going to fix that problem right there before you even commit that. And then you commit, then you push that code back to the PR and the PR analysis is going to run. Now, the solve part is where we have introduced the SonarCube remediation agent.
SPEAKER_00
So let's say that even after all of that, if there are issues that have slipped through your verify stage and it has gone back to the PR stage. So you committed the code, you push a PR, and then you found out SonarCube found that there are issues and your quality gate fails. If that fails, then the remediation agent is right there. It's in open beta where you can just click and then say that I want to fix all of the issues that are there in the PR. Not only that, let's say that you have tech debt, right? A huge amount of tech debt. Now, the solve part is where we have introduced the SonarCube remediation agent.
SPEAKER_00
So let's say that even after all of that, if there are issues that have slipped through your verify stage and it has gone back to the PR stage, right? So you committed the code, you pushed a PR, and then you found out that SonarCube found issues that are there and your quality gate fails. If that fails, then the remediation agent is right there. It's in open beta where you can just click and say that I want to fix all of the issues that are there in the PR. Not only that, let's say that you have tech debt, right? A huge amount of tech debt.
SPEAKER_00
So you go to the SonarCube dashboard and you see all those tech debts that you have, you just click and select all of the issues that you want to select and fix and then say assign it to the agent. And we are going to create each PR per issue and we are going to fix that one, giving it back to the developers. Developers are going to review that. If they're happy, they're going to approve it and then merge it. The beautiful thing that we have built for the remediation agent is that the remediation agent is going to create the fix, run it through the analysis again, run it through the compilation part again,
SPEAKER_00
and see whether that is creating any issues or not. If there is an issue, it is going to discard it. We are not going to give you the code which is going to create a regression. We don't do that, right? So that's the kind of verify loop that we run through. So yeah, this is the entire product portfolio where we are providing guide and verify and then solve. We have all this, 40 plus programming languages and frameworks. All the DevOps are supported. IDEs are supported. We are partnering with this. We are in the marketplace as well. So yeah, that's how we are solving the issues that you are seeing where the LLMs are generating the code, but we are not trusting that, right?
SPEAKER_00
Yeah, so if you want some more info, we are in the expo booth. Come and visit us and maybe we can show you one or two demos as well for the product that we have built. Right? Thank you. Thank you. Developers are going to review that. If they find, if they're happy, they're going to approve it and then merge it. The beautiful thing that we have built for the remediation agent is that the remediation agent is going to create the fix, run it through the analysis again, run it through the compilation part again, and see whether that is creating any issues or not. If there is an issue, it is going to discard it.
SPEAKER_00
We are not going to give you the code which is going to create a regression. We don't do that, right? So that's the kind of a verify loop that we run through. So yeah, this is the entire, I would say the product, this is our product portfolio where we are providing from guide and to verify and then solve. We have all this, you know, 40 plus programming language and framework. All the DevOps are, you know, supported IDs. You know, we are partnering with this. We are in the marketplace as well. So yeah, that's, that's how we are solving the issues that you are seeing where the LLMs are generating the code, but we are not trusting that, right?
SPEAKER_00
Yeah, so if you want some more info, we are in the expo booth. Come and visit us and maybe we can show you one or two demos as well for the product that we have built. Right? Thank you. Thank you.