SPEAKER_02
I am Nupur, I work with Codo. At Codo we do agentic reviews. I have a background in DevSecOps. So I'm coming from an industry where everything was deterministic. The pipelines, they run, they crash, if they crash, we fix them to a place where we are doing agents where nothing is deterministic. So in my last few years I have learned where and how agents fail, what are the learnings and today I will be sharing some of my learnings with you.
SPEAKER_02
So, if you see the evolution of agents, it started with static prompts where it was a 4K context window and we tried to put whatever was important or whatever we deemed important and the AI models will process it and provide you with the results, right. When we started with that, that means that it was on us to tell LLMs what they should look into. That means if we provide wrong inputs, we might not get proper results. And then we thought maybe if the context window grows, if the context size grows, we can do better, we can have more inputs and we started with agentic workflows. So we created an agent, we get them tools like search tool to go into search into documents and do something as a command, then again look into the search and do something, which again created a loop where the tool does not know where to stop. It thinks I need more inputs, again going back and back, it is a loop. To improvise on that, nowadays multi-agents is becoming more popular. Create multi-agents, do a lot of stuff together. When we see it like that, we have a lot of agents working for you. So a security agent trying to figure security concerns, a review agent trying to review the tool, a coding agent trying to fix things. Now again, the more the tools, the more issues you have. Not every agent understands and they have clashed in their understandings where you do not get into the results. So, what do we learn from here? What we see is context is not a problem. Day by day the models are coming where you can dump a lot of context, a lot of data, but does that make sure that the results you are getting is smart enough to give you everything or smart enough to decide what is important. If you see the current LLM models, we see a pattern where it takes the initial inputs you provide, it takes the last inputs, but the in between context is removed. So, they do not focus on the in between context, agents look at the starting point, end point and try to provide you the results. This is like a U curve where some of the things from the start, some of the things from the end make sense, but whatever you are providing in between that, that is not taken up. Yeah.
SPEAKER_02
How do you know this?
SPEAKER_02
This is something which we are working on and we are actually benchmarking things. So, when we create agents, we try to see this is the context we provide to the agents. Does it take this context into effect and also give us the results? So, we are working with multi-agentic architecture where for each of the tasks we do for code reviews, we give the task to an agent and say, okay, give us the result. Now, every time we, for example, code reviews, we try to see can we give all the context? Can we give the whole code base for example and see if we can get the results? But we see that whenever we start working with that, the initial prompt or the initial goal which we start with that is in focus. If we give something at the end as an input that is in focus, but all between contexts like I have JIRA, I have MCPs, can you look into that? The LLMs try to get rid of those things and push them to make sense by themselves. So, to have this or to make a way out of this, how we deal with is creating strategic solution for context optimization. Rather than dumping everything down to the models and asking them to be smart enough to find out what is more important, we usually start to see, okay, what we can do to make better context for the model. There are lots of solutions in place if you see currently and context engine is a buzzword. Everybody wants to create context engine and everybody wants to provide that. But context engine is like a bouncer, right? So, your high speed car is going and it acts as a bouncer and tells you this is more important. Now, if you have a large messy code base, it makes sense to create a context engine because it creates a search pattern, it creates a ranking logic so that whenever you ask for a task, it looks for those rankings and say this is more important for you, take it and work with it. The problem is the indexing part takes moderate effort, but the scaling is a challenge. Like if you start talking about 600 repositories or 700 repositories, the mapping and the indexing starts to slow down and it becomes again unpredictable to find or create a context engine if you are not actually into making context engine only.
SPEAKER_02
There are lots of areas where agent can get more context instead of investing highly on context engine. Hierarchical summarization where instead of creating or going through everything, a summary is created for each file and folder so that when the agents try to find, they can try to read the summary and see if that is more important to us or not, can be a good one. The only thing is that you need a lot of LLM processing. So, every time a file is created or changed, some of the agents need to go and create a mapping for that. So, it is a high upfront context on LLM processing that is needed.
SPEAKER_02
Another way is knowledge graph. Now, knowledge graph is complex, but it works wonder when you have logical dependencies. For example, you have one file which impacts another file, which again impacts other files. You can create a graph DB hosting. It is the initial input needed by the developer is very high. It takes a lot of time to create that, but if you have complex logics or you have more dependencies on multiple reports, that works wonder. For me, I think for most of the task, if you are not a product company, but if you are building agents for yourself or your processes, iterative retrieval works really good, because instead of even creating a summary, it creates an index. So, it is like a library card which you give to your agents and see this is the topic. If that is relatable to you, you can look deep into the code and look for the results. Again, it has quite cost impacts, but you do not have to invest a lot of energy. The input required by the developers to provide to the LLM is low, and it provides better results. There is also option of self correction where you ask the LLMs to do something and there is a critic node which looks and say if that is relevant to your initial goal or not. In that case, if the context is lost, you can again ask the agent to do it again, retry it again because the critic node said, this is not the right way. It takes a little bit more time because it adds the latency of running the agents and again, but it does not require a lot of input initially from the developers to create something.
SPEAKER_02
Another challenge which I have seen people getting into when they create invest a lot of energy. The input required by the developers to provide to the LLM is low, and it provides better results. There is also option of self correction where you ask the LLMs to do something and there is a critic node which looks and says if that is relevant to your initial goal or not. In that case, if the context is lost, you can again ask the agent to do it again, retry it again because the critic node said, this is not the right way. It takes a little bit more time because it adds the latency of running the
SPEAKER_02
agents again, but it does not require a lot of input initially from the developers to create something. Another challenge which I have seen people getting into when they create these agents is the orchestration paradox. Now, what it does is that LLMs are becoming more and more smart. So, when you give them the task, they think, okay, I should use this tool. Maybe I can do better. I should research more on what should I use. It goes into a loop where instead of actually looking to solve the problem, they look for the method to solve the problem. They hop from one method to another method and most of the API tokens are wasted on finding a way
SPEAKER_02
to do it rather than doing it. So, they just go into research mode. For example, if you use Opus, the latest and greatest, they will try to see what is the best method to do it and challenging themselves again and again. Maybe not this, another way, another way and it just goes into a loop of trying to do something rather than doing something. To resolve this, we worked with an 80-20 hybrid approach. I think this is one of the most interesting outcomes I have seen or the way to resolve this in a finite loop. What our teams are doing is giving the latest and greatest models or
SPEAKER_02
giving the agents power to research 80% of the time. So, you give them the goal and say, okay, try to do whatever you can. But the 20% of the task where you need final validation, you want summarization, those are not something which is free-flowing. Those are more restricted. Those are more hard gates. For example, if I get x results, I want y. It is more deterministic so that the research which is coming from the 80% can be lowered down. Now, when you see, you can always say that the 80% tool can still go on and go into infinite loops. We have mechanisms to work on that. For some organizations,
SPEAKER_02
we do counter mechanisms where after four or five counters, you have to work with whatever was the last result. For some of them, they have timeout counters and after five minutes, whatever is the last tool or whatever is the last decision, you work with that and then go back if the results are not good. But you can restrict that 80%. But in short, if you are using anything like discovery or you are trying to see which tool to use, you are trying to plan, those 80% research models are really good. But if you are trying to create a summarization, you are trying to see, this is the research I have
SPEAKER_02
got. Now, I have to make a result out of it. The 20% works really well. Now, for 80% usually, you use high reasoning models, the latest and greatest, but you do not need a high reasoning model for the 20% because those 20% things are doing deterministic tasks. They are telling you what is needed. For example, the critic node, which we talked about, they do not need to research, they do not need to find out what is the best thing to do. They just need to see what was your goal, what was the result you are trying to achieve and how to provide or how to summarize that for that. Also, things like if you think about
SPEAKER_02
what would be the next possible action, I have this result, what should I go and look for, those are things that can be done by the 80% dynamic models. Whereas, I have all the results from the 80% models, but what is the proper way or what are the proper results that the user is looking for, those decisions can be taken by the 20% model. Finally, this is an interesting failure which we have seen, where as the context grows, teams think we can do everything in one agent because the context window is quite large. We can
SPEAKER_00
[SPEAKER_02] put everything, we can ask an agent to do the testing part, we can ask the agent to do the review part,
SPEAKER_02
we can ask the agent all kinds of things because the context is the same and they can provide us the results. But make sure that when the agent is going forward, it gets overwhelmed with the inputs. And again, it tries to start losing what was the original task. So, maybe you give four tasks to the agent and somewhere down the line, it focuses on two tasks. So, you get great results for the two tasks, but the other two just get lost in the middle. For that particular purpose, we have something called mixture of agents and that's where you hear a lot of buzz about multiple agents or multi-agentic architecture where instead of one big agent,
SPEAKER_02
we create issue-specific expert agents. We create small agents which are doing great in a specific task which they have been provided. Now, to build on top of that, each of the agents come up with their own interesting ideas or results. How to make sure those results combine and make sense together? Because for example, I am trying to search for a vacation, I give an agent to find the best hotel, another agent to find the best location, another agent trying to find the best flights, but all three of them give me different results. The hotel is in Greece, the flight is from Amsterdam to
SPEAKER_02
maybe Portugal and everything just doesn't make sense, right? So, for that particular purpose, there is a concept called a judge agent. What it does is try to get all the results and see if they can make sense together. So, now you are doing all the greatest things from different agents, getting the best results from their part, but a judge agent helps us to combine these and make one sense out of it instead of getting so many things which don't make sense together. Something similar is implemented by us. So, this is an R architecture, Codos architecture, where for code reviews, we are using the same formula. So, as part of a PR review,
SPEAKER_02
we have a context collector, which actually goes and collects context from the PRs. It could collect context from the context engine, it collects context from the tools, but then it does not start working and giving you the reviews. It actually bifurcates all the context it has provided and passes it on to different agents. Now, what these agents do is basically specialize in what they are supposed to. For example, there will be a security agent trying to find security flaws. There might be an agent trying to find code differences, there might be an agent trying
SPEAKER_02
to find the JIRA issues. Once all these agents give us back results, a judge agent actually looks at the results and says, okay, these are interesting enough, but is it relevant to you? We can again go back to the context engine, look into the PRs and see out of the 10 things which were provided by you, how many of them start working and giving you the reviews. It actually bifurcates all the context it has provided and pass it on to different agents. Now, what these agents do is specialize in what they are supposed to. For example, there will be a security agent trying to find security flaws.
SPEAKER_02
There might be an agent trying to code, there might be code differences, there might be an agent trying to find the JIRA issues. Once all these agents give us back, a judge agent actually looks at the results and say, okay, these are interesting enough, but is it relevant to you? We can again go back in the context engine, look into the PRs and see out of the 10 things which are provided by you, how many of them actually make sense for your thing. So again, refining the results to make sense to you. Yeah, I think that was it from my side. Any questions? Yes. In practice, how do you let the swarm communicate with each other? You are talking about the agents?
SPEAKER_02
So you have agent A and agent B and the judge. I can imagine they write to a file system or do you have some kind of proprietary tool? We use a long chain at the bottom and that is being used to communicate and build infrastructure for different agents. Do you know what long chain uses for that? Is it just collects the responses and then shoves it back into the prompt of the next agent? Yes. That's what we do. So what we do is we try to get the results and create a prompt for the next agent. And if there are multiple things, again, there is an agent just to collect the results and create a better prompt, which is refined for the next agent.
SPEAKER_02
Have you thought about a calibration step for each agent? When you say calibration, can you tell me more? Calibration, right? So when you do a code review with an agent, right, you need at least what I heard today might have done some calibration, right? That you actually tell me what is good and what is bad.
SPEAKER_02
Yes. So when you say it that way, let me know if that makes sense for your question or not, we do calibration in a form where we check what we have as a context. So for example, when we get the code reviews, LLM does not know what is important for you or how you work. So for example, an LLM when it gets input, it gets input from healthcare industry, it gets from retail industry, it gets from finance industry, and all of them can use the same Java framework in different ways. Different things are important for them, the rest of them doesn't make sense. So what we do is we give you two different options to tell the agents how to perform or what to work on.
SPEAKER_02
One part is we give them the PR history. So we index all your PRs and see when was the last time something like this was identified and compare it with the current subversion. If that is... And you do that in the context standpoint? Yes. We do that. So the changes you make to the code, we look into if we can see something similar in the past. That again is transferred to the context twice. First is when we are actually giving context to the sub-agents to find things for you, and another time to the judge agent. So that when I get 15 different recommendations for a code review,
SPEAKER_02
my judge agent can look into what was there before, how your reviewer commented, how your developers commented, and based on that decide if that is worth providing to your developers or not. And the other part is... And this happens for every agent? For every agent, yes. And if I understood you right, you don't share the context between the agents, right? So you have a specific context for every agent? Yes, we are trying to resolve that part. Instead of dumping everything to the LLM, we take the part which is more important. We use the context engine for that. We take the part which is more important and only provide that particular part to that particular agent.
SPEAKER_02
Very good. Yes, sir. But then for me, it's not clear how you bridge the gap, right? So you have a code quality agent and you have a framework-specific coding review agent, right? And then you basically, as I understood, you only share the specific information to each agent, right? And then basically each agent runs autonomously, right? Yeah. And it doesn't have a full picture, right? And at least when, as a human, right? And if you do code reviews, it's always good if you have a full perspective, right? I would say that kind of methodology works for simple things, does it use linting? How does it use linting? Are tests implemented? But when you think about
SPEAKER_02
the overall architecture, for example, to make architecture decisions that cover security, because everything is a balance and you have to weigh that somehow. How do you solve that? Yes. So I think if you look at the older version of code reviews, you should have had a senior engineer who knows your code, who knows what kind of packages you're using, and they can comment on whether the developer has done something similar to what you're used to, or did something totally weird, right? Then you used to have a security person who would see if you are providing all kinds of security,
SPEAKER_02
you are not hard coding your APIs, or you are not putting any SQL injections. So all those kinds of things, the security experts know. On the other hand, if you are working with ESO compliance or SOC 2 compliance, there might be an auditor who might ask the team lead or the senior engineer, is your code being logged? Is it logging the changes and so on? So previously as well, there were lots of people having specialized knowledge, looking into those kinds of areas specifically. Now, when the context is provided, it's always like these are my security concerns, which I always have to look into, these
SPEAKER_02
are my architectural concerns, an architect might look from the architecture perspective. We can do something similar with the agents as well, because for example, architecture and security concerns, we have a web portal where architects can provide their guidelines, compliance people can provide their guidelines and an agent can look into all these guidelines and say, is it validated or not? So, if you see the initial picture, the context collector knows everything and then it provides relevant context to the agents. So, you basically enforce your customer to upload this kind of document, right? Is that a kind of requirement then?
SPEAKER_02
Or because, I mean, the system will have completely different results, right? If you don't share them. Exactly. So, that's something which depends on the organization. There are some people who say, we don't work with any rules or regulations, so just give me out of the box. That also means don't expect the agents to find something very specific to your working unless and until we have certain So, if you see the initial picture, the context collector knows everything and then it provides relevant context to the agents. So, you enforce your customer to upload this kind of document, right? Is that a kind of a requirement then?
SPEAKER_02
Or because the system will have completely different result, right? If you don't share them. Exactly. So, that's something which depends upon organization. There are some people who say, we don't work with any rules or regulations, so just give me out of the box. That also means that don't expect the agents to find something very specific to your working unless we have certain PR history because then the PR history kicks in and tells you what is relevant even when you don't provide. I'm not sure if the PR history is really the best source.
SPEAKER_02
It can be one. It can be one. It can. So, that's why there are various sources, right? So, it's PR history. It's your resource. It's your... And somebody's cooking food at my home. Yeah, but it can be one of the sources and that's why... And my question is, yeah, I mean, of course, it can be, right? Yeah. At the end, you need to decide, how much you weight, right? So, the new documents for the engineering principles, you take your principles and compare them to your... Yeah, yeah, yeah. ...study, right? But they can be completely out of the balance, right?
SPEAKER_02
Oh, it depends. It depends. So, I think it's... If you look at it from one perspective, it's difficult to decide. But if you're getting the context from many angles, for example, PR history, that's one part. But when you do the compliance and you tell in the compliance portal, this is really important. So, we have various segments of it's an error, it's a recommendation. All those kinds of things add weight to a feedback to say if it's good or not. And every time your developer accepts a suggestion, it gets more weighted for the next one. If they do not accept the suggestion, it gets less weight. So, it's all about indexing and making sure those weights are banished somewhere.
SPEAKER_02
Yeah. Two ways. One is by when you give your recommendations, does your developer actually accept it or not? We index that. Another way is from the past few years, we try to find out similar issues identified and if your developer actually implemented them. So, for example, some people are used to hard coding their API keys and I literally had a tough argument with the developer, but this is how we do it. No, this should not be the way. Yeah, but...
SPEAKER_02
That's a nice thing. It nicely meant also what I meant, right? So, if you look in the history, if that happens, it doesn't mean it's good. It's good. Yeah, and that's the way where the system tries to tell you this is not good, this is not good and then it's up to you and...
SPEAKER_02
But if you provide the guidance, right? No, so there is something called bug fixes and there is something called rules. So, if you provide them as a rule, it will get highlighted, doesn't matter if you want it or not. And then there are bugs where agent try to tell you there's something wrong and if the reviewer also agrees with it and did not implement 10 times, the reviewer might get it less weighted and give you. Cool. Thank you so much. Thank you. we take the part which is more important. We use the context engine for that. We take the part which is more important and only provide that particular part to that particular agent. Very wonderful. Yes, sir.
SPEAKER_02
But then for me, it's not clear how you bridge the gap, right? So you have a code quality agent and you have a framework specific coding review agent, right? And then you basically, you only, as I understood, you only share the specific information to each agent, right? And then basically each agent runs atomic, autonomously, right? Yeah.
SPEAKER_02
And it doesn't have a full picture, right? And at least when, let's say as a human, right? And if you do code reviews, it's always good if you have a full prospect, right? I would say like that kind of methodology, it works for simple things like, does it use linting? Who does it use? I don't know. Are tests implemented? But when you think about, at least I think when you think about the overall architecture, for example, to make architecture decisions that covers security, because everything is a balance and you have to weight that somehow. How do you solve that?
SPEAKER_02
Yes. So I think if you look into the older version of code reviews, you should have, you used to have a senior engineer who knows your code, who knows what kind of packages you're using, and they can comment on if the developer has done something similar to what you're used to, or did something totally weird, right? Then you used to have a security person who used to see if you are providing all kinds of security, you are not hard coding your APIs, or you are not putting any SQL injections. So all those kind of things, the security experts know. On the other hand, if you are working with ESO compliance or SOC 2 compliance,
SPEAKER_02
there might be an auditor who might ask the team lead or the senior engineer, is your code being, you know, logged? Is it logging the changes and so on? So previously as well, there were lots of people having specialized knowledge, looking into those kind of areas specifically. Now, when the context is provided, it's always like these are my security concerns, which I always have to look into, these are my architectural concerns, an architect might look from the architecture perspective. We can do that something similar with the agents as well, because for example, architecture security
SPEAKER_02
concerns, we have a web portal where architects can provide their guidelines, compliance people can provide their guidelines and an agent can look into all these guidelines and say, is it validated or not? So, if you see the initial picture, the context collector knows everything and then it provides relevant context to the agents. So, you enforce basically your customer to upload this kind of document, right? Is that a kind of a requirement then? Or because, I mean, the system will have completely, let's say, different result, right? If you don't share them. Exactly. So, that's something which depends upon organization. There are some people who say,
SPEAKER_02
we don't work with any rules or regulations, so just give me out of the box. That also means that don't expect the agents to find something very specific to your working until and unless we have certain PR history because then again, the PR history kicks in and tells you what is relevant even when you don't provide. I'm not sure if the PR history is really the best source. It can be one. It can be one. It can. So, that's why there are various sources, right? So, it's PR history. It's your resource. It's your... And somebody's cooking food at my home. Yeah, but it can be one of the sources and that's why...
SPEAKER_02
And my question is like, yeah, I mean, of course, it can be, right? Yeah. At the end, you need to decide, let's say, also in the jury, right, how much you weight, right? So, like, the new documents for, let's say, the engineering principles, like, you take your principles and compare them to, let's say, your... Yeah, yeah, yeah. ...study, right? But they can be completely out of the balance, right? Oh, it depends. It depends. So, again, I think it's... If you look into from one perspective, it's difficult to decide. But if you're getting the context from many angles, for example, PR history,
SPEAKER_02
that's one part. But when you do the compliance and you tell in the compliance portal, this is really important. So, we have various segments of it's an error, it's a recommendation. All those kind of things adds weight to a feedback to say if it's good or not. And every time your developer expects a suggestion, it gets more weighted for the next one. If it does not accept the suggestion, it gets a less weight. So, it's all about indexing and making sure those weights are banished somewhere. Yeah. Two ways. One is by when you give your recommendations, does your developer actually
SPEAKER_02
accept it or not? We index that. Another way is from the past few years, we try to find out similar issues identified and if your developer actually implemented them. So, for example, some people are used to hard coding their API keys and I literally had a tough argument with the developer, but this is how we do it. No, this should not be the way. Yeah, but... That's a nice thing. It nicely meant also what I meant, right? So, if you look in the history, if that happens, it doesn't mean it's good. It's good. Yeah, and that's the way where the system tries to tell you this is not good, this is not good and then it's up to you and...
SPEAKER_02
But if you provide the guidance, right? No, so there is something called bug fixes and there is something called rules. So, if you provide them as a rule, it will get highlighted, doesn't matter if you want it or not. And then there are bugs where agent try to tell you there's something wrong and if the reviewer also agrees with it and did not implement 10 times, the reviewer might get it less weighted and give you. Cool. Thank you so much. Thank you.