Why More Context Makes Your Agent Dumber and What to Do About It — Nupur Sharma, Qodo
Description
Give an agent your full codebase and it will attend to the start and the end, then quietly drop the middle. Nupur from Qodo calls this the U curve and builds the whole talk around it: why growing the context window did not fix the problem, and what actually does. She runs through iterative retrieval, hierarchical summarization, and self correction with honest cost tradeoffs for each. The second half covers the orchestration paradox: capable models burn most of their tokens deciding how to solve a problem rather than solving it. Her team's fix is an 80/20 split, using high reasoning models for open ended discovery and lighter deterministic models for validation. Qodo's code review architecture runs this live: a context collector feeds specialized agents, a judge node recombines the results and weighs them against PR history, and every accepted or rejected suggestion shifts the weights for the next run. Speaker info: - https://www.linkedin.com/in/nupursh/
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: More context in agentic systems causes LLMs to lose focus on the middle of large inputs (U-curve attention), requiring strategic context optimization and multi-agent architectures instead of dumping everything into one prompt.
- Why it matters: Practical patterns from production code review agents showing how to prevent agents from getting overwhelmed, looping infinitely, or losing task focus as context grows—directly relevant for anyone building multi-step agent systems.
- Best use: Study the 80-20 hybrid approach, mixture-of-agents pattern with judge nodes, and context optimization strategies; apply to any agent workflow that risks infinite loops or context overload.
Executive Summary
Nupur Sharma from Qodo (agentic code review product) presents hard-won lessons from building production agents in a domain that moved from deterministic DevSecOps pipelines to nondeterministic agentic workflows. The central finding: larger context windows do not make agents smarter—they cause a U-curve attention problem where LLMs focus on the start and end of inputs but ignore the middle. This means dumping entire codebases, JIRA tickets, and documentation into a single prompt leads to lost context and poor results.
The talk walks through the evolution from static prompts → single-agent tool loops → multi-agent architectures, explaining why each step failed at scale. Key failure modes include agents entering infinite research loops (the 'orchestration paradox'), losing track of original goals when given too many tasks, and producing non-coherent results when multiple agents work in isolation. Qodo's solutions include an 80-20 hybrid model (80% high-reasoning exploration, 20% deterministic validation with cheaper models), iterative retrieval over heavy context engines for most use cases, and a mixture-of-agents architecture with a judge agent to unify outputs.
The Qodo architecture for code review is detailed: a context collector gathers PR history, JIRA, compliance docs, and codebase snippets; specialized agents (security, code quality, compliance) each receive only relevant context slices; a judge agent reconciles results against past PR feedback and explicit customer rules (weighted by developer acceptance). This prevents any single agent from being overwhelmed and ensures coherent, prioritized feedback. The system continuously recalibrates weights based on whether developers accept or reject suggestions, making the agent smarter over time without manual retraining.
Practical takeaways include avoiding context engines unless you have 600+ repos and complex dependencies (knowledge graphs work for logical dependencies but require high upfront dev input); using iterative retrieval or hierarchical summarization for most tasks; employing self-correction critic nodes to catch context loss; and reserving high-reasoning models for discovery/planning (80%) while using cheaper models for deterministic summarization/validation (20%). The talk is candid about tradeoffs: some approaches have high LLM processing costs, others require developers to upload guidelines, and calibration relies on PR history quality. No timestamps were provided in the transcript.
Key Takeaways
- Claim: LLMs exhibit a U-curve attention pattern: they focus on the beginning and end of large context windows but ignore the middle, causing agents to lose important context even when the window size supports it. | Evidence: Qodo benchmarked this by providing full codebase context to agents and observing that intermediate inputs (JIRA tickets, MCPs) were ignored while initial goals and final prompts remained in focus. Agents 'try to get rid of those things and push them to make sense by themselves.' | Caveat: This is Qodo's observed pattern in their multi-agent code review system; the speaker does not cite external research or specify which models exhibit this behavior most strongly. | Implication: You cannot simply dump all context and trust the model to prioritize; you must engineer context selection and retrieval strategies or risk agents missing critical information in the middle of long inputs. | Timestamp: timestamp unavailable
- Claim: The orchestration paradox: giving agents too much freedom with high-reasoning models causes them to enter infinite loops researching the best method instead of solving the problem, wasting tokens on meta-decisions. | Evidence: Using Opus (the 'latest and greatest'), agents 'go into research mode' and 'hop from one method to another,' endlessly challenging their approach rather than executing. Nupur describes this as agents 'trying to do something rather than doing something.' | Caveat: No quantitative data on token waste or loop frequency is provided; the claim is based on observed behavior in Qodo's systems. | Implication: High-reasoning models need guardrails (counter limits, timeouts, or deterministic validation steps) to prevent runaway deliberation; reserve them for discovery/planning phases, not execution. | Timestamp: timestamp unavailable
- Claim: An 80-20 hybrid approach resolves the orchestration paradox: use high-reasoning models for 80% exploration/discovery, then switch to cheaper, deterministic models for the final 20% (validation, summarization) to impose hard gates. | Evidence: Qodo gives agents power to research for 80% of the task (e.g., deciding which tool to use, planning), but the final 20% (e.g., critic node validation, summarization) is deterministic and does not require high reasoning. Counter mechanisms (4-5 attempts) or timeout counters (5 minutes) force the agent to work with the last result and move on. | Caveat: The 80-20 split is not empirically justified; it is a heuristic that worked for Qodo's use case. The approach requires explicit task segmentation and may not fit all workflows. | Implication: You can reduce costs and prevent loops by reserving expensive models for open-ended exploration and using fast, cheap models for deterministic steps; design your agent flow to clearly separate these phases. | Timestamp: timestamp unavailable
- Claim: Dumping all tasks into one agent causes it to get overwhelmed and drop some goals, even with large context windows; mixture-of-agents with a judge node prevents this by distributing tasks to specialized experts and reconciling outputs. | Evidence: When Qodo tried giving one agent multiple tasks (testing, review, security), the agent 'focuses on two tasks' and the other two 'get lost in the middle.' Qodo's solution: separate agents for security, code quality, JIRA compliance, etc., plus a judge agent that checks if results 'make sense together' (e.g., hotel in Greece, flight to Portugal, no coherence). | Caveat: The speaker does not quantify how often tasks are dropped or provide before/after accuracy metrics; the claim is illustrated with analogies (vacation planning) rather than code review data. | Implication: For complex workflows with multiple objectives, split responsibilities across specialized agents and use a judge/aggregator to ensure coherent, prioritized outputs instead of overwhelming a single agent. | Timestamp: timestamp unavailable
- Claim: Context engines (indexing, ranking, search) are only worth building if you have 600-700+ repositories; for most use cases, iterative retrieval (library-card-style index) or hierarchical summarization is cheaper and faster. | Evidence: Context engines require 'moderate effort' to index but 'scaling is a challenge' at 600-700 repos; indexing becomes 'unpredictable' if you are not solely focused on building a context engine. Iterative retrieval creates a topic index (like a library card) that agents can use to decide whether to look deeper, with 'low input required by developers' and 'better results.' | Caveat: No cost or latency numbers are provided; the comparison is qualitative. Hierarchical summarization requires 'high upfront LLM processing' every time a file changes, and knowledge graphs require 'very high initial input from developers' but work well for logical dependencies. | Implication: Do not over-invest in context engines unless you operate at large scale; start with iterative retrieval or summarization and only build a custom context engine if you hit scaling limits with hundreds of repositories. | Timestamp: timestamp unavailable
- Claim: Self-correction with a critic node can recover lost context: if the agent drifts from the original goal, a critic node checks relevance and forces a retry, adding latency but requiring low upfront developer input. | Evidence: A critic node 'looks and says if that is relevant to your initial goal or not'; if not, 'you can again ask the agent to do it again, retry it again.' This 'takes a little bit more time because it adds the latency of running the agents again' but does not require developers to create complex upfront structures. | Caveat: The speaker does not specify how often retries occur, how many retries are allowed, or whether this causes compounding latency issues in production. | Implication: Add a lightweight validation layer (critic node) to catch and correct context drift; this is a cheap way to improve reliability without heavy upfront engineering, though it will increase total runtime. | Timestamp: timestamp unavailable
- Claim: Qodo's judge agent reconciles specialized agent outputs by weighting feedback based on past PR acceptance, explicit compliance rules, and developer behavior, continuously recalibrating without retraining. | Evidence: The judge agent checks if results from security, code quality, and JIRA agents 'make sense for your thing' by consulting PR history (similar issues, how reviewers/developers commented) and compliance rules (marked as 'error' vs. 'recommendation'). 'Every time your developer accepts a suggestion, it gets more weighted for the next one. If they do not accept, it gets less weight.' Developers who ignore a bug fix 10 times see that issue get less weight in future reviews. | Caveat: This relies on PR history quality and assumes developers provide feedback (accept/reject); new repos or teams without history will see less effective calibration initially. The weighting mechanism is not detailed (linear, exponential, etc.). | Implication: You can build agents that learn from implicit user feedback (accept/reject actions) without fine-tuning models; design a judge/aggregator that tracks feedback and adjusts weights to align outputs with actual user preferences over time. | Timestamp: timestamp unavailable
Detailed Brief
The U-curve attention problem and why more context makes agents dumber
- Claims: LLMs focus on the start and end of large context windows but ignore the middle; Dumping entire codebases, JIRA tickets, and docs into a single prompt causes agents to lose critical context; This is an observed pattern in Qodo's benchmarking of multi-agent code review systems
- Evidence: Qodo tested giving full codebase context to agents; agents focused on initial goals and final prompts but discarded intermediate inputs like JIRA and MCPs; Agents 'try to get rid of those things and push them to make sense by themselves' when overloaded with middle context; The evolution from 4K static prompts → agentic tool loops → multi-agents shows that context size alone does not solve the problem
- Caveats: No external research or model-specific data is cited; this is Qodo's internal finding; The speaker does not quantify how much middle context is lost or specify which models (GPT-4, Claude, etc.) exhibit this most strongly
- Implications: You must engineer context selection and retrieval strategies instead of relying on large context windows; Assume that anything in the middle of a long prompt may be ignored; structure prompts to front-load or back-load critical info, or split context across agents
The orchestration paradox and the 80-20 hybrid approach
- Claims: High-reasoning models (e.g., Opus) enter infinite loops researching methods instead of solving problems, wasting tokens on meta-decisions; The 80-20 hybrid approach resolves this: 80% high-reasoning exploration, 20% deterministic validation with cheaper models; Counter mechanisms (4-5 retries) or timeout counters (5 minutes) prevent runaway loops
- Evidence: Opus and similar models 'go into research mode' and 'hop from one method to another,' endlessly challenging their approach; The 80% phase uses high-reasoning models for discovery, planning, deciding which tool to use; the 20% phase uses deterministic models for summarization, validation, critic nodes; If the agent does not converge, counters force it to 'work with whatever was the last result' and move on
- Caveats: The 80-20 split is a heuristic, not empirically derived; it worked for Qodo but may not generalize; No quantitative data on token savings or success rate improvements is provided; The approach requires explicit task segmentation, which may not fit all workflows
- Implications: Reserve expensive, high-reasoning models for open-ended exploration; use fast, cheap models for deterministic validation and summarization; Implement counters or timeouts to prevent infinite deliberation; design flows that clearly separate exploration from execution phases; This pattern can reduce costs and improve reliability in multi-step agent systems
Context optimization strategies: when to use context engines vs. iterative retrieval vs. knowledge graphs
- Claims: Context engines (indexing, ranking, search) are only worth building at 600-700+ repos; they slow down and become unpredictable at scale; Iterative retrieval (library-card-style index) works best for most use cases: low dev input, cost impact is manageable, provides better results; Hierarchical summarization requires high upfront LLM processing (summarize each file/folder on change) but reduces search space; Knowledge graphs are complex, require high dev input, but work wonders for logical dependencies across multiple repos
- Evidence: Context engines require 'moderate effort' to index but 'scaling is a challenge' at 600-700 repos; 'mapping and indexing start to slow down'; Iterative retrieval creates a topic index ('like a library card') that agents use to decide whether to look deeper; 'low input required by developers'; Hierarchical summarization needs agents to re-summarize files every time they change, 'high upfront cost on LLM processing'; Knowledge graphs require 'very high initial input from developers' but excel when one file impacts another in complex dependency chains
- Caveats: No cost or latency numbers are provided; comparisons are qualitative; The speaker assumes most teams are not building a dedicated context engine product; Iterative retrieval has 'quite cost impacts' but is not quantified
- Implications: Do not over-invest in context engines unless you operate at large scale (hundreds of repos); start with iterative retrieval or summarization; If you have complex logical dependencies (e.g., multi-repo microservices), invest in a knowledge graph; otherwise, it is overkill; Hierarchical summarization is useful if you can absorb the LLM processing cost on every file change
Mixture-of-agents architecture with a judge node to prevent task overload and incoherent outputs
- Claims: Giving one agent multiple tasks (testing, review, security) causes it to get overwhelmed and drop some goals, even with large context windows; Mixture-of-agents splits tasks across specialized experts (security agent, code quality agent, JIRA agent, etc.); A judge agent reconciles outputs to ensure coherence (e.g., hotel in Greece, flight to Portugal, location mismatch); Qodo's architecture: context collector → specialized agents → judge agent → refined results
- Evidence: When one agent is given four tasks, 'it focuses on two tasks' and the other two 'get lost in the middle'; Specialized agents each receive only relevant context slices (security agent gets compliance docs, code quality agent gets PR history, etc.); The judge agent 'looks at the results and says, okay, these are interesting enough, but is it relevant to you?' and can re-query the context engine or PR history; Qodo's system splits code reviews into security, code differences, JIRA issues, etc., then the judge refines results to 'make sense for your thing'
- Caveats: The speaker does not quantify how often tasks are dropped or provide before/after accuracy metrics; The judge agent adds latency (another LLM call); no timing data is provided; This architecture requires orchestration infrastructure (Qodo uses LangChain to pass results between agents via prompts)
- Implications: For complex workflows, split responsibilities across specialized agents instead of overwhelming a single agent; Use a judge/aggregator to ensure outputs are coherent and prioritized; this prevents nonsensical results when agents work in isolation; Design your context collector to know everything, then slice context to agents; the judge re-consults the full context to validate
Self-correction with critic nodes and continuous calibration via feedback loops
- Claims: A critic node can catch context drift by checking if agent output is relevant to the original goal; if not, force a retry; This adds latency but requires low upfront developer input; Qodo continuously recalibrates by weighting feedback based on developer acceptance (accept = higher weight, reject = lower weight); Explicit rules (compliance, architecture guidelines) always get highlighted, regardless of PR history
- Evidence: A critic node 'looks and says if that is relevant to your initial goal or not'; if irrelevant, 'you can again ask the agent to do it again'; Qodo tracks whether developers accept or reject suggestions: 'Every time your developer accepts a suggestion, it gets more weighted for the next one'; If a reviewer ignores a bug fix 10 times, 'the reviewer might get it less weighted'; Compliance rules are marked 'error' vs. 'recommendation'; errors are always surfaced, recommendations are weighted by past behavior
- Caveats: No data on retry frequency, how many retries are allowed, or whether this causes compounding latency; Calibration relies on PR history quality; new repos or teams without history see less effective results initially; The weighting mechanism is not detailed (linear, exponential, etc.)
- Implications: Add a lightweight validation layer (critic node) to catch context drift without heavy upfront engineering; Build agents that learn from implicit user feedback (accept/reject actions) without fine-tuning models; Design a feedback loop that adjusts weights over time to align outputs with actual user preferences
Notable Concepts & Terms
- U-curve attention pattern: LLMs focus on the beginning and end of large context windows but ignore the middle; a key reason why more context makes agents dumber in practice (Qodo's observed failure mode)
- Orchestration paradox: Giving agents too much freedom with high-reasoning models causes them to enter infinite loops researching methods instead of solving problems, wasting tokens on meta-decisions
- 80-20 hybrid approach: Use high-reasoning models for 80% of the task (exploration, planning) and deterministic, cheaper models for the final 20% (validation, summarization) to prevent runaway loops and reduce costs
- Mixture-of-agents (MoA) with judge node: Distribute tasks to specialized expert agents, then use a judge agent to reconcile outputs and ensure coherence; prevents any single agent from being overwhelmed and dropping goals
- Context collector: Qodo's component that gathers all context (PR history, JIRA, compliance docs, codebase) and then slices it to specialized agents; the judge re-consults it for validation
- Iterative retrieval: Creating a library-card-style index (topic summaries) that agents use to decide whether to look deeper into code; lower dev input and cost than full context engines for most use cases
- Hierarchical summarization: Summarizing each file/folder so agents can read summaries first and decide relevance; requires high upfront LLM processing on every file change
- Knowledge graph: Graph DB hosting logical dependencies (file A impacts file B impacts file C); complex, high dev input, but works wonders for multi-repo dependency chains
- Critic node: A validation step that checks if agent output is relevant to the original goal; forces a retry if context is lost; adds latency but requires low upfront dev input
- Continuous calibration via feedback loops: Qodo adjusts weights based on developer acceptance (accept = higher weight, reject = lower weight) without retraining models; makes agents smarter over time
Operator Notes / Why Ken Should Care
- If you are building multi-step agent workflows, this talk provides production-tested patterns to prevent infinite loops, context overload, and incoherent outputs—directly applicable to any domain, not just code review.
- The 80-20 hybrid approach (high-reasoning for exploration, cheap models for validation) is a cost-reduction and reliability pattern you can apply immediately to reduce runaway token usage.
- The mixture-of-agents architecture with a judge node is a general solution for complex workflows with multiple objectives; it prevents task overload and ensures coherent results.
- Context optimization strategies (iterative retrieval vs. context engines vs. knowledge graphs) help you decide where to invest engineering effort based on scale and complexity.
- The continuous calibration via feedback loops (accept/reject tracking) is a lightweight way to make agents smarter over time without fine-tuning models—useful for any agent that interacts with users.
- The U-curve attention finding is a critical design constraint: you cannot rely on models to prioritize middle context, so you must engineer context selection or split it across agents.
- If you are investing in context engines, Qodo's experience suggests they are only worth building at 600-700+ repos; start with simpler strategies unless you operate at that scale.
- The talk is candid about tradeoffs (LLM processing costs, dev input requirements, latency) and does not oversell any solution—useful for realistic planning.
Watch Map
- timestamp unavailable: Introduction: background in DevSecOps, transition to nondeterministic agentic systems, overview of agent evolution
- timestamp unavailable: The U-curve attention problem: why dumping more context into large windows makes agents lose focus on the middle
- timestamp unavailable: Context optimization strategies: when to use context engines vs. iterative retrieval vs. hierarchical summarization vs. knowledge graphs
- timestamp unavailable: The orchestration paradox: high-reasoning models enter infinite research loops; the 80-20 hybrid approach (exploration vs. validation)
- timestamp unavailable: Mixture-of-agents architecture: specialized expert agents + judge node to prevent task overload and ensure coherent outputs
- timestamp unavailable: Qodo's code review architecture: context collector → specialized agents (security, code quality, JIRA) → judge agent → refined results
- timestamp unavailable: Continuous calibration via feedback loops: weighting suggestions based on developer acceptance, PR history, and explicit compliance rules
- timestamp unavailable: Q&A: how agents communicate (LangChain, prompt passing), calibration details, weighting bug fixes vs. rules, handling PR history quality
Source/Metadata
- Title: Why More Context Makes Your Agent Dumber and What to Do About It — Nupur Sharma, Qodo
- Transcript words: 5552
- Duration seconds: 1587
- Timestamp note: No timestamps or chapter markers were present in the transcript; watch_map entries reflect logical flow, not specific timings
Transcript
I am Nupur, I work with Codo. At Codo we do agentic reviews. I have a background in DevSecOps. So I'm coming from an industry where everything was deterministic. The pipelines, they run, they crash, if they crash, we fix them to a place where we are doing agents where nothing is deterministic. So in my last few years I have learned where and how agents fail, what are the learnings and today I will be sharing some of my learnings with you. So, if you see the evolution of agents, it started with static prompts where it was a 4K context window and we tried to put whatever was important or whatever we deemed important and the AI models will process it and provide you with the results, right. When we started with that, that means that it was on us to tell LLMs what they should look into. That means if we provide wrong inputs, we might not get proper results. And then we thought maybe if the context window grows, if the context size grows, we can do better, we can have more inputs and we started with agentic workflows. So we created an agent, we get them tools like search tool to go into search into documents and do something as a command, then again look into the search and do something, which again created a loop where the tool does not know where to stop. It thinks I need more inputs, again going back and back, it is a loop. To improvise on that, nowadays multi-agents is becoming more popular. Create multi-agents, do a lot of stuff together. When we see it like that, we have a lot of agents working for you. So a security agent trying to figure security concerns, a review agent trying to review the tool, a coding agent trying to fix things. Now again, the more the tools, the more issues you have. Not every agent understands and they have clashed in their understandings where you do not get into the results. So, what do we learn from here? What we see is context is not a problem. Day by day the models are coming where you can dump a lot of context, a lot of data, but does that make sure that the results you are getting is smart enough to give you everything or smart enough to decide what is important. If you see the current LLM models, we see a pattern where it takes the initial inputs you provide, it takes the last inputs, but the in between context is removed. So, they do not focus on the in between context, agents look at the starting point, end point and try to provide you the results. This is like a U curve where some of the things from the start, some of the things from the end make sense, but whatever you are providing in between that, that is not taken up. Yeah. How do you know this? This is something which we are working on and we are actually benchmarking things. So, when we create agents, we try to see this is the context we provide to the agents. Does it take this context into effect and also give us the results? So, we are working with multi-agentic architecture where for each of the tasks we do for code reviews, we give the task to an agent and say, okay, give us the result. Now, every time we, for example, code reviews, we try to see can we give all the context? Can we give the whole code base for example and see if we can get the results? But we see that whenever we start working with that, the initial prompt or the initial goal which we start with that is in focus. If we give something at the end as an input that is in focus, but all between contexts like I have JIRA, I have MCPs, can you look into that? The LLMs try to get rid of those things and push them to make sense by themselves. So, to have this or to make a way out of this, how we deal with is creating strategic solution for context optimization. Rather than dumping everything down to the models and asking them to be smart enough to find out what is more important, we usually start to see, okay, what we can do to make better context for the model. There are lots of solutions in place if you see currently and context engine is a buzzword. Everybody wants to create context engine and everybody wants to provide that. But context engine is like a bouncer, right? So, your high speed car is going and it acts as a bouncer and tells you this is more important. Now, if you have a large messy code base, it makes sense to create a context engine because it creates a search pattern, it creates a ranking logic so that whenever you ask for a task, it looks for those rankings and say this is more important for you, take it and work with it. The problem is the indexing part takes moderate effort, but the scaling is a challenge. Like if you start talking about 600 repositories or 700 repositories, the mapping and the indexing starts to slow down and it becomes again unpredictable to find or create a context engine if you are not actually into making context engine only. There are lots of areas where agent can get more context instead of investing highly on context engine. Hierarchical summarization where instead of creating or going through everything, a summary is created for each file and folder so that when the agents try to find, they can try to read the summary and see if that is more important to us or not, can be a good one. The only thing is that you need a lot of LLM processing. So, every time a file is created or changed, some of the agents need to go and create a mapping for that. So, it is a high upfront context on LLM processing that is needed. Another way is knowledge graph. Now, knowledge graph is complex, but it works wonder when you have logical dependencies. For example, you have one file which impacts another file, which again impacts other files. You can create a graph DB hosting. It is the initial input needed by the developer is very high. It takes a lot of time to create that, but if you have complex logics or you have more dependencies on multiple reports, that works wonder. For me, I think for most of the task, if you are not a product company, but if you are building agents for yourself or your processes, iterative retrieval works really good, because instead of even creating a summary, it creates an index. So, it is like a library card which you give to your agents and see this is the topic. If that is relatable to you, you can look deep into the code and look for the results. Again, it has quite cost impacts, but you do not have to invest a lot of energy. The input required by the developers to provide to the LLM is low, and it provides better results. There is also option of self correction where you ask the LLMs to do something and there is a critic node which looks and say if that is relevant to your initial goal or not. In that case, if the context is lost, you can again ask the agent to do it again, retry it again because the critic node said, this is not the right way. It takes a little bit more time because it adds the latency of running the agents and again, but it does not require a lot of input initially from the developers to create something. Another challenge which I have seen people getting into when they create invest a lot of energy. The input required by the developers to provide to the LLM is low, and it provides better results. There is also option of self correction where you ask the LLMs to do something and there is a critic node which looks and says if that is relevant to your initial goal or not. In that case, if the context is lost, you can again ask the agent to do it again, retry it again because the critic node said, this is not the right way. It takes a little bit more time because it adds the latency of running the agents again, but it does not require a lot of input initially from the developers to create something. Another challenge which I have seen people getting into when they create these agents is the orchestration paradox. Now, what it does is that LLMs are becoming more and more smart. So, when you give them the task, they think, okay, I should use this tool. Maybe I can do better. I should research more on what should I use. It goes into a loop where instead of actually looking to solve the problem, they look for the method to solve the problem. They hop from one method to another method and most of the API tokens are wasted on finding a way to do it rather than doing it. So, they just go into research mode. For example, if you use Opus, the latest and greatest, they will try to see what is the best method to do it and challenging themselves again and again. Maybe not this, another way, another way and it just goes into a loop of trying to do something rather than doing something. To resolve this, we worked with an 80-20 hybrid approach. I think this is one of the most interesting outcomes I have seen or the way to resolve this in a finite loop. What our teams are doing is giving the latest and greatest models or giving the agents power to research 80% of the time. So, you give them the goal and say, okay, try to do whatever you can. But the 20% of the task where you need final validation, you want summarization, those are not something which is free-flowing. Those are more restricted. Those are more hard gates. For example, if I get x results, I want y. It is more deterministic so that the research which is coming from the 80% can be lowered down. Now, when you see, you can always say that the 80% tool can still go on and go into infinite loops. We have mechanisms to work on that. For some organizations, we do counter mechanisms where after four or five counters, you have to work with whatever was the last result. For some of them, they have timeout counters and after five minutes, whatever is the last tool or whatever is the last decision, you work with that and then go back if the results are not good. But you can restrict that 80%. But in short, if you are using anything like discovery or you are trying to see which tool to use, you are trying to plan, those 80% research models are really good. But if you are trying to create a summarization, you are trying to see, this is the research I have got. Now, I have to make a result out of it. The 20% works really well. Now, for 80% usually, you use high reasoning models, the latest and greatest, but you do not need a high reasoning model for the 20% because those 20% things are doing deterministic tasks. They are telling you what is needed. For example, the critic node, which we talked about, they do not need to research, they do not need to find out what is the best thing to do. They just need to see what was your goal, what was the result you are trying to achieve and how to provide or how to summarize that for that. Also, things like if you think about what would be the next possible action, I have this result, what should I go and look for, those are things that can be done by the 80% dynamic models. Whereas, I have all the results from the 80% models, but what is the proper way or what are the proper results that the user is looking for, those decisions can be taken by the 20% model. Finally, this is an interesting failure which we have seen, where as the context grows, teams think we can do everything in one agent because the context window is quite large. We can [SPEAKER_02] put everything, we can ask an agent to do the testing part, we can ask the agent to do the review part, we can ask the agent all kinds of things because the context is the same and they can provide us the results. But make sure that when the agent is going forward, it gets overwhelmed with the inputs. And again, it tries to start losing what was the original task. So, maybe you give four tasks to the agent and somewhere down the line, it focuses on two tasks. So, you get great results for the two tasks, but the other two just get lost in the middle. For that particular purpose, we have something called mixture of agents and that's where you hear a lot of buzz about multiple agents or multi-agentic architecture where instead of one big agent, we create issue-specific expert agents. We create small agents which are doing great in a specific task which they have been provided. Now, to build on top of that, each of the agents come up with their own interesting ideas or results. How to make sure those results combine and make sense together? Because for example, I am trying to search for a vacation, I give an agent to find the best hotel, another agent to find the best location, another agent trying to find the best flights, but all three of them give me different results. The hotel is in Greece, the flight is from Amsterdam to maybe Portugal and everything just doesn't make sense, right? So, for that particular purpose, there is a concept called a judge agent. What it does is try to get all the results and see if they can make sense together. So, now you are doing all the greatest things from different agents, getting the best results from their part, but a judge agent helps us to combine these and make one sense out of it instead of getting so many things which don't make sense together. Something similar is implemented by us. So, this is an R architecture, Codos architecture, where for code reviews, we are using the same formula. So, as part of a PR review, we have a context collector, which actually goes and collects context from the PRs. It could collect context from the context engine, it collects context from the tools, but then it does not start working and giving you the reviews. It actually bifurcates all the context it has provided and passes it on to different agents. Now, what these agents do is basically specialize in what they are supposed to. For example, there will be a security agent trying to find security flaws. There might be an agent trying to find code differences, there might be an agent trying to find the JIRA issues. Once all these agents give us back results, a judge agent actually looks at the results and says, okay, these are interesting enough, but is it relevant to you? We can again go back to the context engine, look into the PRs and see out of the 10 things which were provided by you, how many of them start working and giving you the reviews. It actually bifurcates all the context it has provided and pass it on to different agents. Now, what these agents do is specialize in what they are supposed to. For example, there will be a security agent trying to find security flaws. There might be an agent trying to code, there might be code differences, there might be an agent trying to find the JIRA issues. Once all these agents give us back, a judge agent actually looks at the results and say, okay, these are interesting enough, but is it relevant to you? We can again go back in the context engine, look into the PRs and see out of the 10 things which are provided by you, how many of them actually make sense for your thing. So again, refining the results to make sense to you. Yeah, I think that was it from my side. Any questions? Yes. In practice, how do you let the swarm communicate with each other? You are talking about the agents? So you have agent A and agent B and the judge. I can imagine they write to a file system or do you have some kind of proprietary tool? We use a long chain at the bottom and that is being used to communicate and build infrastructure for different agents. Do you know what long chain uses for that? Is it just collects the responses and then shoves it back into the prompt of the next agent? Yes. That's what we do. So what we do is we try to get the results and create a prompt for the next agent. And if there are multiple things, again, there is an agent just to collect the results and create a better prompt, which is refined for the next agent. Have you thought about a calibration step for each agent? When you say calibration, can you tell me more? Calibration, right? So when you do a code review with an agent, right, you need at least what I heard today might have done some calibration, right? That you actually tell me what is good and what is bad. Yes. So when you say it that way, let me know if that makes sense for your question or not, we do calibration in a form where we check what we have as a context. So for example, when we get the code reviews, LLM does not know what is important for you or how you work. So for example, an LLM when it gets input, it gets input from healthcare industry, it gets from retail industry, it gets from finance industry, and all of them can use the same Java framework in different ways. Different things are important for them, the rest of them doesn't make sense. So what we do is we give you two different options to tell the agents how to perform or what to work on. One part is we give them the PR history. So we index all your PRs and see when was the last time something like this was identified and compare it with the current subversion. If that is... And you do that in the context standpoint? Yes. We do that. So the changes you make to the code, we look into if we can see something similar in the past. That again is transferred to the context twice. First is when we are actually giving context to the sub-agents to find things for you, and another time to the judge agent. So that when I get 15 different recommendations for a code review, my judge agent can look into what was there before, how your reviewer commented, how your developers commented, and based on that decide if that is worth providing to your developers or not. And the other part is... And this happens for every agent? For every agent, yes. And if I understood you right, you don't share the context between the agents, right? So you have a specific context for every agent? Yes, we are trying to resolve that part. Instead of dumping everything to the LLM, we take the part which is more important. We use the context engine for that. We take the part which is more important and only provide that particular part to that particular agent. Very good. Yes, sir. But then for me, it's not clear how you bridge the gap, right? So you have a code quality agent and you have a framework-specific coding review agent, right? And then you basically, as I understood, you only share the specific information to each agent, right? And then basically each agent runs autonomously, right? Yeah. And it doesn't have a full picture, right? And at least when, as a human, right? And if you do code reviews, it's always good if you have a full perspective, right? I would say that kind of methodology works for simple things, does it use linting? How does it use linting? Are tests implemented? But when you think about the overall architecture, for example, to make architecture decisions that cover security, because everything is a balance and you have to weigh that somehow. How do you solve that? Yes. So I think if you look at the older version of code reviews, you should have had a senior engineer who knows your code, who knows what kind of packages you're using, and they can comment on whether the developer has done something similar to what you're used to, or did something totally weird, right? Then you used to have a security person who would see if you are providing all kinds of security, you are not hard coding your APIs, or you are not putting any SQL injections. So all those kinds of things, the security experts know. On the other hand, if you are working with ESO compliance or SOC 2 compliance, there might be an auditor who might ask the team lead or the senior engineer, is your code being logged? Is it logging the changes and so on? So previously as well, there were lots of people having specialized knowledge, looking into those kinds of areas specifically. Now, when the context is provided, it's always like these are my security concerns, which I always have to look into, these are my architectural concerns, an architect might look from the architecture perspective. We can do something similar with the agents as well, because for example, architecture and security concerns, we have a web portal where architects can provide their guidelines, compliance people can provide their guidelines and an agent can look into all these guidelines and say, is it validated or not? So, if you see the initial picture, the context collector knows everything and then it provides relevant context to the agents. So, you basically enforce your customer to upload this kind of document, right? Is that a kind of requirement then? Or because, I mean, the system will have completely different results, right? If you don't share them. Exactly. So, that's something which depends on the organization. There are some people who say, we don't work with any rules or regulations, so just give me out of the box. That also means don't expect the agents to find something very specific to your working unless and until we have certain So, if you see the initial picture, the context collector knows everything and then it provides relevant context to the agents. So, you enforce your customer to upload this kind of document, right? Is that a kind of a requirement then? Or because the system will have completely different result, right? If you don't share them. Exactly. So, that's something which depends upon organization. There are some people who say, we don't work with any rules or regulations, so just give me out of the box. That also means that don't expect the agents to find something very specific to your working unless we have certain PR history because then the PR history kicks in and tells you what is relevant even when you don't provide. I'm not sure if the PR history is really the best source. It can be one. It can be one. It can. So, that's why there are various sources, right? So, it's PR history. It's your resource. It's your... And somebody's cooking food at my home. Yeah, but it can be one of the sources and that's why... And my question is, yeah, I mean, of course, it can be, right? Yeah. At the end, you need to decide, how much you weight, right? So, the new documents for the engineering principles, you take your principles and compare them to your... Yeah, yeah, yeah. ...study, right? But they can be completely out of the balance, right? Oh, it depends. It depends. So, I think it's... If you look at it from one perspective, it's difficult to decide. But if you're getting the context from many angles, for example, PR history, that's one part. But when you do the compliance and you tell in the compliance portal, this is really important. So, we have various segments of it's an error, it's a recommendation. All those kinds of things add weight to a feedback to say if it's good or not. And every time your developer accepts a suggestion, it gets more weighted for the next one. If they do not accept the suggestion, it gets less weight. So, it's all about indexing and making sure those weights are banished somewhere. Yeah. Two ways. One is by when you give your recommendations, does your developer actually accept it or not? We index that. Another way is from the past few years, we try to find out similar issues identified and if your developer actually implemented them. So, for example, some people are used to hard coding their API keys and I literally had a tough argument with the developer, but this is how we do it. No, this should not be the way. Yeah, but... That's a nice thing. It nicely meant also what I meant, right? So, if you look in the history, if that happens, it doesn't mean it's good. It's good. Yeah, and that's the way where the system tries to tell you this is not good, this is not good and then it's up to you and... But if you provide the guidance, right? No, so there is something called bug fixes and there is something called rules. So, if you provide them as a rule, it will get highlighted, doesn't matter if you want it or not. And then there are bugs where agent try to tell you there's something wrong and if the reviewer also agrees with it and did not implement 10 times, the reviewer might get it less weighted and give you. Cool. Thank you so much. Thank you. we take the part which is more important. We use the context engine for that. We take the part which is more important and only provide that particular part to that particular agent. Very wonderful. Yes, sir. But then for me, it's not clear how you bridge the gap, right? So you have a code quality agent and you have a framework specific coding review agent, right? And then you basically, you only, as I understood, you only share the specific information to each agent, right? And then basically each agent runs atomic, autonomously, right? Yeah. And it doesn't have a full picture, right? And at least when, let's say as a human, right? And if you do code reviews, it's always good if you have a full prospect, right? I would say like that kind of methodology, it works for simple things like, does it use linting? Who does it use? I don't know. Are tests implemented? But when you think about, at least I think when you think about the overall architecture, for example, to make architecture decisions that covers security, because everything is a balance and you have to weight that somehow. How do you solve that? Yes. So I think if you look into the older version of code reviews, you should have, you used to have a senior engineer who knows your code, who knows what kind of packages you're using, and they can comment on if the developer has done something similar to what you're used to, or did something totally weird, right? Then you used to have a security person who used to see if you are providing all kinds of security, you are not hard coding your APIs, or you are not putting any SQL injections. So all those kind of things, the security experts know. On the other hand, if you are working with ESO compliance or SOC 2 compliance, there might be an auditor who might ask the team lead or the senior engineer, is your code being, you know, logged? Is it logging the changes and so on? So previously as well, there were lots of people having specialized knowledge, looking into those kind of areas specifically. Now, when the context is provided, it's always like these are my security concerns, which I always have to look into, these are my architectural concerns, an architect might look from the architecture perspective. We can do that something similar with the agents as well, because for example, architecture security concerns, we have a web portal where architects can provide their guidelines, compliance people can provide their guidelines and an agent can look into all these guidelines and say, is it validated or not? So, if you see the initial picture, the context collector knows everything and then it provides relevant context to the agents. So, you enforce basically your customer to upload this kind of document, right? Is that a kind of a requirement then? Or because, I mean, the system will have completely, let's say, different result, right? If you don't share them. Exactly. So, that's something which depends upon organization. There are some people who say, we don't work with any rules or regulations, so just give me out of the box. That also means that don't expect the agents to find something very specific to your working until and unless we have certain PR history because then again, the PR history kicks in and tells you what is relevant even when you don't provide. I'm not sure if the PR history is really the best source. It can be one. It can be one. It can. So, that's why there are various sources, right? So, it's PR history. It's your resource. It's your... And somebody's cooking food at my home. Yeah, but it can be one of the sources and that's why... And my question is like, yeah, I mean, of course, it can be, right? Yeah. At the end, you need to decide, let's say, also in the jury, right, how much you weight, right? So, like, the new documents for, let's say, the engineering principles, like, you take your principles and compare them to, let's say, your... Yeah, yeah, yeah. ...study, right? But they can be completely out of the balance, right? Oh, it depends. It depends. So, again, I think it's... If you look into from one perspective, it's difficult to decide. But if you're getting the context from many angles, for example, PR history, that's one part. But when you do the compliance and you tell in the compliance portal, this is really important. So, we have various segments of it's an error, it's a recommendation. All those kind of things adds weight to a feedback to say if it's good or not. And every time your developer expects a suggestion, it gets more weighted for the next one. If it does not accept the suggestion, it gets a less weight. So, it's all about indexing and making sure those weights are banished somewhere. Yeah. Two ways. One is by when you give your recommendations, does your developer actually accept it or not? We index that. Another way is from the past few years, we try to find out similar issues identified and if your developer actually implemented them. So, for example, some people are used to hard coding their API keys and I literally had a tough argument with the developer, but this is how we do it. No, this should not be the way. Yeah, but... That's a nice thing. It nicely meant also what I meant, right? So, if you look in the history, if that happens, it doesn't mean it's good. It's good. Yeah, and that's the way where the system tries to tell you this is not good, this is not good and then it's up to you and... But if you provide the guidance, right? No, so there is something called bug fixes and there is something called rules. So, if you provide them as a rule, it will get highlighted, doesn't matter if you want it or not. And then there are bugs where agent try to tell you there's something wrong and if the reviewer also agrees with it and did not implement 10 times, the reviewer might get it less weighted and give you. Cool. Thank you so much. Thank you.