SPEAKER_00
Hey everyone, I'm Soyim. I'm the CEO and co-founder of Starlight Search and today my topic is user signals die at retrieval boundaries. So let's look into what our agents do, why agents fail, what is the cause of fails in retrieval particularly, and how to make signals cross the retrieval boundary and how to make your agent outcome fair.
SPEAKER_00
So let's get started. What is an agent? An agent has an LLM that has agency to reason, invoke tools, interact with the real world, retrieve memory to complete the task. One major loop here is missing: learning. It should also learn from what worked and what didn't work. Suppose I have to explain what an agent is, I can explain it with ReAct agent. So if I have to explain what an agent is, I will explain it with ReAct agent. So basically use it from the agent and execute it in a loop similar to a retrieval search and then pause when the task is complete. This is a very basic ReAct architecture. One thing that is missing is how to make an agent learn from the outcome. So agents keep failing at the same task. That can represent 85% of AI analytics failures in applications. So it's in McKinsey's 2025 report. The problem turned out to be that most of the time retrieval is static. 73% of our pipeline fails because of retrieval non-generation and context stuffing. So a recent post from Ram Sriharsha, the new CTO of Pinecone, said we have been optimizing for the wrong thing. You are paying a note for your agent's memory. This is probably broken and we have been optimizing for the wrong things. We made wrong answers appear faster and cheaper, but we forgot to make retrieval learn.
SPEAKER_00
So why does this matter? Again, the third problem is agents are not outcome-fair.
SPEAKER_00
So there's a missing layer between evals and action. Your observability has all the traces, all these tags that capture every tool call, every LLM completion, every exception. Your evals judge whether the final output was correct or wrong, basically pass or fail. But these evals are not reflected in agent context, skills, MD files, or agent action in any way. So the agent doesn't have any access to why yesterday's runs passed or failed. The evals signal dies in the dashboard. This is a missing layer—a system that consumes traces, absorbs eval, and converts both into retrieval guidance for future runs.
SPEAKER_00
So there's manual improvement and engineers actually have to sit and see if the email and ability to perform well. Relight the prompt, redeploy it, upgrade either to an expensive model, restructure to or harness or fine-tune custom models. Why are current memories failing? Why memories was designed to address this, but it's not. So let's see what we have in the current system. Current memory basically stores user preferences, profile, conversational history, or long-lived personalization. Chart experiences know itself improving learning systems for production.
SPEAKER_00
If you see the already existing approach in the market, there's Memzero, which does extracted fact preferences. So users' retrieval signal is embedding similarity. Does it learn from our code? No.
SPEAKER_00
So we have come up with something called utility score, which is similarity weighted by how useful it is for the agent to execute the task. It has history of past processes and past outcomes. So we came up with agent architects, and that is agents with runtime experience. It's a runtime layer that lets the agent improve from experiences without retraining, fine tuning, or manual prompt training. It's different from compile time like DSPy because you bake in all the lessons in the prompt. Here it's actually improving while it is executing the task. So let's introduce utility score. You do not retrieve by keyword. You retrieve by semantic similarity to the current task rated by whether those memories have historically helped or hurt the execution or the outcome. The eval outcome becomes a first class signal in the retrieval reranking and not just for.
SPEAKER_00
One of the key things is it should get memory as reasoning, not as static facts with no context and no history, but reasoning. Like suppose if there's a customer support bot looking for a refund, it will not only say, hey, user prefers dark theme or user prefers to be called by a shorter name. It actually reasons about the query. Like if someone asks for a refund, you should check the settlement before refunding it so that the customer doesn't get paid refund twice. So rerank based on usefulness. Context is updated based on tasks. So this is a very big thing because most agents fail with context stuffing and this has been brought up in the past and learned from history and reasoning, right?
SPEAKER_00
Talking about benchmarks. So we have benchmarked our memory system reflect with on towel bench, which essentially measures if agents have followed the policy well or not. So we have seen the performance improve from 66 to 76% without baking in skills and with skills, reflect performance at 80. So once there are enough memories, like 10 memories, what we do is we bake in the reasoning and the understanding into skills so that your agent always remains updated.
SPEAKER_00
What happens most of the time—we have seen suppose you have a product SQL agent and there's a column in system prompt. Even though that column is not useful anymore, it remains in the system prompt. So there's no system right now that can update that—hey, there's no column called this right now. So maybe probably you shouldn't entertain it in the future. And this is possible with skills because it always uses calls that skill, updated skill all the time. And similar behavior has been seen in GPT 5.4. We have also benchmarked on agentic tasks, which essentially test a model's ability to reason, plan and use tools over extended multi-step workflows rather than measuring static Q&A. So you can see here, suppose the human last exam with rock, you get 47.5. If it is starting from the baseline 35.7, with the other memory system, it gets to 58.2, but with the refilled memory system, it gets to 61.3%. So this trend is shown in another agentic benchmarks as well, like big code bench, like Long TV, et cetera.
SPEAKER_00
So of course there are limitations to this approach. First of all, there's cold start. In the beginning, it's pure semantic search until enough reviews have been activated. There's utility drift. Maybe sometimes similar memories could come up. There are a lot of problems that could come at scale with this experience, but we have come back most of them. There's a review quality. So noisy labels can make the utility noisy as well. And there is a hyperparameter called lambda that is associated with credit and reranking. So we have built reflect in such a way that most of these problems and most of these limits are now reduced, except for cold start, which we cannot do much about it. Let's now get into the demo. So let's check this demo.
SPEAKER_00
So it's basically a product SQL demo. I'll ask it to search some product in a SQL database and let's see if it is able to fill it out. So I gave it find me a gaming mouse. Zero memories retrieved. I couldn't find a gaming mouse. Okay. So maybe let's just go and see what's happening in the dashboard. Okay. It came out that—
SPEAKER_00
Coming on another benchmark that with the agent bench with actions essentially measures the agent has actually done the right reasoning, planning and have followed the agent has actually done the right reasoning. But essentially test a model's ability to reason, plan and use tools or extended multi-step workflows rather than measuring a static Q&A. So you can see here, suppose the human last exam with the memory system, it gets to 58.2, but with the refilled memory system, it gets to 61.3%. So this trend is shown in another agentic benchmarks as well, like big code bench, like Long TV, et cetera.
SPEAKER_00
So wait, let's check, it couldn't find. I couldn't find any gaming mouse in the product catalog. So suppose I want to mark it fail and tell it, give any control of that. Because there is a wireless mouse in the database and I have submitted the failure. This was the input, this was the response. I couldn't find any, and this was the trajectory and tool calls it took. So you can see the product is empty with this product or with this query. So let's check again what happens now.
SPEAKER_00
Let's see now, what happens. So it searched a wireless mouse. Let's see what happens in the dashboard now. Finally, I'm giving mouse as input. The response was, I found a product related to mouse. We must have to find any kind of relevant. Let, but the most important thing is how the tool call evolved. So you can see that previously, the tool calls or the trajectory used to look just with one search at the search products, it called, and the product was empty.
SPEAKER_00
Now the trajectory has changed in production and it found something in the product that. And it is still calling search product to call, but it is getting some answer that we fed in as feedback. So that's the demo. And the most important thing that is happening over here is it's forming memories, which is retrieved based on the utility score. That is the score, which basically keeps improving, keeps reranking itself on the basis of how useful that memory was. So you can see from the traces, past traces, which memory, so this memory scores keep changing. And after a certain while, after suppose five reviews, you can bake in these findings or these new updates in a skill without changing any system prompt. You can actually update certain things that agent draws a lot from. So you can update the skill, which is very cool. So I hope everything was, I hope you get to try this new feature, this new runtime experience that we have learned. If you want more details, you can visit our website and you can also contact me at cinema3styles.com or visit me. Thank you.
SPEAKER_00
Why are current memories failing? Why memories was designed to actually address this, but, uh, it's not. So let's see, uh, what we have in as a current system, uh, and current memory is that they basically store user preferences, profile, conversational history, or wrong-lived personalization. So, chart experiences know itself improving learning systems for production. If you see, uh, the already existing approach in the, uh, market, there's a, uh, launching, uh, there's Memzero, which does extracted fact preferences. So, uh, users retrieval signal is, uh, embedding similarity. Does it learn from our code? No.
SPEAKER_00
Uh, so we have come up with something called utility score, which is a similarity weighted by how useful it is for the agent to execute the task. It has actually the history of past process and past outcomes. So we came up with agent architects and that is agents with runtime experience. It's a runtime layer that let the agent improve from experiences without retraining, fine tuning, or manual prompt training. It's a bit different from compile time like DSPy because, uh, you bake in all the lessons in the prompt. Here it's actually improving, uh, while it is executing. The task. So let's, uh, again, introduce us, uh, utility score. So you do not retrieve by keyword.
SPEAKER_00
You'd retrieve by semantic similarity to the current task rated by whether those memories have historically helped or hurt the execution or the outcome. The evil outcome becomes a first class signal in the retrieval relanking and not just for. One of the key things is it should get memory as reasoning, not as facts, static facts with no context and no history, but reasoning. Like suppose, uh, if, uh, if there's a, if there is a customer support bot looking for a refund, it will not only say, Hey, user prefers, uh, uh, uh, dark theme or user prefers, uh, to be called by a shorter name. It's actually reason about the query. Like, uh, if some,
SPEAKER_00
someone asks for refund, you should check the settlement before refunding it so that the, the customer doesn't get paid, uh, refund twice. So relank based on usefulness. Context is updated based on tasks. So this is a very big thing, uh, because most of the agents fail with contact stuffing and this has been brought up in the past and learned from history and reasoning, right? Talking about benchmarks. So, uh, we have benchmarked, uh, our memory system reflect with on towel bench, which essentially measure if agents have followed the policy well or not. So we have seen, uh, the, uh,
SPEAKER_00
uh, the performance improve from 66 to 76% without baking in, uh, skills and with skills, uh, reflect performance at 80. So, uh, once there are, uh, enough memory, like 10, uh, memories, uh, what we do is we baking the, uh, reasoning and the understanding into skills so that your agent always remains updated. What happens most of the time we have seen, suppose you have a product SQL agent and there's a system, there's a column in system prompt, even though that column is not useful anymore, it remains as the system prompt. So there's no system right now that can update that. Hey, there's no column, uh,
SPEAKER_00
uh, right now called this. So maybe probably you shouldn't entertain it in the future. And this is possible with skills that, uh, because it is always uses, uh, calls that skill, uh, updated skill all the time. And the similar behavior has been seen in GPT 5.4. Uh, we have also benchmarked on agentic tasks, which essentially test a model's ability to reason, plan and use tools overextended with multi-step workflows rather than measuring the static Q and a. So you can see here, uh, with the suppose the human last exam, uh, with rock, you get 47.5. If it is starting from the baseline
SPEAKER_00
35.7, uh, with the other memory system, it gets to 58.2, but with the refilled memory system, uh, it gets to 61.3%. So this, uh, this kind of trend is shown in another, uh, uh, agentic benchmarks, uh, as well, like a big coat bench, like long TV, et cetera. So of course there are limitations to this, uh, the, this approach. First of all, there's a cool strut. So in the beginning, it's pure semantic search and the, uh, enough reviews have been activated. Um, there's a utility drift. Maybe sometimes, uh, similar memories could come up. There are a lot of problems that could come at scale with this
SPEAKER_00
experience, but we have come back most of them. Uh, there's a review quality. So noisy labels can make the utility noisy as well. And there is a hyper parameter called lambda that is associated with credit and re-ranking. So we have built reflect in such a way that most of these problems and most of these, uh, limits are now reduced, uh, except for cold start, which we cannot do much about it. Uh, let's now get into the demo. So let's check this demo.
SPEAKER_00
So it's a basically a product SQL demo, uh, I'll give, ask it to search some product in a SQL database and let's see if it is able to fill it out.
SPEAKER_00
So I gave it find me a gaming mouse. Zero memories retrieved. I couldn't find a gaming mouse. Uh, okay. So maybe let's just go and see what's happening in the dashboard.
SPEAKER_00
Okay. It came out that.
SPEAKER_00
Coming on another benchmark that with the agent bench with actions essentially measures. The agent has actually done the right reasoning planning and, uh, have followed. The agent has actually done the right reasoning. But essentially test a model's ability to reason, plan and use tools or extended multi-step workflows rather than measuring a static Q and A. So you can see here, uh, with the, suppose the human last exam, uh, with the, uh, with the, uh, with the, uh, with the, uh, with the, uh, with the, uh, with the memory system, it gets to 58.2, but with the refilled memory system, uh, it gets to 61.3%.
SPEAKER_00
So this, uh, this kind of trend is shown in another, uh, agentic benchmarks, uh, as well, like, uh, big code bench, like long TV, et cetera.
SPEAKER_00
So, uh, wait, let's check, uh, it couldn't find, I couldn't find any gaming mouse in the product catalog. So suppose I want to mark it fail and tell it, give any control of that. Because there is, uh, there is a wireless mouse in the database and I have submitted the failure. Uh, this was the input, this was the response. I couldn't find any, and this was a trajectory and who call it took. So you can see the product is empty with this product or with this query. So let's, uh, let's check again what happens now.
SPEAKER_00
Let's see now, now what happens.
SPEAKER_00
I.T. Uh, So it searched a wireless mouse.
SPEAKER_00
Let's see what happens in the dashboard now.
SPEAKER_00
Finally, I'm giving mouse was input. The response was, I found a product related to mouse as we must have to find any kind of relevant. Let, but the most important thing is how the tool call evolved. So, uh, you can see that previously, uh, the tool calls, uh, or the trajectory, uh, used to look just with one search at the search products, uh, it called, uh, and the product was empty. Now the trajectory has changed in production and, uh, it found something in the product that. And it is still calling search product to call, but it is getting some answer that we fed in as a feedback. So, uh, that's the demo. And the most important thing that is happening over here
SPEAKER_00
is it's forming memories, which is retrieved based on the utility score. That is the score, which basically keeps improving, uh, keeps re-ranking itself on the basis of how useful that memory was. So, uh, you can see from the traces, past traces, which memory, uh, so this memory scores keep changing. And after certain while, after suppose, uh, five reviews need, you can bake in these findings or these, uh, new updates in a skill without, uh, so without changing any system prompt, you can actually update certain things that agent, uh, uh, draws, uh, a lot from. So, uh, uh, you can update the skill,
SPEAKER_00
uh, which is very cool. So, uh, I hope everything was, I hope you get to try, uh, this new, uh, feature, this new, uh, runtime experience that we have learned. If you want more details, you can visit our website and, uh, you can also contact me at cinema3styles.com or visit me. Thank you.