Open Reader

User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch

completed 15:37 Jun 28, 2026 Watch on YouTube

Current Status

completed

Video ID

Jx4ZFEAq6bY

RAG / Chat

Enabled
User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch
Description

Utility is all you need! Closing the Agent Learning Loop with Utility-Ranked Memory Most production agent systems have a fatal flaw: they start every run from a blank slate. You have traces in your observability stack and pass/fail judgments in your eval suite, but the agent that runs tomorrow has no memory of why yesterday's runs succeeded or failed. This talk exposes the gap between observation and action and shows how to close it. We'll examine why current memory approaches stall: conversation buffers that only remember recency, semantic systems that retrieve what sounds similar rather than what helped, and reflection-based methods that capture lessons but don't learn which ones actually work. The core idea: utility-ranked memory. Treat memories like a credit score. When a memory is retrieved and the run passes, its utility rises. When the run fails, its utility falls. The ranking formula combines semantic similarity with outcome history. There is also a demo with an example of the product SQL agent, of how it updates the context for the right outcome, everything happening at runtime. Speakers: - Sonam Pankaj (StarlightSearch Inc): Sonam is the CEO and Co-Founder of StarlightSearch. She is also the co-creator of embedanything, which is a Rust pipeline for RAG, which got contributions from Elastic, Milvus, and Qdrant, and has over 450k+ downloads. Prior to Starlight Search, Sonam spent years in developer tools and AI infrastructure, and has worked as a generative AI Evangelist, GTM lead at Articul8, a spin-off of Intel, and AI Researcher at Saama. She has been presenting talks for the last 10 years, and loves to interact with developers. She has been constantly speaking at Berlin Buzzwords, Europe's largest search conference, PyCon DE, and PyData. She also got an opportunity to present her work at Google, Deutsche Bank, and JetBrains. X/Twitter: https://x.com/sonam_pankaj_ LinkedIn: https://www.linkedin.com/in/sonam-pankaj/ GitHub: https://github.com/s

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI agents fail primarily because retrieval doesn't learn from outcomes—eval feedback dies in dashboards instead of updating agent memory, and StarlightSearch's 'utility score' approach makes retrieval rerank memories based on whether they historically helped or hurt task execution.
  • Why it matters: 73% of agentic pipeline failures stem from static retrieval; this runtime learning layer claims 10–15% benchmark improvements without retraining or prompt engineering, directly addressing the eval-to-action gap.
  • Best use: Study for understanding why agent memory fails in production, the utility-score reranking concept, benchmark comparisons (TAU-bench, AgentBench), and the cold-start/drift tradeoffs of runtime experience layers.

Executive Summary

Sonam Pankaj argues that most agent failures trace to retrieval boundaries where user signals and eval outcomes never feed back into future runs. Current memory systems (Mem0, etc.) store static facts or preferences and retrieve by embedding similarity alone, with no mechanism to learn whether a memory helped or hurt task completion. Observability captures traces, evals judge pass/fail, but that signal 'dies in the dashboard'—engineers manually tune prompts or upgrade models instead of the system improving itself at runtime.

StarlightSearch's solution is a 'utility score': semantic similarity weighted by historical usefulness for task execution. Memories are reranked based on whether they contributed to successful outcomes, turning eval feedback into a first-class retrieval signal. After ~10 memory activations, the system distills reasoning into 'skills' that update agent context without changing system prompts. This is runtime improvement (vs. DSPy's compile-time baking) and claims TAU-bench improvement from 66% to 80% with skills, and AgentBench gains from baseline 35.7% to 61.3%.

The live demo shows a product-SQL agent failing to find a 'gaming mouse,' user marking it as failure with feedback ('wireless mouse exists'), and the next query adapting its tool-call trajectory to search more broadly and succeed. Utility scores of memories shift over time as the agent learns which retrieval strategies work. Limitations include cold start (pure semantic search until evals accumulate), utility drift at scale, noisy labels polluting scores, and a hyperparameter (lambda) for credit assignment—most mitigated except cold start.

The pitch positions this as the missing layer between observability and action: a runtime experience store that consumes traces, absorbs evals, and updates retrieval guidance without retraining, fine-tuning, or manual prompt loops. McKinsey 2025 cited 85% of AI analytics failures; Pinecone's new CTO blogged 'we optimized for the wrong thing'—making wrong answers faster/cheaper instead of making retrieval learn. StarlightSearch claims to close that loop by making agents outcome-aware at the retrieval layer.

Key Takeaways

  • Claim: 73% of agentic pipeline failures are caused by retrieval rather than generation or context stuffing. | Evidence: Sonam cites this statistic and references a recent post from Ram Sriharsha (Pinecone CTO) stating 'we have been optimizing for the wrong thing—making wrong answers faster and cheaper but forgetting to make retrieval learn.' | Caveat: No source link or methodology detail provided for the 73% figure; Pinecone CTO post not named/linked. | Implication: If retrieval is the choke point, throwing bigger models or prompt-tuning at the problem won't fix systemic learning gaps—runtime feedback loops are the leverage point. | Timestamp: timestamp unavailable
  • Claim: Current memory systems (Mem0, etc.) retrieve by embedding similarity alone and do not learn from task outcomes or eval feedback. | Evidence: Mem0 example: stores extracted facts/preferences, retrieval signal is embedding similarity, 'Does it learn from outcomes? No.' Observability has traces, evals have pass/fail, but evals 'die in the dashboard' with no path to agent context. | Caveat: No acknowledgment that some systems attempt learned reranking (e.g., LlamaIndex rerankers, RAG-fusion); claim is broad. | Implication: Static retrieval means agents repeat the same mistakes; Ken's agent systems need a closed loop from eval → memory → retrieval to self-improve. | Timestamp: timestamp unavailable
  • Claim: StarlightSearch's 'utility score' reranks memories by semantic similarity weighted by historical usefulness for task execution, turning eval outcomes into a first-class retrieval signal. | Evidence: Utility score = similarity × whether memories helped/hurt execution. After enough activations (~10), system distills reasoning into 'skills' that update agent context automatically. Demo showed agent changing tool-call trajectory after one failure feedback. | Caveat: Cold start problem acknowledged: pure semantic search until evals accumulate. Utility drift and noisy labels can corrupt scores. Hyperparameter lambda for credit assignment is a tuning burden. | Implication: For Ken's agent ops: this is a runtime learning layer, not compile-time (DSPy) or manual prompt-tuning. Tradeoff is cold-start latency vs. self-improvement once feedback loops kick in. | Timestamp: timestamp unavailable
  • Claim: TAU-bench (agent policy compliance) improved from 66% → 76% without skills, 80% with skills. AgentBench with Actions (multi-step reasoning/planning) improved from baseline 35.7% → 58.2% (other memory) → 61.3% (Reflect memory). | Evidence: Named benchmarks: TAU-bench for policy adherence, AgentBench for multi-step workflows. Similar trends claimed for BigCodeBench and LongTV. Human baseline on one task: 47.5%. | Caveat: No details on eval setup, dataset size, task variety, or comparison baselines beyond 'other memory system.' Gains are 10–15% absolute but context-dependent. | Implication: Gains are modest but significant for production agents; Ken should ask if utility-score overhead (compute, latency, label quality) justifies 10–15% improvement vs. simpler rerankers. | Timestamp: timestamp unavailable
  • Claim: Skills auto-update agent context (e.g., remove obsolete SQL columns from system prompt) without manual prompt engineering, once enough memories accumulate. | Evidence: Example: product SQL agent has a column in system prompt that's no longer useful; skills mechanism can update/remove it automatically. 'Similar behavior seen in GPT 5.4' (likely typo for GPT-4o or o1). | Caveat: No detail on how skills are distilled from memories, conflict resolution if memories contradict, or how to audit/roll back bad skill updates. GPT version reference unclear. | Implication: For Ken's GTM/ops: this is auto-tuning of agent prompts/context via feedback loops, reducing manual iteration cycles but adding black-box risk if skill updates go wrong. | Timestamp: timestamp unavailable
  • Claim: Demo: agent failed to find 'gaming mouse,' user marked failure with feedback 'wireless mouse exists,' next query adapted tool-call trajectory and succeeded. | Evidence: First run: 'find gaming mouse' → empty result, zero memories retrieved. User submitted failure feedback. Second run: 'mouse' → found product, tool-call trajectory changed to search more broadly, memory score updated. | Caveat: Demo is single-run anecdote; no multi-user or adversarial-feedback scenarios shown. Utility score evolution over time not quantified beyond 'scores keep changing.' | Implication: Proof-of-concept for runtime learning, but Ken should ask about guardrails for bad feedback (e.g., user marks correct answer as failure), multi-agent drift, and latency overhead. | Timestamp: timestamp unavailable

Detailed Brief

Why agents fail: the retrieval boundary problem

  • Claims: Agents are ReAct loops (reason, act, observe) but missing the learning loop—they don't learn from what worked/didn't work.; 85% of AI analytics application failures cited in McKinsey 2025 report.; 73% of pipeline failures due to retrieval, not generation or context stuffing.; Agents are not 'outcome-aware'—there's a missing layer between evals and action.
  • Evidence: McKinsey 2025 report reference (not linked).; Ram Sriharsha (Pinecone CTO) post: 'optimizing for the wrong thing—wrong answers faster/cheaper, forgot to make retrieval learn.'; Observability: traces, tags, tool calls, exceptions. Evals: pass/fail. But evals don't update agent context, skills, MD files, or actions—signal dies in dashboard.
  • Caveats: McKinsey statistic not sourced in detail; 73% figure not methodology-explained.; No discussion of other failure modes (prompt design, model capability, tool reliability).
  • Implications: Retrieval is the leverage point, not model size or prompt-tuning.; Ken's agent systems need closed-loop feedback from evals → memory → retrieval to self-improve at runtime.; Current manual loops (engineer sees failure, rewrites prompt, redeploys) are unbounded iteration costs.

Current memory systems and their limits

  • Claims: Existing memory systems (Mem0, etc.) store user preferences, profile, conversational history, long-lived personalization.; Retrieval signal is embedding similarity alone—no outcome learning.; Memories are static facts with no context, no history, no reasoning trail.
  • Evidence: Mem0 example: extracted fact preferences, embedding similarity retrieval, 'Does it learn from outcomes? No.'; Customer support bot example: current memory says 'user prefers dark theme,' but doesn't reason 'if refund request, check settlement first to avoid double refund.'
  • Caveats: No acknowledgment of hybrid retrieval systems (keyword + semantic), learned rerankers, or RAG-fusion approaches.; Claim is broad; some systems do attempt outcome-aware retrieval (e.g., LlamaIndex RouterQueryEngine with feedback).
  • Implications: Static retrieval = agents repeat mistakes; dynamic retrieval = self-improvement.; Ken's systems should distinguish memory-as-facts vs. memory-as-reasoning-trails.; Utility score is positioning as differentiation vs. Mem0 and similar.

Utility score and runtime experience architecture

  • Claims: Utility score = semantic similarity weighted by historical usefulness for task execution.; Memories retrieved by semantic similarity to current task, reranked by whether they helped/hurt past outcomes.; Eval outcomes become first-class signal in retrieval reranking.; After ~10 memories, system distills reasoning into 'skills' that auto-update agent context without changing system prompts.; Runtime improvement (vs. DSPy compile-time baking).
  • Evidence: Utility score formula implicit: similarity × outcome-weight.; Skills example: SQL agent has obsolete column in system prompt; skills auto-remove it when memories show it's no longer useful.; Demo: agent changed tool-call trajectory after one failure feedback loop.
  • Caveats: Cold start: pure semantic search until evals accumulate.; Utility drift: similar memories may conflict, scores can drift at scale.; Noisy labels: bad user feedback corrupts utility scores.; Hyperparameter lambda for credit assignment is a tuning burden.; No detail on skill distillation algorithm, conflict resolution, or rollback mechanism.
  • Implications: For Ken: runtime learning means agents improve in production without retraining or manual prompt loops.; Tradeoff: cold-start period where agent is no better than semantic search alone.; Noisy feedback risk: if users mark correct answers as failures, utility scores degrade; guardrails needed.; Lambda tuning may require per-domain experimentation; not one-size-fits-all.

Benchmarks and performance claims

  • Claims: TAU-bench (policy compliance): 66% → 76% without skills, 80% with skills.; AgentBench with Actions (multi-step reasoning): baseline 35.7% → 58.2% (other memory) → 61.3% (Reflect).; Similar trends on BigCodeBench, LongTV.; Human baseline on one task: 47.5%.
  • Evidence: TAU-bench measures if agents follow policy.; AgentBench measures reasoning, planning, tool use over multi-step workflows vs. static Q&A.; Reflect memory outperforms 'other memory system' by 3–5% absolute on AgentBench.
  • Caveats: No eval setup details: dataset size, task variety, baseline comparison method.; 'Other memory system' not named—could be Mem0, could be naive semantic search.; 10–15% absolute gains are context-dependent; real-world tasks may differ.; Human baseline 47.5% suggests task difficulty, but no error analysis or failure mode breakdown.
  • Implications: Gains are modest but production-relevant for agentic systems.; Ken should ask: does 10–15% improvement justify utility-score overhead (compute, latency, label collection)?; Benchmark choice (TAU-bench, AgentBench) focuses on policy adherence and multi-step reasoning—good proxies for production agent tasks.; No cost/latency benchmarks provided; runtime reranking may add inference overhead.

Live demo and system behavior

  • Claims: Product SQL agent demo: 'find gaming mouse' failed (empty result), user marked failure with feedback 'wireless mouse exists,' next query 'mouse' succeeded with updated tool-call trajectory.; Utility scores of memories change over time as agent learns.; After 5+ reviews, findings baked into skills without changing system prompt.; Dashboard shows traces, tool calls, memory retrieval, utility score evolution.
  • Evidence: First run: query 'gaming mouse,' zero memories retrieved, SQL search returned empty, agent said 'couldn't find.'; User submitted failure with feedback.; Second run: query 'mouse,' agent retrieved memory, tool-call trajectory changed to search more broadly, found 'wireless mouse.'; Memory utility scores visible in dashboard, evolving across runs.
  • Caveats: Single-run anecdote; no multi-user, adversarial-feedback, or long-tail query scenarios.; No quantification of utility score convergence time or stability.; Skills distillation shown conceptually but not step-by-step in demo.
  • Implications: Proof-of-concept for runtime learning at retrieval layer.; Ken should ask: how does system handle bad feedback (e.g., user marks correct answer as failure)? Multi-agent environments? Latency overhead?; Dashboard as observability + feedback loop is key UX—agents need human-in-the-loop for label quality.; Skills auto-update is powerful but needs audit/rollback if bad skill emerges.

Limitations and open questions

  • Claims: Cold start: pure semantic search until enough evals accumulate.; Utility drift: similar memories may conflict or scores drift at scale.; Review quality: noisy labels corrupt utility scores.; Hyperparameter lambda: credit assignment tuning required.; Most limits mitigated except cold start.
  • Evidence: Sonam explicitly lists these as limitations.; Cold start 'cannot do much about it' per talk.
  • Caveats: No detail on mitigation strategies for drift, noisy labels, or lambda tuning.; No discussion of multi-tenancy, privacy, or memory isolation across users/tasks.
  • Implications: For Ken: cold start means agents won't improve until evals accumulate—early-run failures still happen.; Noisy labels are a production risk; need label-quality guardrails (e.g., confidence thresholds, anomaly detection).; Lambda tuning may require per-domain experimentation; not plug-and-play.; Drift mitigation unclear—Ken should ask about memory decay, version control, or conflict resolution mechanisms.

Notable Concepts & Terms

  • Utility Score: Semantic similarity weighted by historical usefulness for task execution; core innovation reranking memories by whether they helped/hurt past outcomes, turning eval feedback into retrieval signal.
  • Runtime Experience / Reflect: StarlightSearch's runtime learning layer that improves agent retrieval from experiences without retraining, fine-tuning, or manual prompt-tuning; contrast to DSPy (compile-time) or manual loops.
  • Skills: Distilled reasoning/context updates auto-generated after ~10 memory activations; update agent context (e.g., remove obsolete SQL columns) without changing system prompts; auto-tuning mechanism.
  • TAU-bench: Benchmark measuring if agents follow policy/compliance; StarlightSearch claims 66% → 80% improvement with skills.
  • AgentBench with Actions: Benchmark testing reasoning, planning, tool use over multi-step workflows (vs. static Q&A); baseline 35.7% → 61.3% with Reflect memory.
  • Retrieval Boundary: The interface where user signals, eval outcomes, and feedback fail to cross into agent memory/context; where learning loops break in current systems.
  • Cold Start Problem: Period where utility-score system defaults to pure semantic search until enough eval feedback accumulates; acknowledged unsolved limitation.
  • Lambda (hyperparameter): Credit assignment weight for utility-score reranking; tuning burden for balancing similarity vs. outcome-based ranking.

Operator Notes / Why Ken Should Care

  • Core insight: retrieval is the agentic failure choke point (73%), not model capability or prompt design—runtime feedback loops at retrieval layer are the leverage point Ken should focus on for self-improving agents.
  • Utility score is a reranking innovation: similarity × outcome-usefulness, turning evals into first-class retrieval signals. This closes the eval-to-action gap that kills most agent learning loops.
  • Skills auto-update agent context (e.g., remove obsolete SQL columns) without manual prompt engineering—auto-tuning mechanism for production agents, reducing iteration cycles but adding black-box risk.
  • Benchmarks (TAU-bench, AgentBench) show 10–15% absolute gains, modest but production-relevant. Ken should ask: does utility-score overhead (compute, latency, label collection) justify gains vs. simpler rerankers?
  • Cold start is acknowledged unsolved: agents won't improve until evals accumulate, early-run failures still happen. Noisy labels corrupt utility scores; need label-quality guardrails (confidence thresholds, anomaly detection).
  • Demo proves runtime learning concept but is single-run anecdote; Ken should probe multi-user, adversarial-feedback, long-tail query, and multi-agent drift scenarios before production adoption.
  • Positioning vs. Mem0 and static memory systems; differentiation is outcome-aware retrieval vs. fact storage. Ken's agent ops should distinguish memory-as-facts vs. memory-as-reasoning-trails.
  • No cost/latency benchmarks; runtime reranking may add inference overhead. No discussion of multi-tenancy, privacy, memory isolation, or conflict resolution—Ken should ask about these for production.
  • Skills distillation algorithm, rollback mechanism, and lambda tuning details missing; Ken needs to understand how to audit/rollback bad skill updates and tune lambda per domain.
  • McKinsey 85% failure stat and Pinecone CTO quote are strong positioning hooks but not detailed; Ken should verify sources for diligence.
  • Reflect is runtime (vs. DSPy compile-time); this is key distinction for Ken's agent systems—improvement happens in production vs. pre-deployment baking.
  • For Ken's GTM/content: 'evals signal dies in the dashboard' is a sharp framing of the eval-to-action gap; 'retrieval boundary' as metaphor for where learning breaks is strong messaging.
  • For Ken's investing: StarlightSearch is early-stage (demo-level product), benchmark gains are modest but directionally correct; cold start and noisy labels are production risks to price in.
  • For Ken's agent systems: utility-score approach is worth testing in production agentic workflows where eval feedback is available and retrieval is a known bottleneck (e.g., customer support, SQL agents).
  • Open questions Ken should ask: How does system handle bad feedback? Multi-agent drift? Latency overhead? Memory decay/version control? Audit/rollback for bad skills? Lambda tuning cookbook?

Watch Map

  • timestamp unavailable: Timestamps unavailable; transcript is concatenated/repeated content from live talk. Key sections: problem statement (retrieval boundary), utility-score concept, benchmarks (TAU-bench, AgentBench), live product-SQL demo, limitations (cold start, drift, noisy labels).

Source/Metadata

  • Title: User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch
  • Transcript words: 3232
  • Duration seconds: 937
  • Timestamp note: No timestamps present in transcript; content is repeated/concatenated from live presentation. Video duration 937 seconds (~15.5 min).

Transcript

1807 words en Processed in 79.8s

Hey everyone, I'm Soyim. I'm the CEO and co-founder of Starlight Search and today my topic is user signals die at retrieval boundaries. So let's look into what our agents do, why agents fail, what is the cause of fails in retrieval particularly, and how to make signals cross the retrieval boundary and how to make your agent outcome fair. So let's get started. What is an agent? An agent has an LLM that has agency to reason, invoke tools, interact with the real world, retrieve memory to complete the task. One major loop here is missing: learning. It should also learn from what worked and what didn't work. Suppose I have to explain what an agent is, I can explain it with ReAct agent. So if I have to explain what an agent is, I will explain it with ReAct agent. So basically use it from the agent and execute it in a loop similar to a retrieval search and then pause when the task is complete. This is a very basic ReAct architecture. One thing that is missing is how to make an agent learn from the outcome. So agents keep failing at the same task. That can represent 85% of AI analytics failures in applications. So it's in McKinsey's 2025 report. The problem turned out to be that most of the time retrieval is static. 73% of our pipeline fails because of retrieval non-generation and context stuffing. So a recent post from Ram Sriharsha, the new CTO of Pinecone, said we have been optimizing for the wrong thing. You are paying a note for your agent's memory. This is probably broken and we have been optimizing for the wrong things. We made wrong answers appear faster and cheaper, but we forgot to make retrieval learn. So why does this matter? Again, the third problem is agents are not outcome-fair. So there's a missing layer between evals and action. Your observability has all the traces, all these tags that capture every tool call, every LLM completion, every exception. Your evals judge whether the final output was correct or wrong, basically pass or fail. But these evals are not reflected in agent context, skills, MD files, or agent action in any way. So the agent doesn't have any access to why yesterday's runs passed or failed. The evals signal dies in the dashboard. This is a missing layer—a system that consumes traces, absorbs eval, and converts both into retrieval guidance for future runs. So there's manual improvement and engineers actually have to sit and see if the email and ability to perform well. Relight the prompt, redeploy it, upgrade either to an expensive model, restructure to or harness or fine-tune custom models. Why are current memories failing? Why memories was designed to address this, but it's not. So let's see what we have in the current system. Current memory basically stores user preferences, profile, conversational history, or long-lived personalization. Chart experiences know itself improving learning systems for production. If you see the already existing approach in the market, there's Memzero, which does extracted fact preferences. So users' retrieval signal is embedding similarity. Does it learn from our code? No. So we have come up with something called utility score, which is similarity weighted by how useful it is for the agent to execute the task. It has history of past processes and past outcomes. So we came up with agent architects, and that is agents with runtime experience. It's a runtime layer that lets the agent improve from experiences without retraining, fine tuning, or manual prompt training. It's different from compile time like DSPy because you bake in all the lessons in the prompt. Here it's actually improving while it is executing the task. So let's introduce utility score. You do not retrieve by keyword. You retrieve by semantic similarity to the current task rated by whether those memories have historically helped or hurt the execution or the outcome. The eval outcome becomes a first class signal in the retrieval reranking and not just for. One of the key things is it should get memory as reasoning, not as static facts with no context and no history, but reasoning. Like suppose if there's a customer support bot looking for a refund, it will not only say, hey, user prefers dark theme or user prefers to be called by a shorter name. It actually reasons about the query. Like if someone asks for a refund, you should check the settlement before refunding it so that the customer doesn't get paid refund twice. So rerank based on usefulness. Context is updated based on tasks. So this is a very big thing because most agents fail with context stuffing and this has been brought up in the past and learned from history and reasoning, right? Talking about benchmarks. So we have benchmarked our memory system reflect with on towel bench, which essentially measures if agents have followed the policy well or not. So we have seen the performance improve from 66 to 76% without baking in skills and with skills, reflect performance at 80. So once there are enough memories, like 10 memories, what we do is we bake in the reasoning and the understanding into skills so that your agent always remains updated. What happens most of the time—we have seen suppose you have a product SQL agent and there's a column in system prompt. Even though that column is not useful anymore, it remains in the system prompt. So there's no system right now that can update that—hey, there's no column called this right now. So maybe probably you shouldn't entertain it in the future. And this is possible with skills because it always uses calls that skill, updated skill all the time. And similar behavior has been seen in GPT 5.4. We have also benchmarked on agentic tasks, which essentially test a model's ability to reason, plan and use tools over extended multi-step workflows rather than measuring static Q&A. So you can see here, suppose the human last exam with rock, you get 47.5. If it is starting from the baseline 35.7, with the other memory system, it gets to 58.2, but with the refilled memory system, it gets to 61.3%. So this trend is shown in another agentic benchmarks as well, like big code bench, like Long TV, et cetera. So of course there are limitations to this approach. First of all, there's cold start. In the beginning, it's pure semantic search until enough reviews have been activated. There's utility drift. Maybe sometimes similar memories could come up. There are a lot of problems that could come at scale with this experience, but we have come back most of them. There's a review quality. So noisy labels can make the utility noisy as well. And there is a hyperparameter called lambda that is associated with credit and reranking. So we have built reflect in such a way that most of these problems and most of these limits are now reduced, except for cold start, which we cannot do much about it. Let's now get into the demo. So let's check this demo. So it's basically a product SQL demo. I'll ask it to search some product in a SQL database and let's see if it is able to fill it out. So I gave it find me a gaming mouse. Zero memories retrieved. I couldn't find a gaming mouse. Okay. So maybe let's just go and see what's happening in the dashboard. Okay. It came out that— Coming on another benchmark that with the agent bench with actions essentially measures the agent has actually done the right reasoning, planning and have followed the agent has actually done the right reasoning. But essentially test a model's ability to reason, plan and use tools or extended multi-step workflows rather than measuring a static Q&A. So you can see here, suppose the human last exam with the memory system, it gets to 58.2, but with the refilled memory system, it gets to 61.3%. So this trend is shown in another agentic benchmarks as well, like big code bench, like Long TV, et cetera. So wait, let's check, it couldn't find. I couldn't find any gaming mouse in the product catalog. So suppose I want to mark it fail and tell it, give any control of that. Because there is a wireless mouse in the database and I have submitted the failure. This was the input, this was the response. I couldn't find any, and this was the trajectory and tool calls it took. So you can see the product is empty with this product or with this query. So let's check again what happens now. Let's see now, what happens. So it searched a wireless mouse. Let's see what happens in the dashboard now. Finally, I'm giving mouse as input. The response was, I found a product related to mouse. We must have to find any kind of relevant. Let, but the most important thing is how the tool call evolved. So you can see that previously, the tool calls or the trajectory used to look just with one search at the search products, it called, and the product was empty. Now the trajectory has changed in production and it found something in the product that. And it is still calling search product to call, but it is getting some answer that we fed in as feedback. So that's the demo. And the most important thing that is happening over here is it's forming memories, which is retrieved based on the utility score. That is the score, which basically keeps improving, keeps reranking itself on the basis of how useful that memory was. So you can see from the traces, past traces, which memory, so this memory scores keep changing. And after a certain while, after suppose five reviews, you can bake in these findings or these new updates in a skill without changing any system prompt. You can actually update certain things that agent draws a lot from. So you can update the skill, which is very cool. So I hope everything was, I hope you get to try this new feature, this new runtime experience that we have learned. If you want more details, you can visit our website and you can also contact me at cinema3styles.com or visit me. Thank you. Why are current memories failing? Why memories was designed to actually address this, but, uh, it's not. So let's see, uh, what we have in as a current system, uh, and current memory is that they basically store user preferences, profile, conversational history, or wrong-lived personalization. So, chart experiences know itself improving learning systems for production. If you see, uh, the already existing approach in the, uh, market, there's a, uh, launching, uh, there's Memzero, which does extracted fact preferences. So, uh, users retrieval signal is, uh, embedding similarity. Does it learn from our code? No. Uh, so we have come up with something called utility score, which is a similarity weighted by how useful it is for the agent to execute the task. It has actually the history of past process and past outcomes. So we came up with agent architects and that is agents with runtime experience. It's a runtime layer that let the agent improve from experiences without retraining, fine tuning, or manual prompt training. It's a bit different from compile time like DSPy because, uh, you bake in all the lessons in the prompt. Here it's actually improving, uh, while it is executing. The task. So let's, uh, again, introduce us, uh, utility score. So you do not retrieve by keyword. You'd retrieve by semantic similarity to the current task rated by whether those memories have historically helped or hurt the execution or the outcome. The evil outcome becomes a first class signal in the retrieval relanking and not just for. One of the key things is it should get memory as reasoning, not as facts, static facts with no context and no history, but reasoning. Like suppose, uh, if, uh, if there's a, if there is a customer support bot looking for a refund, it will not only say, Hey, user prefers, uh, uh, uh, dark theme or user prefers, uh, to be called by a shorter name. It's actually reason about the query. Like, uh, if some, someone asks for refund, you should check the settlement before refunding it so that the, the customer doesn't get paid, uh, refund twice. So relank based on usefulness. Context is updated based on tasks. So this is a very big thing, uh, because most of the agents fail with contact stuffing and this has been brought up in the past and learned from history and reasoning, right? Talking about benchmarks. So, uh, we have benchmarked, uh, our memory system reflect with on towel bench, which essentially measure if agents have followed the policy well or not. So we have seen, uh, the, uh, uh, the performance improve from 66 to 76% without baking in, uh, skills and with skills, uh, reflect performance at 80. So, uh, once there are, uh, enough memory, like 10, uh, memories, uh, what we do is we baking the, uh, reasoning and the understanding into skills so that your agent always remains updated. What happens most of the time we have seen, suppose you have a product SQL agent and there's a system, there's a column in system prompt, even though that column is not useful anymore, it remains as the system prompt. So there's no system right now that can update that. Hey, there's no column, uh, uh, right now called this. So maybe probably you shouldn't entertain it in the future. And this is possible with skills that, uh, because it is always uses, uh, calls that skill, uh, updated skill all the time. And the similar behavior has been seen in GPT 5.4. Uh, we have also benchmarked on agentic tasks, which essentially test a model's ability to reason, plan and use tools overextended with multi-step workflows rather than measuring the static Q and a. So you can see here, uh, with the suppose the human last exam, uh, with rock, you get 47.5. If it is starting from the baseline 35.7, uh, with the other memory system, it gets to 58.2, but with the refilled memory system, uh, it gets to 61.3%. So this, uh, this kind of trend is shown in another, uh, uh, agentic benchmarks, uh, as well, like a big coat bench, like long TV, et cetera. So of course there are limitations to this, uh, the, this approach. First of all, there's a cool strut. So in the beginning, it's pure semantic search and the, uh, enough reviews have been activated. Um, there's a utility drift. Maybe sometimes, uh, similar memories could come up. There are a lot of problems that could come at scale with this experience, but we have come back most of them. Uh, there's a review quality. So noisy labels can make the utility noisy as well. And there is a hyper parameter called lambda that is associated with credit and re-ranking. So we have built reflect in such a way that most of these problems and most of these, uh, limits are now reduced, uh, except for cold start, which we cannot do much about it. Uh, let's now get into the demo. So let's check this demo. So it's a basically a product SQL demo, uh, I'll give, ask it to search some product in a SQL database and let's see if it is able to fill it out. So I gave it find me a gaming mouse. Zero memories retrieved. I couldn't find a gaming mouse. Uh, okay. So maybe let's just go and see what's happening in the dashboard. Okay. It came out that. Coming on another benchmark that with the agent bench with actions essentially measures. The agent has actually done the right reasoning planning and, uh, have followed. The agent has actually done the right reasoning. But essentially test a model's ability to reason, plan and use tools or extended multi-step workflows rather than measuring a static Q and A. So you can see here, uh, with the, suppose the human last exam, uh, with the, uh, with the, uh, with the, uh, with the, uh, with the, uh, with the, uh, with the memory system, it gets to 58.2, but with the refilled memory system, uh, it gets to 61.3%. So this, uh, this kind of trend is shown in another, uh, agentic benchmarks, uh, as well, like, uh, big code bench, like long TV, et cetera. So, uh, wait, let's check, uh, it couldn't find, I couldn't find any gaming mouse in the product catalog. So suppose I want to mark it fail and tell it, give any control of that. Because there is, uh, there is a wireless mouse in the database and I have submitted the failure. Uh, this was the input, this was the response. I couldn't find any, and this was a trajectory and who call it took. So you can see the product is empty with this product or with this query. So let's, uh, let's check again what happens now. Let's see now, now what happens. I.T. Uh, So it searched a wireless mouse. Let's see what happens in the dashboard now. Finally, I'm giving mouse was input. The response was, I found a product related to mouse as we must have to find any kind of relevant. Let, but the most important thing is how the tool call evolved. So, uh, you can see that previously, uh, the tool calls, uh, or the trajectory, uh, used to look just with one search at the search products, uh, it called, uh, and the product was empty. Now the trajectory has changed in production and, uh, it found something in the product that. And it is still calling search product to call, but it is getting some answer that we fed in as a feedback. So, uh, that's the demo. And the most important thing that is happening over here is it's forming memories, which is retrieved based on the utility score. That is the score, which basically keeps improving, uh, keeps re-ranking itself on the basis of how useful that memory was. So, uh, you can see from the traces, past traces, which memory, uh, so this memory scores keep changing. And after certain while, after suppose, uh, five reviews need, you can bake in these findings or these, uh, new updates in a skill without, uh, so without changing any system prompt, you can actually update certain things that agent, uh, uh, draws, uh, a lot from. So, uh, uh, you can update the skill, uh, which is very cool. So, uh, I hope everything was, I hope you get to try, uh, this new, uh, feature, this new, uh, runtime experience that we have learned. If you want more details, you can visit our website and, uh, you can also contact me at cinema3styles.com or visit me. Thank you.