Open Reader

Continual Learning for AI Agents: From Failures to Durable Improvements - Soheil Feizi, RELAI

completed 22:35 Jul 05, 2026 Watch on YouTube

Current Status

completed

Video ID

2IxD9OB3XuQ

RAG / Chat

Enabled
Continual Learning for AI Agents: From Failures to Durable Improvements - Soheil Feizi, RELAI
Description

Agents fail in production in ways that static benchmarks cannot fully capture. The key question is whether they can learn from those experiences without drifting or breaking prior capabilities. This talk introduces verifiable continual learning for AI agents: a framework for converting traces, failures, and feedback into testable, regression-aware improvements. I will discuss four core requirements: turning failures into replayable learning environments, preserving prior capabilities during updates, routing repairs to the right layer of the agent stack, and keeping the learning loop efficient enough to run continuously. We will use these principles to examine today’s approaches, including prompt optimizers, memory consolidation, coding-agent harness repair, and trace-to-harness systems. I will then discuss the remaining gap: a holistic, lifelong, verifiable learning loop with online regression control. Speakers: - Soheil Feizi (RELAI): Dr. Soheil Feizi is the Founder and CSO of RELAI and an Associate Professor of Computer Science at the University of Maryland, College Park, whose work focuses on the reliability, safety, and optimization of AI systems. X/Twitter: https://x.com/FeiziSoheil LinkedIn: https://www.linkedin.com/in/soheil-feizi-b14a4895/

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Production AI agents need verifiable continual learning—turning failures/logs into replayable test environments and making regression-aware improvements across memory/harness/model layers, not just model fine-tuning.
  • Why it matters: Most teams treat agent improvement as model tuning or prompt iteration; Feizi presents a framework (and product, RELi) that makes every fix testable, measurable, and regression-proof, addressing the core problem of agents forgetting or breaking past successes when patching new failures.
  • Best use: Ken should watch fully for the four-principle framework (replayability, holisticness, lifelongness, efficiency) and understand how to convert production logs into learning environments—critical for AI ops, agentic systems, and any team building custom agents that need to improve in production without catastrophic forgetting.

Executive Summary

Soheil Feizi (founder/CSO at RELi, UMD professor) argues that agent continual learning is fundamentally different from model fine-tuning. Most teams have production logs and feedback but no way to replay, test, and verify improvements. His central claim: logs alone are not learning environments—you must infer replayable simulations (with mock/real tools, synthetic users, and evaluators) from one-shot observations, then optimize agents holistically across three layers: memory (cheapest/fastest, e.g. Letta, Mem0), harness (prompts, tools, code; methods like GEPA, trace-to-harness), and model (most expensive; SFT, DPO, RLHF). The problem is that current methods either lack testability (white-box trace-to-harness edits) or require upfront benchmarks (GEPA, RL methods). Neither handles regression: fixing a new failure often breaks past successes.

Feizi introduces Verifiable Continual Learning (VCL) with four principles: (1) Replayability—turn one-off failures into executable tests; (2) Holisticness—route fixes to the right layer (memory/harness/model) with the smallest durable change; (3) Lifelongness—regression-aware optimization, where new fixes are constrained to preserve past learning environments, not just post-hoc checked; (4) Efficiency—the loop runs frequently, and updates scale sub-linearly with the number of past environments. RELi's product embeds this loop: signals (logs, feedback, instructions) → lift to learning environments → root-cause analysis → regression-aware optimize → reviewable PR with explanation of what changed and why. He shows a demo on a fictional support agent benchmark with regression traps, achieving 10% average improvement per loop (78% → 97% in one iteration) without breaking prior tasks.

Key takeaways: (1) Agent continual learning ≠ model fine-tuning—most wins happen in harness and memory. (2) Production logs are not learning environments—you must infer testable simulations. (3) The frontier is regression-aware continual improvement—verify that new fixes don't forget old ones. RELi claims you can add VCL to any agent in two lines of code (one-time setup, then create learning environments and call optimize). Feizi does not discuss pricing, data requirements for lifting logs to environments, or how well the system handles tool/API changes, policy drift, or multi-agent scenarios. The talk is a product pitch wrapped in research framing, but the framework and demo are concrete enough to be operationally useful.

Key Takeaways

  • Claim: Production logs with feedback are not sufficient for agent learning—you need replayable learning environments with simulation and evaluators. | Evidence: Feizi shows a session log example where the user is unhappy but there's no explicit test. He distinguishes logs (one observation) from learning environments (a distribution you can replay with defined grading). He says you must infer how tools behave (mock vs. real, what data to use), synthetic users, and success evaluators from a single log instance. | Caveat: The talk does not explain how RELi's system handles ambiguous logs, missing context, or when tool behavior changes between log capture and replay (e.g., API updates, data policy shifts). It's unclear how much manual curation or data volume is needed to lift logs reliably. | Implication: Ken should treat raw logs as starting material, not training data. For AI ops teams, this means investing in a 'log-to-simulation' pipeline before attempting any continual learning. For product/agent teams, this is a make-or-break step: without replayability, you can't test fixes or measure regressions. | Timestamp: 02:40
  • Claim: Agent improvements should happen across three layers—memory (cheapest/fastest), harness (prompts, tools, code), and model (most expensive)—and the system should choose the smallest durable change. | Evidence: Feizi categorizes memory layer (methods like Letta, Mem0 for storing facts/skills), harness layer (GEPA for prompt search, trace-to-harness for code edits), and model layer (SFT, DPO, RLHF for weight updates). He argues that memory writes are unverified, trace-to-harness is white-box and untestable, and GEPA/RL methods require benchmarks. Example: a stale policy failure could stem from memory, prompt, tool normalization, workflow, or model—root cause analysis should route the fix to the right layer. | Caveat: The talk doesn't provide metrics on how RELi's root-cause analysis decides which layer to target, or how often it misroutes. It's also unclear whether the system can make cross-layer fixes (e.g., update memory and harness together) or if that requires multiple loops. | Implication: For Ken's operator lens: most teams over-index on model fine-tuning when harness/memory changes are faster and sufficient. The framework suggests a triage system: try memory first, then harness, then model only if necessary. For AI ops, this means building observability into which layer caused a failure, not just that a failure occurred. | Timestamp: 06:30
  • Claim: Regression-aware optimization must be integrated into the learning loop, not treated as a post-hoc check, to ensure new fixes don't break past successes. | Evidence: Feizi says the naive approach optimizes only on the new learning environment E(k+1), risking regression on k past environments. RELi's approach is 'regression-aware learning': fix new failures subject to no regression on past environments, and do this efficiently (sub-linearly with k). He built a benchmark with 'regression traps' to test this—optimizers that overfit on the latest fix fail on prior tasks. Demo shows 10% average improvement (78% → 97%) on the new task without breaking past ones. | Caveat: Feizi does not specify the algorithm or efficiency mechanism (e.g., does RELi use sampling, constraint optimization, multi-objective RL?). He also doesn't discuss how the system prioritizes when a new fix improves the new task but slightly degrades an old one (trade-off scenario). It's unclear how many past environments are actively tested per optimization loop. | Implication: For Ken: this is the core technical innovation and the main reason to watch. Without regression awareness in the loop, continual learning degrades into catastrophic forgetting. For product teams, this means every agent update should ship with a regression report, not just a 'new task fixed' claim. For investing/GTM, this is a key differentiator for agent platforms—ask vendors if their learning systems are regression-aware by design or only post-hoc. | Timestamp: 12:00
  • Claim: RELi's system lifts signals (logs, feedback, instructions) into learning environments automatically, then outputs reviewable PRs explaining what changed in the agent and why. | Evidence: Feizi demos creating a learning environment from a single instruction ('caller is rude and adversarial') → system generates personas, intent, mock/real tools, and evaluators. After optimization, RELi outputs a PR with the agent changes. Example: 'keep fast eligible refunds but do not generalize generosity beyond refund thresholds' (feedback on a log) → lifted to environment → optimized → reviewable update. | Caveat: The demo is on a fictional support agent benchmark with a 'single source of truth' for policies and deterministic evaluators—production agents often lack clean ground truth or deterministic tool behavior. It's unclear how much human review is needed for the PR or whether the system can produce incorrect/unsafe changes that look plausible. | Implication: For Ken's AI ops angle: the 'reviewable PR' is the UX hook—engineers can treat agent improvements like code reviews. For product teams, this means the learning loop is auditable and explainable, which matters for compliance and trust. For GTM, the two-line integration claim (setup + optimize) is a strong adoption wedge if true—Ken should test whether this holds for custom agents or only RELi-native harnesses. | Timestamp: 16:30
  • Claim: The four principles of practical verifiable continual learning are replayability, holisticness, lifelongness, and efficiency. | Evidence: Feizi defines each: (1) Replayability—turn one-off failures into rerunnable tests; (2) Holisticness—one failure may have multiple causes/repairs, route to the right layer; (3) Lifelongness—new fixes must not break past successes, regression-aware by design; (4) Efficiency—the loop runs frequently, updates are cheap where possible, and optimization scales sub-linearly with k past environments. He says these are the foundation of RELi's system. | Caveat: These are design principles, not provable guarantees. Feizi doesn't provide failure modes (e.g., when replayability is infeasible, when holistic analysis picks the wrong layer, when efficiency trades off with correctness). He also doesn't compare to alternative frameworks (e.g., Anthropic's constitutional AI, OpenAI's RLHF, Meta's ReAct variants). | Implication: For Ken: use these four as a checklist when evaluating any agent learning system. If a vendor claims continual learning but can't demonstrate replayability or regression-awareness, it's likely just fine-tuning with extra steps. For product/ops, these principles translate to operational requirements: can you replay, do you know which layer to change, do you preserve past behavior, and can you run this loop daily/weekly? | Timestamp: 10:00

Detailed Brief

The Two Fundamental Challenges of Continual Learning for Agents

  • Claims: Challenge 1: How to get feedback—how do we know if the agent did well, and if not, what should it have done?; Challenge 2: How to act on that feedback—which layer/component to change, and how to optimize without forgetting?; Easy case: benchmarks with evaluators during development. Hard case: production logs without explicit feedback.
  • Evidence: Feizi contrasts development time (curated benchmarks, explicit evaluators) with production (logs, no explicit feedback).; Two approaches to get feedback on logs: (1) automatic (models/LLMs analyze logs, agent self-critiques—scalable); (2) human experts (domain feedback on select logs—low volume but critical).; Even with logs + feedback, it's not enough—must lift to replayable learning environments with simulation and evaluators.
  • Caveats: Automatic feedback can be noisy or misaligned. Human feedback is expensive and doesn't scale.; Lifting logs to environments is an inference problem (inferring a distribution from one observation)—no discussion of when this fails or requires manual intervention.
  • Implications: Ken's takeaway: feedback is a pipeline problem, not a data annotation problem. Teams need tooling to convert logs → environments, not just logs → labels.; For AI ops, this means building feedback loops that are executable/testable, not just logged/monitored.

Three Layers of Agent Optimization: Model, Harness, Memory

  • Claims: Model layer: change weights via SFT, DPO, GRPO, RLHF (expensive, needs benchmarks, explicit evaluators).; Harness layer: rewrite prompts, learn skills, change tools/code. Methods: GEPA (prompt search via evolutionary algorithms), trace-to-harness (coding agent edits). GEPA is testable but needs benchmarks; trace-to-harness works on logs but is white-box, unverified.; Memory layer: write facts, distill skills. Methods: Letta, Mem0 (cheapest/fastest, works on logs but unverified).; Good learning asks for the smallest durable change at the right layer.
  • Evidence: Feizi provides method examples: SFT imitates correct trajectories, RL samples/scores against reward, LoRA limits parameter changes.; Example of holistic analysis: agent uses stale policy and skips escalation—cause could be memory (stale fact), prompt (not optimized), tool (doesn't normalize policy), workflow (missing escalation gate), or model (weak reasoning).
  • Caveats: Trace-to-harness is white-box and untestable—no guarantee the change works even for that sample, let alone others.; Memory writes are unverified—writing a fact doesn't prove it resolves the issue or doesn't cause regressions.; Model fine-tuning is expensive and requires benchmarks—not practical for every production failure.
  • Implications: For Ken: the layer taxonomy is useful for reasoning about agent costs and risks. Memory is fast but unverified; harness is flexible but fragile; model is durable but expensive.; For product/ops, this implies a decision tree: try memory first, harness second, model last. For investing, ask vendors if they optimize at one layer or orchestrate across all three.

Verifiable Continual Learning (VCL) Framework and RELi's Implementation

  • Claims: VCL goal: improve agent from experience where every fix is proven to help and proven to break nothing that already worked.; Three steps: (1) executable test (failure → replayable task); (2) measure delta (score before/after); (3) regression test (prior tests still pass).; Four principles: replayability, holisticness, lifelongness, efficiency.; RELi's loop: signals (logs/feedback/instructions) → lift to learning environments → root-cause analysis → route to layer → regression-aware optimize → reviewable PR.
  • Evidence: Feizi demo: fictional support agent benchmark with regression traps. One loop improves 10% on average (78% → 97%). System works on instructions ('caller is rude') or logs with feedback ('keep fast refunds, don't generalize generosity').; Claims two-line integration: (1) one-time setup (create learning harness, use own LLM, works on major agent frameworks); (2) create environments + call RelyOptimize.; Output is a PR explaining what changed and why, preserving past successes.
  • Caveats: Demo is on a synthetic benchmark with deterministic evaluators and a single source of truth—production agents often lack this. Real-world tool behavior, user variability, and policy drift are not addressed.; No discussion of false positives (system thinks it fixed something but didn't), false negatives (system misses a fix), or how it handles conflicting feedback.; Efficiency claim (sub-linear scaling with k past environments) is stated but not proven or measured.; Pricing, data requirements, and integration complexity are not discussed.
  • Implications: For Ken: this is the operational payoff. If RELi's claims hold, it's a 10x workflow improvement over current agent tuning (manual prompt tweaking, ad-hoc fine-tuning, no regression testing).; For product/ops, the key question is whether the two-line integration is real or requires heavy customization. Ken should test this on a real agent to validate.; For investing/GTM, the 'reviewable PR' UX is a strong adoption wedge for engineering teams used to code review workflows. The regression-aware optimization is the technical moat.

Continual Learning Benchmark and Regression Traps

  • Claims: Feizi built a reproducible testbed for continual learning on a fictional support agent with deterministic evaluators and regression traps.; Regression traps: if optimizer overfits on latest fix, it breaks what worked on prior tasks.; Used to validate that RELi's system preserves past successes while fixing new failures.
  • Evidence: Demo shows agent scoring 78% on 'rude/adversarial caller' scenario, improved to 97% in one loop without breaking prior tasks.; Benchmark includes policies, tool behavior, and evaluators that define success.
  • Caveats: Fictional benchmark with a single source of truth—production agents face ambiguity, incomplete policies, and evolving tools.; No public benchmark or reproducibility details provided—can't verify claims independently.
  • Implications: For Ken: the benchmark validates the framework's internal consistency but not real-world robustness. Ken should ask RELi for customer case studies or production metrics.; For product/ops, the regression trap concept is valuable—teams should build their own regression tests, even if not using RELi.

Notable Concepts & Terms

  • Verifiable Continual Learning (VCL): Framework where every agent improvement is testable, measurable, and regression-proof. Feizi's core contribution and RELi's product positioning. Distinguishes from naive fine-tuning or prompt iteration.
  • Replayability principle: Turn one-off failures into executable tests. Lift logs to learning environments with simulation and evaluators so you can rerun and measure improvements. Without this, you can't verify fixes.
  • Holisticness principle: One failure may have multiple causes (memory, harness, model). Root-cause analysis should route the fix to the right layer with the smallest durable change, not default to model fine-tuning.
  • Lifelongness principle: Regression-aware optimization: new fixes must preserve past successes. Not post-hoc checking, but integrated into the optimization loop (e.g., constrained optimization on k past environments).
  • Efficiency principle: The continual learning loop runs frequently, and optimization scales sub-linearly with the number of past environments. Updates should be cheap where possible (memory first, model last).
  • Learning environment: A replayable simulation and evaluation setup inferred from logs/feedback. Includes how tools behave (mock or real, what data), synthetic users, and evaluators that define success. Not the same as a log or benchmark.
  • Regression traps: Test scenarios where optimizing for a new task breaks past successes. Feizi built these into his benchmark to validate that RELi's system doesn't catastrophically forget.
  • Trace-to-harness: Method where a coding agent analyzes a log/feedback and edits the agent harness (prompts, tools, code). White-box, works on logs, but unverified (no guarantee the change works or doesn't regress).
  • GEPA: Prompt search method using evolutionary algorithms to mutate prompts, score candidates, and keep winners. Testable but requires benchmarks and explicit evaluators.
  • SFT, DPO, GRPO, RLHF: Model fine-tuning methods. SFT (supervised fine-tuning) imitates correct trajectories. DPO, GRPO, RLHF are RL-based post-training that sample/score against reward or preference signals. Expensive, need benchmarks.
  • Letta, Mem0: Memory layer methods for storing facts or skills in agent memory (session or persistent). Cheapest/fastest updates but unverified (no test that writing resolves the issue or avoids regressions).
  • LoRA (Low-Rank Adaptation): Fine-tuning approach that limits the set of parameters that can change, making model updates cheaper and safer. Feizi mentions it as a variant of model-layer optimization.

Operator Notes / Why Ken Should Care

  • For AI ops: The 'log → learning environment' pipeline is the critical infrastructure. Without replayability, continual learning is guesswork. Ken should prioritize tooling that turns logs into executable tests, not just dashboards.
  • For agentic systems: The three-layer taxonomy (memory/harness/model) is a useful mental model for debugging and optimization. Most teams over-index on model fine-tuning when harness/memory fixes are faster and sufficient.
  • For product/business: The 'reviewable PR' UX is the adoption wedge—engineers can treat agent improvements like code reviews. This makes learning auditable and explainable, which matters for compliance and trust.
  • For investing: Regression-aware optimization is the technical moat. Most agent platforms lack this, making them prone to catastrophic forgetting. Ken should ask vendors if their learning systems are regression-aware by design or post-hoc.
  • For GTM: RELi's two-line integration claim is strong if true. Ken should test whether it holds for custom agents or only RELi-native harnesses. The fictional benchmark demo is compelling but needs production validation.
  • For workflow: This talk argues that agent improvement should be a continuous, automated loop (signals → environments → optimize → PR), not a quarterly fine-tuning project. Ken should think about how his agent systems can adopt this loop.

Watch Map

  • 00:00: Intro: continual learning for agents, from failures to durable improvements. Humans learn from experience; agents should too.
  • 01:20: Big picture: agent interacts with world/users/tools/data policies, improves across model/harness/memory layers.
  • 02:40: Challenge 1: How to get feedback? Easy case (benchmarks) vs. hard case (production logs). Automatic vs. human feedback. Logs ≠ learning environments.
  • 04:10: What is a learning environment? Infer simulation/evaluation from one observation. Output is executable, testable, replayable.
  • 05:20: Challenge 2: How to act on feedback? Three layers: model (expensive, SFT/RL), harness (GEPA, trace-to-harness), memory (cheap, unverified). Ask for smallest durable change.
  • 06:30: Deep dive: model layer (SFT, DPO, GRPO, RLHF, LoRA). Need benchmarks. Can't apply directly to logs unless lifted to environments.
  • 07:30: Deep dive: harness layer. Trace-to-harness (white-box, untestable). GEPA (testable, needs benchmarks). Memory layer (write facts, distill skills, unverified).
  • 08:40: Introduce Verifiable Continual Learning (VCL). Goal: every fix proven to help, proven to break nothing. Three steps: executable test, measure delta, regression test.
  • 10:00: Four principles of practical VCL: replayability, holisticness, lifelongness, efficiency. Detailed explanation of each.
  • 12:00: Lifelongness principle: regression-aware optimization. Fix new failure subject to no regression on k past environments. Efficient (sub-linear with k).
  • 13:30: RELi's learning loop: signals → lift to environments → root-cause analysis → route to layer → regression-aware optimize → reviewable PR.
  • 14:40: Add VCL to your agent in two lines: one-time setup, then create environments + call optimize. Output is PR explaining changes.
  • 15:20: Demo: continual learning benchmark on fictional support agent with regression traps. Deterministic evaluators.
  • 16:30: Demo: create environment from instruction ('rude caller'). Simulate agent, score 78%. Call optimize, improve to 97% (10% avg gain) without breaking past.
  • 18:00: Demo: production log with feedback ('keep fast refunds, don't generalize'). Lift to environment, optimize, generate PR. Lifelong and compounding.
  • 19:10: Three key takeaways: (1) continual learning ≠ model fine-tuning; (2) logs ≠ learning environments; (3) frontier is regression-aware improvement.
  • 20:00: Conclusion: VCL built on four principles. Try it at rely.ai. Q&A.

Source/Metadata

  • Title: Continual Learning for AI Agents: From Failures to Durable Improvements - Soheil Feizi, RELAI
  • Transcript words: 4033
  • Duration seconds: 1355
  • Timestamp note: Timestamps manually mapped from transcript structure; some are approximate due to repetition in transcript.

Transcript

3152 words en Processed in 220.5s

Hi everyone, my name is Sohel Feizzi. I'm founder and CSO at RELi. I'm also an associate professor in the computer science department at University of Maryland. Today I'm going to talk about continual learning for AI agents. How we can go from failures to durable improvements. And if you're interested in any of the tools that I'll be talking in this presentation, you can visit our website RELi.ai. Let's get started. Humans learn mainly from experience by interacting with the world and getting feedback. The goal of continual learning is to imitate the same for agents so they can also learn from experience by acting, getting feedback and improving without forgetting. So here's a bigger picture of how continual learning for agents looks like. An agent interacts with the world, with diverse users, with complex tools, with various data policies. And as I mentioned, the goal is to continuously improve the agent from its experience without forgetting. And this learning can happen in different layers of the agent. It can happen in the model layer where potentially we can change weights of LLMs or other models used in the agent or use different types of models in the agent. It can happen in the harness layer where it brings the proper context to the LLM with components like prompts, skills, tools, code, workflow. And it can also happen in the memory, either in session memory or persistent memory of the agent. So there are two fundamental challenges in continual learning for agents. The first challenge is how to get feedback. How do we know if the agent did well? And if not, what should it have done instead? That's the first part, getting the feedback. And the second part is how we can act upon that feedback, how we can optimize and improve the agent and learn from that feedback. Which layer, which component do we need to change? And also how? I'll be talking about these two challenges, current approaches in order to deal with them and also provide some perspective of how we think about these two problems. So let's get started with the first problem. Where does the feedback come from? The easy case is when we have a benchmark and some evaluators on that benchmark. The agent can run tasks from the benchmark. Now we have the evaluators in order to score and we can get grades like pass, fail or reward, as well as potentially some feedback on the agent behavior and agent performance. This is usually what is happening during development time where different teams curate benchmarks in order to understand the performance of the agent in certain applications. But in production, we don't have such benchmarks. We have logs. Here's an example of a session log where a user is interacting with the agent. Maybe the user is not very happy with the way the agent is behaving, but we don't have any explicit feedback. So there are two ways of getting such feedback on such session logs. One is automatic using some other models or LLMs or code in order to analyze the log and provide feedback. In some cases, even the agent itself can look at its log and provide some criticism of it. It is automatic and scalable. And the second approach is where we have human experts to look at some handful of these logs and provide domain expert feedback on those agent outputs. This is lowering the volume, but it is critical because it provides expert knowledge on the behavior of the agent and its alignment with the way that we want the agent to behave in those applications. Either way, now we have session log plus some feedback on those logs. Is it enough? The answer is no, because it is still not testable. Here we have log and feedback, but what we really need is a replayable learning environment, a simulation that we can rerun with defined grading on what success looks like, not one instance of what happened and the feedback on top of it. So what is a learning environment? Here we are inferring a distribution from one observation that replaces what happened and what success means. The input is what we have, some session logs and feedback. This is one observation of what happened. And now we want to create a simulation and evaluation environment from that information. That involves, for example, understanding how tools in the agent behavior in the agent log should behave. Should we use real tools? Should we use mock tools? And if so, what kind of data should be brought to the mocking process of those tools? If the agent is interacting with some users, how we can infer synthetic users from that data, and also how success looks like in that learning environment, what are the evaluators that need to be inferred? So there are lots of technical challenges in any of these components. But if we could do this successfully, the good news is the output is executable. We can now run different candidates of the agents against such learning environments, understand the behavior and the performance of the agents in those scenarios and in those patterns. And we can fix the issues based on the feedback based on the information that we observe because not everything becomes testable and verifiable. All right, so the second problem is now we have this feedback, how we can act upon it, how we can optimize the agent. And from a high level point of view, there are three layers that we can improve the agent. There's a model layer where we can change the weights of the model and there are methods like SFT, supervised fine tuning, RL based post training in order to make those changes. And these are usually expensive because that requires more intensive compute in terms of changing the model weights. We can change the model weights. The second layer is the harness layer, harness engineering, context engineering, where we can potentially rewrite prompts, maybe learn some skills, change tools or add tools, maybe change code around the LLM. And there are different methods like GEPA, trace to harness, which provides a lot of flexibility in terms of learning from that feedback. And the last layer is the memory layer where we store facts and learn skills in order to not repeat those issues and failures in the future. But good learning is not going to be focusing on any of these components exclusively. A good learning engine should ask for the smallest durable change at the right layer of the agent. All right. So let me dig a little bit deeper into each of these layers. First, in terms of updating the model weights, there are various approaches in order to do that, including SFT supervised fine tuning, where we imitate correct trajectories. We often need labeled samples in order to fit the model to those samples. Other approaches are based on RL reinforcement learning post-training, like DPO, GRPO, RLHF, where we sample and score against reward or preference signals and reinforce what wins. And there are some categories based on LoRA and adaptation that limits the set of parameters that can potentially change. It makes learning in this layer cheaper and also safer in terms of the updates. But these methods, they usually need benchmarks and explicit evaluators. They cannot be directly applied on, for instance, if you have a log and feedback, unless we turn those into replayable learning environments. So let's look at the next step. We often need labeled samples in order to fit those samples, fit the model to those samples. Other approaches are based on RL reinforcement learning, post-training, like DPO, GRPO, RLVR, where we sample the score against the reward or preference signals and reinforce what wins. And there are some categories based on LoRa, LoRa, LoRa, and adaptation that limits the set of parameters that can potentially change. It makes this, the learning in this layer cheaper and also safer in terms of the updates. But these methods, they usually need benchmarks and explicit evaluators. They cannot be directly applied on, let's say, if you have a log and feedback, unless we turn those into replayable learning environments. So let's look at the next step. In terms of updating the harness, I would highlight the two categories in this domain, in this layer. One is trace to harness approaches. Let's say you observe a log, you have some feedback on top of it, you can effectively ask a coding agent in order to analyze the log and improve the agent. So this works on the case where we have log and feedback, but it is white box based. We don't know if even for that particular sample, if the change is effective, because it is not testable. And we don't know what is the impact of it on other samples and other scenarios, what might have been working previously, but these changes might not work properly and create some hidden regressions. The other category that I want to highlight is methods like GEPA and PROMPT SEARCH, where they mutate prompts, they score different candidates and they keep the winners using some search algorithms like evolutionary algorithms. These methods are testable, but they need benchmarks and explicit evaluators in order to have those scorings. And in the memory layer, we effectively write down facts and distill skills so the agent doesn't rediscover them. It can happen in the information memory layer, methods like Letta and Mem0, where they can effectively store a fact or correction. And we have also methods through skill distillation that compresses successful trajectory into reusable how-to scale for the agent. So this layer in terms of the update is cheapest and fastest. It works directly on the cases where you only have log and feedback, but usually it is unverified because you don't have a way in order to test whether or not writing in the memory will resolve the issues that you have dealt with and whether or not it can potentially create some regressions on some other cases. With that, let me introduce a new subcategory of continual learning called verifiable continual learning. In a verifiable continual learning, the goal is to improve an agent from its own experience where every fix is proven to help and proven to break nothing that already worked. And usually it involves three steps where we need to have an executable test where the failure becomes a task you can replay and create. Then we need to have a measure delta where the update is scored on the test before and after. And then we have a regression test. So prior tests still pass even after we make such changes to the agent. So let's think about what are the principles of a practical verifiable continual learning first. And I will argue there are four important principles that we need to keep in mind. So the first principle is replayability. We need to turn a one-off failure into a test that we can rerun. Here, as I had mentioned previously, many cases we have log and feedback, but that is not testable. We need to lift it in a learning environment to simulate and evaluate the agent on a similar pattern, on a similar scenario. So everything becomes testable based on that simulation and evaluation environment. So that's the first principle that we need to have. The second principle is holisticness. One failure may have several causes and several possible repairs. Let me give you an example. Let's say you have an agent that uses a stale policy and skips the required escalation. The issue might come from the memory where you have some stale fact. It can come from not optimized prompt. It might come from a tool that doesn't normalize the policy. It might come from the workflow that we need to add escalation gate before we find it. It might come from the model. Maybe the model is not a good model in order to have a strong reasoning. So here we need to route the fix to the right layers that explains the failure with the smallest durable change to the agent. And that is the principle of holisticness in verifiable continual learning. The third principle is lifelongness. A new fix must improve the new case without breaking the past. Next, let's consider this setup where we already optimize the agent on K past learning environments. And a new failure comes and we turn that into a learning environment, E K plus one. What do you want to change? So the first approach is, okay, just focus on this new learning environment. But that can create regression on the past behavior on the past learning environments that the agent was successful. A better approach is a regression-aware learning where the regression is not treated as a post-hoc approach, but as a mechanism within the optimization itself. So here we are fixing the recent failures subject to having no regression of the past learning environments. And obviously this needs to be done in an efficient manner. So it doesn't scale even linearly with K because K can grow and the complexity of this approach can be very high. And the last but not least principle is efficiency. This continual learning loop needs to run frequently. And we need to have efficiency in different layers in updates to the agent. So sometimes the change can be cheap, like writing something in the memory can be medium in terms of the complexity by changing the prompt or harness. And sometimes it can be very expensive by changing the weight of the model. Also efficiency should be in the optimization loop itself, especially when we have regression-aware optimization. And regression is treated within the loop, not as a post-hoc approach. To sum up, these are the four principles of a practical verifiable continual learning. Replayability, holisticness, lifelongness, and efficiency. And this is what we have been working on at Rely to create a verifiable continual learning engine for AI agents based on these four principles. In particular, here's how Rely's learning loop runs. You can start with some signals to this loop. It can be logs, feedback, or even instructions and prompts. We lift those signals to replayable learning environments. So that's based on the replayability principle. This makes everything that follows testable and verifiable. Then we do root cause analysis and route the fixes to the right layer of the agent. It can be memory, it can be model, or it can be harness. So that touches the holisticness principle that I described. We have regression-aware optimization. Regression is not being treated as a post-hoc approach. So that touches the lifelongness principle that I mentioned. And obviously, this loop should run efficiently. That touches the efficiency principle. So the output of this is a reviewable version update to the agent, explaining what changes in the agent during this loop and why those changes are improving the agent without creating regression. So the beauty of it is you can add a VCL, verifiable continual learning, to your current agent in just two comments. So the first one is a one-time setup. So that touches the holisticness principle that I described. We have regression-aware optimization. Regression is not being treated as a post-hoc approach. So that touches the lifelongness principle that I mentioned. And obviously, this loop should run efficiently. That touches the efficiency principle. So the output of this is a reviewable version update to the agent, explaining what changes in the agent during this loop and why those changes are improving the agent without creating regression. So the beauty of it is you can add a VCL, verifiable continual learning, to your current agent in just two comments. So the first one is a one-time setup. You create a learning harness in your agent. You can use your own LLM and your agent can be built on top of any of available major agent frameworks. And then after that, you need two commands in order to activate this learning group. You can create learning environments using various type signals that you can have either log feedback or some instructions. And then you can call RelyOptimize in order to use holistic lifelong optimizer to improve the agent. And the output is an optimized version pull request that you can review and you can use it in order to improve your agent. So let's look at how it actually works in practice. We build a continual learning benchmark on a fictional support agent case, where we have reproducible testbeds for continual learning in a tool using support agent. So we have a single source of truth and the policies are interacting for this agent to be handling. So we have deterministic evaluators and we also build this benchmark in a way that it has some regression traps. So if the optimizer focuses on overfitting on the latest fix, it can potentially break what the agent was previously successful on other tasks. So let's say we have an agent. We don't even have logs or anything. And we want to just see how the agent is behaving. Let's say when a caller is rude and adversarial. So simply we can create a learning environment using such instruction. And what it will do, it will create a learning environment to simulate and evaluate the agent. The simulator will include personas, intent, mock or real tools. And also the learning environment contain evaluators in order to define success metrics. All of these are produced from just one interactive comment. And after that, we can just simulate the agent using this learning environment. See how it behaves. Okay, the score is not too high. It is 78%. And in particular, there are two evaluators that show very low scores of agent in this environment. So these are some of the failures that we observe. How to improve such failures. We can do that by calling rely optimize with certain number of rollouts. And as you can see, the average improvement can be quite high. It is 10% improvement on average, just with one loop. And the score increases to 97% from 87%. Okay. Now, let's consider the case that the agent is in production. Now you have a log, you have an agent session that is not desired and you have feedback. For example, you can say keep fast eligible refunds, but do not generalize generosity beyond refund thresholds. So that's feedback on the agent behavior. Again, the flow is the same. So we lift it into a replayable learning environment and we can rely optimize in order to mitigate this issue. Use this feedback without creating regression of the agent behavior in past environments. [SPEAKER_00] And this is lifelong. So you can keep doing that to improve the agent without breaking what already works. And it is compounding. This is verifiable continual learning in practice, where each update is tested, every gain is measured, and nothing that already works breaks during this optimization. So that's it for today and for this talk. So there are three key takeaways that I want to highlight. The first one is agent continual learning is not necessarily model fine tuning. The updates and many useful updates can happen in the harness and memory layer. So the second takeaway is production logs are not learning environments. We need to transform them into replayable learning environments to simulate and evaluate the agent on the same patterns and scenarios. And the third takeaway is that the frontier is regression-aware continual improvement. Where when fixing the new failure, we verify that we don't forget the old ones. We don't create regression. So that's verifiable continual learning built on four principles, replayability, holisticness, lifelongness, and efficiency. And if you want to try VCL and apply to your agent, you can use it today at rely.ai. Thank you. And then after that, you need two commands in order to activate this learning group. You can create learning environments using various type signals that you can have either log, log feedback or some instructions. And then you can call RelyOptimize in order to use holistic lifelong optimizer to improve the agent. And the output is an optimized version, pull request that you can review and you can use it in order to improve your agent. So let's look at how it actually works in practice. We build a continual learning benchmark on a fictional support agent case, where we have reproducible testbeds for continual learning in a tool using support agent. So we have a single source of truth and the policies are interacting for this agent to be handling. So we have deterministic evaluators and we also build this benchmark in a way that it has some regression traps. So if the optimizer focuses on overfitting on the latest fix, it can potentially break what the agent was previously successful on other tasks. So let's say we have an agent. We don't even have logs or anything. And we want to just see how the agent is behaving. Let's say when a caller is rude and adversarial. So simply we can create a learning environment using such instruction. And what it will do, it will create a learning environment to simulate and evaluate the agent. The simulator will include personas, intent, mock or real tools. And also the learning environment contain evaluators in order to define success metrics. All of these are produced from just one interactive comment. And after that, we can just simulate the agent using this learning environment. See how it behaves. Okay, the score is not too high. It is 78%. And in particular, there are two evaluators that basically show very low scores of agent in this environment. So these are some of the failures that we observe. How to improve such failures. We can do that by calling rely optimize with certain number of rollouts. And as you can see, the average improvement can be quite high. It is 10% improvement on average, just with one loop. And the score increases to 97% from 87%. Okay. Now, let's consider the case that no, the agent is in production. Now you have a log, you have an agent session that is not desired and you have a feedback. For example, you can say keep fast eligible refunds, but do not generalize generosity beyond refund thresholds. So that's a feedback on the agent behavior. Again, the flow is the same. So we lift it into a replayable learning environment and we can rely optimize in order to mitigate this issue. Use this feedback without creating regression of the agent behavior in past environments. And this is lifelong. So you can keep doing that to improve the agent without breaking what already works. And it is compounding. This is verifiable continual learning in practice, where each update is tested, every gain is measured, and nothing that already works breaks during this optimization. So that's it for today and for this talk. So there are three key takeaways that I want to highlight. The first one is agent continual learning is not necessarily model fine tuning. The updates and many useful updates can happen in the harness and memory layer. So the second takeaway is production logs are not learning environments. We need to transform them into replayable learning environments to simulate and evaluate the agent on the same patterns and scenarios. And the third takeaway is that the frontier is regression-aware continual improvement. Where when fixing the new failure, we verify that we don't forget the old ones. We don't create regression. So that's verifiable continual learning built on four principles, replayability, holisticness, lifelongness, and efficiency. And if you want to try or VCL and apply to your agent, you can use it today at rely.ai. Thank you.