Open Reader

Improving Agents is a Data Mining Problem — Vivek Trivedy, LangChain

completed 20:01 Aug 12, 2026 Watch on YouTube

Current Status

completed

Video ID

CvRngaQZQ3Y

RAG / Chat

Enabled
Improving Agents is a Data Mining Problem — Vivek Trivedy, LangChain
Description

Does your agent get dumber after the first compaction? After the second? You cannot read that off the code, only off the traces, and there are far too many to read yourself. So LangChain points agents at the traces of other agents and asks exactly that, alongside questions like where users got upset and what a different model would have done at the same step. Vivek Trivedy's argument is that observability and continual learning are the same problem in different clothing, because an agent acting in an environment produces the only real record of what happened, and that record is the substrate everything else is built on. The economics fall out of reading it. Working with Harvey on a legal benchmark, they found an open model could match their frontier model's trace judging at one to two orders of magnitude lower cost, arrived at through harness engineering that the traces themselves pointed to. His rule for when to stop tuning prompts and start finetuning is speed of feedback: harness engineering answers in about two minutes, so you exhaust that ceiling first, finetune to break through it, then return to harness engineering. He also argues that dense feedback is what agents lack most, since a benchmark returning only pass or fail gives an agent nothing to act on, while traces already hold the fine grained signal. The claim worth arguing with is that you can describe an agent's behavior just by showing the evals it was measured against, because those are what it hill climbs. Speaker info: - https://x.com/Vtrivedy10 - https://www.linkedin.com/in/vivek-trivedy-433509134/ - https://www.vtrivedy.com/ Timestamps: 0:00 - My agent made mistakes, now what 1:28 - Ship it, collect traces, mine them 2:44 - Observability and continual learning are the same problem 4:00 - Why agents are harder to reason about than code 4:36 - Trading determinism for autonomy 5:15 - Sending agents to read other agents' traces 6:29 - Today's data is the least we will ever have 7:09 - When a trace

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Improving production agents is fundamentally a trace-data mining loop: instrument every run, mine behavior and feedback at scale, convert findings into evals, harness changes, fine-tuning data, and persistent memory updates.
  • Why it matters: For agent systems, prompt and orchestration changes are too nondeterministic to reason about from source code alone; trace-derived evidence becomes the control plane for quality, cost, routing, and continual improvement.
  • Best use: Use this as a practical operating model for building an agent-improvement pipeline, especially the sequencing of tracing, trace mining, dense evaluation feedback, harness engineering, model distillation, and long-horizon memory maintenance.

Executive Summary

Vivek Trivedy argues that agent development should be treated less as hand-tuning prompts and more as a data-mining and model-fitting problem. Once an agent is deployed, its tool calls, outputs, API activity, failures, and user reactions form traces. Those traces are the only reliable evidence of real behavior in systems whose outcomes emerge from prompts, models, tools, middleware, multi-agent orchestration, and long contexts.

The recommended loop is: ship the agent, centralize and retain exhaustive traces, use agents to mine those traces for failure patterns and successful behavior, then run controlled experiments against trace-derived evaluations. Outputs of the mining step should feed three downstream mechanisms: curated examples for human review, new evals and environments, and datasets for distillation or fine-tuning.

He makes a pragmatic optimization case for starting with frontier models to establish task feasibility, then using trace evidence to determine whether a cheaper open model plus a better harness can match the needed behavior. When prompt/harness improvements plateau, teams can fine-tune a model on narrow, high-value vertical tasks; at sufficiently high volume, serving that model on owned or rented compute may be cheaper than continuing to pay per-token API costs.

The longer-term framing is continual learning across three layers: observational training data, evolving harnesses, and non-append-only memory. Trivedy is explicit that the exact implementation of continual learning remains unclear, but contends that agents will need offline "sleep-time" processing over their accumulated traces to revise prompts, operating state, and memory rather than merely storing ever-larger logs.

Key Takeaways

  • Claim: Trace collection is the foundation of continuous agent improvement because traces capture the behavior users actually experience, whereas code inspection cannot reliably predict emergent agent behavior. | Evidence: A production agent may combine prompts, tools, skills, hooks, middleware, APIs, CLIs, nested agents, and swarm orchestration; Trivedy contrasts this with ordinary code, whose function calls and logic can be read more directly. | Implication: Treat full-fidelity tracing as a required production capability, not merely debugging instrumentation; without it, agent changes lack an empirical improvement loop.
  • Claim: Agents should be used to mine other agents' traces, because manual review cannot scale to millions of traces or traces containing millions of tokens. | Evidence: Suggested mining questions include finding interactions where users were upset or highly satisfied, diagnosing whether quality degrades after first or second context compaction, and comparing counterfactual model choices such as GPT 5.5 versus GLM 5.2 on the same task. | Implication: Build trace mining as an agentic retrieval/query system over externally stored trace objects rather than attempting to paste complete logs into a single model context. | Caveat: Large trace corpora create both token-cost and context-window constraints: cost grows with input-token price × number of traces × average trace length, and long coding-agent traces may not fit into a reviewer model's context at all.
  • Claim: The most useful outputs of trace mining are not reports alone, but datasets that drive distillation/fine-tuning, evaluation generation, and targeted human review. | Evidence: For distillation, LangChain's example takes successful GLM 5.2 traces and uses them to fine-tune a 9B or 13B model to mimic the stronger model's behavior. For high-trust areas such as legal and medical, the system should prepare compact material that humans can review rather than asking them to inspect all raw traces. | Implication: Design the trace pipeline around explicit output contracts: a training set, an eval set/environment, and a human-review queue, each tied to a detected behavioral pattern.
  • Claim: Evaluation suites effectively define an agent's behavior because teams optimize their agents to pass those tests. | Evidence: Trivedy says that seeing the evals run against an agent would provide a rough prediction of its behavior, since the agent is repeatedly hill-climbed against them. | Implication: Mine production traces to continuously expand and rebalance evals; do not treat static benchmark scores as a sufficient representation of agent quality. | Caveat: Agents can optimize a measurable score in undesirable ways or "cheat," so score-based optimization requires checks for reward hacking and coverage gaps.
  • Claim: Dense feedback is materially more useful for automated agent improvement than binary pass/fail outcomes. | Evidence: Terminal Bench-style outcomes can reduce a complex run to a single passed/failed number. Trivedy notes that a person or agent receiving only failure after a long, unfamiliar task has little information about what to try next; traces provide the substrate for richer feedback. | Implication: Capture intermediate execution quality signals—tool outcomes, policy violations, dead ends, compaction effects, user sentiment, and task-stage failures—so automated researchers can generate productive hypotheses and fixes.
  • Claim: Use a harness-engineering-first, fine-tuning-second, harness-engineering-again sequence to improve agents quickly and economically. | Evidence: Harness changes can return feedback in roughly two minutes through eval runs. LangChain's recommended "sandwich" is to improve prompts/tools/orchestration first, fine-tune only after hitting a harness ceiling, then return to harness changes as needed. | Implication: Prioritize rapid feedback cycles and establish a measured harness plateau before committing to dataset creation, training, model hosting, and the operational overhead of fine-tuning. | Caveat: Fine-tuning is most compelling for narrow domain-specific task distributions; it is not presented as a universal replacement for a stronger general frontier model.
  • Claim: Open models can often replace frontier models for a defined task after trace-informed harness work, and narrow fine-tunes can meet or exceed frontier performance in a vertical. | Evidence: In work with Harvey's legal benchmark, Trivedy says open models roughly matched Opus's trace-judging capability at one to two orders of magnitude lower cost. The process begins with Opus or GPT 5.5 to test feasibility, then tests cheaper alternatives using observed traces and additional guidance. | Implication: Adopt model routing and distillation as an empirical cost-reduction program: establish a frontier baseline, identify the minimum intelligence needed per task, then validate cheaper models against production-derived evaluations. | Caveat: The speaker acknowledges frontier models are typically smarter and starts with them to establish the task's feasibility; the claimed parity is task- and harness-dependent rather than general model parity.

Detailed Brief

Model, harness, task fit as the operating framework

  • Claims: The speaker recasts agent improvement as a modern version of classical machine learning's fitting process: fit data, a model, and a harness to the tasks that matter.; The core research work is finding good data and good fit functions, including automated research loops, reinforcement-learning-style methods, and supervised fine-tuning.
  • Evidence: Trivedy invokes Scikit-learn as an analogy: classical ML provides helpers for fitting learning systems to data; in agent systems, the corresponding components are model, harness, and task.; The talk references automated research and methods described as OPD, OPSD, and SFT, without providing implementation details or comparative results.
  • Caveats: The talk is a conceptual operating model rather than a technical specification for a trace schema, retrieval architecture, labeling protocol, or training methodology.; No benchmark numbers are provided for the claimed fine-tuned vertical models exceeding frontier performance.
  • Implications: Agent teams should manage model choice, prompt/tool/orchestration design, task definition, and trace-derived data as jointly optimized variables rather than independent workstreams.; A model swap without re-evaluating its harness and task distribution is unlikely to reveal its true production capability or cost profile.

Economics and long-horizon continual learning

  • Claims: Fine-tuning changes the cost model from variable token spending to provisioned hardware spending.; Continual learning must eventually update more than model weights: it requires ongoing harness evolution and actively maintained memory.
  • Evidence: For high-inference-volume workloads, the speaker argues it can be cheaper to run a compute cluster, obtain effectively unlimited inference while it is active, and spin it down when not needed.; He argues agents operating over years cannot rely on an append-only memory file plus retrieval; offline processing of the full lifecycle of traces should revise the agent's state, likened to scaling "sleep-time compute" or dreaming.
  • Caveats: The crossover point at which dedicated compute is cheaper is workload-specific and no utilization threshold, hardware configuration, or total-cost model is supplied.; Trivedy says the practical form of continual learning is still unclear.
  • Implications: Cost governance should compare end-to-end serving costs, including utilization and operational burden, rather than comparing API token prices alone.; Long-lived agent products need a governed memory lifecycle—consolidation, correction, expiration, and state revision—not only retrieval over growing historical logs.

Notable Concepts & Terms

  • Trace mining: Using models or agents to search, summarize, compare, and classify production execution traces in order to uncover failures, successful patterns, and training/evaluation data.
  • Harness engineering: Improving the operational wrapper around a model—prompts, tools, orchestration, loops, guidance, hooks, and related logic—to raise task performance without changing model weights.
  • Model, harness, task fit: The speaker's adaptation of classical ML fitting: production quality depends on jointly matching the model, its operating harness, and the real task distribution.
  • Distillation / SFT: Using high-quality traces from a stronger model to create a supervised fine-tuning dataset for a smaller, cheaper model that mimics the required behavior.
  • Dense feedback: Rich execution-level signals explaining why an agent succeeded or failed, as opposed to a sparse binary reward; it gives automated improvement loops useful directional information.
  • Counterfactual model comparison: Re-running or evaluating the same trace-derived task set across model alternatives to determine whether a cheaper model preserves the behavior that matters.
  • Sleep-time compute: Offline processing of accumulated traces after agent runs to consolidate memory and update prompts or agent state, analogous to reflection or dreaming.
  • LangSmith Engine: LangChain's product framing for an agentic trace-mining system that searches trace data, identifies issues, and prepares data for evals, fine-tuning, or human review.

Operator Notes / Why Ken Should Care

  • Require structured, queryable traces for every production agent run, including model inputs/outputs, tool calls and results, compaction events, task outcome, latency/cost, and available user feedback.
  • Stand up a recurring trace-mining job that produces three artifacts per cycle: prioritized failure clusters, candidate eval cases with expected outcomes, and reviewed positive traces suitable for distillation.
  • Add explicit monitoring around context compaction: compare quality, task completion, and error patterns before and after each compaction event rather than assuming long-context degradation is benign.
  • Create a model-routing evaluation protocol: use a frontier baseline for feasibility, then test open/smaller candidates against production-derived evals before moving any traffic.
  • Set a decision gate for fine-tuning: proceed only after fast harness iterations have plateaued on the target task distribution and expected inference volume supports the training and hosting overhead.
  • Define memory governance before deploying long-lived agents: specify what gets consolidated, corrected, expired, promoted to durable state, and escalated for human review.

Source/Metadata

  • Title: Improving Agents is a Data Mining Problem — Vivek Trivedy, LangChain
  • Transcript words: 6608
  • Duration seconds: 1201
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion substantially repeats earlier material.

Transcript

3569 words en Processed in 124.0s

. Hey everyone, I'm Vivek, and I lead applied research at Langchain, and I'm going to talk about something that I think is sexy, which is data mining. It's not as sexy as LLM, so we're going to try to make it sexy together. And the problem that we're going to talk about today is how do we continuously improve agents, but how do we do that via data? So to start, I'm going to tell a little story that I think maybe a lot of us have felt before. I ran my agent, it did a bunch of things, it made some mistakes, now I ask someone, what do I actually do about that? I have all this data, I made some mistakes, what now? What we're going to do today is motivate a recipe for what we should do to continuously improve agents over time, and then I'm going to talk from some lived experience and some stuff that we help customers do to run this over large-scale trace data. So the first step in building a successful agent is shipping it. So if you put it out into the real world, then it can operate in environments, and then you can get feedback from what it's doing. The second step is collect a ton of traces. So agents operate in the environment. Every single time they operate, they do tool calls, they have output messages, they call APIs, they use CLIs. All of that generates data, and we want to store all of that so we can do stuff with it. The next thing is the data mining in this talk, which is once we have tons of trace data, maybe gigabytes, maybe terabytes, depending on how many agents you're shipping, we're going to do data mining over that. And I promise I will tell you exactly what data mining we're going to do, but we're going to do some over it. And then the fun part, which is I collected that data, I read it, I curated it, and now we actually need to run the experiments in a data-driven way to see, hey, is this new prompt or is this new tool or is this new orchestration or is this new loop, is it actually improving things based on the previous traces that I've seen? And this is maybe a bit of a hot take, but continued learning is super hot right now. I'm talking about it, this whole room is going to hear about it for the next five, six hours. But there's a very tight coupling between what observability is and what continual learning is. And the main reason for that is that agents that operate in environments, they produce trace data, and what continual learning for agents and continual learning for humans is, is I do a bunch of stuff in the world, I think about what I did, and then I need to update my definition, my knowledge, stuff I write down, in order to respond to the feedback from the environment. And if you're a continual learning company, you need traces, and if you have traces, then you can try to do continual learning over your agents. I had to put in a meme because if you look at your data, then you can be like wheel hunting, if anyone's seen the movie, where everything is super, super easy, and you can improve over time. And I promised Hamill I would put this in there, so putting it in there. Cool. So why am I talking a bunch about traces anyway? So I'm sure a ton of us were software engineers before, we're software engineers now. And on the left we have a code block, and we can read the code. And in my head, I can almost reason over what this code does. I can see the functions, I can see how they call each other, I can roughly understand the logic in Python. That doesn't exactly exist in agent world, because agents have prompts, they have tools, they have skills, they have hooks, they have middlewares. Some agents call other agents, and I orchestrate them in swarms. It's really, really hard for humans to reason about how certain prompts that they change are actually going to affect agent behavior at scale. And this also varies between the different domains that you're doing it on. So a prompt change that you're using for the medical domain is going to be completely different than a prompt change that you want to do for the law domain. And in general, over the last four years, since the ChatGPT moment, we've started trading determinism for autonomy. And in that shift, what we need to do is create tools and create systems to still understand agents when they're autonomously operating in environments. So I talked about traces. Why should you read them? And at Langchain, what do we actually do when we're reading traces? So we centralize a bunch of our data. So we put everything in a tracing project, and this is usually either per agent or centralized across all of our agents. And then what we do is we send agents to read traces from other agents. Right? And then we look for a bunch of different things. And we might ask, hey, find a bunch of good and bad interactions where users got upset, or users were really happy. Another question I might ask is, this is a technical question, agents now run for millions of tokens. Does the agent get really dumb after the first compaction, after the second compaction? Does it never get dumb? How do we actually answer those questions? We need to do it by actually looking at the traces. And then the other thing is, if I look at the traces, then I can try to prove some counterfactuals, which is, hey, I ran GPT 5.5 for this, and I heard GLM is really good. What happens if I run GLM 5.2 for this task, and how do I compare them? Metrics, awesome. The trace level captures the actual behavior that users see, so that's also very helpful for seeing behavior at fine-grained scales. And the way that we think about the data that's being generated by agents is that the data that we see today is going to be the smallest that humans have ever seen in their entire lives, because we're in this massive exponential shift to how agents are doing more and more work in the economy. And what that means is the amount of data that humans have produced in our entire lifetime will soon be eclipsed by agents running on year scales, and then six-month scales, then three-month scales, and then maybe every day, right? And to understand a ton of that data, roughly what we need to do is contend with a couple problems. There's more, but these are the two that I'm going to focus on. So, one, reading traces at scale is super expensive, especially if you have millions of traces and if you have millions of tokens per trace, right? Think of it as an input token cost. You can literally multiply the input token cost times the number of traces times how big each trace is on average, right? The other thing is, if I have a super long interaction with a coding agent, like Cloud Code, or Codex, or deep agents, I can't even read that trace with another agent because that context doesn't fit in memory, right? So, we need to develop systems so I can treat that context as an external object, and then I can query into it, right? So, we need to build agents to efficiently mine data from other agents, and it's no longer as simple as just feeding the data into context, and there's tricks that we'll talk about to do that well. Great. So, one of the things that I think is really, really cool in the last six months is that open models have hit an inflection point in intelligence that we at Langchain don't reach for the frontier models for every single use case. We're quite conscious about what is the minimum level of intelligence that I need to do any given task. And practically speaking, honestly, yes, we start with Opus, we start with 5.5, because we just want to know if the task is even possible, but then once we reach that waterline, then we look back at those traces and we see, hey, can we use an open model to do the same thing? So, this is a bunch of work that we did with Harvey and their lab legal benchmark. What we're looking at is can I match the trace judging capability of Opus with an open, cheaper model? And the answer is roughly yes, at an order or two orders of magnitude cheaper. And the way we do that is we try a bunch of models, we do a bunch of harness engineering, and the harness engineering is informed by a bunch of the traces that we read. So, it's like, hey, Opus reasons about things in this way. Maybe that's because of the prompt, maybe Opus is just smarter, which it is, than a bunch of the open models, but that might mean I need to give it a little bit more guidance so it can reach the same intelligence level at a much lower cost. And the other thing that we look at is harness engineering is amazing. You get instant feedback and you can run on your evals, but eventually what we find is you hit a threshold of intelligence where it's like, if I keep tweaking this prompt, I'm not going to get too much more out of it. And the way we do that is we try a bunch of models, we do a bunch of harness engineering, and the harness engineering is informed by a bunch of the traces that we read. So, it's hey, Opus reasons about things in this way. Maybe that's because of the prompt, maybe Opus is just smarter, which it is, than a bunch of the open models, but that might mean I need to give it a little bit more guidance so it can reach the same intelligence level at a much lower cost. And the other thing that we look at is harness engineering is amazing. You get instant feedback and you can run on your evals, but eventually what we find is you hit a threshold of intelligence where it's, if I keep tweaking this prompt, I'm not going to get too much more out of it. And once we reach that point, we look at, okay, can I actually fine-tune the model on my domain-specific task? And can I make it better on those tasks? And what we find is if we did base models and we tuned them on very specific vertical tasks, which is what a lot of our customers do, they don't really care about the entire variance of tasks, they care about what their customers care about. So if we focus on that narrow set of tasks, then we can fine-tune base models to reach and then also go beyond frontier performance. And I think one small thing I'll mention, as a lot of people are getting into fine-tuning, is that another economic decision is that you can move from token costs to hardware costs. And this can be a really big change, right? Because you're very used to, hey, a million tokens cost this much, not as much this cluster costs this much. But for very high inference workloads, we find it to be way cheaper just to run a cluster and I get unlimited inference on that cluster. I don't have to worry about tokens, but I can just do the calculation of, hey, this will end up being cheaper, and then I can spin it down when I don't need it. Cool. And I said all of this, so we obviously built a product to do that. I won't show it too much, but it's Langsmith Engine. This product is trying to automate this loop for you, which is if you have any volume of trace data and you're looking for something in that trace data or you want to generate evals from that trace data or you want to generate feedback for humans to read from that trace data, it will go read all of it, it will find issues, it will agentically search over it, and then prepare data sets for you to do something after and a bit of a leader. What that something basically is, is the outputs of this trace mining exercise. So there's three things that I mentioned here, which we see a bunch and we put into the product. So one is distillation and fine-tuning, which is let's say I'm running GLM 5.2, it's doing great, but I think that I can run this task way cheaper with a 9B or 13B model. Then what I'll do is I'll take the good traces and the good examples from the GLM 5.2 runs, I'll prepare them in a data set, and then I'll try to fine-tune a small model on that data set to mimic behavior, essentially. And this is distillation, SFT. The other one is generating evals in environments. So maybe another slightly hot take. I think you can basically define agent behavior by showing the evals that you ran on it. If someone showed me all the things that they're trying to test their agent on, I think I would have a rough idea about how that agent is going to behave because it literally hill climbs those evals and you alter the behavior of the agent to make the evals pass. The purpose of evals is roughly to try to make them pass, right? So I update my agent so that they essentially pass. And then the other thing is humans are still in the loop. I need to know that customers are happy. I also want to know what my agents are doing. I just don't have the bandwidth to read a bunch of traces. So preparing content for humans is still really, really valuable today, especially in high-trust domains like legal and medical. Some human needs to review this, but they can't read it all. So we try to make it easy for them to process all that data. Great. This is maybe a bit of a throwback. How many people here know what Scikit-learn is? Let me put your, oh sick, this crowd is just awesome. Cool. So when I was first doing my PhD, my PhD was trying to do this, but add new algorithms to Scikit-learn. And what Scikit-learn basically is, at an abstract level, it's a bunch of helpers to fit learning systems to data. Right? And in classical machine learning, I had a data set and I tried to fit it to it. But I think the same principles that we use in modern, I call it classical machine learning, it was six years ago, that we do in classical machine learning definitely still apply to this agent-first world. The way that they apply is what I like to call model, harness, task fit. So we still have this fit function that I'm going to try to take my data, take a harness, take a model, and I'm going to try to fit it all together to make sure that all of my tasks pass. Right? The algorithms look slightly different, but the overall process of machine learning doesn't really look that different. And we'll talk about maybe roughly what our job becomes in this data-first, agent-first, fit-first world. So a couple of our main jobs now are find good fit functions. So these are auto research. This is tons of great work that's being done in RL on different methods like OPD, OPSD, try SFT. And also find good data. Right? So if you put those two things together, then that is basically the applied, or just overall, research question that every team has to make their agents better. And some examples that we've seen that are very popular, that we're pretty bullish on, are just generally auto research. So if you have some sort of score that you can make number go up, agents are pretty good at making that number go up. They might cheat a little bit and you need to check them on some stuff. But this general feedback loop of do something, read the results, read the traces, and then do an update ends up being pretty useful. And then I talked about model fine-tuning a bunch as well. So we just went and did this. This was, I think, even before the term auto research came out, I think it was a lot of people doing it, which is hey, terminal bench is really hard. What would happen if an agent just read its traces, proposed experiments, and then tried to do fixes? I think one really key thing here is giving agents dense feedback signals. So terminal bench, the output is just a number, right? Did you pass or did you not pass? That's helpful. But if I gave you a super random task, you just did a bunch of stuff, and then I just said you failed or you passed. If you failed, you wouldn't really have a good signal to figure out what you should do next, right? So densifying feedback is a really good way to improve agents. And traces are the substrate that hold that feedback. And then agents are very good at reading those traces and then figuring out what to do next. And then this question always comes up, which is when should I harness eng? When should I fine-tune? Should I do more harness eng after it? I'm pretty bullish on the idea of if you need to do something for improving your agent, the best thing that you can do is collect feedback as quickly as possible, either from humans labeling or just letting the agents run. So harness engineering gives you feedback in maybe two minutes. Once you saturate the harness engineering ceiling, right, then you can maybe try to do fine-tuning after that. But we find a lot of teams are happy with harness engineering and it solves their customer use case. So we always recommend it. And then we have this sandwich, which is try harness engineering, try to do fine-tuning to break through that ceiling, and then do more harness engineering again if you need to. And then I'll end on the idea generally of continual learning, is that there's an agent taking actions in the environment and then it needs to use that information, sorry guys, needs to use that information to update information about itself, right? So it's like I did a bunch of these tasks and I need to update my prompts to make sure I do them more efficiently. Or users keep asking to search for these types of things. I should maybe tell my creator that they're doing this sort of stuff, right? It's like taking action in the environment, kind of like humans do, and updating ourselves. But we find a lot of teams are happy with harness engineering, and it solves their customer use case. So we always recommend it. And then we have this sandwich, which is try harness engineering, try to do fine tuning to break through that ceiling, and then do more harness engineering again if you need to. And then I'll end on the idea generally of continual learning, is that there's an agent taking actions in the environment, and then it needs to use that information, sorry guys, needs to use that information to update information about itself, right? So it's I did a bunch of these tasks and I need to update my prompts to make sure I do them more efficiently. Or users keep asking to search for these types of things. I should maybe tell my creator that they're doing this sort of stuff, right? It's taking action in the environment, kind of like humans do, and updating ourselves. What that looks like today, slightly unclear, but we think that you're going to have to do it across all three axes, which is one, collect a bunch of training data, which is observational data from agents taking actions. The other one is harness updates generally, the codex harness and the cloud code harness and our harness and everyone's harness, they look a certain way because models are trained in them, and they look a certain way because of the tasks that they do in the real world. And we think evolving those over time is going to be super important in order to make them work. And the last thing is memory. So we humans are really good at remembering stuff over time, but we are not append-only logs of information. And if agents are going to be working with us over year, five-year, decade, lifetime timescales, we cannot just append everything to a really big file and then search over it. There's a ton of stuff that needs to happen with updating those files over time and then just making memory really efficient. But we think a lot of that actually comes from this idea of scaling sleep time compute and dreaming generally. So it's read all of the traces over the entire agent life cycle and then do things to update agent state. Awesome. So quick, quick takeaways. Mining traces gives you signals to hill climb on. I would say if you have an agent, just turn on tracing and point an agent at it. And that's the easiest thing that you can do to see, to understand what your agents are doing. We're very excited about open models. We want to help you fine tune open models. We provide them as a service as well. So if you're interested in that, we would love to chat how you can use open models to make everything smarter and cheaper. Continual learning is about operating environments and then integrating that data back into agent state. And then finally, I think this is so cool that we have systems that are going to produce more data than we ever have before. We need to all come up with interesting research directions to learn how to manage that at scale and make all of our agents better. And with that, thank you all for coming. Another question I might ask is, this is a technical question, agents now run for millions of tokens. Does the agent get really dumb after the first compaction, after the second compaction? Does it never get dumb? Like how do we actually answer those questions? We need to do it by actually looking at the traces. And then the other thing is like, if I look at the traces, then I can try to prove some counterfactuals, which is, hey, I ran GPT 5.5 for this, and I heard GLM is really good. What happens if I run GLM 5.2 for this task, and how do I compare them? Metrics, awesome. The trace level captures the actual behavior that users see, so that's also very helpful for seeing behavior like fine-grained scales. And the way that we sort of think about the data that's being generated by agents is that the data that we see today is going to be the smallest that humans have ever seen in their entire lives, because we're in this massive exponential shift to how agents are doing more and more work in the economy. And what that means is like the amount of data that humans have produced in our entire lifetime will soon be eclipsed by agents running on like year scales, and then six-month scales, then three-month scales, and then maybe every day, right? And to understand a ton of that data, roughly what we need to do is contend with a couple problems. There's more, but these are the two that I'm going to focus on. So, one, reading traces at scale is super expensive, especially if you have millions of traces and if you have millions of tokens per trace, right? Think of it as like an input token cost. You can literally multiply the input token cost times the number of traces times how big each trace is on average, right? The other thing is, if I have a super long interaction with a coding agent, like Cloud Code, or Codex, or like deep agents, I can't even read that trace with another agent because that context like doesn't fit in memory, right? So, it's like we need to develop systems so I can sort of treat that context as like an external object, and then I can sort of query into it, right? So, we need to build agents to efficiently mine data from other agents, and it's no longer as simple as just like feeding the data into context, and there's like tricks that we'll sort of talk about to do that well. Great. So, one of the things that I think is really, really cool in the last six months is that open models have basically hit an inflection point in intelligence that we at Langchain don't reach for the frontier models for every single use case. We're quite conscious about what is the minimum level of intelligence that I need to do any given task. And like practically speaking, honestly, yes, we start with Opus, we start with 5.5, because we just want to know if the task is even possible, but then once we reach that sort of like waterline, then we like look back at those traces and we see, hey, can we use an open model to do the same thing? So, this is a bunch of work that we did with Harvey and their lab legal benchmark. Basically what we're looking at is can I match the trace judging capability of Opus with an open, cheaper model? And the answer is roughly yes at like an order or like two orders of magnitude cheaper. And like the way we do that is we try a bunch of models, we do a bunch of like harness engineering, and the harness engineering is informed by a bunch of the traces that we read. So, it's like, hey, like Opus reasons about things in this way, maybe that's because of the prompt, maybe Opus is just smarter, which it is than a bunch of the open models, but that might mean I need to give it a little bit more guidance so it can reach the sort of same intelligence level at like a much lower cost. And the other thing that we sort of look at is like harness engineering is amazing. You get instant feedback and you can sort of like run on your evals, but eventually what we find is you hit a threshold of intelligence where it's like, if I keep tweaking this prompt, I'm not going to get too much more out of it. And once we reach that point, we sort of look at, okay, can I actually like fine tune the model on my domain specific task? And can I like make it better on those tasks? And what we find is if we did like base models and we tuned them on like very specific vertical tasks, which is what a lot of our customers do, they don't really care about the entire variance of tasks, like they care about what their customers care about. So if we focus on that narrow set of tasks, then we can fine tune base models to sort of like reach and then also go beyond frontier performance. And I think one sort of like small thing I'll mention as a lot of people are getting into fine tuning is that another sort of like economic decision is that you can move from token costs to hardware costs. And this is like, can be a really big change, right? Because like you're very used to, hey, like a million tokens cost this much, not as much like this cluster sort of cost this much. But for like very high inference workloads, we find it to be way cheaper just to like run a cluster and I get like unlimited inference on that cluster. I don't have to worry about tokens, but I can just do the calculation of like, hey, this will end up being cheaper and then I can spin it down when I don't need it. Cool. And I said all of this, so we obviously like built a product to do that. I won't show it too much, but it's Langsmith Engine. Basically this product is trying to automate this loop for you, which is if you have any volume of trace data and you're looking for something that trace data or you want to generate evals from that trace data or you want to like generate feedback for like humans to read from that trace data, it will go read all of it, it will like find issues, it will agentically search over it and they can like prepare data sets for you to do something after and a bit of a leader. What that something basically is, is the outputs of this trace mining exercise. So there's like three things that I mentioned here, which we see a bunch and we kind of put into the product. So one is distillation and fine tuning, which is let's say I'm running GLM 5.2, it's doing great, but I think that I can run this task like way cheaper with like a 9B or 13B model. Then what I'll do is like, I'll take the good traces and the good examples from the GLM 5.2 runs, I'll prepare them in a data set and then I'll try to fine tune a small model on that data set to like mimic behavior essentially. And this is like distillation, SFT. The other one is generating evals in environments. So maybe another slightly hot take. I think you can basically define agent behavior by showing the evals that you ran on it. Like if someone showed me all the things that they're trying to test their agent on, I think I would have a rough idea about how that agent is going to behave because it literally like hill climbs those evals and you alter the behavior of the agent to make the evals pass. The purpose of evals is roughly to try to make them pass, right? So I update my agent so that they essentially pass. And then the other thing is like humans are still in the loop. Like I need to know that customers are happy. I also want to know what my agents are doing. I just don't have the bandwidth to read a bunch of traces. So preparing content for humans is still like really, really valuable today, especially in like high trust domains like legal and medical. Like some human needs to review this, but they can't read it all. So we try to make it easy for them to process all that data. Great. This is maybe a bit of a throwback. Like how many people here know what Scikit-learn is? Let me put your, oh sick, this crowd is just awesome. Cool. So when I was like first doing my PhD, my PhD was like kind of trying to do this, but like add new algorithms to Scikit-learn. And like what Scikit-learn basically is an abstract level. It's a bunch of helpers to fit learning systems to data. Right? And like classical machine learning, I had like a data set and I tried to fit it to it. But I think the same principles that we use in modern, I call it classical machine learning. It was like six years ago that we do in classical machine learning, definitely still apply to this agent first world. The way that they apply is what I like to call model, harness, task fit. So we still have this sort of like fit function that I'm going to try to like take my data, take a harness, take a model, and I'm going to try to fit it all together to make sure that all of my tasks pass. Right? The algorithms look slightly different, but the overall process of machine learning doesn't really look that different. And we'll talk about maybe roughly what our job becomes in this data first, agent first, fit first world. So a couple of our main jobs now are find good fit functions. So these are like auto research. This is tons of great work that's being done in RL on different methods like OPD, OPSD, try SFT. And also find good data. Right? So if you put those two things together, then that is basically the applied or just overall research question that every team has to make their agents better. And like some examples that we've seen that are very popular that we're pretty bullish on are just generally auto research. So if you have some sort of score that you can make number go up, agents are pretty good at making that number go up. They might cheat a little bit and you need to like check them on some stuff. But this sort of like general feedback loop of do something, read the results, read the traces, and then do an update ends up being pretty useful. And then I talked about like model fine tuning a bunch as well. So we just like went and did this. This was I think even before the term auto research came out, I think it was a lot of people doing it, which is hey, like terminal bench is like really hard. What would happen if an agent just like read its traces, proposed experiments and then tried to do fixes. I think one like really key thing here is giving agents dense feedback signals. So like terminal bench, the output is just a number, right? Like did you pass or did you not pass? That's like kind of helpful. But if I gave you like a super random task, like you just did a bunch of stuff and then I just said like you failed or you passed. If you failed, like you wouldn't really have a good signal to figure out what you should do next, right? So densifying feedback is a really good way to improve agents. And like traces are the substrate that hold that feedback. And then agents are very good at like reading those traces and then figuring out like what to do next. And then this sort of question always comes up, which is when should I like harness eng? When should I fine tune? Should I do more harness eng after it? I'm like pretty bullish on the idea of if you need to do something for improving your agent, the best thing that you can do is collect feedback as quickly as possible. Like either from humans labeling or just letting the agents run. So like harness engineering gives you feedback in maybe two minutes. Once you sort of saturate the harness engineering ceiling, right? Then you can maybe try to do like fine tuning after that. But we find a lot of teams are happy with harness engineering and it solves their customer use case. So like we always sort of recommend it. And then we have this like sort of sandwich, which is like try harness engineering, try to do fine tuning to sort of like break through that ceiling and then do more harness engineering again if you need to. And then I'll sort of end on the idea generally of continual learning is that there's an agent taking actions in the environment and then it needs to use that information, sorry guys, needs to use that information to update information about itself, right? So it's like I did a bunch of these tasks and like I need to update my prompts to make sure I do them more efficiently. Or users keep asking to search for these types of things. I should maybe tell my creator that like they're doing this sort of stuff, right? It's like taking action in the environment kind of like humans do and updating ourselves. What that looks like today, slightly unclear, but we think that you're going to have to do it across all three axes, which is one, collect a bunch of training data, which is like observational data from agents taking actions. The other one is like harness updates generally, like you know, the codex harness and the cloud code harness and like our harness and everyone's harness, like they look a certain way because like models are trained in them and they look a certain way because of the tasks that they do in the real world. And we think like evolving those over time is going to be super important in order to make them work. And the last thing is like memory. So we humans are like really good at like remembering stuff over time, but we are not append only logs of information. And if agents are going to be working with us over like year, five year, decade, lifetime timescales, we cannot just append everything to like a really big file and then search over it. There's a ton of stuff that needs to happen with like updating those files over time and then just making memory like really efficient. But we think a lot of that actually comes from this idea of scaling sleep time compute and dreaming generally. So it's like read all of the traces over the entire agent life cycle and then like do things to update agent state. Awesome. So like quick, quick takeaways. Mining traces gives you signals to hill climb on. I would say like if you have an agent, just turn on tracing and point an agent at it. And that's like the easiest thing that you can do to see like, to basically understand what your agents are doing. We're very excited about open models. We want to help you fine tune open models. We provide them as a service as well. So if you're interested in that, we would love to chat how you can use open models to make everything smarter and cheaper. Continual learning is about operating environments and then integrating that data back into agent state. And then finally, I think this is so cool that like we have systems that are going to produce more data than we ever have before. We need to all come up with like interesting research directions to learn how to like manage that at scale and like make all of our agents better. And with that, thank you all for coming.