AI Engineer

Memory Harnesses for Long-Running Research Agents — Stefania Druga, Sakana.ai

1832 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: For long-running agents, memory should be treated as an explicit write-manage-read control loop, with a ranked decision-ledger recall policy outperforming no memory, generic vector RAG, and simple memory-use gating on the speaker's long-horizon evaluations.
  • Why it matters: This offers a concrete evaluation framing for agent memory: retrieval quality and recall policy—not merely the existence of a vector store—determine both task accuracy and token cost once relevant history exceeds the context window.
  • Best use: Use it to shape a memory-harness experiment for Ken's long-running agents: compare no-memory, similarity retrieval, ranked structured memory, and an oracle-retrieval upper bound while logging whether retrieved memory is actually used.

Executive Summary

Stefania Druga argues that context-window growth does not eliminate the operational problem of agent memory. Long-horizon agents still contradict prior work, repeat completed tasks, and drift from the original question when relevant information falls outside active context. Her proposed remedy is a “memory harness”: an external control loop responsible for writing, managing, and retrieving memory around an otherwise stateless research agent.

Her harness separates always-visible core traces, a recall layer, and archival storage across sessions. She holds the underlying model constant and changes only recall policy: no memory, vector RAG, a ranked decisions ledger, and an oracle that supplies the correct memory item. The key result is conditional: on literature-review tasks where all source material fits in context, memory did not improve performance and only increased cost.

On XBench-style long-horizon tasks, where the answer at step 124 must be recovered when queried at step 500, a rank-only decision ledger produced the strongest results across more than 68 questions, multiple seeds, Qwen 27B 4-bit, DeepSeek V4 Flash, and a reported Spider V2 test. It also beat a policy that merely decided whether memory should be used. The framing is important: a model can still fail even when an oracle supplies the right memory, because it may ignore, misuse, or be confused by that information.

The practical message is to promote recall policy to a first-class system metric. Bad recall is doubly harmful: it consumes tokens and sends the agent down incorrect paths. Local deployment gives strong control over traces, data, and evaluation, but her setup also exposes the throughput trade-off: on an M3 Ultra, DeepSeek V4 Flash could only be evaluated serially rather than in batches.

Key Takeaways

  • Claim: Memory is not simply a storage layer; it is a write-manage-read control loop that must be designed around the model. | Evidence: Druga's harness has an always-visible core of traces, a recall block that selects what to inject into the agent context, and an archival block that preserves information across sessions. | Implication: Ken should specify memory-writing criteria, retention rules, ranking logic, and read/injection behavior as separate components rather than treating a vector database as the memory system.
  • Claim: Do not add a memory harness by default when the full task and relevant evidence already fit inside the model context. | Evidence: In a literature-review task involving a Nature paper's claim of 742,000 promising materials and its less-prominent later retraction, memory and no-memory conditions performed the same because the corpus fit in context; memory only raised cost. | Implication: Route short, context-contained tasks to a clean-context workflow and reserve durable-memory retrieval for cases with measurable context overflow or cross-session state requirements. | Caveat: This finding applies to tasks where all relevant material is genuinely available in active context; it does not establish that memory is unnecessary for multi-session work, large corpora, or stateful workflows.
  • Claim: For long-horizon recall, a ranked decisions ledger outperformed generic recall alternatives in the reported experiments. | Evidence: On XBench, the agent was asked at step 500 for an answer originating at step 124, outside its context window. Across over 68 questions, multiple cells and seeds, the rank-only ledger was the best-performing condition; it also outperformed simple gating of whether to use memory. | Implication: For agent workflows, preserve structured decision artifacts—decisions, commitments, conclusions, and their relevance signals—and rank those before relying on embedding similarity alone. | Caveat: The transcript does not provide absolute accuracy, confidence intervals, exact ledger schema, or a full comparison table, so this is a directional architecture result rather than a production-ready benchmark claim.
  • Claim: Perfect retrieval is not sufficient for correct agent behavior because the model may fail to use retrieved evidence correctly. | Evidence: The oracle condition supplied the ground-truth memory item for each loop but did not reach maximum performance; the model could still choose incorrect information, ignore the retrieved item, or become confused. | Implication: Evaluate memory systems in two stages: retrieval correctness and downstream answer/action faithfulness. A high retrieval hit rate alone can conceal an agent integration or reasoning failure.
  • Claim: Poor memory policy increases both cost and error, while structural recall policies can improve the cost-quality trade-off. | Evidence: Druga reports that the ranked approach not only improved recall but cost less across the tested long-horizon settings, including Qwen 27B and DeepSeek V4 Flash, with an additional reported Spider V2 test. Her stated heuristic is that bad memory wastes tokens and misdirects the agent. | Implication: Track retrieved-token volume, irrelevant-memory rate, retrieval-to-action adherence, and task success jointly; optimizing memory solely for recall can increase spend and degrade trajectories. | Caveat: No token totals, latency figures, or per-model cost breakdowns are provided.
  • Claim: Local models can support controlled memory-harness research and agentic tool-use experiments, but local evaluation throughput may be a material constraint. | Evidence: The setup used an M3 Ultra with 96 GB memory and 28 CPU cores, running Qwen 27B at 4-bit quantization and DeepSeek V4 Flash. Druga says DeepSeek V4 Flash lacked batch querying in her configuration, forcing serial evaluation over days. | Implication: Local execution is valuable where sovereignty, trace control, and reproducibility matter, but Ken should budget separately for batch-serving capability and evaluation turnaround time. | Caveat: The claim is based on this specific hardware/software setup; it should not be generalized to all local serving stacks or model variants.

Detailed Brief

Experimental framing and comparison ladder

  • Claims: The speaker intentionally begins with small research agents that have zero durable memory so that all retained state is attributable to the harness rather than hidden model or application state.; The core experimental intervention is recall policy while the model remains fixed, making the experiment an attempt to isolate the value of memory selection rather than model capability.; The recall ladder moves from no recall, to vector similarity retrieval, to a decision ledger with prioritization, to oracle retrieval as a reference condition.
  • Evidence: The oracle is defined as ground truth telling the harness which memory should be retrieved on each loop.; Ablations included giving arbitrary examples, the wrong historical step, and the most recent step; the ranked-policy condition remained best in the speaker's account.
  • Caveats: The transcript does not explain the ranking features, the decision extraction process, the number of items injected, or whether the ledger was manually structured versus model-generated.; Because the oracle does not compel model use, oracle performance is an upper bound on retrieval availability rather than a ceiling on end-to-end task accuracy.
  • Implications: A useful internal benchmark should distinguish retrieval selection, prompt injection, and the agent's ability to ground its answer in the injected state.; Decision memory may be more robust than raw episodic transcript retrieval when the agent needs to recover prior commitments rather than merely find semantically similar text.

Memory landscape and sovereignty positioning

  • Claims: Druga characterizes the memory design space as broad, spanning simple filesystem retrieval through trained memory models, with varying degrees of structure.; She positions local execution as a form of sovereignty because it permits control over data, compute traces, and evaluations.
  • Evidence: She references more than 30 runnable memory cookbooks in an open-source repository from Diamond.; Her rationale for pursuing local models includes better routing, caching, cleaner context, and visibility into task usage; she cites a recent Coinbase CEO post as an example of lowering AI spend while increasing use through such practices.
  • Caveats: The Coinbase reference is anecdotal in this presentation and is not evidence that local deployment alone lowers cost.; The talk is an early experiment rather than a comprehensive survey or standardized production benchmark.
  • Implications: Memory architecture should be chosen on an operational spectrum: start with the minimum structure that solves demonstrated context loss, then add learned or more elaborate memory only when evaluation justifies it.; For sensitive or sovereignty-constrained deployments, local harnesses can improve inspectability, but performance engineering of batching and scheduling remains necessary.

Notable Concepts & Terms

  • Memory harness: An external system around an agent that controls memory writing, management, retrieval, and cross-session archival rather than delegating memory to the model context alone.
  • Write-manage-read loop: The speaker's operational model for memory: decide what state to record, maintain and rank it, then retrieve the relevant subset at the correct point in an agent trajectory.
  • Decisions ledger: A structured per-turn record of agent decisions that can be prioritized for recall; it is the reported best recall mechanism in the experiments.
  • Rank-only ledger: The specific leading condition in which ledger memories are ranked for recall, rather than relying only on a gate deciding whether memory is needed.
  • Oracle retrieval: A ground-truth condition that supplies the correct memory item to assess the gap between retrieval availability and the model's actual ability to use that evidence.
  • XBench: An established long-horizon memory benchmark used here to test whether an agent can retrieve information from an earlier step after it has fallen outside the context window.
  • Context blow / context rot: Failure modes in extended agent runs: contradiction, repeated work, and goal drift caused by losing or mishandling prior context.
  • Sovereign AI: The ability to control models, data, compute traces, and evaluation pipelines locally; presented as valuable despite slower serial evaluation.

Operator Notes / Why Ken Should Care

  • Create a memory-harness evaluation matrix for one representative long-running workflow: no memory, vector RAG, ranked structured ledger, and an oracle-retrieval condition.
  • Instrument two separate failure labels in agent traces: retrieval failure (needed memory was absent or wrong) and utilization failure (correct memory was present but ignored or misapplied).
  • Define a compact durable schema for decisions, commitments, unresolved questions, source provenance, and supersession/retraction status; test ranking rules before expanding storage volume.
  • Add a context-fit gate: skip durable-memory retrieval when the relevant task state is already safely contained in active context, and verify the gate with token-cost and success-rate telemetry.
  • If evaluating local models, confirm batch-query support and throughput before committing to a large experimental sweep; serial-only evaluation can dominate iteration time.

Source/Metadata

  • Title: Memory Harnesses for Long-Running Research Agents — Stefania Druga, Sakana.ai
  • Transcript words: 2998
  • Duration seconds: 784
  • Timestamp note: No timestamps or chapters were present in the supplied transcript; the latter portion substantially repeats earlier material.
Full transcript 1821 words · 13 min read
0:12

Hello, welcome. This is a big room, so if you're in the back, don't hesitate to come closer. My name is Stefania Druga. I'm a research scientist at Sakana AI in Tokyo. I used to be based here, and AI Engineering is home community for me before Hyperloop. So it's very good to be back. Today I'm going to talk to you about memory harnesses for long-running research agents on device.

0:34

If you work with long-horizon tasks, you probably ran into this issue of context blow, right? When the model starts contradicting itself, or it has to redo the work because it forgot it did that task in the first place. Or it starts to drift from your questions because it forgot them. This matters now more than ever because, from these recent projections from Meter, we see that the trend is to solve longer and longer horizon tasks and also that we're getting fewer and fewer model releases. At some point later this year, we're going to have this convergence, right? Where we'll get many more long-term horizon tasks and fewer model releases.

0:42

That makes this issue of dealing with context rot a priority. Why did I want to tackle this problem on local models and with a local harness? Maybe some of you have seen this tweet. It's only two days old. The CEO of Coinbase actually shared how their company managed to reduce their AI spend while actually increasing AI usage. The way they did that was by transitioning to use many more local models, but also having better practices like using better routing, better caching, keeping the context clean, and then having better visibility for what people are using it for, what kind of task.

1:01

We are seeing the local models crossing the line, right? GLM is on everyone's minds, especially with Fable going away. Deep Seek v4 Flash can now be run on M3 Ultra. There's still a bottleneck for RAM. It's tricky. But these local models are starting to be useful for agentic tasks and for tool use.

1:10

I wanted to show you what has been my setup for the experiments I'm going to share with you today. This is my Mac. It's still running evaluations right now, back at my desk in Tokyo. And I'm controlling it from my phone. After running evals non-stop for a couple of days, it started to get hot, so I had my husband put fans around it. We're running out of fans. But the machine is still running, and the evals are still giving results. On this M3 Ultra with 96 gigabytes and 28 core CPUs, I'm using two models. I'm using the QN 27B quantized at 4-bit and the Deep Seek v4 Flash.

1:22

Before I show you how I built the memory harness on this machine, I wanted to tell you what this loop is an example of. Memory, when we design a harness for memory, this is the mental model I want you to have in mind. You can think of memory as a write, manage, read loop. It's not just the database store. It's actually this control loop around the model.

1:36

More concretely, how did I take that loop and customize it? This is my harness design. I started with research agents that are the small agents because they have zero durable memory. I wanted all the memory to come from the harness. In the middle, I have a core, which is always shown to the agent of traces. Then I have a recall block where I'm testing different modes, and an archival block where I'm keeping track of information across different sessions.

1:42

In that recall block, I'm actually going through a ladder of modes that I'm testing. The baseline is not to use memory at all, no recall at all. So I'm testing for that. Next is to use RAG, vector, vector RAG, just to see whatever the harness would pull in terms of similarity. Then is to use a decisions ledger, where I actually keep track of what decisions are being made for every turn. Then I can prioritize them. Last but not least, and this piece is very important, I have what I call an oracle. This is the ground truth. This is telling the harness for every loop what the correct memory that needs to be retrieved is.

1:49

The model is fixed across all the different tasks. So the only things that I'm changing are these different variables in the recall block.

1:57

I wanted to give you an example of a first task that I tested. I wanted to see if I give the agent a task of doing literature review, and I'm including a lot of papers in the corpus where there was a big scientific claim. This is actually a Nature paper where they said they discovered 742,000 promising materials. It was a very big claim, which got retracted later. But the retraction is a much smaller haystack needle in that corpus than the headlines and the citations. So I wanted to see if the system can retrieve the right answer for these types of questions.

2:05

What I found was because, for these tasks, all the papers and all the information fit into the context, the memory actually didn't add more capability. It was the same performance with memory and without memory. It only added more cost. So when your task fits in context, the harness doesn't add much. However, if I start to run tasks that are longer-term horizon and the entire task and the relevant context don't fit, then having a good memory harness really starts to pay off.

2:21

This is another example of a task that I ran. This is actually from an established benchmark for long-horizon task memory. It's called XBench. This is an example of a question, right? I'm asking a question, and the right answer is step 124. But the moment when I ask the question, I'm asking it at step 500. So it's completely outside of the context window. The model needs to use the memory harness to retrieve the specific answer from the right step. I'm testing this by changing the different policy ladder that I explained before, with memory off, by deploying recall, different types of recall, and by using the oracle as a reference.

2:39

What I found was that with the ranked recall, the model gets the right answer more frequently than without. Here's a breakdown of the decomposition of performance on this XBench task. I ran over 68 questions. For each of these questions, there were multiple cells and lots of different seeds. What I found was that the rank-only ledger performed the best. It performed better than just gating the harness by saying, do you need to use memory or do you not need to use memory?

2:53

You're probably going to ask why the oracle is not hitting the max, and I'm going to explain that too. The oracle, what it does, is provide the right information, the right memory to the model, but it doesn't force it to use it. So the model can get the right memory but still retrieve the wrong information, or choose to ignore it or be confused. That's why the oracle, in this case, doesn't hit the max performance.

3:01

I've done lots of ablations on these tasks to see what happens if I give arbitrary examples, what happens if I give it the wrong step, what happens if I give it the most recent step. I still found that the best-performing condition was the one with the ranked policy for recall. This actually works on several models, not only on the QN27B but also on the DS4 Flash, and it also works across different benchmarks. I also tried it on the Spider V2 benchmark.

3:16

It's not just that it gives you better recall, it actually costs less. Maybe a good heuristic to have here is that bad memory is expensive because it spends more token and it can send the agent the wrong way. But having a good structural policy for recall can save you a lot of tokens and budget.

3:27

One thing that I want to encourage you from this experiment is to consider the recall policy as a first-class metric, and to start to think about how you might use it in your systems. What are the types of memories that you want to store? How do you rank them? How do you design your recall function? Then, what survives when you run this over and over and over, and multiple sessions, multiple runs? This is just a simple first experiment. But the memory technique landscape is very rich. There's over 30 runnable cookbooks that are shared in this open-source repository from Diamond.

3:42

Memory is complex. We have short-term, long-term, different cognitive techniques. We can start to use evaluation results as well. Right now, there's actually a pretty broad landscape of solutions, right? Going from simple file system retrieval to training memory models, there's a wide spectrum of solutions from less structural to completely structured. I think there's a lot of research we're going to see in this space.

3:48

It's important. It becomes more and more relevant. For me, it's been super fun to test this on local models because I got to control everything. I got to control the data I was using, the entire traces of compute and evaluations. I see that as an example of sovereignty. It comes at a cost. I didn't tell you that these local models, I can only run them in serial. They don't support batch querying for the Deep Seek v4 Flash. That's why I'm still running evaluations back on my computer in Tokyo. Or I was doing it on the flight on my way here, because it takes a long time.

4:00

But I still think it's very powerful, and it's a very good test for what memory can do when you can control every single step of the pipeline. Sovereign capability is part of a bigger ecosystem that is very important for us at Sakana AI in Japan. We believe in the importance of sovereign AI today more than ever. We are also hiring. If you're interested and want to hear more about this, and if you want to come join us in Japan, come talk to me. Thank you very much. always shown to the agent of traces. And then I have a recall block where I'm testing different modes. And an archival block where I'm keeping track of information across different sessions.

4:28

And in that recall block, I'm actually going through a ladder of modes that I'm testing. The baseline is, like, not to use memory at all. No recall at all. So I'm testing for that. Next is to use rag, vector, vector rag, just to see whatever, like, the harness would pull in terms of similarity. Then is to use a decisions ledger, where I actually keep track of what decisions are being made for every turn. And then I can prioritize them. And last but not least, and this piece is very important, I have what I call an oracle. But basically, this is the ground truth. So this is, like, telling the harness for every loop what the correct memory that needs to be retrieved is.

5:18

And the model is fixed across all the different tasks. So the only things that I'm changing is, like, these different variables in the recall block. And I wanted to give you an example of a first task that I tested. So I wanted to see if I give the agent a task of doing literature review. And I'm including a lot of papers in the corpus where there was a big scientific claim. Like, this is actually a nature paper where they said they discovered 742,000 promising materials. Like, it was a very big claim, which got retracted later. But the retraction, it's as much smaller like haystack needle in that corpus

6:03

than the headlines and the citations. So I wanted to see if the system can retrieve the right answer answer for these type of questions. And what I found was because, like, for these tasks, all the papers and all the information fit into the context, the memory actually didn't add more capability. It was the same performance with memory and without memory. And it only added more cost. So when your task fits in context, the harness doesn't add much. However, if I start to run tasks that are longer-term horizon and the entire task and the relevant context doesn't fit, then having a good memory harness really starts to pay off.

6:54

So this is another example of a task that I ran. This is actually from an established benchmark for a long horizon task memory. It's called XBench. And this is an example of a question, right? So I'm asking a question, and then, like, the right answer is, you know, like, step 124. But the moment when I ask the question, I'm asking it, like, at step 500. So it's completely outside of the context window. And the model needs to use the memory harness to retrieve the specific answer from the right step. So I'm testing this by changing the different policy ladder that I explained before with memory off,

7:41

by deploying recall, different types of recall, and by using the oracle as a reference. And what I found was that with the ranked recall, the model gets the right answer more frequently than without. And here's a breakdown of the decomposition of performance on this XBench tasks. So I ran over 68 questions. And for each of these questions, there were, like, multiple multiple cells and lots of different seeds. And what I found was that the rank-only ledger performed the best. And it performed better than, like, just gating the harness by saying, do you need to use memory or do you not need to use memory?

8:32

And you're probably going to ask, like, why is the oracle not hitting, like, the max? And I'm going to explain that too. So the oracle, what it does, it provides the right information, the right memory to the model, but it doesn't force it to use it. So the model can get the right memory, but still retrieve the wrong information, or choose to ignore it or be confused. So that's why the oracle, in this case, doesn't hit the max performance. And I've done lots of ablations on these tasks to see, like, what happens if I give arbitrary examples? What happens if I give it the wrong step? What happens if I give it the most recent step? And I still found that the best performing

9:17

condition was the one with the ranked policy for recall. And this actually works on several models, not only on the QN27B, but also on the DS4 Flash, and it also works across different benchmarks. I also tried it on the Spider V2 benchmark. And it's not just that it gives you better recall, it actually costs less. So maybe a good heuristic to have here is that bad memory is expensive because it spends more token and it can send agent the wrong way. But having, like, a good structural policy for recall can save you a lot of tokens and budget. So one thing that I want to encourage you from this experiment is to consider the recall policy as a

10:09

first-class metric. And to start to think about how you might use it in your systems. Like, what are the type of memories that you want to store? How do you rank them? Like, how do you design your recall function? And then, what are the type, what survives when you run this over and over and over? And multiple sessions, multiple runs. And this is just a simple first kind of experiment. But the memory technique landscape is very rich. So there's over 30 runnable cookbooks that are shared in this open source repository from Diamond. And memory is complex. We have short-term, long-term, different cognitive techniques. We can start to use evaluation results as well.

11:02

And right now, there's actually a pretty broad landscape of solutions, right? So going from simple file system retrieval to training memory models, there's a wide spectrum of solutions from less structural to completely structured. So I think there's a lot of research we're going to see in this space. It's important. It becomes more and more relevant. And for me, it's been super fun to test this on local models because I got to control everything. I got to control the data I was using, the entire traces of compute and evaluations. And yeah, I see that as an example of sovereignty.

11:49

And it comes at a cost. I didn't tell you that these local models, I can only run them in serial. Like they don't support batch querying for the DeepSync v4 flash. So that's why I'm still running evaluations back on my computer in Tokyo. Or I was doing it on the flight on my way here because it takes a long time. But I still think it's very powerful. And it's a very good test for what memory can do when you can control every single step of the pipeline. Sovereign capability is part of a bigger ecosystem that is very important for us at Sakana AI in Japan. We believe in the importance of sovereign AI today more than ever. And we are also hiring. So

12:35

if you're interested and want to hear more about this, and if you want to come join us in Japan, come talk to me. Thank you very much.

12:51

So So So So So So So So So So So

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note