.
All right. Good afternoon, everyone. Thank you for joining and not watching the game. I hope it will be a bit more interesting, or at least you will learn something compared to hopefully Germany winning or some... Anyways. All right.
Is it fine? Okay. All right. So I'm here to talk about... We are here to talk about context engineering in 2026. And more specifically, we are here because we've all lived that situation where you try to do things with an agent, and ultimately it does exactly the thing that you do and you don't want it to do. And in my case, it usually ends up like this, where I'm super mad and I just type back, hoping it learns. And usually the problem here is not that the model just got dumber and you need to switch to cloud or to codex or whatever harness that you're using, but it's more that the context is filling up and it's getting worse and worse.
The results are getting worse because of it. In our case, this is important because we build courses and trainings for AI engineers specifically. And one of the features that we provide is an AI tutor to help answer questions based on our lessons.
And if the interaction is just like the one before and they are super mad at us, they might just ask for a refund and could end up like this. So that's not what we want. And so what we did for this workshop, and just for the AI tutor in general, is to run many different experiments in order to figure out how, in our case, we can fix context rot or at least improve the AI tutor as much as possible and reduce the cost as well of running the tutor. The QR code here is a link to a Hugging Face Space where you have all these experiments that you can see. And also the AI tutor is open source. We will share another code for the repo, but it's also linked on the Hugging Face.
So everything is open source. You can access everything and even see the experiments online and use the AI tutor online as well. So in the next 80 minutes, I will start talking about compaction, memory, retrieval, and everything that you can do in 2026 that usually works. And then my colleagues will jump in with the architecture of our AI tutor, our decisions, what we built, and the evaluations that we built, how we built them, and what we decided to evaluate, and then the results and what we took out of this.
So, of course, it's applied to our use case and our AI tutor, but hopefully you can get away with some interesting insights at least from this and some best practices that we learned throughout. More specifically, we is Towards AI. I founded the company with my partners a few years ago, and we've always been focused around education. Obviously, back in the day, it was more about computer vision and more basic machine learning. Now it's Towards AI engineering and agents, anything that works for the industry. And I'm joined by my colleagues helping me develop this AI tutor and our courses, Omar and Samridi, who will jump in later on.
And here, more specifically, Towards AI is quite large. But one of the things that we do is our academy, so the Towards AI Academy, where we build courses, technical courses, for AI engineers to upskill towards AI engineering. And as I said, we provide an AI tutor for the students. So, the AI tutor specifically will be our baseline for all our experiments. We will use that to test all the different features based on real user interactions. And to have the best results possible, we had five requirements we wanted to ensure that the chatbot follows. The first one, obviously, we want the answers of the chatbot to be grounded in our content, not just its own knowledge.
We need the tutor to be based on the current students and current lesson, because we have multiple courses, so we just need to ensure that it answers from this course content. Then it needs to hold long help sessions in case the student is debugging or just iterating a lot with the tutor. And, obviously, it handles code because it's for AI engineers, so we just code a lot. And it needs to have somewhat low latency to not be frustrating to use. And all of this is related to context engineering. And here we'll be talking about what context engineering is in 2026, or at least what we figured out from these experiments.
And since everything here is in the context of models, it creates two problems. First, the context window is finite. Everything the model will see, from instructions to lessons to code, will lie in the same space that is limited. And the more you pile things in this space, the worse the results will be. And the more expensive it will be because you pay for more tokens. So that's one of the main problems we are trying to fix. And the second problem is that the model is stateless. So when a student reopens the AI tutor, typically if you don't build anything around it, the model will have no idea what's going on. It's just starting from zero.
So this creates two things we have to work on: the context management, which means within one session, and the memory aspect of these models, which means across sessions. In these experiments and in this workshop, we focus on the first part, the context management, because you cannot have multiple sessions if one session is shitty. So we try to really optimize for this. And maybe in a future workshop, we will do one about memory, hopefully. So for context management, what does the tutor see?
In our case, it sees many things, from the system prompt for being a tutor and the context of our courses and stuff, to tool definitions, so which tools it can use, how to use them, when to use them, the chat history, if it's an ongoing discussion, any old tool outputs that were called, course chunks that we retrieved to answer the students, and finally, the user questions. So it's not just the user question that we send, obviously. And typically it's the smallest part, but it can also contain a lot of code to debug and to help our error logs. So it can be large as well. And all of this together just ends up costing more and more to us.
So we really want to optimize this context management aspect. And what we've seen, just quickly, is that the main bottleneck or the main problem scaling the context is the old tool outputs, which contain any old chunks retrieved, all the tool calls and tool result pairs from all the tools called over the many turns that one session could have, and any file or searches that it did in its own memory. And the main problem is not that it's more costly. It's also that the quality degrades. It's a lot of things, as we know, with a problem called context rot because of the way large language models are trained to handle longer context.
We just inject facts into a large corpus, which doesn't tell them to manage the whole context together or understand the global context. And so one of the reasons is to help with quality. We want to reduce the context as much as possible. And the others are that if you re-query the model in a discussion, you re-send every previous token, so you pay for them again as well, which is far from ideal. And you increase the latency, so the time to first token, or TTFT, which creates a very bad user experience.
So you may want to, in our case, manage our context for speed and for spending, not just because the quality drops, as we will see in our experiments and in the results in the near future. So how do we do that to manage spending and speed? How do we manage our context? We use compaction, and the idea is just very simple. It is just to try to have the smallest context possible that contains the information to be able to answer the question. And you drop or save the rest somewhere. And to do compaction, even good compaction, you don't necessarily have to have large language models. You can start quite cheap with trivial tools.
Like if you use tools in your system, like a web search or just executing code, you can automatically truncate outliers. So if you have one tool that produced 300 lines instead of 10 usually, you can just truncate almost everything except the head and tail, the beginning and the end, and just write that it's truncated so that the model in the future can recall the tool if it feels it lacks context. You can use the simplest approach that works the best, to just use a sliding window or just trim, use the last N number of turns that the user sent, which you need to determine based on your own system and your users.
And you can clear, for some specific tools, depending on those that you implement, most of the outputs. So that's not even using language models. And then you can use language models to basically spend tokens to save even more tokens. And in this case, you don't even have to use large language models. You can even use smaller ones or even super small local ones that run on one MacBook to do a few techniques. There are many techniques that exist for compacting. Those that had the most impact in our experiments were selective retention, where the language models will just decide, based on where the discussion is going, what to keep, what to discard.
Then the simplest one here is summarization. So just summarize, continuously summarize some previous turns, depending on your own application. use the last N number of turns that the user sent, which you need to determine based on your own system and your users, and you can clear, for some specific tools, depending on those that you implement, you can just always clear most of the outputs. So that's for when that's not even using language models. And then you can use language models to spend tokens to save even more tokens. And in this case, you don't even have to use large language models. You can even use smaller ones or even super small local ones
that run on one MacBook to do a few techniques. There are many techniques that exist for compacting. Those that had the most impact in our experiments were selective retention, where the language models will just decide based on where the discussion is going, what to keep, what to discard. Then the simplest one here is summarization. So just summarize, continuously summarize some previous terms, depending on your own application. And in the end, you can do here what Cloud Code does. So when it reaches the limit, you just produce a summary of everything and reset completely with that summary.
And there are many more techniques. I highlighted here those that work the best. We will discuss them later on in this workshop and the actual results and experiment setup. But those are the most successful techniques. And I want to highlight also delta summarization that Cloud Code uses. That is very useful when you spawn subagents, when you use subagents. It just means to keep a summary and update the summary based on the new summary that you produce over time. And then the subagent will just give that to the main agent. But in our case, we don't use subagents because the tutor works really well with just one main, so we don't need to add complexity for that.
And lastly, after spending tokens to save more tokens, you can also offload things. Right now, obviously, memory and scales are super popular. So you can, of course, offload to your memory. So just saving text in your documentation. And you can use what's been there for years now, retrieval augmented generation, which is very powerful. And as a side note, we also compared with GraphRAG. So everything, even if I don't mention it, we compared. And I highlight just the best results here. So we compared GraphRAG with RAG here. And in our case, it just ended up being way costlier to set up and just tied on the results.
Because it basically was 100% based on our real user evaluations. So we don't need to use GraphRAG. But it depends on your own case. If you have a very large data set with relations and interconnected topics and things, it might be worth implementing. So we definitely want to still test it. And speaking of memory and offloading to files, this is, as I said, just saving them locally or on the server for you to use. Which means it's fully reversible because you don't lose anything. You don't lose any ongoing discussion. You just save it ready to be referred to in the future if the same student comes back and asks related questions.
And it's the expertise idea of the LLM wiki. Which, if you link with some sort of chunks, so a version of RAG, it's pretty powerful. And it makes your system become quite cheap and durable. And it's easy to inspect from both humans and agents. So it's really interesting. More specifically, it looks like this in our case. So we have chunks. We save everything into chunks. And we cross-link the chunks with pointers. And then these chunks are linked to raw data files from where they come from. Then we have one index that will just map all these chunks. So just a link to all the chunks and some context of what it is about. And the agent will just see this index.
So it sees, I think it was 450 tokens. So it's very small. And it just sees that index. And then, based on the user question, if it seems to be related to some user-specific question that may exist in the memory, it will scan the index. It will go back to the chunk. If it's enough, it will answer based on the chunk retrieved. If it's not enough, it can even go back to the raw data to have even more information. So it's just the best way to pull context according to the task complexity. So if the task is complex, you will pull more. And if it's simple, you will pull less. And a parenthesis on this:
What we've seen working with our clients and just building this in general is that right now everyone is converging towards having more and smaller skills. So it's way better to build small, very precise skills that refer to each other's skills to save on context and just load skills one by one. And even be able to spawn a sub-agent with one dedicated skill context, just to save context. It's the idea of progressive disclosure, so you just load what you need right now. Okay, just to go back now on compaction. This is what the talk is about because memory, we made some experiments, but it couldn't really fit here. It was a bit too much.
And when talking about compaction, there's an important problem or solution that appeared recently. Well, not recently, but it was way more popularized. The main problem is that, first, when you ask a follow-up question, you need to recompute all previous tokens every time. So you just end up paying twice for the same tokens, or three times, or four times if the conversation is going, which is obviously far from ideal. So what providers do nowadays is offer prompt caching. So they will save the embedding and KV cache, and I won't enter into the details, but they will precompute, they will have saved a lot of the compute for some tokens, and you can just reload them.
And what's interesting for us is that these already sent tokens that we reuse are much, much, much cheaper. Specifically, it can go up to 50 times cheaper with some API like DeepSeq, which we will discuss in the experiments. And what that means is that if you send a very long context, you will pay just 1.5 of the price per token and just pay the full price of the new user question or the new interaction. And that's a problem for compaction because when you are compacting, summarizing, doing any transformation to this context, the provider cannot use the cache because it's a new context. The model is not intelligent enough to understand it's the same topic change.
It just cannot use the cache. So you will pay full price for these new transformed tokens. So here what it means is that for compaction to be worthwhile, you need to compress the context by more than 50 times. So it can be quite difficult in some cases without losing quality. So caching is truly a game changer, especially because nowadays almost all APIs offer it, and it's very easy to use. And typically they save costs on 90% of the cost when you use caching. But in some cases, as I said, it can go up to way more than that. And not only does it help with cost, but cache tokens are also already computed. So it's way faster to get the answer back.
So ultimately, it means that summarization is potentially a trap. You may not want to use it at all, or you may want to use it very specifically. Which is what the most serious harnesses do nowadays. Cloud Code, Codex, and all of them use context caching, but also use a different method of compaction. And we know that because obviously of the leak, and then because Codex is open source. And then the APIs also provide ways to manage context directly and caching directly. So when you build yourself a harness like we do with the AI tutor, it becomes really interesting to understand when to use which technique and test them, obviously. So where is this going?
It means that you don't want to just compact. You don't want to summarize everything anytime. Because it may kill the cache. And you won't be able to use it. So some best guidance that we found is that obviously when the user seems to talk about a very different topic, you may want to refresh the session to clean it, to clear it. When you scope your files, use what I described with progressive disclosure to just show the smallest amount of context possible. You want to clear every old tool output that is not useful anymore. You may want to compact in some cases. We will see that in the experiments in a few minutes. And you may want to optimize for cache hits.
So just having the model be able to use its cache. And regarding that, providers are constantly improving their feature set to manage cache. So you just need to stay current and follow what the API allows nowadays. But they all provide different methods. And you will have access to the slides. But I put a link earlier in the first few slides on a very interesting article regarding prompt caching that I recommend checking out. It's on the sixth or seventh slide. But you will have the link to the slides. And lastly, you may want to use a model router to optimize especially cost on various tasks.
And most importantly, and what the majority of people don't do, you want to log everything. It's super easy. You want to clear every old tool outputs that are not useful anymore. You may want to compact in some cases. We will see that in the experiments in a few minutes. And you may want to optimize for cache hits. So just having the model be able to use its cache. And regarding that, providers are constantly improving their feature set to manage cache. So you just need to stay current and follow what the API allows nowadays. But they all provide different methods. And you will have access to the slides.
But I put a link earlier in the first few slides on a very interesting article regarding prompt caching that I recommend checking out. It's on the sixth or seventh slide. But you will have the link to the slides. And lastly, you may want to use a model router to optimize especially cost on various tasks. And most importantly, and what the majority of people don't do, you want to log everything. It's super easy. You just ask Cloud to implement OPIC and track everything. You don't have anything to do. So it's definitely worthwhile to implement. And you can track cache iterate. You can track user frustration, which we've seen that Cloud does.
So I keep telling it when I'm not happy. And you may want to log for some abnormally long outputs or any weird behavior that a small language model could detect. And so all of this together is context engineering, which means to decide what the model sees every time you call it. And to us, for the AI tutor, it means to decide what to keep in the current context window in order to optimize the caching. What to drop or compact and when to do that. And those two first things are exactly what we studied in many experiments that my colleague Omar will share with you right now. And so you can follow along with the experiments on the QR code link.
It's the Hugging Face Space that I mentioned earlier. And in there, there's also a link to the repo and everything. All right. So welcome to this second part of the workshop. So now that we have an overall idea of what context engineering is and what are the different techniques that we can apply to our agents, now what we want to do here is see how they actually perform in our case for our AI tutor. I will first start by describing a little bit about how the AI tutor works. So the system design, and then I will follow up with the initial experiments that we did. All right. So the AI tutor is actually very simple.
It's one agent that is a ReAct type of agent that just loops over tool calls and thinking blocks. And here we just create a very simple one using the LangChain library. So we use the create agent method with the in-memory saver because we, in this case, don't save past conversations. We just use the current history. And to customize this agent, we use the middleware feature of LangChain where we can add different features that can change the behavior of the agent at runtime. So in this case, we want to summarize, for example, and clear the outputs. Or also, in our case, have the user be able to choose the sources, like which lessons, which courses to use to answer.
To do that, we also add two different tools. So we have the first one, the retrieve tutor context, which uses a very classic hybrid search pipeline. So semantic search along with keyword search. And we combine both results to get the best possible list of chunks. And then recently, we also added the second one, which is letting the agent actually browse the file system, the knowledge base, just like we can, just like coding agents can in browsing your code base, for example. And we also borrowed from the idea of Carpathic by creating a wiki and helping the agent browse this knowledge base more easily. I will come back to these tools in a few moments.
We also add a FastAPI app to wrap the whole system with an endpoint. And we add a Next.js UI to let students use the DII Tutor. So I will talk a little bit about the first tool. So we have a large corpus. So we have all of the lessons from all the different courses that we created over the past two years. And also documentation from various public open source libraries like LangChain, Llama Index, and even documentation from OpenAI. So they make available all of the markdown files from how to use the OpenAPI, how to use Codex. And we also have Cloud Code documentation. So it's a very big corpus that has over 8 million tokens.
And of course, this cannot fit in a single context window. So we have to store it in a way where we can retrieve the most important or most relevant information.
So in this case, the agent receives a question and the user can then choose... Beforehand, the user can choose a specific source. So we can filter this knowledge base. It makes it better to improve recall, so precision, to get the most relevant information. Then we do hybrid search. So this is very classic hybrid search. We use embedding model. In this case, it's a coherent model with BM25 for the keyword index to get the most relevant top 30 chunks. Then we merge the two results from the semantic similarity and the keyword search. And we then re-rank to the top five most relevant chunks. And that's what we return to the agent.
We also have a limit of 100,000 tokens. So we don't want to... So let's say, for example, we go over. We just remove the last few chunks. The last chunks that make it so that we don't cross that threshold. So these numbers, this configuration, actually are not random. We did experiments to optimize this pipeline. I'm not going to talk about it, but it's in one of the courses that we share. Basically, we just want to try as many configurations as possible and improve recall. So we measure did we retrieve the correct page in our knowledge base, and we just choose the best settings. So like I said, in this case, it's very precise, it's very good.
But what if the agent needs to browse the whole knowledge base to get the best possible answer? So let's say I want to learn about Codex and I want to learn about Cloud Code. How can I best use those two tools? And also, if we need to have various documentation pages in context, is this the best possible tool? And there's a paper, I link it here in the slides, but it's a paper that shows that letting the agent browse the knowledge base can be very beneficial. So you can look at it if you want afterwards. Very slow to load.
So that's why we ended up creating this second tool, the run knowledge base command, where the agent can browse the knowledge base, the file system, using bash commands. So first of all, what we did first is create these three different folders. So we have the raw folder with all the different markdown files. So all the lessons from the different courses, the documentation from the different open source libraries. We also have this generated folder with, this is generated automatically, with basically just the titles of each of the markdown files. So it's easier for the agent to find relevant information.
And then we also have the wiki where, in this case, it was Cloud Code that created this. It basically reads all of the raw files in the folder, in the raw folder, and it creates a very concise set of files. So topics, frameworks, from the different sources. So for example, I can have a topic related to fine tuning. So in this case, the agent would be able to find all of the different raw files related to fine tuning, for example. Here, we let the agent only read. So this is when we actually deploy it. So this is the one you can try on the space. Here, the agent can only read the knowledge base. It cannot modify it.
So we only allow these bash commands that basically cannot modify the file system. We also put some limits around this. So for example, if a command lasts over eight seconds, we can just return an error or let the agent execute something else because it's taking too much time. And we also cap the tool outputs to 40,000 characters. So in this case, for example, if a lesson is over this amount, what the agent can do is then, OK, so I just got the first 40,000. Let me do a follow-up command to get the last piece of the lesson, for example. We also limit the number of commands.
We actually never see the agent go over this limit of 20 commands per turn, but this is just a fallback in case it takes too much time to answer. And we also sandbox, of course, the agent to only browse the knowledge base, this specific folder. Like I said, we create this offline. So the three different folders, we create this once. Or every time we want to add a new course, for example, we tell Cloud Code, can you add this to the raw folder and also create new topics around this new content. So we do this once and then we deploy it so that the agent can browse it. With our current system prompt, we can see that it's used almost every time.
So for almost 90% of the turns, we can tweak this to make it use it less or more. We didn't optimize for this specifically. And I guess the most interesting aspect of doing this is that we measured the precision, the recall of using this tool, actually turning it off. And we also sandbox, of course, the agent to only browse the knowledge base, this specific folder. I said, we create this offline. So the three different folders, we create this once. Or every time we want to add a new course, for example, we tell Cloud Code, can you add a new, can you add this to the raw folder and also create new topics around this new content.
So we do this once and then we deploy it so that the agent can browse it. With our current system prompt, we can see that it's used almost every time. So for almost 90% of the turns, we can tweak this to make it use it less or more. We didn't optimize for this specifically. And I guess the most interesting aspect of doing this is that we measured the precision, the recall of using this tool, actually turning it off. And we actually got the same amount of recall. So just using the first tool was enough to get all the relevant information. And just having this second tool was just 50% slower. It's faster without this tool because it does fewer tool calls.
And yeah, we didn't see any improvement on answering with the correct documentation. It was fun to add, but we didn't see any benefit. And one reason for that is that we tested using real-world, the questions we get from students. And those questions weren't complex enough, we guess, to actually benefit from using this new setup. So this is how initially our AI tutor managed its context. We started with this because it looked fast, it looked good. We didn't actually measure anything. It was just, oh, it looks good. Okay, we will just set it like this. And so we have these three different context engineering techniques
where we clear the outputs after 5,000 tokens, but we at least keep the last five ones. So this was a way to keep a very small context during a conversation. We also added the capacity to summarize. So after 30,000 tokens, the system summarizes the history, but we at least keep the last 20 messages to make sure that those messages are very accurate. And we also have the source preference. So that just allows people to choose what sources to use when the tutor answers. But I said these are unproven defaults, and we actually want to know what actually works best. So I guess you can, Louis showed this at the beginning, but you can access the tutor live.
I'm just going to show this very quickly. This is the Hugging Face Space. So this is a separate UI just to show you. We also have the chat bubble one on the lessons, on the course themselves. And here on the left, you can choose the different sources to use, enable them or disable them. And then you can send your query.
So as you can see here, I can just send a new request. And we see Gemini, in this case, Gemini 3.5 Flash, use its reasoning and use its capacity to do tool calling and to answer. And I guess here there's a, I think it's just the internet bugging. But we should have a response. And yeah, the code is open source, so you can use your favorite coding agent to explore the code base and learn about how we implemented this specifically. So we have the different activity, the activity that the model did. So for tool calls, it thought four times and it used 10 sources, and we have the final answer. So every time the student uses this chatbot, we actually log everything.
So for every single turn, we have the input tokens, the output tokens, how many of them were cached, what was the cost, what time it took to get the first token, how many tool calls it did, and if the system actually did something around summarization. So this is very useful. And that's what we are going to use when measuring the different techniques. So why do we need to measure? Because the techniques that we showed all sound very smart. So you might think that they are very useful, but sometimes they're not. And as we discovered with our experiments, actually, it might be detrimental. So because of the way APIs cache the tokens when they are sent.
And it's also difficult to know in advance what is best to use. So before I go into the experiments, I just want to define a few words, because I'm going to use these words throughout the presentation. So a preset is the way the AI tutor was set up in the experiments. So, for example, it can be in this preset, we did summarization at this amount of tokens, or in this preset, we used sliding window, for example.
So that's what the preset is. We have the different tasks. So a task type, for these initial experiments, we only did two tasks, a single turn and multiple turns, also called sessions.
And one run is just running one preset on a task, and then you get the run. And then the bundle is just the result. So it's just a JSON file with all the different metrics that we save to the disk. So first task, single turn. So these are question and answers. And we didn't generate this. It's not synthetic. I had Codex scrape all of the questions and responses from our website, where students can ask questions and get answers from members of the staff. And that's how I got this initial, we got this initial data set. We cleaned the data set and only used 60 pairs, because we saw that some of the questions weren't good for the type of task.
For example, there were old questions about previous versions of some libraries and right now, if the tutor answers, it's not going to be using this old version of the library. So we just removed some of the questions, some duplicate ones, and we got this first data set. And what we measure is the retrieval. So did we retrieve the correct, did the tutor retrieve the correct lesson? For example, this is done automatically. We can see just by looking at the code, did we use the correct lesson or not? And we also look at, do we have the correct facts in the answer or the right kind of response in the answer? These two are actually graded using an LLM.
You can use APIs to do it, but right now I think the best way to do it is to use your code subscription or your Codex subscriptions because it's cheaper than using the APIs. We also have the second task, the session. So multi-turn conversations back and forth. Here, what we want to know is, is the AI tutor able to recall facts after multiple turns? So this is a bit, in this case, we do use some generated content. So we generate facts that we put at the beginning. So we have a student, like a fake student, state a fact. So, for example, I want to learn about RAG. And then we stuff the conversation with filler messages because we just want to have a lot of messages.
And then we have a probe, which is just the student asking a question again. So, for example, it can be, what should I learn today? And since the, I'm going to go through it in the next slide, but let's, the example. So, for example, here at turn one, we have the student, the fake student, state a fact. So the student wants to learn about the RAG evaluation. And that's basically the fact. Then we just add a lot of messages, filler messages. And then, at some point, we have the student say, for example, what topic should I learn about today? And then what we expect the AI tutor to say is that the student should learn about RAG evaluation.
So more specifically hit rate and MRR, for example. We also have a gate part in the evaluation where, let's say we are testing the summarization technique. We actually just want to know, did summarization actually happen or not. So this is one example of one task in this session, that asset in the sessions task. Now we have the evaluation hardness. So the main hardness, I guess the main function is the run, the run task, the run battery function that just runs this task. We have the grading, like I said, it can either be a code check. So did we retrieve the correct lesson or not? I learn about today?
And then what we expect the AI tutor to say is that the student should learn about reg evaluation. So more specifically, heat rate and MRR, for example. We also have a gate part in the evaluation where, let's say, we are testing the summarization technique. We actually just want to know, did summarization actually happen or not? So this is one example of one task in this session, that asset in the session's task. Now we have the evaluation hardness. So the main hardness, I guess, the main function is the run task, the run battery function that just runs this task.
We have the grading. Like I said, it can either be a code check. So, did we retrieve the correct lesson or not? And we can also have the LLM as a judge. And in this case, we use the subscription of code code. We also have the check triggers aspect. So this is just a check to see if the evaluation went good or not. Did we actually compact or not? This is just to make sure that the run is actually good and we can save it. And then we have a generated report to see what was the latency, what was the time to first token, and every metric that we can measure. So yeah, we can evaluate everything and then grade it afterwards. We can run it once and grade it afterwards.
So what we run, so we run 11 presets, and we change them for each experiment. So we have the full history. So these are the main ones, the full history. So this is the case where we don't touch the context. We leave everything as is in the history. And then we also have the production that I showed at the beginning, the defaults that we have. So these are the reference points. And then we have these six techniques that I want to compare. So sliding window from compression, selective retention, and the other ones. And what I want to see is just what memory recall do I get if I keep everything else fixed?
So I use the same model, the same prompt, the same tools, the same data set. What's the difference? Now, just doing this was a bit expensive. I didn't expect this to get over $500, but it did. And that's one of the reasons we did follow-up experiments afterwards using cheaper models. But my colleague Sam Brady will talk about this. So what preset actually won? And so these are our results. And as you can see, we weren't expecting this, but basically not touching the context was actually the best strategy for recovering this fact over time, over multiple messages. You can see that the production, so the defaults that we thought were good enough, were actually not the best.
Not doing anything is actually better. We did two different experiments where one was just one trial and the second was two trials. So we have more of a statistic. So I guess it's better, but the numbers might not be accurate because it's just one trial and two trials. But I guess the most interesting thing is just the order in which the techniques ended up being in the table. So in this case, keeping everything wins on the memory side. But what about the cost? This is what the production cost was for the single turn and the session. So almost 50 cents for a single turn and 24 cents for the multi-turn for each turn.
We actually had very good memory recall for pretty much all the techniques in the single turn, because in single turn, you don't have enough tokens to actually fire up the different strategies. So for one response, you don't need to do summarization, compaction, or anything like that. So that's why you see high numbers. But as you can see, after the multi-turn task, you can see that the quality degraded to 38%. And if we compare this to the full history, so here we don't touch the context, the history of the model, we can see that not touching is actually cheaper, it's faster, and we have better recall overall. So keeping everything wins on all of these three fronts.
So why is it actually, why do we have less latency? We wanted to understand that. And it's basically because if you remove the tool outputs consistently, then the agent needs to re-retrieve afterwards information it already had. So you're just making the agent do more tool calls. And that's why it ended up costing more and using more tokens and having less memory recall. So these are the results we initially got using Gemini 3.5 with this data set of 11 to 13 turns. It's not huge. There's not many messages. That's why we now want, in the follow-up part, to do more different experiments that my colleague, Samaridi, will show you. So yeah, let me introduce you to Samaridi.
So this is going to be my part. And as we ended on the note where Umar just said that it cost us almost $600 to run the evals that we ran, one question we were trying to answer with this extended evaluation was, so when does compaction actually matter? Or should you actually compact or not? Because we clearly saw that when we have the full answers in the window, it works really well, but it costs a lot of money. So I tried to do this evaluation in three sorts of contexts. The first one was cash chats. So, when you're chatting with Gemini, you'd be able to see that is the cash chat option. The second version is going to be document plus tools.
So if you are pasting a long document in the AI tutor, or if there's just tool output, what happens then? And finally, if you go local or if you scale this evaluation, how well is it going to work out? So, should you ever compact? So before my part, Umar just showed that on Gemini 3.5 Flash, keeping everything won. But why did we come up with that question? It was because full history on a frontier model like Gemini would be very, very expensive. So with all this extended experiment, we are trying to figure out whether it was actually worth it or not. But before we start that, I wanted to just talk about the different contexts we see in our AI tutor application.
So the first one is going to be a long chat history where we have a long chat, but all of the details that the students are asking or they're chatting about can get buried in it. The second is going to be a pasted document. So we do have a lot of students who are going to just copy-paste a lot of documents there. And also because we have a limited context window, how does that fit? And then it's going to be different tools that we use internally. But this is just going to be a bunch of logs that the tools have. And since these are different contexts, each of them needs some sort of different fix. So we had to evaluate it for all different contexts that we had.
So the first obvious thing, looking at the cost, we were like, okay, let's try out a cheaper model and see, does it do better? What sort of techniques work on that? Does compaction work on it or not? So DeepSeq V4 Flash was an obvious option. And we were like, we will try this out on this and see how well it works out. And then it's going to be different tools that we use internally. But this is just going to be a bunch of logs that the tools have. And since these are different contexts, each of them needs some sort of different fix. So we had to evaluate it for all different contexts that we had.
So the first obvious thing, looking at the cost, we were like, okay, let's try out a cheaper model and see, does it do better? What sort of techniques work on that? Does compaction work on it or not? So DeepSeq V4 Flash was an obvious option. And we were like, we will try this out on this and see how well it works out. And you can see there is a drastic cost difference between Gemini and DeepSeq here. So here we can see there's a drastic difference between cost when we checked the performance on DeepSeq V4. And the main reason was that we were getting a cash discount. So the cash discount on DeepSeq was 50x as compared to Gemini.
And even in this setting, we saw that keeping all of the context still won. So we were just getting the best performance in that case for DeepSeq as well. But the last thing from this experiment was that we figured out a cheaper model. But now we wanted to see that even though keeping everything makes it cheaper, does it remember better? Does keeping all of the context remember all of the details the student might be asking us? So this is how I tested out the memory of the model. So we have conversations within our system where students are asking questions about their setup, about the errors that they're seeing, whatever they've already tried.
So I just took these chats and asked questions about specific details just to see if the model is able to figure that out, if it is able to give me an output for that or not. And the results I saw were that 95% of the time, the model was able to give us the right exact details that I was trying to look for. And even when it is keeping all of the details. Whereas if I summarize first or if I compact the context I had, it only gave me the answer back 32% of the time. And if you think about it, the reason for that is that when you summarize, you remove all of the necessary details. So you're not able to keep all of those details, and the model keeps on missing those out.
So we are able to see that keeping everything wins in terms of cost. If you have a model like DeepSeq, then it also is remembering things. So it's not that if you have a long conversation, it is not able to remember things. So it is correct 95% of the time. So the next thing we wanted to see is the cost part of it. How does it actually, does it cost most? Does it cost the least? What happens in terms of tokens? So on DeepSeq, we saw the setup that was sending the most tokens is actually the cheapest to run.
So the full history setup that we had was sending the most tokens, but we were still getting the best results out of it because 97% of the tokens that we had were cached. And as we just saw in the previous slide, cached tokens are really cheap. They are charged separately. So summarizing works the other way, and every turn it makes the model read and write new tokens. And this is something that I ran on a 36-turn conversation, and it was about 1.78 million tokens. And keeping everything still came out ahead. And so this answers the question that it wasn't just Gemini. It works the same on DeepSeq as well.
So the cheaper model is what brought the cost down, but the results are the same. And in respect of this, the next thing I wanted to ask is what happens when the conversations really grow? Because all of these were tested out on short conversations. If I really increase the length of the conversation, how does that work out? So when I tried to run this entire experiment on a longer context, I saw that even if I'm pulling one specific detail out of it, the model performed really well. So on the top part, the green line, those are all of the distinctive, distinctive facts that the model is able to find out.
So we can see that up until 800k tokens as well, the model was not missing out on those facts. It was giving me good and consistent results for some ambiguous facts. The performance dropped to half of what I observed on the distinctive facts. But overall, for our AI tutor, this result was really good. So we saw that the model, even when you're not compacting anything, holds really well, even if you have a really long conversation. But so far, everything that I've talked about is only per turn. So every cost has been per turn, but does it help when we scale it? Because a chatbot is not something that is a one-turn situation.
The tutor is a long-term service where students are asking questions on a massive scale. So let's say if we have 100,000 to a million turns of questions every day, what would be the cost on DeepSeek? For DeepSeek, the cost was approximately somewhere from 18,000 to 180,000 a month. And even though we don't see that sort of volume as of now, paying per token starts to add up. And so the alternative for that was going over to a local model, just to see if we are getting the same sort of performance on a local model or not. And so one way we were thinking was that because local models cache as well, can we use the same sort of setup on a local model?
And will we get the same result? So because we had hardware limitations, we just tested it out on a MacBook. The maximum context window that we could go up to was 32K. And so we thought that can we do that locally now? But we can't because the sort of lessons that we had are bigger than a 32K context window on their own. And once the conversation doesn't fit in the window, caching was no longer helpful for us. And we have to make the context smaller either by compressing it or by retrieving only the parts that we need. So the next question we were trying to answer is, once you have to compact locally, what actually works?
So for the chat memory, going local and trying to keep everything stops winning because you can't keep everything. And a simple question here would be why can't you just keep increasing the length of the model? Why can't you use a bigger model? Because, of course, we have hardware limitations. For us, it was a MacBook, but GPUs also have hardware limitations. But we went from a 7B model, 8B model, to a 32B model. So, the next question we were trying to answer is, can we, once you have to compact locally, what actually works? So, for the chat memory, going local and trying to keep everything stops winning because you can't keep everything.
And a simple question here would be, why can't you just keep increasing the length of the model? Why can't you use a bigger model? Because, of course, we have hardware limitations. For us, it was a MacBook, but GPUs also have hardware limitations. But we went from a 7B model, 8B model to a 32B model. But here we landed on the conclusion that even though you keep increasing the length of the model, it is not going to increase your context window. You cannot repair that part. You have to make a choice there. But this is only for the chat history, and what happens when you are trying to deal with documents locally.
So, this was a little surprising because if you are retrieving results with local documents, it was really good. We got 100% accuracy in that case. So, when students are pasting something which is too big, rag is a good option there. You can use that and it can help you retrieve the exact data that you're looking for. Also, the processing time in this case was anywhere from 25 to 65 seconds. So, which is pretty good in terms of the output that we are getting. But if you are trying to stuff the window with more context than you have, we saw that it took us approximately 340 seconds to get the output when our conversations were really long.
And the output that we got was a single token. So, you are not getting anything, but you're also wasting a lot of time when you are trying to do it. So, you have to be careful about what option you choose in this case. So, you also have different types of retrieval strategies that you could use. The default retrieval is, of course, semantic search, where you're just trying to match the meaning of the text. And that is the dense rag heading that you can see on the chart. It mostly works, but we tried to make it work from 50K token to 200K. And we saw that dense rag worked really well. It was 80%.
But when we increased it to 400K tokens, it was not able to fax that were buried in the middle. And it started giving us zero percent recall. Whereas something like BM25, it still got 100% every time. So, semantic search on its own is not enough. And that's why, when Umar talked about our setup in the AI tutor, we're actually using a hybrid search. We're using a mix of both dense and BM25. We're using a combination of both those. So, after all of this, we came up with a lot of data. We had one other question, which was, how does all of this local setup compare to cloud? Because that's the real way we'll see the result.
So we wanted to put them side by side just to see what is the output. And for chat, the local setup was not up to the part of the cloud setup. On cloud, keeping everything scores somewhere from 92 to 95%. But locally, it was stuck at 33%. And the context window was a limitation here. Also, you can see that local models actually work because there is no cost. Like, you cannot see because you already own the hardware. Though there's a throughput limitation there. But if you use a technique like retrieval, you get good accuracy even on a local setup.
So, in our case, what we found is that on memory, keeping the whole chat recalled about 95% of the details that we were providing it. It was able to give us correct answer 95% of the time. Versus it was just 32% if you summarize it. On long context, finding a single fact is easy for the model. We went up to 800k tokens and we did not see any sort of context rot in that case. On cost per turn, we saw that the cheapest run is actually the one which is sending the most tokens. Because caching makes resending the same context very cheap. And on no cost at scale, it scales up for, let's say, if we have 1,000 students. Gemini costs us about $40,000 a month.
Whereas DeepSeq, DeepSeq was around $1,900 a month. So, going local saves us on cost a bit more. So the main thing to take away is that do not compact by default. You have to name the constraint that you have and then look for a better alternative. So what did we finally decide after all of these different experiments that we ran? So we decided on DeepSeq, V4, Flashcase. We had hardware limitations. So for us, the cloud structure worked out really well. It is also the cheapest considering the current intake of students we have. So, it works out well for us. And we are using, on top of that model, we are using a mix of, we're using hybrid retrieval to get good results.
For memory, we have chosen to keep everything. We have got a default limit that after 30K tokens, we are gonna have compaction. But up until that, we're planning to keep everything. And that is the tutor setup that we have. So, because we had time limitations, these are all the evaluations and experiments that I could pack into this time. But if you would like to learn more about these evaluations, or you would want to build a tutor yourself, this is the full-stack AI engineering course on academy.towardsai.net. So you can go to this link and go through the course. But thank you so much, everyone. And now we can take any questions. Thank you. Thank you.
But up until that, we're planning to, like, keep everything. And, uh, you know, that is the tutor setup that we have. So, you know, because we had, like, time limitations, so these are all the evaluations and experiments that I could pack into this time. But if you would like to, you know, learn more about these evaluations, or you would want to build a tutor yourself, um, this is the, uh, full-stack AI engineering course, uh, on academy.towardsai.net. Um, so you can go to this link and, uh, you know, go through the course. Uh, but thank you so much, everyone. Uh, and now we can take any questions. Thank you. Thank you.