Hi, everybody. Thank you for coming today. Welcome to a talk about retrieval. My name is Yuval. I work at AI21, which is an AI research lab. And today I want to talk to you about something that most people don't want to talk about, which is chunking. And I hope to convince you by the end that chunking isn't dead and there is something to do with it. And really, if you are at X, LinkedIn, wherever, you've probably seen that RAG is dead, right? I think people also killed MCP lately. And RAG is dead again. Long live the agentic retrieval, agentic search. And there comes a time where you have to ask yourself, how many times can RAG die?
Right? And even when someone says, well, RAG isn't dead, like Jerry, the CEO of Llama Index, they still have to kill something. And apparently this something is chunking. Don't invest in it. Don't do it. And this is the reason that people said that chunking is dead, because everybody is using agentic search now, right? You have greps, you have ls, you have finds. All of these are great, but these are still not enough if you have a lot of data and you have various amounts of queries.
And I think that the main reason that a lot of people don't like to talk about chunking is because it's not the fun part, right? In every RAG or file system, we have two stages. The first stage is the boring one. The one you do in the beginning: you have a lot of data, you have to preprocess it, you have to decide on the chunk size, and then you have to store everything in a vector DB.
The other part is the retrieval part, essentially the one that happens per query. This is something which is much easier to do, right? It's much easier to optimize. You can use all your queries, and then you can play with the max K, top K. You can play with hybrid search, maybe. Much more fun to do retrieval tuning, right?
So I will claim that if we have to kill something, if something has to be dead, then it's probably retrieval tuning. And yes, agentic search probably killed that. But still, agentic search, even if we can accept the fact that it killed retrieval tuning, it's still not good enough when you have a lot of scale, a lot of data. It costs a lot of money. I don't think I have to mention that anymore. Token maxing is something that everybody's talking about. And the thing underneath, which is if the data itself is not ordered in the right way, in your folders, in your directories, you still get something which is inefficient.
So let's try to think of a timely example, right? The FIFA World Cup is now. And let's imagine that we have a dataset that contains all the FIFA World Cups. So every directory is the 1998 one, the 2002 one, and so on. But if your query asks which team won the most World Cups, you can't just go to a folder and ask that. You have to go to every folder, see who won, and then aggregate this together, which is very inefficient. The answer, by the way, is Brazil, I hope, at least according to the time that this conversation is happening.
So retrieval didn't actually die. We're not killing anything in this lecture. It got devoted into plumbing. And I think that everybody who worked on any RAG system knows the feeling. Day one or week one or maybe even month one, if you're very thorough, you're picking some sort of a chunk size. Let's say 512. And maybe you're probably putting some overlap, right? 10, 20%, indexing everything and forget all about it.
And you can, right, we talk a lot about the fixed chunking strategies where if your chunk is something which is too big, right, so you get the whole picture, which is nice, but you're losing a lot of the nuance, and all the chunks will not get meaningful embeddings. Where if you choose your chunks to be too small, you're getting the big picture lost. And it won't be as efficient. So what this tells us is that chunking is essentially a lossy compression. No matter what we're doing, we're losing something. And I will claim that there is no right chunk size.
And a lot of you who worked on data will say, no, but we have this corpus, we have this dataset, and we really optimized our system to work very well on this data. And we thought so too. We had a lot of experience with different types of agents and systems and workflows. You think about benchmarks, how easy it is to overfit your model to a benchmark. But not with RAG. It doesn't happen there. And you cannot really optimize it per dataset, and I will claim that it is query dependent. And how can I be so sure? How can I claim such a thing? Because we ran experiments and we tested, and now I'm going to present it to you.
So what we did, instead of saying what is the best chunk size per data, let's actually take a dataset and duplicate this dataset several times. In this case, six times. In every duplication, in every instance, the chunk size is different. So we have a database with a chunk size of 2,000, a database with a chunk size of 1,000, and so on. And we did it with several datasets: QMSum, which is a meeting transcript dataset; Narrative QA, which is question answering on novels; and Seinfeld dataset, which is trivia about the transcripts of Seinfeld. It's a trivia dataset that we built in-house. We also published it if anybody wants the link at the end. And we tested on all of them to see what happens.
And first of all, we just wanted to see, for every dataset, which chunk size is the best. And what we're seeing here is an example from the Seinfeld dataset, where essentially two queries, which are different by nature, get different results based on that chunk size. So the first question: what is the name for Jerry's favorite shirt? You can see this is a very focused question, very specific question. The answer to it is probably very contained, and this is something that a smaller chunk size will do best in. And you can see rank one versus rank below 50, between 100 tokens fixed chunk size to 100.
Whereas a question like, who does Jerry describe as his nemesis and pure evil, which I'm not that big of a Seinfeld fan, and I know it's Newman, but if you look at the transcript, it's not something you can find easily. And you can see that it really changes, right? If you use small chunk size, you will not get the answer.
And what we did to really, after we ran all of these things and we noticed that, we said, what if we had an oracle, or a genie, if you want, that can tell us for every query what is the best chunk size to do retrieval for? This essentially is the oracle experiment. This is what we wanted to know to see the potential. This is not right, we already have the answers, so we're not actually building a system here. We just want to see what is the potential that we have here.
And what you can see here, okay, in this graph, all the blue, first of all, the y-axis is the recall, higher is better, the x-axis is the number of retrieved chunks, so it's recall at k versus k. You can see all the blue lines are probably indistinguishable, but each of them is the performance for a fixed chunk size, whereas the orange one is the oracle line. This is, for every query, we took the best one out of these. And you can see it happens across several datasets.
In a lot of them, you can actually see that the blue lines intersect with each other, meaning that indeed for a lot of the datasets, no chunk size actually dominates. And what's more interesting is that there is a lot of potential. The gap, which you can see between the orange line and all the blue lines, is big. And when I say big, it's something like 20 to 40% just from doing strategy on chunking, and very simple strategy. And this gap, this is what the choice of 512, or 1,000, or whatever, right? This number is just arbitrary. This is what it costs you.
And I think that the problem here is it's a bit tricky because it's an information problem that we don't have the information that we need at every stage. And what do I mean by that? If I'm looking at the indexing part, where I do have control over the chunk size, I don't know what the queries will be. I can guess. I can estimate. I can try. But I don't know what the queries will be, so I cannot adjust my chunk size accordingly. And the retrieval part, where I do have my queries, I cannot control the chunk size, right? It's already fixed, and I obviously will not do the entire process per query from the beginning.
So we looked at prior works such as Entropic, contextual retrieval, where they enrich every chunk, and others that essentially try to improve the latent space of every chunk, but this is not the direction that we went. All of them just stayed in the model of let's work with a fixed chunk size, whereas we took a different approach, and we said, why commit to one where we can commit to several?
And we call it multiscale indexing. Essentially, we're doing what we've seen before. We're checking the database, we duplicate it and chunk it with several chunk sizes or window sizes, and then at retrieval time, we are querying all of them. So if we had n duplicates of the database, n window sizes, we now have to run n different retrieval calls per query. And how do we combine them? We obviously cannot use the oracle, right? The oracle is something that we have just for potential. In real life, we don't know the answer. But what we can do is find some sort of merging algorithm.
Now, you would say, when we look at it like this, what can be the issue? The fact that we have n rankings, but the rankings are for chunks, and chunks with different sizes are not really comparable, right? So instead, we opted to do something which is pretty popular these days, and a lot of the RAG systems actually work like this: instead of just retrieving the chunk, when we're getting a chunk, we're retrieving the entire document, right? When context window grows, we want to give more and more context. And now, in this case, we have n rankings of the same documents because they're not chunks anymore, and this we can compare.
And in this case, you can think of retrieval as essentially just voting. We have n different ranks of the relevant documents, and we want to aggregate them all into one. That's why we are using something called RRF, reciprocal rank fusion, which is pretty much a simple formula. We tried several things. This worked the best. And as you can see, it's not a model. It's not something that you have to do specifically. This is just a simple script that takes really no time. And this is how the full system looks like. So we have the indexing n times. Then we query each query from every database, and we're using RRF to combine them all.
And the results, you can guess that they're good. Otherwise, I would not be standing here being way too confident. But you can see we tested across several datasets: QMSum, Narrative QA, Seinfeld, and also FinanceBench. We took all of them, and it matches or beats the best fixed size. Let's see it in a graph. Every row here is chunk size, so you can see 50, 100, and so on. The bottom row is our method, the one that combines all of them. And every column is recall at something. So recall at one, two, three, up until 10.
What you can see here is two things, right? First of all, across recall at whatever, our method still wins, which you can think is very easy, but the fact that you have to combine all of them is not something which is very trivial. And also, you can see that the quality actually increases. The heat map, where you can see it becomes much greener.
And here you can see all four of the datasets where we do achieve better results, really quite like 20, 30, 40% even in a lot of the things. Also, there are results that I did not show you here which are on MTab. You can see in our blog, I will put the link later. We're getting there also a lot of improvements, somewhere between 10 to 40%, depending on the dataset.
Now, I'm not naive. I'm not going to claim here that this costs nothing. Obviously, there is a cost, right? No free lunch. Everything has to come with something. And yes, this costs with extra memory. It costs something between two to five, a constant of additional memory where you have to keep all of those copies of the database. However, if you think about it, latency-wise, it doesn't really affect that because you can do all the retrieval part in parallel. And also, the RRF part doesn't really take a lot of time.
I will say that this was a very nice research project that we did and we got really cool results. There are things to do, right? There are places to improve. There is future work to do. More precisely, we want to understand how many chunk sizes do we want and which. The fact that we worked with 50, 100, 200, and so on was pretty arbitrary, to be honest. So we do need to figure out how to compute this and how to know how many copies exactly you need. Also, go beyond RRF, right? The fact that we're using RRF is because it worked the best from the methods that we used, but it doesn't mean that there is no better method.
And if I need to leave you with something, I would say that agents didn't kill retrieval. Nothing died. It's just infrastructure. And the bad part is that it's infrastructure from 2022. And with really simple methods, you can take your RAG system or anything that has to do with storing data and then retrieve it with 20 to 40 percent improvement, without anything too sophisticated. So if you want to hear more about it, read more about it, you can read the blog. There is also example code there and the Seinfeld dataset. And that's it. I'm Yuval. Thank you so much for being here. you're picking some sort of a chunk size. Let's say 512.
And maybe you're probably putting some overlap, right? 10, 20%, so on, indexing everything and forget all about it. And you can, right, we talk a lot about the fixed chunking strategies where if your chunk is something which is too big, right, so you get the whole picture, which is nice, but you're losing a lot of the nuance, and all the chunks will not get meaningful embeddings, where if you will choose your chunks to be too small, you're getting the big picture lost. And really it won't be as efficient. So what this tells us is that chunking is essentially a lossy compression. No matter what we're doing, we're losing something.
And I will claim that there is no right chunk size. And a lot of you who worked on data will say, no, but we have this corpus, we have this data set, and we really used and we optimized our system to work very, very well on this data. And we thought so too. We had a lot of experience with it, with a lot of different types of agents and systems and workflows that you can really, and right, you think about benchmarks, how easy it is to overfit your model to a benchmark. But not with RUG. It doesn't happen there. And you cannot really optimize it per data set, and I will claim that it is query dependent. And how can I be so sure? How can I claim such a thing?
Because we ran experiments, and we tested, and now I'm going to present it to you. So what we did, instead of saying what is the best chunk size per data, let's find out, let's actually take a data set and duplicate this data set several times. In this case, six times. In every duplication, in every instance, the chunk size is different. So we have a database with a chunk size of 2,000, a database with a chunk size of 1,000, and so on, and so on. And we did it with several data sets, so QMSum, which is a meeting transcript data set, narrative QA, which is question answering on novels, and Seinfeld data set, which is a trivia about nothing. Not really.
It's trivia questions about the transcripts of Seinfeld. It's kind of a trolling data set that we build in-house. We also published it if anybody wants the link at the end. And we tested on all of them to see what happens. And first of all, we just wanted to see, for every data set, which chunk size is the best. And what we're seeing here is an example from the Seinfeld data set, where essentially two queries, which are different by nature, get different results based on that chunk size. So the first question, what is the name for Jerry's favorite shirt? You can see this is a very focused question, very specific question. The answer to it is probably very contained,
and this is something that a smaller chunk size will do best in. And you can see rank one versus rank below 50, between 100 tokens fixed chunk size to 100. Whereas a question like, who does Jerry describe as his nemesis and pure evil, which I'm not even that big of a Seinfeld fan, and I know it's Newman, but if you look at the transcript, it's not something you can find that easily. And you can see that it really changes, right? If you use small chunk size, you will not get the answer. And what we did to really, after we ran all of these things and we've noticed that, we said, what if we had an oracle, or a genie, if you want, that can tell us for every query
what is the best chunk size to do retrieval for? This essentially is the oracle experiment. This is what we wanted to know to see the potential. This is not, right, we already have the answers, so we're not actually building a system here. We just want to see what is the potential that we have here. And what you can see here, okay, in this graph, all the blue, first of all, the y-axis is the recall, higher is better, the x-axis is the number of retrieved chunks, so it's recall at k versus k. You can see all the blue lines, probably indistinguishable, but each of them is the performance for a fixed chunk size, whereas the orange one is the oracle line.
This is, for every query, we took the best one out of these. And you can see it happens across several datasets. In a lot of them, you can actually see that the blue lines intersect with each other, meaning that indeed for a lot of the datasets, no chunk size actually dominates. And what's more interesting is that there is a lot of potential. The gap, which you can see between the orange line and all the blue lines, is big. And when I say big, it's something like 20 to 40% just from doing strategy on chunking, and very simple strategy. May I add? And this is, like, this gap, this is what the choice of 512, or 1,000, or whatever, right? This number is just arbitrary.
This is what it costs you. And I think that the problem here is, like, it's a bit tricky because it's kind of like an information problem that we don't have the information that we need at every stage. And what do I mean by that? If I'm looking at the indexing part, where I do have control over the chunk size, I don't know what the queries will be. I can guess. I can maybe estimate. I can try. But I don't know what the queries will be, so I cannot adjust my chunk size accordingly. And the retrieval part, where I do have my queries, I cannot control the chunk size, right? It's already fixed, and I obviously will not do the entire process per query from the beginning.
So we looked at prior works such as notably entropic, contextual retrieval, where they enrich every chunk, and others that essentially try to improve the latent space of every chunk, but this is not the direction that we went. All of them just stayed in the model of let's work with a fixed chunk size, whereas we took a different approach, and we said, why commit to one where we can commit to several? And we call it the multiscale indexing. Essentially, we're just doing what we've seen before. So we're checking the database, we duplicate it and chunk it with several chunk sizes or window sizes, and then, sorry, and then this is what happens at the indexing,
and then at retrieval time, we are querying all of them. So if we had n duplicates of the database, n window sizes, we now have to run six different retrieval calls per query. Sorry, six as n. And how do we combine them? We obviously cannot use the oracle, right? The oracle is something that we have just for potential. In real life, we don't know the answer, but what we can do is to find some sort of merging algorithm. Now, you would say, when we look at it like this, what can be the issue? The fact that we have n ranking, but the rankings are for chunks, and chunks with different sizes are not really comparable, right? So instead, we opted to do something
which is pretty popular these days, and a lot of the RUG systems actually work like this, that instead of just retrieving the chunk, when we're getting a chunk, we're retrieving the entire document, right? When context window grows, we want to give more and more context. And now, in this case, we have n, right, n rankings of the same documents because they're not chunks anymore, and this we can compare. And in this case, you can think of retrieval as essentially just voting. All right, so it's not purely ranking. We don't have run ranking, and then we're doing it re-rank. We're having n different ranks of the relevant documents, and we want to aggregate them all into one.
That's why we are using something called RRF, reciprocal rank fusion, which is pretty much a simple formula. We tried several things. This worked the best. And as you can see, it's not a model. It's not something that you have to do specifically. Like, especially, this is just a simple script that takes really no time. And this is how the full system looks like. So we have the indexing n times. Then we query each query from every database, and we're using RRF to combine them all. And the results, you can guess that they're good. Otherwise, I would not be standing here and being way too much confident. All right, but you can see we tested across several data sets,
QMSum, NarrativeQA, Seinfeld, and also FinanceBench. We took all of them, and it matches the best, or bits the best fixed size. Let's see it in a graph. It's a bit hard to see here, so I'll walk it slowly. Every row here is chunk size, so you can see 50, 100, and so on. The bottom row is our method, this one, the one that you do from all of them and then combine. And every column is recall at something. So recall at one, two, three, up until 10. What you can see here is that two things, right? First of all, that across recall at whatever, our method still wins, which you can think is very easy, but the fact that you have to combine all of them
is not something which is very trivial. And also, you can see that the quality actually increases. The heat map, where you can see it's become much greener. And again, this was just something that I wanted to show in large. Here you can see all four of the data sets where we do achieve better results, really quite like 20, 30, 40% even in a lot of the things. Also, there are results that I did not show you here which are on MTab. You can see in our blog, I will put the link later. We're getting there also a lot of improvements, somewhere between 10 to 40%, depending on the data set. Now, I'm not naive. I'm not going to claim here that this costs nothing.
Obviously, there is a cost, right? No free launches. Everything has to come with something. And yes, this costs with extra memory. It costs something between two to five, two, all of one, right? A constant of additional memory where you have to keep all of those copies of the database. However, if you think about it, latency-wise, it doesn't really affect that because you can do all the retrieval part parallelly. And also, the RRF part doesn't really take a lot of time. I will say that this was a very nice research project that we did and we got really, really cool results. There are things to do, right? There are places to improve. There are future work to do.
More precisely, we want to understand how many chunk sizes do we want and which. The fact that we worked with 50, 100, 200, and so on was pretty arbitrary, to be honest. So we do need to figure out how to compute this and how to know how many copies exactly do you need. Also, go beyond RRF, right? The fact that we're using RRF is because it worked the best from the methods that we used, but it doesn't mean that there is no better method. And if I need to leave you with something, I would say that agents didn't kill retrieval. Nothing died. Come on. It's just infrastructure. And the bad part is that it's infrastructure from 2022. And with really simple methods,
you can take your RUG system or anything that has to do with storing data and then retrieve it with 20 to 40 percent. Again, without any something too sophisticated. So if you want to hear more about, read more about it, you can read the blog. There is also an example code there and the Seinfeld dataset. And that's it. I'm Yuval. Thank you so much for being here. Woo!
My name is Yuval. My name is Yuval. you