Open Reader

Stop Chunking Like It's 2022 — Yuval Belfer, AI21 Labs

completed 18:00 Sep 16, 2026 Watch on YouTube

Current Status

completed

Video ID

r9OwPx_HoV0

RAG / Chat

Enabled
Stop Chunking Like It's 2022 — Yuval Belfer, AI21 Labs
Description

Ask a corpus of Seinfeld transcripts for the name of Jerry's favorite church and a 100 token chunk returns it at rank one, while every larger window buries it below rank fifty. Ask the same corpus who Jerry calls his nemesis and pure evil and the small chunks fail completely, because the answer is spread across a scene rather than sitting in a sentence. Same data, same index, opposite requirements. Yuval Belfer uses that pair to make a claim most retrieval teams have quietly assumed away: there is no correct chunk size, because the correct size is a property of the query, and you pick it at indexing time when you do not yet have any queries. That is the trap in one sentence. At indexing you control the window and know nothing about the questions. At retrieval you have the question and the window is already frozen. To size the cost, his team duplicated several datasets at six different chunk sizes and ran an oracle experiment, choosing per query the size that happened to work best. The gap between that oracle and any single fixed choice ran 20 to 40 percent of recall, which is what an arbitrary 512 has been quietly costing. Their fix refuses the premise rather than tuning it. Index the corpus at every window size, query all of them, and because chunks of different sizes cannot be compared, return whole documents so the rankings become commensurable, then merge them by reciprocal rank fusion. It is a short script rather than a model. The honest accounting is at the end: two to five times the memory, and almost no added latency. Speaker info: - https://x.com/yuvalinthedeep - https://linkedin.com/in/yuval-belfer Timestamps: 0:00 - A talk about nothing, and why chunking 1:43 - Indexing is boring, retrieval tuning is fun 3:23 - A World Cup directory that cannot answer the query 4:17 - Chunking as lossy compression 6:13 - Six copies of the same dataset 7:07 - Two Seinfeld questions, opposite answers 8:12 - The oracle experiment 9:59 - An information problem at both ends

Summary

Generated by claude-sonnet-4-5-20250929

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Chunking strategy is query-dependent, not corpus-dependent; multiscale indexing (duplicating databases at multiple chunk sizes and merging results via RRF) yields 20-40% recall improvement over fixed chunk sizes with minimal latency cost.
  • Why it matters: Ken's agent systems and OpenClaw orchestration depend on retrieval quality; this addresses a foundational infrastructure problem (chunk size selection) that most teams set once and forget, leaving 20-40% performance on the table. Directly applicable to any RAG-based workflow or control plane needing to query structured/unstructured data.
  • Best use: Watch the oracle experiment explanation and RRF merging logic; use as blueprint for upgrading retrieval infrastructure in Ken's systems where context window usage and recall quality matter. Consider for any agent workflow using vector DBs or file-based retrieval.

Executive Summary

Yuval Belfer from AI21 Labs argues that chunking is not dead despite recent discourse claiming RAG is obsolete. The core problem is that fixed chunk sizes are a lossy compression that cannot adapt to different query types. Focused queries (e.g., 'What is Jerry's favorite shirt?') need small chunks to avoid noise, while broader queries (e.g., 'Who is Jerry's nemesis?') require larger chunks to capture context. No single chunk size dominates across query types, creating an information asymmetry: at indexing time you control chunk size but don't know future queries; at retrieval time you know the query but chunk size is already fixed.

AI21 ran experiments duplicating datasets (QMSum, NarrativeQA, Seinfeld trivia, FinanceBench) at six different chunk sizes (50, 100, 200, etc. tokens). They built an oracle that selected the optimal chunk size per query, showing 20-40% recall improvement over any fixed size. This gap represents the cost of choosing an arbitrary fixed size like 512 tokens. The oracle is not deployable (requires ground truth) but proved the potential.

Their solution: multiscale indexing. Duplicate the database n times with different chunk sizes, query all copies in parallel at retrieval time, and merge rankings using Reciprocal Rank Fusion (RRF). RRF works because they retrieve full documents (not just chunks), making rankings across chunk sizes comparable. This approach matched or beat the best fixed size across all datasets, achieving 20-40% improvements. Cost: 2-5x memory overhead (constant factor) but no latency penalty since retrieval is parallelizable and RRF is a lightweight script.

Future work includes determining optimal number and values of chunk sizes (currently arbitrary), exploring better merging algorithms beyond RRF, and applying this to larger-scale systems. The method is infrastructure-level, requires no sophisticated ML, and is immediately applicable to any RAG or file-retrieval system. Yuval released example code and the Seinfeld dataset.

Key Takeaways

  • Claim: Chunk size choice is query-dependent, not corpus-dependent; no fixed chunk size dominates across all query types in a dataset. | Evidence: Experiments on Seinfeld dataset showed 'What is Jerry's favorite shirt?' ranked first with 100-token chunks but below rank 50 with larger chunks, while 'Who does Jerry describe as his nemesis?' failed with small chunks but succeeded with large chunks. Blue lines (fixed chunk sizes) intersected across datasets, meaning no single size won everywhere. | Implication: Optimizing chunk size per corpus is a false optimization; you must account for query distribution at runtime or accept 20-40% recall loss.
  • Claim: An oracle selecting optimal chunk size per query yields 20-40% recall improvement over best fixed chunk size. | Evidence: Oracle experiment on four datasets (QMSum, NarrativeQA, Seinfeld, FinanceBench) showed orange oracle line 20-40% above best blue fixed-size line in recall@k graphs. The gap represents performance left on the table by arbitrary fixed sizes like 512. | Implication: There is massive headroom in retrieval quality if you can approximate oracle behavior at runtime without needing answers. | Caveat: Oracle is not deployable (requires ground truth answers); it only proves potential, not a production system.
  • Claim: Multiscale indexing (duplicate database at n chunk sizes, query all, merge via RRF) matches or beats best fixed chunk size with 20-40% gains. | Evidence: Tested on QMSum, NarrativeQA, Seinfeld, FinanceBench, and MTab. Heat maps showed multiscale (bottom row) consistently greener (higher recall) across recall@1 through recall@10. MTab results showed 10-40% improvement depending on dataset. | Implication: If Ken's systems can afford 2-5x storage (often feasible for vector DBs or file systems), this is a plug-and-play infrastructure upgrade with no latency penalty. | Caveat: Cost is 2-5x memory overhead (constant factor) to store n copies of the database. Method requires parallelizable retrieval infrastructure.
  • Claim: Reciprocal Rank Fusion (RRF) enables merging rankings from different chunk sizes by retrieving full documents instead of chunks. | Evidence: Instead of merging chunk-level rankings (incomparable across sizes), retrieve full documents for each chunk hit, producing n document rankings. RRF is a simple formula (not a model) that aggregates rankings like voting. Tested multiple merging methods; RRF worked best. | Implication: This is a non-ML, low-complexity solution; Ken can implement RRF as a lightweight script without retraining or model deployment.
  • Claim: Latency is not significantly affected because retrieval across n databases can be parallelized and RRF is fast. | Evidence: RRF is 'just a simple script that takes really no time.' Retrieval calls can run in parallel across n duplicates. No additional per-query indexing required. | Implication: Multiscale indexing is a memory-for-quality tradeoff with minimal latency cost, making it viable for production systems where recall matters more than storage.
  • Claim: Chunk size selection (50, 100, 200, etc.) was arbitrary; optimal number and values of chunk sizes are unknown. | Evidence: Yuval explicitly stated the choice of 50, 100, 200 tokens was 'pretty arbitrary, to be honest.' Future work includes determining how many chunk sizes are needed and which values to use. | Implication: Ken should treat this as a starting heuristic, not a final recipe. There is likely room to optimize n and chunk size distribution per use case. | Caveat: Current results are proof-of-concept; production systems may need fewer or different chunk sizes for cost/performance balance.
  • Claim: Agentic search did not kill retrieval; it only killed retrieval tuning (e.g., playing with top-k, hybrid search per query). | Evidence: Yuval argued agentic search (greps, ls, finds) is useful but not sufficient at scale with diverse queries. Retrieval tuning (post-query optimization) is easy and fun, but indexing strategy (pre-query) is harder and ignored. Agentic search does not solve the chunking problem when data is not well-organized or queries span multiple contexts. | Implication: For Ken's systems, agentic search is a complement to retrieval infrastructure, not a replacement. Chunking strategy remains a foundational concern for large-scale data.

Detailed Brief

Experimental Design and Oracle Setup

  • Claims: Datasets used: QMSum (meeting transcripts), NarrativeQA (novels Q&A), Seinfeld trivia (in-house, published), FinanceBench, MTab.; Each dataset duplicated six times with chunk sizes 50, 100, 200, 500, 1000, 2000 tokens.; Oracle experiment: for each query, select the chunk size that gives best recall; graph shows oracle (orange) vs. fixed sizes (blue).
  • Evidence: Seinfeld dataset is 'trivia about nothing' (transcripts of Seinfeld episodes), built in-house, available via blog link.; Oracle line shows 20-40% improvement; blue lines intersect, meaning no fixed size dominates.
  • Caveats: Oracle is not a production method; it requires knowing the answer to each query in advance.
  • Implications: The oracle proves the ceiling of performance gains from dynamic chunk size selection; any practical method that approximates it will capture a fraction of this 20-40% gain.

Multiscale Indexing Implementation Details

  • Claims: Index n copies of database at different chunk sizes.; At retrieval, run n parallel queries (one per chunk size copy).; Retrieve full documents (not just chunks) to make rankings comparable across chunk sizes.; Merge n document rankings using RRF (reciprocal rank fusion), a non-ML voting formula.
  • Evidence: RRF formula tested against multiple merging methods; RRF performed best.; Example code and blog post published by AI21.
  • Implications: Implementation is straightforward: duplicate data, parallelize queries, run RRF script. No retraining or model tuning required.

Cost and Tradeoffs

  • Claims: Memory overhead: 2-5x (constant factor) to store n database copies.; Latency: negligible if retrieval is parallelizable; RRF is fast.; No inference cost increase (no model calls added).
  • Evidence: Yuval stated 'if you think about it, latency-wise, it doesn't really affect that because you can do all the retrieval part parallelly.'
  • Caveats: Requires infrastructure capable of parallel queries across n vector DBs or file indexes.; Memory cost may be prohibitive for very large corpora unless storage is cheap.
  • Implications: For Ken's systems, evaluate whether 2-5x storage is acceptable tradeoff for 20-40% recall improvement. Likely yes if storage is cheap relative to LLM inference or if recall quality is critical (e.g., agent decision-making, compliance, high-stakes retrieval).

Future Work and Open Questions

  • Claims: How many chunk sizes (n) are optimal? Currently arbitrary.; Which chunk size values should be used? 50/100/200/etc. was arbitrary.; Can merging algorithms better than RRF be found?; Can this generalize to other retrieval methods (e.g., hybrid search, sparse+dense)?
  • Evidence: Yuval noted 'we do need to figure out how to compute this and how to know how many copies exactly do you need.'
  • Implications: This is research-grade work; production deployment will require per-system tuning of n and chunk size distribution. Ken should expect to run experiments to find the right balance for OpenClaw or other systems.

Notable Concepts & Terms

  • Multiscale indexing: Duplicating a database at n different chunk sizes, querying all copies in parallel, and merging results to approximate oracle-level performance without knowing the answer in advance.
  • Reciprocal Rank Fusion (RRF): A lightweight, non-ML formula for merging multiple rankings of the same items (here, documents) by treating each ranking as a vote. Aggregates n retrieval results from different chunk sizes.
  • Oracle experiment: A theoretical upper bound experiment where, for each query, the system cheats by selecting the optimal chunk size using ground truth answers. Used to measure the potential of dynamic chunk size selection.
  • Lossy compression (chunking): Chunking is inherently lossy: large chunks lose nuance/meaningful embeddings, small chunks lose context. No fixed size avoids loss for all query types.
  • Query-dependent chunk size: The optimal chunk size varies by query type (focused vs. broad), not by corpus. This invalidates corpus-level chunk size optimization.
  • Information asymmetry (indexing vs. retrieval): At indexing time you control chunk size but don't know queries; at retrieval time you know queries but chunk size is fixed. Multiscale indexing resolves this by deferring the choice to retrieval.
  • Agentic search vs. retrieval tuning: Agentic search (greps, ls, finds, dynamic queries) killed retrieval tuning (post-query optimizations like top-k, hybrid search), but did not solve the indexing/chunking problem at scale.

Operator Notes / Why Ken Should Care

  • Evaluate whether 2-5x storage overhead is acceptable for Ken's agent systems (e.g., OpenClaw, GTM systems, workflow engines) in exchange for 20-40% recall improvement. If storage is cheap relative to LLM inference or if retrieval quality is a bottleneck, this is a high-ROI infrastructure upgrade.
  • Implement multiscale indexing as a quick proof-of-concept: duplicate one vector DB or file index at 3-5 chunk sizes (e.g., 100, 250, 500, 1000 tokens), parallelize queries, merge via RRF, and benchmark recall improvement on representative queries.
  • If deploying multiscale indexing, start with 3-4 chunk sizes (e.g., 100, 250, 500, 1000) as a heuristic and run experiments to tune n and values per use case. Do not treat 50/100/200 as gospel; Yuval noted it was arbitrary.
  • Monitor whether RRF is sufficient or whether a learned merging model could improve results. RRF is lightweight and good enough for initial deployment, but future optimization may require domain-specific tuning.
  • Consider this method for any RAG-based workflow, agent control plane, or file-system retrieval in Ken's stack where recall quality matters more than storage cost. Especially relevant for compliance, high-stakes decision-making, or orchestration workflows that depend on accurate context retrieval.
  • Review AI21's blog post and example code (released with this talk) for implementation details. Yuval also released the Seinfeld dataset, which may be useful for benchmarking retrieval systems on trivia-style queries.
  • Do not assume agentic search eliminates the need for chunking strategy. Agentic search is a query-time optimization; chunking is an indexing-time foundational choice. For large-scale or diverse-query systems, both are needed.
  • If Ken's systems already use fixed chunk sizes (e.g., 512 tokens with 10-20% overlap), treat this talk as a signal that 20-40% recall improvement is on the table with minimal engineering complexity. Prioritize this if retrieval quality is currently a pain point.

Source/Metadata

  • Title: Stop Chunking Like It's 2022 — Yuval Belfer, AI21 Labs
  • Transcript words: 4686
  • Duration seconds: 1080
  • Timestamp note: Timestamps provided in transcript (MM:SS format); used for key takeaways.

Transcript

2715 words en Processed in 114.4s

Hi, everybody. Thank you for coming today. Welcome to a talk about retrieval. My name is Yuval. I work at AI21, which is an AI research lab. And today I want to talk to you about something that most people don't want to talk about, which is chunking. And I hope to convince you by the end that chunking isn't dead and there is something to do with it. And really, if you are at X, LinkedIn, wherever, you've probably seen that RAG is dead, right? I think people also killed MCP lately. And RAG is dead again. Long live the agentic retrieval, agentic search. And there comes a time where you have to ask yourself, how many times can RAG die? Right? And even when someone says, well, RAG isn't dead, like Jerry, the CEO of Llama Index, they still have to kill something. And apparently this something is chunking. Don't invest in it. Don't do it. And this is the reason that people said that chunking is dead, because everybody is using agentic search now, right? You have greps, you have ls, you have finds. All of these are great, but these are still not enough if you have a lot of data and you have various amounts of queries. And I think that the main reason that a lot of people don't like to talk about chunking is because it's not the fun part, right? In every RAG or file system, we have two stages. The first stage is the boring one. The one you do in the beginning: you have a lot of data, you have to preprocess it, you have to decide on the chunk size, and then you have to store everything in a vector DB. The other part is the retrieval part, essentially the one that happens per query. This is something which is much easier to do, right? It's much easier to optimize. You can use all your queries, and then you can play with the max K, top K. You can play with hybrid search, maybe. Much more fun to do retrieval tuning, right? So I will claim that if we have to kill something, if something has to be dead, then it's probably retrieval tuning. And yes, agentic search probably killed that. But still, agentic search, even if we can accept the fact that it killed retrieval tuning, it's still not good enough when you have a lot of scale, a lot of data. It costs a lot of money. I don't think I have to mention that anymore. Token maxing is something that everybody's talking about. And the thing underneath, which is if the data itself is not ordered in the right way, in your folders, in your directories, you still get something which is inefficient. So let's try to think of a timely example, right? The FIFA World Cup is now. And let's imagine that we have a dataset that contains all the FIFA World Cups. So every directory is the 1998 one, the 2002 one, and so on. But if your query asks which team won the most World Cups, you can't just go to a folder and ask that. You have to go to every folder, see who won, and then aggregate this together, which is very inefficient. The answer, by the way, is Brazil, I hope, at least according to the time that this conversation is happening. So retrieval didn't actually die. We're not killing anything in this lecture. It got devoted into plumbing. And I think that everybody who worked on any RAG system knows the feeling. Day one or week one or maybe even month one, if you're very thorough, you're picking some sort of a chunk size. Let's say 512. And maybe you're probably putting some overlap, right? 10, 20%, indexing everything and forget all about it. And you can, right, we talk a lot about the fixed chunking strategies where if your chunk is something which is too big, right, so you get the whole picture, which is nice, but you're losing a lot of the nuance, and all the chunks will not get meaningful embeddings. Where if you choose your chunks to be too small, you're getting the big picture lost. And it won't be as efficient. So what this tells us is that chunking is essentially a lossy compression. No matter what we're doing, we're losing something. And I will claim that there is no right chunk size. And a lot of you who worked on data will say, no, but we have this corpus, we have this dataset, and we really optimized our system to work very well on this data. And we thought so too. We had a lot of experience with different types of agents and systems and workflows. You think about benchmarks, how easy it is to overfit your model to a benchmark. But not with RAG. It doesn't happen there. And you cannot really optimize it per dataset, and I will claim that it is query dependent. And how can I be so sure? How can I claim such a thing? Because we ran experiments and we tested, and now I'm going to present it to you. So what we did, instead of saying what is the best chunk size per data, let's actually take a dataset and duplicate this dataset several times. In this case, six times. In every duplication, in every instance, the chunk size is different. So we have a database with a chunk size of 2,000, a database with a chunk size of 1,000, and so on. And we did it with several datasets: QMSum, which is a meeting transcript dataset; Narrative QA, which is question answering on novels; and Seinfeld dataset, which is trivia about the transcripts of Seinfeld. It's a trivia dataset that we built in-house. We also published it if anybody wants the link at the end. And we tested on all of them to see what happens. And first of all, we just wanted to see, for every dataset, which chunk size is the best. And what we're seeing here is an example from the Seinfeld dataset, where essentially two queries, which are different by nature, get different results based on that chunk size. So the first question: what is the name for Jerry's favorite shirt? You can see this is a very focused question, very specific question. The answer to it is probably very contained, and this is something that a smaller chunk size will do best in. And you can see rank one versus rank below 50, between 100 tokens fixed chunk size to 100. Whereas a question like, who does Jerry describe as his nemesis and pure evil, which I'm not that big of a Seinfeld fan, and I know it's Newman, but if you look at the transcript, it's not something you can find easily. And you can see that it really changes, right? If you use small chunk size, you will not get the answer. And what we did to really, after we ran all of these things and we noticed that, we said, what if we had an oracle, or a genie, if you want, that can tell us for every query what is the best chunk size to do retrieval for? This essentially is the oracle experiment. This is what we wanted to know to see the potential. This is not right, we already have the answers, so we're not actually building a system here. We just want to see what is the potential that we have here. And what you can see here, okay, in this graph, all the blue, first of all, the y-axis is the recall, higher is better, the x-axis is the number of retrieved chunks, so it's recall at k versus k. You can see all the blue lines are probably indistinguishable, but each of them is the performance for a fixed chunk size, whereas the orange one is the oracle line. This is, for every query, we took the best one out of these. And you can see it happens across several datasets. In a lot of them, you can actually see that the blue lines intersect with each other, meaning that indeed for a lot of the datasets, no chunk size actually dominates. And what's more interesting is that there is a lot of potential. The gap, which you can see between the orange line and all the blue lines, is big. And when I say big, it's something like 20 to 40% just from doing strategy on chunking, and very simple strategy. And this gap, this is what the choice of 512, or 1,000, or whatever, right? This number is just arbitrary. This is what it costs you. And I think that the problem here is it's a bit tricky because it's an information problem that we don't have the information that we need at every stage. And what do I mean by that? If I'm looking at the indexing part, where I do have control over the chunk size, I don't know what the queries will be. I can guess. I can estimate. I can try. But I don't know what the queries will be, so I cannot adjust my chunk size accordingly. And the retrieval part, where I do have my queries, I cannot control the chunk size, right? It's already fixed, and I obviously will not do the entire process per query from the beginning. So we looked at prior works such as Entropic, contextual retrieval, where they enrich every chunk, and others that essentially try to improve the latent space of every chunk, but this is not the direction that we went. All of them just stayed in the model of let's work with a fixed chunk size, whereas we took a different approach, and we said, why commit to one where we can commit to several? And we call it multiscale indexing. Essentially, we're doing what we've seen before. We're checking the database, we duplicate it and chunk it with several chunk sizes or window sizes, and then at retrieval time, we are querying all of them. So if we had n duplicates of the database, n window sizes, we now have to run n different retrieval calls per query. And how do we combine them? We obviously cannot use the oracle, right? The oracle is something that we have just for potential. In real life, we don't know the answer. But what we can do is find some sort of merging algorithm. Now, you would say, when we look at it like this, what can be the issue? The fact that we have n rankings, but the rankings are for chunks, and chunks with different sizes are not really comparable, right? So instead, we opted to do something which is pretty popular these days, and a lot of the RAG systems actually work like this: instead of just retrieving the chunk, when we're getting a chunk, we're retrieving the entire document, right? When context window grows, we want to give more and more context. And now, in this case, we have n rankings of the same documents because they're not chunks anymore, and this we can compare. And in this case, you can think of retrieval as essentially just voting. We have n different ranks of the relevant documents, and we want to aggregate them all into one. That's why we are using something called RRF, reciprocal rank fusion, which is pretty much a simple formula. We tried several things. This worked the best. And as you can see, it's not a model. It's not something that you have to do specifically. This is just a simple script that takes really no time. And this is how the full system looks like. So we have the indexing n times. Then we query each query from every database, and we're using RRF to combine them all. And the results, you can guess that they're good. Otherwise, I would not be standing here being way too confident. But you can see we tested across several datasets: QMSum, Narrative QA, Seinfeld, and also FinanceBench. We took all of them, and it matches or beats the best fixed size. Let's see it in a graph. Every row here is chunk size, so you can see 50, 100, and so on. The bottom row is our method, the one that combines all of them. And every column is recall at something. So recall at one, two, three, up until 10. What you can see here is two things, right? First of all, across recall at whatever, our method still wins, which you can think is very easy, but the fact that you have to combine all of them is not something which is very trivial. And also, you can see that the quality actually increases. The heat map, where you can see it becomes much greener. And here you can see all four of the datasets where we do achieve better results, really quite like 20, 30, 40% even in a lot of the things. Also, there are results that I did not show you here which are on MTab. You can see in our blog, I will put the link later. We're getting there also a lot of improvements, somewhere between 10 to 40%, depending on the dataset. Now, I'm not naive. I'm not going to claim here that this costs nothing. Obviously, there is a cost, right? No free lunch. Everything has to come with something. And yes, this costs with extra memory. It costs something between two to five, a constant of additional memory where you have to keep all of those copies of the database. However, if you think about it, latency-wise, it doesn't really affect that because you can do all the retrieval part in parallel. And also, the RRF part doesn't really take a lot of time. I will say that this was a very nice research project that we did and we got really cool results. There are things to do, right? There are places to improve. There is future work to do. More precisely, we want to understand how many chunk sizes do we want and which. The fact that we worked with 50, 100, 200, and so on was pretty arbitrary, to be honest. So we do need to figure out how to compute this and how to know how many copies exactly you need. Also, go beyond RRF, right? The fact that we're using RRF is because it worked the best from the methods that we used, but it doesn't mean that there is no better method. And if I need to leave you with something, I would say that agents didn't kill retrieval. Nothing died. It's just infrastructure. And the bad part is that it's infrastructure from 2022. And with really simple methods, you can take your RAG system or anything that has to do with storing data and then retrieve it with 20 to 40 percent improvement, without anything too sophisticated. So if you want to hear more about it, read more about it, you can read the blog. There is also example code there and the Seinfeld dataset. And that's it. I'm Yuval. Thank you so much for being here. you're picking some sort of a chunk size. Let's say 512. And maybe you're probably putting some overlap, right? 10, 20%, so on, indexing everything and forget all about it. And you can, right, we talk a lot about the fixed chunking strategies where if your chunk is something which is too big, right, so you get the whole picture, which is nice, but you're losing a lot of the nuance, and all the chunks will not get meaningful embeddings, where if you will choose your chunks to be too small, you're getting the big picture lost. And really it won't be as efficient. So what this tells us is that chunking is essentially a lossy compression. No matter what we're doing, we're losing something. And I will claim that there is no right chunk size. And a lot of you who worked on data will say, no, but we have this corpus, we have this data set, and we really used and we optimized our system to work very, very well on this data. And we thought so too. We had a lot of experience with it, with a lot of different types of agents and systems and workflows that you can really, and right, you think about benchmarks, how easy it is to overfit your model to a benchmark. But not with RUG. It doesn't happen there. And you cannot really optimize it per data set, and I will claim that it is query dependent. And how can I be so sure? How can I claim such a thing? Because we ran experiments, and we tested, and now I'm going to present it to you. So what we did, instead of saying what is the best chunk size per data, let's find out, let's actually take a data set and duplicate this data set several times. In this case, six times. In every duplication, in every instance, the chunk size is different. So we have a database with a chunk size of 2,000, a database with a chunk size of 1,000, and so on, and so on. And we did it with several data sets, so QMSum, which is a meeting transcript data set, narrative QA, which is question answering on novels, and Seinfeld data set, which is a trivia about nothing. Not really. It's trivia questions about the transcripts of Seinfeld. It's kind of a trolling data set that we build in-house. We also published it if anybody wants the link at the end. And we tested on all of them to see what happens. And first of all, we just wanted to see, for every data set, which chunk size is the best. And what we're seeing here is an example from the Seinfeld data set, where essentially two queries, which are different by nature, get different results based on that chunk size. So the first question, what is the name for Jerry's favorite shirt? You can see this is a very focused question, very specific question. The answer to it is probably very contained, and this is something that a smaller chunk size will do best in. And you can see rank one versus rank below 50, between 100 tokens fixed chunk size to 100. Whereas a question like, who does Jerry describe as his nemesis and pure evil, which I'm not even that big of a Seinfeld fan, and I know it's Newman, but if you look at the transcript, it's not something you can find that easily. And you can see that it really changes, right? If you use small chunk size, you will not get the answer. And what we did to really, after we ran all of these things and we've noticed that, we said, what if we had an oracle, or a genie, if you want, that can tell us for every query what is the best chunk size to do retrieval for? This essentially is the oracle experiment. This is what we wanted to know to see the potential. This is not, right, we already have the answers, so we're not actually building a system here. We just want to see what is the potential that we have here. And what you can see here, okay, in this graph, all the blue, first of all, the y-axis is the recall, higher is better, the x-axis is the number of retrieved chunks, so it's recall at k versus k. You can see all the blue lines, probably indistinguishable, but each of them is the performance for a fixed chunk size, whereas the orange one is the oracle line. This is, for every query, we took the best one out of these. And you can see it happens across several datasets. In a lot of them, you can actually see that the blue lines intersect with each other, meaning that indeed for a lot of the datasets, no chunk size actually dominates. And what's more interesting is that there is a lot of potential. The gap, which you can see between the orange line and all the blue lines, is big. And when I say big, it's something like 20 to 40% just from doing strategy on chunking, and very simple strategy. May I add? And this is, like, this gap, this is what the choice of 512, or 1,000, or whatever, right? This number is just arbitrary. This is what it costs you. And I think that the problem here is, like, it's a bit tricky because it's kind of like an information problem that we don't have the information that we need at every stage. And what do I mean by that? If I'm looking at the indexing part, where I do have control over the chunk size, I don't know what the queries will be. I can guess. I can maybe estimate. I can try. But I don't know what the queries will be, so I cannot adjust my chunk size accordingly. And the retrieval part, where I do have my queries, I cannot control the chunk size, right? It's already fixed, and I obviously will not do the entire process per query from the beginning. So we looked at prior works such as notably entropic, contextual retrieval, where they enrich every chunk, and others that essentially try to improve the latent space of every chunk, but this is not the direction that we went. All of them just stayed in the model of let's work with a fixed chunk size, whereas we took a different approach, and we said, why commit to one where we can commit to several? And we call it the multiscale indexing. Essentially, we're just doing what we've seen before. So we're checking the database, we duplicate it and chunk it with several chunk sizes or window sizes, and then, sorry, and then this is what happens at the indexing, and then at retrieval time, we are querying all of them. So if we had n duplicates of the database, n window sizes, we now have to run six different retrieval calls per query. Sorry, six as n. And how do we combine them? We obviously cannot use the oracle, right? The oracle is something that we have just for potential. In real life, we don't know the answer, but what we can do is to find some sort of merging algorithm. Now, you would say, when we look at it like this, what can be the issue? The fact that we have n ranking, but the rankings are for chunks, and chunks with different sizes are not really comparable, right? So instead, we opted to do something which is pretty popular these days, and a lot of the RUG systems actually work like this, that instead of just retrieving the chunk, when we're getting a chunk, we're retrieving the entire document, right? When context window grows, we want to give more and more context. And now, in this case, we have n, right, n rankings of the same documents because they're not chunks anymore, and this we can compare. And in this case, you can think of retrieval as essentially just voting. All right, so it's not purely ranking. We don't have run ranking, and then we're doing it re-rank. We're having n different ranks of the relevant documents, and we want to aggregate them all into one. That's why we are using something called RRF, reciprocal rank fusion, which is pretty much a simple formula. We tried several things. This worked the best. And as you can see, it's not a model. It's not something that you have to do specifically. Like, especially, this is just a simple script that takes really no time. And this is how the full system looks like. So we have the indexing n times. Then we query each query from every database, and we're using RRF to combine them all. And the results, you can guess that they're good. Otherwise, I would not be standing here and being way too much confident. All right, but you can see we tested across several data sets, QMSum, NarrativeQA, Seinfeld, and also FinanceBench. We took all of them, and it matches the best, or bits the best fixed size. Let's see it in a graph. It's a bit hard to see here, so I'll walk it slowly. Every row here is chunk size, so you can see 50, 100, and so on. The bottom row is our method, this one, the one that you do from all of them and then combine. And every column is recall at something. So recall at one, two, three, up until 10. What you can see here is that two things, right? First of all, that across recall at whatever, our method still wins, which you can think is very easy, but the fact that you have to combine all of them is not something which is very trivial. And also, you can see that the quality actually increases. The heat map, where you can see it's become much greener. And again, this was just something that I wanted to show in large. Here you can see all four of the data sets where we do achieve better results, really quite like 20, 30, 40% even in a lot of the things. Also, there are results that I did not show you here which are on MTab. You can see in our blog, I will put the link later. We're getting there also a lot of improvements, somewhere between 10 to 40%, depending on the data set. Now, I'm not naive. I'm not going to claim here that this costs nothing. Obviously, there is a cost, right? No free launches. Everything has to come with something. And yes, this costs with extra memory. It costs something between two to five, two, all of one, right? A constant of additional memory where you have to keep all of those copies of the database. However, if you think about it, latency-wise, it doesn't really affect that because you can do all the retrieval part parallelly. And also, the RRF part doesn't really take a lot of time. I will say that this was a very nice research project that we did and we got really, really cool results. There are things to do, right? There are places to improve. There are future work to do. More precisely, we want to understand how many chunk sizes do we want and which. The fact that we worked with 50, 100, 200, and so on was pretty arbitrary, to be honest. So we do need to figure out how to compute this and how to know how many copies exactly do you need. Also, go beyond RRF, right? The fact that we're using RRF is because it worked the best from the methods that we used, but it doesn't mean that there is no better method. And if I need to leave you with something, I would say that agents didn't kill retrieval. Nothing died. Come on. It's just infrastructure. And the bad part is that it's infrastructure from 2022. And with really simple methods, you can take your RUG system or anything that has to do with storing data and then retrieve it with 20 to 40 percent. Again, without any something too sophisticated. So if you want to hear more about, read more about it, you can read the blog. There is also an example code there and the Seinfeld dataset. And that's it. I'm Yuval. Thank you so much for being here. Woo! My name is Yuval. My name is Yuval. you