RAG is dead, right?? — Kuba Rogut, Turbopuffer
Description
Cursor added semantic search and measured a 24% increase in answer accuracy on their composer model, a 2.6% gain in code retention in large codebases, and a 2.2% drop in dissatisfied user requests. Those numbers look small until you factor in that semantic search does not fire on every query. Meanwhile Google search volume for RAG hit a new inflection point in mid 2025 and went through the roof. The Twitter "RAG is dead" discourse and the actual usage curve are moving in opposite directions. Kuba Rogut's argument is that the problem was never retrieval, it was the narrow definition of it. RAG is not just a vector search call. It is vector search, full text search, glob, regex, and filters used iteratively by an agent that keeps searching until it has what it needs. He contrasts Claude Code (grep per session, no index, repeat cost every run) with Cursor (one time upfront indexing, lightweight tool calls at runtime). Claude Code's approach is not wrong, it is a deliberate tradeoff. The frame that clarifies it: embeddings are cached compute, and whether to cache depends on query volume. Jeff Dean's version: you do not need a trillion tokens at once, you need the right million. Speaker info: - socials - related links
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Traditional one-shot RAG (vector search → LLM) is being replaced by iterative, multi-tool agentic retrieval that combines semantic search, full-text search, and file-system tools to progressively narrow context for better performance and cost efficiency.
- Why it matters: Demonstrates architectural shift from simple embedding-retrieval patterns to agent-driven, cached-compute paradigms that dramatically improve answer accuracy (up to 24% in Cursor's case) and reduce token costs through strategic, multi-step retrieval.
- Best use: Watch fully to understand why leading AI code tools architect retrieval as iterative agent workflows rather than single vector searches, especially the Cursor indexing strategy and Jeff Dean's 'right million tokens' principle.
Executive Summary
Kuba Rogut from TurboPuffer argues that simple RAG—embedding content once and querying via vector search—is obsolete for serious agentic systems. Despite Twitter discourse claiming 'RAG is dead,' Google search volume for RAG spiked mid-2025, indicating growing sophistication rather than abandonment. The real evolution is from one-shot retrieval to agentic search: agents equipped with multiple tools (vector search, BM25 full-text, grep, regex, filters) that iteratively fetch and reason over context until they satisfy task requirements.
Cursor, one of TurboPuffer's first customers, exemplifies this approach. They index codebases using Merkle trees to share embeddings across teams, avoiding redundant chunking/embedding when 99% of code is identical. Their semantic search yields a 12.5-13.5% average accuracy boost (24% for Composer model), 2.6% code retention increase, and 2.2% decrease in dissatisfied requests in A/B tests. These gains occur despite semantic search being invoked only on applicable queries, not every request.
Claude Code (Anthropic's tool) initially used local vector DBs but found them insufficient, opting instead for per-session file-system grep. However, this costs ~6,000 tokens per sub-step per session, repeating the same discovery work across developers and days. Cursor's cached-compute model—upfront indexing cost amortized over many queries—proves faster and cheaper at runtime. Jeff Dean's principle 'you don't need a trillion tokens at once, you need the right million' encapsulates the staged-retrieval philosophy: lightweight mechanisms progressively narrow trillions of tokens to the relevant subset for context windows.
The shift is from 'retrieval as a single DB call' to 'retrieval as iterative agent reasoning.' Sophisticated systems now perform multiple searches (semantic, full-text, filters) across steps, fetching only what each step requires. This unlocks new product experiences (faster Composer 2, better code understanding) and fundamentally changes vector DB usage patterns—from one-time queries to orchestrated, multi-tool retrieval pipelines.
Key Takeaways
- Claim: Twitter's 'RAG is dead' narrative contradicts actual search volume, which spiked mid-2025, showing growing sophistication in retrieval rather than obsolescence. | Evidence: Google search volume graph displayed showing RAG interest doubling from mid-2024 plateau to mid-2025 inflection point. | Caveat: Social media discourse often reflects attention cycles rather than technical adoption; search volume is a lagging indicator and doesn't distinguish between novice interest and advanced practice. | Implication: Ken should interpret 'RAG is dead' as shorthand for 'naive one-shot RAG is dead'—the actual trend is toward more complex, agent-orchestrated retrieval, not less retrieval. | Timestamp: timestamp unavailable
- Claim: Cursor's Merkle-tree-based code embedding strategy yields 12.5-13.5% average answer accuracy improvement, up to 24% for Composer model, via semantic search. | Evidence: Cursor blog post data cited: 24% accuracy boost for Composer on internal benchmark; 2.6% code retention increase and 2.2% dissatisfied-request decrease in online A/B test. Merkle trees allow Cursor to copy shared embeddings across team members' identical codebases, re-embedding only changed files. | Caveat: Gains measured on Cursor's internal benchmark, not public. A/B test improvements (2.6%, 2.2%) appear small because semantic search isn't invoked for every query—only applicable ones benefit. | Implication: Ken should note that upfront indexing infrastructure (Merkle trees, TurboPuffer) pays off in accuracy and retention. Small percentage gains at scale (millions of users) translate to major product differentiation and revenue. The 'not every query needs it' caveat suggests intelligent routing/tool-selection is critical. | Timestamp: timestamp unavailable
- Claim: Claude Code abandoned local vector DBs in favor of per-session file-system grep, but this costs ~6,000 tokens per discovery sub-step and repeats work across sessions. | Evidence: Boris Cherny (founding engineer) tweet stating early Claude Code used RAG/local vector DB but found it didn't work. Kuba's trace comparison: Claude Code grep-read-assess loop costs 6,000 tokens per sub-step; 10 agents across 10 days repeating same question incur 10x that cost. | Caveat: No public Claude Code token cost data provided; 6,000 token figure is Kuba's illustrative estimate. Claude Code may have optimized since Boris's tweet or found grep acceptable for their use case (smaller codebases, different accuracy requirements). | Implication: Ken should recognize the cached-compute argument: upfront indexing (Cursor) amortizes cost over many queries, while per-session discovery (Claude Code) scales cost linearly with queries. For multi-agent, multi-user, or repeated-task scenarios, cached embeddings win on speed and cost. This is an architectural principle for agent systems, not just code tools. | Timestamp: timestamp unavailable
- Claim: Jeff Dean's principle: 'You don't need a trillion tokens at once, you need the right million' defines the staged-retrieval paradigm. | Evidence: Jeff Dean quote from podcast/show discussing Gemini's large context windows, emphasizing lightweight mechanisms to progressively narrow trillions to millions. | Caveat: No specific show/podcast cited; quote is paraphrased. Trillion-token context windows are hypothetical—current Gemini models cap at ~2M tokens. The principle is aspirational/directional, not a technical spec. | Implication: Ken should internalize this as a design heuristic: for massive corpora (TurboPuffer customers have trillions of tokens), success isn't about fitting everything in context but about multi-stage filtering to surface the 'right' subset. This justifies investing in hybrid retrieval (vector + full-text + filters) and agent-orchestration layers, not just bigger context windows. | Timestamp: timestamp unavailable
- Claim: Agentic search means giving agents a set of tools (vector, full-text, grep, regex, filters) to iteratively fetch and reason over context, not just file-system grep. | Evidence: Definition contrast: popular view conflates agentic search with Claude Code's grep; TurboPuffer's definition is tool-rich, iterative, progressive retrieval. Examples: Cursor uses vector search + file filters; agents perform multiple searches per task, fetching only what each step requires. | Caveat: Term 'agentic search' lacks industry-standard definition; TurboPuffer's framing is self-serving (they sell vector + full-text search). However, Cursor's success and Jeff Dean's quote corroborate the multi-tool, staged approach. | Implication: Ken should architect agent retrieval as orchestrated tool use, not monolithic RAG or grep-only. This means: (1) multiple retrieval primitives (vector, BM25, filters, SQL), (2) agent decides which tool + when, (3) iterative loops until sufficient context gathered. This is the new default for 'serious' agentic systems per the talk. | Timestamp: timestamp unavailable
- Claim: Embeddings/semantic search are 'cached compute'—upfront indexing cost amortized over many runtime queries, delivering speed and cost advantages. | Evidence: Cursor trace: upfront parse/embed/index cost is one-time; runtime query is lightweight. Contrast with Claude Code: per-session discovery repeats compute each time. Kuba notes TurboPuffer team members switching from Claude Code to Cursor for speed. | Caveat: Cached-compute framing assumes query volume justifies upfront cost. For one-off tasks or rapidly changing corpora, per-session discovery might be acceptable. No quantitative ROI threshold provided. | Implication: Ken should evaluate retrieval architectures by amortization: if agents/users query the same corpus repeatedly, invest in indexing infrastructure. If corpus changes constantly or queries are one-off, lighter-weight, just-in-time approaches may suffice. This is a cost-benefit decision, not a universal rule. | Timestamp: timestamp unavailable
Detailed Brief
The 'RAG is dead' discourse vs. actual adoption trends
- Claims: Twitter/X filled with 'RAG is dead, agentic file search is all we need' posts (late 2025, early 2026).; Google search volume for RAG shows 2023 spike, 2024 plateau, mid-2025 inflection to new highs.
- Evidence: Screenshot of tweets declaring RAG dead.; Google Trends-style graph showing RAG search volume doubling from mid-2024 to mid-2025.
- Caveats: Social media amplifies contrarian takes; search volume is a proxy, not direct adoption metric.; Graph doesn't distinguish between developers learning RAG basics vs. enterprises deploying advanced agentic retrieval.
- Implications: The narrative 'RAG is dead' is actually 'naive one-shot RAG is dead'—interest and sophistication are growing, not declining.; Ken should interpret this as a maturation signal: early hype cycle (2023) → plateau of productivity (2024) → next wave of advanced patterns (2025+).; For business/investment: RAG infrastructure (vector DBs, embedding APIs, orchestration layers) is growing TAM, not shrinking.
Cursor's Merkle-tree embedding strategy and performance gains
- Claims: Cursor was one of TurboPuffer's first customers.; When users open a codebase, Cursor parses, chunks, embeds, and indexes it for semantic search.; Teams of ~100 engineers typically work on 1-2 shared codebases; re-embedding identical code is wasteful.; Cursor uses Merkle trees (crypto hash trees) to detect similarities; if codebases match, copy embeddings and re-embed only changed files.; Semantic search yields 12.5-13.5% average accuracy boost; 24% for Composer model on internal benchmark.; Online A/B test: 2.6% code retention increase, 2.2% dissatisfied-request decrease.
- Evidence: Two Cursor blog posts cited (indexing strategy, semantic search impact).; Merkle tree technique named specifically.; Numerical results: 13.5% average, 24% Composer, 2.6% retention, 2.2% dissatisfaction drop.; TurboPuffer used for secure embedding storage/retrieval.
- Caveats: Internal benchmark, not public—results may not generalize.; A/B test gains look small (2.6%, 2.2%) because semantic search isn't invoked for every query, only applicable ones.; No data on upfront indexing cost or latency impact of Merkle-tree diffing.
- Implications: For Ken's agent systems: upfront indexing infrastructure (Merkle diffing, shared embeddings) is a competitive moat—faster onboarding, lower per-user cost.; For product: even small percentage gains (2.6% retention) at scale translate to major revenue/churn impact.; For AI ops: the Merkle-tree pattern is generalizable to any multi-tenant, shared-corpus scenario (documentation, legal docs, customer data).; Cursor's move to Composer 2 + semantic search is part of their speed advantage over Claude Code, per TurboPuffer team experience.
Claude Code's grep-based approach and token costs
- Claims: Claude Code doesn't use vector search (Boris Cherny tweet).; Early iterations used RAG + local vector DB, but it didn't work for them.; Current approach: per-session file-system grep—agent reads, assesses, repeats until it finds needed context.; Example task ('understand how metadata filtering works'): ~6,000 tokens per sub-step discovery.; 10 agents across 10 days asking the same question incur 10x redundant compute.
- Evidence: Boris Cherny (founding engineer) tweet stating early Claude Code tried RAG, abandoned it.; Trace diagram showing Claude Code: grep → read → assess → repeat, costing 6,000 tokens.; Contrast trace: Cursor's upfront indexing + lightweight runtime query.
- Caveats: 6,000 token figure is Kuba's illustrative estimate, not official Claude/Anthropic data.; Boris's tweet doesn't explain why RAG didn't work—could be accuracy, latency, cost, or corpus characteristics.; Claude Code may have optimized since; grep approach might be acceptable for their user base (smaller codebases, different accuracy tolerance).; No direct comparison of Claude Code vs. Cursor accuracy/speed on identical tasks.
- Implications: Ken should recognize the cached-compute vs. just-in-time tradeoff: upfront indexing (Cursor) amortizes cost over queries; per-session discovery (Claude Code) scales linearly with queries.; For multi-agent or repeated-query scenarios, cached embeddings win on speed and cost. For one-off or rapidly changing corpora, grep may suffice.; The fact that TurboPuffer team members switched from Claude Code to Cursor for speed suggests cached-compute is winning in practice for code-heavy workflows.; This is an architectural principle for Ken's agent systems: if agents repeatedly query the same knowledge base, invest in indexing; if queries are unique, optimize for just-in-time retrieval.
Jeff Dean's 'right million tokens' principle and staged retrieval
- Claims: Jeff Dean (Google) discussed Gemini's large context windows on a podcast/show.; Even with trillion-token context windows, you don't need all trillion at once—you need lightweight staged retrieval to narrow down to the right million.; Quote: 'You don't need a trillion at once, you need the right million.'; TurboPuffer customers have trillions of tokens embedded; the key challenge is narrowing to the right 100K-1M for context windows.
- Evidence: Jeff Dean quote (paraphrased from unnamed podcast/show).; Kuba's assertion that TurboPuffer customers embed trillions of tokens.
- Caveats: No specific show/podcast cited; quote is paraphrased, not verbatim.; Trillion-token context windows are hypothetical—current Gemini models cap at ~2M tokens.; No technical detail on what 'staged retrieval' means in Google's implementation.; TurboPuffer's 'trillions of tokens' claim is aggregate across customers, not per-customer corpus size.
- Implications: Ken should internalize this as a design heuristic: success isn't about fitting everything in context, but about multi-stage filtering to surface the relevant subset.; For massive corpora (enterprise knowledge bases, legal docs, customer data), hybrid retrieval (vector + full-text + filters) + agent orchestration are essential, not nice-to-haves.; This justifies investing in retrieval infrastructure (indexes, caching, routing logic) rather than waiting for 100M-token context windows—even infinite context requires smart narrowing.; For Ken's agent systems: architect retrieval as a funnel (coarse filters → semantic search → fine-grained extraction), not a single query.; For business/investing: companies solving staged retrieval (vector DBs, orchestration layers, query routers) are solving the bottleneck, not the context window providers.
From one-shot RAG to iterative, tool-rich agentic retrieval
- Claims: Traditional 'Twitter RAG': embed corpus, query via vector search, pass to LLM—one-shot.; TurboPuffer's definition of retrieval: vector search, BM25 full-text, grep, glob, regex, filters—multi-tool.; Agentic search (popular definition): file-system grep like Claude Code.; TurboPuffer's definition of agentic search: giving agents a set of tools to progressively/iteratively find and reason over context.; Sophisticated customers no longer do one-shot RAG; they orchestrate multiple tool calls, reasoning through steps, fetching only what each step requires.; Retrieval is now iterative—agents search to understand, then search again based on findings.
- Evidence: Kuba's definitions contrasting popular vs. TurboPuffer views.; Cursor example: semantic search + file filters + iterative queries.; Claude Code example: grep → read → assess loop.; Jeff Dean quote supporting staged/progressive retrieval.
- Caveats: 'Agentic search' lacks industry-standard definition; TurboPuffer's framing is self-serving (they sell vector + full-text).; No quantitative data on how many queries typical agents make per task, or how retrieval accuracy improves with iteration.; Examples are code-specific (Cursor, Claude Code); unclear how this applies to other domains (e.g., customer support, legal research).
- Implications: Ken should architect agent retrieval as orchestrated tool use: (1) multiple primitives (vector, BM25, filters, SQL), (2) agent decides which tool + when, (3) iterative loops until sufficient context.; This is the new default for 'serious' agentic systems per the talk—not a future vision, but current best practice at Cursor and TurboPuffer's customer base.; For AI ops: instrumentation/observability must track multi-step retrieval (which tools used, in what order, token costs per step) to optimize agent performance.; For product: expose iterative retrieval to users as 'thinking' or 'searching' indicators to set latency expectations.; For investing: companies building agent orchestration layers (LangChain, LlamaIndex, custom frameworks) are enabling this shift; pure vector DB plays may be commoditized unless they support multi-tool workflows.
Notable Concepts & Terms
- Cached compute (embeddings as cached compute): Upfront cost of parsing, embedding, and indexing content is amortized over many runtime queries, delivering speed and cost advantages vs. per-query discovery. Cursor's Merkle-tree strategy exemplifies this.
- Merkle trees (in code embedding context): Crypto hash trees used by Cursor to detect identical/similar codebases across team members, enabling embedding reuse and re-embedding only changed files. Reduces redundant indexing costs.
- Staged retrieval: Jeff Dean's principle: progressively narrowing trillions of tokens to millions/thousands via lightweight mechanisms (vector search, filters, etc.) rather than fitting all tokens in context at once.
- Agentic search (TurboPuffer definition): Giving agents a set of tools (vector, BM25, grep, regex, filters) to iteratively fetch and reason over context until task requirements are met, not just file-system grep.
- BM25: Full-text search ranking algorithm (part of TurboPuffer's retrieval toolkit) used alongside vector search in hybrid retrieval systems.
- Hybrid tool-rich retrieval: The talk's proposed paradigm: combining vector search, full-text (BM25), grep, regex, filters in agent workflows, orchestrated iteratively to fetch only what's needed per step.
Operator Notes / Why Ken Should Care
- For Ken's agent systems: architect retrieval as multi-tool, iterative orchestration—not single vector queries. Cursor's 24% accuracy boost and TurboPuffer team's switch from Claude Code to Cursor for speed validate this approach.
- Cached-compute principle (upfront indexing, runtime lightweight queries) is critical for multi-agent, multi-user, or repeated-query scenarios. For one-off/rapidly changing corpora, per-session discovery may suffice—this is a cost-benefit decision.
- Jeff Dean's 'right million tokens' heuristic should guide Ken's architecture: build multi-stage funnels (coarse filters → semantic search → fine extraction), not single-query retrieval. Even infinite context windows require smart narrowing.
- Instrumentation/observability must track multi-step retrieval (tools used, order, token costs per step) to optimize agent performance and cost. This is now table stakes for 'serious' agentic systems.
- For investing: companies solving staged/hybrid retrieval (orchestration layers, query routers, multi-modal indexes) are solving the bottleneck. Pure vector DB plays may commoditize unless they support multi-tool workflows and agent integration.
- The 'RAG is dead' narrative is marketing noise—actual trend is toward more sophisticated retrieval (agentic, multi-tool, iterative), not less. TAM for retrieval infrastructure is growing, not shrinking.
- Cursor's Merkle-tree embedding strategy is generalizable to any multi-tenant, shared-corpus scenario (docs, legal, customer data)—a tactical pattern for Ken's systems if relevant.
Watch Map
- timestamp unavailable: Transcript lacks timestamps; video is a single 11-minute talk without clear chapter markers.
- ~0:00-2:00 (estimated): Intro: 'RAG is dead' discourse vs. Google search volume spike mid-2025.
- ~2:00-4:00 (estimated): Definitions: RAG (multi-tool retrieval, not just vector) vs. agentic search (iterative tool use, not just grep).
- ~4:00-7:00 (estimated): Cursor case study: Merkle-tree indexing, 12.5-24% accuracy gains, 2.6% retention boost.
- ~7:00-9:00 (estimated): Claude Code grep approach, cached-compute argument, Cursor vs. Claude Code token costs.
- ~9:00-11:00 (estimated): Jeff Dean's 'right million tokens' principle, staged retrieval, conclusion.
Source/Metadata
- Title: RAG is dead, right?? — Kuba Rogut, Turbopuffer
- Transcript words: 2697
- Duration seconds: 673
- Timestamp note: Transcript contains no timestamps; estimated watch_map based on 673-second duration and logical flow.
Transcript
KUBBA D. Hi. Welcome, everyone. Thanks for coming out. I see it's a full room. So I appreciate everyone coming out. So welcome to the talk about Rag is Dead. So my name is Kuba. I'm a deployed engineer at Turbo Puffer. So for those that don't know what Turbo Puffer is, we are a full text search and vector search database built from first principles on top of object storage. If you would love to learn more, just come find me after the talk if you have any questions. So let's get started. So this talk is about how RAG is dead, how hybrid tool, tool-rich retrieval is becoming a default for serious agentic search. So if you guys have been on Twitter or other social media platforms, or I guess X as they call it now, you might have seen a lot of tweets about how RAG is dead. You can see there's lots of tweets, especially in the last, end of 2025 and in the early of this year about how RAG is dead, agentic file search is all we need, and there's a lot of tweet and a lot of content about this now. But interestingly, if we were to look at something like the Google search volume over the last two years, or the last couple of years, you can see that in 2023, as AI starts, we have this increase, it caps out a little bit in 2024, settles down for about a year, and about midway through 2025, we hit this new inflection point, where search volume just goes through the roof. So take that, Twitter. So let's clarify first. What is RAG and what is agentic search? These are two terms a lot of people are throwing out these days. So RAG, what a lot of people think RAG is, is just simple vector search. They just think that this is simply embedding a bunch of a corpus of contents, passing an embedding vector, and getting it back, passing it through your LLM. And at TurboPover, what we think this actually means, if we break down RAG into retrieval, augmented generation, retrieval is not just vector search. It's a lot of different things. It could be vector search, full-text search, using stuff like BM25, grepping, globbing, using RegEx, using other just basic filters. And then you augment the generation is obviously just passing it into your LLM of choice. And then agentic search. This is the terms people are throwing out a lot these days. And generally when people start talking about agentic search, what they usually talk about is essentially just file system grep. So if you guys are familiar with something like Cloud Code, that's what, or Cloud Code Codex, a lot of people call this agentic search. And this essentially is grepping through your file system, and this is why these terms are so correlated. And what we actually believe it is, and the definition we want to give it is, it's really giving the agents a set of tools to progressively and iteratively find and reason over context. So with Cloud Code, you can, if you guys are familiar with it, it can read your file, you start grepping through your file system, read a file, decide that it hasn't found what it needed, what it needed to actually complete the task, and it will find something again. And then keep doing this until it's happy, it's reached a happy state where it can continue on with the task. So we're going to take a step back and talk about one of the companies that use Turbo Puffer that we believe is doing an excellent job with agentic search. This is a company called Cursor, you might have heard of them. Fun fact, they're actually one of Turbo Puffer's very first customers. And they have this excellent blog post that came out in the beginning of 2026 about how they index code bases. So for those unaware, when you open up a new code base or new branch in Cursor, what happens is that Cursor will start embedding your code base. So what they'll do is, chunk out your parse, chunk and embed your code base, and make it available for semantic search. And this blog post goes into excellent technical detail of how they do this. Just to give you the gist, essentially, the cool thing they do is that they found that most people working on a team, let's say there's a hundred engineers, when they open up code bases, they're normally the same code base, 99% of the time, because you can have a team of 100 people most of the time working on one, two, maybe a few code bases, right? And it's really expensive to have to re-chunk, re-embed, and re-upload these code bases every single time. So they essentially use Merkle trees, which is this crypto hash tree, to calculate similarities between code bases people open on the same team. And if they're similar enough, they will essentially copy over the data and then only update, re-chunk and re-embed the files that have changed, and use TurboPuffer in order to make sure this is done securely. And yes, just excellent blog posts. They do some really cool stuff. And you may think this is a lot of work. Why do they do this? Well, the reason they do this is also covered in a different blog post about how they use semantic search. Again, they use TurboPuffer for this. And what they find is, on average, across models, I think it's a 12.5% or 13.5% increase in answer accuracy. This is across their internal cursor context benchmark. and use TurboPuffer in order to make sure this is done securely. And, yeah, just excellent blog posts. They do some really cool stuff. And you may think this is a lot of work. Why do they do this? Well, the reason they do this is also covered in a different blog post about how they use semantic search. Again, they use TurboPuffer for this. And what they find is, on average, across models, I think it's a 12.5% or 13.5% increase in answer accuracy. This is across their internal cursor context benchmark. So not a public benchmark, but you can trust the numbers they give us. And you can see on the right side, their Composer model, so this is before Composer 2, they had almost a 24% increase in answer accuracy. So giving semantic search to these tools and to these models is really can drive real performance gains. And you can see on the bottom right, this is from an online A-B test they did, which is also covered in their thing, in their blog post, about how it's almost a 2.6% code retention in large code bases. And there's a 2.2% decrease in dissatisfied user requests. And you might be thinking, oh, well, these numbers aren't that big, 2.6%, 2.2%, not that large. But they also cover that semantic search isn't used in every single query. So in their online A-B test, if you give it, if you give these tools to 100 random queries, not every 100 query will actually benefit from the existence of a semantic search tool. So that's why these numbers look small. And now let's talk a little bit about Cloud Code. So Cloud Code doesn't use vector search, as covered by the suite from Boris Cherny. So those unfamiliar with Boris, he's essentially the founding father of Cloud Code. And he says that in early iterations of Cloud Code, they actually did use RAG in a local vector DB. But they found that it just didn't really work out for them. But this is something that is important to understand. It's something we've taken on a lot internally of understanding here at Turbo Puffer is this idea that embeddings and semantic search are cached compute. And you may be thinking, cached compute, throwing out a lot of terms at me right now. I don't know exactly what that means. And I think it's best to walk through an example of essentially almost a Cloud Code looking trace and a cursor looking trace of how some agents will understand your code base. So on the left is a per session discovery of Cloud Code. So for example, if we were to ask the agent to understand how metadata filtering works, what it would have to do is grep, read, assess, and repeat, and try to find the files it needs in order to gain this understanding on a per session basis. So what this means is you can have 10 agents on 10 different days across 10 developers, and they can be asking the same question multiple times, every day. Every time the agent's gonna have to repeat these same exact steps to gain the same understanding of this code base. And this could cost quite a few tokens. You know, 6,000 doesn't seem like a lot here, but just remember this is one sub step of an agent. And then on the right is a more cursor looking trace where there's this upfront cost of indexing, but then we're able to allow for this lightweight tool to help the agent retrieve this information at runtime. So obviously there's this upfront cost of parsing code base, embedding it, and making it available. But this is a one-time cost. And then at runtime, the agent can just query something like how is metadata filtered. It can get some simple results and it would save a lot of tokens, a lot of time, and a lot of money. And this just helps the agent to become a lot faster. You know, a lot of people on the team now that maybe were big cloud code users here at TurboPuffer, they've actually started switching to cursor just because of how fast it's becoming, especially with their Composer 2 models and also the semantic understanding. It just become, we're finding really good. So from RAG to agentic retrieval. So what we're finding now is that a lot of people are no longer doing the simple RAG, the Twitter quote-unquote RAG of just doing a search once and throwing it into the context windows. What we're finding is that this work, back in 2023, early 2024, the beginnings of AI, but a lot of the more sophisticated customers are doing agentic search and is giving real big performance gains and unlocking new products. And what we're finding is they're doing a ton of calls. They're reasoning, these agents are reasoning through several steps. They're searching semantically or through full text, et cetera, as needed. And they're only fetching what's needed for that specific use case. But the important thing to know is that retrieval is no longer just this simple one-time call to VectorDB. It's becoming super iterative and these agents are really understanding what they're searching and searching to understand more in a sense. And it's interesting. You know, Google's Jeff Dean, he went on a, I forget if it was a show or a podcast or whatever, and he had this really good quote that we like to use, that we also thought was super interesting. He was talking a little bit, I believe, about how Gemini's models were having these really big context windows. And I forget the exact question the host asked them. But he was saying, you know, big context windows, it doesn't matter if you get to a trillion context window size, what you really need is staged retrieval, a lightweight mechanism to narrow down these trillion tokens into essentially millions at a time. And the exact quote is, you don't need a trillion at once, you need the right million. This is something we think a lot about here at TurboPuffer. You know, we have customers that embed, have trillions of tokens inside TurboPuffer. And as we see, the really important part is just getting down to this right 100,000, right 10,000, right million in order to pass into these context windows. So that's about it for the talk. A trillion context window size, what you really need is stage retrieval, a lightweight mechanism to narrow down these trillion tokens into essentially millions at a time. And the exact quote is, you don't need a trillion at once, you need the right million. This is something we think a lot about here at TurboPuffer. We have customers that embed trillions of tokens inside TurboPuffer. And as we see, the really important part is just getting down to this right 100,000, right 10,000, right million in order to pass into these context windows. So that's about it for a talk. If you have any questions about any specifics, I'd love to either have them asked now, or you can find me after the talk. But appreciate you guys coming out. We'll be right back. But the important thing to know is that, you know, retrieval is no longer just this like simple one-time call to VectorDB. It's becoming super iterative and these agents are really understanding what they're searching and searching to understand more in a sense. And it's kind of like interesting loop. You know, Google's Jeff Dean, he went on a, I forget if it was a show or a podcast or whatever, and he had this really good quote that we like to use, that we also thought was super interesting. He was talking a little bit, I believe, about how Gemini's models were kind of having these really big context windows. And I forget the exact question the host asked them. But he was saying, you know, big context windows, it doesn't matter if you get to a trillion context window size, what you really need is stage retrieval, like a lightweight mechanism to narrow down these trillion tokens into essentially millions at a time. And the exact quote is, you don't need a trillion at once, you need the right million. This is something we think a lot about here at TurboPuffer. You know, we have customers that embed, you know, have trillions of tokens inside TurboPuffer. And as we see, like the really important part is just getting down to this right 100,000, right 10,000, right million in order to pass into these context windows. So that's about it for a talk. If you have any questions about any specifics, I'd love to, you know, either have them asked now, or you can find me after the talk. But appreciate you guys coming out. So, just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just to just We'll be right back.