If we want them to do Knowledge Work, design them as Knowledge Agents — Benjamin Clavié, Mixedbread
Description
An agent searching a contract for "30 days" has no way to know whether it has found a deadline, a grace period, or a retention rule, or wandered into a medication schedule entirely. Benjamin Clavié uses that to explain why coding agents turned out to be the easy case. Code carries durable cues: identifiers, file paths, method definitions, things you can literally grep for. It also comes with narrow instructions, because nobody asks a coding agent to invent a paradigm, they ask it to close a ticket. Knowledge work outside code is the opposite on every axis. The meaning is implicit, the same string means different things in different domains, and the search has to start from an intent the agent must reconstruct on its own. His argument is that the industry generalized from the one domain that least resembles the rest. The good news is that humans solved this already, repeatedly. Clavié traces a tool loop running from the first catalogue of the Library of Alexandria through bibliographies to search engines, and an organization loop running from the single gifted polymath through monasteries and universities to the modern specialist firm. They are really one loop: better tools create new roles, new roles generate knowledge, new knowledge demands better tools. His benchmark numbers make the practical version concrete. A badly tuned retrieval tool lands near 60 percent accuracy, which nobody would trust enough to keep using. A well tuned one reaches the same ceiling as its rivals while making 20 percent fewer tool calls, and adding a layer of searcher agents that return short memos closes roughly 40 percent of the remaining gap to a human. Speaker info: - https://x.com/bclavie - https://mixedbread.com - https://ben.clavie.eu Timestamps: 0:00 - Coding agents are a special case of knowledge work 3:48 - Why code is unusually easy to search 5:05 - Thirty days, and when one clue means four things 5:55 - The tool loop and the organization loop 8:06 - Tooling is never a neutr
Summary
Generated by claude-sonnet-4-5-20250929At-a-Glance
- Verdict: Watch fully
- Core thesis: Agent systems should be designed like knowledge-work organizations (lawyer + paralegals, doctor + assistants) rather than monolithic coding agents, co-designing hierarchical orchestration with search primitives to overcome ceiling effects.
- Why it matters: Ken's agent systems face the same orchestration and retrieval challenges as legal/medical knowledge work: ambiguous queries, contextual meaning, and finite context windows require task decomposition and specialized sub-agents, not just better models.
- Best use: Extract the multi-agent orchestration pattern (manager breaks down query → specialized searchers retrieve → manager synthesizes) and the tool co-design principle (train agents on search primitives beyond grep/BM25) for OpenClaw's control plane and agentic workflows.
Executive Summary
Benjamin Clavié argues that the agent community has over-indexed on coding agents as the template for all agentic work, when in fact code is a special case of knowledge work with unusually durable cues (identifiers, paths, keywords) and narrowly scoped tasks (features, tickets). Most knowledge work—legal research, medical diagnosis, financial analysis—requires handling ambiguous, context-dependent information (e.g., "30 days" could be a deadline, grace period, or retention rule) and open-ended intent-driven queries that agents must decompose themselves. Human organizations solved this centuries ago: senior experts (partners, doctors) break down problems and delegate research to trained specialists (paralegals, nurses) who use appropriate tools and return memos, enabling the senior agent to synthesize and act.
Clavié presents two co-evolving loops: a tool loop (oral tradition → writing → library catalogs → search engines) and an organization loop (polymaths → monasteries → universities → specialized firms). These are actually one loop: new knowledge demands better tools, better tools enable new specialized roles, new roles create more knowledge. Tools are not neutral upgrades; they determine whether tasks are economically scalable. He demonstrates this with BrowsGym Plus (200K text documents): unoptimized BM25 gets 60% accuracy and 25 tool calls, optimized hybrid search gets 90.2% accuracy with 8 calls (5% of the cost). On MatQA (PDF-based enterprise benchmark), even a human with unlimited BM25 searches plateaus at the same accuracy as Gemini 1.5, proving tool choice sets the ceiling.
The breakthrough comes from organizational design: a single Gemini agent with multimodal search reaches 88.9% on MatQA, but a manager-searcher architecture (manager decomposes the question into sub-queries, searcher agents retrieve and return memos, manager synthesizes) reaches 92.4%, closing 40% of the gap to human performance. Clavié's takeaways: (1) copy human knowledge-work patterns for agent design, (2) use tools to overcome ceilings not as default paths, (3) co-design tools and agents so models know when to use semantic search vs. BM25 vs. grep as primitives, and (4) orchestrate hierarchically because context will always be finite and expensive, even at 100M tokens.
Key Takeaways
- Claim: Code is a special case of knowledge work with durable cues and narrow tasks; non-code knowledge work is contextual, meaning-driven, and intent-based, requiring agents to decompose open-ended queries themselves. | Evidence: In coding, tasks are pre-scoped ("implement this feature") and cues are stable (function names, file paths). In legal/medical knowledge work, the same phrase ("30 days") can mean deadline, grace period, or retention rule depending on domain context, and agents must infer conditional chains ("which international norms apply, then do they apply here?") from client intent. | Implication: Ken's agents for OpenClaw or GTM workflows cannot assume coding-agent patterns will transfer; they need hierarchical orchestration and contextual reasoning layers to handle ambiguous, multi-domain queries.
- Claim: Human knowledge work evolved two co-dependent loops—tool development and organizational specialization—that are actually one self-optimizing loop: new knowledge → better tools → new specialized roles → more knowledge. | Evidence: Tool loop: oral tradition → library catalogs (Pinakes in Alexandria) → Dewey system → search engines. Organization loop: polymaths → monasteries → universities → specialized firms (doctor/nurse hierarchy). Each tool change forced retraining and role creation (e.g., teaching librarians to use Google vs. physical catalogs). | Implication: When Ken introduces new retrieval or reasoning tools in agent systems, he must co-design the orchestration roles (manager, researcher, synthesizer) and retrain agents on the new primitives, or the tools won't scale.
- Claim: Tools are not neutral upgrades; they determine whether tasks are economically scalable, not just whether they are possible. | Evidence: Finding a manuscript in pre-catalog Library of Alexandria was possible but took weeks; with a catalog it took 10 minutes. On BrowsGym Plus, unoptimized BM25 hits 60% accuracy with 25 tool calls; optimized hybrid search hits 90.2% with 8 calls, costing 5% of the baseline. On MatQA, even humans with unlimited BM25 searches plateau at Gemini 1.5's accuracy, proving the tool sets the ceiling. | Implication: Ken should benchmark retrieval tools (BM25, semantic, hybrid, multimodal) on his actual tasks before defaulting to the training-data-common grep/BM25 path, and optimize baselines before claiming a tool is bad.
- Claim: A manager-searcher multi-agent architecture (manager decomposes query → searchers retrieve → manager synthesizes memos) closes 40% of the gap to human performance vs. a single-agent system. | Evidence: On MatQA (PDF enterprise benchmark), a single Gemini agent with Mixedbread multimodal search reached 88.9% (humans: 99.4%, 10.5-point gap). The Mixedbread search agent (manager writes sub-queries, searchers return memos) reached 92.4%, reducing the oracle gap from 10 points to 6 points (40% mistake reduction). | Implication: Ken should architect OpenClaw's retrieval and reasoning layers as hierarchical teams: a control-plane agent that decomposes tasks and delegates to specialized retrieval/analysis agents, then synthesizes their outputs, rather than a monolithic agent with tool access.
- Claim: Agents must be co-designed with tools and trained on search primitives (grep, BM25, semantic search) as distinct capabilities, not just grep/BM25 from training data. | Evidence: Agents default to writing grep queries because grep and BM25 dominate training data, but PDFs require semantic or multimodal search. Models need explicit training on when to use each primitive (grep for exact matches, BM25 for lexical, semantic for meaning-driven). | Implication: Ken should build or fine-tune models/harnesses that know when to route to different retrieval primitives and expose those as first-class tool options, not assume GPT-4/Claude will infer the right search strategy.
- Claim: Context is a finite resource even at 100M tokens, so hierarchical orchestration and task decomposition are always necessary, not just scaling context windows. | Evidence: 100M token context still cannot hold even half of one U.S. state's legal code, let alone federal, international, and specialist case law. Cost and coverage both demand breaking down tasks rather than stuffing everything into context. | Implication: Ken's agent systems should prioritize orchestration and retrieval architecture over waiting for larger context windows; even with infinite context, decomposition and specialized roles remain more efficient and cost-effective.
Detailed Brief
Benchmark-Driven Tool Optimization
- Claims: BrowsGym Plus (200K documents, text-only, deep research queries) shows optimized hybrid retrieval reaches 90.2% accuracy with 8 tool calls vs. 60% accuracy with 25 calls for unoptimized BM25.; MatQA (PDF-based enterprise benchmark with OCR fallback) demonstrates that tool choice sets the ceiling: humans with BM25 plateau at Gemini 1.5's accuracy, but multimodal search (Mixedbread) enables another jump to 88.9% for single agents and 92.4% for multi-agent systems.
- Evidence: 20% fewer tool calls at same accuracy = 20% cost reduction; optimized hybrid vs. baseline = 5% of original cost.; Oracle gap (perfect docs vs. system) reduced from 10 points to 6 points (40% mistake reduction) by adding manager-searcher orchestration.
- Caveats: BrowsGym Plus is text-only and not open-ended; MatQA uses PDFs but still has ground-truth queries, so real-world ambiguity may require even more decomposition.
- Implications: Ken should use task-specific benchmarks (not just generic evals) to justify tool choices and measure orchestration gains before deploying agents to production.
Real-World Knowledge Work Patterns
- Claims: In legal/medical/actuarial work, senior experts (partners, doctors) meet clients to understand open-ended problems, decompose them into research tasks, delegate to trained specialists (paralegals, nurses), review memos, and synthesize responses.; Coding agents skip this because the user pre-decomposes the task into a ticket or feature spec; knowledge agents must do the decomposition themselves.
- Evidence: Legal partner workflow: client describes problem → partner identifies relevant laws and conditional chains → assistants research using specialized tools → partner reviews memos and responds.; Contrast: coding agent workflow: user writes ticket → agent implements ticket directly.
- Implications: Ken's agent UX for OpenClaw or GTM workflows should accept high-level intent ("find all relevant contracts for this client scenario") and expose the decomposition step, not assume the user will write the sub-queries.
Notable Concepts & Terms
- Knowledge work: Work where the main input is information (not physical materials) and the output is actionable judgment or decision; includes legal, medical, financial, actuarial, and coding domains.
- Tautological definition of knowledge work: If you need search, it's a knowledge problem; if it's a knowledge problem, you need search.
- Tool loop vs. organization loop: Human history shows parallel evolution: better tools (writing, catalogs, search engines) and specialized roles (polymaths, universities, firms) co-evolve in one self-optimizing feedback loop.
- Durable cues: In code, identifiers/paths/keywords are stable and greppable; in knowledge work, the same phrase ("30 days") has multiple context-dependent meanings.
- Oracle gap: The performance difference between a system with perfect ground-truth documents and the actual retrieval system; on MatQA, hierarchical agents reduced oracle gap from 10 points to 6 (40% mistake reduction).
- Search primitives: Grep (exact match), BM25 (lexical search), semantic search (embedding-based), and multimodal search (vision + text) are distinct tool capabilities that agents must be trained to use appropriately.
- Manager-searcher architecture: Hierarchical agent pattern: manager decomposes query into sub-queries, searcher agents retrieve relevant docs/memos, manager synthesizes final answer; mirrors legal partner/paralegal structure.
- BrowsGym Plus: Leaderboard/benchmark for evaluating search tool quality on deep research tasks with 200K text documents and convoluted queries.
- MatQA: PDF-based enterprise knowledge benchmark (works with Snowflake) that uses OCR and multimodal retrieval to test real-world document QA; shows humans plateau at ~99.4% accuracy.
Operator Notes / Why Ken Should Care
- Copy the manager-searcher pattern for OpenClaw's control plane: have a decomposition agent that writes sub-queries and a set of retrieval agents that return structured memos, then synthesize in the manager.
- Benchmark retrieval tools (BM25, semantic, hybrid, multimodal) on Ken's actual document corpus before defaulting to grep/BM25 from training data; always optimize baselines before claiming a tool is insufficient.
- Build or fine-tune models/harnesses to expose search primitives (grep, BM25, semantic, multimodal) as first-class tool options and train them on when to use each primitive for different query types.
- Design agent UX to accept high-level intent and expose the decomposition step rather than assuming users will write sub-queries; make decomposition transparent so Ken can audit and refine the breakdown logic.
- Monitor oracle gap (perfect docs vs. system performance) as the key metric for retrieval quality, not just accuracy; a large oracle gap means the bottleneck is retrieval, not reasoning.
- Do not wait for 100M-token context windows to solve knowledge work; even at infinite context, hierarchical orchestration is more cost-effective and scalable than stuffing everything into one prompt.
- When introducing a new retrieval or reasoning tool, co-design the organizational roles (who uses it, who reviews it, who synthesizes) and retrain agents on the new primitives; tools alone do not scale without workflow changes.
Source/Metadata
- Title: If we want them to do Knowledge Work, design them as Knowledge Agents — Benjamin Clavié, Mixedbread
- Transcript words: 6348
- Duration seconds: 1074
- Timestamp note: Timestamps present and usable
Transcript
OK, so hi, everyone. I'm going to give the quickest introduction to myself. I'm Ben Clavier. I work at Mixed Bread where we do retrieval. I'm French and I live in Tokyo. And today I'm going to talk to you about the fact that agents should do knowledge work. And so we should design them like knowledge workers. We should design them like knowledge agents and not coding agents. And I'm going to explain the difference and why I think that's important. It's a bit of a hot tech talk, but let's start now. So the first thing is early agents give us fun trivia. And I'm talking early agents, 2022 agents, back when all you had was RAG, but agents couldn't even do tool calls back then. So all you had is you had the IF statement, you did search, you got results, you could talk to your PDF. That was the very first form of agentic work. It was not very useful when I talked about that for long. What came next was programming agents. And that's been all the rage. Once agents started being able to search properly, actually properly understand things, carry tasks out, call tools, we started designing coding agents. And coding agents are a big thing. I don't think there's anyone in this room that does not use coding agents. We all use Code, Codex, et cetera. And that's a form of knowledge work. But agents were not knowledge workers at the time. Agents were coding agents. And now they're becoming knowledge workers. And by knowledge worker, I mean that knowledge work is a big superset. Every workflow you've thought of before is a form of knowledge agents, just because of the nature of knowledge. So if you have an agent that's a lawyer that's looking for legal documents, a financial agent, if you're looking for medication information, all of that is knowledge agents. It's trying to find knowledge, it's trying to make use of knowledge. And coding's part of that, of course, and even the small RAG bit that we talked about in the first slide is part of that. But that's a very, very small proportion of the actual full thing. There's so much more to knowledge than any one domain. And why is knowledge work important? Because I'm saying that it's important. They do knowledge work. Coding is knowledge work. And I think there's two ways to define it in my opinion. One of them is knowledge work is work where your main input is information. Your main input is not an actual physical material. It's not something that you can touch. It's knowledge, it's information. And the nature of knowledge work is that you process this information, which is by nature very ambiguous, very diffuse. And the main output you get from that is something actionable. It's a judgment, it's a decision. If it's a lawyer, you're going to get the findings on your case and they might plead for you. You're going to get an actual actionable thinking item, something still not tangible, but that exists as knowledge. There's also a tautological definition, which makes sense here: if you need search, it's a knowledge problem. And if it's a knowledge problem, you need search. It's very easy, self-defined. And in the real world, that's basically most of the work that we see in the service economy is a form of knowledge work: lawyers, knowledge workers, academics, knowledge workers, actuaries, knowledge workers, software engineers, researchers, also knowledge workers. And the fact that there's so much knowledge work in society has contributed to a never-improving structuring of knowledge work. There's actually very, very well-defined workflows for how we should do knowledge work, for how knowledge works in itself and how we evolve that. But agentics so far has focused on the special case and has tried to generalize from it. And that special case is coding and software engineering. And the thing is, code is knowledge, but not all knowledge is code. And code is a very, very unique form of knowledge because it has very durable cues. In a code base, there's going to be a lot of references to an identifier or a file or a path. When we update code, that can change, but most of the time it's not going to change all that much. All the things are very, very durable. Then you've got the surface that's obvious and grippable. There's keywords, there's method definitions. There's a lot of things that by definition you can grip in code. And the task—and this one's actually very important and we don't talk about it a lot—but people ask, why is grip good enough for programming? Or why can an agent do programming? And then you're telling me it can do deep research for a legal question. One of the reasons is because we don't realize it, but when we interact with coding agents, we are giving them extremely narrow tasks. We're not actually expecting that much from them. Everything's always about a feature, about a given ticket. There's a task at hand. You're not going to tell the agent to discover a new programming paradigm and then implement it in this new app. But in knowledge work, that's often the case. First of all, you don't have those durable cues. The meaning is always implicit. And more importantly, the same cue can mean a lot of different things. We don't have function definitions in knowledge work. If you see 30 days, if your agent's looking for 30 days, is it a deadline? Is it a grace period? Is it a retention rule? Is it even in the same domain? Are you searching for 30 days on the contract and you're getting medication? There's a lot of contextual information here. But more importantly, the search starts from an intent. Even if you're doing a legal example, if you're asking about a specific rule that you want to apply to a specific domain, you're going to need to look at the international norms that apply and then do they apply in this case. There's a lot of conditional information that is not predefined in the task. That's all up for the agent to find. And so non-code knowledge is very contextual and meaning-driven, which is much harder than code. And that's led to the fact that none of what I'm saying is new. People have been doing knowledge work for a very, very long time. And that's resulted in two endless loops. So you have a tool loop, which is at the start we were talking. Then at some point someone was like, we should write stuff down. Then in Alexandria we had the Pinakes, which was the curator of the library of Alexandria, came up with an idea that maybe we should have a way to catalog all of the books we have. Then we developed writing. Then we developed bibliographies. Then we ended up with the current version of the Dewey system for libraries. And nowadays we have search engines. But we also had an organization loop, which is giant births disjoint from the tool one. It used to be the one gifted expert. We've all heard of the polymaths of the past, the person who just knew everything about one domain or all domains. And you just went to them if you had information. But that doesn't scale. So we ended up with monasteries, which were guardians of knowledge. And then we had universities and then we ended up creating bureaucracies. And now we ended up creating the modern organization of work where we have very specialized firms. Like at hospitals, you've got the doctor, you've got the senior doctor, you've got the nurse practitioner, the nurses, the healthcare assistants, and all of them specialize on different levels of tasks. And that's a really good form of optimization. But the thing is that it's actually just one loop. I'm showing two loops here, but they're actually just one loop, which is we have new knowledge and new knowledge means that we need better tools. And better tools mean that we end up creating new workflows, new roles. We need people that are trained to use those tools, people that understand what the new tool does. If you have a person that knows how to go to the library and you're like, okay, use Google, you need the knowledge of what Google is. That person needs to be taught that it's a search engine. You can just type stuff in it. There's no need to physically go there. And that means you retrain, you get new knowledge workers who are more efficient, so they create more knowledge. So we need new tools and so on and so on. So both the tool loop and the organizational loop are actually just one self-optimizing loop that kind of triggers the other endlessly. And the thing about tooling and optimization is that they're not neutral add-ons. which is we have new knowledge and new knowledge means that we need better tools. And better tools mean that we end up creating new workflows, new roles. We need people that are trained to use those tools, people that understand what the new tool does. If you have a guy that knows how to go to the library and you're like, okay, use Google, you need the knowledge of what Google is. That person needs to be taught that's a search engine. You can just type stuff in it. There's no need to physically go there. And that means you retrain, you get new knowledge workers who are more efficient, so they create more knowledge. So we need new tools and so on and so on. So both the tool loop and the organizational loop are actually just the one self-optimizing loop that kind of triggers the other endlessly. And the thing about tooling and optimization is that they're not neutral add-ons. I said that we keep optimizing tools and things come up and we create new things out of those tools. But that's never actually a neutral thing. Tooling is not just, oh, my search is 5% better. The fact that we have a tool or the fact that we don't have a tool is what decides if a task—not if the task is possible because you can do things without the right tool—but if the task is actually scalable and can be carried out cheaply because something being cheap means it can scale. It's yes, of course, if you go to the library of Alexandria before the Pinacotheca, you can find your manuscript somewhere. Whatever you're looking for is there. It's probably going to take two or three weeks. So you're going to really need that knowledge. But if there's a library catalog, it's going to take you 10 minutes, and now it's way easier to just, oh, okay, I need to know something more about this. I'm going to search for it. Likewise, if you have a map directory or if you even have a map in the first place, which in itself is a tool for information, then exploring the world is a much better idea. You're not going to rely on randomly discovering America on your way to the Indies. You know where you're going. And likewise, if you have a multimodal search platform, then you can search millions of PDFs in a way that we couldn't before. So now there's a lot of use cases where you're, oh, it's in the archives. I'm not going to touch that. That becomes actually useful. And in practice, this looks like that, and I'm getting into the more technical stuff here, which is on a simple deep research task. So this is the Browscom Plus leaderboard, which is made to evaluate the quality of search tools on a very abundant deep research task. You have 200 found documents and you have specific queries. We all talked about this this morning, and it's a really useful benchmark to analyze queries. And what we see is that a bad tool—so that's the thing that people often rant about. You will see that there's two BM25 here. There's two like optimized and unoptimized, and that's because quite often people will tell you BM25 is not great. And the reason they'll tell you BM25 is not great is because there's not one BM25, there's hundreds of them. It's a way to do a lexical search. You should always optimize your baselines. You should always optimize what you're beating. And so what you see here is a badly optimized tool is useless, 60% accuracy. You're not going to trust someone that's right 60% of the time. You're just going to do it yourself. When you start optimizing the tools, you can see we go up to 70, 80, and then the actual best is a hybrid harness. It gets to 90. But that's maybe not the most interesting part because we start plateauing at one point. The jump from 89.8 to 90.2 is in-run variance. That doesn't matter. What matters here, however, is that 90.2% accuracy, you reach it with 20% fewer tool calls. And that's huge because in practice that's 20% fewer tokens, 20% fewer resources that you use. That's basically 20% free cash. And if you compare it to the unoptimized baseline, you're spending 5% of what you were spending in the first place. So the tool is actually what makes the task worth doing. Nobody would keep using that tool if it takes 25 calls. But if it takes eight calls, you're like, oh, yeah, that's a workflow I can introduce. And the second part, which goes with tooling, and I think is just as important because Browscom Plus in the previous slide is interesting, but it's easy. It's 100,000 documents. It's just text. It's just the one question. And it's not really that open-ended. It's just a bit convoluted. But when you're actually doing real-life knowledge work, there is that workflow that you see that I try to do that, which is you have a client. They come to the big shot. They come to the lawyer. That's the partner of the agency. And they're like, okay, this is my situation. That's my problem. And they meet together. But then what the partner does is they're not going to be the ones doing all the legal research. They're not going to be the ones doing every single step of the problem. What they'll do is understand that, okay, this person has this problem. That's going to cause them that. Those are the facts. I need the relevant laws to this, that, and other aspects. And then they've got party goals. They've got assistants. And the assistants are going to be doing this research. They're going to be using the size of tools they've been trained to use. And they're going to produce memos and notes. And then they're going to give that back to the big shots. Maybe they'll research one clarification point, but they mostly rely on what their searcher agents, if you want, their assistants have done for them. And that's the response that you're going to get. And that's echoing the point I made before, which is in code, when you're using code code, you're doing that work yourself. You've already broken down the query. You know what you want to do. You've got a linear ticket. You've got something that you're giving the agent. In the real world, you've got a client that's got a very open-ended problem, and you need to break it down yourself, and your agent needs to break it down himself, and then needs to use sub-agents that do this research. And this is how better tools and organization work together, because this one is MatQA, which is another form of knowledge benchmark. MatQA is something that's working fast in Snowflake, and it's PDF-based, enterprise task. It's got PDFs, and it's got OCR versions of the PDF. And the current state of the leaderboard really shows that both tools and organizations are necessary. And you can see that in the fact that with BM25, however optimized it gets, the human in Gemini 3 reached the same setting. And that doesn't mean that Gemini 3 is as good as a human. That means that even a human cannot get the right information given unlimited searches with BM25. So you have the tool setting, and you need better tools to go forward. That's the tool optimization part of the loop. Thankfully, we've got better tools. We've got models that can handle PDFs. We've got vision. We don't need to rely on OCR text. And what we see with that is that we get another jump, which is Gemini and the mixed-bread search tool, which is fully multimodal, so it can read the PDF. You get the tables. You get all that nice stuff in your own search. And that gets us a big jump in accuracy. But the interesting part is that that doesn't work well. That doesn't work. That does work. But that doesn't work as well as we would like because why is my agent getting 88.9 if the human is getting 99.4? That's 10% I'm leaving on the table here. But I don't understand that's an agent. It gets, I think, 10 tons in the benchmark. So it's a fully agentic system. It gets to think about its results. And yet, it's missing performance. And that's where we introduce the mixed-bread search agent, which is exactly that breaking down of work we saw earlier, where we basically tell the manager, the one answering the question, being like, OK, that's a big topic. There's thousands of PDFs. You're not going to search yourself. Just please break down the problem for me. Please write queries about the aspects that you think are important to answer the actual query. And we get searchers that go off on their own, and they find the right results, and they The human is getting 99.4? That's 10% I'm leaving on the table here. But I don't understand that's an agent. It gets, I think, 10 tons in the benchmark. So it's a fully agentic system. It gets to think about its results. And yet, it's missing performance. And that's where we introduce the mixed-breed search agent, which is exactly that breaking down of work we saw earlier, where we basically tell the manager, the one answering the question, "OK, that's a big topic. There's thousands of PDFs. You're not going to search yourself. Just please break down the problem for me. Please write queries about the aspects that you think are important to answer the actual query." And we get searchers that go off on their own, and they find the right results, and they bring a little memo to your agent, and then your agent actually answers that. And that gets the accuracy by 3.5 points. And that doesn't sound like a lot, but I like to think of it as an Oracle gap. And the Oracle gap is the difference between perfect documents and your search system. And the Oracle gap here is about 10 points before using the agents, and it goes down to six points after using the agents. So that means we have about a 40% reduction in mistakes. The gap between humans and agents goes down by 40% just by having a better architecture to search through it. And I think this is the end of my slide because I'm running out of time, and that's perfect, because that's my takeaway slide. And what I want you to get from this talk is that we know how to design better knowledge work for humans. And AI agents really benefit from this pattern. We have designed this. We know how to do this. Humans have worked on this for centuries. People have always needed more knowledge. Empires used to have librarians. We have paralegals. We've got legal firms. We know exactly how the legal industry has figured out paralegals. The medical industry has figured it out. And none of it looks like programming. Programming has a very different system because it's a very specific use case, and we should really learn from the knowledge world to know how to design agents that will do work for the knowledge world. And then you must not overfeed on tools because tools don't exist as a way to do things by themselves. Tools exist as a way to overcome ceilings. You want a better tool when you see that you're hitting a ceiling, that your performance is not where you want it to be. So we designed better tools to overcome those ceilings. And more importantly, the tools need to be co-designed with the agents. The agents need to know how to use tools because one thing you would often say is agents will try to write grep queries because grep's everywhere in the training data. BM25's everywhere in the query in the data. And that's not always what you need. Sometimes you need semantic search of a PDF, and you can grep a PDF. You can BM25 a PDF. You need to write a better query. So it's very important that your agentic harnesses or even your agentic models know that they have got more than one tool, and it's about primitives. grep's a primitive. BM25's a primitive. And semantic search is a primitive. And all of those need to be very well trained. The models need to know about all of them. And the last one is that the right orchestration of search will get you much better results because context is a finite resource. And even if we get to a model that's got 100 million token context, A, that's going to cost you a lot of money. And B, that's still nothing. You're not even getting half of one state's legal code, let alone the U.S., let alone International Law, let alone the Specialist Course, et cetera. So you need to have a way to break down your task, and you need to have your orchestrator, your main agents, and people that can actually organize the knowledge for them. Yeah, so we've got two minutes for questions. Thank you. There's a lot of contextual information here. But more importantly, the search starts from an intent. Like even if you're doing, again, legal example, if you're asking about like a specific rule that you want to apply to a specific domain, you're going to need to look at the international norms that apply and then do they apply in this case. There's a lot of conditional information that is not predefined in the task. Like that's all up for the agent to find. And so non-code knowledge is very contextual and meaning-driven, which is much harder than code. And that's led to the fact that none of what I'm saying is new. Like people have been doing knowledge work for a very, very long time. And that's resulted in like two endless loops. So you have a tool loop, which is at the start we were like talking. Then at some point some guy was like, we should write stuff down. Then in Alexandria we had the Pinakes, which was the curator of the library of Alexandria, came up with an idea that maybe we should have a way to catalog all of the books we have. Then we developed writing. Then we developed bibliographies. Then we ended up with like the current version of the like Dewey system for libraries. And nowadays we have search engines. But we also had an organization loop, which is giant births or disjoint from the tool one, which is it used to be the one gifted expert. Like we've all heard of the polymath of the past, the person who just knew everything about one domain or all domains. And you just went to them if you had information. But that doesn't scale. So we ended up with like monasteries, which were like guardians of knowledge. And then we had universities and then we ended up creating the bureaucracies. And now we ended up creating the modern organization of work where we have very specialized firms. Like at hospitals, you've got the doctor, you've got the senior doctor, you've got the nurse practitioner, the nurses, the healthcare assistants, and all of them kind of like specialize on different levels of tasks. And that's a really good form of optimization. But the thing is that it's actually just the one loop. Like I'm showing two loops here, but they're actually just the one loop, which is we have new knowledge and new knowledge means that we need better tools. And better tools mean that we end up creating new workflows, new roles. Like we need people that are trained to use those tools, people that understand what the new tool does. If you have a guy that knows how to go to the library and you're like, okay, use Google, you need the knowledge of what Google is. Like that person needs to be taught that's a search engine. You can just type stuff in it. There's no need to physically go there. And that means you retrain, you get new knowledge workers who are more efficient, so they create more knowledge. So we need new tools and so on and so on. So both the tool loop and the organizational loop are actually just the one self-optimizing loop that kind of triggers the other endlessly. And the thing about tooling and optimization is that they're not neutral add-ons. Like I said that we keep optimizing tools and things come up and we create new things out of those tools. But that's never actually a neutral thing. Like tooling is not just, oh, my search is 5% better. The fact that we have a tool or the fact that we don't have a tool is what decides if a task, not if the task is possible because you can do things without the right tool, but if the task is actually scalable and can be carried out cheaply because something being cheap means it can scale. And it's like, yes, of course, if you go to the library of Alexandria before the Pinacchis, you can find your manuscript somewhere. Like whatever you're looking for is there. It's probably going to take two or three weeks. So you're going to really, really, really need that knowledge. But if there's a library catalog, it's going to take you 10 minutes, and now it's way easier to just, oh, okay, I need to know something more about this. I'm going to search for it. Likewise, if you have a map directory or if you even have a map in the first place, which in itself is a tool for information, then exploring the world is a much better idea. Like you're not going to rely on randomly discovering America on your way to the Indies. You know where you're going. And likewise, if you have like a multimodal search platform, then you can search millions of PDF in a way that we couldn't before. So now there's a lot of use cases where you're like, oh, it's in the archives. I'm not going to touch that. That becomes actually useful. And in practice, this kind of looks like that, and I'm getting into the more technical stuff here, which is on a simple deep research task. So this is the Browscom Plus leaderboard, which is made to evaluate the quality of search tools on a very abundant deep research task. You have 200 found documents and you have like specific queries. We all talked about this this morning, and it's a really useful benchmark to like analyze queries. And what we see is that like, okay, a bad tool. So that's the thing that people often rant about. You will see that there's two BM25 here. There's two like optimized and unoptimized, and that's because quite often people will tell you BM25 is not great. And the reason they'll tell you BM25 is not great is because there's not one BM25, there's hundreds of them. It's a way to do a lexical search. You should always optimize your baselines. You should always like optimize what you're beating. And so what you see here is like a badly optimized tool is useless, like 60% accuracy. You're not going to trust someone that's right 60% of the time. You're just going to do it yourself. When you start optimizing the tools, you can see we go up to 70, 80, and then the actual best is a hybrid harness. It gets to 90. But that's maybe not the most interesting part because we start kind of plateauing at one point. Like the jump from 89.8 to 90.2 is in-run variance. That doesn't matter. What matters here, however, is that 90.2% accuracy, you reach it with 20% fewer tool calls. And that's huge because in practice that's 20% fewer tokens, 20% fewer resources that you use. That's basically 20% free cash. And if you compare it to the unoptimized baseline, you're spending like 5% of what you were spending in the first place. So the tool is actually what makes the task worth doing. Nobody would keep using that tool if it takes 25 calls. But if it takes eight calls, you're like, oh, yeah, cool. That's a workflow I can introduce. And the second part, which goes with tooling, and I think is just as important because Browscom Plus in the previous slide is interesting, but it's easy. It's 100,000 documents. It's just text. It's just the one question. And it's not really that open-ended. It's just a bit convoluted. But when you're actually doing like real-life knowledge work, there is that workflow that you see that I try to do that, which is you have a client. They come to the big shot. They come to the lawyer. That's the partner of the agency. And they're like, okay, this is my situation. That's my problem. And they meet together. But then what the partner does is they're not going to be the ones doing all the legal research. They're not going to be the ones doing every single step of the problem. What they'll do is kind of understand that, like, okay, this person has this problem. That's going to cause them that. Those are the facts. I need the relevant laws to this, that, and so on aspects. And then they've got party goals. They've got assistants. And the assistants are going to be doing this research. They're going to be using the size of tools they've been trained to use. And they're going to produce memo and notes. And then they're going to give that back to the big shots. Maybe they'll research, like, one clarification point, but they mostly rely on what their searcher agents, if you want, like their assistants are fun for them. And that's the response that you're going to get. And that's echoing the point I made before, which is in code, when you're using code code, you're kind of doing that work yourself. You've already broken down the query. You know what you want to do. You've got a linear ticket. You've got something that you're giving the agent. In the real world, you've got a client that's got a very open-ended problem, and you need to break it down yourself, and your agent needs to break it down himself, and then needs to use sub-agents that do this research. And this is how better tools and organization work together, because this one is MatQA, which is another form of knowledge benchmark. MatQA is something that's working fast in Snowflake, and it's PDF-based, enterprise task. It's got PDFs, and it's got OCR versions of the PDF. And the current state of the leaderboard really shows that both tools and organizations are necessary. And you can see that in the fact that with BM25, however optimized it gets, the human in Gemini 3 reached the same setting. And that doesn't mean that Gemini 3 is as good as a human. That means that even a human cannot get the right information given unlimited searches with BM25. So you have the tool setting, and you need better tools to go forward. That's the tool optimization part of the loop. Thankfully, we've got better tools. We've got models that can handle PDFs. We've got vision. We don't need to rely on OCR text. And what we see with that is that we get another jump, which is Gemini and the mixed-bread search tool, which is fully multimodal, so it can read the PDF. You get the tables. You get all that nice stuff in your own search. And that gets us a big jump in accuracy. But the interesting part is that, like, that doesn't work. Well, that doesn't work. That does work. But that doesn't work as well as we would like because why is my agent getting 88.9 if the human is getting 99.4? Like, that's 10% I'm leaving on the table here. But I don't understand that's an agent. It gets, I think, 10 tons in the benchmark. So it's a fully agentic system. It gets to think about its results. And yet, it's missing performance. And that's where we introduce the mixed-bread search agent, which is exactly that breaking down of work we saw earlier, where we basically tell the manager, the one answering the question, being like, OK, that's a big topic. There's, like, thousands of PDFs. You're not going to, like, search yourself. Just please break down the problem for me. Please write queries about the aspects that you think are important to answer the actual query. And we get searchers that go off on their own, and they find the right results, and they bring you, like, a little memo to your agent, and then your agent actually answers that. And that gets the accuracy by 3.5 points. And that doesn't sound like a lot, but I like to think of it as, like, an Oracle gap. And the Oracle gap is the difference between perfect documents and your search system. And the Oracle gap here is about 10 points before using the agents, and it goes down to six points after using the agents. So that means that we have about a 40% reduction in mistakes. Like, the gap between humans and agents goes down by 40% just by, like, having a better architecture to search through it. And I think this is the end of my slide because I'm running out of time, and that's perfect, because that's my takeaway slide. And what I want you to, like, get from this talk is that we know how to design better knowledge work for humans. And AI agents really benefit from this pattern. Like, we have designed this. We know how to do this. Humans have worked on this for centuries. Like, people have always needed more knowledge. The empires used to have librarians. We have paralegal. We've got legal firms. We know exactly how the legal industry has figured out paralegal. The medical industry has figured it out. And none of it looks like programming. Programming has a very different system because it's a very specific use case, and we should really learn from the knowledge world to know how to design agents that will do work for the knowledge world. And then you must not, like, overfeed on tools because tools don't exist as a way to do things by themselves. Whereas tools exist as a way to overcome ceilings. You want a better tool when you see that you're hitting a ceiling, that your performance is not where you want it to be. So we designed better tools to overcome that ceilings. And more importantly, the tools need to be co-designed with the agents. Like, the agents need to know how to use tools because one thing you would often say is agents will try to write grep queries because grep's everywhere in the training data. BM25's everywhere in the query in the data. And that's not always what you need. Sometimes you need semantic search of our PDF, and you can grep a PDF. You can BM25 a PDF. You need to, like, write a better query. So it's very important that your agentic harnesses or even your agentic models know that they have got more than one tool, and it's about primitives. grep's a primitive. BM25's a primitive. And semantic search is a primitive. And all of those need to be, like, very well trained. Like, the models need to know about all of them. And the last one is that the right orchestration of search will get you much better results because context is a finite resource. And even if we get to a model that's got, like, 100 million token context, A, that's going to cost you a lot of money. And B, that's still nothing. You're not even getting half of, like, one state's legal code, let alone the U.S., let alone the International Law, let alone the Specialist Course, et cetera. So you need to have a way to break down your task, and you need to have your orchestrator, your main agents, and people that can actually organize the knowledge for them. And, yeah, so we've got two minutes for questions. Thank you.