On AI and Knowledge — Pablo Castro, Distinguished Engineer & CVP for AI Knowledge, Microsoft
Description
Pablo Castro explores AI and knowledge systems for building better applications and agents. Speaker: Pablo Castro —Distinguished Engineer and CVP, Microsoft, leads the AI Knowledge team in Microsoft's CoreAI division, where he focuses on state-of-the-art information understanding and retrieval systems for AI applications and agents, including Foundry IQ, Azure AI Search, and Azure Content Understanding. LinkedIn: https://www.linkedin.com/in/pabloc Timestamps: 0:00 Introduction and speaker background 1:14 Defining the nature of knowledge: Intrinsic, Extrinsic, and Learned 1:27 Intrinsic knowledge and the history of AI coding tools 4:38 Extrinsic knowledge and corporate data grounding 7:06 Evolution of retrieval systems and Foundry IQ 9:56 Foundry IQ demo: Building a knowledge base 13:08 Learned knowledge: The agent learning loop 14:25 Foundry agent optimization demo 16:49 Closing remarks and resources Key quotes Intrinsic Knowledge Perspective: This knowledge represents the foundational parametric memory of models. "Intrinsic knowledge is just the knowledge that comes with the models... it's what started many of the scenarios that then grew on all the things we're doing with agents today." (1:27 - 1:48) "I would argue that GitHub Copilot and ChatGPT, those sort of experiences, were heavily grounded on this intrinsic memory—what the models already knew." (2:59 - 3:04) Extrinsic Knowledge Perspective: To be truly useful in an organization, agents must access private, ambient data through sophisticated retrieval. "Intrinsic model got us here, but it only gets you so far if you're building a system that or an agent that needs to participate in what's happening in an organization." (4:41 - 4:49) "The trick is how do you build a platform that allows you to combine all these building blocks without putting the complexity right in front of you." (8:02 - 8:10) "For more sophisticated cases you do want a system that can reflect on what's in the data set and decide whet
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Effective enterprise agents require a layered knowledge strategy: leverage model-native knowledge, ground agents in governed organizational and web data through hybrid and agentic retrieval, then continuously improve agent behavior from evaluation and production traces.
- Why it matters: The talk provides a concrete control-plane view of knowledge for agents—retrieval, context assembly, evaluation, optimization, and deployment—that maps directly to building reliable, cost-efficient agent systems.
- Best use: Use it as a product-informed architecture briefing on Microsoft's Foundry IQ and Agent Optimizer approach, while extracting the general patterns for retrieval and closed-loop agent improvement rather than treating the vendor claims as independent benchmarks.
Executive Summary
Pablo Castro frames AI knowledge in three layers. Intrinsic knowledge is what models absorb during training and store in parametric memory; it powered the initial leap from code completion to Copilot-like generation. But intrinsic knowledge alone is insufficient for an agent operating inside a company, where answers and actions must reflect current documents, communications, analytics, and external information.
For that organizational context, Castro argues that RAG has evolved from basic vector-database lookup into context engineering: a system must connect to multiple enterprise data domains and combine retrieval methods rather than relying on embedding similarity alone. Microsoft positions its IQ portfolio as a unified access layer across work data, analytics assets, agent-owned data, and the public web, with Foundry IQ exposing either managed defaults or low-level index and retrieval controls.
The central technical argument is that hybrid retrieval and agentic retrieval outperform isolated, single-shot methods on hard questions. Agentic retrieval can inspect whether gathered evidence satisfies the request, perform additional reasoning over the corpus, and return more complete evidence, but it introduces an explicit quality-versus-latency trade-off. The platform goal is to hide this complexity for standard cases while preserving expert controls over indexing, quantization, lexical retrieval, and other retrieval internals.
Finally, Castro presents learned knowledge as the operational learning accumulated by people and agents. Foundry's Agent Optimizer turns this into a loop: establish an evaluation baseline, generate candidate configurations, evaluate them against a task-adherence rubric, and apply a better configuration. The key implementation prerequisite is externalizing agent instructions, tool definitions, skills, and related configuration so optimization can safely replace configurations rather than rewrite application code.
Key Takeaways
- Claim: Model-intrinsic knowledge created the initial generative-AI productivity inflection, but it is only one of three knowledge layers needed for useful agents. | Evidence: Castro traces coding assistance from IntelliSense in 1996, to ML-ranked completion 22 years later, to GitHub Copilot three years after that, arguing that early Copilot and ChatGPT experiences were heavily grounded in what models already knew. | Implication: Treat base-model capability and reasoning as a foundation, not as the enterprise knowledge architecture; agent design must deliberately add external grounding and a mechanism to learn from execution. | Caveat: Parametric knowledge is inherently insufficient when an agent must reflect an organization's current and private operating context.
- Claim: Enterprise grounding must span both the agent's purpose-built knowledge and ambient organizational data. | Evidence: Castro distinguishes agent-specific knowledge from ambient sources such as SharePoint documents, email, calendar, chat threads, people connections, data warehouses, data lakes, and Power BI reports; he also includes web information where public sources complete the agent's worldview. | Implication: Knowledge access should be designed as a federated, permission-aware layer across systems of record rather than as a collection of isolated document indexes maintained separately by each agent. | Caveat: Broad access to ambient data is presented as a capability need, but the talk does not address permission propagation, tenant boundaries, or data-governance controls in detail.
- Claim: Vector similarity alone is not a sufficient retrieval strategy; combining retrieval methods produces better real-world results. | Evidence: Castro says industry evaluations repeatedly show that individual retrieval methods underperform combined approaches, citing Azure AI Search evaluations in customer scenarios and contrasting hybrid methods with a prior belief that increasingly accurate cosine similarity would solve retrieval. | Implication: Build retrieval evaluation around answer quality and evidence coverage, then test hybrid lexical/vector/ranking approaches instead of optimizing embedding similarity in isolation. | Caveat: No benchmark values, datasets, or comparison methodology are supplied in the transcript, so the magnitude and generality of the claimed improvement cannot be independently assessed here.
- Claim: Retrieval infrastructure should support a managed-to-expert spectrum instead of forcing every team to choose between black-box convenience and bespoke search engineering. | Evidence: In the Foundry IQ demonstration, a user can submit PDFs or images and have the system handle chunking, vectorization, relevance, ranking, and agentic retrieval; advanced users can instead control vector indexes, quantization, lexical retrieval, indexing algorithms, and inspect chunk organization. | Implication: Standardize common ingestion and retrieval defaults for speed, but retain escape hatches for domain-specific evaluation failures, cost tuning, and compliance-driven retrieval behavior.
- Claim: Agentic retrieval is most justified for difficult information needs, where the system must assess evidence sufficiency rather than issue a single retrieval query. | Evidence: Castro contrasts quick, single-shot retrieval for easy cases with an agentic stack that reflects on the dataset and determines whether the input's information need has been met; Microsoft reports better evidence recall and answer completeness in its evaluations for difficult cases. | Implication: Route requests by difficulty: use low-latency retrieval for straightforward lookup, and reserve iterative retrieval/reasoning loops for multi-hop, ambiguous, or completeness-sensitive work. | Caveat: The speaker explicitly presents effort as a latency-versus-quality setting, and does not claim agentic retrieval is necessary for simple requests.
- Claim: Knowledge-base design should be reusable across agent runtimes, with retrieval assets exposed as interoperable tools rather than embedded as application-specific glue. | Evidence: After creating a knowledge base spanning blob-storage PDFs, structured Parquet tables, and web grounding, Castro says it can be attached to a Foundry agent or used independently because every knowledge base is exposed as an MCP server. | Implication: Separate knowledge assets from individual agent implementations and expose them through stable interfaces, allowing the same governed retrieval capability to serve multiple agents and evaluation harnesses.
- Claim: Agent behavior can be systematically improved through an evaluation-driven optimization loop, provided the agent's configurable behavior is externalized. | Evidence: Foundry Agent Optimizer can generate a task-adherence evaluation from agent instructions and traces, establish a baseline, generate candidates, and hill-climb against the rubric. In the demo, an optimization run took roughly 45 minutes and produced non-handwritten instructions that could be applied by swapping externalized configuration. | Implication: Version prompts, instructions, tool definitions, and skills as deployable configuration, then gate automated agent changes through representative offline evaluations and controlled production rollout. | Caveat: Optimization quality is bounded by the evaluation dataset and rubric; the transcript does not establish that generated changes generalize beyond the evaluated tasks.
Detailed Brief
Microsoft product architecture presented in the talk
- Claims: Microsoft positions GitHub as the build environment, Azure/Cloud as the contextualization layer, and Foundry as the hosting, observability, management, and model-access layer for agents.; Foundry's model catalog is described as containing thousands of models, enabling task-specific model selection within the same agent platform.; The IQ naming scheme separates organizational work content, analytics content, agent-ingested content, and web content into Work IQ, Fabric IQ, Foundry IQ, and Web IQ.
- Evidence: Castro states that Claude in Microsoft Foundry became generally available the day before the talk.; Work IQ is described as connecting documents, email, calendars, chats, and people relationships; Fabric IQ covers data warehouses, lakes, and Power BI reports.
- Caveats: This is a Microsoft product presentation, so the component taxonomy and capability claims should be validated against current product documentation, pricing, regional availability, and security features before making platform commitments.
- Implications: A unified agent platform can reduce integration overhead, but platform consolidation should be weighed against portability requirements and dependence on a provider's retrieval and optimization abstractions.
Cost and observability considerations
- Claims: Token efficiency is treated as a first-class retrieval objective: the desired result is the most information-dense response context with the fewest tokens.; Agent traces are not only an observability artifact; they become input to evaluation generation and behavior optimization.
- Evidence: The demo exposes a user-selected effort level that controls a latency-versus-quality trade-off.; The optimizer is described as reflecting on actual agent traces in addition to instructions, tools, and skills.
- Caveats: The transcript offers no token-reduction figures, cost benchmarks, or description of how trace quality, privacy, and retention are governed.
- Implications: Instrument retrieval and agent execution so teams can measure evidence quality, latency, token consumption, and task success together rather than optimizing any metric independently.
Notable Concepts & Terms
- Intrinsic knowledge: Knowledge encoded in model parameters through training; it explains strong zero- or few-shot capability but cannot reliably represent current enterprise state.
- Extrinsic knowledge: Knowledge supplied outside model weights through organizational systems, agent-managed sources, or the web; this is the grounding layer for current, attributable agent work.
- Learned knowledge: Organizational know-how accumulated from work processes and agent execution, captured through evaluation and optimization loops.
- Context engineering: The evolved form of RAG described here: assembling the right multi-source, retrieval-ranked context for an agent rather than merely fetching nearest document chunks.
- Agentic retrieval: An iterative retrieval approach in which the system assesses whether evidence satisfies the request and can continue searching or reasoning; aimed at difficult cases.
- MCP server: A Model Context Protocol endpoint; Castro uses it to make a knowledge base reusable by agents or external harnesses without custom glue code.
- Agent Optimizer: Foundry component that generates or uses evaluations, creates candidate configurations, scores them, and applies an improved externalized agent configuration.
- JEPA-style loop: Castro's label for the iterative candidate-generation and evaluation process used in the optimizer demo; the transcript does not define the acronym or mechanics further.
Operator Notes / Why Ken Should Care
- Adopt a retrieval test suite that measures evidence recall, answer completeness, latency, and token use; compare vector-only retrieval against hybrid and iterative retrieval on the actual hard queries your agents face.
- Classify agent requests into simple lookup versus difficult/completeness-sensitive work, and set explicit latency and spend budgets for each route.
- Externalize and version instructions, tool schemas, skills, retrieval settings, and model-routing configuration before introducing any automated optimization workflow.
- Require representative holdout tasks and rollout gates before applying optimizer-generated prompt or tool changes; do not promote candidates solely because they improve the optimization rubric.
- Design reusable knowledge services with stable interfaces such as MCP, while separately auditing source permissions, trace retention, and cross-system access controls.
Source/Metadata
- Title: On AI and Knowledge — Pablo Castro, Distinguished Engineer & CVP for AI Knowledge, Microsoft
- Transcript words: 2904
- Duration seconds: 1055
- Timestamp note: No timestamps or chapters were present in the supplied transcript.
Transcript
[SPEAKER_00] Now taking the stage is CVP and distinguished engineer at Microsoft, Pablo Castro. [SPEAKER_01] Hello everyone. Hello everyone, good morning. It's great to be back here at the AI Engineer World's Fair. Now, my job at Microsoft is to connect the dots between AI and knowledge. As an information retrieval nerd, that's great for me. I spend a lot of time looking at knowledge representation, extraction, search, and whatnot. And thinking about agents and knowledge really invites reflection on what it means to know something. And the nature of how do we get things done based on what we know. Next slide. [SPEAKER_00] All right. [SPEAKER_01] There. So this morning, what I thought we would do is spend a little bit of time talking about the nature of knowledge and split it into these three categories of intrinsic, extrinsic, and learned. Intrinsic knowledge is just the knowledge that comes with the models. It's what we train the models on, the training data, and what's stored in the models, parametric memory. And while it's the obvious thing, I would argue this is the knowledge that actually threw us into the exponential we are in today. It's what started many of the scenarios that then grew into all the things we're doing with agents today. Let me give you an example with code. So I wrote these two pieces of code about 25 years apart. And yet the process to put this thing together was surprisingly similar. I had to sit down with what I knew or what I had to go look up and then just write it up. And while I'm illustrating this with knowledge, you could say the same thing about writing an email or creating a summary of a document. Now, you can see this exponential at play in tasks like these where I'm sure you can go further back. But an interesting point in time to start looking at this would be when Microsoft introduced IntelliSense. That was in 96. And it was great. You didn't have to remember function signatures anymore and whatnot. It took 22 years from there to go to the next step, where machine learning helps us actually rank the options we give you in IntelliSense so it's quicker to pick the right choice. Just three years after that, GitHub Copilot launches. And that was one key inflection point. This was even before ChatGPT was announced. And I would argue that GitHub Copilot, ChatGPT, those sort of experiences were heavily grounded on this intrinsic memory, what the models already knew. From there, of course, things shifted a couple of years later. Courser launches, GitHub Copilot X launches, and how we do things evolves really quickly. Which takes us to late last year, Opus 4.5 ships, and then in rapid succession, GPT, Opus, and other models keep getting better and better at coding. Which takes us to early this year, where incredibly successful software like OpenClaw comes into existence with not a single line of code written by it. And then we're going to have a single line of code written by hand. So this is the shape of the exponential we're in. And a lot of this was powered by the intrinsic knowledge in models and, of course, their ability to reason. Now, in the context of Microsoft, we want to make available all these models and make it easy for you to integrate them into the agents you're building. We do this from our agent platform that starts in GitHub, where we all go and build. Cloud has a contextualization system so you can ground your agents. And when it comes to agent hosting, observability, and management, we do all of these in Foundry. Microsoft Foundry is also where we offer thousands of models in our model catalog. So you can pick whatever is the right model for the right task. And we keep adding more every day. In fact, just yesterday we announced that Cloud in Microsoft Foundry is generally available. So you can use all the capabilities of Cloud in the context of the unified experience in Foundry. So you get the best of both worlds. Now, intrinsic models got us here, but they only get you so far if you're building a system or an agent that needs to participate in what's happening in an organization or a company. And as an industry we realized this early, and we saw the RAG pattern emerge. That started as a pretty low-tech technique, but quickly evolved into what we do today with context engineering, and it became a pretty sophisticated system for connecting agents and the knowledge they need to get their job done. Of the many dimensions in which this got complicated, I'm going to pick on two. One is the evolution from simple and isolated data sets to whole company-wide grounding. And the other one is how we started with simple vector search and whatnot, and we really saw this evolve into fairly complicated retrieval systems. So let's start with company grounding. At Microsoft, spending time with customers, one of the things we saw early was that whenever you build an agent, you always have the knowledge you care about for that agent and you manage that yourself, but you also need to ground the agent often on the ambient data of your organization whenever the agent leaves. This includes maybe your documents, your emails, your chat threads, or the information in your data warehouse and whatnot. So we built Microsoft IQ as a way to give you a single entry point into all this ambient data that agents need to get their job done in addition to the specific information that you build into the agent. Microsoft IQ is not one feature. It's more like a set of capabilities that goes from Work IQ, which connects your agents to all the documents in, say, SharePoint, all the emails, calendar, your chats, and the connections between people, to Fabric IQ, that gives you access to all your analytics assets, from data warehouses and data lakes to Power BI reports, and Foundry IQ, which is what you use for your agents where you can push your own data and then use it for grounding. And of course, sometimes your agents need to go out to the web to ground on data. Maybe not yours, it's public information, but you need to use it to complete the picture of what the agent world view is, and for that we have Web IQ. Now, this first part allows agents to ground on this ambient data. Now, the second dimension I mentioned before is the evolution of the actual retrieval systems. When RAG first emerged, I think what we saw was an initial adoption of vector databases that really unblocked us from getting a lot of these systems off the ground. And that was great. I think for a hot second as an industry, we thought that if we could get really, really, really good at computing cosine similarity between vectors, we were all set for retrieval. It turns out things are never that easy. So what evaluations show over and over again is how, if you combine methods, you just get better results. In this case, this is an evaluation from Azure AI Search, the search technology behind Foundry IQ. And you can see how individual methods don't do as well as combined methods, When RUG first emerged, I think what we saw is an initial adoption for vector databases that really unblocked us from getting a lot of these systems off the ground. And that was great. I think for a hot second as an industry, we thought that if we could get really, really, really, good at computing cosine similarity between vectors, we were all set for retrieval. It turns out things are never that easy. So what evaluations show over and over again is how, if you combine methods, you just get better results. In this case, this is an evaluation from Azure AI Search, the search technology behind Foundry IQ. And you can see how individual methods don't do as well as combined methods, particularly when you apply them to real-world customer scenarios. Now, the trick is how you build a platform that allows you to combine all these building blocks without putting the complexity right in front of you. It lets you opt into it when you need control. But when you have a scenario that is clear, then you can have an easy system. So, in Foundry IQ, that was one of our core design goals. And the way we do this is we actually layer the system. So, you can start at the top, you can go to Foundry and say, hey, I have a bunch of, I don't know, PDFs or pictures over there. Just deal with them. And then we'll do everything under the covers. We'll do chunking, vectorization, deal with relevance and ranking, deal with agentic retrieval and whatnot. Now, if you're an expert and you want control, you can also do that. You can go to the bottom of the stack. You want to build vector indexes and tell us how to quantize the vectors or control lexical retrieval and whatnot. You can do all of that. And you can do it in the same stack, which means you can go up and down as your needs change. Now, on top of the core retrieval system, we also introduced an agentic retrieval stack because we see that for easy cases, quick single-shot retrieval is great. But for more sophisticated cases, you do want a system that can reflect on what's in the dataset and decide whether or not we've satisfied the information needed stated in the input before we come back with results. Of course, we see a lot of patterns like this emerge. And always the question is, is this actually useful? Are the results better? Our experience in our own evaluations is for difficult cases, agentic retrieval can make a difference. Across the many metrics that we track, things like the actual evidence recall or answer completeness, we see the agentic retrieval approach continuously does better than simple, than individual simple parts. Now, let me show you some of these in action, if we can go to the laptop. Can we switch to the laptop? There you go. Okay. Sure. Sure. Sure. Sure. Sure. And I can say how much effort you want the model to make or the system to make. And this is effectively a trade-off between latency and quality. I can configure a number of other things, but critically, I want to say where the data I want to ground [SPEAKER_01] is coming from, and I can start from scratch. [SPEAKER_01] Or in this case, I have a bunch of unstructured data, [SPEAKER_01] like PDFs and whatnot in blob storage. [SPEAKER_01] I have structured parquet tables with statistics, and I also want to ground on the web. So if I take these three steps, and then I save this knowledge base, now I have this asset, this knowledge base, that I can connect to a Foundry agent right here, and it'll take a second. But also it's a standalone asset that if I already have a harness that I'm using in other places, every knowledge base is an MCP server, so you can just connect to it without having to write any glue code in the middle. Now, a knowledge base like this has a bunch of parts. Some of them, for example, this storage content, you usually build indexes and you vectorize these things and whatnot, and if you want control over that, if you don't, you can just use it here. But if you do, let me just switch to Azure and show you the service behind that particular instance, where if I go to knowledge bases, this is the knowledge base we just created a second ago, and I can go peek inside. For example, I can go fish out the indexes that back this particular piece of content, and in that index, I can see what is the structure of the index. If I'm opinionated about, I don't know, maybe the quantization approach I want to use, or which indexing algorithm I want for my vectors, I can say all of that, and of course, I can actually go and explore the data and see what's inside, how chunks were organized and whatnot. So the goal of this is to, again, give you a highly productive environment when you don't need the sophistication, and when you need it, to make sure you have it, to get your job done. Go back to slides. And of course, the other aspect of this is, top of mind these days for all of us, token efficiency. And so we carefully evaluate this system to make sure that we give you the most information-dense answer that has the fewest tokens, so that your consumption of tokens has a high value when it comes to all retrieval tasks. The last category of knowledge I wanted to talk about is learned knowledge. Now, learned knowledge is the result of us doing the work we do as individuals and as organizations every day. And the idea that we can actually observe the processes and get better at them by reflecting and improving every step of it is something that has really changed now that we have agents doing the work and we can go tune the agents automatically. Satya wrote about this recently and reflected on the fact that people and agents can really compound in how they do the work and how they can create this learning loop that effectively captures what's unique about the company or the organization you're working on and inputs that to work to differentiate the work that you do. Now, in Foundry, we wanted to offer a materialized version of this that you can use today. So we built a component called the Agent Optimizer that effectively goes through this process and allows you to evaluate the baseline, generate candidates, Satya wrote about this recently. and reflected on the fact that people and agents can really compound in how they do the work and how they can create this learning loop that effectively captures what's unique about the company or the organization you're working on and inputs that to work to differentiate the work that you do. Now, in Foundry, we wanted to offer a materialized version of this that you can use today. So we built a component called the Agent Optimizer that effectively goes through this process and allows you to evaluate the baseline, generate candidates, and then evaluate the new candidates. If we have a strong result, then deploy that to production. Let me give you a quick flavor of what this looks like if we can switch back to the laptop. All right. So here I'm in VS Code. I have the Foundry Toolkit installed. And I have a simple agent. It doesn't matter how you write your agent, as long as you externalize configuration, your instructions, tool definitions, skills, and whatnot. So once you have one of those, it takes two key steps to do this. So first... Whoops. I can actually... So usually you have an evaluation already, but if you don't, you can actually say, eval generate, and what we'll do is we'll look at what we know about the agent, traces and instructions and whatnot, and we'll produce a task adherence-focused evaluation for you. In this case, I ran this a little bit earlier. So just to give you a flavor of what this looks like, you have a bunch of tasks, and then the questions and the criteria and whatnot. Once you have a dataset you can evaluate, then the next step is you can say optimize. And I could just run optimize on its own, and that will run, in this case. This ran for maybe 45 minutes or so, and you get an optimized version by effectively hill-climbing the metric that's established from the evaluation. So I ran this earlier, and so let me show you the output for this particular one, where you can see that we established the baseline first, and then we kept iterating on candidates using different combinations, using a JEPA-style loop, and looking for options that perform better, given the rubric that we have. And the interesting thing is that once you found one that is better, then you can simply just say optimize, apply, and what this does is, since you externalize the configuration, it allows you to swap one configuration for the other. If we look here, you can see that, for example, I have a baseline and the one we just applied, and just to pick on instructions, these are just the trivial instructions for this example agent. But if I look at the optimized one, then you can see a bunch of instructions that are not handwritten, but that emerged out of the hill-climbing process to make this particular agent better, given what we have in terms of instructions and skills and tools, but also based on reflecting on the actual traces from the agent