[SPEAKER_00] Now taking the stage is CVP and distinguished engineer at Microsoft, Pablo Castro.
SPEAKER_00
[SPEAKER_01] Hello everyone.
SPEAKER_01
Hello everyone, good morning. It's great to be back here at the AI Engineer World's Fair. Now, my job at Microsoft is to connect the dots between AI and knowledge. As an information retrieval nerd, that's great for me. I spend a lot of time looking at knowledge representation, extraction, search, and whatnot. And thinking about agents and knowledge really invites reflection on what it means to know something. And the nature of how do we get things done based on what we know. Next slide. [SPEAKER_00] All right.
SPEAKER_00
[SPEAKER_01] There.
SPEAKER_01
So this morning, what I thought we would do is spend a little bit of time talking about the nature of knowledge and split it into these three categories of intrinsic, extrinsic, and learned. Intrinsic knowledge is just the knowledge that comes with the models. It's what we train the models on, the training data, and what's stored in the models, parametric memory. And while it's the obvious thing, I would argue this is the knowledge that actually threw us into the exponential we are in today. It's what started many of the scenarios that then grew into all the things we're doing with agents today. Let me give you an example with code.
SPEAKER_01
So I wrote these two pieces of code about 25 years apart. And yet the process to put this thing together was surprisingly similar. I had to sit down with what I knew or what I had to go look up and then just write it up. And while I'm illustrating this with knowledge, you could say the same thing about writing an email or creating a summary of a document. Now, you can see this exponential at play in tasks like these where I'm sure you can go further back. But an interesting point in time to start looking at this would be when Microsoft introduced IntelliSense. That was in 96. And it was great. You didn't have to remember function signatures anymore and whatnot.
SPEAKER_01
It took 22 years from there to go to the next step, where machine learning helps us actually rank the options we give you in IntelliSense so it's quicker to pick the right choice. Just three years after that, GitHub Copilot launches. And that was one key inflection point. This was even before ChatGPT was announced. And I would argue that GitHub Copilot, ChatGPT, those sort of experiences were heavily grounded on this intrinsic memory, what the models already knew. From there, of course, things shifted a couple of years later. Courser launches, GitHub Copilot X launches, and how we do things evolves really quickly.
SPEAKER_01
Which takes us to late last year, Opus 4.5 ships, and then in rapid succession, GPT, Opus, and other models keep getting better and better at coding. Which takes us to early this year, where incredibly successful software like OpenClaw comes into existence with not a single line of code written by it. And then we're going to have a single line of code written by hand. So this is the shape of the exponential we're in. And a lot of this was powered by the intrinsic knowledge in models and, of course, their ability to reason.
SPEAKER_01
Now, in the context of Microsoft, we want to make available all these models and make it easy for you to integrate them into the agents you're building. We do this from our agent platform that starts in GitHub, where we all go and build. Cloud has a contextualization system so you can ground your agents. And when it comes to agent hosting, observability, and management, we do all of these in Foundry. Microsoft Foundry is also where we offer thousands of models in our model catalog. So you can pick whatever is the right model for the right task. And we keep adding more every day. In fact, just yesterday we announced that Cloud in Microsoft Foundry is generally available.
SPEAKER_01
So you can use all the capabilities of Cloud in the context of the unified experience in Foundry. So you get the best of both worlds. Now, intrinsic models got us here, but they only get you so far if you're building a system or an agent that needs to participate in what's happening in an organization or a company.
SPEAKER_01
And as an industry we realized this early, and we saw the RAG pattern emerge. That started as a pretty low-tech technique, but quickly evolved into what we do today with context engineering, and it became a pretty sophisticated system for connecting agents and the knowledge they need to get their job done. Of the many dimensions in which this got complicated, I'm going to pick on two. One is the evolution from simple and isolated data sets to whole company-wide grounding. And the other one is how we started with simple vector search and whatnot, and we really saw this evolve into fairly complicated retrieval systems. So let's start with company grounding.
SPEAKER_01
At Microsoft, spending time with customers, one of the things we saw early was that whenever you build an agent, you always have the knowledge you care about for that agent and you manage that yourself, but you also need to ground the agent often on the ambient data of your organization whenever the agent leaves. This includes maybe your documents, your emails, your chat threads, or the information in your data warehouse and whatnot. So we built Microsoft IQ as a way to give you a single entry point into all this ambient data that agents need to get their job done in addition to the specific information that you build into the agent.
SPEAKER_01
Microsoft IQ is not one feature. It's more like a set of capabilities that goes from Work IQ, which connects your agents to all the documents in, say, SharePoint, all the emails, calendar, your chats, and the connections between people, to Fabric IQ, that gives you access to all your analytics assets, from data warehouses and data lakes to Power BI reports, and Foundry IQ, which is what you use for your agents where you can push your own data and then use it for grounding. And of course, sometimes your agents need to go out to the web to ground on data.
SPEAKER_01
Maybe not yours, it's public information, but you need to use it to complete the picture of what the agent world view is, and for that we have Web IQ. Now, this first part allows agents to ground on this ambient data. Now, the second dimension I mentioned before is the evolution of the actual retrieval systems. When RAG first emerged, I think what we saw was an initial adoption of vector databases that really unblocked us from getting a lot of these systems off the ground. And that was great. I think for a hot second as an industry, we thought that if we could get really, really, really
SPEAKER_01
good at computing cosine similarity between vectors, we were all set for retrieval. It turns out things are never that easy. So what evaluations show over and over again is how, if you combine methods, you just get better results. In this case, this is an evaluation from Azure AI Search, the search technology behind Foundry IQ. And you can see how individual methods don't do as well as combined methods, When RUG first emerged, I think what we saw is an initial adoption for vector databases that really unblocked us from getting a lot of these systems off the ground.
SPEAKER_01
And that was great. I think for a hot second as an industry, we thought that if we could get really, really, really, good at computing cosine similarity between vectors, we were all set for retrieval. It turns out things are never that easy. So what evaluations show over and over again is how, if you combine methods, you just get better results. In this case, this is an evaluation from Azure AI Search, the search technology behind Foundry IQ. And you can see how individual methods don't do as well as combined methods,
SPEAKER_01
particularly when you apply them to real-world customer scenarios. Now, the trick is how you build a platform that allows you to combine all these building blocks without putting the complexity right in front of you. It lets you opt into it when you need control. But when you have a scenario that is clear, then you can have an easy system. So, in Foundry IQ, that was one of our core design goals. And the way we do this is we actually layer the system. So, you can start at the top, you can go to Foundry and say, hey, I have a bunch of, I don't know, PDFs or pictures over there. Just deal with them. And then we'll do everything under the covers.
SPEAKER_01
We'll do chunking, vectorization, deal with relevance and ranking, deal with agentic retrieval and whatnot. Now, if you're an expert and you want control, you can also do that. You can go to the bottom of the stack. You want to build vector indexes and tell us how to quantize the vectors or control lexical retrieval and whatnot. You can do all of that. And you can do it in the same stack, which means you can go up and down as your needs change. Now, on top of the core retrieval system, we also introduced an agentic retrieval stack because we see that for easy cases, quick single-shot retrieval is great.
SPEAKER_01
But for more sophisticated cases, you do want a system that can reflect on what's in the dataset and decide whether or not we've satisfied the information needed stated in the input before we come back with results.
SPEAKER_01
Of course, we see a lot of patterns like this emerge. And always the question is, is this actually useful? Are the results better?
SPEAKER_01
Our experience in our own evaluations is for difficult cases, agentic retrieval can make a difference. Across the many metrics that we track, things like the actual evidence recall or answer completeness, we see the agentic retrieval approach continuously does better than simple, than individual simple parts. Now, let me show you some of these in action, if we can go to the laptop. Can we switch to the laptop? There you go. Okay. Sure.
SPEAKER_01
Sure. Sure. Sure. Sure. And I can say how much effort you want the model to make or the system to make. And this is effectively a trade-off between latency and quality. I can configure a number of other things, but critically, I want to say where the data I want to ground [SPEAKER_01] is coming from, and I can start from scratch. [SPEAKER_01] Or in this case, I have a bunch of unstructured data, [SPEAKER_01] like PDFs and whatnot in blob storage. [SPEAKER_01] I have structured parquet tables with statistics, and I also want to ground on the web. So if I take these three steps, and then I save this knowledge base, now I have this asset, this knowledge base,
SPEAKER_01
that I can connect to a Foundry agent right here, and it'll take a second. But also it's a standalone asset that if I already have a harness that I'm using in other places, every knowledge base is an MCP server, so you can just connect to it without having to write any glue code in the middle. Now, a knowledge base like this has a bunch of parts. Some of them, for example, this storage content, you usually build indexes and you vectorize these things and whatnot, and if you want control over that, if you don't, you can just use it here. But if you do, let me just switch to Azure and show you the service behind that particular instance, where if I go to knowledge bases,
SPEAKER_01
this is the knowledge base we just created a second ago, and I can go peek inside. For example, I can go fish out the indexes that back this particular piece of content, and in that index, I can see what is the structure of the index. If I'm opinionated about, I don't know, maybe the quantization approach I want to use, or which indexing algorithm I want for my vectors, I can say all of that, and of course, I can actually go and explore the data and see what's inside, how chunks were organized and whatnot. So the goal of this is to, again, give you a highly productive environment when you don't need the sophistication, and when you need it, to make sure you have it,
SPEAKER_01
to get your job done. Go back to slides. And of course, the other aspect of this is, top of mind these days for all of us, token efficiency. And so we carefully evaluate this system to make sure that we give you the most information-dense answer that has the fewest tokens, so that your consumption of tokens has a high value when it comes to all retrieval tasks. The last category of knowledge I wanted to talk about is learned knowledge. Now, learned knowledge is the result of us doing the work we do as individuals and as organizations every day. And the idea that we can actually observe the processes and get better at them by reflecting and improving every step of it
SPEAKER_01
is something that has really changed now that we have agents doing the work and we can go tune the agents automatically. Satya wrote about this recently and reflected on the fact that people and agents can really compound in how they do the work and how they can create this learning loop that effectively captures what's unique about the company or the organization you're working on and inputs that to work to differentiate the work that you do. Now, in Foundry, we wanted to offer a materialized version of this that you can use today. So we built a component called the Agent Optimizer that effectively goes through this process and allows you to evaluate the baseline,
SPEAKER_01
generate candidates, Satya wrote about this recently. and reflected on the fact that people and agents can really compound in how they do the work and how they can create this learning loop that effectively captures what's unique about the company or the organization you're working on and inputs that to work to differentiate the work that you do. Now, in Foundry, we wanted to offer a materialized version of this that you can use today. So we built a component called the Agent Optimizer that effectively goes through this process and allows you to evaluate the baseline, generate candidates, and then evaluate the new candidates. If we have a strong result,
SPEAKER_01
then deploy that to production. Let me give you a quick flavor of what this looks like if we can switch back to the laptop. All right. So here I'm in VS Code. I have the Foundry Toolkit installed. And I have a simple agent. It doesn't matter how you write your agent, as long as you externalize configuration, your instructions, tool definitions, skills, and whatnot. So once you have one of those, it takes two key steps to do this. So first... Whoops. I can actually... So usually you have an evaluation already, but if you don't, you can actually say, eval generate, and what we'll do is we'll look at what we know about the agent, traces and instructions and whatnot,
SPEAKER_01
and we'll produce a task adherence-focused evaluation for you. In this case, I ran this a little bit earlier. So just to give you a flavor of what this looks like, you have a bunch of tasks, and then the questions and the criteria and whatnot. Once you have a dataset you can evaluate, then the next step is you can say optimize. And I could just run optimize on its own, and that will run, in this case. This ran for maybe 45 minutes or so, and you get an optimized version by effectively hill-climbing the metric that's established from the evaluation. So I ran this earlier, and so let me show you the output for this particular one, where you can see that
SPEAKER_01
we established the baseline first, and then we kept iterating on candidates using different combinations, using a JEPA-style loop, and looking for options that perform better, given the rubric that we have. And the interesting thing is that once you found one that is better, then you can simply just say optimize, apply, and what this does is, since you externalize the configuration, it allows you to swap one configuration for the other. If we look here, you can see that, for example, I have a baseline and the one we just applied, and just to pick on instructions, these are just the trivial instructions for this example agent. But if I look at the optimized one,
SPEAKER_01
then you can see a bunch of instructions that are not handwritten, but that emerged out of the hill-climbing process to make this particular agent better, given what we have in terms of instructions and skills and tools, but also based on reflecting on the actual traces from the agent