AI Engineer

Your agents lack context: Here's how to fix "You're absolutely right!" — Brandon Waselnuk, Unblocked

1821 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent quality at enterprise scale is constrained less by model intelligence than by a governed context engine that retrieves, resolves, and delivers the right organizational knowledge at the moment of work.
  • Why it matters: Without reliable context, agents generate superficially valid code while missing business logic, deployment procedures, ownership knowledge, and recent operational decisions—creating rework, review burden, token waste, and production risk.
  • Best use: Use this as an architecture and evaluation framework for an agent context/control plane, particularly for brownfield engineering environments with distributed knowledge across code, Slack, incidents, and internal documentation.

Executive Summary

Brandon Waselnuk argues that organizations have treated AI agents as if access to tools or documents were equivalent to understanding the company. It is not. Human engineers accumulate operational context over years through code reviews, incidents, meetings, and informal questions; a newly started agent session has none of that accumulated knowledge. As teams move from autocomplete to semi-autonomous and background agents, an initial context failure compounds into agent “doom loops,” wasted search tokens, rework, and a larger human review tax.

The proposed answer is a context engine: a system that ingests data from engineering and adjacent business systems, understands the requesting user or agent and its permissions, identifies relevant information, resolves contradictions between sources, and returns token-efficient context tailored to the workflow. The emphasis is not merely broad search. The engine must distinguish an outdated architecture record from a more recent Slack decision, respect access controls, and support both fast targeted retrieval and deeper research.

Waselnuk rejects two common stopping points. Curated Markdown or repository context files can improve an agent temporarily, but they are hard to distribute, become stale, and require an unrealistic central curator. MCP tool access is useful but insufficient because agents may fail to invoke a tool or stop after finding the first plausible answer, a behavior he calls satisfaction-of-search bias. His key distinction is that access to information is not the same as synthesized understanding.

The talk is partly an Unblocked product pitch, but it contains practical patterns and open-source starting points: mapping a GitHub-derived social/expert network, indexing and deduplicating repository rule files, and combining RAG with relational querying for questions that vector retrieval alone cannot answer. The most actionable lesson for Ken is to design context as a first-class, permissioned agent subsystem rather than a folder of prompts, documents, or disconnected MCP servers.

Key Takeaways

  • Claim: Context defects become more expensive as agent autonomy rises, so organizations should solve context quality before expanding agent scope. | Evidence: The speaker frames an adoption curve from tab completion to agents that work with less human oversight, then parallel agents and background agents. Poor context produces repeated correction loops, wasted search tokens, rework, and an increasing review tax. | Implication: Gate higher-autonomy workflows on demonstrated context coverage, retrieval quality, and operational correctness—not merely model benchmark performance. | Caveat: The talk offers a directional operating model rather than a quantified causal study of autonomy versus failure rates.
  • Claim: Static curated context repositories are a local optimum, not a durable enterprise context strategy. | Evidence: Waselnuk describes the common pattern of placing Markdown/project knowledge in a virtual or local filesystem for an agent to grep. He says this creates distribution burden, documentation rot, and a need for an “omnipotent” curator with the judgment to keep every repository’s context current. | Implication: Treat repository instructions as governed source material within a broader continuously refreshed system, and explicitly monitor staleness, duplication, and ownership. | Caveat: Static rules and project documentation can still be valuable inputs; the argument is that they should not be the whole context layer.
  • Claim: MCP-based access to source systems does not guarantee that an agent will find or correctly weigh the needed information. | Evidence: The speaker says tool-server and tool descriptions can cause agents not to call an MCP when they should. Even when retrieval occurs, satisfaction-of-search bias can make an agent stop after finding a seemingly valid architecture record while missing a newer Slack discussion that reverses the decision. | Implication: Do not evaluate an agent integration solely by whether every system is exposed through MCP; test whether the orchestration layer searches sufficiently, ranks sources by recency/authority, and detects conflicting guidance. | Caveat: MCP remains useful for authenticated access and scoped permissions; its limitation is that it is an access mechanism rather than a complete reasoning, retrieval-planning, or conflict-resolution layer.
  • Claim: A production-grade context engine requires six capabilities: unified context, targeted retrieval, conflict resolution, personalized relevance, token optimization, and permission enforcement. | Evidence: Waselnuk lists organization-wide source ingestion; quick link/document unfurling alongside deep research; adjudication when sources disagree; awareness of who the requester is and where they work; compact machine-to-machine outputs; and OAuth/SSO-based prevention of unauthorized disclosure. | Implication: Use these six dimensions as an architecture checklist and as evaluation criteria for any internal context platform, agent memory system, or knowledge-retrieval vendor. | Caveat: The presentation does not specify a formal conflict-resolution policy or evidence-scoring model, which is one of the hardest implementation choices.
  • Claim: Relevant context can materially reduce both execution time and token consumption while improving answer quality. | Evidence: In a reported same-prompt, same-model test, the task used about 21 million tokens without context versus 10.8 million with context, with roughly two hours of wall-clock savings. The speaker summarizes the result as approximately 50% fewer tokens, faster triage, and better answers. | Implication: Instrument baseline-versus-context runs on representative internal tasks, measuring not only token and latency savings but also correctness, incident risk, reviewer edits, and successful merge rates. | Caveat: This is a vendor-presented test with no disclosed task design, model configuration, or independent replication, so the exact savings should not be generalized as a benchmark.
  • Claim: RAG alone cannot answer many operational questions; agents need relational and deterministic querying over organizational data as well. | Evidence: The speaker contrasts semantic retrieval with a question such as “What are the open PRs that I worked on in the last week with authentication?” He says this requires queries over structured/relational data, and describes a workshop that teaches schema-less discovery followed by deterministic query generation. | Implication: Build a hybrid retrieval stack: semantic search for unstructured explanations and documents, plus structured graph/database access for entities, ownership, status, time ranges, and relationships. | Caveat: Schema-less query discovery may improve flexibility, but production use still needs validation, authorization checks, and protections against incorrect query generation.

Detailed Brief

Open-source implementation starting points

  • Claims: The speaker offers tools intended to help teams construct the organizational signals that a context engine needs rather than relying exclusively on a proprietary knowledge layer.; GitHub activity can be used to infer an expertise and collaboration graph that helps focus retrieval around likely owners, reviewers, and relevant teams.
  • Evidence: The “social comment network” tool deterministically analyzes GitHub to show who commits where, who reviews whom, and a distilled experts graph; optional OpenAI or Anthropic labeling can infer team groupings.; The “repo rules agent” searches for rule files throughout a codebase, reports their severities and issues, identifies duplicate/conflicting material, and makes the result available as an index for retrieval.; A relational-context-engine workshop provides a workbook with six stacked pull requests intended to demonstrate construction of the system from scratch.
  • Caveats: GitHub-derived expertise is an incomplete proxy for actual decision authority, particularly where ownership, architecture, security, and customer knowledge reside outside code review history.; Automated rule discovery can expose inconsistency, but resolving competing rules still needs explicit ownership and governance.
  • Implications: A practical first data model for an engineering context plane is a graph linking people, repositories, commits, reviews, rules, PRs, services, and incidents.; Treat context hygiene as an ongoing data-quality program: detect duplicate rules, identify stale guidance, and assign owners for adjudication.

Context engine uses beyond software generation

  • Claims: The speaker positions context infrastructure as a cross-functional capability rather than only a coding-assistant feature.; The same organizational retrieval layer can support customer-success issue resolution and sales conversations by surfacing relevant company knowledge while work is occurring.
  • Evidence: Waselnuk cites customer-success personnel resolving tickets as they arrive and salespeople querying the Unblocked context engine while in the field to accelerate deals.
  • Caveats: Cross-functional expansion raises the stakes for entitlement design because sensitive engineering, product, customer, and strategic data may coexist in the same context plane.
  • Implications: If building shared context infrastructure, design tenancy, auditability, source-level permissions, and audience-specific response policies from the first implementation rather than retrofitting them after adoption.

Notable Concepts & Terms

  • Context engineering: The discipline of supplying agents with the organization-specific information, relationships, constraints, and operational history needed to act correctly.
  • Context engine: The proposed system that ingests organizational sources, determines relevant context for a requester, resolves conflicts, applies permissions, and returns optimized outputs to people or agents.
  • Satisfaction-of-search bias: An agent behavior in which it stops looking after finding the first plausible answer, potentially missing newer or more authoritative information.
  • MCP plateau: The point at which exposing systems through Model Context Protocol tools stops delivering gains because tool availability does not ensure correct invocation, broad enough retrieval, or source adjudication.
  • Brownfield codebase: An established production codebase with legacy behavior, business logic, operational procedures, and revenue dependence—the environment where missing context is especially costly.
  • Token optimization: Returning only the context needed for the current task so agents avoid repeatedly searching or consuming large context windows with irrelevant material.
  • Relational context: Structured knowledge about entities and relationships—such as people, PRs, repositories, dates, ownership, and status—that must be queried rather than only semantically retrieved.

Operator Notes / Why Ken Should Care

  • Create a context-readiness scorecard for agent workflows using the six capabilities named in the talk: source coverage, targeted retrieval, conflict resolution, identity relevance, token efficiency, and authorization enforcement.
  • Run a controlled internal benchmark on several brownfield tasks: compare no-context, static repository instructions, MCP-only access, and hybrid semantic-plus-relational context; record tokens, elapsed time, reviewer changes, and operational correctness.
  • Require every retrieved recommendation used for consequential changes to carry source provenance, recency, authority/owner metadata, and an explicit conflict state when sources disagree.
  • Build or assess an organizational graph connecting engineers, teams, repositories, code reviews, services, incidents, and rule files; use it to route agent retrieval toward likely owners and applicable constraints.
  • Keep semantic RAG for explanatory material, but add validated structured-query tools for time-bounded, ownership-based, status-based, and relationship-based questions.
  • Treat cross-functional context access as a security architecture project: enforce source-level entitlements, log retrieval and response exposure, and test for leakage of restricted projects or customer information.

Source/Metadata

  • Title: Your agents lack context: Here's how to fix "You're absolutely right!" — Brandon Waselnuk, Unblocked
  • Transcript words: 2775
  • Duration seconds: 849
  • Timestamp note: No timestamps or chapter markers were present in the supplied transcript.
Full transcript 2768 words · 12 min read
0:00

Good afternoon.

0:12

I hope you're all having a lovely day here at AIE. We've had great weather, though the UV has been nine, so hopefully you put your sunscreen on and are being appropriate adults. I'm here to talk to you about context engineering, and I have the good fortune of following AJ from LinkedIn, because he talked a lot about the system that we actually design and sell to other solutions, and I'm going to give you a bunch of open-source tools. So if you watched that last talk just before me, you're going to get a bunch of tool chains so you can go mess around yourself, and I'll teach you a bunch of techniques today. The goal, of course, is to fix your absolutely right.

0:39

I think they've taken that out of the prompts now, so it just says you're right or other things, but I'm sure you've all been there. So I'm Brandon. I work at Unblocked. Yes, I have a coconut. We've been giving these away for fresh context, fresh coconuts. But the thing that I want to talk to you about is with these models, especially with Mytho class models— I think Fable 5 is coming back today, so they say you can watch my grain call recording and try to book this. We'll ignore it. But what I want you to do is think about the fact that with these tools, AI-generated code should feel like it was written by someone who's been on your team for years.

1:16

So to get in the right headspace, for years you have to consider that you have been the context engine.

1:21

How did you do that? You built context by going to work and asking questions, shipping PRs and getting them rejected, going to meetings, and all this slowly over time built up the engine that is your brain. You understand how it works here. You know how stuff gets shipped. You were on call that night when you took prod down and why that happened. The problem is that these agents have this exact same problem. Every time you create a new terminal session with an agent in it, it's very intelligent, but it doesn't have any context on how your company operates. So it needs to get that somehow. The problem is as you move these agents up in scale,

2:01

that cost compounds if you get it incorrect at the beginning. The leverage off context and content. We're just going to fix this because I think people want to take some photos. Perfect. That context issue will compound. So at the far left, we all remember the age-old time of two years ago where we had tab complete models that were pretty cool. What happened is it popped up and said, hey, do you want to tab this? And quickly in your head with your context, you either go, no, that's bad, or you went, oh, sweet. You hit tab. Nice. As we move along the agentic adoption curve, what happens is you are moving into more situations in which you have agents running

2:39

without a human in the loop, or at least you wish you didn't have to be in the loop. What they need is some way to be able to ask the questions they need when they hit walls in order to write code or solve or fix the issue and ultimately output code that's mergeable into your code base. So we have a lot of people here who actually work in brownfield code bases that have been around for a long time that run real revenue across them, not just greenfield fund projects. So that cost of bad context compounding at the beginning is cheap. If you think shift left, finding a defect or a bug, you want to find it as early as possible. It's the same with context.

3:19

Because as you move across, you get into doom loops. You usually ask your agent to do something. It's like, hey, I did it. And you're like, no, man. And then you correct and correct and correct. That's wasted search tokens. It's also wasted rework time. And that is not acceptable with the tokenomics we have coming. And then as you move into parallel agents, you start hitting a review tax. So these AI code reviewers we're trying to use. But again, key context is important there. So those code reviews are able to understand how the operations of the business are. So it knows the business logic and more.

3:52

And then finally, if your hope is to move all the way out of the loop, you're like background agents, get it done, make no mistakes. You really need to make sure that you have a context engine. So those agents can query it and get all the answers they need so they can keep operating in an effective way. There are some common approaches that don't work. They're basically local maxima. Two of the ones we see the most with our hundreds of enterprise clients and mid-market sized businesses is the curated context. If you've ever sat down and taken a virtual file system or a local file system, you put some markdown files in it.

4:22

And you're, here's all the context of this project. It's how it works. You then allow your agent to grep over that. And it gets a bunch of good data. And then it will perform better. The issue is first, now you have to distribute that. So maybe you throw it up in the GitHub and your team can grab it. But then the next is that repo is going to rot just like all the other docs you wrote down. And then who at your org is the omnipotent one who has the taste to curate this file? And then you have to do a repo for literally everyone in the org. So you start to hit these issues. The next is the MCP plateau. This one is pretty clear. We have MCPs. They're great.

5:02

You can give it to your agent. And now it can basically get information from another source system. The problem is, of course, based on how you write the server description, the tool descriptions, your agent may never call it, even though it should have. Or if it does, there's a known bias called the satisfaction of search bias. What that means is the agent, when it finds the first piece of information that it thinks is correct, it goes, oh, I have what I need, and it proceeds. In most organizations, there's a Slack conversation from last night that says you should be doing A instead of doing B. And the agent will never find it if it found some architecture record first.

5:34

So it doesn't actually consider all of the context. The problem here is access to information is not understanding. So to deliver understanding to a model, you have to do other techniques. What I'm basically trying to say is what your agent can't see is everything below the waterline. It can get code that compiles, but that code that compiles is taking down prod and you have a P0 at one in the morning. Because it missed the fact that you have a certain rollout procedure, you're supposed to turn off a feature flag, whatever it might be. So your team needs a context engine because what it should do is understand who you are and where you work in an organization.

6:12

So if I say to you, I want to get auth stood up, it knows where I work, it knows where my git commits are, it knows who reviews those commits, and it understands that my context, it can focus me and then use that as a trigger point to find the rest of the information. It resolves conflicts, as mentioned, an old architecture diagram and last night's Slack convo with the CTO. Which one is right? You need to use a bunch of techniques to determine that. It respects permissions and governance, of course. MCP allows us to use OAuth and other scopes and SSO, but if someone asks a question over here who's not supposed to know about secret project A,

6:47

you need to make sure that doesn't leak into the response. And then finally, deliver the right context at the right time to the model in a token optimized way. We have multiple surface areas because human engineers still talk to Unblocked all the time to get information they need in Slack or otherwise. But then you want token optimized responses if you're just speaking machine to machine in order to not waste a bunch of bold classes on your token spend. This is how an engine works. I'm going to be brief on this, but basically on the left hand side, you see all the data sources that are coming in.

7:21

For us, we focus on engineering teams and that's who uses us, as well as the technically light teams around it, like support, sales, and otherwise. You ingest all that data, you get real-time data from tools like your incident management toolchain. It comes into the engine, where that engine thinks. At the bottom, I'll expand on that slide in a moment. But basically, it uses these six key characteristics. And then on the right, you output the context to the exact workflow in the manner that it is needed. Those six key points, as mentioned, unified system context. You have to go across the whole thing.

7:55

At large orgs, companies like LinkedIn scale, Workday, General Motors, whatever, they need this type of data. They need to understand everything that's happening. And Tharik this morning, actually talking about Fable coming out potentially later today, he mentioned that you need to actually provide a map and then let Fable discover the territory. The way to help confine that is making sure that these models have access to all of the context because they will find your unknown unknowns. There are definitely things going on in your company that you're just unaware of, but would be really helpful for the task you're trying to do.

8:29

That will move faster, but the targeted retrieval, you should be able to, if you provide a link, quickly unfurl it, get that document back and move along. So two tasks: deep research, go long, that's fine. But you also need speed when speed is required. Conflict resolution, we already talked about that, but one thing says do A, one thing says do B, who is right. Personalized relevance, who am I, where do I work, what am I working on. That token optimization, making sure the response is good and effective and doesn't bloat the window. And then permission enforcement, of course. OAuth, you shouldn't see it, you shouldn't see it.

9:04

What we did with some tests is we actually ran the exact same prompt to the same model and one with context and one without. This is the wall clock time savings and then two hours, which is great. And then the tokens savings. So it was a sizable task, it took about 21 million tokens without and then 10.8 million tokens with it. This is the type of experience that you typically see when you're using a context engine, because the majority of those wasted search tokens where it has to grep at the beginning of every session to understand and discover things are no longer there when it's hydrated with context. And then as you move forward, you get these types of outcomes.

9:45

50% fewer tokens, faster triage, and the answer quality is actually better because it knew what was going on inside of the business. Now this next part, you'll probably want a photo. If you don't know, you can actually take a picture of a QR code and then later in photos tap on it and then load the link so you don't need to float here because I'm going to give you three QR codes. This first one is for the social comment network. I'll pop that up so you can take a photo. But this is an open source tool that we've got that actually using all deterministic programming goes over your GitHub and understands who works on your team. This is my real team.

10:14

We called Rasheen the machine because he ships constantly. But on the right, you can see who he commits, where he commits, who's reviewing his work. And then in those tabs, you can find a distilled experts graph. You get full coverage of what's going on in your business. And if you optionally add one of the API keys for either OpenAI or Anthropic, it'll determine what your teams are by doing some labeling for you. It's a really cool tool to understand where your team works and get that social network in there in order to focus a context engine if you're going to be building these tools yourself. The next is called the repo rules agent.

10:44

This is a sample from our real code base. I'm going to pop that up anyway so you don't need to talk to the thing. But in short, what it does is discover all the places your team has written rules files, checks them all, and then tells you what severities you've given, what other things you've given. It tells you, hey, it's good to meet you all. Basically, it will find all the rules that are inside of your repo and then tell you if you have duplicate issues or other problems, and then you can grep over it as an index. So that index can be called and you can dedupe, and it'll help improve your retrieval of context.

11:13

And then finally, on Monday, we delivered this workshop, which was going beyond RAG, and taught how to build a relational context engine from scratch. So if you scan that, you'll get the full workbook. It has six PRs stacked that teach you how to walk through doing this. But in short, RAG is an incredible technique, and you want that. But the other half of the problem is what people actually ask is, what are the open PRs that I worked on in the last week with authentication? RAG cannot answer that question alone. You need queries. So this shows you how to do a schema-less lookup that allows the agent to discover a schema and then write queries against it.

11:58

And then you can see how to do it deterministically in order to get that type of relational data out. Very useful technique. Use cases of a context engine, of course, go beyond code generation. This is where we live a lot. A lot of our customers spend their time. But it's amazing to see what happens when a bunch of other people around the business start picking up these tools. Customer success people solving tickets right at the time that it comes in from a customer. We've got salespeople closing deals earlier in their quarter because they're able to query the unblocked context engine on the fly while in the field. And so many more.

12:31

What you can also do is if you saw that curve chart earlier where I talked about the levels, we've built a fun little tool where basically an LLM will quiz you and ask you about what's going on. And then it will map you to exactly where you are and then tell you some techniques about how to level up through that if you are looking to compound your capabilities and ship with AI tools at scale. It's readiness.getunblocked.com. The gap is not intelligence any longer. It's context. We will continue to get incredible models like Mythos as it's been grown by Anthropic. And I'm sure Sol, once I'm allowed to see it, I will get it. Happy Canada Day, by the way.

13:14

But what's happening is it's about the context you surround these models with in order for them to be effective and token efficient inside of your organization. So I have a question slide, but I'm not sure I'm allowed. Nope. So what you'll do is come meet me at booth P16.

13:36

You can look for the coconut. It would be great to hang out with all of you and get into details here if you need it. Thank you for your time. Thank you. Thank you. Thank you for your time.

14:02

Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note