AI Engineer

WTF Is the Context Layer? The Missing Infrastructure for Production Agents — Prukalpa Sankar

1888 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Production agents need a shared, portable context layer that converts company knowledge, expertise, norms, and operational history into governed, machine-usable context rather than embedding separate context inside every agent.
  • Why it matters: As agent fleets grow, fragmented memory, inconsistent definitions, stale skills, hidden dependencies, and weak governance become larger constraints than model intelligence and can cause autonomous systems to make contradictory or unsafe decisions.
  • Best use: Use the talk as an architecture and governance framework for designing OpenClaw or enterprise agent systems, especially the shared context plane, skill lifecycle, trace-based learning loop, and separation of context from agent runtimes.

Executive Summary

Sankar argues that model intelligence is improving much faster than real-world AI usefulness because models lack the situated knowledge required to operate inside a specific business. That context includes formal definitions and data relationships, accumulated diagnostic expertise, and organizational norms about scope, communication, and decision rights.

Athlan initially built narrow agents for individual customer-experience jobs. Agent construction became easy, but supplying reliable context remained slow and fragile. The agents developed isolated memories, failed to absorb changes made elsewhere in the organization, were difficult to debug, and trapped context inside successive platforms such as Relevance, Google ADK, Glean, Claude Code, and Codex.

Athlan's second approach was a shared company brain between business systems and multiple general-purpose or external agents. Domain experts contributed reusable skills, while the shared layer also held data relationships, metric definitions, entities, organizational structure, and playbooks. Its marketing implementation reportedly grew to approximately 300 skills and 40 agents in six months.

That scale exposed the next problem: context must be operated like code. Skills need owners, versioning, dependencies, approvals, testing, security controls, and portable retrieval interfaces. Sankar's proposed context layer continuously mines business systems, serves context through mechanisms such as MCP, SQL, vector retrieval, and hybrid assembly, and turns agent traces into proposed improvements reviewed by maintainers. Her strategic conclusion is that company context is proprietary operating IP and will differentiate agents when competitors have access to the same underlying models.

Key Takeaways

  • Claim: The main bottleneck to useful enterprise AI is increasingly business context rather than raw model intelligence. | Evidence: Sankar contrasts rapid benchmark improvement with claims that only one in five AI use cases reaches production, 56% of CEOs see zero financial benefit from AI, and IQ explains only 10% of job-performance variance. She frames performance as intelligence plus context learned on the job. | Implication: Ken should evaluate agent platforms on how well they acquire, govern, update, and apply company-specific context—not primarily on which frontier model they expose. | Caveat: The talk does not cite the sources or methodologies behind these headline statistics, so they should be treated as framing evidence rather than independently validated benchmarks.
  • Claim: Useful context has at least three distinct components: business knowledge, accumulated expertise, and organizational norms. | Evidence: In the fictional drive-through analysis, the agent must know the metric definition, weekly cutoff, time zone, and requester; apply diagnostic knowledge about seasonality and a recent product launch; and tailor the scope and explanation to the audience. | Implication: A context system limited to document retrieval or vector search will be incomplete; Ken's architecture needs semantic definitions, procedural playbooks, historical lessons, personas, and decision policies.
  • Claim: Embedding context separately inside specialist agents creates stale behavior, inconsistent truth, weak observability, and platform lock-in. | Evidence: Athlan's website SDR continued using old positioning after marketing changed it. Separate agents learned through separate memory systems, and failures were difficult to attribute to the model, agent logic, or context. Context also became trapped as the team moved from Relevance to Google ADK, Glean, Claude Code, and a roughly 50-50 Claude/Codex setup. | Implication: OpenClaw should keep authoritative context and learning history outside individual agents so models and orchestration frameworks can be changed without rebuilding institutional memory.
  • Claim: A shared context layer can act as a company brain serving many agents while allowing domain experts to maintain reusable skills. | Evidence: Athlan's marketing team placed a common repository between its data, community, advertising, and analytics systems and agents including Claude Code, Codex, an internally deployed Claude, Qualified, and Artisan. Its best SEO and competitive-intelligence practitioners authored corresponding skills; the system expanded to about 300 skills and 40 agents over six months. | Implication: Ken can use this as a reference pattern for a model-neutral control plane: business systems feed a governed context repository, domain owners maintain reusable skills, and multiple agents retrieve from the same source. | Caveat: The talk reports the scale of the deployment but provides no controlled accuracy, productivity, cost, or financial outcome measurements.
  • Claim: Context and skills require software-style lifecycle management because changes can break downstream behavior. | Evidence: Athlan's competitive-intelligence skill feeds category positioning, which then feeds sales battle cards; improvements to upstream skills could break downstream outputs or cause drift. The team also encountered unclear ownership, hardcoded secrets in .env files, and public skill repositories being downloaded. | Implication: Ken should treat skills as governed production artifacts with dependency graphs, version control, approvers, maintainers, contributors, security posture, regression tests, and rollback capability.
  • Claim: Agent traces can create a compounding learning loop if they are converted into reviewable context updates rather than left as passive logs. | Evidence: Sankar describes a specialized harness that reads traces, reverse-constructs potential improvements, and returns them to maintainers for approve-or-reject decisions before the shared context evolves. | Implication: Tracing in Ken's systems should support context improvement proposals and human approval workflows, not merely debugging and observability. | Caveat: The talk proposes the operating loop but does not specify evaluation thresholds, trace schemas, conflict resolution, or protections against agents learning from erroneous outputs.
  • Claim: Company context is proprietary IP and becomes a strategic differentiator when competitors can access the same models. | Evidence: Sankar asks what would distinguish an American Express support agent from an Amazon support agent if both use similar intelligence, answering that the difference is each company's culture, norms, expertise, and way of operating. She warns that hardcoded context could also reproduce the familiar conflict in which sales and finance report different revenue numbers across autonomous systems. | Implication: Ken should regard context assets as durable intellectual property requiring security, provenance, and investment—not as disposable prompts attached to individual applications.

Detailed Brief

Selecting work for initial agent deployment

  • Claims: Athlan began with a jobs-to-be-done map of the customer-experience function and assigned different scaling potential to each task.; The initial operating assumption was that AI could scale documentation and meeting preparation sooner than relationship management.; Narrow agents can create early value, but their context burden increases rapidly once they interact with other functions or shared business facts.
  • Evidence: The team created named specialist agents including Hermione, a health-intelligence lead, and Moneypenny, a financial-risk analyst.; By the middle of the prior year, building an agent reportedly took about five minutes, while supplying enough business context for trustworthy accuracy took much longer.
  • Caveats: No scoring rubric for the scaling factor or results from individual task categories is provided.; The talk does not separate failures caused by task selection from those caused by context architecture.
  • Implications: The strongest first pilots are bounded, information-heavy jobs with explicit outputs and limited relationship or authority requirements.; Fast agent creation should not be mistaken for production readiness; context preparation and validation are the larger delivery workload.

Context assembly and initialization

  • Claims: The proposed layer must support both global company context and more localized context appropriate to a function, task, or agent.; It should continuously mine connected operational systems rather than depend entirely on manually authored prompts and documents.; The resulting context must be available through multiple retrieval modes because structured facts, semantic knowledge, and procedural skills require different access patterns.
  • Evidence: Sankar names MCP, SQL, vector retrieval, and hybrid assembly as possible interfaces.; She suggests connecting systems such as Salesforce and HubSpot to the data warehouse and application layer, then reverse-constructing the relationships across those hops to bootstrap a first company brain.; The proposed repository includes a data graph, metric semantics such as ARR and qualified-lead definitions, organizational structure, entities, and a skill library.
  • Caveats: Reverse-constructing relationships from existing systems may also import inconsistent definitions, obsolete processes, and access-control mistakes.; The presentation offers an architectural direction rather than a concrete schema, protocol, or reference implementation.
  • Implications: A practical context plane will likely combine a semantic or graph layer with structured query access, retrieval, skill packaging, and policy-aware assembly rather than rely on one database technology.; Existing system lineage can accelerate initialization, but authoritative owners still need to resolve contradictions before agents receive autonomous authority.

Notable Concepts & Terms

  • Context layer: A shared system that turns company knowledge, expertise, norms, and history into governed, machine-usable context for multiple AI systems.
  • Company brain: The common repository of definitions, relationships, entities, playbooks, skills, and organizational memory that agents query instead of maintaining isolated truth.
  • Context engineering: The work required to supply an agent with enough relevant business information and operating guidance to produce trustworthy results; Athlan found this much harder than creating the agent itself.
  • Skill lifecycle: The versioning, ownership, testing, approvals, dependency management, security, and updating required for reusable agent skills.
  • Context portability: Keeping institutional knowledge usable across model and agent frameworks so migrations do not strand context inside a vendor-specific runtime.
  • Compounding learning loop: A process in which agent interactions and traces generate proposed improvements that maintainers review and fold back into shared context.
  • Hybrid assembly: Constructing an agent's working context from multiple retrieval methods, potentially combining structured queries, graphs, vector search, skills, and policies.
  • Context as IP: The thesis that encoded culture, expertise, norms, and operating practices—not commodity access to frontier models—will be a durable source of competitive differentiation.

Operator Notes / Why Ken Should Care

  • Create a context-asset inventory for OpenClaw covering metric definitions, entity relationships, playbooks, decision rights, historical lessons, personas, and source systems.
  • Select one cross-functional workflow and map every context dependency before adding more agents; prioritize a workflow where stale or conflicting information is already visible.
  • Establish a skill registry with an owner, maintainer, version, dependencies, evaluation suite, access policy, and rollback path for every production skill.
  • Separate context storage and learning history from model-specific runtimes so Claude, Codex, or future agent frameworks can be swapped without losing institutional memory.
  • Add dependency-aware regression testing so an update to competitive intelligence, positioning, pricing, or metric semantics automatically tests downstream skills and agents.
  • Prohibit secrets in skill files and public repositories; apply least-privilege retrieval and provenance controls before granting agents write or execution authority.
  • Instrument traces to generate proposed context changes, but require human approval and evaluation before those changes enter the authoritative layer.
  • Define a conflict-resolution process for contradictory facts across CRM, finance, warehouse, and application systems before using automated lineage reconstruction to bootstrap the context graph.

Source/Metadata

  • Title: WTF Is the Context Layer? The Missing Infrastructure for Production Agents — Prukalpa Sankar
  • Transcript words: 5566
  • Duration seconds: 1253
  • Timestamp note: No timestamps or chapters were present. The supplied transcript also contains substantial duplicated passages and a repeated closing section.
Full transcript 3196 words · 25 min read
0:00

[SPEAKER_00] Hi everyone, my name is Prakalpa.

0:12

SPEAKER_00

I'm the founder of Athlan. And today I'm going to talk about this thing where context is having its moment. And so my goal today is to talk about what is the context layer. Just before I start, and I promise this is the last time, I don't know if the clicker's working. It's working. Athlan, it's working, yeah, thank you. The problem we solve is we say AI doesn't know your business, we fix that. We work with an incredible group of companies around the world, ranging from GitLab and Zoom and Discord and Affirm to large enterprises like MasterCard and General Motors.

0:36

SPEAKER_00

And about a year ago, my co-founder and I went on stage and we said, at the dawn of the internet era, Bill Gates had written this very famous blog post and it said, "Content is king." And as we are at the dawn of the agentic era, context will be king. Since then, it feels like 2026 is the year of context. Context graphs, anyone? You know, every two days you see some version of context popping up. And so what is going on?

1:04

SPEAKER_00

I believe the answer to this is in this reality distortion field that we live in. I live here in the Bay Area. Every day or two I have conversations with people which go like, how far are we from AGI? And we have a debate, and we're like, well, one year, three years, so on. There is no doubt that the models are getting exponentially smarter by the day. Two years ago, they couldn't pass the bar. Today, if they were to take the bar, they're in the top 1% of test scores.

1:10

SPEAKER_00

On the other hand, they're not exponentially more useful by any benchmark. One out of five AI use cases actually make it to production. You know, 56% of CEOs say that there's zero financial benefit from AI today. So what's going on? I believe hidden in plain sight is actually how performance is measured in the human world. Cognitive intelligence doesn't really determine real world effectiveness. In fact, only 10% of job performance variance is explained by IQ.

1:20

SPEAKER_00

Just think about it. Would you say your smartest teammate who scored the highest on the SATs is also your best teammate? Or would you say no, it's the person who works the most, takes the most feedback, and learns the fastest? In the real world, we care about performance, and performance is outcomes that you deliver in the real world. And performance is a function of two things. It's a function of intelligence, which is cognitive horsepower. That's what the model benchmarks measure every day. But it's also a function of context. This is what they say in the human world as learning on the job, right? Knowledge and skills and expertise that you learn over time.

1:30

SPEAKER_00

And in the last decade, we have compounded on one of those parameters. Intelligence has thousand X'd in the last decade. Just in the last six months, we have two X'd on that axis. On the other hand, context, the situated knowledge of your business, that's barely moved. We've moved some data to the cloud, but that's about it. It's otherwise logged in dashboards and Slack threads and the head of that analyst who might be leaving next week. And so the question ahead of us, and I really believe this is the next frontier, is how do we help AI build context about our business?

1:39

SPEAKER_00

And every time I'm faced with a question about how do we help AI do this, I always go back and understand how did we help humans do this? I'm going to take you into the life of an exemplar employee, Maya. Let's say she's a data analyst at Met Context Burgers, because I thought I was going to be creative, and I'm not very creative. And let's say she's that analyst that everybody pings in your company, right? She's the person that everybody sends a message to every morning when they're trying to solve a problem. So let's say this morning, there's a franchisee owner who sends her a message and says, why is my drive-through time up this week? Why is this metric up this week?

1:50

SPEAKER_00

Sounds like a really simple question. But it's actually a really complicated question to ask. Just to answer this one very simple question, Maya first needs to know what is drive-through time? And who's asking? Is it finance or is it my ops team? And it might mean different things. But not just that. What does this week mean? Is the cutoff period Monday to Sunday? Is it Pacific time? Is it Eastern time? That's knowledge. That's facts. That's the map of the business.

2:02

SPEAKER_00

But not just that, there's expertise, right? There's a diagnostic playbook. What does a great analyst do? They know that quarter three is a seasonal quarter because of weather patterns. And they know to go check if the reason there's a spike is because of seasonality. They also know that the company launched a product just that previous quarter. And so they know to check if that's why the root cause analysis failed. This is expertise and skills that people pick up over time as they learn on the job. And then there's norms, right? There's persona, scoping, who's asking the question, how do I answer this question? And Maya, she's one of those cool people. She nails it.

2:13

SPEAKER_00

And they know to go check if the reason there's a spike is because of seasonality. They also know that the company launched a product just that previous quarter. And so they know to check if that's why the root cause analysis failed. This is expertise and skills that people pick up over time as they learn on the job. And then there's norms, right? There's persona, scoping, who's asking the question, how do I answer this question? And Maya, she's one of those cool people. She nails it. She sends an answer, not just with the answer, but with the why and the root cause, and she finds the reason for it. How did Maya learn to do this? She just joined the company a year ago.

2:41

SPEAKER_00

First, Maya has joined and she got some training, as all of us do. But that's not where any of us learn, right? In our companies. How do we learn? We learn because we shadow the best teammate. And then you see why they're doing something. And then you learn from that. And then you make a mistake. Who here has learned more from a mistake than anything else? Right? You make a mistake and then you learn.

3:11

SPEAKER_00

Your manager gives you feedback and you learn not to do that again. You deal with an edge case. And then you learn from that. That's how all of us humans learn at work. And so then the question is how do you help build the agentic Maya? And now I want to walk you through our experiments and learnings as we've built this at Athlon. Era one, and this was roughly about 18 months ago now. We started on the track of bootstrapping agents. And the way we went about it was, and we started this with our customer experience team, and we did this jobs to be done analysis map. Right?

3:35

SPEAKER_00

And so we said, hey, if you are someone on our customer experience team, what are all the things that you do on a day-to-day basis? And then we made some hypothesis. We said, for example, one part of the job is documentation and meeting prep. We said, well, AI could probably do that job pretty well. And so we built a scaling factor. So on the other hand, relationship management is something that our customer experience team does. And we said, that doesn't sound like something AI is going to be able to do anytime soon. And so we built a scaling factor. And then we basically started bootstrapping these individual agents that were built for that specific topic.

4:06

SPEAKER_00

Our team got creative, so we had Hermione, who was our health intelligence lead, and then we had Moneypenny, who was our financial risk analyst. And we just made that particular agent really good at doing that one thing. And that worked for some time. But then we realized there were some challenges with this approach. The first, context engineering. We got to the point by middle of last year where building an agent was really easy, took about five minutes. But giving it the business context that it took to actually get it to be accurate took forever. Quality of the agent often dependent on the quality of context engineering.

4:29

SPEAKER_00

And that led to a lot of weird lost trust cases with our stakeholders. Then, as we started taking this into production, we started seeing that these agents basically were on their own island. Now imagine, for example, if you're in a human team, and your marketing changes positioning and then they come to the town hall and they tell you that they changed positioning. And so then the SDR on your team, or your sales development rep, they know that they should use that new positioning. This is the infrastructure that we've built for humans inside our organizations. Agents didn't have them. So our marketing team had these agents, and they started making changes to that.

4:48

SPEAKER_00

And then our SDR agent on our website was still pitching the old version. We had no idea how any of these things were even connected. So we didn't even know how to run this as a team of agents. When an agent gets something wrong, this is hard. It was really hard to trace back what happened. Was it the model? Was it the agent? Was it the context? How do we even go back and fix this? And over time, we started dealing with context problem. We had the hard part about this was agents all had their own memory systems to a certain extent. So they were learning. They were all learning separately, and they were learning differently.

5:23

SPEAKER_00

It became very, very difficult, very quickly to say, okay, what does the single version of truth here look like? And then over time, we actually went through in the last 12 months, we've gone through cycles at the agent declare. About 12 months ago, we were using one of these no code type builders called Relevance. We went from there into Google ADK, then we tried Glean. Start of this year, we moved to Claude Code. Now we are about 50-50 Claude and Codex. And every single time as these changes happened, our contacts got trapped in each of these individual systems.

5:31

SPEAKER_00

So, started this year as general purpose agents started to become a thing, we said, what if there was a different approach with general purpose agents? Again, going back to the human world, well, Maya, she's not an individual star. She's part of a team, right? And you talk about these dream teams like Maya and someone who runs customer support. Start of this year, we moved to Claude Code. Now we are about 50-50 Claude and Codex. And every single time as these changes happened, our contacts got trapped in each of these individual systems.

5:45

SPEAKER_00

So we started this year as general purpose agents started to become a thing. We said, what if there was a different approach with general purpose agents? Again, going back to the human world, well, Maya, she's not an individual star. She's part of a team, right? And you talk about these dream teams like Maya and someone who runs customer support and someone who launches ads. These people work really well together. And often these dream teams are built on shared context. They have a shared language. They have a shared picture of what's true today. They have shared playbooks. They have shared norms, who's allowed to make what decision.

5:56

SPEAKER_00

And then they learn together. I think this is the most important part of it. They have compounding learning loops of what good looks like. And they have shared memory that, oh, we launched this thing last quarter, and it was terrible, and we're not going to make that mistake again, right? And so we said, is there a way to bring that into the way we think about AI in our companies? And so the mental model we started working on was we said, okay, we have these teams of humans, and they're across the board, and can these people essentially start building domain skills?

6:07

SPEAKER_00

So each of them is responsible for a certain set of skills. All of this goes into this common one place, which is this one company brain of sorts, right? I like to think of this as the context layer. And then this has a bunch of retrieval mechanisms, which then talks to the general purpose agent across the ecosystem. So then we started an experiment. This is some version of what our marketing team ended up building. So you'll see on the left, those are all the systems that our marketing team uses. So data systems, our social and community platforms, our ad platforms, our analytics platforms.

6:18

SPEAKER_00

And then you'll see this agent block. We build this very specifically for having openness. So we had Claude Code and Codex. We also had our own Claude that we deployed, which has essentially stocks in our Slack channels. And then we used some external products, like Qualified and Artisan. In the middle is this context layer that our team started building. So think of it as our best SEO person was building the SEO skill. Our best competitive intel person was building the best competitive intel skill. And that became this common repo that we were building into and pulling out from. This became our living brain.

6:30

SPEAKER_00

Over time we realized there were some things that we needed in this brain, right? We realized we needed a data graph. Like if, for example, our autonomous ads agent, we realized it needs to do analysis on a daily basis. So which table should I go pull from? We needed a library of skills. We also needed some other things: semantics, metrics, what is ARR, how do you measure that, what is a qualified lead in our company, and org structure, entities, things like that. Over the last six months we ended up creating about 300 skills and 40 agents in this team, which has been incredible. But then, with this approach too, we realized that there were some challenges.

6:42

SPEAKER_00

We realized that context needs to be managed like code. So some challenges. Let's pick skills. Dependency management became really complicated. So for example, we have this competitive intelligence skill. And it learns from the market on what's changing in the market. And it improves. It feeds our category positioning skill, which then feeds our sales battle card skill. Now each of these skills is learning and evolving. But every time they learn and evolve, it breaks something downstream. And these skills very quickly start getting outdated and start drifting. Who owns skill quality became another thing. Like who eventually owns the quality of this.

6:55

SPEAKER_00

Security and governance was a nightmare. We had secrets hardcoded in .env files. People were downloading these public skill repos. The whole thing was a nightmare. And then I talked about context portability across all these multi-agent systems. I started this talk by saying what is a context layer. These are the problems that a context layer is meant to solve. The question I'd like to ask is what does the GitHub for context look like? Few thoughts. Company context needs lifecycle management, collaboration, and versioning. Just like code does. You know, there's questions like what's local context? What's global context? How do I keep this updated? And so on.

7:06

SPEAKER_00

Some thoughts in this. Can skills have a profile just like code does? Can that have a self learning loop that's baked into it? What does quality management look like? Can you have security and posture management associated with that? That's really the first step. I see this as having something that has built in versioning and quality and dependency management. So you should be able to say hey, this thing impacts all these other things. This is the approver. This is the maintainer. These are the contributors. How do you build human plus AI workspaces that these skills are managed via? Second thing, every AI interaction creates more context and harnessing this is gold.

7:23

SPEAKER_00

There's been a lot of talks about self improving loops. We have found that with traces, deploying a specific harness that actually is specialized in being able to go and reverse construct from that. So think of it as AI that's reading through all your traces.

7:33

SPEAKER_00

So you should be able to say hey, this thing impacts all these other things. This is the approver. This is the maintainer. These are the contributors. How do you build human plus AI workspaces that these skills are managed via? Second thing, every AI interaction creates more context and harnessing this is gold. There's been a lot of talks about self improving loops. We have found that with traces, deploying a specific harness that actually is specialized in being able to go and reverse construct from that. So think of it as AI that's reading through all your traces and almost brings it back to your maintainer loop and says approve, reject, approve, reject, improve this over time. That's the compounding learning loop. And the third, often a lot of people ask me this question, which is how do I start? Because my business is really disparate and I have all these 60 systems and how do I even start? One of the biggest learnings we've had is context is hidden in business systems and across this context quality can really compound. So for example, if you're able to connect your sales force and your HubSpot to your data warehouse, to your application layer, and then you're able to reverse construct how these things are actually connected one to another, context today gets lost in every one of those hops. But if you can reverse construct that and then deploy AI on top of it, we've seen incredible accuracy in being able to reverse construct the first version of your company brain. So I'll end with this. The way I think about a context layer is it's a system that turns knowledge and expertise and norms that we talked about that Maya knows into a machine usable context for AI systems. At a very high level, the way I like to think of it is it looks like this. It continually is mining context from your business systems. It's feeding this into that one company brain. It's harnessing this in skills and context development life cycles as your teams go and deploy these agents. And then it has a bunch of ways you can retrieve it. So MCP, SQL, vector retrieval, hybrid assembly, all these different ways that you retrieve it and pull back from traces and build this compounding learning loop. Today, we're largely building agents by hard coding context. The scale of this problem, I truly believe, is underhyped because with scale, this can become really unsustainable and a little dangerous. All of us know this old joke, which is if you ask sales and finance the revenue number, you're going to get two different numbers. We're fast approaching a moment of starting to deploy autonomous systems where the same thing is starting to happen. So I'll end with one last thing. I started this presentation by saying context is king. I'd like to end it by saying context is also IP. Something I think a lot about is in a world where you and your competitor have access to the same models and the same intelligence, what differentiates a company? What differentiates a customer support agent at American Express versus Amazon? That's how you do business. That's what makes your company special. Context is how we take and encode our culture and our norms into something that we will be proud of as we build autonomous Frontier forms. And that's all I had. You can find me at Trucalpa on Twitter or write to me. We are actively working with folks on the Frontier ongoing and shipping and building company brands. So if you'd like to talk to us, feel free to reach out. Thank you.

7:35

SPEAKER_00

and we started this with our customer experience team, and we did this jobs to be done analysis map. Right? And so we said, hey, if you are someone on our customer experience team, what are all the things that you do on a day-to-day basis? And then we made some hypothesis. We said, you know, for example, one part of the job is documentation and meeting prep. We said, well, AI could probably do that job pretty well. And so we built a scaling factor. So on the other hand, relationship management is something that our customer experience team does. And we said, hmm, that doesn't sound like something AI is going to be able to do anytime soon. And so we built a scaling factor.

8:13

SPEAKER_00

And then we basically started bootstrapping these individual agents that were built for that specific topic. Our team got creative, so we had Hermione, who was our health intelligence lead, and then we had Moneypenny, who was our financial risk analyst. And we just made that particular agent really good at doing that one thing. And that worked for some time. But then we realized there were some challenges with this approach. The first, context engineering. We got to the point by middle of last year where building an agent was really easy, took like five minutes. But giving it the business context that it took to actually get it to be accurate took forever.

9:00

SPEAKER_00

Quality of the agent often dependent on the quality of context engineering. And that led to a lot of weird lost trust cases with our stakeholders. Then, as we started taking this into production, we started seeing that these agents basically were kind of like living on their own island. Now imagine, for example, if you're in a human team, and your marketing changes positioning on your, you know, and then they come to the town hall, and they tell you that they changed positioning. And so then, you know, the SDR on your team, or your sales development rep, they know that they should use that new positioning. This is like the infrastructure that we've built for humans

9:38

SPEAKER_00

inside our organizations. Agents didn't have them. So our marketing team had these agents, and they started making changes to that. And then our SDR agent on our website was still pitching the old version. We had no idea how any of these things were even connected. So we didn't even know how to like run this as a team of agents. When an agent gets something wrong, this is hard. It was really hard to like trace back what happened. Was it the model? Was it the agent? Was it the context? Like, how do we even go back and fix this? And over time, we started dealing with context problem. We had, the hard part about this was agents

10:19

SPEAKER_00

all had their own memory systems to a certain extent. So they were learning. They were all learning separately, and they were learning differently. It became very, very difficult, very quickly to say, okay, what does the single version of truth here look like? And then over time, we actually went through, in the last 12 months, we've gone through cycles of at the agent declare. About 12 months ago, we were using one of these no code type builders called relevance. We went from there into Google ADK, then we tried Glean. Start of this year, we moved to Claude code. Now we are kinda like 50-50 Claude and Codex. And every single time as these changes happened,

11:01

SPEAKER_00

our contacts got trapped in each of these individual systems. So, started this year as general purpose agents started to become a thing, we said, what if there was a different approach with general purpose agents? Again, going back to the human world, well, Maya, she's not an individual star. She's part of a team, right? And you talk about these dream teams like Maya and someone who runs customer support and someone who launches ads. These people work really well together. And often these dream teams are built on shared context. They have a shared language. They have a shared picture of what's true today. They have shared playbooks.

11:42

SPEAKER_00

They have shared norms, who's allowed to make what decision. And then they learn together. I think this is the most important part of it. They have compounding learning loops of what good looks like. And they have shared memory that, you know, oh, we launched this thing last quarter, and it like was terrible, and we're not going to make that mistake again, right? And so we said, is there a way to bring that into the way we think about AI in our companies? And so the mental model we started working on was we said, okay, we have these teams of humans, and they're across the board, and can these people essentially start building domain skills?

12:19

SPEAKER_00

So each of them is responsible for a certain set of skills. All of this goes into this common one place, which is this one company brain of sorts, right? I like to think of this as the context layer. And then this has a bunch of retrieval mechanisms, which then talks to the general purpose agent across the ecosystem. So then we started an experiment. This is some version of what our marketing team ended up building. So you'll see on the left, those are all the systems that our marketing team uses. So data systems, our social and community platforms, our ad platforms, our analytics platforms. And then you'll see this agent block.

13:01

SPEAKER_00

We build this very specifically for having openness. So we had Claude Code and Covark. We also had our own Claude that we deployed, which has essentially stocks in our Slack channels. And then we used some external products, like Qualified and Artisan. In the middle is kind of this context layer that our team started building. So think of it as our best SEO person was building the SEO skill. Our best competitive intel person was building the best competitive intel skill. And that kind of became this common repo that we were building into and pulling out from. This sort of became our living brain. Over time we realized there were some things

13:48

SPEAKER_00

that we needed in this brain, right? We realized we needed a data graph. Like if, for example, our autonomous ads agent, we realized it needs to do analysis on a daily basis. So like which table should I go pull from? We needed a library of skills. We also needed some other things, semantics, metrics, what is ARR, how do you measure that, what is a qualified lead in our company, and org structure, entities, things like that. Over the last six months we ended up creating about 300 skills and 40 agents in this team, which has been incredible. But then, with this approach too, we realized that there were some challenges. We realized that context kind of needs to be managed

14:33

SPEAKER_00

like code. So some challenges. Let's pick skills. Dependency management became really complicated. So for example, we have this competitive intelligence skill. And it learns from the market on what's changing in the market. And it improves. It feeds our category positioning skill, which then feeds our sales battle card skill. Now each of these skills is learning and evolving. But every time they learn and evolve, it breaks something downstream. And these skills very quickly start getting outdated and start drifting. Who owns skill quality became another thing. Like who eventually owns the quality of this. Security and governance was a nightmare.

15:16

SPEAKER_00

We had secrets hardcoded in .env files. People were downloading these public skill repos. The whole thing was like a nightmare. And then I talked about context portability across all these multi-agent systems. I started this talk by saying WTF is a context layer. These are the problems that a context layer is meant to solve. The question I'd like to ask is what does the GitHub for context look like? Few thoughts. Company context needs lifecycle management, collaboration, and versioning. Just like code does. You know, there's questions like what's local context? What's global context? How do I keep this updated? So on. Some thoughts in this.

16:03

SPEAKER_00

Can skills have a profile just like code does? What's, can that have a self learning loop that's baked into it? What does quality management look like? Can you have security and posture management associated with that? That's really like the first step. I see this as like having something that has built in versioning and quality and dependency management. So you should be able to say hey, this thing impacts all these other things. This is the approver. This is the maintainer. These are the contributors. How do you build like kind of human plus AI workspaces that these skills are managed via? Second thing, every AI interaction creates more context

16:46

SPEAKER_00

and harnessing this is gold. There's been, I know, a lot of talks about self improving loops. We have found that with traces, deploying a specific harness that actually is specialized in being able to go and reverse construct from that. So think of it as AI that's reading through all your traces and almost brings it back to your maintainer loop and says approve, reject, approve, reject, improve this over time. That's the compounding learning loop. And the third, often a lot of people ask me this question, which is like how do I start? Because my business is really disparate and I have all these like 60 systems and how do I even start?

17:26

SPEAKER_00

One of the biggest learnings we've had is context is hidden in business systems and across this context quality can really compound. So for example, if you're able to connect your sales force and your HubSpot to your data warehouse, to your application layer, and then you're able to reverse construct how these things are actually connected one to another, context today gets lost in every one of those hops. But if you can reverse construct that and then deploy AI on top of it, we've seen incredible accuracy in being able to reverse construct the first version of your company brain.

18:05

SPEAKER_00

So I'll end with this. The way I think about a context layer is it's a system that turns knowledge and expertise and norms that we talked about that Maya knows into a machine usable context for AI systems.

18:22

SPEAKER_00

At a very high level, the way I like to think of it is it looks like this. It continually is mining context from your business systems. It's feeding this into that one company brain. It's harnessing this in skills and context development life cycles as your teams go and deploy these agents. And then it has a bunch of ways you can retrieve it. So MCP, SQL, vector retrieval, hybrid assembly, all these different ways that you retrieve it and pull back from traces and build this compounding learning loop.

18:58

SPEAKER_00

Today, we're largely building agents by hard coding context.

19:06

SPEAKER_00

The scale of this problem, I truly believe, is underhyped because with scale, this can become really unsustainable and a little dangerous. Like all of us know this old joke, which is if you ask sales and finance the revenue number, you're going to get two different numbers. We're fast approaching a moment of starting to deploy autonomous systems where the same thing is starting to happen. So I'll end with one last thing. I started this presentation by saying context is king. I'd like to end it by saying context is also IP. Something I think a lot about is in a world where you and your competitor have access to the same models

19:49

SPEAKER_00

and the same intelligence, what differentiates a company? What differentiates a customer support agent at American Express versus Amazon? That's how you do business. That's what makes your company special. Context is how we take and encode our culture and our norms into something that we will be proud of as we build autonomous Frontier forms. And that's all I had. You can find me at Trucalpa on Twitter or write to me. We are actively working with folks on the Frontier ongoing and shipping and building company brands. So if you'd like to talk to us, feel free to reach out. Thank you. Haha Haha Haha Haha Haha Haha Haha Haha Haha Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note