AI Engineer

Mergeable by default: Building the context engine to save time and tokens — Peter Werry, Unblocked

845 summary words 4 min summary Watch video

Start with the signal

4 min read

Summary

Summary: Building Context Engines for AI Agents

Main Topics

  • Context Engines: Systems that intelligently supply relevant organizational context to AI agents to improve task execution
  • Three Myths About Context Engines: Debunking common misconceptions about implementing context engines
  • Organizational Knowledge Building: How context engines understand team structure, expertise, and decision history
  • Data Integration & Governance: Challenges of pulling together multiple data sources while respecting access controls
  • Social Graph Architecture: Building relationship maps to identify organizational experts and team structures
  • Practical Implementation Workshop: Hands-on coding session building a social engineering graph tool

Key Points

The Evolution of AI Context

  • Four years ago: Developers manually provided all context to AI systems (8K token windows)
  • Today: Larger context windows but organizations have exponentially more data than can fit
  • Future trend: Autonomous background agents requiring highly optimized, pre-curated context

Three Myths Debunked

  • "Naive RAG over my docs is a context engine"
  • Vector search alone causes satisfaction of search problem (agents stop at first "good enough" answer)
  • Requires personalization and conflict resolution, not just keyword matching
  • "Bigger context windows solve this"
  • Even with million-token windows, models struggle with reasoning across diverse data sources
  • Organizations have more than a million tokens of relevant context anyway
  • Larger windows don't improve understanding of organizational intent
  • "Just wire up MCP servers"
  • Access ≠ understanding; connecting tools doesn't build semantic relationships
  • Misses organizational history, decision rationale, and best practices
  • Agents need guidance on why things are done certain ways

Core Context Engine Principles

Unified System Context

  • Builds relationships between data sources (not just storing them)
  • Distills historical patterns into organizational memories
  • Tracks decision evolution and rejected approaches

Conflict Resolution

  • Not just recency-based; must weigh against code, expert opinion, and organizational direction
  • Should surface unresolvable conflicts rather than hiding them
  • Learn from human feedback to improve future decisions

Data Governance

  • Respect access controls (private Slack channels, secret projects)
  • Compartmentalize synthesis to avoid crossing permission boundaries
  • Tagging mechanisms to ensure information reaches only authorized users

Targeted Retrieval

  • Personalize based on user's typical work areas (PR contribution patterns)
  • Bias toward relevant repositories over exhaustive search
  • Focus on expert conversations over noisy general discussions

The "Satisfaction of Search" Problem

Borrowed from radiology—agents stop searching as soon as they find something that looks right, missing deeper insights. Context engines must surface critical information proactively (incident reports, past failed attempts, organizational experts).

Real Performance Gains

| Metric | Without Context Engine | With Context Engine |

|--------|----------------------|-------------------|

| Time | 2.5 hours | 25 minutes |

| Token usage | 21 million | 10 million |

| Task completion quality | ~20% accurate | ~80% accurate |

| Doom loops | Frequent | Minimal |

Key insight: 90% of agent execution time is context collection; 10% is actual code writing

Notable Quotes

> "AI generated code should just feel like it was written by someone that's been in your team for 20 years."

> "Access doesn't equal understanding."

> "The puck is going down the line towards background agents for sure."

> "Without context, you'll probably end up in doom loops."

> "Satisfaction of search is a real problem—agents will stop as soon as they find what looks like the answer."

> "The only AI step in this is to label the teams. You don't have to run the AI step if you don't want to."

> "As we move towards a more autonomous universe, response times for MCP servers are actually less important. The more important thing is that they get the answer absolutely bang on."

Takeaways

For Teams Building Context Engines

  • Start with organizational structure: Build social graphs to identify experts, not just code metrics
  • Respect permissions: Don't expose data across organizational boundaries
  • Enable feedback loops: Surface conflicts and learn from human corrections
  • Optimize for understanding: Use multiple retrieval layers (vector search → pre-built memories → expert distillation)
  • Track sentiment, not just metrics: "Vibes" matter—measure user satisfaction with agent interactions

For Using AI Agents Effectively

  • Invest in context curation: Better up-front context dramatically reduces iteration loops
  • Use context engines in planning & review phases: Biggest ROI for agent work
  • Consider use cases: Ticket enrichment, triage, incident management, customer support see major benefits
  • Budget time for learning: The model learns from feedback; correct it when wrong

Practical Implementation

  • Social graphs can be built with procedural algorithms (PageRank-style) + LLM distillation
  • Best practices distillation happens at ingestion time (weekly cadence is sufficient)
  • Can run against your own repo locally without uploading data
  • MCP servers + API + CLI provide multiple integration points
  • Free enterprise trial available for evaluation

Future Direction

  • Autonomous agents as the primary use case
  • API-first architecture for flexibility across tools
  • Reduced response times while maintaining accuracy
  • Extended benchmark tracking (currently sentiment-based at ~0.75-0.8 normalized score)
Full transcript 13518 words · 77 min read
0:01

[SPEAKER_06] All right, thanks everyone.

0:14

SPEAKER_06

Sorry about the wait. This is gonna be a bit of a strange session because there is a workshop component to this, so everyone will be coding on their laps. Sorry. But anyway, sorry, I'm Peter and this is my colleague Brandon. So we're gonna break this session into two different parts. One is a talk that I'm gonna give about what context engines are useful for and how you might go about building one, what to think about. And then we'll launch into the second part of it. So just briefly, quick agenda. We're gonna talk about three myths that are circulating right now about context engines.

0:46

SPEAKER_06

And then I'll go over a couple of lessons or a few lessons that we learned along the way building one of these things. And then finally, we'll do this. So we're gonna build a social engineering graph. This is a component that is super useful in a context engine. And first, just a show of hands, does everyone know what I mean by context engine? Or does anyone want clarification on that? Yes, ma'am. So typically what happens is the first thing they do is they start to rip around your code base based on the task that you give them to try to gain some understanding, background understanding before they start to do their task.

1:24

SPEAKER_06

So context engineering is the art of supplying all the context that you need and most importantly, none of the context that you don't need in a highly optimized way so that when the agent starts to run, it executes the task in a streamlined way that's in line with your organization's best practices and expectations and so on.

1:29

SPEAKER_06

Okay. So we'll get to this after. So not long ago, as in four years ago or less, you were the context engine. Okay. So when your agent needed something, you would prompt it, you'd grab the issue ticket, you'd hand it all of the information that it needed to start its task. And in many cases, even when it was ripping around and getting background context, when it got to the end of its task, sometimes it got it wrong. In fact, many times it did. And you'd have to reset it, re-guide it towards the solution that you were thinking of. Or if it completely missed the mark, you'd have to say, no, not the JavaScript dummy. It's the Python source code that I want you to look at.

2:34

SPEAKER_06

So let's just remember how you built context in an organization. So we're taking AI out of the picture for just a sec. And I just want you to pretend that pre-AI, you just joined an organization. Let's remember how we built it up. So over time, you would accumulate this context through experience, right? You would start a job, maybe start code spelunking a little bit to figure out how things work. You'd maybe latch on to a mentor. And eventually, you'd experience real things like incidents and outages and things like that.

3:12

SPEAKER_06

Those are the pain things that stick with you. Those are the battle scars, right? And that is what constitutes organizational context. It's the learnings along the way, the why did we do things the way we did it. And now you're good at your job because after all of that experience of pain, now you know what questions to ask. You know where to look when an incident happens. And this is the goal. This is what we want to get for our AI agents. So I'm gonna just lift this adoption curve from Basim Illedath. And I may have butchered his last name, but sorry, Basim, if you see this. So let's start at the beginning here. This was four years ago in 2022.

4:12

SPEAKER_06

Everyone remembers Fancy Autocomplete, right? So back in those days, context windows in AI were pretty limited. I'm not sure if everyone even remembers this, but it was 8 kilobytes or 8K tokens, I should say. And that's not a ton. And so tokens were highly optimized and agentic IDEs like cursor focused just on the code that surrounded the code that you wanted to go in and autocomplete.

4:31

SPEAKER_06

So basically they took some code before, they took some code after, they put it into a model and they said, this user's working on this piece of code, what's the most likely next thing? And that's what was printed out. It got progressively better than that. ASTs were integrated, language servers, and then you were able to basically pull callers of source code and pull all that into context. And then the LLMs were really good at completing code. So at those levels, you were the context engine. And in many circumstances here, this is where most people are here. They're at the parallel agents hooked up with MCP and skills. Okay?

5:18

SPEAKER_06

Just super curious, has anyone gone beyond curated context into the last few degrees of agentic freedom, where you have background agents running in the cloud doing stuff in YOLO mode? Is anyone experimenting with that? Okay, cool. That's very cool. That's bleeding edge. But let's just take a moment to recognize that bleeding edge today is yesterday's news in six months.

5:49

SPEAKER_06

Okay? So the puck, I'm Canadian, so I'm gonna say this. The puck is going down the line towards background agents for sure. And one of the things that we run into right now is this, we're becoming the bottleneck as humans, right? I'm not sure if people have tried managing parallel agents and working on several tasks at once, but everyone's starting to feel this cognitive disconnect because you're context switching all the time. And it's just really, really painful.

6:25

SPEAKER_06

It is very difficult to move from that mode where you're the human managing context into the background agents mode unless you have some kind of context engine that knows how your code operates, how your organization works, and understands the motivations for historical changes and things like that. So Andre, he nailed it. Systems are intelligent. We're reaching the exponential on intelligence for code pretty soon. Everyone's seen the release about Mythics. Even though we all haven't had a chance to really try it out yet, the promise is that from a code intelligence perspective, this thing is pretty much close to perfect. But so now the bottleneck is context, of course.

6:43

SPEAKER_06

Without context, I'm just gonna reemphasize this point. You'll probably end up in doom loops. Does everyone know what a doom loop is? A doom loop is when you're struggling with the agent. It's not quite doing what you want. And you have to keep iterating on it. We're reaching the exponential on intelligence for code pretty soon. Everyone's seen the release about mythos.

7:31

SPEAKER_06

Even though we all haven't had a chance to really try it out yet, the promise is that from a code intelligence perspective, this thing is pretty much close to perfect. But so now the bottleneck is context, of course. Without context, I'm going to reemphasize this point. You'll probably end up in doom loops. Does everyone know what a doom loop is? A doom loop is when you're struggling with the agent. It's not quite doing what you want.

8:01

SPEAKER_06

And you have to keep iterating on it. The worst case scenario is you run this thing in Yolo mode and it finishes the entire task and it's completely wrong. And you have to go back and correct various stages. So when you have a context engine, you can get there faster. The problem is that access doesn't equal understanding. So we have customers that are on various parts of this journey. And one of the interesting things that we've noted is that people feel that they understand their organization best. So when it comes to feeding the right context to these agents, people will try to build some semblance of what a context engine actually is.

8:28

SPEAKER_06

They'll maybe build a rag system or they'll build some way to feed organizational data to an agent. So unfortunately though, access doesn't mean understanding. So what that means is you could just wire up a bunch of MCP servers and it's not going to be able to understand what the relationships are between all that data. How it got there and why it is the way it is. And then there's another problem which I'll talk about a little bit later called satisfaction of search. So just remember that term. I'll come back to it. Okay. So I just wanted to show you this. This was something that we actually implemented and we did it in two parts.

9:31

SPEAKER_06

One was just without any context engine, but wired up to a bunch of MCP servers. It did a pretty good job. But then when we reached the end, it missed the fact that we had some legacy stuff that depended on this old method of intelligence size to Anthropics. So they have adaptive thinking now, but it used to be you had to supply a token budget.

9:42

SPEAKER_06

And that's how you could increase the size of the thinking window. So we had some code that depended on this. And there were reasons for that, that the agent didn't understand or see. And so it just basically clobbered all that code. But when we added the context engine, then it saw all those reasons and implemented it the right way. So it made the appropriate changes in the right places, included backwards compatibility for the code that was using the old method. Okay. So now for the myths. Myth one, naive rag over my docs is a context engine. So if you implement, say vector search, or just a couple of search methods, you're going to run into this.

10:31

SPEAKER_06

You're going to run into a few issues. One is the satisfaction of search problem, where the agent will search crazy, consume your tokens. And then in the worst case, you'll reach compaction. Okay. So without being able to find the end game. There are a few different other techniques.

10:57

SPEAKER_06

You need to have personalization when you build a retrieval system. Because if you just rag all your data, especially for very large organizations, there's going to be things like conflicts that you have to resolve in the data. It won't be focused on the task that you're trying to perform. It might pull in relevant code from other parts of your organization that, especially if you have a really big org and you've got tons of different repos, it's just going to create a huge mess. So you need to have some element of personalization. And then here again, connect a bunch of MCPs. I'm going to reiterate this point. I'm done. Nope, definitely not.

11:39

SPEAKER_06

So that really puts an emphasis on the satisfaction of search point. And I'll explain that in a sec. And finally, a bigger context window will solve this. So way back, when the models were starting to get big, people were really excited about a million tokens in your context window. The first models that tried this, I think it might have been Claude actually. Was it Claude? I think it was Codex. [SPEAKER_07] Oh. Oh. Or OpenAI. Okay. Gemini. Yes, sorry. I'm so sorry.

12:34

SPEAKER_07

[SPEAKER_06] So yeah, Gemini first model that tried this.

12:35

SPEAKER_06

And it was really good at finding needle in the haystack. So you could feed a huge document to it. And as long as you knew what you were looking for ahead of time, it could find it. But it wasn't good at all at reasoning across different data sources. Understanding the real meaning behind a problem and then recommending the appropriate solution. So none of that was possible. Obviously things have gotten much better. Now the problem is most organizations have more than a million tokens worth of context. So trying to fit all that into their context window isn't going to work anyways.

12:49

SPEAKER_06

Let's project out to the future and imagine that you could fit 10 million tokens, 50 million tokens. At the current rate of memory consumption, just to operate the models, that's not going to be possible for a really long time. Even if it was, and you fit all that context into your context window, you're still going to run into problems with understanding what's true, what's false. How to select the right information. Okay? So now I'm going to come back to this second point here. Satisfaction of search. This is a term that actually comes out of the medical field in radiology.

13:25

SPEAKER_06

And the idea is that when techs are looking at x-rays and they're looking for the cause of symptoms, they might find something on the x-ray that explains those symptoms and then they stop. And that's a dangerous thing medically because there might be other indicators for things like cancer that get missed. So satisfaction of search is a real problem in radiology and there's lots of protocols to prevent just stopping as soon as you find the first thing. This is what happens with agents. Satisfaction of search. This is a term that actually comes out of the medical field in radiology.

13:52

SPEAKER_06

And the idea is that when techs are looking at x-rays and they're looking for the cause of symptoms, they might find something on the x-ray that explains those symptoms and then they stop. And that's a dangerous thing medically because there might be other indicators for things like cancer that get missed. So satisfaction of search is a real problem in radiology and there's lots of protocols to prevent just stopping as soon as you find the first thing. This is what happens with agents. When they search around in, say, Notion and your code, Confluence, they'll stumble across what looks like the thing they're looking for and they'll stop. And then they'll proceed.

14:34

SPEAKER_06

But the real golden nuggets of information might be in a different place that the agent wouldn't think to look, in a past Slack conversation or an incident report, something like that. So here's the classic iceberg meme. Code that compiles. That's the baseline. Does the agent produce code that compiles? But everything that is actually important happens underneath here. So understanding the user's original intent, what was rejected in the past by the team and tried before but failed. How are you going to surface that kind of content just by looking at docs and code and stuff? So you need to understand that somehow.

15:10

SPEAKER_06

And even worse, it's sometimes hard to know when things were deleted, in the absence of information. So you need history as well, leading up to decisions. So this is why we think you need a context engine. A context engine understands who you are, what team you work on, who you work with, who the experts are in your organization, and what the decisions were that led up to the current iteration of your code base. It's able to resolve conflicts. This is a true and false type situation. What's true, what's not.

15:48

SPEAKER_06

Sometimes that truthiness is a gray area, right? So the context engine needs to also understand when to instruct the agent that it wasn't able to resolve the conflicts and then learn from additional user input. This third point is super important, of course, in any large organization or enterprise. There's often repositories that not everybody can access, secret projects, that sort of thing. So it's really important that you flow the access controls up. We have an example that everyone will appreciate, which is Slack. Our context engine integrates with Slack or Microsoft Teams. And when you have private channels, that's really highly sensitive, right?

16:33

SPEAKER_06

You could be discussing HR information or maybe something that you just really don't want everyone else to see. And so when Unblocked answers questions, it will use private channel information, but it will only use that information if the person that's asking the question has access to it. And then those answers are not public. So they're private to you. And then finally, of course, delivering the right context at the right time. And this is about token efficiency. It's about getting to the answer as quickly as possible. So here's a high level overview of how a context engine might work.

17:12

SPEAKER_06

On the left, we've got data source inputs. Things like planning tools, docs, conversations, code, PRs—basically anything that's relevant to getting work done at the engineering level. And then on the right side, we have the output. This all can flow to coding agents, CLI tools. You can custom build apps through the API. We have a code review component that just plugs right into your SCM and provides code reviews. And of course, integrations with social messaging apps. These are the broad requirements that we think are important. There's actually much more than this, but these are the high level things. So again, unified system context.

17:46

SPEAKER_06

This is about building relationships between data. But it's more than just recognizing when one piece of data is related to another. For example, in Slack, you might have conversations about PRs. That's an easy linkage because you're posting links back and forth. What's less easy is understanding the reason why decisions were made or your organization's best practices, right? So to understand that you have to go a little deeper. Do things like distill historical pull request comments on PRs and try to distill those down to their core essence. And then when you see repeated patterns, you can pull those patterns together and store them as memories.

18:18

SPEAKER_06

So that when someone is working on a similar piece of code, you can load those memories and then the agent can see that and go, oh yeah, right.

18:24

SPEAKER_06

This is the way this organization does this particular thing. Conflict resolution. Super important. We took initially a naive approach to this at first and based it just on recency, right? So we would bias towards newer stuff. Unfortunately when you have the fullness of all your context, recency is not enough. Often you have people writing documents or chatting in their messaging platforms and they might be saying things that are not completely aligned with how the system works. So then we started to bias towards code.

18:53

SPEAKER_06

We had recency and we said the main branch is definitely your source of truth, but not always because sometimes what's important is what happens next, not the way a system currently works. When you're working on a task, what you really want is for the agent to understand where you're going, not necessarily where you've been. Where you've been helps it understand what not to do. Where you're going helps it understand what you should do. So in the Slack case, looking at the conversations that your organization's experts are having is more important than just understanding what every random engineer is talking about.

19:10

SPEAKER_06

Targeted retrieval and personal relevance are very related. So I'll just talk about them together briefly. When you're working on a task, what you really want is for the agent to understand where you're going, not necessarily where you've been. Where you've been helps it understand what not to do. Where you're going helps it understand what you should do. Okay. So in the Slack case, looking at the conversations that your organization's experts are having is more important than just understanding what every random engineer is talking about. Targeted retrieval and personal relevance are very related. So I'll just talk about them together briefly.

19:45

SPEAKER_06

So again, when you're pulling context in, it's important that you're only pulling context in for the relevant task at hand and probably relevant to you. So here's a technique that's interesting. You can understand what repos a person works on most by the number of PRs that they submit contributions. And then if you're doing vector retrieval, you can do a deep retrieval on those focused repositories and then a wider retrieval on the rest of the source code and then bias the selection towards the focused repositories. That's more likely where someone's going to be working and spending their time.

19:55

SPEAKER_06

And then we've talked about data governance, so I don't think I need to go over that again. Super important though. This was just an experiment that we ran with a larger task. I fully admit that some of these numbers are a bit wonky. This is basically Claude outputting numbers. So don't trust it. Just trust the vibe of the thing and not necessarily the numbers.

20:05

SPEAKER_06

Basically what it's saying is that when we started out without the MCP server, sorry, without the context engine active, it really missed the mark on a lot of stuff. That's just because it didn't understand how the existing implementation really worked and why it was the way it was, what was tried before and failed. And so it made a lot of those same mistakes. With a context engine turned on, obviously it nailed it. The key numbers though are the time and the tokens that it took.

20:26

SPEAKER_06

So without the context engine, it took two and a half hours to finish this task with 21 million tokens, which is a lot of tokens. But with the context engine, it took only 25 minutes and 10 million tokens. So it's a pretty dramatic difference. Okay. So the hard lessons, these are just samples by the way, but these are ones that we thought were interesting. So first of all, initially we optimized for access, not understanding. Our first premise was if we just wire up a bunch of tools and provide a knowledge graph, it will be able to traverse the knowledge graph and execute a bunch of retrieval specific tools for particular integrations and so on and figure everything out.

20:55

SPEAKER_06

That does not work. So you'll have to go a little bit deeper than that. Second one is we hid conflicts instead of surfacing them. So by hiding conflicts, I don't mean that we just ignored the conflicts. What we did instead was we tried to resolve those conflicts using naive strategies and we didn't surface the conflicts that we weren't able to resolve. So this was a really good learning. A context engine, I mean we'll get there eventually probably, but it can't always tell what the truth elements are. And when it can't, you should surface that and learn from it. That's the key thing.

21:14

SPEAKER_06

And then finally, I think a lot of folks tried this. This is a really bad idea. So when a context engine supplies an answer, do not cache the answer and try to serve that same answer up again to a similar question. The reason is fairly obvious in retrospect, but everything changes constantly, right? Code changes, docs change, the reason for things change. So this just doesn't work. The other thing is if you try to use the previous answers as context for new answers, you regress towards a mean. So if the model is misbehaving or doing something bad and you continuously bring that into context, you're obviously going to pollute the context. And this is what happens.

21:33

SPEAKER_06

Okay. So let's now talk about where AI forward teams are using and taking advantage of context engines. Definitely, and especially during the planning phase. This is where you get the biggest bang for buck unquestionably. Get the context engine involved, use a skill to bring it in, connect it to the MCP server and watch it do its thing. This is where you get the biggest bang for buck. It's also useful to do this during review. So you get planning and review at the end. Because if you get an agent to do review, it's basically just going to pay attention to the code and try to understand where the breakpoints are. Security concerns, that kind of thing.

22:03

SPEAKER_06

But without the organizational context, it doesn't understand the motivation for it. So that's the really important thing. Ticket enrichment. This is a super cool use case. So you create a ticket for a new feature and then you just ask the agent that's connected to a context engine to fill in the blanks. Works.

22:13

SPEAKER_06

Triage. I use this all the time. When I see an issue in production, I just whack it into it. And my agent connects to the context engine and it just instantly brings up all the past issues related to this and starts operating right away. Increasingly, we're seeing this one is incident management. Okay. So we just wired up Datadog and Century. Sorry. Century and Datadog. And this is already proving a super cool use case. It can see the signals and then it can act on all the signals and relate that to code, relate it to past incidents that you and discussions that you've had in Slack. Having all those things come together at once is almost magical.

22:26

SPEAKER_06

And finally, I think this one's actually my favorite one and it's the one that customers use the most, is customer success and sales and engineering support. So what a lot of big teams do is they have engineering support channels where other teams can come in and ask questions. If you put a context engine into one of these, you can have it automatically answer a lot of questions and save engineers a ton of time. It can see the signals and then it can act on all the signals and relate that to code, relate it to past incidents that you've had in Slack discussions. Having all those things come together at once is almost magical.

23:00

SPEAKER_06

And finally, I think this one's actually my favorite one and it's the one that customers use the most, is customer success and sales and engineering support. So what a lot of big teams do is they have engineering support channels where other teams can come in and ask questions. If you put a context engine into one of these things, you can have it automatically answer a lot of questions and save engineers a ton of time. All right. So how teams make a context engine their own skills. So definitely build skills that you can use to curate context in the GitHub repo. And you can build other skills around it like typing ticket enrich, give it the issue ID.

23:36

SPEAKER_06

And then it can use the context engine to build the enrichment. Workflows like this one prepare issue ID. Incident timeline. And then you can just send it off to your agent again, context engine, blah, blah, blah.

24:03

SPEAKER_06

It brings everything together. Magical. And this thing here, you can wire this up to all kinds of agents. I've got one of the things that a lot of customers like to do is wire this up to cloud code in their CI system. We actually do have a code review component. So you don't have to do this if you're using unblocked. But people use this for other things, not just code review. As soon as you wire up a context engine in the background, give it an API key, let it run on its own. It can do some pretty insane stuff. So I'm just going to show a quick example of what wiring up a context engine can do. So this is a PR that my colleague wrote.

25:07

SPEAKER_06

And unblocked went through and provided a kind of review to this thing. And at the bottom of this review, here's the review part, you can see that Richie, who was the author of this PR, was very cool. This is something I would say.

25:31

SPEAKER_06

Now the reason for the comment, which was you've basically duplicated a bunch of tests.

25:36

SPEAKER_06

You can dry that up a little bit. Is because this was a best practice that was distilled from a bunch of other PRs. And the funny part is that the author of those PRs was Richie. So he's the one that actually instilled the best practice in the organization. So that was just a cool little moment when we discovered that. Here's another example. So this was a fairly long transcript. I'm not going to show the whole thing, but we sent it on a mission to do a big, large task. Without unblocked, it took quite a while. You can see the transcript is quite long. And it missed a whole bunch of stuff. With unblocked, it was a lot more compact.

26:32

SPEAKER_06

It got to the answer very quickly and correctly. And just because we're now AI forward and lazy, we took both of those transcripts and ran them into Claude. And just said, hey Claude, why don't you do an analysis of both of these things and give us your result. So it went through, I won't bore you with the details, but just to say that at the end, the verdict is that the context engine plan is what I'd ship with. This other one is good for a prototype, but it's missing a whole bunch of stuff that is important to this organization that was previously discussed. Okay. So this is essentially what I've been trying to say.

26:52

SPEAKER_06

AI generated code should just feel like it was written by someone that's been in your team for 20 years. Okay. If it doesn't yet, that's fine. It will. If you wire up unblocked, you'll see a huge difference in performance of agents. And if you're building one of these things, absolutely take all these things and build and let's see where that goes. So just before we get into the workshop component, maybe we'll have five, ten minutes of Q and A. What up? I'm Brandon. And this is Brandon. So he'll help with Q. I swear I'm on this. [SPEAKER_07] Oh, maybe. [SPEAKER_07] Is that the report? [SPEAKER_07] Yes. Oh, thank you. Sorry. I'll give you a mic. Thanks.

28:03

SPEAKER_06

[SPEAKER_08] So it's clear what it does for you and what kind of problems it solves. [SPEAKER_08] But to me, a big question mark is what is the thing? [SPEAKER_08] What is the artifact that fits the bill? [SPEAKER_08] Is it a program you install, an API that's hosted remotely, or an MCP server? [SPEAKER_08] What is it? It's all of those things. So a context engine, I'll explain what Unblocked is. Maybe I can just show a quick demo of it. So broadly speaking, there's a bunch of different surfaces to a context engine. You want to get it into your agent flow and you can do that with an MCP server. You can do that with a CLI tool, for example.

28:51

SPEAKER_06

We also have this dashboard surface where you can ask questions about your code. This is a pretty basic one. But you can see it understands who I am and what I've been working on. And then we have Slack connectivity as well. So you can bring Unblocked into Slack, drive it in conversations and have it auto answer things. Does that make sense? Did I answer your question or? [SPEAKER_08] I think my question will get the answer in the demo. [SPEAKER_08] Okay. [SPEAKER_08] But yes, API, CLI, MCP. [SPEAKER_07] Yeah. And products. Mostly API. [SPEAKER_07] Yeah. Oh, sorry. I got carried away. [SPEAKER_07] Thank you.

30:03

SPEAKER_06

[SPEAKER_11] And so my question is, as far as I understand, it's a knowledge management and retrieval application. [SPEAKER_11] Yeah. [SPEAKER_10] And does this relate somehow to things like LLM Wiki, like it was made popular recently by Andre Carpatti, or the decision traces and context graphs, which was discussed a lot a few months ago. [SPEAKER_11] Yeah. I think my question will get the answer in the demo. [SPEAKER_08] Okay. [SPEAKER_08] But yes, API, CLI, MCP. [SPEAKER_07] Yeah. And products. Mostly API. [SPEAKER_07] Yeah. Oh, sorry. I got carried away. [SPEAKER_07] Thank you.

31:15

SPEAKER_06

[SPEAKER_11] And so my question is, as far as I understand, it's a knowledge management and retrieval application. [SPEAKER_11] Yeah. [SPEAKER_10] And does this relate somehow to things like LLM Wiki, like it was made popular recently by Andre Carpatti, or the decision traces and context graphs, which was discussed a lot a few months ago. [SPEAKER_11] Yeah. So you can think of all of those things as useful components to a context engine. A context engine has to do much more than that because agents are really good at recursing through a wiki, for example.

31:38

SPEAKER_06

It depends on how you build this wiki, because there's a bunch of things like organizational memories, best practices, experts in your organization, and those are used as pivot points for context retrieval. So a wiki doesn't solve those problems unless it has a structure with it. And I think Carpatti discovered that if you treat a wiki as a file system, you can break it down and have the agent go through it like a file system. Agents are, by the way, highly optimized for file system traversal.

31:56

SPEAKER_06

The compilation step. [SPEAKER_06] Yeah. The compilation step. Exactly. Yeah.

32:01

Yeah. I'm sorry. Everyone else is ready. Okay. Sorry. [SPEAKER_01] Maybe the same question, but is it a general purpose context engine or is it targeted against code? Because it would be useful as, say, a business domain expert, or building up a business domain and then having this context engine use my information, so all my other agents could use this as context for their business. Or would you say that is more for just the code part of it? [SPEAKER_01] It's definitely engineering focused. [SPEAKER_06] The integrations are focused on engineering activities. So, you know, SEM integrations and other tools that engineers use.

32:23

SPEAKER_08

We are increasingly seeing customers using it for other purposes. So, business intelligence is a key thing.

32:26

SPEAKER_06

And that's usually useful when people in business functions are trying to get an understanding of the product and its function. We don't have, say, Salesforce integrations wired up for that. So you couldn't use it to understand anything that's sales related. It's really just primarily an engineering focused context engine. That's not to say that won't change. More questions. Yeah. [SPEAKER_03] The governance thing. Yeah. If you're respecting access rights, how can it do synthesis across stuff and then develop new knowledge internally that it could then surface to people? So yes, you're correct to point that out. The synthesis is compartmentalized.

33:12

SPEAKER_06

So there are places that are compartmentalized like individual repositories.

33:17

SPEAKER_08

That's the level of access. So if you can synthesize historical data based off of that, and then correlate that with public Slack information, then that's one way to do synthesis without crossing organizational boundaries.

33:21

SPEAKER_07

The other way is to look at and tag when synthesized information crosses those organizational boundaries.

33:21

And you can take something like a group ID approach to that problem where you attach group ID tags to the synthesized information and then only retrieve it if the person that has access to that can build it out. So first take the compartmentalized approach because that's where you'll get the most mileage and then you build up from there. I mean, this is the core problem with using a technology like graph rag, right?

33:25

Because graph rag is a pyramid where it builds up in layers and then summarizes at each layer, but that unavoidably crosses permissions boundaries. So you have to create compartmentalized pockets. Yeah.

33:30

SPEAKER_11

More questions? [SPEAKER_10] That's a good question.

33:39

SPEAKER_10

[SPEAKER_06] Yeah.

33:56

SPEAKER_11

[SPEAKER_10] Yeah.

33:57

SPEAKER_06

Yeah. [SPEAKER_07] So you've talked a lot about all the different sources of information that you consume and putting them all together. When it's synthesizing those down, is that still naive rag vector search, all that stuff under the hood or is it agents deciding what is appropriate? Or probably combinations of all of them, but what is that step? [SPEAKER_07] Yeah, you're right. It is a combination of all of them. So knowledge graph buildup happens in a bunch of different ways. The PR thing that I showed you, for example, is first you build a naive knowledge graph procedurally.

34:47

SPEAKER_06

And then from there you can use an LLM to distill down and summarize and build up those types of techniques. Our context engine builds first a knowledge graph from the base using, trying to leverage all the different entities. It's a page rank thing where it builds up the relationships procedurally. And then of course it vectorizes data. And then there are procedural tools that fetch data at runtime. A lot of the distillation for conflict resolution happens in two places. So one is during data ingestion time, there are tags that relate data to each other so that we can see if we can de-conflict at that level. And then rank against each other at that level.

34:54

SPEAKER_01

And then at runtime you have to pass the things to a judge with the criteria. And then it does additional de-confliction in real time. Does that make sense? Yeah.

35:27

SPEAKER_06

Okay. I have one more question. So I was curious. You said conflicts, but at some point you get conflicts that something means revenue for one company and means revenue for another company. [SPEAKER_09] It's a totally different meaning how you can recognize that. So how do you get humans in the loop? How do you use their ontologies and how can you do it when you run into it? So I'm very curious about that actually. Yeah. So if you, I can show you just a quick thing here. So you'll notice that at the bottom, the references that were used for answers are delivered both to the human in this interface, but also to the agent.

36:17

SPEAKER_03

[SPEAKER_06] Yeah. [SPEAKER_06] You said conflicts, but at some point you get conflicts that something means revenue for one company and means revenue for another company. [SPEAKER_09] It's in a totally different meaning how you can recognize that. [SPEAKER_09] So how do you get humans in the loop?

36:35

SPEAKER_06

[SPEAKER_09] How do you use their ontologies and how can you do you use it when you run into it? [SPEAKER_09] So I'm very curious about that actually. [SPEAKER_09] Yeah. So if you, I can show you just a quick thing here. So you'll notice that at the bottom, the references that were used for answers are delivered both to the human in this interface, but also to the agent. Yeah. So if the agent, if the context engine isn't able to do the de-confliction, then at this point here, the human can step in and guide the agent when there are enough signals. So you can literally just reply and say like, that's not correct. Or you can come here and.

37:39

SPEAKER_06

Oh, yeah. [SPEAKER_07] Sorry. [SPEAKER_10] Yeah. [SPEAKER_10] Or you can do this not helpful and give the reason why.

37:58

SPEAKER_10

[SPEAKER_06] It is a bit of a manual process at this stage, but the signals that build up over time.

38:02

SPEAKER_06

It's funny, right?

38:03

[SPEAKER_09] So you get some production, but some companies you might have, you catch a lot of human intelligence by this, right? [SPEAKER_09] Yeah. [SPEAKER_09] That's amazing. [SPEAKER_09] Yeah. [SPEAKER_09] Yeah. [SPEAKER_09] For a typical customer, how much do you have this volume of that and metric?

38:27

SPEAKER_06

[SPEAKER_09] Oh, it's huge. It's amazing. I was actually really surprised by how willing people are to give feedback. Yeah, no, it's. Can you give an example, one project that you catch with thousands or hundreds of these human feedback things? [SPEAKER_09] At small team size, it's in the hundreds at so small team size being like 20, 30 people. Yeah. At large team size, a hundred to 200 people, it's hundreds and hundreds of feedback. Oh, wow. [SPEAKER_09] Cool. People just really like to interact with agents and tell them in natural language what's wrong. It's just a totally natural thing to do. Yeah. Good purpose. Cool.

39:42

SPEAKER_06

All right. Are we good for Q and A? And then we can get on. Yeah, let's get onto the workshop part of this.

39:52

SPEAKER_09

[SPEAKER_06] So we have created a, for actually what I'll do is I'll just do this first. [SPEAKER_06] So you can do this now if you'd like. [SPEAKER_06] I will come back to the slide in a sec. [SPEAKER_06] So the idea here is we're gonna get everyone to join a Slack workspace that we created. [SPEAKER_06] And then we're gonna get everyone into a repo where this sample code lives.

40:07

SPEAKER_06

And then we'll just start hacking away on it together. Hopefully this is working. [SPEAKER_07] Yeah. Light mode. [SPEAKER_07] Thank you. [SPEAKER_07] One time I muted my mic. [SPEAKER_07] Hopefully that's working.

40:41

SPEAKER_07

[SPEAKER_06] Oh, I got some people coming in here.

40:42

SPEAKER_10

[SPEAKER_06] Nice. [SPEAKER_06] When you drop in, you'll see an AI engineering London channel.

40:46

SPEAKER_06

[SPEAKER_07] Hopefully. [SPEAKER_07] Yeah.

40:54

SPEAKER_09

[SPEAKER_07] There's a link to click there. [SPEAKER_07] And please, in this thread, drop your GitHub, your name and that will bite. [SPEAKER_07] The unblocked link will not work until you do step two. [SPEAKER_06] Yeah. [SPEAKER_06] To get into the GitHub org. [SPEAKER_06] Oh no. [SPEAKER_06] I've got many people coming in so I'm really hopeful.

41:09

SPEAKER_06

[SPEAKER_07] Is it network? Yeah. It might be network. We will find out.

41:23

[SPEAKER_07] Okay. [SPEAKER_06] Well, while folks are doing that, I'm just going to show you what we're getting into here. So this is the GitHub organization. What we're working on is a social graph builder. So what this is going to do is look at a source code repository. So you can run this on your own repo.

41:39

[SPEAKER_06] It's not going to upload anything. [SPEAKER_06] It's all local. So that you can see this thing building up against your own organization. And it's going to, you know, we're going to do a bunch of things. We're going to get basically a social graph out of it and I'll show you what that looks like. And we're going to understand who the experts are and which parts of the code they work on. [SPEAKER_06] And then there's going to be a little interactive visualization thing. So what the goal of this exercise is, is to get this thing up and running and start hacking away on it. So start submitting PRs as soon as we get this going. So this is what it looks like.

41:58

SPEAKER_06

This graph here is our organization unblocked. Unblocked. And what you're seeing here is a relationship graph that shows who's reviewing whose PRs, and who's getting reviewed essentially. This thing is a distillation of all the different teams within unblocked. So this is roughly accurate actually. Well, not roughly, it is pretty accurate. I did this all the way back to the start of 2025. When you run the thing, I'd recommend maybe doing it for a shorter timeline because it will be a little bit slow.

42:34

SPEAKER_07

[SPEAKER_06] If you go all the way back to 2025, it could take like 15 minutes. [SPEAKER_06] But it's effectively distilled who the teams are. [SPEAKER_06] And the only AI step in this is to label the teams.

42:38

SPEAKER_06

You don't have to run the AI step if you don't want to.

42:39

SPEAKER_07

[SPEAKER_06] It'll just use the parts of the code that people work on the most. [SPEAKER_06] This tab here will show the experts in the organization and what they work on.

42:47

[SPEAKER_06] So this is just broken down by project area and path. [SPEAKER_06] And shows what areas of the code have good coverage. [SPEAKER_06] Coverage is defined mostly by whether a high contributing organizational expert is present. If you go all the way back to 25, it could take 15 minutes.

42:56

SPEAKER_07

[SPEAKER_06] But it's effectively distilled who the teams are. [SPEAKER_06] And the only AI step in this is to label the teams. [SPEAKER_06] You don't have to run the AI step if you don't want to. [SPEAKER_06] It'll just use the parts of the code that people work on the most. [SPEAKER_06] This tab here will show the experts in the organization and what they work on. [SPEAKER_06] So this is just broken down by project area and path. [SPEAKER_06] And shows what areas of the code have good coverage. [SPEAKER_06] Coverage is defined mostly by whether a high contributing organizational expert is present.

42:56

SPEAKER_07

[SPEAKER_06] And whether it's an actively contributed to part of the code. [SPEAKER_06] And then finally, we'll have this interactive graph that breaks things down by team area. [SPEAKER_06] And we'll show who the major contributors are. [SPEAKER_06] I'm over here on the AI team. [SPEAKER_06] Yeah. [SPEAKER_06] So that's it. [SPEAKER_06] Let's get everybody in and we'll start hacking away at this. [SPEAKER_06] Yes, absolutely. [SPEAKER_06] Yeah.

43:10

SPEAKER_06

Many of you should have an invite, but put your GitHub already in. [SPEAKER_07] So please give it a check.

43:17

SPEAKER_06

[SPEAKER_07] GitHub is the worst. [SPEAKER_07] Yeah.

43:22

SPEAKER_07

[SPEAKER_06] We will, it is an MIT license.

43:23

SPEAKER_06

[SPEAKER_07] We will be making it public later, but for now we needed it locked down. [SPEAKER_07] Oh, did you, you still wanted to fly. [SPEAKER_07] That's crazy. All right. What's your app? Just what's your name? I'll invite you.

43:26

SPEAKER_07

KAS.

43:26

SPEAKER_06

[SPEAKER_07] KAS. [SPEAKER_07] If you do, you do. Okay. Yeah. Yeah. Yeah. Oh, thank you. I'll show you. [SPEAKER_07] I'm coming over. I will put it in. What have you done? Let me type this out. It's a bit weird. I keep trying to type kiwi, but it's not the right. Yeah, that was good for him. [SPEAKER_10] This is Dutch. [SPEAKER_10] No, no, no. [SPEAKER_10] Thank you. [SPEAKER_10] I also have to slide a different dog with my water, so I just took it down. [SPEAKER_10] It's your time. All right, you've got the mic. I'm at Julian, and that was. I think the rest of this session is going to be now just hacking away.

43:51

SPEAKER_06

In a sec here, I think I'll take this down if everyone's got it so that Brandon and I can concentrate on working with you guys to build features. Oh, when you submit PRs, by the way, you'll notice that Unblocked is sitting there as a coder viewer, so don't feel badly if it sprays on your PR a little bit. It depends on how much slop you send. [SPEAKER_07] It depends on how much slop you send. [SPEAKER_10] Is everyone good with this? I take it down? Okay, cool. That is it. Brandon, you're on top of the invites? Okay, cool. There's a few more. I'm on to Chris. [SPEAKER_07] Chris has given two. [SPEAKER_10] That's okay. [SPEAKER_07] I'm going to send both.

44:34

SPEAKER_06

[SPEAKER_07] Don't worry. [SPEAKER_10] Amazing. [SPEAKER_07] That's just where I am in this list. [SPEAKER_07] Yes. Thank you. Oh, forgot to mention a couple of things here actually. Yeah. Coming back live? Yeah, good. Just a couple of things. So if you're looking for something to implement and starting with coming up with ideas and stuff, there is a set of predefined issues that you can hack away on. So you can just grab one of these, whack it into Clyde and see how it does when it's connected to the context engine. The MCP server for unblocked is here.

45:19

SPEAKER_06

So if you want instructions on how to wire this up to a Cloud Code or another agent, then you can grab it from the instructions from here. All right. [SPEAKER_07] So I'm at Lars. [SPEAKER_07] There's two more in here. [SPEAKER_07] So I'm still going by the way for those just adding their names. [SPEAKER_07] How are those questions. [SPEAKER_07] Yeah. [SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions.

46:21

SPEAKER_07

[SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions. [SPEAKER_10] How are those questions. [SPEAKER_10] Thank you.

46:27

SPEAKER_06

[SPEAKER_07] Thank you.

46:29

SPEAKER_07

All right. [SPEAKER_05] Kat, you should have an invite. [SPEAKER_05] Andra, you're coming next. [SPEAKER_05] I'm just too behind. [SPEAKER_10] Christopher, did you not get his invite yet? [SPEAKER_10] No, no, I didn't. [SPEAKER_10] Oh, that's weird. [SPEAKER_10] Let me double check.

46:45

[SPEAKER_07] You should have one, but... [SPEAKER_07] Yeah, you should have an email. [SPEAKER_07] I'm up to like 31 of you. [SPEAKER_07] So... [SPEAKER_07] It's like five clicks to add a member. [SPEAKER_07] I'm like, GitHub. [SPEAKER_07] We need to talk.

47:19

[SPEAKER_07] I was going to say, you should be able to use the CLI for this. What's going on? [SPEAKER_07] You don't have an agent for it yet. [SPEAKER_07] That's right. [SPEAKER_07] That's right. [SPEAKER_07] No, they keep putting it in my PR, and I don't want it there. [SPEAKER_07] Co-pilot's going to review for me. [SPEAKER_07] Very poorly, but it will review. Andra... Is that there from... Let me double check. Ready? [SPEAKER_07] Yeah, for sure. [SPEAKER_06] Hold on. [SPEAKER_06] Let me grab you the mic. [SPEAKER_06] Hopefully that's on.

47:48

[SPEAKER_06] Does it work?

47:54

SPEAKER_06

[SPEAKER_11] Yes. [SPEAKER_07] What's going on? [SPEAKER_07] You don't have an agent for it yet.

48:24

SPEAKER_06

[SPEAKER_07] That's right. [SPEAKER_07] That's right. [SPEAKER_07] No, they keep putting it in my PR, and I don't want it there.

48:41

SPEAKER_07

Co-pilot's going to review for me.

48:53

SPEAKER_10

[SPEAKER_07] Very poorly, but it will review. Andra...

49:23

SPEAKER_06

Is that there from... [SPEAKER_10] Let me double check. Ready?

49:36

SPEAKER_06

Yeah, for sure. Hold on. Let me grab you the mic. [SPEAKER_06] Hopefully that's on.

49:45

SPEAKER_07

[SPEAKER_06] Does it work?

49:48

SPEAKER_10

[SPEAKER_11] Yes.

49:48

SPEAKER_07

[SPEAKER_11] So I guess that context engine works very well for asynchronous agents so that you don't need to specify things on your keyboard because they can fetch what they need. [SPEAKER_11] That's one of the main uses, I guess.

49:50

SPEAKER_10

[SPEAKER_11] So it plays very well, I think, with agents like Co-pilot on GitHub.

49:52

SPEAKER_07

[SPEAKER_11] Do you see, if you can share it, which agents are used most with Unblocked?

50:02

SPEAKER_07

[SPEAKER_11] Because on the wild, as developers with our laptops, I think Cloud Code is much more used than Co-pilot, but maybe you see a different picture.

50:13

[SPEAKER_06] Okay, so I'm going to take this off the screen for a sec and try to see if I can pull that up for you.

50:42

SPEAKER_06

[SPEAKER_06] But the answer is yes, we do know roughly what that breakdown looks like. So let me grab that.

50:53

SPEAKER_06

Okay. [SPEAKER_04] Okay. Okay. Okay. Okay. Okay. Okay.

51:26

SPEAKER_06

Okay. Okay. [SPEAKER_00] Okay.

51:52

SPEAKER_06

Okay.

51:53

SPEAKER_07

Okay. Okay. Okay. [SPEAKER_06] I think this gives you the rough picture. [SPEAKER_06] Okay.

52:00

SPEAKER_10

[SPEAKER_06] So this is the rough picture here. [SPEAKER_06] Unfortunately, because of the way this is, I should probably extend the screen. [SPEAKER_06] But I'll just step over here. [SPEAKER_06] So cloud code is by far the most used. [SPEAKER_06] This is the next one is cursor. [SPEAKER_06] So that seems fairly obvious. [SPEAKER_06] This last one here is a catch-all. [SPEAKER_06] But what's really interesting is that a lot of people use cloud desktop, which was very unexpected. [SPEAKER_06] But this is the case. [SPEAKER_06] So and then VS Code and Codex account for a much smaller component.

52:26

[SPEAKER_06] But yeah, it seems like everyone's using either cursor or cloud code. [SPEAKER_06] I would have expected more totally asynchronous agents, something that people would just run from a PR. [SPEAKER_11] Yeah. [SPEAKER_11] Okay, you can run cloud from a PR, but it's less common. [SPEAKER_11] Maybe sometimes you use copilot because it's built in. [SPEAKER_11] Yeah, actually, this one here, cloud code, may capture some of that traffic. [SPEAKER_06] So that's probably what you're seeing because people will wire up cloud code in CI and do things like that. [SPEAKER_06] Thanks. [SPEAKER_11] No problem. [SPEAKER_06] I've got a potentially dumb question.

52:28

[SPEAKER_05] There's no dumb questions. [SPEAKER_05] Well, we'll see. [SPEAKER_05] Actually, I had a teacher in grade three that used to tell me there are no dumb questions, only dumb people. [SPEAKER_06] Go on. [SPEAKER_06] I could be one of them. [SPEAKER_05] From your point of view, you can use sub-agents from an exploratory standpoint.

52:36

[SPEAKER_05] Yeah.

53:23

[SPEAKER_05] How does that plus memory plus storing snippets of information that might be able to...

53:33

[SPEAKER_05] I'm thinking of the social graph that you just showed, right? [SPEAKER_05] Yeah. [SPEAKER_06] Even in an organization that's several thousand people, you would be able to store that in a very small file? No.

53:43

SPEAKER_05

You would as the graph that you showed.

53:47

SPEAKER_10

[SPEAKER_05] Oh, I see. [SPEAKER_06] The social graph component, yes. [SPEAKER_06] Yeah. [SPEAKER_05] It can be compacted.

53:52

SPEAKER_07

[SPEAKER_05] I'm trying to understand how this compares.

53:57

SPEAKER_07

[SPEAKER_05] What's the USP compared to the exploratory agents and repeating that? [SPEAKER_05] I see what you're saying. [SPEAKER_06] Okay. [SPEAKER_06] So there are two components to that. [SPEAKER_06] One is that an exploratory agent would have to do this every time. [SPEAKER_06] So when it starts from ground zero, yes, it might be possible for it to reconstitute a social graph hierarchy. [SPEAKER_06] But it would have to do two things in order to do that. [SPEAKER_06] One is it would actually have to write code in order to constitute the graph. [SPEAKER_06] At least the way that agents are today or the way that the models are today.

54:15

SPEAKER_07

[SPEAKER_06] You wouldn't be able to just have it run basic tools around the organization and figure out the who's who. [SPEAKER_06] It would have to write what that social graph algorithm is, run it, and then get the distillation out the back end. [SPEAKER_06] So at that point, you're basically getting close to that component anyways. [SPEAKER_06] So you short circuit it and just run it and use it. [SPEAKER_06] Maybe I should explain some of the motivation for that thing, actually.

54:29

SPEAKER_10

[SPEAKER_06] I realize now that I may not have done that effectively. [SPEAKER_06] Social graph is not just about conveying information about who the experts are. [SPEAKER_06] It's used within the context engine as a pivot point into more important context. [SPEAKER_06] So understanding who the experts are in a particular code area acts as a jump point because another part of a context engine which happens at the ingestion and processing layer is distilling the—we call it bottling the expert.

54:35

[SPEAKER_06] But it's essentially distilling what that individual has worked on in the past, where they sit in the hierarchy of the organization, the decisions that they've made based on Slack conversations that they've had, based on their PR comments, all this stuff.

54:42

SPEAKER_06

[SPEAKER_06] When you distill that down, it's—and you pass it to the agent, then what happens is, let's say that I'm a new employee and I'm coming to work on a particular area of code. There are a bunch of different ways of loading context for that code. One is a semantic search via vector search, right?

54:53

SPEAKER_06

So that's layer one. Another layer is pre-built memories.

54:56

SPEAKER_11

[SPEAKER_06] And then the third layer is bottling—un-bottling the expert for that area of code. [SPEAKER_06] And getting that expert's learnings into context is a really powerful mechanism. [SPEAKER_06] It helps drive the rest of the retrieval in an agentic loop. [SPEAKER_06] And it helps the agent directionally, where to go next. [SPEAKER_06] Does that make sense? [SPEAKER_06] Right. I think everybody's in now. [SPEAKER_06] Awesome. [SPEAKER_07] So let's— [SPEAKER_07] Okay. [SPEAKER_07] So I think we're—if we're all in, then the next thing here is—once I get this back up on the screen. [SPEAKER_06] Another layer is pre-built memories.

55:41

SPEAKER_06

And then the third layer is bottling un-bottling the expert for that area of code.

55:46

SPEAKER_06

And getting that expert's learnings into context is a really powerful mechanism. It helps drive the rest of the retrieval in an agentic loop.

55:58

SPEAKER_06

And it helps the agent directionally where to go next.

55:59

SPEAKER_04

[SPEAKER_06] Does that make sense?

56:03

[SPEAKER_06] Right. I think everybody's in now. [SPEAKER_06] Awesome. [SPEAKER_07] So let's— [SPEAKER_07] Okay. [SPEAKER_07] So I think we're—if we're all in, then the next thing here is—once I get this back up on the screen. [SPEAKER_06] I'm still sending in lights. [SPEAKER_06] I saw someone just—so please keep coming and we can keep going as— [SPEAKER_07] Yep. [SPEAKER_07] So feel free to just fire this repo at your agent and get it to run it. [SPEAKER_06] If you literally just say to Cloud Code, run this against my repo, be sure to give it a time range or a PR limit. [SPEAKER_06] Otherwise, it'll go off the rails and take a really long time to finish.

56:14

SPEAKER_00

[SPEAKER_06] So just say, process the last 300 PRs or process up till September 2025 or something like that. [SPEAKER_06] There's enough information in the readme that it should be able to just do it and just run it against your repo. [SPEAKER_06] This morning I get cloned, said read the readme and make it happen, Cloud, and it did. [SPEAKER_06] Yep. [SPEAKER_07] Can I ask another question? [SPEAKER_06] What's your roadmap? If you have a roadmap, what's your plans for the coming year or something? For unblocked? [SPEAKER_09] You're in AI time. [SPEAKER_06] Sorry. [SPEAKER_06] I haven't seen you this long.

56:21

[SPEAKER_10] Is it about unblocked or about this sort of side project? [SPEAKER_10] I mean unblocked. [SPEAKER_06] Yeah. [SPEAKER_09] Yeah. [SPEAKER_06] So I mean, I've alluded to this before, but where the puck is going is with fully autonomous agents. [SPEAKER_09] So we're very focused on making sure that autonomous agent flows are highly optimized. [SPEAKER_06] You, as I was saying at the beginning of the conversation, you cannot run those things effectively without very finely tuned context. [SPEAKER_06] Yeah. [SPEAKER_06] So when you think about it, it's agent swarms. It's going to be input, output.

56:22

[SPEAKER_07] There will be an IOPS at some point on a complex and then shaping that token efficiency. [SPEAKER_07] I read things like tracing. What do agents, and you get runbooks out of those. [SPEAKER_09] Is that the path you're investing in or what is it? [SPEAKER_09] Retrieval, what? Are you talking specifically about incident management then or? [SPEAKER_09] Sorry? [SPEAKER_06] Are you speaking specifically about incident management and that sort of thing?

56:22

[SPEAKER_06] No, I'm speaking about your, I'm thinking actually more from a business perspective. How can we extract business knowledge that's really deeply embedded into systems that nobody knows anymore and some people think they know but they don't know? [SPEAKER_09] Yeah. [SPEAKER_09] And then documents, human knowledge, right? [SPEAKER_06] Tested knowledge. [SPEAKER_09] So I mean, there's two ways of surfacing that. Either at the product level or through the context engine itself.

56:01

[SPEAKER_06] And increasingly what we see is that people leverage agents to do their work even at that level. So they'll go to Cloud Code, they'll connect the unblocked context engine, it'll be like, do this thing for me and then the context engine will find all the things that it needs to do that task and then it'll surface that data. [SPEAKER_06] Yeah. [SPEAKER_06] For us that means the first near-term roadmap is API. [SPEAKER_04] Yes. [SPEAKER_07] It's CLI. [SPEAKER_07] CLI API. [SPEAKER_07] And points on API layer to back to this. [SPEAKER_06] Yeah. [SPEAKER_07] And some first class products that you saw via QA experiences in Slack or dashboard.

56:01

[SPEAKER_07] And then the code review is one of our first class apps as well. [SPEAKER_07] But when you use these in your org, it's just good to basically make sure that that access is available to obviously QA and agents. Most are. But it's over in API in some way. [SPEAKER_07] Cool. [SPEAKER_07] I'm gonna lift this off again. [SPEAKER_06] Hopefully people will start submitting some PRs and we can— [SPEAKER_06] Yeah. [SPEAKER_06] Rip some PRs. [SPEAKER_06] That's a good question. [SPEAKER_07] Test the agent.

56:39

SPEAKER_06

[SPEAKER_10] Once you're in that GitHub org, let me actually repost it in that Slack channel. Because that link will— [SPEAKER_07] So this org will stay up until the end of the week. At which point we'll basically bring it down and release this as open source.

56:55

SPEAKER_06

And everyone that contributes obviously is gonna get credited. So your name will be on it. Should we set up the repo locally and only then start doing work? What's the first thing we should do?

57:01

SPEAKER_06

[SPEAKER_02] So I just finished setting up. [SPEAKER_02] Yeah. [SPEAKER_02] Just clone the repo. The easiest thing to do is to take an agent like Claude and point it at— Just launch it from that repo. From that directory and just say, please bootstrap and launch this product and away it will go. So if you guys run into any kind of technical things, we're here obviously. Yeah. One of the questions— Backtracking of the—The—The—Let's—Hold on. Let's get you the mic. Oh, you got it. I've got a lapel now, so—

57:43

SPEAKER_11

[SPEAKER_06] Awesome. [SPEAKER_07] Cool. [SPEAKER_07] Yeah. [SPEAKER_06] Can you hear me? [SPEAKER_07] Yeah, perfect.

57:58

SPEAKER_06

[SPEAKER_04] So on the slide where you had the performance and you guys were 80% and without Unblocked it was 20%. [SPEAKER_04] Yeah. [SPEAKER_04] And now I see that while you are basically hooking up Unblocked to ClaudeCode. So in a way is it a fair comparison to say I will use vanilla ClaudeCode with access to the MCP and to the skills.

58:04

SPEAKER_11

[SPEAKER_04] Yeah.

58:20

SPEAKER_06

[SPEAKER_04] And then I will use ClaudeCode hooked with Unblocked with the same MCPs and the same skills.

58:22

SPEAKER_05

[SPEAKER_04] Yeah. [SPEAKER_04] And here you can do the performance comparison and here you still have a lot of alpha from I guess whatever you are cooking inside Unblocked. Was it the comparison that was done or was it done without—Was it done with a vanilla ClaudeCode but without context? [SPEAKER_04] No, it was done with MCP servers like GitHub and Slack wired up.

58:33

SPEAKER_06

[SPEAKER_04] Yeah. Yeah. Yeah.

58:39

SPEAKER_05

[SPEAKER_04] So in a way is it a fair comparison to say I will use vanilla ClaudeCode with access to the MCP and to the skills. [SPEAKER_04] Yeah. [SPEAKER_04] And then I will use ClaudeCode hooked with Unblocked with the same MCPs and the same skills. [SPEAKER_04] Yeah. [SPEAKER_04] And here you can do the performance comparison and here you still have a lot of alpha from whatever you are cooking inside Unblocked. [SPEAKER_04] Was it the comparison that was done or was it done without...

58:58

SPEAKER_06

[SPEAKER_04] Was it done with vanilla ClaudeCode but without context?

59:02

SPEAKER_05

[SPEAKER_04] No, it was done with MCP servers like GitHub and Slack wired up. [SPEAKER_04] Yeah. [SPEAKER_06] Yeah. [SPEAKER_06] Yeah.

59:09

SPEAKER_06

We basically got parity with all the MCP servers of every SaaS vendor in one. It was vanilla Claude all MCPs and the other one was Claude with Unblocked only and then do the task.

59:12

SPEAKER_05

[SPEAKER_07] And same context? [SPEAKER_07] Same context files? [SPEAKER_04] Same. Same prompt. [SPEAKER_07] And same access.

59:22

SPEAKER_06

[SPEAKER_07] Yeah.

59:25

SPEAKER_06

[SPEAKER_07] Yeah. [SPEAKER_07] It's pretty fun. [SPEAKER_07] Yeah. [SPEAKER_07] Oh, thank you. [SPEAKER_07] Yeah. [SPEAKER_07] Maybe two questions. [SPEAKER_07] So one is I see that a lot of these social graphs are built with the traditional network kind of calculation and the statistical aspects of networks. [SPEAKER_02] Is this the approach that you began with and it already worked the best or did you try something else? [SPEAKER_02] Because most of the memory systems they work more on filtering out episodic memory, something else, something else. [SPEAKER_02] And this is a really nice scoring system. [SPEAKER_02] Yeah. [SPEAKER_02] That's the first question.

1:00:15

SPEAKER_06

[SPEAKER_02] Is it also with the Unblocked? [SPEAKER_02] Yeah. [SPEAKER_02] The second question, you mentioned that it works with teams, Microsoft environment. [SPEAKER_02] I wonder what differences did you observe between building social graphs for different environments? [SPEAKER_02] Because on GitHub, I imagine it's very different than on SharePoint teams, and so on. [SPEAKER_02] Yeah. [SPEAKER_02] Is it also these network stats based or is it something different? [SPEAKER_02] So our first implementation was incredibly naive, right? [SPEAKER_02] It was just using the numbers of PR contributions and comparing that directly with the number of PRs reviewed by each person.

1:01:05

SPEAKER_06

So just a simple numbers game. That didn't produce accurate team clusters. So then we got onto the algorithms that you see here. Unblock does a little bit more than this. So this is a middle road. Another strategy that Unblock uses is experts by vector clusters. So when we ingest the source code and vectorize it, we understand who the most contributors are for that piece of source code. So when we look up individuals, we can see what they've been working on and what the clusters in proximity are and then relate people based on their cluster proximity. So that's more of an ML type approach.

1:01:47

SPEAKER_06

And then there's a final layer, which is an AI LLM heavy layer that does distillations of a whole bunch of different context elements. Things that people have worked on in the past, conversations that they've been having in Slack. And then when you take all that and you weigh it against the procedurally generated graph, you get a much more accurate distillation. This one here, you'll notice some people will get pulled into team clusters that you know are operating across many different teams, for example. And this won't account for that. Yeah.

1:02:05

SPEAKER_07

[SPEAKER_06] And then for the differences between different environments like Microsoft versus Slack, do you see, do you need different algorithms, different ways, let's say that you ascribe there? [SPEAKER_06] I mean, I don't want to take all the things. [SPEAKER_02] No, this algorithm is purely SEM based.

1:02:16

SPEAKER_06

So the algorithms for, you're right, Slack teams, they're quite a bit different because you don't have these review points. So then it becomes who's the most active in particular channels? And then you need a distillation or a summary of what that channel is about.

1:02:24

SPEAKER_07

[SPEAKER_06] And then you need to vectorize that. [SPEAKER_06] And then you need to score it against the most frequent contributors.

1:02:33

SPEAKER_06

But it's not enough. You have to relate that back to the SEM data in order to figure out who the real experts are. One of the problems that I personally experienced in some organizations I've worked at is that you get the noisy junior engineer, right? So they're very noisy. They love to talk. But the signal to noise ratio is not great. And just because someone's not saying a lot of things doesn't mean that their messages are not impactful.

1:03:05

SPEAKER_06

So part of this game is about assessing the impact of when people say certain things, you know, how does that relate to the data? The PRs that get spawned off as a consequence, how many of those PRs get merged, that sort of thing.

1:03:12

SPEAKER_07

[SPEAKER_06] Yeah.

1:03:14

Yeah. [SPEAKER_10] I just sent a message to Slack, but I don't think there's write access to the repo. Is there not? [SPEAKER_03] There should be.

1:03:22

SPEAKER_06

[SPEAKER_07] Okay. [SPEAKER_07] Check that.

1:03:25

SPEAKER_10

[SPEAKER_07] Well, you should be able to open a pull request. [SPEAKER_07] You can't push to main.

1:03:34

SPEAKER_06

[SPEAKER_07] Okay.

1:03:34

SPEAKER_09

[SPEAKER_07] So if that's the situation, I mean, we'll check.

1:03:36

SPEAKER_06

Yeah.

1:03:42

SPEAKER_09

[SPEAKER_07] You should be able to create a branch.

1:03:44

Oh, he can't fork the repo either. [SPEAKER_07] Oh, yeah. [SPEAKER_07] Forks might be disabled. This will be open source at the end of the week. And your contributions will be on it. [SPEAKER_07] What's really fun is using that social graph tool later against your own repo and showing your team. [SPEAKER_07] Yeah. Oh, sorry. [SPEAKER_10] Oh, I'll do it. I like that Unblock tried to answer you for that question.

1:04:21

SPEAKER_09

Oh, you see that? [SPEAKER_07] The Slack auto response. [SPEAKER_07] Yeah. [SPEAKER_07] Sorry to speak. [SPEAKER_10] Are you in here?

1:04:32

SPEAKER_06

[SPEAKER_10] That's a camera. [SPEAKER_07] Sorry. [SPEAKER_10] Sorry.

1:04:45

SPEAKER_09

[SPEAKER_10] Sorry. [SPEAKER_10] Oh, it's okay. [SPEAKER_07] I was just thinking.

1:04:52

SPEAKER_06

[SPEAKER_07] How are you.

1:04:53

SPEAKER_09

[SPEAKER_07] How are you. [SPEAKER_07] How are you.

1:05:01

SPEAKER_06

[SPEAKER_07] How are you. [SPEAKER_07] How are you. [SPEAKER_07] How are you. [SPEAKER_07] How are you. [SPEAKER_07] Oh, okay, let me check to see. That should not be the case.

1:05:26

SPEAKER_04

[SPEAKER_06] Okay.

1:05:27

SPEAKER_07

[SPEAKER_06] Let me know if you still need to get an invite. Just check the members. I think there might be an issue here.

1:05:31

SPEAKER_06

Just a second.

1:05:37

SPEAKER_07

[SPEAKER_06] Okay, okay, okay, okay. [SPEAKER_06] Oh, these were direct assignments. [SPEAKER_10] Sorry. [SPEAKER_10] Sorry. [SPEAKER_10] Oh, it's okay. I was just thinking. How are you? [SPEAKER_06] Okay. [SPEAKER_06] Let me know if you still need to get an invite.

1:06:00

SPEAKER_06

[SPEAKER_07] Just check the members. [SPEAKER_07] I think there might be an issue here. Just a second. Oh, these were direct assignments.

1:06:06

SPEAKER_07

[SPEAKER_06] So I think we have to pull people into the whole project.

1:06:09

SPEAKER_10

[SPEAKER_06] Right.

1:06:14

SPEAKER_07

[SPEAKER_06] Because they're not org-assign. [SPEAKER_06] So GitHub, I love you. Zero nines of uptime. Yeah.

1:06:28

SPEAKER_06

[SPEAKER_07] We'll fix this one here. [SPEAKER_10] Yeah.

1:06:49

SPEAKER_06

[SPEAKER_10] Slam everybody in. Unblocked, you all have right access.

1:06:56

SPEAKER_02

It is the name of the company. [SPEAKER_07] Just validate that for us if you would. [SPEAKER_07] Yeah, please. Let me know.

1:07:00

SPEAKER_06

Perfect. [SPEAKER_07] All right. [SPEAKER_07] We're getting real PRs now. There we go. Nice. [SPEAKER_07] Hell yeah. [SPEAKER_07] Now let's do fun things. [SPEAKER_07] Looks good to me. [SPEAKER_07] I think we have our first approved PR. [SPEAKER_10] I'm just sending ridiculous chats to Unblocked so you can see it try to answer questions in Slack as PRs come up. [SPEAKER_06] I'm going to see what it says about this. [SPEAKER_06] It's like, oh, let me think about it. [SPEAKER_06] Oh, did you ask it about the PR? [SPEAKER_07] Yeah, but the PR I think you accepted so we'll see what happens. [SPEAKER_07] I mean, it did approve it so Unblocked was down.

1:07:28

SPEAKER_06

[SPEAKER_07] Unblocked was down. [SPEAKER_07] Only visible to you.

1:07:30

SPEAKER_07

[SPEAKER_06] Oh, no. Oh, it's such a good answer though.

1:07:32

SPEAKER_06

Nice PR.

1:07:33

SPEAKER_07

[SPEAKER_06] Good job, Unblocked.

1:07:34

SPEAKER_04

[SPEAKER_06] Great answer. Yep. [SPEAKER_07] Sorry. [SPEAKER_07] I just joined in later. [SPEAKER_07] Yeah. [SPEAKER_06] I'll put it back up. [SPEAKER_10] One sec. [SPEAKER_10] Where did it go actually? [SPEAKER_06] I lost the... [SPEAKER_06] So this thing that I showed before, it is the project that exists in that repo. [SPEAKER_06] Oh, so the idea is think about features that you want to add or things that you want to fix, like new components, and then just hack away at it and submit a PR. [SPEAKER_06] Sorry, I didn't get there.

1:08:27

SPEAKER_06

Yeah, all good. Do you want to open up a terminal session and show the MCP? Oh, sure, yeah. [SPEAKER_11] Because people can obviously use it, but they don't have all our source.

1:08:40

SPEAKER_07

Yeah. Explain where I'm coming from.

1:08:42

[SPEAKER_07] Yeah. [SPEAKER_07] I'm a consultant and I really like the idea, but I wanted to try it and propose it to a client. [SPEAKER_07] I cannot show the context you have or maybe get an idea of how it works, the internal answer. [SPEAKER_11] Because I really like the idea. [SPEAKER_11] This is a big problem. [SPEAKER_11] Yeah. [SPEAKER_11] And without seeing the product, it's hard. [SPEAKER_11] Well, I mean, one thing that you could do if you're visiting clients is you can ask them if they run the tool on their repo, and then it will generate this result for them so they can see on their own project what the value is, right?

1:08:48

SPEAKER_07

[SPEAKER_11] I think Peter, I think he's just asking about our product, specifically not those. [SPEAKER_11] Oh, Unblocked. [SPEAKER_06] Yeah. [SPEAKER_06] You're asking about Unblocked.

1:09:04

SPEAKER_02

[SPEAKER_06] My bad, man. [SPEAKER_07] We're driving this way. [SPEAKER_06] Yeah. [SPEAKER_07] Sorry. [SPEAKER_07] Single track mind. [SPEAKER_07] OK. [SPEAKER_06] So your question is, how can you demonstrate the value of Unblocked to customers? [SPEAKER_06] See the value. [SPEAKER_06] Exactly. [SPEAKER_06] Yeah. [SPEAKER_06] I guess that one big question for customers have is how to use this information and enable, for example, if there's a conflict. [SPEAKER_10] Tell me. [SPEAKER_06] Maybe they have some conflict information. [SPEAKER_10] Yeah.

1:09:59

SPEAKER_06

[SPEAKER_10] You can make conflicts emerge in your app. [SPEAKER_10] But then there is the compliance layer, which is very interesting for corporate clients. [SPEAKER_10] I was thinking how this is translated to a UX, because many people are not, I understand it's mainly for coding. [SPEAKER_11] Yeah. [SPEAKER_11] And whether this is for technical people or maybe people overseeing some engineers or the engineer itself, I mean, just see how your platform works. [SPEAKER_11] But if it is, it's out of context, I mean, it's okay. [SPEAKER_11] No, no. [SPEAKER_11] That's totally fine. [SPEAKER_11] So this dashboard is a sort of front-end customer interface to the product.

1:10:45

SPEAKER_06

[SPEAKER_11] So you come in here and you can ask any question about your code base or your organization and get an answer for it here. This is right now attached to... Sorry. I lost my cursor. This is attached to this test org that we have, but I could use it against unblocked, and I could say, you know... [SPEAKER_11] But if it is, it's out of context. It's OK. [SPEAKER_11] I... [SPEAKER_11] No, no.

1:11:44

[SPEAKER_11] That's totally fine. [SPEAKER_11] So this dashboard is the front-end customer interface to the product. [SPEAKER_11] So you come in here and you can ask any question about your code base or your organization and get an answer for it here. This is right now attached to... Sorry. I lost my cursor. This is attached to this test org that we have, but I could use it against unblocked, and I could say... I have a little hot thing here that I can show. Yeah. Yeah. Yeah. Oops. Yeah. So the source mark engine is an internal component that we use to track source code changes through time, including where changes move between files and so on.

1:12:48

SPEAKER_06

So as a demonstration, you can show off... I mean, you can book your customers into a demo with us and we can demonstrate this, or you can wire it up to your own organization and demonstrate this flow to customers and try to find use cases where data sources conflict and demonstrate that.

1:13:00

SPEAKER_06

The challenge with context engines is that it's really hard to demonstrate the value to someone without actually wiring it up.

1:13:00

[SPEAKER_06] Okay. [SPEAKER_06] So people have to connect it to all of their integrations. [SPEAKER_06] Now, the good thing is unblocked has a free enterprise trial period, so people can try out the products in its fullest form before paying for it. [SPEAKER_06] Yeah.

1:13:08

SPEAKER_07

[SPEAKER_06] Sure. [SPEAKER_06] So if some of that information is incorrect, you can just reply in the chatbot or flag it in the references? [SPEAKER_06] Yes, exactly. [SPEAKER_11] Yeah. [SPEAKER_06] Yeah. [SPEAKER_11] So you can just... [SPEAKER_11] You can reply here or you can say not helpful and explain why and then it will distill it for the next round. [SPEAKER_06] So you could adjust some weights or confidence scores internally?

1:13:18

SPEAKER_06

Well, internally, what it does is it constructs task memory. So it looks for those repeated signals and this is actually where the expert's graph comes in.

1:13:22

SPEAKER_07

[SPEAKER_06] It's used a lot. [SPEAKER_11] The expert's graph provides weight.

1:13:26

SPEAKER_06

So when an expert comes in and says that's not correct, it's going to get some more weight and distill a memory for it. If it's just a new engineer that says that's not right, then that's not really a trustworthy source yet.

1:13:29

SPEAKER_07

[SPEAKER_06] So you have to have a trustworthy source to base that on.

1:13:39

SPEAKER_07

[SPEAKER_06] Does that make sense? [SPEAKER_06] Yeah, it makes a lot of sense. [SPEAKER_06] It's a social network. [SPEAKER_06] Exactly.

1:13:47

SPEAKER_10

[SPEAKER_06] Somehow.

1:13:48

SPEAKER_07

[SPEAKER_06] Yeah.

1:14:12

SPEAKER_07

[SPEAKER_11] Yeah.

1:14:25

SPEAKER_07

[SPEAKER_11] Thanks. [SPEAKER_06] No problem. [SPEAKER_06] Sweet. [SPEAKER_06] Cool.

1:14:34

SPEAKER_10

[SPEAKER_06] What is your memory? [SPEAKER_06] Is that as files?

1:14:39

SPEAKER_07

[SPEAKER_06] Oh, under the hood?

1:14:39

SPEAKER_10

[SPEAKER_06] Yeah. Well, when it's presented to the AI, it's presented as files. But under the hood, we store it in database tables and such.

1:14:42

SPEAKER_07

The memories are constituted from a bunch of different sources. [SPEAKER_06] So they're not just flat file based. [SPEAKER_06] The whole memory construct will be hydrated at runtime. [SPEAKER_06] So, yeah.

1:14:53

SPEAKER_07

[SPEAKER_06] And will you just give your data tools to vary your database based on whatever criteria the users are? [SPEAKER_06] Yeah. [SPEAKER_06] Well, so yes, there are a bunch of tools for data retrieval. [SPEAKER_06] For memories specifically, you can't really leave it up to the agent to do memory hydration because that's part of the seed context. [SPEAKER_07] In order to get the agent to go in the right direction, you have to seed it with the appropriate data. [SPEAKER_06] And experts context is a good jump off point for the agent. [SPEAKER_06] So, yeah. [SPEAKER_06] Thank you. [SPEAKER_06] Yep. [SPEAKER_06] Thank you.

1:15:04

[SPEAKER_06] Is there any official benchmark that tracks the type of value you try to bring? [SPEAKER_06] Like- Yeah. [SPEAKER_09] Yeah, because it feels like it's not exactly coding. [SPEAKER_06] Well, it is. [SPEAKER_04] Yeah. [SPEAKER_04] But I'm curious if there's any public things that you're tracking yourself against. [SPEAKER_04] So we do have some internal benchmarks. [SPEAKER_04] You're right. [SPEAKER_04] It's a little bit squishy. [SPEAKER_04] So, have you heard Boris Cherney talk at Cloud Code? [SPEAKER_04] He's the creator of Cloud Code. [SPEAKER_06] The creator of Cloud Code. [SPEAKER_06] Yeah.

1:15:04

[SPEAKER_06] So he did this interview where they were talking about how they measure success for Cloud Code internally. [SPEAKER_07] This may have changed because there's a lot of benchmarks now that they have. [SPEAKER_06] They have the shit talk benchmark. [SPEAKER_06] You guys have probably seen that one. [SPEAKER_06] But what that really distills down to is vibes. [SPEAKER_06] And so the most important thing in systems like this is to capture sentiment. [SPEAKER_06] And so if your sentiment is trending upwards, then that's a good thing. [SPEAKER_06] Our sentiment right now is on a scale of minus 100 to 100, somewhere around 60 scores.

1:15:04

[SPEAKER_06] So on a normalized scale, that's 0.75 to 0.8. [SPEAKER_06] So the vibe would be captured by something like maybe less back and forth on the PRs or maybe you having less back and forth with Cloud to get your stuff done? [SPEAKER_06] Yeah.

1:15:16

[SPEAKER_06] So the vibes are, are people satisfied, right? [SPEAKER_06] So satisfaction can come from a lot of different sources and dissatisfaction can come from a lot of different sources.

1:15:29

SPEAKER_06

[SPEAKER_04] So the way to think about that is that it encodes all of those things. [SPEAKER_04] But you can capture specific metrics, and we do, how long things take.

1:15:35

SPEAKER_07

[SPEAKER_04] And we're actually currently working really hard to bring the response times down because even though agents are... [SPEAKER_06] Here's the interesting thing.

1:15:39

SPEAKER_06

As we move towards a more autonomous universe, response times for MCP servers are actually less and less important. The more important thing is that they get the answer absolutely bang on.

1:16:02

SPEAKER_06

Yeah. And the reason is because the amount of time that a context engine spends collecting all that information and distilling it is a microcosm of what the full task takes to implement and to traverse. So if you can spend a little bit more time and cut the implementation down by 60, 70, 80%, that's a huge win. All right. And...

1:16:16

SPEAKER_06

[SPEAKER_04] And we're actually currently working really hard to bring the response times down because even though agents are...

1:16:19

SPEAKER_07

[SPEAKER_06] Here's the interesting thing. [SPEAKER_06] As we move towards a more autonomous universe, response times for MCP servers are actually less and less important.

1:16:21

SPEAKER_10

[SPEAKER_06] The more important thing is that they get the answer absolutely bang on. [SPEAKER_06] Yeah.

1:16:22

SPEAKER_07

[SPEAKER_06] And the reason is because the amount of time that a context engine spends collecting all that information and distilling it is a microcosm of what the full task takes to implement and to traverse. [SPEAKER_06] So if you can spend a little bit more time and cut the implementation down by 60, 70, 80%, that's a huge win. [SPEAKER_06] All right. And... [SPEAKER_06] Go ahead. [SPEAKER_06] Yeah. Sorry. [SPEAKER_06] Very small follow-up. [SPEAKER_06] Actually, I'm curious. [SPEAKER_06] Do you have any rough numbers on how much time does it spend retrieving context versus executing the task? [SPEAKER_06] To your point, is it 10% right now, 90%?

1:16:22

[SPEAKER_04] Or is it... [SPEAKER_04] I have no idea.

1:16:33

[SPEAKER_04] I mean, I have my own experience. [SPEAKER_04] It's like... [SPEAKER_04] Yeah. [SPEAKER_04] Agent context collection is probably close to that number. [SPEAKER_04] It's about 90%. [SPEAKER_04] The actual code writing part is really, really fast. [SPEAKER_04] If you can even just watch what an agent's doing. [SPEAKER_06] When it writes the code that... [SPEAKER_06] Output tokens are, by the way, the thing that drags down the performance. [SPEAKER_06] Everyone used to think it was input tokens. [SPEAKER_06] We've run tons of experiments with this. [SPEAKER_06] You can bring the input token size up. [SPEAKER_06] And time to first output token now is pretty good.

1:16:33

[SPEAKER_06] It's pretty highly optimized. [SPEAKER_06] The thing that really impacts performance is output tokens. [SPEAKER_06] So you have to be judicious with the way that you collect and supply context back to the agent, so that it remains tight on its output loops as well. [SPEAKER_06] For one benchmark that Peter mentioned in the talk, we gave an ambitious task, because obviously it's prompt-dependent how much time you're adding in with a context engine.

1:16:33

[SPEAKER_06] But the ambitious task we gave was to implement the new adaptive thinking mode in Anthropix toolchain when they introduced that, which, as mentioned, it went from a 25-minute wall clock time with unblock with a context engine. [SPEAKER_06] The other case without was two and a half hours. [SPEAKER_07] It was two hours and 25 minutes. [SPEAKER_07] But the main reason for that was we gave it all the data, we ran the prompt, and then its first output was totally wrong. [SPEAKER_07] So you had to... the human had to loop again and be like, no, no, this, this, this. [SPEAKER_07] And the next output was wrong and the next output.

1:16:33

[SPEAKER_07] So once you do four loops, you have a two and a half hour wall clock time, versus obviously the 25 minutes when it did not need that, when there's no corrections required. [SPEAKER_07] Yeah. [SPEAKER_07] So as mentioned, it's think of it as a waterfall. [SPEAKER_07] The more high-quality, correct, high-signal context you have up front, the better every single thing the agent's gonna do until it says it's done, whether it got it right or not. [SPEAKER_07] Yeah. [SPEAKER_07] Yeah? [SPEAKER_07] Man's got it. [SPEAKER_07] Man, you also mentioned that the token usage on tool calls and information search really decreased.

1:16:33

[SPEAKER_07] So I know that a lot of these tools that provide or aggregators for tool use, they have insane token usage. [SPEAKER_07] So maybe have some estimations on how... [SPEAKER_02] like let's say I need a Slack conversation, some summary from one conversation to another, or how people interact there would be like 60K tokens on Composio. [SPEAKER_02] I wonder how many tokens it would be using unblocked. [SPEAKER_02] Yeah. [SPEAKER_02] Lower. [SPEAKER_02] We're still very vibes there.

1:16:41

[SPEAKER_02] It's hard to get real data from other customers or people in the market. [SPEAKER_07] But again, with that same, I'm going to keep talking to the same task as ZZ.

1:16:44

[SPEAKER_07] That one went from 21 million token total usage to 10 million token with the context engine. [SPEAKER_07] So a part of that though is because you didn't have to doom loop.

1:16:50

[SPEAKER_07] So when, of course, that increased a lot of the tokens expense.

1:17:21

[SPEAKER_07] So we did drop it by 50% on a large task. [SPEAKER_07] Again, obviously if you're like, yo, center a div, you're not going to get a lot of game. [SPEAKER_07] It's probably in the training data. But yeah, any feature fix, so a lot of what people are putting through unblocked or what an engineer is doing every day, it's very rare that you're doing a task that's so minor that... I mean, then again, I've asked Claude to do git push.

1:17:44

SPEAKER_06

[SPEAKER_07] So I'm not the only one I bet.

1:17:46

SPEAKER_07

I was like, you do it.

1:18:00

SPEAKER_07

It's like, why did that cost me 30 cents? I don't know.

1:18:10

SPEAKER_07

Just run it in high place. I did all the effort to put my GBG keys in the right place.

1:18:20

SPEAKER_07

So I'm like Claude.

1:18:45

SPEAKER_07

Go. Any more questions while y'all ship? Any confusion, anything I can unblock for you? It's my purpose in life.

1:18:55

[SPEAKER_07] Sorry, you may have answered this question already.

1:19:09

SPEAKER_10

[SPEAKER_07] But are you using knowledge base rag on unblocked?

1:19:19

SPEAKER_06

[SPEAKER_07] Or what exactly is the tech that you are surfacing? [SPEAKER_07] Oh, so many things. [SPEAKER_07] I can come talk to you at the side. [SPEAKER_00] I'll take my mic off.

1:19:25

SPEAKER_07

[SPEAKER_00] I'm just going to answer that question. [SPEAKER_00] Sure. That was just this.

1:19:37

SPEAKER_07

Sorry.

1:19:38

SPEAKER_06

[SPEAKER_07] What's the churn rate on the data?

1:19:41

SPEAKER_07

Oh, it's real time, basically.

1:19:47

SPEAKER_06

[SPEAKER_07] So I guess there's two parts to that question. [SPEAKER_07] One is how frequently unblocked updates the data on the backend. [SPEAKER_07] So it's real time for many of the integrations and then on a cron job for others because for those particular integrations, they don't have webhooks, basically.

1:20:00

SPEAKER_04

[SPEAKER_06] Yeah, that's a question.

1:20:01

SPEAKER_07

[SPEAKER_06] Yeah. [SPEAKER_06] But the disk, so that means that rebuilding the graph data has to happen on a very frequent basis. [SPEAKER_00] Sure. That was just this.

1:20:05

[SPEAKER_07] Sorry. [SPEAKER_07] What's the churn rate on the data? [SPEAKER_07] Oh, it's real time. [SPEAKER_07] So I guess there's two parts to that question.

1:20:10

SPEAKER_06

[SPEAKER_07] One is how frequently unblocked updates the data on the backend. [SPEAKER_07] So it's real time for many of the integrations and then on a cron job for others because for those particular integrations, they don't have web hooks. Yeah, that's a question.

1:20:23

SPEAKER_06

Yeah. But the disk, so that means that rebuilding the graph data has to happen on a very frequent basis. Yeah. [SPEAKER_06] Yeah. [SPEAKER_06] Yeah. [SPEAKER_10] Yeah. [SPEAKER_10] Yeah. [SPEAKER_10] Yeah. [SPEAKER_10] Yeah. [SPEAKER_10] No, it's incremental. [SPEAKER_10] So our social graph build algorithm has an incremental component to it. [SPEAKER_06] So we don't have to rerun the whole thing. [SPEAKER_10] But also social graphs are less sensitive to frequent changes in data because it's unlikely that a single change is going to make a huge impact on the experts graph unless your organization is brand new. [SPEAKER_06] So for... [SPEAKER_06] Yeah.

1:20:42

[SPEAKER_06] Yeah. [SPEAKER_06] Yeah. [SPEAKER_06] Yeah. [SPEAKER_10] So as an example, we do best practices distillation on a much lower cadence, like week by week, because it just doesn't change that much. [SPEAKER_10] Yeah. [SPEAKER_10] Yeah. [SPEAKER_06] Yeah. [SPEAKER_06] Yeah. [SPEAKER_06] Well, the... [SPEAKER_06] Oh, yeah.

1:20:59

SPEAKER_07

[SPEAKER_06] Repeat your question.

1:21:00

SPEAKER_10

[SPEAKER_06] That's a good question. [SPEAKER_06] So I want to make sure we get that one done. [SPEAKER_06] Yeah.

1:21:02

[SPEAKER_06] It's attention. [SPEAKER_06] Oh. [SPEAKER_06] And in terms of customer privacy data retention... [SPEAKER_06] Yeah. [SPEAKER_10] From my point of view, I'm thinking of enterprise SaaS or even on-prem type deployments, which I'm not suggesting that you... [SPEAKER_05] I'm just thinking of that customer modality. [SPEAKER_05] Yeah. [SPEAKER_05] Yeah. [SPEAKER_05] Do you get pushback? [SPEAKER_05] How do they feel about you holding data? [SPEAKER_05] It's another processor in the loop. [SPEAKER_10] Well, so the privacy discussions happen at the organizational level. [SPEAKER_05] We don't actually run into a lot of friction.

1:21:06

SPEAKER_06

[SPEAKER_05] There are definitely environments like in government and at banks that have super sensitive needs. [SPEAKER_06] And so for those needs, we have an on-prem solution. [SPEAKER_06] But it's definitely not the path that I would recommend. [SPEAKER_06] Staying cloud-based. [SPEAKER_06] We have very large enterprise organizations that are entirely cloud-based, fully cloud-based. [SPEAKER_06] The secret sauce is now less encoded in source code and more encoded in the reasoning. So organizations tend to be a little bit more sensitive around things like Slack data, for instance.

1:21:29

SPEAKER_06

But the way that we store data, we have a whole white paper about how we protect customer data. And it's never been a problem. Yeah. Yeah.

1:21:36

SPEAKER_11

[SPEAKER_06] Pardon me? [SPEAKER_06] I missed that wasn't recommended.

1:21:40

SPEAKER_07

[SPEAKER_06] Oh, why it's not recommended? [SPEAKER_03] Well, the cloud-based integrations get updated more frequently. [SPEAKER_03] And so there's software patches. [SPEAKER_03] It's a little bit harder to maintain within an organization. [SPEAKER_06] There's one customer, a bank, where administering the platform becomes quite difficult because they have network isolation. [SPEAKER_06] And so now one of us has to sit within that network and administer the platform or we have to train individuals within the company to administer the platform.

1:21:50

SPEAKER_11

[SPEAKER_06] So it's just more of a maintenance and hand-holding exercise. [SPEAKER_06] But yeah. [SPEAKER_06] Yeah. [SPEAKER_06] Exactly. [SPEAKER_06] That's exactly right. [SPEAKER_06] Yeah. [SPEAKER_06] Thank you so much for the talk. [SPEAKER_10] Yeah. [SPEAKER_06] Thank you. [SPEAKER_06] Thanks for coming.

1:22:26

SPEAKER_06

Thanks. [SPEAKER_09] Thanks.

1:22:30

SPEAKER_07

[SPEAKER_06] Thanks.

1:22:31

SPEAKER_06

Yeah.

1:22:31

SPEAKER_07

Sorry. Sorry. Single track mind. OK.

1:22:35

SPEAKER_06

So your question is, how can you demonstrate the value of Unblocked to customers? See the value. Exactly. Yeah. I guess that one big question for customers have is how to use this information and enable, for example, if there's a conflict .

1:22:50

SPEAKER_10

Tell me.

1:22:52

Maybe . They have some conflict information. Yeah. And I understood that you can make this. Sorry. Yeah. You can make conflicts emerge in your app. But then there is the compliance layer, which is very interesting for corporate clients. Yeah. I was thinking how this is translated to a UX, because many people are not, I understand it's mainly for coding.

1:23:20

SPEAKER_11

Yeah. And whether this is for technical people or maybe people overseeing some engineers or the engineer itself, I mean, just see how your platform works. But if it is, it's out of context, I mean, it's OK. I... No, no. That's totally fine. So this dashboard is kind of like the sort of front-end customer interface to the product. So you come in here and you can ask any question about your code base or your organization and get an answer for it here.

1:23:56

SPEAKER_06

This is right now attached to... Sorry. I lost my cursor. This is attached to this test org that we have, but I could use it against unblocked, and I could say, like, you know... I have a little hot thing here that I can show. Yeah. Yeah. Yeah. Oops. Yeah. So the source mark engine is an internal component that we use to track source code changes through time, including, like, where changes move between files and so on. So as a demonstration, you know, you can show off... I mean, you can book your customers into a demo with us and we can demonstrate this, or you can wire it up to your own organization and demonstrate this flow to customers and

1:24:46

SPEAKER_06

try to find, you know, use cases where data sources conflict and demonstrate that. That the challenge with context engines is that it's really hard to demonstrate the value to someone without actually wiring it up.

1:25:01

SPEAKER_06

Okay. So, you know, where people have to connect it to all of their integrations. Now, the good thing is, unblocked has a free enterprise trial period, so people can try out the products in its fullest form before paying for it. Yeah. Sure. So if some of that information is incorrect, you can just reply in the chatbot or flag it in the references? Yes, exactly.

1:25:25

SPEAKER_11

Yeah.

1:25:27

SPEAKER_06

Yeah.

1:25:29

SPEAKER_11

So you can just... You can reply here or you can say not helpful and explain why and then it will distill it

1:25:36

SPEAKER_06

for the next round. So you could adjust some weights or confidence scores internally? Well, internally, what it does is it constructs task memory. So it looks for those kind of repeated signals and this is actually where the expert's graph comes in. It's used a lot.

1:25:55

SPEAKER_11

The expert's graph provides, like, weight.

1:25:58

SPEAKER_06

So when an expert comes in and says that's not correct, it's gonna get some more weight and distill a memory for it. If it's just a new engineer that says that's not right, then that's not really a trustworthy source yet. So you have to have a trustworthy source to base that on. Does that make sense? Yeah, it makes a lot of sense. It's like social network. Exactly. Somehow. Yeah.

1:26:25

SPEAKER_11

Yeah. Thanks.

1:26:26

SPEAKER_06

No problem.

1:26:29

SPEAKER_06

Sweet. Cool.

1:26:38

SPEAKER_06

What is your memory? Is that as files? Oh, under the hood? Yeah. Well, when it's presented to the AI, it's presented as files. But under the hood, we store it in, you know, database tables and stuff. Like the memories are constituted from a bunch of different sources. So they're not just, like, flat file based. You know, they'll be- the whole memory construct will be hydrated at runtime. So, yeah. And will you just give your data tools to, like, vary your database based on whatever criteria the users are? Yeah. Well, for- so, yes, there are a bunch of tools for data retrieval. For memories specifically, you can't really leave it up to the agent to do memory hydration

1:27:24

SPEAKER_07

because that's kind of, like, part of the seed context. In order to get the agent to go in the right direction, you have to seed it with the appropriate

1:27:34

data. And experts context is a good jump off point for the agent. So, yeah. Thank you. Yep. Thank you. Is there any official benchmark that kind of track the type of value you try to bring? Like- Yeah. Yeah, because it feel like it's not exactly coding.

1:27:55

SPEAKER_06

Well, it is.

1:27:56

SPEAKER_04

Yeah. But, yeah, I'm curious if there's any public things that you're tracking yourself against. So we do have some internal benchmarks. You're right. It's a little bit squishy. So, and Throt- have you heard Boris Cherney talk at Cloud Code? He's, like, the creator of Cloud Code?

1:28:13

SPEAKER_06

The creator of Cloud Code. Yeah. So he did this interview where they were talking about, like, how they measure success for Cloud Code internally.

1:28:23

SPEAKER_07

This may have changed because there's a lot of benchmarks now that they have.

1:28:26

SPEAKER_06

Like, they have, like, the shit talk benchmark. You guys have probably seen that one. But what that really distills down to is vibes. And so the most important thing in systems like this is to capture sentiment. And so if your sentiment is trending upwards, then that's a good thing. Our sentiment right now is on a scale of minus 100 to 100, somewhere around 60 scores. So on a normalized scale, that's, like, 0.75 to 0.8. So the vibe would be captured by something like maybe less back and forth on the PRs or maybe, I don't know, you having less back and forth with Cloud to get your stuff done? Yeah. So the vibes are, like, they're people satisfied, right?

1:29:18

SPEAKER_06

So satisfaction can come from a lot of different sources and dissatisfaction can come from a lot of different sources.

1:29:23

SPEAKER_04

So the way to think about that is that it encodes all of those things. But you can capture specific metrics, and we do, how long things take. And we're actually currently working really hard to bring the response times down because, you know, even though agents are...

1:29:43

SPEAKER_06

Here's the interesting thing. As we move towards a more autonomous universe, response times for MCP servers are actually less and less important. The more important thing is that they get the answer absolutely bang on. Yeah. And the reason is because the amount of time that a context engine spends collecting all that information and distilling it is a microcosm of what the full task takes to implement and to traverse. So if you can spend a little bit more time and cut the implementation down by, like, 60, 70, 80%, that's a huge win. All right. And... Go ahead. Yeah. Sorry. Very small follow-up. Actually, I'm curious.

1:30:27

SPEAKER_06

Do you have any rough numbers on how much time does it spend retrieving context versus executing the task? To your point, like, is it 10% right now, 90%?

1:30:39

SPEAKER_04

Or is it... I have no idea. I mean, I have my own experience. It's like... Yeah. Agent context collection is probably close to that number. It's like 90%. The actual code writing part is really, really fast. If you can even just watch what an agent's doing.

1:30:57

SPEAKER_06

When it writes the code that... Output tokens are, by the way, the thing that drags down the performance. Everyone used to think it was input tokens. We've run tons of experience with this. You can bring the input token size up. And, you know, time to first output token now is pretty good. Like, it's pretty highly optimized. The thing that really impacts performance is output tokens. So, you have to be, like, judicious with the way that you collect and supply context back to the agent, so that it remains tight on its output loops as well. For one benchmark that Peter mentioned in the talk, we gave an ambitious task,

1:31:40

SPEAKER_06

because obviously it's prompt-dependent, how much time you're adding in, like, with a context engine. But the ambitious task we gave was to implement the new adaptive thinking mode in Anthropix toolchain when they introduced that, which, as mentioned, it went from a 25-minute wall clock time with unblock with a context engine. The other case without was two and a half hours.

1:32:01

SPEAKER_07

It was two hours and 25 minutes. But the main reason for that was we gave it all the data, we ran the prompt, and then its first output was, like, totally wrong. So, you had to... the human had to loop again and be like, no, no, this, this, this. And the next output was wrong and the next output. So, once you do four loops, you have, like, a two and a half hour wall clock time, versus, obviously, the 25-minute when it did not need that, when there's no corrections required. Yeah. So, as mentioned, it's... think of it as a waterfall. The more high-quality, correct, like, high-signal context you have up front,

1:32:33

SPEAKER_07

the better every single thing the agent's gonna do until it says it's done, whether it got it right or not. Yeah. Yeah? Man's got it. Man, you also mentioned that, uh, the, uh, token usage on tool calls, and, like, just information, search, really, decreased. So, I know that, that, a lot of these tools, that, provides, uh, or, aggregators for tool use, they have insane, like, uh, token usage. So maybe have some estimations on how,

1:33:06

SPEAKER_02

like let's say I need a Slack conversation, some summary from one conversation to another, or how people interact there would be like 60K tokens on Composio. I wonder how many tokens it would be using unblocked. Yeah. Lower. We're still very vibes there. It's hard to get real data from other customers or people in the market.

1:33:33

SPEAKER_07

But again, with that same, I'm gonna keep talking to the same task as ZZ. That one went from 21 million token total usage to 10 million token with the context engine. So a part of that though is because you didn't have to doom loop. So when, of course, like that increased a lot of the tokens expense, like so we did drop it by 50% on a large task. Again, obviously if you're like, yo, center a div, you're not gonna get a lot of game. It's like probably in the training data. But yeah, like any feature fix, like so a lot of like, again, a lot of what people are putting through unblocked or what an engineer is doing every day,

1:34:07

SPEAKER_07

it's very rare that you're doing a task that's like so, I don't know, minor that like, I mean, then again, I've asked, I've asked Claude to do git push. So I'm not the only one I bet. I was like, you do it. It's like, why did that cost me 30 cents? I don't know. Just run it in high place. I did all the effort to put my GBG keys in the right place. So I'm like Claude. Go.

1:34:33

SPEAKER_07

Any more questions while y'all ship? Any confusion, anything I can unblock for you? It's my purpose in life. Sorry, you may have answered this question already. But so are you using knowledge base rag on unblocked? Or what exactly is the tech that you are surfacing? Oh, so many things. I can come talk to you at the side.

1:34:57

SPEAKER_00

I'll take my mic off. I'm just going to answer that question. Sure.

1:35:00

SPEAKER_07

That was just this. Sorry.

1:35:39

SPEAKER_07

What's the churn rate on the data?

1:35:46

Oh, it's real time, basically. So I guess there's two parts to that question. One is how frequently unblocked updates the data on the backend. So it's real time for many of the integrations and then on a cron job for others. Because for those particular integrations, they don't have web hooks, basically. Yeah, that's a question. Yeah. But the disk, so that means that rebuilding the graph data has to happen on a very frequent basis.

1:36:23

SPEAKER_06

Yeah. Yeah. Yeah.

1:36:39

SPEAKER_10

Yeah. Yeah. Yeah. Yeah. Yeah. Yeah.

1:36:47

SPEAKER_10

No, it's incremental. So our social graph build it algorithm has an incremental component to it.

1:36:55

SPEAKER_06

So we don't have to rerun the whole thing.

1:36:57

SPEAKER_10

But also social graphs are less sensitive to frequent changes in data because it's unlikely that a single change is going to make a huge impact on the experts graph unless your organization

1:37:13

SPEAKER_06

is brand new. So for...

1:37:20

SPEAKER_06

Yeah. Yeah. Yeah. Yeah.

1:37:26

SPEAKER_10

So as an example, we do best practices distillation on a much lower cadence. Like basically week by week because yeah, it just doesn't change that much. Yeah. Yeah.

1:37:49

SPEAKER_06

Yeah.

1:37:58

Yeah. Yeah.

1:38:01

SPEAKER_06

Yeah. Well, the... Oh, yeah. Repeat your question. That's a good question. So I want to make sure we get that one done. Yeah. It's attention. Oh. And in terms of customer privacy data retention, kind of... Yeah.

1:38:17

SPEAKER_10

From my point of view, I'm thinking of like enterprise SAS or even like on premis type deployments, which I'm not suggesting that you...

1:38:25

SPEAKER_05

I'm just thinking of that customer kind of modality. Yeah. Yeah. Do you get pushback? How do they feel about you holding data? It's another processor in the loop.

1:38:36

SPEAKER_10

Well, so the privacy discussions happen

1:38:40

SPEAKER_05

at the organizational level.

1:38:44

SPEAKER_05

We don't actually run into a lot of friction. There are definitely environments like in government and at banks that have super sensitive needs.

1:38:54

SPEAKER_06

And so for those needs, we have an on-prem solution. But it's definitely not the path that I would recommend. Like staying cloud-based. Like we have very large enterprise organizations that are entirely cloud-based, like fully cloud-based. That, you know, the secret sauce is kind of like less encoded in source code now and more encoded in the reasoning. So organizations tend to be a little bit more sensitive around things like Slack data, for instance. But the way that we store data, like we have a whole white paper about how we protect customer data. And it's never been a problem. Yeah. Yeah.

1:39:56

SPEAKER_06

Pardon me? I missed that wasn't recommended. Oh, why it's not recommended?

1:40:00

Well, the cloud-based integrations, you know, get updated more frequently. And so there's software patches. It's a little bit harder to maintain within an organization. There's one customer. It's a bank where administering the platform becomes quite difficult because they have network isolation. And so like now one of us has to, you know, sit within that network and administer the platform or we have to train individuals within the company to administer the platform. So it's just, it's more of a maintenance and hand-holding exercise. But yeah. Yeah.

1:40:49

SPEAKER_06

Exactly. That's exactly right.

1:40:57

SPEAKER_06

Yeah. Thank you so much for the talk.

1:41:02

SPEAKER_10

Yeah.

1:41:02

SPEAKER_06

Thank you. Thanks for coming. Thanks. Thanks. Thanks.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note