AI Engineer

Let AI Agents Tell You What They Need — Raj Navakoti, IKEA

956 summary words 4 min summary Watch video

Start with the signal

4 min read

Summary

Let AI Agents Tell You What They Need — Raj Navakoti, IKEA

Main Topics

  • AI Agents and Context Management - The core challenge of providing institutional knowledge to AI agents
  • Enterprise AI Delivery Gap - Why companies see minimal value despite AI capabilities
  • Demand-Driven Context Approach - A novel framework for building curated knowledge bases
  • Knowledge Base Architecture - Transforming monolithic documentation into structured context blocks
  • Practical Implementation - Tools, automation, and scaling strategies

Key Points

The Core Problem

  • The "Memento" Analogy: Like the movie protagonist with 15-minute memory loss, current AI agents excel at reasoning and code generation but struggle with institutional/domain knowledge
  • Enterprise Reality: 88% of companies use AI, but only 6% see actual value creation (McKinsey data)
  • Knowledge Gap: AI performs well on general knowledge (documented), struggles with institutional knowledge (tribal, undocumented)

Knowledge Base Reality

Enterprise institutional knowledge typically breaks down as:

  • 20% - Outdated information
  • 20% - Unreliable
  • 10% - Duplicated across systems
  • 40% - Tribal knowledge (never documented, exists only in people's heads)

The Demand-Driven Context Solution

Core Principle: Use a pull approach rather than push approach

  • Instead of building all MCP servers/RAGs first, then giving to agents
  • Assign work items to agents → let them fail → identify knowledge gaps → fulfill gaps → document → repeat

Why It Works:

  • Similar to onboarding new employees: give them tasks, let them ask questions, knowledge emerges organically
  • Comparable to TDD: write failing tests first, then code to pass them

Implementation Framework

One Cycle Process:

  • Agent receives a problem/work item
  • Agent attempts to solve using available knowledge base
  • Agent fails and identifies missing information (as a checklist)
  • Domain experts fulfill the checklist
  • Agent documents the new knowledge in structured blocks
  • Cycle repeats with new problems

Results from Testing:

  • Started with 56 entities in knowledge base
  • One incident surfaced 6 never-documented entities
  • After 14 incidents: confidence increased from 1.4 to 4.4 (on 5-point scale)
  • Agent gradually becomes more autonomous with better curated knowledge

Key Advantages

  • Surfacing the Unknown: Discovers what was never documented (tribal knowledge) automatically
  • Knowledge as Managed Resource: Agents become knowledge managers, not just consumers
  • Prioritization: Identifies what's critical vs. nice-to-have based on actual problem frequency
  • Scalability Through Automation: Can run across historical work items (Jira tickets, incidents, support tickets)

Technical Recommendations

Storage Medium: GitHub repository

  • Benefits: Built-in PR/review processes, version control, conflict resolution
  • Can publish to other platforms (Confluence, Slack) after curation

Meta Model Addition: Graph-based relationship mapping showing how:

  • Business processes relate to systems
  • Systems relate to APIs
  • Business/technical jargon connects to entities
  • Helps agents navigate contextually

Context Management:

  • Average domain requires ~96k tokens (well within modern context windows)
  • Full context loading better than RAG for most cases
  • Graph RAG only needed for very large contexts (1M+ tokens)

Notable Quotes

> "Every 15 minutes, he has to take his notebook, watch his tattoos...figure out what was I doing before. If you relate to the AI and AI agents things, it actually fits exactly how the agents are right now."

> "Unless it picks a task, if it picks a ticket, it has to fulfill all of them. It is so good with green ones and orange ones, but it struggles with the red one, with the institutional knowledge."

> "So unless we break down that monolith knowledge base into some kind of context blocks, which are useful for agents, then only we can actually make it useful for them."

> "I prefer to have it on GitHub...because eventually somebody will come up with a 20 million seed funded SaaS solution for you. But before that, I prefer to put it in GitHub."

> "Don't do it at an operational level, at a real-time level, but do it before retrieval itself."

> "The whole premise is to move agents from consumer to knowledge manager. You don't just consume from me. The whole knowledge management is your job too."

Takeaways

For Organizations

  • Address the Real Problem: Focus on knowledge quality, not just agent capabilities—this is where 80% of enterprise failure occurs
  • Systematic Knowledge Discovery: Use work items from archives to automatically surface documentation gaps
  • Kanban-Based Documentation: Treat missing knowledge items as Jira tickets with critical/high/medium prioritization
  • Cost-Effective: Running context gap scanner costs minimal tokens (~100k per domain); far cheaper than meetings and onboarding

Implementation Roadmap

  • Phase 1: Scope to team level, fix 60-80% of documentation in 2-3 weeks using automated gap scanning
  • Phase 2: Establish meta model to show relationships between entities
  • Phase 3: Move to operational mode where agents continuously refine knowledge while solving problems
  • Phase 4: Scale across multiple domains with GitHub as central repository

For Developers/Architects

  • Try the Context Gap Scanner immediately (available with presets, minimal cost)
  • Or use simple prompt approach: Take any Jira ticket/incident, ask agent to assess knowledge base quality against it
  • GitHub repository with starter guides available for implementation
  • ArXiv preprint available for academic details

Critical Success Factors

  • Domain expert involvement: Required to fulfill discovered gaps
  • Small scope initially: Start with one team, not whole organization
  • Accept incompleteness: Aim for 80-90% quality, then maintain incrementally
  • Continuous monitoring: Track document staleness via update dates and duplication metrics

Long-Term Vision

The future of enterprise AI depends on agents as knowledge managers, not just task executors. Organizations that systematically curate institutional knowledge will see the true 6% → 50%+ value realization possible from AI.

Full transcript 10439 words · 60 min read
0:15

SPEAKER_01

Thank you, maybe we can get started. First of all, thank you so much for coming for the workshop, especially ones who didn't get the seat. I promise you I'll do my best to make it entertaining, especially for you. Thank you so much. Actually, it makes sense. So now I know why the tickets got sold out, right? Which workshop actually sold out the tickets? So let's start with my introduction. So I'm Raj. I work as a staff software engineer at IKEA. I work for a domain called Deliverance Services. We are almost more than 100 engineers and six teams all together. It's a mini company within the company itself. I'm very interested in architecture, neuroscience, and linguistics. And now AI. So if anyone has some cool projects, because everybody is building cool projects these days, please find me after this meeting. So quick pulse check with the audience. Who is visiting London for the first time? Okay, cool. Welcome to London. Who is here from the engineering background? Also with white coding and prototyping? No, all of them. Okay. So who actively uses agents like copilot? Okay. Okay. This is going to be tough for me then. So extensions? Okay, so everybody's for okay, fine. So much tension now. So you're going to sit here in this hot room for more than an hour. So first, I will give a bit of introduction of what I'm going to present today. It's basically on agent and context management. I divide it into three parts. One is the situation, which all of you already know. So I'll keep it tight and short like five minutes. Then I will talk about the problem. This is where I'll spend a bit more time on the slides because I think nobody is actually seriously looking into the problem. And I see how many of you have seen this movie. Okay, okay, okay, okay, cool. So I'll give the gist of the movie. So this guy is very skilled, very talented. The only problem he has is he can't hold memory more than 15 minutes. So every 15 minutes, he has to take his notebook, watch his tattoos that he put on his arm and figure out, okay, what was I doing before the 15 minutes. And he does it again and again. If you relate to the AI and AI agents things, it actually fits exactly how the movies and how the agents are right now. If you go and watch this movie, you don't need to watch YouTube or blogs to understand agents and MCP. This movie actually tells you about everything, literally.

0:22

SPEAKER_01

And as this guy has a memory problem, in the same way, the AI that we got introduced a couple of years ago is very good with reasoning, computation, code generation. Its benchmark is above par. The only problem is the institutional knowledge, right? The domain knowledge that you have. That's the only thing that could be a bit problematic.

0:29

SPEAKER_01

So from AI to agent, if you look at the evolution, it exploded. So it started from prompt engineering, first of all. Then there was RAG, MCPs, then multi-agent. Now it is deep agents. I recently found out using Replit, you can build a full stack app in 10 minutes. That means by the time you make the instant noodles, you already have a million dollar app working on your laptop. So we got to this point, extraordinarily good.

0:39

SPEAKER_01

Now that's AI and agents. So let's talk about enterprise AI. Okay. I don't know how many of you have this question, but most of the enterprises, I see the question is, okay, AI is pretty smart. It's doing code generation, full stack apps, reviewing your PRs, doing incident management, all those things. So if I was doing that much, why are the Jira tickets or epics not moving on the dashboard, right? Why don't I see the delivery actually? So everybody's speaking about, look at like three minutes, everything is ready. Yeah. Okay, fine. Why are my Jira epics not moving?

0:56

SPEAKER_01

Because that defines the business delivery and that defines the rate of investment, right? And as you see, it's from McKinsey this year, like 88% of all companies use AI, but they only see 6% of value creation. Okay. So I think this is the problem that we have. I have four Jira tickets, different ones, sample ones. And you can see the green ones that I have marked as basically which LLMs are already trained on, like APS standards or things they already know. It's general knowledge, right? That's fine. Those tasks from the ticket, it can pick up and it can do it. Now, there are second part orange ones, which we have to teach them actually. So you know this, but do this in this way. So all this orange color things will fit into your agent extension like skills, but red ones, that's what the institutional knowledge is, which sits within the company and within the people. So unless it picks a task, if it picks a ticket, it has to fulfill all of them. It is so good with green ones and orange ones, but it struggles with the red one, with the institutional knowledge. And what I believe is, right now the coding agents are getting so, so better. I feel like if there is an AGI coming, the first AGI will be a coding agent for sure.

1:02

SPEAKER_01

So to fix the giving the institutional knowledge to the agents, we have an industry solution already. So this is basically your return of investment on AI pipeline will look like, right? So you have LLM model quality, you have agents and you have agent harness and your institutional knowledge sits under Confluence, Jira, SharePoint, GitHub, all those things. And basically the retrieval layer is what industry is telling us will fix that issue. So you build that retrieval layer and then it will fetch all those things and give it to an agent and agent should be able to do it, right?

1:08

SPEAKER_01

So basically, 40 percent factual accuracy can be achieved through RAG or knowledge graphs actually, but with the documented knowledge base. Now, if you build a retrieval layer, it has to work, right? Now, let me ask you a question. How many have you built retrieval layer things like RAGs and MCPs? Okay, cool. All of them. Okay. Now the question is, how many MCPs did you build? How many have you built more than 20 MCP servers? Okay. Okay. So nobody beats my record then.

1:23

SPEAKER_01

So what I see is mostly in the enterprise organizations, people are building like 10 to 15 or like 20 MCP servers or like RAGs, knowledge graphs on top of their institutional knowledge and to the agent, right? So the assumption is, if I can build all this MCP servers and give that agent, I do not need to work on anything. It will do it. But the thing is, when you plug in this MCP servers, basically all this data coming out is mostly undeterministic, it is unreliable and it is untested, right? So especially in engineering, nobody does evals. Actually, it's more like a data machine learning concept, but we do not do evals. So for me, if you plug an MCP or a RAG, we see whether the output is coming or not rather than is it really valuable? Actually, is it really solving the problem or not? There's a main problem that I see. I'm not saying pointing other people because I was that person. I was like, okay, let me build all MCP servers, plug in my institutional knowledge. I'm going to prove the point that agents can semi-autonomously,

1:36

SPEAKER_01

It is unreliable and it is untested, right? So especially in engineering, nobody does evals. Actually, it is more like a data machine learning concept, but we do not do evals. So for me, if you plug, sorry, if you plug an MCP or a RAG and all, we see whether the output is coming or not rather than is it really valuable? Actually, is it really solving the problem or not? There's a main problem that I see.

1:52

SPEAKER_01

I'm not saying pointing other people because I was that person. I was, okay, let me build all MCP servers, plug in my institutional knowledge. I'm going to prove the point that agents can semi-autonomously continue and fill those zero tickets and finish it, right? But every time when I build those MCP servers, 10%, 20%, 30% of time, it was accurate, but the rest of the time, I was doing the data entry job for them, actually. So I was filling the gaps, asking the questions. So I'm doing more work than actually doing less work. So I think this is the main problem. And I actually was in this fourth stage where I literally started to write the domain context with handwritten, actually. So okay, let me write everything and prove the point, but I got really exhausted of doing it.

2:02

SPEAKER_01

Okay. So I don't know how many can you relate with this pie chart, but most of the enterprise, the institutional knowledge is something like this. So 20% if you see, it's outdated. 20% it's unreliable. 20%, 10% is always duplicated with different places. And the major problem is, 40% of the knowledge is always tribal knowledge, which means people know how things work. So it's never documented, actually. So in this situation of an enterprise, and you build 100 MCP servers and plug into that monolith, it doesn't matter how many you build it, it won't work. Because your whole institutional knowledge is a monolith. I think, because you're all from the engineering background, so you already know the transformation of monolith legacy system to microservices, right? So in the same way, unless we break down that monolith knowledge base into some kind of context blocks, which are useful for agents, then only we can actually make it useful for them, for the agents, and actually make them semi-autonomously do the tasks. So that said, we are going to talk in this workshop mostly on that monolith, how to break it, what is the approach to break it, and how it is useful once we break it. And this is a job we need to do, because the LLM providers will focus on the model quality, the agents will focus on the harness things, and there is a big retrieval market of 9 billion, they're focusing on retrieval. But nobody is going to come to your company and fix your knowledge base, you have to fix it yourself, right? So how can we do it?

2:09

SPEAKER_01

Okay, so the demand-driven context is what the solution I was trying to propose. So basically, if I have to give an abstract of it, what it is, we have monolith services, and we have this process of breaking them into microservices. We have waterfall model which we transform into Agile, in the same way when you have a monolith of institutional knowledge, how do you transform into context blocks using an approach? So this is an approach of how we can do it. Before starting, just not an idea. So we already tried with some data sets and tried to prove this approach works. And in March, we have published a preprint in arXiv. So if anyone interested in reading academic papers, you can find it with the demand-driven context, or I can also give you a link after the workshop.

2:16

SPEAKER_01

Okay. So how does it work, actually? When we are giving institutional knowledge to agents, basically what we are trying to do is we are trying to do a push strategy, right? So we build everything and we push it to it. So in this approach, it's more pull approach, which means, for example, let's say a new joinee has joined your company, right? How do you onboard a person? So you onboard them for one, two days, you give some initial orientation, and then you tell them, okay, these are the confluence links, these are the GitHub, this is some kind of documentation, you have to follow things and all. Then you just assign a task to the person. But you're not going to tell, okay, go and get graduated on this knowledge and come back, then I will give you work, right? So you will just assign the work item. And when you assign the work item, the person will start asking questions, fill the gaps. If the person is very much into documentation, he will also fill the documentation for you. He gradually get his knowledge of the institutional knowledge, right? In the same way, we don't push all the knowledge to the agent rather than we start giving problems to agents like work items, and let them actually pull the information from us. And once you pull the information, also ask them to document it in a better way, rather than in a monolithic structure. So that's the four layers. So you have a monolith, a framework, and it actually pulls and actually creates a good, better context blocks. You can actually relate it to more into a legacy monolith to microservices directly if you have to have an analogy of it. So this is how it works. So this is one cycle of how the problem to an agent, and in the first attempt, the agent will fail to do it. So it will say, you know what, you gave me a problem, but most of the documentation, I couldn't able to find anything, I couldn't able to do it. Then these are the things I need to do to finish this task, and it gives a checklist of things. So we fulfill the checklist. So once it is given, the problem is solved, it will take the knowledge, and also it will update, that means curate the knowledge in a particular place, so that it can reuse or other agents can also reuse. This is one cycle. So the idea is, if we can do it in multiple sessions, with multiple problems, so it will gradually curate your knowledge, monolithic knowledge base, and also document it for you.

2:21

SPEAKER_01

You can also relate it to TDD. So how many are doing TDD, or nobody hates TDD, right? Before I, yeah.

2:48

SPEAKER_01

Set them. Okay. Okay. So in the same way, right? In a TDD approach, what we do, we just write the failed test cases, we don't build the product, first of all, we just write the failed test cases, we see what is a code that is missing, for the failed test case to pass, and we just give the code, and we gradually build a product based on the failed test cases. In the same way, we give problems that agent will definitely fail, and we gradually fill those gaps, and at a certain point, it becomes semi-autonomous with a good institutional knowledge already. Okay.

3:09

SPEAKER_01

So I think I can jump into some kind of a demo already. So I will use Terminal, so don't hate me. I think all of you are from engineering background, so I think you will like Terminal. Let me switch to Terminal. Okay. Okay. So on the far right, what you see is how under the hood it works, actually. For the failed test case to pass, we just give the code, and we gradually build a product based on the failed test cases. In the same way, we give problems that agent will definitely fail, and we gradually fill those gaps, and at a certain point, it becomes semi-autonomous with a good institutional knowledge already. Okay.

3:56

SPEAKER_01

So I think I can jump into some kind of a demo already. So I will use Terminal, so don't hate me. I think all of you are from engineering background, so I think you will like Terminal. Let me switch to Terminal. Okay. Okay. So on the far right, what you see is how under the hood it works, actually. So when you have given a problem, how does the agent fail? How does it demand for the knowledge that the problem has to be solved? And a human, a domain expert and all, fills those gaps, and then it will curate a new knowledge base for you, which is much better. And then the agent succeeds, and you repeat on the next problems. So that is how one cycle things are done.

5:08

SPEAKER_01

So how can it be implemented? It can be implemented using any agent. There is no, it can be implemented on Cloud or Copilot, because it's an approach you can do in any way you want. At work, I use Copilot. So I implemented this using Copilot. But because everybody, I believe, loves more Cloud code. So I created this demo with Cloud code. And you can see, it's just a combination of skills, rules, agents, and hooks, and some kind of a place to save the knowledge base.

5:43

SPEAKER_01

On the middle pane, what you're seeing on the top is your monolith, basically. This is a representation of your Confluence, Slack, GitHub, and all. But just for the sake of demo, I just put some flat files that look like them. So that is how your monolith knowledge base will look like. On the down, what you're seeing is on a live. So when it's solving a problem, how it is actually adding the new knowledge to it. So this is how, so let me, what I'm going to do is I'm going to go to an agent. Okay, I'm going to basically give an incident problem to do the root cause analysis, right?

6:23

SPEAKER_01

So, okay, what I did is, you remember in the previous slide, there is a geographic examples that I showed, right? It's a combination of knowledge that is documented, not documented things at all. So this incident also represent the same kind of combination.

6:56

SPEAKER_01

So there is some knowledge that is documented already on your monolith. Some, it is not there or outdated or things like that. And most of it, it couldn't, wouldn't be able to find because it's never written down, actually. So when I gave this problem, it uses those skills that I have developed using this approach. And it will try to actually first go to your monolith, actually, on the knowledge base and try to find information on what is there. So think about it like this. So first part is retrieval. That means it's already doing what RAG and MCPA is doing first part. But what else it is doing is after it fixes the data, what it will do with the data, actually.

8:08

SPEAKER_01

So that is a missing part. For example, when you give a new Confluence links to a new employee, the employee goes there, looks into it, but doesn't find information. But he doesn't stop there, actually. He continues asking questions to solve the problem, then just adding more knowledge and things and all right. Those are the next steps missing right now. It's, we just stopped at retrieval. So this is the next three steps it does, actually. So you can see the confidence score is almost one to five because it says, these are the particular terminologies. I don't understand, actually, these terminologies.

9:19

SPEAKER_01

And these business logics is not needed. So one thing you need to look at here is whatever it has said, this is the undocumented information. That means it was never written down. So unless you don't do this way, you will never know what is not documented. For example, if somebody says, okay, there is documentation missing, we need to write. Okay, what do you want me to write, actually? So there is so much in the people's head.

10:12

SPEAKER_01

I can't write so much. Somehow it has to surface. So when you give a problem, it actually surfaces what is not documented. And it tells me, okay, this is missing. I need to have a new information there. So what I will do is, so it does all the three steps. So then what I will do is, I already have a pre-prepared answer. Very high level prepared answer, I gave it to it, of what is the missing information. So, okay, this is the missing information, you asked me to solve this problem. Can you solve this problem now?

11:19

SPEAKER_01

I didn't expect it this one, okay. Notifications is, yes, I will just say yes. If that's also what fictional name should I just put it? Okay, I didn't see. See, when I did the test and it didn't ask the questions. Yeah.

12:02

SPEAKER_01

Let's see, it knows it is a demo and, yeah, I tried it too. Okay. No, it is already, what it is doing is already. So you can see on the live, it is already adding the entities that the new knowledge base has come into the place. So the knowledge base is managed as a file system, right? Also, as a system files. For the demo I'm just showing it as a file system, but it's basically your MCP servers, the data will stay in Confluence, Slack, or things. You can just plug in and use the same MCP servers or RAG and all.

13:00

SPEAKER_01

So it no need to be a flat file. So you treat this as your systems layer for this area, right? It's a mem, do you use any memory tool for this? Yeah, I will show you on the next look at how I'm going to see it. Okay. So it started from 56 entities or something with this one, right? Now, one problem actually surfaced six entities that are never been documented. And when I gave that information to it, it is able to actually discover, curate another five or six new entities that were never documented.

14:12

SPEAKER_01

So it does discovery of the gaps. It also gets information from me and also stores information, new information and all. So this is one, okay. Next, let's see. This is a busy window. I tried to actually do things, but it didn't work out.

15:18

SPEAKER_01

Okay. So what you are seeing on the window is 14 incidents. You have seen one problem that I solved with an agent, right?

15:49

SPEAKER_01

Now, one problem actually surfaced six entities that have never been documented. And when I gave that information to it, it is able to actually discover and curate another five or six new entities that were never documented. So it does discovery of the gaps. It also gets information from me and stores new information as well. So this is one. Next, let's see. This is a busy window. I tried to actually do things, but it didn't work out.

16:55

SPEAKER_01

Okay. So what you are seeing on the window is 14 incidents. You have seen one problem that I solved with an agent, right? The communication. What if I took 14 incidents and I just go and have 14 cycles of this thing and see how it does? So if you see on the left side, it was the first incident, right? So right now it has 1.5 confidence and everything is critical.

18:09

SPEAKER_01

So basically nothing is documented. So everything is critical. The data is missing. So I started giving answers on the first incident. Then I repeated for the second and third and continuously for 14 incidents. But on the 14th incident, it basically actually able to go to a confidence level of 4.4, because first it discovered on every incident.

19:11

SPEAKER_01

It got the list of answers for me and also it documented everything for me. So it gradually improved from 1.4 to almost the 5 range of knowledge. So if you look at the traditional way, in traditional way, what we do is we solve all the context problem, right? We have to deal with it first. Then I have to give it to agent. In this one, we are moving agent from consumer to a knowledge manager. So you just do not consume from me. I am going to tell you, but the whole knowledge management is also your job and you have to do it for me. Okay. I think we can get back to the slides a bit. Okay.

20:04

SPEAKER_01

So what we have seen is I have run one cycle and also I have shown how it looks like when I run 15 or 16 different cycles, right? But if you want to do it manually, it would be really painful because I tried it after 15 cycles, and nobody would like to actually sit with an agent and have an incident, but you won't be sitting with your agent and keep actually asking questions and telling you about your problems, right? So that is super painful. But the thing is we can automate this process. So this is where actually it's really good and gets interesting. So here is the thing. You all, we already have all the work items, right?

20:27

SPEAKER_01

We have Jira, we have incidents, we have customer support tickets and all those kinds of work items already there, right? Sitting in the archive. So why can't we take them and actually use the framework and validate across your monolith database, run an automation and see actually what is the state of your content actually right now. Okay. Let me see, let me show you how it looks like.

20:50

SPEAKER_01

Okay. So rather than actually doing it manually at scale, if we do this approach, so how does it look like? So the demo that you're seeing is almost like everything is preset. For example, I have the demo, I have a platform operations agent and I'm saying, okay, these are the recent incidents. Let's say I have 20 recent past incidents I have, an MD file or a JSON file, right? It has all the details of description, things and all comments and everything. And the rest of the files are your knowledge base. So it's a file system, but you can also actually connect with the same way, Confluence and things and all. Just for a demo purpose and just showing it as flat files.

21:17

SPEAKER_01

Now, what I'm trying to do is, I'm going to take all these incidents and validate each incident across my knowledge base and ask the agent, okay, tell me how much of the document is good, how much of the documentation I can't trust or is old or outdated and how much is actually missing, not documented as per these incidents. So let me run it, okay, it will take some time. So it will take three steps actually. So one is, it generates probes, which means a basic test it will write to actually test your knowledge. Then it will run those tests and then analyze the gaps actually. Okay. It's a little bit hot in the room actually. Okay.

21:49

SPEAKER_01

For example, let's say you have an incident called the notification service is not sending customer messages to SMS service, right? So the notification service then you mentioned that the agent sees, is there documentation related to notification service. So it doesn't find that means you never wrote documentation on notification service.

22:04

SPEAKER_01

I do understand what is a customer SMS and things and all. The customer notification service when you mentioned, it's a gap because it's never documented.

22:13

SPEAKER_01

Or it takes a customer notification service, goes to Confluence and sees how old the documentation is. If it is, say it's like one year old, it will tell you, look, I looked into it. It's like one year old, I don't know whether I need to trust this documentation or not or it's incomplete documentation. So if you see it's code each incident, it took it and it looked at all the knowledge base that you have connected and have consolidated list of scoring of, okay, partially the agents can handle the basic edge cases of the incidents that you give because your knowledge base is not complete.

22:21

SPEAKER_01

And it will show you how much of the tribal knowledge is missing, system information, business process, what actually is missing from your institutional knowledge when whatever is not documented. These are the probes.

22:30

SPEAKER_01

And it will also identify what is critical and what is high. This is really important because let's say there is some kind of an example of notifications of which I mentioned, right.

22:48

SPEAKER_01

It is repeatedly appearing in 20, 30 incidents. You don't have it. This is the first that you need to fix as per your documentation. So it will also help us actually understand when you're breaking down your knowledge base, you need to understand what is critical actually, what I need to focus on first, what makes value for me. So we organize into critical, high, medium. So this is what I showed with the flat files but you can also connect it to the various data sources that we have.

23:10

SPEAKER_01

This is really important because let's say there is some kind of an example of notifications of which I mentioned, right. It is repeatedly appearing in 20, 30 seconds. You don't have, this is the first that you need to fix as per your documentation. So it will also help us actually understand when you're breaking down your knowledge base, you need to understand what is critical actually, what I need to focus on first, what makes value for me. So we organize into critical, high, medium. So this is what I showed the flat files but you can also connect it to the various data sources that we have. So step one is basically what it does, demand extraction, that means every incident, it will extract the checklist of information what is missing. So it will create systems and APS and all and how many are clean, how many are stale, which is incomplete, what is entirely missing, something is tribal, those kinds of classification it will also do. And it will create a Kanban board for you. So what happens is, so if you want to fix your institutional knowledge base basically, you just like Jira tickets, we finish it, we actually have to document these missing pieces and all. And the moment you started to, so it also saves in the context lake, so it also has to build its own knowledge base. And you can see the performance, so once you're fixing those tickets on the Kanban, you can see the institutional knowledge. So that how, so what we seen is, one is the approach first of all, which means not the push approach but the pull approach how to do it, one cycle or multiple cycle how it look like, but if you put it in a scale of automation, then how much valuable it would be. Okay, now, the important question is, double mic, it's cutting out a bit, so I'm just going to put that one on as well. Should I do a voice over later actually?

23:14

SPEAKER_01

Yeah. Rock and roll. I think you need to repeat that one. From the beginning. I have the patience, who has the patience for a second? So, the question is, so I was all, all the time I was talking about, okay, it receives [SPEAKER_06] the context, we give the information, it will store it, but the question is where does it actually store it?

23:29

SPEAKER_01

So, I have a very opinionated opinion, hear me out. I prefer it has to go to a GitHub repository because eventually somebody will actually come up with a 20 million seed funded SaaS solution for you. [SPEAKER_03] But before that, I prefer to actually put it in GitHub as a repository. Why? Because if you look at it as scale, if you want to do this, there will be multiple agents, multiple teams actually contributing to the same knowledge base. And there will be conflicts and resolutions, right?

23:40

SPEAKER_01

So, the easiest way to do is using GitHub because it actually comes with inbuilt PR processes, review processes, things and all. So, if multiple domain experts are sitting and uploading the files or agents are contributing to it, the most efficient way to manage is in a GitHub, something like a structure like this. So, the other advantage is also if you put it on GitHub, you can also publish it to Confluence later or Slack, wherever you want to publish it to another solution that you want to use. So, I prefer to have it on GitHub, but if you want to directly integrate it to Confluence and all, you can also insert into it. Next is a meta model. How many are aware of the word meta model and the, okay, maybe I can quickly show you how does it look like? So, meta model is basically something like this, right? So, and how does your domain actually structured around is a business process or how it is related to systems, how systems are related to an APIs and how is this business jargon or tech jargons are actually linked to which one. So, these kinds of relationship meta model is really important. It's not necessary for the approach that I have proposed, but it's an add-on. And why you need to have this one is right now, think of it as a map. Right now, your agents doesn't have any map actually to navigate with your knowledge base. Basically, what you're doing is you're dumping these many number of files and you need to figure out which file I need, right? But your file structure is actually a representation of your meta model. It actually knows how to navigate. For example, let's say, can you fix this system? It will understand if I make changes, which business processes will be affected and which APIs I need to change or touch these kinds of things. So, it's also important to have a meta model. If you have it, then it will produce more value. So, I strongly prefer to have a meta model along with this approach. Okay. So, the last part is what is the value it created? So, there's a lot of slides that you have seen, a lot of demos that you've seen. So, personally, I need to also share what is the value that I see when we, I was using it or the other people who I shared with already were using it, came back with the feedback and told me. First, the most valuable thing is knowing the unknown. So, what is never documented is something can be surfaced only by this approach actually. Otherwise, you will just end up in an endless Miro board of putting tickets on, okay, this is missing, this is missing, I need to add it, I need to add it and keep on doing it. So, this is the fastest and better way to discover with your previous work items and all, what is never documented things and all. Second is basically, I can now give work to agents rather than I do all those things. Rather than I give the agents all this information, let it manage my knowledge management. I don't want to be the knowledge manager of it. So, let it do it. So, those are the two big values that I have seen. If you want to use it, I think you will also see those two as the most valuable. But these are the other things what I've seen. Now, okay, so I also need to tell you what is the drawbacks of also using it, right. First of all, if you are coming from a small team or if you say no, no, no, my documentation, my knowledge base is really good. I am super happy for you.

23:49

SPEAKER_01

I don't want to be the knowledge manager of it. So, let it do it.

23:59

SPEAKER_01

So, those are the two big values that I have seen. If you want to use it, I think you will also see those two as the most valuable. But these are the other things what I've seen. Now, okay, so I also need to tell you what is the drawbacks of also using it, right. First of all, if you are coming from a small team or if you say my documentation, my knowledge base is really good. I am super happy for you. You are the lucky ones in this world right now with agents. For you, it might not be really relevant unless you have a very complicated documentation that you have. Second is I already mentioned the manually doing is very painful. I don't prefer anyone to do it.

24:29

SPEAKER_01

If you want to just try it for testing purposes, you can also do it. But automation is the best way to actually use this one. This is very early, this approach. So, by tomorrow morning on YouTube, somebody would have already posted something differently, better than me. So, in the era of AI, nobody knows how long a thesis or an approach or an app product want to survive. So, for now, I see this is the best approach. Okay, so, the whole workshop. So, we started with one pipeline right on the ROI. And so, the demand-driven context actually sits between this monolith and also the retrieval layer actually.

25:03

SPEAKER_01

And what it does is, it actually helps you build curated context blocks for you. You can also think of it as a cache database that you have. So, every time your agent doesn't need to go and boil the ocean for fixing an issue, rather than if you have a good context block of information, most of the time, 80% of the time, that can be usable. Because what I also believe is, it's always the 80-20% rule. So, 20% of your documentation is most useful. 80% is some corner cases you have to look into it. So, rather than giving 100% of things, you need to figure out what is my 20% that is super helpful for agent. And have it as a cache database, the context block of it, using it.

25:41

SPEAKER_01

And the rest of it, you can leave it as links. So, whenever agent feels I need more information, then only it can go and check the whole monolithical institution knowledge. Okay. So, from here, what you can take from this workshop is three things. One, I hope I made sense of this approach. So, there is a GitHub repo, which I detailed it out, and also a starter guide on it. If you want to go home and try with it, you can try it. You already know how the framework works. So, you want to go home and just remix the whole approach, you can do it and let me also know. I'll leave this one and I'll join with you for contribution.

26:21

SPEAKER_01

You have a context gap scanner that I showed you, which is live already with presets. I think I added $20 on it. So, hit it as much as possible. You already tried? Okay. Okay. So, after $20, so, first time first serve. So, all these three you can use, you can take away from this workshop. Okay.

27:02

SPEAKER_01

So, because this is a workshop, so I also would like to want you to try something. What you can try is three things. One is either if you say, you know what, I am so tired already. It's almost like four. It's almost about to go for a party. I don't want to do it. So, you can just go to the context gap scanner. Everything is a preset here. You can just try it out, hit it and see how it works. Otherwise, if you think it can be done better, let me know so that we can work two ways. Or otherwise, let's say, no, I am very technical. I want to know how it works under the hood. This is a GitHub repository. It's under, maybe I will just take this out.

27:53

SPEAKER_01

This is a GitHub repository and it has all the information. Plus, there is a starter guide also if you want to try it out. But if you still feel no, I want much more simpler. You can also try this one. So, you don't need to do anything. Take this prompt. Take one of your Jira ticket or incident that you have right now.

28:36

SPEAKER_01

If you already built MCP servers or any other kind of things, you just use that prompt.

28:40

SPEAKER_01

Give it to your agent with the incident or a Jira ticket and ask it, give me the quality of the knowledge base that I have as per this incident or Jira ticket in this way.

28:52

SPEAKER_01

And see how much of it comes in the red, which is never documented. So, you can try also the simple one. I will just leave it like this. Maybe you can take a picture. Maybe I can switch to the slide if anyone wants to. Ok. Ok. [SPEAKER_00] Ok. [SPEAKER_04] Ok. Ok. [SPEAKER_04] Google likes in this approach, I find it very interesting, so my first question is have you already used this way of working at scale or because we've seen multiple toy examples right? Yep, yep. Ok. Ok. Ok. Ok. Google likes in this approach, I find it very interesting, so my first question is have you already used this way of working at scale or because we've seen multiple toy examples right?

30:02

SPEAKER_01

Yep, yep. That's why it's so fun. I used it not at a scale, I started a bit simpler because you also need to see what is the scope of it. Let's say I have an enterprise and I try it at an enterprise level, I can't do it because it's multiple domains, things and all.

30:22

SPEAKER_01

Even if I do it at domain level, I need to understand, I tried it at domain level. Then even at a domain level, there is so much domain expertise I need to fill it up and fill those gaps. So again, I cut down into maybe what is the smallest team that I have and the smallest team's Jira tickets, the smallest team's incidents and the team's confluence page with a bit of a scope. So then if I drill down the scope, then I feel it's more fast, more useful. But if I do it at a bigger scope, what happens is not one person has the whole domain expertise. So it again becomes that somebody has to come and five or six people have to sit down and start doing these things.

30:53

SPEAKER_01

I'm a bit concerned that this might denial of service attack your team members in a certain way because our LLMs are fine-tuned to keep eliciting information, to keep getting more information out of us to ask follow-up questions. So I think it will be hard on the engineers that have to do the question answering. And secondly, the scanner is nice, but that still builds on the assumption that all of your team members and the rest of the enterprise are still using your enterprise IT as planned, that they're actually filling in their tickets with all the details and et cetera.

31:06

SPEAKER_01

And I know from practice that that is most of the time not the case. That is true. That is true. I agree with you. Even if I go, my assumption is also, even if I go to a leadership to buy in, hey, can you give me a bandwidth or I need these people to actually sit and fix the context. I don't think right at this point of time, anybody will do. But I think it will happen. Because slowly, I think we are slowly moving towards agent managers where agents are becoming semi-autonomous or autonomous.

31:27

SPEAKER_01

And we manage them. But at this certain point of time, somebody has to fix that knowledge because it's not going to come from anywhere. You have to. So then the enterprise focus will shift towards the gap. That's what I started saying. I don't think anybody is looking into the problem yet. Everybody is very focused with agent. How good the agent is, how good the retrieval is, but how good the context is, you're not solving it. I think down the line in a year or so, I think people will realize the importance of it. And the Kanban board will definitely come into reality very soon. Thanks.

31:59

SPEAKER_01

I think actually dealing with the same points. I think when we look at large, which applies actually the source of truth is not actually documentation, it's actually code. I'm just wondering how you applied it to the code base. I did. I also applied the code base. But I got a mixed result when I... So here is the thing.

32:22

SPEAKER_01

What happened is when I only use code base, it is particularly good. Or when I only use confluence or textual data, it gives good results.

32:31

SPEAKER_01

But when I combine it, somehow actually it conflicts because it creates a theory out of the GitHub repository. But the same GitHub repository documentation is also on confluence. So there it gets a conflict of, okay, what is the source of truth? Code says this. Should I implement it this way as per the documentation?

32:54

SPEAKER_01

So then again, I need to create an additional skill or rules, okay, what is the ranking that you need to give?

33:04

SPEAKER_01

If you see it in GitHub, that means that is the source of truth. Or if you see it, if you don't see it, then you have to look the information in confluence things and all. But those are still, I'm trying to fix those things actually. So seeing the gaps and fixes it. But I definitely see the issue combining those two. And the second question is, interestingly, is actually applying the same approach or skills. Because what we find out is actually you have your kind of your process that's running agents, which is bringing a context and identify the right skills you need to use, right? Mm-hmm. Then you go into the task and you fail. Mm-hmm.

33:22

SPEAKER_01

The moment you fail, you identify what you need to resolve.

33:28

SPEAKER_06

[SPEAKER_01] Mm-hmm.

33:33

SPEAKER_01

You go back, you curate, you fix. And then you fix in the knowledge, kind of things. Once we find out actually if you go back and fix the skill. Okay. Increment of the next iteration. Sorry. So I'm not sure if you had some skill. Yep.

33:49

SPEAKER_03

[SPEAKER_01] Is it still part of the iteration loop or possibly?

33:54

SPEAKER_01

I think right now the skill that I have built is static. But what you are proposing, if I'm not wrong, it's evolving skill, right? If the skill fails, it has to evolve, right? I agree with you. I never tried it. But I think it has to be like that.

34:00

SPEAKER_06

[SPEAKER_01] Because I'm also more concentrating on how to do it at scale. The reason is also, I want the context to be fixed before retrieval itself. [SPEAKER_01] Mm-hmm.

34:05

SPEAKER_01

Not during operational. So first when I started with it, I started doing it when operational, which means, oh, I have a work item, I will assign to it, it will fail, then I will start giving context and all. But it takes a lot of time, it takes a lot of patience for me. So rather than doing it, I'm going to fix the context before retrieval. So if I can, while I was answering her question, if you take a team, the context that you need to fix is very small.

34:21

SPEAKER_06

[SPEAKER_01] So you can use a context gap scanner kind of a thing. And maybe if you're good, have a good domain expert, I think in a couple of weeks, you can actually fix your documentation. [SPEAKER_01] Not 100%, at least 60, 70, 80% of good quality that you can already build.

34:27

SPEAKER_01

So my proposal would always be don't do it at an operational level, at a real-time level, but do it before retrieval itself. [SPEAKER_04] That is much better in this approach. Yep. So I think that's a question. I don't know, five or six different docs. So you can use a context gap scanner. And maybe if you're good, have a good domain expert, I think a couple of weeks, you can actually fix your documentation. Not 100%, at least 60, 70, 80% of good quality that you can already build. So my proposal would always be don't do it at an operational level, at a real-time level, but do it before retrieval itself. [SPEAKER_04] That is much better in this approach. Yep.

35:00

SPEAKER_01

So I think that's a question. I don't know, five or six different docs. So you will have, I am going there reading all the docs. This takes time for retrieval information and that takes a lot of context. Right now, after cloud code, I know it's one million tokens in the context window. I don't have any problems. So I calculated it. At an average, it's 96k tokens, because I tried with different domains actually. Per domain, I see around 96k tokens, if I consolidate everything like confluence things and all.

35:33

SPEAKER_01

So it easily fits in the context window actually. I tried to do some experimentation around a graph rag, put them there rather than just take all the files, use a graph rag, understand the intent. But for me, just putting the whole context right now in the window gives you more results than actually doing a rag. Unless you have a very big, almost around a million tokens of context that you want to fit in, maybe then you have to use a bit more retrieval mechanisms between it. But otherwise, I think it should be fine.

35:49

SPEAKER_01

I have a question. I opened your paper and could you explain this graph like comparison between different techniques like domain knowledge, strategy, knowledge access? Which one? Yep, sure. Okay, so I also did the citations from other papers. Okay. So not directly related, but you have the paper of ACE, which also does a similar thing. Okay. So, but ACE is not exactly into discovery and curation actually. If I remember it correctly, maybe I need to refresh my memory. What do you mean between the difference between domain knowledge and strategy knowledge?

36:25

SPEAKER_01

Okay. So, strategic knowledge, okay. So, what ACE and all are doing is when you are trying to have a conversation with AI, you can see in the cloud code and all, it updates its memory or the relationship with you or things like that, right? So, and also from the chat history to understand what is the most important context I need to remember, those kinds of things. So, when you are in communication with it, that operational conversations with AI improvement, they propose. So, what I propose is not based upon your conversation with AI, but rather on your domain knowledge which is documented actually. Somebody else has a question. Sorry.

36:48

SPEAKER_01

So, when I wrote a pipeline for extracting from Confluence, it also allows actually to give you a date and also last update. So, when I wrote a pipeline for extracting from Confluence, it also allows actually to give you a date and also last update. Which one is stale and which one is not. You don't have an intermediate layer where you're in the repo, you store this is still. Okay, so when it is curating the context, also it updates with a date and also the state of the document, like stale, active, and clean.

37:14

SPEAKER_01

[SPEAKER_05] So it also looks into, okay, this is stale, I'm not going to touch it, and I'll just go to, just look for any of the new other documents are there in this one. Okay, thank you. Do you think about how to manage access or permissions later to this knowledge? Like if you have some knowledge in the company that all the specific people can get access to it, [SPEAKER_05] is it correct that you just have all the knowledge and everything that's accessible? Okay, so because it's not a product or a SaaS solution, it's just basically GitHub.

37:39

SPEAKER_01

For me, right now, permissions and things are not difficult to implement because GitHub out of the box gives me who I can give the permission to is GitHub, who can have write, read, access things and all, who can merge, those things and all. But in case if it evolves into a product, and for example, context Gap Scanner as a product, and I want you to test it, [SPEAKER_05] because the reason why I was using presets for this workshop, not actually asking you to upload the files, is because I don't want to take your IP data on this one, right?

37:57

SPEAKER_01

So unless it becomes a product, you don't have any problem, GitHub and all, but if you have a SaaS solution for this one, then it's between how the SaaS solution will manage it right now. [SPEAKER_05] But the approach has nothing to do with access things and all. So how you implement access on the knowledge is up to you. You have a question? Yes. So this is of course about documentation, but did you give any consideration about using it on some central tooling that a company would use? Like if you have a platform team and you have a CLI that the different teams are using, and so now it's used by different agents, right? Okay.

38:30

SPEAKER_01

So the agents can also be like, well, this action is available for a resource, but I don't want to do 500 calls, yes, because I have at least 500 resources. Okay. It would be nice if the tool could do that. I don't know if you've given any consideration to me. [SPEAKER_05] I think that is how it has to work in an organization. You need to have a central solution for it. But how you want to do the solution is up to the organization. [SPEAKER_05] For example, we are doing Agile, right? [SPEAKER_05] So Agile can be done by Scrum, Kanban or Lean or something, and also you can do different apps to do it.

39:22

SPEAKER_01

[SPEAKER_05] The process is the same, but how you do it, which method you will choose, and which app that you will choose in your organization is different. In the same way, what we have discussed is the approach. If you want to put it in the organization, you can use the approach and you can do it in whichever way you want. My point is more like, so with these, you can identify gaps in your. Uh-huh. Right. Could you use it to identify gaps when you're tooling? [SPEAKER_05] OK. [SPEAKER_05] When you said tooling, it's the agent.

40:15

SPEAKER_01

[SPEAKER_05] Internal tools that, I don't know, Which method will you choose? [SPEAKER_05] And which app that you will choose in your organization is different. In the same way, what we have discussed is the approach. If you want to put it in the organization, you can use the approach and you can do it in whichever way you want. My point is more, with these, you can identify gaps in your. Uh-huh. Right. Could you use it to identify gaps when you're tooling? [SPEAKER_05] OK. [SPEAKER_05] When you said tooling, it's the agent. [SPEAKER_05] Internal tools that, I don't know, maybe a team is building for the rest of the company.

40:58

SPEAKER_01

[SPEAKER_05] So infrastructure in general, right? Maybe, yeah. Can you give me an example of how it could? Let's say that, I don't know, you build some sort of abstraction on top of Kubernetes. Uh-huh. OK. You don't want your developers to necessarily know what to do with that. OK. Then you have a different CLI or you have something, right? But then, I don't know, maybe you thought that they would list one of your custom applications or corporate applications one by one. But a team has grown into using more of that and suddenly they have a lot and they don't want to do that many calls. Or perhaps even the agent is saying, well, this is inefficient.

41:35

SPEAKER_01

I would like this internal tool to work in a different way. OK. And in that way, you would identify a gap in the tool or a performance improvement.

41:48

SPEAKER_01

[SPEAKER_04] This task for documentation could be extended, actually. [SPEAKER_04] So because we have seen the business processes also, right? [SPEAKER_04] So it can also document business process. The business process is nothing but how the process and the application actually runs and does things, right?

42:05

SPEAKER_01

So you can extend it to also find out the gaps in the business process or how it works. It could be an extension to it. Yeah.

42:30

SPEAKER_01

All right. Thank you. How do you ensure that many things won't kill you? Sorry? How do you ensure that many things won't kill you? Because the knowledge of our company changes over time. Yeah. So if the answer for a question today is B, tomorrow it would be one. On Friday, it could be C, right? Mm-hm. [SPEAKER_04] You have to at C, you have to identify B1 and B and override them.

43:22

SPEAKER_01

[SPEAKER_04] If it's stored as text, that's a problem. [SPEAKER_04] OK.

43:30

SPEAKER_01

[SPEAKER_04] So when I showed the context gap scanner, you also saw an indicator of duplication, right?

43:33

SPEAKER_00

[SPEAKER_04] So if today you have a document, tomorrow you have version 2.0, and something else, it will find out the same information that the duplication is having in three different places. It also will find it.

43:34

SPEAKER_01

[SPEAKER_04] If you have only one and it is changed, it will take the latest updated one because, as a human, you changed it. So it will take it as a source of truth, right? [SPEAKER_04] But if you have three versions of the same document, that's a duplication, and it will flag it as duplicate. [SPEAKER_04] Right.

43:34

SPEAKER_04

But it's a search problem.

43:34

SPEAKER_01

[SPEAKER_04] How do you ensure performance?

43:34

SPEAKER_04

[SPEAKER_01] Well, let's say you have a document of 100,000 words. [SPEAKER_01] Uh-huh. [SPEAKER_01] If you just change a word, a passport, for example, right? [SPEAKER_01] It won't be there, but it has changed.

43:35

SPEAKER_01

OK. So you have to find this specific word, compare those three documents, type them and replace them, right? How is that feasible? Which tools would you use to ensure it will not be cost-wise? I didn't quite get the question actually. Is it the token usage you're worried about? How many tokens we used? What is the cost-saving? Yeah, precisely. If it will grow and if there are small changes, you have to maintain that whole database, right? So you have to have a structure, maybe a graph or whatever, right? My question is, what you presented is a happy path where you have a gap, you do it, and then you reuse it.

43:35

SPEAKER_01

But in a while, you will have a bigger problem where you pretend to have that gap served, but actually it doesn't contain up-to-date information. [SPEAKER_03] It contains wrong information, right? [SPEAKER_03] So you want to preserve that. [SPEAKER_03] Okay, okay. [SPEAKER_03] So it can flag as per when it is created or the last updated. You can set such kind of filters. But let's say you have a latest document which has wrong information, right? But people know that, right? No, that's true. But, for example, as a human being, right?

43:35

SPEAKER_01

So you go and look into documentation, you told somebody to look into documentation, and the person looked into the documentation and asked about the documentation. This is being permitted in this way. The person will do it, right? It's not an agent or a human issue. Way too late, right? [SPEAKER_06] Okay. [SPEAKER_06] You solved one ticket, one jury ticket, and the agent is trying to solve the second one. The assumption is the knowledge is pristine. Okay. So it's solving it, but the solution is wrong, not too late, right? Okay. So you say there is a scanner, fine. Do you run it daily? How much will it cost?

43:35

SPEAKER_01

You can have another process that would try to see last updated, try to see cadence, try to run a process that will update the knowledge. [SPEAKER_06] Fantastic. How much will it cost? For example, you're saying that less than doing more meetings every day to onboard someone or update the knowledge. That's a true claim. But I don't think it will cost that much. It's the whole premise of AI. [SPEAKER_08] As I said, when I tested it, there is none of the domains which cross more than 100k tokens, actually. So I don't think we will, for example, contact gap scanner, right? [SPEAKER_06] Try to run a process that will update the knowledge. Fantastic. How much will it cost?

43:57

SPEAKER_01

For example, sorry, you're saying that less than doing more meetings every day to onboard someone or update the knowledge. That's a true claim. But I don't think it will cost that much. It's the whole premise of AI. [SPEAKER_08] As I said, when I tested it, there is none of the domains which cross more than 100k tokens, actually. So I don't think we will, for example, contact gap scanner, right? I don't think you have to do it on a daily basis or anything. Even if you run daily basis like 100 tokens and do one scan. [SPEAKER_08] For example, right, if you try to start hitting all of them, all of you, the contact gap scanner, I think you can't even burn like $1.

45:11

SPEAKER_01

[SPEAKER_08] I think so, if I'm not wrong. [SPEAKER_08] It already had, oh, you're ready. Okay. I'll cancel the subscription. But I depend on two things. One is the complex organization insurance is all. I believe that airline industry we're talking about prepositioning systems documentation for parties. [SPEAKER_09] So I agree with you, you're trying specific domain. [SPEAKER_09] So the moment you scale, you will have to solve this cost question.

46:18

SPEAKER_01

[SPEAKER_09] Which is going to depend also how fast the data change. [SPEAKER_09] Which I think most of the time, not that much. [SPEAKER_09] The moment you get to like 80%, 90%, you just continue to evolve in line. [SPEAKER_07] Yep. I see it's different use case. Use case by use case. Yep. Any other questions? Yeah, I was curious, so I ran this camera for a bit. Okay. It has a bunch of recommendations. How do I know that it's enough?

47:31

SPEAKER_01

It actually makes it easy to start. How did you accept that? It actually tries to detail out as much as possible. Right now, I haven't actually exposed everything what it did, just for the UI purpose. But all the per ticket what it actually found, it writes like 100 or 150 lines, it's a list of markdown files and save it somewhere. So that gives you more details in case if you want to know actually. So for the demo I just put the nice UX stuff on top of it. But you also have detailed information at the desk. So I think it's a good question. [SPEAKER_07] Anyone else has any questions? [SPEAKER_07] Can I have another one? [SPEAKER_07] Oh, yeah, sure. [SPEAKER_07] Go on.

48:36

SPEAKER_01

[SPEAKER_07] Right. [SPEAKER_07] So you're using an easy one, huh? [SPEAKER_07] I'm just trying. [SPEAKER_07] Yeah, I mean, you started with a job where claims... [SPEAKER_07] Yeah. [SPEAKER_07] And if I got in the presentation right, your claim is not just... [SPEAKER_07] I use that a couple of weeks, right? A couple of weeks that you will have not fill the gaps, but discover. [SPEAKER_07] But if you scope down to a team, within weeks you can do it. Right.

49:48

SPEAKER_01

Because I mean, if you wouldn't fill it and you would have to keep asking those questions, then it doesn't make sense, right? So at some point you have the knowledge base that's greater, available and delivered by.

50:04

SPEAKER_04

[SPEAKER_01] Yeah.

50:07

SPEAKER_01

So you do it one time, first of all, or multiple times at first.

50:19

SPEAKER_01

See the whole picture, first of all. What is the state of your knowledge base?

50:35

SPEAKER_01

First, fix it at that level. Then you go into operations, right? You still can actually also continue doing it with the agent with skills. But at some point you were squeezing that for example, right? [SPEAKER_07] Yep. [SPEAKER_07] And I think the same process of providing that knowledge because the knowledge you say is in both heads, right? Uh-huh. [SPEAKER_07] So what is the thing that happens pretty frequently when a new member joins the organization?

51:46

SPEAKER_01

[SPEAKER_07] So wouldn't a replacement to that be just get Zoom calls, transfer them and use them as a source, assuming all the calls for all the knowledge transfers or maybe introductions happen over Zoom or Teams or whatever the communication, if there is a new member joining asking questions, and you have access to that, you also have all sorts of things.

51:51

SPEAKER_01

[SPEAKER_07] Yeah, yeah. [SPEAKER_07] Same problem. Okay. So you mean you can also give all the transcripts rather than doing the cycle, you mean? [SPEAKER_07] Yeah. That can also be done if you only have all the time you have a discussion in the meeting, everything is documented in meetings transcript itself. [SPEAKER_07] But I don't think this same case for everyone at least sure, but that should be easier to solve, right? [SPEAKER_07] Because if I came across something that I don't know or I'm sure about, I will call someone.

52:42

SPEAKER_01

[SPEAKER_07] I think the amount of time people spending in Teams, if you do use the transcripts actually, those are the ones actually who have more tokens actually. [SPEAKER_07] There are so many useless meetings, that transcripts actually could be, again, it depends upon institution to institution, right? [SPEAKER_07] So are you more into meetings, have solving problems within the conversations, and those conversations has the data or your Confluence or things has the data.

53:10

SPEAKER_01

If you have it, use those transcripts as your knowledge base.

53:23

SPEAKER_01

And at the same time, the compression actually that works actually, that is more useful. Yeah.

53:49

Anyone else? Any questions? No? All good. Then, thank you so much for adding this option. Thank you. also it updates with a date and also the state of the document, like stale, active, and clean. So it also looks into, okay, this is stale,

54:28

I'm not gonna touch it, and I'll just go to, just look for any of the new other documents are there in this one. Okay, thank you. Do you think about how to manage access or permissions later to this knowledge? Like if you have some knowledge in the company that all the specific people can get access to it, is it correct that you just have all the knowledge and everything that's accessible? Okay, so because it's not a product or a SaaS solution, it's just basically GitHub. For me, right now, permissions and things are not difficult to implement because GitHub out of the box gives me who I can give the permission to is GitHub,

55:04

SPEAKER_05

who can have write, read, access things and all, who can merge, those things and all. But in case if it evolves into a product, and for example, context, Gap Scanner as a product, and I want you to test it, because the reason why I was using presets for this workshop, not actually asking you to upload the files, is because I don't want to take your IP data on this one, right? So unless it becomes a product, you don't have any problem, GitHub and all, but if you have a SaaS solution for this one, then it's between how the SaaS solution will manage it right now. But the approach has nothing to do with access things and all.

55:43

So how you implement those access on the knowledge is up to you. You have a question? Yes. So this is of course about documentation, but did you give any consideration about using it on some central tooling that a company would use? Like let's say that you have a platform team and you have a CLI that the different teams are using, and so now it's used by different agents, right? Okay. So the agents can also be like, well, this action is available for a resource, but I don't want to do 500 calls, yes, because I have at least 500 resource. Okay. It would be nice if the tool could do that. I don't know if you've given any consideration to me.

56:21

SPEAKER_05

I think that is how it has to work in an organism. You need to have a central solution for it. But how you want to do the solution is up to the organization. For example, we are doing Agile, right? So Agile can do by Scrum, Kanban or like Lean or something, and also you can do different apps to do it. The process is the same, but how you do it,

56:42

SPEAKER_01

which method you will choose,

56:43

SPEAKER_05

and which app that will you choose in your organization is different.

56:47

SPEAKER_01

In the same way, what we have discussed is the approach. If you want to put it in the organization, you can use the approach and you can do it in whichever way you want. My point is more like, so with these, you can identify gaps in your . Uh-huh. Right. Could you use it to identify gaps when you're tooling?

57:09

SPEAKER_05

OK. When you said tooling, it's the agent. Internal tools that, I don't know, maybe a team is building for the rest of the company. So infrastructure in general, right?

57:18

SPEAKER_01

Maybe, yeah. Can you give me an example of like how it could. Let's say that, I don't know, you build some sort of abstraction on top of Kubernetes. Uh-huh. OK. You don't want your developers to necessarily know what to do with that. OK. Then you have a different CLI or you have something, right? But then, like I say, I don't know, maybe you thought that they would list one of your custom applications, or corporate applications, one by one. But a team has grown into using more of that and suddenly they have a lot and they don't want to do that many calls. Or perhaps even the agent is like, well, this is inefficient.

57:52

I would like this internal tool to work in a different way. OK. And in that way, you would identify like gap in the tool or a performance improvement, kind of like this task for documentation. Could be extended, actually. So because we have seen the business processes also, right? So it can also document business process.

58:13

SPEAKER_01

The business process is nothing but how the process and the application, it actually runs and does things, right? So you can extend it to also find out the gaps in the business process or like how it works. It could be an extension to it. Yeah. All right. Thank you. How do you ensure that many things won't kill you? Sorry? How do you ensure that many things won't kill you? Because the knowledge of our company changes the time. Yeah. So if the answer for a question today is B, tomorrow would be one. On Friday, it could be C, right? Mm-hm.

58:49

SPEAKER_04

You have to at C, you have to identify B1 and B and override them. If it's stored as text, that's a problem. OK. So when I showed the context gap scanner, you also saw like an indicator of duplication, right? So if today you have a document, tomorrow you have version 2.0, and something else actually, it will find out the same information that the duplication is having in three different, it also will find it. If you have only one, it is changed, it will take the latest updated one because as a human, you changed it.

59:22

SPEAKER_01

So it will take it as a source of truth, right?

59:25

SPEAKER_04

But if you have three versions of the same document, that's a duplication, and it will flag it as duplicate. Right. But it's a search problem. How do you ensure performance?

59:39

SPEAKER_01

Well, let's say you have a document of 100,000 words. Uh-huh. If you just send, you just change a word, like a passport, let's say. It won't be there, but for example, right? It has changed. OK. So you have to find this specific word, compare those three documents, type them and replace them, right? How is feasible? Which tools would you use to ensure it will not be cost-wise? I didn't quite get to the question actually. Is it like the token usage you're worried about, like that many tokens that we used? How is, what is the cost-saving? Yeah, precisely, if it will grow, and if there are little changes, you have to maintain that whole database, right?

1:00:31

SPEAKER_01

So you have to have a structure, I don't know, maybe grab or whatever, right? My question is, well, because what you presented is sort of a happy path, where you have a gap, you do it, and then you reuse it. But in a while, you will have a bigger problem where you pretend to have that gap served, but actually it doesn't contain up-to-date information to be done,

1:00:59

SPEAKER_03

it contains wrong information, right? So you want to preserve that. Okay, okay. So it can flag as per when it is created or the last updated.

1:01:12

SPEAKER_01

You can set such kind of filters. But let's say you have a latest document which has the wrong information, right? But people know that, right? No, that's true. But for example, as a human being, right? So you go and look into a documentation, you told somebody to look into documentation, and the person looked into the documentation and asked for the documentation, this is being permitted in this way, the person will do it, right? It's not an agent or a human issue. Way too late, right?

1:01:37

SPEAKER_06

Okay. You solved one ticket, one juror ticket, and the agent is trying to solve the second one.

1:01:43

SPEAKER_01

The assumption is the knowledge is pristine. Okay. So it's solving it, but the solution is wrong, not too late, right? Okay. So you say there is a scanner, fine. Do you run it daily? How much will it cost? You can have another process that would try to see last updated, try to see also cadence,

1:02:05

SPEAKER_06

try to run a process that will update the knowledge.

1:02:08

SPEAKER_01

Fantastic. How much will it cost? For example, sorry, you're saying that. Less than doing more meetings every day to onboard someone or update the knowledge. That's a true claim. But I don't think it will cost that much. It's the whole premise of AI.

1:02:26

SPEAKER_08

As I said, like, when I tested it, there is none of the domains which cross more than 100k tokens, actually.

1:02:33

SPEAKER_01

So I don't think we will, for example, contact gap scanner, right? You, I don't think like you have to do it like on a daily basis or anything. Even if you run daily basis like 100 tokens and do one scan.

1:02:45

SPEAKER_08

For example, right, if you try to start hitting all of them, all of you, the contact gap scanner, I think you can't even burn like $1. I think so, if I'm not wrong. It already had like, oh, you're ready.

1:02:58

SPEAKER_01

Okay. I'll cancel the subscription. But I- Because I go back and depend on two things. One is the complex organization insurance is all. Like, I believe that like, airline industry. We're talking about like, prepositoring systems documentation for parties.

1:03:16

SPEAKER_09

So I agree with you, you're trying to specific domain. So the moment you scale, you will have to solve this cost question. Okay. Which is going to depend also how fast the, to this point, how fast the data change. Which I think most of the time, not that much. The moment you get to like 80%, 90%, you just continue to evolve in the line.

1:03:39

SPEAKER_07

Yep.

1:03:39

SPEAKER_01

I see it's different use case. I'm going to use case. Yep. Use case by use case. Yep. Any other questions? Yeah, I was curious, so I ran this camera for like . Okay. It has a bunch of recommendations. How do I know that it's enough? It actually makes it easy to start. How did you accept that? It actually tries to detail out as much as possible. Right now, I haven't actually exposed everything what it did, just for the UI purpose. But all the per ticket what it actually found, like it writes like 100 or like 150 lines, it's a list of markdown files and save it somewhere. So, that gives you more details in case if you want to know actually.

1:04:21

SPEAKER_01

So, for the demo I just put the, you know, nice UX stuff on top of it. But you also have detailed information at the desk. So, I think it's a good question.

1:04:30

SPEAKER_07

Anyone else has any questions? Can I have another one? Oh, yeah, sure. Go on. Right. So, you're using- Easy one, huh? I'm just trying. Yeah, I mean, you started with a job where claims or . Yeah. And if I got in the presentation right, your claim is .

1:04:56

SPEAKER_07

Not just- I use that a couple of weeks, right?

1:04:58

SPEAKER_01

A couple of weeks that you will have .

1:05:01

SPEAKER_07

Not fill the gaps, but discover. But if you scope down to a team, within weeks you can do it.

1:05:08

SPEAKER_01

Right. Because, I mean, if you wouldn't fill it and you would have to keep asking those questions, then it doesn't make sense, right? So, at some point you have the knowledge base that's greater, available and delivered by. Yeah. So, you do it one time, first of all, or like multiple times at first. See the whole picture, first of all. What is the state of your knowledge base? First, fix it at that level. Then you go into operations, right? You still can actually also-

1:05:37

SPEAKER_07

I mean- You can still continue doing it with the agent with skills.

1:05:41

SPEAKER_01

But at some point you were kind of squeeze that for example, right?

1:05:46

SPEAKER_07

Yep. And I think the same process of kind of providing that knowledge because the knowledge you say

1:05:53

SPEAKER_01

is in both heads, right? Uh-huh.

1:05:55

SPEAKER_07

So, what is the thing that happens pretty frequently when a new member joins the organization? So, wouldn't a replacement to that be just get Zoom calls, transfer them and use them as a source, assuming all the calls for all the knowledge transfers or maybe introductions happen over Zoom? Or teams or whatever the communication, like if there is a new member joining asking questions,

1:06:25

SPEAKER_01

and you have access to that, you also have the start, all sorts of things.

1:06:29

SPEAKER_07

Yeah, yeah. Same problem.

1:06:30

SPEAKER_01

Okay. So, you mean like you can also give all the transcripts rather than doing the cycle, you mean?

1:06:37

SPEAKER_07

Yeah.

1:06:38

SPEAKER_01

That can also be done if you only have the- all the time you have a discussion in the meeting,

1:06:45

SPEAKER_07

everything is documented in meetings, transcript itself. But I don't think like this same case for everyone at least- Sure, but that should be easier to solve, right? Because if I came home across a church that I don't know or I'm sure about, I will call someone. I will call someone. I think the amount of time people spending in teams, if you do use the transcripts actually, those are the ones actually who which have more tokens actually. There are so many useless meetings, that transcripts actually- Could be, again, it depends upon institution to institution, right? So, are you like more into meetings, have solving problems within the conversations,

1:07:33

SPEAKER_07

and those conversations has the data or like your confluence or things has the data.

1:07:37

SPEAKER_01

If you have it, use those transcripts as your knowledge base. And at the same time, like the compression actually that works actually, that is more useful. Yeah.

1:07:49

SPEAKER_01

Anyone else? Any questions? No? All good.

1:07:55

SPEAKER_01

Then, thank you so much for adding this option.

1:08:02

Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note