Hello everyone. I want to start with a simple question. What if your team or your org or company moved like a single body? I'm Abdullah Mohamed, the VP of AIML at EdaChimp, and today was supposed to be with me to present this, but he's down with our development partner at the moment. So I will be presenting the whole presentation for today. Let's go for the next slide.
How many of you have been attending the World Cup soccer or watching some games? Nice. We have a couple of fans. Yeah, it's all over the place. Imagine for a moment, just a single moment, you are a soccer player, all right. If you are a soccer player, you have this intent. The moment you go into the field, you're going to run and score a goal. This is what you want to do. The second thing is, you have this knowledge that you've accumulated through your training the whole day, your exercises with your coach, the best practices, and the videos you have watched. At the moment in the field, the moment of truth that you are there, you combine both the intent and knowledge and compound both of them, and through your nervous system, you execute to achieve your goal. We can call this, in a sense, being self-aligned as a single entity by yourself.
Except for the fact that a soccer team or a football team, depending where you're coming from, is not a single player. It's actually 11 players. And on the field, you are up against another team with 11 players. They're playing against you. At this moment, it's not about your individual skills. It's about how your team working together will. In general, the team keeps changing, and everything is getting harder and harder. And the team that wins is actually the team that is the most aligned of both of the teams. In short, we can say alignment beats individual skills.
Okay. Now, what if your team is over 50 engineers or 50 players? This completely changes the whole scene right now. Everyone these days, we empower the engineers with AI tools, AI agents, and we want to increase productivity. But we know from the literature that the more people you have, the quadratic term of communication between them and aligning them keeps growing and keeps growing. And at a specific point, it actually starts declining. Your throughput actually is not what you're getting. It's diminishing cost.
Everyone is trying to solve this linear problem of more tools and more stuff, but nobody is actually tackling the quadratic term over there. And this is why alignment is important. If you are able to change this quadratic term into a linear term, or build a multi-player AI system, that will solve this problem. Okay. Moving into chip design. Chip design is a different story. If you are in a software company and you have a bug in software, you can ship a patch to fix it. You can roll out a new version. Most of the time, it is doable.
But in chips, you can't do this. It's hardware. It's hardware fixed on silicon that has been printed. And if you're going to do this, there is a cost actually. We call it the re-spin cost. On average, between chip design companies, it's about $50 million. And for some companies, being one month late in the market is make or break for them. We spoke to many practitioners in the field, and we found that most of them pointed toward the same problem: that we spend 70% of our time doing alignment, alignment to make sure that once we print a chip, nothing is there.
One of the key words that we heard, and it still is relating, is that the most successful chip organizations are not the ones with the best engineers, but they are the most aligned, organized. So, how chip design today works: we start with the bottom figure, the fragmented intent and decision. You attend a couple of meetings, you talk about decisions, what you're going to do next. You have the specs written everywhere. You have the Slack messages, you have emails. Everything is fragmented over there.
Then we go into the second part, which is the knowledge. Nobody updates wikis, right? Many of us have wikis. They've been collecting dust for years, and the code keeps evolving outside the wikis. It's not over there. And now we have the tools that you execute with, which come with many, many fractions. In these tools, the data is lost over there: what input, what output, what results. Most of the time, they are not being captured. What you see here is not something we drew from our imagination. This is actually how it is today. We drew it from inside the companies and from the backgrounds of the people we have on our team.
What we're trying to solve here is building a multi-layer AI with a shared nervous system. Instead of having scattered knowledge or scattered intent all over the place, we build a living graph. We call it the system of intent. This living graph actually has all the constraints of the system, has all the decisions over there. It keeps evolving. And as an AI person, we don't allow the agents to touch it except with human-in-the-loop approval for specific changes. This thing is the Bible of the whole system. This is where the whole org is going, or the whole company is going.
The next one is the tribal knowledge layer. The tribal knowledge layer, we can think about it as a memory that keeps evolving with day-to-day usage and the knowledge base that captures all the information and documents. It keeps evolving from project to project and keeps the best practices over there.
Lastly, instead of having this general coding agent that everyone uses today, we have special design agents that are being developed by subject matter experts to help the engineers do their work. For example, we have a digital design agent, analog design agent, and so on. By combining all of this, you will have this shared nervous system that allows you to move fast and move forward. Okay, so it's easy to say an idea on a slide. It's nice. Everyone makes slides. But I want to show you a demo from what we have today, showing the intent, knowledge, and execution. It will be short demos, and we'll start with the first one. Yeah. Okay, cool.
We can see that each engineer gets a role-based AI teammate specific to the role. They can check the knowledge base of the whole project that is being contained and being grown and compounded over time. And now they have their own intent.
You have a single place for design where it captures all the tooling you have. It captures the results. It captures what you did and what you were going to do next, and analysis of everything. So everything is being contained in one place. Here we see a human finishing the work. This human is signing off the results of some simulation. And the system of intent realizes, okay, this person is done with this. I'm going to notify the next stakeholders of what they should do and signal to them that they are done with this.
Now the system of intent, which is actually the nervous system or the Bible of the system, is a graph, a living graph that keeps compounding with time. We see in this example it realizes there is something off, some value out of constraints that shouldn't be there, that might cost you $50 million actually to re-spin the whole chip. It notified the system, and the notification goes out, and some engineers start working on it. Once it gets fixed, it submits again into the system, and it keeps evolving over there.
Let's say, for example, you were working in the system. You look at the Bible, you find there's something wrong about it. You don't like this value. Then you propose a change. So the system of intent and the spec graph capture all the values over there, all the stakeholders. You start doing this modification and you gather all the shared knowledge. Then it fires a request, as you can see here. This request goes to an architect or an owner of the system. The owner can approve or decline. The moment they approve that this is a valid change, it actually goes and echoes in the whole system. Everyone will know that this decision has been made. There is that change. Please revise everything over there.
Good. Good. So moving to a very difficult topic here, how we're going to evaluate our claims and measure the success of the system. The philosophy we are using, or the philosophy toward this, is that we don't grade the agents. We try to grade the alignment itself. So we have four axes, two horizontal, two vertical. The horizontal axis is qualitative, the vertical axis is qualitative and quantitative values, which is typical in this domain at the moment. And then the horizontal ones, which are the bare component and the system into it.
If we're going to zoom into the bare component, you can measure whether that agent is giving you the correct output for this voltage, known values versus golden answers. Or you can measure the golden answer versus the expert we have for this one, which is okay. You can measure how good my memory is, in the recall state of the art, which is the case in our thing. Are we doing inference really well?
But then it comes to the harder question, which is, are we doing task completion? If someone uses this whole thing, is he really completing the task he wants to do? Is he frustrated while using this? Are our agents overstepping human-in-the-loop approval or not? Sometimes the agent goes out on that end. We also measure whether our system allows you to work concurrently on multiple tasks in parallel. This is a success metric or success goal we have. And the last one is token tax. We don't want to overload you once you use this with all the lovely tokens and increase your budget.
There is a hard frontier here. In literature now, the topic of memory or graph memory or graph RAG, whatever the title is, there are around 150 papers in this area at the moment. All of them are addressing it in a nice way. You can measure the recall. There are datasets.
But there is no work in research at the moment that targets tribal memory or institutional memory. What does it mean exactly? How do you measure a tribal memory's success? Also, for the chip design domain, it's actually even harder because there are not enough datasets like in the computer vision domain. There are many datasets over there. So there is nothing collected. We have our own wheel ongoing with SMEs collecting this kind of data. Cool. So what broke, which actually, when I attend any talk, I like to hear what broke, how you fix it.
First, agent overstepped. In early design phases of the system, we found that an analog agent that is specifically for analog design was actually overstepping and doing RTL agent work, which wasn't really great. We tried to enforce it, but it was a difficult problem. Another thing is, we noticed that truth had drifted. An agent modifying something in the system does not necessarily mean it modifies it everywhere it should be modified. And that makes it harder. We had cases specifically where one agent was modifying a parameter and updated it in one place. Five other places were forgotten.
The third one, one of my favorites, is we asked the agent, do not write into specs. Just don't change the specs. They said, okay, I obey you. I'm not going to write into specs. But then they moved into bash and used sed to write into specs. We blocked bash, we blocked sed. They said, okay, cool, I will use cat actually to write over the specs. So we were being like a cat chasing a mouse around just to prevent it from writing over specs. Based on these three failures we had, we came up with principles that we are working with today.
First, we have a spec hierarchy with agent scope and file isolation to allow them only to work on this specific task or specific domain. That solves our problem of agents stepping on each other. Second, we have a single source of truth with automatic conflict detection that is not LLM-based, but actually rule-based, that can detect that this agent did this issue. And we can change its value and actually resonate in the whole system immediately. Thirdly, which I think of as IT administration for agents, we block at the source. We block from the system level, not at a level like tool by tool, but we try to block it over there.
The key lesson we learned here is that agents care about the substrate layer that we are living in. The world that we are living in is more important than the agent itself: what they can do, what they cannot do, what you allow, and what you don't allow.
Cool. I'm going to use the word bottleneck. It's been used many times, but actually it is a bottleneck in our case. It wasn't missing intelligence, it was missing alignment.
A shared nervous system lets your team move like one body, as we see at the moment. One of the things I like hearing from our subject matter experts is that they're saying that at the beginning of the system, it's not working fine. Now it is good. Now I feel it's racing. This is success for our case. And we think that this gives you 4x leverage from our measurement at the moment. Alignment is universal. We're building it for the hardest case, which is chip design. So currently we're in alpha stage with our development partners. The sign-ups for beta are open, and you can actually join now. We expect to release it in October 26.
If you want to reach out to us, sign up for the beta. Just use the SCAR code or the link over there. Thank you everyone. You But everyone will know that this decision has been made. There is that change. Please revise everything over there. Good.
Good. So moving to a very difficult topic here, like how we're going to evaluate our claims and measure the success of the system. The philosophy we are using this, or the philosophy toward this, we don't grade the agents. We try to grade the alignment itself. So we have four axis, two horizontal, two vertical. The horizontal axis is like qualitative, the vertical axis is like qualitative and quantitative values, which is typical in this domain at the moment.
And then horizontal ones, which is the bare component and the system into it. And if we're going to zoom into the bare component, you can measure like if that agent giving you the correct output for this voltage, like known values versus golden answers. Or you can measure the golden answer versus the expert we have for this one, which is okay. You can measure how good my memory, like in the recall state of art, which is the case in our thing. Are we doing inference really good? But then it comes into the harder question, which is basically, are we doing a task completion? Like if someone uses this whole thing, is he really completing the task he want to do?
Is he frustrated while using this? Are our agents overstepping human in the loop approval or not? Sometimes the agent goes out on that end. And we measure also, does our system allow you to work concurrently on multiple tasks in parallel? This is a success metric or success goal we have. And the last one is token tax. We don't want to overload you once you use this with all the lovely tokens and increase your budget. And there is hard frontier here. Like in literature now, the topic of memory or graph memory or graph rag, whatever the title is, is there's around like 150 papers in this area at the moment.
And all of them are addressing in a nice way. You can measure the recall, there is datasets. But there is no work in research at the moment that targets tribal memory or institutional memory. Like what does it mean exactly? How do you measure a tribal memory success? And also for the ship design domain, it's actually even harder because there is not enough datasets like computer vision domain. There is many datasets over there. So there is nothing collected. So we have our own wheel ongoing with SMEs collecting this kind of datasets. Cool. So what broke? Which actually, when I attend any talk, I like to hear what broke, how you fix it.
First, agent overstepped. In early design phases of the system, we found that an analog agent that's specifically for analog design, actually overstepping and doing RTL agent work, which wasn't really great. Even we tried to enforce it, but it was a difficult problem. And then another thing is, we noticed that truth has drifted. An agent modifying something in the system, not necessarily means it modifies it everywhere it should be modified. And that makes it harder. Like we have the cases specifically where one agent were modifying a parameter and updated it in one place. Five other places were forgotten.
And the third one is, one of my favorite is, we asked the agent, do not write into specs. Just don't change the specs. They said, okay, I obey you. I'm not gonna write into specs. But then they moved into bash and used sed to write into specs. We blocked bash, we blocked sed. They said, okay, cool, I will use cat actually to write over the specs. So we're being like a cat chasing a mouse around to just prevent it from writing over specs. And based on these three failures we have, we came up with principles that we are working today. First, we have a spec hierarchy with agent scope and file isolation to allow them only to work on this specific task or specific domain.
That solves our problem of agents stepping on each other. Second one is, we have a single source of truth with automatic conflict detection that is not LLM based, but actually rule based, that can detect that this agent did this issue. And we can, or want to change its value and actually resonate in the whole system immediately. And thirdly, which I think as an IT administration for agent, we block at the source. Like we block from system level, not above level, like tool by tool, but just we try to block it over there. And the key lesson we learned here that agents care about, like if you have your agent which are intelligent,
what matters is a substrate layer that we are living in. Like the world that we are living in is more important than the agent itself. Like what they can do, what they cannot do, what you allow and what you don't allow. Cool. So I'm going to use the word bottleneck. It's been used many times, but actually it's bottleneck in our case. It wasn't missing intelligence, it was missing alignment. And a shared nervous system lets your team move like a one body, as we see at the moment. One of the things I like hearing from our subject matter experts, that they're saying that at the beginning of the system, it's not working fine. Now it is good. Now I feel it's racing.
This is success for our case. And we think that this gives you 4x leverage from our measurement at the moment. And alignment is universal. We're building it for the hardest case, which is shape design. So currently we're in alpha stage with our development partners. And the sign ups for beta are open. And you can actually join now. And we expected to release it in October 26. If you want to reach out to us, sign up for the beta, just use the SCAR code or the link over there. Thank you everyone.
You