[SPEAKER_01] Hi there.
SPEAKER_01
Very nice to see you guys. It's so interesting to see. It's been such a long day and you guys still showed up. And I'm always very flattered when I'm trying to present something and people are there. It just makes you feel like what you do matters. So today the topic that I want to talk about is don't build slop, four levels of AI agent maturity. And what I'm trying to do here is that I'm going to talk about something that a lot of people brought this up to me. And there's a mass psychosis problem. So it's something where every person feels like there's this giant set of robots around all of you. They're just breezing through, doing a lot of things.
SPEAKER_01
And you're in the middle and you're so confused. You're so confused. It's like, what should I do? Should I have 15 agents ripping through all the time and I'm just vibing? Or should I just be a peasant and just read through every line of code? And it's very hard to make up your mind. I feel like you come to a place like this and you're always getting to the formal, you get a panic attack or something. But I think I want to pull back out of this. And I want to say, guys, let's slow down. Let's just slow down. And let's think through what are the problems that we can actually solve that will actually help you build really useful agents.
SPEAKER_01
And I want to take care of every end of the spectrum of your necessity to build things really fast. But at the same time, your necessity to at a certain point build things slow and take things to production. So that's basically the goal. To give an example of mass psychosis, there's a lot of things that are very similar. So I'll give you an example of three UIs. And these are three frontier labs. And I guarantee you not one of you can predict which one is which. One of them is factory. One of them is codex. And one of them is cursor. I am positive none of you know which one is which. Even I don't know which one is which. But yeah, so that's basically the point.
SPEAKER_01
Everything is the same. You want to do your own thing. So to build agents in particular, here's what I would present here. I want to break down the problem of building agents into four different parts. The first part is just trying to somehow figure out a way to see if this even makes sense. If this even works. That's very user framework. Second part is when you actually are serious, okay I actually want to do something here. Now you build things by yourself, as a state machine, building an actual agent. The third part is the UX workflow where you use Kanban, which I suggest is a great form factor to be able to work with agents.
SPEAKER_01
And the fourth part would be shipping to cloud. So this is mostly an outline read of how we want to be able to work with agents. But I think it will give you a good heuristic of how you would solve this problem. So level one of building agents is literally just use a framework. And I think there's a couple different frameworks like Langchain, Langgraph. I wouldn't, I don't use them. I work with client. I'm supposed to write all the agent experience stuff by myself. So why would I use a framework? I won't be the best person to give you advice.
SPEAKER_01
But I do think that if you're trying to find PMF, if you have a problem where I want to be able to I don't know, aggregate emails or do something rudimentary. And I think an AI agent would probably be helpful here. And probably don't care about the best model. I just want something that works. Any of the agent frameworks can give you something that just works in half an hour. You can just wipe code this. It's a great thing to get started. It's a great thing to see that agents actually do work and that you can build them yourself. There are a lot of pitfalls with using frameworks.
SPEAKER_01
And one of the biggest ones that I think is that if you really want to take things to production, if you really want to build something serious, very quickly you will learn that the level of customizability, the level of futuristicness, the level of modularity that you need. You just won't find them in a framework. I know a lot of people who disagree with me and a lot of those people are wrong. Anyway. So level two is building your agents yourself. Right? So here's how you build. I feel this is a very intricate problem. I can't sum this down. So I'll just give you five rules of when you actually write code to build your agents. There are five rules that you could use.
SPEAKER_01
And these rules will give you a rough outline of how to build and write code for your own agent. So the first one is you always want to think of every agent as a state machine. State machine is a sophomore year, freshman year topic. But state machine is basically every agent is at the end of the day a recursive loop. That's basically every hypercycle, whatever you think, it's at the end of the day, it's all a recursive while loop. It's a while loop with a few conditions. And no matter what your agent is doing, it always doesn't matter if it's cursor, clock code, whatever. It is a while loop with a few conditions and a few end states.
SPEAKER_01
And what you want to be able to do is that at any point of time, you want to be able to have a mental motto of which point in the state it is. So let's say you want to be able to read a few files and explain them through clock code. So it starts from the user task at the top where you ask it to read a few files. It goes to the state of reading a few files and the action tool where it reads the file. Code, whatever. It is a while loop with a few conditions and a few end states. And what you want to be able to do is that at any point of time, you want to be able to have a mental motto of which point in the state it is.
SPEAKER_01
So let's say you want to be able to say, okay, so let's say you want to be able to read a few files and explain them through clock code. So it starts from the user task at the top where you ask it to read a few files. It goes to the state of reading a few files and the action tool where it reads the file. Then it realized, oh, I have read the file. It makes sense. And then it will call the completion tool and just complete. And the red thing at the bottom, task complete, is when the state machine finishes. You can take this in a very complex way where you can rip through the whole state machine for eight to ten minutes or even hours if you wanted to.
SPEAKER_01
But essentially every agent is a state machine. If you can visualize that as a mental model, then every time you're building an agent it will be so much easier for you to think through that. The second rule is every single thing you add to an agent risks making it worse. I think this is the hardest thing that we've learned. And this has been a very important lesson that a lot of agent builders have learned. Which is that large system prompts, lots of different edge cases, lots of different fancy if-else logic, all of that for frontier models just makes them worse. Just get out of the way of the model is the lesson that we learned.
SPEAKER_01
Which is that frontier models are so good at their job that the less instructions you give them, they actually perform better. A classic example is if you go through the codex repo, the prompt for GPT-5 versus the prompt for GPT-5.3 is one third of the size. Part of the reason for that is that the newer models are so good at their job that giving them too many instructions and longer system prompts leads to sensory overload where they get so many instructions that they get overwhelmed and can't figure out what the right thing to do is. So I think the simpler it is, it's always better. And you have to think that every single thing I'm adding, I'll be very careful.
SPEAKER_01
Then I hope to God I'm not making it worse. And just always start to prune it down. We took it so far where we literally rewrote the entirety of Klein because we realized there was so much junk from the older versions of Klein. Cloud code has been written I think at least seven times from scratch. But again, the people of the team might know better. The third rule is that you want to be able to make agent an easy part of a pseudo RL pipeline. And this is a tricky one, but basically what this means is that anytime you're building an agent, you want to be able to have some sort of a CLI.
SPEAKER_01
The reason for that is that as long as you have something that can build and test the agent really well in the form of a CLI, you want to be able to build things that are very easy to build and test with other coding agents. So right now there's this interactive dance that's happening between AI and humans where in the back of the day, humans used to guide AI like do this. And I think at this point, we're at a point where humans are being guided by AI. And this is a part of that where if you're as a human, you want to be able to build things in such a way that the AI can work very easily through it.
SPEAKER_01
So that might entail writing agents.md the right way, using the right skills. And building a CLI or CD such that the agent, other coding agents can easily build your agent, test it, make changes to it, and then test it end to end is a very critical part. Because that way, if you want to make changes, you can just let a long running agent run through in a parallel thread. It will make all those changes, test it, and you have the whole thing running. But if it's harder to build and test, it would also be harder for you to use agents to work on your agent. Super meta. Rule number four, don't build slop. Guys, don't, please, for the love of God, don't build slop.
SPEAKER_01
I think that there's so much throughput that you can get, and so many tokens that can go through really fast. I think that the best lessons that we've learned as real engineers is that it's really worth spending some time just thinking through the architecture, thinking through the design and the outline of what your agent's supposed to do, making sure it actually makes sense, and just actually, at least spend some time reading the code, even if you don't write everything by hand. I think that is super critical because the architecture point of building an agent has to be done by a human and has to be done very thoughtfully.
SPEAKER_01
Even if you're using an agent to have a conversation with it, spend a lot of time trying to think through what the architecture, what the state machine is going to be. Don't just let other models rip through the code. And then rule five is frontier labs kind of want to lock you down. And this is a tricky one. A lot of people will disagree, but basically what's happening is that a lot of times when you are working with the APIs of frontier labs, the APIs are trying to lock you down and they make the interchangeability harder.
SPEAKER_01
So to give a very precise example, the new set of models that have come out, say, Opus 4.6, Gemini 3.1 Pro, 5.3 Codex, they have this thing called reasoning traces. And reasoning traces are a part of the cache and they're also part of the reasoning test time compute loop that the model does. And when you have conversations and back and forth conversations with those models, you want to send the reasoning traces in the exact precise format that is expected. If you don't, the response would still work except that the performance would be degraded and you would have no way of knowing.
SPEAKER_01
And a lot of people are missing out on the massive performance gains that are coming from the new models because they're just not using the APIs in the exact precise way that they're supposed to be used. And there are asymmetries in the API. So some would argue that maybe you could use open router, but I don't think that's enough. I think you really want to be careful that if you're using different frontier labs and different frontier lab APIs, you have very carefully thought through if the API is working correctly and you have actually tested it.
SPEAKER_01
If you don't, the response would still work except that the performance would be degraded and you would have no way of knowing. And a lot of people are missing on the massive performance gains that are coming from the new models because they're just not using the APIs in the exact precise way that they're supposed to be used. And there are asymmetries in the API. So some would argue that maybe you could use open router, but I don't think that's enough. I think you really want to be careful that if you're using different frontier labs and different frontier lab APIs, you have very carefully thought through if the API is working correctly and you have actually tested it.
SPEAKER_01
So the form factor is the next step is how do you visualize the agents, right? So I think originally I came back to one of the previous slides, I tried to show you guys the thing where Codex and cursor and others were all looking the same. And I think I have a different claim. So on March 26, I made a tweet where I said people should use Kanban. So Kanban boards, I think if anyone has used linear, I'm sure all of you are familiar with Kanban boards. So my argument is that Kanban boards are the thing to use. Henzen in the audience was gracious enough to offer me his thoughts as well.
SPEAKER_01
Thank you so much. So Kanban boards are this idea that if you're working through an agent, you're always in front of the team. You're inference bound. A lot of you are working through Codex or Opus and it's working for eight to ten minutes at a time. When one of the agents is working for eight to ten minutes, what do you do? You could doom scroll, but you can only doom scroll so long. So then you run another agent, right? So that way you have at least two or three agents running in parallel at all times because you're inference bound. And they're all mutating the same source potentially. So you want to isolate the thing that they're mutating.
SPEAKER_01
To take care of that, the isolation of state and the inference bounding, the best UX form factor to me is Kanban board. Mainly because it takes care of it, it gives you the ability to be an engineering manager which can look at all your agents. And it gives you a headline level view of what they are. And it also helps you build flows with them where, okay, if these two tasks finish first, then I'll do the other task. And that helps you become an engineering manager and all your agents are your ICs. And you can look at them through this. So I was making this claim on March 26 and 10 hours ago, Cloud Code came out with the same thing. So I believe I was right.
SPEAKER_01
Okay, so you can use that through Cloud Code, Klein, wherever you can use Klein as well, whatever works for you. So you want to think of Kanban as basically an engineering manager, which you would be. Then there's a final step of, okay, you have an agent, you've tested it, you've made it, and you have a good UX form factor to interface with it and look at it. What do you do then? How do you have an agent that's useful, it works well, but it scales? It really scales for millions of tasks, millions of users. If you're working for a company which has 8,000 people, how do you make sure that all of them can very easily interface with this?
SPEAKER_01
And I think that rather than making people install and having these complex workflows on machines and stuff, it's so much jank and so hard. Just take it all out, put the hard work once, just take it all to the cloud. So there are many benefits of cloud agents, but the primary one is that you can completely parallelize and have a separate machine for each of them. There are no local dependencies in the cloud. The agent can set up the environment, do all the UX tasks, because right now, one of the most missing pieces, the thing that a lot of people are not using, that I use fairly extensively, is cloud agents, because they can really run for a long time.
SPEAKER_01
So I on my phone often would send tasks that would run on a cloud agent for 15 to 20 minutes. So let's say it's a UX change of, go build this VS Code extension. And this VS Code extension, I want you to sign in here. I want you to click on settings. I want you to change the settings to pick this theme. And then I want you to test this thing in the terminal. And the cloud agents are so good that they will actually manually do all the clicks of the Q&A testing that I just described on their own, figure out if they worked. And if they didn't work, it will just keep iterating, keep trying. And this thing could easily take 50 to 60 minutes.
SPEAKER_01
But if you send lots of these tasks in parallel, through your phone or through your laptop or whatever, I think you have this really easy, customizable, extensible thing. And it helps to scale really fast. And then you could send all these tasks running in the cloud machine. And you could come back to your laptop, and then it's, oh, you can just pull down the PR, and then you've got the whole thing. The other aspect of cloud agents is that if you're working with a lot of different people, I think it just helps you build a common setup that so many other people can share and so many people can mutate. So I think bringing that together would be the final form factor.
SPEAKER_01
So my claim is that there exists a future where most of the UX of working with agents would be Kanban, and most of the actual compute that's involved with agents would be on the cloud. I think these four levels of frameworks are the last part of Kanban and shipping to cloud, those things are very difficult and very intricate problems. And I think that if I were you, I would just use this as rough heuristics and start with, okay, let me just do the bare minimum easy thing. And then depending on how much effort I want to put in, I will slide up and down these levels.
SPEAKER_01
So that's it for me. And so I made a lot of hard takes, I feel. I think I left a lot of open-ended questions here. If you have any questions, if you have any thoughts, this is my Twitter. You're very welcome to reach out to me, send me any questions. And it was very kind of you, guys, to give me your time. Thank you so much. May I take a photo of you? All right. Okay. All right, thank you. Do you guys have any questions? I've got a minute. All right. All right. This is it. Oh, yeah. How do you planning inside of a Kanban? I find the back and forth requirements gathering, getting the agent to figure out what it wants from me to be the most useful part of it.
SPEAKER_01
Oh, yeah, yeah, yeah, yeah. And yeah, it was very kind of you, guys, to give me your time. Thank you so much. May I take a photo of you? All right. Okay. All right, thank you. Do you guys have any questions? I've got a minute. All right. All right. This is it. Oh, yeah. How do you plan inside of a kind of—I find the back and forth requirements gathering, getting the agent to figure out what it wants from me, to be the most useful part of it. Oh, yeah, yeah, yeah, yeah.
SPEAKER_01
Oh, yeah, let me show you. Let me show you. So I think my interpretation of Kanban is that it's just like, if you see the screen, you can go into any task and it will give you the entire trace of the task. And at that point, this interface that you're looking at is the actual CLI of Codex. And I think what I usually do is have a conversation here. And then I know at a certain point that, bro, you can go out on your own, do your thing. At that point, I'll just pull out and focus on other things. And then does it transition state when it needs your review or asks you what to do? Yes. Yes. Yes.
SPEAKER_01
So initially that is, let's say I would say if I say read a few files or whatever, right? So it will be initially in the in-progress state. And then when it needs my input, it will go through review. Yeah. All right. Thank you. All right. All right.
SPEAKER_01
And I think that if I were you, I would just like, use this like, as like rough heuristics, and start with like, okay, let me just like, let me just do the bare minimum easy thing. And then just like, depending on how much effort I want to put in, I will like slide up and down, up and down these levels. So yeah, so that's it for me. And yeah, so I made a lot of hard takes, I feel like I think I left a lot of like, open ended questions here. If you have any questions, if you have any thoughts, this is my Twitter, you're very welcome to reach out to me, send me any questions. And yeah, it was very, very kind of you, guys, to give me your time. Thank you so much.
SPEAKER_01
May I take a photo of you? All right. Okay. All right, thank you. Do you guys have any questions? It's like I've got a minute.
SPEAKER_01
All right. All right. This is it. Oh, yeah. How do you planning inside of a kind of like, I find the sort of like, back and forth requirements gathering, you know, the, the, getting the agent to figure out what it wants from me to be the most useful part of it. Oh, yeah, yeah, yeah, yeah. Oh, yeah, let me, let me show you. Let me show you. So basically, like, I think my interpretation of, of, of Kanban is that like, it's just like, if you see the screen, like, you can go into any task and it will give you like the entire trace of the task. And that point, it's like, so this, this interface that you're looking at is basically like, the actual CLI of Codex.
SPEAKER_01
And I think I would, I, what I, what I usually do is just like, have a conversation here. And then I know at a certain point that like, bro, you can go out on your own, do your thing. At that point, I'll just pull out and just focus on other things. And then does it transition state when it's either need your review or ask you what to do? Yes. Yes. Yes. So, so initially that is like, it's like, let's say I would say like, if I say read a few files or whatever, right? So it will be like initially it's in, in the in progress state. And then when it's, when it's like, it needs my input, it will go through review. Yeah.
SPEAKER_01
All right. Thank you.
SPEAKER_01
All right. All right.