SPEAKER_00
Hi everyone, so I'm Johan from Poolside. If you haven't heard of us, we are one of the handful of companies that are making their own foundational model from scratch, their own LLM and coding agents. Check us out if you're not aware of some school stuff, and you should hear more soon. But what I want to talk about today is this. You have people seeing AI and using AI and getting vastly different experiences. If you're on Reddit, if you're on Twitter, you're going to see people that say, oh yeah, I'm never touching code anymore. The AI is doing everything for me. It's fantastic. And others that say, what are you talking about? I'm trying it in my production app. It produces absolute garbage. What are you guys working on to do apps? What are you doing? And there are multiple ways to try to understand what's going on there. One way is to say, oh, this guy is a shill from OpenAI trying to sell you AI. Or this one is an entity that doesn't care about anything and is just lying. He didn't even try it. Another is to say, well, maybe someone is working on a Greenfield app, and that's nice and easy. Whereas someone else is working on Brownfield, on a legacy application. That's complicated. And agents are not there yet. But personally, I think that doesn't quite hold up. We've seen people using AI and legacy applications with good success. I have myself, so at least on my own experience, I know that can work. So what's the difference? What's the difference really between Brownfield and Greenfield? The difference is that with Greenfield, the agent's intuition is correct. The agent's writing the code and expects, if I write the components here, if I write the service, it's going to work. I think that's going to be fine. And he's right, because it has very good intuition. On Brownfield, however, there be dragons. You're going to have things that the agent is not expecting. Maybe dead ends, code that's not used anymore. Things that he's not aware of in different parts of the code that he hasn't even looked at. And that's where the big difference between those two is the feedback loop. And everybody is somewhat talked about it in the background of the talks we've seen over the past three days. But I think that's the difference with getting these results. So you've got your agent that says, yeah, I've implemented the new OAuth flow, and it's all working perfectly. What the agent really means is, well, to the best of the capabilities of my capabilities, to the best of what you have given me, that sounds like it should work. Maybe the agent was able to verify its work. Maybe it wasn't. But as far as it knows, it's working. If you're a skeptic of AI, you're going to see that first quote, see that it's indeed not working, and just see, I'm a liar, I'm incompetent, or you cannot trust the AI. And that's, I think, where this cleavage is between those two types of users. The first category is going to see, oh, yeah, actually, you know what, it's not working, agent. Can you try again? Check those logs or whatever. Whereas those ones are just going to give up. But I think we can make this still better. At Poolside, I've created a little CLI tool called Poolside. Yeah, I might be good at programming. I'm not good at naming things. That basically allows it to test our applications effectively. Some of the stuff you've already seen in things like Gistag, for instance, being able to take screenshots of the applications, being able to take a snapshot, that is to say, a very token compressed version of what's going on on a web page and use it. But our application is not a web page. It's an extension within VS Code. So already, it takes an extra step to get there.
SPEAKER_00
But with that, our AI can interface with it just like it would with a normal web page. But we take it further. We have things to extract logs from different services, from the back end, from the front end, ways to restart different services, high level commands. Can you access a specific menu? Can you go to this page? Can you send a message to the agent, wait for it to reply, send another message, upload an image and do that and it can stack things, somewhat efficiently a bit like we've talked with coding tools this morning. And that's pretty useful because then the agent is able to test what it's doing. If it's working on a bug, it can actually reproduce the bug before it starts working on that. The agents are very eager. Yeah, I know what's going on. You just need a margin there. You just need to go and add this call. But until it reproduces the bug, I don't trust you. And that's the big thing. It's maybe without that, the agent is able to still have good intuition and fix the issues. But I don't trust it. And then I'm going to have to go and verify it myself. And I'm wasting time. And I cannot then take that agent and start running it overnight, for instance. This is a first step to trust. And the point of this is not Poolside. It's not something you're going to find on GitHub to use for yourself. It's to build your own. I think, as engineers, that's our new role. We are going to have to focus less on the product and more on trying to make the AI work on the product. How can we make it easy for it? That can be those tools. That can be improving the code base so that it's easier to work on. That can be improving knowledge bases. You can implement this as a CLI, as a skill, as an MCP. In my case, it's a CLI because I like things simple. But there are many different variations of it. And I think it's going to be different from people to people and problem to problem. But yeah, I think in 2025, we had product engineers that were very focused on doing everything with the product. But now that AI is getting quite good, you want to focus more on making sure that that velocity is not a trap, that you're not going to multiply errors or compound errors, that you're going to actually verify what you're doing, making it easy for you to verify as well with presenting the work that the AI done and everything around that. And so I think we're all going to become AI ex-engineers, essentially. And yeah, it's a bit like when you're in an airline and they say, oh, you put your mask on yourself before you feed, you put your mask on your children. It's the same way they are. You need to put the mask on the AI. You need to make sure that it's self-served before you try to work on features. Even if it slows you down right now, it's an investment that pays off as soon as you start multiplying agents and running things over time. And that's me done with two minutes to spare. So thank you very much. Any questions from anybody? Yes?
SPEAKER_00
I think it was awesome. When you're thinking about what to put in the CLI, something like you have a budget-based primitive, for example. Where do you draw a micro-visual model for building the mechanical primitive versus putting together, checking in, for example, unit tests or integration tests? How ephemeral are those things that are on the other side? At the limit, checking them in and all the time?
SPEAKER_00
I think they're quite ephemeral in the sense, and that, I guess, is a personal preference, but I do feel like automated tests can be sometimes a bit too rigid and hard to predict and hard to work over time. So I like some things that mimics the way I would test it, the way a human would go and test the application. The way I find those is in the first stage where I work with the AI, even though, you know, I can see that the button is a bit to the left and I want to tell the AI that, what I want is the AI to realize that by itself. So whenever I can't realize that sort of problem, I take a step back and try to think on how to make the AI realize the problem by itself. Another thing is very attractive loop. After you've done that many times, look back, past over your logs, ask an AI, oh, did you notice any issues, any stink? Are you doing sleep, for instance, calling sleep parentheses 15 everywhere is a thing that, like, there are some things that should be weighed for whatever, like a comment that could be there. Is there anything that you keep doing over and over, running storybooks? And another thing is really think about your own product. If you're making, say, a game in Unity, do you want an ASCII representation of your 3D world for your AI? If you're making something with lots of permissions, different logins that your AI can take very easily. It's really your new role to think about all of that. That's what I think anyway. And I think we're at time. Thank you very much.
SPEAKER_00
AI and getting vastly different experiences. If you're on Reddit, if you're on Twitter, you're going to see people that say, oh yeah, I'm never touching code anymore. The AI is doing everything for me. It's fantastic. And others that say, what are you talking about? I'm trying it in my production app. It produces absolute garbage. What are you guys working on to do apps? What are you doing? And there is multiple ways to try to understand what's going on there. One way is to say, oh, this guy is a shield from OpenAI trying to sell you AI. Or this one is an entity that doesn't care about anything and is just lying. He didn't
SPEAKER_00
even try it. Another is to say, well, maybe someone is working on a Greenfield app, and that's nice and easy. Whereas someone else is working on Brownfield, on a legacy application. That's complicated. And agents are not there yet. But personally, I think that doesn't quite hold up. We've seen people using AI and legacy applications with good success. I have myself, so at least on my own experience, I know that can work. So what's the difference? What's the difference really between Brownfield and Greenfield? The difference is that with Greenfield, the agent's intuition is correct. The agent's writing the code and expect, you know, if I write the components
SPEAKER_00
here, if I write the service, it's going to work. I think that's going to be fine. And he's right, because it has very good intuition. On Brownfield, however, there be dragons. You're going to have things that the agent is not expecting. Maybe, you know, like dead ends, code that's not used anymore. Things that he's not aware of in different parts of the code that he hasn't even looked at. And that's where the big difference between those two is the feedback loop. And everybody is somewhat talked about it in the background of the talks we've seen over the past three days. But I think that's the difference with getting these results. So you've got your agent that
SPEAKER_00
says, yeah, I've implemented the new OA flow, and it's all working perfectly. What the agent means really is, well, to the best of the capa of my capabilities, to the best of what you have given me, that sounds like it should work. Maybe the agent was able to verify its work. Maybe it wasn't. But as far as it knows, it's working. If you're a skeptic of AI, you're going to see that first quote, sees that it's indeed not working, and just sees, you know, I'm a liar, I'm incompetent, or like, you cannot trust the AI. And that's, I think, where this cleavage in between those two types of users. Like, the
SPEAKER_00
first category is going to see some things. Oh, yeah, actually, you know what, it's not working, agent. Can you try again? Check those logs or whatever. Whereas those ones are just going to give up. But I think we can make this still better. At Poolside, I've created a little CLI tool called Poolside. Yeah, I might be good at programming. I'm not good at naming things. That basically allows it to test our applications effectively. Some of the stuff you've already seen in things like Gistag, for instance, being able to take screenshots of the applications, being able to take a snapshot, that is to say, like, a very token
SPEAKER_00
compressed version of what's going on on a web page and use it. But our application is not a web page. It's an extension within VS Code. So already, like, it takes an extra step to get there. But with that, our AI can interface with it just like it would with a normal web page. But we take it further. We have things to extract logs from different services, from the back end, from the front end, ways to restart different services, high level commands. Can you access a specific menu? Can you go to this page? Can you send a message to the agent, wait for it to reply, send another message, upload an
SPEAKER_00
image and do that and it can stack things, like, somewhat efficiently a bit like we've talked with coding tools this morning. And that's pretty useful because then the agent is able to test what it's doing. If it's working on a bug, it can actually reproduce the bug before it starts working on that. The agents are very eager. Yeah, I know what's going on. You just need a margin there. You just need to go and add this call. But until it reproduces the bug, I don't trust you. And that's the big thing. It's maybe without that, the agent is able to still have good intuition and fix the issues. But I don't
SPEAKER_00
trust it. And then I'm going to have to go and verify it myself. And I'm wasting time. And I cannot then take that agent and start running it overnight, for instance. This is a first step to trust. And the point of this is not full side. It's not something you're going to find on GitHub to use for yourself. It's to build your own. I think, as engineers, that's our new role. We are going to have to focus less on the product and more on trying to make the AI work on the product. How can we make it easy for it? That can be those tools. That can be improving the code base so that it's easier to work
SPEAKER_00
to work on that can be improving knowledge bases. You can implement this as a CLI, as a skill, as an MCP. In my case, it's a CLI because I like things simple. But there are many different variations of it. And I think it's going to be different from people to people and problem to problem. But yeah, I think in 2025, we had product engineers that were very focused on doing everything with the product. But now that AI is getting quite good, you want to focus more on making sure that that velocity is not a trap, that you're not going to multiply errors or compound errors, that you're going to actually verify what you're doing,
SPEAKER_00
making it easy for you to verify as well with presenting the work that the AI done and everything around that. And so I think we're all going to become AI ex-engineers, essentially. And yeah, it's a bit like when you're in an airline and they say, oh, you put your mask on yourself before you feed, you put your mask on your children. It's the same way they are. You need to put the mask on the AI. You need to make sure that it's self-served before you try to work on features. Even if it slows you down right now, it's an investment that pays off as soon as you start multiplying agents and running things
SPEAKER_00
over time. And that's me done with two minutes to spares. So thank you very much. Any questions from anybody? Yes? I think it was awesome. When you're thinking about what to put in the CLI, something like you have a budget-based primitive, for example, .
SPEAKER_00
Where do you draw a micro-visual model for building the mechanical primitive versus putting together, like checking in, for example, unit tests or integration tests? Like, how ephemeral are you things that are on the other side? At the limit, like checking them in and all the time? I think they're quite ephemeral in the sense, and that, I guess, is a personal preference, but I do feel like automated tests can be sometimes a bit too rigid and hard to predict and hard to, like, work over time. So I like some things that mimics the way I would test it, like a human would go and test the application. The way I find those is in the first stage where I work with the AI,
SPEAKER_00
even though, you know, like I can see that the button is a bit to the left and I want to tell the AI that, what I want is the AI to realize that by itself. So whenever I can't realize that sort of problem, I take a step back and try to think on how to make the AI realize the problem by itself. Another thing is very attractive loop. After you've done that many times, look back, past over your logs, ask an AI, oh, did you notice any issues, any stink? Are you doing sleep, for instance, calling sleep parentheses 15 everywhere is a thing that, like, there is some things that should be weighed for whatever, like a comment that could be there. Is there anything
SPEAKER_00
that you keep doing over and over, like running storybooks? And another thing is really think about like your own product. If you're making, say, a game in Unity, like, do you want an ASCII representation of your 3D world for your AI? If you're making something with lots of permissions, different logins that your AI can take very easily. It's really your new role to think about all of that. That's what I think anyway. And I think we're it for time. Thank you very much.