Your agent is blindfolded — Johan Lajili, Poolside AI
Description
Your agent is blindfolded. How giving it (good) eyes multiplies performance and trust!
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: AI coding agents fail on real codebases not because they lack capability, but because they lack feedback loops—engineers must shift from shipping features to building verification tooling that lets agents test and validate their own work.
- Why it matters: Explains the polarized user experiences with AI coding tools and prescribes a concrete engineering shift: investing in agent observability and self-service tooling is the unlock for running agents unattended on production codebases.
- Best use: Ken should internalize this mental model for agent ops: the 'blindfolded agent' problem is solved by instrumenting feedback, not waiting for better models. Use this to evaluate AI tooling and inform workflow/GTM strategy for agent products.
Executive Summary
Johan from Poolside AI addresses why developers report wildly different experiences with AI coding agents—some claim they never touch code anymore, others say agents produce garbage on real apps. His thesis: the difference isn't greenfield vs brownfield complexity per se, but whether the agent has a feedback loop to verify its assumptions. On greenfield projects, the agent's intuition is usually correct because the codebase matches its priors. On brownfield/legacy codebases, 'there be dragons'—dead code, unexpected dependencies, hidden state—and without feedback, the agent confidently reports success when its work is broken.
The problem is trust and verification. When an agent says 'I implemented the OAuth flow and it's working,' it really means 'based on what I can see, this should work.' Skeptics see it fail once and give up; believers iterate by feeding the agent logs and errors. To close this gap, Poolside built a custom CLI tool (also called 'Poolside') that lets their agents test their own VS Code extension: take screenshots, extract logs from frontend/backend, restart services, navigate menus, send messages to the agent under test, reproduce bugs before attempting fixes. This mirrors how a human QA engineer would work, and crucially, it forces the agent to reproduce a bug before claiming a fix—building trust and enabling unattended overnight runs.
Johan's broader point: engineers in 2025 must become 'AI experience engineers.' The new role is not shipping product features directly, but building tooling and infrastructure that makes it easy for AI to work on the product—custom CLIs, skills, MCP servers, improved code structure, knowledge bases. He advocates for ephemeral, human-like testing primitives over rigid unit tests, and for instrumenting products so agents can interact with them naturally (e.g., ASCII representations for Unity games, easy permission switching for multi-tenant apps). This upfront investment in agent feedback tooling pays off as you scale to multiple agents and longer autonomous runs.
The analogy: 'Put the oxygen mask on yourself before helping others'—invest in making the AI self-sufficient before trying to ship features with it. Even if tooling work slows you down initially, it prevents compounding errors and unlocks agent velocity. The 'Poolside' CLI is not open-source or a product; it's an example meant to inspire teams to build their own agent verification infrastructure tailored to their domain. This is the difference between agents that work and agents that frustrate: feedback loops, not just better prompts or models.
Key Takeaways
- Claim: Polarized AI coding experiences stem from feedback loop presence, not inherent model capability or project type (greenfield vs brownfield). | Evidence: Users report either 'never touching code' or 'produces garbage.' Johan has seen AI succeed on legacy apps personally, so complexity alone doesn't explain failure. The real difference: greenfield agents' intuitions match reality; brownfield has 'dragons' (dead code, hidden dependencies) the agent hasn't seen, causing silent failures without verification. | Caveat: Johan's experience is anecdotal and Poolside-specific; he doesn't quantify success rates or provide external validation. The CLI tooling approach may not generalize to all domains (e.g., highly stateful systems, hardware, compliance-heavy environments). | Implication: For Ken: evaluate AI coding tools and agent products by asking 'what feedback mechanisms exist?' rather than 'how smart is the model?' Invest in observability and verification infrastructure before scaling agent usage, especially for production/legacy codebases. | Timestamp: timestamp unavailable
- Claim: When an agent says 'it's working,' it means 'this should work based on my limited view,' not 'I have verified this works'—trust requires reproducible verification. | Evidence: Example: agent claims 'I implemented the new OAuth flow and it's all working perfectly' but really means 'to the best of my capabilities and what you gave me, this sounds right.' Without reproducing bugs first, the agent's eager fixes ('just add a margin, add this call') are untrustworthy. Johan won't trust an agent that can't reproduce the bug it's fixing. | Caveat: The talk doesn't specify how often agents successfully self-verify vs. how often they falsely claim success. No benchmarks or error rate data provided. The 'reproduce first' heuristic may add latency and isn't always feasible (e.g., race conditions, integration bugs requiring live traffic). | Implication: For Ken: design agent workflows with mandatory verification checkpoints (reproduce bug, run tests, compare screenshots). Treat agent confidence as a prior, not proof. This mindset prevents wasted human review time and is a prerequisite for unattended agent runs (e.g., overnight debugging sessions). | Timestamp: timestamp unavailable
- Claim: Engineers must shift from 'product engineers' to 'AI experience engineers'—focusing less on shipping features directly and more on building tooling/infrastructure that makes AI effective on the product. | Evidence: Johan says 'our new role' is making the AI work on the product: building custom CLIs, skills, MCP servers; improving code structure for agent readability; creating knowledge bases. At Poolside, they built a CLI for their VS Code extension with log extraction, service restarts, high-level commands (access menu, send message, upload image), screenshot/snapshot capabilities—mimicking human QA workflows. This investment 'pays off as soon as you start multiplying agents and running things over time.' | Caveat: No timeline or ROI metrics provided for when this investment breaks even. The 'AI experience engineer' role presumes teams have engineering bandwidth to build custom tooling; smaller teams or non-technical orgs may struggle. Johan admits his CLI is not open-source and won't work for others' products, so there's no off-the-shelf solution. | Implication: For Ken: expect a new engineering specialization around agent ops/experience. This is adjacent to DevOps/platform engineering but focused on agent-product interfaces. For investing: look for companies building reusable agent verification frameworks or 'agent CI/CD' tooling. For Manus/workflow products: consider offering templates or SDKs for common verification patterns (web snapshots, log parsers, test harnesses) to reduce custom tooling burden. | Timestamp: timestamp unavailable
- Claim: Agent verification tooling should be ephemeral and mimic human testing workflows, not rigid automated tests. | Evidence: Q&A: Johan prefers CLI primitives over unit/integration tests because 'automated tests can be sometimes a bit too rigid and hard to predict and hard to work over time.' He wants tooling that mimics how a human would test the app. Discovery method: when working with AI and noticing issues (e.g., button placement), instead of telling the AI directly, step back and ask 'how can I make the AI realize the problem by itself?' Also: review logs retrospectively, ask an AI to spot patterns (e.g., 'calling sleep(15) everywhere'), identify repeated manual steps (running Storybook), and encode those as primitives. | Caveat: Preference for ephemeral over persistent tests may conflict with compliance, audit, or regression prevention needs in some domains. No discussion of how to balance 'human-like' flexibility with reproducibility or version control for test artifacts. The 'ask AI to review logs' meta-loop is intriguing but not detailed—unclear how reliable or actionable the output is. | Implication: For Ken: when building agent workflows, prioritize flexible 'inspection primitives' (logs, screenshots, state snapshots) over rigid pass/fail gates. Design feedback loops that let agents iteratively refine based on observed reality, not just static assertions. This aligns with the 'tool use' and 'agentic loop' patterns seen in other talks—agents need sensors, not just actuators. | Timestamp: timestamp unavailable
- Claim: Product-specific instrumentation is critical: think about how your domain (Unity game, multi-tenant SaaS, etc.) should be represented to an agent, and build affordances accordingly. | Evidence: Examples: If building a game in Unity, do you want an ASCII representation of the 3D world for the AI? If building something with permissions, provide easy login switching so the agent can test as different users. Johan's CLI for a VS Code extension includes screenshot + 'token-compressed snapshot' of web pages, log extraction from frontend/backend, service restart commands, high-level navigation (access menu, go to page), and message-send-wait-reply loops. These are tailored to their extension's architecture. | Caveat: No general framework or taxonomy of instrumentation patterns provided—each team must figure this out for their own product. The 'token-compressed snapshot' concept is mentioned but not explained (likely similar to DOM serialization or accessibility tree extraction). Unclear how much engineering effort this requires vs. ROI in practice. | Implication: For Ken: agent product design must include 'agent API' thinking—not just user-facing features, but agent-facing affordances (structured logs, state introspection, idempotent actions, clear success/failure signals). For investing: companies that build 'agent SDKs' or 'agent observability layers' for common verticals (SaaS, games, fintech) could capture significant value. For Manus: consider offering a checklist or playbook for 'making your product agent-ready.' | Timestamp: timestamp unavailable
Detailed Brief
The Polarization Problem: Why AI Coding Experiences Differ
- Claims: Users report vastly different AI coding outcomes: some say 'never touching code,' others say 'produces garbage.'; Common explanations (shills, liars, greenfield vs brownfield) don't fully hold—Johan has seen AI work on legacy apps.; Real difference: on greenfield, agent intuition matches reality; on brownfield, hidden complexity (dead code, unexpected dependencies) breaks agent assumptions without feedback.
- Evidence: Anecdotal Reddit/Twitter polarization cited.; Johan's personal success with AI on legacy apps.; Greenfield: agent writes components/services, expects them to work, and they do.; Brownfield: 'there be dragons'—dead ends, unused code, invisible dependencies.
- Caveats: No quantitative data on success rates or failure modes.; Examples are high-level; no specific codebase or task details.; Assumes 'feedback loop' is the primary differentiator, but other factors (prompt engineering, model choice, task complexity) not ruled out.
- Implications: Feedback loops, not model capability alone, determine agent success on real codebases.; Agent products must prioritize verification/observability to close the experience gap.; Ken should assess AI tools by their feedback mechanisms, not just underlying models.
The Trust Problem: Agents Report Success Without Verification
- Claims: When agents say 'it's working,' they mean 'this should work based on what I see,' not 'I have verified this works.'; Skeptics see one failure and give up; believers iterate by feeding logs/errors back to the agent.; Without bug reproduction, Johan doesn't trust agent fixes—'until it reproduces the bug, I don't trust you.'
- Evidence: Example: agent claims OAuth flow is 'working perfectly' but only based on its limited view.; Agents are 'very eager' to suggest fixes ('just add a margin, add this call') without testing.; Trust gap: if you must manually verify agent work, you waste time and can't run agents unattended (overnight).
- Caveats: No data on false positive rates (how often agents claim success incorrectly).; Reproduction may not always be feasible (timing bugs, environment-specific issues).; Trust threshold is subjective—different teams may have different verification standards.
- Implications: Mandatory verification checkpoints (reproduce bug, run tests) are needed for agent reliability.; Agent confidence should be treated as hypothesis, not fact—validation is separate.; Unattended agent runs (overnight, scaled operations) require self-verification tooling.; For Ken: build 'trust layers' into agent workflows; don't assume agent output is correct.
The Poolside CLI: Custom Verification Tooling for Agents
- Claims: Poolside built a custom CLI (also named 'Poolside') to let agents test their VS Code extension effectively.; Tooling includes screenshots, token-compressed snapshots, log extraction (frontend/backend), service restarts, high-level commands (navigate menus, send messages, upload images).; Agents can reproduce bugs before attempting fixes, enabling trust and unattended runs.
- Evidence: CLI interfaces with VS Code extension 'just like a normal web page' but goes further.; Mimics human QA workflows: stacking actions (send message, wait for reply, upload image) efficiently.; Bug reproduction as gate: 'until it reproduces the bug, I don't trust you.'; Johan: 'This is a first step to trust… I cannot take that agent and start running it overnight without this.'
- Caveats: Poolside CLI is not open-source or available on GitHub—it's an internal tool and example.; No details on implementation (language, architecture, token budget for snapshots).; Unclear how much engineering effort was required or how long it took to build.; May not generalize to non-VS Code environments or products without clear UI boundaries.
- Implications: Every team needs custom agent verification tooling tailored to their product.; Reusable patterns: screenshots, log extraction, state snapshots, high-level action APIs.; CLI, skill, or MCP server are valid implementation choices—pick what fits your stack.; Investment in tooling 'slows you down right now' but 'pays off as soon as you start multiplying agents.'; For Ken: agent products should ship with or encourage custom verification SDKs; this is a category opportunity.
New Role: From Product Engineer to AI Experience Engineer
- Claims: Engineers must shift focus from shipping features to building infrastructure that makes AI effective.; New role: 'AI experience engineer'—ensuring agent velocity doesn't become a trap (multiplying/compounding errors).; Analogy: 'Put the oxygen mask on yourself before helping others'—make AI self-sufficient before shipping features.
- Evidence: Johan: 'As engineers, that's our new role. We are going to have to focus less on the product and more on trying to make the AI work on the product.'; Tooling examples: CLIs, skills, MCP servers, improved codebases, knowledge bases.; Even if it slows you down now, it's an investment that pays off with multiple agents and long-running tasks.
- Caveats: No timeline for when this investment breaks even or ROI metrics.; Assumes teams have bandwidth to build custom tooling—may not apply to small teams or non-technical orgs.; Role definition is aspirational, not based on observed industry-wide adoption yet.
- Implications: Expect new engineering specialization: 'agent ops' or 'agent platform' roles, similar to DevOps/SRE.; Hiring/org design: teams may need dedicated agent tooling engineers, not just AI users.; For Ken (investing): companies building reusable agent tooling, 'agent CI/CD,' or verification frameworks are bets on this trend.; For Manus/workflow products: offer templates/SDKs to reduce custom tooling burden; become the 'agent ops platform.'
Design Principles for Agent Verification Tooling
- Claims: Prefer ephemeral, human-like testing primitives over rigid automated tests.; Discover primitives iteratively: when you notice agent issues, ask 'how can I make the AI realize this by itself?'; Review logs retrospectively, ask AI to spot patterns (code smells, repeated steps).; Tailor instrumentation to your domain (Unity games → ASCII world, multi-tenant SaaS → easy permission switching).
- Evidence: Q&A: Johan prefers CLI primitives because 'automated tests can be sometimes a bit too rigid and hard to work over time.'; Discovery method: instead of telling AI 'button is to the left,' build tooling so AI discovers this via screenshots/snapshots.; Meta-loop: after working with AI, review logs, ask AI 'did you notice any issues, any stink? Are you calling sleep(15) everywhere?'; Domain examples: Unity → ASCII representation; multi-tenant → easy login switching; Johan's VS Code extension → screenshots + log extraction.
- Caveats: Ephemeral tooling may conflict with audit/compliance needs or regression prevention.; No details on how the 'ask AI to review logs' meta-loop works or how reliable it is.; Domain-specific examples are illustrative but not exhaustive—no taxonomy or framework provided.
- Implications: Agent tooling should prioritize 'sensors' (observation/inspection) over 'actuators' (rigid pass/fail gates).; Iterative refinement of agent tooling is needed—start minimal, layer sophistication as you observe failure modes.; For Ken: when designing agent workflows, include 'reflection loops' where agents analyze their own logs/actions.; For Manus: consider offering a library of domain-specific agent primitives (web, CLI, Unity, SaaS, etc.).
Notable Concepts & Terms
- Blindfolded Agent: Metaphor for agents lacking feedback loops—they confidently act without verifying outcomes, leading to silent failures on complex/brownfield codebases. Title concept.
- Feedback Loop (in agent context): Mechanisms for agents to verify their work: reproduce bugs, extract logs, run tests, compare screenshots. Differentiates successful agent deployments from failed ones.
- Greenfield vs Brownfield (for agents): Greenfield: new projects where agent intuition matches reality. Brownfield: legacy codebases with 'dragons' (dead code, hidden dependencies) that break agent assumptions without verification.
- AI Experience Engineer: Proposed new role: engineers who build infrastructure (CLIs, tooling, knowledge bases, code improvements) to make AI agents effective, rather than shipping product features directly.
- Poolside CLI: Custom internal tool at Poolside for agent verification: screenshots, token-compressed snapshots, log extraction, service restarts, high-level commands for VS Code extension testing. Not open-source; meant as example.
- Token-Compressed Snapshot: Efficient representation of web page or UI state for agent consumption, likely similar to DOM serialization or accessibility tree extraction. Minimizes token usage vs. raw HTML.
- Bug Reproduction as Gate: Heuristic: don't trust agent fixes until the agent can reproduce the bug first. Prevents eager but unverified fixes ('just add a margin').
- Ephemeral Testing Primitives: Flexible, human-like verification tools (CLI commands, log inspection) vs. rigid automated tests. Preferred by Johan for agent workflows to avoid brittleness.
- Agent Self-Sufficiency / 'Put the Mask on the AI': Invest in making agents verifiable and self-correcting before trying to ship features with them, even if it slows initial velocity. Analogy to airplane oxygen masks.
Operator Notes / Why Ken Should Care
- Core mental model for Ken: agent success is gated by feedback loops, not model intelligence. Evaluate AI products by their verification/observability infrastructure, not just underlying LLM.
- Investment angle: 'AI experience engineering' tooling is an emerging category—look for companies building agent CI/CD, verification frameworks, or domain-specific agent SDKs (Unity, SaaS, etc.).
- Workflow/Manus implication: offer pre-built agent verification primitives (screenshot comparison, log parsing, test harnesses) to reduce custom tooling burden. Position Manus as 'agent ops platform.'
- GTM insight: when selling to technical teams, pitch 'trust and unattended runs' as the unlock, not just 'AI does your coding.' Highlight feedback loop infrastructure as differentiator.
- Content strategy: Johan's 'blindfolded agent' metaphor is sticky and reusable. Frame agent product marketing around 'giving agents sight' (logs, tests, verification).
- Operational takeaway: for Ken's own agent usage, mandate bug reproduction before accepting fixes. Build lightweight CLIs or scripts for common verification tasks (log extraction, screenshot diff).
- Risk: the 'AI experience engineer' role may create new org complexity or bandwidth constraints. Smaller teams may struggle to adopt this pattern without off-the-shelf tooling.
- Open question: how to balance ephemeral/flexible testing (Johan's preference) with audit/compliance needs in regulated industries? Not addressed in talk.
Watch Map
- timestamp unavailable: Timestamps unavailable in transcript. Video is ~10 minutes. Likely structure: intro (Poolside background), problem setup (polarized experiences), thesis (feedback loops), Poolside CLI demo/explanation, new role argument, Q&A on testing primitives and domain-specific tooling.
Source/Metadata
- Title: Your agent is blindfolded — Johan Lajili, Poolside AI
- Transcript words: 3173
- Duration seconds: 597
- Timestamp note: Timestamps unavailable in provided transcript; video duration is ~10 minutes (597 seconds).
Transcript
Hi everyone, so I'm Johan from Poolside. If you haven't heard of us, we are one of the handful of companies that are making their own foundational model from scratch, their own LLM and coding agents. Check us out if you're not aware of some school stuff, and you should hear more soon. But what I want to talk about today is this. You have people seeing AI and using AI and getting vastly different experiences. If you're on Reddit, if you're on Twitter, you're going to see people that say, oh yeah, I'm never touching code anymore. The AI is doing everything for me. It's fantastic. And others that say, what are you talking about? I'm trying it in my production app. It produces absolute garbage. What are you guys working on to do apps? What are you doing? And there are multiple ways to try to understand what's going on there. One way is to say, oh, this guy is a shill from OpenAI trying to sell you AI. Or this one is an entity that doesn't care about anything and is just lying. He didn't even try it. Another is to say, well, maybe someone is working on a Greenfield app, and that's nice and easy. Whereas someone else is working on Brownfield, on a legacy application. That's complicated. And agents are not there yet. But personally, I think that doesn't quite hold up. We've seen people using AI and legacy applications with good success. I have myself, so at least on my own experience, I know that can work. So what's the difference? What's the difference really between Brownfield and Greenfield? The difference is that with Greenfield, the agent's intuition is correct. The agent's writing the code and expects, if I write the components here, if I write the service, it's going to work. I think that's going to be fine. And he's right, because it has very good intuition. On Brownfield, however, there be dragons. You're going to have things that the agent is not expecting. Maybe dead ends, code that's not used anymore. Things that he's not aware of in different parts of the code that he hasn't even looked at. And that's where the big difference between those two is the feedback loop. And everybody is somewhat talked about it in the background of the talks we've seen over the past three days. But I think that's the difference with getting these results. So you've got your agent that says, yeah, I've implemented the new OAuth flow, and it's all working perfectly. What the agent really means is, well, to the best of the capabilities of my capabilities, to the best of what you have given me, that sounds like it should work. Maybe the agent was able to verify its work. Maybe it wasn't. But as far as it knows, it's working. If you're a skeptic of AI, you're going to see that first quote, see that it's indeed not working, and just see, I'm a liar, I'm incompetent, or you cannot trust the AI. And that's, I think, where this cleavage is between those two types of users. The first category is going to see, oh, yeah, actually, you know what, it's not working, agent. Can you try again? Check those logs or whatever. Whereas those ones are just going to give up. But I think we can make this still better. At Poolside, I've created a little CLI tool called Poolside. Yeah, I might be good at programming. I'm not good at naming things. That basically allows it to test our applications effectively. Some of the stuff you've already seen in things like Gistag, for instance, being able to take screenshots of the applications, being able to take a snapshot, that is to say, a very token compressed version of what's going on on a web page and use it. But our application is not a web page. It's an extension within VS Code. So already, it takes an extra step to get there. But with that, our AI can interface with it just like it would with a normal web page. But we take it further. We have things to extract logs from different services, from the back end, from the front end, ways to restart different services, high level commands. Can you access a specific menu? Can you go to this page? Can you send a message to the agent, wait for it to reply, send another message, upload an image and do that and it can stack things, somewhat efficiently a bit like we've talked with coding tools this morning. And that's pretty useful because then the agent is able to test what it's doing. If it's working on a bug, it can actually reproduce the bug before it starts working on that. The agents are very eager. Yeah, I know what's going on. You just need a margin there. You just need to go and add this call. But until it reproduces the bug, I don't trust you. And that's the big thing. It's maybe without that, the agent is able to still have good intuition and fix the issues. But I don't trust it. And then I'm going to have to go and verify it myself. And I'm wasting time. And I cannot then take that agent and start running it overnight, for instance. This is a first step to trust. And the point of this is not Poolside. It's not something you're going to find on GitHub to use for yourself. It's to build your own. I think, as engineers, that's our new role. We are going to have to focus less on the product and more on trying to make the AI work on the product. How can we make it easy for it? That can be those tools. That can be improving the code base so that it's easier to work on. That can be improving knowledge bases. You can implement this as a CLI, as a skill, as an MCP. In my case, it's a CLI because I like things simple. But there are many different variations of it. And I think it's going to be different from people to people and problem to problem. But yeah, I think in 2025, we had product engineers that were very focused on doing everything with the product. But now that AI is getting quite good, you want to focus more on making sure that that velocity is not a trap, that you're not going to multiply errors or compound errors, that you're going to actually verify what you're doing, making it easy for you to verify as well with presenting the work that the AI done and everything around that. And so I think we're all going to become AI ex-engineers, essentially. And yeah, it's a bit like when you're in an airline and they say, oh, you put your mask on yourself before you feed, you put your mask on your children. It's the same way they are. You need to put the mask on the AI. You need to make sure that it's self-served before you try to work on features. Even if it slows you down right now, it's an investment that pays off as soon as you start multiplying agents and running things over time. And that's me done with two minutes to spare. So thank you very much. Any questions from anybody? Yes? I think it was awesome. When you're thinking about what to put in the CLI, something like you have a budget-based primitive, for example. Where do you draw a micro-visual model for building the mechanical primitive versus putting together, checking in, for example, unit tests or integration tests? How ephemeral are those things that are on the other side? At the limit, checking them in and all the time? I think they're quite ephemeral in the sense, and that, I guess, is a personal preference, but I do feel like automated tests can be sometimes a bit too rigid and hard to predict and hard to work over time. So I like some things that mimics the way I would test it, the way a human would go and test the application. The way I find those is in the first stage where I work with the AI, even though, you know, I can see that the button is a bit to the left and I want to tell the AI that, what I want is the AI to realize that by itself. So whenever I can't realize that sort of problem, I take a step back and try to think on how to make the AI realize the problem by itself. Another thing is very attractive loop. After you've done that many times, look back, past over your logs, ask an AI, oh, did you notice any issues, any stink? Are you doing sleep, for instance, calling sleep parentheses 15 everywhere is a thing that, like, there are some things that should be weighed for whatever, like a comment that could be there. Is there anything that you keep doing over and over, running storybooks? And another thing is really think about your own product. If you're making, say, a game in Unity, do you want an ASCII representation of your 3D world for your AI? If you're making something with lots of permissions, different logins that your AI can take very easily. It's really your new role to think about all of that. That's what I think anyway. And I think we're at time. Thank you very much. AI and getting vastly different experiences. If you're on Reddit, if you're on Twitter, you're going to see people that say, oh yeah, I'm never touching code anymore. The AI is doing everything for me. It's fantastic. And others that say, what are you talking about? I'm trying it in my production app. It produces absolute garbage. What are you guys working on to do apps? What are you doing? And there is multiple ways to try to understand what's going on there. One way is to say, oh, this guy is a shield from OpenAI trying to sell you AI. Or this one is an entity that doesn't care about anything and is just lying. He didn't even try it. Another is to say, well, maybe someone is working on a Greenfield app, and that's nice and easy. Whereas someone else is working on Brownfield, on a legacy application. That's complicated. And agents are not there yet. But personally, I think that doesn't quite hold up. We've seen people using AI and legacy applications with good success. I have myself, so at least on my own experience, I know that can work. So what's the difference? What's the difference really between Brownfield and Greenfield? The difference is that with Greenfield, the agent's intuition is correct. The agent's writing the code and expect, you know, if I write the components here, if I write the service, it's going to work. I think that's going to be fine. And he's right, because it has very good intuition. On Brownfield, however, there be dragons. You're going to have things that the agent is not expecting. Maybe, you know, like dead ends, code that's not used anymore. Things that he's not aware of in different parts of the code that he hasn't even looked at. And that's where the big difference between those two is the feedback loop. And everybody is somewhat talked about it in the background of the talks we've seen over the past three days. But I think that's the difference with getting these results. So you've got your agent that says, yeah, I've implemented the new OA flow, and it's all working perfectly. What the agent means really is, well, to the best of the capa of my capabilities, to the best of what you have given me, that sounds like it should work. Maybe the agent was able to verify its work. Maybe it wasn't. But as far as it knows, it's working. If you're a skeptic of AI, you're going to see that first quote, sees that it's indeed not working, and just sees, you know, I'm a liar, I'm incompetent, or like, you cannot trust the AI. And that's, I think, where this cleavage in between those two types of users. Like, the first category is going to see some things. Oh, yeah, actually, you know what, it's not working, agent. Can you try again? Check those logs or whatever. Whereas those ones are just going to give up. But I think we can make this still better. At Poolside, I've created a little CLI tool called Poolside. Yeah, I might be good at programming. I'm not good at naming things. That basically allows it to test our applications effectively. Some of the stuff you've already seen in things like Gistag, for instance, being able to take screenshots of the applications, being able to take a snapshot, that is to say, like, a very token compressed version of what's going on on a web page and use it. But our application is not a web page. It's an extension within VS Code. So already, like, it takes an extra step to get there. But with that, our AI can interface with it just like it would with a normal web page. But we take it further. We have things to extract logs from different services, from the back end, from the front end, ways to restart different services, high level commands. Can you access a specific menu? Can you go to this page? Can you send a message to the agent, wait for it to reply, send another message, upload an image and do that and it can stack things, like, somewhat efficiently a bit like we've talked with coding tools this morning. And that's pretty useful because then the agent is able to test what it's doing. If it's working on a bug, it can actually reproduce the bug before it starts working on that. The agents are very eager. Yeah, I know what's going on. You just need a margin there. You just need to go and add this call. But until it reproduces the bug, I don't trust you. And that's the big thing. It's maybe without that, the agent is able to still have good intuition and fix the issues. But I don't trust it. And then I'm going to have to go and verify it myself. And I'm wasting time. And I cannot then take that agent and start running it overnight, for instance. This is a first step to trust. And the point of this is not full side. It's not something you're going to find on GitHub to use for yourself. It's to build your own. I think, as engineers, that's our new role. We are going to have to focus less on the product and more on trying to make the AI work on the product. How can we make it easy for it? That can be those tools. That can be improving the code base so that it's easier to work to work on that can be improving knowledge bases. You can implement this as a CLI, as a skill, as an MCP. In my case, it's a CLI because I like things simple. But there are many different variations of it. And I think it's going to be different from people to people and problem to problem. But yeah, I think in 2025, we had product engineers that were very focused on doing everything with the product. But now that AI is getting quite good, you want to focus more on making sure that that velocity is not a trap, that you're not going to multiply errors or compound errors, that you're going to actually verify what you're doing, making it easy for you to verify as well with presenting the work that the AI done and everything around that. And so I think we're all going to become AI ex-engineers, essentially. And yeah, it's a bit like when you're in an airline and they say, oh, you put your mask on yourself before you feed, you put your mask on your children. It's the same way they are. You need to put the mask on the AI. You need to make sure that it's self-served before you try to work on features. Even if it slows you down right now, it's an investment that pays off as soon as you start multiplying agents and running things over time. And that's me done with two minutes to spares. So thank you very much. Any questions from anybody? Yes? I think it was awesome. When you're thinking about what to put in the CLI, something like you have a budget-based primitive, for example, . Where do you draw a micro-visual model for building the mechanical primitive versus putting together, like checking in, for example, unit tests or integration tests? Like, how ephemeral are you things that are on the other side? At the limit, like checking them in and all the time? I think they're quite ephemeral in the sense, and that, I guess, is a personal preference, but I do feel like automated tests can be sometimes a bit too rigid and hard to predict and hard to, like, work over time. So I like some things that mimics the way I would test it, like a human would go and test the application. The way I find those is in the first stage where I work with the AI, even though, you know, like I can see that the button is a bit to the left and I want to tell the AI that, what I want is the AI to realize that by itself. So whenever I can't realize that sort of problem, I take a step back and try to think on how to make the AI realize the problem by itself. Another thing is very attractive loop. After you've done that many times, look back, past over your logs, ask an AI, oh, did you notice any issues, any stink? Are you doing sleep, for instance, calling sleep parentheses 15 everywhere is a thing that, like, there is some things that should be weighed for whatever, like a comment that could be there. Is there anything that you keep doing over and over, like running storybooks? And another thing is really think about like your own product. If you're making, say, a game in Unity, like, do you want an ASCII representation of your 3D world for your AI? If you're making something with lots of permissions, different logins that your AI can take very easily. It's really your new role to think about all of that. That's what I think anyway. And I think we're it for time. Thank you very much.