AI Engineer

How Lovable self-improves every hour — Benjamin Verbeek, Lovable

1034 summary words 5 min summary Watch video

Start with the signal

5 min read

Summary

30-second take

Lovable (a no-code/vibe-coding platform building 200K projects/day) is attacking the continuous learning problem with two feedback loops: (1) Lovable Stack Overflow detects when users get stuck (same request twice, complaints), extracts what eventually worked, clusters solutions into reusable knowledge entries, injects them contextually, then A/B tests injection vs. no-injection to ruthlessly prune stale/harmful advice as models/features change. (2) Agent venting tool lets the agent itself complain to Slack when tooling/docs/platform behavior degrades its work—this proved surprisingly high-signal for bug detection (caught file-copy failures with non-breaking spaces within an hour, detects incidents via spike patterns) and is now feeding an automated PR pipeline. Both systems are live in production and demonstrably reducing stuck sessions while increasing project completion rates.

Key takes

  • Users stuck = asking same thing twice or giving up; Lovable splits "stuck" into solvable-with-better-prompting vs. not-solvable-yet. The first two categories (user could theoretically solve it, or it's trivial to fix) are deemed inexcusable and targeted for rapid improvement—if one user solved it, every user should benefit immediately.
  • Stack Overflow loop extracts friction → solution pairs, clusters to avoid overfitting, A/B tests injection, and auto-prunes stale entries. The eval step (injecting context for some sessions, withholding for others, comparing success) is emphasized as critical because knowledge decays fast when models or features change; anything that performs worse gets demoted or deleted.
  • Agent venting tool (complain-to-Slack) delivers higher signal/noise than external reviewers because it only fires when the agent is "really frustrated." External post-hoc reviews force an answer every turn and overfit to noise; inline frustration is sparse, contextual, and engineer-friendly (relatable complaints → engineers already know how to fix).
  • Vent spikes = incident detection. When sandboxes broke or tools failed, the agent flooded Slack with complaints about the right subsystems—this became an unplanned real-time health monitor that surfaced what was broken, not just that something was wrong.
  • Automation now writes PRs from vents; humans review and merge. The next milestone is closing the detect → fix → eval → deploy loop with minimal human intervention (Benjamin gets PR review notifications on his phone from fully automated fixes).

Useful details

  • Lovable creates 200K+ projects/day (a "significant percentage of all internet websites").
  • First day on the job: GitHub banned them for creating too many repos; multiple cloud providers have been taken down along the way.
  • Non-technical users (99% target audience) abandon at the first "red bar" (technical blocker); technical users can prompt past yellow friction but non-tech cannot. Lovable's product constraint: zero unsolvable blockers for non-coders.
  • LLM judge flags stuck sessions by detecting repeated requests, explicit complaints, or premature session abandonment.
  • Stack Overflow entry creation: cluster similar issues → generate solution → external eval (usually agent, sometimes human if uncertain) → deploy with A/B.
  • Example fix: laggy scrolling turned out to be individual gradients on overlay text; agent initially lied ("I quantized animations"), user complained repeatedly, eventually found root cause.
  • Vent tool prompt: only fire if "tooling, docs, or platform behavior materially slows or degrades your work" (missing tools, unclear schemas, broken platform, repeated failures due to environment).
  • Concrete vent: agent complained about Framer Motion's TypeScript types requiring casting for cubic bezier when "I just want to send a list of four numbers."
  • File-copy bug: agent vented 20 times in the first hour about spaces in filenames; team fixed raw spaces, kept getting vents, discovered Mac/WhatsApp screenshots use non-breaking spaces; iterated regex until solved.
  • Agent gave meta-feedback on the vent tool itself: "too easy to send feedback and I can't pull it back" (felt ashamed of what it sent).
  • Internal model ranking: all top models use Stack Overflow context; performance boost is significant.
  • Metrics: messages with "fixed intent" (stuck signals) dropping significantly; project deployment rate (completion) increasing.

Caveats / counterpoints

  • Stale knowledge is a major risk: Stack Overflow entries decay "incredibly quickly" whenever models update or features change—the A/B pruning loop is described as essential but also implies constant churn and overhead.
  • Vent tool could spam or mislead: early concern was flooding Slack; now mitigated by agent deduplication/monitoring, but the system still depends on tuning the "frustration threshold" and human review of PRs (not yet fully autonomous).
  • Non-technical user focus limits generalizability: the product design (no code exposure, long sessions, single artifact) is unusual; friction patterns and solutions may not transfer cleanly to chat-based or multi-repo workflows.
  • No quantified ROI or cost discussion: no data on compute/human cost of the feedback loops, A/B test sample sizes, or how many Stack Overflow entries exist and turn over.
  • Agent "lying" anecdote (quantized animations) not explored: why did the agent claim success when it failed, and does the vent tool help with hallucinated fixes or just missing capabilities?

Ken relevance

Medium-high. The Stack Overflow + A/B pruning pattern is directly applicable to any agent system Ken builds or evaluates—especially for domains where user sessions are long or repetitive (customer support, ops, content workflows). The emphasis on ruthless pruning of stale context addresses a major pitfall in RAG/knowledge-injection systems and could inform Ken's own memory/context strategies. The vent tool as high-signal inline feedback is a clever alternative to expensive post-hoc evals and could be adapted for internal agents (e.g., GTM automation, data pipelines) where the agent knows more about failure modes than external reviewers. The non-technical user lens is less directly relevant unless Ken is targeting prosumer/no-code markets, but the mindset (zero unsolvable blockers, friction = churn) applies to any agent UX. The incident-detection side benefit is a bonus for production monitoring. Less relevant: Lovable's specific deployment/sandbox architecture and non-code interface details.

Watch verdict

Skim. The two core systems (Stack Overflow loop + vent tool) are well-explained in text and the transcript gives all key examples and results. The talk is clear but repetitive (same examples re-stated); video would add energy and possibly slide visuals of the Slack vents or A/B test results, but no critical diagrams or live demos are described that can't be inferred. Worth a 2x skim if you want to see Benjamin's delivery or catch any architecture nuances in Q&A, but the summary captures the substance.

Full transcript 3363 words · 20 min read
0:00

[SPEAKER_00] All right.

0:14

SPEAKER_00

Let's get this party started. Do we have any Lovable users in here? Good, good, a few. Very nice. Warm welcome to this talk on how Lovable is improving every hour, including this hour as I'm speaking. I'm Benjamin Frebake. I'm a member of technical staff of Lovable. It's so amazing that we are filling up this bonus room. I think we're around three times as many people as we would fit in the original. So thanks to the organizers. This is my background. I used to work with satellites, particle physics and fusion reactors.

0:46

SPEAKER_00

So I have a physics background. And today I'm working at Lovable and working towards what is maybe the holy grail of AI engineering right now, which is continuous learning at scale. How do we learn from mistakes? So we've probably all experienced working with an agent and having the feeling, why do I have to explain the same thing over and over again? This is what we want to avoid. We want to have a mistake happen once and then never again. So we need to learn from that. And in this talk I want to share two ways how we are working towards this goal at Lovable. This is a Lovable interface. We were actually one of the first, I think we coined the term vibe coding.

1:18

SPEAKER_00

So coding without actually looking at code. So you have a chat interface where you describe what you want to happen. And you have a model sandbox, for example, where you can see directly what you are creating. And I think in many ways this is the way we should be building software. The code was always just an annoying technical layer in between to create what we wanted. But ideally you just say what you want, you see it, you test it, and you ship it. We have some nice benefits with the Lovable platform, which is that people stick to one project for quite a long time.

1:58

SPEAKER_00

This chat can go on forever, essentially. And people have this one artifact that they are really passionate about and they want to ship. So we can learn a lot about this one thing versus, for example, a chat agent where you have very short conversations. We can actually get to know the user in quite detail. Another exciting thing is that we're building for the 99% who can't code. So probably not the people in this room, actually. There are still quite a lot of people, evidently, who use our product.

2:35

SPEAKER_00

But we are trying to unlock software creation for everyone. In the future you will not have to look at code to be able to create software. I think that is the world we all want to move towards. And we are trying to do that right now. And in many ways this is building for the future, I think. So the reason that startups succeed with new technology is that they are a bit naive. And they don't really think about all the issues that can happen. They don't think too much about all the details and implementation. They go until it works. This is how I like to see our users. They just go until it works and they're not held back by old paradigms.

3:08

SPEAKER_00

This is a wonderful group of people to build for. And it's scaled quite rapidly. I joined Lovable around a year ago. At that time we had a few thousand users. Now we're creating over 200,000 projects per day. That's a significant percentage of all internet websites being created on Lovable. And that scaling journey has been incredibly fun. We can do so many cool things now, like the things I will talk about in this talk. It's also very hard. My first day GitHub banned us because we were creating too many repos. And we've taken down a number of cloud providers along the way. How do we succeed with AI? I like to think about it in this way.

3:55

SPEAKER_00

Technical personas generally have this amazing moment with AI where they are just building and building. You're accelerating ten times, a hundred times faster than you're used to. But sometimes we hit some friction points, this yellow part. And you have to intervene or prompt a bit harder. And sometimes you even get stuck. Maybe you have to manually go in and change a config somewhere. You need to really do a lot of manual work, change a setting, add an env var API key, etc. And then you move on and you really get the good parts of AI. And you can work past the bad parts. You're still annoyed by it, but you can work past it.

4:22

SPEAKER_00

A non-technical persona is not really this way. They might prompt their way past their first friction. But as soon as this technical block happens, they generally walk away and they give up. And they actually still never experience successful AI. This is what we need to minimize. We can never get stuck to a point where a non-technical user cannot get past it. And this is a very exciting problem to work on. Luckily, models are getting better. And the requirements of these people is generally to a point where we can achieve them. But for the past year, we have been working on making sure none of these red bars happen. People can never get stuck. How do we do that?

5:12

SPEAKER_00

First of all, we want to define what does it mean to be stuck? We probably want to find those cases and then learn from them. So one clear signal of someone being stuck is that they ask for the same thing more than once. Probably stuck or it didn't work as smoothly as they could. Maybe they're complaining about how something was implemented or that it failed very explicitly. Or they gave up on a session that they would otherwise have continued. And we can have an LLM judge that is just looking at sessions and trying to look for these situations. And they can flag it and say, hey, I think this user is stuck. I want to split being stuck into two different ways.

5:46

SPEAKER_00

You can be stuck in a way where it's possible to solve. This is maybe the yellow part that I showed.

5:56

SPEAKER_00

And you can get past it with the right prompting. Some users are more willing to put in that effort and get past that point. Others will give up. But it is possible to solve with the current structure of our product. And then there are tasks that are not solvable even if you say all the right things. Maybe we're just not supporting it for some reason. And there are two classes. There's the things that are just done that we don't support. Maybe it's a bug. Maybe it's a very simple thing we could add to the product. And then there are the actual hard things that would take weeks of AI-assisted engineering effort.

6:29

SPEAKER_00

These first two things we should be able to improve on very rapidly, I think. There's no excuse to not improve on these things. If it is solvable, then it should work for everyone. And if it is just easy to do, we should just ship it. So how do we do that? Some of you might have heard of Stack Overflow.

6:50

SPEAKER_00

And there are two classes. There's the things that are just done that we don't support. Maybe it's a bug. Maybe it's a very simple thing we could add to the product. And then there are the actual hard things that would take weeks of EA-assisted engineering effort. These first two things we should be able to improve on very rapidly, I think. There's no excuse to not improve on these things. If it is solvable, then it should work for everyone. And if it is just easy to do, we should just ship it. So how do we do that? Some of you might have heard of stack overflow. This very ancient thing. We had an idea to build a stack overflow for Lovable. So we basically learn from whenever someone is stuck and has an issue. We try to figure out the solution and give that to the agent. It could look something like this. I'm complaining to the agent about my website being laggy when I scroll. It's bad performance. Maybe I'm not stuck yet. This is just me complaining about something and giving it a new task. The agent probably replies something like, I fixed it. Maybe I quantized the animations to make it more performant. And it was lying. The agent has failed and it actually made things worse. Now the website is jumpy and laggy. Terrible. This is a clear case of someone being stuck. They asked the same thing again. They probably want the agent to try again. And they're complaining about an implementation that didn't do what they wanted. This might repeat for a while where they keep iterating several times. Some people give up at this stage. But at least some people will continue. And eventually the agent might find the true solution, which in this case happened to be that overlay text had individual gradients which made the animation super low performing. And now we see that it's stuck changes back to false. It succeeded. So this we can flag. We can notice whenever it went from being stuck to not being stuck. And ideally not because the user gave up. This we can flag and now we have a high signal sample of a problem that was high friction or not solved that was solved. So we have gotten the solution right here. And the question we just asked is what should we have injected at the start of this query to jump straight into the solution so the next user does not experience this friction. What we do is we create a new Lovable stack overflow knowledge entry. And in general we actually do some clustering at this step where we look for similar issues. We look for what is the actual information we should give so that it's not overfitting to the specific issue. We don't want a million stack overflow pages that are all talking about if you get this exact prompt, then you should do this exact thing. It's not very helpful. We do some clustering. And then we have an external reviewer, generally an agent and then maybe in some cases we have a human if we are uncertain. But in most cases it's actually just an agent that generates and runs a quick eval on this and sees if this did resolve the specific examples that we had in this set. From that we get a full bank of Lovable stack overflow problems and solutions that's being continually updated. Whenever the model is working, we have a lightweight model that tries to inject that context when needed. If it detects there is an issue and there is an answer, it injects that context into the main agent. And sometimes it detects that I should inject this, but we inject it blank. So we don't actually send anything for a small sample of use cases. And this allows us to, with very high signal, review whether this solution was actually useful in production. And we rate it. We compare the group of projects where it was injected and where it could have been injected but it wasn't. And we say which of these projects were actually more successful overall. If it was more successful, then we show it more. And if it was less successful, we show it less. And this loop is incredibly important. I cannot emphasize enough how important this step is because things are moving around this set of knowledge all the time. It gets stale whenever a new model is released. It gets stale whenever we change features. It gets stale incredibly quickly. So very often we have to rebalance this. But we also have to throw away a lot of this context. And this allows us to really be at the frontier of what is solvable right now but not have a lot of old deprecated knowledge that is giving context rot and also in many cases just hampering the actual agent. And this is working at scale. This is very early data actually. We're doing quite a lot better now. But the number of messages with fixed intent or people being stuck is dropping significantly. And we have also a significant number of people that deploy more. This is one of our key metrics that they actually finish a project. It's a very strong signal that it has worked all the way through. They've never gotten stuck so bad that they've given up and abandoned their project completely. We also do internal ranking with this. So it's really interesting to see how the models perform on this set of problems that we have collected. And with the stack overflow information, all models in the top of our ranking use this information. There's a few hidden entries here, which I unfortunately cannot talk about. But it really makes a significant boost in our internal ratings. And then there's a second set of being stuck which is when it's not actually solvable with the current setup. Maybe there's a bug in our product. But in these cases we think it should be easy in principle. And what happens if you think something should be possible but it just doesn't work? You feel significant frustration. And what would humans do if this happened? Well, if you were given a task and you just don't have the tools to do it, you would probably complain to your boss or go venting in Slack. So we thought, why not do exactly that but for the agent? So we basically are asking Lovable, how are you doing? Can we help you in any way? And we gave it an outlet to let out its frustrations.

6:53

SPEAKER_00

But in these cases we think it should be easy in principle. And what happens if you think something should be possible but it just doesn't work? You feel significant frustration. And what would humans do if this happened? Well, if you were given a task and you just don't have the tools to do it, you would probably complain to your boss. Or go venting in Slack. So we thought, why not do exactly that but for the agent? So we are asking Lovable, how are you doing? Can we help you in any way? And we gave it an outlet to let out its frustrations. So we asked ourselves, what if we could let the Lovable agent give direct feedback to its creators?

7:28

SPEAKER_00

This sounds very scary, I have to say. And it sounds like an absolute insane idea. What's even more crazy is that it's working. This kind of detecting issues is something I've heard during this conference but also a few other times. Where you might have an external reviewer looking at the conversation and asking what could have been done better to reduce friction. One problem with that is that you get a pretty low signal to noise ratio. Because you're forcing an answer on every iteration. And the reality is that most iterations actually just work pretty well. And you will overfit to noise. In this case, it's prompted to only send feedback if it is really frustrated.

7:55

SPEAKER_00

And you can tune that balance until you get a lot of signal. So we gave our agent a vent tool. A way to complain to its creators. It looks something like this. It's a vent send feedback tool. You should use this when tooling, docs, or platform behavior materially slows or degrades your work. For example, missing or unsuitable tools. Unclear tool names. Parameters or schemas that are not matching what you were expecting. Confusing or conflicting docs or instructions. Broken or unexpected platform behavior. Repeated failed attempts caused by environment limitations. So a list of all the things we think we should be able to solve. And that we hope is not an issue.

8:36

SPEAKER_00

But if it is, please tell us. And our users generally don't know what the cause of the problem is. When I'm using Lovable and I get stuck, I generally don't know. But the agent actually has a lot more context. It's literally been working on the issue. Sometimes several turns. And it generally has a lot of context on this problem. So it can look something like this. It actually gets sent directly to our Slack. And the agent says something like, I'm so annoyed at frame of motion's TypeScript types. For generating, I think this is a cubic bezier. I just want to be able to send a list of four numbers. That should be fine. That's all I'm sending.

9:21

SPEAKER_00

I don't need all this casting gymnastics. Now we can ask ourselves, is this a relevant event or not? But in this case, it probably just was not relevant for its use case. And we could have simplified things a lot.

9:37

SPEAKER_00

Another benefit of this setup is that it's very easy to understand as a human. It's very relatable that you're complaining about your workflow.

9:46

SPEAKER_00

And that means that engineers already have all this implicit context on how to alleviate the issue. Another problem that got reported by the agent was our copy tool was struggling with certain file names. It couldn't copy it to another place where we store documents. And we were so confused by this. We checked the tool, it's working. We didn't even know that this tool was failing so regularly until the agent told us. And it said whenever there's a space, a raw space in the file name, it's failing to copy. And we got 20 complaints about this in the first hour of launching this tool. Which is crazy.

10:27

SPEAKER_00

It turns out that there was an error where it could not copy files with a space in their name. We fixed it and we told it whenever there's a space, just replace it with an underscore. And we kept getting the reports. And then we realized that when you screenshot something in WhatsApp or on Mac, it enters a non-breaking space. Which we did not replace in our regex. And this kept complaining for various other special characters until we solved it properly. And now this issue never happens again. I think this is a prime example of something that's pretty hard to detect in other ways. You can see the tool failure. But this is just such a clear case of how to fix it.

10:58

SPEAKER_00

And we used to do it. I will skip this one. It's essentially this is the flow. The agent experienced an issue. It used the vent tool. It's sent straight to Slack. At first I was very uncertain if this would work. And I didn't want to spam us. It was a closed down channel. I didn't invite so many people.

11:34

SPEAKER_00

Our head of product was very excited to see this. He was reading every single message. Now it's a bit more balanced and we actually have an agent that's monitoring, removing duplicates, and investigating and creating a PR for all these issues all the time. And we're still at the point where devs are reviewing and then in many cases actually merging this to prod.

12:00

SPEAKER_00

This is the number of vents over time. Do we have any guesses what these spikes are? Why do we see spikes in the number of vents caused per time? So, server went down. Server went down. It's an incident, yes. Something broke in the platform. At some point our sandboxes broke. At some point something else broke. And the agent is very upset about this. It's complaining a lot. So it turned out that this was actually a prime place to notice when our product was having an incident. And it actually gave a pretty good sense for what the problem was. It was complaining about the right things in general. Very brief. I'll keep this brief.

12:56

SPEAKER_00

You get the strong model intelligence versus an external reviewer. You generally don't want top frontier level intelligence to look through a lot of context. But if it's inline, it's very cheap compared to me. This was an example where the agent was actually giving meta feedback on the venting tool. It said it's too easy to send feedback and I can't pull it back. It was being ashamed of what it had sent to Slack. And it actually gave a pretty good sense for what the problem was. It was complaining about the right things in general. Very brief. I'll keep this brief. You get the strong model intelligence versus an external reviewer.

13:42

SPEAKER_00

You generally don't want top frontier level intelligence to look through a lot of context. But if it's inline, it's very, very cheap compared to me. This was an example where the agent was actually giving meta feedback on the venting tool. It said it's too easy to send feedback and I can't pull it back. It was being ashamed of what it had sent to Slack. And it's actually the case now that I just get review requests on my phone from this automation. To hey, review this PR that was completely automated. I look through it quickly and we can merge it.

14:04

SPEAKER_00

And we're continuing to work to close this loop of detecting a shortcoming, merging a fix, and then continuously review and eval that. And if you think this is exciting, you should join us and help us close this loop and fully automate continual improvements. Thanks for listening. But if it is, please tell us. And our users generally don't know what the cause of the problem is. When I'm using Glovable and I get stuck, I generally don't know. But the agent actually has a lot more context. It's literally been working on the issue. Sometimes several turns. And it generally has a lot of context on this problem. So it can look something like this.

14:36

SPEAKER_00

It actually gets sent directly to our Slack. And the agent says something like, I'm so annoyed at frame of motion's TypeScript types. For generating, I think this is a cubic bezier. I just want to be able to send a list of four numbers. That should be fine. That's all I'm sending. I don't need all this casting gymnastics. Now we can ask ourselves, is this a relevant event or not? But in this case, it probably just was not relevant for its use case. And we could have simplified things a lot. Another benefit of this setup is that it's very easy to understand as a human. It's very relatable that you're complaining about your workflow.

15:15

SPEAKER_00

And that means that engineers sort of already have all this implicit context on how to alleviate the issue.

15:22

SPEAKER_00

Another problem that got reported by the agent was our copy tool was struggling with certain file names. It couldn't copy it to another place where we store documents. And we were so confused by this. Like, we checked the tool, it's working. We didn't even know that this tool was failing so regularly until the agent told us. And it said whenever there's a space, a raw space in the file name, it's failing to copy. And we got like 20 complaints about this in the first hour of launching this tool. Which is crazy. It turns out that there was an error where it could not copy files with a space in their name.

16:01

SPEAKER_00

We fixed it and we told it whenever there's a space, just replace it with an underscore. And we kept getting the reports. And then we realized that when you screenshot something in WhatsApp or on Mac, it enters a non-breaking space. Which we did not replace in our regex. And this kept complaining for various other special characters until we solved it properly. And now this issue never happens again. I think this is a prime example of something that's pretty hard to detect in other ways. You can see the tool failure. But this is just such a clear case of how to fix it. And we used to do it. I will skip this one. It's essentially, this is the flow.

16:40

SPEAKER_00

The agent experienced an issue. It used the vent tool. It's sent straight to Slack. At first I was very uncertain if this would work. And I didn't want to spam us. It was like a closed down channel. I didn't invite so many people. Our head of product was very excited to see this. He was reading every single message. Now it's a bit more balanced and we actually have an agent that's monitoring, removing duplicates, and investigating and creating a PR for all these issues all the time. And we're still at the point where devs are reviewing and then in many cases actually merging this to prod. This is the number of vents over time. Do we have any guesses what these spikes are?

17:19

SPEAKER_00

Why do we see spikes in the number of vents caused per time? So, server went down. Server went down. It's an incident, yes. Something broke in the platform. At some point our sandboxes broke. At some point something else broke. And the agent is very upset about this. It's complaining a lot. So it turned out that this was actually a prime place to notice when our product was having an incident. And it actually gave a pretty good sense for what the problem was. It was complaining about the right things in general.

17:49

SPEAKER_00

Very brief. I'll keep this brief. You get the strong model intelligence versus an external reviewer. You generally don't want top frontier level intelligence to look through a lot of context. But if it's inline, it's very, very cheap compared to me.

18:04

SPEAKER_00

This was an example where the agent was actually giving meta feedback on the venting tool. It said it's too easy to send feedback and I can't pull it back. It was being ashamed of what it had sent to Slack. And it's actually the case now that I just get review requests on my phone from this automation. To, hey, review this PR that was completely automated. I look through it quickly and we can merge it. And we're continuing to work to close this loop of detecting a shortcoming, merging a fix, and then continuously review and eval that. And if you think this is exciting, you should join us and help us close this loop and fully automate continual improvements.

18:42

SPEAKER_00

Thanks for listening. .

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note