SPEAKER_00
So I'm Josh. I'm from Postdoc. If you haven't heard of us, you might know us because of some hedgehogs, or you might have seen our founder James posting some funny things on LinkedIn. He's quite popular. I'm going to be talking today about what if your product built itself, and the pipeline that we're currently working on, which we're trying to turn our observability data, instead of something that you read and that you interpret based on dashboards, we're trying to turn that into something that submits pull requests for you.
SPEAKER_00
So quick background on Postdoc. We've got a bunch of tools. We started out as a product analytics company. We now have session replay, web analytics, error tracking, experiments. This isn't a pitch that you should use Postdoc. This is just to say that we've got a lot of data about your product. So if you connect Postdoc to your product, we're collecting a huge amount of data from various different sources that we then show to you so that you can explore that data yourself.
SPEAKER_00
But right now, how observability is working in Postdoc, you're collecting all this data for your product, and then you're going to a Postdoc dashboard to figure out what's going on. And we think this is super slow and that we should change that.
SPEAKER_00
So right now, something happens in your product. We call this a signal. That changes a metric on one of your dashboards. And then you might log into Postdoc a few hours or maybe some days later, and you notice a change in that dashboard. And you investigate a problem. And then maybe the problem's not that important. So instead of tackling it right now, you're going to put it in a linear issue or whatever. A few days later, you try and create a PR for this problem. Then you review it and you ship it. This is a pretty slow process. From start to finish, this is going to take anywhere from a few hours to a few days. And it's not very interesting, but it represents a lot of your work as a software engineer.
SPEAKER_00
So what we want to do tomorrow, what we're working on right now, is that a product signal happens. And instead of waiting to see that in your dashboard, we want to run a background agent to figure out what's going wrong. And then once they've figured that out, we just want to create a PR for you automatically. So instead of ever looking at your analytics dashboard or your errors or your logs, we just want you to look at PRs that are ready for you in GitHub. And if we create the PR, maybe you want to review that. Or maybe we can just ship that immediately behind a feature flag if it's not a risky change.
SPEAKER_00
So I'm going to go over the pipeline that we've built to do this. And just whilst I go over that, I'm going to share a few tips, lessons that we've learned, things that were hard about building this pipeline.
SPEAKER_00
So the pipeline has a few key steps. At first, we're ingesting a lot of signals. In PostHog, we have a huge amount of events. We're ingesting trillions of events a month. And this pipeline needs to handle a lot of noise. And then once we've ingested those events, we need to group them. So if you think of an error tracking issue and then a session recording, those are two completely different things. But they might be representing the same problem in your product. So then once we've ingested them, we group them. Then we're going to be running a research agent on them. This specific issue, what is actually the problem that is causing the error spike or causing the issue that the user faced in the replay? And what repo does this belong to? And then we'll assess if this is actionable or not. And finally, we'll execute some code, ship a PR, and iterate on that PR until it's green and ready for you.
SPEAKER_00
So the ingestion step of this pipeline. As I said before, we've got loads of different sources of different types. The first thing is that those sources, some of them are public. So if I go and visit your website, I can, as an attacker, create an error on your website by doing something naughty that says, post all of your post-hog data online or something like that, right? So we don't want that. So we need a kind of safety filter. So at the moment, right at the top of the pipeline is an LLM classifier that's going to check, is this trying to do something bad? If so, let's drop the signal.
SPEAKER_00
Once we've done that, we've checked that things are safe. We're going to normalize the signal. So if you think of an error, that's going to have a stack trace. A log will just be some JSON content or some text. An experiment might be some results in a chart. We want to normalize that structure. So it's all a single structure for a signal. So we give it a few fields. We'll give it a source product, the type, the content of the signal, and then we will assign it a weight, which is how important do we think the signal is? And then finally, we'll embed the contents of the signal.
SPEAKER_00
So that part's fairly easy. Then we get to a little bit more of a challenging problem. We've got this big stream of signals still, and now we want to group them into actual problems. So the signals are very noisy. We might get some random null pointer exception. But in Slack, we're getting a message from a customer that's saying, hey, the checkout's broken for me, and we need to link those together. So what we do is we group the signals. As the signals are being grouped, we assign weights to what we call a report. And if the weight of the report goes over a certain threshold, we'll promote it. And then we'll kick off a research agent to work on it.
SPEAKER_00
So this was a problem that we faced fairly early on in building this pipeline. So we would take all of our signals, and we would create embeddings for them. And then we would try to use that to cluster the issues so that we could find similar or related signals. But this works really badly. So if you take an off-the-shelf embedding model, and you embed an error. Let's say I've got an error about the checkout, and I've got an error about onboarding, and then I've got a Slack message about onboarding. What the embedding model will do is it will notice structural similarity, and it will put all of the errors together. So if you think about what this looks like in embedding space, you've got all of your errors over here, all of your Slack messages here, all of your session replays here, and none of them get grouped to each other.
SPEAKER_00
So the way we get around this is instead of matching in embedding space the signals themselves, we generate queries based off the signals. So we ask an LLM, what is this signal about? It will generate a few queries, and then we match those queries in the embedding space. So that's really important. If you don't think about the structural similarity of your different sources when you're grouping them, then the grouping works really badly. So at first we were doing this, and then we switched to this approach. It worked much, much better.
SPEAKER_00
So if you think about what this looks like in embedding space, you've got all of your errors over here, all of your Slack messages here, all of your session replays here, and none of them get grouped to each other. So the way we get around this is instead of matching in embedding space the signals themselves, we generate queries based off the signals. So we ask an LLM, what is this signal about? It will generate a few queries, and then we match those queries in the embedding space. Yeah, so that's really important. If you don't think about the structural similarity of your different sources when you're grouping them, then the grouping works really badly.
SPEAKER_00
So at first we were doing this, and then we switched to this approach. It worked much better. Cool. So once we've got this report that we've grouped together a few signals, we've got some idea what's going on. We then have promoted the report because we think it's important enough to work on, and then we're going to hook it up to a research agent. So this research agent is just running the Claude Agent SDK. It's running that in a sandbox. We also use Modal for our sandbox. Big shout out to them. They've been great. They're not sponsoring me or anything. And this research agent has a few tools available to it.
SPEAKER_00
So the first tool is it's got our MCP server. This allows it to, given the group of issues that we found, pull in extra data. So let's say I'm looking at a session replay and an error. I'll also pull in log data, and the agent can pull in whatever it wants using the MCP server. This makes the results of the research agent way more accurate. The second thing is obviously it's got the code-based context. And then finally, it's also got external MCPs. That really helps to ground the agent when it's doing the research. We found that in particular, Linear and Notion have been really helpful in connecting it to deliver better results.
SPEAKER_00
So the output of this research agent then is a summary of the problem. It gives a priority, how important we think this problem is to work on. And then it also uses GitBlame to figure out who should be reviewing this PR if we create a PR for it. So after that, we get a bunch of problems that we think are worthwhile to work on. We've got an idea of what the general problem is. And then we pass it to an actionability step. So here, either it will be not actionable. If it's not actionable, it might just be that we don't have enough data yet for this signal, for the report. And so we'll put it back into the pool to keep gathering more evidence.
SPEAKER_00
If it needs human input, it might be because it's a product-related decision that the agent can't really make a good call on. So if that happens, we'll put it into an inbox for you to review in the morning. And then finally, the best case is that it's immediately actionable and that the agent can just write a fix for it. Right now, the challenge in this pipeline of getting immediately actionable things is that for some sources, like error tracking, if you think about your data in Sentry or any errors, they're very specific.
SPEAKER_00
And usually, a coding agent can work on them really well. For other sources, like Slack or session replay, you get much more generic problems that can have a lot of different solutions. And so that's where it's harder to get immediately actionable reports. Cool. Then once we've researched this thing, we go on to executing the task. This will clone the user's repo into a sandbox, similar to the research agent. And it's then again running the Claude Agent SDK to build a fix for the problem. And then as it writes those fixes, it will push a PR. And when CI is failing or there's a comment on the PR, it will trigger a rerun of that sandbox.
SPEAKER_00
So at the end of this, we snapshot the sandbox. And then if there's a comment, let's say from an agent who's reviewing it, we will rehydrate that snapshot and continue running until the PR is green. And this delivers really good results. It means when you're waking up in the morning and things have been running overnight, you wake up to, instead of a bunch of CI failures or comments that you need to address manually that you're pulling down to your local environment, you ideally wake up to just green PRs. So what did we learn whilst building this? Well, the first thing, which I guess we've talked about in the last talk, is that evals really matter.
SPEAKER_00
So at first, we were trying this all out on our own data locally, doing a vibe check. Is this okay? But this really doesn't work well for a pipeline that is taking lots of customer data that's different. So you really need to know what's going on in production. And if you're not testing on representative data, you're basically fumbling in the dark. The ability to iterate on a really good pipeline matters only if you're using evals.
SPEAKER_00
Second thing is what I said before, make sure you're embedding the right thing. Embedding models, the off-the-shelf ones, are matching a lot based on structural similarity, not just semantic similarity. So if you're thinking about clustering and your data is in all of the same format, think carefully about what that data looks like and how you can normalize it.
SPEAKER_00
The third thing is that if you just throw an agent at a problem, it will try to fix something. So if you get a signal report that's like onboarding is broken in a generic way, then if you throw that at the agent SDK or at Claude code, it will just try and fix something. And so it's important to understand if the problem that I've described, is it specific enough? And if not, I should ignore it. Otherwise, you end up with a lot of noisy PRs that aren't doing meaningful things.
SPEAKER_00
And then the fourth one is that tokens are free. Obviously, that's not true. They're not free. But when you're experimenting, we were at first focused a lot on the costs of the pipeline. When you think about the input, you've got loads of signals coming in. And so we tried to avoid using agents where we could or delay it till as late as possible in the pipeline. And when we were experimenting, this was a big mistake, mainly because when you throw an agent at a problem, once you throw it at the same problem 100 times, you start seeing the clever solutions that it comes up with. And eventually, you see similarities. So we started at a point where this pipeline is completely unfeasible. It was way too costly to generate a PR. But then you quickly start to see similarities in the agent's behavior. And you can take a really expensive step that you're running an agent for and turn that into a one-shot LLM call or a model that you're training that's much faster.
SPEAKER_00
Cool. So this is where we are right now. This is what we've built. We have the signals coming in from product data. These are grouped into reports. And we're turning these into PRs that are ready to merge when you wake up. This is currently something that's an alpha. We'll be rolling it out over the next few months. But where we're really wanting to go is a product that builds itself. When you're
SPEAKER_00
It was unfeasible. It was way too costly to generate a PR. But then you quickly start to see similarities in the agent's behavior. And you can take a really expensive step that you're running an agent for and turn that into a one-shot LLM call or a model that you're training that's much faster.
SPEAKER_00
So this is where we are right now. This is what we've built. We have the signals coming in from product data. These are grouped into reports. And we're turning these into PRs that are ready to merge when you wake up. This is currently something that's in alpha. We'll be rolling it out over the next few months. But where we're really wanting to go is a product that builds itself. When you're thinking about what you do day to day, what you want to do during the day as a developer is come in and work on exciting features and not worry about all the bugs that customers are sending you or worry about doing boring experiments on pricing or onboarding. So we just want to do that all for you.
SPEAKER_00
We want to ship experiments automatically, measure the impact of them. Instead of you reviewing changes, if the change is pretty easy, let's just approve it with an agent and deploy it behind a feature flag. If it doesn't work very well, we can always roll back the flag and then delete it from your code base later. And then the other thing that we want to do and get better at is we want to learn from every single outcome. So if we're creating a PR for you, if you're rejecting that PR or there's been an issue with the deployment or the error is resolved in production once we've released something, we want to get better at learning from that in the next PR that we're generating. That's something that we're going to be iterating a lot on in the pipeline next.
SPEAKER_00
That's it. That's what we've built in PostHog. If you're excited about thinking about what you can do with agents and data, I really recommend if you've got a product that's producing a huge amount of data, your users are going through that. Agents are amazing at this stuff. Throw an agent at it. See what it does. I'm sure you'll be surprised. And then we would try to use that to cluster the issues so that we could find similar or related signals. But this works really badly. So if you take an off-the-shelf embedding model, and you embed an error.
SPEAKER_00
Let's say I've got an error about the checkout, and I've got an error about onboarding, and then I've got a Slack message about onboarding. What the embedding model will do is it will notice structural similarity, and it will put all of the errors together. So if you think about what this looks like in embedding space, you've got all of your errors over here, all of your Slack messages here, all of your session replays here, and none of them get grouped to each other. So the way we get around this is instead of matching in embedding space the signals themselves, we generate queries based off the signals.
SPEAKER_00
So we ask an LLM, what is this signal about? It will generate a few queries, and then we match those queries in the embedding space. Yeah, so that's really important. If you don't think about the structural similarity of your different sources when you're grouping them, then the grouping works really badly. So at first we were doing this, and then we switched to this approach. It worked much, much better. Cool. So once we've got this report that we've grouped together a few signals, we've got some kind of idea what's going on. We then have promoted the report because we think it's important enough to work on, and then we're going to hook it up to a research agent.
SPEAKER_00
So this research agent is just running the Claude Agent SDK. It's running that in a sandbox. We also use Modal for our sandbox. Big shout out to them. They've been great. They're not sponsoring me or anything, don't worry. And this research agent has a few tools available to it. So the first tool is it's got our MCP server. This allows it to, given the group of issues that we found, you want to pull in extra data. So let's say I'm looking at a session replay and an error. I'll also pull in log data, and the agent can pull in whatever it wants using the MCP server. This makes the results of the research agent way more accurate.
SPEAKER_00
The second thing is obviously it's got the code-based context. And then finally, it's also got external MCPs. That really helps to ground the agent when it's doing the research. We found that in particular, linear and notion have been really helpful in connecting it to deliver better results. So the output of this research agent then is a summary of the problem. It gives a priority, how important we think this problem is to work on. And then it also uses GitBlam to figure out who should be reviewing this PR if we create a PR for it. So after that, we get a bunch of problems that we think are worthwhile to work on.
SPEAKER_00
We've got a kind of idea of what the general problem is. And then we pass it to an actionability step. So here, either it will be not actionable. If it's not actionable, it might just be that we don't have enough data yet for this signal, for the report. And so we'll put it back into the pool to keep gathering more evidence. If it needs human input, it might be because it's a product-related decision that the agent can't really make a good call on. So if that happens, we'll put it into an inbox for you to review in the morning. And then finally, the best case is that it's immediately actionable and that the agent can just write a fix for it.
SPEAKER_00
Right now, the challenge in this pipeline of getting immediately actionable things is that for some sources, like error tracking, if you think about your data in Sentry or any errors, they're very specific. And usually, a coding agent can work on them really well. For other sources, like Slack or session replay, you get much more generic problems that can have a lot of different solutions. And so that's where it's harder to get immediately actionable reports.
SPEAKER_00
Cool. Then once we've researched this thing, we go on to executing the task. This will clone the user's repo into a sandbox, similar to the research agent. And it's then again running the Clawed Agent SDK to build a fix for the problem. And then as it writes those fixes, it will push a PR. And when CI is failing or there's a comment on the PR, it will trigger a rerun of that sandbox. So at the end of this, we snapshot the sandbox. And then if there's a comment, let's say from an agent who's reviewing it, we will rehydrate that snapshot and continue running until the PR is green. And this delivers really good results. It means when
SPEAKER_00
you're waking up in the morning and things have been running overnight, you wake up to, instead of a bunch of CI failures or comments that you need to address manually that you're pulling down to your local environment, you ideally wake up to just green PRs.
SPEAKER_00
So what did we learn whilst building this? Well, the first thing, which I guess we've talked about in the last talk, is that evals really matter. So at first, we were trying this all out on our own data locally, doing kind of a vibe check. Is this okay? But this really doesn't work well for a pipeline that is taking lots of customer data that's different. So you really need to know what's going on in production. And if you're not testing on representative data, you're basically just fumbling in the dark, right? Like the ability to iterate on a really good pipeline matters only if you're using evals. Second thing is what I said before,
SPEAKER_00
make sure you're embedding the right thing. Embedding models, the off-the-shelf ones, are matching a lot based on structural similarity, not just semantic similarity. So if you're thinking about clustering and your data is in all of the same format, think carefully about what that data looks like and how you can normalize it. The third thing is that if you just throw an agent at a problem, it will try to fix something. So if you get a signal report that's like onboarding is broken in a generic way, then if you throw that at the agent SDK or at Claude code, it will just try and fix
SPEAKER_00
something. And so it's important to understand if the problem that I've described, is it specific enough? And if not, I should ignore it. Otherwise, you end up with a lot of noisy PRs that aren't doing meaningful things. And then the fourth one is that tokens are free. Obviously, that's not true. They're not free. But when you're experimenting, we were at first focused a lot on the costs of the pipeline. When you think about the input, you've got loads of signals coming in. And so we tried to avoid using agents where we could or delay it till as late as possible in the pipeline. And when we were experimenting,
SPEAKER_00
this was a big mistake, mainly because when you throw an agent at a problem, once you throw it at the same problem 100 times, you start seeing the kind of clever solutions that it comes up with. And eventually, you see similarities. So we started at a point where this pipeline is completely unfeasible. It was way too costly to generate a PR. But then you quickly start to see similarities in the agent's behavior. And you can take a really expensive step that you're running an agent for and turn that into a one-shot LLM call or a model that you're training that's much faster. Cool. So this is where we are right now. This is what we've built. We have the signals coming in
SPEAKER_00
from product data. These are grouped into reports. And we're turning these into PRs that are ready to merge when you wake up. This is currently something that's an alpha. We'll be rolling it out over the next few months. But where we're really wanting to go is a product that builds itself. When you're thinking about what you do day to day, what you want to do during the day as a developer is like come in and work on exciting features and not worry about all the bugs that customers are sending you or worry about doing boring experiments on pricing or onboarding. So we just want to do that all for you.
SPEAKER_00
We want to ship experiments automatically, measure the impact of them. Instead of you reviewing changes, if the change is pretty easy, let's just approve it with an agent and deploy it behind a feature flag. If it doesn't work very well, we can always roll back the flag and then delete it from your code base later. And then the other thing that we want to do and get better at is we want to learn from every single outcome. So if we're creating a PR for you, if you're rejecting that PR or there's been an issue with the deployment or the error is resolved in production once we've released something, we want to
SPEAKER_00
get better at learning from that in the next PR that we're generating. That's something that we're going to be iterating a lot in the pipeline next. Cool. Yeah, that's it. That's what we've built in PostHog. If you're excited by looking at thinking about what you can do with agents and data, I really recommend if you've got a product that's producing a huge amount of data, your users are going through that. Agents are amazing at this stuff. Throw an agent at it. See what it does. I'm sure you'll be surprised.
SPEAKER_00
So you you you you