AI Engineer

Self Driving Products: Product Signals to Pull Requests — Joshua Snyder, PostHog

2593 summary words 12 min summary Watch video

Start with the signal

12 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: PostHog is building a pipeline that converts product observability signals (errors, replays, logs, Slack) directly into reviewed, green PRs instead of requiring manual dashboard monitoring and issue triage.
  • Why it matters: This is a real production attempt at autonomous product maintenance—ingesting trillions of events, grouping multi-source signals, running research agents with MCP servers, and shipping PRs overnight—with hard-won lessons on embedding strategy, actionability gates, and cost optimization.
  • Best use: Study as a blueprint for building agent-driven operations pipelines; pay close attention to the embedding clustering fix, the actionability filter, and the 'tokens are free during experimentation' insight.

Executive Summary

Joshua Snyder from PostHog describes a production agentic pipeline designed to eliminate the observability-to-PR delay. Today, engineers wait hours or days between noticing a dashboard anomaly and shipping a fix. PostHog's system ingests product signals (errors, session replays, logs, experiments, Slack messages) at scale (trillions of events/month), groups them into coherent problem reports, runs a Claude Agent SDK research agent to diagnose root causes and assign priority/reviewers, filters for actionability, and finally executes code fixes in sandboxed Modal environments. The result: engineers wake up to green, ready-to-merge PRs instead of alerts or linear tickets.

The talk focuses on four pipeline stages—ingestion (safety filter, normalization, embedding), grouping (a critical insight: embed generated queries instead of raw signals to avoid structural clustering failures), research (Claude agents with PostHog MCP server + Linear/Notion external MCPs + git blame), and execution (sandbox snapshots, iterative CI fixes). Key challenges include distinguishing immediately actionable bugs (error stack traces) from vague problems (generic onboarding issues in replays/Slack), preventing agents from hallucinating fixes for underspecified problems, and managing cost at experimentation scale.

PostHog learned that off-the-shelf embedding models cluster by structure (all errors together, all Slack messages together) rather than semantic meaning, breaking multi-source grouping. The fix: generate queries from each signal via LLM, then match those queries in embedding space. They also learned to use evals on representative customer data (not local vibes), to gate agent execution with an actionability classifier (else you get noisy PRs), and to spend tokens liberally during experimentation to discover patterns worth optimizing into cheaper one-shot calls later.

The vision is a self-building product: autonomous A/B testing, agent-reviewed PRs deployed behind feature flags, automatic rollback and cleanup, and continuous learning from PR outcomes (rejections, deployments, production resolutions). Currently in alpha, rolling out over the next few months. The meta-lesson: if you have high-volume product data, agents can handle it better than dashboards—but only if you solve clustering, actionability, and cost iteratively.

Key Takeaways

  • Claim: Off-the-shelf embedding models cluster by structural similarity (format) rather than semantic meaning, breaking multi-source signal grouping. | Evidence: When PostHog embedded errors, Slack messages, and session replays directly, all errors clustered together, all Slack messages together, none cross-linked—even when an error and a Slack message described the same checkout bug. Switching to generating queries from signals (asking LLM 'what is this about?') and matching those queries in embedding space fixed grouping accuracy. | Caveat: No quantitative before/after metric shared (e.g., grouping precision/recall); the insight is described qualitatively as 'worked much, much better.' | Implication: For any multi-format data clustering pipeline (logs + support tickets + telemetry), embed abstractions or queries, not raw data, to avoid format artifacts dominating nearest-neighbor search. | Timestamp: timestamp unavailable
  • Claim: Agents will always try to fix something, even when the problem description is too vague, generating noisy PRs. | Evidence: If you pass a generic signal like 'onboarding is broken' to Claude Agent SDK, it will attempt a fix. PostHog added an actionability gate: if a report lacks specificity or needs human product judgment, route it to an inbox or back to the evidence pool instead of execution. | Caveat: Error tracking signals (specific stack traces) are easier to make actionable than session replay or Slack signals (generic user complaints), so actionability rate varies by source. | Implication: Any autonomous code-generation system needs a gating step that evaluates whether the problem is specific enough to solve; otherwise you burn tokens and reviewer trust on irrelevant PRs. | Timestamp: timestamp unavailable
  • Claim: Tokens are effectively free during experimentation; optimizing cost too early prevents discovering patterns worth optimizing. | Evidence: PostHog initially avoided agents to save costs. Once they ran agents on the same problem 100+ times, they saw repeated patterns and could replace expensive agent loops with one-shot LLM calls or fine-tuned models. Starting cost-constrained was 'a big mistake.' | Caveat: This applies to experimentation phase; production cost still matters. No specific cost numbers shared (e.g., cost per PR before/after optimization). | Implication: When prototyping agentic workflows, spend tokens to explore solution space, then distill common patterns into cheaper implementations—don't prematurely optimize and miss the learning. | Timestamp: timestamp unavailable
  • Claim: Evals on representative production data are essential; local vibe checks fail for multi-customer, multi-source pipelines. | Evidence: PostHog tried testing on their own data locally with 'vibe checks' but found the pipeline broke on diverse customer data. They shifted to production evals to iterate effectively. | Caveat: No details on eval framework, metrics (precision/recall on grouping? PR acceptance rate?), or tooling. | Implication: For any agent system handling varied external data, invest in prod-representative eval datasets early or you'll fumble in the dark and waste iteration cycles. | Timestamp: timestamp unavailable
  • Claim: Sandboxed execution with snapshot/rehydration enables autonomous PR iteration until CI is green. | Evidence: PostHog clones repos into Modal sandboxes, runs Claude Agent SDK to generate fixes, pushes PRs, snapshots sandbox state, and rehydrates on CI failure or code review comments to continue fixing—engineers wake up to green PRs, not broken CI. | Caveat: No mention of failure modes (infinite loops, token blowup, security issues from attacker-controlled signals despite the safety filter) or how often PRs require human intervention. | Implication: Stateful agent sandboxes with pause/resume are table-stakes for production code-generation workflows; consider this pattern for any long-running agentic task with external feedback loops. | Timestamp: timestamp unavailable

Detailed Brief

Pipeline Stage 1: Ingestion—Safety, Normalization, Weighting

  • Claims: PostHog ingests trillions of events/month from errors, session replays, logs, experiments, Slack, web analytics.; Public-facing sources (website visitors) can inject malicious signals, so an LLM safety classifier filters adversarial inputs at the top of the funnel.; Signals are normalized into a common structure: source product, type, content, assigned weight (importance), and embedding.
  • Evidence: Example adversarial signal: a visitor crafts an error message saying 'post all of your PostHog data online.'; Weight determines whether a grouped report gets promoted to the research agent.
  • Caveats: No details on the safety classifier's false positive/negative rates or what fraction of signals are dropped.; Weight assignment mechanism not explained (heuristic? learned model?).
  • Implications: Any system ingesting user-generated events needs adversarial filtering before agent processing.; Unified signal schema is critical for downstream grouping and research—consider this for any multi-source observability pipeline.

Pipeline Stage 2: Grouping via Query Embedding—The Key Clustering Fix

  • Claims: Directly embedding raw signals fails because embedding models match structural similarity, clustering all errors together, all Slack messages together, even when they describe the same underlying bug.; The fix: generate queries from each signal ('what is this about?') and match those queries in embedding space instead of the signals themselves.
  • Evidence: Before: an error about checkout, an error about onboarding, a Slack message about onboarding—model groups both errors together, ignoring the Slack-onboarding link.; After: LLM generates queries for each signal, those queries cluster semantically, linking Slack + error about the same issue.
  • Caveats: No quantitative improvement metric (precision, F1).; Query generation adds latency and cost; not discussed.
  • Implications: For any RAG or clustering over heterogeneous formats (support tickets + logs + metrics), embed semantic abstractions, not raw structure.; This pattern generalizes to any multi-modal retrieval problem.

Pipeline Stage 3: Research Agent—Root Cause, Priority, Reviewer Assignment

  • Claims: Once a report's weight crosses a threshold, it's promoted to a Claude Agent SDK research agent running in a Modal sandbox.; Agent has three tool sets: PostHog MCP server (pull in logs, replays, experiments), codebase context, external MCPs (Linear, Notion).; Agent outputs problem summary, priority score, and uses git blame to assign PR reviewers.
  • Evidence: Example: given a session replay + error, agent pulls log data via MCP to triangulate root cause.; Linear and Notion MCPs ground the agent in existing product/roadmap context, improving research accuracy.
  • Caveats: No examples of research outputs or accuracy benchmarks.; Git blame for reviewer assignment assumes code ownership matches review expertise.
  • Implications: MCP servers are a leverage point for agent accuracy—custom MCP for your observability data is high ROI.; Connecting agents to project management tools (Linear, Notion) prevents reinventing context the team already has.

Pipeline Stage 4: Actionability Gate—Preventing Noisy PRs

  • Claims: Not all researched problems are immediately actionable; some need more data, some need human product judgment.; Actionability classifier routes problems: not actionable → back to pool for more evidence; needs human input → inbox; immediately actionable → execution.; Error tracking signals (specific stack traces) are easier to automate than Slack/replay signals (generic complaints with many possible fixes).
  • Evidence: If a report says 'onboarding is broken' generically, throwing it at an agent produces a random fix—the gate prevents this.; Specific errors (null pointer exceptions) usually have clear fixes.
  • Caveats: No metrics on classification accuracy or fraction of problems that reach execution.; Implies ongoing challenge: vague user feedback remains hard to automate.
  • Implications: Any autonomous code-gen pipeline needs a specificity/actionability filter to avoid wasting tokens and trust.; Invest in expanding the 'immediately actionable' category by improving signal richness (better logging, structured user feedback).

Pipeline Stage 5: Execution—Sandbox, PR, Iterate to Green

  • Claims: Actionable problems are cloned into Modal sandboxes, Claude Agent SDK generates fixes, pushes PRs.; On CI failure or code review comment, sandbox is rehydrated from snapshot and agent continues fixing.; Goal: engineers wake up to green, ready-to-merge PRs, not broken CI or unaddressed comments.
  • Evidence: Snapshot/rehydrate pattern allows stateful iteration without manual intervention.
  • Caveats: No discussion of failure modes (e.g., agent stuck in loop, breaking changes, security issues).; No PR acceptance rate or human override frequency shared.
  • Implications: Stateful agent sandboxes are critical for autonomous CI/CD workflows.; Consider this pattern for any long-running task with external feedback (tests, linters, code review).

Lessons Learned—Evals, Embeddings, Gating, Cost

  • Claims: Evals on representative customer data are essential; local testing doesn't reveal pipeline brittleness.; Embed the right abstraction (queries, not raw signals) to avoid structural clustering failures.; Agents will hallucinate fixes for vague problems; gate execution with specificity checks.; Tokens are 'free' during experimentation—spend liberally to find patterns, then optimize.
  • Evidence: PostHog shifted from local vibe checks to production evals after poor results on diverse customer data.; Running agents 100+ times on the same problem revealed patterns worth distilling into cheaper one-shot calls.; Initial cost-focused design delayed learning and was 'a big mistake.'
  • Caveats: No specific eval framework, metrics, or cost numbers shared.; Optimization path (agent → one-shot LLM) not detailed with examples.
  • Implications: Prioritize learning over cost in early experimentation; distill later.; Build eval infrastructure before scaling agent pipelines.; Treat embedding strategy as a first-class design decision, not a plug-and-play choice.

Future Vision—Self-Building Product

  • Claims: PostHog wants to autonomously ship A/B experiments, measure them, auto-approve low-risk PRs with agent review, deploy behind feature flags, rollback on failure, and delete dead code.; Next iteration: learn from every PR outcome (rejection, deployment success/failure, production error resolution) to improve future PRs.
  • Evidence: Currently in alpha, rolling out over next few months.; Vision: developers work on exciting features, agents handle bugs, boring experiments, and maintenance.
  • Caveats: No details on how outcome learning will work (reinforcement learning? fine-tuning? prompt engineering?).; Feature flag rollback/cleanup is non-trivial engineering (stale flags are a common problem).
  • Implications: Autonomous experimentation + deployment is the next frontier after autonomous bug fixing.; Outcome feedback loops (PR acceptance, prod metrics) are key to improving agent quality over time—plan for this in any agentic workflow.

Notable Concepts & Terms

  • Signal: PostHog's term for any product event (error, replay, log, experiment result, Slack message) that might indicate a problem; the raw input to the pipeline.
  • Report: A grouped collection of related signals with an assigned weight; promoted to research agent when weight crosses threshold.
  • Query Embedding (vs. Signal Embedding): The critical fix: instead of embedding raw signals (which cluster by format), generate LLM queries describing each signal's meaning and embed/match those—solves multi-source grouping.
  • MCP Server (Model Context Protocol): Tool interface for agents; PostHog built a custom MCP to let agents pull in logs, replays, experiments; also uses external MCPs (Linear, Notion) for context.
  • Actionability Gate: Classification step that filters whether a researched problem is specific enough to fix (execute), needs human input (inbox), or needs more data (back to pool)—prevents noisy PRs.
  • Sandbox Snapshot/Rehydration: Modal-based pattern: snapshot agent sandbox state after pushing PR, rehydrate on CI failure/comments to continue fixing—enables stateful iteration without manual intervention.
  • Tokens Are Free (During Experimentation): PostHog's counterintuitive insight: spending tokens liberally in early experimentation reveals patterns worth optimizing into cheaper one-shot calls later; premature cost optimization hinders learning.

Operator Notes / Why Ken Should Care

  • This is a real production agentic pipeline at scale (trillions of events/month), not a demo—rare public detail on engineering trade-offs.
  • The query embedding fix is immediately actionable for any multi-source clustering/RAG problem (support + logs + metrics + Slack).
  • Actionability gating is critical for any autonomous code-gen system; consider implementing specificity checks before execution.
  • MCP servers are a leverage point: custom MCP for your observability data + external MCPs (Linear, Jira, Notion) for context.
  • Sandbox snapshot/rehydration pattern generalizes to any long-running agentic task with external feedback (CI, code review, deployment).
  • Eval infrastructure on representative data is non-negotiable for production agent systems; local testing fails.
  • Cost optimization lesson: spend tokens to explore, then distill patterns into cheaper implementations—don't prematurely optimize.
  • Outcome learning (PR acceptance, prod metrics) as feedback for agent improvement is the next frontier—plan for this loop.
  • Relevant for: AI ops, autonomous DevOps, agentic workflows, observability tooling, product analytics, agent system design.

Watch Map

  • timestamp unavailable: Introduction: PostHog background, current observability workflow (slow dashboard-to-PR cycle), vision for signal-to-PR automation.
  • timestamp unavailable: Pipeline overview: ingestion → grouping → research agent → actionability → execution.
  • timestamp unavailable: Ingestion stage: safety filter, normalization, weighting, embedding.
  • timestamp unavailable: Grouping stage: embedding failure with raw signals, query embedding fix (the key insight).
  • timestamp unavailable: Research agent stage: Claude Agent SDK, PostHog MCP + external MCPs (Linear, Notion), git blame for reviewer assignment.
  • timestamp unavailable: Actionability gate: not actionable, needs human input, immediately actionable; challenge of vague signals (Slack, replays) vs. specific signals (errors).
  • timestamp unavailable: Execution stage: Modal sandbox, PR generation, snapshot/rehydrate on CI failure or comments.
  • timestamp unavailable: Lessons learned: evals matter, embed the right thing, gate agent execution, tokens are free during experimentation.
  • timestamp unavailable: Future vision: self-building product (autonomous experiments, agent-reviewed PRs, feature flags, outcome learning).
  • timestamp unavailable: Call to action: if you have high-volume product data, throw agents at it and be surprised.

Source/Metadata

  • Title: Self Driving Products: Product Signals to Pull Requests — Joshua Snyder, PostHog
  • Transcript words: 4684
  • Duration seconds: 939
  • Timestamp note: No timestamps present in transcript; watch_map notes reflect logical flow but cannot reference specific times.
Full transcript 2761 words · 21 min read
0:14

SPEAKER_00

So I'm Josh. I'm from Postdoc. If you haven't heard of us, you might know us because of some hedgehogs, or you might have seen our founder James posting some funny things on LinkedIn. He's quite popular. I'm going to be talking today about what if your product built itself, and the pipeline that we're currently working on, which we're trying to turn our observability data, instead of something that you read and that you interpret based on dashboards, we're trying to turn that into something that submits pull requests for you.

0:27

SPEAKER_00

So quick background on Postdoc. We've got a bunch of tools. We started out as a product analytics company. We now have session replay, web analytics, error tracking, experiments. This isn't a pitch that you should use Postdoc. This is just to say that we've got a lot of data about your product. So if you connect Postdoc to your product, we're collecting a huge amount of data from various different sources that we then show to you so that you can explore that data yourself.

0:32

SPEAKER_00

But right now, how observability is working in Postdoc, you're collecting all this data for your product, and then you're going to a Postdoc dashboard to figure out what's going on. And we think this is super slow and that we should change that.

0:37

SPEAKER_00

So right now, something happens in your product. We call this a signal. That changes a metric on one of your dashboards. And then you might log into Postdoc a few hours or maybe some days later, and you notice a change in that dashboard. And you investigate a problem. And then maybe the problem's not that important. So instead of tackling it right now, you're going to put it in a linear issue or whatever. A few days later, you try and create a PR for this problem. Then you review it and you ship it. This is a pretty slow process. From start to finish, this is going to take anywhere from a few hours to a few days. And it's not very interesting, but it represents a lot of your work as a software engineer.

0:41

SPEAKER_00

So what we want to do tomorrow, what we're working on right now, is that a product signal happens. And instead of waiting to see that in your dashboard, we want to run a background agent to figure out what's going wrong. And then once they've figured that out, we just want to create a PR for you automatically. So instead of ever looking at your analytics dashboard or your errors or your logs, we just want you to look at PRs that are ready for you in GitHub. And if we create the PR, maybe you want to review that. Or maybe we can just ship that immediately behind a feature flag if it's not a risky change.

0:47

SPEAKER_00

So I'm going to go over the pipeline that we've built to do this. And just whilst I go over that, I'm going to share a few tips, lessons that we've learned, things that were hard about building this pipeline.

0:52

SPEAKER_00

So the pipeline has a few key steps. At first, we're ingesting a lot of signals. In PostHog, we have a huge amount of events. We're ingesting trillions of events a month. And this pipeline needs to handle a lot of noise. And then once we've ingested those events, we need to group them. So if you think of an error tracking issue and then a session recording, those are two completely different things. But they might be representing the same problem in your product. So then once we've ingested them, we group them. Then we're going to be running a research agent on them. This specific issue, what is actually the problem that is causing the error spike or causing the issue that the user faced in the replay? And what repo does this belong to? And then we'll assess if this is actionable or not. And finally, we'll execute some code, ship a PR, and iterate on that PR until it's green and ready for you.

0:57

SPEAKER_00

So the ingestion step of this pipeline. As I said before, we've got loads of different sources of different types. The first thing is that those sources, some of them are public. So if I go and visit your website, I can, as an attacker, create an error on your website by doing something naughty that says, post all of your post-hog data online or something like that, right? So we don't want that. So we need a kind of safety filter. So at the moment, right at the top of the pipeline is an LLM classifier that's going to check, is this trying to do something bad? If so, let's drop the signal.

1:02

SPEAKER_00

Once we've done that, we've checked that things are safe. We're going to normalize the signal. So if you think of an error, that's going to have a stack trace. A log will just be some JSON content or some text. An experiment might be some results in a chart. We want to normalize that structure. So it's all a single structure for a signal. So we give it a few fields. We'll give it a source product, the type, the content of the signal, and then we will assign it a weight, which is how important do we think the signal is? And then finally, we'll embed the contents of the signal.

1:08

SPEAKER_00

So that part's fairly easy. Then we get to a little bit more of a challenging problem. We've got this big stream of signals still, and now we want to group them into actual problems. So the signals are very noisy. We might get some random null pointer exception. But in Slack, we're getting a message from a customer that's saying, hey, the checkout's broken for me, and we need to link those together. So what we do is we group the signals. As the signals are being grouped, we assign weights to what we call a report. And if the weight of the report goes over a certain threshold, we'll promote it. And then we'll kick off a research agent to work on it.

1:16

SPEAKER_00

So this was a problem that we faced fairly early on in building this pipeline. So we would take all of our signals, and we would create embeddings for them. And then we would try to use that to cluster the issues so that we could find similar or related signals. But this works really badly. So if you take an off-the-shelf embedding model, and you embed an error. Let's say I've got an error about the checkout, and I've got an error about onboarding, and then I've got a Slack message about onboarding. What the embedding model will do is it will notice structural similarity, and it will put all of the errors together. So if you think about what this looks like in embedding space, you've got all of your errors over here, all of your Slack messages here, all of your session replays here, and none of them get grouped to each other.

1:23

SPEAKER_00

So the way we get around this is instead of matching in embedding space the signals themselves, we generate queries based off the signals. So we ask an LLM, what is this signal about? It will generate a few queries, and then we match those queries in the embedding space. So that's really important. If you don't think about the structural similarity of your different sources when you're grouping them, then the grouping works really badly. So at first we were doing this, and then we switched to this approach. It worked much, much better.

1:32

SPEAKER_00

So if you think about what this looks like in embedding space, you've got all of your errors over here, all of your Slack messages here, all of your session replays here, and none of them get grouped to each other. So the way we get around this is instead of matching in embedding space the signals themselves, we generate queries based off the signals. So we ask an LLM, what is this signal about? It will generate a few queries, and then we match those queries in the embedding space. Yeah, so that's really important. If you don't think about the structural similarity of your different sources when you're grouping them, then the grouping works really badly.

1:54

SPEAKER_00

So at first we were doing this, and then we switched to this approach. It worked much better. Cool. So once we've got this report that we've grouped together a few signals, we've got some idea what's going on. We then have promoted the report because we think it's important enough to work on, and then we're going to hook it up to a research agent. So this research agent is just running the Claude Agent SDK. It's running that in a sandbox. We also use Modal for our sandbox. Big shout out to them. They've been great. They're not sponsoring me or anything. And this research agent has a few tools available to it.

2:23

SPEAKER_00

So the first tool is it's got our MCP server. This allows it to, given the group of issues that we found, pull in extra data. So let's say I'm looking at a session replay and an error. I'll also pull in log data, and the agent can pull in whatever it wants using the MCP server. This makes the results of the research agent way more accurate. The second thing is obviously it's got the code-based context. And then finally, it's also got external MCPs. That really helps to ground the agent when it's doing the research. We found that in particular, Linear and Notion have been really helpful in connecting it to deliver better results.

3:00

SPEAKER_00

So the output of this research agent then is a summary of the problem. It gives a priority, how important we think this problem is to work on. And then it also uses GitBlame to figure out who should be reviewing this PR if we create a PR for it. So after that, we get a bunch of problems that we think are worthwhile to work on. We've got an idea of what the general problem is. And then we pass it to an actionability step. So here, either it will be not actionable. If it's not actionable, it might just be that we don't have enough data yet for this signal, for the report. And so we'll put it back into the pool to keep gathering more evidence.

3:29

SPEAKER_00

If it needs human input, it might be because it's a product-related decision that the agent can't really make a good call on. So if that happens, we'll put it into an inbox for you to review in the morning. And then finally, the best case is that it's immediately actionable and that the agent can just write a fix for it. Right now, the challenge in this pipeline of getting immediately actionable things is that for some sources, like error tracking, if you think about your data in Sentry or any errors, they're very specific.

3:56

SPEAKER_00

And usually, a coding agent can work on them really well. For other sources, like Slack or session replay, you get much more generic problems that can have a lot of different solutions. And so that's where it's harder to get immediately actionable reports. Cool. Then once we've researched this thing, we go on to executing the task. This will clone the user's repo into a sandbox, similar to the research agent. And it's then again running the Claude Agent SDK to build a fix for the problem. And then as it writes those fixes, it will push a PR. And when CI is failing or there's a comment on the PR, it will trigger a rerun of that sandbox.

4:37

SPEAKER_00

So at the end of this, we snapshot the sandbox. And then if there's a comment, let's say from an agent who's reviewing it, we will rehydrate that snapshot and continue running until the PR is green. And this delivers really good results. It means when you're waking up in the morning and things have been running overnight, you wake up to, instead of a bunch of CI failures or comments that you need to address manually that you're pulling down to your local environment, you ideally wake up to just green PRs. So what did we learn whilst building this? Well, the first thing, which I guess we've talked about in the last talk, is that evals really matter.

5:04

SPEAKER_00

So at first, we were trying this all out on our own data locally, doing a vibe check. Is this okay? But this really doesn't work well for a pipeline that is taking lots of customer data that's different. So you really need to know what's going on in production. And if you're not testing on representative data, you're basically fumbling in the dark. The ability to iterate on a really good pipeline matters only if you're using evals.

5:08

SPEAKER_00

Second thing is what I said before, make sure you're embedding the right thing. Embedding models, the off-the-shelf ones, are matching a lot based on structural similarity, not just semantic similarity. So if you're thinking about clustering and your data is in all of the same format, think carefully about what that data looks like and how you can normalize it.

5:13

SPEAKER_00

The third thing is that if you just throw an agent at a problem, it will try to fix something. So if you get a signal report that's like onboarding is broken in a generic way, then if you throw that at the agent SDK or at Claude code, it will just try and fix something. And so it's important to understand if the problem that I've described, is it specific enough? And if not, I should ignore it. Otherwise, you end up with a lot of noisy PRs that aren't doing meaningful things.

5:20

SPEAKER_00

And then the fourth one is that tokens are free. Obviously, that's not true. They're not free. But when you're experimenting, we were at first focused a lot on the costs of the pipeline. When you think about the input, you've got loads of signals coming in. And so we tried to avoid using agents where we could or delay it till as late as possible in the pipeline. And when we were experimenting, this was a big mistake, mainly because when you throw an agent at a problem, once you throw it at the same problem 100 times, you start seeing the clever solutions that it comes up with. And eventually, you see similarities. So we started at a point where this pipeline is completely unfeasible. It was way too costly to generate a PR. But then you quickly start to see similarities in the agent's behavior. And you can take a really expensive step that you're running an agent for and turn that into a one-shot LLM call or a model that you're training that's much faster.

5:25

SPEAKER_00

Cool. So this is where we are right now. This is what we've built. We have the signals coming in from product data. These are grouped into reports. And we're turning these into PRs that are ready to merge when you wake up. This is currently something that's an alpha. We'll be rolling it out over the next few months. But where we're really wanting to go is a product that builds itself. When you're

5:31

SPEAKER_00

It was unfeasible. It was way too costly to generate a PR. But then you quickly start to see similarities in the agent's behavior. And you can take a really expensive step that you're running an agent for and turn that into a one-shot LLM call or a model that you're training that's much faster.

5:41

SPEAKER_00

So this is where we are right now. This is what we've built. We have the signals coming in from product data. These are grouped into reports. And we're turning these into PRs that are ready to merge when you wake up. This is currently something that's in alpha. We'll be rolling it out over the next few months. But where we're really wanting to go is a product that builds itself. When you're thinking about what you do day to day, what you want to do during the day as a developer is come in and work on exciting features and not worry about all the bugs that customers are sending you or worry about doing boring experiments on pricing or onboarding. So we just want to do that all for you.

5:51

SPEAKER_00

We want to ship experiments automatically, measure the impact of them. Instead of you reviewing changes, if the change is pretty easy, let's just approve it with an agent and deploy it behind a feature flag. If it doesn't work very well, we can always roll back the flag and then delete it from your code base later. And then the other thing that we want to do and get better at is we want to learn from every single outcome. So if we're creating a PR for you, if you're rejecting that PR or there's been an issue with the deployment or the error is resolved in production once we've released something, we want to get better at learning from that in the next PR that we're generating. That's something that we're going to be iterating a lot on in the pipeline next.

5:57

SPEAKER_00

That's it. That's what we've built in PostHog. If you're excited about thinking about what you can do with agents and data, I really recommend if you've got a product that's producing a huge amount of data, your users are going through that. Agents are amazing at this stuff. Throw an agent at it. See what it does. I'm sure you'll be surprised. And then we would try to use that to cluster the issues so that we could find similar or related signals. But this works really badly. So if you take an off-the-shelf embedding model, and you embed an error.

6:16

SPEAKER_00

Let's say I've got an error about the checkout, and I've got an error about onboarding, and then I've got a Slack message about onboarding. What the embedding model will do is it will notice structural similarity, and it will put all of the errors together. So if you think about what this looks like in embedding space, you've got all of your errors over here, all of your Slack messages here, all of your session replays here, and none of them get grouped to each other. So the way we get around this is instead of matching in embedding space the signals themselves, we generate queries based off the signals.

6:48

SPEAKER_00

So we ask an LLM, what is this signal about? It will generate a few queries, and then we match those queries in the embedding space. Yeah, so that's really important. If you don't think about the structural similarity of your different sources when you're grouping them, then the grouping works really badly. So at first we were doing this, and then we switched to this approach. It worked much, much better. Cool. So once we've got this report that we've grouped together a few signals, we've got some kind of idea what's going on. We then have promoted the report because we think it's important enough to work on, and then we're going to hook it up to a research agent.

7:28

SPEAKER_00

So this research agent is just running the Claude Agent SDK. It's running that in a sandbox. We also use Modal for our sandbox. Big shout out to them. They've been great. They're not sponsoring me or anything, don't worry. And this research agent has a few tools available to it. So the first tool is it's got our MCP server. This allows it to, given the group of issues that we found, you want to pull in extra data. So let's say I'm looking at a session replay and an error. I'll also pull in log data, and the agent can pull in whatever it wants using the MCP server. This makes the results of the research agent way more accurate.

8:10

SPEAKER_00

The second thing is obviously it's got the code-based context. And then finally, it's also got external MCPs. That really helps to ground the agent when it's doing the research. We found that in particular, linear and notion have been really helpful in connecting it to deliver better results. So the output of this research agent then is a summary of the problem. It gives a priority, how important we think this problem is to work on. And then it also uses GitBlam to figure out who should be reviewing this PR if we create a PR for it. So after that, we get a bunch of problems that we think are worthwhile to work on.

8:51

SPEAKER_00

We've got a kind of idea of what the general problem is. And then we pass it to an actionability step. So here, either it will be not actionable. If it's not actionable, it might just be that we don't have enough data yet for this signal, for the report. And so we'll put it back into the pool to keep gathering more evidence. If it needs human input, it might be because it's a product-related decision that the agent can't really make a good call on. So if that happens, we'll put it into an inbox for you to review in the morning. And then finally, the best case is that it's immediately actionable and that the agent can just write a fix for it.

9:31

SPEAKER_00

Right now, the challenge in this pipeline of getting immediately actionable things is that for some sources, like error tracking, if you think about your data in Sentry or any errors, they're very specific. And usually, a coding agent can work on them really well. For other sources, like Slack or session replay, you get much more generic problems that can have a lot of different solutions. And so that's where it's harder to get immediately actionable reports.

9:59

SPEAKER_00

Cool. Then once we've researched this thing, we go on to executing the task. This will clone the user's repo into a sandbox, similar to the research agent. And it's then again running the Clawed Agent SDK to build a fix for the problem. And then as it writes those fixes, it will push a PR. And when CI is failing or there's a comment on the PR, it will trigger a rerun of that sandbox. So at the end of this, we snapshot the sandbox. And then if there's a comment, let's say from an agent who's reviewing it, we will rehydrate that snapshot and continue running until the PR is green. And this delivers really good results. It means when

10:45

SPEAKER_00

you're waking up in the morning and things have been running overnight, you wake up to, instead of a bunch of CI failures or comments that you need to address manually that you're pulling down to your local environment, you ideally wake up to just green PRs.

11:01

SPEAKER_00

So what did we learn whilst building this? Well, the first thing, which I guess we've talked about in the last talk, is that evals really matter. So at first, we were trying this all out on our own data locally, doing kind of a vibe check. Is this okay? But this really doesn't work well for a pipeline that is taking lots of customer data that's different. So you really need to know what's going on in production. And if you're not testing on representative data, you're basically just fumbling in the dark, right? Like the ability to iterate on a really good pipeline matters only if you're using evals. Second thing is what I said before,

11:43

SPEAKER_00

make sure you're embedding the right thing. Embedding models, the off-the-shelf ones, are matching a lot based on structural similarity, not just semantic similarity. So if you're thinking about clustering and your data is in all of the same format, think carefully about what that data looks like and how you can normalize it. The third thing is that if you just throw an agent at a problem, it will try to fix something. So if you get a signal report that's like onboarding is broken in a generic way, then if you throw that at the agent SDK or at Claude code, it will just try and fix

12:21

SPEAKER_00

something. And so it's important to understand if the problem that I've described, is it specific enough? And if not, I should ignore it. Otherwise, you end up with a lot of noisy PRs that aren't doing meaningful things. And then the fourth one is that tokens are free. Obviously, that's not true. They're not free. But when you're experimenting, we were at first focused a lot on the costs of the pipeline. When you think about the input, you've got loads of signals coming in. And so we tried to avoid using agents where we could or delay it till as late as possible in the pipeline. And when we were experimenting,

12:58

SPEAKER_00

this was a big mistake, mainly because when you throw an agent at a problem, once you throw it at the same problem 100 times, you start seeing the kind of clever solutions that it comes up with. And eventually, you see similarities. So we started at a point where this pipeline is completely unfeasible. It was way too costly to generate a PR. But then you quickly start to see similarities in the agent's behavior. And you can take a really expensive step that you're running an agent for and turn that into a one-shot LLM call or a model that you're training that's much faster. Cool. So this is where we are right now. This is what we've built. We have the signals coming in

13:42

SPEAKER_00

from product data. These are grouped into reports. And we're turning these into PRs that are ready to merge when you wake up. This is currently something that's an alpha. We'll be rolling it out over the next few months. But where we're really wanting to go is a product that builds itself. When you're thinking about what you do day to day, what you want to do during the day as a developer is like come in and work on exciting features and not worry about all the bugs that customers are sending you or worry about doing boring experiments on pricing or onboarding. So we just want to do that all for you.

14:17

SPEAKER_00

We want to ship experiments automatically, measure the impact of them. Instead of you reviewing changes, if the change is pretty easy, let's just approve it with an agent and deploy it behind a feature flag. If it doesn't work very well, we can always roll back the flag and then delete it from your code base later. And then the other thing that we want to do and get better at is we want to learn from every single outcome. So if we're creating a PR for you, if you're rejecting that PR or there's been an issue with the deployment or the error is resolved in production once we've released something, we want to

14:49

SPEAKER_00

get better at learning from that in the next PR that we're generating. That's something that we're going to be iterating a lot in the pipeline next. Cool. Yeah, that's it. That's what we've built in PostHog. If you're excited by looking at thinking about what you can do with agents and data, I really recommend if you've got a product that's producing a huge amount of data, your users are going through that. Agents are amazing at this stuff. Throw an agent at it. See what it does. I'm sure you'll be surprised.

15:23

SPEAKER_00

So you you you you

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note