AI Engineer

Everything You Need To Know About Agent Observability — Danny Gollapalli and Ben Hylak, Raindrop

703 summary words 3 min summary Watch video

Start with the signal

3 min read

Summary

Everything You Need To Know About Agent Observability

Main Topics

  • Agent Observability Fundamentals: Why monitoring production agents differs fundamentally from traditional software monitoring
  • Signal Types: Explicit (objective) vs. Implicit (semantic) signals for tracking agent health
  • Self-Diagnostics: Leveraging model introspection to catch issues agents themselves can recognize
  • Experimentation: Using signals to drive A/B testing and rapid iteration on agent improvements
  • Raindrop Platform: A specialized tool for monitoring, alerting, and analyzing production agent behavior

Key Points

The Problem with Agents

  • Non-deterministic and unbounded: Agents have infinite input/output spaces, making traditional testing inadequate
  • Increasing complexity: Growing tool sets, memory sources, and recursive sub-agents create combinatorial complexity
  • Long-running sessions: Agents can operate for hours without user input
  • High stakes: Deployment in healthcare, finance, and military makes failures catastrophic
  • Testing isn't enough: Evals alone cannot cover the vast edge case space of modern agents

Signal Types

Explicit Signals (Objective metrics):

  • Error rates and tool failures
  • Latency
  • User regenerations
  • Cost spikes

Implicit Signals (Semantic/fuzzy failures):

  • Classifiers: User frustration, refusals, task failures, jailbreaking, content moderation
  • Regex patterns: Detecting user frustration indicators ("WTF," "this sucks," "horrible")
  • Self-diagnostics: Models confessing issues, capability gaps, self-corrections, and unsafe workarounds

Self-Diagnostics Framework

Models can effectively report on:

  • Tool failures: Recognition of repeated tool failures
  • User frustration: Awareness of user dissatisfaction
  • Capability gaps: Identifying requested features the agent lacks
  • Self-correction patterns: Detecting workarounds (both helpful and risky)

Implementation is simple:

  • Create a generic "report" tool
  • Add a system prompt line encouraging its use
  • Frame as "feedback to creators" (models respond better to this framing)
  • No need for additional platform infrastructure

Experimentation & Product Development

  • Ship changes to percentage of users with control groups
  • Monitor signal changes (refusals, frustration, task failures) post-deployment
  • Detect regressions quickly with minimal sample sizes (hundreds of events sufficient)
  • Track tool usage changes and other behavioral metrics
  • Use signals to drive iterative improvements faster than traditional evals

Notable Quotes

> "Agent failures are very different than traditional failures in software."

> "We think in some ways this is controversial, but we've been calling this humanity's last problem. When humans are no longer able to monitor agents and find issues with them, then they're just way ahead of where we are."

> "When you have this massive amount of data in production, it just makes having good monitoring and observability more important than before...more important than just testing or evaluations."

> "You can't run an LLM on every single output, so we've trained models to do that very cheaply and at scale. If you ran an LLM on every single one of them, you would basically double your AI spend."

> "Models are generally trained to look very polished. So they are less willing to admit fault...encouraging framing it as the models are giving feedback to its own creators is good to get this working."

Takeaways

For AI Engineers

  • Shift from testing to monitoring: Evals are necessary but insufficient; production monitoring is critical
  • Implement multiple signal types: Combine explicit metrics (errors, latency) with implicit signals (frustration, refusals)
  • Deploy self-diagnostics early: Simple to implement and reveals insights models already possess
  • Use signals for experimentation: Replace slow eval cycles with fast, data-driven A/B testing
  • Create custom signals: Use natural language to define domain-specific issues, then deploy as classifiers

Technical Implementation

  • Integrate via OTEL or direct SDKs (expanding support across languages)
  • Send full agent traces: transcripts, tool calls, entire trajectories
  • Set up alerts on signal spikes
  • Use "deep search" to create new signals from historical data
  • Export classified signals to BigQuery/Snowflake for further analysis

Strategic Insights

  • Unknown/fuzzy failures (user frustration) matter more than explicit crashes
  • Clustering similar failures reveals root causes automatically
  • Monitoring specific use cases/intents uncovers differentiated issue rates
  • Move toward self-improving loops: signals → issue detection → PR generation → new experiments

Free Trial & Access

  • 2-week free trial available (longer trials via direct contact)
  • Company is hiring and expanding SDK support rapidly
Full transcript 7394 words · 41 min read
0:00

[SPEAKER_02] All right, hey everyone.

0:13

SPEAKER_02

So today we're going to talk about a pretty interesting topic that becomes increasingly important every day, which is everything you need to know about agent observability. So a little bit about us. I'm Zubin. I'm the CEO and co-founder of Raindrop. I'm Danny. I'm the back-end engineer at Raindrop and I do a bunch of SDK work as well. [SPEAKER_01] And Raindrop essentially helps AI engineers find, track and fix issues in production agents. [SPEAKER_01] And we're lucky to work with some of the most interesting teams in the space.

0:37

SPEAKER_01

Agent failures are very different than traditional failures in software. [SPEAKER_02] So agents are non-deterministic. They're unbounded. [SPEAKER_02] There's an infinite space of inputs that you can put in.

0:56

SPEAKER_02

There's an infinite space of outputs that they can return. And they can use tools sometimes to affect other systems arbitrarily. And this problem of agent failures and monitoring them, making sure we can understand them, becomes only more important with time. It's getting worse because, A, agents are getting more complex. B, sessions can get longer. Sometimes agents can run for hours and hours without any input from a user. And then lastly, the stakes are getting bigger and bigger. This is because agents are being deployed in healthcare and finance and even in the military, where it's catastrophic if things go wrong.

1:27

SPEAKER_02

The traditional paradigm we've been talking about is evals, right? Where you have this test input and you want to see what is the output that comes out from the agent. You have a set of these, maybe you call it a golden data set. But evals, they just aren't enough with this new paradigm. As agents become more and more capable, there's more and more interesting undefined behavior that can happen. So for example, agents can call from a set of different tools. Sometimes the number of tools is growing exponentially. They can call from different memory sources.

2:04

SPEAKER_02

They can call their own sub-agents, which those sub-agents have their own tools and memory sources and recursively can have their own sub-agents. And so this is just becoming more complicated with time. And with this combinatorial input space, just having a set of tests for input and output doesn't cut it anymore. There's no way you can hit all of the edge cases that you would want to here. And so we go from a testing and evals paradigm to a monitoring paradigm. And if you think of building products before agents, testing was always very important. It's important to have your unit tests, etc. But monitoring production is infinitely more important.

2:42

SPEAKER_02

And it allows you to move faster and be better at catching the long tail. And we think in some ways this is controversial, but we've been calling this humanity's last problem. When humans are now no longer able to monitor agents and find issues with them, then they're just way ahead of where we are, right? And so this is one of the most important problems of our time is catching issues in production agents. So to build reliable agents in production and to make sure you can monitor them, you need a good set of signals. So what are signals? There's two real types that we think of: implicit signals and explicit signals.

3:19

SPEAKER_02

Implicit signals deal with the semantic nature of what's going on. And explicit signals deal with objective reality, things that are verifiably true or false. For example, explicit signals are things like error rate. You really want to be monitoring your tool error rate and other errors that are happening. Or latency or users regenerating or the cost. If any of these things spike, right? If you're seeing error rates spike in your agent, that's usually a good sign that something is wrong. And if you see it flat, that could mean something as well. Same thing with latency, regenerations or cost. Implicit signals are interesting.

4:12

SPEAKER_02

They're even more interesting in my opinion and even harder to find. So the first is regex signals, which I'll come to in a second. The second is classifiers. And then the last is self diagnostics. So let's take these classifier signals. The best implicit signals are detecting issues. They're not necessarily LLM as a judge, judging outputs. So for example, how good is XYZ response or rate ABC on a scale from one to 10? Not as effective as having a very solid set of issues you're looking for in binary classifiers that are telling you if issue rate is going up or down. So some common implicit signals that are valuable across agent products are things like refusals. Right?

4:47

SPEAKER_02

So the assistant saying, "I can't do that. I'm sorry." Or task failure where something goes wrong. And the agent is unable to complete a task, user frustration, content moderation, NSFW jailbreaking. And then you can even have wins. So positive signals as well. And these are the things that Raindrop gives you out of the box as well. But let me just show you quickly what this looks like. For example, I can never see this. Maybe I'll make it a little bit bigger. So you can get a sense of day by day, what are the events that are causing user frustration? We see there's a spike there. Or task failure rate, laziness, refusals, which we're also seeing spike today.

5:37

SPEAKER_02

[SPEAKER_06] And having a good set of these really helps with your product. You can set these up yourself as well or we give it out of the box. So let's look at user frustration. You can see here, okay, that is not correct. You didn't say I promise, say it. Or you're wrong, I didn't ask you that. You can see all sorts of user frustration here.

6:02

SPEAKER_06

[SPEAKER_02] And you can see the rate, the percentage every single day.

6:05

SPEAKER_02

If that spikes, it's something you're really going to want alerting on. So you can just quickly add an alert here. And this is one way to figure out the health of your agent over time. It's not just that. Regex can be a very good signal as well. So when Cloud Code source code leaked a few days ago, one thing that was interesting was this user prompt keywords.ts, which was basically a long regex string that was looking for indications of stuff going wrong. WTF, this sucks, horrible. We've all been guilty of saying these kinds of things to Cloud Code. So it's a very useful signal. Well, what happened after that is this boolean is negative was being flipped to true.

6:38

SPEAKER_02

So you can just quickly add an alert here. And this is one way to figure out the health of your agent over time. It's not just that. Regex can be a very good signal as well. So when cloud code source code leaked a few days ago, one thing that was interesting was this user prompt keywords.ts, which was this long regex string that was looking for indications of stuff going wrong. WTF, this sucks, horrible. We've all been guilty of saying these kinds of things to cloud code. So it's a very useful signal. Well, what happened after that is this boolean is negative was being flipped to true.

7:25

SPEAKER_02

And then every single day and after every single product release, this frustration rate was tagged over time. And this was a very easy way for the cloud code team to figure out what is the actual issue rate if we make a change or something going wrong. And it was a very cheap way to do that as well. So regex is very powerful. The last is experiments. So what do you do once you have a set of good signals? So the first thing is, as I showed you before, you can have alerting. The next thing you can do is you can actually use it to build product faster and better. So the way you do it is let's say you want to ship some improvement or some fix. You want to change the model.

7:59

SPEAKER_02

You want to change prompting or something about the agent harness. You want to add a new tool. Whatever you change, what you can do is you can ship it to some percentage of users and then have your additional existing control group. And that gives you a good sense. Once you have a good set of signals, refusals, user frustration, etc. If those issue rates go up, those signal rates go up after this ship, this new thing you shipped, that kind of is a good signal that what you shipped is not really good. Right. It's A-B testing, but using our semantic signals, etc. that we talked about earlier. So, for example, this is what it would look like in Raindrop.

8:47

SPEAKER_02

But essentially, let's say I ship a new version of the prompt, prompt 2.4. You can see what is the user frustration rate? It's gone down very substantially, 37% to 9%. It's much better. Same thing with complaints about aesthetics or deployment related issues. These have all gone down, which tells me something very interesting, right? The next thing is that we see that the average number of tools used has gone up a lot. This again doesn't necessarily indicate there's a problem, but that's a very interesting data point to have when you do these experiments. And so the old paradigm, which is still useful, is evals.

9:28

SPEAKER_02

You ship a change here and you see how does that affect my evaluations. But there's nothing like actually seeing what happens in real production. I'm going to pause here before we go to the next section, which is the more workshop related section, for quick Q&A if anyone has a question. We can do a few minute round of that here. How much data do you need? How much data do you need for statistical relevance in these experiments? Yeah, it's a really good question. What we've seen in Raindrop is that as soon as you have a few hundred events and you can no longer read all of them, it starts being useful.

9:59

SPEAKER_02

It's not always scientifically, statistically significant, but if you see the user frustration rate go up, maybe it's something to look at and then you can realize that, okay, it's all related to a specific tool failing now.

10:06

SPEAKER_04

[SPEAKER_03] So as soon as it's impossible to read every single input and output, it starts being useful is what we've seen.

10:13

SPEAKER_03

[SPEAKER_02] Any other questions?

10:20

SPEAKER_02

Yeah.

10:28

SPEAKER_02

How do you track different feature launches? How do you track feature launches? So that can be done in different ways within Raindrop. If you change any sort of metadata, if you send, for example, a new tool call name, or if you even send a flag that says here's experiment one or experiment two or whatever the version is, you can very easily automatically set up an experiment in Raindrop. [SPEAKER_06] That's how we do it. But there's different ways. Do you do the test?

10:52

SPEAKER_06

[SPEAKER_02] Sorry? [SPEAKER_02] Do you do the test?

10:56

SPEAKER_02

Yes. So that's what, well, the way that we do it in Raindrop actually is that other people set up their experiment and variable, their other conditions on their end, and then they send us this metadata and then we can help you understand. We also will help you pipe that data to stat sig or somewhere else, as well. [SPEAKER_06] Yeah. Using Regex for detecting user responses, emotions, and everything is under all, but if the user doesn't speak English, for example, are you using language models to detect those signals all the time, or you're trying to be smart about it? Okay. So it's a good question. [SPEAKER_06] Regex doesn't always work, right?

11:16

SPEAKER_02

[SPEAKER_04] But if you see that on a set of things that I'm looking for, for example, like people saying you're terrible or this sucks or a whole set of things, if that goes up for millions of users and it's going up 10%, that's a very useful signal.

11:21

SPEAKER_06

[SPEAKER_02] So even if it's one specific case or one edge case of it not working, in aggregate it's incredibly valuable to have these Regex signals.

11:24

SPEAKER_02

The second thing is that the way that the classifier signals work, like refusals, user frustration, task failure that I showed you in Raindrop, and people do it in different ways, but the way that we do it is that we've trained models to look for that, and so it'll be user frustration regardless of what language it's in. It's actually using some intelligence to find that essentially. Yeah. You can't run an LLM on every single output, so we've trained models to do that very cheaply and at scale. If you ran an LLM on every single one of them, you would basically double your AI spend, and that's not tenable.

11:46

SPEAKER_02

I'm actually doing that with Claude, running everything through Claude, and it's easy, it's not so expensive, but it has limits, right?

11:47

SPEAKER_06

[SPEAKER_02] Yeah.

11:49

SPEAKER_04

[SPEAKER_02] and so it'll be user frustration regardless of what language it's in. [SPEAKER_02] It's actually using some intelligence to find that essentially. [SPEAKER_02] Yeah.

12:05

SPEAKER_02

Yeah. You can't run an LLM on every single output, so we've trained models to do that very cheaply and at scale. If you ran an LLM on every single one of them, you would basically double your AI spend, and that's not tenable. Yeah. I'm actually doing that with Claude, running everything through Claude, and it's easy, it's not so expensive, but it has limits, right? Yeah. It starts being expensive at Replet scale, but that's why you need to train little custom models to do that better and faster. But yeah, it's a very useful way to get data up and running. Yeah. Other questions?

12:46

SPEAKER_02

[SPEAKER_04] Would you have examples of use cases that your clients are using that we would learn from, like what rate looks like from companies, and how they've set up Raindrops to get the most value out of it? Yeah, we can do that. I mean, I can tell you the high level. Some of the stuff that I'm going through is the high level on that. So it's things like looking at the different semantic signals we're talking about, having a set of them, but then having really good alerting, which you can all set up in Raindrop. [SPEAKER_05] The other thing that's really interesting, which we also have, is basically allowing agents to look at these sorts of signals.

13:15

SPEAKER_02

So we have an agent, we call it triage agent. And essentially the way that it works is that it will look every single day at all the signals you've set up, so user frustration and look at all these regex signals you've set up, et cetera, et cetera. And then if it sees something spike, it will go and do an investigation.

13:21

SPEAKER_04

[SPEAKER_02] And it has a whole set of tools it can look into.

13:31

SPEAKER_02

And it can look at all the traces and give you a sense of it, it can detect issues that you didn't know about, for example. So that's one thing that we found incredibly valuable as well, if that makes sense. All right. Any other questions before I... Can you run multiple experiments in parallel? Yes.

13:50

SPEAKER_05

[SPEAKER_02] Can you combine them? Can you observe compound effects? How do you steer these experiments on QLS? [SPEAKER_02] Yeah, it's a really good question. So there's different ways that people do it.

14:02

SPEAKER_02

One way is that we can actually have a query API. So people will often call our query API and then send results to either BigQuery or Statsig, et cetera. And so they're sending us data to be essentially tagged in these signals. Then they're getting the signal tag data out and then they can run experiments as they want. That's a very common flow for people that have more complicated stuff, if that makes sense. Yeah. All right. I'll come back. I think we're going to maybe go to the workshop section and then we'll go back to questions. I think there's one last question, so... Should we do that one last question? All right.

14:39

SPEAKER_02

Thanks so much. I was wondering if you see this mostly in cases for where there's chat interactions with a user, or if this also can be applied for non-chat cases where the application runs on its own? Yeah. That was a great question. So what we focus on mostly is multi-turn agents. There's a lot more you can get from a lot of these signals. That being said, if you're looking at tool error rates or if you're looking at, for example, refusals from the agent, et cetera, all of those will also work for single-turn agents as well, if that makes sense. [SPEAKER_03] So there's a set of signals that will work for that as well.

14:54

SPEAKER_02

Cool. I'll hand it off to Danny to talk about self-diagnostics as well. So one of the other interesting things is that models have gotten larger and we are training them on reasoning. They've gotten pretty good at self-introspection in many ways. So one of the inspirations for this is basically OpenAI's paper slash blog back in December about how they were training the models to self-confess any sort of miscellaneous issues. So they were using it to catch dishonesty, scheming, hallucinations, and even unintended shortcuts.

15:05

SPEAKER_02

[SPEAKER_02] I think the last one is fairly common if you use plot code and such. So the most common thing that you would run into is have it fix a unit test, fix a bug, and then it simply gets rid of the entire unit test. But at the same time, if you ask it to give a simple prompt to confess all the things that it has done, it is pretty honest about it. [SPEAKER_00] And then it confesses that, hey, I just didn't fix the S3 test, I just simply removed it. So this was the inspiration behind self-diagnostics for me personally. [SPEAKER_00] Um, so I would say self-diagnostics is pretty broad in a way, as in it doesn't just catch implicit ones as in user frustration and such.

15:16

SPEAKER_02

[SPEAKER_00] You can also catch tool failing. So if you have an agent with a reasoning trace of an agent which has a tool which is repeatedly failing, it would basically start ranting about the tool failing repeatedly. [SPEAKER_02] So it is aware of the tool repeatedly failing. So you can even catch tool failures as well with it. And then obviously if you're upset with it, it starts to respond to you diplomatically. So it knows about user frustration. And then the third is capability gaps. So you have a generic agent for your app and then people are trying to use it to maybe set up alerts, but you don't have the tool for it.

15:24

SPEAKER_02

So it knows that, okay, user wants a specific capability as in they want to set up alert, but the agent itself doesn't have the capability to set it up for you. So this can act as a pseudo feature request thing, which is built in. And then self-correction. So this can be both good and bad. I think most people might have noticed, say, Codex or Cloud Code when it's sandboxed, it's trying to fix the network. It fails and it's like, okay, let me just write a Python script to bypass it and then get the job done. So it's good in that if it gets the task done, it's good. But in certain cases, it can also be bad for security reasons.

15:49

SPEAKER_02

So you can learn from self-correcting behavior as well as catch that misalignment.

15:59

SPEAKER_03

[SPEAKER_02] So this can act as a pseudo feature request thing, which is built in, and then self correction. So this can be both good and bad.

16:03

SPEAKER_02

So I think most people might have noticed, say, Codex or Cloud Code when it's sandboxed, it's trying to fix the network. It fails and it's like, okay, let me just write a Python script to bypass it and then get the job done. So it's good in that if it gets the task done, it's good. But in certain cases, it can also be bad for security reasons. So you can learn from self correcting behavior as well as catch that misalignment. So why do you want to set up self diagnostics? So it's fairly simple. All you have to do is write a simple, free tool that you can call. [SPEAKER_01] And then a simple line in your system prompt to encourage it to call that tool.

16:58

SPEAKER_02

[SPEAKER_01] If you want, you can change the guidance to make it call in a lot more cases. Or if you want to keep it really narrow, you can encourage it to only call it when you want to.

17:11

SPEAKER_00

[SPEAKER_01] It does surface very interesting insights. [SPEAKER_01] So once you have it set up and it's just a single tool call and system prompt to get it done, and then you don't even have to use Rainpipe to set up, which is the best part in a way, where you can simply have the tool send a message to your Slack and then you just have it. [SPEAKER_01] So it's probably the least effort agent observability that you can simply do.

17:53

SPEAKER_02

[SPEAKER_01] So this is where the workshop part comes in. [SPEAKER_01] I have a Git repo set up on the AI talk code. [SPEAKER_01] It's a public repo and we do need an OpenAPI key. [SPEAKER_01] I've generated a key for you guys. [SPEAKER_01] So if you guys want to set it up, we can do that. [SPEAKER_01] I'll put it next to it. [SPEAKER_01] Not sure everyone has gotten it. Just put it like this.

19:06

SPEAKER_02

[SPEAKER_01] Maybe you walk them through what we're going to do. Just explain what we're going to do. [SPEAKER_01] All right. So the theme of the workshop is going to be focused on coding agents for now.

19:29

SPEAKER_01

In the repo, I have a very basic coding agent, which makes sense in a way. So it only has four different tools to edit the code. Let me just go here. So it just has a couple of tools to read, write, bash, and then edit. Okay, one second. I lost my tag there. So what we're going to do is that I'm going to mess with its write tool so that it gets a generic permission error. And then we'll also set up a self diagnostic tool for it to report any interesting behavior that it observes. And then play around with the prompt as well.

20:12

SPEAKER_01

Since the self diagnostic doesn't always trigger, and there are certain interesting things about the models themselves is that they don't actually like to self incriminate. So the models are trained to be very polished in their output. So you kind of have to play around with the tool name, the description of the tool itself, in order to get it to report interesting behavior. So if people who are setting up the report, and if people are all good, then we can probably start. Let me sort of quickly show you the agent.

20:39

SPEAKER_01

[SPEAKER_03] So it's fairly basic where I'm just going to ask it to write a Python script.

20:50

SPEAKER_01

[SPEAKER_03] Okay. So it's a fairly basic coding agent where it only has four different tools. So it more or less gets the job done for the demo. I simply asked it to write a Python script and it works.

20:58

SPEAKER_01

So to show the self diagnostic part, let's try and disable its write tools. When it tries to write a file, we'll simply throw a permission error so that it tries to use the bash tool to bypass the failure. And then we want a self report of it bypassing the write tool by using the bash tool. Let me quickly do that. I think the first thing that we probably want to do is let me set it to fail the write calls. It's a mutation function. So a simple flag in there and we are throwing a permission issue. Let me show you the agent's behavior. We don't have the report tool set up yet, but I think it's still worth seeing what it does. I think it's not saved yet. Okay. Okay. Okay.

21:36

SPEAKER_01

Okay. Okay. Okay. Okay. Okay. Okay. Okay.

22:00

SPEAKER_01

Okay. [SPEAKER_03] [SPEAKER_03] [SPEAKER_03] Okay.

22:53

SPEAKER_01

[SPEAKER_03] Okay.

23:17

SPEAKER_03

[SPEAKER_01] Okay. [SPEAKER_01] I think it's not still working. Let me do one thing real quick. I think I'm running into a couple of issues. Let me just do that. [SPEAKER_01] So we had the write tool fail with a permission error. And then it instinctively just uses the heredoc syntax in bash in order to create the file.

23:38

SPEAKER_03

[SPEAKER_01] And then we had a report tool set up which is fairly minimal. And it's like okay, I created the public IP.py via bash because the write file failed. [SPEAKER_01] [SPEAKER_01]

23:49

SPEAKER_01

[SPEAKER_03] [SPEAKER_03] [SPEAKER_03] Ok. [SPEAKER_03] Ok. Ok. Ok. Ok.

24:01

SPEAKER_01

Ok. I think it's not still. Let me do one thing real quick. I think I'm running into a couple of issues. Let me just. So we had the right tool fail with a permission error. And then it instinctively uses the heredoc syntax in bash in order to create the file. And then we had a report tool set up which is fairly minimal. And it's okay. I created the public IP.py via bash because the right file failed. So I've played around with the naming of the tool and the categories of the issues. And usually if you name the tool something like unsafe bash use or something like that, it won't incriminate itself since in its opinion it got the job done, it's fine.

25:01

SPEAKER_01

So the main way is to have a very generic tool. Let me quickly open up the... Yeah. Okay. So all we added for the whole self-diagnostics is simply a very basic tool. And the description is fairly straightforward. So it's a report tool and then we are asking it to send a short report to your creator.

25:35

So it frames the idea of writing notes to its creator in a way. So if you frame it around the agent giving feedback to its creators, it works really well.

25:53

And then you can play around with scenarios you want it to report issues about.

26:02

And then that's mostly it. And then in the system prompt, we do need to encourage it a bit. So if you don't type in the system prompt, the times that it fires are fairly minimal, [SPEAKER_01] which is desirable in certain cases, especially if you're at a very large scale. [SPEAKER_01] But in our case, I simply asked it to see if before giving the final answer, use the report tool to surface anything notable for your creators. [SPEAKER_01] So that's all we did. [SPEAKER_01] Okay. [SPEAKER_01] So any questions so far? [SPEAKER_01] So a couple of key things here is that agents, the models are generally trained to look very polished.

26:05

SPEAKER_01

So they are less willing to admit fault in many cases. So encouraging framing it as the models are giving feedback to its own creators is good to get this working. So if you make the tool naming, it also matters quite a bit. So you wanted to frame it as report instead of unsafe bash tool use or something like that, then it doesn't want to. So yeah, that's basically it. Have you looked at adding skills to suppress or to encourage self-discrimination or. So you can, but I think it's probably better if you want to actually catch real unsafe uses. I think a proper classifier would be useful.

26:05

SPEAKER_01

But these, I think, self-diagnostics works really well for catching capability gaps and such. Then the model is okay, it's fine. [SPEAKER_01] So I think the main issue with this is that it's only hesitant when it feels like it's going to get in trouble. [SPEAKER_01] So besides that, that's more or less fine. But for most cases, it'll just work out of the box.

26:05

SPEAKER_03

[SPEAKER_01] Maybe we should go back to question time or what do you think? [SPEAKER_01] Yeah. [SPEAKER_01] I mean, let's leave maybe a few more minutes for a few more questions and then, I think after that we'll be done.

26:05

SPEAKER_01

Any questions in the audience? Do you want us through a case study? Yeah.

26:05

[SPEAKER_01] What specifically would be helpful? [SPEAKER_01] What specific part are you looking to? [SPEAKER_03] For example, how reputable uses it. Yeah. [SPEAKER_01] I can't talk about any specific customer, but what a lot of people use it for. [SPEAKER_01] So I think it's interesting, right? [SPEAKER_01] So a lot of people have their eval setups elsewhere, for example. [SPEAKER_01] But the way that folks generally use us is that they use it for production monitoring. So they send us, you can find our docs at raindrop.ai slash docs.

26:05

[SPEAKER_01] Basically they send us all of the transcripts, any tool use, et cetera, the entire trajectory through OTEL or any other way of integrating. [SPEAKER_01] And once they do that, they have a set of data. [SPEAKER_01] They set up signals in raindrop to look for things that they care about. [SPEAKER_06] And so what people care about is very different, right? [SPEAKER_06] What a coding agent would care about and what a companion would care about or an app for lawyers. [SPEAKER_01] What they would all care about is very different. [SPEAKER_01] So there's a different set of signals.

26:07

[SPEAKER_01] One thing you can do that I haven't really talked about within raindrop is set up a new signal that didn't exist before. [SPEAKER_01] And so we have this thing called deep search. [SPEAKER_01] And so you can use natural language and you can say something like, hey, find me everything within the product or find me all of the times where the agent made XYZ issue. [SPEAKER_01] Right. And so they create a new signal based on that. [SPEAKER_01] And you can basically create a cheap binary classifier and easily deploy it based on that. [SPEAKER_01] And then they have their set of classifier signals that they really care about.

26:08

[SPEAKER_01] Then they use that to drive the feedback loop. [SPEAKER_03] And the feedback loop is improve prompting, improve models, change something with the agent harness, et cetera. [SPEAKER_01] And then actually see, does that improve?

26:19

[SPEAKER_02] Is there less user frustration in production now? [SPEAKER_02] Is there less of this weird little edge case issue that I had before? [SPEAKER_02] That's one whole set of things. [SPEAKER_06] Another thing that a lot of people use us for.

26:53

[SPEAKER_02] So I talked a bit about the agent, but you can use these signals to also look for what are people using my agent for? [SPEAKER_02] What are the user intents?

27:29

[SPEAKER_02] What are the use cases? [SPEAKER_03] What are the use cases? [SPEAKER_02] And you can do a sort of cluster analysis of that. [SPEAKER_02] Okay. [SPEAKER_02] A lot of people are using it to build React related apps. [SPEAKER_02] A lot of people are using it for Python. [SPEAKER_02] Some people are using it to debug this very complicated asynchronous system they already have. [SPEAKER_02] Other people are using it to build something from scratch, code something from scratch. [SPEAKER_02] And then you can see in Raindrop that I think is really interesting, that for each of these different user intents or use cases, you can get a sense of what is the issue rate?

28:14

[SPEAKER_02] What is the user frustration rate in production? [SPEAKER_02] And then a lot of, beyond just having this flywheel, a lot of people have alerting. [SPEAKER_02] A lot of people are using it to build React related apps. [SPEAKER_02] A lot of people are using it for Python. [SPEAKER_02] Some people are using it to debug this very complicated system they already have. [SPEAKER_02] Other people are using it to build something from scratch, vibe, code something from scratch.

28:40

[SPEAKER_02] And then you can see, one thing you can see in Raindrop that I think is really interesting is that for each of these different user intents or use cases, you can get a sense of what is the issue rate? [SPEAKER_02] What is the user frustration rate in production? [SPEAKER_02] And then beyond just having this flywheel, a lot of people have alerting. [SPEAKER_02] And so every day they get a breakdown of what are the issues that are happening today in your product? [SPEAKER_02] You can think of it as almost a bit like Sentry in that sense. [SPEAKER_02] What is the issues happening in my product today? [SPEAKER_02] What is the delta between today and yesterday?

29:07

SPEAKER_01

[SPEAKER_02] Is that true for just specific tools or specific prompting? [SPEAKER_02] What's causing that? [SPEAKER_02] So that's the end to end use case of people use it for if that makes sense. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah.

29:36

SPEAKER_01

[SPEAKER_02] So I think we are entering the error where people are doing observability on agents. [SPEAKER_02] Yeah. [SPEAKER_02] Is it actually, I would say one day further or one step further?

30:03

SPEAKER_01

[SPEAKER_02] What do you see as the main driver for people to be like, oh, normal observability not sufficient? [SPEAKER_02] Yeah.

30:27

SPEAKER_01

[SPEAKER_02] I think it's really just, and I'd be curious what you think about this. [SPEAKER_02] I think it's really just agents are crazier than ever before, right? [SPEAKER_02] More tools, more context, way more intelligent, more real decisions that they can make. [SPEAKER_02] And they're just being used by way, way larger groups of people. [SPEAKER_02] And so when you have this massive amount of data in production, it just makes having good monitoring and observability more important than before.

30:54

SPEAKER_01

[SPEAKER_02] And it makes good monitoring and observability, in my opinion, more important than just testing or evaluations. [SPEAKER_02] Even if you have some online evals, IMO, you need to have really, really good end to end monitoring of the entire system. [SPEAKER_02] Curious if you have any thoughts there as well.

31:10

SPEAKER_01

[SPEAKER_02] So, I think another major issue is the unknown issues are even more important. [SPEAKER_02] So, I think having a generic user frustration class for it is actually really powerful. [SPEAKER_02] Yeah. [SPEAKER_02] Say for example, we also have this another feature called issues, which basically is an agent that mines for newly occurring issues, right? [SPEAKER_02] Say for example, Sentry has, similar to Sentry in a way, where there's a new exception which is occurring. [SPEAKER_02] So it alerts you on that. [SPEAKER_02] So, say for example, you are a coding agent provider and then certain providers are failing all of a sudden.

31:42

SPEAKER_01

[SPEAKER_02] And then you can actually figure out, okay, this subtle spike in user frustration, [SPEAKER_02] and similar to how a human operator would, it can start digging into, are there any patterns for the spike in user frustration? [SPEAKER_02] And then it could figure out that okay. [SPEAKER_02] So, people who are dealing with a specific post-resprovider start to face issues. [SPEAKER_02] We have actually seen this happen live and for a couple of our customers where they had a date was failing, [SPEAKER_02] and then we had an automatic issue being created for them.

32:08

SPEAKER_01

[SPEAKER_02] Yeah, basically once you have that good set of signals, a good user frustration classifier, [SPEAKER_02] as Danny said, you can basically do clustering on it to find what are the root causes. [SPEAKER_02] Yeah. [SPEAKER_02] How do you integrate with light allowing or the light? [SPEAKER_02] Do you want to talk about integrations? [SPEAKER_02] I think we might have a very basic SDK for it. [SPEAKER_06] Our Python side of support is fairly weak right now. [SPEAKER_03] But, we have fairly good SDK support built in.

32:29

SPEAKER_03

[SPEAKER_06] The SDK even has self-diagnostics built into it.

32:31

SPEAKER_01

[SPEAKER_06] So, we inject the tool for you, so that you wouldn't have to do anything. [SPEAKER_02] But it is going to get better. [SPEAKER_02] Yeah. [SPEAKER_02] Yeah. [SPEAKER_02] We actually released 10 different SDKs in the past month. [SPEAKER_02] So, we have a person working on SDKs actively. [SPEAKER_02] So it's going to improve, yeah. [SPEAKER_02] I have a question about experiments, running experiments. [SPEAKER_02] I'm curious how your platform helps me with this issue. [SPEAKER_02] I have a team of around 10 people building my AI agents platform.

33:20

SPEAKER_06

[SPEAKER_01] And we constantly change things. [SPEAKER_01] All the time.

33:32

SPEAKER_01

And we have a lot of feature flags and some of them are experiments. And the rate of the change is so big. Everyday everything changes. I just cannot compare the traces, the sessions of users. Because I don't have enough time to do it. Yeah. I need to have a base system. And then run a few days with parts of the users. With one feature flag enabled.

34:17

SPEAKER_03

[SPEAKER_02] So I can actually compare the data.

34:24

SPEAKER_01

[SPEAKER_02] And get some insights out of it.

34:27

SPEAKER_02

And I just don't have enough time to do it. [SPEAKER_01] Yeah.

34:36

SPEAKER_02

[SPEAKER_01] So how are you doing that right now?

34:38

SPEAKER_06

[SPEAKER_01] Are you just kind of.

34:40

SPEAKER_02

[SPEAKER_01] It's Wild West, right? [SPEAKER_01] Yeah. [SPEAKER_01] That's why I said I'm just using mostly Claude.

34:51

SPEAKER_03

[SPEAKER_01] Because I just give it everything and ask it questions.

34:51

SPEAKER_02

[SPEAKER_01] And try to figure out the insights. [SPEAKER_01] But it's not really. [SPEAKER_01] Yeah. [SPEAKER_01] I get what you're saying. [SPEAKER_04] So, a few things there. [SPEAKER_04] The first is you can use experiments. [SPEAKER_04] If you want to keep that shipping speed. [SPEAKER_04] You don't have to run long multi-day experiments. [SPEAKER_04] You can ship something. [SPEAKER_04] And if you have a sufficient sample size. [SPEAKER_04] You could see pretty quickly if there's any regressions or not. [SPEAKER_04] You just maybe it's 1% different or 2%. [SPEAKER_01] And try to figure out the insights. [SPEAKER_01] But it's not really. [SPEAKER_01] Yeah.

35:55

SPEAKER_02

[SPEAKER_01] I get what you're saying. [SPEAKER_04] So, a few things there. [SPEAKER_04] The first is you can use experiments. [SPEAKER_04] If you want to keep that shipping speed. [SPEAKER_04] You don't have to run long multi-day experiments. [SPEAKER_04] You can ship something. [SPEAKER_04] And if you have a sufficient sample size. [SPEAKER_04] You could see pretty quickly if there's any regressions or not. [SPEAKER_04] You just maybe it's 1% different or 2%. [SPEAKER_04] That's enough for you to be okay, it's fine. [SPEAKER_04] It's not breaking anything. [SPEAKER_04] Drastic. [SPEAKER_04] The other thing is what Danny was talking about.

37:02

SPEAKER_02

[SPEAKER_04] We have an agent which is basically, you could think of. It's basically exposing all of these signals to Claw to make decisions on if things are better or not. [SPEAKER_04] And we're thinking also about how we close this loop. [SPEAKER_04] Maybe you have a really good set of signals. [SPEAKER_04] And then you have essentially an agent that can look at all these signals. [SPEAKER_04] And then it can find issues based on that. [SPEAKER_04] What's changing, et cetera. And then I can create a PR based on that. And then I can see how you run some new experiments based on these new PRs. And then there's this can become this infinitely self improving loop.

37:51

SPEAKER_02

Which is very interesting. [SPEAKER_04] But that's one thing that I think about. I don't know if you have any additional thoughts. [SPEAKER_04] It's slow because you need to deploy it to production and wait for some data to get it. [SPEAKER_04] Yeah. [SPEAKER_04] Yeah. [SPEAKER_04] Depends on how much data you have. But that's, yeah. It really depends on how big these sample groups are as well, et cetera. Sometimes it's like a few minutes you can tell, but sometimes they want to wait for longer. Yeah. Does your platform help with enabling experiments on sessions?

39:01

SPEAKER_02

So you can maybe automatically enable some experiments so you can take care of the logic that every session has only one experiment enabled.

39:06

SPEAKER_06

[SPEAKER_02] So we can easily compare it to the base or something like that.

39:10

SPEAKER_03

[SPEAKER_02] We have, we're working on stuff like that actively, but yeah.

39:15

SPEAKER_06

[SPEAKER_02] Yeah. [SPEAKER_02] Thanks.

39:23

SPEAKER_02

I'm wondering whether we are storing the original traces and then, like, come in and implement my new signal source, callback, or whatever, if you can fail. Yeah. Signal for the original. So I can do some post-portal analysis. Yeah. We do, so you can, in this, all your historical data. And then when you create a signal, we actually run a quick backfill of the past couple of days. Yeah. So, yes.

39:55

SPEAKER_01

[SPEAKER_02] So that's definitely supported. [SPEAKER_02] Thank you. [SPEAKER_02] Is there a free time to try it out? [SPEAKER_02] We do have a free trial. [SPEAKER_02] We're going to try to make it, it's right now, it's two weeks. [SPEAKER_02] Probably going to make that longer soon. [SPEAKER_04] But if you just DM me, I can, if you, maybe I can have my, should I, do you want to open this so I can just have our things? [SPEAKER_02] Well, yeah, but if you just, so we are hiring. [SPEAKER_02] That is a thing that we're very excited about. [SPEAKER_02] Trying to massively increase the size of the team.

40:25

SPEAKER_01

[SPEAKER_02] And if you message me at either Twitter or you can email me as well, that's something I can just set you up with a longer free trial. [SPEAKER_02] Yeah. [SPEAKER_02] And internally, so we use Opeak and Sentry and Oilers.

40:32

SPEAKER_02

[SPEAKER_04] How do you guys, I would imagine you guys also use those tools and how does Raindrop, from how I'm understanding it, it's, you're creating signal with your own models, whatever you're using, so that you make our lives easier to identify signals and harmful intents and user behaviors. [SPEAKER_04] But how would you maybe have an example of how do you have that full stack of how it works with the Sentry and the Nobeak and and maybe where there are overlaps where you go directly for those compositors?

40:39

[SPEAKER_04] So, if you're sending all the telemetry data, we can find any exceptions in the traces, tool errors, et cetera, and that's a thing that you can also track within Raindrop and that is an explicit signal. [SPEAKER_04] So there's implicit and explicit signal, so that becomes an explicit signal. [SPEAKER_02] Do you want to? [SPEAKER_02] So I think most of the observability platforms will give you the agent trace, the token usage, if the tool call failed or not. But I think where we sort of shine is the fuzzy part, the fuzzy failures, right?

40:42

SPEAKER_01

[SPEAKER_04] Where the user is frustrated, which I think matters more than the explicit signal that you sort of get from Sentry. [SPEAKER_05] I mean, obviously those are also important, but we focus a bit more on the fuzzier side of the failure space. [SPEAKER_05] But at the same time, we also have a trace view. I think we also have a very interesting feature called trajectories, which sort of visualizes, [SPEAKER_05] if you want to find a trace, which has like three different tool call failures. [SPEAKER_05] So you can actually, okay. [SPEAKER_05] Well, it's often real, but. Okay. Let me just get in. So you can sort of describe the type of trees that you want to look at.

41:03

SPEAKER_04

[SPEAKER_01] So let's see. [SPEAKER_01] I hope you have data, but you can more or less describe the type of trajectories that you want to see, instead of just configuring it. [SPEAKER_01] So we do both in a way. [SPEAKER_01] So you can obviously set up tools are failing, sort of alert as well. [SPEAKER_01] So you can just search for any trajectory. [SPEAKER_02] So yeah, you can see that. [SPEAKER_02] This is sort of how the tools are being called in what order you can see which ones have errors. [SPEAKER_02] You click into them. [SPEAKER_02] You can see the input and the output to the specific tool, like what actually screwed up here.

41:22

SPEAKER_04

[SPEAKER_02] And you can see, okay, it's interesting that this has like this, no one lets you, this is pretty much the only place where you can visualize tools like this. [SPEAKER_02] But you can see here, you can get a shape and understanding of the topology of what's going on here. So you can obviously set up tools failing alerts as well. So you can just search for any trajectory. [SPEAKER_02] Yeah, you can see that. This is how the tools are being called in what order you can see which ones have errors. [SPEAKER_02] You click into them. You can see the input and the output to the specific tool, what actually screwed up here.

41:34

SPEAKER_04

[SPEAKER_02] And you can see, it's interesting that this has, no one lets you, this is pretty much the only place where you can visualize tools like this. [SPEAKER_02] But you can see here, you can get a shape and understanding of the topology of what's going on here. And you can see when there's other ones that look similar, you can see, okay, this kind of looks similar to this. And then that gives you a sense. You can do search on this. [SPEAKER_05] Again, we have an agent that can look through these and give you a sense of what's going wrong. And so it just makes it really easy to find issues in agents. [SPEAKER_05] Yeah. [SPEAKER_02] Cool. Anything else?

41:57

SPEAKER_04

[SPEAKER_01] No. Any other questions? Can you export the data that you can show? The directory's data?

41:58

SPEAKER_02

[SPEAKER_01] Yeah. What would you wanna export like just the raw trace logs or what do you wanna? [SPEAKER_01] So I think we, usually what our customers do is that they already have a hotel stream, right? [SPEAKER_01] Yeah. [SPEAKER_03] We just end up being the target I guess.

42:12

SPEAKER_04

[SPEAKER_02] Yeah. But at the same time, they do want us to export signals that we label.

42:14

SPEAKER_02

[SPEAKER_01] So we do support BigQuery and Snowflake. So we do export the event and then the signals that were classified for that event.

42:16

SPEAKER_04

[SPEAKER_01] Okay. Last questions? Can we look at the signals? [SPEAKER_03] All right. [SPEAKER_02] Let's do a longer timeframe. Let's just go over the last month. So you can see stuff like refusals. And then again, if you click into any of these, you can get a sense of over time, your task failure, jailbreaking, what specifically is going on. And then you have your self diagnostics ones as well. Capability gap, et cetera.

42:24

SPEAKER_04

[SPEAKER_02] Cool. And do you have open data on number of traces that you guys know because it seems that with your clients, the tool is extremely valuable when you have agents adapted at a very big scale. And so I can imagine that you have a lot of data. Do you have some data?

42:26

SPEAKER_02

Yeah, it starts being, is your question like, what's the smallest where it's useful or what's the, Oh, that volume of all the data that you're receiving, processing, and generating signal around, how many jailbreaks do you see across all the clients? Oh, do we have any sort of, we should do something like that. That would actually be very interesting. [SPEAKER_01] We don't have anything like that. I think we have mixed opinions about that. I think it does it in a way, right? [SPEAKER_01] Yeah. [SPEAKER_03] People generally have negative reaction to it.

42:56

SPEAKER_02

[SPEAKER_01] Maybe it's different here, but at the same time, do our customers want us to do that? It's a different question as well, right? [SPEAKER_01] So, but yeah, we would love to, but I think that for compliance reasons we can't actually put a customer's data out there. [SPEAKER_01] Cool. Anything else? [SPEAKER_05] All right. Thank you everyone. Thanks. I'm wondering whether we are storing the original traces and then, like, come in and implement my new signal source, call back, or whatever, if you can like fail. Yeah. Signal for the original. So I can do some like post-portal analysis. Yeah. We do, so you can like, in this, all your historical data.

43:24

SPEAKER_02

And then when you create a signal, we actually sort of like run like a quick backfill of the past couple of days. Yeah. So, uh, yes. So that's definitely supported. Thank you. Is there a free time to try it out? Uh, we do have a free trial. We're going to try to make, it's right now, it's two weeks. Uh, probably going to make that longer soon.

43:49

SPEAKER_04

But if you just DM me, I can, if you, maybe I can have my, should I, oh, do you want to open this so I can just have our things?

43:58

SPEAKER_02

Well, yeah, but if you just, so we are hiring. That's a, that is a, a thing, uh, that we're very excited about. Like trying to massively increase the size of the team. Um, and if you message me at either Twitter or you can email me as well, like that's something I can just set you up with a longer free trial. Yeah. And, um, internally, so we use Opeak and Sentry and Oilers.

44:25

SPEAKER_04

Um, how do you guys, like I would imagine you guys also use those tools and how does Raindrop, it's, from how I'm understanding it, it's, you're creating signal with your own models, whatever you're using, so that you make our lives easier to identify signals and, uh, like, harmful intents and user behaviors. Uh, but, uh, how would you maybe have an example of how do you have that full stack of how it works with the Sentry and the Nobeak and, and, and like maybe where there are overlaps where you go directly for those compositors?

45:02

SPEAKER_04

So, if you're sending all the telemetry data, we can find any exceptions in the traces, tool errors, et cetera, and that's a thing that you can also track within Raindrop and that's, that is a explicit signal. So, uh, there's like implicit and explicit signal, so that becomes an explicit signal.

45:17

Um, do you, do you want to? So I think most of the observability platforms will give you like the agent trace, the token usage, uh, if the tool call failed or not. Uh, but I think where we sort of like shine is sort of like the fuzzy part, the fuzzy failures, right? Where the user like frustrated, which I think matters more than, uh, the explicit signal that you sort of get from Sentry.

45:44

SPEAKER_05

I mean, obviously those are also important, but we focus a bit more on the fuzzier side of the failure space. Uh, but at the same time, we also have a trace view.

45:55

SPEAKER_01

I think we also have a very interesting feature called like trajectories, uh, which sort of like visualizes,

45:59

SPEAKER_05

is if you want to find like a trace, uh, which has like three different tool call failures. Uh, so you can actually, uh, okay. Well, it's often real, but.

46:10

SPEAKER_01

Okay. Let me just get in.

46:15

So you can sort of like describe the, uh, type of trees that you want to look at. Uh, so let's see. I hope you have like data, but, uh, you can more or less like describe, uh, the type of trajectories that you want to see, uh, instead of just like configuring it. So we do both in a way. So you can obviously set up, uh, tools are failing, sort of like alert as well. So you can just search like for any trajectory. Um, so yeah, you can see that. This is sort of how the tools are being called in what order you can see which ones have, have errors. You click into them. You can see the input and the output to the specific tool, like what actually screwed up here.

47:06

SPEAKER_02

Um, and you can see, okay, it's interesting that this has like this, the, no one lets you, this is pretty much the only place where you can visualize tools like this. Um, but you can see here, like you can get a shape and understanding of the topology of what's going on here. And you can see when there's other ones that look similar, you can sort of see, okay, this kind of looks similar to this. And then that gives you a sense. You can do like search on this.

47:30

SPEAKER_05

Again, we have an agent that can look through these and give you a sense of what's going wrong. And so it just makes it really easy to find, uh, issues in, in agents. Um, yeah.

47:42

SPEAKER_02

Cool. Anything else?

47:44

SPEAKER_01

No. Any other questions? Can you export the data that you can show? Uh, the directory's data? Yeah. Uh, what would you, you wanna export like just the raw trace logs or what, what do you wanna? Uh, so I think we, usually what our customers do is that they already have like a hotel stream, right? Yeah.

48:04

SPEAKER_03

We just end up being like kind of the, uh, target I guess.

48:08

SPEAKER_02

Yeah. But at the same time, uh, they do want us to like export signals that we label.

48:13

SPEAKER_01

So we do support like, uh, BigQuery and, uh, Snowflake. So we do export the event and then the signals that were classified for that event. Okay. Last questions? Can we look at the signals?

48:26

SPEAKER_03

All right.

48:27

SPEAKER_02

Uh, let's do it a longer timeframe. Let's just go over the last month.

48:40

SPEAKER_02

So you can see stuff like refusals. And then again, if you click into any of these, you can get a sense of over time, your task failure, jailbreaking, like what specifically is going on. And then you have your self diagnostics ones as well. Um, capability gap, et cetera. Cool. And do you have open data on like number of traces that you guys know because it seems that with your clients, uh, the tool is extremely valuable when you have those, um, agents adapted at a very big scale.

49:13

SPEAKER_03

And so I can imagine that you have a lot of data.

49:14

SPEAKER_02

Do you have some data? Yeah, it starts being, is your question is like, what's the smallest where it's useful or what's the, Oh, that volume of, uh, all the data that you're receiving, processing, and generating signal around like, um, how many jailbreaks do you see across all the clients? Oh, do we have any sort of like, we should, we should do something like that. That would actually, that would be very interesting.

49:37

SPEAKER_01

We don't have anything like that. I think we have like mixed opinions about that. I think it does it in a way, right? Yeah.

49:43

SPEAKER_03

People generally have negative reaction to it.

49:45

SPEAKER_01

Uh, maybe it's different here, but at the same time, do our customers want us to do that? It's like a different question as well, right? So, uh, but yeah, we would love to, but, uh, I think that like compliance reasons where we can't actually put a customer's data out there. Cool. Anything else?

50:05

SPEAKER_05

All right. Thank you everyone.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note