Hello, hello. Hey everybody, welcome to this talk, Always-On Agents Run Production. Without the on-call text, my name is Justin Smith, one of the founding product engineers at Resolve AI. I've been in the space for about 15-plus years in the monitoring, observability, how-do-you-operate-production-systems space. I was at Splunk for a while, was one of the architects on the observability suite there, spent a good tenure at VMware, and really enjoy product design and front-end architecture. How do people experience a product or a use case or something like that? That's the stuff I like to dabble in.
But I want to talk a little bit about the first wave of AI, and it's been a fun one. I think the first big wave, and I'm sure we've all experienced this, is just how we build software. But there's some net effects of that. It's a lot of bigger PRs that are coming through. We definitely see a lot of this more frequently. So people are shipping code at a much faster rate. And we're beginning to see maybe even non-developers that don't actually know the code or what it's doing, or the operating principles behind it. But we're getting developer productivity. And that's good, right? That's a good thing, that we're all able to produce more and faster.
What we actually found out, and this was a survey study done, is that 70% of the time for an engineer is not focused just on writing code. It's actually spent on running the code that is shipped into production: maintaining all the platforms, scaling the infrastructure, debugging all the incidents and being on call, shipping hot fixes, dealing with alerts, updating all the runbooks and operating procedures, restoring services, dealing with escalations, dealing with questions from other teams, and things like that.
So really, coding was never the big bottleneck, right? A lot of it was really around, thank you, Granola, how do we actually run these things in production? And that's getting harder and harder and harder. AI is creating a lot more issues in production as AI code goes through. It's not clear we have the right structures in place to deal with the amount of changes that are coming through. Unlimited tokens is coming to an end. The token max, right, they're starting to clamp down. Prices are going up. Companies are getting a lot more stringent on what's being used for AI. We need full-stack AI. It's not just about the models anymore. It's about the context around the models and what the models can do inside of a specific domain. These become the problem areas that we need to focus and tackle. And this is true today. So it's creating more complexity inside of our environment.
The reality is that systems have always been complex. That's why we have these big tools that can try to give us insights into these systems. There are multiple teams. There are multiple systems that are all having to work together. And they all have their own goals that they're trying to deliver toward. But you have organizational goals, and how do you keep all of this in balance, right? How do you pull all of this stuff together in a way that actually helps you and facilitates your organization?
And the answer is, you've got to use AI inside of production to deal with the amount of increased complexity that AI is putting into your product or into your system. And so that's where Resolve, this was our hypothesis from the beginning, was we're going to see an influx in issues coming out of coding, just the increase in coding velocity. There's going to be more need for AI to actually operate and run these systems. We're lucky to work with some world-class engineering teams that are solving really difficult problems at crazy scale. And that gives us insight into how bigger organizations are having to deal with the influx of AI, et cetera.
Resolve itself hosts a bunch of different capabilities. We have a number of agents that you get to experience. One of them is just an on-call agent, and this is where we started, right? So for every alert that comes in, we can do a triage of that alert. We can do a full root cause investigation of that alert. And this is for anybody that's had to be on call before. On call is a nightmare, right? You're often only going on call every few weeks. You maybe don't fully understand all the changes that have come in. You don't fully understand all the different systems that you're having to interact with. And so the complexity is already there. And having an AI agent that's able to come support you and pull that context together is incredibly valuable.
And so that can often grow from just getting a single page into a much larger incident across many different teams across an organization, and we have agents there to support the much larger activity of all of this cross-collaboration, et cetera, keeping everybody in sync and aligned on where the incident is happening, what the impact of that is, et cetera.
And then we also focus a lot on background agents, and this is covering the long tail of what happens when there's not a fire brewing or going on at any one point. There's still lots of operational work that you as an engineer or an engineering team have to do, and a lot of ceremonies of passing context off or dealing with one-off issues or having to scratch that itch in the back of your head of, is that part of the system okay or not okay? And you're constantly having to bounce across all these different things.
Underneath all that, we have an agent architecture that deals with models and context and reasoning and actions. Learning is, I'll half pause on that one. I think some of the biggest issues that we've seen, it's not that a model by itself is not smart or whatever. Models have gotten incredibly capable over the last year, let's say, especially over the last six months or so. But the idea of understanding, truly understanding, your environment and the way that your services interact and where the hotspots are, keeping track of all of that understanding is incredibly difficult. But it's incredibly important for any model to be successful at the task that it needs to do. It has to have an underlying learning system to be able to capture that knowledge and that understanding of how your system operates.
So we spend a lot of time thinking about how do we have systems that not just can understand your environment at any one point, but grow as your system evolves because, again, your system is evolving faster and faster. We need to keep up with learning about what's the current state, what's the current causal change that we need to be keeping an eye on. And then, of course, all the enterprise stuff underneath.
And so this is that same view packed out. So today, we do a lot at Resolve, the on-call and the incident stuff. I'm going to focus a lot more on the background agent stuff. So how do we deal with the things that maybe aren't immediate fires? If you have questions about the immediate fire stuff, we have a booth down in the expo. Please come check it out. Our team would love to demo to you, et cetera. But today we're going to focus on the background agent.
So, a little pop quiz. Feel free to raise your hands. Is anybody using agents as part of your daily workflow? Maybe outside of coding. I'm assuming everybody's doing coding agents these days. Is anybody actually running agents that are helping in other ways? Okay. Oh, decent. Any good examples? Any fun stuff that anybody has? You can just yell it out. Meeting reviews. Meeting reviews? Yeah, meeting reviews. I just had my Granola show up. Market research. Market research. I do a lot of that. Let's have a good conversation about it. Yeah, yeah, I do that all the time. Any other ones? Maybe one more. What's the one? Therapy. I can't hear it. Therapy. Therapy.
That's a fantastic one, actually. We are humans here today. This is very important. That's actually a very good one. Okay. So people are having some stuff going on. And this recaps a little bit again. Maybe outside of the coding. I'm assuming everybody's doing coding agents these days. Is anybody doing actually running agents that are helping in other ways? Okay. Oh, decent. Any good examples? Any fun stuff that anybody has? You can just yell it out. Meeting reviews. Meeting reviews? Yeah, meeting reviews. I just had my granola show up. Market research. Market research. I do a lot of.
Let's have a good conversation about it. Yeah, yeah, I do that all the time. Any other ones? Maybe one more. What's the one? Therapy. I can't hear it. Therapy. Therapy. That's a fantastic one, actually. We are humans here today. This is very important. That's actually a very good one.
Okay. So people are having some stuff going on.
So this recaps a little bit again. A lot of production work is not about, there's not a big ceremony that everyone is focused on for the type of work that we have to do. On call, you've got a page that goes off. You know somebody's going to receive that. Incidents, you create a bridge. You invite people in. That's great. But there's just a long tail of other things that we are accountable for that doesn't have a thing that's going to show up in your job description of, this is what you're going to be responsible for. Watching deploys that go out and make sure that they're actually getting out healthy.
A morning report or incident digest of just, what's the state of my system today so that we're all on the same page? Hey, that P99 drift came back. Is somebody looking at that or not? And this is pulling people in to try to figure out what's going on. This may not be paging, right? Because we're not going to alert on everything. Produce the capacity report, right? Are we tracking okay, right? This is maybe a company goal this quarter. Are we tracking against that? Somebody's going to have to be responsible for doing that. The recurring health check and just checking and making sure things are running okay and not waiting for a customer to complain.
So this work doesn't have an obvious, oh, this now needs to go be done. But it's work that we end up having to do. So what is the task? Task is just execution and the context to understand how to actually execute the task. Execution is very, very important. It's understanding what to do and being able to execute that. Maybe having access to the tools, et cetera, right? Obviously very important to do. But we think the production context is just way more important because it's one thing to go check a dashboard. It's another thing to say that metric smells off. And the execution is, can load the dashboard. It's the production context that's going to say, this feels wrong.
And I don't know if I can even explain why it feels wrong. It just feels wrong. And I want to dig into the next layer of understanding of that. And so really, if we start talking about background agents and being able to perform tasks, you need both of these. You need the execution engine. That's great. But you really need that production context that tells you, is this important or not important?
So every background agent, there are a few different principles that we like to think about with our background agents. When does it work? How does it work? How does it know what to go do? When does the agent work? It can work in a bunch of different ways. It can just do it on a schedule. Maybe this is the morning report, et cetera. Just do some summarization for me on an ongoing basis. Maybe it's a weekly event, right? We do an on-call handover every Thursday. And so a lot of the work that our agent does is prepare. What are the interesting trends from the last week that the next on-caller needs to understand as they pick up the rotation? Event streams.
There's lots of systems that will push events as key things happen. So deployments go through ICD pipeline. There are other Slack-based, right? We get a lot of Slack things, messages coming through, et cetera. And these are things that we can pick up and trigger and say, oh, if this event happens, let me understand what that event is and go do some work. And then message-based. So I can just tell it, hey, go do some work. And it will go do some work. That's fantastic. How does it run? It always runs. It's in the cloud. So if you close your laptop, it's okay. It runs inside of a sandbox. So it has a file system underneath it.
This allows it to self-organize a lot of its work, et cetera, as it's doing things. And then obviously back to the learning loop, right? So that idea of knowledge and a memory system underneath that to really understand your systems. And as it's doing a task, able to reflect on that task and do a better job next time. Or the things that it learned from one task, it can apply into a different task. Because again, this shared knowledge system works across all the different tasks that we have. So how does the agent know what to do? It has a task system. It can pull in all the skills that you have in other systems. That's fine. You can connect those.
And it's got, obviously, the integrations that it's going to plug into. So let's talk a little bit about what types of things you can hand over. And we've got four workloads that we're going to talk about. But if you think about the previous couple slides, these are very basic primitives that we've built into the system. You can get very creative. We have a number of people inside of Resolve that have gotten very creative with the type of background activities that they have. So I want you to use these as, these are things we've seen be very successful inside of Resolve, but also with a number of our customers. But sky's the limit, and you can get really creative.
So deployment monitoring. This is a big one. Any change inside of your environment is an opportunity for something to go wonky. And so having an agent that's able to watch as all these change events come in, just to do a sanity check of is everything stable, is incredibly, incredibly important. And a lot of people have a decent CI CD system. This is tried and true stuff that we've had as an industry for quite a while. But we notice a couple of gaps from most of our customers. Typically the checks that it does are good. They're good baselines, but it's not exhaustive based on the type of changes that are going in, et cetera.
There are certain signals you'd want us to watch or not want to watch. And so every rollout is a bit unique. Oftentimes you have change systems that you're not piping through a CI CD system, like a feature flag or maybe some infra changes that might happen, which maybe don't get any monitoring at all. And you're just trusting that an alert might fire and an on-caller will wake up and say, who changed what, right? And so deployment monitoring is actually a really big use case that we suggest people go through. And I'll show some examples of that in a second. Okay. Scheduled health and anomaly checks. So this is the ongoing periodic checking of some of your systems.
And this is maybe something where it's like, go check my general dashboards on a routine basis. Maybe every morning just do a casual check to make sure there's nothing weird from last night that I might need to be aware of. But this can also just be a time-based thing. Like I made a change in part of our system. I'm worried about this third-party service that I'm interacting with. Let me just set an agent to watch that maybe for the next week just to make sure everything is stable. And then that agent can stop its job. Operational reports and handoffs. I talked a little bit about this.
These are the ceremonial things that we might want to do just to spread information, summarize things, bring things to the foreword. And then a first responder to engineering questions. And this one's kind of fun because the trigger for this is actually just a Slack message. And I will say one of my biggest responsibilities as an engineer is watching all my Slack channels and trying to make sure everyone's happy and that nobody has any burning questions or anything like that. I'm worried about this third-party service that I'm interacting with. Let me just set an agent to watch that maybe for the next week, just to make sure everything is stable.
And then that agent can stop its job.
Operational reports and handoffs. I talked a little bit about this. These are the ceremonial things that we might want to do just to spread information, summarize things, bring things to the foreword. And then a first responder to engineering questions. And this one's fun a little bit because the trigger for this is actually just a Slack message. And I will say one of my biggest responsibilities as an engineer is watching all my Slack channels and trying to make sure everyone's happy. And that nobody has any burning questions or anything like that.
And so I can be heads down trying to build something, and then the eventual Slack notification comes in that this channel, somebody asked this important question. And I just need to jump in there and try to provide context, et cetera. It's not hard work. It's not hard for me to go answer questions, but it's disrupting me. And if I don't go answer it, they won't get an answer for a while. And what we found is our agent actually has access to a lot of information that people ask questions about. At least this is true internally.
And so we actually have an agent that can watch all of these critical channels and determine whether it has enough confidence to answer the question or not. And one of the fun things is our agents have access to Slack DMs and things like that. And so you can have an agent that will DM you to say, I think I know the answer to this, but I'm not sure. Can you confirm this for me before I respond back? So this emergent behavior gets fun and interesting as you build these things out. Okay. So I'm going to flip over and hope all of this works. Cool. So let me see if I can find the one that I wanted to show. So this is our demo application running in our demo Slack environment.
And what I wanted to show off was some of our deployment stuff and talk a little bit more about what's going on under the hood. So this is a fake environment just to showcase some things. So here, any time somebody posts a GitHub tag, our agent's going to see that and say, oh, that's a release. Is that a release? Yes, that is a release. Let me go watch that. But it's not just going to watch it. It's going to do something that's slightly more intelligent. Because, like I said before, everybody has a CI CD system. It will do the standard checks on certain KPIs.
But what the agent is able to do is actually look at the changes that are going in, understand what telemetry might help us evaluate whether those changes are good or not good, or are putting the system in an abnormal state.
And build a customized plan that it's going to check just for this specific release. And this is why I go back to our goal is not to sit here and say we're going to replace an entire CI CD pipeline. You've spent time organizing that. But this can patch a lot of parts of your system that may not be as robust as they should be. And it would be great if you had a single engineer just focused on watching all the things on every release. But that's really expensive. There's a lot of cognitive load. You'd rather have them doing other things.
So now the agent can come and actually do a lot of that dynamic understanding of this is the change, so I'm going to check for these things. And so here, the checkout replaces currency service. We'll monitor the checkout latency and the error rates. Well, let's take a look at the Kafka pipeline because that involves this causal chain. I want to make sure it is healthy. And it'll check that. And it can check it not just once but on an ongoing basis. And none of this is hard coded in. It's not like, oh, let's just wait for 15 minutes and then try this again and then we'll be done.
Again, the agent has a bit more autonomy, and you get to guide it a bit on how much autonomy you want it to have. But it could decide I want to wait for another hour because this type of issue might only hit every so often. So I really want to spend a little bit more time focused on this. Maybe I'll come back in three days and say, is this deploy still healthy? Are we seeing the change and the effect that I expected to see out of this? So this is the type of thing that we can bring. And this, again, works for feature flags, for interchanges, any eventing system that you can think of. The on-call handoff reports. Let's scroll up just a little bit.
So this is just a summarization of all of the work that was done over the last day that the agent is summarizing up. And this one's a little bit verbose, but you can see a bunch of different investigation summaries that we did, some notable changes, et cetera, work completed, et cetera, critical open. I guess that's the on-call handoff. But one of the nice things, and I don't know if I'll be able to watch it go all the way, but you can always just come back in this thread and, you know, this is too verbose. Verbose. Make it shorter. And I'm not going to be able, unfortunately, with no time, to watch this actually go, but this works.
The agent is able to update its task underneath and able to give you the answer, like update, so that the next time it fires, it's not going to be as verbose. And I can tell it explicitly what I want, et cetera. I was just giving you an example. This is the more fun one. So here I'm just posting different problems. I'm not having to know that resolve exists. I don't have to at-mention resolve, whatever. I've set up this agent to passively watch this channel. If you see something that you think you have an answer for, that somebody is digging into, go ahead and respond. Otherwise, don't. So here's a message that I posted that it's decided I don't need to respond to this.
So again, a very flexible system that can adapt to a bunch of different things. Uh-oh. Where'd my slides? Oh. We have, in the UI, a bunch of stuff that you can do. You can always go down and inspect all the different tasks that you have and view their reports, view previous runs. You can see all the work that the agent has done underneath to accomplish that task. So you get a lot of visibility into what the agent is doing. But we think the surface area being where you live, right? So Slack, or MS Teams if you're on MS Teams, as this first-party experience to integrate the agent into, is incredibly important. And so, how would you get this stuff set up?
It's really just through talking with the agent. And so here, this is me saying, hey, I want to do a new recurring health summary for my team. So the agent's going to take a look at my environment. It's going to explore my environment a little bit.
And eventually, it'll likely come back and ask me a couple questions about what I want to see, what kind of report I want, how verbose I want it, et cetera.
The agent's going to go ahead and do all that. And it's going to set up that initial thing for me so that I can test it out and make sure it's working. And then share it with the rest of my team. In the interest of time, I don't think we'll get to this. But come by the booth and you can see more. Cool. So that's background agents. And again, I lean back on be creative, right? Everyone has unique work. As a company, we believe that every company is a unique place. That's why we spend so much time on our knowledge system, et cetera. Truly understanding what your environment looks like, what your needs are, et cetera.
That begins to tell you where the biggest benefit from having these agents begin to pick up work would be. The agent's going to go ahead and do all that. And it's going to set up that initial thing for me so that I can test it out and make sure it's working. And then share it with the rest of my team. In the interest of time, I don't think we'll get to this. But come by the booth and you can see more. Cool. So, that's background agents. And again, I lean back on be creative, right? Everyone has unique work. As a company, we believe that every company is a unique place. That's why we spend so much time on our knowledge system, et cetera.
Truly understanding what your environment looks like, what your needs are, et cetera. That begins to tell you where the biggest benefit from having these agents begin to pick up work would be. Also, just to call out, if you have an agent harness, internally if you're building your own, everything that I showed is accessible through MCP servers, et cetera. So, you can really graph resolve into any system that you have as an extension of learning to do deeper work or to pull production context a bit more efficiently or even augmented with all that learning stuff that we've done. And then, obviously, bring your own skills along for the ride.
Don't go duplicate a bunch of stuff. So, really, the biggest things to take away, cost of operational work. It's not navigating, it's not just in the task execution. It's in the environment complexity, right? That's where the biggest issue is going to happen. Background agents, they run on schedules. They run on triggers. They're very composable. You can graph them into lots of different use cases. It's fun to see people explore that. And, yeah, you open resolve just to see what the top findings are, et cetera. But, ideally, a lot of your interaction is in the places that you're already doing work.
So, if you have any questions, you can find me down at the booth or just meet me out in the hall. But thanks for coming. Appreciate it. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. This is kind of the more fun one. So, you know, here I'm just like posting different problems. I'm not having to know that resolve exists. I don't have to like at mention resolve, whatever. I've set up this agent to sort of passively watch this channel. If you see something that you think you have an answer for that somebody is kind of, you know, digging into, go ahead and respond. Otherwise, don't.
So, you know, here's a message that I posted that it's decided I don't need to respond to this. So, again, very kind of flexible system that can kind of adapt to a bunch of different things.
Uh-oh. Where'd my slides? Oh.
We have, in the UI, there's a bunch of stuff that you can do. You know, you can always go down and inspect all the different tasks that you have and view their reports, view previous runs. You can see all the work that the agent has done underneath to sort of accomplish that task. So, you get a lot of visibility into what the agent is doing. But we think the surface area being where you live, right? So, Slack is this like kind of, or MS Teams if you're on MS Teams, as this kind of first party experience to sort of integrate the agent into is incredibly important. And so, you know, how would you get this stuff sort of set up?
It's really just through talking with the agent. And so, here, this is me sort of saying, hey, I want to do a new recurring health summary for my team. So, the agent's going to take a look at my environment. It's going to explore my environment a little bit. And eventually, you'll likely come back and ask me a couple questions about what I want to see, what kind of report do I want, how verbose do I want it, et cetera. The agent's going to go ahead and do all that. And it's going to set up that sort of initial thing for me so that I can test it out and make sure it's working. And then share it with the rest of my team.
In the interest of time, I don't think we'll get to this. But come by the booth and you can see more.
Cool.
So, that's background agents. And again, I sort of lean back on be creative, right? Like everyone has unique work. I mean, as a company, we believe that every company is a unique place. That's why we spend so much time on our knowledge system, et cetera. Truly understanding what your environment looks like, what your needs are, et cetera. That begins to kind of tell you where the biggest benefit from having these agents begin to pick up work would be. Also, just to call out, if you have an agent harness, internally if you're building your own, everything that I showed is accessible through kind of MCP servers, et cetera.
So, you can really graph resolve into kind of any system that you have as just kind of an extension of learning to kind of do deeper work or to sort of pull production context a bit more efficiently or even augmented with all that learning stuff that we've done. And then, obviously, bring your own skills along for the ride. Don't go duplicate a bunch of stuff. So, really, the biggest things to take away, cost of operational work. It's not navigating, you know, it's not just in the task execution. It's in the environment complexity, right? That's where the biggest issue is going to happen. Background agents, they run on schedules. They run on triggers.
They're very composable. You can sort of graph them into lots of different use cases. It's fun to see people explore that. And, yeah, you open resolve just to kind of see what the top findings are, et cetera. But, you know, ideally, a lot of your interaction is kind of in the places that you're already kind of doing work. So, if you have any questions, you can find me down at the booth or, you know, just meet me out in the hall. But thanks for coming. Appreciate it.
Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.