SPEAKER_00
Hi everyone, my name is Gabe DeMesa. I'm an engineer here at OpenGov and today we're going to be talking about agents in production, specifically how OpenGov built and scaled OG Assist. So this presentation is going to be jam-packed with good stuff. We're going to talk about AI agents, we're going to talk about our harness, we're going to talk about evals, observability, traces, we're going to talk about tools and skills. We're going to talk to you guys about what we do at OpenGov and how we operate at the scale that we operate at in production so you'll be able to see a real use case and workload with AI agents. So without further ado, let's get started.
SPEAKER_00
Okay, agenda. So just really quickly going to go through high level what we're going to talk about today. I'm going to tell you guys a little bit about OG Assist and what OpenGov is. I'm going to tell you guys the origin story of how this all came to be. We're going to talk about OG Assist's big bet on Effect. A little bit into our core agent loop, we're going to talk about the A2A protocol, evals and sandboxing. We're going to talk about how we manage long context. We're going to talk about monitoring observability, how we collect feedback and how we iterate on that feedback. We're going to lastly also talk about tools and skills and how at OpenGov we use AI not only externally that we serve to customers but also internally to improve our development workflows.
SPEAKER_00
Just a little bit about me before we go any further. My name is Gabe. I'm a software engineer here at OpenGov. I work on the AI agents team and I'm one of the folks that helped build OG Assist and some of the systems that you guys will be seeing today. So a little bit about OpenGov. OpenGov is a software company on a mission to power more effective and accountable government. So OpenGov sells ERP software that's things like budgeting, procurement, asset management and permitting and we were founded about 14 years ago and what's cool is we have this thing called OG Assist and OG Assist is this little button on the top of all of our products in the navigation bar. And what's cool is all of our product suites and product teams have built tools and skills in order to power this button. So for example, if I open up this, if I click this button and I open up OG Assist, it says, hey, I'm going to ask about rate codes, which is very specific to utility billing, the current product that I'm in. And you can see that inside of this chat interface, I'm able to speak to an agent and the agent is able to make tool calls in order to look up information against data inside of that suite. So it's really cool to be able to first party create these experiences through the capability that we've built called OG Assist.
SPEAKER_00
Okay, so just a quick story about how this all came to be. So a little while back, we saw that AI was really starting to take off and a principal spun up this new team called the AI agents team and asked me to join. And instantly I said yes. And OG Assist started to grow and we started to integrate OG Assist into all our products and not only our back end capabilities, but also our front end capabilities as well. So you'll see that one of the capabilities that we give the agent is it's able to see what's on the screen and see and take action on what's on the page. So you could see that I'm asking the agent here, hey, what's the screen? Can you maybe highlight some of the next steps that I could take? So you can see that the agent here is thinking it's saying, okay, what tools do I have available to use? And hey, let me go and highlight something that you could actually click on and tell you more about it. So just another capability of OG Assist and just a little short story about how this all came to be. So the big bet on Effect. So I really wanted to include this slide because here on the agents team, we made a huge bet to bet on Effect. And suffice to say, it's paid off in dividends. We write Effect. So Effect is this library for TypeScript. It's open source and it helps you write better TypeScript code. It's got a lot of stuff baked in and a schema similar to Zod, if you've ever used that, it's also got things for error handling, for logging, for traces, for it's just got so much in there. It really helps write better code and structure your code better and helps with architecture, spinning up new services for and for us on the agents team, really helping design and build the core agent loop. So you'll see throughout this presentation sprinkled in how Effect on our team has paid off in dividends. So we really love Effect here at OpenGov and we encourage other folks to try it out and let's keep going.
SPEAKER_00
The Effect native loop. So originally we were on LangGraph and that was fine until the team really started to scale and our use cases started to evolve. So we decided to move over to our own Effect native agent loop to have full regency over this agent loop such that if we have complex use cases or features that we need to build, we could get in, we had full control of the agent loop and not only that but now we're fully on Effect. So all the cool things you get with Effect is now propagated throughout the entire agent loop like the tracing, structured concurrency, the logging, everything is more fine-grained control and it really allows us to unlock the full potential having our own agent loop from the ground up. So another thing I wanted to mention is on the left side you'll see a code example. This is really the basics of the Effect loop that we're using. We're using this thing called the Effect AI package and in that package there's this thing called there's a chat and a language model. So with the chat you can instantiate a chat for example and then you could stream text using that stream text function. You could pass in a prompt and what's cool is with a language model under the hood of since we're doing dependency injection we could pass in a different language model if we were to hot swap to another one for example. So really just having full control of our own agent loop just gives us all the levers and it really just unlocks the full capabilities of the model and for the team as well to have full agency over this loop.
SPEAKER_00
Another thing I wanted to mention is the agent to agent protocol. So here on the agents team we've had a lot of success with this protocol. So this protocol being the protocol that Google created, an open protocol for agents to intercommunicate but we found this very useful for defining our agent routes like for example in the back end and our model and our schema to follow this agent protocol. So we modeled so for example there's this thing called an agent card which you see here and it's got the name of the agent a description etc and having this rigorous protocol, this rigorous spec really helped drive our development and drive alignment because all we had to do was align with this spec and
SPEAKER_00
Another thing I wanted to mention is the agent to agent protocol. So here on the agents team we've had a lot of success with this protocol. This protocol is the protocol that Google created, an open protocol for agents to intercommunicate, but we found this very useful for defining our agent routes, like for example in the back end and our model and our schema to follow this agent protocol. So we modeled, for example, there's this thing called an agent card which you see here and it's got the name of the agent, a description, etc., and having this rigorous protocol, this rigorous spec really helped drive our development and drive alignment because all we had to do was align with this spec and follow this spec and we knew that this was the contract that our front end and back end would both consume and produce. So this I would say has also been very helpful for us, and what's really cool is A2A has a lot of extensions, so you could extend the protocol, add in metadata. There's also A2UI, so lots of fun stuff with A2A protocol, but this is what's worked for us, so sharing that with you folks. Feedback and evals. So here the quote is shipping is the start, not the finish. So what we do here on the agents team is we have multiple ways we do evals and collect feedback. Obviously we'll have folks call in or email us or let us know, but the main way is we have this thumbs up and thumbs down mechanism, and here someone is able to tell us "this worked really well, this was a great response" or "that wasn't a great response," and that signal we take and we're able to iterate on and we can take it back and help improve the response in the future. We also have automated evals, so in the RCI we have evals that run against real completion, so we could test a prompt against "did it hit some tools, did it do what it's supposed to do," and that also helps with our accuracy. So those automated evals in conjunction with collecting feedback really help us improve our tools, our skills, our harness, and that's really how we're able to iterate so fast and so quickly. Humans in the loop. So this is a really cool feature we built where we deterministically interrupt the agent loop if there is a tool call approval required. So if an agent tries to make a tool call that it needs human approval for, it'll show this UI and the human can click accept or reject, explicitly rejecting or explicitly accepting the action that the agent is trying to make, and this ensures that we're building trust and also ensuring that we're being safe, especially when the agent is trying to do a mutating operation, and always making sure that humans are in the driver's seat. Sandboxing. So another thing that we worked on, similar to the safety slide we just saw, was whenever an agent tries to execute code or tries to create files, it does so in a sandbox. So we gave our agents sandboxes such that it could spin up these sandboxes on demand and it could use those sandboxes to write code, execute code, create files, and it's this safe, ephemeral, isolated space such that the agent can take action there and we don't have to worry about any risk to our production systems. It's really cool because they also get teared down at the end. So in this example, I said, "Hey, create a PDF for the folks of the AI Engineer Conference 2026 and allow me to download it so I can share it with them," and you can see that the agent created this really cool PDF inside that sandbox. So I just really wanted to cover this sandbox feature and give you a brief overview of sandboxing. Long context. So inside of OG Assist we have hit many hurdles, especially with legacy models, with token limits or just way too much, completely overloaded with context, especially as conversations get longer. So we found that having some sort of rolling summarization was more effective than always stuffing in the latest and most recent messages. Rather, give a running summary after n number of messages, and maybe you only want the n minus five most recent messages or n minus ten most recent messages, right? And it may be that you're only talking about a specific topic now, but you may want to refer to context earlier, like a hundred messages above. Then that's where the memory component comes in, because when you have this rolling summary of a really long conversation, then you could do recall over that summarization, and if you ask the agent, "Hey, remember that thing that we talked about?" then the agent within the thread will be like, "Yeah, I do know what you were talking about. I have this short tidbit," and it can follow up and do more with that rolling summary in mind. So that's how we handled long context and memory, and it's worked pretty well for us. I just wanted to share a little bit about that and how we've solved the long context problem. UI on the fly. So in this example, I said to the agent, "Hey, generate me a long essay but give me some examples about what the essay could be about." So what's really cool is the agent had this primitive registered of this form and it was able to build out this form for me at runtime and give me some options of what I could choose from. So it feels very personal and very in the moment that it's able to give me these options at runtime. So this is a short thing I wanted to include here about generative UI and how we are able to render UIs on the fly. You can't scale what you can't see. So this section is about tracing and observability. What's cool about Effect is you get tracing out of the box. When you use these Effect functions, they all get tagged automatically with these spans, and the span gets picked up and feeds into these traces so that you can get these drill downs of these function calls. So here is an example of a trace from the Effect team. I have it linked. You could see that when you hit this API, it goes to this endpoint, to this handler, etc., and it takes, and what's really cool is you could profile all your traces, so this takes a total of this many seconds and you can see where the bottleneck is. If there's a failure, you can cross reference it across services. So really important, especially working in agentic systems where we're integrating with other teams and other APIs and other platform capabilities. So what's cool with Effect is you get all this tracing out of the box, and it really makes building this agentic experience, debugging it, and maintaining it just a breeze. Tools and skills. So not only did we make a big bet on Effect, but we also made a big bet on tools and skills. So we believe that tools and skills are really all you need, and in this case you can see on the left we have this tool called get dad joke, and this is the Effect way and the building blocks of how we do things here at Open Gov, but this is pulled from the Effect website.
SPEAKER_00
agentic systems where we're integrating with other teams and other APIs and other platform capabilities. So what's cool with Effect is you get all this tracing out of the box, and it really makes building this agentic experience, debugging it, and maintaining it just a breeze.
SPEAKER_00
Tools and skills. So not only did we make a big bet on Effect, but we also made a big bet on tools and skills. So we believe that tools and skills are really all you need. In this case, you can see on the left we have this tool called get_dad_joke, and this is the Effect way and the building blocks of how we do things here at Open Gov. But this is pulled from the Effect website, but you can see hey, this is how you make a tool, and then you add it to a toolkit, which is a collection of tools, and then you can register this toolkit with the language model. So for example, if you had a prompt that said, "Hey, generate some dad jokes about pirates," well, the agent has a tool that can help get a dad joke. So really, this is the building blocks of how we did tools and eventually skills, and it has paid off wonderfully for our organization. So we really recommend trying out this Effect AI package from Effect and trying out building out your own tools and skills.
SPEAKER_00
Developer velocity. So not only do we build agents for our customers, but we also use agents internally here in Open Gov. So we use a lot of Claude and Cursor, and it's been a real game changer for our team. It's funny because we're building tools and skills for customer-facing agents, and that has been great, but we're also building them internally as well to help accelerate our development workflows. So things like Claude, Cursor, Cloud agents—they really help accelerate how we read, write, review code, and ship. So it's been such an accelerant. So definitely wanted to mention that.
SPEAKER_00
Before we wrap up, that's it. Thanks so much for watching. You've made it to the end. Let's build agents that ship to production. Assist's big bet on effect. A little bit into our core agent loop, we're going to talk about the A2A protocol, evals and sandboxing. We're going to talk about how we manage long context. We're going to talk about monitoring observability, how we collect feedback and how we iterate on that feedback. We're going to lastly also talk about tools and skills and how at OpenGov we use AI not only externally that we serve to customers but also internally to improve our development workflows.
SPEAKER_00
Just a little bit about me before we go any further. My name is Gabe. I'm a software engineer here at OpenGov. I work on the AI agents team and I'm one of the folks that helped build OG Assist and some of the systems that you guys will be seeing today. So a little bit about OpenGov. OpenGov is a software company on a mission to power more effective and accountable government. So OpenGov sells ERP software that's things like budgeting, procurement, asset management and permitting and we were founded about 14 years ago and what's cool is we have this thing called OG Assist and OG Assist is this little
SPEAKER_00
button on the top of all of our products in the navigation bar. And what's cool is all of our product suites and product teams have built tools and skills in order to power this button. So for example, if I open up this, if I click this button and I open up OG Assist, it says, hey, I'm going to ask about rate codes, which is very specific to utility billing, the current product that I'm in. And you can see that inside of this kind of chat interface, I'm able to speak to an agent and the agent is able to make tool calls in order to look up information against data inside of that suite. So it's really cool to be
SPEAKER_00
able to kind of first party create these experiences through the capability that we've built called OG Assist. Okay, so just a quick story about how this all came to be. So a little while back, we saw that AI was really starting to take off and a principal spun up this new team called the AI agents team and asked me to join. And instantly I said yes. And OG Assist started to grow and we started to integrate OG Assist into all our products and not only our back end capabilities, but also our front end capabilities as well. So you'll see that one of the capabilities that we give the agent is it's able to
SPEAKER_00
see what's on the screen and see and take action on what's on the page. So you could see that I'm asking the agent here, hey, what's the screen? Can you maybe highlight some of the next steps that I could take? So you can see that the agent here is thinking it's saying, okay, what tools do I have available to use? And hey, let me go and highlight something that you could actually click on and tell you more about it. So just another capability of OG Assist and just a little short story about how this all came to be. So the big bet on effect. So I really wanted to include this slide because
SPEAKER_00
here on the agents team, we made a huge bet to bet on effect. And suffice to say, it's paid off in dividends. We write effect. So effect is this library for TypeScript. It's open source and it helps you write better TypeScript code. You know, it's got a lot of stuff baked in and like a schema similar to like Zod, if you've ever used that, it's also got things for error handling, for logging, for traces, for it's just got so much in there. It really helps write better code and structure your code better and helps with architecture, spinning up new services for and for us on the agents team, really helping
SPEAKER_00
design and build the core agent loop. So you'll see throughout this presentation sprinkled in how effect on our team has paid off in dividends. So we really love effect here at OpenGov and we encourage other folks to try it out and yeah, let's keep going. The effect native loop. So originally we were on Landgraf and that was fine until the team really started to scale and our use cases started to evolve. So we decided to move over to our own kind of effect native agent loop to have full regency over this agent loop such that if we have complex use cases or features that we need to build, we could kind of get in, we had full control of the agent loop and not
SPEAKER_00
only that but now we're fully on effect. So all the cool things you get with effect is now propagated throughout the entire agent loop like the tracing, structured concurrency, the logging, everything is more fine-grained control and it really allows us to really unlock the full potential having our own agent loop from the ground up. So another thing I wanted to mention is on the left side you'll see a code example. This is really the basics of the effect loop that we're using. We're using this thing called the effect AI package and in that package there's this thing called there's a chat and a language model. So with the chat you can instantiate like a chat for example
SPEAKER_00
and then you could stream text using that kind of stream text function. You could pass in a prompt and what's cool is with a language model under the hood of since we're kind of doing dependency injection we could pass in a different language model if we were to hot swap to another one for example. So really just having full control of our own agent loop just kind of gives us all the levers and it really just unlocks the full capabilities of the model and for the team as well to have full agency over this loop. Another thing I wanted to mention is the agent to agent protocol. So here on the agents team we've had
SPEAKER_00
a lot of success with this protocol. So this protocol being the protocol that Google created kind of an open protocol for agents to intercommunicate but we found this very useful for defining our agent routes like for example in the back end and our model and our schema to follow this kind of agent protocol. So we modeled so for example there's this thing called an agent card which you see here and it's got the name of the agent a description etc right and having this kind of rigorous protocol this rigorous spec really helped drive our development and drive alignment because you know all we had to do was align with this spec and
SPEAKER_00
follow this spec and we knew that this was kind of the contract that our front end and back end would both consume and and produce. So this I would say also has been very helpful for us and and what's really cool is A2A has a lot of extensions right so you could extend the protocol add in like metadata there's also A2UI so lots of fun stuff with A2A protocol but this is kind of what's worked for us so sharing that with with you folks. Feedback and evals. So here the quote is shipping is the start not the finish. So what we do here on the agents team is we have kind of multiple ways we do evals and collect feedback. Obviously you
SPEAKER_00
know we'll have folks call in or email us or just let us know and tell us but the main way is we have this thumbs up and thumbs down mechanism and here someone is able to tell us hey this this worked really well this was a great response or that wasn't a great response and that signal we take and we're able to iterate on and we can take it back and help improve you know the response in the future. We also have automated evals so in in the in RCI we we have evals that run against real completion so we could test a prompt against hey did it hit some tools did it do what it's supposed to do
SPEAKER_00
and that also helps with our accuracy. So those automated evals in conjunction with collecting feedback really help us improve our our our tools our skills our harness and and that's really how we're able to iterate so fast and so quickly. Humans in the loop. So this is a really cool feature we built where we deterministically interrupt the agent loop if there is a tool call approval required. So if an agent tries to make a tool call that it needs human approval for it'll show this UI and the human can click accept or reject so explicitly rejecting or explicitly accepting the action that the agent is
SPEAKER_00
trying to make and this ensures that you know we're building trust and also ensuring that you know we're being safe especially when the agent is trying to do a mutating operation and always always always making sure that humans are in the driver's seat. Sandboxing. So another thing that we worked on kind of similar to the safety slide we just saw was whenever an agent tries to execute code or tries to create files it does so in a sandbox so we gave our agents sandboxes such that it could spin up these sandboxes on demand and it could use those sandboxes to honestly write code execute code create files and it's kind of this safe
SPEAKER_00
ephemeral isolated space such that the agent can can can take action in there and not and we don't have to worry about any risk to you know our production systems and it's really cool because they also get uh tight uh teared down at the end so um in this example i said hey create a pdf uh for the folks of the ai engineer conference 2026 and um allow me to download it so i can share it with them and you can see that the agent created this really cool pdf inside of that sandbox so just really wanted to cover this sandbox feature and just give you a brief kind of overview of of sandboxing
SPEAKER_00
um long context so um inside of og assist we have hit many hurdles like with especially with legacy models now like with uh token limits or just way too much just completely overloaded with context um especially as conversations get longer so we found that having some sort of um rolling summarization was more effective than you know always stuffing in the latest and most recent uh messages uh rather just you know give like a running summary after n number of messages and uh maybe you only want the like n minus five most recent messages or n minus 10 most recent messages right and it may be that
SPEAKER_00
you're only talking about a specific topic now but you may want to refer to uh context earlier like 100 messages above then um it uh that's where kind of the the memory component comes in because when you have this rolling summary of a really long conversation then you could do recall over that uh summarization and um you know if you ask the agent hey remember that thing that we talked about then the agent within the thread will be like yeah i do know what you were talking about i have kind of this short little tidbit and it can you know follow up and and do more kind of with that uh and summary rolling summary in mind so
SPEAKER_00
that's kind of how we handled long context and memory and it's worked pretty well for us so uh just wanted to share a little bit about that and how we've uh solved the long context problem um ui on the fly so um in this example i said to the agent hey generate me a long essay but give me some examples about what the essay could be about so what's really cool is the agent had this primitive registered of this form and it was able to build out this form for me at runtime and give me some options of what i could choose from so it feels very personal and it feels very kind of in the moment that
SPEAKER_00
it's able to to give me these options just that runtime so um this is kind of just a short little um thing i wanted to include here about generative ui and how we are able to render uis on the fly you can't scale what you can't see so um this uh this kind of section is about tracing uh and observability really what's cool about effect is you kind of get tracing out of the box um you know when you use these effect functions they all get kind of tagged automatically with like these spans and kind of the span gets picked up and feeds into these traces so that you can kind of get these kind of
SPEAKER_00
drill downs of these function calls so here is an example of a trace from the effect team um i have it linked uh you could see that like hey when you hit this api it goes to this endpoint to this handler and uh you know etc and it takes and what's really cool is you could profile all your traces right so this takes a total of this many seconds and you can see where the bottleneck is here if there's a failure you can cross reference it across services so really really important especially working in agentic systems where we're we're integrating with other teams and other apis and other um
SPEAKER_00
other platform capabilities so uh what's cool with um effect is you get all this tracing out of the box and it really makes building this agentic experience debugging it and maintaining it just the breeze tools and skills so uh not only did we make a big bet on effect but we also made a big bet on tools and skills so we believe that tools and skills are really all you need and in this case you can see on the left we have we have this tool called get dad joke and this is kind of the effect way and kind of the building blocks of how we do things here at open gov but you know this is pulled from the effect website
SPEAKER_00
but you can see hey this is how you make a tool and then you add it to a toolkit which is a collection of tools and then you can register this this toolkit with the with the language model so for example if you had a prompt that said hey generate some dad jokes about pirates well guess what the agent has a tool uh that can help get a dad joke so um really this is the the building blocks of how we did tools uh and eventually skills and really just is has paid off wonderfully for for our organization so um really recommend trying out this effect ai package from uh effect and um just trying out building out your own tools and skills
SPEAKER_00
developer velocity so not only do we build agents for our customers but we also use agents um here internally uh in open gov so we use a lot of cloud and cursor uh and it's just really been a game changer for our team um it's funny because we're building tools and skills for customer facing uh agents and and that has been great but we're also building them internally as well to help accelerate our development workflows so things like claude cursor cloud agents they really help accelerate how we read write review code and and ship um so it's it's just been such an accelerant so definitely just wanted to mention that
SPEAKER_00
uh before we wrap up uh before we wrap up that's it thanks so much for watching you've made it to the end let's build agents that ship to production