AI Engineer

Agents in Production: How OpenGov Built and Scaled OG Assist - Gabe De Mesa, OpenGov

3492 summary words 16 min summary Watch video

Start with the signal

16 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: OpenGov successfully scaled production AI agents (OG Assist) across their government ERP platform by building a custom Effect-native agent loop, standardizing on the A2A protocol, and implementing comprehensive tooling/eval infrastructure—demonstrating that framework independence, rigorous observability, and human-in-the-loop safeguards are essential for enterprise agentic systems.
  • Why it matters: Real production case study showing how a traditional enterprise (government ERP) integrated agents into multiple product suites at scale, with specific architectural decisions, safety mechanisms, and team practices that actually shipped and survived customer use.
  • Best use: Study the architectural choices (Effect loop vs LangGraph, A2A protocol adoption, sandbox/approval gates) and operational practices (rolling summarization, thumbs up/down feedback loops, CI evals) as a blueprint for building production-grade agentic systems; use as reference for agent infrastructure decisions and team organization.

Executive Summary

Gabe De Mesa from OpenGov presents a detailed case study of OG Assist, an AI agent deployed across all of OpenGov's government ERP products (budgeting, procurement, permitting, utility billing). OpenGov made two critical architectural bets: migrating from LangGraph to a custom Effect-native agent loop for full control and observability, and standardizing on Google's Agent-to-Agent (A2A) protocol to align backend/frontend contracts. The agent is embedded as a button in every product suite, with product teams building suite-specific tools and skills that power domain-specific interactions (e.g., utility billing rate codes).

The production system includes multiple safety and quality layers: deterministic human-in-the-loop approval for mutating operations, sandboxed code execution environments, rolling conversation summarization to manage long context, and dual feedback mechanisms (thumbs up/down plus automated CI evals). The Effect framework provides built-in tracing, structured logging, and dependency injection, which proved essential for debugging multi-service agentic workflows. OpenGov also demonstrates visual capabilities—the agent can inspect page state and generate UI components (forms, highlights) at runtime.

OpenGov's philosophy centers on 'tools and skills are all you need,' building composable toolkits using Effect AI packages. They collect user feedback via thumbs up/down to iterate on prompts and tool behavior, running automated evals in CI against real completions to test tool invocation accuracy. The team uses agents internally (Claude, Cursor) to accelerate their own development, creating a flywheel where building customer-facing agents improves their internal workflows.

Key production lessons: framework independence provides full control as use cases evolve; rigorous protocol adoption (A2A) drives alignment across teams; observability (Effect tracing) is non-negotiable for debugging distributed agentic systems; human approval gates and sandboxing are critical for trust in enterprise environments; and 'shipping is the start, not the finish'—continuous feedback loops and automated evals enable rapid iteration post-deployment.

Key Takeaways

  • Claim: OpenGov migrated from LangGraph to a custom Effect-native agent loop to gain full control over the agent execution cycle as their use cases scaled and evolved. | Evidence: Team originally used LangGraph but found it insufficient as they scaled; built custom loop using Effect AI package (chat, language model, stream text functions) with dependency injection for model swapping; gained fine-grained control over tracing, structured concurrency, logging throughout the entire loop. | Caveat: Migration requires upfront investment in building core loop primitives; only justified when framework limitations block critical features or when deep observability/control becomes necessary; Effect has a learning curve for teams unfamiliar with functional TypeScript patterns. | Implication: For operators building production agents, evaluate framework lock-in early—if your use cases are evolving rapidly or you need granular observability/customization, building on a composable runtime (Effect, raw OpenAI SDK, etc.) may be more sustainable than higher-level frameworks like LangGraph or LangChain. | Timestamp: timestamp unavailable
  • Claim: Standardizing on Google's Agent-to-Agent (A2A) protocol as the contract between frontend and backend drove development alignment and enabled consistent agent routing across OpenGov's product suites. | Evidence: A2A protocol defines agent cards (name, description, tools); OpenGov modeled backend routes and schemas to match A2A spec; frontend and backend both consume/produce A2A contracts; protocol is extensible (metadata, A2UI extensions). | Caveat: A2A is relatively new and adoption outside Google ecosystem is still emerging; requires buy-in across teams and upfront schema design work; may not fit all use cases if your agent architecture diverges significantly from the protocol's assumptions. | Implication: If you're building multi-agent systems or coordinating agents across teams/products, adopting a shared protocol (A2A or custom) early can prevent schema drift and integration headaches; the rigidity of a spec accelerates development by reducing ambiguity. | Timestamp: timestamp unavailable
  • Claim: OpenGov uses deterministic human-in-the-loop approval interrupts for tool calls requiring authorization, ensuring humans remain in control and building user trust for mutating operations. | Evidence: When agent attempts a tool call flagged as requiring approval, loop interrupts and shows UI with accept/reject buttons; explicitly requires human authorization for mutating actions (e.g., deletes, writes). | Caveat: Adds latency and friction to agent workflows; requires clear criteria for which tools need approval vs. which can auto-execute; UI/UX must be seamless to avoid user frustration; may not scale for high-frequency agent actions. | Implication: For enterprise/high-stakes environments, approval gates are table stakes for deploying agents with write access; design tool metadata/flags upfront to declaratively mark approval requirements; consider async approval flows for non-blocking workflows. | Timestamp: timestamp unavailable
  • Claim: Agents execute code and create files in ephemeral sandboxes to isolate risk from production systems, enabling safe experimentation and code generation. | Evidence: Agents spin up on-demand sandboxes; can write/execute code and create files (e.g., PDFs) inside isolated environments; sandboxes are torn down after use; example shown: agent creates a PDF for AI Engineer Conference 2026 and allows download. | Caveat: Sandboxing adds infrastructure overhead (spinning up/tearing down environments); may introduce latency; requires secure sandbox implementation to prevent escapes; file persistence/sharing requires explicit download/export mechanisms. | Implication: If your agents generate or execute code, sandboxing (E2B, Modal, Docker, etc.) is critical for production safety; design for ephemeral, isolated execution from day one; consider cost/latency tradeoffs for frequent sandbox usage. | Timestamp: timestamp unavailable
  • Claim: OpenGov handles long conversations and token limits via rolling summarization (summary after n messages, keep only n-5 most recent) plus memory-based recall over summaries. | Evidence: Rolling summarization triggers after n messages; recent messages (e.g., last 5-10) kept verbatim; older context compressed into running summary; agent can recall earlier topics by searching over summary when user asks 'remember that thing we talked about?' | Caveat: Summarization loses detail and nuance; recall accuracy depends on summary quality; may miss context that seemed unimportant during summarization but becomes relevant later; requires tuning n threshold per use case. | Implication: For long-running agent sessions (support, project management), implement summarization + recall early to avoid context window blowouts; test summary quality on real conversations to ensure recall works; consider hybrid approaches (summaries + semantic search over full history). | Timestamp: timestamp unavailable
  • Claim: OpenGov collects feedback via thumbs up/down on agent responses and runs automated CI evals against real completions to iterate on prompts, tools, and skills. | Evidence: Users can thumbs up/down responses in chat UI; team uses signal to improve prompts and tool behavior; automated evals in CI test prompts against 'did it hit the right tools, did it do what it's supposed to do'; quote: 'shipping is the start, not the finish.' | Caveat: Thumbs down may not provide root cause detail without follow-up questions; users may not give feedback consistently; automated evals require maintaining test cases and expected behaviors; eval quality depends on coverage of real-world scenarios. | Implication: Implement dual feedback loops from day one: qualitative (user ratings) to identify issues and quantitative (automated evals) to prevent regressions; use thumbs up/down to prioritize eval creation; instrument CI to block deploys on eval failures for critical paths. | Timestamp: timestamp unavailable
  • Claim: Effect provides built-in tracing, structured logging, and span instrumentation that proved essential for debugging multi-service agentic workflows at OpenGov. | Evidence: Effect functions auto-tagged with spans; traces feed into observability tools showing function call drill-downs, bottleneck profiling (e.g., 'this endpoint took X seconds'), and cross-service failure correlation; quote: 'you can't scale what you can't see.' | Caveat: Effect requires team ramp-up on functional TypeScript patterns; built-in tracing may need integration with external observability platforms (Datadog, Honeycomb, etc.) for production dashboards; tracing overhead can impact performance at high scale. | Implication: For agentic systems integrating multiple services/APIs, invest in structured tracing early—either via Effect, OpenTelemetry, or vendor SDKs; prioritize traces that show LLM calls, tool invocations, and latency breakdowns; use traces to identify cascading failures across services. | Timestamp: timestamp unavailable
  • Claim: OpenGov's agent can inspect page state and generate UI components (forms, highlights) at runtime, enabling context-aware and personalized interactions. | Evidence: Agent has tools to 'see what's on the screen' and take action on page elements (e.g., highlight next steps user can click); example: agent generates a form at runtime with essay topic options when user asks for essay ideas; uses registered UI primitives (forms) to render components on the fly. | Caveat: Requires tight integration between agent backend and frontend rendering engine; UI generation quality depends on prompt quality and available primitives; may produce inconsistent UX if not constrained; accessibility and responsiveness of generated UI must be validated. | Implication: For agent-assisted applications (SaaS, productivity tools), consider building a library of generative UI primitives (forms, modals, highlights) that agents can compose; enable agents to inspect DOM/state for context-aware suggestions; validate generated UI for accessibility/usability before rendering. | Timestamp: timestamp unavailable

Detailed Brief

OpenGov's OG Assist Architecture and Evolution

  • Claims: OG Assist is embedded as a button across all OpenGov ERP product suites (budgeting, procurement, permitting, utility billing); Each product team builds suite-specific tools and skills to power the agent for their domain; Originally built on LangGraph, migrated to custom Effect-native loop for full control as team and use cases scaled; Core loop uses Effect AI package (chat, language model, stream text) with dependency injection for model swapping
  • Evidence: OG Assist button appears in navigation bar of every product; chat interface allows domain-specific queries (e.g., utility billing rate codes); Migration decision driven by need for 'full regency over agent loop' as use cases evolved and team scaled; Effect provides tracing, structured concurrency, logging, error handling, schema (Zod-like) out of the box; Code example shows instantiating chat, streaming text with prompt, passing in language model via dependency injection
  • Caveats: Migration required building core loop primitives from scratch; Effect has learning curve for teams not familiar with functional TypeScript; Framework-specific bets (Effect vs. LangGraph) carry long-term maintenance and hiring implications
  • Implications: Organizations should evaluate framework flexibility early—if scaling or custom features are anticipated, lower-level libraries (Effect, raw SDKs) may be more sustainable; Effect's opinionated patterns (dependency injection, structured concurrency) can improve code quality but require team training; For operators, consider whether your agentic use cases are stable enough for high-level frameworks or will require custom loop logic

Safety, Trust, and Risk Management in Production

  • Claims: Deterministic human-in-the-loop interrupts for tool calls requiring approval (mutating operations); Agents execute code and create files in ephemeral sandboxes, not production systems; Sandbox environments are spun up on demand and torn down after use; Approval UI shows accept/reject buttons, ensuring explicit human authorization
  • Evidence: When agent tries tool call needing approval, loop interrupts with UI; user clicks accept/reject; Sandboxes allow agents to write/execute code, create files (example: PDF generation for AI Engineer Conference 2026); Sandboxes are 'safe, ephemeral, isolated spaces' that eliminate risk to production; Philosophy: 'always making sure humans are in the driver's seat'
  • Caveats: Approval gates add latency and may frustrate users if overused; Sandboxing requires infrastructure investment and secure implementation to prevent escapes; Requires clear criteria for which tools need approval vs. auto-execution; File persistence/sharing from sandboxes requires explicit export mechanisms
  • Implications: For enterprise deployments, approval gates are mandatory for write/delete operations; design tool metadata to declaratively flag approval needs; Sandboxing (E2B, Modal, Docker) should be architected from day one for code-generating agents; consider cost/latency tradeoffs; Trust-building with users requires visible control mechanisms—design UI/UX carefully to communicate agent limitations and human control; Security teams will require sandbox audit trails and escape prevention measures before production approval

Observability, Feedback Loops, and Continuous Improvement

  • Claims: Effect provides built-in tracing with automatic span tagging for all Effect functions; Traces show drill-downs of function calls, bottleneck profiling (time per endpoint/handler), and cross-service failures; Dual feedback mechanisms: user thumbs up/down on responses, plus automated CI evals against real completions; Quote: 'shipping is the start, not the finish'—emphasis on post-deploy iteration; Evals test 'did it hit the right tools, did it do what it's supposed to do'
  • Evidence: Trace example from Effect team shows API hit → endpoint → handler with time profiling per step; Thumbs up/down UI in chat interface; team uses signal to improve prompts/tools; CI evals run against real completions to validate tool invocation and behavior; Feedback loops enable team to 'iterate so fast and so quickly' on harness, tools, skills
  • Caveats: Thumbs down lacks detail without follow-up questions; users may not provide consistent feedback; Automated evals require maintaining test cases and expected behaviors; eval quality depends on scenario coverage; Tracing overhead can impact performance at high scale; may need integration with external observability platforms; Effect tracing is powerful but requires team to learn how to read/interpret traces
  • Implications: Operators should instrument both qualitative (user feedback) and quantitative (evals) loops from day one; Use thumbs up/down to prioritize which eval scenarios to build; treat user feedback as backlog for eval creation; For distributed agentic systems, invest in structured tracing (Effect, OpenTelemetry) early; prioritize traces showing LLM calls, tool invocations, latency; Post-launch iteration velocity depends on observability depth—skimping on tracing/logging infrastructure will bottleneck debugging; Consider blocking deploys on critical eval failures to prevent regressions in production

Context Management, UI Generation, and Developer Experience

  • Claims: Rolling summarization after n messages, keeping only n-5 most recent messages verbatim; Memory/recall over summarization allows agent to reference earlier topics ('remember that thing we talked about?'); Agent can generate UI components (forms, highlights) at runtime using registered primitives; Team uses AI agents internally (Claude, Cursor, Cloud agents) to accelerate development workflows; Internal agents help with reading, writing, reviewing code, and shipping—'such an accelerant'
  • Evidence: Rolling summary solves token limit issues with legacy models and overloaded context as conversations grow; Example: agent generates essay topic form at runtime with options to choose from; renders form using registered primitive; Page inspection capability: agent can highlight next steps user can click on page; Quote: 'tools and skills are all you need'—composable toolkits using Effect AI package (get_dad_joke example)
  • Caveats: Summarization loses detail; recall accuracy depends on summary quality; may miss context that becomes relevant later; Requires tuning n threshold per use case; no one-size-fits-all; UI generation quality depends on prompt and available primitives; may produce inconsistent UX if not constrained; Generated UI must be validated for accessibility and responsiveness; Internal agent usage (Claude, Cursor) requires budget for developer seats and LLM costs
  • Implications: For long-running agent sessions (support, project management), implement summarization + recall early; test on real conversations; Consider hybrid context strategies: summaries for old context + semantic search over full history for critical details; If building agent-assisted SaaS, invest in a library of composable UI primitives agents can render; enable DOM inspection for context; Internal agent adoption creates flywheel: building customer agents improves team's own workflows; budget for internal LLM usage as productivity investment; For operators, generative UI requires frontend rendering engine tightly integrated with agent backend; validate UX before rendering

A2A Protocol and Tools/Skills Philosophy

  • Claims: OpenGov standardized on Google's Agent-to-Agent (A2A) protocol for agent routing and schema contracts; A2A defines agent cards (name, description, tools); both frontend and backend consume/produce A2A contracts; Philosophy: 'tools and skills are all you need'—big bet on composable toolkits; Tools built using Effect AI package; toolkits are collections of tools registered with language model; A2A protocol is extensible (metadata, A2UI extensions)
  • Evidence: A2A agent cards shown in presentation; rigorous spec 'helped drive development and drive alignment'; Quote: 'all we had to do was align with this spec and follow this spec'—reduced frontend/backend ambiguity; Code example: get_dad_joke tool → add to toolkit → register with language model; agent can invoke tool based on prompt (e.g., 'generate dad jokes about pirates'); Product teams build suite-specific tools/skills; composability enables cross-product agent capabilities
  • Caveats: A2A adoption outside Google ecosystem is still emerging; may not fit all architectures; Requires upfront schema design work and team buy-in across frontend/backend; Tools/skills approach can proliferate tool count; requires governance to avoid tool bloat and ambiguity; Extensibility of A2A is powerful but may lead to schema drift if not carefully managed
  • Implications: If coordinating agents across teams/products, adopt a shared protocol (A2A or custom) early to prevent integration friction; Rigidity of a spec accelerates development by reducing ambiguity—tradeoff is less flexibility; For operators, composable tools/skills are powerful but require tooling governance: naming conventions, deprecation policies, documentation; Effect AI package provides solid primitives for tool building; evaluate whether Effect's patterns fit your team's skillset; A2A may become de facto standard if Google/others push adoption; monitor ecosystem for tooling/library support

Notable Concepts & Terms

  • Effect (TypeScript library): Open-source functional TypeScript library providing schema (Zod-like), error handling, logging, tracing, structured concurrency, dependency injection; OpenGov's core runtime for agent loop and services; 'paid off in dividends'
  • Effect AI package: Effect's module for building AI agents; provides chat, language model, stream text functions, and composable toolkit primitives for tool building; OpenGov's foundation for tools/skills
  • Agent-to-Agent (A2A) protocol: Google's open protocol for agent intercommunication; defines agent cards (name, description, tools) and contracts; OpenGov uses it to align frontend/backend schemas and agent routing; extensible with metadata and A2UI
  • Rolling summarization: Context management strategy: after n messages, summarize older messages into running summary, keep only n-5 most recent verbatim; enables memory/recall over long conversations without hitting token limits
  • Human-in-the-loop approval: Deterministic interrupts in agent loop when tool call requires authorization; shows UI with accept/reject buttons; ensures humans control mutating operations (deletes, writes)
  • Sandboxing (for agents): Ephemeral, isolated execution environments where agents can write/execute code and create files without risk to production systems; spun up on demand, torn down after use
  • Tools and skills: OpenGov's composable building blocks for agent capabilities; tools are individual functions (e.g., get_dad_joke), skills are collections (toolkits) registered with language model; philosophy: 'all you need'
  • Generative UI: Agent capability to render UI components (forms, highlights) at runtime using registered primitives; enables personalized, context-aware interactions (e.g., essay topic form, page highlights)
  • OG Assist: OpenGov's production AI agent embedded across all ERP product suites; chat interface accessible via button in navigation bar; powered by suite-specific tools/skills built by product teams
  • Structured concurrency: Effect pattern for managing concurrent operations with strict parent-child lifecycle guarantees; ensures resources are cleaned up and failures propagate correctly; critical for reliable agent loops

Operator Notes / Why Ken Should Care

  • Real production case study of enterprise agent deployment (government ERP, multiple product suites); validates that agents can ship and scale in traditional enterprise contexts, not just startups
  • Migration from LangGraph to custom Effect loop demonstrates when framework independence pays off—evaluate your own framework choices against anticipated scaling/customization needs
  • A2A protocol adoption shows value of rigorous contracts in multi-team agent development; consider adopting A2A or similar protocol if coordinating agents across teams/products
  • Dual feedback loops (user thumbs up/down + CI evals) are production table stakes; implement both from day one to enable rapid iteration post-deploy
  • Human-in-the-loop approval and sandboxing are non-negotiable for enterprise trust; design approval gates and sandbox infrastructure early, not as retrofits
  • Effect's built-in tracing proved critical for debugging multi-service agentic workflows; invest in observability (Effect, OpenTelemetry, vendor tools) upfront
  • Rolling summarization is practical solution for long conversations; test summary quality on real sessions and consider hybrid approaches (summaries + semantic search)
  • Generative UI (forms, highlights at runtime) requires tight frontend-backend integration and library of composable primitives; plan for this if building agent-assisted SaaS
  • Internal agent usage (Claude, Cursor) creates flywheel: building customer agents improves team velocity; budget for internal LLM costs as productivity investment
  • Tools/skills composability is powerful but requires governance: naming conventions, deprecation policies, documentation to prevent tool bloat and ambiguity
  • OpenGov's philosophy 'shipping is the start, not the finish' aligns with operator reality: post-launch iteration velocity depends on feedback infrastructure and observability depth

Watch Map

  • timestamp unavailable: Timestamps not provided in transcript; key sections: OG Assist demo and architecture, Effect bet rationale, A2A protocol adoption, safety mechanisms (approval gates, sandboxing), observability with Effect tracing, feedback loops and evals, context management (rolling summarization), generative UI examples, tools/skills philosophy, internal agent usage for developer velocity

Source/Metadata

  • Title: Agents in Production: How OpenGov Built and Scaled OG Assist - Gabe De Mesa, OpenGov
  • Transcript words: 5706
  • Duration seconds: 1109
  • Timestamp note: Timestamps were not present in the transcript; watch_map provides thematic sections but not specific minute markers
Full transcript 2802 words · 25 min read
0:00

SPEAKER_00

Hi everyone, my name is Gabe DeMesa. I'm an engineer here at OpenGov and today we're going to be talking about agents in production, specifically how OpenGov built and scaled OG Assist. So this presentation is going to be jam-packed with good stuff. We're going to talk about AI agents, we're going to talk about our harness, we're going to talk about evals, observability, traces, we're going to talk about tools and skills. We're going to talk to you guys about what we do at OpenGov and how we operate at the scale that we operate at in production so you'll be able to see a real use case and workload with AI agents. So without further ado, let's get started.

0:06

SPEAKER_00

Okay, agenda. So just really quickly going to go through high level what we're going to talk about today. I'm going to tell you guys a little bit about OG Assist and what OpenGov is. I'm going to tell you guys the origin story of how this all came to be. We're going to talk about OG Assist's big bet on Effect. A little bit into our core agent loop, we're going to talk about the A2A protocol, evals and sandboxing. We're going to talk about how we manage long context. We're going to talk about monitoring observability, how we collect feedback and how we iterate on that feedback. We're going to lastly also talk about tools and skills and how at OpenGov we use AI not only externally that we serve to customers but also internally to improve our development workflows.

0:15

SPEAKER_00

Just a little bit about me before we go any further. My name is Gabe. I'm a software engineer here at OpenGov. I work on the AI agents team and I'm one of the folks that helped build OG Assist and some of the systems that you guys will be seeing today. So a little bit about OpenGov. OpenGov is a software company on a mission to power more effective and accountable government. So OpenGov sells ERP software that's things like budgeting, procurement, asset management and permitting and we were founded about 14 years ago and what's cool is we have this thing called OG Assist and OG Assist is this little button on the top of all of our products in the navigation bar. And what's cool is all of our product suites and product teams have built tools and skills in order to power this button. So for example, if I open up this, if I click this button and I open up OG Assist, it says, hey, I'm going to ask about rate codes, which is very specific to utility billing, the current product that I'm in. And you can see that inside of this chat interface, I'm able to speak to an agent and the agent is able to make tool calls in order to look up information against data inside of that suite. So it's really cool to be able to first party create these experiences through the capability that we've built called OG Assist.

0:22

SPEAKER_00

Okay, so just a quick story about how this all came to be. So a little while back, we saw that AI was really starting to take off and a principal spun up this new team called the AI agents team and asked me to join. And instantly I said yes. And OG Assist started to grow and we started to integrate OG Assist into all our products and not only our back end capabilities, but also our front end capabilities as well. So you'll see that one of the capabilities that we give the agent is it's able to see what's on the screen and see and take action on what's on the page. So you could see that I'm asking the agent here, hey, what's the screen? Can you maybe highlight some of the next steps that I could take? So you can see that the agent here is thinking it's saying, okay, what tools do I have available to use? And hey, let me go and highlight something that you could actually click on and tell you more about it. So just another capability of OG Assist and just a little short story about how this all came to be. So the big bet on Effect. So I really wanted to include this slide because here on the agents team, we made a huge bet to bet on Effect. And suffice to say, it's paid off in dividends. We write Effect. So Effect is this library for TypeScript. It's open source and it helps you write better TypeScript code. It's got a lot of stuff baked in and a schema similar to Zod, if you've ever used that, it's also got things for error handling, for logging, for traces, for it's just got so much in there. It really helps write better code and structure your code better and helps with architecture, spinning up new services for and for us on the agents team, really helping design and build the core agent loop. So you'll see throughout this presentation sprinkled in how Effect on our team has paid off in dividends. So we really love Effect here at OpenGov and we encourage other folks to try it out and let's keep going.

0:28

SPEAKER_00

The Effect native loop. So originally we were on LangGraph and that was fine until the team really started to scale and our use cases started to evolve. So we decided to move over to our own Effect native agent loop to have full regency over this agent loop such that if we have complex use cases or features that we need to build, we could get in, we had full control of the agent loop and not only that but now we're fully on Effect. So all the cool things you get with Effect is now propagated throughout the entire agent loop like the tracing, structured concurrency, the logging, everything is more fine-grained control and it really allows us to unlock the full potential having our own agent loop from the ground up. So another thing I wanted to mention is on the left side you'll see a code example. This is really the basics of the Effect loop that we're using. We're using this thing called the Effect AI package and in that package there's this thing called there's a chat and a language model. So with the chat you can instantiate a chat for example and then you could stream text using that stream text function. You could pass in a prompt and what's cool is with a language model under the hood of since we're doing dependency injection we could pass in a different language model if we were to hot swap to another one for example. So really just having full control of our own agent loop just gives us all the levers and it really just unlocks the full capabilities of the model and for the team as well to have full agency over this loop.

0:36

SPEAKER_00

Another thing I wanted to mention is the agent to agent protocol. So here on the agents team we've had a lot of success with this protocol. So this protocol being the protocol that Google created, an open protocol for agents to intercommunicate but we found this very useful for defining our agent routes like for example in the back end and our model and our schema to follow this agent protocol. So we modeled so for example there's this thing called an agent card which you see here and it's got the name of the agent a description etc and having this rigorous protocol, this rigorous spec really helped drive our development and drive alignment because all we had to do was align with this spec and

0:42

SPEAKER_00

Another thing I wanted to mention is the agent to agent protocol. So here on the agents team we've had a lot of success with this protocol. This protocol is the protocol that Google created, an open protocol for agents to intercommunicate, but we found this very useful for defining our agent routes, like for example in the back end and our model and our schema to follow this agent protocol. So we modeled, for example, there's this thing called an agent card which you see here and it's got the name of the agent, a description, etc., and having this rigorous protocol, this rigorous spec really helped drive our development and drive alignment because all we had to do was align with this spec and follow this spec and we knew that this was the contract that our front end and back end would both consume and produce. So this I would say has also been very helpful for us, and what's really cool is A2A has a lot of extensions, so you could extend the protocol, add in metadata. There's also A2UI, so lots of fun stuff with A2A protocol, but this is what's worked for us, so sharing that with you folks. Feedback and evals. So here the quote is shipping is the start, not the finish. So what we do here on the agents team is we have multiple ways we do evals and collect feedback. Obviously we'll have folks call in or email us or let us know, but the main way is we have this thumbs up and thumbs down mechanism, and here someone is able to tell us "this worked really well, this was a great response" or "that wasn't a great response," and that signal we take and we're able to iterate on and we can take it back and help improve the response in the future. We also have automated evals, so in the RCI we have evals that run against real completion, so we could test a prompt against "did it hit some tools, did it do what it's supposed to do," and that also helps with our accuracy. So those automated evals in conjunction with collecting feedback really help us improve our tools, our skills, our harness, and that's really how we're able to iterate so fast and so quickly. Humans in the loop. So this is a really cool feature we built where we deterministically interrupt the agent loop if there is a tool call approval required. So if an agent tries to make a tool call that it needs human approval for, it'll show this UI and the human can click accept or reject, explicitly rejecting or explicitly accepting the action that the agent is trying to make, and this ensures that we're building trust and also ensuring that we're being safe, especially when the agent is trying to do a mutating operation, and always making sure that humans are in the driver's seat. Sandboxing. So another thing that we worked on, similar to the safety slide we just saw, was whenever an agent tries to execute code or tries to create files, it does so in a sandbox. So we gave our agents sandboxes such that it could spin up these sandboxes on demand and it could use those sandboxes to write code, execute code, create files, and it's this safe, ephemeral, isolated space such that the agent can take action there and we don't have to worry about any risk to our production systems. It's really cool because they also get teared down at the end. So in this example, I said, "Hey, create a PDF for the folks of the AI Engineer Conference 2026 and allow me to download it so I can share it with them," and you can see that the agent created this really cool PDF inside that sandbox. So I just really wanted to cover this sandbox feature and give you a brief overview of sandboxing. Long context. So inside of OG Assist we have hit many hurdles, especially with legacy models, with token limits or just way too much, completely overloaded with context, especially as conversations get longer. So we found that having some sort of rolling summarization was more effective than always stuffing in the latest and most recent messages. Rather, give a running summary after n number of messages, and maybe you only want the n minus five most recent messages or n minus ten most recent messages, right? And it may be that you're only talking about a specific topic now, but you may want to refer to context earlier, like a hundred messages above. Then that's where the memory component comes in, because when you have this rolling summary of a really long conversation, then you could do recall over that summarization, and if you ask the agent, "Hey, remember that thing that we talked about?" then the agent within the thread will be like, "Yeah, I do know what you were talking about. I have this short tidbit," and it can follow up and do more with that rolling summary in mind. So that's how we handled long context and memory, and it's worked pretty well for us. I just wanted to share a little bit about that and how we've solved the long context problem. UI on the fly. So in this example, I said to the agent, "Hey, generate me a long essay but give me some examples about what the essay could be about." So what's really cool is the agent had this primitive registered of this form and it was able to build out this form for me at runtime and give me some options of what I could choose from. So it feels very personal and very in the moment that it's able to give me these options at runtime. So this is a short thing I wanted to include here about generative UI and how we are able to render UIs on the fly. You can't scale what you can't see. So this section is about tracing and observability. What's cool about Effect is you get tracing out of the box. When you use these Effect functions, they all get tagged automatically with these spans, and the span gets picked up and feeds into these traces so that you can get these drill downs of these function calls. So here is an example of a trace from the Effect team. I have it linked. You could see that when you hit this API, it goes to this endpoint, to this handler, etc., and it takes, and what's really cool is you could profile all your traces, so this takes a total of this many seconds and you can see where the bottleneck is. If there's a failure, you can cross reference it across services. So really important, especially working in agentic systems where we're integrating with other teams and other APIs and other platform capabilities. So what's cool with Effect is you get all this tracing out of the box, and it really makes building this agentic experience, debugging it, and maintaining it just a breeze. Tools and skills. So not only did we make a big bet on Effect, but we also made a big bet on tools and skills. So we believe that tools and skills are really all you need, and in this case you can see on the left we have this tool called get dad joke, and this is the Effect way and the building blocks of how we do things here at Open Gov, but this is pulled from the Effect website.

0:50

SPEAKER_00

agentic systems where we're integrating with other teams and other APIs and other platform capabilities. So what's cool with Effect is you get all this tracing out of the box, and it really makes building this agentic experience, debugging it, and maintaining it just a breeze.

0:56

SPEAKER_00

Tools and skills. So not only did we make a big bet on Effect, but we also made a big bet on tools and skills. So we believe that tools and skills are really all you need. In this case, you can see on the left we have this tool called get_dad_joke, and this is the Effect way and the building blocks of how we do things here at Open Gov. But this is pulled from the Effect website, but you can see hey, this is how you make a tool, and then you add it to a toolkit, which is a collection of tools, and then you can register this toolkit with the language model. So for example, if you had a prompt that said, "Hey, generate some dad jokes about pirates," well, the agent has a tool that can help get a dad joke. So really, this is the building blocks of how we did tools and eventually skills, and it has paid off wonderfully for our organization. So we really recommend trying out this Effect AI package from Effect and trying out building out your own tools and skills.

1:01

SPEAKER_00

Developer velocity. So not only do we build agents for our customers, but we also use agents internally here in Open Gov. So we use a lot of Claude and Cursor, and it's been a real game changer for our team. It's funny because we're building tools and skills for customer-facing agents, and that has been great, but we're also building them internally as well to help accelerate our development workflows. So things like Claude, Cursor, Cloud agents—they really help accelerate how we read, write, review code, and ship. So it's been such an accelerant. So definitely wanted to mention that.

1:07

SPEAKER_00

Before we wrap up, that's it. Thanks so much for watching. You've made it to the end. Let's build agents that ship to production. Assist's big bet on effect. A little bit into our core agent loop, we're going to talk about the A2A protocol, evals and sandboxing. We're going to talk about how we manage long context. We're going to talk about monitoring observability, how we collect feedback and how we iterate on that feedback. We're going to lastly also talk about tools and skills and how at OpenGov we use AI not only externally that we serve to customers but also internally to improve our development workflows.

1:51

SPEAKER_00

Just a little bit about me before we go any further. My name is Gabe. I'm a software engineer here at OpenGov. I work on the AI agents team and I'm one of the folks that helped build OG Assist and some of the systems that you guys will be seeing today. So a little bit about OpenGov. OpenGov is a software company on a mission to power more effective and accountable government. So OpenGov sells ERP software that's things like budgeting, procurement, asset management and permitting and we were founded about 14 years ago and what's cool is we have this thing called OG Assist and OG Assist is this little

2:33

SPEAKER_00

button on the top of all of our products in the navigation bar. And what's cool is all of our product suites and product teams have built tools and skills in order to power this button. So for example, if I open up this, if I click this button and I open up OG Assist, it says, hey, I'm going to ask about rate codes, which is very specific to utility billing, the current product that I'm in. And you can see that inside of this kind of chat interface, I'm able to speak to an agent and the agent is able to make tool calls in order to look up information against data inside of that suite. So it's really cool to be

3:18

SPEAKER_00

able to kind of first party create these experiences through the capability that we've built called OG Assist. Okay, so just a quick story about how this all came to be. So a little while back, we saw that AI was really starting to take off and a principal spun up this new team called the AI agents team and asked me to join. And instantly I said yes. And OG Assist started to grow and we started to integrate OG Assist into all our products and not only our back end capabilities, but also our front end capabilities as well. So you'll see that one of the capabilities that we give the agent is it's able to

4:01

SPEAKER_00

see what's on the screen and see and take action on what's on the page. So you could see that I'm asking the agent here, hey, what's the screen? Can you maybe highlight some of the next steps that I could take? So you can see that the agent here is thinking it's saying, okay, what tools do I have available to use? And hey, let me go and highlight something that you could actually click on and tell you more about it. So just another capability of OG Assist and just a little short story about how this all came to be. So the big bet on effect. So I really wanted to include this slide because

4:41

SPEAKER_00

here on the agents team, we made a huge bet to bet on effect. And suffice to say, it's paid off in dividends. We write effect. So effect is this library for TypeScript. It's open source and it helps you write better TypeScript code. You know, it's got a lot of stuff baked in and like a schema similar to like Zod, if you've ever used that, it's also got things for error handling, for logging, for traces, for it's just got so much in there. It really helps write better code and structure your code better and helps with architecture, spinning up new services for and for us on the agents team, really helping

5:28

SPEAKER_00

design and build the core agent loop. So you'll see throughout this presentation sprinkled in how effect on our team has paid off in dividends. So we really love effect here at OpenGov and we encourage other folks to try it out and yeah, let's keep going. The effect native loop. So originally we were on Landgraf and that was fine until the team really started to scale and our use cases started to evolve. So we decided to move over to our own kind of effect native agent loop to have full regency over this agent loop such that if we have complex use cases or features that we need to build, we could kind of get in, we had full control of the agent loop and not

6:23

SPEAKER_00

only that but now we're fully on effect. So all the cool things you get with effect is now propagated throughout the entire agent loop like the tracing, structured concurrency, the logging, everything is more fine-grained control and it really allows us to really unlock the full potential having our own agent loop from the ground up. So another thing I wanted to mention is on the left side you'll see a code example. This is really the basics of the effect loop that we're using. We're using this thing called the effect AI package and in that package there's this thing called there's a chat and a language model. So with the chat you can instantiate like a chat for example

7:08

SPEAKER_00

and then you could stream text using that kind of stream text function. You could pass in a prompt and what's cool is with a language model under the hood of since we're kind of doing dependency injection we could pass in a different language model if we were to hot swap to another one for example. So really just having full control of our own agent loop just kind of gives us all the levers and it really just unlocks the full capabilities of the model and for the team as well to have full agency over this loop. Another thing I wanted to mention is the agent to agent protocol. So here on the agents team we've had

7:52

SPEAKER_00

a lot of success with this protocol. So this protocol being the protocol that Google created kind of an open protocol for agents to intercommunicate but we found this very useful for defining our agent routes like for example in the back end and our model and our schema to follow this kind of agent protocol. So we modeled so for example there's this thing called an agent card which you see here and it's got the name of the agent a description etc right and having this kind of rigorous protocol this rigorous spec really helped drive our development and drive alignment because you know all we had to do was align with this spec and

8:38

SPEAKER_00

follow this spec and we knew that this was kind of the contract that our front end and back end would both consume and and produce. So this I would say also has been very helpful for us and and what's really cool is A2A has a lot of extensions right so you could extend the protocol add in like metadata there's also A2UI so lots of fun stuff with A2A protocol but this is kind of what's worked for us so sharing that with with you folks. Feedback and evals. So here the quote is shipping is the start not the finish. So what we do here on the agents team is we have kind of multiple ways we do evals and collect feedback. Obviously you

9:31

SPEAKER_00

know we'll have folks call in or email us or just let us know and tell us but the main way is we have this thumbs up and thumbs down mechanism and here someone is able to tell us hey this this worked really well this was a great response or that wasn't a great response and that signal we take and we're able to iterate on and we can take it back and help improve you know the response in the future. We also have automated evals so in in the in RCI we we have evals that run against real completion so we could test a prompt against hey did it hit some tools did it do what it's supposed to do

10:09

SPEAKER_00

and that also helps with our accuracy. So those automated evals in conjunction with collecting feedback really help us improve our our our tools our skills our harness and and that's really how we're able to iterate so fast and so quickly. Humans in the loop. So this is a really cool feature we built where we deterministically interrupt the agent loop if there is a tool call approval required. So if an agent tries to make a tool call that it needs human approval for it'll show this UI and the human can click accept or reject so explicitly rejecting or explicitly accepting the action that the agent is

10:55

SPEAKER_00

trying to make and this ensures that you know we're building trust and also ensuring that you know we're being safe especially when the agent is trying to do a mutating operation and always always always making sure that humans are in the driver's seat. Sandboxing. So another thing that we worked on kind of similar to the safety slide we just saw was whenever an agent tries to execute code or tries to create files it does so in a sandbox so we gave our agents sandboxes such that it could spin up these sandboxes on demand and it could use those sandboxes to honestly write code execute code create files and it's kind of this safe

11:43

SPEAKER_00

ephemeral isolated space such that the agent can can can take action in there and not and we don't have to worry about any risk to you know our production systems and it's really cool because they also get uh tight uh teared down at the end so um in this example i said hey create a pdf uh for the folks of the ai engineer conference 2026 and um allow me to download it so i can share it with them and you can see that the agent created this really cool pdf inside of that sandbox so just really wanted to cover this sandbox feature and just give you a brief kind of overview of of sandboxing

12:25

SPEAKER_00

um long context so um inside of og assist we have hit many hurdles like with especially with legacy models now like with uh token limits or just way too much just completely overloaded with context um especially as conversations get longer so we found that having some sort of um rolling summarization was more effective than you know always stuffing in the latest and most recent uh messages uh rather just you know give like a running summary after n number of messages and uh maybe you only want the like n minus five most recent messages or n minus 10 most recent messages right and it may be that

13:19

SPEAKER_00

you're only talking about a specific topic now but you may want to refer to uh context earlier like 100 messages above then um it uh that's where kind of the the memory component comes in because when you have this rolling summary of a really long conversation then you could do recall over that uh summarization and um you know if you ask the agent hey remember that thing that we talked about then the agent within the thread will be like yeah i do know what you were talking about i have kind of this short little tidbit and it can you know follow up and and do more kind of with that uh and summary rolling summary in mind so

13:58

SPEAKER_00

that's kind of how we handled long context and memory and it's worked pretty well for us so uh just wanted to share a little bit about that and how we've uh solved the long context problem um ui on the fly so um in this example i said to the agent hey generate me a long essay but give me some examples about what the essay could be about so what's really cool is the agent had this primitive registered of this form and it was able to build out this form for me at runtime and give me some options of what i could choose from so it feels very personal and it feels very kind of in the moment that

14:35

SPEAKER_00

it's able to to give me these options just that runtime so um this is kind of just a short little um thing i wanted to include here about generative ui and how we are able to render uis on the fly you can't scale what you can't see so um this uh this kind of section is about tracing uh and observability really what's cool about effect is you kind of get tracing out of the box um you know when you use these effect functions they all get kind of tagged automatically with like these spans and kind of the span gets picked up and feeds into these traces so that you can kind of get these kind of

15:24

SPEAKER_00

drill downs of these function calls so here is an example of a trace from the effect team um i have it linked uh you could see that like hey when you hit this api it goes to this endpoint to this handler and uh you know etc and it takes and what's really cool is you could profile all your traces right so this takes a total of this many seconds and you can see where the bottleneck is here if there's a failure you can cross reference it across services so really really important especially working in agentic systems where we're we're integrating with other teams and other apis and other um

16:04

SPEAKER_00

other platform capabilities so uh what's cool with um effect is you get all this tracing out of the box and it really makes building this agentic experience debugging it and maintaining it just the breeze tools and skills so uh not only did we make a big bet on effect but we also made a big bet on tools and skills so we believe that tools and skills are really all you need and in this case you can see on the left we have we have this tool called get dad joke and this is kind of the effect way and kind of the building blocks of how we do things here at open gov but you know this is pulled from the effect website

16:47

SPEAKER_00

but you can see hey this is how you make a tool and then you add it to a toolkit which is a collection of tools and then you can register this this toolkit with the with the language model so for example if you had a prompt that said hey generate some dad jokes about pirates well guess what the agent has a tool uh that can help get a dad joke so um really this is the the building blocks of how we did tools uh and eventually skills and really just is has paid off wonderfully for for our organization so um really recommend trying out this effect ai package from uh effect and um just trying out building out your own tools and skills

17:37

SPEAKER_00

developer velocity so not only do we build agents for our customers but we also use agents um here internally uh in open gov so we use a lot of cloud and cursor uh and it's just really been a game changer for our team um it's funny because we're building tools and skills for customer facing uh agents and and that has been great but we're also building them internally as well to help accelerate our development workflows so things like claude cursor cloud agents they really help accelerate how we read write review code and and ship um so it's it's just been such an accelerant so definitely just wanted to mention that

18:18

SPEAKER_00

uh before we wrap up uh before we wrap up that's it thanks so much for watching you've made it to the end let's build agents that ship to production

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note