The Missing Layer After Launch - Raphael Kalandadze, Wandero AI
Description
We run a production system of agents for real customers. The team that keeps it healthy is also made of agents. Operating an agent product isn't like operating software. When our agent fails a customer — a dropped constraint, a stale price, a confident wrong answer — nothing crashes and no log lights up. The failure is in the conversation, not the stack trace. So we put agents on the operations: - One monitors production conversations and judges where the agent actually let a customer down — across thousands of live sessions, not a sampled few. - One watches logs and system health and traces real problems back into the code. - One writes and runs tests, because "green CI" means nothing for a non-deterministic agent. - One reviews every PR — human or agent-authored — against a single question: root cause, or just the symptom? Humans stay at the merge and approval boundaries. The agents do the watching, judging, testing, and drafting that no human team could keep up with at this volume. This talk is the honest version: what each operating agent actually checks, where we trust it and where we don't, what breaks, and why operating an agent system is becoming its own engineering discipline — done, increasingly, by agents. Speakers: - Raphael Kalandadze (Wandero AI): Co-founder and CTO of Wandero AI, an agent-native operating system for travel and hospitality, and co-founder of Tbilisi AI Lab, where we build the first Georgian large language model. X/Twitter: @RaphaelKalan LinkedIn: https://www.linkedin.com/in/rapael-kalandadze/ GitHub: https://github.com/RRaphaellRaphaelKalan
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Shipping an AI agent is just the beginning; production success requires a 'missing layer'—an automated feedback loop using meta-agents to monitor, diagnose, and fix agents in production at scale.
- Why it matters: Most agent talks focus on building and shipping, but the real work begins post-launch when you lose intuitive feel for system behavior across thousands of non-deterministic conversations.
- Best use: Essential blueprint for anyone building production agents; watch for the operational architecture (log monitoring agent, session analyzer, review agent, computer-use agent) and how Wandero closes the loop with automated PRs and health dashboards.
Executive Summary
Raphael Kalandadze from Wandero AI argues that the industry fixates on building and shipping agents, but the 'missing layer' is what happens after launch. Because agents are non-deterministic, have endless coverage, and can silently fail or succeed while delivering poor user outcomes, traditional monitoring (unit tests, regex checks, scripts) provides only a thin slice of visibility. The real problem is that operators lose intuitive 'feel' for their own system once it scales beyond a few demos.
His solution is a 'meta-harness': a system of specialized agents that continuously monitor, diagnose, and improve the production agent. Wandero runs four core agent types: (1) a log monitoring agent (every 15–60 min) that deep-dives traces, sends PRs or Slack alerts for critical issues; (2) a review agent with fresh context that scores and critiques PRs to prevent eager, low-quality fixes; (3) a session analyzer that scores every conversation, detects patterns, and provides a high-level health dashboard; and (4) a computer-use agent that simulates user journeys in the browser to catch UI issues logs can't reveal.
The architecture emphasizes closing the loop fast. The log monitoring + review agents generate 10x more PRs daily than Wandero's three-person team, with PRs ready in ~30 minutes. Humans remain in the loop for now (reviewing agent-generated PRs), but Raphael notes the trend is to remove that bottleneck once the loop is reliable. The session analyzer runs weekly to surface systemic patterns, entity analysis, sentiment trends, and AI-generated insights connecting dots across hundreds of sessions.
Key insight: operating an agent is itself an agent problem. These meta-agents need access to logs, trajectories, metrics, databases, UI, and codebase—the same tools a human operator would use—so they can reason about root causes vs. symptoms and avoid guessing. Raphael's central claim is that the internal feedback system matters as much as (or more than) the product itself, and everyone has access to the same frontier models, so competitive edge comes from the operational layer.
Key Takeaways
- Claim: Shipping is the moment when the real work begins; tight post-launch feedback loops are at least as important as the product itself, sometimes more. | Evidence: Wandero can build a product in days/weeks using latest models, but maintaining 'feel' for system health across hundreds of non-deterministic conversations requires continuous agent-driven monitoring and PRs. The loop enables daily product improvements. | Caveat: The speaker acknowledges the system may not be optimal and is still calibrating; humans are currently a bottleneck (reviewing 10x more agent PRs than they write themselves). | Implication: Ken should prioritize building post-launch feedback infrastructure early, not treat it as an afterthought. The operational layer is a moat when models commoditize. | Timestamp: 00:30
- Claim: Agents have a scary failure mode: they can struggle mid-task, recover via luck/workarounds, and finish without red alerts—but the hidden problem remains in the codebase as an early warning. | Evidence: Anthropic blog post noted agents love to mark features complete without checking if they worked. Wandero example: user asks for itinerary, agent books wrong service/misprices, task marked successful but user unhappy. | Caveat: No specific numbers given on how often silent failures occur or how quickly they compound. | Implication: Traditional success metrics (task completion, no errors) are insufficient; Ken needs semantic/outcome-based scoring and trajectory analysis to catch silent drift. | Timestamp: 03:45
- Claim: Operating an agent is itself an agent problem; logs are the source of truth, and machines are better than humans at exploring them, but diagnosis still requires deep reasoning. | Evidence: Log monitoring agent gets traces, trajectories, codebase access, runs every 15–60 min, deep-dives to differentiate real bugs from noise, symptoms from root causes, and sends PRs. Review agent with fresh context scores PRs, requires changes, or closes them. | Caveat: Requires calibration to trust the loop; speaker admits it's not yet optimal and humans still review PRs. | Implication: Ken should build a meta-agent harness early, give it full tooling access (logs, DB, UI, codebase), and accept that calibration is part of the process. Don't wait until scale forces it. | Timestamp: 06:00
- Claim: The log monitoring + review agent loop generates 10x more PRs than Wandero's three-person team daily, with PRs ready in ~30 minutes. | Evidence: Speaker states the agents send 10x more PRs than the three humans; PRs include description, metadata, diagrams (Mermaid, ASCII, HTML), and are ready in half an hour after detection. | Caveat: Humans are still the bottleneck reviewing these PRs; speaker may remove human loop in future but hasn't yet. No quantification of PR acceptance rate or false positive rate. | Implication: Ken should expect meta-agents to vastly outpace human output; design for that scale (good PR templates, review tooling). Human-in-loop can be temporary scaffolding, not permanent architecture. | Timestamp: 08:15
- Claim: Session analyzer provides high-level system health by scoring every conversation and detecting patterns across hundreds/thousands of sessions; runs weekly, not hourly. | Evidence: Dashboard shows: sessions analyzed, cost, avg scores, success rate, trends, AI insights (critical patterns with description, root cause, affected sessions, fix recommendations), sentiment, entities, tool call analytics. Built in-house because speaker knows what to look for. | Caveat: Token-intensive (runs many sub-agents per conversation). Randomized demo data shown, not real production screenshots. Frequency is weekly/bi-weekly, so slower than log monitoring loop. | Implication: Ken needs both fast (hourly) and slow (weekly) feedback loops. The session analyzer is for zoomed-out health monitoring and pattern detection, not immediate bug fixes. In-house build gives control over what to measure. | Timestamp: 10:00
- Claim: Computer-use agent simulates user perspective by opening browser, logging in, and checking UI; catches issues logs/code can't reveal, but is slow and token-heavy. | Evidence: Agent uses browser (initially slow), then built domain-specific skill for Wandero's site (much faster). Can open sessions, expand messages, check artifacts. Has access to trajectories and DB to diagnose problems detected in UI. | Caveat: Very token-intensive. Speaker built a custom skill to speed it up, implying the generic approach was impractically slow. | Implication: Ken should plan for user-perspective testing as a third angle (beyond logs and sessions), but optimize with domain-specific skills/tools rather than raw browser automation. Budget tokens accordingly. | Timestamp: 11:30
- Claim: The 'meta-harness' must have access to all tools a human operator would use (logs, trajectories, metrics, DB, UI, codebase) so agents reason from real data, not guesses. | Evidence: Example: computer-use agent detects UI issue, then queries DB and analyzes trajectories to confirm root cause. PRs depend on real problems, not hallucinated diagnosis. | Caveat: No discussion of security/access-control risks when giving agents DB/codebase write access. | Implication: Ken should treat the meta-harness as a first-class operator with full tooling; don't artificially limit access. Also design guardrails (review agent, human approval) to mitigate risk of bad automated changes. | Timestamp: 12:45
Detailed Brief
Why post-launch monitoring is the new hard problem
- Claims: Agents are not normal software: no predefined flow, endless coverage (like GPT/Codex), users do unpredictable things.; LLMs are non-deterministic; same input can yield different paths, slight input changes can trigger different trajectories.; Unit tests, regex, simulation scripts only slice the problem; production is where you learn what to test.; The deepest issue: you lose intuitive 'feel' for your own system after launch.
- Evidence: Harrison (LangChain) quote: 'You don't know what your agent will do until it is in production.'; Speaker tried unit tests, regex, customer conversation scripts—all helped but didn't prevent real-world surprises.; Agents can run hundreds/thousands of tool calls, use sub-agents, summarizations, terminal, third-party APIs—impossible to pre-test all paths.
- Caveats: No quantification of how often traditional tests catch issues vs. miss them.; Speaker acknowledges their approach may not be optimal.
- Implications: Ken should budget serious engineering time for post-launch ops, not just pre-launch dev.; Accept that production is the real test environment for agents; design for rapid feedback, not exhaustive pre-testing.
The log monitoring agent: fastest loop for local fixes
- Claims: Runs every 15–60 minutes, analyzes one-hour window of logs/traces.; Deep-dives to differentiate bugs from noise, symptoms from root causes.; Sends PRs for fixes or Slack alerts for critical issues.; Needs codebase access and reasoning capability.
- Evidence: Example PR: description, metadata, diagrams (Mermaid/ASCII/HTML) to give at-a-glance understanding.; Example Slack alert: heads-up for warnings or critical problems requiring immediate attention.; Speaker says PRs ready in ~30 minutes.
- Caveats: Only sees one-hour window, no high-level understanding.; Requires calibration to be reliable and trusted.; No data on false positive rate or how many PRs are rejected.
- Implications: Ken should implement hourly monitoring as the 'fast loop' for tactical fixes.; Invest in good PR templates and artifacts to make human review fast.; Plan for a review agent to filter low-quality fixes before human review.
The review agent: preventing eager, low-quality PRs
- Claims: Separate agent with fresh context (not biased by the problem).; Scores PRs, critiques from different angle, runs focused tests.; Can require changes, close PR directly, or approve for human review.; Acts as filter to reduce human bottleneck.
- Evidence: Speaker notes most agents are 'pretty eager to send the PR, they love to fix problems.'; Review agent feedback example: summary, diagrams, risks, edge cases, change requests.; After iterations, marks PR as ready or closes it.
- Caveats: No metrics on how often review agent rejects vs. approves.; Humans still review approved PRs (10x more agent PRs than human PRs daily).
- Implications: Ken should build a review agent as a quality gate, not just a log monitor.; Fresh context matters—don't reuse the same agent for diagnosis and review.; Expect agent output to vastly exceed human capacity; design review UX accordingly.
The session analyzer: high-level health and pattern detection
- Claims: Scores every conversation, runs many sub-agents, expands many tokens.; Detects patterns, connects dots, provides AI insights on critical issues.; Dashboard shows: health score, sessions analyzed, cost, avg scores, success rate, trends, sentiment, entities, tool call analytics.; AI insights include: description, why it matters, root cause, affected sessions, fix recommendations.; Runs weekly or bi-weekly, not hourly.
- Evidence: Speaker built in-house because he knows what he's looking for (vs. third-party tools).; Dashboard demo (randomized data from real conversations) shows detailed session ranking, score distribution, entity analysis.
- Caveats: Token-intensive.; Slower cadence (weekly) means not for urgent issues.; No discussion of how insights are prioritized or fed back into product roadmap.
- Implications: Ken needs a 'slow loop' for strategic pattern detection, separate from tactical bug fixes.; In-house build gives control over metrics; worth the engineering if off-the-shelf tools don't fit.; Session-level scoring is key to maintaining system 'feel' at scale.
The computer-use agent: user-perspective testing
- Claims: Opens browser, logs in, simulates customer journey.; Checks UI for issues logs/code can't reveal.; Has access to trajectories and DB to diagnose root cause when UI issue detected.; Initially slow; built domain-specific skill for speed.
- Evidence: Example: opens website, logs in, opens session, expands messages, checks artifacts.; Speaker notes it's 'pretty slow' and 'will spend a lot of tokens.'
- Caveats: Very token-intensive.; Requires domain-specific optimization to be practical.; No frequency mentioned (presumably less frequent than log monitoring).
- Implications: Ken should add user-perspective testing as third angle, but optimize with custom skills, not raw automation.; Useful for catching presentation/UI bugs that don't show in logs.; Budget tokens and accept slower cadence.
The meta-harness: full tooling access and closing the loop
- Claims: Meta-agents need logs, trajectories, metrics, DB, UI, codebase—same as human operators.; This ensures PRs and diagnoses depend on real problems, not guesses.; The internal feedback system is as important as the product, sometimes more.; Everyone has access to same models; competitive edge is the operational layer.
- Evidence: Computer-use agent example: detects UI issue, then queries DB and analyzes trajectories for root cause.; Speaker emphasizes 'operating an agent is an agent problem.'
- Caveats: No discussion of security, access control, or risk of agents making bad DB writes or code changes.; Human-in-loop currently required; future may remove it.
- Implications: Ken should design the meta-harness as a first-class platform with full access, not a bolt-on dashboard.; Security/guardrails (review agent, approval flows) are critical when giving agents write access.; The feedback loop is the moat; invest in it as a core competency, not an ops afterthought.
Notable Concepts & Terms
- The missing layer: Post-launch operational infrastructure for monitoring, diagnosing, and improving agents in production—the layer most teams neglect while focusing on building and shipping.
- Meta-harness: A system of specialized agents (log monitor, review agent, session analyzer, computer-use agent) that continuously watch, understand, and improve the production agent, with access to all operator tools (logs, DB, codebase, UI).
- Silent failure mode: When an agent struggles mid-task, recovers via luck/workarounds, and finishes without errors—hiding a latent bug that will cause future failures. Anthropic noted agents often mark tasks complete without verifying success.
- Fresh context review: Using a separate agent (not the one that diagnosed the problem) to review PRs from a different angle, avoiding bias and catching eager/low-quality fixes before human review.
- Session analyzer: High-level health monitoring system that scores every conversation, detects patterns across hundreds/thousands of sessions, and provides AI insights on systemic issues—runs weekly/bi-weekly, not hourly.
- Computer-use agent: Agent that simulates user perspective by navigating the UI in a browser, catching presentation issues logs/code can't reveal, then querying DB/trajectories for root cause diagnosis.
- Closing the loop: Establishing fast feedback from production to diagnosis to fix (PR) to deployment, ideally automated. The speaker argues the loop is more important than the initial product when models commoditize.
Operator Notes / Why Ken Should Care
- For Ken's agent systems: This is the operational playbook for production agents. The meta-harness architecture (log monitor, review agent, session analyzer, computer-use agent) is a concrete pattern to copy. Prioritize the fast loop (hourly log monitoring) first, then add session analyzer for weekly health checks.
- For AI ops: Wandero's agents generate 10x more PRs than humans daily. Ken should design for this scale now—good PR templates, review UX, approval workflows. Human-in-loop can be temporary; plan to remove it once loop is calibrated.
- For content/business strategy: The 'missing layer' framing is the talk's hook. It's a gap in the market (most tools focus on building, not operating) and a moat (everyone has same models, but not same feedback loops). Ken could position agent ops tooling or consulting around this.
- For investing: Companies building the operational layer (observability, diagnostics, auto-remediation for agents) are addressing the real bottleneck post-launch. Look for teams that treat agent ops as first-class product, not afterthought.
- For GTM: The insight that 'operating an agent is an agent problem' justifies selling meta-agents or agent ops platforms to teams struggling post-launch. The pain is losing 'feel' for the system at scale; the solution is automated feedback loops.
- For workflow: Ken should implement this for his own agents—start with hourly log monitoring and weekly session analysis. Budget tokens for the session analyzer (it's expensive but provides the 'pulse' at scale). Build domain-specific skills for computer-use testing to avoid slowness.
Watch Map
- 00:00: Intro: the missing layer thesis—shipping is when real work begins
- 01:30: Why this is hard: agents aren't normal software, non-deterministic, endless coverage
- 03:45: Scary failure mode: silent struggles that recover via luck, hiding bugs
- 06:00: Log monitoring agent: fastest loop, hourly, sends PRs in 30 min
- 08:15: Review agent: fresh context, scores PRs, filters low-quality fixes, 10x human output
- 10:00: Session analyzer: high-level health dashboard, scores all conversations, weekly cadence
- 11:30: Computer-use agent: user-perspective testing, slow/token-heavy, needs domain skills
- 12:45: Meta-harness: full tooling access (logs, DB, UI, codebase), competitive moat
- 14:00: Closing: the loop is as important as the product; production is where you learn
Source/Metadata
- Title: The Missing Layer After Launch - Raphael Kalandadze, Wandero AI
- Transcript words: 6624
- Duration seconds: 1173
- Timestamp note: Timestamps estimated from 1173-second duration; transcript did not include explicit chapter markers, so timestamps are approximated based on content flow.
Transcript
All right. So we built an agent, you launched it, everything works pretty well in the demo, everyone is happy. But now let me ask you a few simple questions. So how do you know if it's actually working out there? How do you watch across hundreds or thousands of real conversations every day? How do you feel or understand the health of the system? How do you make it better? How do you find the holes that you don't know are there yet? And that's the thing, right? So most of the talks about agents focus on the moment when you ship, so we build it, it works the end. But I think shipping is the moment when the real work begins. And somehow only a few people are talking about that, and I'm calling it the missing layer. So let's dive into it. So that's the world that we are living in. You can create the whole product, you can create a whole startup in a couple of days, in a couple of weeks, you can write hundreds of thousands of lines of code, you can spend a lot of tokens. And to be honest, the easiest part today with the help of the latest models. But I think shipping is the moment when the real work begins, because you need to close the loop as soon as possible. So after your launch, you need to have some control and understanding of the system. And from my experience, the loop is at least as important as the product itself, sometimes even more, because the tight feedback is the one that helps you to make the product better every single day. And that's the missing layer. And that's what the rest of this talk is all about. So what happens after your launch? And this is not something surprising. We had the same questions in classical old software. You need to monitor what is happening, you need to understand how it behaves, you need to have some logs, detect the problems and fix them. And for agentic systems, each one of those is even harder. And sometimes they turn into something entirely new. So let's talk about why this is hard and why this is hard now. So the agent is not normal software, right? You don't have a few features, several buttons, you don't have a predefined flow that you can test before you go live. And the coverage is endless. So think about GPT Code or Codex, they can do a giant range of stuff, whatever the user needs and most agents do the same, right? So you give the instructions and they can handle it. And you cannot write all the conversations in advance. So this leads to the deepest part of the problem, the part that keeps me up all night, which is you lose the feel for your own system. So after you build the product, you need to have some kind of understanding—does it get better or worse? So you need to monitor, understand what is happening. And the problem is that the normal safety nets don't save you here. And Bill and me, we tried a few things. We built unit tests, we have some regex, some rule-based checks, we even created some scripts to simulate customer conversations. And yes, it helps in some ways. But at the end of the day, it is only one slice of the whole problem. Because customers always do something different. You cannot write it all down, there are too many, and they are all different. And Harrison put it the same way. So you don't know what your agent will do until it is in production. So why is this a new problem? So as you know, LMs are not deterministic. The same input can have a different path, even slight modification of the input can call a different trajectory and coverage is endless—you cannot list it all down, you cannot pre-test until you go into production. The second and the scariest one is the failure mode itself. So let's think about what happens when the agent is running for a long time. The agent is struggling in the middle of the task. It has some problems, but it got lucky. It recovered, it finds some workarounds, tries some other tool calls or whatever. And you did not get any red alerts, any problems on the dashboard, everything looks fine. But you know, this is an early warning for you. This is a problem that is hidden and that lives in your code base. And you need to fix it as soon as possible. Because if you're talking about reliable agents, each will depend on their luck, right? And also, as Anthropic mentioned in a blog post, sometimes the agent loves to mark the feature as complete without checking if it actually worked. The next one is the tool calling, right? So the tool surface is huge. Long-running tasks need hundreds of tools, even thousands of tools. Sometimes it has several summarizations in the middle, it uses some sub-agent, writes a lot of code, it uses a terminal for sure, it calls some other company services, third-party libraries, and they work in different ways all the time. So actually, you don't know what you're looking for. So sometimes finish doesn't mean it is helpful for the user. Maybe there was any, there wasn't any problem. So it finished, everything looks good. The answer was successful. But what happens, for example, for our use case, sometimes a user asked to build an itinerary. Agent runs the flow, it builds a trip, but it booked a different service, it made a lot of mistakes in calculating the price, or the user is not happy. So technically, it's successful, but still failing the task. And as I mentioned, the unit tests don't save you here. And production is the place when you learn what you need to test in the first place. All right, so what happens after you launch the product, right? So you need to monitor, you need to understand and improve the system. And what's the source of truth, right? This is logs. Everyone has logs, everyone loves logs. You have structured information. And you know, those machines are the best to understand and explore the logs, much better than any human, much faster. They can write some scripts to filter the giant walls. And this is the most obvious way that you can hand it to the agent. But as soon as you start working on that, maybe you build an agent or skill or whatever, you quickly understand this is not as easy as we imagine, because it needs a lot of reasoning. You need to understand the problem to differentiate if it's a real bug or noise, a tip time in a trace or trajectory, understand if it's a symptom or a root cause. So actually, you'll find out that operating an agent itself is an agent problem. So this is a loop end to end. So we have traces, you have trajectories, you give the agent access to the code base, it diagnoses the problem, understands what was happening and sends the PR. And then you can have a skill or sub-agent or wherever that controls the PR. Because you know, most agents are pretty eager to send the PR, they love to fix problems. We prefer to have a separate agent, which has fresh context. It tried to check the PR from a different angle, different view, run the focus test. And it is not biased by the problem itself. It tries to criticize, score the PR. And sometimes it So this is a loop end to end. So we have traces, you have trajectories, you give the agent access to the code base, it diagnoses the problem, understands what was happening and sends the PR. And then you can have a skill or sub agent that controls the PR. Because most agents are pretty eager to send the PR, they love to fix the problems. We prefer to have a separate agent which has fresh context, it tries to check the PR on a different angle, on a different view, run the focused test. And it is not biased by the problem itself. It tries to criticize, score the PR. And sometimes it requires changes, maybe it closes the PR directly and helps us to filter those problems. And then you have the human in the loop, sometimes you don't. And we can talk about this later. But this is the most obvious thing that you can hand to the agent, which seems pretty obvious for most people. And a lot of teams use the same practice. But I think people don't appreciate how important it is. And you need to spend some time on that. You need to calibrate, you need to make it reliable to trust the loop. And this is the fastest loop that we ever had. So after you build this simple system, you already have the feel and understanding of how it behaves, what is happening, you detect the problems, some local fixes as soon as possible. And you can have the PR ready in half an hour. And you can easily understand what is happening. So let me walk you through how I'm handling this. Maybe this is not optimal. But I think this will help you to get some points. So actually, we have two main flows. The first one is the fastest loop that helps you to detect and fix the problems as soon as possible. And another one is more zoomed out, that helps you to have a high level understanding. And it helps you to have a hand on the pulse. So the first one is a log monitoring agent, the one that I already mentioned. So you have trajectories, you have logs, you have access to the code base, and it runs every hour or every 15 minutes. And it tries to understand the problem with deep dives in the logs, understand that the user ended up in the stack and send the PR or sometimes send a Slack alert. And this is pretty important because sometimes the problem is so critical that we need to fix it as soon as possible. And it works pretty well. So this is one example of how the PR can look like. So you have a nice description, a short explanation of what is happening, you have some metadata, you have nice diagrams, mermaid or ASCII tables, maybe you will have some HTML, as far as it helps you to give you a glance of the problem and quickly understand what this PR is for. So this is an example of how a Slack notification can look like. It helps you to quickly give you feedback, detect if there are some critical problems. Sometimes it's just a heads up so you know there are some warnings, there are some problems, so you need to check them when you have some time. And as I mentioned, the review agent is pretty critical because it tries to check the problem from a different angle. It always tries to criticize the problem, score the problem, and as soon as it's ready, it sends the PR to the human. And sometimes people talk about whether we need to have the human in the loop because the human is still the bottleneck in this case. In our case, the PR agent and the review agent send 10 times more PRs than three of us every day. So I need to have some clean system for how we're going to handle this to not be a bottleneck in this system. But from my experience, as I mentioned, this PR and the review agent helps me to have a nice description, some artifacts that help me to quickly understand the problem. Maybe I spend a few minutes to at least understand what is happening. So I think at this time it's okay, but maybe we will remove it in the future. And people are talking about this. So you don't need to be a bottleneck in this problem. Some of them prefer to remove the humans from this loop, but the trend is that you need to close the loop first. So let's solve the problem when you are the bottleneck and then you can remove yourself pretty easily. So this is an example of how the review agent feedback can look like. As you know, it requires some changes, you have some summary, you have some diagrams, explanation, where there are some risks, some edge cases, and sometimes after a few iterations or they go in the loop where the key is ready. So you don't know, you don't need to have any more changes. So after that, the human will jump into the loop. For the previous agent, as I mentioned, that is especially good when you want to fix some local problems. For our case, it runs every hour. It checks only a one hour window of the logs and it doesn't have any high level understanding of the problem. So we found out that we need to have another system that helps us to get a more high level explanation, some kind of a health of the system. And visibility is the easiest piece. Before the agent, it was impossible. So you cannot deep dive or summarize hundreds of conversations. But right now we can have a system that helps you to score every conversation, understand what was happening and give you a high level zoom out of the system. So this is the session analyzer. And the main goal is to just give me the score of the system health, right? So we try to check over every conversation, run through a lot of sub agents, expand a lot of tokens. But it helps us to detect some patterns, connect the dots and try to understand some high level patterns. And it also detects some classical problems. What is the cause, how many tool calls are used, how many sub agents, how many summaries happened or wherever, but also it gives some AI insights, some entities. So we have an understanding of the problem. As I mentioned, one of the main problems is that you lose control after you launch the production. So you need to have a system that helps you to control it, at least to have some understanding of what is happening, how it looks, is it healthy or not. So actually, let me show you one example of how it looks like. So actually we built it ourselves. So there are a lot of other companies and tools that provide the same kind of system, but I prefer to build it myself because I How many summaries happened or wherever, but also it gives some AI insights, some entities. So we have an understanding of the problem. Right. As I mentioned, one of the main problems is that you lose control after you launch the production. So you need to have a system that helps you to control. It's at least to have some understanding of what is happening, how it looks like, is it healthy or not. So actually, let me show you one example of how it looks like. So actually we built it ourselves. So there are a lot of other companies and tools that provide the same kind of system, but I prefer to build it myself because I know what I'm interested in, what I'm looking for. So as you can see, we have the health of the system. We have number of sessions that were analyzed, the cost. We have the average scores and success rate. We have some trends. We have the AI insights, and this is the most important part. When the agent tried to connect the thoughts, find the patterns. If there are some critical ones, you have a description for each of them. What is this? Why it matters? What are the root costs? All the sessions affected and some recommendations to fix. So right now, this is just the randomized data, but it comes from real conversations. So we have a score distribution, each company, number of sessions or the cost. You have the sentiment analysis, some entities. We have analytics about tool cost or the success rate, how many were rejected and why. And you have detailed analysis of each session. So it ranks each one of them by scores. It is how many times was running, how many messages, how many tool calls, how many summaries, and you have a detailed explanation for each of them. What was the problem? Is it something that you need to understand. And the limit helps a lot and it helps you to check and watch across hundreds of conversations. So the main goal of this dashboard and the system is not to fix some specific problems or bugs. This is more on a high level that helps you to watch across hundreds or thousands of real conversations. So you can run it once a week, two times a week or something like that. All right. So the next one is that the problem of the previous agent that I mentioned is like the first one helps you to detect the specific problems. The second one is to give you a high level understanding, but both of them work from an angle of the logs or a code or a session itself, but you need to have some kind of user perspective, right? So that's why we have some computer use agent that helps you to open the browser, to log in and try to simulate the customer itself because sometimes you have some problems in the UI. You need to check if everything looks good. And yeah, sometimes code and logs don't help you in this way. So this is an example of how it looks like. So you use codex, sometimes, so it used the browser itself. Actually this is pretty slow and we tried to build a specific skill that knows our website, our domain, how it looks like, and that is much faster. And you can open the website, log in, open the session, expand the messages, check what is happening, and also check how it looks like, what are some artifacts and yeah, it was pretty good, but yeah, you need to know that it will spend a lot of tokens. And also, for all of those problems, you need to remember that you need to give access to all kinds of tools that are needed. So you need to have logs, trajectories, metrics, database, UI, so as humans need, right? So you need to give all context, all possibilities to understand what is happening. For example, for the computer use agents, when they detect some problems, it should be able to analyze the trajectories, check the database to understand what happened. And that's why I'm calling it the meta harness, the whole system, when everything is connected and the PR or the answer from the agents depends on the real problem and they're not guessing what is happening. So the most important is not the model alone and you need to build the agent or a system or a harness around it, which watches itself, understands, improves and helps you at least to monitor what is happening, but also in an ideal case to close the loop, send automatic PR notifications wherever and help you to speed up the process. So shipping is the easiest part today. If you want to build a production agent, you need to close the loop first because somehow people are talking about how important it is, what happens after you launch. So everyone can have the same model. Everyone can have the same agent or harness, but you need to have some internal system that helps you in this process, detect the problems, and give you a sense of what is happening and helps to make the product better. Thank you. So you can create the whole product, you can create a whole startup in a couple of days, in a couple of weeks, you can write hundreds of thousands of lines of code, you can spend a lot of tokens. And to be honest, the easiest part today with the help of the latest models. But I think the shipping is the moment when the real work begins, because you need to close the loop as soon as possible. So after your launch, you need to have some control and understanding of the system. And from my experience, the loop is at least as important as the product itself, sometimes even more, because the tight feedback is the one that helps you to make product better every single day. And that's the missing layer. And that's what the rest of this talk is all about. So what happens after your launch? And this is not something surprising. We had the same questions in the classical old software, you need to monitor what is happening, you need to understand how it behaves, you need to have some logs, detect the problems and fix them. And for agentic systems, each one of those are even harder. And sometimes they turn into something generally new. So let's talk about why this is hard and why this is hard now. So the agent is on a normal software, right? You don't have a few features, several buttons, you don't have a predefined flow that you can test before you go to the live. And the coverage is endless. So think about like plot code or codex, they can do a giant range of stuff, wherever the user needs and most of the agent to the same, right? So you give the instructions and they can handle it. And you cannot write all the conversations in advance. So this leads to the deepest part of the problem, the part that keeps me up all night, which is you lose the feel for your own system. So after you build the product, you need to have some kind of understanding, does it get better or worse? So you need to monitor, understand what is happening. And the problem is that the normal safety nets don't save you here. And Bill and me, we try a few stuff, we build a unit test, we have some reg X, some rule based checks, we even create some scripts to simulate the customer conversation. And yeah, it helps in some ways. But at the end of the day, it is like only one slice of the whole problem. Because customers always do something different. You cannot write it all down, they are too many, and they are all different. And the Harrison put it in the same way. So you don't know what your agent will do until it is in the production. So why this is a new problem? So as you know, LMs are not deterministic, deterministic, the same input can have a different path, even slight modification of the input can call a different trajectory and coverage endless, you cannot list it all down, you cannot pre test until you go into production. The second and the scariest one is the failure height itself. So let's think about what happens when the agent is running for a long time, the agent is struggling in the middle of the task. It has some problems, but he was lucky, he was recovered, he finds some workarounds, try some other tool calls or whatever. And you did not get any red alerts, any problems on the dashboard, everything looks fine. But you know, this is an early warning for you. This is a problem that is hidden and that lives in your code base. And you need to fix it as soon as possible. Because if you're talking about reliable agents, each will depend on their luck, right? And also, as an anthropic method in the blog post, sometimes the agent loves to make mark the feature as complete without checking if they actually worked. The next one is the tool calling, right? So the tool surface is huge. So long running task needs hundreds of tools, even thousands of tools. Sometimes it has several summarizations in the middle, it uses some subagent, writes a lot of code, it uses a terminal for sure, it calls some other company services, third party libraries, and uh, they work for a different way all the time. So actually, you don't know what you're looking for. So sometimes finish doesn't mean it is helpful for the user, maybe there was any, there wasn't any problem. So it ain't finished, everything looks good. The answer was successful. But what happens, like, for example, for our use case, sometimes a user asked to build the itinerary agent run the flow, it builds a trip, but it booked a different service, it made a lot of mistakes in calculating the price, or user is not happy. So technically, it's successful, but still failing the task. And as I mentioned, the unit tests don't save you here. And production is the place when you learn what you need to, what you need to test on the first place. All right, so what happens after you launch the product, right? So you need to monitor, you need to understand and improve the system. And what's the source of the truth, right? This is a logs, everyone has the logs, everyone loves the logs, you have structured information. And, you know, those machines are the best to understand and explore the logs, much better than any human much faster, they can write some scripts to filter the giant of walls. And this is the most obvious way that you can hand it to the agent. But as soon as you start working on that, maybe you build an agent or skill or whatever, you quickly understand this is not as easy as we imagine, because it needs a lot of reasoning, you need to understand the problem to differentiate if it's a real bug or a noise to tip time in a trace or trajectory, understand if it's a symptoms or a root cause. So actually, you'll find out that the operating an agent itself is an agent problems. So this is a loop like end to end. So we have a traces, you have trajectories, you give the agent to access the code base, it diagnose the problem, understand what was happening and send the PR. And then you can have a skill or sub agent or or wherever that controls the PR. Because you know, most of the agents are pretty eager to send the PR, they love to fix the problems, we prefer to have a separate agent, which has like, fresh context, it tried to check the PR on a different angle on different view, run the focus test. And it is not biased of the problem itself, it tries to criticize, score the PR. And sometimes it requires the changes, maybe it close the PR directly and help us to to filter those problems. And uh, then you have the human in the loop, sometimes you don't. And we can talk about this later. But you know, this is the most obvious thing that you can hand it to the agent, which is which seems like a pretty obvious for most of the people. And a lot of teams use the same practice. But I think people don't appreciate how important it is. And you need to spend some time on that you need to calibrate, you need to make it reliable to trust the loop. And this is the the fastest loop that we ever had. So you know, after you build this simple, simple system, you already have the feel and understanding how it behaves, what is happening, you detect the problems, some local fixes as soon as possible. And you can have the PR in PR ready in half an hour. And you can easily understand what is happening. So let me walk you through uh, how I'm handling this. Maybe this is not optimal. But I think this will help you to to get some point. So actually, we have a two main flow. The first one is that the fastest loop that helps you to to detect and uh, fix the problems as soon as possible. And another one is like more on a zoom out that how that helps you to have a high level understanding. And it helps you to have hand on a pulse. So the first one is a log monitoring agent, the one that I already mentioned. So you have trajectories, you have logs, you have the access to a code base, and it runs every hour or every few 15 minutes. And it tries to understand the problem with deep ties in there in the logs, understand that the user end up the stack and send the PR or sometimes send the Slack alert. And this is this is pretty important because sometimes the problem is uh, uh, so critical. So we need to fix them as soon as possible. And uh, yeah, it works pretty well. So this is the one example how the PR can look like. So you have a nice description, a short explanation what is happening, you have some metadata, you have a nice diagrams, mermaid or ascii tables, maybe you will have some HTML, as far as helps you to give you a glance of the problem and quickly um, understand what is this PR for. So this is the example how a stack notification can look like. It helps you to quickly give you a feedback, detect if there are some critical problems. Sometimes it is just like heads up so you know there are some warnings, there are some problems, so you need to check them when you have some time. And as I mentioned, the review agent is uh, pretty critical because it tries to check the problem in a different angle. It always tries to criticize the problem, score the problem, and uh, as soon as it's ready, it sends the PR to the human. And sometimes people talk about that if we need to have the human in the loop because you know still human in the bottleneck in this case. Uh, in our case, uh, the, the, the PR agent and the review agent send 10 times more PR than three of us uh, every day. So I need to have some clean system how we're gonna handle this to not be a bottleneck in this system. But from my experience, as I mentioned, this PR and the review agent helps me to have a nice description, uh, some artifacts that helps me to quickly understand the problem. Uh, man, maybe I spend a few minutes to at least understand what is happening. So I think at this time it's okay, but maybe we will remove it in the future. And yeah, people are talking about that. So you need to, you need, you don't need to be a bottleneck in this problem. Some of them prefer, uh, to remove the humans in this loop, but you know, the trend is that you need to close the loop first. So let's make the problem when you are the bottleneck and then you can remove yourself pretty easily, I think. So this is some examples how the review agent feedback can look like. As you know, uh, it requires some changes, you have some summary, you have some diagrams, explanation, where are there some risks, some edge cases, and sometimes, uh, after a few iterations, or they go in the loop, uh, where the key, where the key is ready. So we don't know, you don't need to, uh, have any more changes. So after that, the human will, uh, jump, uh, into the loop. For the previous agent, as I mentioned, that is especially good when you want to fix some local problems, uh, for our case, it runs every hour. It checks only a one hour window of the logs and, uh, it doesn't have any, uh, high level understanding of the problem. So we found, found out that we need to have another system that helps us to, to get, uh, to get, uh, um, like more than a high level explanation, some kind of a health of the system. And, you know, visibility is the easiest piece. Um, before the agent, it was impossible. So you cannot, uh, deep dive or summarize hundreds of conversations. But right now we can have a system that helps you, uh, to, to score every conversations, understand what was happening and give you a high level zoom out of the system. So this is the session analyzer. And the main goal is to just give me the score of the health system, right? So we try to check over every conversation, run through a lot of sub agents, expand a lot of tokens. Uh, but yeah, it helps us to detect some patterns connecting the dots and, uh, try to understand some high level patterns. And, uh, yeah, it also detects some, uh, um, uh, like classical problems. What is the cause, uh, how many tool costs are used, how many sub agents, uh, how many summary happened or wherever, but also it gives some AI insights, some entities. So we have an understanding of the problem. So, right. As I mentioned, one of the main problem is that you lose the control after you launch the production. So you need to have the system that helps you to control. It's at least to have some understanding or is happening, how it looks like, is healthy or what, or not. So, um, actually, let me, let me show you, uh, one of the example, how, how it looks like. So actually we build it ourselves. So there are a lot of other, um, companies and tools that provide the same kind of system, but I prefer to build it myself because I know what I'm interested for, what I'm looking for. So as you can see, we have the health of the system. We have number of sessions that was analyzed, the cost. We have the average scores and success rate. We have some trends. We have the AI insights, and this is the most important part. When the AI, when the agent tried to connect the thoughts, find the patterns. If there are some critical ones, you have a description for each of them. What is this? Why it matters? What are the root costs? All the sessions, uh, affected and some recommendation fixed. So right now, this is just, uh, the randomized data, but it comes from a, from a real conversations. So we have a score distribution, each company, um, number of sessions or the cost. You have, uh, the sentiment analysis, some entities. We have analytics about a tool cost or the success rate, how many rejected and why. And you have, uh, detailed, uh, analysis of each session. So it ranks each one of them scores. It is, um, how many times was running, how many messages, how many tool calls, how many summary, and you have a detailed explanation for each of them. What was a problem? It is something that you need to understand. And, uh, the limit helps a lot and it helps you to check and watch across hundreds of, uh, conversations. So the, the main goal of this dashboard and the system is not to fix some, uh, um, uh, specific problems or bug. This is more on a high level that helps you to watch across hundreds or thousands of real conversations. So you can run it once in a week, two times in a week or something like that. All right. So the next one is that, um, so the problem of the previous, uh, agent that I mentioned is like the first one helps you to detect the specific problems. The second one is to give you a high level understanding, but both of them work from a angle of the logs or a, uh, code or a session itself, but you need to have some kind of user perspective, right? So that's why I we have some computer use agent that helps you to, uh, to open the browser, uh, to log in and try to simulate the customer itself because sometimes you have some problems in the UI. You need to, um, check if everything looks good. Uh, and, uh, yeah, sometimes code and logs don't help you, uh, in this way. So this is, uh, example, how it looks like. So you use codex, uh, sometimes, so it used, uh, uh, the browser itself. Uh, actually this is pretty slow, uh, and we tried to build the specific skill that know our, uh, website, our dome, how it looks like, and that is much faster. And you can, uh, open the website, log in, open the session, extend the messages, check what is happening, uh, and also check, uh, how it looks like, what are the, some artifacts and, uh, yeah, it was pretty good, but yeah, you, you need to know that it, it, it will spend a lot of, a lot of tokens. And also, uh, for all of those problems, you need to remember that, uh, you need to give access to all kinds of tools, uh, that is needed. So you need to have a log, projectories, a metrics, database, UI, uh, so as the humans need, right? So you need to give all context, all possibilities to understand what is happening. For example, for the computer use agents, when detect some problems, it, uh, should be able to, uh, analyze the trajectories, check the, uh, database to understand what happened. And, you know, that's why I'm calling it the meta harness, the whole system, when everything is connected and, uh, the PR or the answer from the agents are, uh, depend on the, depend on the real problem and they're not guessing or is happening. So the most, uh, is not the model alone and you need to build, uh, the agent or a system or a harness around it, uh, which, uh, which, uh, watch itself, understand, improve and help you at least, uh, at least to monitor what is happening, but also any ideal case to close the loop, send, uh, automatic PR notifications wherever and, uh, help you to speed up the process. So shipping is the easiest part today. Uh, if you want to, uh, if you want to build the production agent, you need to close the loop first because somehow people are talking about how important it is, uh, what happens after you launch. So everyone can have the same model. Everyone can have the same agent or harness, but you need to have some internal system that helps you in this process, uh, detect the problems, uh, and, uh, give you a sense of what is happening and helps to make the product better. Thank you.