Open Reader

The Agentic AI Engineer - Benedikt Sanftl, Mutagent

completed 34:49 Jun 29, 2026 Watch on YouTube

Current Status

completed

Video ID

pSto5YaNGUo

RAG / Chat

Enabled
The Agentic AI Engineer - Benedikt Sanftl, Mutagent
Description

In this video we introduce the concept of the agentic ai engineer. similar to coding agent loops for agents we build a system that build AI Agents in an Eval-Driven Developement Loop. The Agentic AI Enginner is a collection of a multi-agent team, steared by an orchestrator and combines, spec, build, evaluate, diagnose, monitor, optimse. We round up the talk with a live demo from one of our agents in research preview. Speakers: - Benedikt Sanftl (Mutagent): Bene is the CEO and Co-Founder of Mutagent. The platform for Agentic AI Engineering. LinkedIn: https://www.linkedin.com/in/benedikt-sanftl-294a6039a/

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully if building production AI agents at scale; Skim if exploring agent development tooling
  • Core thesis: Apply the same agentic loop paradigm used to build software to building AI agents themselves—automate the spec→build→eval→ship→monitor→diagnose→optimize cycle with specialized agents to achieve production reliability at scale
  • Why it matters: Manual iteration on AI agents doesn't scale beyond a handful of features. An "agentic AI engineer" (orchestrated eval, diagnostics, and optimization agents) compresses feedback loops and handles the human bottleneck in agent lifecycle management
  • Best use: Teams deploying multiple production agents who face bottlenecks in eval creation, trace analysis, and iterative debugging

Executive Summary

Benedikt Sanftl (CEO) and Burak (CTO) of Mutagent argue that building AI agents at scale requires meta-automation: using agents to build, evaluate, diagnose, and optimize other agents. Their talk centers on two loops: an offline loop (iterative development with evals and optimization before production) and an online loop (monitoring live traces, diagnosing failures, feeding insights back into the offline loop). Today, human review of traces and manual iteration creates a throughput ceiling—"the human becomes the bottleneck"—especially when organizations plan to deploy hundreds of agents.

Mutagent's solution is an orchestrated suite of specialist agents (evaluator, diagnostics, optimizer) that automate the tedious, time-intensive parts of the agent lifecycle. The workflow begins with a spec-driven design (analogous to software PRDs) that is decoupled from the target harness (Langchain, Hermes, etc.), enabling framework portability. A coding agent then builds the initial implementation. An evaluator agent constructs an eval suite (metrics, criteria, datasets) that evolves through production feedback, functioning like "unit tests for agents." Once deployed, a diagnostics agent ingests traces from observability platforms, segments them intelligently (to avoid reading millions of traces), performs root-cause analysis, clusters failure modes, and proposes remedies. These remedies feed back into the offline loop as new evals and optimizations, creating a continuous improvement cycle.

The product is in research preview and runs locally or in customer environments. It integrates with existing observability tools (Langfuse, custom JSON-L exports) and version control (GitHub PRs). The diagnostics agent generates HTML reports with failure modes, recursive "why" chains, assumptions (to catch incorrect inferences when code isn't available), and multi-choice remedies that can be fed directly to a coding agent. The evaluator emphasizes binary, actionable evals over score-based LLM-as-judge to reduce variance and provide clear call-to-action on failures.

Mutagent's pitch: stop manual debugging, let agents handle the tedious work, and scale agent deployment without scaling headcount. The framework is designed for production reliability, not exploratory prototyping.

Key Takeaways

  • Claim: Manual agent iteration doesn't scale beyond ~5-10 features | Evidence: Human review of traces and eval creation becomes the bottleneck; each cycle can take hours or days | Caveat: Only a problem if you're deploying many agents or handling high trace volume | Implication: Organizations planning large-scale agent rollouts need automation at the lifecycle management layer, not just at the task execution layer
  • Claim: Spec-driven development for agents enables harness portability and future-proofing | Evidence: The spec (context, tools, jobs-to-be-done, constraints) is separated from implementation; coding agents can retarget to Langchain, Hermes, DeepAgent, etc. | Caveat: Assumes the spec captures enough detail and doesn't become stale; framework-specific optimizations may be lost in translation | Implication: Teams won't be locked into a single agent framework as the ecosystem rapidly evolves; reduces migration risk
  • Claim: Eval-driven development for agents mirrors test-driven development for software | Evidence: Evals provide termination conditions and actionable feedback; the suite grows through production failures and edge case discovery | Caveat: Requires upfront investment in eval design and LLM-as-judge variance management; score-based evals without rubrics are too vague | Implication: Binary evals (pass/fail criteria) are preferred over numeric scores because they tell you what to fix, not just that something is broken
  • Claim: Diagnosis agents reduce trace analysis cost and time via intelligent segmentation | Evidence: Reading all traces at scale costs more than execution; multi-tier filtering identifies representative samples and learned indicators (tool call sequences, content patterns) | Caveat: Depends on having structured trace data and observability integrations; "assumptions block" in reports hints at potential misdiagnosis when code access is limited | Implication: Trace analysis becomes economically feasible at scale; failure mode clustering and root-cause "why chains" replace manual log spelunking
  • Claim: The online→offline loop compounds agent quality over time | Evidence: Production failures generate new evals and dataset cases; agents "learn" failure modes and build up historical indicators | Caveat: Garbage-in-garbage-out if production data isn't representative or if user feedback is poor; requires continuous monitoring triggers | Implication: Agents improve in the wild, not just in dev; quality becomes a function of deployment time and trace volume, similar to reinforcement learning from human feedback
  • Claim: Mutagent's orchestrator handles dispatch and workflow sequencing for specialist agents | Evidence: Demo showed slash commands (/diagnose) in Claude Code triggering multi-stage workflows with HTML artifact generation | Caveat: Runs in customer environments (local or cloud); no managed service yet; integrations require setup | Implication: Teams can integrate into existing dev workflows (IDE, CI/CD, Slack incident reports) without vendor lock-in; orchestrator pattern is reusable for other agentic workflows

Detailed Brief

The Offline Loop: Build → Eval → Optimize

  • Claims: The offline loop (before production) includes specking, building, evaluating, and optimizing agents in an iterative cycle analogous to software development
  • Evidence: Spec defines context, tools, jobs-to-be-done, constraints; coding agents translate specs into target harnesses (Langchain, etc.); eval suite acts as "unit tests" with pass/fail criteria; optimizer generates mutations or remedies based on eval failures
  • Caveats: Assumes teams can articulate clear specs upfront (often difficult for novel workflows); eval completeness is discovered over time, not upfront; optimization may require multiple cycles and doesn't guarantee convergence
  • Implications: Offline loop compresses iteration time by running evals autonomously; enables parallelism (test multiple variants simultaneously); reduces manual QA burden; harness portability means teams can swap frameworks without rewriting agents from scratch

The Online Loop: Monitor → Diagnose → Feedback

  • Claims: The online loop (post-deployment) continuously monitors live traces, diagnoses failures via root-cause analysis, clusters failure modes, and feeds insights back into the offline loop to generate new evals and optimizations
  • Evidence: Diagnostics agent segments traces intelligently to avoid reading millions of executions; produces HTML reports with failure modes, recursive "why" chains, and remedies; assumptions block catches incorrect inferences when code isn't accessible; markdown task definitions auto-generate for coding agents
  • Caveats: Requires instrumentation and observability integrations (Langfuse, JSON-L exports); trace quality matters (incomplete context logs weaken diagnosis); human-in-the-loop still needed for remedy selection and assumption validation; volume triggers (weekly/daily jobs) may miss time-sensitive failures
  • Implications: Online loop makes agent quality a compounding asset; failure modes become learned artifacts that speed up future diagnosis; reduces operational toil for SRE/AI ops teams; incident response time shrinks from hours to minutes

Eval Philosophy: Binary over Score-Based

  • Claims: Binary pass/fail evals with clear criteria are superior to score-based LLM-as-judge evals for agent development because they provide actionable feedback and reduce variance
  • Evidence: Score-based evals require well-defined rubrics and stable LLM behavior; variance in judge outputs makes A/B testing inconclusive; binary evals (e.g., "did the agent use the correct tool?") have clear call-to-action on failure
  • Caveats: Binary evals may oversimplify nuanced quality dimensions (e.g., tone, creativity); require more granular criteria to cover complex behaviors; LLM-as-judge still needed for subjective criteria but must be carefully designed
  • Implications: Teams should decompose agent success into multiple binary checks rather than a single quality score; eval suites grow large but remain interpretable; reduces debugging ambiguity ("this failed because criterion X was violated" vs. "score dropped from 8 to 7")

Product Architecture: Orchestrator + Specialist Agents

  • Claims: Mutagent's platform is an orchestrator plus two specialist agents (evaluator, diagnostics) that run in customer environments and integrate with existing tools
  • Evidence: Demo showed Claude Code integration with slash commands; HTML artifact generation for diagnostics reports; markdown task definitions for coding agents; connectors to observability platforms (Langfuse, JSON-L); GitHub PR generation for remedies
  • Caveats: Research preview status; no managed service yet; setup requires environment configuration and connector plumbing; assumptions block hints at diagnosis errors when trace context is incomplete
  • Implications: Architecture is decentralized and privacy-preserving (runs locally); orchestrator pattern is extensible to other specialist agents (optimizer, deployment, monitoring); product is workflow glue, not a replacement for agent frameworks or LLMs

Why Framework Portability Matters

  • Claims: Agent frameworks evolve rapidly (Langchain → Hermes → DeepAgent) and hit bottlenecks; spec-driven development future-proofs against framework churn
  • Evidence: Teams encounter roadblocks in underlying frameworks and must wait for fixes or migrate; spec-to-implementation translation allows retargeting to new harnesses without rewriting agents; examples: Hermes (defined agent loop runtime), DeepAgent
  • Caveats: Translation may lose framework-specific optimizations (e.g., Langchain's specific memory handling); assumes specs are expressive enough to capture all requirements; migration isn't zero-cost
  • Implications: Reduces lock-in risk; teams can adopt "best-of-breed" harnesses as they emerge; agent logic becomes an asset independent of tooling; specs act as living documentation

Notable Concepts & Terms

  • Agentic AI Engineer: The meta-agent suite (evaluator, diagnostics, optimizer, orchestrator) that builds, tests, and improves other AI agents autonomously—analogous to how coding agents build software
  • Offline vs. Online Loop: Offline = pre-production iteration (spec, build, eval, optimize); Online = post-deployment feedback (monitor, diagnose, feed back into offline loop). Both loops must run agentic to scale
  • Spec-Driven Development: Defining agent requirements (context, tools, jobs-to-be-done, constraints) in a declarative spec that is decoupled from the target harness; enables framework portability
  • Eval-Driven Development: Constructing an evaluation suite (metrics, criteria, datasets) that evolves through production failures; acts as "unit tests" for agents with pass/fail gates
  • Binary Evals vs. Score-Based Evals: Binary = pass/fail criteria (actionable, low variance); Score-based = numeric quality ratings (requires rubrics, higher variance). Binary preferred for agent development
  • Learned Failure Modes: Clusters of production failures with known root causes, stored as historical artifacts; include code-checkable indicators (tool call sequences, content patterns) that speed up future diagnosis
  • Trajectory Evaluation: Assessing the full chain of agent actions (context completeness, tool call correctness, decision tree adherence) rather than just the final output; critical for multi-step agents
  • Multi-Tier Trace Segmentation: Intelligent filtering strategy to pick representative trace samples for diagnosis rather than reading all traces; reduces cost and LLM inference load
  • Assumptions Block: Section in diagnostics reports where the agent declares inferences made when code or full context wasn't available; allows humans to catch and correct misdiagnosis
  • Recursive Why Chain: Root-cause analysis output that traces a failure back through multiple causal layers (e.g., "agent failed → tool returned null → missing API key → config not loaded")
  • Remedies / Mutations: Proposed fixes for diagnosed failures, generated by the diagnostics agent and fed to a coding agent for implementation; can be multi-choice
  • Orchestrator Pattern: Central dispatch agent that sequences and manages specialist agents (evaluator, diagnostics, optimizer) based on triggers (slash commands, CI/CD hooks, incident reports)

Operator Notes / Why Ken Should Care

High relevance for Ken's agent systems and AI ops work. Mutagent tackles a problem Ken likely faces or will face: scaling agent development beyond a few handcrafted features. The insight that "loops for building agents should themselves be agentic" is the natural extension of agentic software engineering. If Ken's team is deploying multiple agents or considering agent platforms for customers, Mutagent's orchestrator + specialist agent pattern is a blueprint.

Specific hooks:

  • Agent Ops / Reliability: Ken's likely drowning in trace analysis. Diagnostics agent + learned failure modes is a direct solution. The multi-tier segmentation strategy (don't read all traces) is cost-efficient and operationally sound.
  • Eval Infrastructure: If Ken's building evals manually, this is painful and doesn't scale. Eval-driven development + binary criteria philosophy is immediately actionable. The "evals are discovered, not designed upfront" insight is gold—Ken should steal this for his own workflows.
  • Framework Hedging: Spec-driven development = insurance against framework churn. If Ken's betting on Langchain today, he should care that Hermes or DeepAgent might be better in 6 months. Portability reduces switching cost.
  • Business / GTM: Mutagent is positioning as infrastructure for "hundreds of agents" enterprises. If Ken's building agent platforms or consulting on agent deployment, this is competitive intelligence. The research preview + local deployment model is customer-friendly (no data egress).
  • Investing Lens: Mutagent is pre-revenue (research preview) but solving a real pain point. Market timing is good (enterprises are moving from "one agent experiment" to "agent portfolio"). Risk: small team, early product, unclear moat (can hyperscalers bundle this?). Opportunity: if they nail the orchestrator UX and integrations, they become the "GitHub Actions for agents."

Caveats for Ken:

  • Product is immature (research preview). No pricing, no managed service, unclear scale limits.
  • Diagnostics agent's "assumptions block" is a red flag for reliability—hallucination risk in root-cause analysis could send teams down wrong paths.
  • Eval philosophy (binary > score) is opinionated but not universally applicable (e.g., creative tasks need nuanced evals).
  • Competitive pressure from observability platforms (Langfuse, Braintrust, Weights & Biases) adding agent-specific features; Mutagent needs to differentiate on the orchestration layer, not just diagnosis.

Action items:

  • If Ken's team has >5 agents in production, trial Mutagent's diagnostics agent on a subset of traces. Measure time saved vs. manual log review.
  • Steal the "learned failure modes" concept for Ken's own agent ops. Build a lightweight version (trace clustering + root-cause checklist) in-house.
  • Watch Mutagent's trajectory for 6 months: if they ship a managed service + good integrations, they could be an acqui-hire or partnership target for a hyperscaler.

Watch Map

Note: Timestamps were not available in the provided transcript. Below is a logical chapter map based on content flow.

  • unavailable: Introduction to Mutagent and team; overview of offline vs. online loops for agent development
  • unavailable: Problem statement—manual agent iteration doesn't scale; human review becomes the bottleneck
  • unavailable: Deep dive on offline loop stages: Spec → Build → Eval → Optimize; spec-driven development rationale
  • unavailable: Build stage: coding agents translate specs to target harnesses; framework portability argument
  • unavailable: Eval stage: eval-driven development philosophy; binary evals vs. score-based evals; trajectory evaluation
  • unavailable: Ship to production; transition to online loop: Monitor → Diagnose → Feedback
  • unavailable: Diagnosis stage: root-cause analysis, failure mode clustering, learned indicators, multi-tier trace segmentation
  • unavailable: Optimization stage: generating remedies/mutations, feeding back into offline loop
  • unavailable: Product demo: orchestrator in Claude Code; diagnostics agent workflow; HTML report walkthrough; assumptions block; markdown task definitions
  • unavailable: Closing remarks: stop debugging, let agents do tedious work, reach out for questions

Source/Metadata

  • Title: The Agentic AI Engineer - Benedikt Sanftl, Mutagent
  • Transcript words: 4,762
  • Video duration: 2,089 seconds (~35 minutes)
  • Timestamp note: Timestamps/chapters were not available in the provided transcript. Content flow reconstructed logically from speaker transitions and topic shifts.

Transcript

4004 words en Processed in 219.5s

Hi everybody. Welcome to our talk, the Agentic AI Engineer. I'm Bene, CEO and co-founder of Mutagent, and I'm here with my colleague. Hi, I'm Burak. I'm the CTO of Mutagent, and today we're going to talk about loops and how the Agentic AI Engineer works. So, as you're all aware of now, loops is the hot topic, how you build software in an agentic loop, and we apply the same loop to the building of AI agents. And as you're all aware, there's two concepts here. One is the offline loop, where while you build, you iterate on your agent, you test it, you evaluate it, you improve it, and you go on. And then you have a second loop, which we call the online loop, where once your agent is deployed to production, you monitor its traces, your diagnosis, and then you feed it back into your optimization loop to iterate and have multiple versions of your agents. Yeah, until now, what we did was doing this loop manually. It's quite slow. The lifecycle is basically, you have an issue, you want to change something to your agent, you implement the change, you maybe implement the change if you use coding agents for it. Yeah, you generate some samples for this new feature or issue to test it. Yeah, then you look at the result, you look through the traces, how does the outcome look, then you maybe ship it, you do A-B testing, and all your feedback takes very long. Yeah, and the bottleneck basically becomes the human review and the human building time. And that you can't scale, especially if in your organization, you are now planning to roll out hundreds of agents, etc. Yeah, and this is why we think the agentic AI engineer is the natural next step to build agents. And I'll have Burak deep dive into how we improve timing and the road to production reliability with the agentic AI engineer. So yeah, the key thing here is basically once you reach a certain number of agents or AI based features, the human performing this loop cannot really scale in enough time. So this is why doing this agentically is the key to increasing the throughput, because then you can fit many more cycles into the same time window. And now how that loop works is basically we have a few stages. So this is when you are starting from scratch, like the current software development practices, you first create a spec for your agent, or your skill in this case. And here you need to define all the functions that agents need to handle the decisions that it has to make on certain conditions. And here again, this is when you are starting from scratch, and here again, this is only the definition stage. Once you define your agents requirements, then you can finally go on to the build and build is where you then realize that spec in a specific harness or agent framework, or in these days, you could even build it as a cloud code or a codex agent. Then comes the next step. Then comes the next step. This is where you define clear evaluations to evaluate your agent's performance, because these are the key metrics then where you can say, hey, my agent is functional or not. Think of essentially equivalent to unit tests for coding. This is how you verify your agent works. Then after evaluation, if everything looks fine, usually have the ship, basically where you deploy this agent to production. Again, this can be a code update. This can be a direct update on any agent platform or again, your local harness agents. Then comes the online part. This is where then the agent is continuously monitored for issues and based on certain trigger conditions, then you can start automatic diagnostics. Again, this can be based on the volume of traces that your agent generates or weekly or daily jobs. Then we go into diagnosis stage. This is where you collect all the failures for your agents and do structured root cause analysis to then understand where the failures are coming from. Once you understand and categorize the failures, then you can finally go on to the optimization stage. This is where then you create, let's say, very specific changes or mutations for your agents to deal with the found failure modes. And then the whole cycle repeats again. You evaluate and if everything looks good, then you can deploy again. Now we will maybe do a deep dive on each stage, what that entails. So before we continue, Burak, we have two passes here. One is the cold start path and one is basically existing features. Should we dive deeper for a half a minute on what the difference is here, what we see and why this is important? Yeah. So today, if you again sit down to build an agent, there are two options. Either you already have an agent, it's built, it's already running somewhere with a certain accuracy. However, the other option is again, you are creating from scratch. So when you create from scratch, obviously you design with the spec and the conceptualization stage. If you have an existing feature, the most likely again, the agent is there, but then you are optimizing over something that already exists. Okay, cool. And yeah, let's dive deeper into each phase. I mean, you just mentioned the specking of new building agents. So this seems to be an important artifact. So let's dive into it. Right. So with the spec driven development also prevalent for building software artifacts these days or any kind of coding agent workflows. Basically, the goal is to capture the requirements for the agent and especially the success criteria. So yeah, the reason being apart from coding agents, the agents in other domains, they handle specific processes and this differs from company to company. So it's very specialized to the environment where the agent is in. And then the key point here is to clearly define which context requirements that the agent has again, which integrations and tools it then needs to have. What are the jobs to be done or the responsibilities that the agent will handle and what it will not. [SPEAKER_01] And then finally, the constraints and in general, the boundaries for the agent. And then the solution is to be done. [SPEAKER_00] Now, once we can define a clear spec, as I said, it becomes a blueprint for any future development, which then the implementation is held against. [SPEAKER_00] Okay. [SPEAKER_00] And let's dive into how we would then build an agent from the spec, right? Right. So basically, then the spec tells your coding agent what to build. And then here the target platform choice is entirely yours. [SPEAKER_01] Because as I mentioned, the agent space is changing very rapidly these days. [SPEAKER_01] So the framework you are using today, you might want to change in a year or so. [SPEAKER_01] So essentially because of that spec is isolated from the implementation detail. So once you decide to build, you can pick any target here than your coding agent of your choice will take that spec and give you an initial version of the target. So that agent, which is then basically customized to run on any platform that you see fit. [SPEAKER_00] Why would I want to change the platform in a year down the line? [SPEAKER_00] So what did we learn out of experience here? [SPEAKER_00] Yeah. Yeah. [SPEAKER_01] Yeah. Yeah. Yeah. So occasionally you hit a bottleneck or a roadblock, and then you have to rely on the underlying framework to get rid of that. [SPEAKER_01] So that agent, which is then customized to run on any platform that you see fit. Why would I want to change the platform in a year down the line? So what did we learn out of experience here? [SPEAKER_00] Yeah. [SPEAKER_00] Yeah. [SPEAKER_00] Yeah. Yeah. [SPEAKER_01] Yeah. [SPEAKER_00] Yeah. [SPEAKER_01] Yeah. Yeah. So occasionally you hit a bottleneck or a roadblock, and then you have to rely on the underlying framework to get rid of that. And this can sometimes take a while. The key here is to be flexible because, yes, essentially, you want to pick the best harness or the framework that can fulfill your requirements. We've all seen it in the last months, how new harnesses shipped and how everything went from agents building code towards defined as an agent loop runtime. We've seen Hermes coming up, deep agents and all of these frameworks, right? [SPEAKER_00] Yeah. Okay, let's continue. [SPEAKER_01] What happens after build? Now, after you build for agents, you essentially go into the eval-driven development loop, which I would call. And this is equivalent to test-driven development for building software with agents because the agent needs a termination condition, right? So when is an AI feature or an agent good enough? So here, there are two ways to create your eval suite, which is composed of the evals, the metrics and the criteria you evaluate. And the datasets themselves, which usually contain the cases you have to satisfy. Now, in the beginning, the original option is that you can sit with your domain experts and try to write down eval metrics and criteria that would cover the feature or the agent you want to build. But most teams working on that will already know that this is difficult, as in you cannot pre-guess the entire evaluations suite from the beginning. And secondly, you can always start with historical data or synthesized data from a known sample of the data that you would like the agent to be tested on. But essentially, the real and the complete eval suite is a product of discovery. What that means is over time from user feedback, from production failures, you collect the metrics and criteria plus the additional dataset cases, which is often representative for edge cases or hard cases that the agent needs to deal with. And with that, you finally have an evaluation suite where you can run the agent against and exactly know where it fails. Okay, before we continue, why does the agentic AI engineer help us here so much? So obviously we're talking about the evaluator agent. So why is it so good to use this concept or thought process in the building? Yes, one issue is, imagine you have a dataset item of 200 here without automated evals running this and evaluating by human eyes takes quite a while. You would have to scroll through an observability dashboard and logs and this in turn increases your loop time per eval state. So as soon as you have a lot of features that you need to evaluate an experiment on, it suddenly becomes impossible to do this quickly or in parallel. Then the human essentially becomes the bottleneck. Okay. So we just have an agent do this work in sifting through traces. Yes, as the current era says, you design loops for your agents so they can autonomously work as many of these things in the background. And then your job becomes designing these loops with a clear eval or termination gate. Okay. Understood. So let's look at how good eval is supposed to be constructed and what it evaluates. Right. So in the context of agents in general, mainly the trajectory is important because agents receive input or intent or tasks, and then they have a specific system prompt or a decision tree that they operate over. And when doing agent evaluations, we check, hey, was the context complete as in did the agent have all the required context to perform the task end to end. Then this includes chain checking every tool output in the trajectory because every wrong tool output in session can in the end lead to a wrong output as the final output. And apart from that, there are different things you can evaluate on. So these days the harness that the agent is operating on has quite drastic effects on the agent behavior. And this is another vector of optimization, but in the end, when you evaluate an agent, you would like to evaluate all of these things and not something in just isolation. [SPEAKER_01] Right. And in general here, what makes an eval useful? [SPEAKER_00] So you can always have metrics or evals that are working in LLM as a judge fashion and give you some score. [SPEAKER_00] But in order to make an eval useful, it has to provide actionable feedback. As in, when you use score based evals, unless your rubric is very well defined, this does not exactly tell you what to fix. In such cases, using binary type of evals or criteria is preferred because there you have a call to action. [SPEAKER_01] If an eval or a criteria fails, you know exactly what happened and how to deal with the problem. [SPEAKER_01] name [SPEAKER_00] name [SPEAKER_00] name [SPEAKER_01] name name [SPEAKER_01] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] [SPEAKER_00] you have to make sure your LLM as a judge solution deals with this variance problem. Otherwise, it's hard to run experiments and conclusively say, hey, my improved version is better than my initial version of my agent. Right. After evaluation, we go live and this is where collecting learned failures, failure modes or primary signals from the executions takes place. Here, the idea is, when an agent encounters a problem over time, there will most likely be multiple occurrences of this. So it starts by identifying what failure mode the agent has encountered and then grouping essentially these failure modes by the root causes and where they originate from. Again, this could be a section in the agent prompt. This could be missing tools or malfunctioning tools. But after your diagnostics results are categorized, you can finally generate new evaluations. First to detect these problems. And then second, you can generate improvements and remedies based on these problems. And with that over time, you have a buildup of these learned failure modes. So every agent over time gathers this historical data that it can always check against. When you start diagnosing, there's an upfront cost in the beginning because you often need to deep read the LLM traces to find out what's going on. Over time, you can collect code checkable indicators per failure mode. What that means is there are maybe specific pieces of content [SPEAKER_00] or there's a specific tool call sequence where you know that the agent will encounter an issue. [SPEAKER_00] And then this later on helps you to diagnose problems in your traces [SPEAKER_00] without actually having to read through all your traces. [SPEAKER_00] Now, why is that a problem? Then you can finally generate new evaluations. First to detect these problems. And then second, you can generate improvements and remedies based on these problems. And with that over time, you have a build up of these learned failure modes. So then every agent over time gathers these historical data that it can always check against. When you start diagnosing, there's a bit of an upfront cost in the beginning because you often need to deep read the LLM traces to find out what's going on. Over time, you can collect code checkable indicators per failure mode. What that means is there are maybe specific pieces of content [SPEAKER_00] or there's a specific tool called sequence where you know that the agent will encounter an issue. [SPEAKER_00] And then this later on helps you to diagnose problems in your traces [SPEAKER_00] without actually having to read through all your traces. [SPEAKER_00] Now, why is that a problem? [SPEAKER_00] Because if you have, let's say, millions of agent traces, [SPEAKER_00] trying to read all of these actually costs more than the execution itself. [SPEAKER_00] So it's not the most efficient way of diagnosing the problems. [SPEAKER_00] What you want to aim for is try to pick a representative sample [SPEAKER_00] from your whole traces by intelligent segmentation strategies [SPEAKER_00] and using the learned indicators. [SPEAKER_00] Okay, cool. [SPEAKER_00] And then with that, you can finally start building the autonomous optimization loop [SPEAKER_00] because here then given that you have an eval suite, you can [SPEAKER_00] vary your feature again, update certain sections or even run auto research style experiments. [SPEAKER_00] To see whether you can reach the desired or the target scores for your evals, [SPEAKER_00] as long as you can reach that, then it's automatically shipped to production. [SPEAKER_00] Then from production, you have your outer loop. [SPEAKER_00] Again, every issue that's found then on diagnostics gives you an improvement or a remedy, [SPEAKER_00] which then you can optimize on. [SPEAKER_00] And then as long as the eval suite is green, then this again gets deployed to production. [SPEAKER_00] Let's continue now with the whole life cycle. [SPEAKER_00] Yeah. [SPEAKER_00] As we learned now, the details of each, everything starts with a spec. [SPEAKER_00] You define and design it. [SPEAKER_00] You build your agent. [SPEAKER_00] And all of this can be done, obviously, agentic. [SPEAKER_00] So you define agents that do this work. [SPEAKER_00] And then they all together become the agentic AI engineer. [SPEAKER_00] You go through the offline loop where once you build your evaluation system, you evaluate, you optimize, you test, you improve your accuracy on your test data set. Once you feel ready, you deploy it to production. And that's when the online loop starts. Here you get real feedback from users. [SPEAKER_01] Yeah. Real test cases. You continuously monitor them. You look at your life traces. You diagnose them. As we learned before with the diagnose agent, you can do root cause analysis. You can cluster your failure modes. You'll derive new evals criterias from it. They become part of the spec. They become part of the agent. [SPEAKER_01] So they continuously grow with you while you use your agent in production. [SPEAKER_01] And this becomes the online loop. [SPEAKER_01] And the more use cases you collect, the more production data you see, [SPEAKER_01] the better your agent becomes and the better scoring your agent becomes. [SPEAKER_01] And this altogether now becomes one end-to-end loop that you can run agentic. [SPEAKER_01] And from here on, you can now see the combination of software-driven development with coding agents. [SPEAKER_01] You can transport this concept also for AI engineering and building AI agents. [SPEAKER_01] And now, as we at Mutation work on this, we're going to show you now how this looks like as a product. [SPEAKER_01] This is how it looks like. [SPEAKER_01] For now, everything runs in your environment. [SPEAKER_01] And as cloud and local, we'll offer and manage service down the line. [SPEAKER_01] So it's a set of agents. [SPEAKER_01] We have two in research preview. [SPEAKER_01] The ones we talk deeper about this in our talk. [SPEAKER_01] First one is the evaluator agent that helps you build an eval set, a good data set, [SPEAKER_01] because this is the core of the optimization, the eval-driven development loop. [SPEAKER_01] The second one we have is the diagnostics agent. [SPEAKER_01] It helps you diagnose your traces you already have in productions, [SPEAKER_01] because we learn from our users that reading through these traces took them a lot of time, and having an agent they can just spin off from their coding environment to analyze the traces is very helpful to them. [SPEAKER_01] And as you can see here, both of these agents are connected through an orchestrator, [SPEAKER_01] which runs in your coding environment. [SPEAKER_01] And what you need basically is connectors to sources where you basically have all your traces, [SPEAKER_01] where you get your incidents from. [SPEAKER_01] This could also be a ticketing system, Slack, where people report failures of your agents. [SPEAKER_01] Then you connect them. [SPEAKER_01] And then obviously, as we mentioned before, there's different target platforms. [SPEAKER_01] This could become a PR that automatically gets erased in GitHub. [SPEAKER_01] This could be just the adjustment of your agents in MD files. Different frameworks we target, or being deployed to manage services. This is how our platform works. We spin off different agents. They are connected to an orchestra. And I guess Burak, let's show our first agent in research preview to the audience. In practice, as I said, you have an orchestrator, which handles the dispatch for all the other sub-agents or stages we have. And then you can use it in any coding harness or agent you want. In this case, I'm using Cloud Code. But then what it will do is when you first boot up, it will show you a dashboard of which stages you have. And also the existing things in, let's say, your code base, your agents and the prior configurations. Now, here I have some star commands. These are basically trigger commands for the workflows or specific stages. And if I type diagnose here, this will basically start the diagnostic stage for the root cause analysis. And here, you can point the diagnostics into a specific scope. This can be an agent. This can be a skill as well. So you could basically diagnose all invocations of a particular skill. Here, when you start the diagnostics, it will retrieve traces from your configured source platform. In this case, it can be something like length use. It can be your local cloud transcripts or another JSON-L format that's exported from one of the observability platforms. Now, the diagnostics run itself takes quite a while to finish. So I will show you instead a pre-generated version of what it does and how that looks like in the end. Okay. So what do we see here now? So basically, once you run the diagnostics agent on your agents, you will get a generated HTML artifact, which basically shows you the details of your agent, which tools you have. Here, when you start the diagnostics, it will retrieve traces from your configured source platform. In this case, it can be something like Languse. It can be your local cloud transcripts or another JSON-L format that's exported from one of the observability platforms. Now, the diagnostics run itself takes quite a while to finish. So I will show you instead a pre-generated version of what it does and how that looks like in the end. Okay. So what do we see here now? So once you run the diagnostics agent, as I said, on your agents, you will get a generated HTML artifact, which shows you the details of your agent, which tools you have. Then in general, if you have code access, it will also tell you which harness or the framework that the agent is using. And then here in general, it will go through your traces and then select primary signals or failure modes, depending on how frequent they occur in the trace samples. Now, as I said, most of the time, you don't want to read all of your traces because this is not cost efficient. [SPEAKER_00] So there the diagnostics use a multi-tier filtering or segmentation to pick a representative sample. [SPEAKER_00] So originally a few pieces of the traces or a portion will be read by the LLM to see if there are any detectable obvious problems. [SPEAKER_00] And based on that, then the diagnostics decides to focus on a particular failure mode or a signal. [SPEAKER_00] The other option is to trigger the diagnostics. [SPEAKER_00] You can also specify an issue that you have seen yourself or an issue that you're looking for. [SPEAKER_00] This could be something that the users reported. [SPEAKER_00] In this case, it's more of a guided search, meaning, hey, if you tell the diagnostics agent you have a particular problem you're interested in, then it will try to find all incidents or occurrences regarding the same problem. [SPEAKER_00] Okay. [SPEAKER_00] And what do we see here in full? [SPEAKER_00] So this is in general the overview, which tells you again which issues were detected and shows you the frequency given the timeframe. [SPEAKER_00] Then for each tab, you will have a failure mode and an explanation of the problem. [SPEAKER_00] And then here you can see where that problem comes from or why it happens. [SPEAKER_00] The agent will provide a recursive why chain to tell you, hey, this is where the issue is originating from. [SPEAKER_00] Now, in addition, it will offer you remedies or fixes to deal with that. [SPEAKER_00] And then here you will see an assumptions block. This is important because when you are reading traces, sometimes you don't always have access to the code. So this helps us detect the LLM or the diagnostics agent makes certain assumptions that are not correct. [SPEAKER_00] And we can see and correct them here. [SPEAKER_00] But in the end, you will be presented with certain corrections or remedies, which can fix your problem. [SPEAKER_00] You can pick either the recommended or as many as you need as it's multi choice. [SPEAKER_00] And once you go through all your problems, at the end, you get to a decisions page. And here again, for those using text to speech or speech to text AI features like Whisperflow, there's always a general feedback box so that you can talk into your mic. And finally, when you are happy with all the decisions, you get a markdown task definition for your coding agent. So all the fixes or the remedies can then be directly applied once you go back to your terminal and to your coding agent. So thanks for the demo, Burak. So last words on our product. [SPEAKER_01] Thanks for listening to our talk. [SPEAKER_01] I hope we could show you some good insights on how we envision the future of agent building with the agentic AI engineer. [SPEAKER_01] Stop the debugging. Have agents do the tedious work and reach out to us if you have further questions. And yeah, have a great conference. So this is in general, the overview, which tells you again, which issues kind of were detected and kind of shows you like the frequency, giving the given timeframe. Then for each tab, you will have a failure mode and kind of like an explanation of the problem. And then here you can kind of see where that problem comes from or why it happens. The agent will provide kind of like a recursive Y chain, let's say, to then tell you, hey, this is where the issue is originating from. Now, in addition, it will offer you kind of remedies or fixes to deal with that. And then here you will see an assumptions block here. This is important because when you are reading traces, sometimes you don't always have access to the code. So this helps us detect the LLM or let's say the diagnostics agent makes certain assumptions that are also not correct. And we can see and correct them here. But in the end, you will be presented by certain corrections or remedies, which then kind of can fix your problem. You can, you know, pick either the recommended or, you know, as many as you need as it's multi choice. And once you basically go through all your problems, at the end, you kind of get to a decisions page. And here again, for those using, you know, text to speech or speech to text AI features like Whisperflow, there's always a general feedback box so that you can always talk into your mic. And finally, when you are happy with all the decisions, you get a markdown, let's say, task definition for your coding agent. So all the fixes or the remedies can then be directly applied once you go back to your terminal and to your coding agent. So thanks for the demo, Burak. So last words on our product. Yeah. So thanks for listening into our talk. Yeah. I hope we could show you some good insights on how we envision the future of agent building with the agentic AI engineer. Yeah. Stop the debugging. Have agents do the tedious work and reach out to us if you have further questions. And yeah, have a great conference.