Open Reader

LLM Observability, Evaluation, Experimentation Platform — Dat Ngo, Arize

completed 16:32 Jun 07, 2026 Watch on YouTube

Current Status

completed

Video ID

JsCCrBF7F1g

RAG / Chat

Enabled
LLM Observability, Evaluation, Experimentation Platform — Dat Ngo, Arize
Description

Your agent called tool B before tool A, and B has a dependency on A. You did not catch it because nothing in your code audits agents. The telemetry does. Dat from Arize AI walks through what observability actually means when the system you are debugging is nondeterministic and the execution path changes with every run. The talk covers the five flavors of eval signal (LLM as judge, human feedback, golden datasets, deterministic checks, business metrics), what scope to run them at (single span, multispan, trajectory, session), and where this is heading. Arize Phoenix is open source, runs as a single container, no Kubernetes required. The enterprise product adds an AI layer called Alex that scans traces, surfaces high latency and errors, and creates evals automatically. The stated goal: automate you out of the observability loop entirely. Speaker info: - https://www.linkedin.com/in/datdarylngo/ - https://x.com/dat_attacked

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI agent/LLM production success requires a comprehensive observability-evaluation-experimentation flywheel that can be automated via AI, moving from manual dashboards to agent-driven system improvement
  • Why it matters: Addresses the critical gap between building AI agents and running them reliably in production at enterprise scale, with concrete frameworks for evaluation scopes, automated debugging, and regression detection
  • Best use: Study for building production-grade agent infrastructure; note the evaluation taxonomy (span/multi-span/trajectory/session), then examine how Arize's automation approach (Alex AI) could inform Ken's own agent system architectures

Executive Summary

Dat Ngo, an AI architect at Arize AI working with major enterprises (Uber, Booking, Reddit), presents a practitioner's framework for making AI agents work in production. He burned 100B-1T tokens last year and distills lessons into three pillars: observability (what's happening), evaluation (is it working), and experimentation (how to improve). His core insight: 'AI feels like magic but it's just software reimagined'—meaning the same engineering patterns apply, but non-determinism creates new failure modes where fixing one issue often introduces 2-3 regressions you didn't anticipate.

The observability layer uses OpenTelemetry traces/spans as the fundamental audit record since 'code doesn't audit agents, telemetry does.' Arize provides distributional views across all agent instantiations to answer questions like 'what percentage of traffic goes down branch A vs B?' and 'which component in this path causes latency spikes?' They expose session-level state tracking (inspired by Anthropic's recent managed agents paper) and trajectory visualization showing all possible agent paths/loops, not just individual executions.

Evaluation taxonomy is the talk's most actionable content: (1) span evals = single component input/output, (2) multi-span evals = cross-component data flow (e.g., how well agents pass data), (3) trajectory evals = did agent call things in correct order to complete business process, (4) session evals = state machine evaluation (user frustration, question coverage). Evaluation sources include LLM-as-judge (tuned on golden datasets), human feedback (end-users + domain experts), deterministic logic (JSON schema validation), and business metrics (revenue/cost/time). Key caveat: 'Just because you can eval something doesn't mean you should'—minimize eval set due to cost.

The vision is full automation: Arize's 'Alex' AI system plans and executes the entire flywheel without human intervention—detecting issues, creating evals on-the-fly, running experiments (prompt/model/orchestration changes), and proposing fixes. They expose CLI tools and primitives so coding agents (Claude Code, Cursor) can drive the platform programmatically. Two products: Phoenix (open-source, single container, no Kubernetes) for engineering teams; AX (enterprise) for largest deployments. The roadmap is 'automate you out of the process'—pull down the ecosystem and 'it just works.'

Key Takeaways

  • Claim: Non-deterministic AI systems create a new failure mode: perceived fixes often introduce 2-3 regressions you didn't know about | Evidence: Dat states this as a pattern observed across enterprise AI teams—when you 'fix the thing you thought you fixed, you might have actually produced two or three regressions that you didn't really know about' | Caveat: No specific examples or metrics quantifying regression rates; based on Arize's customer observation across unspecified number of enterprises | Implication: Ken's agent systems need automated regression detection as a first-class concern, not an afterthought—every change requires distributional evaluation across all agent paths/branches | Timestamp: timestamp unavailable
  • Claim: OpenTelemetry traces/spans are the fundamental audit primitive for agents because 'code doesn't audit agents or harnesses, it's actually the telemetry that does that' | Evidence: Arize built their platform 'OpenTelemetry-first' with auto-instrumenters that add one line of code to see what's happening in any framework/SDK and generate traces/spans | Caveat: Assumes teams adopt OTel conventions; no discussion of trace storage costs, sampling strategies, or trace data volume at scale | Implication: For Ken's operator workflow, instrumenting via OTel provides vendor-agnostic agent observability and enables switching between observability backends without rewriting instrumentation | Timestamp: timestamp unavailable
  • Claim: Evaluation scope matters as much as evaluation method—four distinct scopes (span, multi-span, trajectory, session) solve different production problems | Evidence: Span = single component I/O; multi-span = 'how well are agents passing data back and forth to each other'; trajectory = 'did we call things in the right trajectory to finish the business process'; session = state machine evaluation like 'was the user ever frustrated' | Caveat: No guidance on which scope to prioritize or how to balance cost/signal tradeoffs; implies teams should implement all four but earlier stated 'minimal set of evals' | Implication: Ken should map his agent system's failure modes to these scopes: if agents fail due to incorrect tool call sequencing, trajectory evals are critical; if user dissatisfaction is the issue, session-level evals matter more | Timestamp: timestamp unavailable
  • Claim: Distributional agent views answer critical production questions that single-trace debugging cannot: traffic distribution across branches, latency attribution to specific paths, and out-of-order execution detection | Evidence: Arize AX shows 'views into the distribution of your agent—what are all the possible paths and branches, also loops' enabling questions like 'what percentage of my traffic goes down one branch versus another' and 'was there a particular component in that particular branch that we took that caused a significant amount of latency' | Caveat: No details on sampling methodology, statistical significance thresholds, or how to handle agents with exponentially branching paths; assumes sufficient traffic volume for distribution analysis | Implication: Ken's agent systems need aggregate/distributional monitoring, not just per-invocation tracing—build dashboards that show branch-taking frequency and latency p95/p99 by path, not just mean latency | Timestamp: timestamp unavailable
  • Claim: The production AI flywheel (observability → evaluation → experimentation → improvement) is 'very much automatable' via AI agents that detect issues, create evals on-the-fly, and propose fixes without human intervention | Evidence: Arize's 'Alex' AI system takes a query like 'do you see any issues with my application' and 'will go in and plan and run these tasks' including detecting high latency and errors; Dat claims 'our ultimate goal as a company is actually to automate you out of this process' with the vision 'you pull the ecosystem down and then it just works' | Caveat: Alex is demonstrated in demo/aspirational mode; no specifics on accuracy, false positive rates, or production deployment maturity; no customer case studies showing successful autonomous debugging at scale | Implication: This is the eventual end-state for Ken's operator systems: the meta-layer (agent observability/evals/experiments) should itself be agent-driven, but the tech is still early—invest in primitives (OTel, eval frameworks, experiment harnesses) that an AI can later orchestrate | Timestamp: timestamp unavailable
  • Claim: Enterprise AI teams separate into two personas: technical builders (AI engineers automating/coding) and domain experts (product managers/SMEs who understand desired AI experience), and platforms must serve both without forcing non-technical users to code | Evidence: Arize allows 'folks to be able to run evals in just a non-technical way' via UI where users 'select their model, run some out-of-the-box template, or customize some eval' while technical users can 'attach evals and run them programmatically' | Caveat: No discussion of how to handle conflicts when technical and non-technical personas disagree on eval criteria or system behavior; assumes clean separation of concerns | Implication: Ken's agent platforms should expose both programmatic APIs (for engineers building agent systems) and no-code interfaces (for domain experts tuning behavior)—don't assume all users can/want to write Python | Timestamp: timestamp unavailable
  • Claim: LLM-as-judge evaluation requires tuning on golden datasets labeled by domain experts, not just using an LLM out-of-the-box, to approximate trusted human judgment | Evidence: Dat describes the process: 'you'll run techniques like... I'm going to run my LLM as a judge on some golden dataset so that you can tune your LLM as a judge... can I get my LLM to approximate this thing or this person or this dataset that I trust' | Caveat: No specifics on tuning methodology (few-shot prompting vs fine-tuning), dataset size requirements, or when LLM-as-judge approximation is good enough; assumes golden datasets exist | Implication: For Ken's systems, don't treat LLM-as-judge as plug-and-play—invest in curating domain-specific golden datasets and measure judge accuracy against ground truth before trusting it for production evals | Timestamp: timestamp unavailable

Detailed Brief

Observability Architecture: OpenTelemetry-First Approach and Multi-Layer Views

  • Claims: Code does not audit agents; telemetry (traces/spans) provides the audit record of agent actions; OpenTelemetry auto-instrumentation requires one line of code to instrument any framework/SDK; Four observability layers: traces/spans (component-level), sessions (back-and-forth state), runs (executions), and distributional views (aggregate patterns); Distributional views answer questions like traffic distribution across branches, latency attribution by path, and loop detection; Analytics dashboards remain relevant for real-time agent monitoring in enterprises
  • Evidence: Arize uses OpenTelemetry as foundational layer, works across harnesses/agents/model setups; Traces/spans show 'audit record of what did my agent do'; Session tracking inspired by Anthropic's 'managed agents' paper from two days prior to talk; Arize AX visualizes 'all possible paths and branches, also loops' across agent instantiations; Example questions enabled: 'what percentage of my traffic goes down one branch versus another' and 'was there a particular component in that particular branch that caused significant latency'
  • Caveats: No discussion of trace storage costs, sampling strategies, or data volume management at scale; Distributional views require sufficient traffic volume for statistical validity (not quantified); OTel adoption assumes teams are willing to instrument code and integrate telemetry pipelines
  • Implications: Ken should prioritize OTel-compatible instrumentation for vendor neutrality and future-proofing; Distributional monitoring is critical for production agents—single-trace debugging misses systemic issues like branch imbalance or path-specific latency; Session-level tracking becomes essential for stateful agents or multi-turn conversations; Real-time dashboards still matter for enterprise users even in an 'everything is automated' future

Evaluation Taxonomy: Scope, Methodology, and Cost Optimization

  • Claims: Four evaluation scopes: span (single component), multi-span (cross-component), trajectory (execution order), session (state machine); Five evaluation methodologies: LLM-as-judge, human feedback, golden datasets, deterministic logic, business metrics; LLM-as-judge must be tuned on golden datasets to approximate trusted human judgment; Minimize evaluation set because 'there's a cost associated with this stuff'—don't eval everything just because you can; Business metrics reduce to three categories: make money, save money, save time
  • Evidence: Span eval: 'input and output of one part of an LLM call'; Multi-span eval: 'how well are agents passing data back and forth to each other' requiring data from every agent; Trajectory eval: 'did we call things in the right trajectory to finish the business process'; Session eval: 'was the user ever frustrated? Did we answer all of their questions?'; Golden datasets: 'you trust the person who labeled this data because they know the domain'
  • Caveats: No guidance on how to choose minimum viable eval set or balance cost vs. signal; No specifics on LLM-as-judge tuning methodology (few-shot vs. fine-tuning) or required dataset size; Assumes golden datasets exist or can be created; no discussion of cost to produce them; Trajectory evals require understanding correct execution order, which may not be well-defined for exploratory agents
  • Implications: Ken should map agent failure modes to eval scopes: data passing errors → multi-span; incorrect tool sequencing → trajectory; user dissatisfaction → session; Start with deterministic evals (schema validation, null checks) before adding costly LLM-as-judge layers; Invest in golden dataset creation early—it's the foundation for tuning LLM judges; For each eval, explicitly justify the cost with a business metric outcome (revenue/cost/time impact)

Experimentation and Continuous Improvement: Manual to Automated Workflows

  • Claims: Experiments consist of changes to prompts, models, orchestration, or configurations; Teams can start with traces (filter by signal, collect into datasets) or upload input/output pairs directly; The future is automation: 'most people don't want to live in dashboards or buttons or manual things'; Arize exposes all primitives via CLI and tools/skills for coding agents (Claude Code, Cursor) to orchestrate; Alex AI system automates the full flywheel: detect issues, create evals on-the-fly, run experiments, propose fixes; Ultimate goal: 'automate you out of this process'—pull down ecosystem and 'it just works'
  • Evidence: UI and programmatic interfaces for running experiments on datasets; Arize allows 'coding agent' orchestration: 'we expose all the primitives via the CLI and a set of tools and skills'; Alex demo: user asks 'do you see any issues with my application' and Alex 'will go in and plan and run these tasks' detecting 'high latency' and 'errors'; Dat claims: 'we believe a lot of this stuff can be automated' and 'AI should have context of here's the traces, here's what's happening, let me create evals on the fly'
  • Caveats: Alex AI is demonstrated in demo/aspirational mode; no production case studies or accuracy metrics provided; No discussion of false positives, incorrect diagnoses, or when autonomous debugging fails; Automation assumes clean telemetry, well-defined eval criteria, and stable system behavior; The 'it just works' vision is future roadmap, not current reality (based on tone/phrasing)
  • Implications: For Ken's agent systems, build toward agent-orchestratable primitives: every manual workflow should have a programmatic equivalent; The meta-layer (observability/evals/experiments) is becoming agentic itself—design systems that an AI can introspect and modify; Near-term: invest in high-quality telemetry and eval harnesses that an AI can later orchestrate; long-term: let the AI drive the improvement loop; Watch Arize's Alex AI maturity as a leading indicator of when autonomous debugging becomes production-ready

Enterprise Deployment Patterns and Product Strategy

  • Claims: Two products: Phoenix (open-source, single container, local deployment, no Kubernetes) and AX (enterprise SaaS for Uber, Booking, Reddit scale); Dual-persona design: technical users (AI engineers, developers) and non-technical users (product managers, SMEs); Technical users want programmatic control; non-technical users want UI-based eval creation without coding; Arize works with 'world's largest companies and enterprises' giving unique vantage point into industry trends
  • Evidence: Phoenix: 'single container, you can deploy it locally, doesn't require a Kubernetes layer'; AX: 'reserved for largest enterprises between Uber and Booking and Reddit'; Non-technical workflow: 'select their model, run some out-of-the-box template, or customize some eval'; Technical workflow: 'attach evals and run them programmatically'; Dat burned '100 billion to one trillion tokens last year' working with enterprise customers
  • Caveats: No pricing, adoption metrics, or competitive positioning vs. alternatives (Langfuse, Helicone, Braintrust, etc.); No discussion of data privacy, tenant isolation, or compliance for enterprise deployments; Open-source Phoenix vs. commercial AX feature parity unclear—what capabilities require enterprise?
  • Implications: For Ken building agent infrastructure: consider dual-track strategy (OSS for developers, enterprise SaaS for large deployments); Serve both technical and non-technical personas from day one—don't assume all users will code; Phoenix's single-container architecture is attractive for local development and edge deployments where Kubernetes overhead is prohibitive; Arize's enterprise customer base (Uber, Booking, Reddit) makes them a good signal source for what production patterns scale

Notable Concepts & Terms

  • Span eval vs. multi-span eval vs. trajectory eval vs. session eval: Four-level evaluation taxonomy by scope: span = single component I/O, multi-span = cross-component data flow, trajectory = execution order correctness, session = state machine/conversation quality
  • Distributional agent views: Aggregate visualization of all possible agent paths/branches across instantiations, not just single traces—enables traffic distribution, latency attribution, and loop detection questions
  • LLM-as-judge tuning on golden datasets: Process of training/prompting an LLM evaluator to approximate domain expert judgment by running it on trusted labeled data, not using LLM-as-judge out-of-the-box
  • Harness vs. agent: Framework/orchestration layer (harness) that manages agent components—Dat uses terms interchangeably but implies harness is the runtime/framework wrapping one or more agents
  • OpenTelemetry (OTel) traces and spans: Industry-standard distributed tracing format where traces = end-to-end execution and spans = individual operations; Arize built platform 'OTel-first' for framework/vendor neutrality
  • Alex (Arize AI system): Autonomous AI agent that observes telemetry, detects issues, creates evals on-the-fly, and runs experiments to improve agent systems without human intervention—aspirational 'automate you out' vision
  • Phoenix (Arize open-source product): Single-container observability/eval platform requiring no Kubernetes; designed for local development and engineering-first teams
  • Arize AX (enterprise product): Enterprise SaaS platform for largest deployments (Uber, Booking, Reddit scale) with full observability/eval/experimentation capabilities

Operator Notes / Why Ken Should Care

  • For Ken's agent infrastructure work: the evaluation scope taxonomy (span/multi-span/trajectory/session) is immediately actionable—map your agent failure modes to these scopes and build the minimal eval set for each layer before adding complexity
  • Distributional agent monitoring is a blind spot for most teams—build aggregate views (branch-taking frequency, latency by path, loop detection) before over-investing in per-trace debugging
  • The 'automate you out of the process' vision (Alex AI) is aspirational but directionally correct—invest in agent-orchestratable primitives (OTel, programmatic eval APIs, experiment harnesses) that an AI can later drive, but don't wait for full autonomy
  • Arize's dual product strategy (Phoenix OSS for developers, AX enterprise for scale) is a proven go-to-market pattern—consider for your own agent tooling if serving both small teams and large enterprises
  • The non-deterministic regression problem ('fix one thing, break 2-3 others') requires automated regression detection as a first-class concern—every agent change needs distributional evaluation across all branches, not just happy-path testing
  • For content/business: Dat's token burn (100B-1T last year) and enterprise customer vantage point make Arize a strong signal source for production AI patterns—follow their blog/research for early indicators of what scales
  • For investing: Arize is well-positioned in the 'observability for AI' wave with strong enterprise traction (Uber, Booking, Reddit), OTel-first architecture (vendor-neutral moat), and clear vision for autonomous debugging—but Alex AI maturity is unproven; watch for case studies showing autonomous improvement at scale

Watch Map

  • timestamp unavailable: Introduction: Dat's background, Arize overview, three-pillar framework (observability, evaluation, experimentation)
  • timestamp unavailable: Observability deep dive: OTel traces/spans, session tracking, distributional agent views, analytics dashboards
  • timestamp unavailable: Evaluation taxonomy: span/multi-span/trajectory/session scopes + five methodologies (LLM-as-judge, human, golden datasets, deterministic, business metrics)
  • timestamp unavailable: Experimentation and improvement: dataset collection, experiment workflows, automation vision
  • timestamp unavailable: Alex AI demo: autonomous issue detection, eval creation on-the-fly, 'automate you out of the process' vision
  • timestamp unavailable: Product overview: Phoenix (OSS, single container) vs. AX (enterprise SaaS), dual-persona design

Source/Metadata

  • Title: LLM Observability, Evaluation, Experimentation Platform — Dat Ngo, Arize
  • Transcript words: 4111
  • Duration seconds: 992
  • Timestamp note: Transcript does not contain granular timestamps; watch_map reflects logical sections rather than exact timecodes

Transcript

2858 words en Processed in 398.6s

I don't know if there's a cut scene but okay so really nice to meet you all my name is Dat. I work at Arise AI so I'll talk a little bit about what that is, a little bit about me, and what I want to share today is I work very deeply in the space. I'm an AI architect. I work with a lot of the largest enterprises across the world to talk about observability, evaluation, experimentation, but really it's just how do you make AI work right? So I do spend a lot of tokens in this space. This is last opening I dev day I think I made it to probably somewhere between a hundred billion and one trillion tokens last year. I do know the space really well. We work with some of the world's largest companies and enterprises so we get to see their transformation into this space and really what I wanted to share today was what do I see in the industry? So I think we have a very unique vantage point being the company that we are. So we get to see what every team is building, how they're building it, what are the biggest pains that they face, and really how they're trying to fix those things. If I had to really distill down what we do in a nutshell, it's really these three things. Maybe by show of hands, who's built agents? Who's building agents? Who's productionized agents? Who's building harnesses and who has no idea what a harness is? Okay, we're all pretty cracked. So that's good. I think it's really funny that the AI space really just feels like software reimagined. It's really the same set of patterns just maybe a different flavor coming out and it feels like magic but it's not magic right? It's all just engineering. So really what we're going to cover today is three things. The first one is observability, which answers the question of what's happening in the thing that I've built and what does that look like whether it's a harness or an agent? Then we'll go into evals. Evals is just simply how do I drive signal from my systems in some form or fashion right? And then as we talk about how we make improvements in this new non-deterministic world, you'll come to find out that when you make what you perceive as a fix and you fix the thing that you thought you fixed, you might have actually produced two or three regressions that you didn't really know about. So in our world, we talk about observability. Everything that we do is something that we're super proud about. Hotel is a really strong pattern for those in the engineering space but everything we do is through open telemetry. So it doesn't really matter what particular type of harness, agent, model setup that you have. The good news is that being hotel first, we're really prepped for a lot of these use cases. So whether it's an auto instrumenter, basically you add one line of code that one line of code will see what's happening around in that particular framework or SDK, create open telemetry traces and spans and produce these views. So if you've ever seen a trace or span, it's basically the audit record of what did my agent do? Because now we know that code doesn't audit agents or harnesses, it's actually the telemetry that does that. So traces is a big fundamental part of observability. Now there are many other different parts of observability that you should be thinking outside of traces and spans. You can think about sessions too. So I don't know if any of you read the Anthropic paper managed agents that came out two days ago, pretty awesome read, but sessions is another one about state. So what are the back and forth conversations? What are the back and forth states that are happening? You can think runs. Sessions for a lot of people in the enterprise, they may want to end up running. I realize this isn't the easiest to see so let me change over to light mode. A lot of folks would like to understand, hey, what are those back and forth conversations that are being had? So that will be something like a session. So it's like, hey, what are those back and forth conversations that are being had? Great. Now, people like to eval those. So in the enterprise, you'll see a lot of folks being like, I don't really care at the deep level, you know, the agent did this tool call, that tool call. They may not care about that as much as like, hey, was the end user satisfied? Were all their questions kind of answered? Now, one unique thing, I'll change this over to light mode too. One unique thing about what we do here at Arise, this is Arise AX, is that sometimes when you think about your agent as a non-deterministic call, right, you want to be able to see, hey, what did my agent do? There's different paths that your agent could take, right? So different branches. But what if you wanted to look over all instantiations of your agent and get a more distributional view of what's happening? These are kind of like views into the distribution of your agent. So what are all the possible paths and branches, also loops? It allows you to answer questions like, what percentage of my traffic goes down one branch versus another, right? Was there a particular component in that particular branch that we took that caused a significant amount of latency? When we start to talk about agents or different paths, you may think about trajectory evals. And so trajectory could be like, hey, I went down this one path and everything was really good. But for some reason, when I go down this path, the evals or the signal that I'm collecting is dropping. Why? What's the root cause? Oh, the root cause issue was that these two components are actually out of order. I did B before A, and actually, B has a dependency on A. So it turns out the way my LLM decided to call these things was mismatched, right? We need to put some context in there to say, hey, actually, before you do this, you need to do that. And so there's many different views into observability. Of course, things like analytics aren't dead either. So a lot of the folks in the enterprise end up looking at is they just want to build views on what their agents look like in real time. And so being able to customize and build those views out is super, super helpful. to call these things was mismatched, right? We need to put some context in there to say, hey, actually, before you do this, you need to do that. And so there's many different views into observability. Of course, things like analytics aren't dead either. So what a lot of the folks in the enterprise end up looking at is just, they just want to build views on what their agents look like in real time. And so being able to customize and build those views out, super, super helpful. And that's observability in a nutshell. It's can I see all the different layers? We have many, many different types of layers here. The next thing is okay, so I have observability. That's step one. It's the same thing that happened in software, right? Now you have to determine signal, right? And so signal comes in actually many various forms as well. The way I like to break it down is these five flavors of signal. I think everyone here in the room has heard of LLM as a judge. And so it may seem like a simple concept, but in actuality, it could actually get quite complex. And we'll go through all of that. Now you can't forget about your humans. When you think about humans, whether it's the end users using your product, it's extremely valuable signal. So whether you're a product manager or someone technical or non-technical, you do care about this signal. We've all heard of golden datasets. They're extremely valuable because if the third column here represents quality, you trust the person who labeled this data because they know the domain. Then you'll run techniques like, hey, I'm going to run my LLM as a judge on some golden dataset so that you can tune your LLM as a judge. You basically say, hey, can I get my LLM to approximate this thing or this person or this dataset that I trust? And then, of course, we're all thinking about costs as well. So when we think about costs, you don't always have to use an LLM call or even humans. Determinism is super nice. So think about logic or deterministic-based evals. If I go from paragraph to JSON payload, does this JSON, is it a valid JSON? Does it have this schema? Does it have these fields that are non-null? And then, of course, we're all building these things for one of three purposes, I think. So the business metrics you care about are either some form of how do I make more money, how do I save money, or how do I save time, right? And so what you'll notice, as you start to build really good AI products, you'll start to have two types of personas that end up coming together. So obviously, you have your technical users, right? These are your AI engineers, your developers of the world. These people are extremely good at building and automating things, right? They're good at frameworking. But then you have folks who are maybe less technical, but they understand what the AI experience should be, right? These are the subject matter experts, the product managers of the world. These folks end up, you want to relegate the work of, hey, this is how the prompt engineering should go. Here's the evals that I care about. Because you want people who can code coding, and you want people who know the domain to work in that domain. And so in our world, what that looks like is something like this. We allow folks to be able to run evals in just a non-technical way. Of course, if you are technical, you can attach your evals and run them programmatically, if you want. But in our world, we want to be able to say, hey, I want to be able to allow a user to be able to select their model, be able to run some out-of-the-box template, or customize some eval here. And when we talk about complexity on the eval side, right, imagine for a second that you have built some application, some agent, some harness. That harness has got components in it. They may be called deterministically or non-deterministically, whatever. So evals can be run on, you can think, a single component. We call that a span eval. Let me come here. So the scope would be one single input and output. I'll pull up a more complex view of this. But let me close this. But you can think of the simple span input and output as, hey, I want to look at the input and output of one part of an LLM call. So that's, most people understand that, and that's really, really simple. Now, we also have multi-span evals. So you can think of that as, hey, in order to run the eval that I want, it actually requires data across many different components in the system. So if I want to say, hey, how well are agents passing data back and forth to each other? Well, it turns out I need the data from every single agent and how they pass data. So that's a multi-span eval, and it allows you to run more complexity. If you want to look over all of the spans in total, that's something like a trajectory eval. Did we call things in the right trajectory to finish the business process? And then there's that session level eval, right? It's zooming out and saying, hey, what does the state machine, if I want to evaluate that state machine of, hey, let me turn this to light mode. Hey, in this conversation, was the user ever frustrated? Did we answer all of their questions? So think of that as, I want to evaluate the state machine. So as you're thinking about evals, it's not generally, it's also, hey, what flavor of eval do we want to run? But at what scope and depth? So you can get very granular, and then you can also zoom out. And just because you can't eval something doesn't mean you always should. It's not this exhaustive thing. That state machine of, hey, let me turn this to light mode. Hey, in this conversation, was the user ever frustrated? Did we answer all of their questions? So think of that as, I want to evaluate the state machine. So as you're thinking about evals, it's not generally also, hey, what flavor of eval do we want to run? But at what scope and depth? So you can get very granular, and then you can also zoom out. And just because you can't eval something doesn't mean you always should. It's not this exhaustive thing. You want to see, hey, what are the minimal set of evals I can get away with? To understand signal of, is my application working as intended? Because there's a cost associated with this stuff, right? And so TLDR, that's observability and evals in a nutshell. We'll talk about experimentation and improvement. So not everyone starts with traces. If you do start with traces, you can take them and do really cool things, say, hey, show me where some signal is, whatever, bad. Where am I missing stuff, right? Then you can find those things, collect them up into a data set. Also, if you don't have traces, you can just upload a data set outright, input output pairs. And then from here, you can do things like grab that data set, which is just rows and columns of data. Then you can start to run experiments. Experiments can be changes. So as you think about how do I make things better for my agent or harness, it's generally changes. Changes to prompts, changes to models, changes to orchestration, changes to configurations. Think that way. We allow folks to be able to test these things in a UI or programmatically. But one thing I always like to share with our customers is, where is the space going? What we quickly realized at Arise here is that most people don't want to live in dashboards or buttons or manual things. And we very much recognize this. As you think about where the future is going, software will compress. It's going to be easier to build, easier to customize. So everything I just showed you was just the nice manual way to see it. But everything we've done, we've allowed you to be able to do this through your coding agent. We realize people are very comfortable with their cloud code and codexes. So we expose all the primitives via the CLI and a set of tools and skills. So that's opinionated. And then there's also an AI system built into all of this, too. Meaning cloud code, your AI system can end up calling our system. So that you can do things like, if you don't want to just figure out these things on your own, we believe a lot of this stuff can be automated. So I can go here and ask Alex. Obviously cloud code or something outside of the system can call Alex. And you can just simply say, hey, do you see any issues with my application? And because we have all the data, because we have the hooks and everything else, Alex will go in and plan and run these tasks. So our ultimate goal as a company is actually to automate you out of this process. Observability, evals, experimentation and improvement. We think the whole flywheel is very much automatable. Meaning it's not magic, again, but it should feel like magic. So our main goal is one day you work with Arise and you pull the ecosystem down. And then it just works. So, Alex, we are very heavy believers that you shouldn't even have to choose your evals. An AI should have context of, hey, here's the traces, here's what's happening. Let me create evals on the fly and think about them for you. Or, hey, something has changed. I know I need a new eval. But you'll notice Alex is already getting to work with, hey, what's happening here? It looks like we have some high latency. We have some errors detected, things like that. And so in a nutshell, this is where we're going and what we're after. And so as we think about this world, we actually have two products out today. So for more of the engineering first folks, we have Arise Phoenix, which is open source. The really nice thing about Phoenix is it's single container. You can deploy it locally. It doesn't require a Kubernetes layer. And then for our largest enterprises, they use Arise AX, which is generally reserved for some of the largest enterprises between like Uber and Booking and Reddit. You guys couldn't tell we love dark mode, so we probably should try to go light mode in some stuff. But yeah, in a nutshell, that's what we do and who we are. And if you guys want to chat about anything past that, super excited. But yeah, thank you very much for your time today. Thank you. So the scope would be one single input and output. I'll pull up like a more complex view of this. But, oops, let me close this. But you can think of the simple span input and output as, hey, I want to look at the input and output of one part of an LLM call. So that's, most people understand that, and that's really, really simple. Now, we also have multi-span evals. So you can think of that as, like, hey, in order to run the eval that I want, it actually requires data across many different components in the system. So if I want to say, hey, how well are agents passing data back and forth to each other? Well, it turns out I need the data from every single agent and how they pass data. So that's a multi-span eval, and it allows you to run more complexity. If you want to look over all of the spans in total, that's something like a trajectory eval. Did we call things in the right trajectory to finish the business process? And then there's that session level eval, right? It's like zooming out and saying, hey, what does the state machine, like if I want to evaluate that state machine of, hey, let me turn this to light mode. Hey, in this conversation, was the user ever frustrated? Did we answer all of their questions? So think of that as, I want to evaluate the state machine. So as you're thinking about evals, it's not generally, you know, it's also like, hey, what flavor of eval do we want to run? But at what scope and depth? So you can get very granular, and then you can also zoom out. And just because you can't eval something doesn't mean you always should. It's not this exhaustive thing. You want to see, like, hey, what are the minimal set of evals I can get away with? To understand signal of, like, is my application working as intended? Because there's a cost associated with this stuff, right? And so, you know, TLDR, you know, that's observability and evals in a nutshell. We'll talk about experimentation and improvement. So not everyone starts with traces. If you do start with traces, you can take them and do really cool things, like, say, like, hey, show me where some signal is, you know, whatever, bad. Where am I missing stuff, right? Then you can find those things, collect them up into a data set. Also, if you don't have traces, you can just upload a data set outright, input output pairs. And then from here, you can do things like grab that data set, for example, which is just rows and columns of data. Then you can start to run experiments. Experiments can be changes. So as you think about how do I make things better for my agent or harness, it's generally changes. Changes to prompts, changes to models, changes to orchestration, changes to configurations. Think that way. We allow folks to be able to test these things in a UI or programmatically. But, you know, one thing I always like to share with our customers is, like, where is the space going? What we quickly realized at Arise here is that, like, most people don't want to live in dashboards or buttons or manual things. And we very much recognize this. As you think about where the future is going, software will compress. It's going to be easier to build, easier to customize. So everything I just showed you was just, like, the nice manual way to see it. But everything we've done, we've allowed you to be able to do this, like, through your coding agent. We realize people are very comfortable with their cloud code and codexes. So we expose all the primitives via the CLI and a set of tools and skills. So that's kind of opinionated. And then there's also an AI system built into all of this, too. Meaning cloud code, your AI system can end up calling our system. So that you can do things like, you know, if you don't want to just figure out these things on your own, we believe a lot of this stuff can be automated. So I can go here and ask Alex. Obviously cloud code or something outside of the system can call Alex. And you can just simply say, like, hey, do you see any issues with my application? And, you know, because we have all the data, because we have the hooks and everything else, you know, Alex will go in and plan and run these tasks. So our ultimate goal as a company is actually to automate you out of this process. Observability, evals, experimentation and improvement. We think the whole flywheel is very much automatable. Meaning it's not magic, again, but it should feel like magic. So our main goal is like one day you work with Arise and you pull the ecosystem down. And then it just works. So, Alex, we are very heavy believers that you shouldn't even have to choose your evals. Like, an AI should have context of like, hey, here's the traces, here's what's happening. Let me create evals on the fly and think about them for you. Or like, hey, something has changed. I know I need a new eval. But you'll notice Alex is already getting to work with like, hey, what's happening here? It looks like we have some high latency. We have, you know, some errors detected, things like that. And so in a nutshell, this is kind of where we're going and what we're after. And so as we think about, you know, this world, we actually have two products out today. So for more of the engineering first folks, we have Arise Phoenix, which is open source. The really nice thing about Phoenix is it's single container. You can deploy it locally. It doesn't require a Kubernetes layer. And then for our largest enterprises, they use Arise AX, which is kind of generally reserved for, you know, some of the largest enterprises between like Uber and Booking and Reddit. You guys couldn't tell we love dark mode, so we probably should try to go light mode in some stuff. But yeah, in a nutshell, that's kind of what we do and who we are. And, you know, if you guys want to chat about anything past that, super excited. But yeah, thank you very much for your time today. Thank you.