Open Reader

The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks

completed 37:05 Jun 18, 2026 Watch on YouTube

Current Status

completed

Video ID

ObTPqBGsEbA

RAG / Chat

Enabled
The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks
Description

A retail bank spent £85,000 over six months on a chatbot PoC that could not reach production. No one could explain why it was failing. When Sandipan Bhaumik's team got involved, they picked the model in week seven of an eight-week engagement — the first six weeks went to evaluation data, tracing infrastructure, and a measurement pipeline. Six weeks post launch, when the bank updated its interest rate policy and customer satisfaction dropped, the tracing system caught the cause: the new policy document had not been reembedded and the agent was serving stale answers. The talk covers the five pillars he built from that and similar engagements: evaluation (define success numerically before touching code), observability (trace every agent decision — European regulators require it), data foundation (agents do not forgive bad data the way humans do), multi agent orchestration patterns, and governance (47 PII breaches caught in testing before launch). The evaluation data set is a living system, not a fixed benchmark. The production incident playbook connects all five. Speaker info: - https://www.linkedin.com/in/sandipanbhaumik

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Enterprises fail to productionize AI agents because they focus on model selection and demos rather than building the five foundational pillars: evaluation systems, observability/tracing, data infrastructure, multi-agent orchestration, and governance—and the speaker provides a tested framework to fix this, drawn from AWS/Databricks production deployments.
  • Why it matters: This talk addresses the gap between AI demos and production at enterprise scale with concrete patterns, real incident playbooks, and cost/risk trade-offs gleaned from regulated industries (financial services, retail banking) where AI accountability is mandatory.
  • Best use: Use as the operational playbook for production AI systems: adopt the five-pillar framework, integrate the evaluation pipeline and tracing architecture, reference the production incident playbook, and apply the three advanced lessons (test case governance, prompt versioning rigor, behavioral eval cost control) before scaling agents.

Executive Summary

Sandipan Bhaumik (Sandy), a technical lead at Databricks with prior principal architect experience at AWS, presents a production AI deployment framework distilled from years of working with enterprise customers in B2B software and regulated industries. He opens by diagnosing a common failure pattern: organizations rush to choose models, build demos with controlled data, get leadership approval, then face mysterious failures in production because they lack observability (can't see what the AI does), evaluation (can't measure success), and governance (no accountability when things break). This cycle wastes money and effort without ROI.

To solve this, Sandy proposes five pillars that must be designed before writing code, ideally in sequence but often implemented in parallel: (1) Evaluation—define success numerically and build automated test pipelines with golden datasets from domain experts; (2) Observability—trace every agent decision for regulators, debugging, and performance; (3) Data Foundation—architect both question data (for answering queries) and tracking data (tracing logs) with proper governance; (4) Orchestration—manage complexity when scaling from one agent to many, choosing patterns like orchestrator-worker, choreography, or human-in-the-loop; (5) Governance—establish audit trails, PII validation, prompt versioning as code, and model change management. He illustrates these pillars with a retail banking chatbot case study where the team selected the model in week 7 of an 8-week POC, after building evaluation and data layers first.

The case study customer had 20,000 monthly calls, wanted to deflect 60% of simple queries to an AI agent, and had already spent $85K on a failed 6-month POC. By implementing Sandy's framework—starting with 200 human-agent answer examples as the eval dataset, defining 85% accuracy and 60% deflection targets, building an automated evaluation pipeline, setting up tracing, and only then testing models—they launched successfully. Post-launch, they caught a production issue (outdated policy document in the vector database causing wrong answers) within weeks because tracing and CSAT monitoring revealed the problem, which would have been invisible without the system. Sandy shares practical lessons: treat the test case library as a living, governed system with categorized rows; version prompts with detailed commit messages explaining the failure being fixed; and control costs of expensive behavioral evals (layer 3) by running subsets during CI and full tests only on main branch merges.

Sandy emphasizes that enterprises won't run AI in one framework or cloud—they'll use LangChain, CrewAI, AWS, Azure, etc.—so a centralized tracing data strategy (e.g., Databricks Unity Catalog, MLflow, Delta Lake) is essential to serve dashboards, first-line support, auditors, and online monitoring. He provides a production incident playbook (detect via dashboards, diagnose via tracing, contain via prompt rollback or circuit breakers, fix via test case library, then add new test cases) and offers downloadable artifacts including evaluation checklists, tracing setup guides, and templates. The talk concludes with three easily missed lessons: govern the growing test case library with owners and categories, treat prompt commits like code changes with rigorous documentation, and architect CI pipelines to avoid runaway costs from re-running behavioral evals on large datasets.

Key Takeaways

  • Claim: Enterprises fail at production AI because they start with model selection and demos instead of building evaluation, observability, and governance systems first. | Evidence: Sandy observed a repeating pattern across customers: teams debate GPT vs Claude, build features in controlled environments with predictable data, get leadership sign-off on demos, then face mysterious production failures with no ROI and wasted money. A retail banking client spent $85K over 6 months on a failed POC before engaging Sandy's team. | Caveat: The speaker does not quantify how many organizations exhibit this pattern or provide comparative success rates, though he states it appeared 'in every customer conversation' over two years at AWS and Databricks. | Implication: Operators must design the five-pillar framework (evaluation, observability, data, orchestration, governance) before touching code or choosing models, ideally in sequence, to avoid costly demo-to-production failures and ensure measurable business outcomes. | Timestamp: 01:30
  • Claim: The three critical gaps blocking production AI are observability (can't see what AI does), evaluation (can't measure success), and governance (no accountability when AI fails). | Evidence: In customer engagements, Sandy found teams couldn't trace agent decisions for regulators, lacked systems to continuously measure business-relevant metrics (accuracy, deflection, latency), and had no playbook for 3am incidents or data ownership when agents misbehaved. | Caveat: The speaker focuses on regulated industries (financial services) and B2B software; less-regulated contexts may face different governance pressures, though observability and evaluation gaps apply universally. | Implication: Production AI systems require mandatory tracing infrastructure (especially in Europe and regulated sectors), automated evaluation pipelines tied to business KPIs, and incident playbooks that define who owns what when failures occur. | Timestamp: 03:45
  • Claim: Evaluation is the specification for your AI system: define success numerically (e.g., 85% accuracy, 60% deflection) and build a golden dataset from domain experts before any coding. | Evidence: In the banking chatbot project, the team collected 200 real human-agent answer examples, defined 85% accuracy and 60% query deflection as success metrics, then built an automated pipeline comparing agent responses to the golden dataset. This living dataset grew post-launch as new edge cases emerged. | Caveat: The speaker does not specify how to determine the 'right' number of initial test cases (he says 'there's no correct number, but we started with 200') or how to balance coverage versus cost when scaling the dataset. | Implication: Operators should treat evaluation datasets as living systems that grow with production use, categorize test cases by failure type (security, logic, PII) for governance, and use domain experts—not just engineers—to capture real-world edge cases and gray areas. | Timestamp: 06:20
  • Claim: Evaluation must be layered across three tiers: deterministic checks (formats, regex), semantic checks (LLM-as-judge for groundedness/relevance), and behavioral checks (tool call efficiency, duplicate API calls, loops). | Evidence: Sandy describes layer 1 as 'easy, cheap stuff' (email format, PII detection with classic ML), layer 2 as LLM-as-judge prompts checking safety and groundedness, and layer 3 as behavioral analysis catching duplicate database calls that are fine in demos but expensive at scale (thousands of daily queries). The banking chatbot made 3 API calls to answer one balance query—acceptable in testing, costly in production. | Caveat: Layer 3 behavioral evals can become very expensive as the test dataset grows (300-500 rows); Sandy recommends running subsets during CI and full tests only on main branch merges, but does not provide cost benchmarks or thresholds. | Implication: Operators should implement all three layers but architect CI pipelines to control layer 3 costs, prioritize deterministic checks to reduce LLM judge load, and monitor behavioral patterns (API call duplication, retries) to prevent runaway cloud spend at scale. | Timestamp: 08:50
  • Claim: Observability via tracing is mandatory for regulators and debugging: enterprises must capture every agent decision (intent classification, API calls, retrieval, reasoning, guardrails) with timestamps and confidence scores. | Evidence: Sandy shows a simplified trace of a banking chatbot handling an overdraft dispute: intent classification (with confidence score and latency), customer DB API call, policy document retrieval from vector DB, reasoning step, final guardrail checks. Without this, customer disputes have no audit trail, and regulators (especially in Europe) block production deployment. | Caveat: The speaker notes real tracing data 'is not as beautiful as this slide'—he simplified it for presentation. He does not discuss schema design, storage costs, or query performance for high-volume tracing data at scale. | Implication: Operators must implement structured tracing (e.g., MLflow, OpenTelemetry) from day one, design tracing schemas to support auditors/regulators and online monitoring, and use traces to detect issues like duplicate calls or stale data retrieval before customer satisfaction drops. | Timestamp: 11:40
  • Claim: Data foundation must cover both 'question data' (for answering queries) and 'tracking data' (tracing logs for monitoring and audits), with a centralized strategy because enterprises use multiple frameworks and clouds. | Evidence: Sandy estimates 60% of project time goes to data foundation because 'data was built for humans, agents don't forgive.' Agents confidently return wrong answers from stale data. In the banking case, an outdated policy document in the vector DB caused wrong answers; tracing revealed embeddings weren't updated. Databricks Unity Catalog centralizes metadata tagging, permissions, and discovery so agents query tables with context (PII tags, column descriptions). | Caveat: The speaker is a Databricks employee and heavily promotes Databricks tools (Unity Catalog, Delta Lake, MLflow); he does not compare alternatives or discuss vendor lock-in, though he mentions open-source foundations (Spark, Delta, MLflow) and multi-cloud support. | Implication: Operators should architect a centralized data catalog with metadata tagging for agent context, design tracking data pipelines that unify tracing from multiple frameworks (LangChain, CrewAI, etc.) and clouds, and implement incremental data loading with versioning to catch stale embeddings before production failures. | Timestamp: 14:30
  • Claim: Multi-agent orchestration requires choosing patterns based on dependencies: orchestrator-worker for centralized control, choreography for parallel independence, human-in-the-loop for low-confidence thresholds. | Evidence: Sandy contrasts orchestrator-worker (one orchestrator distributes work to specialized agents, all requests flow through central logs) with choreography (autonomous agents listen to a message bus, work in parallel—e.g., mortgage app agents for customer details and approval running independently to reduce latency). He states one agent is simple, but five agents increase complexity exponentially. | Caveat: The speaker does not provide decision criteria for choosing patterns or discuss failure modes (e.g., orchestrator as single point of failure, message bus ordering/idempotency in choreography). He refers to a separate deep-dive video on fault tolerance patterns (saga, compensation, circuit breaker) but doesn't detail them here. | Implication: Operators should select orchestration patterns based on task dependencies and latency needs, implement fault-tolerance patterns (circuit breaker to limit retries, compensation for rollbacks) as agent count grows, and use human-in-the-loop fallback when confidence drops below thresholds to maintain CSAT. | Timestamp: 18:15
  • Claim: Governance for production AI means audit trails for every action, pre-validation of PII, prompt versioning as code with detailed commit messages, and model change management independent of vendor benchmarks. | Evidence: In the banking project, pre-validation with NER and regex caught 47 PII breaches during testing. Sandy emphasizes prompt changes must go through enterprise change management (not just git commit), with commit messages explaining what failure the change fixes. Model provider benchmark boards don't reflect enterprise-specific performance, so teams must test new model versions against their own eval datasets before switching. | Caveat: The speaker does not specify how to enforce detailed commit message standards or integrate prompt versioning with existing ITSM systems, though he mentions integration is necessary for alerting the right person during incidents. | Implication: Operators must treat prompts as code with rigorous versioning and change justification, validate PII at ingestion to avoid compliance breaches, test model upgrades against internal eval datasets (not public benchmarks), and maintain flexibility to switch models to reduce vendor risk and regulatory exposure. | Timestamp: 20:45
  • Claim: In the retail banking case study, the team deferred model selection to week 7 of an 8-week POC—after building eval, data, and tracing systems—and successfully launched with measurable results. | Evidence: The customer had 20K monthly calls, wanted 60% deflection of simple queries, spent $85K on a failed 6-month POC. Sandy's team spent weeks 1-2 collecting 200 human-agent answer examples and defining success (85% accuracy, 60% deflection, latency targets). Weeks 3-6 built data foundation (API connections, distributed storage, tracing) and tested for duplicate calls. Week 7 they ran models against the eval dataset to pick the best. Post-launch, they caught a policy document staleness issue via tracing and CSAT monitoring. | Caveat: The speaker does not disclose the final cost, timeline, or quantitative post-launch CSAT/deflection numbers ('six weeks post-launch, we calculated operational metrics'), so ROI and improvement magnitude are unclear. | Implication: Operators should resist pressure to choose models early, invest first in eval pipelines and data infrastructure so model selection is data-driven, and plan for living eval datasets that catch issues like stale retrieval sources after launch. | Timestamp: 23:10
  • Claim: Three commonly missed lessons are: govern the growing test case library with owners and categories, version prompts with detailed failure-context commit messages, and control layer 3 behavioral eval costs via CI subset testing. | Evidence: Sandy notes test case libraries grow over time (from 200 cases to hundreds more) and need categorization (security, login, logic) so teams can trace changes. Prompt commit messages are often generic ('simple commit messages'), but must document what failure caused the change. Layer 3 evals (tool call checks) cost money at scale; he recommends running subsets during CI, full tests only on main branch merges. | Caveat: The speaker does not provide cost benchmarks, subset sizing heuristics, or tooling for test case categorization/governance, leaving operators to design their own systems. | Implication: Operators should assign test case library owners, tag cases by failure type for efficient root cause analysis, enforce prompt commit standards (document failure and expected fix), and architect CI eval budgets to run cheap deterministic checks always, expensive behavioral checks selectively, to avoid runaway cloud costs as datasets scale. | Timestamp: 30:20

Detailed Brief

The Failure Pattern: Why AI Demos Don't Reach Production

  • Claims: Organizations start with model selection debates (GPT vs Claude), build features in controlled environments, get leadership buy-in on demos, then face mysterious production failures without ROI.; This pattern appeared in every customer conversation over two years across B2B software and regulated industries.; A retail banking client spent $85K over 6 months on a failed POC before Sandy's team intervened.
  • Evidence: Teams debate models, build demos with predictable data, leadership signs off, then after a few weeks people ask why AI isn't answering correctly—leading to wasted money and effort.; No one knew why the banking POC failed: couldn't see what AI was doing, couldn't measure success, couldn't assign accountability when things broke.
  • Caveats: The speaker does not quantify failure rates, provide comparative success metrics, or name specific customer organizations (likely under NDA).; Focus is on regulated industries and B2B; less-regulated consumer contexts may have different risk/governance profiles.
  • Implications: Operators must resist the market's model-first narrative and design systems for observability, evaluation, and governance before writing code.; Leadership pressure to 'do something with AI' drives demo culture; aligning exec expectations with production requirements (tracing, eval pipelines) is critical to avoid wasted investment.

Pillar 1: Evaluation as Specification

  • Claims: Evaluation is the AI system's specification: define success numerically (accuracy %, deflection rate, latency), not vaguely ('good accuracy').; Build a golden dataset from domain experts showing how humans answer real queries, including gray areas and edge cases.; Automate evaluation pipelines to compare agent responses to the golden dataset continuously in production.
  • Evidence: Banking chatbot example: defined 85% accuracy, 60% deflection of simple queries, collected 200 human-agent answer examples as ground truth.; Three evaluation layers: (1) deterministic (formats, regex, PII detection), (2) semantic (LLM-as-judge for groundedness/relevance), (3) behavioral (tool call efficiency, duplicate API calls, loops).; Layer 3 caught the banking agent making 3 database calls for one balance query—fine in demos, expensive at scale with thousands of daily queries.
  • Caveats: No 'correct number' of initial test cases specified (team started with 200); coverage vs cost trade-off is left to operator judgment.; Layer 3 behavioral evals are expensive; running full evals on large datasets (300-500 rows) during every CI run can cause runaway costs.; LLM-as-judge accuracy depends on prompt quality and judge model selection; speaker does not discuss judge model failures or adversarial cases.
  • Implications: Operators should co-design eval datasets with domain experts (support agents, compliance teams) to capture real-world nuance, not just engineer intuition.; Treat eval datasets as living systems: add new test cases after every production incident, categorize by failure type (security, logic, PII) for governance.; Architect CI pipelines to run cheap deterministic checks always, LLM judges selectively, and expensive behavioral checks only on main branch merges to control cloud spend.; Define business-relevant KPIs (deflection rate, CSAT) as eval targets, not just technical metrics, so success aligns with ROI.

Pillar 2: Observability and Tracing

  • Claims: Tracing every agent decision (intent classification, API calls, retrieval, reasoning, guardrails) is mandatory for regulators, especially in Europe and regulated industries.; Without tracing, customer disputes have no audit trail, and teams can't debug production failures (e.g., why AI gave wrong answers).; Tracing enables detection of issues like duplicate API calls, stale data retrieval, and policy document mismatches before customer satisfaction drops.
  • Evidence: Banking chatbot trace example: user asks to waive overdraft fee, agent does intent classification (confidence score, latency), calls customer DB API, retrieves policy docs from vector DB, reasons, applies guardrails, responds.; Without tracing, when customer raised dispute, team had 'nowhere to go' and would have to give discounts to placate customer.; Post-launch, tracing caught an issue where agent queried outdated policy doc (embeddings not updated) causing wrong answers and CSAT drop.
  • Caveats: Real tracing data 'is not as beautiful as this slide'—speaker simplified for presentation, does not discuss schema design, storage costs, or query performance at scale.; No mention of tracing data retention policies, GDPR/privacy implications of storing detailed user interaction logs, or anonymization strategies.
  • Implications: Operators must implement structured tracing (MLflow, OpenTelemetry, vendor SDKs) from day one, not as an afterthought.; Design tracing schemas to support multiple consumers: auditors/regulators, first-line support dashboards, online monitoring/alerting, LLM judges running over traces.; Use tracing to detect behavioral patterns (retries, duplicate calls, tool call loops) that are invisible in metrics dashboards but expensive at scale.; In regulated industries, tracing is a blocker for production approval—build it before pilots, not after.

Pillar 3: Data Foundation Strategy

  • Claims: Data foundation has two parts: 'question data' (for answering queries) and 'tracking data' (tracing logs for monitoring/audits).; 60% of project time goes to data foundation because 'data was built for humans, agents don't forgive'—agents return wrong answers confidently from stale/bad data.; Enterprises use multiple frameworks (LangChain, CrewAI, etc.) and clouds (AWS, Azure, GCP), requiring a centralized data strategy to unify tracing and serve multiple teams.
  • Evidence: Banking case: outdated policy document in vector DB caused wrong answers; tracing revealed embeddings weren't updated, so agent retrieved stale data.; Databricks Unity Catalog example: centralizes metadata (table/column descriptions, PII tags), permissions, and delta sharing; agents query tables with context, improving groundedness.; Architecture: raw data on cloud storage → Delta Lake (database-like properties, incremental loading) → Unity Catalog (permissions, discovery) → apps (Mosaic AI, Genie text-to-SQL, BI).; Centralized tracking data strategy: collect traces from multiple frameworks/clouds, serve to operational dashboards, first-line support, auditors, online monitoring, LLM judges.
  • Caveats: Speaker is Databricks employee; heavy promotion of Unity Catalog, Delta Lake, MLflow without comparing alternatives (Snowflake, BigQuery, AWS Glue, etc.).; No discussion of vendor lock-in, migration complexity, or cost at scale; mentions open-source foundations (Spark, Delta, MLflow) and multi-cloud support but doesn't detail portability.; Schema design, storage costs, and query performance for high-volume tracing data not addressed.
  • Implications: Operators should invest in metadata cataloging (table/column descriptions, PII tags) so agents have semantic context when querying, reducing hallucinations.; Architect tracking data pipelines early: unify tracing from multiple agent frameworks and clouds into one queryable store (Delta Lake, data warehouse) to serve auditors, dashboards, and eval pipelines.; Plan for incremental data loading and versioning to catch stale embeddings or outdated retrieval sources before they cause production failures.; Consider data governance (who owns customer embeddings, policy docs, tracing logs) and privacy (PII in traces) as part of the foundation, not bolt-on later.

Pillar 4: Multi-Agent Orchestration Patterns

  • Claims: One agent is simple; five agents increase complexity exponentially due to coordination, dependencies, and failure modes.; Three orchestration patterns: (1) orchestrator-worker (centralized control, all requests through orchestrator), (2) choreography (autonomous agents, message bus, parallel execution), (3) human-in-the-loop (fallback when confidence drops below threshold).; Pattern choice depends on task dependencies and latency needs; choreography reduces latency for independent parallel tasks, orchestrator provides central logging for debugging.
  • Evidence: Orchestrator-worker: one orchestrator distributes work to specialized agents, all logs go to orchestrator, good for debugging but potential bottleneck.; Choreography: agents listen to message bus for events (e.g., mortgage app triggers), work in parallel (one agent gets customer details, another approval details), reduces latency but harder to trace.; Human-in-the-loop: when agent confidence falls below threshold, route to human to review agent's work, take action, avoid low CSAT.
  • Caveats: Speaker does not provide decision criteria or failure mode analysis (orchestrator single point of failure, message bus ordering/idempotency).; Refers to separate deep-dive video on fault tolerance patterns (saga, compensation, circuit breaker) but doesn't detail them here.; No discussion of state management complexity, distributed transaction consistency, or observability challenges in choreography pattern.
  • Implications: Operators should choose orchestrator-worker for early-stage multi-agent systems where centralized debugging is critical, then migrate to choreography for latency-sensitive parallel tasks.; Implement fault-tolerance patterns (circuit breaker to limit retries, compensation for rollback, saga for long transactions) as agent count scales to prevent cascading failures.; Use human-in-the-loop as safety valve for low-confidence responses, not as excuse to skip evals—monitor handoff rate to detect agent degradation.; Plan for state management (where is conversation state stored, how long is it retained, how do agents share context) and tracing across distributed agents before deploying multiple agents.

Pillar 5: Governance (Regulatory, Security, Change Management)

  • Claims: Governance is not just data governance; it's audit trails for every action, PII validation, prompt versioning as code, model change management.; Pre-validation (NER, regex for PII) caught 47 breaches during testing in the banking project.; Prompt changes must go through enterprise change management like code, not casual git commits, with detailed messages explaining what failure is being fixed.; Model providers' benchmark boards don't reflect enterprise-specific performance; teams must test new model versions on their own eval datasets before switching.
  • Evidence: Banking case: named entity recognition and regex for PII detected 47 breaches before production.; Prompt versioning: commit messages are often generic, but must document failure cause and expected fix so teams can trace why changes were made.; Model change management: can't rely on one model due to risk; must have flexibility to switch models and test them on internal eval datasets, not public benchmarks.
  • Caveats: No specifics on enforcing detailed commit message standards or integrating prompt versioning with ITSM tools.; Doesn't discuss model fallback strategies, A/B testing for model changes, or rollback procedures if new model degrades performance.; Regulatory focus (audit trails, PII) may be less critical in non-regulated industries, though observability and accountability still apply.
  • Implications: Operators must treat prompts as first-class code artifacts with version control, change approval workflows, and documentation linking changes to incidents.; Implement PII validation at data ingestion (not just at agent output) to catch compliance breaches early and avoid costly remediations.; Test model upgrades (e.g., GPT-4 to GPT-4.5) against internal eval datasets before production cutover; vendor benchmarks (MMLU, HumanEval) are not proxies for business-specific accuracy.; Design incident playbooks (who to alert, ITSM integration, fallback strategies) before production, not during 3am outages, to reduce mean time to recovery and reputation damage.

Case Study: Retail Banking Chatbot (Model Selection in Week 7)

  • Claims: Customer had 20K monthly calls, wanted to deflect 60% of simple queries to AI agent, spent $85K on failed 6-month POC before Sandy's team.; New approach: weeks 1-2 collect 200 human-agent answer examples, define 85% accuracy and 60% deflection targets; weeks 3-6 build data foundation and tracing; week 7 test models on eval dataset.; Post-launch success: caught policy document staleness issue via tracing and CSAT drop within weeks, fixed by updating vector DB embeddings.
  • Evidence: Failed POC symptoms: no visibility into AI decisions, no measurement system, no accountability.; New POC structure: eval system built first, automated pipeline comparing agent responses to golden dataset, rating system with human review for low scores, test cases added to dataset for next iteration.; Foundation work: API connections traced, distributed storage set up, duplicate API calls detected during testing (agent made 3 calls for one balance query).; Post-launch issue: interest rate policy changed, notifications sent to customers, but agent retrieved outdated policy doc from vector DB, causing wrong answers and CSAT drop—detected via tracing, fixed by updating embeddings.
  • Caveats: No final cost, timeline, or quantitative post-launch metrics disclosed (CSAT improvement, deflection %, cost savings).; Speaker does not explain how they chose 200 as initial test case count or how they prioritized which 60% of queries to target.; Customer org and POC vendor not named (likely NDAs), limiting ability to verify claims or compare approaches.
  • Implications: Operators should defer model selection until eval and data systems are in place, so the decision is data-driven, not driven by vendor hype or executive preference.; Living eval datasets are critical: the policy doc issue would not have been caught without tracing + CSAT monitoring + eval pipeline that could quickly add new test cases.; Detecting issues post-launch requires the full system (tracing, CSAT tracking, automated eval) to be operational, not bolted on after complaints escalate.; Spending weeks on eval and data before coding is counterintuitive but de-risks production: the 'slower' start enables faster, safer scaling.

Three Commonly Missed Lessons

  • Claims: Test case library is a living, growing system requiring governance: owners, categorization by failure type (security, logic, PII), documentation.; Prompt versioning commit messages must be detailed, explaining what failure caused the change and what it will fix, not generic 'updated prompt' messages.; Layer 3 behavioral evals (tool call checks) are expensive at scale; run subsets during CI, full tests only on main branch merges to control costs.
  • Evidence: Test case library grows from 200 to hundreds as production uncovers edge cases; categorizing (e.g., security issues, login failures) helps trace what changed.; Prompt changes in git often have simple commit messages, but teams need to understand why the change was made—what failure it addresses—for root cause analysis.; Running layer 3 evals on 300-500 test cases during every CI run costs money; selecting subsets (e.g., 50 cases) for quick checks, full suite for merges, controls cloud spend.
  • Caveats: No cost benchmarks, subset sizing heuristics, or tooling recommendations for test case governance.; Doesn't discuss how to enforce commit message standards (code review, linting, templates) or integrate with existing change management systems.; Layer 3 eval cost control strategy (subsets in CI) may miss edge cases; no guidance on which subset to run or how to rotate coverage.
  • Implications: Operators should assign test case library owners, use tagging or folder structure to categorize by failure type, and review growth patterns to identify systemic issues.; Enforce prompt commit standards via templates, code review, or linting; link commits to incident tickets (Jira, ITSM) so changes are traceable to business impact.; Design CI eval budgets: run deterministic checks (layer 1) always, LLM judges (layer 2) selectively, behavioral checks (layer 3) on main merges or scheduled nightly runs, to avoid cloud bill surprises.; Living datasets + version control + cost controls together enable continuous improvement without runaway costs or lost context.

Production Incident Playbook

  • Claims: When AI fails in production, follow: detect (via dashboards), diagnose (via tracing), contain (rollback prompts, fallback to human, apply circuit breakers), fix (using test case library and LLM judge reports), then add test cases to eval dataset.; Playbook must integrate with ITSM systems (PagerDuty, ServiceNow) to alert the right person at the right time and prevent downstream system impact.; Without this playbook, teams lack accountability and waste time during 3am incidents.
  • Evidence: Detect: use dashboards monitoring CSAT, latency, deflection rate.; Diagnose: trace logs show which step failed (intent classification, API call, retrieval, reasoning, guardrails).; Contain: version prompts so you can roll back; use fault-tolerance patterns (saga, compensation, circuit breaker) from multi-agent orchestration video; deflect to human if confidence drops.; Fix: review LLM judge reports, eval dataset results, identify root cause, apply fix (prompt change, data refresh, tool call logic update).; Test: add incident as new test case in eval dataset so it's caught in future CI runs.
  • Caveats: Speaker does not detail ITSM integration mechanics, alerting thresholds, or escalation paths.; Fault-tolerance patterns (saga, compensation, circuit breaker) mentioned but not explained; referred to separate video.; No discussion of incident retrospectives, blameless postmortems, or knowledge sharing across teams.
  • Implications: Operators must define the playbook before production, not during incidents: who is on-call, what dashboards to check, what fallback strategies to apply.; Integrate AI monitoring with existing ITSM/alerting systems so AI incidents trigger the same escalation paths as infrastructure/app failures.; Treat every production incident as a test case: document root cause, apply fix, add to eval dataset, verify fix with automated tests, close the loop.; Without the playbook, teams will panic during incidents, apply ad-hoc fixes, and repeat the same failures—defeating the purpose of observability and eval systems.

Notable Concepts & Terms

  • LLM-as-Judge: Using a secondary LLM to evaluate the primary LLM's output for safety, groundedness, relevance, by providing a judge prompt specifying criteria; automates semantic evaluation at scale (layer 2 evals).
  • Layer 3 Evals (Behavioral Evaluation): Beyond deterministic (formats) and semantic (groundedness) checks, behavioral evals analyze agent tool call efficiency, duplicate API calls, loops—critical for cost control at production scale but expensive to run.
  • Golden Dataset / Eval Dataset: Collection of real-world questions with expected answers from domain experts (e.g., 200 human-agent responses), used to benchmark agent performance; treated as a living system that grows with production incidents.
  • Orchestrator-Worker Pattern: Multi-agent orchestration where one orchestrator distributes work to specialized agents, provides centralized control and logging, but can be a bottleneck; contrasts with choreography (message bus, parallel agents).
  • Choreography Pattern: Multi-agent orchestration where autonomous agents listen to a message bus for events, work in parallel without central orchestrator, reduces latency but harder to trace and debug.
  • Unity Catalog (Databricks): Centralized data catalog providing permissions, metadata tagging, discovery, delta sharing; enables agents to query tables with semantic context (PII tags, column descriptions) to improve groundedness.
  • Delta Lake: Open-source storage layer that brings database-like ACID properties (incremental loading, versioning) on top of raw data files (Parquet, JSON) in cloud storage; foundation for Databricks' data management.
  • Circuit Breaker Pattern: Fault-tolerance pattern that limits retries (e.g., max 3 attempts) to prevent cascading failures or runaway API calls; mentioned as part of multi-agent orchestration but detailed in separate video.
  • Question Data vs Tracking Data: Sandy's distinction: question data = data needed for agents to answer user queries (pre-training, retrieval, APIs); tracking data = tracing logs for observability, audits, monitoring—both need data strategy.
  • Deflection Rate: Percentage of user queries handled by AI agent without human escalation; key business KPI for chatbots (e.g., 60% deflection target = 60% of simple queries answered by agent, reducing support costs).
  • Agent Bricks: Databricks product for building production AI agent applications, providing out-of-the-box LLM judges, tracing, evaluation, monitoring; Sandy positions it as operationalizing the five-pillar framework.
  • Production Incident Playbook: Framework for handling AI failures: detect (dashboards), diagnose (tracing), contain (rollback/fallback), fix (test case library), close loop (add test case to eval dataset)—must integrate with ITSM.

Operator Notes / Why Ken Should Care

  • This is a practitioner's playbook, not vendor fluff: Sandy spent 60% of project time on data foundation, deferred model selection to week 7 of 8, and caught 47 PII breaches pre-launch—signals real trench experience.
  • For agent system operators: adopt the five-pillar sequence (eval → observability → data → orchestration → governance) as your design checklist before writing code; resist exec pressure to ship demos without tracing and eval pipelines.
  • For AI ops teams: the three-layer eval framework (deterministic → semantic → behavioral) is actionable; implement layer 1 (regex, PII) and layer 2 (LLM-as-judge) first, then add layer 3 (tool call checks) with cost controls (CI subsets, main branch full tests).
  • For content/workflow ops: the living eval dataset concept applies beyond AI—treat every production incident as a test case, categorize by failure type, version control with detailed commit messages linking to incidents.
  • For investing/due diligence: ask SaaS/AI companies if they version prompts as code, have incident playbooks, separate question data from tracking data strategies, and test model upgrades on internal evals (not public benchmarks)—gaps indicate production risk.
  • For GTM/product teams: deflection rate and CSAT as KPIs (not just accuracy) align AI metrics with business outcomes; tracing and eval systems are table stakes for enterprise buyers, especially in regulated industries (Europe, finserv).
  • Critical gotcha: layer 3 behavioral evals (tool call checks) can cause runaway cloud costs if run on full datasets during every CI run—architect eval budgets and subset strategies before scaling.
  • Multi-agent orchestration complexity is exponential: one agent is simple, five agents require state management, fault tolerance (circuit breakers, saga, compensation), and centralized tracing—don't skip pillar 4.
  • Data quality is the bottleneck: 'agents don't forgive' stale/bad data; invest in metadata catalogs (table/column descriptions, PII tags) and versioning (embeddings, policy docs) so retrieval is traceable and agents have semantic context.
  • Governance is not optional: regulators (especially Europe) block production without tracing/audit trails; prompt versioning and model change management reduce vendor lock-in and enable rapid response to model drift or failures.

Watch Map

  • 00:00: Intro: speaker background (Databricks, AWS principal architect, 5 years data/AI at scale), session goal (production AI playbook from trenches)
  • 01:30: The failure pattern: orgs start with model selection, build demos, get leadership buy-in, then face mysterious production failures (no ROI, wasted $85K example)
  • 03:45: Three critical gaps: observability (can't see AI decisions), evaluation (can't measure success), governance (no accountability at 3am)
  • 05:00: Five-pillar framework intro: evaluation → observability → data → orchestration → governance (design before code, ideally in sequence)
  • 06:20: Pillar 1: Evaluation as specification—define success numerically (85% accuracy, 60% deflection), build golden dataset from domain experts, automate eval pipeline
  • 08:50: Three eval layers: deterministic (formats, regex), semantic (LLM-as-judge), behavioral (tool calls, duplicate APIs)—layer 3 catches expensive patterns at scale
  • 11:40: Pillar 2: Observability/tracing—capture every agent decision (intent, API calls, retrieval, reasoning, guardrails) for regulators and debugging; banking chatbot trace example
  • 14:30: Pillar 3: Data foundation—60% of project time; question data (answering queries) vs tracking data (tracing logs); agents don't forgive stale data; Unity Catalog example
  • 18:15: Pillar 4: Multi-agent orchestration—one agent simple, five agents exponential complexity; orchestrator-worker, choreography, human-in-the-loop patterns; fault tolerance (circuit breaker, saga, compensation) in separate video
  • 20:45: Pillar 5: Governance—audit trails, PII validation (47 breaches caught in banking case), prompt versioning as code, model change management independent of vendor benchmarks
  • 23:10: Case study: retail banking chatbot—20K calls/month, 60% deflection goal, $85K failed POC; new approach: eval first (weeks 1-2), data foundation (3-6), model selection week 7
  • 26:30: Case study results: post-launch, tracing caught policy doc staleness (embeddings not updated), causing wrong answers and CSAT drop—fixed via eval pipeline and data refresh
  • 30:20: Three missed lessons: (1) govern growing test case library with owners/categories, (2) prompt commit messages must explain failure context, (3) control layer 3 eval costs via CI subsets
  • 33:00: Production incident playbook: detect → diagnose → contain → fix → add test case; integrate with ITSM; treat every incident as eval dataset entry
  • 35:00: What to do tomorrow: define success, collect example answers, build automated eval pipeline; QR codes for downloadable artifacts (eval checklists, tracing setup, templates) and LinkedIn newsletter

Source/Metadata

  • Title: The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks
  • Transcript words: 7212
  • Duration seconds: 2225
  • Timestamp note: Timestamps manually reconstructed from transcript structure; actual video timestamps may differ slightly. QR codes mentioned for artifacts (eval templates, tracing guides, LinkedIn newsletter).

Transcript

6002 words en Processed in 359.4s

All right, thank you for joining my session. Thank you, mate. I'm Sandy. I'm a technical lead for data and AI at Databricks. Prior to working in Databricks, I worked in Amazon Web Services for five years as a principal architect for data and AI. In the past few years, I've worked extensively building and scaling data and AI platforms using distributed systems and technology. And in the past couple of years, specifically, I've been working with customers trying to figure out what we do with this new AI technology. When I say new AI, AI has been here for a long time, but we all started experimenting quite exponentially in the past couple of years. And I have learned a great deal of lessons from building demos and how to take those demos to production, working with different customers in B2B software industries and then regulating industries like financial services. So in this session, I want to share a playbook, a framework that I put together from lessons that I have learned working in the trenches that you can take and apply when you think about how to put your AI systems into production. And I think this session is nicely placed in the afternoon because what you can do now is fit the different knowledge that you have gathered attending these different sessions throughout the day and see where they fit in each of these elements in the framework. So when I started two years ago, this is the pattern I noticed in every customer conversation. Everyone wanted to do something with AI. There was immense pressure from the top to do something, to build a demo. And every conversation started with, let's choose the model. And it was nobody's fault because the market was like that. We were talking about models. The models were new technology for us. And every conversation started, shall we use GPT? Shall we use Claude? There was huge debate within organizations. Then you would choose a model, you'd build some features, debates over what features to build for that application. You would build it in a controlled environment with predictable data sets and limited scenarios. And then it looked great as a demo. And then leadership would get happy. They would sign it off and they'd put it into a production environment. Then after a few weeks, people would start asking questions about what the AI is doing. Why is it not answering the questions the way we expected it to answer when we were doing the demos? It would result in not only no realization in return on investment, but also loss of money and effort in building these demos that can never scale to production. Throughout these meetings, I gathered three insights that connect to everything that we are talking about when thinking about taking AI to production. The first one is the observability gap. When we use AI and put it into production, if we can't see what it is actually doing, if we can't trace every decision that it's making, it's no use in production. Second is the evaluation gap. A lot of these conversations that we were doing, we were not actually thinking about what is that one thing that we are measuring? Yes, we talk about accuracy. We talk about latency. We talk about groundedness. But we were not defining what is that exact thing that matters to the business and how can we build a system that can continuously measure that, whether it's improving or not improving. What is that system that we need to build? And that was the evaluation gap that I noticed. And the third is the governance gap. We were not actually thinking what happens when AI fails in production? Who's accountable? Who do I go to when something happens at 3am in the morning? Who needs to own the data assets that feed some AI responses? What happens if AI talks nonsense to a customer? So there is no accountability, no governance around it. And these three insights led me to build a framework on how I think AI should be taken to production. And this has been implemented across multiple customer organizations. And I think this is something that you can pick up from here. These are the five pillars. And these are absolutely what you need to think about even before starting a project. Then you start and build them gradually, preferably in sequence, but in real life I know that this sequence doesn't work. But these are the pillars that you have to know about and you have to think about when you start building. First one is evaluation. Before touching any code, before discussing any models or features, you have to think about when we build this system, how do we measure? What does success look like? What is that system that will help us continuously measure what success looks like for us? Second is how do we trace each and every decision that AI makes? It's not only important for the performance of the AI system. It is also important for the regulators. In Europe, in a lot of companies, especially in regulated industry, you cannot even onboard AI into production without having tracing and observability in place. So this is a must have. The third is the data foundation. I think of data foundation in two ways. One is the question data, so that is the data needed for the AI to answer questions that users ask to it. So it could be your pre-training data, post-training data, data that you use APIs to hook onto and get to the answer that the user needs. The other one is the tracking data related to the tracing data in observability. But when you think from the data foundation and data strategy perspective, this needs to be handled in this pillar because you need a whole data strategy now with tracing data, especially when you run hundreds of agents in your organization. Fourth is orchestration. One agent would work very well. You don't need to think about orchestration. But when you onboard five agents, the complexity increases exponentially. You will have multiple coordination patterns between these agents. They will need to talk to each other in multiple different ways. They will need to each wait for each other's responses. There's a lot of complexity that comes in. And that's where orchestration patterns and thinking about how you will orchestrate your agents in a particular system becomes really important. especially when you run hundreds of agents in your organization. Fourth is orchestration. One agent would work very well. You don't need to think about orchestration. But when you onboard five agents, the complexity increases exponentially. You will have multiple coordination patterns between these agents. They will need to talk to each other in multiple different ways. They will need to each wait for each other's responses. There's a lot of complexity that comes in. And that's where orchestration patterns and thinking about how you will orchestrate your agents in a particular system becomes really important. Fifth is governance. This is where you think about what happens when something fails. Who's accountable? How do we govern data? How do we secure it? How do we secure our systems? How do we make sure that no one injects into our agent and leads to misbehavior or loss of repetition? So in the rest of the session, I will dive a bit deeper into each of these pillars and tell you how you can think about when you start working with them. The first one is evaluation. Evaluation is the specification for your AI system. You define success. As I mentioned, it's not about talking about accuracy. You have to define it with numbers, what accuracy is good for your business use case. Define it in numbers. Define what kind of false positives you can handle. What should be the deflection? This is an example from a retail chatbot, a banking chatbot, where when you implement a chatbot with an AI agent, one of the main goals is to deflect simple queries to the agent so that a human agent doesn't need to deal with them. So you need to track those queries and track those numbers and put that system in place. Second is building those test cases like the evaluation data set. You have heard about golden data sets in evaluation. Talk with the domain experts and find what is actually happening in real life on the ground. What answer would a human support agent give to a customer on a particular question? Collect that information. What happens in gray areas, in edge cases? What happens when a human sees a customer asking a confusing question? Collect those into a data set and then automate your AI testing. So you put a question to AI, it answers, take that answer, compare against the test set, and automate this whole pipeline so that when you put AI in production, that pipeline can actually take live responses and evaluate against the test data set that you are building and then give you the result in terms of how AI is performing against those numbers and the goals that you have defined. When we talk about evaluation, there are three main layers that I see appear across organizations. And this is an architectural decision that you need to make when you build these evaluation systems. The first layer is deterministic. These are the easy stuff, checking formats, checking email formats, phone formats, the regular expression things that we have already been doing with our coding systems. The other is you could use classic ML models for name entity recognition for intent classification, for understanding what is a first name, last name, PII detection, et cetera. So these are easy stuff, cheap stuff. You should get them out of the way. We have already been doing this for years. The second layer is the non-deterministic, semantic stuff. This is where groundedness comes in. This is where we implement technologies like LLM as judges. We all know what LLM as judges are. Okay, I see a lot of nods. So again, this is a pretty simple version of how a prompt would look for an LLM as a judge. With LLM as a judge, you use a separate LLM from the primary LLM to judge the response of the primary model. And when you do that, you tell the secondary, the judge model on how it should judge the primary model's output. So it could be around safety, groundedness, relevance to the answer, et cetera, et cetera. Again, that can feed from a lot of the evaluation data set that you have created. To look at what are the expected answers, and then it can check against that. There is a sample prompt on how these things work, but I'm sure you have attended some of these sessions where you have seen vendors doing this automatically at scale. For example, in Databricks, in MLflow, you will find automatic LLM as judge, where you can create custom LLM as judges that run automatically on traces. That's your second layer. The third layer is behavioral. This is where you think about a tool cause. Are agents calling the right tool? Are they getting into loops? For example, the first layer, you can have a user ask the question, what is my account balance? And you could go and check that, okay, there is no deterministic problem with it. The agent answered right that your account balance is this many dollars. And that was right. And you can say this is right. But when you go into the behavioral checks, you will see that the agent actually made three calls to the database to find that answer. And that is because it was making duplicate calls for whatever reason. Calls failed. Calls did not work. It went and retried and stuff. Now, three API calls in demo environment is fine. But in production, when you get thousands of queries from users every day and there's duplication in API calls, that's an expensive operation. And that's where you need to think about behavioral evaluation. And this layer is very, very important. I see a lot of organizations or a lot of teams miss them when talking about this. The second layer is observability. And in this pillar, what we are talking about is tracing. So you collect all the decisions that an agent is making. I want to explain this with a scenario here. And this is a scenario from an actual project I worked on with a retail banking chatbot. Now, if you have seen tracing data, it's not as beautiful as this slide. And that's where you need to think about behavioral evaluation. And this layer is very, very important. I see a lot of organizations or a lot of teams miss them when talking about this. The second layer is observability. Right? And in this pillar, what we are talking about is tracing. Right? So you collect all the decisions that an agent is making. So I want to explain this with a scenario here. And this is a scenario from an actual project I worked on with a retail banking chatbot. Now, obviously, if you have seen tracing data, it's not as beautiful as this slide. Right? So I have simplified it and made it beautiful for this slide. But what this slide says is a user comes in and says I have been charged an overdraft fee. Can you waive it for me? Because the user thinks that the customer thinks that is not legitimate. So the agent does an intent classification. And you all know about this because you have enabled observability. You're capturing traces and you're actually seeing what the agent is doing. Right? What AI is doing. Intent classification. It has done this. It took this many seconds. This was confidence score. Then it goes and connects to the customer's account, maybe in a database, a customer database. Calls an API, connects to the customer database, gets the account details. It retrieves policy documents. It checks from a rag vector database. It says, what is the policy around overdraft? Right? [SPEAKER_00] Is what the customer claiming legitimate? So it checks for policy documents. Then it goes and does a reasoning on what should be responded to the customer. And then it does final guardrail checks and responds to the customer. Now, if you did not set up a system that helps you look and visualize all of these traces, when the customer comes to you and raises a dispute, you have no way to check what the AI did. Right? You have nowhere to go. And you end up saying that I don't have any idea. Let's give the customer a discount or something and then make them happy. So this is why you need this. And this is why regulators are basically mandating. Because otherwise there is no production system if you cannot do this kind of stuff. So this is where you detect the example that I gave around duplicate API calls. This is where you start detecting this stuff. So when you enable these traces, you can actually go and see duplicate calls and then take relevant actions based on that. Not only that, you can actually do that in online monitoring. So when it's happening in production, you can set up online monitoring and at that point if it is doing duplicate calls, you can apply fallback strategies. Or even if it is doing a call that is failing, you can actually go and apply a strategy where it will say, okay, go and retry for three times, not more than three times. If it is more than three times, then report somewhere or pass it to a human to take some action. The third pillar is the most important pillar in my opinion is the data foundation, right? [SPEAKER_00] In my typical projects, I spend 60% of my time and I see a lot of organizations spending a lot of time here. Because no one expected agents to come suddenly in the market and start acquiring data. Data was always built for humans and humans are always forgiving. You find the wrong data in a report, you just go and ask someone to correct it. Agents don't forgive you, right? Agents will go find it wrong, they'll give you the wrong answer confidently. And you wouldn't know what's happening. And this is why data quality, setting the right data strategy has become so important for enterprises now. I divide it into two sections. One is the question data as I was explaining, data needed for actually serving the AI's outcome. And the other one is the tracking data. This is the observability data, the tracing data I was talking about earlier. You need a proper plan on how you collect this tracing data and how you serve it to auditors, to regulators, to do online monitoring, to run LLMS judges on the tracing and everything else. So it needs a proper strategy on how you structure the schema on the tracing data. On Databricks, we create a robust data foundation for our customers using some of the technologies that we provide. If you don't know Databricks, Databricks has been built on some open source technologies like Apache Spark, MLflow, and Delta Lake. We provide a bunch of capabilities on top of it. So the blue layer at the bottom is basically your cloud storage. Databricks works on the three major clouds, Google, AWS, Azure. And then you go into the data. Okay. I thought it was for me. So once you store raw data on your cloud storage, the data is then brought in a layer called the Delta Lake layer, which basically brings in database-like properties on top of your raw data. So you have got images, text files, video files, whatever. We help you create this table-like structure on top of it using manifest files. Right? And we help you to incrementally load data, do all of those data management tasks in a structured way. On top of that, we bring in Unity Catalog, which is a data catalog. With Unity Catalog, you can centrally apply permissions on top of the data. You can share the data using Delta sharing. But also, what happens with Unity Catalog is you can enable discovery and ownership, metadata tagging capabilities at the catalog level. What that means is when you apply table description, column description, tag columns, like PII columns with metadata, it becomes really easy for AI to then get that context when it queries these tables on top of the Unity Catalog. So everything is governed at one layer through Unity Catalog. And on top of that, we bring in different applications. So whether it's AI through Mosaic AI, to build LLM, tune LLM, or even build AI applications. We bring in data warehousing capabilities, BI capabilities, and some of the other text-to-SQL capabilities. We have got Genie that helps you write natural language to do SQL query, et cetera. And one application of that in the observability and tracking data that I was showing is this. So basically, think about when I was talking about the tracking data strategy. Organizations, especially enterprises, will not be running AI in just one framework. They'll be using different frameworks, QAI, Langchain, et cetera. They'll be using different cloud platforms. And once they do that, you need a centralized layer of collecting that tracing data so that you can serve several use cases on the right-hand side. We have got Genie that helps you write natural language to do SQL query, et cetera. And one application of that in the observability and tracking data that I was showing is this. Think about when I was talking about the tracking data strategy. Organizations, especially enterprises, will not be running AI in just one framework. They'll be using different frameworks, QAI, Langchain, et cetera, et cetera. They'll be using different cloud platforms. And once they do that, you need a centralized layer of collecting that tracing data so that you can serve several use cases on the right-hand side. So whether it's for operational dashboarding, for first-line support, a lot of this first-line of defense teams need health monitoring dashboards, right? These teams can also write SQL using Databricks Genie to do text-to-SQL. But they can also build Databricks apps using coding agents to create common workspaces or custom UIs that customers might need for different use cases. And then we've got Agent Bricks and MLflow that serves out-of-the-box LLMS judges and proactively monitored drift. The idea is no matter where your AI runs, you can create this strategy, bringing in data in one common place and serving different teams from one shared location. The fourth pillar is multi-agent orchestration patterns. As I said, one agent is good, multiple agents increases complexity. That's where you start thinking about what pattern is good for my use case. The first one I describe here is the orchestrator-worker pattern, where you have one orchestrator which orchestrates all the work, which controls all the work from a centralized plane, and then distributes this work to different agents based on their specialized skills. And then every request goes through the orchestrator, so you have got central control. If something goes wrong, you can go to the orchestrator logs and look into them and see what has happened. So that's the orchestrator pattern. There is this choreography pattern where each agent is independent, they're autonomous, they don't depend on an orchestrator. [SPEAKER_00] All of them talk to a message bus, and they listen to the events that they are interested in. So think about agents that are independent of each other, right? They can run parallely. So they are not sequential. One agent is not dependent on another. So they run parallely, they listen to the message bus for the events that they're interested in. Maybe it's a trigger for a mortgage application. And it says the mortgage application agent, one of the agents looks at customer details, right? The other agent looks at approval details and everything else, right? They can work in parallel, and the advantage it brings you is the latency is reduced because they are not dependent on an orchestrator and sending messages back and forth. So this is the choreography pattern. And the third one is human in the loop, which is where when an agent crosses a threshold or serves below threshold, a confidence threshold, then a human is calling the workflow to look into the pattern. So look into what the agent has done and then take action based on that. I have done a dive deep video on multi-agent orchestration pattern for the online track of this conference. It's already on YouTube, so you can look into it. I talk about the real implications of when you think about multi-agent patterns. One is state management. The other is fault tolerance, what happens when things fail. How do you manage them? I talk about different patterns. And then I talk about how you think about scaling them in large scale enterprises. Pillar five is governance. Now here I'm not talking about data governance at all. That's given. We need that. From AI perspective, what are we thinking about? Regulatory. Audit trails. Have we got the trail of every action, every user connection, every request, everything that happens in the system? Are we capturing everything? Are we doing pre-validation of personal information? Are we using named entity recognition, the easy stuff, the rejects and all of those things? In our example, the work that I was doing with the customer that I mentioned, we already detected 47 PII breaches during the testing phase by applying this layer. So that's really important. Fourth is prompt versioning. You have to treat prompt versioning as change management in enterprise-grade solution. It cannot be just a change to a prompt and commit to get. It has to go to proper change management processes as you do with code. So treating prompt as code. Third is model change management. So as models change, the model providers upgrade these models. You have to have a system to understand whether that upgraded model will be good for your use case, for your data. Model providers will put evaluation benchmarks on three benchmark boards. But those are not really useful when you put them in your context in your enterprise. So that's where these evaluation data sets come in handy, where you try these different models on this evaluation data set and try to understand which one performs better. And that management needs to be done. Because from a risk perspective, you cannot really rely on one single model. You have to have the flexibility to switch to different models and also test them on your own data. That management needs to be done. In Databricks, we have taken all of these points, these pillars that I have been talking about into Agent Bricks. We are building Agent Bricks to make all of these operations out of the box for you. So that it's easy to implement production-grade AI applications in your enterprises. So I wanted to quickly touch upon a case study just to give you a flavor of how these things go. So when I was working with this client, they were retail banking, they were building a retail banking chatbot, one and a half, 18 months ago. Their problem, the problem they wanted to solve is they had around 20,000 odd calls per month from customers on their chatbot. They wanted to deflect, they saw that there were 60% of them were simple queries. What is my account balance? What do I do with my overdraft and all of those stuff, right? That can be answered simply. So they wanted to reduce reliance on human agents for those answers. So they identified those queries and they wanted to automate them, right? They spent around 85K in six months doing a POC, which did not succeed. When we got involved, we found those insights that I was talking about like no one knew why things were failing when it was in production. No one could actually measure why it's not succeeding and no one could actually understand who is accountable for what when things go wrong, right? What is my account balance? What do I do with my overdraft and all of those things, right? That can be answered simply. So they wanted the reliance on human agents for those answers. So they identified those queries and they wanted to automate them, right? They spent around 85K in six months doing a POC, which did not succeed. When we got involved, we found those insights that I was talking about. No one knew why things were failing when it was in production. No one could actually measure why it's not succeeding and no one could actually understand who is accountable for what when things go wrong, right? So the goal we set for them is AI agent handles 60% of user queries, right, which were simple user queries and then a way to identify and track them. The key difference in this project that we did is that we selected the model in week seven, in the eight weeks POC, right? And this is how it turned out. For week one and two, we built the evaluation layer. We collected 200 cases on their actual human agents answering to their customers on simple queries and understand how they are responding to them. We created that database. Then we defined the success metrics. What does success look like to you? So out of, let's say, 100 queries, you need 60 queries. That's 60% of the queries that are simple queries to be handled by the agent, right? They needed some accuracy. So 85%, it was around 85% accuracy target. They needed latency. All of the operational targets that you need, they were there. So we created this automated evaluation pipeline for them. And what I mean by that is an automated system where you can capture a user's question and the AI agent's response. You take that, compare that against your evaluation data set. You rate that. And if the rating is below certain threshold, you get it checked by a human. And if something goes wrong, you make sure that you find the solution. So it could be a change to the prompt. It could be changed to a tool calling system or something else. Once you have done that, you add that test case in the test data set. So that when it happens next time, the test cases catch them. So the summary of that story is that your evaluation data set is a living system. You start with 200. [SPEAKER_00] Maybe there's no correct number here. But once you start, as you start building in production, this is a living system. This will keep growing. And the bigger it grows, the better your system will be. In the second we talked about the foundation earlier. So the question data. We thought that, okay, if you have to call the database, have you got the API connections right? Have you got a system in place that can trace the API connections? Are those secure? We were not talking about MCP at that time. It was just direct API calls to database to run queries. Have you got the distributed storage? Are you collecting traces? And this is where when we started testing after building the systems, we could catch those duplicate API calls, right? We could catch why customer satisfaction was dropping and stuff like that. And then comes in week seven to eight, we started talking about models. Now that we had the evaluation data set, we could run different models on that data set to see the responses, compare them against the expected responses, and calculate a number on the accuracy. Right? That helped us to decide which model to use. Now that decision didn't take long, right? We, as I explained in the introduction, spent weeks debating on which model to use. But when you took the other approach, you can actually do that in a very quick way. So, once that's done, we stitched everything that I was talking about around observability, evaluation, the layers of evaluation. Once we had that system that can make AI visible, measurable, and accountable, that's when we started launching it to production. And that's when, so this is the result, six weeks post-launch, we calculated the operational metrics, of course, the accuracy, the deflection rate, the response time, the customer CSAT. But what's important here is, in a few weeks' time, when there was a problem with, so one of the things that happened was that the bank changed some interest rate-related policies. So, when they changed the policy, they actually sent emails to customers or notifications in the application, in the mobile banking app, about the policy change. But when the customers came and queried on the chat board for further questions, they couldn't get the right answers. And they were putting thumbs down on the answers, so they were getting this feedback, right? So, feedback decreased. The problem with this kind of system, if you did not have this measurement system, is that you couldn't actually know what's happening. But because we had the measurement system in place, the drop in CSAT was detected, right? Because we were getting negative feedback from customers. We could actually look into the tracing decisions and see that the agent was looking at a policy document that was outdated. So, the new policy document was not updated in the vector database. The embeddings did not come through. Because it did not come through, it was giving stale answers. And that's when we went and fixed that. But it was all possible because they built those systems that led us to detect this. Before you go, I generally in these sessions, I share different artifacts that you can take away. I have a QR code at the end for you to download. And you will find multiple artifacts. One of the important artifacts that I want to talk about is the production incident playbook. This is something that a lot of us tend to miss when we work in AI projects. And this playbook is a definition of what needs to happen when things fail in production. First, you detect using a valid dashboard. Then you diagnose using tracing, as I explained. Then you contain. So you are versioning your prompts. If there is a problem with the prompt, you take that prompt out, right? And start the changes. Deflect it to a human. Or in my multi-agent orchestration video, I've talked about multiple fault tolerance, failure recovery patterns, around saga pattern, compensation pattern, and circuit breaker pattern. But look into the video. I've explained them in detail on how you can handle them. And then you use the test case library to fix. So, you look into LLMS judge reports. You look into your evaluation data set reports. Then you fix your problem. Once you fix your problem, you put those test cases in your data set, right? And create that eval suite that is a living system that will keep growing. Deflect it to a human. Or in my multi-agent orchestration video, I've talked about multiple fault tolerance, failure recovery patterns around saga pattern, compensation pattern, and circuit breaker pattern. Look into the video. I've explained them in detail on how you can handle them. And then you use the test case library to fix. So you look into LLMs judge reports. You look into your evaluation data set reports. Then you fix your problem. Once you fix your problem, you put those test cases in your data set, right? And create that eval suite that is a living system that will keep growing. And you keep improving your AI system based on that, right? But this playbook needs to be in place. When it runs in production, you will need to integrate it with your ITSM system so that it alerts the right person at the right time. A lot of organizations would have existing ITSM systems, right? Which is useful for alerting and making sure that the downstream systems don't get affected, etc. So once you have this in place, you can stitch it together to other systems. What can you do tomorrow, right? Start with, if you have a project in mind, start with defining success. Success not from the technical sense, from the business sense, what it means for the business, right? Come up with a few examples of what good answers look like and create a data set of that. And then build that pipeline using simple Python code. See if you can automate that so that when you run AI and get some response, it can go and compare the answer against that data set, and then that can be delivered to the customer. Now, these are three lessons that I have learnt while doing these things, which you might easily miss. The test case library, as I explained, is a growing system. It will grow over time. And because it grows over time, you need some sort of governance around it. You need an owner, right? You need to figure out which test cases relate to what kind of problem so that whenever you go back to it, you can relate your answers to those problems. If it is a security issue, if it's a login, you can say that the agent did not ask for login credentials when the customer asked to answer. And all those kind of issues can be put under a security category within that data set. So categorize the rows in your data set so that you can pick up what changed and compare it with them. The second is prompt versioning. When you start versioning prompts using Git, we all know when you put a Git message, the commit messages tend to be simple commit messages. But you have to put governance around what kind of commit messages you are putting in when you are changing these prompts. Because you need to understand when a prompt was changed and for exactly what reason it was changed, right? What was the failure that caused this prompt to be changed? What kind of failure would it address and what would it correct in the next version? That needs to be documented. Otherwise, it becomes difficult because when you go back and look into prompt versioning and look at different versions and you cannot trace why those changes were made, then it becomes difficult to track what's happening. The third is layer three evals, right? The behavioral evals that I was talking about around tool calls and stuff like that can be really expensive as you grow your eval data set as well. So when you have a wrong tool call, for example, and you want to correct that system, when you correct it and run it against the eval data set, you have to run it against, let's say, if you have 300, 400, 500 rows in the data set, you have to run it against them. And you do all the testing again and again and again, that can cost you a lot of money. So you have to put some governance around that. For example, when in your continuous integration pipeline, when you do the prompt change, you can put some checks around just selecting a small subset of the eval data set to do the testing. And you only do the full test when you merge to the main branch. So you can put these kind of decisions in place so that you can reduce cost around expensive eval decisions. If you scan this QR code, it'll take you to a Google Drive link where I have put some examples on how these templates look like, what an evaluation checklist would look like. I've given you some guide on setting up tracing with open source technologies so that you can quickly set up some tracing and start testing in the test environment before you decide on what kind of tools you want to use. Thank you very much for listening to me. This QR code will take you to my LinkedIn profile. I have a newsletter where I share this kind of topic every week. So if you're interested, you can join. It's free. I basically share what I learn in the field working with customers. So it might be useful for you. Thank you very much. it can become, it can go and compare, you can go and compare the answer against that data set, and then that can be delivered to the customer. Now, these are three lessons that I have learnt while doing these things, which you might easily miss. The test case library, as I explained, is a growing system. It will grow over time. And because it grows over time, you need some sort of governance around it. You need an owner, right? You need to figure out which test cases relate to what kind of problem. So that whenever you go back to it, you can relate your answers to those sort of problems. If it is a security, if it's a login, so you can say that the agent did not ask for login credentials when the customer asked to answer. And all those kind of issues can be put under a security category within that data set. So to categorize the rows in your data set so that you can pick up what changed and compare it with them. The second is prompt versioning. Now, when you start versioning prompts using Git, you know, we all know when you put Git message, the commit messages tend to be simple commit messages. But you have to put governance around what kind of commit messages you are putting in when you are changing these prompts. Because you need to understand when a prompt was changed for exact what reason it was changed, right? What was the failure that caused this prompt to be changed? What kind of failure would it address and what would it correct, right? In the next version. That needs to be documented. Otherwise, it becomes difficult because when you go back and look into prompt versioning and look at different versions and you cannot trace why those changes were made, then it becomes difficult to track what's happening. The third, the layer three evals, right? So the behavioral evals that I was talking about around tool calls and stuff like that, they can be really expensive as you grow your eval data set as well. So when you have a wrong tool call, for example, and you want to correct that system, when you correct it and run it against the eval data set, you have to basically run it against, let's say, if you have 300, 400, 500 rows in the data set, you have to run it against them. And you do all the testing again and again and again and again and again, that can cost you a lot of money. So you have to put some governance around that. So for example, when in your continuous integration pipeline, when you do the prompt change, you can actually put some checks around just selecting a small subset of the eval data set to do the testing. And you only do the full test when you merge to the main branch. So you can put these kind of decisions in place so that you can reduce cost around expensive eval decision. If you scan this QR code, it'll take you to a Google Drive link where I have put some examples on some of these, how these templates look like, what an evaluation checklist would look like. So I've given you some guide on setting up tracing with open source technologies so that you can quickly set up some tracing and start testing in the test environment before you decide on what kind of tools you want to use. Thank you very much for listening to me. This is this QR code will take you to my LinkedIn profile. So I share I have a newsletter where I share this kind of topics every week. So if you're interested, you can join. It's free. I basically share what I learn in the field working with customers. So it might be useful for you. Thank you very much. Thank you very much.