AI Engineer

Shipping complex AI applications — Braintrust & Trainline

1185 summary words 5 min summary Watch video

Start with the signal

5 min read

Summary

30-second take

This is a hands-on workshop teaching attendees to operationalize AI applications using Braintrust's observability and evaluation platform, delivered by Braintrust and Trainline engineers at AI Engineer Europe 2026. The core thesis: most AI POCs fail in production not because models aren't smart enough, but because teams lack the operational rigor to ship, monitor, and iterate on non-deterministic systems at scale. The workshop walks through building a multi-stage support triage agent from scratch, instrumenting it with tracing, creating evaluation datasets (the "golden set"), deploying prompts/tools to managed infrastructure, scoring production traffic, and fixing discovered failure modes—demonstrating Braintrust's "flywheel" of continuous improvement. Trainline shares real-world context: they serve 27M users, run agentic travel assistants handling refunds and rebooking at scale, and use Braintrust to evaluate model switches, ship features faster, and enable cross-functional teams (product, non-technical SMEs) to self-serve insights without engineering babysitting.

Key takes

  • POCs fail due to ops gaps, not model quality: The gap between demo and production is operational rigor—treating logs as observability, patching prompts reactively, and lacking versioning/reproducibility. Traditional software engineering is deterministic (1+1=2); LLM systems require evaluation frameworks because outputs are probabilistic.
  • Multi-stage agents beat monoliths: Breaking agentic workflows into explicit stages (context collection → triage → policy review → reply drafting → escalation) makes debugging easier and mirrors microservices decomposition. More stages = more failure points, but also more precise remediation.
  • Trainline's scale demands structured evaluation: With 6.3B tickets sold and millions of AI conversations, Trainline uses Braintrust to evaluate model switches (cost/token optimization), run offline evals before launch, and track online performance post-deployment. This prevented regressions when switching cheaper models and accelerated feature shipping.
  • Golden datasets bootstrap the flywheel: Start with edge cases and expected outputs (even if imperfect) rather than waiting for perfect production data. Use deterministic scorers (cheap, always-on) for schema checks and LLM-as-judge scorers (expensive, sample 5-10% once stable) for nuanced quality like tone or policy adherence.
  • Managed prompts enable non-technical collaboration: Offloading prompts, tools, and parameters to Braintrust's managed mode lets PMs/SMEs change models or prompts via UI without code changes. Trainline cited this as critical for cross-functional teams to self-serve without blocking engineers.
  • Production traces reveal real failure modes: Online scoring on live traffic (100% sampling initially, then 5-10% for LLM judges) identifies edge cases golden datasets miss. Example: "Not urgent, but CFO can't export invoices before board meeting" misclassified as low-priority until production trace caught it.
  • Reproducibility and versioning unlock regulated industries: Braintrust's version history, diff views, and audit trails satisfy compliance needs (right to be forgotten, stress test reproducibility) critical in banking/capital markets. Prompts and evals are treated as code but managed collaboratively.

Useful details

  • Workshop artifact: GitHub repo with step-by-step cheat sheet, runnable checkpoints at each git tag, Slack channel AI Engineer Europe 2026 Braintrust-Workshop for support.
  • Tech stack: Node.js v22, PNPM, OpenAI API (GPT-4o-mini default), Braintrust SDK (supports Python, Go, Ruby, .NET). Uses make commands or raw pnpm scripts.
  • Agent architecture: 5-stage pipeline → context collector (deterministic) → triage specialist → policy reviewer → reply writer → escalation decider → final packager.
  • Scoring functions: Deterministic (category check, schema validation, escalation presence) run always; LLM-as-judge (customer tone rubric) samples 100% initially, 5-10% in steady state.
  • Trainline-specific use cases: Travel assistant handles refunds/rebooking with multi-agent workflow; disruption predictor (classic ML model) forecasts train delays. Both use Braintrust for evals and observability.
  • Braintrust scale: Customers push tens of millions of traces in short periods; near-instantaneous write/read via custom "Brainstorm" database built for semi-structured, rapidly evolving AI trace data.
  • Managed mode workflow: Run make setup-braintrust to sync local prompts/tools/scorers to managed infra. Change params in UI (e.g., swap GPT-4o-mini → GPT-4-mini), run make ticket-runtime-mode=managed to execute against updated config without code edits.
  • Remediation example: CFO invoice bug caught in traces → tighten prompt ("consider two facts...") → re-run eval → scores improve. Diff view shows exact prompt delta and performance lift.
  • Auto-instrumentation: Braintrust CLI announced for auto-tracing; SDKs wrap OpenAI/Anthropic calls with one-line wrap() function; decorators for Python.
  • Hiring callout: Trainline is hiring AI engineers.

Caveats / counterpoints

  • No discussion of cost trade-offs for managed mode: Workshop doesn't address whether Braintrust's hosted prompts/tools/evals introduce vendor lock-in risk or cost overhead vs. self-hosting.
  • Limited scope on agentic framework integration: Uses raw OpenAI SDK for simplicity; doesn't demo LangChain, AutoGen, or other agent frameworks despite claiming tool-agnostic support.
  • Golden dataset bootstrapping is hand-waved: The workshop assumes you can intuit 10 edge cases upfront. No guidance on systematic edge case discovery for truly greenfield apps.
  • Sampling rate guidance is vague: "5-10% in steady state" lacks detail on when/how to adjust, how to stratify sampling by severity/user tier, or what happens if you undersample rare but critical failures.
  • Reproducibility claims need regulatory validation: Audit trails are mentioned for banking compliance, but no specifics on SOC 2, GDPR, or stress test audit support.
  • No mention of alternative observability tools: Doesn't compare/contrast LangSmith, Weights & Biases, or Arize—positioning is "build vs. buy" without acknowledging competitive landscape.
  • Workshop friction: Multiple attendees joined late/lacked OpenAI keys; reliance on Slack channel suggests onboarding could be smoother. Some confusion around automation vs. automated setup.

Ken relevance

  • Agent ops framework maps to your workflow: The flywheel (trace → eval → remediate → monitor → repeat) is directly applicable to your agent systems. If you're not tracing production agents end-to-end with nested spans, you're flying blind.
  • Managed prompts solve collaboration pain: If non-technical teammates (or you in rapid iteration mode) need to tweak prompts/models without redeploys, Braintrust's managed mode is a forcing function for operational discipline.
  • Evaluation urgency for content/GTM: Your AI-generated content workflows need deterministic scorers (tone, schema, policy) and LLM judges (helpfulness, brand voice). Start with golden datasets of edge cases you've already hit.
  • Trainline's scale insight for portfolio/investment: 27M users, billions in GMV, agentic systems handling refunds at scale = proof LLM apps can work in high-stakes, high-volume settings if ops infra exists. Look for companies solving eval/observability bottlenecks.
  • Cost optimization via model switching: Trainline's use case (eval to validate cheaper model swaps without regressions) is a direct cost-saving lever. If you're burning OpenAI credits, Braintrust evals de-risk downgrades.
  • Regulated industry positioning: If you're building for finance/healthcare, the versioning/audit story is a wedge. Braintrust's founder previously built ML at Figma (Impera acquisition), so pedigree is credible.
  • Fast-follow on CLI/auto-instrumentation: Eric's fast-tech work (mentioned briefly) suggests Braintrust is pushing toward zero-config onboarding. Worth tracking for your own stack.

Watch verdict

Skim. The workshop format is valuable if you're hands-on building agents right now, but the core ideas (trace everything, break agents into stages, eval with golden sets + LLM judges, manage prompts collaboratively, score production traffic) are well-summarized in slides and can be absorbed without live coding. Trainline's real-world examples (CFO invoice bug, model cost optimization) are the most concrete value. Skip the verbose setup/troubleshooting sections unless you're actively implementing Braintrust. The slide deck (available on GitHub Pages) is likely more efficient for reference than rewatching.

Full transcript 13792 words · 72 min read
0:14

SPEAKER_03

Good afternoon everybody. Welcome to sunny London. Is everyone's first time here? I think it is because this is the first conference. Amazing, amazing. Well, thank you very much for joining today's session. Hopefully you are in the right session, but for those who need to double click, this is a hands-on workshop to delivering quality AI applications with Brain Trust. And we're also partnering with our colleagues at Train On, which I'll introduce to you shortly. So probably you'll be wondering who's this guy, doesn't even go here, what's his name? Introduction to myself is, you can call me Duran, a bit like the band, Duran Duran with a G.

0:48

SPEAKER_03

So hopefully folks are a Simon Le Bon fan. And I've had a bit of part of helping organizations and enterprises help scale adoption of mission critical systems. And now I'm moving into the age of AI. Background in full mathematics. So I know there's all the rage going from machine learning to data science and AI engineering. And this comes at a very topical time. Feel free to connect me with LinkedIn if you want to stalk or just want to do general chit-chat around this area. I'm also joined by my two friends over at Train Line. If you want to come over and introduce yourselves, team. Of course, if the mic is working. All good? Yes.

1:24

SPEAKER_03

Hello, everyone. Thanks for coming to this workshop, especially after that lunch break and the sunny time outside. Thank you very much for coming here. So my name is Osama. I'm a senior AI ML engineer at Train Line. For those who know me, yeah, one person. Yeah, I was a staff platform engineer before and I have a background in computer vision. And I was doing also mobile apps on the side. So nothing to do with AI at some point. But yeah, now we are doing AI together with Mayank at Train Line. Hi, everyone. My name is Mayank. I'm a senior AI engineer at Train Line. I have background, research background in LLMs. But my research is in the pre-big LLMs.

2:10

SPEAKER_03

So if you've heard of BERT or the older LLMs. So we are continuing building a state of the art agentic products at Train Line. If you have come here from abroad, I'm sure you must have used Train Line. If not, please do. Because it's the state of the art app for buying tickets and other things. So yeah, welcome. And we are very excited to host you at this workshop. And yeah, we'll guide you through the hands-on experience. Fantastic. I'm also joined by some of my colleagues as well from Braintrust coming over the pond. So Phil, Eric, and Rose, if you just raise your hands. So don't worry, folks, you're in safe hands if you're ever stuck.

2:47

SPEAKER_00

So just holler and they can help. Fantastic. Just want to do a little bit of housekeeping as well. If everybody could join into the AI engineer Slack channel. Well, there's an AI engineer organization. But there's also a specific Slack channel which we'll be using today to help progress the workshop. So if you are stuck, we can use the Slack channel to help each other out. Again, we're already on there. But as we progress and get to the hands-on element, and if you are stuck, this will really help.

3:16

SPEAKER_00

We're also providing a cheat sheet. So if you are encountering a particular hurdle, there's step-by-step instructions which help you just get to where we need to go to the workshop without having to feel like you're falling behind. Again, a lot of the assets which we share today is publicly available. Again, we can have any particular follow-ups if needed. But we'll give you a few minutes to make sure you get onto that and then join the channel AI Engineer Europe 2026 Braintrust-Workshop. It's a public-facing channel.

3:49

SPEAKER_00

All right, team. I guess we can start. Yep, just for the people who are in the back.

3:59

SPEAKER_00

We'll need you to join the Slack channel.

4:02

SPEAKER_02

So if you've seen Piers, then we'll be there. So, with that team, thank you, everybody. Let's proceed. Okay, so just to help orient today's workshop, we'll be breaking into three main sections. We'll spend a little bit of time just establishing the background while we're here and why this workshop is relevant. Hopefully, we'll set the context for when we go into the workshop, building the system, talking about this, how do we ship AI quality. And then we'll wrap up for the key takeaways and we'll be around to answer any open questions and answers that we have on the field. Okay. So, hopefully today there's going to be a lot of people from very different backgrounds.

4:46

SPEAKER_02

We intended this workshop to be catered for, again, probably everyone here knows what an LLM is. I don't want to insult anybody. But hopefully you're starting to explore your journey in terms of maturing your operations in terms of building AI systems. So, whether you're an AI product engineer, probably an applied team or come up from a traditional machine learning folks,

5:07

SPEAKER_03

maybe you might be a platform operation infrastructure, this is really appropriate for you. Just a show of hands here, who here comes from, let's say, a traditional data science background, perhaps? Okay. Interesting, interesting. Who here perhaps comes from maybe software engineering or then pivoting into AI? Okay. So, this is definitely the right room for you. So, hopefully we'll be able to accelerate things as we progress. Okay. To start with the context here, I think this is not an uncommon expression, which we are seeing more broadly across the industry. Again, a show of hands, who here has done a machine learning or AI POC?

5:45

SPEAKER_03

Let's say, more specifically, a generative AI POC, but then has failed to take that into production. Okay. We've got a few hands here. I'd be worried if everyone didn't, but this is a key thing which, speaking to me on my custom is, executives, top-level folks, all the range, they're thinking about this new technology. It's not necessarily new, but it's newer for a lot of folks, especially in more enterprise and regulated industries. And then they're trying to take this to delivering value to their customers. But unfortunately, there's a big hurdle between taking what you might develop locally on your machine and then industrializing it and making sure it works in anger.

6:23

SPEAKER_03

But the key thing that we've seen from all of the research out there, it's not that the models aren't particularly smart. We've got very sophisticated models, whether you're building something in-house or you're using a top-of-the-shelf, commercial LLM provider out there. What we do see more broadly speaking is the type of operational rigor when it comes to delivering these systems to scale has not kept up. Because traditional software engineering, very deterministic, one plus one equals two, great. LLM systems, as my five-year-old would say, two plus two equals ten, daddy. So, it's having to adjust it and make sure that we're delivering that to scale.

7:00

SPEAKER_03

We've got very sophisticated models, whether you're building something in-house or you're using a top commercial LLM provider out there. What we do see more broadly speaking is the type of operational rigor when it comes to delivering these systems to scale has not kept up. Because traditional software engineering is very deterministic, one plus one equals two, great. LLM systems, as my five-year-old would say, two plus two equals ten, daddy. So it's having to adjust and make sure that we're delivering that to scale. And what we're seeing when shipping these things is the fact that people think the demo state is suitable.

7:29

SPEAKER_01

[SPEAKER_03] But clearly, doing two to three to five demos is great.

7:35

SPEAKER_03

Putting it in production, everything goes awry. Treating some of your logs as observability, and this is really critical to how we work at Braintrust: logs will tell you what has happened. But sometimes you need to go deep into the system and understand its behavior. And this is really where observability comes into play. [SPEAKER_01] Something as well is, it works on my machine, fails in production. I try to patch the prompt. And then it's operational until the next issue happens or the next failure mode. But how do we keep track of that? Especially if you don't have a system in place, irrespective of what tooling you use, this can have a categorical effect.

8:03

SPEAKER_03

So a lot of what we see is not to do with the tooling technology. It's down to operational workflows. And this is really what we're aiming to help you with in today's workshop. Okay? So, as I mentioned, it's not the prototype. It's getting to a state where we're knowing exactly what's changed in the system. How do we interact with that? And then how do we systematically put a set of rigor so that we can get better and better? Remember, our target is not 100% coverage. It's getting as close as possible while maintaining and fixing the gaps that might have existed. And something that we see time and time again is a one prompt might work.

8:54

SPEAKER_03

But as you move and industrialize it, you want to probably do things like breaking this down into each individual sector of responsibility. So if you do come from the software engineering background, you know about these. You're breaking down the monolith into microservices. We'll outline a very similar approach here when we talk about building these systems. Again, making sure that we're understanding these changes and putting a set of systems in place. This is really what we want to do in today's workshop. Okay? In terms of today, hopefully, as I mentioned, this is a hands-on workshop.

9:31

SPEAKER_03

So we're going to be going into the terminal, we're going to be going into the UI, and we're going to be going through step-by-step and guiding along the way. So we'll be doing a staged AI system with multi-stage tool calling, which really allows us to see this more agentic flow. We'll then use Braintrust to instrument and see how the performance of the application is working. We then also want to take a look at creating what we call identifying new failure modes using a golden set. So we'll push that through. We also then want to talk about how to industrialize it.

9:58

SPEAKER_03

So moving from, hey, it just works on my machine to something that you can use in production and have it managed forward with a system in place. And the key thing is identifying those edge cases. Because you can create a test data set, but ultimately, there's no substitute for real-world data. So we'll be able to show you how we take those real signals in and evaluate and complete the loop. Just a bit of introduction here today as well. Who here has heard of Braintrust or played around with Braintrust? Can I get a show of hands? Okay, great, fantastic. So just a bit of introduction to Braintrust. We're a company now that's just shy of three years old.

10:32

SPEAKER_03

We're approximately a Series B company. We just announced that a few months ago. We raised $80 million at an $800 million valuation. We have investors such as Iconic, AZ16, as well as Greylock, to really talk about helping organizations ship quality AI at scale. So we're the platform for AI observability. We've got a heavy user base globally, but we're also expanding our presence heavily in Europe.

11:00

SPEAKER_03

So I'm one of the first engineers to join and help build out our go-to-market function here. And we're very excited with our customers and our friends over at Trainline to do that. Some of our local customers include Lovable and Dr. Lib as well, which are really pushing the forefront of these AI systems. When it comes to using Braintrust, I want to get your hands on with it, but where we really distinguish ourselves is being able to do this at scale. So our founder, Ankur Goel, this is actually his third time building Braintrust.

11:32

SPEAKER_03

He's an expert when it comes to database systems, and he's built a company called Impera previously that was acquired by Figma that talks about document extraction. And he's led the ML machine learning team out there because he realized building these evaluations is hard. Understanding production traces is hard. So he was having this issue, and I'm sure there are other organizations out there which are doing the same. And so he founded Braintrust to really help do this at scale. As we're doing and understanding these traces which are coming in and being very highly semi-structured data which changes, he realized that traditional analytical systems weren't fit for purpose.

12:01

SPEAKER_03

So we've created a new category of databases called Brainstorm, which really helps identify and accelerate this at scale. With us, we're tool-agnostic and platform agnostic. So irrespective of the agent framework you're using or the LLM providers out there, we're intending to help you deliver value irrespective of that. Okay. One of the things I will talk about as we progress in this workshop is the concept of a flywheel. So if you ever come from agile development, you know perfection is the enemy of good. We want to start somewhere.

12:35

SPEAKER_03

So even if it's the case that it's a new application and you don't know how it's going to be in production, we can start up with an evaluation set. If you have an existing application to instrument, that's great. We can pull that information in and identify the failure modes here. So, irrespective of the agent framework you're using or the LLM providers out there, we're intending to help you deliver value irrespective of that. One of the things I will talk about as we progress in this workshop is the concept of a flywheel.

12:57

SPEAKER_03

So, if you ever come from agile development, perfection is the enemy of good. We want to start somewhere. So, even if it's a case that it's a new application and you don't know how it's going to be in production, we can start up with an evaluation set. If you have an existing application to instrument, that's great. We can pull that information in and identify the failure modes here. So, the key thing is get information into the system, identify those modes, remediate, ship it out, and then monitor, and complete the flywheel again and again, so you get to where you need.

13:07

SPEAKER_03

You've heard a lot from me. So, one thing I do is I'm going to provide my colleagues over at Trainline to maybe just share their experiences prior to Braintrust and how they're helping us. So, Osama, I'll give it to you. Thank you very much. And hello again for people who are joining us just now.

13:10

SPEAKER_03

As my colleague introduced, Trainline is a company that's actually helping people get on trains. Trains are different than planes, if you don't know. There is a worldwide system, central system for all planes around the world. It's not the case for trains. And in Europe and the UK, it's very hard, I would say. If you would like to install an app for each carrier, it will take the whole space on your mobile phone for sure. So, Trainline is actually being that platform to help you book tickets. Mobile app agnostic, platform agnostic, carrier agnostic. You do it on one app and you can book a train from Paris to London, from Lyon to Milano, whenever you like. Every carrier in the EU. We sell almost 6.3 billion tickets on trains. 27 million active users and counting. And the other interesting one probably for this conference is how many AI conversations we actually have with our travel assistant. So, we do have a travel assistant that is exposed to people. And it's not just a chatbot. It's actually a multi-agent system that can handle refunds for you, can handle changing trains for you. So, it is very proactive, an agentic system, and not just a chatbot.

13:12

SPEAKER_03

One of the benefits of having 27 million active customers is you've got a huge space of how you can serve agentic applications live to the customers. One of the examples of that is what Osama talked about, a travel assistant, which you can get to from a ticket window in the application. It's an agentic system, which is something we want to talk about a little bit later in the slide.

13:16

SPEAKER_03

Which brings us to the next point. I will keep it short. Selling train tickets and being a train ticket company, how come we are doing machine learning? Of course, we can. There's so much to do and to help people in terms of their journeys, getting their tickets, getting their trains, getting back home. We do two things. The classic ML part, which is actually building models, we do that. We build ML models inside Trainline from scratch, from data to model. This is something we do. And we also do the multi-agentic generative AI systems that we are now familiar with, on top of LLMs that we love, all the tooling and context engineering and all of that. We do that. So, we do both sides of the story at Trainline.

13:25

SPEAKER_03

[SPEAKER_00] And these are two examples of what users are actually using on top of those systems. On the left side, this is what you can think of as your weather application, but for train disruptions. So, you have a ticket for a train. You get there. We know if this train would be disrupted or not, if it will probably be late or not. We know that based on huge data that we have, which was a machine learning model that was trained and can actually predict train disruptions, being late, and all of that. So, this is the classic ML part. The other one is the travel assistant that I told you about. It is, as I said, a very advanced multi-agent system. It can show you alternative trains if your train is cancelled or something is wrong with it. Good luck doing that yourself, even with ChatGPT. And the other one is handling refunds. So, you can actually get a refund on your ticket if your train is late or something like that. And it can give you all of that without a handover and also can do the handover to actual human customer support in our customer support team. Which means that if you are doing this at production level and at Trainline scale, it means that you can ask the question: are we breaking things at Trainline? Of course, we don't want to do that. And we are moving fast because technology is definitely moving fast. And this is why we are here. We are here to show you what we are doing to move fast without breaking things in terms of AI.

13:28

SPEAKER_03

[SPEAKER_02] And we can do that on top of handling the complex software systems that we love from APIs at scale and serving millions of users. So, how we would like to think about it is this scale. We know for sure that whatever we have as software systems, this is the deterministic side of the story that we have. And on the other side, building ML models, this is the non-deterministic side of the story. And we know for sure that the agentic systems are in between. There are parts of them that are deterministic. There are parts of them that are not deterministic. And this is how we have the framework of thinking that we have.

13:33

SPEAKER_03

[SPEAKER_00] In terms of quality, how are we handling that? That's the question. For the ML models, we do care about the quality of data that we use for training the models. But also, we have on top of it, of course, ML machine learning evaluations, whether offline or online. Offline means before going to production, you need to do your evaluations. And online is the data from production. You get it and evaluate your model if it's doing well or not. In our case, for instance, for that weather forecast, is it predicting the state of the train if it's disrupted or not? Is the model correct in its prediction? So that's one example.

13:35

SPEAKER_03

[SPEAKER_00] In terms of quality, how are we handling that? [SPEAKER_00] That's the question. [SPEAKER_00] For the ML models, we do care about the quality of data that we use for training the models. [SPEAKER_00] But also, we have on top of it, of course, ML, machine learning evaluations, whether offline or online. [SPEAKER_00] Offline means before going to production, you need to do your evaluations. [SPEAKER_00] And online is the data from production, you get it, and evaluate your model if it's doing well or not. [SPEAKER_00] In our case, for instance, for that weather forecast, is it predicting the state of the train if it's disrupted or not?

14:16

SPEAKER_03

[SPEAKER_00] Is the model correct in its prediction? [SPEAKER_00] So that's one example. [SPEAKER_00] On the other hand, for people who are familiar with software engineering, we do all quality checks and diagrams or whatever, and the tooling that we have for handling quality at scale, at a very large scale at training line. [SPEAKER_00] And for those systems, I think you already guessed it, it's a combination of both. [SPEAKER_00] It's not one without another, for sure. [SPEAKER_00] We do everything that we do quality-wise in terms of deterministic systems, but we also use the evaluation side from the non-deterministic systems.

14:47

SPEAKER_03

[SPEAKER_00] So that's the framework of thinking that we have at TrainingLine, thinking about these systems. [SPEAKER_00] And definitely, Braintrust is helping with that. [SPEAKER_00] And by the way, disclaimer, we are not paid.

15:12

SPEAKER_03

[SPEAKER_00] We are here just because we are convinced that Braintrust works for us. [SPEAKER_00] So that's 100%. [SPEAKER_00] We are happy customers. [SPEAKER_00] We have been with Braintrust for a long time. [SPEAKER_00] And we use some of Braintrust features. [SPEAKER_00] So for instance, for the ML AI evaluations, we do that. [SPEAKER_00] And we follow the scoring of Travel Assistant on many levels, from tone of voice to actual helpfulness when it comes to tickets. [SPEAKER_00] And tickets are really complex in terms of reasoning and what you should get, what you should not get, depending on if the train is late or not.

15:37

SPEAKER_03

[SPEAKER_00] The type of ticket is a return ticket or advance. [SPEAKER_00] So many complex cases. [SPEAKER_00] We do follow that evaluation side. [SPEAKER_00] Probably Mayur would like to add more about that. [SPEAKER_00] Can I get a raise of hand from people who have struggled with LLM costs, number of tokens, switching models, problems like this? [SPEAKER_00] Right.

16:04

SPEAKER_00

So it was a problem with TrainingLine as well, because we do this at scale. And the amount of OpenAI and Anthropic that we pay is just high. So we have to keep switching models, which is the best model for our use case, cheaper models, models with efficient tokens. Now, any time you want to switch models, you want to make sure that it is performing at least at the same level as your current model. Before BrainTrust, we had no specific way of doing it because we didn't have scores set up, we didn't have evaluation. With the usage of BrainTrust, what it has enabled us to do is simulate how the performance of the lower model would look.

16:43

SPEAKER_00

And we've used BrainTrust extensively to run offline evaluation and see what the effect is going to be. And also online evaluation to see that the intended effect that we observed in offline is as expected. That's one of the use cases. The other one is more generic, which is it would have taken us a lot longer to evaluate a new feature shipping into Travel Assistant. But with BrainTrust, we have been able to make sure that we are assured the new experience for the user is going to be good. So BrainTrust has helped us a lot in shipping fast. And also, the other part is observability for sure. We do use BrainTrust for observability. Yes.

17:09

SPEAKER_00

This is a true example from BrainTrust. We are spoiling the workshop, whatever you are going to see there. We can track everything regarding tool calling, the other agent, in terms of number, in terms of quality. Those kind of insights, you will need them to get from proof of concept to actual system in production or in your company or for something actually that's been used out there. And users find it a true working product and not just any slow AI-generated system. We will need that, definitely, for those cases. Would you like to add something here?

17:42

SPEAKER_00

I was just going to say that BrainTrust enables you to look inside complex agentic workflows up to tool call level, token level, which is very insightful and helps you debug a lot of things in production and before you deploy it.

17:48

SPEAKER_02

[SPEAKER_00] And one last thing, probably, is the cross-functional friendly point. [SPEAKER_00] That's something that we have discovered along the way. [SPEAKER_00] Because we are building those systems from travel assistant to the models, at some point, we need to have what we need is that we are a big company and we have people from product, people with non-technical backgrounds, and we need to communicate and we need to share many things.

18:12

SPEAKER_00

And they also need to self-serve some of those things. [SPEAKER_00] We cannot babysit any people and say, hey, you should do this, should you do this, let me get you the logs, download it and send them that. That does not work. At scale, we need a way where we can help work cross-functional and let people be free to do whatever they want and self-serve with data and insights. So this is what we have also discovered along the way. [SPEAKER_02] And Braintrust helps us with many requests that we had. [SPEAKER_02] We really appreciate it, for sure. [SPEAKER_02] And, of course, more. [SPEAKER_02] We are using just a part of that system. [SPEAKER_02] We have our own things.

18:54

SPEAKER_00

[SPEAKER_02] Braintrust, I think you are building even more tooling. [SPEAKER_02] So definitely more to come. [SPEAKER_02] I think the workshop will definitely help you discover all those things. [SPEAKER_02] And, of course, have fun building during the workshop. [SPEAKER_02] Laptops out. [SPEAKER_02] And if you are interested, Trainline is hiring in the AI engineering side, of course. [SPEAKER_02] And I give it to you, Sean. [SPEAKER_02] Perfect. [SPEAKER_02] Well, thank you so much for that, team. And just to say, again, an impartial thank you, an extension of my thanks. We thank you so much for the split-saver functionality. So giving a close to that.

20:07

SPEAKER_00

All right, then, team. All right, let's proceed with the setup. Again, this is going to be a hands-on workshop. Better dust off your bash skills, team. [SPEAKER_02] And if you are interested, Trainline is hiring in the AI engineering side, of course. [SPEAKER_02] And give it to you, Sean. [SPEAKER_02] Perfect. [SPEAKER_02] Well, thank you so much for that, team. [SPEAKER_02] Yeah. And just to say, again, an impartial thank you, an extension of my wife. We thank you so much for the split-saver functionality. So give a close to that. All right, then, team. All right, let's proceed with the setup. So, again, this is going to be a hands-on workshop.

21:05

SPEAKER_00

Better dust off your bash skills, team. Jokes. I've got a nice Rafa Comonst to help out with that. So hopefully everybody, or at least most folks, have joined the Slack organization for AI engineer. And then also joined the channel. [SPEAKER_02] I've put a link to this. [SPEAKER_02] But there's also a QR code to the repository which we'll be using. So I'll give it a few minutes. Really key things is signing up to a free brain trust account. If you use Gmail, you can use a little plus sign trick to just create your account. If you are an existing brain trust user today. So hopefully it doesn't pollute your existing one. You will also need access to an open AI API key.

22:20

SPEAKER_00

That should be easy for you to generate. If for some reason you're not able to generate a key, then please let my colleagues know. We will be able to send a DM you on Slack to use for this specific session today. And I think obviously as an engineering conference, and especially in AI, if you are using AI coding assistance, feel free to use that with the IDE or in the terminal to guide you along the way. Ask questions around the code base and so forth. I've tried to simplify a lot of the scaffolding in place. So I'm using Mise to manage Node and PNPM on the machine. But alternatively you can download Node V22 and the specific version there. It should be fine.

23:28

SPEAKER_00

But again, I just want to caveat, especially for this workshop, I fixed that specific runtime. I'm using make to just, again, submit syntactic sugar to wrap the commands. If for a reason you don't have make installed, if you're on a Windows machine, you can then just run the PNPM commands, which are wrappers for the package.json file. Hello. Yeah, there today. So yeah, we'll give it a few minutes for folks to scan and clone. I just want to point out as well, I have a sheet today I can see lots of folks have already gone through it. But yeah, we provide a step-by-step guide.

23:56

SPEAKER_00

So as mentioned, as we're proceeding with the workshop today, if you are stuck, please ask questions on the Slack channel. But also you refer to this, as I mentioned, this will stay. This is a public facing asset. [SPEAKER_03] So even if you, outside of this workshop, are stuck, you can go to it at your own leisure.

24:14

SPEAKER_02

[SPEAKER_03] Just show of hands, who here is not able to get an open AI API key? [SPEAKER_03] Hopefully folks can be able to provision that. [SPEAKER_03] It's really critical to this. [SPEAKER_04] Yeah. [SPEAKER_03] Some people joined late who don't have the QR code for the Slack. [SPEAKER_03] For the Slack, yeah. [SPEAKER_03] I can do that. [SPEAKER_03] Yeah. [SPEAKER_03] Okay, folks, I think we did have a few late starters. [SPEAKER_03] So, yeah, I don't want to spend too long with this. [SPEAKER_03] But if you did join late, please scan that QR code, which will give you access to the AI Slack engineering organization.

24:48

SPEAKER_02

[SPEAKER_03] And then join the AI channel, AI Engineer Europe 2026 Brain Trust Workshop. [SPEAKER_03] It's a public facing channel. [SPEAKER_03] Through that channel, you'll be able to get access to the pinned content. [SPEAKER_03] So, the repository as well as a cheat sheet, which we'll be following as well.

24:52

SPEAKER_00

[SPEAKER_03] Just a bit of a sense check, folks. [SPEAKER_03] Hopefully most of you have been able to join the Slack channel and are able to clone the repository, have access to that. [SPEAKER_03] All right, then. [SPEAKER_03] As I mentioned at the start of the agenda as well, what we'll be doing is creating an example support triage agent. [SPEAKER_03] So, hopefully this exemplifies a lot of systems that folks maybe have already built or are trying to build. [SPEAKER_03] It is a fictitious application designed specifically for this workshop. [SPEAKER_03] Please do not use it in production.

25:12

SPEAKER_00

[SPEAKER_03] It's really around to teach us how do we build these AI patterns at scale. [SPEAKER_03] So, the idea there is given a ticket raised in the particular system. [SPEAKER_03] We have a set of agents that goes through a pipeline process, a stage process of tool calling that produces a set of information that can be emitted into downstream systems. [SPEAKER_03] Okay.

25:24

SPEAKER_02

[SPEAKER_03] Just to help visualize what the system does, again, we're taking the ticket input. [SPEAKER_03] We have a first step, which is collecting the context.

25:31

SPEAKER_00

[SPEAKER_03] It's a deterministic way of extracting information. [SPEAKER_03] We then proceed with the agentic portion where we've broken those into three stages. [SPEAKER_02] So, we have an LLM and tool calls to triage a particular issue. We then want to do a policy reviewer to make sure the output is correct. [SPEAKER_03] We then want to create and draft a customer-facing reply with reply writer. [SPEAKER_03] And then we want to package this together. [SPEAKER_03] And depending on how severe this particular ticket is, we can invoke another tool call to determine whether we need to escalate this to a human in the loop. [SPEAKER_03] And then draft the final result.

26:01

SPEAKER_00

[SPEAKER_03] So, again, not too complex, but it gives you an idea of the system that we'll be trying to build and operationalize in the session. [SPEAKER_03] One thing to point out is, as we begin to use Braintrust, where does it fit into helping you deploy and manage this at scale? [SPEAKER_03] I just isolated this here where everything we're going to be doing towards the end, especially if we progress, we'll trace it end to end, as my colleagues talked about. [SPEAKER_03] We'll be able to use manage mode for prompts, so offloading what you'll do locally into a secure environment. [SPEAKER_03] Tool calls as well will be managed.

26:31

SPEAKER_00

[SPEAKER_03] So, how do you talk about getting to external systems? [SPEAKER_03] And lastly around evaluations and scorers, which will be run through our Braintrust infrastructure. [SPEAKER_03] It's worth pointing out as well, if you do go into the GitHub repository, there's a GitHub Pages version of this as well. [SPEAKER_02] So, the slides will be readily available for you to use and consume. [SPEAKER_03] Okay. [SPEAKER_03] Key thing as well with this is I've tried to help you do each phase and checkpoint this. [SPEAKER_03] We'll be able to use manage mode for prompt, so offloading what you'll do locally into a secure environment.

27:07

SPEAKER_00

[SPEAKER_03] Tool calls as well will be managed. [SPEAKER_03] So, again, how do you talk about getting to external systems? [SPEAKER_03] And lastly around evaluations and scorers, which will be run through our Braintrust infrastructure. [SPEAKER_03] It's worth pointing out as well, if you do go into the GitHub repository, there's a GitHub Pages version of this as well. [SPEAKER_02] So, again, the slides will be readily available for you to use and consume. [SPEAKER_03] Okay. [SPEAKER_03] Key thing as well with this is I've tried to help you do each phase and checkpoint this.

27:39

SPEAKER_00

[SPEAKER_03] So, if you're ever stuck, the idea is just use git checkout to a specific tag or the branch, but in this case, each individual tag or branch is fully runnable at that stage.

27:40

SPEAKER_03

So, if you're ever stuck, again, git checkout to that branch.

28:05

SPEAKER_03

So, you may install, make setup, run the commands, and then you should be able to get near identical output every time. So, again, this is all documented in the readme, but it's also included in the cheat sheet as well for you to progress. Okay.

28:22

SPEAKER_04

[SPEAKER_03] So, just to give you the sequence of events, we'll be doing a scaffold and setup.

28:23

SPEAKER_03

I'll talk about building a basic agent. Again, this workshop isn't really designed to talk about building agents. It's about, okay, once you do have an agent, how do you operationalize that? So, but just for brevity, we'll do it step by step. And then we'll get into the core fit where we'll add the tracing, talk about the evaluations, talking about that golden set and identifying where we can improve.

28:37

SPEAKER_03

And then again, tying it all together and using managed infrastructure within Braintrust to help operationalize this out. And again, talk about that collaboration that my colleagues at Trainline talked about. Also, the key thing, which is, once you do identify a production failure, it's then how do we apply a fix to this and then complete that flywheel and give you a finalized asset. So, in step one, we're going to be talking about building the agent. So, in this checkpoint here, again, we need to start somewhere, right? So, what do we do? We will take an initial call to an LLM, we'll set up a prompt, one shot, one in, one out and get an output.

29:13

SPEAKER_03

So, again, if you're doing a proof of concept, this might be great as part of that initial spike, but we know there's work to be done. But, for context, we're just going to be stepping through this. But, as I mentioned, just because it works in the demo doesn't mean it's necessarily going to work in production. I've put some pseudocode here, just to articulate here, but you can see here, as with all these language models, we provide a set of system prompt, we provide the user and especially the text, and we pass that output back into our application. So, one function, one model call, we have the structured output, which we want to receive as part of the output.

29:44

SPEAKER_03

So, what we'll do as well, when we do the scaffolding and checking out to the first branch, we can run a set of pre-built tickets. So, it's a case that's already done in code, or you can use the command line. So, I'll just demonstrate that shortly, what it would look like with the outputs, and you'll get something similar to that in JSON format. So, hopefully, everyone can see this at the moment. I have the application here. Okay. So, I'm just going to go to basic checkout. Okay. And, what's this? So, if I take a look here, I'm now checked out into that first type, so building the basic agent.

30:16

SPEAKER_03

I've got the application here, which is, again, very similar to the prompt code. We're creating the client, calling OpenAI SDK. Again, for brevity, I didn't use any agent SDKs, but we fully support that as well as we proceed with the workshop. So, this is really available, and you'll see, again, from makefile, if you don't have make installed, then you can just execute pnpm scripts.

30:34

[SPEAKER_03] Or, if you use npm or yarn, that's also possible as well. [SPEAKER_03] I haven't tested it, but theoretically, it should work, because it's just using package.json. So, in this case, let's say you want to do something, say, my password needs to be reset. In this case, I've provided some defaults, so I'll just enter, enter, enter. And now it's just making a call to OpenAI. I'm using, you probably would have seen from the environment variable file, I'm using GPT-Pi Mini. You can switch it if you want, but just for the purpose of this, we'll keep it simple. So, you can see here, I've got the ticket, and it's provided some output here as well.

31:22

SPEAKER_03

But again, that's a singular shot. Right. So, based on this output, it looks fairly plausible. But again, it's not going to account for a lot of the edge cases that we want, especially if you've got a lot of nuance to the organization, which we're trying to build into the logic here. I can even do things like make demo.

31:53

SPEAKER_03

So, make demo.

32:01

SPEAKER_02

[SPEAKER_03] For the scripts.

32:06

SPEAKER_03

Yeah, it's the same thing. I'm just calling the same function, but I've got it codified as JSON here. So, the JSON fields are available to see. Awesome. Okay. So, that's fairly straightforward. I don't want to draw on that too much. So, the next thing I want to do then is talk about adding the local tools. So, we probably want to say, look, let's try to make this a bit more deterministic, even though a prompt might be very well structured. I may want to bring in different ways of how it might operate. So, in this case, I'm calling three different tools to look at relevant help desk articles around that. This could be both internal and external.

33:04

SPEAKER_03

I may want to look at certain things that have happened to the account. So, let's say a customer might have done a certain migration and that might have an impact on backend systems. And that's probably a reason why I want to create an escalation here. In this case, what I've done is I've made it a bit more deterministic.

33:35

SPEAKER_03

For the purposes of this workshop, again, this is treated as code. But, in reality, you are probably going to be interfacing with external systems like Vector Search, MCP, CLI, and other types of interfaces to build out this capability. This could be both internal and external. I may want to look at certain things that have happened to the account. So, let's say a customer might have done a certain migration and that might have an impact on backend systems. And that's probably a reason why I want to create an escalation here. In the case, what I've done is I've made it a bit more deterministic. For the purposes of this workshop, again, this is treated as code.

34:21

SPEAKER_03

But, in reality, you are probably going to be interfacing with external systems like Vector Search, MCP, CLI, and other types of interfaces to build out this capability. And, again, key thing here is the more things that you add, the number of ways that it can fail will also increase. [SPEAKER_04] So, again, this is why tracing, as we begin to go, will become more important. [SPEAKER_04] So, just one here. [SPEAKER_04] So, again, read me here.

35:00

SPEAKER_03

Next thing we'll do is go to add local tools.

35:10

SPEAKER_03

So, what if...

35:18

SPEAKER_03

Check.

35:26

SPEAKER_03

All right. So, in this case, you have the tools that I created. Are then available here. But, again, they just checked in as code for simplicity's sake. And, again, I can do the same thing where you make a ticket. Say password needs reset account locked. Okay. Now, see, I just provided a little bit more information. You can see it's a little bit more verbose because we've introduced tool calls into this. And it's giving more context to the LLM. So, if you are pointing out, if you are feeling stuck, folks, that a lot of these work... Or the tags of the workshop branches are built sequentially.

36:36

SPEAKER_03

So, if you let's go into, let's say, git tag number six, it's going to include everything as part of that. So, don't feel like you have to go through each one. If you are feeling stuck and you want to skip, you can do that as well. Okay.

36:53

SPEAKER_03

Let's get on to tools.

36:56

SPEAKER_03

So, again, I've already shown the code anyway. But this is an idea to pseudocode what it would look like.

37:20

SPEAKER_03

Stages. So, I think this is the next thing where, again, we're breaking down that monolithic call for an LLM. We're now introducing tools. And the next thing is even further drilling down to special stages of how the LLM should behave. So, you would have seen from that sequence diagram, we've done effectively five stages for this.

37:44

SPEAKER_03

So, I'm now setting up things to collect the context, triage, determining if it's meeting brand policies, providing a customer-friendly reply, and also something internally for our systems, and then finalizing the result for downstream systems.

37:50

SPEAKER_03

Again, it's just coming from traditional software engineering. You're breaking down your problem.

38:03

SPEAKER_03

You can see exactly where something's going wrong on the stack, and you'll be able to remediate. So, again, where possible, try to be more explicit and break it down into challenges that you can work on. So, yeah, just a bit of pseudocode. This is what it would look like. Probably going to put a debugger within your IDE and take a look at that. Okay, so I'm going to be doing this in part, but we've already started with the start of the point, and then I'm going to talk about doing the specialist stage here. So, let's take a look at the next part of the readme. So, we can check out the specialist stages here.

38:33

SPEAKER_03

So, if I take a look now, in the source folder, it's got the individual functions which are pieced out. The prompts associated with that is being used. So, let's say here is a triage from what we're using. And if I go down to the application, you can see it's done here. So, I'm using asynchronous functions to execute. And similarly as well, if I do something like make tickets, maybe I'll do something different. What's another classical problem?

39:12

SPEAKER_04

[SPEAKER_03] I'll just say, I need to upgrade my plan from pro to enterprise.

39:20

SPEAKER_04

[SPEAKER_03] But the website is not working. [SPEAKER_03] Getting 500 errors.

39:24

SPEAKER_03

[SPEAKER_04] So, in this case, I'm in customer tier two. We're talking about billing here. And my account is actually at count number three. So, again, just to show you that this is live.

39:35

SPEAKER_03

It's not doing something that's hard coded in the system.

39:39

SPEAKER_03

[SPEAKER_04] So, let's go ahead and see. So, again, you can see that this is a visual one-shot LLM into sequential calls. So, it is expected to take a little bit longer. But, again, that's this part of building out the sophisticated flow. And now you can see, again, it's a bit more verbose.

40:28

SPEAKER_03

But there's a lot more thought into this agent here. So, I guess you can see, again, we talked about a billing issue. You can see that it believes that it's quite high because a customer wants to do that upgrade from a different tier. But there's obviously an impact to this from a revenue perspective.

40:50

SPEAKER_03

So, in this case, we should escalate to the appropriate people on our side. So, you can see here what the escalation region should be. And both the internal and the customer are facing reply as well. So, we've included a confidence score to say, you know, if this is a true issue. But, again, as mentioned, tool calls will be able to pull information which is happening across different systems to provide a greater level of confidence.

41:17

SPEAKER_03

Okay. Perfect. All right. With that in mind. So, again, we've now hopefully shown how we can build and take an agent, break it down, build something that's multi-stage, introducing tool calling to give us what we need. The next thing is then to provide that information and start tracing it so we can actually see what's happening, down to the individual details. [SPEAKER_03] And this is where observability comes into play. So, what we want to do in this section is break down the full execution path. We know that the stuff is very nested in structure. So, there's tool calls. There's additional function calls behind it. We do want to track some of the key things.

42:12

SPEAKER_03

And I think Mayim pointed out the early struggles that they had to talk about latency, cost, tokens, account, especially for time to first token is a very important metric that we see many of our customers trying to identify. [SPEAKER_04] Also, what were the inputs? What were the outputs?

42:36

SPEAKER_03

Metadata associated with this. And then also including additional types of fields so that, again, when we talk about monitoring and observability, We know that the stuff is very nested in structure. So there's tool calls. There's additional function calls behind it. We do want to track some of the key things. And I think Mayim pointed out the early struggles that they had to talk about latency, cost, tokens, account, especially for time to first token is a very important metric that we see many of our customers trying to identify. [SPEAKER_04] Also, what were the inputs?

43:30

What were the outputs? Metadata associated with this. And then also including additional types of fields so that, again, when we talk about monitoring and observability, we can query it within the UI on the fly and set up alerts if needed. [SPEAKER_03] So, again, having an output is not enough. [SPEAKER_03] We need to understand the full execution part. [SPEAKER_03] And that's what tracing allows us to do.

44:05

SPEAKER_03

[SPEAKER_04] Yeah. [SPEAKER_04] Just to give you an idea of the context. [SPEAKER_04] And I know it's a very contentious topic to someone because, again, I come from a background in full stack development. So I split both coins.

44:20

SPEAKER_04

[SPEAKER_03] It turns both Python, Go, as well as TypeScript.

44:46

SPEAKER_03

But, yeah, our SDKs are multilingual. So Ruby, Go, I think even .NET, we've got some folks who are using it. And that's all good. But, yeah, I've just kept TypeScript for simplicity here today. But, yes, our SDKs do cover a range of different languages out there that you can start tracing your application with. So, yeah, a real key thing to this and one of my challenges that I do see with our customers broadly is they will do an individual interaction against a singular parent span. And this even might work with, let's say, multi-turn, multi-conversational agents.

45:13

SPEAKER_03

What we want to be able to do is trace that into a nested structure so you can see everything but in one interaction, let's say one conversation in a parent span. So it's really critical as we start instrumenting our application to make sure that we're getting the right structures in place. Otherwise, you're not going to be able to see the full effect of where things might go wrong in your application. Okay. Again, this will just come down to reading the trace. Hopefully, folks, especially, I know some of you've got more software engineers in the room that are using the traditional observability tools out there.

45:53

SPEAKER_03

This is not too dissimilar. But for folks who may be new to the screen, a trace just allows us to find out not what has happened but what's currently happening with your application in real time. So, again, tracking every single call, bringing that metadata. And, again, depending where the failure mode is, we can then identify, okay, what might be the course of remediation that we need to take. Okay. So in the next stage, what we're going to do is add tracing to this. And then we're actually going to run it and then go into Braintrust and hopefully we'll see it happening. But one thing I do want to bear in mind before I do that. So let's go to our readme.

46:27

SPEAKER_03

It's telling us to add tracing. If I can spell. Okay. Now we're great. I'm going to keep it there. Okay.

46:53

SPEAKER_04

[SPEAKER_03] So I've introduced a helper script here called tracing.

46:55

SPEAKER_03

So what we do quite well within our SDK is, again, we want to not introduce more complexity where it's needed. So if you are using, again, the standard LLM providers SDK, you can simply just wrap that function. We provide that out of the box. But then when it comes down to the individual calls, again, we can set up some helper scripts here to help just wrap up as needed. If you are using something like Python, then you can also have a decorator function which helps as well. In this case, I've got a nice little function which helps with the parent and child.

47:14

SPEAKER_03

And then when it comes down to the actual application tracing, you can see here I've got a child span which has been executed throughout this.

47:19

SPEAKER_04

[SPEAKER_03] Again, I don't want to do too many code dives at this point because the code is expanding pretty heavily. [SPEAKER_03] But, yeah, I just want to walk through the key concepts for this. [SPEAKER_03] Okay.

47:29

SPEAKER_03

Before I run the application, I do want to go into the Braintrust UI. Sometimes it might create a, if you've already signed up for a free trial account, so hopefully you would have done that earlier today. It'll create a test project. That's totally fine. We will create a new project as we execute this. A really key thing as well, if you haven't done already, top left-hand corner, you go to your profile. Oh, okay, let me just create one.

47:57

SPEAKER_03

Okay. Let's create that just for now. Go to your profile. And then where it says API keys, enter your name for the key, generate that in, and then use that within your environment, your .env file. Or secure that in a key vault if you have that already.

48:37

SPEAKER_03

There's also the OpenAI key, which we're hoping you generated. That should also be used in AI providers. So I've set this at the organization level. You can also do this per project, depending if you want to segregate it by a particular team or environment. That's also fully supported. But in this case, we'll need this key for later when it comes down to the managed and online scoring. So, but just a tidbit, that key that you generated, make sure you put it into your Braintrust organization here to be used later. Okay. Right. With that in mind, I'm going to run the demo.

49:23

SPEAKER_03

So, and just point of reference, if you are ever stuck with dependencies, you should just use make setup, make sure everything's in place. In this case, then we want to do make demo. And just be going to run and execute those tickets.

49:43

SPEAKER_03

I can even do it using the make ticker command as well. Okay. So while it's running in the background, oh, sorry.

49:54

SPEAKER_03

While it's running in the background, if you go back to your application, you would see there's a project called helper workshop. So that's the one that will be created as part of this workshop today. [SPEAKER_04] If you navigate to the logs tab, you'll start to see this coming through in real time. [SPEAKER_04] It's worth pointing out, again, thanks to our capabilities of Braintrust, we have near instantaneous write into our system and read available shortly after. So, especially, it's a non-blocking function. So a lot of our customers, especially the more sophisticated ones, are really using Braintrust at scale and really pushing the envelope. Okay.

50:28

SPEAKER_03

So while it's running in the background, sorry. While it's running in the background, if you go back to your application, you would see there's a project called helper workshop. So that's the one that will be created as part of this workshop today. [SPEAKER_04] If you navigate to the logs tab, you'll start to see this coming through in real time. [SPEAKER_04] It's worth pointing out, again, thanks to our capabilities of BrainTrust, we have near instantaneous write into our system and read available shortly after. So, especially, it's a non-blocking function.

51:10

SPEAKER_03

So a lot of our customers, especially the more sophisticated ones, are really using BrainTrust at scale and really pushing the envelope. So it's not the case, hey, I've got a thousand traces. They're pushing tens of millions of traces at a time across a very short period. And they're going to be able to aggregate this. And again, as you build more sophistication in your application, you're sending this out to more users, that's going to grow up pretty quickly. And you need to be able to have a system that can handle this at scale. So once you start to see the logs that are coming in, they come in reverse order, I'm just going to take a look at the first one here.

51:38

SPEAKER_03

And we'll start to see the instrumented application going to be traced. There's a particular button here that allows you to view this in full screen, which I think is quite helpful. So as I mentioned, because we've got this nested structure in place where we're going through each in the tree, this interaction at the top is the demo ticket. You can see the very top level, we're saying how long it took to actually run that invocation, the prompt up tokens in, out, cost and latency associated with that as well. I've also included some metadata because I want to be able to extract and filter that out as needed, and I'll show you how that will work.

52:15

SPEAKER_03

And everything is available here to view. Metadata is there. If I also want to take a different look at this, we can even look at individual steps. In this case, I want to look at what's happened with a tree-out specialist down to the actual invocation to the LLM. So again, the SDK provides a lot of flexibility around this, so you can see what was put in, [SPEAKER_04] the reasoning behind it, and the information that we set output. And then coming down to the last, should we escalate or not.

52:51

SPEAKER_03

Another view that we will see as part of this is taking a look at the timeline, so it just gives you an idea of a waterfall methodology to see if there's a particular step which is taking longer, do we need to remediate, it's all possible to do. So if I take a look at the logs page, get it back out of that, refresh, so I think it was four tickets that were pushed.

53:17

SPEAKER_03

Yeah, this should come through right now. Okay, yep, so that's two tickets at the moment, so again, that same information that we replayed is also available in the console. Okay. Yep, so we've covered tracing.

53:39

SPEAKER_04

[SPEAKER_03] Let's talk about the evaluation portion for this. [SPEAKER_03] It's quite interesting, again, depending where you are in your journey of building this application.

53:55

SPEAKER_03

A lot of customers already built an application, already have it monitored and pull that in. But what happens when you have effectively a cold start problem where you don't know what you're building, right? So what does good look like? And effectively, what does good ship to me? In the case of the support application that we're developing today, can I care non-negotiables? Have we categorized the support case? Have we made sure that there's no low severity outcome for these issues which are blocking? Does the escalation stay in policy with SLAs? Does the structure look sound?

54:43

SPEAKER_03

And if we're making any particular changes, does it actually improve without actually breaking or having a regression to the application? We can do this using evaluation. So an evaluation for those, hopefully many folks will know what they are, but those folks in the room, think of it as a way you have your data set, your input, you have a task, and then you have an outcome which you want to evaluate against, or a scoring function. And this is kind of a little bit different comparing traditional software development into working with AI systems because of this non-deterministic nature.

55:10

SPEAKER_03

So in this particular portion, what we're talking about here is creating what we call a golden dataset. So in this support application, I'll be testing anecdotally, but I want to create a set of edge cases where I think this is really going to help give us at least an initial level of confidence to the business that what we're releasing out into production is fit for purpose. There's always room for improvement, but I want to just think, look, it's not just me releasing this application based on vibes.

55:22

SPEAKER_03

I have a concrete way of saying, okay, this is how it's performed over time. To do this, again, we use two main types of scoring functions. The first of which is deterministic. So I think a lot of folks may have already started using this.

55:43

SPEAKER_04

[SPEAKER_03] Again, coming from traditional software engineering, you know, unit tests, I wouldn't say analogous, but they are quite similar in nature where they're very easy to run, cost effective.

55:47

SPEAKER_03

You're not actually using a model at this point. The secondary type, which is a little bit more sophisticated, is then using an LLM as a judge, another AI system. And this is really helpful when it comes to systems where there's nuance which can't really be determined on deterministic systems alone. So again, creating one to talk about, you know, branding style is meeting, you know, customer satisfaction and so forth.

56:11

SPEAKER_03

So the most important thing is, if you cannot write it in a deterministic way, you want to be using an LLM as a judge where possible. So, and why this is more important, again, it's just making sure that any change that we make is safe to change as we progress this. Okay, so I'm going to pivot into the IDE again.

56:37

SPEAKER_03

Let's take a look at the readme file. And then we're going to do... Well, it looks like a dictionary. Maybe come on. Okay. I'll see. Let's get checkout. There we go. Okay. And then based on this here, make demo, make setup, we should be fine. Okay. So one thing I'd like you to do as well is then to run the seed dataset command. So we do make seed dataset. Okay. So what this is going to do is upload our evaluation test cases into the BrainTrust UI.

57:45

SPEAKER_03

So if I pivot into my datasets now, you'll see it's called helper seed dataset. Again, just for simplicity, I've created 10 inputs. I've also categorized them around there. So the input, what we expect, and some metadata associated with that. There we go. Okay. And then based on this here, make demo, make setup, we should be fine. Okay. So one thing I'd like you to do as well is then to run the seed dataset command. So we do make seed dataset.

58:59

SPEAKER_03

Okay. So what this is going to do is upload our evaluation test cases into the BrainFrost UI. So if I pivot into my datasets now, you'll see it's called helper seed dataset. Again, just for simplicity, I've created 10 inputs. I've also categorized them around there. So the input, what we expect, and some metadata associated with that. So the core structure to creating an evaluation. So as part of this as well, the deterministic and non-deterministic, I've created some scoring functions, which we'll be using as part of this. So it's available here. And then the scoring functions.

1:00:11

SPEAKER_03

So again, checking the category. Is a schema in place? Is an escalation reason when needed? Again, very easy to run in Codify. So we're going to do that.

1:00:16

SPEAKER_03

Yeah. We have a bit more sophistication with the customer rubric. So this is more of the LLM as a judge use case here. All right. And then what we can do is then go back to the WeedMe. So we don't need to do the demo. We already pushed that through. Let's go ahead and then do make eval. So this is then going to run an evaluation.

1:00:46

SPEAKER_03

And this is right here the seed data as well. Again, you don't necessarily have to put it into code, but whether it's coming from a database. So in this case, I've just used a flat file JSON for this just to give you some context. So we are running this evaluation. And we should have, if we go to the UI, the place for experiments. [SPEAKER_04] So I'll just double check here. Yeah.

1:01:28

SPEAKER_03

So that experiment ran against all the JSONs associated with it. [SPEAKER_04] So you should get an output like this, at least in the terminal. [SPEAKER_04] But if we go back to the UI, again, we're starting to track our application against the inputs and outputs for this particular dataset.

1:01:45

SPEAKER_03

[SPEAKER_04] Quite similar to the tracing you would have seen from the online traces. [SPEAKER_04] You get to see a very similar view here and see how it works across your experimentation as well. [SPEAKER_04] So that's the way we're going to go back to the UI.

1:01:55

SPEAKER_03

[SPEAKER_01] And I think pointing out as we progress with the workshop, we'll start to see how we then improve it using the set of difference functions and track that across the UI. So yeah.

1:02:04

SPEAKER_03

If you need a bit more real estate as well, you can collapse the menu, which does help. Okay. Time. Perfect. Okay. All right. Let's proceed with the next checkpoint around deploying and managing this. So as I mentioned, I think a lot of folks, at least anecdotally when I've seen this for a bit of customers, is again, things will work really well on my machine. Okay. Now I'm checking that into code. I want to take this into a place where I can start to collaborate a little bit better.

1:02:45

SPEAKER_03

It's a point of reference. I've got versioning history. You can start to identify again who's made what change. And you need a way to be able to bring these users together.

1:03:02

SPEAKER_03

So I think Osama talked about collaboration.

1:03:24

SPEAKER_03

What's interesting as well is again, changing the prompt on your machine and then trying to ship that code to a repository.

1:03:32

SPEAKER_04

[SPEAKER_03] And there may be somebody who's, let's say, a non-technical SME, a product manager, perhaps.

1:03:35

SPEAKER_03

They want to update the prompts. They can't do it. They have to tap you on the shoulder.

1:03:41

SPEAKER_04

[SPEAKER_03] Maybe that's happened to some of the folks in the room here today. [SPEAKER_03] I know it's happened to me a few times before. [SPEAKER_03] And I understand it can get really frustrating.

1:04:07

SPEAKER_04

[SPEAKER_03] So we actually want a way to be able to pull that together. [SPEAKER_03] A really key thing, especially for those who work in very regulated industries, is reproducibility is a very big key thing.

1:04:25

SPEAKER_04

[SPEAKER_03] And I've worked in a better part of over a decade in both banking and capital markets.

1:04:26

SPEAKER_01

[SPEAKER_03] I know obviously with the regulatory out there, especially things like right to be forgotten, understanding who's made a change, especially in a stress test scenario.

1:04:34

SPEAKER_03

How can we put this into a new system?

1:04:41

SPEAKER_03

This is really key to helping unlock that.

1:04:52

SPEAKER_03

And interesting when it comes to identifying changes before you do that. So again, we don't want to just be making changes, pushing out production and asking what's happened. We need to be able to get that in place. And so for us, we're introducing some capabilities now.

1:05:05

SPEAKER_03

So what you've been running at the moment, you've been running the tools, you've been running the prompts, everything on your local machine. What we want to do is then offload that capability into Braintrust.

1:05:14

SPEAKER_03

So when your application is running in a secure environment, it can refer to Braintrust to pull that information and then help with the path of execution, which again follows the tracing mechanisms that we've done. So by default, when you're running the make commands, the runtime mode was set to local. If you want to use managed mode, just use the prefix managed and whatever the remaining make commands will happen there as well. So what I want to do then is pivot into the IDE. So going back to my readme. Okay. So just taking a look here, we'll do make setup.

1:05:40

SPEAKER_03

And the key thing I do want to emphasize at this point, because we're managing this in Braintrust, use the setup Braintrust command here. So what that's going to do is package up those scoring functions, those tools, those prompts, and push it onto the secured infrastructure. So you should get an output like this. And what that would look like in the UI. [SPEAKER_04] So let's go back to overview to see. [SPEAKER_04] So on the left hand side, if we take a look at prompts, you'll start to see the three prompts that we created as part of that workflow. [SPEAKER_04] So we'll give you an idea here. [SPEAKER_04] If we go to the three-yard specialist. So you can define a slug.

1:06:09

SPEAKER_03

So let's say an immutable ID, which you can then refer to it in the code. This can also be generated if needed. The prompt is again treated as code. We also use some interpolation there if you want to parameterize this. And what that would look like in the UI. [SPEAKER_04] So let's go back to overview to see. [SPEAKER_04] So on the left hand side, if we take a look at prompts, you'll start to see the three prompts that we created as part of that workflow. [SPEAKER_04] So we'll give you an idea here. [SPEAKER_04] If we go to the three-yard specialist. So you can define a slug. So let's say an immutable ID, which you can then refer to it in the code.

1:07:03

SPEAKER_03

This can also be generated if needed. The prompt is, again, treated as code. We also use some interpolation there if you want to parameterize this. And I'll show what that looks like shortly. If I go to the scoring function. So let's click on scorers. Okay, it'll come up in a second.

1:07:56

SPEAKER_03

Take a look at parameters.

1:08:20

SPEAKER_03

So I'm just going to take a look here. And I've created parameters specifically, just to simplify things, changing the baseline model. So maybe just a show of hands, these models get released so quickly. Has anybody had a PM or let's say a non-engineering SME say, hey, by the way, can I change the prompt? And can I change the model or the prompt and see what that would look like?

1:08:56

SPEAKER_04

[SPEAKER_03] Has that happened to you folks? [SPEAKER_03] I think you've got a few hands here. [SPEAKER_03] Great. [SPEAKER_03] Okay.

1:09:16

SPEAKER_03

So the good thing with this now using the managed parameters, those non-technical SMEs, they can come into Braintrust and change the prompt here, write a comment to say, look, let's use a different model. Let me use something like 4mini. Let me just say testing a new model. I'm going to save this version here. Write the comment.

1:09:38

SPEAKER_03

And what I'm going to do as well, just to give you an ending claim, I'm going to do make setup Braintrust. So every time you change something, you can run this command, but I'm just doing it for brevity here. That's more to keep this model in sync. So if I go to prompts, that's the sync. But if I take a look at the parameters in place, it should be there. So what I'm going to do is if I do something like... I think this is... What was a runtime model? I'm going to go to unlock. So in this case, by running the management, I'm just changing the course of the execution. Not to run the model locally, but to follow the path of what Braintrust is set for the model.

1:10:17

SPEAKER_03

[SPEAKER_04] OK. And if everything goes well... So you can see something just now. You can see that I've changed a model here. So again, I didn't have to do any code changes. All I had to do was go into the UI, change the model where I wanted to, or any other parameter, run that, and have that used as an establishing baseline for the evaluations. Again, if you want to, you can also run the demo script to push in the demo tickets.

1:10:47

SPEAKER_03

I'm just skipping that for the workshop today. I think I do see this with many of our different customers is saying, oh, there can be not necessarily a cause of concern. It's saying, hey, by the way, Braintrust is now having access control to these pattern parameters. What I would say is Braintrust is not really intended to replace that rigor. You probably still want to use things like version control systems anyway to track that.

1:11:35

SPEAKER_03

What we're just saying is when it comes to operationalize it and making sure that other users are able to work on a shared system, this is a recommended part that we would take to help out with that.

1:11:55

SPEAKER_03

So again, you would probably still have to have your prompts, your tooling, your parameters in a centralized way, but then provide automation in place for you to synchronize that and work. And that's the best way we've seen customers take advantage of this.

1:12:18

SPEAKER_04

[SPEAKER_03] Okay.

1:12:20

SPEAKER_03

This next portion we're going to talk about online scoring.

1:12:28

SPEAKER_03

So now that we have those evaluations in place, what we're going to do is then apply those scoring to actual live production logs that are coming through in the application.

1:12:32

SPEAKER_03

So, okay, it's great that we've done our test cases. We've got some level of confidence that it's working. But again, there's no substitute for production data. We all know this.

1:12:55

SPEAKER_03

So what we're going to be doing is creating, again, moving that logic into Braintrust and setting up our goal automations that would then track and evaluate this as logs are coming in real time. So, what pointing out when you start your journey, it's probably, especially if you're using LLM as a judge, you want to start with a higher sampling rate.

1:13:08

SPEAKER_03

So, again, as logs are coming in where possible, you want to make sure you identify a baseline. But again, there's a trade-off when these calls can be quite expensive, especially if you're using more sophisticated models and you need higher rate of reasoning. At that point, when you do want to, are happy with the output, you really want to reduce that sampling rate down to 5% to 10%. So, again, you're managing your cost effectively. Deterministic scores, again, they're cheap. Recommend running them all the time. Sorry, what was that question? Sorry, what was that question? Yes, so I'll bring that up in the UI. I'll show you.

1:13:57

SPEAKER_03

Yep. Okay, so if I go back into the IDE, I'll go into README here. Ah, so that's why I missed out the manage tools. [SPEAKER_04] So, let's just say git checkout. [SPEAKER_04] Okay, there. [SPEAKER_04] And as I mentioned, they all build upon each other. So, skipping this is totally fine. So, git checkout. All right. Now, if I do, it'll make setup. Everything should be fine. Make setup.

1:15:12

SPEAKER_03

Brain first. So, in this case, I'm actually taking the tools as well for the production. Okay, this will just really help accelerate things where possible.

1:15:24

SPEAKER_03

Coming down to, if I hit refresh. Right. So, the scoring functions that you saw earlier in code, they're now managing brain trust. So, you can see it's available here. The ones that I want to call out is the tree. So, I've got an LLM as a judge, which has been applied. So, input and output taking this.

1:15:43

SPEAKER_04

[SPEAKER_03] And there's an automation rule in place. [SPEAKER_03] So, root quality online. [SPEAKER_03] Brain first.

1:15:53

SPEAKER_03

So, in this case, I'm actually taking the tools as well for the production. Okay, this will just really help accelerate things where possible. Coming down to, if I hit refresh. Right. So, the scoring functions that you saw earlier in code, they're now managing brain trust. So, you can see it's available here.

1:16:14

SPEAKER_03

The ones that I want to call out is the tree.

1:16:18

SPEAKER_03

So, I've got an LLM as a judge, which has been applied. So, input and output taking this. And there's an automation rule in place.

1:16:36

SPEAKER_03

So, root quality online. So, what I've done here, again, I've automated the setup as code. So, to bootstrap the project. But, again, what we're going to do here is to say, look, depending if it's an individual span, we can run the execution or the entire trace. And this is why metadata is so important because we may only want to trace maybe specific failures that might happen within the code. Again, depending on the use case. My sampling rate, as I mentioned, is set to 100. But, again, for more expensive calls, we want to taper that as well.

1:17:08

SPEAKER_03

So, this is what the automation or how we would do that in brain trust there today. It's packaging the scorers and... No, so automation is more like the execution against incoming logs. So, it's with that. Yeah. But, more realistically, I've applied automation to setting up this environment and scaffolding. Yeah. So, that's... Just want to delineate that here.

1:17:48

SPEAKER_03

Yeah. Sorry, there's a question. Yeah, what kind of things are you scoring if you don't have to delete the data? Could you expand on that, please? Yeah. So, the animations are like all of the judge, right? Yes, correct. [SPEAKER_04] Where do you have to judge the properties?

1:18:13

SPEAKER_03

In... If there's no ground truth, general... Okay. [SPEAKER_04] How can you perform online so that you don't have ground truth? [SPEAKER_04] Yeah. [SPEAKER_04] Yeah. [SPEAKER_04] Well, you probably want to take that as an edge case, push that into a data set, identify it, and then move that back. [SPEAKER_04] So, that would be the approach that I would take for this.

1:18:23

SPEAKER_04

[SPEAKER_03] If you don't have any ground truth already, right?

1:18:27

SPEAKER_03

So, this is why I said, depending where you start, it's better to have some kind of data, and then begin your flywheel around from that. Oh, so, if you do have some data, how do you apply that to the alum as a judge? Oh, that case.

1:18:30

SPEAKER_04

[SPEAKER_03] We can probably put that into a data set, and then replay that through the playground. [SPEAKER_03] That's going to be the way we do that. [SPEAKER_03] Yeah. [SPEAKER_03] Yeah. [SPEAKER_03] All right. [SPEAKER_03] So, we checked up online scoring, cognizant over time.

1:18:44

SPEAKER_03

Okay. Now, to the remediation portion. So, hopefully seeing the delta, or again, why we're here today. So, hopefully this might help you with that particular question. So, again, here's something that might happen as a plausible input to our agentic system. A customer or user might say, "Hey, this isn't urgent, but our CFO can't export the invoices before the board meeting." The model says, "Look, hey, this looks okay for me." Someone says, not urgent. [SPEAKER_01] Come see, come see. But the business is very different, right?

1:19:12

SPEAKER_03

This probably does need immediate attention. CFO shows the end of quarter report that needs to be done. And this is the difference between what we're doing here today is trying to identify what is a proper failure mode, and then remediate that where possible. So, in this case, again, I can run this particular mode here. I've got this in this dataset. Just to kind of give you a play of C where we want to replay the failure, we want a specific evaluation against this, we want to tighten the prompt, run it again and see what that looks like.

1:19:27

SPEAKER_03

We probably want to do this against not just one particular test case, but then clutter our entire test cases as well, just to see if it does work as intended, and we have regressed on something else from that perspective. Okay, so let's go ahead and have two, we have two separate branches for this. So, I'll split out branches A. One's going to have the failure rate in mind. We're going to go to the UI and view that.

1:19:48

SPEAKER_01

[SPEAKER_03] And then we're going to talk about the remediation part there as well.

1:19:50

SPEAKER_03

[SPEAKER_05] Okay. Nice. Okay. So we can even do runtime mode. Okay. So I've got this set of five cases which are the regression or failure modes in this case. So that first one that you saw on the example ticket, that's done as a JSON file here. So take a look here.

1:20:19

SPEAKER_03

[SPEAKER_04] So if I take a look at the failure mode here, I can drill into this. [SPEAKER_04] See the... [SPEAKER_04] And you'll notice as well, now that we set up the online, the managed tools as well as the online scoring, the trace becomes even more sophisticated in the fact that we're executing this against the secure brain trust environment. [SPEAKER_04] So again, moving from local to managed. [SPEAKER_04] On the other hand, let me see what happens. [SPEAKER_04] I'm just going to go to the make file, sorry, .json. [SPEAKER_03] So I've got this set of five cases which are the regression or failure modes in this case.

1:20:45

SPEAKER_03

So that first one that you saw on the example ticket, that's done as a JSON file here. So take a look here. OK. [SPEAKER_04] So if I take a look at the failure mode here, I can drill into this. See the...

1:21:01

[SPEAKER_04] And you'll notice as well, now that we set up the online, the managed tools as well as [SPEAKER_04] the online scoring, the trace becomes even more sophisticated in the fact that we're executing this against the secure brain trust environment.

1:21:27

SPEAKER_03

[SPEAKER_04] So again, moving from local to managed.

1:21:36

SPEAKER_03

[SPEAKER_04] On the other hand, let me see what happens.

1:21:41

SPEAKER_03

I'm just going to go to the make file, sorry, .json.

1:21:43

SPEAKER_03

[SPEAKER_04] Let's go to. I'm just doing an evaluation against a specific scenario. Do you also want to show the monitor?

1:21:47

SPEAKER_05

Yeah, so as we can see, as we're progressing with the application, the experiment, again,

1:21:47

SPEAKER_03

Viewer just allows us to see the progress of our changes. So you can see we've run the latest set. We notice some degradation here, which just allows us to track it.

1:21:47

SPEAKER_04

So the managed data set that I had in mind, the failure rates are being captured.

1:21:47

SPEAKER_03

And you can see the ability to compare it against existing experiments to track the progress and remediate where possible. So the next thing is I'm going to go to the ReadMe and then proceed with the remediation. So I'm going to go to the Get Checkout. So I'm going to go to the Mediation. So I'm going to go to the Prompt, which I've used. So one way to view that is if I take a look at the Prompt and the Change. So you'll see any differences there. So I'm going to go to the Remi file here. So let's say I do one thing. [SPEAKER_04] I think codex here. Let's give us this essay. [SPEAKER_04] Let's run the Pinterest command. It's contents.

1:21:47

SPEAKER_03

Yeah, there's a specific flag, I think. [SPEAKER_04] I don't know if you want to just double check. Yeah, that's it. [SPEAKER_04] If you want to just put that if exists replace in the chat. [SPEAKER_04] Yeah, yeah. [SPEAKER_04] Yeah, no, no, it's fine. [SPEAKER_04] So yeah, just focus if you want to. [SPEAKER_04] In the part of the remediation script, if you set the environment variable [SPEAKER_04] Braintrust underscore if underscore exists, you set that to replace. It's going to push in the updated changes into the environment. So if I take, oh, sorry for that. Okay.

1:22:08

SPEAKER_03

[SPEAKER_04] And then you'll see as well, now that we've updated the Prompt via code and pushed it up, it's available here as well.

1:22:29

SPEAKER_03

So as well in the UI, again, part of the operationalization, we can see who's changed what, but actually what has been changed to that particular prompt in this case.

1:22:33

SPEAKER_03

Taking care. So including two facts.

1:23:09

SPEAKER_03

Print that out.

1:23:16

SPEAKER_04

[SPEAKER_03] Okay. [SPEAKER_03] So coming back to here.

1:23:24

SPEAKER_04

Let's say we didn't do a run this evaluation again. So we made the change to the prompt and now we're running the remediated version to see how that performs. Perfect. [SPEAKER_04] Okay. [SPEAKER_03] So that's around the experiment here with the new changes, hopefully.

1:24:08

SPEAKER_04

[SPEAKER_01] And I'll pivot into the UI.

1:24:18

[SPEAKER_03] I'm going to go to experiments here.

1:24:30

SPEAKER_04

[SPEAKER_03] And it's barely there, but you can just see tapered back up.

1:24:43

[SPEAKER_03] Any improvement going up is that improvement.

1:25:14

[SPEAKER_03] But yeah, just to give you an idea here that now that we've done the evaluation, actually what I can do is do a diff.

1:25:51

[SPEAKER_03] And we can do a comparison in the delta.

1:26:07

[SPEAKER_03] Yeah. [SPEAKER_03] Yeah. [SPEAKER_03] So, yeah, that's come up and improve over time. [SPEAKER_04] Which way, which way, the intended outcome. [SPEAKER_03] Yeah. [SPEAKER_03] Okay. [SPEAKER_03] So, I think we're approaching the end of the content.

1:26:49

[SPEAKER_04] So, I know it's been a number of steps.

1:27:03

[SPEAKER_04] So I really want to thank you for your time and attention to walk through that.

1:27:12

[SPEAKER_04] As mentioned, the artifacts are public. [SPEAKER_04] We've got the cheat sheet there.

1:27:42

[SPEAKER_04] We've got the Slack channel to help if you have any questions, but hopefully just to give you a summary of what you've accomplished today.

1:27:57

Uh, in this order is you went from taking a single shot prompt into building a five stage agentic workflow using tool calls.

1:28:12

[SPEAKER_04] What we were able to then do is then inspect how this works pretty much by diving into this by adding brain trust tracing, making sure everything is recorded.

1:28:27

[SPEAKER_04] We also then want to talk about how do we then evaluate the system from what is not online, something new, creating those effectively golden set with those test cases, which we want to execute against.

1:28:35

SPEAKER_04

[SPEAKER_03] We then deployed those managed prompts, those tools and parameters into the brain trust secure architecture to be able to use.

1:28:44

SPEAKER_04

[SPEAKER_03] And we've also added online scoring to then evaluate the system as it unfolds. And then we picked a particular production failure. [SPEAKER_03] We looked at the trace, we modified the prompt in our case and we saw the delta there and running the evaluation. [SPEAKER_03] Again, we saw it achieve back up to where it needed to be.

1:29:08

SPEAKER_04

[SPEAKER_04] And in this case, completing the full evaluation of building, observing it, deploying it and taking note of that moving forward. [SPEAKER_03] So yeah, just to call it and bring it home.

1:29:25

[SPEAKER_03] So again, hopefully this is not uncommon, but again, what might work in production is not really going to work in prototype.

1:29:45

SPEAKER_04

[SPEAKER_03] We really need to break this down, identify the failure modes and move forward. [SPEAKER_03] And that's where, again, explicit stages become really important, right?

1:30:00

SPEAKER_04

Again, this does introduce more areas where things could go wrong, but it's easier to debug if that's the case.

1:30:04

SPEAKER_04

[SPEAKER_03] Again, there's no substitute for diving into the code and tracking everything. [SPEAKER_03] So I would say observability's table stakes at this point. [SPEAKER_03] So, just to call it out and bring it home: hopefully this is not uncommon, but what might work in production is not really going to work in prototype. We really need to break this down, identify the failure modes, and move forward. That's where explicit stages become really important, right?

1:30:15

SPEAKER_03

[SPEAKER_04] This does introduce more areas where things could go wrong, but it's easier to debug if that's the case.

1:30:28

SPEAKER_03

There's no substitute for diving into the code and tracking everything. So I would say observability is table stakes at this point. If you've got a production application and you're not tracing it, you need to go back to the drawing board and get that done operationally. Hopefully, we can show you how to do that using BrainTrust. There's no substitution for production logs, but it's better to start somewhere.

1:30:38

SPEAKER_03

[SPEAKER_04] If you have an idea of what an issue might be, these are perfect ways to supplement evaluation. Using those failure modes as your test cases. As I mentioned, this is a continuous process, right? Nothing's ever done. If you've ever worked in agile development, constant feedback is important.

1:30:40

SPEAKER_04

[SPEAKER_03] We're bringing this operation model with a newer surface of operating. Hopefully, bring it home to your teams here today. My encouragement to you is: if this is something of interest, pick something that's already operational. It doesn't have to be the entire suite. Maybe start with something that's more critical that you really want to improve operationally. Add the tracing, collect your edge cases from that mode, build automatic scores, and then route everything back as possible. The faster feedback loop you have, the more insight you have, and the more you can improve the overall delivery operations of your system.

1:30:46

SPEAKER_03

So, just a call to action. I know we've thrown a lot of content at you. We're trying to put some systems in place. I appreciate everyone juggling everything to do this. As mentioned, you have to start somewhere. Let's try to accelerate you. We have a list of documentation provided. You can also use our AI agent to search if you have any questions. We also have a cookbook available. I tend to throw that cookbook directly into Cursor, Codecs, or whatever, and say, based on this, take the SDK and start tracing my application. It works pretty effectively. We even have Vuvashi announced a CLI for the BrainTrust application. That allows you to do things like auto-instrumentation. I'm going to plug my colleague, Eric, who's doing some fast-tech work. Please check out his booth. It's amazing. If this is something that interests you and you want to explore more, please reach out to your account team at BrainTrust. We're happy to support where possible. If you're on Discord, feel free to join. I can answer your questions there. Again, I really want to thank you for your time, attention, and energy. I know it's a really sunny day, and I don't want to keep people in here. I want you to get some fresh air. But on behalf of BrainTrust, TrayLine, and myself, thank you so much for your time and attention. It's been an honor, and I look forward to seeing you out there tracing and gaining value from delivering AI in production. Thank you.

1:30:57

SPEAKER_03

So as well in the UI, again, part of the operationalization, we can see who's changed what, but actually what has been changed to that particular prompt in this case.

1:31:13

SPEAKER_03

Taking care. So including two facts. Print that out. Okay. So coming back to here. Let's say we didn't do a, um, run this evaluation again.

1:31:34

SPEAKER_04

So we made the change to the prompt and now we're running, um, the remediated version to see how that performs.

1:32:05

SPEAKER_04

Perfect. Okay.

1:32:06

SPEAKER_03

So that's around the experiment here with the new changes, hopefully.

1:32:11

SPEAKER_01

And I'll pivot into the UI.

1:32:13

SPEAKER_03

I'm going to go to experiments here.

1:32:17

SPEAKER_03

And it's barely there, but you can just see, you know, kind of tapered back up. Just, you know, any improvement going up is that improvement. But yeah, just to give you an idea here that, um, you know, now that we've done the evaluation, uh, actually what I can do is, um, do a diff.

1:32:39

SPEAKER_03

And we can do, uh, do a comparison in the delta. Yeah.

1:32:51

SPEAKER_03

Yeah. So, yeah, that's come up, uh, and improve over time.

1:32:55

SPEAKER_04

Um, uh, which way, which way, uh, the intended outcome.

1:33:00

SPEAKER_03

Yeah.

1:33:03

SPEAKER_03

Okay.

1:33:07

SPEAKER_03

So, I think we're approaching the end of the, the content.

1:33:12

SPEAKER_04

Um, so, I know it's been a number of steps. So I really want to thank you for your time and attention to kind of walk to that. Uh, as mentioned, the, the artifacts are public. We've got the cheat sheet there. We've got the Slack channel to help if you have any, any questions, but, uh, hopefully just to give you a summary of what you've accomplished today. Uh, in this order is you went from taking a single shot prompt into building a five stage, uh, agentic workflow using, uh, tool calls. What we were able to then do is then inspect how this works pretty much into been diving into this by adding brain trust tracing, making sure everything is recorded.

1:33:49

SPEAKER_04

Um, we also then want to talk about, you know, how do we then evaluate the system from, um, um, you know, what it's not online, something new, uh, creating those, those, uh, effectively golden set with those test cases, which we want to execute against.

1:34:03

SPEAKER_03

We then deployed, uh, uh, those, uh, those managed prompts, those tools and parameters into the brain trust secure architecture, uh, to be able to use. And we've also added online scoring to then evaluate the system, uh, as it unfolds.

1:34:15

SPEAKER_04

And then we picked a particular, uh, production failure.

1:34:20

SPEAKER_03

We looked at the trace, uh, we modified the prompt in our, uh, case and we saw the Delta there and running the evaluation. Again, we saw it, uh, achieve, uh, back up to, to where it needed to be.

1:34:31

SPEAKER_04

And in this case, completing the full evaluation of, you know, building, observing it, deploying it and taking note of that moving forward.

1:34:39

SPEAKER_03

Um, so yeah, just to kind of, uh, call it, um, and bring it at home. So again, hopefully this is not uncommon, but again, what might work in production is not really going to work in prototype. Uh, we really need to break this down, identify the failure modes and, and move forward. And that's where, again, explicit stages become really, really important, right?

1:34:59

SPEAKER_04

Um, again, this does introduce more, uh, areas of, of where things could go wrong, but it's easier to debug if that's the case.

1:35:07

SPEAKER_03

Um, again, there's no substitute for diving into the code and tracking everything. So I would say it's, it's observability's table stakes at this point. If you've got a production application and you're not tracing it, you need to go back to the drawing board and get that done operational. And hopefully again, we can show you how to, to be able to do that, um, using brain trust. Um, again, no substitution for production logs, but better to start somewhere from that way.

1:35:33

SPEAKER_04

If you have an idea of what an issue might happen, these are your perfect ways to, to supplement evaluation. So using those, uh, failure modes as, as your test cases. And as I mentioned, this is a continuous process, right? Nothing's ever done. If you've ever worked in agile development, constant feedback is, is important.

1:35:49

SPEAKER_03

And again, we're bringing this, this operation model, but with, uh, a newer surface of, of, of operating. Um, and yeah, just hopefully bring it home to your teams here today. So, um, you know, my encouragement to you is to, if this is something of, of, of interest is pick something that's already operational day. It doesn't have to be the entire suite, maybe start off with something that's maybe a bit more, um, uh, I guess more critical that you really want to improve, um, the operational modes. Add the tracing, collect your edge cases from that, that mode, um, build automatic scores, and then route everything back as possible.

1:36:29

SPEAKER_03

So again, the faster feedback loop that you have, uh, you have more insight and then the more that you can, again, improve the overall, uh, delivery, uh, operations of, of your system. Um, yep. And just kind of call to action. So again, I know we've thrown a lot of content at you. It's, uh, well, obviously trying to get a bit of feedback, we're trying to put some tables in place. So I appreciate everyone's kind of juggling everything to do this. But, um, as mentioned that you have to start somewhere. Let's just try to accelerate you. Uh, we have a list of documentation that's provided.

1:37:00

SPEAKER_03

Um, that's, uh, you can also use our AI agent on the agent to, to search if you have any questions. But we also have a cookbook available. So I tend to throw that cookbook directly into, you know, cursor, codecs, or whatever, and say, you know, based on this, take the SDK, uh, and start tracing my application. It does it pretty effective. We even do have, uh, Vuvashi announced a CLI, um, for our, the BrainTrust, um, application. So that allows you to even do things as auto-instrumentation. I'm just going to plug my colleague, Eric, who's doing some fast-tech work. And please check out his booth. Because it's a, it's, it's amazing.

1:37:38

SPEAKER_03

Um, and again, if you, this is something that's interested you and you want to explore more, then, you know, please reach out to your account team at BrainTrust. Again, we're happy to support where possible. Uh, and if you're on Discord, you know, feel free to, to join. I have to answer your questions there, uh, from that. Um, and, uh, yeah, again, really want to thank you for your time, attention, energy. I know it's a really sunny day, and I don't want to keep people in here. I want you to get some fresh air. But, uh, yeah, just on behalf of BrainTrust, TrayLine, and myself, like, thank you so much for your time and attention.

1:38:07

SPEAKER_03

It's, it's been an honor, and look forward to seeing you out there tracing and, and gaining value from delivering AI and production. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note