Open Reader

Does GenAI "belong" to data scientists? — Phil Hetzel, Braintrust

completed 18:53 May 25, 2026 Watch on YouTube

Current Status

completed

Video ID

NKwIX3CiRgU

RAG / Chat

Enabled
Does GenAI "belong" to data scientists? — Phil Hetzel, Braintrust
Description

At most traditional enterprises, GenAI got handed to the ML platform team because it had AI in the name. Phil Hetzel from Braintrust argues that was the wrong move, not because data scientists lack value, but because Anthropic and OpenAI already ran the data pipeline. What is left is prompt and context engineering, distributed systems, human annotation, and functional evaluation across a much broader surface area than precision and recall. The mistake is isolating it to one team. The answer is a diverse one. Speaker info: - https://www.linkedin.com/in/philliphetzel

Summary

Generated by claude-sonnet-4-5

30-second take

Phil Hetzel (Braintrust Solutions Engineering lead) argues that agent development should not be siloed to data scientists/ML engineers but requires diverse teams including product engineers and domain experts. His core thesis: traditional ML teams excel at model training/testing pipelines, but LLMs are pre-trained APIs where value comes from prompt/context engineering and subject matter expertise—not feature engineering or model retraining. He observes AI-native startups build cross-functional teams with problem proximity, while traditional enterprises wrongly delegate GenAI to existing ML platforms just because "AI" is in the name. The talk positions evals/observability as the new quality discipline for agents in production.

Key takes

  • Agent development is fundamentally different work than traditional ML: The model training/testing pipeline is already done by OpenAI/Anthropic/Mistral; value now comes from prompt engineering, context engineering, and understanding user problems—not feature engineering or data pipelines. This shifts required expertise away from pure stats/math backgrounds.
  • Traditional enterprises make a structural mistake: They hand GenAI to existing ML/data science teams because "AI" is in the name, but these teams often lack proximity to the actual problem the agent should solve and default to irrelevant metrics (precision/recall/F1) instead of functional performance evaluation.
  • AI-native companies structure teams around cross-functionality: Smaller startups building on LLMs have engineers who are "agile enough to grow" and cross-functional across product, engineering, and AI—everyone has proximity to the problem rather than segmented responsibilities.
  • Data scientists still add critical value as "adults in the room": They should provide guardrails (understanding LLMs are token prediction, not knowledge), validate LLM-as-judge evals with traditional metrics, and handle fine-tuning when needed—but shouldn't own the entire agent build.
  • Domain experts/non-technical people should control prompts and labeling: Those closest to the problem need to do human annotation, prompt engineering, and define what "good" looks like—not be blocked by technical gatekeepers. Product engineers handle API integration and complex distributed systems orchestration.
  • Eval surface area is broader than traditional ML: Agent quality requires evaluating functional performance across multi-step reasoning, tool use, and user experience—not just binary classification metrics that ML engineers obsess over.
  • Observability = production confidence: Evals happen pre-production during experimentation; observability ensures agents remain reliable when confronted with real users and edge cases, forming a feedback loop to improve offline eval datasets.

Useful details

  • Phil's background: 12 years consulting, led global Databricks practice at Slalom; saw customers build many GenAI POCs but struggle to reach production—started using Braintrust as customer, then joined ~1 year ago.
  • Braintrust positioning: Agent quality platform with two pillars—evals (experimentation/confidence before production) and observability (monitoring production execution). Human labeling component and agent/prompt playground for domain experts.
  • Concrete team composition recommendation: Data scientists provide guardrails/evals/fine-tuning; product/systems engineers implement requirements and infrastructure; non-technical domain experts do annotation and prompt/context engineering.
  • Example organizational pattern: Traditional enterprise—CEO reads magazine → delegates to CIO → hands to existing ML platform team. AI-native—small cross-functional engineering teams with no legacy structure.
  • LLM-as-judge caveat: Teams are "very tempted to just believe LLM as judges" but these are just prompts/models; data scientists should create labeled datasets and apply traditional precision/recall/F1 to validate the judge itself.
  • Systems complexity example: Complex agents may have supervisor + child/sub-agents running on distributed infrastructure calling different systems—a systems engineering problem, not a stats problem.
  • Feedback loop design: Gather production data → continually add to offline eval dataset → check if evals align with human agreement or diverge over time.

Caveats / counterpoints

  • Phil doesn't address cost/latency tradeoffs: No discussion of when fine-tuning vs. prompt engineering makes economic sense, or how product engineers would make model selection decisions without ML expertise.
  • "Models are already built" oversimplifies: The talk ignores RAG pipelines, embedding model selection, vector DB tuning, and retrieval quality—areas where ML expertise clearly applies to GenAI systems.
  • Non-technical prompt ownership may create versioning chaos: No mention of how to govern prompts, prevent regression, or handle A/B testing when domain experts directly edit production prompts.
  • "Proximity to the problem" assumes small organizations: In large enterprises, domain experts may be far removed from engineering teams; the talk doesn't address how to bridge this at scale.
  • Limited acknowledgment of when data scientists should lead: Fine-tuning is mentioned almost as an afterthought; doesn't explore scenarios where custom model training is strategically valuable.
  • No discussion of security/compliance roles: Agent systems calling external APIs and handling sensitive data likely need specialized oversight Phil doesn't mention.

Ken relevance

High relevance for AI ops and GTM strategy. This directly challenges how you'd staff agent development teams and where to invest in hiring/capability building. Key implications:

  • Team design for your agent products: If building commercial agent systems, this argues for cross-functional pods (not ML team silos) with domain experts controlling prompts and product engineers handling orchestration—aligns with your "builder" philosophy.
  • Investment screening: AI-native startups with cross-functional teams may execute faster than traditional ML companies trying to pivot into agents—a structural advantage worth filtering for.
  • Your own workflow: Suggests you (as non-data-scientist domain expert) should directly control prompts/evals for your personal agent systems rather than delegating to ML specialists—lowers friction for building custom tooling.
  • Content angle: The "agents don't belong to data scientists" thesis is contrarian clickbait but substantive—could be a good podcast topic or essay exploring when this breaks down (e.g., AlphaFold-style agents still need ML depth).
  • Braintrust as tool: The observability + eval platform positioning is relevant if you're operationalizing agents at scale—worth evaluating against Langsmith, Helicone, etc.

Caveat: The talk is Braintrust marketing disguised as org design advice; Phil's incentive is to make agent development seem accessible to non-ML buyers.

Watch verdict

Skim. The core argument (agents need diverse teams, not just data scientists) is valuable and well-structured, but the 20-minute talk doesn't offer enough tactical depth to justify full attention. The slides are clear and scannable. Worth reading transcript for the team composition framework and AI-native vs. traditional enterprise patterns, but the Q&A is low-signal and Phil rushes through slides to hit time. If you're designing agent teams or hiring, read the summary; if you're a solo builder, skip—you already embody the "domain expert controlling prompts" model.

Transcript

3002 words en Processed in 198.7s

What we're going to do today is we're going to talk about whether agents or agent development really belong to data science or machine learning engineers. How many people here would describe themselves as either a data scientist or b machine learning engineer? This is going to be awesome. I'm glad that no one has brought any rotten tomatoes because the answer that I'm going to give is probably not going to be exactly to your liking, but give me a chance to justify why. What we're going to do today is talk through these couple things. I'll introduce myself, introduce the company that I work for, and then we'll get into the topic. I'll probably rip through the slides pretty quickly and hopefully give some time for Q&A. But before I do that, I'll introduce myself. My name is Phil Hetzel. I lead the Solutions Engineering team at Braintrust. Braintrust is the people that allow our customers to get the most value out of the platform as quickly as possible. Prior to Braintrust, I spent 12 years in consulting and systems implementation. My last role in consulting was leading the global Databricks business unit at a company called Slalom Consulting, and I noticed that a lot of my customers were really prolific at creating generative AI proofs of concepts, but not nearly as good at bringing those proofs of concepts to production. So I started using Braintrust as a user first, and I liked the product so much that I applied for a job, and I've been here for about a year since that happened. Outside of work, I like to play chess, but I'm not very good at it. And I like to spend time with my wife and my dachshund, Pistol Pete. He's pictured over there. He's the one in brown and not the one in black. The one in black is me. The one in brown is him. What is Braintrust? Braintrust is an agent quality platform. The way that we perform agent quality is two different pillars. There's evals and observability. Evals are the things that you're doing in experimentation as you're tweaking and building your agent to become confident in your agent's execution once you push it to production. Agent observability to us means that once it is in production, that you remain confident in its execution once it's confronted with real usage and real users. There are other ancillary things that the platform does, but in general, that's what we do. I'm really not going to talk about the product today. If you're interested about the product, you can find me at the booth downstairs. But other than that, very happy to get into the content. Just some observations that I've seen through the last year of watching some of the top teams building agents across many different industries. I think there are two different types of organizations that we work with. There's the traditional enterprise and there is the AI natives. Traditional enterprise approaches agentic development a little bit differently than AI natives. Traditional enterprise: a person in charge, a person of note, CEO or CIO will read something in the CIO or CEO monthly magazine that says that they need to be building agents. And then they'll tell their delegate that you need to be building agents because that's the thing that's going to take us to the AI promise land. And then that will get further delegated to an existing ML or data science platform team who already have a lot of the tooling in place. And since generative AI has AI in the name, it's a pretty natural fit to hand over that capability to an existing AI or data science team. Is anyone in that bucket today where they just got handed generative AI because that seemed like the best fit? Yeah, got it. And there's no judgment or connotation here. Just something that I've observed. There's a whole other set of companies that are more AI native that think less about what already exists because nothing really existed before generative AI started to gain popularity for these companies. In fact, they started building their entire offering around agents. So rather than having an AI ML platform team, they'll just have a small team of engineers that are agile enough to grow at the times. And rather than having very specific segments of things that they do, everyone is very much cross functional across both product engineering and AI engineering. The other thing that's interesting about these AI natives is that since these are typically smaller companies, each person has more proximity to the problem, i.e. they have a better understanding of what the end agent is actually meant to solve. Two different differences between traditional ML and generative AI. The model is already built. So much of what data scientists and machine learning engineers are going through is that data pipeline of training a model. What do we do when the model is already built? And the other interesting thing, other interesting nuance, I should say, is that if you want to add value to these models, then you can add values not necessarily with feature engineering, but with natural language, which could bring in a different skill set to the conversation. So just to make this more clear, this is more complicated than that, but abstracted to a certain level, this is what data scientists and machine learning engineers do: a data pipeline of training and testing, making sure that you're not overfitting, eventually deploying that model where it can be used by some downstream product team. That is what data scientists and machine learning engineers are used to doing. This, though, has already been done. Anthropic and OpenAI and Mistral, they've already done the data process of grabbing that data, putting it through the pipeline, training the underlying LLM, and then deploying it through an endpoint so that their consumers can use it. So the one nuance here is that while Anthropic and OpenAI and Mistral will be doing testing of their own, we still need to, as AI teams, we still need to perform evals after we've implemented those APIs into our product. That needs to be very important to us. So that's really the only nuance here between these two images. How do you change this? This is going back to these two differences before. The second difference, how do we change these predictive applications before and after? Traditional ML, you either add more data to retrain it, or you're performing feature engineering to adjust how the underlying model is performing, and then you're performing a lot of A-B testing to understand how your model changes have provided lift or not provided lift. With generative AI, since that model is already trained, and irrespective of performing any fine tuning on that model, which is pretty rare, the way that you can change that behavior is just by changing the inputs, the prompts, the context that you're giving that model. So there are a lot of folks that will be performing context engineering on top of these models, and those folks could have a better understanding of how real users could be using the agent. They'll have closer proximity to the problem. So I'm going to make the case for and against agents belonging to data scientists. Let's say I was debating the position of it really does belong to data scientists and traditional machine learning engineers. Agents use models. In our organization, models are governed by data scientists. So data scientists will have a lot of underlying knowledge about how neural nets work and thus how LLMs work. Because of that, they're going to have a far better appreciation of the risks inherent with using this very complex technology. So there's a lot of folks that will be performing context engineering on top of these models, and those folks could have a better understanding of how real users could be using the agent. They'll have closer proximity to the problem. So I'm going to make the case for and against agents belonging to data scientists. Let's say I was debating the position of it really does belong to data scientists and traditional machine learning engineers. Agents use models. In our organization, models are governed by data scientists. So data scientists will have a lot of underlying knowledge about how neural nets work and thus how LLMs work. Because of that, they're going to have a far better appreciation of the risks inherent with using this very complex technology. Other thing here is that they'll have very rigorous processes to push models and model assets to production. They will understand some type of testing process that they can use to keep the company safe and make sure that end users are getting the experience that they need. And then number three, very related, just very rigorous mindset around testing. The counterpart to that is, again, models are already built. So we don't necessarily need to do any training and testing. Entirely different pipeline. We're not doing the whole cross validation dance. And this is probably the biggest argument. Does an ML engineer or a data scientist really know what they're testing for? One of the things that I've noticed with some of these teams is they will really lock on to the traditional ML engineer metrics like precision recall F1. And they'll obsess over those metrics because that is what has gotten them there up to that point. But when you're analyzing agents, it is far broader of a surface area that you need to be evaluating. You need to be evaluating the functional performance of that agent rather than just a technical performance across that two box that we're traditionally used to working with. So argument here is let's say that I was arguing that agents belong to non-data scientists, which could be both technical and non-technical experts. We could make the case that LLMs are just APIs. Product engineers are very used to using APIs as they build applications. That is a massive part of what they do is reaching out, grabbing information from another system based upon some payload and bringing that information back in a way that's useful to the end users. That is a thing that product engineers do. The other thing about agents that is unique is that if you have a very complex agent, it could be running across many different types of compute if it's some distributed agent. For example, you have a supervisor agent up here and then it's calling different child or sub-agents that might be running on different infrastructure. And as they're running on different infrastructure, they might be calling different systems as a result. That can be a very complex systems problem that might not be up the alley of someone with more of a statistics or maths background. And then finally, more on the non-technical side, it's really valuable to have subject matter experts or product managers be able to control the actual prompts that we're seeding the agent with. These people are the ones that have the most proximity to the problem that the agent is trying to solve. So there's a lot of lift in having a non-technical person have a lot of say in how the agent performs. Not only that, there can be a very large human annotation workflow that goes into making great agents where as you see these interactions, if you are a non-technical person but has a lot of domain expertise for how the agent is supposed to be performing, that non-technical person can look into an agent trace and describe whether or not the agent is performing well or not performing well, and most importantly, why that's the case. So where I'm landing with all this is not that all of the people that raise their hand that says yes, I am proudly a data scientist or ML engineer. I am not going to stand here and say, well, guess what, you need to completely refresh your skill set. That would be a very silly thing to say because I kind of figure there would be a lot of data scientists in the room and I'm at least smarter than that. But it does make sense to have a very diverse team when you're building these platforms. It makes sense to bring both non-technical and different types of technical people into the fold. How can data scientists add value to building agents? A couple of different ways. Also, all these ways are irrespective of actually helping to build the product itself. I think that's inherent. With the tools that we have available to us now, it's actually quite easy for us to be able to add value to a product, even if you are not coming from a product engineering background. But what I think is really valuable is data scientists can add, to use an overloaded term, add the guardrails to this process. A lot of people are very aggressive in how they implement LLMs. They don't understand how the underlying technology works. They don't come from a stats background. I think data scientists can be the adult in the room during those situations and say, you know, the LLM, this is how it's trained. It's just predicting token after token. It doesn't actually know anything really. It's just a bunch of stats problems at the end of the day. I also think that LLM as judge is a huge part of the eval process when you're building agentic applications. Again, people are very tempted to just believe LLM as judges when they're performing evals. They're just prompts and models at the end of the day. And it's very easy to be able to create some label dataset and perform the traditional recall, precision, and F1 style metrics on those, which data scientists will have expertise in. And the last one, this is the most technical one, of course. If you do need to fine tune an open source model very specifically to your use case, that's probably going to be the most fun and technical thing where data scientists and machine learning engineers can add a ton of value. The ideal mix here, in addition to that top section, we want both product application and systems engineers to be able to implement those requirements into the product itself that the non-technical experts are giving to them. We want to make sure that the systems that we're building around these agents, i.e. where the agents are executing, is such that it's going to lead to a great user experience. And then finally, and this is probably something that the data scientists can pitch into as well, implement actual eval and observability pipelines so you have that feedback loop of what's happening in production and what's happening in experimentation. For non-technical experts, we want them to be performing a ton of human annotation and a lot of prompt and context engineering. They have the closest proximity to their problem. You need to bring them into the fold if you want to have a very relevant agent to your use case. So what's next? Answer is always in the middle. So I hope I didn't fully insult half the people in the room today. If I have, then you can feel free to come to the Brain Trust booth and give me an earful. Implement actual eval and observability pipelines so you have that feedback loop of what's happening in production and what's happening in experimentation. For non-technical experts, we want them to be performing a ton of human annotation and a lot of prompt and context engineering. They have the closest proximity to their problem. You need to bring them into the fold if you want to have a very relevant agent to your use case. So what's next? The answer is always in the middle. So I hope I didn't fully insult half the people in the room today. If I have, then you can feel free to come to the Brain Trust booth and give me an earful. That's completely fine. But the idea here is that there's a ton of value for data scientists. Just make sure that you're bringing more folks into the room as you're building agents. Two minutes for questions. I know that we're keeping a very tight timeline. Yes, sir? Yeah, really good. I like the conclusion. But I do have a question on the framing. So I view agents as a tool. So anybody in the organization could build and own an agent in theory. And they may have more domain expertise in the data science, the engineering side. Rather than thinking about who owns the tool—you know, this is the machine learning team or that's the computing team—how do you view it as thinking about it based on the problem that's being solved and seeing agents as a tool to solve the problem rather than the agents as the— I think you're thinking about it the exact same way. It's a product that a diverse team builds. I think the mistake that I see a lot of typically traditional companies make is they say, "Oh, we're making another predictive model." Yeah. And they isolate it to the ML engineers or data scientists and say, go build these agent things. I think we're actually thinking about it very similarly. Cool. Yeah. Maybe we'll have time for one more question. Yes? Yeah, I really echoed with the. I think I'm just curious about the tool that you are asking about. And I think that, at least from my experience, the missing delta is the tool going to facilitate the enter, a traditional machine learning, but into intra-intra variation in the. Is that something that your interest is thinking about? Yeah. Yeah, how to make it easy for a domain expert to update the system and update the system? Yeah, for sure. Yeah, there's a lot of things that we do to lean into that domain expert persona. We do have a human labeling component as a part of our platform. And we do have an agent and prompt playground where people can experiment with their own prompts and send them to the underlying agents themselves. [SPEAKER_02] Right. [SPEAKER_02] Yeah. [SPEAKER_02] How to understand what is the system of way that you keep the evaluator after the data, and also the system after the data, and understand how you make error analysis, like the error is based on that evaluator? [SPEAKER_02] Mm-hm. [SPEAKER_02] The idea is that we gather data from production to continually add to that offline data set that we're evaluating upon. [SPEAKER_02] And then, hopefully, we're gaining grounded data along the way where we can self-check ourselves to understand if our evals are aligning, starting to align more to human agreement or not, or if they're diverging. Yeah. Okay, everyone, that's my time. I really appreciate the attention today. If there are any more questions, find me downstairs. They have the closest proximity to their problem. You need to bring them into the fold if you want to have a very relevant agent to your use case. So what's next? Answer is always in the middle. So I hope I didn't fully insult half the people in the room today. If I have, then you can feel free to come to the Brain Trust booth and give me an earful. That's completely fine. But the idea here is that a ton of value for data scientists. Just make sure that you're bringing more folks into the room as you're building agents. Two minutes for questions. I know that we're keeping a very tight timeline. Yes, sir? Yeah, really good. I like the conclusion. But I do have a question on the framing. So, you know, this, I view agents as a tool. So anybody in the organization could, you know, build and own an agent in theory. And they may have more domain expertise in the data science, the engineering side. Rather than thinking about who owns the tool, you know, this is the machine learning team, or that's the computing team. How do you view it as thinking about it based on the problem that's being solved and seeing agents as a tool to solve the problem rather than the agents as the- I think you're thinking about it the exact same way. It's a product that a diverse team builds. I think the mistake that I see a lot of typically traditional companies make is they say, oh, we're making another predictive model. Yeah. And they isolate it to the ML engineers or data scientists and say, go build these agent things. I think we're actually thinking about it very similarly. Cool. Yeah. Maybe we'll have time for one more question. Yes? Yeah, I really echoed with, yeah, the . I think, I'm just curious about the, actually, the tool that you are asking about . And I think that, at least from my experience, the missing delta is the tool going to facilitate the enter, like a traditional machine learning, but into intra-intra-like variation in the . Is that something that your interest is, like, thinking about? Yeah. Yeah, like how to make it easy for a domain expert to update the system and update the system? Yeah, for sure. Yeah, there's a lot of things that we do to lean into that domain expert persona. We do have a human labeling component as a part of our platform. And we do have, like, an agent and prompt playground where people can experiment with their own prompts and send them to the underlying agents themselves. Right. Yeah. I mean, I mean, I mean, I mean, I mean, I mean, like, how to, how, what is the system of way that you keep the evaluator after the data, and also the system after the data, and, like, understand how you make, uh, I guess the error analysis, like, the error is based on that evaluator . Mm-hm. The idea is that we, we gather data from production to continually add to that offline data set that we're evaluating upon. And then, hopefully, we're, uh, gaining grounded data along the way where we can kind of self-check ourselves to understand, um, if our evals are aligning, starting to align more to human, uh, agreement, agreement, or, or not, or if they're diverging. Yeah. Okay, uh, everyone, that's my time. I really appreciate the attention today. If there are any more questions, find me downstairs.