Open Reader

AI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash

completed 16:11 Aug 28, 2026 Watch on YouTube

Current Status

completed

Video ID

bMjlRrWjdT0

RAG / Chat

Enabled
AI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash
Description

The people annotating DoorDash's eval data are not engineers, and they build their own annotation tools. Because the GenAI platform team went API first, strategy and operations staff can point a coding agent at those endpoints and vibe code whatever interface their use case needs, whether that is grading restaurant menus or reviewing images. The platform team stopped trying to anticipate every UI, and shipped stable APIs instead. Nachiket Paranjape and Swaroop Chitlur Haridas make the broader case that evals stopped being an engineering harness for them and became a cross functional job. That reframing has an org chart attached. Strategy and operations set the quality bar, product managers turn it into rubrics, operations run the annotations, and engineering supplies telemetry, datasets and judges. Which group actually owns a judge prompt varies by team, and they treat that variation as a sign the org is still learning rather than a problem to standardize away. The loop underneath is deliberately plain: trace, sample down to something a human will really look at, annotate, promote a golden set, calibrate the judge against it, then monitor and go again. Judge calibration runs self serve through a UI, showing the original and optimized prompts side by side so a product manager can see what changed and decide whether to trust it. Per annotation cost fell sharply. Speaker info: Nachiket Paranjape: - https://x.com/nmparanjape - https://www.linkedin.com/in/nachiketparanjape/ Timestamps: 0:00 - The GenAI platform team, and its three forces 2:05 - Why eval became the fourth pillar 3:05 - UI first, then API first, then workflow first 4:01 - Evals as a team sport, not an engineering harness 4:57 - Who owns which part of quality 5:53 - The continuous loop: trace, sample, annotate, calibrate 7:42 - Telemetry and workflow as two surfaces 9:32 - Operators vibe coding their own annotation UIs 11:21 - Calibrating judge prompts, self serve 13:10 - Different teams, different promp

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: DoorDash treats AI evaluation as a cross-functional operating system—not an engineering test harness—built around a self-serve loop from production traces to human annotations, golden datasets, calibrated judges, and continuous monitoring.
  • Why it matters: The talk offers a directly reusable control-plane pattern for scaling agent and LLM quality across many product teams while balancing accuracy, latency, cost, security, and organizational ownership.
  • Best use: Use it to inform the design of an internal eval platform: centralize stable primitives and governance, but push workflow ownership and domain judgment to PM, operations, and labeling teams through APIs, UIs, and coding-agent-assisted workflows.

Executive Summary

DoorDash’s GenAI platform team supports product teams with shared LLM and agent infrastructure, including model switching, tool and agent connectivity, authentication, agent identity, open-weight model hosting, and evaluations. Its central argument is that evals must incorporate the people with domain knowledge—strategy and operations, product managers, and labeling partners—not remain an engineering-only capability.

Their operating loop is production-oriented: capture traces and sessions, sample the cases worth reviewing, annotate them with domain expertise, create golden datasets, calibrate LLM-as-a-judge prompts and workflows against those datasets, monitor results, and repeat. This makes evals a quality-improvement cycle rather than a one-time benchmark.

The platform separates a telemetry layer from a workflow layer. Telemetry holds traces, observations, scores, datasets, APIs, SDK access, and MCP access; the workflow surface enables non-engineering partners to set annotation tasks, review golden datasets, and calibrate judges. DoorDash began UI-first for nontechnical participation, became API-first to avoid central-team bottlenecks, and is now workflow-first, using coding agents to let operators create purpose-built interfaces.

The most practical implementation lesson is architectural: invest in stable APIs beneath the UI, then let teams compose their own workflows. DoorDash found that annotation requirements differed materially across use cases, but common underlying patterns allowed strategy and operations teams to use Codex or Claude Code to build custom annotation UIs themselves. The company reports lower per-annotation spend and faster iteration, though it provides no hard savings figures or validation metrics.

Key Takeaways

  • Claim: A scalable eval program requires cross-functional ownership because the inputs that define AI quality are distributed across operations, product, engineering, and annotation teams. | Evidence: DoorDash assigns strategy and operations teams to set priorities and quality bars; PMs translate requirements into rubrics and workflows; operations runs annotation; and engineering provides APIs, telemetry, datasets, and judges. | Implication: Ken should design eval systems around explicit quality-governance roles and handoffs, rather than assuming engineers can independently define what a good agent outcome is.
  • Claim: The durable eval loop begins with real production behavior and continuously converts human judgment into reusable evaluation assets. | Evidence: DoorDash’s loop is trace/session capture, sampling, annotation, review, golden-dataset creation, judge calibration, monitoring, and repetition. It cites session-level quality needs for a consumer discovery and shopping assistant, scaled human judgment for personalization ML, and trajectory-based evaluation for multi-agent systems. | Implication: For agent systems, Ken should evaluate complete trajectories and sessions where needed—not only isolated final answers—and turn reviewed production failures into a versioned regression set. | Caveat: The transcript describes the process but does not specify sampling policies, statistical thresholds, inter-annotator agreement, or rollout gates.
  • Claim: A central eval platform should expose stable APIs first, with UI and workflow layers built on top, so decentralized teams can move without waiting on the platform team. | Evidence: DoorDash says its APIs power scores, datasets, API access, and SDK access through one plane. It initially prioritized UI participation, then added API-first access for engineers, and now describes its approach as workflow-first. | Implication: Ken should treat telemetry, dataset, score, and evaluation-run APIs as foundational control-plane objects, rather than treating an evaluation dashboard as the product. | Caveat: Stable APIs create leverage only if access controls, schema/version management, and ownership boundaries are equally disciplined; those controls are not detailed in the talk.
  • Claim: Coding agents can make specialized evaluation workflows self-serve for non-engineering operators when the underlying platform primitives are accessible by API. | Evidence: DoorDash found it impractical for a platform team to build a bespoke annotation UI for every use case. With APIs in place, strategy and operations partners used tools such as Codex or Claude Code to create their own annotation interfaces for cases including image annotation and manual testing. | Implication: Ken can enable domain teams to build narrow, fit-for-purpose eval interfaces, but should provide approved templates, permissions, audit trails, and lifecycle ownership rather than allowing unmanaged shadow tooling. | Caveat: The examples demonstrate workflow flexibility, but the talk does not explain review, security, deployment, or maintenance standards for operator-generated applications.
  • Claim: LLM-as-a-judge should be calibrated against human-labeled golden datasets and made inspectable to the people responsible for quality. | Evidence: DoorDash starts with a judge prompt and baseline scores, uses the DSPy library for prompt optimization, and lets partners promote a calibrated prompt to the production judge. Its UI displays the original system prompt alongside the calibrated version and supports model selection, including Gemini, Claude, and OpenAI models. | Implication: Ken should require judge versioning, held-out validation, comparative reviews of prompt changes, and periodic recalibration from newly sampled production behavior. | Caveat: Prompt optimization can improve agreement with a labeled set without proving robustness to distribution shift, reward hacking, or systematic blind spots in the labels.
  • Claim: Self-service evaluation operations can reduce annotation cost and accelerate product iteration at scale. | Evidence: DoorDash says it has thousands of rows annotated weekly and reports that its self-serve annotation platform reduced per-annotation spend while allowing teams to iterate and calibrate judges faster without platform-team dependency. | Implication: The likely ROI case for an eval platform is not only model quality; it is lower coordination overhead and a faster feedback loop across many teams and use cases. | Caveat: No percentage reduction, absolute cost, quality metric, or before/after study is provided, so the economic claim is directional rather than independently assessable.

Detailed Brief

Platform scope: evaluation sits alongside model, agent, and security primitives

  • Claims: DoorDash positions evals as one of four platform pillars alongside LLM access, agent connectivity, and open-weight model hosting.; The platform team’s stated product value is helping product teams balance accuracy, latency, and cost; the speakers argue that this tradeoff applies to agents as well as to model selection.
  • Evidence: Its LLM gateway lets teams switch models and try newer options.; Its agent gateway connects tools and other agents while centralizing authentication and agent identity in a form security can approve.; The team pairs the LLM gateway with open-weight model hosting and says it has already seen significant cost impact, although it defers details to a future presentation.
  • Caveats: The presentation does not quantify the latency, accuracy, cost, or open-weight-hosting results.; It does not explain how model-routing decisions interact with evaluation scores or release policies.
  • Implications: Evaluation data should feed broader routing and deployment decisions, not live in a disconnected quality tool.; Authentication, agent identity, and security approval should be designed into the agent platform before proliferating team-specific tools and workflows.

Flexible organizational ownership of the judge prompt

  • Claims: DoorDash does not prescribe one universal owner for evaluation criteria or judge prompts.; Different teams can assign prompt ownership to strategy and operations, PM, or engineering according to their evolving operating model.
  • Evidence: The speakers explicitly say they have observed all three ownership arrangements internally and built their tooling to support that flexibility.; The calibration UI removes much of the configuration complexity so a product manager or operator can run a calibration loop without engineering mediation.
  • Caveats: Flexibility without a common rubric taxonomy and promotion policy can produce inconsistent quality bars across products; the talk does not describe a formal cross-team governance mechanism.
  • Implications: Separate central ownership of platform standards and auditability from local ownership of domain-specific criteria.; Define who can promote a judge, what evidence is required, and when local evaluation standards must be reconciled with company-wide safety or brand standards.

Notable Concepts & Terms

  • Continuous eval loop: DoorDash’s recurring workflow from traces and sampled cases through annotation, golden datasets, calibrated judges, monitoring, and re-iteration.
  • Golden dataset: Human-reviewed examples used as the reference set for measuring and calibrating evaluators, agents, and workflows.
  • LLM-as-a-judge: Using an LLM prompt as a scoring mechanism for outputs or trajectories, with its prompt calibrated against labeled data.
  • Telemetry layer: The programmatic plane containing traces, observations, scores, and dataset access through APIs, SDKs, and MCP.
  • Workflow layer: The operational surface where PMs and operations teams create annotation tasks, review datasets, and configure or calibrate judges.
  • API-first: The design decision to make stable platform primitives available programmatically before or beneath bespoke user interfaces.
  • Trajectory-based evals: Evaluation of the sequence of actions and tool/agent behavior in a multi-agent workflow, rather than scoring only the final response.
  • DSPy: The prompt-optimization library DoorDash says it uses in its judge-calibration loop.

Operator Notes / Why Ken Should Care

  • Define a minimum eval object model for traces, sessions, observations, scores, annotations, datasets, judge versions, and evaluation runs; make it available through stable APIs before expanding dashboards.
  • Establish a production-to-regression pipeline: sample real agent traces, route them to domain reviewers, adjudicate labels, and promote confirmed cases into versioned golden datasets.
  • Create a judge-promotion policy requiring a held-out set, a documented owner, visible prompt diff, model/version record, and rollback path before an LLM judge can gate releases.
  • Pilot an operator-built annotation workflow for one high-volume use case, but constrain it with approved templates, identity-based permissions, data-handling rules, logging, and maintenance ownership.
  • Track evaluation operations as a business system: annotation throughput, per-row cost, reviewer agreement, time from failure discovery to regression coverage, judge-human agreement, and release-blocking false positives/negatives.
  • Ensure agent trajectory traces include tool calls, identities, auth context, intermediate decisions, and final outcomes so security and quality review can use the same telemetry substrate.

Source/Metadata

  • Title: AI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash
  • Transcript words: 2553
  • Duration seconds: 971
  • Timestamp note: No timestamps or chapters were present in the provided transcript; the closing summary is duplicated in the source text.

Transcript

2392 words en Processed in 106.9s

Good afternoon, everyone. Thanks for coming for a post-lunch talk. Always appreciate that. My name is Svaroop, and here's my teammate Nachiket. We are here on behalf of the DoorDash JNI platform team. We wanted to share our evals journey. It started as evals is another engineering thing, but then it slowly evolved into a cross-functional effort, and we want to share our story here. So what is this team? This team is a Gen AI platform team. We are a horizontal team that helps all other product teams. So product teams at DoorDash build on top of the infrastructure and the primitives that we provide. And we see our USP and the value that we provide is that we help product teams balance these three forces, which is accuracy, latency, and cost. Initially, we applied this in terms of models, but if you think about it, it also applies to agents. And the way we achieve this is we have primitives and building blocks. So, for example, we have an LLM gateway where you can easily switch between different models and try the latest and greatest. We have an agent gateway where you can connect to tools and other agents. And we help solve authentication, agent identity, and other things in a central place, which our security team can bless. Similarly, we pair the LLM gateway with OpenWeights models hosting. Of course, cost is a number one concern these days. And we invested in OpenWeights models and have seen significant impact already. And maybe we'll talk about that in a future conference. The fourth pillar is evals, and that's the part that we would want to share today. When we started talking to product teams internally at DoorDash, there were varying distinct needs across teams. We had a consumer discovery and shopping assistant team. For those who attended Rago Stock earlier today, you will see the need for session-level quality judgments. Then the personalization ML team needed a way to scale up human judgment. And with multi-agent systems, we needed trajectory-based evals. Now, the question is, how do you cater to all these different needs under a common platform? And as we spoke to these teams, we realized we needed to empower the people who are the domain experts. And in our case, that was strategy and operations folks. It was product managers. It was even labeling partners, and not only engineers. So we started with, okay, we have to be UI first, and this was the guidance we had from Andy Fang, our co-founder, as well. So we had UIs for non-engineers to contribute. Then we evolved to also being API first so that engineers can also build and not be blocked on the central platform, and they can build their own systems. And then, of course, with the coding agents, now we have become workflow first, that we empower SNO and PMs to also be able to navigate the platform and run operations as well. So with that context, I'll hand it off to Nachiget to talk about how we went about delivering this. Cool. Thanks, Faroop. And thanks, everyone, for joining us. I know France is playing right now, and I promise you this will be better than that. I'm kidding. So as Faroop was saying, evals is not just an engineering harness. It is a cross-functional effort across different pillars, across different teams, that actually helps us add all the domain-specific knowledge into the quality of the AI itself. So from your traces to your data sets, from scoring mechanisms, this is all a team sport. We all have to play and help improve the quality of AI. So going a little bit deeper into the same aspect, we have different teams at DoorDash who help us actually improve the quality of AI. So you're going to have your strategy and operations folks who are going to set priorities, set the quality bar that you want to aim for. You're going to have your product people who are going to translate these requirements into rubrics, workflows. You're going to have your operations teams running annotations. You're going to have your engineering teams, like us, providing APIs, telemetry, data sets, judges, all the cool things. And combining all these together is what the recipe is for actually making sure that you are shipping quality AI products through an evals platform. So we've tried to boil this down into a continuous iteration loop. So right from tracing, having a tracing solution, viewing your sessions, your traces, to sampling them down to a very small set that you actually want to look at, annotating these with the domain-specific expertise that you bring in with the different teams I mentioned, reviewing those, then creating those golden data sets, which are going to be your golden data sets that you want to measure or calibrate against. And then, of course, monitoring this over a period of time, and then rinse and repeat, go through the whole loop again. So this, in our experience, has been a good continuous loop for shipping quality AI. On the platform level, we have two surfaces. So we have the telemetry layer, where we have all our traces, our scores, observations. That is also the plane where users are able to access these traces using an MCP, using an SDK, using our APIs. And then we have the workflow layer. This is where a lot of our stat ops, our product teams, operate on the platform. So this is where all the annotation tasks are set, this is where they review their golden data sets, create their judges, calibrate their judges, and so on. So maybe today we'll go through these four different modules or pillars of our platform step by step. So again, first one, tracing and sampling, which is actually capturing what your agents, what your LLMs are actually outputting, for lack of better words, and actually viewing those. Now, in order to also power this whole platform, we have, I think as Svarup mentioned, gone in an API-first approach. What that has allowed us to do is have these stable APIs that actually, and then you build UIs on top of that. So all our scores, our data sets, these are all powered by very stable APIs that our team owns. So all your API access, including SDK access, is powered by this single plane. Again, going back, and just refreshing your memory, step one, capture your traces, capture your sessions, measure your scores. Then you want to start adding all your judgment, your context, your domain knowledge, and then calibrating your judges is what we have seen as the whole life cycle. Step two is on the annotation side. So you obviously are capturing a lot of your agentic behavior, your sessions, your traces, but you actually want to see what are some places where things went well and what are some places where things did not go well. This is where you can actually titrate and actually look inside what's actually happening at the session level and annotate these data sets. And as Farooq mentioned, we have a lot of use cases. We talk to multiple different teams who have various ways of annotating their data sets. And it's almost hard for a platform team to build a UI specific for each use case. And, to give you an example, it's usually going to be an annotator who's going to annotate these data sets. So the platform team is in charge of the APIs. We have a strategy and ops person who's actually deciding what to annotate. And then you have an annotator who's actually going to annotate your data set. So we took this approach. Everybody has access to coding agents. And we actually doubled down on that API-first approach. So because we had these APIs, we were actually able to enable our stat ops teams to use something like a codex or a claw code and wipe code their own annotation UIs. So we had different use cases. I think we had a talk from Raghav before. We had image annotation use cases. We had some manual testing use cases. What stood out to us was the underlying patterns were similar. So if we are API-first, we can actually enable our partners to simply wipe code these UIs for annotation. So it's like a very simple example of a wipe-coded UI. Looks pretty clean, does the job, and you get the annotation that you need. This is basically a menu from a restaurant. It's nothing crazy. But the point I want to make here is that what helped us was to give this workflow in the hands of the operators so that they can actually build their own wipe-coded annotation UIs. So moving on, once you have these annotation UIs, you obviously want to calibrate your judge prompts. You obviously have some LM-as-a-judge metric that you're tracking. You want to now start improving that with these golden data sets. In order to do that, we have a pretty simple process. You're going to start with some judge prompt, take a look at what exactly do you want to measure from the output, have something simple. You're going to have your baseline scores, where you're going to simply run those LLM judges on your traces, and then you're going to have that optimization loop. So we use the JEPA library, which is a pretty commonly used library out there for prompt optimization. And once the iteration loop is complete, our partner teams are happy. They're going to then elevate that judge prompt as their LLM-as-a-judge. Now, even while doing that, LLM-as-a-judge, as a concept, the whole prompt calibration concept might be straightforward to a lot of folks, but it is still a pretty new and evolving field. And what we wanted to do was really reduce the friction of back and forth with an engineering team. So we tried to really remove all the complicated logic and make this into a self-serve UI. So the screenshot that you actually see is what actually exists. So a product manager or an operator is going to come to our UI. They're going to set some of these configs on the platform and then actually run the calibration loop themselves. So they don't have to worry about the different settings that they need to worry about, what are the different tweaks that they need to do. And they can actually run a calibration loop using any model of their choice. I think in this example, I have Gemini. They can run it using any of the Claude or the OpenAI models too. The other important piece was actually making this reviewable. Again, a lot of this is a closed box where you can't really, it's hard to see what's actually happening. So the second piece that we built was actually giving them visualization and visibility into what's actually happening. So on the left, you can see, and this is one of the good examples where we saw a significant amount of improvement in the judge prompt. And we actually show the previous, the original system prompt and the calibrated prompt to our partners so that they are also able to gain that trust as we build this. Yeah, I just wanted to add to that. This enables different configurations and different teams. In some teams, we have seen the strategy and operations folks own the prompt. We have seen some teams where the product manager owns the prompt. We have seen some teams where engineering owns the prompt. So this gives the flexibility for teams to design and evolve it because we are all learning. So even the org design is improving, and we are enabling that. Yeah, that's a good point. I think the overall idea was to build something which is as self-serve as possible so that people aren't always necessarily blocked by our team helping them out. And then finally, the quality loop in practice. As we've been going through this exercise, we've seen a lot of improvements happening to our product as well. So, for example, we mentioned we started with the UIs. We are now API and workflow first. We're trying to reuse a lot of the existing infrastructure that already existed at DoorDash. And that's helped us get a long way. Now, we've seen really good results. I think a very good result that we do like to call out is we actually did see a lot of reduction in the spend at per-annotation cost. As you all can imagine, we do have thousands of rows that need to get annotated every week. And it can get pretty expensive at DoorDash scale. And having this self-serve annotation platform really helped us increase the velocity and reduce the cost that we were actually spending with these annotators to annotate the data for us. Obviously, this resulted in faster loops. Teams were able to iterate faster. They were able to calibrate their own judges in a completely self-serve way. And thus, it has resulted in moving with a very, very high velocity. So, finally, I just wanted to quickly touch on this slide again, the eight steps, continuous loop, which is: you have your traces, you want to look at your traces, your sessions, you want to sample it down to a size which you are comfortable with. You want to start annotating your data sets. You really want to start making the data better with the human knowledge that exists and the domain knowledge that exists, and then calibrate your workflows, calibrate your agents, calibrate your LLM judges with this golden data set, and then repeat this whole cycle over a period of time to ship reliably and ship with high quality. Yeah, we have four minutes left. Thank you once again. I think that was the last slide. Thanks for attending, and if there's any questions, we'd be happy to hang out after the talk or even happy to answer them now. Thank you. very high velocity. So, finally, I just wanted to, you know, quickly touch on this slide again, the eight steps, you know, continuous loop, which is, you know, you have your traces, you want to look at your traces, your sessions, you want to sample it down to a size which you are comfortable with. You want to start annotating your data sets. You really want to start making the data better with the human knowledge that exists and the domain knowledge that exists and then calibrate your workflows, calibrate your agents, calibrate your LLM judges with this golden data set and then repeat this whole cycle over a period of time to, you know, to ship reliably and ship with high quality. Yeah, we have four minutes left. Thank you once again. I think that was the last slide. Thanks for attending and if there's any questions, we'd be happy to hang out after the talk or even happy to answer them now. Thank you.