Good afternoon, everyone.
Thanks for coming for a post-lunch talk. Always appreciate that. My name is Svaroop, and here's my teammate Nachiket. We are here on behalf of the DoorDash JNI platform team. We wanted to share our evals journey. It started as evals is another engineering thing, but then it slowly evolved into a cross-functional effort, and we want to share our story here. So what is this team? This team is a Gen AI platform team. We are a horizontal team that helps all other product teams. So product teams at DoorDash build on top of the infrastructure and the primitives that we provide. And we see our USP and the value that we provide
is that we help product teams balance these three forces, which is accuracy, latency, and cost. Initially, we applied this in terms of models, but if you think about it, it also applies to agents. And the way we achieve this is we have primitives and building blocks. So, for example, we have an LLM gateway where you can easily switch between different models and try the latest and greatest. We have an agent gateway where you can connect to tools and other agents. And we help solve authentication, agent identity, and other things in a central place, which our security team can bless. Similarly, we pair the LLM gateway with OpenWeights models hosting.
Of course, cost is a number one concern these days. And we invested in OpenWeights models and have seen significant impact already. And maybe we'll talk about that in a future conference. The fourth pillar is evals, and that's the part that we would want to share today. When we started talking to product teams internally at DoorDash, there were varying distinct needs across teams. We had a consumer discovery and shopping assistant team. For those who attended Rago Stock earlier today, you will see the need for session-level quality judgments. Then the personalization ML team needed a way to scale up human judgment. And with multi-agent systems,
we needed trajectory-based evals. Now, the question is, how do you cater to all these different needs under a common platform? And as we spoke to these teams, we realized we needed to empower the people who are the domain experts. And in our case, that was strategy and operations folks. It was product managers. It was even labeling partners, and not only engineers. So we started with, okay, we have to be UI first, and this was the guidance we had from Andy Fang, our co-founder, as well. So we had UIs for non-engineers to contribute. Then we evolved to also being API first so that engineers can also build and not be blocked on the central platform,
and they can build their own systems. And then, of course, with the coding agents, now we have become workflow first, that we empower SNO and PMs to also be able to navigate the platform and run operations as well. So with that context, I'll hand it off to Nachiget to talk about how we went about delivering this. Cool. Thanks, Faroop. And thanks, everyone, for joining us. I know France is playing right now, and I promise you this will be better than that. I'm kidding. So as Faroop was saying, evals is not just an engineering harness. It is a cross-functional effort across different pillars, across different teams, that actually helps us add
all the domain-specific knowledge into the quality of the AI itself. So from your traces to your data sets, from scoring mechanisms, this is all a team sport. We all have to play and help improve the quality of AI. So going a little bit deeper into the same aspect, we have different teams at DoorDash who help us actually improve the quality of AI. So you're going to have your strategy and operations folks who are going to set priorities, set the quality bar that you want to aim for. You're going to have your product people who are going to translate these requirements into rubrics, workflows. You're going to have your operations teams running annotations.
You're going to have your engineering teams, like us, providing APIs, telemetry, data sets, judges, all the cool things. And combining all these together is what the recipe is for actually making sure that you are shipping quality AI products through an evals platform. So we've tried to boil this down into a continuous iteration loop. So right from tracing, having a tracing solution, viewing your sessions, your traces, to sampling them down to a very small set that you actually want to look at, annotating these with the domain-specific expertise that you bring in with the different teams I mentioned, reviewing those, then creating those golden data sets,
which are going to be your golden data sets that you want to measure or calibrate against. And then, of course, monitoring this over a period of time, and then rinse and repeat, go through the whole loop again. So this, in our experience, has been a good continuous loop for shipping quality AI. On the platform level, we have two surfaces. So we have the telemetry layer, where we have all our traces, our scores, observations. That is also the plane where users are able to access these traces using an MCP, using an SDK, using our APIs. And then we have the workflow layer. This is where a lot of our stat ops, our product teams, operate on the platform. So this is where
all the annotation tasks are set, this is where they review their golden data sets, create their judges, calibrate their judges, and so on. So maybe today we'll go through these four different modules or pillars of our platform step by step. So again, first one, tracing and sampling, which is actually capturing what your agents, what your LLMs are actually outputting, for lack of better words, and actually viewing those. Now, in order to also power this whole platform, we have, I think as Svarup mentioned, gone in an API-first approach. What that has allowed us to do is have these stable APIs that actually, and then you build UIs on top of that. So all our scores,
our data sets, these are all powered by very stable APIs that our team owns. So all your API access, including SDK access, is powered by this single plane. Again, going back, and just refreshing your memory, step one, capture your traces, capture your sessions, measure your scores. Then you want to start adding all your judgment, your context, your domain knowledge, and then calibrating your judges is what we have seen as the whole life cycle. Step two is on the annotation side. So you obviously are capturing a lot of your agentic behavior, your sessions, your traces, but you actually want to see what are some places where things went well and what are some places
where things did not go well. This is where you can actually titrate and actually look inside what's actually happening at the session level and annotate these data sets. And as Farooq mentioned, we have a lot of use cases. We talk to multiple different teams who have various ways of annotating their data sets. And it's almost hard for a platform team to build a UI specific for each use case. And, to give you an example, it's usually going to be an annotator who's going to annotate these data sets. So the platform team is in charge of the APIs. We have a strategy and ops person who's actually deciding what to annotate. And then you have an annotator
who's actually going to annotate your data set. So we took this approach. Everybody has access to coding agents. And we actually doubled down on that API-first approach. So because we had these APIs, we were actually able to enable our stat ops teams to use something like a codex or a claw code and wipe code their own annotation UIs. So we had different use cases. I think we had a talk from Raghav before. We had image annotation use cases. We had some manual testing use cases. What stood out to us was the underlying patterns were similar. So if we are API-first, we can actually enable our partners to simply wipe code these UIs for annotation. So it's like
a very simple example of a wipe-coded UI. Looks pretty clean, does the job, and you get the annotation that you need. This is basically a menu from a restaurant. It's nothing crazy. But the point I want to make here is that what helped us was to give this workflow in the hands of the operators so that they can actually build their own wipe-coded annotation UIs. So moving on, once you have these annotation UIs, you obviously want to calibrate your judge prompts. You obviously have some LM-as-a-judge metric that you're tracking. You want to now start improving that with these golden data sets. In order to do that, we have a pretty simple process. You're going to start
with some judge prompt, take a look at what exactly do you want to measure from the output, have something simple. You're going to have your baseline scores, where you're going to simply run those LLM judges on your traces, and then you're going to have that optimization loop. So we use the JEPA library, which is a pretty commonly used library out there for prompt optimization. And once the iteration loop is complete, our partner teams are happy. They're going to then elevate that judge prompt as their LLM-as-a-judge. Now, even while doing that, LLM-as-a-judge, as a concept, the whole prompt calibration concept might be straightforward to a lot of folks, but it is still
a pretty new and evolving field. And what we wanted to do was really reduce the friction of back and forth with an engineering team. So we tried to really remove all the complicated logic and make this into a self-serve UI. So the screenshot that you actually see is what actually exists. So a product manager or an operator is going to come to our UI. They're going to set some of these configs on the platform and then actually run the calibration loop themselves. So they don't have to worry about the different settings that they need to worry about, what are the different tweaks that they need to do. And they can actually run a calibration loop using any model
of their choice. I think in this example, I have Gemini. They can run it using any of the Claude or the OpenAI models too. The other important piece was actually making this reviewable. Again, a lot of this is a closed box where you can't really, it's hard to see what's actually happening. So the second piece that we built was actually giving them visualization and visibility into what's actually happening. So on the left, you can see, and this is one of the good examples where we saw a significant amount of improvement in the judge prompt. And we actually show the previous, the original system prompt and the calibrated prompt to our partners so that they are also
able to gain that trust as we build this. Yeah, I just wanted to add to that. This enables different configurations and different teams. In some teams, we have seen the strategy and operations folks own the prompt. We have seen some teams where the product manager owns the prompt. We have seen some teams where engineering owns the prompt. So this gives the flexibility for teams to design and evolve it because we are all learning. So even the org design is improving, and we are enabling that. Yeah, that's a good point. I think the overall idea was to build something which is as self-serve as possible so that people aren't always necessarily blocked by our team
helping them out. And then finally, the quality loop in practice. As we've been going through this exercise, we've seen a lot of improvements happening to our product as well. So, for example, we mentioned we started with the UIs. We are now API and workflow first. We're trying to reuse a lot of the existing infrastructure that already existed at DoorDash. And that's helped us get a long way. Now, we've seen really good results. I think a very good result that we do like to call out is we actually did see a lot of reduction in the spend at per-annotation cost. As you all can imagine, we do have thousands of rows that need to get annotated every week. And it can get pretty
expensive at DoorDash scale. And having this self-serve annotation platform really helped us increase the velocity and reduce the cost that we were actually spending with these annotators to annotate the data for us. Obviously, this resulted in faster loops. Teams were able to iterate faster. They were able to calibrate their own judges in a completely self-serve way. And thus, it has resulted in moving with a very, very high velocity. So, finally, I just wanted to quickly touch on this slide again, the eight steps, continuous loop, which is: you have your traces, you want to look at your traces, your sessions, you want to sample it down to a size which you are
comfortable with. You want to start annotating your data sets. You really want to start making the data better with the human knowledge that exists and the domain knowledge that exists, and then calibrate your workflows, calibrate your agents, calibrate your LLM judges with this golden data set, and then repeat this whole cycle over a period of time to ship reliably and ship with high quality. Yeah, we have four minutes left. Thank you once again. I think that was the last slide. Thanks for attending, and if there's any questions, we'd be happy to hang out after the talk or even happy to answer them now. Thank you. very high velocity. So, finally, I just wanted to,
you know, quickly touch on this slide again, the eight steps, you know, continuous loop, which is, you know, you have your traces, you want to look at your traces, your sessions, you want to sample it down to a size which you are comfortable with. You want to start annotating your data sets. You really want to start making the data better with the human knowledge that exists and the domain knowledge that exists and then calibrate your workflows, calibrate your agents, calibrate your LLM judges with this golden data set and then repeat this whole cycle over a period of time to, you know, to ship reliably and ship with high quality. Yeah, we have four minutes left.
Thank you once again. I think that was the last slide. Thanks for attending and if there's any questions, we'd be happy to hang out after the talk or even happy to answer them now. Thank you.