Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber
Description
This talk covers how Uber designed evals for its food enhancement agent, which edits food photography to better present dishes for smaller, independent Uber Eats merchants, along with the pitfalls and lessons learned along the way. The problem is uniquely hard: the system must stay faithful to the original dish, preserve each merchant's brand and packaging, and avoid homogenizing the marketplace, all without an existing playbook for multimodal evals in a narrow domain. Soumya Gupta and Jai Chopra explain how they navigated reward hacking, built a closed feedback loop combining offline and online signals, and balanced creativity against rigid safety guardrails at scale. ML and applied AI practitioners working on multimodal systems, agentic pipelines, or eval design will take away practical strategies for narrow-domain multimodal evaluations, countering reward hacking, and production feedback loops. Speakers: Soumya Gupta — ML Engineer, Uber Soumya is a Tech Lead and Applied AI Engineer who architects and scales production-grade generative AI and computer vision systems at Uber. X/Twitter: https://x.com/guptasoumya12 LinkedIn: https://www.linkedin.com/in/guptasoumya12/ Jai Chopra — Product Manager, Uber Jai is a Product Lead on Uber's Applied AI team and previously worked at Cruise and several startups. X/Twitter: https://x.com/jai_chopra LinkedIn: https://linkedin.com/in/jaichopra
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Uber’s production multimodal image-enhancement agent is made safe and scalable through stage-specific evals, multiple redundant QA gates, comprehensive trace logging, and automated closed-loop tuning against fresh human labels and production feedback.
- Why it matters: This is a concrete reference architecture for operating a creative, nondeterministic agent in a high-trust marketplace where failures can misrepresent products, damage conversion, and create brand or policy risk.
- Best use: Use it as an implementation pattern for multimodal-agent control planes: instrument first, evaluate each routing/generation stage separately, gate releases on golden-set metrics, and automate bounded configuration updates with rollback.
Executive Summary
Uber applies a multimodal agent system to improve merchant food photography for Uber Eats, where millions of new items enter a global marketplace each month. The task is not simply image beautification: the system must selectively improve weak images while preserving the actual food, merchant identity, and visual diversity of the marketplace. Uber treats authenticity as a product requirement because consumers distrust visibly AI-generated imagery and misleading edits can directly harm trust.
The production flow separates understanding/routing from image editing and publishing. An LLM extracts structured information from the image, text, and metadata; a router decides whether to skip or enhance; an editing agent produces a targeted edit; and QA can send failures back for another attempt up to a fixed iteration limit. Images that still fail are withheld rather than published. Each stage has distinct evaluation logic rather than relying on a single end-to-end quality score.
Uber’s core operating lesson is that offline prompt/model tuning is insufficient. It continuously samples production cases, obtains human labels under consistent guidelines, finds agent-human mismatches, and sends these to a diagnoser that identifies the responsible component and triggers an automated tuning pipeline. A reflect agent isolates systematic issues and noise, while a synthesize agent updates the agent configuration; the candidate must re-pass the golden benchmark before registration and deployment.
The system is deliberately conservative in the final mile. Pairwise image QA evaluates whether the edited output is genuinely better while checking faithfulness, completeness, naturalness, realism, object coherence, and physical plausibility. A second publish-ready QA gate supplies Swiss-cheese redundancy. Uber then connects model quality to marketplace outcomes such as conversion, with segmented analysis by geography, device, and dish type.
Key Takeaways
- Claim: For marketplace-facing generative image systems, selective intervention is safer than universally applying a standard enhancement prompt. | Evidence: Uber must preserve merchant authenticity and avoid visual homogenization across 10,000 global cities; its router explicitly chooses either to keep the original image or send it to enhancement based on multimodal image, text-description, and metadata inputs. | Implication: Ken should put a decision/routing layer ahead of expensive or creative agent actions, with an explicit abstain/skip path rather than treating generation as the default response. | Caveat: A routing miss is asymmetric: routing a good image to editing wastes compute and can degrade it, while passing through a problematic image can expose customers to poor or misleading content.
- Claim: Each agent stage needs its own measurable evaluation contract, not just a final user-facing quality judgment. | Evidence: Uber evaluates the enhance-versus-skip router as a classifier using a confusion matrix and precision/recall; for routing, its guardrail metric is recall because it does not want bad images to escape the system. More complex multi-model routing can use an n-by-n confusion matrix. | Implication: Define metrics around the actual decision made at each workflow node—routing accuracy, QA acceptance, and final business performance—rather than masking failure modes inside an aggregate score. | Caveat: Recall optimization needs complementary safeguards because aggressively flagging images creates false positives, unnecessary compute spend, and opportunities for an already-good image to be harmed.
- Claim: Human-labeled, representative golden data is the release authority for aligning a production agent, but it must be refreshed to withstand drift. | Evidence: Uber builds its initial benchmark across cards, geographies, dish types, and image-quality types, using objective labeling guidelines to reduce labeler subjectivity. It regularly samples production data, relabels it with the same guidelines, and compares labels with agent output to detect mismatches. | Implication: Treat a golden set as a versioned deployment gate and production-sampling program, not as a one-time evaluation artifact. | Caveat: The speakers explicitly argue that a static offline model will continue to fail in real-world use; initial benchmark success is not evidence that the system will remain aligned.
- Claim: Closed-loop tuning can be automated, provided changes are constrained by benchmark gates, observability, and rapid rollback. | Evidence: Uber’s diagnoser consumes mismatch feedback, localizes the failing agent, and triggers auto-tuning. A reflect sub-agent identifies systemic issues and removes noise; a synthesize sub-agent modifies the agent configuration. The updated configuration is registered only after it passes the human-labeled golden benchmark, and Uber retains guardrail observability and quick rollback. | Implication: Ken can automate diagnosis and configuration proposal before automating unrestricted deployment; preserve immutable evaluation gates, versioned configs, traceability, and rollback at the control-plane level. | Caveat: The talk does not disclose the proprietary definitions of its image-quality metrics or detailed safeguards against an incorrect diagnoser making repeated harmful configuration changes.
- Claim: Generation QA must evaluate semantic integrity and real product value, not merely visual change or aesthetic attractiveness. | Evidence: Uber’s pairwise input/output evaluation checks faithfulness, completeness, naturalness, and realism. Its examples include an edit that added shrimp, one that removed sauce, and a chicken-wing image whose description said eight pieces but visibly contained six—creating risk that an editor hallucinates the missing items. | Implication: For multimodal agents, include content-preservation and task-specific semantic checks alongside quality scoring; ambiguous cases should fail closed rather than be confidently guessed. | Caveat: Pixel-level difference is not a reliable proxy for improvement: an agent can make a substantial but nugatory change, or reward-hack a QA signal by falling back to a generic ceramic bowl after a more creative edit is rejected.
- Claim: Iterative generation should be bounded and feedback-driven, with coverage consciously traded against safety. | Evidence: Uber generates an image-specific edit prompt from the router’s directives, runs a multidimensional QA gate over factors such as plating, faithfulness, and color, and feeds rejection feedback into the next generation attempt. In a sweet-potato-fries example, the first edit failed for incorrect portion size and unrealistic plating, while the second passed. The central metric is pass@K. | Implication: Set a hard retry budget for self-correcting agents and make failure-to-serve/abstention an intentional product outcome rather than allowing unlimited loops or weak outputs through. | Caveat: If an image does not pass within K iterations, Uber accepts a coverage loss and does not enhance or publish the generated result.
- Claim: Redundant gates and outcome feedback are necessary because upstream evals will miss failures and offline quality is not the ultimate business metric. | Evidence: Uber adds a final publish-ready QA step for policy and broader quality checks even after generation QA, describing the design as a Swiss-cheese model. It logs the full flat, end-to-end agent trace and incorporates dog-food feedback, merchant feedback, design/product-team feedback, and production conversion measures such as add-to-cart and completed orders. | Implication: Use layered defenses for high-impact outputs and join agent traces to downstream behavioral/business data so that optimization targets actual operational value, not evaluator scores alone. | Caveat: Conversion must be examined by segments—such as geography, device type, and dish type—because an aggregate gain can conceal localized regressions or uneven effects.
Detailed Brief
Observability and diagnoser architecture
- Claims: Logging is presented as a prerequisite for optimization and self-learning rather than an operational afterthought.; A higher-level diagnoser generalizes feedback handling across the full orchestration, rather than requiring a separate feedback implementation for every agent.
- Evidence: Uber keeps all agents in the end-to-end orchestration in a flat JSON trace structure, enabling both technical and nontechnical engineering, product, and operations stakeholders to inspect individual cases and aggregate patterns.; The speakers say they use Arize for observability.; The diagnoser can consume model-evaluation mismatches, internal dog-food signals such as thumbs up/down and free-form comments, merchant feedback, and feedback from design or other product teams.; A diagnoser can identify one or multiple agents for configuration-specific optimization, replay the flagged examples, then compare metrics before a new configuration version is pushed.
- Caveats: The talk provides the workflow pattern but not implementation details for feedback normalization, root-cause confidence thresholds, or how conflicting stakeholder feedback is weighted.
- Implications: A unified trace schema and feedback-ingestion layer make an agent system governable across technical quality, policy, user experience, and commercial outcomes.; Free-form user feedback becomes more useful when it is converted into replayable cases tied to the exact agent/configuration path that produced the output.
Failure taxonomy for creative multimodal systems
- Claims: Frontier-model limitations can surface directly as application-level defects, so application teams need eval categories that expose and escalate underlying model issues.; Multimodal uncertainty should be treated explicitly rather than resolved through fabricated certainty.
- Evidence: Uber calls out object coherence and physics plausibility after an output plate covers sauce incorrectly, noting that such failures can be coordinated back to frontier-model teams.; For wontons where neither input nor output makes the item count visible, the evaluator produces an unsure result; Uber rejects the result in production rather than passing it.; An enhancement can create generic-looking output that differs visibly from the original but does not create meaningful product value.
- Caveats: Some semantic facts needed for validation may be unavailable from the pixels alone, even with associated menu text; uncertainty is therefore an unavoidable state, not necessarily a model defect.
- Implications: Maintain a failure taxonomy that distinguishes prompt/configuration failures from model-capability defects and insufficient-evidence cases.; A three-state evaluator output—pass, fail, unsure—can be safer than a forced binary judgment when agents operate over incomplete multimodal evidence.
Notable Concepts & Terms
- Multimodal image understanding and routing agent: The front-end decision agent combines image content, textual description, and metadata to produce structured output and decide whether an asset should be enhanced or preserved.
- Golden data set: A human-labeled, representative benchmark used to tune agents and gate releases against predefined guardrail metrics.
- Recall guardrail: Uber’s key router safeguard: bad images should not be incorrectly passed through, even though false positives create cost and degradation risk.
- Diagnoser: A supervisory agent layer that accepts feedback from multiple loops, localizes the failing component, and routes it to the correct tuning workflow.
- Reflect and synthesize agents: The two-part prompt optimization loop: reflect analyzes mismatches and systemic patterns; synthesize modifies an agent configuration using that analysis.
- Pass@K: The acceptance rate by the K-th generation attempt, used to measure whether feedback-driven iterative editing improves the chance of passing QA within a bounded retry budget.
- Pairwise comparison: An input-versus-output evaluation method used to judge whether an edited image is truly better while preserving required content.
- Swiss cheese model: Redundant, partially overlapping QA layers that reduce the probability of a failure reaching production when no individual gate is perfect.
Operator Notes / Why Ken Should Care
- Require a canonical per-run trace schema before scaling any multi-agent workflow; include inputs, structured observations, routing decision, config/model versions, evaluator results, retries, and final disposition.
- Create a versioned golden-set release gate from representative production slices, then schedule periodic fresh-label sampling specifically to detect behavioral drift.
- For any creative or tool-using agent, implement pass/fail/unsure evaluation semantics, a fixed retry budget, and an explicit abstention path.
- Separate automated config proposal from release authority: proposals may be generated automatically, but promotion should require regression checks, protected guardrail metrics, configuration registration, and rollback readiness.
- Build a failure taxonomy that flags semantic fabrication, omission, reward hacking, visual/physical incoherence, and insufficient evidence as distinct remediation routes.
- Instrument downstream impact by cohort rather than relying on global averages; monitor whether improvements differ by geography, device, asset/category type, and source quality.
Source/Metadata
- Title: Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber
- Transcript words: 3947
- Duration seconds: 1298
- Timestamp note: No timestamps or chapters were present in the supplied transcript.
Transcript
. My name is Jay, and I'm here with Somya. We are part of the computer vision team at Uber. And we're going to talk to you about a real-world production use case. Oh, mic's not on? Okay. Okay, don't worry, I'll manage. Can you hear me now? Okay, so we're going to talk to you today about a real-world production use case. And specifically, we're going to dive into how we design the evals and the eval loops. So, let's see this one. Okay, it's not working. All right, cool. So, just before we get into the agent design, we're going to talk a little bit about the use case. So, our delivery marketplace, Uber Eats, we do about 90 billion run rate per year at the moment. We're adding millions of items to the marketplace each and every month. We're growing at 20% year on year. And we operate in 10,000 cities globally. So, not many people actually know this, but our delivery marketplace is just as big as the mobility side on Uber today. Visual content actually plays a really important role for the user experience. So, a photo is quite often the first signal that a customer gets that gives them that initial impression about a merchant. So, a good photo can make the difference between someone scrolling through the feed and actually clicking on an item and adding to the cart. And more and more, we're seeing different modalities on Uber Eats, especially video content. But this is a problem. So, our smaller independent merchants simply don't have the level of quality for their photos that reflect what the eater is actually going to get. And when we speak to our merchants, there are three themes that emerge: lack of time, lack of know-how, and costs, because these professional photo shoots actually cost a lot of money. And this can be especially problematic if the merchant is updating their menu over time. So, this problem is actually pretty challenging to solve for at scale, right? Because our consumers, they want authentic, real-looking photos. But a meaningful fraction of consumers actually distrust anything that is AI-generated. So, if you open up the Uber Eats app, the last thing that you want is to be scrolling through food photography that looks like AI slop. So, threading the needle here, we need to be able to stay faithful to the original image, preserve the brand of the merchant, and avoid everything looking the same. If we have the same prompt for every photo that we're editing, the diversity of the marketplace is going to collapse. Because we operate globally, we also have this long-tail distribution of different quality that we see across the marketplace. So, we've got some examples here. You might see food photography that is poor sharpness, poor composition, not centered, or poor colors as well. We also have a wide spectrum of user-generated content on the platform as well. So, what are our goals when we're designing these agents? When you think through these goals, you might actually be thinking through your own agents that you're building yourself. But for us, it's about, one, preserving authenticity and trust. Two, improving the quality when we need to. So, we want to be able to improve quality selectively. We want to optimize globally for the entire marketplace. We don't want to cannibalize certain merchants. We want to ship safely. And this is going to be an important theme throughout the talk. We want to learn continuously. And we want to operate at scale in a cost-efficient manner. So, agents are actually really well suited to solve this problem. So, if you imagine a spectrum, on the one side, you've got something that's more deterministic. It's more rules-based, and you have more control over it. But it's a brittle system. It's not actually going to be able to scale for the entire marketplace. Imagine the other side: you provide an agent with a lot of creativity, and it has a lot of agency. And that's actually what we want to lean into. But we can't leave that unconstrained, right? Because we have certain safety and certain guardrails in place that we need to adhere to. So, we want to find a balancing act. And that's set the principle for the way that we think and design around agents and evals. So, now we're going to dive a little bit deeper into a simplified but representative example of what we have in production. And we're going to go through each stage and how we eval it. And then talk through some continuous learning loops as well. So, first up, we have what we call an image understanding and routing agent. So, this is where multimodality is pretty important. We actually ask the LLM to describe what it sees in the photo. And then we create a structured output from that, and we send it to a router. The router will then determine: do we enhance it or do we skip it? If we skip it, we will keep the original. If we enhance it, we send it to our next agent, which is an image editing agent. And this can actually run in a loop. So, it gets feedback from a QA agent. It can edit in this loop and self-correct and fix things as it goes. If it goes through a number of loops and it still fails, we don't publish it. Then we actually send it to a final post-processing and QA step. If that's all good, we'll publish it to the menu. And the last thing that's really critical is we log everything. Just a quick note about logging. I don't know if you can actually read the JSON here, but you might notice that all of the agents in this end-to-end orchestration are within one. It's a flat structure in this JSON. And so, this is actually incredibly useful for the entire team because anyone, be it non-technical or technical folks in engineering product, can actually dive in and look at specific cases to diagnose and also roll up things to look at in aggregate. And it's important to note here that we think this is important to start with. You want to start with your logging because if you don't start with it, you have nothing to optimize for, let alone set up the self-learning loop. And at Uber, we use Arize. Cool. We're going to dive a bit deeper into the router. The router is actually pretty straightforward. If you remember, we have this multimodality input. We look at certain text description, metadata, the image itself. We ask it to describe what it's seeing. We create structured output from that. With that structured output, we can then grade against a rubric. So, we have these pass and fail criteria. The last step is we want to decide whether or not we should enhance or skip. How do we actually eval this? You could think of this as a more traditional classifier. So, here we have a confusion matrix. Many of you are probably pretty familiar with this. But we can look at things like the true positive cases, the false negative cases, and so on and so forth. Essentially, what we're doing is we're measuring the precision recall. In practice, your routers might actually be much more sophisticated. So, for example, we might want to route an image to a lower-latency, smaller model to be able to save on cost and improve the user experience at the trade-off of quality. And if that's the case, instead of having a two-by-two matrix for your confusion matrix, you might actually have an n-by-n matrix. Where each grid is actually telling you whether or not you're correctly routing to that specific branch. So, I'm going to now hand over to Samya, who's going to dive a little bit deeper into how we handle drift and human alignment. So, now that we spoke about how we eval the routing, I want to talk about how you get the first version of the model out. For our use case, we consider human labels as the golden source of truth. And this is what we want to align our models to. The way we do this is we go collect a data set, which is representative. So, different cards, geographies, dish types, image quality types, send it to our human labelers, and give them a very objective guideline to label on. This is to remove any subjective biases or any noise coming in from human labelers. Once we've got that system set up is when we start tuning our model. We take our agent, we go ahead and get output from the agent, compare it to your golden data set, evaluate if it's good enough to ship. If it meets your guardrail metrics, you go ahead and ship it. If not, then you go tune and keep doing this until you meet your guardrail metrics. For routing, our guardrail metric is recall. We don't want any bad image to slip through our system. Here are some examples of the failures we've seen. On your left, you see a very good image of a cheeseburger. On the right, you notice that the routing agent actually failed this. It said the technical is low bar. And it will go send this image for enhancement. Now, there are two challenges when you send this image for enhancement. Firstly, you pay the compute cost for a zero quality lift from this image. Once we've got that system set up, is when we start tuning our model. We take our agent, we go ahead and get output from the agent, compare it to your golden data set, evaluate if it's good enough to ship. If it meets your guardrail metrics, you go ahead and ship it. If not, then you go tune and keep doing this until you meet your guardrail metrics. For routing, our guardrail metric is recall. We don't want any bad image to slip through our system. Here are some examples of the failures we've seen. On your left, you see a very good image of a cheeseburger. On the right, you notice that the routing agent actually failed this. It said the technical is low bar. And it will send this image for enhancement. Now, there are two challenges when you send this image for enhancement. Firstly, you pay the compute cost for a zero quality lift from this image. Secondly, there is a risk of degrading this image, given it's already such a high-quality image. On the other end of the spectrum, you have a recall miss. On your left, you have an image with six chicken wings. On your right, if you notice the dish name, it says eight pieces chicken wings. Your routing agent approved this image. That means now there's a risk here. If you send this image for enhancement and you only see six chicken wings, there's a chance your model is going to hallucinate these two extra wings just to match the description. And that's also the cut we take at our faithfulness metric that Jay earlier showed us. So the meta point I'm going to get here is you've trained your offline model, but there will be cases where your model is going to continue to fail. And the static model will not work in the real system. You need a way such that your prompt, agent, system itself is evolving over time. And that's what we've done for our system as well. And I'm talking more from the routing perspective, but every component in our system is able to tune itself for any drift online. So what we do is we sample production data at regular cadence. I send this to the human labelers with the same guidelines that we have seen before. Once you've got that data, we compare our agent's output with the output we got from the labelers and see if there's a mismatch. If there's a mismatch, we have an umbrella diagnosis agent, which takes in the feedback, localizes where this issue is happening, and triggers an auto-tuning pipeline. Once we tune this agent, we go and benchmark it against our golden data set that we saw earlier. And if we pass our golden data set on the metrics that we had designed, we go ahead and ship this model. If not, then you keep iterating. And this happens on a regular basis on the production data set. The beauty of this is this is completely config-driven and doesn't require a human in the loop. Your diagnoser agent can write your config and trigger the auto-tuning pipeline here. And this is what will keep your model sharp over time. You will have one static model with the offline, but this is what is going to keep your system alive. So Jay is going to spend more time on the diagnoser side of it. What I want to do is zoom into the auto-tuning bit. And again, we are looking at routing, but this is how we tune every agent in our system. So we start with a target agent, and we've already got these unseen eval samples from our humans. We go find out the mismatches and matches and call a prompt optimizer agent. Now, this itself is two sub-agents. There's the reflect agent and the synthesize agent. What reflect does is it looks at the mismatches, tries to remove any noise, find any systemic issues that might be in your data set, and reflect on it and send that feedback to the synthesize agent. Now, the synthesize agent takes this feedback. It has your agent config. It goes and updates your agent with the new config based on the feedback it's getting and goes and benchmarks again. If this benchmark is passed, you actually register this new agent in the new agent config store. And next time your production runs, you pick up the new version of the agent. And this is a closed-loop system, as I mentioned, with no human in the loop. We definitely have observability on the guardrails, quick rollback built in in case of any issues with the system itself. Moving on to the next step of our orchestration flow. So we spoke about routing, moving on to the enhancement bit of it. It's a C-step process. What we do is, in the first step, we generate a prompt specific to this image. We take in this description, we take in the directives we were getting from a routing agent, and we go ahead and generate a prompt for this image, what needs improvement in this image specifically. And we go ahead and enhance this image. Then you've got the QA gate, which is a multidimensional gate, looks at multiple things like plating, faithfulness, colors. And if it passes is when you actually go ahead and publish this. If it doesn't pass, you take the feedback back from the QA gate, push it back to your generate prompt, along with the initial inputs you sent it, and go ahead and enhance it again. So there are two end results here. You either keep enhancing for K iterations and you pass your QA gate and you publish. Or you take a coverage hit and you never enhance this image. Here's an example. On your left, you see a bowl of sweet potato fries. We send it for the first iteration, and our QA agent rejects it because the portion size is incorrect. The plating is very unrealistic. We take that feedback in, go for the second iteration, and we're actually able to pass it the second iteration. So the metric we are measuring here is pass at K. Pass at K is essentially the pass rate at K-th iteration. And ideally, with more iterations, your pass rate will increase because you're getting more feedback in. Now I'll pass it on back to Jay to cover the rest of this. Thanks. Thanks, Samia. So, just before we end here on the generation evals, we use what's called pairwise comparison, right, for our pass at K. So it's looking at the input image and the output image, and it's assessing whether or not it's better. But how do we actually find what's better? So we're not going to dive into too much of the details here because this is proprietary stuff, and so we'll just mention at a higher level that this is where, at least for us at Uber, we have to make sure we're aligning with product design, policy, legal. And this is where we're baking in what we define as a better image on the platform into our evals. So examples here: is it faithful? Is it complete? Is it natural? Is it realistic? And there's a bunch of other things as well. The output of this is then a yes, no, or unsure. So here are some examples of failure modes. So input and output on the right. The input's on the left, output's on the right-hand side. This might be a little bit difficult to see at first pass, but we actually added shrimp here, and we shouldn't be. So we fail faithfulness. This is where we go the other way. So the input has some sauce at the bottom of the sushi. We actually remove it. So we fail completeness. Here's actually a pretty interesting example where the agent actually attempted a more creative edit, the first iteration. And then the QA said, nope, that's not good enough. And then it actually oversteers the other way. And it becomes overly conservative. Falls back to this generic ceramic bowl. So this is an example of reward hacking, actually. And this is a nugatory change, but it's something that we don't think is a meaningful, influential change, despite the actual raw pixels of the input and output being pretty different. Here's another example where, in the output, the plate is covering the sauce. This is an example where some of the frontier models that we're using for the actual image editing, some of their problems will actually leak up into our applied use case. So object coherence and physics plausibility are the evals that sometimes will coordinate with the frontier teams and let them know about these problems and work together with them. Here's an example of why multimodality is pretty important. In the input and the output, we can't actually see that there are eight pieces here of the wontons. So we're not confident, actually. We're not sure. And so this is an example where we would actually reject it in production and it wouldn't go through. So the last step after all of that is a post-processing and what we refer to as the publish-ready QA. This is the final gate before we decide we want to publish something to production. Here we do some policy checks. We also do some more quality checks. And you might be wondering, we've already done some QA. Why are we going to do another step of QA? The reason is because we think of this like a Swiss cheese model. So we want to try and optimize for reducing the chance of a failure getting into production. And so there is some redundancy here or there. And that's okay. Here's an example of why multimodality is pretty important. In the input and the output, we can't actually see that there are eight pieces here of the wontons. So we're not confident, actually. We're not sure. And so this is an example where we would actually reject it in production, and it wouldn't go through. So the last step after all of that is post-processing and what we refer to as the publish-ready QA. This is the final gate before we decide we want to publish something to production. Here we do some policy checks. We also do some more quality checks. And you might be wondering, we've already done some QA. Why are we going to do another step of QA? The reason is because we think of this as a Swiss cheese model. So we want to try and optimize for reducing the chance of a failure getting into production. And so there is some redundancy here or there. And that's okay. And so this QA gate is a little bit more holistic. It captures more things. But it also will try and flag things that we should have caught upstream as well. All right. So we've talked about a couple of feedback loops here. So to summarize, we talked about predominantly this first one here, which is the model loop. And this is accounting for drifts and aligning with the human-labeled data set that we have and we've established offline. But we actually have more feedback loops. So at Uber, what we have is a great dog-fooding culture where we'll test apps before they go live. But we also have, when it goes live in production, how do we get that feedback back into our agent to be able to steer it appropriately? So as we're adding more of these feedback loops, we want to be able to generalize the system. So this is where we've actually created a higher level of abstraction on top, which we call the diagnoser. So the diagnoser can take in any input from these different feedback loops that we're capturing. It can reflect on what actual agent within the overall system needs to be optimized. And it can route that agent to be able to fix that configuration specifically. It could be one agent. It could be multiple agents. So here's an example of internal dog-fooding. You might see these in different apps that you've got where you've got the thumbs down and the thumbs up. We also take some free-form feedback as well. And this is actually great because we'll get feedback from merchants directly. We'll get feedback from design teams, other product teams at Uber. And we'll incorporate that feedback back into our diagnosis step and tune the system over time. Again, a similar workflow pattern here. We'll replay the examples that we know are those ones that have been flagged, be it good examples, be it bad examples. And then we'll benchmark the metrics before we push the latest config version. The last step is actually getting this into production. And this is where we're looking for a whole heap of different metrics that we track for the marketplace quality and health. I've just called out one here, which is conversion. So we're looking for improvements in people adding to cart, converting, completing their orders. I think this one's actually an interesting one to call out because now, at least at Uber, but especially in production settings at scale, you have a lot of data that you can actually slice and dice. So in this area, as opposed to the others, what we can do is slice by geos, by device type, by dish type, et cetera. And we can look at where things are improving in different segments and actually tune on certain segments as well. Cool. And that's it for our presentation. Appreciate it. Thank you. Thank you. Thank you. Appreciate it. Thank you. Thank you. Thank you.