AI Engineer

Distill the LLM, Don't Serve It: Search & Personalization at DoorDash — Raghav Saboo, DoorDash

1737 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: For marketplace discovery, DoorDash treats LLMs primarily as offline reasoning systems that create reusable semantic primitives—graded labels, semantic catalog IDs, user memory, and controlled content—rather than as online serving models.
  • Why it matters: This is a concrete production pattern for converting expensive LLM capability into low-latency, measurable improvements across retrieval, ranking, personalization, and agent experiences.
  • Best use: Use it as an architecture case study for designing an LLM distillation pipeline: establish high-quality offline supervision and shared representations, then deploy conventional models at request time.

Executive Summary

Raghav Saboo argues that marketplace search and recommendation failures are fundamentally failures of semantic understanding, not merely engagement optimization. A popular product may receive clicks yet violate a user's core constraint—for example, regular spaghetti ranking for “gluten-free pasta.” DoorDash uses LLM reasoning to create a more explicit, graded understanding of relevance, then injects that understanding into its existing retrieval and ranking stack.

The operating principle is “distill the LLM, don't serve it.” DoorDash spends expensive reasoning offline to build labels and representations, fine-tunes a lighter model to scale labeling across its catalog, and serves fast production retrieval/ranking models rather than placing an LLM in the online critical path. Its two-stage, hard-negative retrieval training improved relevance NDCG by 2.3%, while semantic IDs improved ranker MRR by 4–5%.

The shared primitives extend beyond search. Hierarchical semantic IDs give billions of store-level items a learnable, catalog-grounded meaning that supports cross-category missions, cold start, and tail-item coverage. Consumer memory combines long-term behavior, live session context, and explicitly stated preferences into text, vectors, and graphs usable by both LLMs and traditional models.

DoorDash applies these inputs to steerable, offline-generated personalized collections. The LLM can synthesize a collection's concept, title, subtitle, and candidate items, while the established serving stack hydrates and ranks actual inventory. In the pets vertical, these tailored collections reportedly increased order rate by nearly 1% and active users by 0.6%.

Key Takeaways

  • Claim: Engagement signals alone are insufficient for discovery because they can reward popular but constraint-violating results; systems need an explicit graded semantic relevance objective. | Evidence: For the query “gluten-free pasta,” popular regular spaghetti can rank from sales history and gluten-free white bread can retrieve from keyword overlap, whereas DoorDash wants a 0–1–2 relevance scale that distinguishes exact matches, reasonable substitutes such as chickpea pasta, and irrelevant results. | Implication: Treat semantic relevance as a first-class target alongside clicks, add-to-cart, and conversion rather than assuming engagement proxies capture user intent. | Caveat: Human labels are high quality but expensive and quickly stale in a rapidly changing catalog; behavioral signals are abundant but biased by exposure, rank position, pricing, promotions, and prior model decisions.
  • Claim: LLM reasoning is most valuable as scalable offline supervision that can be distilled into cheap, fast models shared by retrieval and ranking. | Evidence: DoorDash builds a human-labeled seed set of query-item pairs, audits label/behavior conflicts with a stronger LLM and taxonomy models, then fine-tunes a lightweight model such as GPT-4o mini to label the full catalog. Those resulting graded labels become a shared target for retrieval and ranking. | Implication: For latency- or scale-sensitive products, use frontier-model reasoning to manufacture training data and evaluation signals, not necessarily to answer every live request. | Caveat: The reliability of the distilled labeler depends on maintaining a high-precision golden dataset and actively auditing suspicious disagreements rather than blindly treating LLM outputs as ground truth.
  • Claim: A two-stage contrastive retrieval workflow is needed to separate exact matches from merely related products at e-commerce scale. | Evidence: Stage one trains two-tower encoders with multi-level supervised contrastive loss to shape global embedding geometry. Stage two mines items the model wrongly ranks high and strong positives ranked low, relabels them, and uses curriculum training; DoorDash reports a 2.3% NDCG improvement in relevance. | Implication: Build an error-mining loop around retrieval embeddings; the most valuable training examples are often semantic confusions generated by the current model. | Caveat: Standard embedding similarity tends to collapse distinctions among exact matches, substitutes, and complements, so a generic embedding model without task-specific hard-negative work is not enough.
  • Claim: Semantic IDs provide a learned, hierarchical catalog language that is more flexible than SKUs or human taxonomies and can become a shared interface across models. | Evidence: DoorDash assigns each item a short hierarchical code whose prefixes represent broad neighborhoods and later tokens finer distinctions; hot sauces can share early prefixes while splitting into Mexican, Caribbean, and Korean specialties. Using semantic IDs improved ranker MRR by 4–5%, with associated conversion gains. | Implication: A stable, learned semantic identifier layer can unify language systems, sparse-feature models, retrieval, ranking, and inventory-aware query reformulation. | Caveat: A human taxonomy remains useful for meaningful classification, but it is often too coarse and rigid to capture fine-grained product relationships or shopping missions.
  • Claim: Consumer personalization should be represented as inspectable memory across multiple timescales and materialized in forms that both LLMs and conventional ML can consume. | Evidence: DoorDash combines durable signals from orders, searches, browsing, and support interactions; in-session signals such as cart state and active searches; and user-stated preferences from Ask DoorDash. It represents this as human-readable text, latent vectors, and graph/tree structures within composite memory objects. | Implication: Design user memory as a reusable platform primitive—structured enough to inspect and extend, while also producing vector and graph features for downstream models. | Caveat: Conventional user embeddings are useful but do not readily express why a user has a preference, and LLMs cannot directly make effective use of opaque embeddings without an explicit LLM-native representation.
  • Claim: Steerable LLM content generation can personalize merchandising without putting generation in the live serving path. | Evidence: For store-page collections, DoorDash runs an offline LLM process over user memory and semantic IDs to create titles, subtitles, and item collections; the existing online stack performs inventory hydration and collection ranking. Examples include plant-based pantry rows, cat-food rows for cat-only households, and pantry staples during restocking. Pets tests produced nearly a 1% order-rate lift and 0.6% active-user lift. | Implication: Separate creative/semantic generation from real-time eligibility, retrieval, and ranking so personalization remains controllable, inventory-aware, and operationally reliable. | Caveat: Generated concepts only have product value if they can be grounded in currently available catalog inventory; otherwise even semantically plausible suggestions are useless.

Detailed Brief

How the four primitives connect into a production discovery system

  • Claims: The four primitives are intended as shared representations rather than isolated features: supervision defines what good looks like, semantic IDs define what items mean, memory defines shopper context, and steerable generation turns those inputs into useful interface outputs.; The same semantic foundation supports multiple output shapes, including ranked semantic-ID lists, carousel titles, subtitles, search reformulations, retrieval features, and agent-session personalization.; DoorDash frames these primitives as a way to model a shopping mission that can unfold across query reformulations, collections, complementary items, checkout, multiple sessions, and multiple days.
  • Evidence: The talk's motivating mission is a newly adopted puppy: a user may begin with puppy food, then need crates, leashes, training pads, and other items across an extended journey.; Semantic neighborhoods can associate products in different taxonomy branches—such as chips, salsa, and guacamole—as parts of a joint shopping mission.; Query reformulation is catalog grounded: a query such as “sriracha” can surface related available queries such as chili garlic sauce or sambal oelek instead of generic semantic associations that lack DoorDash inventory.
  • Caveats: The presentation provides directional and selected metrics, but does not disclose experiment design, confidence intervals, baseline definitions, catalog coverage, or the absolute conversion lift associated with retrieval and ranker improvements.; Several claims refer to internal systems and slides not fully available in the transcript, so the exact semantic-ID learning method and memory-extraction pipeline are not specified.
  • Implications: The primary architectural asset is not a single model; it is a representation layer that avoids repeated semantic interpretation and allows new product surfaces to reuse prior work.; Catalog grounding should constrain generative suggestions and reformulations, especially where inventory differs by merchant, geography, or time.

Notable Concepts & Terms

  • Distill the LLM, don't serve it: DoorDash's core production principle: use expensive LLM reasoning offline to create labels and representations, then serve smaller or traditional low-latency models.
  • Graded relevance (0–1–2): A supervision scheme that distinguishes an exact query match, an acceptable substitute, and an irrelevant result rather than treating relevance as binary or inferring it from clicks.
  • Ordinal relevance tower: A ranking-model head trained on LLM-derived relevance levels alongside click, add-to-cart, and conversion heads, allowing a value function to trade off relevance and engagement by surface.
  • Two-stage contrastive training: Retrieval training that first shapes broad embedding geometry, then mines and retrains on model-confused hard negatives and misplaced positives.
  • Semantic IDs: Learned hierarchical item codes whose shared prefixes encode broad semantic neighborhoods and whose later tokens encode finer product distinctions.
  • Memory blocks: Extensible structured representations of consumer preferences and context that can be rendered as text, embeddings, and graph/hierarchical features.
  • Composite memory objects: Unified consumer-memory artifacts supplied to LLMs, retrieval/ranking models, personalized collections, and agent experiences.
  • Steerable content generation: Offline LLM generation of constrained merchandising artifacts—such as collection concepts, copy, and candidates—using catalog semantics and consumer context.

Operator Notes / Why Ken Should Care

  • Adopt an explicit offline “reasoning-to-supervision” pipeline for any high-volume workflow where live LLM inference is costly, slow, or operationally fragile; include human seeds, disagreement audits, stronger-model adjudication, and a lightweight distilled labeler.
  • Instrument a semantic-versus-engagement evaluation split: track constraint satisfaction and graded relevance separately from CTR or conversion so popularity bias cannot masquerade as quality.
  • For OpenClaw or agent systems, make memory portable across text, vector, and structured/graph forms rather than storing it only as uninspectable embeddings or prompt summaries.
  • Require every generated recommendation, query expansion, or collection to pass a real-time grounding/eligibility layer for available inventory, permissions, tool outputs, or current state.
  • Prioritize hard-negative mining from current retrieval/ranking failures; use the error set as the next training and labeling queue rather than relying solely on historical interaction data.

Source/Metadata

  • Title: Distill the LLM, Don't Serve It: Search & Personalization at DoorDash — Raghav Saboo, DoorDash
  • Transcript words: 4990
  • Duration seconds: 1330
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The latter portion substantially repeats earlier content.
Full transcript 2720 words · 22 min read
0:11

Thank you everyone for coming to this talk and Devanj for organizing this. I'm Raghav. I'm a staff machine learning engineer at DoorDash. I work on search and personalization and today I'm going to talk about the ways that we're integrating LLMs for our marketplace discovery, specifically in four pieces and how these primitives are working towards our integration of LLMs into traditional RECSs as well.

0:23

So I'll start off today with a simple claim. The claim is that to build effective discovery, in this case for a marketplace like DoorDash, the real bottleneck is semantic understanding. So that is knowing what items mean for users given some context and what a shopper actually intends to do on an app like DoorDash. And historically we have treated these as engagement optimization problems. However, LLMs give us a genuinely new lever to extend this. So this talk is about the four primitives that have allowed us to start leveraging LLMs for problems at DoorDash.

0:29

As you may know, DoorDash is now grown well beyond restaurants. So we're in grocery, retail, pets, gifting and more. And our goal is to capture every shoppable moment. And we know that across these verticals, the users frequently have very broad shopping missions. For example, a consumer may say when they come to our app is, I just adopted a puppy. What do I need to get started this week? And this consumer may start on searching for puppy food because they just adopted the puppy. That one query could unfold into further query reformulations. They land on collections for other shopping needs, complimentary items such as crates, leashes, training pads, all the way to checkout. And this could happen over multiple sessions, multiple days. And capturing that whole arc is pretty hard because it spans multiple models. And this is where the four primitives I will cover come in.

0:39

Specifically supervision, catalog semantics, semantic personalization, and steerable content generation, as we heard previously as well. It's a system of shared representations that help us map this customer's journey with the use of LLMs.

0:46

Now with supervision, the question we're asking or looking to address is how do we teach our retrieval and ranking systems what good means for different tasks? And here's a simple example. A shopper searches for gluten-free pasta. And if we only rank from engagement, very popular regular spaghettis might show up. And because it sells well, it may continue to show up. Even gluten-free white bread might be retrieved and show up because it overlaps with gluten-free. But neither of these are clearly the right answer. What we actually want is graded relevance. A true gluten-free pasta should be a high-relevance item. Chickpea pasta might be a reasonable substitute. And regular spaghetti, while it's popular, it fails the constraint that the user has.

0:54

Now to solve this, there are obviously two sources of truth. We could look at human annotation. But that's expensive, slow, and becomes quickly stale because our catalog changes quite rapidly. Second is behavioral signals. They're abundant, but obviously, as I mentioned, biased by exposure. Position, price, promotions, and previous model choices. So the missing signal is a scalable reasoning signal. This is where LLMs are very useful as they offer a way to produce that supervision for such tasks.

1:02

So in our case, the pipeline starts with human ground truths. We build a high-quality seed set of example query item pairs on a three-level relevant scale, 0-1-2. We audit the suspicious cases. For example, if a human label says that an item is irrelevant, but that item performs very well on add to cart or conversions, we send that case to a stronger LLM to reevaluate with more granular prompts. We also reconcile with other models. For example, our query to taxonomy or category models. And if a query maps to a set of valid categories and the items belong to those valid categories, we may adjust the label accordingly. And through this process, we're able to achieve a pretty high-precision golden dataset that we then fine-tune a lightweight LLM, let's say, for example, a GPT-4-0 mini, on top of it. And that's when it really pays off. We are able to use that fine-tuned labeler offline across our full catalog to generate full graded query item pairs. And those labels become one shared target for both retrieval and ranking systems.

1:12

This pattern has been pretty important for us over the past few years. Use expensive reasoning once, offline, then distill it into models that can be served cheaply and quickly. Now let's take retrieval first. The challenge is standard embeddings at e-commerce scale quickly collapse the relevance distinction. Items that are merely related sit close to items that actually match. And the model may know two items are related, but not be able to separate exact matches, substitutes, complements. So our fix is a two-stage contrastive method, trained on the same graded labels.

1:29

Stage one, which we call mining, does a global geometry shaping. We use two tower encoders with multi-level supervised contrastive loss. After that, we use that base model to mine harder negatives. And these are items that the model confuses. Hard negatives that are ranked too high or strong positives ranked too low. And the second stage after that is where we really put the model through curriculum training, use those hard negatives, relabel them. And our relevance of the retrieval layer improves sharply after that.

1:36

So after stage one, we see overlaps between relevant and moderately relevant. But after stage two, there is a significant difference. And this has been one of our biggest levers in retrieval improvement across the board. We've improved NDCG of relevance by 2.3%. And that's also been true with our downstream business metrics.

1:46

And we did the same with exercise with our ranking models as well, distilling LLM reasoned labels into our rankers. However, combining this with business and engagement objectives. Here we add a new tower, in this case an ordinal relevance tower, on top of the LLM graded labels. And this prediction sits right alongside existing engagement towers, click, add to card, and conversion. And since they share the same bottom layers, the semantic fit and relevance fit gets distilled in the same backward pass. And our model is able to predict probability of relevance across the different levels. And we are able to blend through a value function on top of this that allows us to fine tune between engagement and relevance for different surfaces.

1:56

Now, these two examples are part of a larger theme of our work across similar projects. Where it seems right, we're not replacing the retrieval and ranking system with an LLM, but rather distilling the reasoning and understanding into some production ranking architecture. And this allows us to quickly test and also scalably improve our models with LLMs while keeping the business and system objectives of these surfaces in mind. So once we have supervision, the next question is, how do we represent a very large and constantly changing catalog in a way that every model, be it language models or traditional retrieval and ranking models, can understand.

2:11

And at DoorDash's scale, we're surfacing items across stores internationally. And that's easily a few billion items at the store item level. And for these items, you obviously have unique SKU IDs, but it says nothing semantically. So we also have a human-curated taxonomy that helps us classify and categorize these items into meaningful spaces. So we have different ways, but often this is too coarse and too rigid. So for example, something like hot sauces may just have sauces, hot sauces in the taxonomy, but it doesn't tell you anything about the relationship between Hoifeng or Franks or Tabasco.

2:18

So that's where we've also adopted semantic IDs. It's been introduced previously already. It's a backbone for a lot of our work now. And what it does is it gives us a short hierarchical code that's analogous to our taxonomy, but the ability to control the fine-grained nature of it. And the result of this learned taxonomy is that each item gets this hierarchical code and the prefix captures broad neighborhoods. And the later tokens capture finer distinctions.

2:27

So, for example, in the map over on the slide, you'll see hot sauces and the structure emerging from the data with zero labels. Each one share the same first and second prefix in this ID sequence, but then split by speciality. So now you've got Mexican, Caribbean, Korean hot sauces being split up. And the important property is that this code that's learned is comparable by prefix, but also usable by all of our downstream models. And that unlocks a few things for us.

2:37

So the first is cross-category comparisons or relationships. Chips, salsa, guacamole may live in different taxonomy branches, but a shared semantic ID neighborhood can indicate that they belong together in a shopping mission. Second, cold start problems. Oftentimes, you get new catalog items added by stores. And with techniques like n-gram and byte-parent coding of these tokens, we're able to scalably add these to our models as sparse ID features. Third, tail coverage. Sparse items inherit a lot of signal from semantically related items. And we don't have to wait for the volume or exposure to consumers for these items. And then lastly, nice to have is the fact that we can actually also do a reverse audit. So we can audit our catalog and how well our human labels agree with the semantically learned labels.

2:45

So a couple of examples of where we're using it today, just to give some real metrics. It has been one of our biggest improvements to our ranker. In this case, it's actually improved our ranker by improving MRR between 4 to 5%. And that's translated to pretty big conversion wins as well. The second one that I'm particularly fond of is query reformulation. So again, going back to the hot sauce examples, sriracha can lead to chili garlic sauce or sambal olek. And because these queries map onto our catalog-grounded semantic neighborhood, that results in much more relevant queries that we're suggesting to users. Ultimately, if these are queries that we don't have inventory for on DoorDash, they're meaningless. And with the addition of semantic IDs in this query graph, we were actually able to see pretty massive MRR gains as well for this piece of work.

2:55

So for now, we've talked about item mapping and relevance. And the next question is consumer context. What does the system know about the shopper? And can that knowledge be reused across models again? And this is where our third primitive comes in, memory.

3:02

So memory in the agent context is pretty well understood now. We apply the same concepts to our recommendation systems. So user embeddings are clearly very useful. We have many representations of our consumers. But they don't necessarily get to why a consumer may have certain intents and why they might have certain preferences. Additionally, LLMs cannot readily use these embeddings. You could tokenize and train these models to learn it. But oftentimes, having some explicit LLM-native counterpart is very useful. And that's what we have found.

3:09

So the idea of memory is to represent the consumer in multiple forms, semantic, inspectable, and reusable. And the way we think about memory is in three timescales. So long-term memory captures durable preferences from orders, searches, browsing, support interactions. Real-time context captures in-session interactions. So cart state, active searches. And then we have stated preferences, which come from agent interactions. So something like Ask DoorDash, for example, where consumers are able to explicitly state their constraints and preferences as well.

3:17

And the way we represent consumers in this long-term memory is through memory blocks. And these are structured in a way that allows us to add new dimensions of the user, from dietary preferences to dining preferences to substitute preferences, as we learn more about the consumer. And that is decoupled from the downstream system that doesn't need to reinterpret this for their own use cases. So each consumer's memory is materialized in multiple forms. So first is text. It's human-readable. It captures the consumer in a compact way. Second are latent vectors, embeddings of those memory blocks that can be fed into retrieval and ranking systems. And third are graph and tree, or hierarchical approaches. So type relationships between consumers and brands, taxonomies, and memory-revealed preferences. And these are assembled into composite memory objects that both ML models and LLMs use downstream.

3:24

So an example is this context graph where we connect consumers with extracted memory concepts. And a graph fits neatly into shopping journeys because oftentimes consumer-item interactions are sparse and multi-hot. And in this case, these context graphs are able to link consumers across memory concepts that we previously did not have relationships for. And it's particularly helpful at fine-grained taxonomy levels. In our case, we're seeing for retrieval using graph-based embeddings outperforming our existing taxonomy-based embeddings.

3:31

And today, this memory framework shows up in three places. The first is personalized collections. I'll talk about that a bit later. Second is agent personalization. So Ask DoorDash, for example, uses this to personalize its sessions for you. And third, in retrieval and ranking models, as I said, we encode these memory blocks and feed them as features in downstream models as well.

3:37

So the fourth primitive is steerable content generation. We heard quite a bit about it in the previous talk. And this slide is my take on how we're piecing all these primitives together. So all of these inputs, semantic IDs, memory blocks, and graded relevance or LLM supervision, feed into multiple models, be it LLMs or small language models or traditional models. And from these inputs, we can generate different output shapes. Ranked semantic ID lists, carousel titles, sub copies. And these show up across different surfaces today on DoorDash.

3:44

And a direct example of this is our personalized collections on store pages. So historically, collections on store pages on DoorDash have been a fixed library or attribute-based. And we've been able to expand that through consumer-level collection generation. And today, this is an offline LLM process where we take in consumer memory, semantic IDs, and synthesize collections all the way from title, subtitle, to the actual items. And the important part is that control and steerability. We're able to react to occasions and moments and build those collections as needed for different consumers.

3:51

At serving time, because this is all batch and generated offline through LLMs, we're able to still use our existing retrieval and ranking stack for item hydration, collection ranking, etc. And this is, for example, what a shopper actually sees. For a shopper with plant-based affinity, the system can generate plant-based pantry rows. For cat-only households, it can generate cat-drive food rows. And if a session indicates that a person is going through pantry restocking, it can bias towards pantry staples.

3:59

And our early tests show that consumers are feeling the benefits of these tailored experiences. An example is within our pets vertical, we've been able to drive close to 1% order rate increases and 0.6% in active users. So to conclude, there are three takeaways I would love for you to take away from this talk and how LLMs fit into search and recommendations. First, discovery is a semantic understanding problem. It's not only about engagement. LLMs give us a way to reason about item meaning and shopper intent, in DoorDash's case, for example.

4:22

Second, distill LLM reasoning into primitives. Capture reasoning offline as labels, semantic IDs, memory, and then let smaller and faster models serve it. The online LLM call is often not the product architecture you need. Third, shared representations create many use cases. Once you have these primitives, they can power retrieval, ranking, content generation, et cetera. So last but not least, thank you. And thank you to all the collaborators at DoorDash who have helped ship a lot of these things as well. Thank you. relevant scale, 0-1-2. We audit the suspicious cases. For example, if a human label says, you know,

5:04

that an item is irrelevant, but that item performs very well on add to cart or conversions, we send that case to a stronger LLM to reevaluate with more granular prompts. We also reconcile with other models. You know, for example, our query to taxonomy or category models. And if a query maps to a set of valid categories and the items belong to those valid categories, we may adjust the label accordingly. And through this process, we're able to achieve a pretty high-precision golden dataset that we then fine-tune a lightweight LLM, let's say, for example, a GPT-4-0 mini, on top of it. And that's when

5:47

it, you know, really pays off. We are able to use that fine-tuned labeler offline across, you know, our full catalog to generate full graded query item pairs. And those labels become one shared target for both retrieval and ranking systems. This pattern has been pretty important for us over the past few years. Use, you know, expensive reasoning once, offline, then distill it into models that can, you know, be served cheaply and quickly. Now let's take retrieval first. The challenge is, you know, standard embeddings at e-commerce scale quickly collapse the relevance distinction. Items that are merely related

6:34

sit close to items that actually match. And the model may know two items are related, but not be able to separate exact matches, substitutes, complements. So our fix is a two-stage contrastive method, trained on the same graded labels. Stage one, which we call mining, does a global geometry shaping. We use two tower encoders with multi-level supervised contrastive loss. After that, we use that base model to mine harder negatives. And these are items that the model confuses. Hard negatives that are ranked too high or strong positives ranked too low. And the second stage after that is where we really put the model

7:22

through curriculum training, use those hard negatives, relabel them. And our, you know, relevance of the retrieval layer improves sharply after that. So after stage one, we see overlaps between, you know, relevant and moderately relevant. But after stage two, there is a significant difference. And this has been one of, like, our biggest levers in retrieval improvement across the board. You know, we've improved NDCG of relevance. We've improved NDCG of relevance by 2.3%. And that's also been true with our downstream business metrics. And we did the same with exercise with our ranking models as well, distilling LLM reasoned labels into our rankers.

8:08

However, combining this with business and engagement objectives. Here we add a new tower, in this case an ordinal relevance tower, on top of the LLM graded labels. And this prediction sits right alongside existing engagement towers, click, add to card, and conversion. And since they share the same bottom layers, the semantic fit and relevance fit gets distilled in the same backward pass. And our model is able to predict, you know, probability of relevance across the different levels. And we are able to blend through a value function on top of this that allows us to fine tune between engagement and relevance for different surfaces.

8:59

Now, these two examples are part of a larger theme of our work across similar projects. Where it seems right, we're not replacing the retrieval and ranking system with an LLM, but rather distilling the reasoning and understanding into some production ranking architecture. And this allows us to kind of quickly test and also scalably improve our models with LLMs while keeping, you know, the business and system objectives of these surfaces in mind. So once we have supervision, the next question is, how do we represent a very large and constantly changing catalog in a way that every model, be it language models or, you know, traditional retrieval and ranking models,

9:45

that they can understand. And at, you know, DoorDash's scale, you can imagine we're kind of surfacing items across stores across, internationally actually, now. And, you know, that's easily a few billion items at the store item level. And for these items, you obviously have unique SKU IDs, but it says nothing semantically. So we also have a human-curated taxonomy that helps us classify and categorize these items into meaningful spaces. So we have a lot of different ways, but often this is too coarse and too rigid.

10:19

So for example, something like hot sauces may just have like sauces, hot sauces, and the taxonomy, but it doesn't tell you anything about the relationship between Hoifeng or Franks or Tabasco. So that's where we've also adapted or adopted semantic IDs. So we've got a lot of different ways. It's been introduced previously already. It's a backbone for a lot of our work now. And what it does is it gives us, you know, a short hierarchical code that's analogous to our taxonomy, but the ability to control the fine-grained nature of it. And the result of this learned taxonomy is that each item gets this hierarchical code and the prefix captures broad neighborhoods.

11:10

And, you know, the later tokens capture finer distinctions. So, for example, in the map over on the slide, you'll see hot sauces and the structure emerging from the data with zero labels. Each one share the same first and second prefix in this ID sequence, but then split by speciality. So now you've got Mexican, Caribbean, Korean hot sauces being split up. And the important property is that this code that's learned is, you know, comparable by, as I said, by prefix, but also usable by all of our downstream models. And that unlocks a few things for us. So the first is cross-category comparisons or relationships.

11:56

You know, chips, salsa, guacamole may live in different taxonomy branches, but a shared semantic ID neighborhood can indicate that they belong together in maybe a shopping mission. Second, cold start problems. Oftentimes, you get new catalog items added by stores. And with techniques like n-gram and byte-parent coding, of these tokens, we're able to scalably add these to our models as sparse ID features. Third, tail coverage. So sparse items inherit a lot of signal from semantically related items. And, you know, we don't have to wait for the volume or exposure to, you know, to consumers for these items.

12:42

And then lastly, nice to have is the fact that we can actually also do a reverse audit. So we can audit our catalog and how well our human labels agree with the semantically learned labels. So a couple of examples of where we're using it today, just to give some real metrics. It has been one of our biggest improvements to our rancor. In this case, it's actually improved our rancor by improving MRR between 4 to 5%. And that's translated to pretty big conversion wins as well. The second one that I'm particularly fond of is query reformulation. So again, going back to the hot sauce examples, riracha can lead to chili garlic sauce or sambal, olek.

13:31

And because these queries map onto, like, our catalog-grounded semantic neighborhood, that results in much more relevant queries that we're suggesting to users. Ultimately, if these are queries that we don't have inventory for on DoorDash, they're meaningless. And with the addition of semantic IDs in this query graph, we were actually able to see pretty massive MRR gains as well for this piece of work. So for now, we've talked about, you know, item mapping and relevance. And the next question is consumer context. What does the system know about the shopper? And can that knowledge be reused across models again? And this is where our third primitive comes in, memory.

14:22

So memory in the agente context is pretty well understood now. We apply the same concepts to our recommendation systems. So user embeddings are clearly very useful. We have many representations of our consumers. But they don't necessarily get to why a consumer may have certain intents and why they might have certain preferences. Additionally, LLMs cannot readily use these embeddings. You could tokenize and, like, train these models to learn it. But oftentimes, having some explicit LLM-native counterpart is very useful. And that's what we have found. So the idea of memory is to represent the consumer in multiple forms, semantic, inspectable, and reusable.

15:16

And the way we think about memory is in three timescales. So long-term memory captures durable preferences from orders, searches, browsing, support interactions. Real-time context captures agentic interactions. Oh, sorry, in-session interactions. So cart state, active searches. And then we have stated preferences, which come from agentic interactions. So something like Ask DoorDash, for example, where consumers are able to explicitly state their constraints and preferences as well. And the way we represent consumers in this long-term memory is through memory blocks.

16:02

And these are structured in a way that allows us to add new dimensions of the user, you know, from dietary preferences to dining preferences to substitute preferences, as we learn more about the consumer. And that is decoupled from, like, the downstream system that doesn't need to reinterpret this for their own use cases. So each consumer's memory is materialized in multiple forms. So first is text. It's human-readable. It captures the consumer in a compact way. Second are latent vectors, embeddings of those memory blocks that can be fed into retrieval and ranking systems. And third are graph and tree, or hierarchical approaches.

16:49

So, you know, type relationships between consumers and brands, taxonomies, and memory-revealed preferences. And these are assembled into, you know, composite memory objects that both ML models and LLMs use downstream. So an example is, you know, this context graph where we connect consumers with extracted memory concepts. And a graph fits neatly into shopping journeys because, you know, oftentimes consumer-item interactions are sparse and multi-hot. And in this case, these context graphs are able to link consumers across memory concepts that we previously did not have relationships for. And it's particularly helpful at fine-grained taxonomy levels.

17:40

In our case, we're seeing for retrieval using graph-based embeddings outperforming our existing taxonomy-based embeddings. And today, this memory framework shows up in three places. The first is personalized collections. I'll talk about that a bit later. Second is agentic personalization. So, Ask DoorDash, for example, uses this to personalize its sessions for you. And third, in retrieval and ranking models, as I said, we encode these memory blocks and feed them as features in downstream models as well.

18:21

So, the fourth primitive is steerable content generation. We heard quite a bit about it in the previous talk. And this slide is my take on how we're piecing all these primitives together. So, all of these inputs, semantic IDs, memory blocks, and graded relevance or LLM supervision, feed into multiple models, be it LLMs or small language models or traditional models. And from these inputs, we can generate different output shapes. You know, ranked semantic ID lists, carousel titles, sub copies. And these show up across different surfaces today on DoorDash. And a direct example of this is our personalized collections on store pages.

19:11

So, historically, collections on store pages on DoorDash have been a fixed library or attribute-based. And we've been able to kind of expand that through consumer-level collection generation. And today, this is an offline LLM process where we take in consumer memory, semantic IDs, and synthesize collections all the way from title, subtitle, to the actual items. And the important part is that control and steerability. We're able to react to occasions and moments and build those collections as needed for different consumers.

19:54

At serving time, because this is all batch and generated offline through LLMs, we're able to still use our existing retrieval and ranking stack for item hydration, collection ranking, etc. And, you know, this is, for example, what a shopper actually sees, right? For a shopper with plant-based affinity, the system can generate plant-based pantry rows. For cat-only households, it can generate cat-drive food rows. And if a session indicates that, you know, a person is going through pantry restocking, it can bias towards pantry staples. And our early tests show that, you know, consumers are feeling the benefits of these tailored experiences.

20:37

An example is within our pets vertical, we've been able to drive close to 1% order rate increases and 0.6% in active users. So, to conclude, there are three takeaways I would love for you to take away from this talk and how LLMs fit into search and recommendations. First, discovery is a semantic understanding problem. It's not only about engagement. LLMs give us a way to reason about item meaning and shopper intent, in DoorDash's case, for example. Second, distill LLM reasoning into primitives. Capture, you know, reasoning offline as labels, semantic IDs, memory, and then let smaller and faster models serve it.

21:25

Second, the online LLM call is often not the product architecture you need. Third, shared representations create many use cases. Once you have these primitives, they can power, you know, retrieval, ranking, content generation, et cetera. So, last but not least, thank you. And thank you to all the collaborators at DoorDash who have helped ship a lot of these things as well. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note