AI Engineer

Why LLM Recommenders Will Be AI's Biggest Consumer App — Devansh Tandon, Meta

2025 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: LLM-native recommender systems will become a dominant consumer-AI application because they can scale with data, compute, and model size while producing engagement far more inference-efficiently than text-generating chat apps.
  • Why it matters: The talk offers a practical architecture for turning a foundation model into a recommendation system—semantic IDs, bilingual pre-training, recommendation post-training, and surface fine-tuning—plus a useful economic lens for evaluating AI products.
  • Best use: Use it to inform investment and product thinking on AI-native discovery, consumer agents, retrieval/ranking architectures, and applications where models select existing supply rather than generate every output token.

Executive Summary

Devansh Tandon argues that recommenders are following the same basic scaling-law dynamic that catalyzed the LLM race: more data, compute, and model capacity produce predictable improvements in offline ranking metrics and, at Meta scale, engagement and revenue. He frames the commercial engine as a "tokens in, engagement out" flywheel: better models drive better recommendations, which increase time spent and monetization, funding the next training cycle.

His central implementation prescription is to convert each item in a content corpus into a compact semantic ID, train a base LLM to connect those IDs with natural-language meaning and behavioral sequences, then post-train it for tasks such as engagement prediction and candidate reranking. Rather than retraining distinct models for every product surface, a shared recommendation foundation model can receive light surface-specific adaptation.

The product shift is from opaque feeds optimized only from implicit behavior to interactive, steerable recommendation systems. A user can express an explicit, temporary goal—such as following the FIFA World Cup—and the model combines that instruction with behavioral history to generate personalized results. Tandon expects recommendation agents eventually to plan, retrieve, rank, critique, and refine results in multi-step loops rather than make a single ranking pass.

The investment argument is economic as much as technical. Content feeds can use a small number of output tokens to emit pointers to creator-supplied content, whereas chat products must generate thousands of tokens of new content per interaction with much larger serving models. Tandon estimates feeds can therefore be 100x or more cheaper per hour of user engagement, positioning LLM recommenders as a potentially enormous and more token-efficient consumer-AI category.

Key Takeaways

  • Claim: Recommendation systems exhibit scaling behavior analogous to LLMs: increasing data, training compute, and model size improves ranking quality and business outcomes. | Evidence: Tandon cites Meta's 2024 HSTU paper and a follow-up paper as showing scaling gains in offline recommendation evaluations. He also links Instagram Reels' reported 30% year-over-year watch-time growth to architecture simplification for efficient scaling, doubled interaction-history length, and richer interaction signals. | Implication: Recommenders should be treated as frontier-model infrastructure rather than mature commodity ranking systems; teams with proprietary behavioral data and distribution can compound an existing advantage through larger shared models. | Caveat: The talk asserts a power-law-like pattern but does not provide the underlying curve values, causal experiment details, or evidence that all recommendation domains scale equally cleanly.
  • Claim: The recommender industry is advancing through four overlapping S-curves: traditional systems, LLM-inspired end-to-end models, LLM-native recommenders, and agentic recommendation loops. | Evidence: Traditional systems include two-tower models, sparse networks, rankers, and embedding scaling. Tandon positions Meta HSTU and OneRec in the LLM-inspired stage; Tiger and PLUM as LLM-native examples; and describes an emerging agentic loop that plans, retrieves, ranks, critiques, and refines recommendations. | Implication: For product and investment evaluation, distinguish a company merely applying transformers to ranking from one building a reusable recommendation foundation model or an agentic orchestration layer. | Caveat: The four-stage model is a strategic framework from the speaker, not a demonstrated universal roadmap; companies will likely retain hybrid conventional retrieval and ranking components for latency, control, and operational reasons.
  • Claim: A deployable LLM recommender can be built in three core steps: tokenize content into a domain language, adapt an LLM to understand that language alongside English, and prompt it with user context to decode recommendations. | Evidence: Tandon's five-layer architecture consists of semantic IDs, a base foundation model, pre-training that bridges English and recommender tokens, post-training for engagement prediction or recommendation reasoning, and light product-surface fine-tuning. | Implication: The reusable asset is not a separate feed model per surface but a common representation and model that can support multiple recommendation tasks and experiences. | Caveat: This recipe describes the model layer, not the full serving stack; the transcript does not address candidate-generation recall, latency budgets, safety policy, online experimentation, or feedback-loop management.
  • Claim: Semantic IDs are the key bridge between unstructured content, long behavioral histories, and LLM reasoning because they provide stable learned representations and aggressively compress items. | Evidence: A three-minute Instagram Reel could consume roughly 10,000 text tokens, while Tandon says it can be compressed to about 10 semantic-ID tokens. Similar tennis Reels share their first three semantic tokens, with a later token distinguishing the individual video. | Implication: Any AI discovery system handling large item catalogs or long user histories should consider learned, hierarchical item tokenization rather than passing raw content into context or relying solely on opaque hash IDs.
  • Claim: LLM-native recommenders enable direct user steering and inspectable recommendation logic, creating product experiences that traditional black-box feeds do not support well. | Evidence: Meta's "Your Algorithm" example lets an Instagram user inspect inferred interests, add or remove interests in natural language, and request a timely objective such as following the FIFA World Cup; the model then combines this instruction with prior history to recommend relevant content. Tandon also cites Spotify prompted playlists, YouTube custom feeds, and Ask DoorDash as parallel industry patterns. | Implication: The durable UX opportunity is not simply conversational search: it is preference control over an ongoing personalized system, including explicit goals that complement noisy implicit signals. | Caveat: Showing chain-of-thought-style rationale may improve usability, but the talk does not establish that visible reasoning is faithful to the actual ranking mechanism or safe to expose directly.
  • Claim: Content-feed recommenders have a structural inference-cost advantage over AI chat because they output pointers to existing creator content instead of generating the content itself. | Evidence: Tandon contrasts feeds serving roughly 1–10 billion active-parameter models and emitting semantic-ID pointers with chat apps serving roughly 10–100 billion active-parameter models and generating a few thousand output tokens per turn. He estimates feeds can be up to 100x or more cheaper per hour of consumer engagement. | Implication: When assessing consumer-AI economics, prioritize products that use models to allocate or orchestrate existing supply—content, products, services, workflows—rather than requiring costly generation for every unit of engagement. | Caveat: The 100x figure is presented as a directional estimate without a cost model, and it depends on content supply remaining abundant, relevant, and inexpensive for the platform.
  • Claim: LLM recommenders target an already massive consumer surface whose engagement and monetization are largely determined by recommendation and advertising models. | Evidence: Tandon states that four of the top ten apps by daily active users are content feeds and argues that growth in engagement and monetization for those feeds is driven primarily by recommender and ads models. | Implication: The likely near-term consumer-AI winners may be incumbent platforms that can upgrade high-frequency recommendation loops, while startups need proprietary supply, a concentrated vertical, or a new interaction surface to build comparable flywheels.

Detailed Brief

Training and serving logic behind the recommendation foundation model

  • Claims: Pre-training should teach a model a translation layer between semantic item tokens and natural language, as well as relationships across sequences of user interactions.; Post-training makes the representation operational for recommendation tasks, including reranking a candidate set based on a user's preferences, engagement style, and creator affinities.; Most compute can be concentrated in a shared core model, with comparatively light fine-tuning for each downstream product surface.
  • Evidence: One pre-training task masks the textual description of a video represented by semantic ID ABC and trains the model to generate the description, such as "a shot that was instantly iconic from Wimbledon."; Another training task masks portions of a user's semantic-ID interaction sequence so the model learns which items co-occur in consumption histories.; The post-training example feeds user data plus 30 candidate videos into an LLM ranker, which selects the top five.
  • Caveats: The talk provides no comparison against a strong hybrid baseline on cost, latency, calibration, diversity, freshness, or online reward metrics.; A shared cross-surface model can create governance complexity when product surfaces have conflicting objectives or safety constraints.
  • Implications: Semantic representation learning could unify retrieval, ranking, explanation, and conversational preference editing around one domain vocabulary.; The most defensible data moat is a combination of item understanding, behavioral sequence data, and repeated online feedback—not merely access to a general-purpose base model.

Economic flywheel and consumer-agent direction

  • Claims: The same loop connects training investment, inference, engagement, and monetization in both feeds and chat applications, but the conversion efficiency differs materially.; The agentic frontier extends beyond improving a single model forward pass: agents can coordinate planning, retrieval, ranking, self-critique, and refinement.
  • Evidence: Tandon compares the agentic recommender direction to coding-agent harnesses: model capability gains resemble moving from one model generation to another, while orchestration gains resemble improvements in Claude Code or Codex-like systems.; Creator-uploaded inventory means a feed can use a recommendation output as an address to an existing asset rather than pay to synthesize the asset token by token.
  • Caveats: Agentic multi-pass recommendation may erode the claimed cost advantage if it adds excessive model calls, and its incremental quality gain is not quantified here.
  • Implications: A consumer agent's viability should be evaluated on total cost per successful or engaged outcome, including the cost of all orchestration calls, not on model quality alone.; Markets with abundant third-party supply and recurring user choice may offer better AI-unit economics than pure chat or content-generation products.

Notable Concepts & Terms

  • Tokens in, engagement out: Tandon's operating and economic framework: training and inference tokens must convert into user engagement and monetization strongly enough to finance subsequent model scaling.
  • Semantic ID (SID): A compact learned token sequence representing a content item; it replaces or complements hash IDs, captures semantic similarity, and makes long user histories tractable in an LLM context.
  • Generative retrieval: A retrieval approach in which a model decodes item identifiers directly, rather than only scoring items from a conventional index; it underlies the idea of generating recommendations as semantic tokens.
  • HSTU: A Meta recommendation-model research line cited as evidence that recommender performance improves with scaled model capacity, compute, and data.
  • LLM-inspired recommender: The current leading paradigm in Tandon's taxonomy: end-to-end scaling of transformer-like recommendation models, without necessarily adapting a general language model to the domain.
  • LLM-native recommender: A model adapted from a base LLM so it can reason over both natural language and domain-specific recommendation tokens.
  • Agentic recommender: An emerging multi-step system in which LLM-based agents plan, retrieve, rank, critique, and refine recommendations rather than relying on one forward pass.
  • Your Algorithm: Meta's example of a steerable-feed interface where users can inspect and edit inferred interests using natural language.

Operator Notes / Why Ken Should Care

  • Add "cost per engaged hour" and "generated output versus pointer-to-existing-supply" to the evaluation template for consumer-AI opportunities.
  • For any OpenClaw or agentic discovery workflow, separate the architecture into candidate retrieval, ranking, critique, and policy layers; measure whether multi-step orchestration earns its added inference cost.
  • Evaluate semantic-ID-like representations for any domain with large catalogs and long interaction histories, especially where raw-item context would be prohibitively large.
  • Treat user-editable preference state as a product and control-plane requirement: preserve explicit instructions, distinguish durable preferences from temporary intent, and provide auditable overrides.
  • Monitor whether recommendation explanations are faithful, privacy-safe, and policy-compliant before exposing model rationales; do not equate fluent explanations with causal transparency.

Source/Metadata

  • Title: Why LLM Recommenders Will Be AI's Biggest Consumer App — Devansh Tandon, Meta
  • Transcript words: 3305
  • Duration seconds: 1080
  • Timestamp note: No timestamps or chapters were provided. The latter portion of the transcript repeats the closing sections on interactive recommendations and token efficiency.
Full transcript 2614 words · 15 min read
0:00

Welcome everyone.

0:12

Thank you for coming out to the LLM Rexxus track. I'll be sharing the first talk. My talk's titled, Tokens and Engagement Out, Training LLM Recommenders. And I want to make two big arguments today. The first is that recommendation systems scale just like LLMs do, and that the field is very early in that scaling curve. And the second is that the LLM recommender's going to be one of the biggest consumer applications of AI. Okay, so quickly about me. I currently work at Meta on research and product.

0:36

I lead a team called Meta Recommendations Research, which is a group that's training frontier models, LLMs and recommenders that power Instagram, Facebook ads, and the Meta family of apps. Before this, I was at Google for a long time, working on a lot of the key ML teams, including DeepMind and YouTube. Last year, I gave a talk at AI Engineer called Teaching Gemini to Speak YouTube about two ideas: semantic IDs and generative retrieval, which we also wrote two papers about, which I've linked here.

0:47

It was really fun, and it led to a lot of discussions and collaborations. The ideas behind semantic ID and generative retrieval have really taken off in the industry over the last year, and they've moved from research to scaled production systems. And I've seen exciting launches and papers from YouTube, Meta, Spotify, DoorDash, across the industry, and we have a couple of examples of that later today. This year, I want to talk about four sections: Recommendation Scaling Curves, a framework of four recommendation paradigm S-curves that we are climbing as an industry, sharing the recommender recipe, and finally, this consumer AI app framework.

1:00

Let's start with scaling curves. I wanted to start with this landmark scaling curve paper from 2020, which feels like a lifetime ago. This is when Dario was still at OpenAI and Anthropic didn't exist yet. But the core idea that this paper shared is the power law of scaling. As you increase model size, data, and the amount of compute flops trained for model training, the loss falls on this log linear scale. And this clean and predictable curve is what really set off the race for the AI frontier, because you can forecast what model quality and capability improvements will look like, and this is what's underwriting the massive CapEx investments and the AI build-out today.

1:18

It turns out that recommendation systems follow a very similar scaling law. In fact, before this wave of LLMs, RECs were the largest production ML models in big tech companies, and they're still some of the largest models that are served at a scale of a billion-plus daily active users. And they follow this similar power law scaling curve. On the x-axis, you have data, compute, and model size. And on the y-axis, you would see falling loss or in this chart an improvement in recommendation quality. In offline evals, it's net entropy or AUC gains.

1:43

And then when it's translated to a real production launch, it's engagement impact, revenue impact, at some of the biggest consumer app scale. Here's a real example from Meta that demonstrates these power law scaling curves. The first is a paper, HSTU from 2024, and the second is a follow-up from this year. Both demonstrate that as we scale model size, compute, and data, we see this clear improvement in offline evals of recommendation quality. These scaling curves aren't just academic research. They're driving real product impact at scale for some of the biggest consumer businesses in the world. Here's a couple of examples I have from Meta's recent earnings reports.

2:17

Instagram reels had a strong quarter with 30% year-on-year watch time, and the optimizations we made to improve the quality of recommendations included simplifying our ranking architecture to enable efficient model scaling and longer interaction histories to identify a person's interests. We doubled the length of user interaction sequences used for training Instagram and increased the richness of each user interaction. So these are direct parallels to the power law scaling curves for LLMs. And I want to introduce this idea of a flywheel of tokens in engagement access. So this is the way that we're going to run out, which is what's powering all of these RECs model scaling.

2:39

You train a model, you then run inference on it, which is the tokens in. That recommendation model results in better content recommendations. It drives consumer engagement, daily active users' time spent. It translates to monetization and ads or subscription, which pays for the next model training run. And so every step on the scaling curve is one loop around this flywheel. And a lot of consumer apps are spinning this core flywheel at the heart of their business. So we have a long way to scale these recommender systems. I want to talk about the four paradigms that I see the industry progressing through.

3:20

The first S-curve was more traditional RECsys, where this S-curve focused more on feature engineering and user and content embeddings. Most production systems are still sitting on this curve. They're running some type of two tower, sparse network, rankers, scaling the embedding models. I don't think this curve is going to go away, but model development here will be accelerated with auto research. And things like feature engineering will be handled by agents rather than real ML engineers. The next curve is kind of LLM-inspired models, where you are scaling models ideally end-to-end. HSTU and one-rec papers are examples in this paradigm.

4:00

I think this is where the leading recommender systems in the industry are largely operating today. I think the next S-curve will be this paradigm of LLM-native, where you adapt a base model that understands and can reason and adapt it for recommendation tasks. The Tiger and Plum papers are some examples of this paradigm. And I think the final paradigm that I start to see emerging is agentic, where LLMs will start to orchestrate REC systems in a loop. I think there's a parallel here with coding agents.

4:24

So the LLM-native models are like improving the core capabilities of the model, going from OPUS 4.5 to 4.8, versus the agentic curve will be like improving the coding harness behind Claude code or codex. And so instead of just having a single forward pass through the recommender, you can imagine a loop where agents plan, retrieve, rank, and then critique the recommendations. They can refine them by calling models again and finally deliver the recommendations. This, I think, is an interesting area of research now. Let me jump into the framework of all the four RECs paradigms. I think companies are scaling across each of these curves in parallel.

4:53

Most of recommendations, I think, lives in LLM-inspired today and is trying to graduate into LLM-native. But then a lot of companies are still using traditional models and climbing that S-curve. Let me shift gears a bit to share the recipe of how to actually build an LLM recommender. I think it's pretty simple. It's three steps. You start with tokenizing your content and creating a language for your domain. Then you want to adapt the LLM so that it understands both English and your domain language and becomes this bilingual model. Finally, you can prompt this model with user information and it will directly decode recommendations from your content corpus.

5:27

Let's go a bit deeper. This is the LLM recommender as a five-layer cake. We'll start at the bottom. That's semantic ID where you're converting your content corpus into tokens that the LLM can understand and reason over. Then you have the base LLM foundation model. This can be an open weights model or an internal first-party model. Then the core training stages. Pre-training is around bridging English and these recommender tokens. Post-training is about steering the model towards recommendation tasks like predicting engagement or reasoning over recommendations. And then finally, you can do some light surface-specific fine-tuning to deploy it on a product surface.

5:57

The exciting thing about this paradigm is most of the compute is shared across all of the product surfaces. So you don't have to train individual models from scratch for every product surface. I'll go a bit deeper into each stage. For semantic IDs, I think this has seen incredible adoption. A lot of teams are replacing their hash ID with the SID in traditional models and seeing good impact. I think there's two big reasons to tokenize content. The first is it gives you a stable representation for models to learn over rather than a constantly shifting hash that the model can only memorize.

6:27

And the second is compression. You want to be reasoning over these long sequences of user interactions. And if you don't compress the content, for example, a three-minute Instagram reel video would be 10,000 tokens. And it will just fill up the context window too quickly. So you have to compress it into about 10 tokens. And so here I have some examples of Instagram reels about tennis. You can see that the semantic token shares the prefix of the first three tokens because they're very similar reels. You can imagine the first token representing sports and the second two tokens representing tennis. And then the final token making these videos individual.

7:08

Once you have a semantic ID, you can train it to understand both English and semantic ID. And so the task I have on the left for pre-training here is an example of where you prompt with a video with semantic ID ABC and the description blank. And the output is a shot that was instantly iconic from Wimbledon. Here you're teaching the model to connect these semantic tokens with synthetic English natural language text. The example on the right is about reasoning over sequences of semantic IDs. And so in a user's interaction history, you can mask some parts of the sequence and the model learns to predict them and understand what videos are watched together in sequence.

7:41

Here I have an example of post-training where we're teaching the model how to re-rank content. So the input is a bunch of user information and 30 candidate videos that are then ranked to be the top five recommendations from this LLM ranker. What's really interesting here is that you can see the chain of thought reasoning of this model. And because this model knows both English and recommendations, you can simply inspect the model and understand why it made the decisions that it did. In this example, the model understands the user's topic interests like comedy, food, DIY, wellness. It understands the engagement style and what creators this user has an affinity towards.

8:18

And then it re-ranks the content based on this chain of thought reasoning. I think this is exciting because once you have a model that can understand both English and recommendations, it opens up new product surfaces and new experiences where users can steer their feed. Here's an example from your algorithm on Instagram where users can talk to the algorithm while they're consuming content. It's a screenshot from scrolling through Reels or when you click in you can understand what the Instagram algorithm thinks about you and your interests. And then you can add or remove interest and talk to it in natural language.

8:59

And so we're going to see this shift, I think, from black box recommendations algorithms to giving users more control over their algorithm and algorithms becoming more interactive and steerable. I'm really excited that users can direct it towards their own goals that are expressed in language rather than just likes or comments. And I think this foundation model can also start to explain its recommendations. And so for this example, I've added an interest that I want to follow the FIFA World Cup at this time.

9:29

And the model would get both my user history and this new input and then be able to decode recommendations that are personalized to me like this free kick that Messi scored recently. I think these interactive recommenders are going to be a really interesting new product surface and we're seeing this across the industry. We have some examples of prompted playlists from Spotify, custom feeds from YouTube, Ask DoorDash, and we'll be hearing more from speakers about these. Finally, I want to talk about the framework of tokens in engagement out. This flywheel that I started with of model training, inference, consumer engagement, and then monetization.

10:06

This is actually the same flywheel that's shared by content feeds and the AI chat apps. And this is a lens that you can use to evaluate any consumer app. What will make an app successful on the training ROI side is how well can it translate compute into a frontier model. On the inference side, how well can the inference tokens translate into engagement and then monetization. And so if you try to compare content feeds and AI chat apps, I think that LLM recommenders are actually structurally more token efficient than the AI chat. So on the left you have content feeds like Instagram, Facebook, TikTok, and YouTube.

10:42

On the right you have the big AI chat apps like Gemini, ChatGPT, and Claude. Content feeds are currently using models that are around 1 to 10 billion active parameters. The chat apps are serving much larger models, 10 to 100 billion active parameters. The output for the content feed is a semantic ID token, which is a pointer or an address to existing content because the content supplied on these content feeds is uploaded by creators. There's a very healthy creator economy. And so the supply of content is effectively free or it's uploaded by creators. For the AI chat apps, they have to decode every token of content themselves.

11:32

And the amount of tokens output in every turn of an LLM chat interaction is a few thousand tokens. And the big difference here is that every token has to be manufactured by the app at inference time. And so what that means is if you compare these two apps on how much inference cost and compute is spent to generate an hour of consumer engagement, there's a huge structural gap where content feeds are significantly cheaper, up to 100 times or more cheaper than AI chat apps, because they're decoding pointers to content rather than content itself. Finally, I want to end with why I think LLM-RECs is one of the most significant consumer AI applications.

12:04

If you look at the top apps by daily active users, these are the top 10 apps, and 4 out of 10 of them are content feeds. And so this is a really significant consumer application. If you look at the content feeds, almost all of the consumer app growth on both the engagement and monetization side is driven by the recommender and ads models. And these are going to be entirely transformed by LLM recommenders. It's a very large and very token efficient application of AI for consumer apps.

12:40

We're going to see a lot of new product experiences come through with steerable and interactive recommendations, explanation of recommendations, and putting more users in control of their experience on these apps. I think we're going to see some really exciting research on consumer agents and recommendation agents that come out over the next year or so. And so this is why I think this is a super exciting area of both research and product at this intersection of LLM and recommendations. That's all. Thank you so much. And I think this foundation model can also start to explain its recommendations.

13:23

And so for this example, I've added an interest that I want to follow the FIFA World Cup at this time. And the model would get both my user history and this new input and then be able to decode recommendations that are personalized to me like this free kick that Messi scored recently.

13:43

I think these interactive recommenders are going to be a really interesting new product surface and we're seeing this across the industry. We have some examples of prompted playlists from Spotify, custom feeds from YouTube, Ask DoorDash, and we'll be hearing more from speakers about these.

14:04

Finally, I want to talk about the framework of tokens in engagement out. This flywheel that I started with of model training, inference, consumer engagement, and then monetization. This is actually the same flywheel that's shared by content feeds and the AI chat apps. And this is a lens that you can use to evaluate any consumer app. What will make an app successful on the training ROI side is how well can it translate compute into a frontier model. On the inference side, how well can the inference tokens translate into engagement and then monetization.

14:45

And so if you try to compare content feeds and AI chat apps, I think that LLM recommenders are actually structurally more token efficient than the AI chat. So on the left you have content feeds like Instagram, Facebook, TikTok, and YouTube. On the right you have the big AI chat apps like Gemini, ChatGPT, and Claude. Content feeds are currently using models that are around 1 to 10 billion active parameters. The chat apps are serving much larger models, 10 to 100 billion active parameters. The output for the content feed is a semantic ID token, which is a pointer or an address to existing content because the content supplied on these content feeds is uploaded by creators.

15:30

There's a very healthy creator economy. And so the supply of content is effectively free or it's uploaded by creators. For the AI chat apps, they have to decode every token of content themselves. And the amount of tokens output in every turn of an LLM chat interaction is a few thousand tokens. And the big difference here is that every token has to be manufactured by the app at inference time. And so what that means is if you compare these two apps on how much inference cost and compute is spent to generate an hour of consumer engagement, there's a huge structural gap where content feeds are significantly cheaper, up to 100 times or more cheaper than AI chat apps,

16:19

because they're decoding pointers to content rather than content itself.

16:25

Finally, I want to kind of end with why I think LLM-REXIS is one of the most significant consumer AI applications. If you look at the top apps by daily active users, these are the top 10 apps, 4 out of 10 of them are content feeds. And so this is a really significant consumer application. If you look at the content feeds, almost all of the consumer app growth on both the engagement and monetization side is driven by the recommender and ads models. And these are going to be entirely transformed by LLM recommenders. It's a very large and very token efficient application of AI for consumer apps.

17:07

We're going to see a lot of new product experiences come through with steerable and interactive recommendations, explanation of recommendations, and just putting more users in control of their experience on these apps. I think we're going to see some really exciting research on consumer agents and recommendation agents that come out over the next year or so. And so this is why I think this is a super exciting area of both research and product at this intersection of LLM and recommendations. That's all. Thank you so much. Thank you very much.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note