Welcome everyone.
Thank you for coming out to the LLM Rexxus track. I'll be sharing the first talk. My talk's titled, Tokens and Engagement Out, Training LLM Recommenders. And I want to make two big arguments today. The first is that recommendation systems scale just like LLMs do, and that the field is very early in that scaling curve. And the second is that the LLM recommender's going to be one of the biggest consumer applications of AI. Okay, so quickly about me. I currently work at Meta on research and product.
I lead a team called Meta Recommendations Research, which is a group that's training frontier models, LLMs and recommenders that power Instagram, Facebook ads, and the Meta family of apps. Before this, I was at Google for a long time, working on a lot of the key ML teams, including DeepMind and YouTube. Last year, I gave a talk at AI Engineer called Teaching Gemini to Speak YouTube about two ideas: semantic IDs and generative retrieval, which we also wrote two papers about, which I've linked here.
It was really fun, and it led to a lot of discussions and collaborations. The ideas behind semantic ID and generative retrieval have really taken off in the industry over the last year, and they've moved from research to scaled production systems. And I've seen exciting launches and papers from YouTube, Meta, Spotify, DoorDash, across the industry, and we have a couple of examples of that later today. This year, I want to talk about four sections: Recommendation Scaling Curves, a framework of four recommendation paradigm S-curves that we are climbing as an industry, sharing the recommender recipe, and finally, this consumer AI app framework.
Let's start with scaling curves. I wanted to start with this landmark scaling curve paper from 2020, which feels like a lifetime ago. This is when Dario was still at OpenAI and Anthropic didn't exist yet. But the core idea that this paper shared is the power law of scaling. As you increase model size, data, and the amount of compute flops trained for model training, the loss falls on this log linear scale. And this clean and predictable curve is what really set off the race for the AI frontier, because you can forecast what model quality and capability improvements will look like, and this is what's underwriting the massive CapEx investments and the AI build-out today.
It turns out that recommendation systems follow a very similar scaling law. In fact, before this wave of LLMs, RECs were the largest production ML models in big tech companies, and they're still some of the largest models that are served at a scale of a billion-plus daily active users. And they follow this similar power law scaling curve. On the x-axis, you have data, compute, and model size. And on the y-axis, you would see falling loss or in this chart an improvement in recommendation quality. In offline evals, it's net entropy or AUC gains.
And then when it's translated to a real production launch, it's engagement impact, revenue impact, at some of the biggest consumer app scale. Here's a real example from Meta that demonstrates these power law scaling curves. The first is a paper, HSTU from 2024, and the second is a follow-up from this year. Both demonstrate that as we scale model size, compute, and data, we see this clear improvement in offline evals of recommendation quality. These scaling curves aren't just academic research. They're driving real product impact at scale for some of the biggest consumer businesses in the world. Here's a couple of examples I have from Meta's recent earnings reports.
Instagram reels had a strong quarter with 30% year-on-year watch time, and the optimizations we made to improve the quality of recommendations included simplifying our ranking architecture to enable efficient model scaling and longer interaction histories to identify a person's interests. We doubled the length of user interaction sequences used for training Instagram and increased the richness of each user interaction. So these are direct parallels to the power law scaling curves for LLMs. And I want to introduce this idea of a flywheel of tokens in engagement access. So this is the way that we're going to run out, which is what's powering all of these RECs model scaling.
You train a model, you then run inference on it, which is the tokens in. That recommendation model results in better content recommendations. It drives consumer engagement, daily active users' time spent. It translates to monetization and ads or subscription, which pays for the next model training run. And so every step on the scaling curve is one loop around this flywheel. And a lot of consumer apps are spinning this core flywheel at the heart of their business. So we have a long way to scale these recommender systems. I want to talk about the four paradigms that I see the industry progressing through.
The first S-curve was more traditional RECsys, where this S-curve focused more on feature engineering and user and content embeddings. Most production systems are still sitting on this curve. They're running some type of two tower, sparse network, rankers, scaling the embedding models. I don't think this curve is going to go away, but model development here will be accelerated with auto research. And things like feature engineering will be handled by agents rather than real ML engineers. The next curve is kind of LLM-inspired models, where you are scaling models ideally end-to-end. HSTU and one-rec papers are examples in this paradigm.
I think this is where the leading recommender systems in the industry are largely operating today. I think the next S-curve will be this paradigm of LLM-native, where you adapt a base model that understands and can reason and adapt it for recommendation tasks. The Tiger and Plum papers are some examples of this paradigm. And I think the final paradigm that I start to see emerging is agentic, where LLMs will start to orchestrate REC systems in a loop. I think there's a parallel here with coding agents.
So the LLM-native models are like improving the core capabilities of the model, going from OPUS 4.5 to 4.8, versus the agentic curve will be like improving the coding harness behind Claude code or codex. And so instead of just having a single forward pass through the recommender, you can imagine a loop where agents plan, retrieve, rank, and then critique the recommendations. They can refine them by calling models again and finally deliver the recommendations. This, I think, is an interesting area of research now. Let me jump into the framework of all the four RECs paradigms. I think companies are scaling across each of these curves in parallel.
Most of recommendations, I think, lives in LLM-inspired today and is trying to graduate into LLM-native. But then a lot of companies are still using traditional models and climbing that S-curve. Let me shift gears a bit to share the recipe of how to actually build an LLM recommender. I think it's pretty simple. It's three steps. You start with tokenizing your content and creating a language for your domain. Then you want to adapt the LLM so that it understands both English and your domain language and becomes this bilingual model. Finally, you can prompt this model with user information and it will directly decode recommendations from your content corpus.
Let's go a bit deeper. This is the LLM recommender as a five-layer cake. We'll start at the bottom. That's semantic ID where you're converting your content corpus into tokens that the LLM can understand and reason over. Then you have the base LLM foundation model. This can be an open weights model or an internal first-party model. Then the core training stages. Pre-training is around bridging English and these recommender tokens. Post-training is about steering the model towards recommendation tasks like predicting engagement or reasoning over recommendations. And then finally, you can do some light surface-specific fine-tuning to deploy it on a product surface.
The exciting thing about this paradigm is most of the compute is shared across all of the product surfaces. So you don't have to train individual models from scratch for every product surface. I'll go a bit deeper into each stage. For semantic IDs, I think this has seen incredible adoption. A lot of teams are replacing their hash ID with the SID in traditional models and seeing good impact. I think there's two big reasons to tokenize content. The first is it gives you a stable representation for models to learn over rather than a constantly shifting hash that the model can only memorize.
And the second is compression. You want to be reasoning over these long sequences of user interactions. And if you don't compress the content, for example, a three-minute Instagram reel video would be 10,000 tokens. And it will just fill up the context window too quickly. So you have to compress it into about 10 tokens. And so here I have some examples of Instagram reels about tennis. You can see that the semantic token shares the prefix of the first three tokens because they're very similar reels. You can imagine the first token representing sports and the second two tokens representing tennis. And then the final token making these videos individual.
Once you have a semantic ID, you can train it to understand both English and semantic ID. And so the task I have on the left for pre-training here is an example of where you prompt with a video with semantic ID ABC and the description blank. And the output is a shot that was instantly iconic from Wimbledon. Here you're teaching the model to connect these semantic tokens with synthetic English natural language text. The example on the right is about reasoning over sequences of semantic IDs. And so in a user's interaction history, you can mask some parts of the sequence and the model learns to predict them and understand what videos are watched together in sequence.
Here I have an example of post-training where we're teaching the model how to re-rank content. So the input is a bunch of user information and 30 candidate videos that are then ranked to be the top five recommendations from this LLM ranker. What's really interesting here is that you can see the chain of thought reasoning of this model. And because this model knows both English and recommendations, you can simply inspect the model and understand why it made the decisions that it did. In this example, the model understands the user's topic interests like comedy, food, DIY, wellness. It understands the engagement style and what creators this user has an affinity towards.
And then it re-ranks the content based on this chain of thought reasoning. I think this is exciting because once you have a model that can understand both English and recommendations, it opens up new product surfaces and new experiences where users can steer their feed. Here's an example from your algorithm on Instagram where users can talk to the algorithm while they're consuming content. It's a screenshot from scrolling through Reels or when you click in you can understand what the Instagram algorithm thinks about you and your interests. And then you can add or remove interest and talk to it in natural language.
And so we're going to see this shift, I think, from black box recommendations algorithms to giving users more control over their algorithm and algorithms becoming more interactive and steerable. I'm really excited that users can direct it towards their own goals that are expressed in language rather than just likes or comments. And I think this foundation model can also start to explain its recommendations. And so for this example, I've added an interest that I want to follow the FIFA World Cup at this time.
And the model would get both my user history and this new input and then be able to decode recommendations that are personalized to me like this free kick that Messi scored recently. I think these interactive recommenders are going to be a really interesting new product surface and we're seeing this across the industry. We have some examples of prompted playlists from Spotify, custom feeds from YouTube, Ask DoorDash, and we'll be hearing more from speakers about these. Finally, I want to talk about the framework of tokens in engagement out. This flywheel that I started with of model training, inference, consumer engagement, and then monetization.
This is actually the same flywheel that's shared by content feeds and the AI chat apps. And this is a lens that you can use to evaluate any consumer app. What will make an app successful on the training ROI side is how well can it translate compute into a frontier model. On the inference side, how well can the inference tokens translate into engagement and then monetization. And so if you try to compare content feeds and AI chat apps, I think that LLM recommenders are actually structurally more token efficient than the AI chat. So on the left you have content feeds like Instagram, Facebook, TikTok, and YouTube.
On the right you have the big AI chat apps like Gemini, ChatGPT, and Claude. Content feeds are currently using models that are around 1 to 10 billion active parameters. The chat apps are serving much larger models, 10 to 100 billion active parameters. The output for the content feed is a semantic ID token, which is a pointer or an address to existing content because the content supplied on these content feeds is uploaded by creators. There's a very healthy creator economy. And so the supply of content is effectively free or it's uploaded by creators. For the AI chat apps, they have to decode every token of content themselves.
And the amount of tokens output in every turn of an LLM chat interaction is a few thousand tokens. And the big difference here is that every token has to be manufactured by the app at inference time. And so what that means is if you compare these two apps on how much inference cost and compute is spent to generate an hour of consumer engagement, there's a huge structural gap where content feeds are significantly cheaper, up to 100 times or more cheaper than AI chat apps, because they're decoding pointers to content rather than content itself. Finally, I want to end with why I think LLM-RECs is one of the most significant consumer AI applications.
If you look at the top apps by daily active users, these are the top 10 apps, and 4 out of 10 of them are content feeds. And so this is a really significant consumer application. If you look at the content feeds, almost all of the consumer app growth on both the engagement and monetization side is driven by the recommender and ads models. And these are going to be entirely transformed by LLM recommenders. It's a very large and very token efficient application of AI for consumer apps.
We're going to see a lot of new product experiences come through with steerable and interactive recommendations, explanation of recommendations, and putting more users in control of their experience on these apps. I think we're going to see some really exciting research on consumer agents and recommendation agents that come out over the next year or so. And so this is why I think this is a super exciting area of both research and product at this intersection of LLM and recommendations. That's all. Thank you so much. And I think this foundation model can also start to explain its recommendations.
And so for this example, I've added an interest that I want to follow the FIFA World Cup at this time. And the model would get both my user history and this new input and then be able to decode recommendations that are personalized to me like this free kick that Messi scored recently.
I think these interactive recommenders are going to be a really interesting new product surface and we're seeing this across the industry. We have some examples of prompted playlists from Spotify, custom feeds from YouTube, Ask DoorDash, and we'll be hearing more from speakers about these.
Finally, I want to talk about the framework of tokens in engagement out. This flywheel that I started with of model training, inference, consumer engagement, and then monetization. This is actually the same flywheel that's shared by content feeds and the AI chat apps. And this is a lens that you can use to evaluate any consumer app. What will make an app successful on the training ROI side is how well can it translate compute into a frontier model. On the inference side, how well can the inference tokens translate into engagement and then monetization.
And so if you try to compare content feeds and AI chat apps, I think that LLM recommenders are actually structurally more token efficient than the AI chat. So on the left you have content feeds like Instagram, Facebook, TikTok, and YouTube. On the right you have the big AI chat apps like Gemini, ChatGPT, and Claude. Content feeds are currently using models that are around 1 to 10 billion active parameters. The chat apps are serving much larger models, 10 to 100 billion active parameters. The output for the content feed is a semantic ID token, which is a pointer or an address to existing content because the content supplied on these content feeds is uploaded by creators.
There's a very healthy creator economy. And so the supply of content is effectively free or it's uploaded by creators. For the AI chat apps, they have to decode every token of content themselves. And the amount of tokens output in every turn of an LLM chat interaction is a few thousand tokens. And the big difference here is that every token has to be manufactured by the app at inference time. And so what that means is if you compare these two apps on how much inference cost and compute is spent to generate an hour of consumer engagement, there's a huge structural gap where content feeds are significantly cheaper, up to 100 times or more cheaper than AI chat apps,
because they're decoding pointers to content rather than content itself.
Finally, I want to kind of end with why I think LLM-REXIS is one of the most significant consumer AI applications. If you look at the top apps by daily active users, these are the top 10 apps, 4 out of 10 of them are content feeds. And so this is a really significant consumer application. If you look at the content feeds, almost all of the consumer app growth on both the engagement and monetization side is driven by the recommender and ads models. And these are going to be entirely transformed by LLM recommenders. It's a very large and very token efficient application of AI for consumer apps.
We're going to see a lot of new product experiences come through with steerable and interactive recommendations, explanation of recommendations, and just putting more users in control of their experience on these apps. I think we're going to see some really exciting research on consumer agents and recommendation agents that come out over the next year or so. And so this is why I think this is a super exciting area of both research and product at this intersection of LLM and recommendations. That's all. Thank you so much. Thank you very much.