Open Reader

Personalization in the Era of LLMs - Shivam Verma, Spotify

completed 20:11 May 19, 2026 Watch on YouTube

Current Status

completed

Video ID

5YSJEP0HWzM

RAG / Chat

Enabled
Personalization in the Era of LLMs - Shivam Verma, Spotify
Description

Spotify represents Ariana Grande and Bruno Mars as sequences of six tokens. The first two are shared because both are pop artists. The remaining tokens diverge to capture what makes each distinct. That is a Semantic ID, and it is how Spotify teaches open-weight LLMs to reason over a catalog of 100 million tracks the same way they reason over words. Shivam Verma from Spotify's AI foundation team walks through the three components they assembled to personalize LLMs at scale without full fine-tuning. User embeddings trained on streaming history across 750 million users form the base. Semantic IDs compress catalog vectors into tokens the model can autoregressively generate, predicting the next song or episode as the next token in a sequence. A soft tokenization layer projects a user's embedding directly into the LLM's token space, giving the frozen model a user-specific token to attend over. Podcast next-episode recommendations are already running on this stack in production. Speaker info: - https://x.com/kaffeinated - https://www.linkedin.com/in/shivam13verma

Summary

Generated by claude-haiku-4-5-20251001

Personalization in the Era of LLMs - Spotify

Main Topics

  • Foundational User Modeling: Building vector representations of users based on their interaction history
  • Catalog Understanding: Teaching LLMs about Spotify's content through semantic IDs and fine-tuning
  • Unified Generative Recommendation System: Combining user and content representations with LLMs for personalized, steerable recommendations
  • Product Applications: Taste Profile, AI DJ, Prompted Playlists, and Next Episode features

Key Points

User Modeling

  • Shift from generalized to foundation models: Moving from autoencoders to transformer-based sequential models for user embeddings
  • Scale: Generate embeddings daily for 1B+ users across Spotify's platform
  • Cross-content embedding space: Users, tracks, and podcast episodes are embedded in the same vector space, enabling discovery and exploration
  • Context as foundation: User interaction history serves as the prompt/context for downstream models

Catalog Understanding

  • Dual knowledge integration: Combines Spotify-specific knowledge (vectors) with world knowledge (from open-weight LLMs like Llama, Qwen)
  • Semantic IDs: Compress high-dimensional content vectors (e.g., 1000-dim) into 4-6 tokens using hierarchical tokenization
  • Domain adaptation: Fine-tune LLMs with Spotify data to teach them to reason about recommendations using semantic IDs
  • Hierarchical structure: Semantic tokens reflect music taxonomy (e.g., shared tokens for pop artists, differentiated tokens for niche preferences)

Unified Recommendation System

  • Soft tokenization: Project user embeddings into the LLM's token space, creating user-specific soft tokens that personalize generation
  • Generative paradigm: Shift from traditional multi-stage pipelines (candidate generation → ranking) to a single autoregressive model
  • Steerability: Users can provide natural language feedback to adjust recommendations via Taste Profile and similar features
  • Explainability: LLM-based systems provide better interpretability compared to traditional recommendation systems

Products & Features

  • Taste Profile: Exposes user taste to users, allows edits to refine recommendations
  • Prompted Playlists: Generate custom playlists from natural language prompts (now supports podcasts)
  • AI DJ: Conversational interface for recommendations
  • Next Episode: Personalized podcast episode recommendations

Notable Quotes

> "This is less about context engineering from the conventional agentic sense. It's going to be more about how we do context engineering on the modeling side."

> "As durable as possible. And I'm going to talk a bit more about that. So generally, you can think about it as going from sequences of actions to vectors. And then once you have the vectors, you can go to tokens."

> "These models are really smart. And the moment you give them information about the user, they can learn to put everything together in a single space, which is really cool."

> "It gives you steerability. It gives you better recommendations. It gives you explainability. So there's a lot of stuff that you get for free when you use language models for recommendations."

Takeaways

  • User representations are foundational: Deep understanding of user taste through transformer-based embeddings enables all downstream personalization
  • Semantic IDs bridge embeddings and LLMs: Compressing high-dimensional vectors into discrete tokens enables LLMs to reason about recommendations autoregressively
  • User control is essential: Products like Taste Profile and Prompted Playlists give users agency, moving beyond opaque recommendation algorithms
  • Unified models beat silos: Moving from product-specific models to a single LLM backbone improves consistency and enables new capabilities
  • Hybrid knowledge works: Combining platform-specific knowledge (collaborative filtering) with LLM world knowledge creates superior recommendations
  • Production-ready today: These techniques are already powering Spotify features like Next Episode, demonstrating real-world viability

Transcript

3398 words en Processed in 240.5s

[SPEAKER_00] Hi, everyone. I am Shivam. I'm from Spotify. And my talk is going to be about how at Spotify we do personalization, especially in the era of LLMs. For those of you who use Spotify, I guess, do we have any Spotify users in the room? Raise your hands. Nice. Yeah, thanks for using Spotify. And a bit like in this talk, this is going to be less about context engineering from the conventional agentic sense. It's going to be more about how we do context engineering on the modeling side. So if you're interested at all in how your Spotify app works, how we recommend you songs, tracks, episodes, et cetera, this talk is going to be really useful for you to contextualize how we use your data for just recommending stuff to you that you like. So a bit about me. I am the tech lead of the user representations team in Spotify's AI Foundation org. So the AI Foundation team builds all of the frontier foundational models that are used across the entire stack of recommendations at Spotify. So we do things like user representations, content representations, as well as adapting open-weight LLMs. We do CPT, SFT, all of the stuff that a lot of the frontier labs are doing, as well as other stuff our competitors are doing. And we try to make sure that we're building the best music recommendation system possible for you guys. So my background is as a machine learning engineer. I used to work at Twitter. And I live in London. So if you're around for a coffee chat, I'm very happy to, after this talk or in general, talk about this stuff because it's really cool. So three things I'm going to talk about today. The first thing is going to be about foundational user modeling. So what is that? That's essentially us trying to understand you, the users. These are the three key aspects that we think are very key aspects of the future of personalization in the era of LLMs. So there's the user modeling component. Then there's the content side. So how can you teach LLMs about the content that you have, the catalog that you have on Spotify or whatever your platform is. And then the last thing, once you have these two pieces of the puzzle, are building these both of these together into something which is durable and personalized. So as durable as possible. And I'm going to talk a bit more about that. So generally, you can think about it as going from sequences of actions to vectors. And then once you have the vectors, you can go to tokens. And then once you have the tokens, you can actually combine the vectors and the tokens with the LLM that you have to get a pipeline where you have something which is as personalized as possible. So a bit about Spotify. For those of you who haven't used it, we have about 750 million users right now. We have a catalog of about 100 million plus tracks. We have, I think, about 350. I think it's more like 400,000 audiobooks now. Millions of podcasts and a lot of video episodes as well. So more and more creators are switching to video as a modality. And we're definitely making sure that we support that. And we're in about 184 markets. So as you can see, we have a lot of users. We have a lot of data and content. How can we combine all of that to build something which is as useful for our users as possible? The way we do that is we obviously have been using machine learning models for at least a decade, if not more. You might have heard of Discover Weekly or you're probably a user of that. That's been around since, I think, 2015, back when I was in grad school. And that was one of the coolest products at that time with my interactions across the tech stack that I was using at the time. Just because it's something which is unique to you. It keeps changing. And over the last decade or so, we've added more and more personalization and products to the app. So we have a bunch of verticals. We have a bunch of new product surfaces. We have this thing called the AIDJ where you can talk to it and it kind of recommends you stuff or it plays stuff for you. We also have a prompted playlist where you can actually prompt the model, Spotify's model, and it kind of generates a custom playlist for you based on your prompts. And as of this week, it also supports podcasts. So if you want, you can just prompt it and it'll create a playlist of episodes for you based on whatever you're looking for. So that is the future that we're heading towards where users have steerability. Users can talk to Spotify in natural language. And we also have something called the Taste Profile. So this is only supported in a few markets, but this is going to be expanded later this year. And the idea is that we want to expose what we know about you. And then we want to let you choose which part of that you want us to keep, which part of that you want us to forget, and just allow you to have as much control as possible. As we work on this, just some context for those of you who haven't worked in the space of recommended systems and machine learning. So what we call TradRex, which used to be the predominant paradigm of building recommended systems up until a few years ago, is essentially a multi-step pipeline where you have a massive catalog of items. You have this candidate generation step, which kind of reduces that item space from millions to a few hundred. And then you have a ranking stage and sometimes you have multiple rankers, which essentially bring that further down and then give you the final list of whatever your top songs that we want to recommend to you. And we use this across different products. So we have home shelf ranking and we have personalized playlists and search and podcasts and ads and a lot of other stuff. And this is generally every team, every product has its own team that has its own model. You have this candidate generation step, which reduces that item space from millions to a few hundred. And then you have a ranking stage and sometimes you have multiple rankers, which essentially bring that further down and then give you the final list of whatever your top songs that we want to recommend to you. And we use this across different products. So we have home shelf ranking and we have personalized playlists and search and podcasts and ads and a lot of other stuff. And this is generally every team, every product has its own team that has its own model. So it's spread across different groups of people. And some models are better than others. Some have different features. So we're moving away from that siloed model of working towards this single unified model, which supports similar to how LLMs work, which supports an LLM backbone, and which allows you to steer it towards the sort of recommendations that you want. And one of the key components of that is the user modeling part. So that is the team that I work with. And what we do is we build user embeddings, which are essentially representations of vectors that tell Spotify about the user's taste across all of the history that we have on you, the user, across the different sessions that you've had with us over the years. And that becomes the foundation of all of the models that are downstream that actually recommend stuff or allow you to search for stuff. And these models are very complex. So the embedding model is we generate embeddings for a billion-plus users because we have a lot of users in general that are also MAUs or that are not MAUs. So we do that every day. So it's a massive pipeline. It's very expensive. And over the years, we've moved away from having these sort of generalized user representations, which was the predominant paradigm in machine learning, where you had these models. In this case, this was a paper from our team last year where we publicly spoke about the user embedding model that we have, which is generally an autoencoder model. If you're familiar with that, what that does is it takes all of your features, it compresses it down to a small vector, and then recreates your features from that. And that sort of compression, decompression process allows the model to learn about you, the user, and just represent user interactions in the form of a vector. So this is fairly standard stuff that a lot of folks do in the NLP computer vision space as well that is also aligned with the way things were in recommendations. We're moving away from that towards the foundation modeling side. So now we have this single sequential model, which, as you can see, with the whole industry moving towards transformers and the whole shift that's happening across not just the regular tech industry, but also the companies that use recommended systems and build recommended systems as their main bread and butter. So we're also a part of that, and we're also moving towards using transformers for that kind of stuff. And the idea is that you want to have the users' interactions as a part of the prompt. So it's the context. When we talk about context engineering, this is the context that we're talking about. And then there's obviously the request context. There is the query. There's the product surface, et cetera, all of that stuff. And then there is the item that you're recommending. So when you add all of this stuff and then you put a transformer layer and a bunch of heads and all that stuff, and you train it over millions or hundreds of millions of users' data, what you get is something really cool. What you get is something like this. So this is an image which is from one of our newer models, which essentially shows you, this is a compressed version of what the model is learning. In this case, we have tracks in the blue and we have episodes, which are podcast episodes in pink. And then users, so that the one in green is me. And some of the other green ones are other folks in my team. So that shows you that we're doing this cross-content modeling. We're embedding users, tracks, and episodes in the same sort of content space, or I guess embedding space. And you're able to visualize, from this image, how on the hypersphere where you live alongside different pieces of content. And what is close to you, what is not close to you, and how can you explore that neighborhood region, and where you live, contextualized by where your friends live and things like that. So this is a visualization of what the model is learning. And as you can see, for me, I'm a machine learning engineer. I care a lot about keeping up with what Anthropic is doing and what's happening in the tech industry and all that stuff. So my specific embedding is really close to this big tech podcast. And on the right, you can see that the point where you see those lines spreading, that is me. And then the ones in pink are streams for tracks. And the blue ones are streams for episodes. Or sorry, it's the reverse. And you can visualize, you can contextualize whatever you're listening, whatever you're not listening to, and how does the embedding space look like for users. So these models are really smart. And the moment you give them information about the user, they can learn to put everything together in a single space, which is really cool. This next part is more about catalog understanding. So now that we have the users, we understand the users, we have a model for them, how do we understand the catalog? Right? So catalog understanding, generally, there's a number of ways that you can understand the catalog. The most common ways you train, similar to the user side, you train a vector to understand the content. So you have a vector that represents the item, whether that's a song, or it's an artist, or it's a podcast, or an episode. And you have that vector. And then alongside that, you have the user vector. So that is the Spotify knowledge. So that's what we know about the content or the different entities that we're dealing with. And then you have the world knowledge. So that's coming not from Spotify, but it's coming from these open-weight LLMs that we're working with. So models like Llama or Quen or other sort of open source models. What we do is we fine tune those models. So you have a vector that represents the item, whether that's a song, or it's an artist, or it's a podcast, or an episode. And you have that vector. And then alongside that, you have the user vector. So that is the Spotify knowledge. So that's what we know about the content or the different entities that we're dealing with. And then you have the world knowledge. So that's coming not from Spotify, but it's coming from these open-weight LLMs that we're working with. So models like Llama or Qwen or other open source models. What we do is we fine tune those models. And then we embed Spotify's knowledge into those models through something that I'm going to talk about. And what that does is it gives you steerability. It gives you better recommendations. It gives you explainability. So there's a lot of stuff that you get for free when you use language models for recommendations. Now, there are trade-offs here. So the model does end up forgetting stuff. Catastrophic forgetting is an issue. But generally, from what we've seen, these models are really good at combining world knowledge with whatever knowledge that you have from your platform and building something which you can holistically use for recommendations. Now, the stuff that I was talking about where on the previous slide, how do we actually teach these LLMs about the content? Let's start with that. So the way that we do that is using something called semantic IDs. Semantic IDs is a fairly new concept. I think there was a paper from Google a few years ago which introduced this concept in the context of YouTube. And what it does is when you have a vector that represents a piece of content, so that's a track or an episode in our case, what we do is we tokenize it, similar to how LLMs tokenize words. And what that does is it compresses that massive, let's say a thousand-dimensional vector into four or six tokens. And that allows us to really use those tokens to train the LLM in the way that LLMs are usually trained. And it allows the LLM to autoregressively generate the next token. In this case, the next token is not a word, but it's the next song or it's the next episode that you're going to be listening to. So that's what we're doing. We're post-training, or we're continually training these LLMs with Spotify's data that we have about the catalog. We use semantic IDs to compress the vectors into semantic IDs. And at the bottom you can see that we have examples of Ariana Grande and Bruno Mars. So we represent them as six tokens. So those numbers are actually token IDs. And the first two tokens for both of them are shared because they're both pop artists and they both share something between them. But then the other tokens are different because those tokens represent more niches. So it's a hierarchical structure where you're compressing the embedding into these six tokens and there's a hierarchy to it. And that allows the model to autoregressively generate the next artist or the next song that you're going to be listening to. So this emphasizes how we do this. We use the user context. In this case we have a user who's Italian. Their listening history, which is tokenized. We send that listening history. We use that in the training data. So we teach the LLM how to talk with semantic IDs. And that is the domain adaptation that I was referring to earlier in the slides. And then the final output is essentially generating the next item. So whether that's going to be an episode or it's going to be a track. What is this person listening to? This is one example, taking the example of the Italian person who maybe listens to an episode, an Italian podcast. This is an example of a prompt that we give that model. And you can see the prompt on the left has the Spotify URI, which is our representation of the item. In this case, the episode. We convert that into a semantic ID, which is the actual tokens that the model is attending to. And that is used to finally predict what is the next episode that the user is going to be listening to. So this is the catalog understanding part of it. So now that we have the user modeling part and the catalog understanding part. The next step is to essentially assemble all of these components to form a single, steerable, personalized generative recommended system. So this is moving away from the traditional recommended system model to this generative model. This is the product that I was talking about that we launched just a few weeks ago. This is called the Taste Profile. The idea is that you have some piece of text that represents who the user is. We expose it to you, the user. And then you're allowed to tell Spotify by chatting or by sending messages or adding some text. So maybe you want to start listening to Justin Bieber more or you don't like this specific podcast that's being recommended to you. And what this does is it allows the model to take that data, that edit, which is going to come back into the generative model. And it's going to upgrade the model's ability to understand you and recommend stuff that's better for you. So essentially we have the content piece of the puzzle, but we don't have the user piece yet. And the user piece doesn't come because ultimately these models are trained on a limited amount of training data. You cannot train them on every 750 million plus users that we have. So there is going to be some level of collaborative filtering. So the model is going to generalize, hopefully, but it also needs to be personalized. The way that we do that is coming back to the idea of user models. So you have the LLM. You have the user representation. What you do is you project the user representation into the space of the LLM. And what that does is it allows you to create what's called a soft token within the model. And it's essentially a token that represents the user, which is obviously contextually changed depending on the user that we're generating this response for. And that allows the model to be personalized. So that is the final piece of the puzzle where you have this vector projection, which is you have the regular LLM. The way that we do that is, again, coming back to the idea of user models. So you have the LLM. You have the user representation. What you do is you project the user representation into the space of the LLM. And what that does is it allows you to create what's called a soft token within the model. And it's essentially a token that represents the user, which is obviously contextually changed depending on the user that we're generating this response for. And that allows the model to be personalized. So that is the final piece of the puzzle where you have this vector projection, which is you have the regular LLM. And then you have a user vector which is projected to the space of the LLM. And that gets inserted into the prompt. And then finally when the model is actually generating a recommendation, the model is personalized because it has the context on you, whoever we're generating the recommendation for. These are some early results that we've had. We've seen some pretty positive results on our internal metrics. If you use Spotify, if you use the next episode feature, if you listen to podcasts on Spotify, this is something which is actually productionized now. So if you're getting a recommendation, it's coming from something like this. And that's the final piece. So we have the embeddings which represent the users. We have the semantic IDs which represent a compressed version of the content. And we have this soft tokenization approach which allows you to project users into the token space of the model. And this moves away from the traditional recommendation system model and it moves towards this sequential modeling framework, which is something that we're very excited about. [SPEAKER_00] And that's definitely going to be really exciting as we go forward and we build this more and more into all of the recommendation systems that we have. [SPEAKER_00] And we're very excited to put this out there pretty soon. [SPEAKER_00] And that's my time. [SPEAKER_00] I'm really happy to connect or if you have any questions feel free to reach out after the talk.