Open Reader

Teaching LLMs to Speak Spotify — Yves Raimond & Jacqueline Wood, Spotify

completed 19:40 Sep 25, 2026 Watch on YouTube

Current Status

completed

Video ID

2LRIAfng7eA

RAG / Chat

Enabled
Teaching LLMs to Speak Spotify — Yves Raimond & Jacqueline Wood, Spotify
Description

One in four US Spotify Premium subscribers now use a single system every day, which Spotify calls the Large Taste Model. Yves Raimond, Spotify's SVP and GM of AI & Personalization, explains how Spotify turned its recommendation system "on its head" to be LLM-native. It went from curated playlists, to recommendations like Discover Weekly, to what Spotify calls generative personalization: a steerable DJ, prompted playlists, an editable taste profile and personal podcasts, across a catalog of more than 100 million tracks plus podcasts and audiobooks. Staff ML Engineer Jacqueline Wood then shows how the models are trained with NEO, Spotify's four-stage recipe. Semantic IDs are added to an open-weight LLM like Qwen. The new tokens are grounded while the backbone stays frozen, which keeps its language ability intact, where continued pre-training wiped it out. Then the model is instruction-tuned across many Spotify tasks, which even helped cold-start audiobook recommendations. She also covers decoding choices (98% of semantic IDs come out valid even without constrained decoding), and how grounding LLM judges with user profiles and behavior raised their agreement with humans by 91% on ambiguous queries. Speaker info: Jacqueline Wood, LinkedIn: https://www.linkedin.com/in/jacquelinewood Related links: Spotify: https://www.spotify.com Timestamps: 0:00 Intro: making LLMs speak Spotify 0:47 Spotify's scale: 760M users and 100M+ tracks 1:52 From curation to recommendations 2:37 Generative personalization 3:02 From guessing to reasoning, from black box to steerable 4:07 A Spotify DJ you can steer 4:37 Prompted playlists 5:32 Taste profile 6:22 Personal podcasts 6:52 The Large Taste Model 7:27 One in four US Premium subscribers use it daily 8:03 Jacqueline Wood: how the models are trained 8:18 Semantic IDs for Spotify's catalog 9:03 Natural-language podcast recommendations 9:33 NEO: a four-stage training recipe 10:02 Domain grounding with a frozen backbone 10:52 Capability ind

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Spotify turns an open-weight LLM into a real-time, language-steerable recommendation engine by representing catalog items as semantic-ID tokens, grounding those tokens without overwriting the base model, then instruction-tuning the model across retrieval and recommendation tasks.
  • Why it matters: This is a production-scale blueprint for replacing isolated rankers and tool-mediated recommendation flows with a unified model that can understand user history, retrieve grounded entities, explain choices, and accept natural-language steering.
  • Best use: Use it as an architecture and evaluation reference for agentic personalization or any domain-specific LLM that must reason over a large private catalog while preserving general language capabilities.

Executive Summary

Spotify frames its shift as a move from conventional recommendation—ranking a fixed set of catalog entities—to "generative personalization": an interactive system that can reason about a request and context, create an experience, explain its choices, and let users actively steer it. Its examples include DJ session steering, prompted playlists, editable taste profiles, and an upcoming personalized daily podcast brief.

The central production system is the "large taste model," used daily by roughly one in four U.S. Premium subscribers. It combines a user’s historical interactions with catalog knowledge and natural-language instructions. Spotify reports gains after deploying it to existing surfaces including autoplay, podcast discovery, and DJ-message interaction, although the talk does not disclose absolute lift values or methodology.

The technical core is Spotify's NEO training recipe. Existing content embeddings are quantized into discrete semantic-ID tokens, which are added to an open-weight LLM vocabulary. Spotify then freezes the LLM backbone while learning alignment between text and semantic IDs, before unfreezing the model for multi-task instruction tuning on tasks such as next-item recommendation and retrieval. This staged approach is intended to retain the base model's language and world knowledge while teaching it catalog-specific grounding.

Spotify's ablations offer practical operating guidance: multi-task training matched or exceeded single-task systems and particularly helped newer catalog types such as audiobooks; continuous pretraining damaged base language/world knowledge; and beam search beat top-p sampling for recommendation accuracy. For evaluation, the speakers argue that interaction metrics are insufficient for generative recommendations and that LLM judges require grounding in user profiles or behavioral data to align reliably with human judgment.

Key Takeaways

  • Claim: A recommendation system becomes materially more useful when users can express intent and correct the system in natural language rather than only accepting or rejecting ranked outputs. | Evidence: Spotify's DJ can be steered mid-session; prompted playlists accept requests ranging from "bands playing in San Francisco tonight" to a phased running playlist; taste profiles let users correct inferred preferences, such as excluding Disney music played by a speaker's children on shared devices. | Implication: For user-facing agents, expose a conversational control layer and durable preference-correction mechanism rather than treating personalization as a hidden ranking function. | Caveat: Steerability depends on the system having an accurate, editable representation of user history and context; a shared-device example shows that raw behavioral logs can encode the wrong user intent.
  • Claim: Semantic IDs let one LLM work directly with both natural language and a proprietary, large-scale content catalog. | Evidence: Spotify converts existing content embeddings, such as podcast-episode embeddings, into discrete tokens through quantization; the model receives listening history as semantic IDs and can return a relevant episode ID plus a natural-language explanation for a request such as a podcast about morality. | Implication: A domain LLM does not have to rely solely on external retrieval tools: structured entities can be represented in the model's token space when low-latency grounded generation is valuable. | Caveat: The approach presupposes useful underlying content embeddings and a catalog representation that can be quantized into meaningful discrete codes.
  • Claim: Spotify's four-stage NEO training recipe is designed to add domain grounding without catastrophic forgetting of the pretrained LLM's general language ability. | Evidence: NEO consists of semantic foundation, frozen-backbone domain grounding through bidirectional text/semantic-ID mapping, full-model multi-task capability induction, and optional post-training such as RL fine-tuning. During domain grounding, Spotify trains only newly added semantic-ID embeddings while freezing original weights and embeddings. | Implication: When adapting an open-weight model to internal entities, separate vocabulary/entity alignment from broad task fine-tuning; do not immediately continuously pretrain the full base model on proprietary domain data. | Caveat: The presentation reports internal ablations rather than enough detail to independently quantify the trade-offs across data sizes, model scales, or catalog domains.
  • Claim: Multi-task tuning across recommendation and retrieval objectives can improve transfer, especially for cold-start or newer content types. | Evidence: Spotify found its multi-task model consistently matched or beat single-task variants; audiobook recommendation showed the clearest benefit, with the explanation that the newer audiobook domain could learn from patterns in more mature content types such as podcasts. | Implication: Train a shared domain model across adjacent tasks and entity types where behavior and semantics transfer, rather than building a separate model for every surface or catalog vertical.
  • Claim: Preserving the pretrained backbone matters more than optimizing only task-specific metrics during domain adaptation. | Evidence: Dropping or combining the frozen domain-grounding stage degraded performance, while initializing from a random backbone caused the largest drop. Continuous pretraining produced only minimal task-specific degradation but reduced the base model's natural-language and world-knowledge capability to essentially zero; the findings were validated on Qwen and LLaMA backbones. | Implication: Maintain explicit regression tests for general capabilities during specialization. A domain system that scores well offline may still fail as an interactive assistant if it loses language comprehension and explanation quality. | Caveat: The phrase "essentially zero" refers to Spotify's undisclosed natural-language/world-knowledge evaluation, not a universal claim about all continual-pretraining setups.
  • Claim: For generating catalog IDs, decoding policy is a product-quality decision: Spotify accepted somewhat higher latency from beam search to gain accuracy. | Evidence: Unconstrained beam search generated valid semantic IDs 98% of the time. Constrained decoding adds latency but is useful when outputs must meet rules such as recommending only new content. Top-p sampling significantly reduced accuracy relative to beam search. | Implication: For agent systems that generate structured actions or entity references, benchmark deterministic or constrained decoding against sampling; reserve probabilistic sampling for cases where diversity outweighs action accuracy. | Caveat: The reported 98% validity concerns syntactically valid semantic IDs, not necessarily relevance or user satisfaction.
  • Claim: LLM-as-judge evaluation must be grounded in behavioral and profile data rather than treated as an inherently reliable evaluator. | Evidence: A judge supplied with textual listening-history profiles reached 75% alignment with human preferences. Adding similar-query behavioral signals increased overall search-judge alignment by 5% and ambiguous-query alignment by 91%. A heavily grounded judge reached 0.87 agreement with human system rankings for scaling Cranfield-style evaluation collections. | Implication: Build evaluation judges around the evidence a human evaluator would need—user state, prior behavior, task context, and candidate set—and separately validate them on ambiguous cases. | Caveat: The gains are task-specific and alignment with human raters does not replace monitoring real user outcomes or checking for systematic judge bias.

Detailed Brief

Product and system scope at Spotify

  • Claims: Spotify describes a progression from manual curation, to recommendation at scale, to generative personalization that can create and explain dynamically shaped experiences.; The large taste model serves multiple product experiences rather than being isolated to a single conversational interface.; The model is positioned as providing low-latency, tool-free inference while handling search, recommendation, and explanation.
  • Evidence: Spotify operates across roughly 760 million monthly active users in 184 markets and a catalog exceeding 100 million music tracks, plus video, podcasts, and audiobooks.; Spotify still has about 10 billion playlists, with many more created each hour; Discover Weekly, launched in 2014, is presented as an earlier recommendation-era product.; The personal-podcast concept extends generative personalization beyond selecting content to generating a daily community-news brief.
  • Caveats: Claims that NEO is the first system to combine grounded catalog items, language steerability, search, recommendation, explanations, and industrial-scale tool-free latency are Spotify's own positioning against related work including semantic-ID retrieval, tool-based recommenders, and PLUM.; The transcript provides no system architecture details on serving infrastructure, model size, latency budget, safety controls, or how generated explanations are verified.
  • Implications: The highest-leverage design is a reusable domain model and shared catalog representation that can power many experiences, while each surface supplies its own interaction design and constraints.; Generated content is a natural extension of personalized retrieval, but it introduces a separate content-quality and provenance problem beyond recommendation relevance.

Evaluation operations beyond engagement metrics

  • Claims: Traditional offline recommendation metrics can reveal whether a user interacted with content, but not whether a recommendation fit the user's intent or whether its explanation was accurate.; Grounded LLM judges can reduce the human-labeling bottleneck in classical pooled-ranking evaluation datasets.
  • Evidence: Spotify describes Cranfield-style collections as pooling candidates from several sources and having humans rank the pool, a costly ranking stage it seeks to scale with a grounded judge.; For podcast discovery, Spotify reports that this model family helped move users beyond habitual listening patterns toward unfamiliar content and produced large online wins, without publishing the size of those wins.
  • Caveats: Breaking habitual consumption patterns can improve discovery but may conflict with short-term engagement, user trust, or diversity goals unless the target objective is explicitly defined.; The talk does not explain the human-rater protocol, the population sampled, or whether judge alignment remains stable across markets and content categories.
  • Implications: Evaluation should distinguish structural validity, relevance, intent satisfaction, explanation faithfulness, novelty, and longer-run user value instead of collapsing them into click or consumption metrics.; Use human judgments strategically to calibrate and audit an LLM judge, then use the judge to expand test coverage rather than eliminate human evaluation.

Notable Concepts & Terms

  • Generative personalization: Spotify's term for personalization that reasons over user intent and context, accepts interactive steering, generates experiences, and can explain selections rather than only emitting a ranked list.
  • Large taste model: Spotify's internal shared model that combines catalog knowledge, historical user interactions, reasoning, and real-time natural-language control across recommendation surfaces.
  • Semantic IDs: Discrete tokens produced by quantizing content embeddings, allowing an LLM to represent and generate catalog entities in the same token-level framework as natural language.
  • NEO: Spotify's four-stage adaptation recipe: semantic foundation, frozen-backbone domain grounding, multi-task capability induction, and optional post-training.
  • Domain grounding: The stage that learns text-to-semantic-ID and semantic-ID-to-text mappings while the pretrained backbone is frozen, intended to preserve general language competence.
  • Capability induction: Full-model or LoRA multi-task instruction tuning on Spotify tasks such as next-item recommendation and retrieval after entity tokens have been grounded.
  • Constrained decoding: Inference-time restrictions that guarantee certain output properties, such as recommending only a specific class of catalog item, at some latency cost.
  • Grounded LLM judge: An LLM evaluator supplied with task-relevant evidence such as user profiles or behavioral signals, making its preference judgments more aligned with human ratings.

Operator Notes / Why Ken Should Care

  • Prototype a semantic-entity token layer only if the product needs a single low-latency model to map natural language directly into a stable internal entity/action space; otherwise compare it against conventional retrieval-plus-tool calling.
  • Adopt a staged adaptation gate: freeze the base model during entity vocabulary alignment, then permit full or LoRA tuning only after running regression tests for general language competence.
  • Create a decoding benchmark for any generated structured output: measure valid-action rate, relevance, latency, and constraint compliance for beam, constrained beam, and sampling strategies.
  • Build an eval set concentrated on ambiguous user requests and shared-account/noisy-history scenarios; use these as release blockers for personalized agents.
  • Require generated rationales to be evaluated for faithfulness separately from recommendation acceptance; persuasive but ungrounded explanations are a material trust risk.
  • Investigate Spotify's published NEO and podcast-discovery papers for missing implementation details before treating this presentation as a directly reproducible design.

Source/Metadata

  • Title: Teaching LLMs to Speak Spotify — Yves Raimond & Jacqueline Wood, Spotify
  • Transcript words: 4214
  • Duration seconds: 1180
  • Timestamp note: No timestamps or chapters were provided. The latter portion of the transcript repeats the technical presentation, so the unique substantive content is shorter than the stated word count.

Transcript

2665 words en Processed in 101.4s

Welcome everyone. I'm really glad and thank you Devenge for inviting us to talk today. So Jackie and I are going to present today about making LLMs speak Spotify and how we turned our recommendation system on its head to be LLM native. Just a few quick words about myself. I joined Spotify a year ago. Before that I was at Google working on personalization for Google Search and prior to that I was working on personalization at Netflix. All right, so let's dive in. First, a couple of numbers to describe the scale of the problem that we have to solve. Spotify has about 760 million active users monthly in about 184 markets. But one thing that makes the Spotify personalization problem particularly challenging is the size of its catalog. Spotify has basically all music ever published, a bit more than 100 million music tracks, but it also has a range of videos, podcasts, and audiobooks as well. So the matching problem is actually surprisingly complicated. So today what we're going to talk about, Jackie and I, is the extent of this matching problem and how we are solving it. I'm going to talk about a little bit of history, the new phase that we're entering, and then Jackie is going to help us go into the guts of these models to understand how those things are trained as well. So a little bit on history. Initially, Spotify personalization was really built around curation. People basically manually assembled playlists that target specific tastes, and that's still a very common use case on Spotify right now. We have about 10 billion playlists with many created every hour. Then Spotify moved into taking these curation signals and other signals and moving into recommendations, so basically being able to turn these curation signals into something that can be applied at scale. A great example of that would be Discover Weekly, for example, which was launched in 2014 as one of the early recommendation use cases on Spotify. But the phase that we're entering now, which we're going to talk about in more detail, is something that we call generative personalization where we are not only solving a matching problem from the user to the content, we are also solving the ability to generate an experience that is interactively and dynamically shaped around each user. And that transition from recommendations to generative personalization involves a couple of different big shifts. One is moving from personalization as guessing, where basically you have a ranking algorithm that spits out a rank list of entities in your catalog, to personalization as reasoning that can introspect these results and really try to understand whether that's indeed the right match for this user in this particular context. The other aspect as well is moving from black box algorithms—it's very difficult to fully introspect a multi-stage ranking system, for example—to transparent and steerable personalization where the user is always fully in control. So basically giving the ability to these models to speak and understand English. The other thing that these systems can do is not stopping just at recommending but also generation of experiences and explaining as well. And we're going to show a couple of examples of that. So one example that we launched a couple of years ago, so pretty early in that journey, was the Spotify DJ. The Spotify DJ is something that you can spin up that will start playing music for you, of course, personalized. And one interesting thing about it is that since last year you can tap that button on the bottom right and steer it in whatever direction you see fit. So at any point you can chime in and let the algorithm know what you want and it's going to steer the session in that direction. Another example of what we call generative personalization is showcased in a prompted playlist, here that Devenge showed a little bit earlier as well. Here you basically have full unfettered access to the recommendation algorithm that Spotify has. And you can prompt it with very high-level prompts or very detailed prompts. On the left-hand side here I have a prompt that tells me "create me a playlist of bands that are playing in San Francisco tonight." So it turns out there's a bunch of good shows if you're excited to check them out. On the right-hand side you see a prompt that is asking for a playlist to accompany me on my run. And what's interesting with the right-hand side as well is that you will see that the experience itself gets dynamically shaped as a function of the request and the user to be able to introduce itself in different phases in my run as well. Another one that I'm really excited about, we launched it in New Zealand a couple of months ago and it's coming soon in more markets, is something called the taste profile. And that basically gives you the ability to introspect in natural language what the Spotify algorithm has understood about you in a way that you can edit and refine. So if you see something that's missing or something that's wrong. So for example, one of my edits is that all Disney music are my kids because we have a bunch of shared devices at home, but please don't recommend that to me. That's not my taste, please. And the algorithm we then take that into account, making sure that we never recommend this in the wrong context. And similarly, you can also use the taste profile to share some more aspiration goals as well. So getting into a new genre, getting into a new topic, learning about a new language, for example, all of these things can be done. Another one that's coming soon, which we announced very recently, is something called personal podcast, where the generative personalization system doesn't stop at just recommending and assembling experiences but also generating content as well. In this particular example, I'm generating a daily brief that's generated on a cadence daily. And I'm going to let it tell me about what's happening in my community. All right. So now to go into the guts of it. There's one big system that controls all of these different applications I mentioned and we call that internally the large taste model. We're not great at naming these internal things. So it has a couple of properties. One is that it understands every historical interaction and piece of content on Spotify. It combines prediction and reasoning to the point that I mentioned earlier, so not only guessing but also reasoning layered on top. And it gives users the ability to shape and generate experiences in real time. So it's fully steerable and promptable by users. And as of today, about one in four US premium subscribers interact with that system on a daily basis as well. So that's pretty exciting. One thing that's exciting as well is that deploying this system across existing recommendation surfaces also led to some gains. We saw gains on autoplay. We saw gains on podcast discovery. We saw gains on users interacting with DJ messages as well. So we saw pretty sizable gains across the board as well by deploying this system. On that note, I'm going to hand it over to Jackie to tell us about what's one of the core components that underpins this whole system. Hi everyone. I'm Jackie, or Jacqueline, a staff machine learning engineer at Spotify. So let's dive a little bit deeper and talk about how these models are actually trained, at least at Spotify. Semantic IDs were presented in the previous talk. But that is how we are embedding these open-weight LLMs with knowledge of Spotify's catalog. They are created by taking existing content embeddings such as podcast episode embeddings and applying a quantization algorithm to convert them to a set of discrete tokens. So we can use the model to be able to understand both natural language as well as these new special tokens, semantic IDs, that represent Spotify catalog entities. So we can empower experiences such as this where the user can ask in natural language for a podcast on morality. And that is passed to the prompt along with their listening history represented as semantic IDs. And the model responds both with a relevant semantic ID podcast episode as well as a natural language description of why they recommended that to this user. So how is this model actually trained? We published a paper linked here describing our training paradigm called NEO, which consists of four distinct stages. The first I already covered, which is the semantic foundation stage where we construct meaningful semantic ID tokens and then add them to an open-weight LLM's vocabulary. The second stage we call domain grounding, in which we align these new semantic ID token embeddings in the original language embedding space. We do this by learning a bidirectional mapping between semantic IDs to text, text to semantic IDs, and any combination. And we actually freeze the LLM backbone at this stage and only train the new semantic ID embeddings. So the original model weights and embeddings are frozen and we just learn those new semantic ID tokens. And this helps us to mitigate catastrophic forgetting of the pre-trained LLM's core language abilities. The third stage we call capability induction, which is multi-task instruction tuning on tasks that Spotify cares about, such as the ones shown here: next item recommendation, retrieval, etc. This is done by unfreezing the whole model, all of its weights and embeddings, and running either full parameter fine-tuning or LoRA fine-tuning on the multiple Spotify tasks. And then there is an optional fourth stage to do post-training, such as RL fine-tuning, etc. So how much of a difference does this four-stage training paradigm actually make? I'll dive into a few of the ablations we've done to investigate this. First, we assessed whether multi-task training is actually hurting performance by comparing the multi-task model against single-task variants. And we consistently saw that across our tasks, the multi-task model can match or actually beat the single-task performance, indicating that there is some positive cross-learning happening across the tasks. This is particularly noticeable for audiobook recommendations, if you look here, which is a newer content type at Spotify, demonstrating that these multi-task models can help with cold-start entities by learning from other items in the catalog, such as podcast recommendations, how to make meaningful audiobook recommendations. Then we did some ablations on the actual training recipe. We evaluated both dropping the frozen backbone domain grounding stage altogether, as well as combining the domain grounding stage with the capability induction stage in a multi-task. So we saw both of those that it degraded performance, but actually the biggest drop in performance was from using a randomly initialized backbone instead of the pre-trained open-weight LLM that we were using. We also evaluated using continuous pre-training for the domain grounding stage. And as you can see, it's minimal actual degradation on the task-specific performance, but where continuous pre-training really hits us is on the natural language and world knowledge capabilities of the pre-trained backbone LLM we are using. After we do continuous pre-training, it goes to essentially zero versus if we do the frozen backbone domain grounding, we retain all of that core language ability and are still able to learn the semantic IDs. I want to call out that these ablations were done with Gwenn, but we also validated that these findings hold with LLaMA, so it is not specific to the model backbone, but actually the training paradigm itself. And then lastly, we did some investigation on different inference strategies and their effect on accuracy versus latency. We tested beam search with both constrained decoding and not, and we saw that even without constrained decoding, we can generate valid semantic IDs 98% of the time. Constrained decoding does add a little latency overhead, but it's also helpful for specific cases where you want to target specific types of content, such as only making new content recommendations, for example. We also compared top-P sampling to beam search and saw that top-P sampling significantly hurts our accuracy. So although beam search is a little more latency intensive, we decided that trade-off worked for us. So what is meaningful about this? There have been lots of work in the industry in this space on generative semantic ID retrieval, tool-based LLM recommenders, the PLUM paper, etc. But NEO is actually the first example of combining all these capabilities into one system that understands grounded catalog items, is naturally language steerable, can do search, recommendation, explanation use cases, as well as low-latency, tool-free inference at an industrial scale. And we are using this in production today. There's a paper linked as well here for how we're using this to power podcast discovery. What we saw is that using a model trained like this, we can break users out of their habitual patterns and get them to listen to more unfamiliar content, and we saw huge wins online with this. So none of this works without meaningful evaluations, so I want to talk about that a little bit. As our recommendations are becoming more generative and explanatory, our original, maybe traditional, offline eval metrics are not sufficient. Yes, they can tell us whether or not the user interacted with that content, but they can't tell us whether or not that recommendation makes sense for the user, whether the explanation is accurate, whether it aligns with the user's intent, etc. So that's where LLM judges really shine. But in our work at Spotify, we really find that you need to invest in grounding your LLM judges in meaningful data, so that they can be reliable evaluators that align with human preferences. So a few examples for evaluating the podcast recommendations that I just mentioned. We create textual user profiles summarizing the user's listening history, and that is passed to the LLM as a judge. And we saw that this corresponded with a 75% alignment between the LLM judge and human preferences. Similarly, you can use actual behavioral signals to ground these LLM judges. So, for example, for a search task, for a given query, you can take similar queries and how the user has interacted with them in the past and pass that to the model. And we saw that overall it increased alignment by 5%, but on ambiguous queries, it actually increased alignment by 91%, showcasing the value this grounding of the LLM judge plays, especially in ambiguous cases where LLM judges tend to struggle. Lastly, we used grounded LLM judges to scale up our Cranfield-style collections. These are evaluation sets that are constructed by taking candidates from multiple different sources, creating a pool, and then using a human to rank that pool. However, that human ranking stage is expensive. So we invested in significant grounding for our LLM judge and are able to have an LLM judge that aligns with human system rankings with an agreement value of 0.87. In summary, as we're going to hear a lot about today in all the talks, there is a new era of personalization among us: this generative personalization. And if you want to power language-steerable personalized recommendations for your users, this is how we taught open-weight LLMs to speak Spotify. Thank you. All right. So now to go into the guts of it. So there's one big system that controls like all of these different applications I mentioned and we call that internally some the large taste model. We're not great at naming these internal things. So it has a couple of properties. One is that it understands every historical interaction piece of content on Spotify. It combines prediction and reasoning to the point that I mentioned earlier. So not only guessing but also reasoning layered on top. And it gives users the ability to shape and generate experiences in real time. So it's fully steerable and promptable by users. And as of today about one in four US premium subscribers interact with that system on a daily basis as well. So that's pretty exciting. One thing that's exciting as well too is that deploying this system across existing recommendation surfaces as well also led to some gains. We saw gains on autoplay. We saw gains on podcast discoveries. We saw gains on users interacting with DJ messages as well. So we saw pretty pretty sizable gains across the board as well by deploying this system. On that note I'm going to hand it over to Jackie to tell to us about what's one of the core components that underpins this whole system. Hi everyone. I'm Jackie or Jacqueline a staff machine learning engineer at Spotify. So let's dive a little bit deeper and talk about how these models are actually trained at least at Spotify. So semantic IDs were presented in the previous talk. But that is how we are embedding these openweight LLMs with knowledge of Spotify's catalog. They are created by taking existing content embeddings such as podcast episode embeddings and applying a quantization algorithm to convert them to a set of discrete tokens. So we can use the model to be able to understand both natural language as well as these new special tokens semantic IDs that represent Spotify catalog entities. So we can empower experiences such as this where the user can ask in natural language for a podcast on morality. And that is passed to the prompt along with their listening history represented as semantic IDs. And the model responds both with a relevant semantic ID podcast episode as well as a natural language description of why they recommended that to this user. So how is this model actually trained? We published a paper linked here describing our training paradigm called NEO, which consists of four distinct stages. The first I already covered, which is the semantic foundation stage where we construct meaningful semantic ID tokens and then add them to an openweight LLMs vocabulary. The second stage we call domain grounding in which we align these new semantic ID token embeddings in the original language embedding space. We do this by learning a bidirectional mapping between semantic IDs to text, text to semantic IDs, and any combination. And we actually freeze the LLM backbone at this stage and only train the new semantic ID embeddings. So the original model weights and embeddings are frozen and we just learn those new semantic ID tokens. And this helps us to mitigate catastrophic forgetting of the pre-trained LLMs core language abilities. The third stage we call capability induction, which is multi-task instruction tuning on tasks that Spotify cares about, such as the ones shown here, next item recommendation, retrieval, etc. This is done by unfreezing the whole model, all of its weights and embeddings, and running either full parameter fine tuning or LoRa fine tuning on the multiple Spotify tasks. And then there is an optional fourth stage to do post training, such as RL fine tuning, etc. So how much of a difference does this four stage training paradigm actually make? I'll dive into a few of the ablations we've done to investigate this. First, we assessed whether multi-task training is actually hurting performance by comparing the multi-task model against single task variants. And we consistently saw that across our tasks, the multi-task model can match or actually beat the single task performance, indicating that there is some positive cross learning happening across the tasks. This is particularly noticeable for audiobook recommendations, if you look here, which is a newer content type at Spotify, demonstrating that these multi-task models can help with cold start entities by learning from other items in the catalog, such as podcast recommendations, how to make meaningful audiobook recommendations. Then we did some ablations on the actual training recipe. We evaluated both dropping the frozen backbone domain grounding stage altogether, as well as combining the domain grounding stage with the capability induction stage in a multi-task. So we saw both of those of those that it degraded performance, but actually the biggest drop in performance was from using a randomly initialized backbone instead of the pre-trained open weight LLM that we were using. We also evaluated using continuous pre-trained using continuous pre-training for the domain grounding stage. And as you can see, it's minimal actual degradation on the task-specific performance, but where continuous pre-training really hits us is on the natural language and world knowledge capabilities of the pre-trained backbone LLM we are using. After we do continuous pre-training, it goes to essentially zero versus if we do the frozen backbone domain grounding, we retain all of that core language ability and are still able to learn the semantic IDs. I want to call out that these ablations were done with Gwenn, but we also validated that these findings hold with LLM, so it is not specific to the model backbone, but actually the training paradigm itself. And then lastly, we did some investigation on different inference strategies and their effect on accuracy versus latency. We tested beam search with both constrained decoding and not, and we saw that even without constrained decoding, we can generate valid semantic IDs 98% of the time. Constrained decoding does add a little latency overhead, but it's also helpful for specific cases where you want to target specific types of content, such as only make new content recommendations, for example. We also compared top piece sampling to beam search and saw that top piece sampling pretty significantly hurts our accuracy. So although beam search is a little more latency intensive, we decided that trade-off worked for us. So what is meaningful about this? There have been lots of work in the industry in this space on generative semantic ID retrieval, tool-based LLM recommenders, the plum paper, etc. But Neo is actually the first example of combining all these capabilities into one system that understands grounded catalog items, is naturally language steerable, can do search recommendation, explanation use cases, as well as low latency tool-free inference at an industrial scale. And we are using this in production today. There's a paper linked as well here for how we're using this to power podcast discovery. What we saw is that using a model train like this, we can break users out of their habitual patterns and get them to listen to more unfamiliar content, and we saw huge wins online with this. So none of this works without meaningful evaluations, so I want to talk about that a little bit. As our recommendations are becoming more generative and explanatory, our original, maybe traditional, offline eval metrics are not sufficient. Yes, they can tell us whether or not the user interacted with that content, but they can't tell us whether or not the user or that recommendation makes sense for the user, whether the explanation is accurate, whether it aligns with the user's intent, etc. So that's where LLM judges really shine. But in our work at Spotify, we really find that you need to invest in grounding your LLM judges in meaningful data, so that they can be reliable evaluators that align with human preferences. So a few examples for evaluating the podcast recommendations that I just mentioned. We create textual user profiles summarizing the user's listening history, and that is passed to the LLM as a judge. And we saw that this corresponded with a 75% alignment between the LLM judge and human preferences. Similarly, you can use actual behavioral signals to ground these LLM judges. So, for example, for a search task, you can, for a given query, you can take similar queries and how the user has interacted with them in the past and pass that to the model. And we saw that overall it increased alignment by 5%, but on ambiguous queries, it actually increased alignment by 91%, showcasing the value this grounding of the LLM judge plays, especially in ambiguous cases where LLM judges tend to struggle. Lastly, we used grounded LLM judges to scale up our Cranfield style collections. These are evaluation sets that are constructed by taking candidates from multiple different sources, creating a pool, and then using a human to rank that pool. However, that human ranking stage is expensive. So we invested in significant grounding for our LLM judge and are able to have an LLM judge that aligns with human system rankings with an agreement value of 0.87. In summary, like we're going to hear a lot about today in all the talks, there is a new era of personalization among us, this generative personalization. And if you want to power language steerable personalized recommendations for your user, this is how we taught OpenWay LLMs to speak Spotify. Thank you.