SPEAKER_00
Hello everyone, I am Alessandro Capelli, I'm co-founder and chief customer officer at AdaptiveML. At AdaptiveML we build an RL ops platform, as in reinforcement learning operation, that allows large enterprises like AT&T, Manulife, CCS, to build, evaluate and serve in production their own specialized large language models. I'm here to show you how reinforcement learning, RL, is not just any other algorithm for post training, but is an algorithm that at its core will bring models to production. Around three years ago, I was part of a team that trained Falcon. Falcon around three years ago was one of the most widely adopted open source models.
SPEAKER_00
And we realized with my team, that is actually the core founding team of AdaptiveML, that the gap that was missing between bringing an open source model to production versus proprietary models like Anthropic Frontier Labs, like OpenAI, was actually reinforcement learning.
SPEAKER_00
95% of Gen AI pilots fail to reach production. Why is that the case? I believe it's what we call the myth of the last mile. Here you see a time description of what it takes to get to production, which I believe is false. The idea is that the hard part is to get to an MVP, to come up with a demo that looks nice in front of stakeholders, in front of your colleagues, and that is the hard part. And then there will be just the last mile, where the last mile will actually get the model into production. The issue is that most MVPs are built on top of proprietary models or are built on top of open source models using instruction fine tuning.
SPEAKER_00
Both of these solutions won't let you systematically improve your solution. They don't integrate in a really systematic and mathematical way what are the defects you might find in the journey to get to production. I will give you an example. If you test your solution and there will be some defects and you're using proprietary models, all you can do is change the system prompt. Now, you change the system prompt in one direction, you might have other defects, and there is no really a scientific, systematic way to improve that system prompt nicely in a way that you can easily monitor.
SPEAKER_00
Likewise, for instruction fine tuning, best you can do is iterate over the data set that might be expensive. And what after production? Will you keep creating a new data set every single week? This is what I believe is a more realistic view of what getting to production and beyond actually looks like. Getting to an MVP is not easy, but it's just the first mile. What actually is the real journey, the real marathon is to get from an MVP to production and beyond. And the secret to do that is to accelerate model life cycle, is to be able to integrate every single feedback you can get from a variety of sources to keep improving your solution.
SPEAKER_00
This continuous retraining, refinement and improvement, driven by real client feedback, business metrics and environmental reward, is unlocked in a systematic way only by reinforcement learning. Reinforcement learning, as I mentioned before, my entire point is that almost by design, by nature, allows to integrate feedback in almost a mathematical way. But reinforcement learning is not just that. Compared to other post-training techniques or steering behavior techniques like prompting and instruction fine tuning, they all reach the same goal, which is steer a model's behavior. But they're not equally effective.
SPEAKER_00
Reinforcement learning is disproportionately more effective than instruction fine tuning and likewise versus prompting. Reinforcement learning unlocks outsize performance. So what does it mean? What you can see in the plot is that you can get the same performance with RL with respect to SFT with a much smaller model. What it unlocks that actually helps you to get to production is a well-enabled scale at adoption.
SPEAKER_00
What do I mean by that? As you train smaller model, those models will be cheaper to serve at scale and the tokenomics of your use case will eventually make sense. When you are a big enterprise like AT&T, any use case, any feature that you want to be a commodity either for internal employees or customer-facing features, you might think at scale will cost you millions of dollars. As an example, AT&T summarizes every single transcript that might happen between a customer and an agent. Just summarizing that costs them millions of dollars. If you can train a model that is much smaller than a Chat GPT or a Sonnet, you will save money.
SPEAKER_00
Another thing you unlock is that smaller models will be faster. Not all use cases require speed, but many of them have a threshold of latency that is not just something nice to have. It's a constraint that will prevent you from getting into production. Let's say you have a model for customer support that is powering a speech-to-speech system. You can't go above half of a second. And I would say half of a second is already weird. When you're talking to someone and you have to wait half of a second, that's already weird. Ideally, it should be a third of a second. And a third of a second is something you will never get if you're using large language models.
SPEAKER_00
You need to use small models. Might be the latest Gemma, the latest Mistral, the latest Qwen, in that 10B family, but you can use much smaller models. The last thing you unlock is ownership. You will own the data that you give to the model because the model will be trained on your own business data and you will own the solution. So you don't need to worry about the latest update of the model that may shift performance underneath your feet. Everything I've said so far is true for any use case you might think.
SPEAKER_00
I've mentioned summarization. Could have been classification. Could have been OCR. Could have been anything you can think of. And reinforcement learning is already the better choice. But now we are in the area of agents. And agents actually make everything more complicated. Agents require more tokens, more complexity. There is less room for errors because now agents will have access to the data. They'll change things in the database connected to internal employees or clients you might have. So all of that raises the standard of what can be brought into production. And it raises further question on whether the tokenomics of an agent actually makes sense or not.
SPEAKER_00
I mentioned before, just for a summarization use case, you might spend millions. Imagine if you scale an agent to actually 10x the number of tokens. An RL advantage that already existed only widens when it comes to training agents. Because RL at its core was actually made to train robots, to train agents to live in environment.
SPEAKER_00
And environments are where agents actually behave. So RL naturally fits in a narrative where you want to train a model to be a good agent. Now there are two scenarios. Either you already have an agent in place. As an example, we work for Manulife, and Manulife already had agents. So they already have an ACCA workflow that has been settled. Imagine if you scale an agent to actually 10x the number of tokens. A rel advantage that already existed only widens when it comes to training agents. Because a rel at its core was actually made to train robots, to train agents to live in environment. And environments are where agents actually behave.
SPEAKER_00
So a rel naturally fits in a narrative where you want to train a model to be a good agent. Now there are two scenarios. Either you already have an agent in place. As an example, we work for Manulife, and Manulife already had agents. So they already have an ACCA workflow that has been settled. We don't need to recreate it on our site. You can directly plug a model that you might train. It can be the latest one, 3.5. And you can directly train the model on an environment that already exists. If such environment doesn't exist, it can still be built. You can still mock the tools, and you can mock, if you need one for this specific case, a mock user.
SPEAKER_00
If you want to create a chat that has access to tools, a mock user might be an LLM. What about the reward? The reward will be any business outcome, any KPIs, any LLM as a judge, that might define what success looks like to you. Was the agent helpful? Was the agent useful? Was the agent using a tone and a vocabulary that is following business guidelines? On this topic, two colleagues of mine, Letizia and João, they recorded a workshop that you might find on AI engineering website that will show you exactly how you can actually train a model by plugging into an existing environment.
SPEAKER_00
When I talk to clients, one of the main sources of doubt on whether they will ever get to an MVP or to production is because they don't have the data to do it. Data was already an issue before agents. After agents is even more of an issue because agents training data doesn't exist in the wild. There's no such data set you can scrape from the web where an agent is using tool. You don't have such a data set. The nice thing is that when you train a model with reinforcement learning and you have an environment and you have a reward in place, you just build what is, as a byproduct of your environment, you created a synthetic data set pipeline.
SPEAKER_00
As you have an environment, you can literally create trajectories that are good because the reward that you put in place will tell you what is good and what is not. So you can do rejection sampling and create a data set that you can use to bootstrap the first training of a model. And the nice thing is that even though many companies don't have the data, the exact data that is required to train agents, they still have a lot of data sets that can be leveraged to improve the entire experience in the environment. Such data might be real transcript between a customer and an agent that can be given to the mock user.
SPEAKER_00
The mock user can even be trained on that to be the actual realistic person that might be annoying, that might ask things three times in a row. We work with customers like medical supply, where people might call them, might be in panic. So the correct behavior might be I will escalate you to a human agent or we'll call 911 for you. And that kind of dirty real conversation is something that can be easily mocked by using proprietary data sets. Where is the human in the loop? RL became famous in the LLM world thanks to ChatGPT, because OpenAI published a blog post where it was saying we did RLHF, so reinforcement learning from human feedback.
SPEAKER_00
But sometimes the human in the loop, which is nice to hear, sometimes what actually hides behind is expensive annotation campaign. So in my experience, nobody wants to run a rotation campaign. It is either expensive or it is really useless, because the reality is that people don't want to do it. But you still want to keep a human in the loop. So where does the human in the loop come in the equation that I just showed you? When you train with RL, the most important thing you want to do is to build a reward signal. A reward signal might come from different sources. It might be a systematic reward when it comes from does the code run? Is the syntax correct?
SPEAKER_00
It can come from direct KPIs or business outcomes. One of our clients, CCS, the medical supply company I was mentioning before, has a customer support system like any other customer support system. What it is trying to maximize is containment rates. How many codes are actually brought end to end by the model. And that reward, that percentage of codes that are actually brought end to end, is something you can directly maximize. Many other things like was the tone correct? Were the business requirements followed? It's a bit of an open-ended question when it comes to systematic reward. But that issue can be solved with LLMs as judges.
SPEAKER_00
So the human in the loop is helping just by defining the rubrics, defining the system prompt of these LLMs as judges, and defining these scenarios. Making sure that it's aligned with what they see. But this activity that the human will do will take from a few minutes to hours. But it will not take weeks, and you don't have to do it iteratively dozens of times. For everything I mentioned before, RL not being just one algorithm, but the one algorithm that industrializes bringing model into production, in the last two years, at Adaptive we built Adaptive Engine. That is an aerial ops platform to evaluate, tune, and serve the best LLMs for your business.
SPEAKER_00
The Adaptive Engine is a holistic platform, where you can observe, train, and serve at once. When I mentioned at the very beginning that the goal is to accelerate the life cycle, that doesn't mean to accelerate training per se. You also want to evaluate, be sure that the model is actually behaving, and you want a systematic way to find defect, pre- and post-production, and to act accordingly. This is something you can do only if you have a systematic, holistic approach. Our models are built on top of the best open source models.
SPEAKER_00
Most open source models you can think of that are available, like the latest Gemma, Gemma 4 that you heard a few days ago, the latest Ministrel, the latest Gwenn, they're all available in your company, and depending on the model of your preference, you can start building on top of it. And finally, what's the catch with RL? The only catch with RL is that reinforcement learning is actually hard. Reinforcement learning is not as easy as changing a system prompt, and it's not as easy as just building a dataset for instruction fine-tuning.
SPEAKER_00
Our models are built on top of the best open source models. Any most open source models you can think of that are available, like the latest Gemma, Gemma 4 that you heard a few days ago, the latest Ministrel, the latest Gwenn, they're all available in your company, and depending on the model of your preference, you can start building on top of it. And finally, what's the catch with REL? The only catch with REL is that reinforcement learning is actually hard. Reinforcement learning is not as easy as changing a system prompt, and it's not as easy as just building a dataset for instruction fine-tuning. One of the most famous REL algorithms, which is PPO, requires orchestrating not one, but four large language models at the same time. That is where adaptive engine shines, because we let you define the rubrics and the rest, but we take care of the complexity of reinforcement learning by exposing a series of pre-built recipes for you. So you don't need to implement the latest algorithm, say GSPO, and you don't need to build the training recipe to run an actual training. So, once again, REL is the one algorithm that will let you bring modeling to production in a systematic and industrialized way, and all of that is possible with the adaptive engine. Thank you very much for your attention.
SPEAKER_00
I have a question about some of the human feedback portion that can get incorporated into REL. [SPEAKER_01] Yeah. For example, last year, cursor had a blog post where they outlined how they take human feedback from production data, such as whether or not the tab completion is accepted or not. [SPEAKER_01] Yeah. [SPEAKER_01] In settings like this where an LLM is at play, like in a more traditional LLM RL style, you would do several rollouts per prompt and pick the ones that work and train on just these samples.
SPEAKER_00
When it's human feedback and there's a single signal, do you do replays to have many variations of outputs for that problem and train on it? Or is it effective to just have a single implicit feedback from production to train on as a reward?
SPEAKER_00
[SPEAKER_01] So I would say what you ask is how do we leverage human feedback. I would say there's two scenarios. There's a scenario where sometimes the human feedback doesn't come from production. It comes from 10 to 20 feedbacks. In that scenario, what we do is that we basically usually use it to improve the LLM as judges, as in that is good, that is bad. How does it fit into the current description of what you are trying to do? And the nice thing is that as you go to production, then you will have thousands of such feedbacks. So what we do is that we usually, rather than using four LLMs as judges, at the very beginning, we just use prompted really big large language models, say QEN 235B. As we move to production, we have so much data that what we do is that we use this data to train reward models so that we can basically scale that human feedback to actually train actively the LLM. And then, with respect to the question when these feedbacks are not as explicit but more implicit, I think it really depends on the specific use case. And then we can build a reward model accordingly. Because we already have the data, we can have two different scenarios to see which training actually gives you the best output, the best performance given a certain evaluation.
SPEAKER_00
Thank you. [SPEAKER_01] You're welcome. Thank you. Thank you very much. Thank you. And they constantly use it to update the model.
SPEAKER_01
Yeah. In settings like this where an LLM is at play, like, in a more traditional LLM RL style, you would do, like, several rollouts per prompt and pick the ones that work and train on just these samples. Yeah.
SPEAKER_00
When it's human feedback and there's a single signal, do you do, like, replays to have many variations of outputs for that problem and train on it?
SPEAKER_01
Or is it effective to just have a single implicit feedback from production to train on as a reward?
SPEAKER_00
So I would say, you know, like, what you ask is how do we, you know, leverage human feedback?
SPEAKER_01
I would say there's two, let's say, scenarios. There's a scenario where, you know, sometimes the human feedback is like, it doesn't come from production, right?
SPEAKER_00
It comes from, you know, like 10 to 20 feedbacks.
SPEAKER_01
In that scenario, what we do is that we basically usually use it to improve the LLM as judges, as in that is good, that is bad. Like, how does it fit into the current description of, you know, of what you are trying to do?
SPEAKER_00
And the nice thing is that as you go to production, then you will have thousands of such feedbacks. So what we do is that we usually, rather than using four LLM as judges, at the very beginning, we just use prompted, really big, large language models, say, QEN 235B. As we move to production, we have so much data that what we do is that we use this data to train reward models so that we can basically scale that human feedback in, like, to actually train actively the LLM. And then, you know, with respect to, you know, to the question when these feedbacks are not as explicit but more like implicit, I think it really depends on the specific use case.
SPEAKER_00
And then we can build a reward model accordingly. You know, because we already have the data, we can have two different scenarios to see which kind of training actually gives you the best output, the best performance given a certain evaluation. Thank you. You're welcome.
SPEAKER_00
Thank you. Thank you very much. Thank you.