Hey, everyone.
Today I will be talking about data and environment curation for post-training LLMs. I am Mahesh Satyamurthy. I am co-founder and CEO at Bespoke Labs. Previously, I was a researcher and engineer at Google DeepMind. Very briefly, I will tell you a little bit about Bespoke. After that, the talk will be mostly around open source work we have done. Bespoke is an applied data research lab with a mission to help enterprises and frontier labs access high-quality data and RLN environments for their post-training needs.
Very briefly, what we do and what we have done is that last year we put out something called Curator, which is a tool for curating synthetic data for post-training with SFD. Right after that, DeepSeek landed and we started an effort to curate reasoning data. That's how we started something called Bespoke Stratos, which eventually formed into the project called Open Thoughts, which some of you hopefully know about. And we have also been core contributors to Terminal Bench. These days we do a lot of research and build and ship RL environments. I was looking forward to the previous talk from Nick, who is also doing something similar.
And the other thing we do is a lot of post-training and help enterprises get their own custom models. That's how we ended up with the Bespoke title for the company. The other thing I want to mention is there are a lot of people in our industry who create data, create RL environments, and then there are the researchers who consume this. But I feel like there is this slight mismatch, and it's beneficial for someone to do both at the same time. In fact, as you're curating data, you want to put yourself in the shoes of the researcher to see what it takes to actually move the metrics on the models. So, that's one of the motivations of how we think about this.
The other thing I want to talk about is how AI has evolved, right? Early on, we used to think about and evaluate models on what they know. For example, this was a very popular benchmark on testing LLMs on various kinds of STEM, humanities, and all that knowledge. And these days, we have all these benchmarks that test how agents are able to do things. We have moved on from knowing to doing, right? So, that's the idea of agents, obviously. And one of the key principles, or one of the key things about agents, is that they are autonomous. And there are, as I was saying, many benchmarks, including SWE Bench, Terminal Bench, and so on.
But ultimately, for many people, what they care about is whether these agents are autonomous for long durations of time. Ross had a great talk on long horizon, right? So, the goal is eventually to make these agents autonomous for maybe a few hours, a few days, or a few weeks. And what is it that's blocking the autonomy of agents? It's reliability, right? At some point, something falls apart. Either they called the wrong tool or they made a mistake, and whatnot, right? And what's one lever to improve reliability? There are, of course, many. Obviously, you can prompt your way to improving the agents' reliability, or you can update the harness, the tools, and whatnot.
But post-training is a very powerful tool to improve reliability, or maybe even pre-train good models, right? So, if you think of Frontier Labs, this is one of their primary mechanisms for improving agents to get better capabilities in various domains, for better benchmark numbers, or better autonomy for longer and longer durations. And for post-training, one of the popular techniques, as you know, is reinforcement learning. And that's something a lot of you are excited about: the notion of RL environments. But ultimately, for post-training, be it SFT or reinforcement learning, data is the bottleneck, right?
So, when I talk about data, RLNs are also something I'm calling data. It's just that the data is now in a very different shape. Again, here, compute is well-defined, good models exist, and the infrastructure for post-training exists. For example, there are various providers like Fireworks, Tinker, or Slime World and whatnot. So, all of those are somewhat well-defined. Most of the places where people struggle, especially enterprises, is that they don't have access to good quality data and RLNs. And this obviously also applies to Frontier Labs, where they have all this infra set up and they need good quality RLNs, right?
So, that's one of the reasons we are thinking about why to invest time in doing data research and RLNs research. And as a side note, one of the other benefits of post-training is that, for example, you can reduce latency or improve cost, throughput, and whatnot. And I'll give one concrete example of post-training work we did with one of the enterprises. In this talk, I will mostly cover some of the work we have done in the open source community. We did some work on curating reasoning data for reasoning models and on curating trajectories and environments for agents.
And recently, we had an engagement with post-training, which I'll very briefly talk about, and some tools on data curation. Open Thoughts is a reasoning data set as well as a paper, right? We started this effort last year. As I was saying, after DeepSea came out, we realized that there is a lack of very high-quality reasoning data in the community. Obviously, the labs have access to good data, but outside, we didn't have access to data, right? So, we at Bespoke started this effort called Bespoke Status. And then we realized that this is actually quite useful.
So, we joined together with various folks at Stanford, UC Berkeley, UW, and so on, to create this consortium called Open Thoughts. And we did a lot of work on basically identifying the curation recipe. We also published this as a paper in iClear this year. And this is the main figure of the paper. So, what it shows is that we figured out a curation recipe, and it shows the scaling law, right? Again, this is last year when Amy and Live Code Bench were some of the popular benchmarks. What we showed is that with this recipe, if you keep scaling up the data set size, it's a scalable recipe, right? The metrics also improve. It's actually very widely used as well.
For example, this is Microsoft CSO tweeting about the work. And this, Alex is my co-founder. He's a chief scientist and also a professor at UC Berkeley. And this is John Schulman talking about Open Thoughts, saying that he and his colleagues have been using it internally at Thinking Machines, right? And some of their blog posts also reference this. So, I'll talk about how we did the curation for Open Thoughts. This is the pipeline that we used. So, you start with a bunch of source questions, right? There are various data sets out there that have the prompt response, and we start with the prompts. These are various sources we have.
And then, if you look at the paper, for any given data point, say if there are 10,000 samples that you want, the question is then how do you choose the questions from all these different data sets so that you have 10,000 for the data point? So, then there is the aspect around how you mix these questions. You can use various methods. The paper talks about, for example, using LLMs to check whether this is a good question, the hardness of a question, and so on. And then you want to filter questions and generate the answers. Again, this is all driven by LLMs, right? So, this is the curation recipe we did for creating this reasoning data set.
And the answer generation is using teacher models. So, you can take other reasoning models such as DeepSeq or Quen-based models or even Gemini and whatnot. And then you can also filter the answers once you have the answers for these questions. And then you can also, given a question, generate multiple answers or a single answer. So, these are various knobs in the curation recipe. And the systematic way of doing this is that you run ablations and figure out what works in each of these stages, and you proceed to the next. So, after doing all of this, you get the final recipe, right? You can read this paper. It has lots of information about how we did the curation.
But here are some of the learnings. Some of them are quite counterintuitive. And some of this was also covered in last year's AI Engineer conference. For example, sampling multiple answers per question works pretty well. This is something that is counterintuitive. As an example, something else we could have done is have many more questions and then just answer them exactly once, versus taking one question and answering it 16 times. I think the reasoning is probably that it gives variety in how reasoning is done. So, during fine-tuning, we also use the reasoning traces, right? So, I think the diversity helps there.
And the other thing we saw is that stronger teachers are not always the best; stronger models are not always the better teachers. And there were a few other counterintuitive aspects around synthetic question generation or question answering working, whereas answer filtering and other aspects not working very well. So, after the Open Thoughts work, which was around data curation for reasoning models such as DeepSeek-type models, we moved on to Open Thoughts Agents, which is very similar. But how do you curate the data and RL environments for training agents now, right, not reasoning models? We have a very similar figure here. Again, we want to establish scaling laws.
So, as you increase the data set size, we want to make sure that the curation recipe actually works. And again, I'm not going to go into details here, but very similarly, there are various ways of choosing different sources, for example, Stack Exchange and whatnot. How do you mix the tests?
How do you filter? Generating the rollouts, choosing the teacher, and so on. And again, these are some of the lessons and learnings. As an example, even here, we saw that stronger models are not necessarily the best teachers, right? So, we found out that some of the QN models were better than, for example, cloud models, I think. And sampling multiple answers, again, helped in this case.
Synthetic rewriting and task augmentation are something we thought would work, but they didn't work very well. And the other thing is that in this whole process of building this Open Thoughts Agent, SFT still contributed a lot to the gains.
RL was very compute intensive, and for the last few percentages, it really helped. But in many situations, for example in enterprises, SFT actually works pretty well. And here's one concrete example I wanted to share on actually deploying something to production by post-training. We have seen a lot of people talk about post-training, but in enterprise settings, we haven't seen a lot of successes, at least. I haven't seen that. Here is a very concrete example with Intuit. There is this app called Credit Karma, which, if you install it, has a place where the app gives you a reasoning as to why a credit card has been recommended.
You can prompt a model to do this, but one of the places where it fails is that it's not always compliant. So, you have to have a long list of rules to make sure the responses are compliant. And that actually blows up the latency. So, the answer here is that you want to curate data and post-train, right? It seems straightforward, but one of the things that we ran into is that the data set can be quite imbalanced. In lots of places, for example, you will have 0% APR, and the model after fine-tuning can hallucinate these numbers. So, this again ties back to what Ross talked about some time back with respect to the tags.
And we created this specific curation recipe where, instead of just having these questions, the prompt response pairs in plain language, we added these tags, which helped the model focus on the kind of form rather than the specific numbers themselves. And that gave a big boost. And we saw that the overall compliance metrics improved, the latency improved, the throughput improved, and eventually they were able to own the model, right? As frontier models improve, they don't need to go and update it. And also, as we see now, the frontier models are getting more and more expensive. This gives them a very good way of owning the model and also lowering the costs.
I think with that, I want to briefly touch upon Curator, the tooling that we built last year, which is for curating reasoning data. What it does is you can basically specify the, you can either go with, say, a Hugging Face data set where you have various prompts, or in many situations, you may have collected logs and you want to get the responses and fine-tune a model. So, this Curator makes it pretty easy to do that. And it comes with integration with Tinker and Fireworks. And this is, again, the tool that we used originally for curating OpenThoughts. And here is a very detailed diagram of what we are building today.
But this, again, connects back to what Ross was talking about, where he was talking about algorithms, environments, and compute, right? So, it feels like we are converging on something very similar. So, if you think about the stack that is needed to not just curate these RL environments, but to post-train models, one of the things you need is obviously a handle on how you build these RL environments. How do you measure the quality? How do you track the different versions? And so on. So, that's one of the layers. And below that, you want various infrastructure to use sandboxes, right? To spin up the rollouts, to spin up the sandboxes to generate rollouts.
And especially if you have long-horizon rollouts, then maybe at some points you need to do checkpointing. And then you need to be able to snapshot or roll back to something else, right? So, that's the lower-level compute and orchestration. And at the top, I have been giving examples on post-training. So, there is all this layer around how you do SFD, how you do RL, and so on. There is also this method called JEPA, which is on prompt optimization. I don't know if you guys have heard of it, but you can use LLMs themselves to optimize the prompts based on reflection. So, that also works pretty well for updating the system prompts and also the harnesses.
So, this is kind of, I feel like, the new architecture or the new reference stack for how at least we are building, and how many others are building, the stack on how to build the RLNs and then also post-training agents. So, I think with that, I'll end the talk and happy to take questions offline. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much.
Thank you very much.
Thank you very much.