Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs
Description
Mahesh Sathiamoorthy's pitch is to stand in the researcher's shoes: the hard part of post-training is not the algorithm but the data and the environments that feed it. As agents get pushed to run autonomously for hours, something eventually falls over, and reinforcement learning is the tool for stretching that reliability, but RL environments are really just data in a different shape. Bespoke Labs works on curating both, from supervised fine-tuning sets to the environments models learn in. He grounds it in OpenThoughts, the widely used reasoning dataset his team built, and the counterintuitive lessons that came out of curating it: diversity of reasoning traces matters, keeping multiple answers per question helps, and the obvious recipe often is not the best one. A favorite example is teaching a model to reason about credit card compliance, where fine-tuning on the right tagged data lifted the compliance metrics that a raw model kept getting wrong. The through line, supported by their Curator tooling, is that a disciplined curation stack, not just more compute, is what turns a base model into a capable post-trained one. Speaker info: - https://x.com/madiator - https://linkedin.com/in/smaheswaran - https://smahesh.com Timestamps: 0:00 - Standing in the researcher's shoes 1:30 - Post-training at Bespoke Labs 3:13 - When agents fall over on long tasks 4:44 - RL environments as data 6:29 - Building OpenThoughts 7:36 - Finding a curation recipe 10:27 - Counterintuitive lessons 13:49 - A credit card compliance example 16:13 - Curating reasoning data with Curator 17:16 - The full curation stack
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: High-quality post-training data and reinforcement-learning environments—not compute or model-serving infrastructure—are the primary bottleneck to making reliable, long-horizon agents, and disciplined curation can also materially improve enterprise cost, latency, and compliance.
- Why it matters: The talk offers reusable evidence and design lessons for improving agent reliability through SFT/RL, building evaluable environments, and deciding when custom post-training is more valuable than continued prompting or reliance on frontier APIs.
- Best use: Use it as a practical reference for designing a post-training data pipeline and RL-environment stack, especially when evaluating whether an agent failure is best solved through prompts, harness changes, SFT, or RL.
Executive Summary
Mahesh Sathiamoorthy argues that AI evaluation has shifted from what models know to what agents can do autonomously over extended periods. The limiting factor for that shift is reliability: agents eventually select a wrong tool, make an unrecoverable mistake, or otherwise fail their task. While prompting, tool design, and agent harnesses can help, he positions post-training as a major lever for improving reliability and extending autonomous operation.
His central operational claim is that post-training infrastructure and compute are increasingly available, while curated data and RL environments are scarce. He treats RL environments as a form of data: rather than static prompt-response examples, they generate interactive trajectories and rewards. Bespoke's approach is to combine data curation with researcher-style measurement, running ablations over source selection, question mixing, difficulty filtering, teacher choice, response filtering, and the number of sampled trajectories before settling on a scalable recipe.
The OpenThoughts and OpenThoughts Agents projects support several counterintuitive findings: multiple independently sampled solutions for a given task can outperform simply collecting more distinct questions; the strongest available model is not necessarily the best teacher; and synthetic rewriting or augmentation may fail to improve agent-training data. The agent result also reinforces that SFT can drive much of the practical gain, while RL is compute-intensive and may be most justified for the final increment of performance.
The enterprise example with Intuit's Credit Karma grounds the argument. Fine-tuning a model to explain card recommendations reduced reliance on a long compliance prompt, but naïve training data was imbalanced enough that the model hallucinated financial attributes such as 0% APR. Bespoke addressed this with tagged representations that taught output form separately from specific values, improving compliance as well as latency and throughput. The strategic outcome was a lower-cost model the customer could own rather than continuously replace as frontier APIs change.
Key Takeaways
- Claim: Agent progress is now constrained less by benchmark knowledge and more by reliable execution over long horizons, making post-training a core capability lever. | Evidence: The speaker contrasts earlier knowledge benchmarks with action-oriented evaluations such as SWE-Bench and Terminal-Bench, and frames the goal as autonomous agents operating for hours, days, or weeks before failures in tool use or reasoning cause breakdowns. | Implication: For an agent system, diagnose failures by whether they are primarily prompt, tool/harness, or learned-behavior problems; recurring behavioral failures justify investing in post-training rather than endlessly expanding instructions. | Caveat: Post-training is not the only reliability lever; prompting, tool design, and harness updates can also improve outcomes.
- Claim: The practical bottleneck in post-training is high-quality curated data and RL environments, not merely compute or model-training infrastructure. | Evidence: Sathiamoorthy says compute, capable base models, and providers such as Fireworks, Tinker, and Slime are relatively established, whereas enterprises and frontier labs struggle to obtain high-quality task data and RL environments. | Implication: A post-training program should prioritize a data and environment asset strategy—task sourcing, evaluators, rollout generation, and versioning—before treating training infrastructure as the differentiator.
- Claim: Reasoning-data curation should be treated as an empirical recipe optimization problem, not as a one-time collection exercise. | Evidence: OpenThoughts starts from source prompts, then experiments with source mixing, LLM-based question-quality and hardness filtering, teacher-generated answers, answer filtering, and one-versus-many response sampling; the resulting recipe showed improving benchmark results as dataset size scaled. | Implication: Build an ablation loop around each curation choice and validate scaling behavior before committing to large-scale generation; a larger dataset only helps if the underlying curation recipe remains sound. | Caveat: The speaker does not provide the full winning recipe or quantitative benchmark values in this talk; those are referred to in the OpenThoughts paper.
- Claim: More solution diversity per task can be more valuable than broader task coverage, and teacher-model selection must be measured rather than inferred from model rankings. | Evidence: For OpenThoughts, generating 16 answers for one question performed well versus answering many more questions once, which the speaker attributes to diversity in reasoning traces. In both reasoning and agent work, stronger teacher models were not always superior; some Qwen models reportedly outperformed Claude models for the agent-data setting. | Implication: Budget generation experiments for multi-rollout diversity and evaluate teachers on downstream student performance, trajectory quality, and task fit—not only on the teacher's headline benchmark score. | Caveat: These findings are task- and curation-pipeline dependent, so they should not be generalized into a universal teacher-model ranking.
- Claim: For agent post-training, SFT can deliver much of the useful performance improvement, while RL is expensive and may be best reserved for the remaining hard-to-capture gains. | Evidence: In OpenThoughts Agents, the team found that SFT still contributed substantially to gains; RL was compute-intensive and primarily helped with the last few percentage points. | Implication: Start with curated demonstrations and SFT for enterprise use cases, then introduce RL only where a measured capability gap remains and the environment can produce trustworthy rewards. | Caveat: RL remains important where interactive feedback, long-horizon behavior, or edge-case optimization cannot be adequately represented in offline trajectories.
- Claim: Enterprise post-training can replace brittle, latency-heavy compliance prompting when the training representation explicitly separates output structure from volatile factual values. | Evidence: For Credit Karma's credit-card recommendation explanations, a long prompt was needed to enforce compliance and hurt latency. Plain-language fine-tuning data was imbalanced around attributes such as 0% APR, leading to hallucinated numbers; adding tags to emphasize the required form rather than individual values improved compliance, latency, and throughput. | Implication: For regulated or policy-bound generation, encode structured constraints and variable fields in the training representation, then test specifically for rare-value hallucination rather than relying on aggregate quality metrics. | Caveat: The presentation does not disclose the exact compliance improvement, latency reduction, training-set size, or the full tagging schema.
- Claim: A durable post-training stack requires three coordinated layers: environment lifecycle management, sandboxed rollout infrastructure, and optimization methods spanning SFT, RL, and prompt/harness improvement. | Evidence: The proposed reference stack includes building and measuring RL environments, quality measurement and version tracking; sandboxes for rollouts; checkpointing, snapshots, and rollback for long-horizon tasks; and upper-layer methods including SFT, RL, and a method the speaker calls JEPA for LLM-driven prompt optimization through reflection. | Implication: Treat environments, rollout execution, and training/prompt optimization as one control plane with reproducibility and recovery built in; long-horizon agent improvement cannot be managed solely as a model-training workflow. | Caveat: The speaker presents this as an emerging reference architecture rather than a validated standard, and does not specify interfaces, security controls, or operational cost tradeoffs.
Detailed Brief
What did not reliably improve the curated agent data
- Claims: Synthetic rewriting and task augmentation were expected to help OpenThoughts Agents but did not perform well in the reported experiments.; Answer filtering and other intuitive curation interventions did not consistently help in the reasoning-data work.
- Evidence: The speaker explicitly characterizes these results as counterintuitive and contrasts them with the more consistently useful strategy of sampling multiple responses.; The curation process was driven by stage-by-stage ablations rather than assumed best practices.
- Caveats: No exact ablation results, dataset splits, or definitions of the failed augmentation methods are presented, so the negative findings are directional rather than directly reproducible from the talk.
- Implications: Do not assume that synthetic expansion increases training value; require downstream evaluations against a fixed baseline before adding generation stages that create more volume but may reduce signal.
Why model ownership was part of the enterprise value proposition
- Claims: A custom post-trained model can reduce dependency on changing frontier models and potentially lower operating costs as proprietary model APIs become more expensive.; Post-training is presented as an operational optimization tool, not only a way to improve capability benchmarks.
- Evidence: In the Credit Karma example, the reported outcomes included improved compliance, lower latency, higher throughput, and the customer's ability to own the resulting model.; The speaker identifies lower latency, cost, and throughput as additional benefits of post-training generally.
- Caveats: The case does not establish total cost of ownership after accounting for data curation, evaluation, retraining, monitoring, and model-hosting operations.
- Implications: Assess custom-model opportunities against a full lifecycle business case: recurring inference spend and latency requirements versus the ongoing cost of maintaining proprietary data, evaluations, and retraining.
Notable Concepts & Terms
- OpenThoughts: An open reasoning-data project and paper created with collaborators from institutions including Stanford, UC Berkeley, and UW; it seeks a scalable recipe for curated reasoning post-training data.
- Bespoke Stratos: Bespoke's earlier effort to curate reasoning data after DeepSeek's release, which evolved into the OpenThoughts project.
- OpenThoughts Agents: The agent-focused extension of the OpenThoughts work, applying similar source selection, filtering, teacher selection, and rollout curation principles to agent trajectories and environments.
- RL environments / RLNs: Interactive training environments treated as a form of post-training data because they generate rollouts, feedback, and rewards needed to improve agent behavior.
- Terminal-Bench: An agent benchmark to which Bespoke has contributed, used as an example of evaluation moving from static knowledge to practical task execution.
- Curator: Bespoke's data-curation tool for generating and organizing post-training data from sources such as Hugging Face datasets or collected logs, with integrations to Tinker and Fireworks.
- Multi-answer sampling: Generating several reasoning traces or solutions for the same underlying task; the speaker reports it can outperform using the same generation budget to collect only more unique prompts.
- JEPA: The name used by the speaker for an LLM-driven reflective prompt-optimization method intended to improve system prompts and agent harnesses alongside model post-training.
Operator Notes / Why Ken Should Care
- Establish a failure taxonomy for production agents that records whether each failure is attributable to prompt design, tool/harness behavior, missing demonstrations, or an interactive-policy gap; use it to gate SFT and RL investment.
- Run a small curation matrix before scaling data generation: compare one-versus-many rollouts per task, multiple teacher models, source mixtures, and filtering policies using downstream task success rather than synthetic-data aesthetics.
- For regulated outputs, create adversarial test sets around sparse or high-risk attributes and evaluate factual-field hallucination independently from general response quality.
- Require every RL environment to have versioned evaluators, reproducible sandbox state, rollout logging, and checkpoint/rollback support before using it for long-horizon training.
- Build the economics case for custom post-training using full lifecycle costs, including curation, evaluation, retraining, hosting, and compliance monitoring—not only savings from reduced API use.
Source/Metadata
- Title: Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs
- Transcript words: 2642
- Duration seconds: 1151
- Timestamp note: No timestamps or chapter markers were present in the supplied transcript.
Transcript
Hey, everyone. Today I will be talking about data and environment curation for post-training LLMs. I am Mahesh Satyamurthy. I am co-founder and CEO at Bespoke Labs. Previously, I was a researcher and engineer at Google DeepMind. Very briefly, I will tell you a little bit about Bespoke. After that, the talk will be mostly around open source work we have done. Bespoke is an applied data research lab with a mission to help enterprises and frontier labs access high-quality data and RLN environments for their post-training needs. Very briefly, what we do and what we have done is that last year we put out something called Curator, which is a tool for curating synthetic data for post-training with SFD. Right after that, DeepSeek landed and we started an effort to curate reasoning data. That's how we started something called Bespoke Stratos, which eventually formed into the project called Open Thoughts, which some of you hopefully know about. And we have also been core contributors to Terminal Bench. These days we do a lot of research and build and ship RL environments. I was looking forward to the previous talk from Nick, who is also doing something similar. And the other thing we do is a lot of post-training and help enterprises get their own custom models. That's how we ended up with the Bespoke title for the company. The other thing I want to mention is there are a lot of people in our industry who create data, create RL environments, and then there are the researchers who consume this. But I feel like there is this slight mismatch, and it's beneficial for someone to do both at the same time. In fact, as you're curating data, you want to put yourself in the shoes of the researcher to see what it takes to actually move the metrics on the models. So, that's one of the motivations of how we think about this. The other thing I want to talk about is how AI has evolved, right? Early on, we used to think about and evaluate models on what they know. For example, this was a very popular benchmark on testing LLMs on various kinds of STEM, humanities, and all that knowledge. And these days, we have all these benchmarks that test how agents are able to do things. We have moved on from knowing to doing, right? So, that's the idea of agents, obviously. And one of the key principles, or one of the key things about agents, is that they are autonomous. And there are, as I was saying, many benchmarks, including SWE Bench, Terminal Bench, and so on. But ultimately, for many people, what they care about is whether these agents are autonomous for long durations of time. Ross had a great talk on long horizon, right? So, the goal is eventually to make these agents autonomous for maybe a few hours, a few days, or a few weeks. And what is it that's blocking the autonomy of agents? It's reliability, right? At some point, something falls apart. Either they called the wrong tool or they made a mistake, and whatnot, right? And what's one lever to improve reliability? There are, of course, many. Obviously, you can prompt your way to improving the agents' reliability, or you can update the harness, the tools, and whatnot. But post-training is a very powerful tool to improve reliability, or maybe even pre-train good models, right? So, if you think of Frontier Labs, this is one of their primary mechanisms for improving agents to get better capabilities in various domains, for better benchmark numbers, or better autonomy for longer and longer durations. And for post-training, one of the popular techniques, as you know, is reinforcement learning. And that's something a lot of you are excited about: the notion of RL environments. But ultimately, for post-training, be it SFT or reinforcement learning, data is the bottleneck, right? So, when I talk about data, RLNs are also something I'm calling data. It's just that the data is now in a very different shape. Again, here, compute is well-defined, good models exist, and the infrastructure for post-training exists. For example, there are various providers like Fireworks, Tinker, or Slime World and whatnot. So, all of those are somewhat well-defined. Most of the places where people struggle, especially enterprises, is that they don't have access to good quality data and RLNs. And this obviously also applies to Frontier Labs, where they have all this infra set up and they need good quality RLNs, right? So, that's one of the reasons we are thinking about why to invest time in doing data research and RLNs research. And as a side note, one of the other benefits of post-training is that, for example, you can reduce latency or improve cost, throughput, and whatnot. And I'll give one concrete example of post-training work we did with one of the enterprises. In this talk, I will mostly cover some of the work we have done in the open source community. We did some work on curating reasoning data for reasoning models and on curating trajectories and environments for agents. And recently, we had an engagement with post-training, which I'll very briefly talk about, and some tools on data curation. Open Thoughts is a reasoning data set as well as a paper, right? We started this effort last year. As I was saying, after DeepSea came out, we realized that there is a lack of very high-quality reasoning data in the community. Obviously, the labs have access to good data, but outside, we didn't have access to data, right? So, we at Bespoke started this effort called Bespoke Status. And then we realized that this is actually quite useful. So, we joined together with various folks at Stanford, UC Berkeley, UW, and so on, to create this consortium called Open Thoughts. And we did a lot of work on basically identifying the curation recipe. We also published this as a paper in iClear this year. And this is the main figure of the paper. So, what it shows is that we figured out a curation recipe, and it shows the scaling law, right? Again, this is last year when Amy and Live Code Bench were some of the popular benchmarks. What we showed is that with this recipe, if you keep scaling up the data set size, it's a scalable recipe, right? The metrics also improve. It's actually very widely used as well. For example, this is Microsoft CSO tweeting about the work. And this, Alex is my co-founder. He's a chief scientist and also a professor at UC Berkeley. And this is John Schulman talking about Open Thoughts, saying that he and his colleagues have been using it internally at Thinking Machines, right? And some of their blog posts also reference this. So, I'll talk about how we did the curation for Open Thoughts. This is the pipeline that we used. So, you start with a bunch of source questions, right? There are various data sets out there that have the prompt response, and we start with the prompts. These are various sources we have. And then, if you look at the paper, for any given data point, say if there are 10,000 samples that you want, the question is then how do you choose the questions from all these different data sets so that you have 10,000 for the data point? So, then there is the aspect around how you mix these questions. You can use various methods. The paper talks about, for example, using LLMs to check whether this is a good question, the hardness of a question, and so on. And then you want to filter questions and generate the answers. Again, this is all driven by LLMs, right? So, this is the curation recipe we did for creating this reasoning data set. And the answer generation is using teacher models. So, you can take other reasoning models such as DeepSeq or Quen-based models or even Gemini and whatnot. And then you can also filter the answers once you have the answers for these questions. And then you can also, given a question, generate multiple answers or a single answer. So, these are various knobs in the curation recipe. And the systematic way of doing this is that you run ablations and figure out what works in each of these stages, and you proceed to the next. So, after doing all of this, you get the final recipe, right? You can read this paper. It has lots of information about how we did the curation. But here are some of the learnings. Some of them are quite counterintuitive. And some of this was also covered in last year's AI Engineer conference. For example, sampling multiple answers per question works pretty well. This is something that is counterintuitive. As an example, something else we could have done is have many more questions and then just answer them exactly once, versus taking one question and answering it 16 times. I think the reasoning is probably that it gives variety in how reasoning is done. So, during fine-tuning, we also use the reasoning traces, right? So, I think the diversity helps there. And the other thing we saw is that stronger teachers are not always the best; stronger models are not always the better teachers. And there were a few other counterintuitive aspects around synthetic question generation or question answering working, whereas answer filtering and other aspects not working very well. So, after the Open Thoughts work, which was around data curation for reasoning models such as DeepSeek-type models, we moved on to Open Thoughts Agents, which is very similar. But how do you curate the data and RL environments for training agents now, right, not reasoning models? We have a very similar figure here. Again, we want to establish scaling laws. So, as you increase the data set size, we want to make sure that the curation recipe actually works. And again, I'm not going to go into details here, but very similarly, there are various ways of choosing different sources, for example, Stack Exchange and whatnot. How do you mix the tests? How do you filter? Generating the rollouts, choosing the teacher, and so on. And again, these are some of the lessons and learnings. As an example, even here, we saw that stronger models are not necessarily the best teachers, right? So, we found out that some of the QN models were better than, for example, cloud models, I think. And sampling multiple answers, again, helped in this case. Synthetic rewriting and task augmentation are something we thought would work, but they didn't work very well. And the other thing is that in this whole process of building this Open Thoughts Agent, SFT still contributed a lot to the gains. RL was very compute intensive, and for the last few percentages, it really helped. But in many situations, for example in enterprises, SFT actually works pretty well. And here's one concrete example I wanted to share on actually deploying something to production by post-training. We have seen a lot of people talk about post-training, but in enterprise settings, we haven't seen a lot of successes, at least. I haven't seen that. Here is a very concrete example with Intuit. There is this app called Credit Karma, which, if you install it, has a place where the app gives you a reasoning as to why a credit card has been recommended. You can prompt a model to do this, but one of the places where it fails is that it's not always compliant. So, you have to have a long list of rules to make sure the responses are compliant. And that actually blows up the latency. So, the answer here is that you want to curate data and post-train, right? It seems straightforward, but one of the things that we ran into is that the data set can be quite imbalanced. In lots of places, for example, you will have 0% APR, and the model after fine-tuning can hallucinate these numbers. So, this again ties back to what Ross talked about some time back with respect to the tags. And we created this specific curation recipe where, instead of just having these questions, the prompt response pairs in plain language, we added these tags, which helped the model focus on the kind of form rather than the specific numbers themselves. And that gave a big boost. And we saw that the overall compliance metrics improved, the latency improved, the throughput improved, and eventually they were able to own the model, right? As frontier models improve, they don't need to go and update it. And also, as we see now, the frontier models are getting more and more expensive. This gives them a very good way of owning the model and also lowering the costs. I think with that, I want to briefly touch upon Curator, the tooling that we built last year, which is for curating reasoning data. What it does is you can basically specify the, you can either go with, say, a Hugging Face data set where you have various prompts, or in many situations, you may have collected logs and you want to get the responses and fine-tune a model. So, this Curator makes it pretty easy to do that. And it comes with integration with Tinker and Fireworks. And this is, again, the tool that we used originally for curating OpenThoughts. And here is a very detailed diagram of what we are building today. But this, again, connects back to what Ross was talking about, where he was talking about algorithms, environments, and compute, right? So, it feels like we are converging on something very similar. So, if you think about the stack that is needed to not just curate these RL environments, but to post-train models, one of the things you need is obviously a handle on how you build these RL environments. How do you measure the quality? How do you track the different versions? And so on. So, that's one of the layers. And below that, you want various infrastructure to use sandboxes, right? To spin up the rollouts, to spin up the sandboxes to generate rollouts. And especially if you have long-horizon rollouts, then maybe at some points you need to do checkpointing. And then you need to be able to snapshot or roll back to something else, right? So, that's the lower-level compute and orchestration. And at the top, I have been giving examples on post-training. So, there is all this layer around how you do SFD, how you do RL, and so on. There is also this method called JEPA, which is on prompt optimization. I don't know if you guys have heard of it, but you can use LLMs themselves to optimize the prompts based on reflection. So, that also works pretty well for updating the system prompts and also the harnesses. So, this is kind of, I feel like, the new architecture or the new reference stack for how at least we are building, and how many others are building, the stack on how to build the RLNs and then also post-training agents. So, I think with that, I'll end the talk and happy to take questions offline. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much. Thank you very much.