AI Engineer

Stop Making Models Bigger, Make Them Behave — Kobie Crawdord, Snorkel

3284 summary words 15 min summary Watch video

Start with the signal

15 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: A 4B parameter model fine-tuned with RL on high-quality tool-use data can outperform a 235B parameter model on financial analysis tasks for under $500 in training costs because the core bottleneck is disciplined tool behavior, not reasoning ability.
  • Why it matters: This demonstrates that production deployment cost/latency problems can be solved with targeted RL on smaller models instead of scaling to larger foundation models, especially when the failure mode is procedural discipline rather than knowledge or reasoning gaps.
  • Best use: Reference for production deployment strategy discussions, agent systems design, and making build-vs-buy decisions for on-prem enterprise AI applications where cost/security constraints matter.

Executive Summary

Snorkel AI, in partnership with UC Berkeley's RLLM/Agentica research group, demonstrated that a 4 billion parameter model fine-tuned with reinforcement learning can outperform a 235 billion parameter reasoning model (Qwen3) on financial analysis tool-use tasks. The training cost was under $500 for a 21-hour job using GRPO (Group Relative Policy Optimization). The breakthrough came from identifying that the core failure mode wasn't reasoning capability but tool-use discipline—specifically, the ability to inspect available tools, query schemas, handle errors, and self-correct.

The 235B model hallucinated answers after failing to query non-existent tables, while the fine-tuned 4B model systematically used getTableName and getTableInfo tools, corrected SQL errors by observing feedback, and arrived at correct answers. Surprisingly, training only on single-table queries (not multi-table curriculum learning) yielded the best performance and generalized to harder multi-table benchmarks, doubling pass@1 from 13.9% to 26.6% on the FinQA Reasoning benchmark.

Snorkel's approach emphasizes expert-in-the-loop data generation (PhD-level domain experts and industry practitioners) and verification to ensure high-quality training data. They advocate using rubric-based evals to decompose model failures into specific behavioral dimensions, then targeting RL training to fix the identified failure mode. Their FinQA environment is open-sourced on GitHub/HuggingFace/OpenEnv and available on Prime Intellect infrastructure, designed as a self-contained deployment with no external dependencies—critical for on-prem enterprise use cases in finance and healthcare.

The work challenges the default pattern of 'model not performing well → drop in a bigger model' and shows that for enterprise production use cases requiring cost efficiency, on-prem deployment, and data security, targeted RL on smaller models can be more effective than scaling parameter count. The presenter frames this as the 'Terence Tao effect'—you don't need a Fields Medallist to do SQL queries and arithmetic; you need someone who follows the right procedures consistently.

Key Takeaways

  • Claim: A 4B parameter model with RL fine-tuning outperformed a 235B parameter model on financial tool-use tasks | Evidence: Pass@1 doubled from ~14% to ~27% on FinQA Reasoning benchmark; the 4B model correctly used getTableName, getTableInfo, corrected SQL errors, while the 235B Qwen3 model hallucinated after querying non-existent tables | Caveat: Results are specific to tool-use tasks in financial analysis domain; not tested on general reasoning or knowledge-intensive tasks where larger models may still hold advantage | Implication: For production deployment of agents requiring tool discipline (API calls, database queries, error handling), Ken should prioritize targeted RL on smaller models over defaulting to frontier-scale models, especially when cost and on-prem deployment matter | Timestamp: timestamp unavailable
  • Claim: The training cost was under $500 for a 21-hour job using GRPO on a 4B model | Evidence: Speaker explicitly states: 'total cost of running that job was under $500 per run' and 'RL does not have to be a very expensive thing to be able to get non-trivial performance gains' | Caveat: No breakdown of compute infrastructure provided; cost likely assumes access to appropriate GPU infrastructure already; does not include data generation or verification costs | Implication: RL fine-tuning for production use cases is economically tractable even for mid-size teams; Ken should consider this for custom agent workflows where behavior needs to be shaped, not just knowledge added | Timestamp: timestamp unavailable
  • Claim: Training only on single-table queries outperformed multi-table training and curriculum learning approaches | Evidence: Ablation study compared single-table only, mixed single+multi-table, and curriculum learning; single-table training yielded best uplift and generalized to multi-table benchmark (13.9% → 26.6%) | Caveat: This is a surprising/counterintuitive result that may be specific to the FinQA domain and tool-use failure mode; mechanism for why single-table training generalizes better is not fully explained | Implication: Simpler, more focused training data can outperform complex curriculum strategies when the core skill being taught is procedural discipline rather than compositional reasoning; Ken should test minimal viable training sets before scaling complexity | Timestamp: timestamp unavailable
  • Claim: The core failure mode was tool-use discipline (inspecting environment, following procedures), not reasoning ability | Evidence: 235B model failed by querying non-existent tables and hallucinating; 4B model succeeded by calling getTableName to discover tables, getTableInfo to inspect schemas, and self-correcting SQL errors based on feedback | Caveat: This finding is specific to the FinQA environment where tools are well-defined and discoverable; may not apply to more ambiguous or poorly-documented tool environments | Implication: Ken should diagnose agent failures using behavioral rubrics to identify whether the issue is knowledge/reasoning vs. procedural discipline, then choose training approach accordingly; most production tool-use failures may be fixable with behavior shaping rather than model scaling | Timestamp: timestamp unavailable
  • Claim: Snorkel uses expert-in-the-loop data generation with PhD-level contributors and industry practitioners | Evidence: Speaker repeatedly emphasized 'we always make sure to have an expert in the loop' and 'we solicit the work and support of experts on various tasks... people at the PhD level for their domains of expertise' and 'people who are deep in the industry' | Caveat: No details on expert recruitment process, compensation, quality control mechanisms, or how expert disagreements are resolved; cost and scalability of expert-in-the-loop approach not addressed | Implication: For Ken's content/business use cases requiring domain-specific quality, human expert data may be necessary; for AI ops/agent systems, consider where expert verification is critical vs. where synthetic data suffices | Timestamp: timestamp unavailable
  • Claim: Rubric-based evals help identify specific behavioral failure modes to target with RL | Evidence: Speaker describes breaking down 'rightness or wrongness of a model's response into a full list of different questions' to 'find where the actual problem is among all the multiple possible arenas' before deciding which data sets to generate | Caveat: No concrete example of a rubric shown; method for translating rubric analysis into data generation priorities not detailed; how rubric output feeds into GRPO's single-value reward signal not explained | Implication: Ken should implement structured failure analysis for agent evals rather than binary pass/fail; decompose failures into schema discovery, error handling, query formation, etc., then target training to the weakest dimension | Timestamp: timestamp unavailable
  • Claim: The FinQA environment is self-contained with no external dependencies, critical for on-prem enterprise deployment | Evidence: Speaker states: 'everything is built into the environment, so there's no external dependencies that might be in some remote data center'; environment available on GitHub/OpenEnv/HuggingFace/Prime Intellect | Caveat: No details on environment complexity, setup requirements, or whether it supports arbitrary SQL databases or just a fixed schema; generalizability to real enterprise data environments unclear | Implication: For Ken's investing/GTM analysis, this pattern of self-contained evaluation environments is important for B2B/enterprise AI products where data sovereignty and offline operation are table-stakes; should look for similar architecture in agent frameworks | Timestamp: timestamp unavailable

Detailed Brief

Research Motivation and Enterprise Context

  • Claims: Enterprise production use cases require smaller models due to cost, speed, security, and on-prem deployment constraints; Common anti-pattern is 'POC works with big model → need to productionize → just throw bigger model at it'; Financial and healthcare domains require data control and cannot export to external APIs; The 'Terence Tao effect': you don't need a Fields Medal mathematician to do SQL queries and arithmetic
  • Evidence: Speaker describes typical enterprise workflow: build POC with large model, then face deployment challenges when trying to productionize; References to on-prem requirements, data export concerns, security constraints as drivers for smaller models; RLLM team coined the 'Terence Tao effect' metaphor: 'that much brilliance might not necessarily be what a financial analyst actually has to have'; Snorkel positions itself as 'Frontier AI Data Lab' focused on expert-in-the-loop data quality for top labs
  • Caveats: No quantification of actual enterprise deployment cost differences between 4B vs 235B models; Security/compliance arguments are asserted but not substantiated with specific regulatory requirements; The Terence Tao metaphor is evocative but doesn't explain why reasoning models fail at procedural tasks
  • Implications: Production AI strategy should start with smallest model that can be made to work via RL, not largest model available; Data control and deployment flexibility may be more important than raw capability for enterprise buyers; Ken should evaluate agent frameworks/products on whether they support on-prem deployment and model flexibility

Experimental Setup and Environment Design

  • Claims: FinQA environment provides 290 standard samples and 79 harder multi-table reasoning samples; Environment is self-contained with built-in tools (getTableName, getTableInfo, SQL query execution); Environment modeled after OpenGym/Harbor patterns for reproducibility; Available on Prime Intellect, OpenEnv, GitHub, and HuggingFace Spaces; Training used GRPO (Group Relative Policy Optimization) on a 4B parameter base model
  • Evidence: Speaker explicitly mentions '290 samples' and '79 samples called FinQA reasoning that requires multi-table queries'; Tools demonstrated: getTableName to discover tables, getTableInfo to inspect schemas, SQL query execution with error feedback; Multiple distribution channels listed: Prime Intellect infrastructure, OpenEnv repo on GitHub, HuggingFace spaces via PyTorch/HuggingFace collaboration; 21-hour training job, under $500 cost, using GRPO algorithm
  • Caveats: No details on hardware used for $500 training run—cost may not be replicable without equivalent infrastructure; Environment complexity not quantified—number of tables, schema depth, query complexity not specified; No comparison to other tool-use benchmarks (WebShop, ToolBench, etc.) to assess difficulty level; Data generation cost not included in $500 figure—expert-in-the-loop verification is likely significant additional cost
  • Implications: Ken can experiment with this environment as a reference for financial agent workflows; Self-contained environment design is a best practice for reproducible RL training and enterprise deployment; Open-sourcing the environment builds community credibility for Snorkel and enables third-party validation; GRPO appears to be effective for tool-use behavior shaping; worth investigating for agent systems work

Behavioral Analysis: What the Models Actually Did

  • Claims: The 235B Qwen3 model queried non-existent tables, got no results, then hallucinated an answer; The 4B fine-tuned model first called getTableName to discover available tables; The 4B model then called getTableInfo to inspect schema before querying; The 4B model made a SQL error (queried for 'revenue' column that didn't exist), observed the error, and self-corrected to the correct column name; Tool-use discipline (inspect environment, handle errors) was the differentiator, not reasoning ability
  • Evidence: Speaker shows example: 'What is the year over year growth rate of YouTube ads revenue from 23 to 24' → 235B model 'threw a query out without' inspecting tables, 'didn't get anything back... falls back to just hallucinating an answer'; 4B model demonstrated sequence: getTableName → getTableInfo → SQL query → error → correction → correct answer; Speaker concludes: 'it wasn't the reasoning that was the issue, it was the tool use' and 'the tool discipline... turned out to be a bigger deal than anything else'; Both models had access to the same tools; only the fine-tuned 4B model learned to use them systematically
  • Caveats: Only one example trajectory shown for each model—no statistical analysis of failure modes across the full benchmark; No discussion of whether the 235B model could be prompted to use tools correctly (system prompt engineering not attempted); Error correction behavior may be specific to SQL query tasks with clear error messages; generalizability uncertain; No analysis of whether the 4B model learned genuine error-handling heuristics or just memorized correction patterns from training data
  • Implications: Ken should treat tool-use discipline as a distinct skill from reasoning and prioritize it in agent design; Prompt engineering may be insufficient for tool-use tasks—behavior shaping via RL may be necessary; Error-handling loops (observe feedback → correct → retry) should be explicit design patterns in agent workflows; For production agents, systematic environment inspection (list available tools/APIs, check schemas) should be enforced at architecture level, not left to model discretion

Ablation Study Results and Generalization

  • Claims: Three training regimes tested: single-table only, mixed single+multi-table, curriculum learning (single → multi); Single-table only training yielded the best performance; Single-table training generalized to multi-table benchmark, doubling performance from 13.9% to 26.6%; Tool discipline learned from simple examples transferred to complex examples
  • Evidence: Speaker describes ablation: 'train with single table only... train with multi-table mixed in... or try to do some curriculum learning'; Result: 'single table only training was actually the one that yielded the greatest uplift'; Performance on harder FinQA Reasoning benchmark: '13.9 to 26.6 percentage jump after this training'; Speaker emphasizes surprise: 'even though the single table only training regime was the best training regime, the uplift... on that harder benchmark... was a similar doubling'
  • Caveats: No explanation for why single-table training outperformed curriculum learning—mechanism unclear; No performance numbers given for the 290-sample standard FinQA benchmark, only the 79-sample harder set; Generalization from single-table to multi-table may be specific to SQL tool use where the tool-calling pattern is similar regardless of query complexity; No analysis of whether multi-table queries require different reasoning or just the same tool discipline applied multiple times
  • Implications: Ken should start with simplest training examples when shaping agent behavior, even if target use case is more complex; Curriculum learning may be overrated for procedural skill acquisition—focused repetition on basics may be more effective; Generalization from simple to complex may work when the underlying behavior pattern is the same (inspect → query → handle errors) regardless of task complexity; For agent training, prioritize teaching foundational behaviors (tool discovery, schema inspection, error handling) over task-specific knowledge

Data Generation Philosophy and Rubric-Based Evals

  • Claims: Snorkel uses expert-in-the-loop data generation with PhD-level and industry expert contributors; Verification step ensures tasks are well-formed and have verifiable answers; Rubric-based evals decompose model responses into multiple binary questions to diagnose failure modes; Rubrics inform which behaviors to target with data generation, while RL uses single-value rewards
  • Evidence: Speaker states: 'we have a platform that we've used internally... we solicit the work and support of experts... people at the PhD level... people who are deep in the industry'; Process described: generate data → verification step to ensure tasks are 'appropriately fitted' and 'verifiable answer'; Rubric approach: 'breaking down the rightness or wrongness... into a full list of different questions... you can then start to use the rubric as a way to find... where the actual problem is'; Separation of rubric (analysis) and RL (training): 'the RL still gets your single value, as a GRPO just usually works with a single value'
  • Caveats: No concrete rubric example shown—unclear what dimensions are evaluated or how granular the decomposition is; Expert recruitment, quality control, and cost not discussed—scalability unknown; How rubric analysis translates into data generation priorities not detailed—process is described but not operationalized; No discussion of inter-rater reliability or how expert disagreements are resolved
  • Implications: Ken should implement multi-dimensional failure analysis for agent evals, not just pass/fail metrics; For high-stakes domains (finance, healthcare), expert data may be necessary; for lower-stakes use cases, consider whether synthetic data with verification is sufficient; Rubrics can inform what to train on, but training itself may still require simplified reward signals (GRPO's single value); The separation of analysis (rubric) and training (single reward) is an important design pattern for RL workflows

Notable Concepts & Terms

  • GRPO (Group Relative Policy Optimization): The RL algorithm used for fine-tuning the 4B model; speaker mentions it 'works with a single value' reward signal, implying it's simpler than methods requiring detailed feedback
  • Terence Tao Effect: Metaphor from RLLM team: using a reasoning genius (Terence Tao, Fields Medal mathematician) for basic tasks (SQL queries, arithmetic) is overkill; highlights that tool-use tasks need procedural discipline, not reasoning brilliance
  • FinQA Environment: Snorkel's open-source self-contained evaluation environment for financial tool-use tasks; includes 290 standard samples and 79 harder multi-table reasoning samples; available on GitHub/HuggingFace/Prime Intellect
  • Rubric-Based Evals: Snorkel's approach to decomposing model responses into multiple binary evaluation questions to diagnose specific failure modes (e.g., did it inspect schema? handle errors? use correct tool?); used to inform data generation strategy
  • Tool-Use Discipline: The speaker's term for systematic, procedural behavior with tools: discovering available tools, inspecting schemas/documentation, handling errors, self-correcting; distinct from reasoning ability or knowledge
  • Expert-in-the-Loop Data Generation: Snorkel's core methodology: using PhD-level domain experts and industry practitioners to generate and verify training data, with emphasis on quality over scale
  • RLLM / Agentica Project: UC Berkeley research group that partnered with Snorkel on this work; developers of the RLLM framework used for the RL training loop
  • Self-Contained Environment: Architecture pattern where all dependencies (databases, tools, schemas) are bundled with the environment; critical for on-prem enterprise deployment and reproducible training with no external API dependencies

Operator Notes / Why Ken Should Care

  • For Ken's agent systems work: Tool-use discipline (systematic environment inspection, error handling, self-correction) is a distinct skill from reasoning and should be explicitly architected, not just prompted. Consider whether RL behavior shaping is needed vs. prompt engineering.
  • For AI ops: Under-$500 RL fine-tuning runs are economically tractable for mid-size teams. GRPO appears effective for behavior shaping on 4B models. Consider this for custom agent workflows where behavior needs to be shaped, not just knowledge added.
  • For content/business applications: Expert-in-the-loop data generation is Snorkel's moat. Evaluate whether your use case requires expert data (high-stakes, domain-specific) or whether synthetic data with verification suffices (lower-stakes, general knowledge).
  • For investing/GTM analysis: On-prem deployment and data sovereignty are table-stakes for enterprise AI in regulated industries (finance, healthcare). Look for agent frameworks/products that support self-contained deployment and model flexibility. The 'POC with big model → productionize with same big model' pattern may be a red flag for GTM fit.
  • For workflow optimization: Rubric-based failure analysis should replace binary pass/fail evals in production agent systems. Decompose failures into dimensions (tool discovery, schema inspection, error handling, query formation) and target the weakest dimension with training or architectural changes.
  • For research/learning: The single-table training outperforming curriculum learning is surprising and worth deeper investigation. May indicate that procedural skills benefit more from focused repetition than progressive complexity. Test this hypothesis in your own agent training workflows.
  • For open-source strategy: Snorkel open-sourced the FinQA environment on GitHub/HuggingFace/Prime Intellect to build credibility and enable third-party validation. This is a good pattern for enterprise AI products that want to establish trust and reproducibility.

Watch Map

  • timestamp unavailable: Conference context and Snorkel positioning (Frontier AI Data Lab, expert-in-the-loop approach)
  • timestamp unavailable: Research objective: 4B model outperforming 235B model on financial tool-use
  • timestamp unavailable: Enterprise context and 'Terence Tao effect' motivation
  • timestamp unavailable: Example: 235B Qwen3 model hallucinating on YouTube ad revenue question
  • timestamp unavailable: RL approach: GRPO, under $500 training cost, 21-hour job
  • timestamp unavailable: FinQA environment design and distribution (GitHub, HuggingFace, Prime Intellect)
  • timestamp unavailable: Behavioral comparison: 4B model using getTableName, getTableInfo, error correction
  • timestamp unavailable: Ablation study: single-table training outperforming curriculum learning
  • timestamp unavailable: Generalization: single-table training doubling performance on multi-table benchmark
  • timestamp unavailable: Tool discipline as the key differentiator, not reasoning ability
  • timestamp unavailable: Rubric-based evals for diagnosing failure modes and informing data generation
  • timestamp unavailable: Summary and links to blog post and UC Berkeley partner post

Source/Metadata

  • Title: Stop Making Models Bigger, Make Them Behave — Kobie Crawdord, Snorkel
  • Transcript words: 5534
  • Duration seconds: 1255
  • Timestamp note: No timestamps present in transcript; video duration is 1255 seconds (~21 minutes)
Full transcript 3756 words · 24 min read
0:00

[SPEAKER_00] This is the last presentation I have to give this conference, so I'm feeling already a little bit of the euphoria of, ah, it's all done.

0:16

SPEAKER_00

I know we're at a close to the end of the whole sequence. I keep finding these conferences to be some of the highest signal that I get, wherever I go. So generally, is that people feel they're getting what they came for here? I'm just curious because we're at Snorkel, we put in a sponsorship and we want to know that people are getting what they want, they know they're going to come back because we want to sponsor next time, we want to know if people are happy about it. So, did you guys see what you wanted to see? Yeah? Yeah, really good. Brilliant, brilliant. So, now that it is 3:45, I'm going to go ahead and start the official thing.

0:55

SPEAKER_00

So, my name is Coby Crawford, I'm a developer advocate at Snorkel. We call ourselves the Frontier AI Data Lab and what we're doing right now, our main thing is starting from the research-backed work that Snorkel's been doing since its inception, we've been working on a variety of things about data quality and then at this point now where we're focused is actually providing data sets where we assure a certain level of quality. We're very attentive to being very motivated about making sure the data is high quality. And part of how we get to high quality is we always make sure to have an expert in the loop as part of the process.

1:20

SPEAKER_00

So, we have expert contributors that we work with and we bring people in to provide their expertise to make sure that the data that we generate is of top quality. And then for the top labs that want to use our data to improve their models and get the hill climbing done in the right way, that's what we do at Snorkel. Because of that, a lot of what goes on is still more research and this is a talk that's talking about some of the work that we did, that our research team did. And one of the keys in this research is that we're looking at how the best quality data can be best applied and where it is that we need to be looking for opportunities to get that done.

1:48

SPEAKER_00

So, in this particular case, talking about stop making models bigger, it's a nice punchy title.

2:00

SPEAKER_00

Of course, we don't really mean that models shouldn't be large intrinsically, but the point broadly speaking is that sometimes we find great wins to be had with the right data applied to the right problem statement. And so this is something we're going to talk about, a specific use case that our research team discovered and in partnership with the RLLM team, which is a research group, part of UC Berkeley. And so the UC Berkeley team over there, RLLM, the Agentica Project, their lab partnered with us on this particular work. So the goal, as I said, is making a four billion parameter model outperform a 235 billion parameter model on tool use tasks for financial analysis.

2:42

SPEAKER_00

So we'll start with the research objective and then we'll iterate through talking about the approach that was used for this particular process and then talk about the results and happy to report that we got what we were looking for. So good things to be had. So a couple of quick level setting backgrounds of what we're talking about here. First is that as we see enterprise use cases take on greater complexity, we obviously have the massive explosion of what people are doing in terms of personal assistance.

3:14

SPEAKER_00

And as people are working in the context of enterprise, a lot of times you still need a more constrained choice about how to implement something and make sure that it's reliable. When you're looking for things that are going to be done for enterprise production use cases, you also have to make sure there's a lot of safety and security things done. So these other priorities that fold into what people typically want to do, we're looking at these things and saying, okay, well, these are the enterprise use cases that people have.

3:35

SPEAKER_00

And as people try to solve the problems of making the models perform at the level that makes it acceptable for actually being deployed as a production service, we see very often that people choose, well, okay, we didn't get the performance that we wanted with this right now. We'll just drop in a larger model. It'll be smarter. It has greater reasoning skills. And we'll just expect that the performance will improve commensurate with the additional load of the size of the model and the greater inference cost that goes along with that. And in some cases that might not always be the right thing.

4:30

SPEAKER_00

So we see people saying, you know, let's just go get a bigger model that'll solve the problem. And sometimes maybe that isn't quite the answer. In this case, what we're trying to do is to say, can we take a smaller model and then use RL with the right data to yield the kind of performance gains that we're looking for and to deliver the kind of application functionality that we want. And so that's the target here. And again, for these various reasons, cost, speed, security, and then the idea that in general, you start with a really big model and make your POC and make it work. And then everybody's happy that it works.

4:49

SPEAKER_00

And it's like, okay, now what do we do to productionize it?

4:55

SPEAKER_00

And you want to roll the production. You want to think about how you're able to deploy that. Do you need to keep everything on premise? Do you have the ability to deploy and run your service yourself so that you don't have to have external dependencies and worry about the data export aspects and data control, especially in the context of financial data and healthcare and other domains? People have to be concerned about those aspects as well. So for getting a smaller model to be able to perform as well as larger models, we feel that in the particular case of talking about tool use for financial analysis, RL is the right time to be making the kind of training.

5:24

SPEAKER_00

You're talking about changing the behavior. And so that's more of a behavior thing. Then RL is better for behavior. Then say you're talking about changing the core data and knowledge that's inside of that. So that's an intuition about how we've approached it. And that's part of what's going on here. So a larger model, sometimes it's more like taking a sledgehammer to crack a walnut. It's like just adding all of this capability is this. And the RLLM team that we worked with, they talked about this as a, and their description of it was the Terence Tao effect.

6:04

SPEAKER_00

Terence Tao, the famous mathematician who's, I forget what awards he's won and whatnot, but well known for being generally brilliant about mathematics across the board. And therefore could approach and manage any kind of mathematical problem. But that much brilliance might not necessarily be what a financial analyst actually has to have. They don't have to know all the kinds of math. They don't have to do late in digital virtual A algorithm stuff to talk about doing a SQL query and getting some math, getting some data back and then doing some addition and subtraction, right? and their description of it was the Terence Tao effect.

6:24

SPEAKER_00

Terence Tao, the famous mathematician, I forget what awards he's won, but well known for being generally brilliant about mathematics across the board. And therefore could approach and manage any kind of mathematical problem. But that much brilliance might not necessarily be what a financial analyst actually has to have. They don't have to know all the kinds of math. They don't have to do late in digital virtual algorithm stuff to talk about doing a SQL query and getting some math, getting some data back and then doing some addition and subtraction, right? So the idea that you must always get to a much smarter model to do something or deeper reasoning to get something done well, as the thing we're challenging here. So here is that 235 billion Quen3 model responding to the question in this environment that we built. I'm going to talk about the environment a little bit more in detail later. But I point this out to show, here's a reasoning model, a smarter model, and its response in the context of needing to actually use tools. So the response that it generated to the question, what is the year over year growth rate of YouTube ads revenue from 23 to 24, began with making a query to find some existing values, but the query it chose was to a non-existent table. The table didn't exist. It didn't actually inspect the environment and inspect the tools to find out what tables it could query. It just threw a query out without doing that. So the table wasn't there and didn't get anything back. It guesses again, still doesn't get anything back. And then having not gotten anything back in either of those two attempts, it falls back to just hallucinating an answer. And so out comes this hallucinated answer. It's completely—don't know what the weights told it to say—but that's what came out and it's not very useful. So even though the model is incredible in terms of being much better at reasoning than a much smaller model would be, that greater reasoning did not help it when it needed to use the tools. We're going to come back to the same question again against the model that we fine-tuned that's only four billion parameters and you're going to see the difference and we'll talk a little bit more about those differences later. So put a pin in that, come back, we'll see that year over year question from come back. So here, this is what we're talking about, summarizing it again. No discipline in tool use, even though it has all the abilities to reason that it has. Moving forward to then what we did for this attempt to use RL to make the smaller model work well. The first thing is to generate a high quality data set. At Snorkel, our general approach is again to have experts in the loop. I don't know if I've said it again now, but I don't think I've already said it, right. We have experts in the loop for the data that we do. The way we generate data and the way we work on it is we have a platform that we've used internally for interacting with things. We solicit the work and support of experts on various tasks and various topics. So if we need somebody who is in the financial analysis space already, then we get them and pull them in. We'll work with people at the PhD level for their domains of expertise. And also, of course, people who are deep in the industry and have been working for some time and they know their space well. The process of doing that, that's one of the things that we put an emphasis on, is how we work at Snorkel for our data generation. And then broadly speaking, naturally that can be augmented with other kinds of things. But that's really key about what we want to do in terms of emphasizing quality as a core element. So we have the data set. And then we go through and make sure there's a verification step done to make sure that the tasks that are defined from that data set are actually appropriately fitted to the task and are actually good tasks that it can be queried. You know that you're going to get the results that you need from it and that we should be able to have a verifiable answer that we're looking for. So we do all the verification steps to make sure that everything's correct on that front. And that's another part of what it means to put together the data set and have it ready for use. And data quality, again, is a big emphasis for us. So we need to make sure that's the key. And then it was time to do RL with it. And the way that this was done, we're talking about very few surprises in terms of what you've seen the state of the art in this space. GRPO, again, we started with a 4 billion parameter model. And then the environment that we use, the RLLM framework, again, through the UC Berkeley partnership, they're the developers of that framework. And we have our FinQA environment that we've built. And we're going to talk a little bit more about the details about that environment in just a moment. But then this is something that was able to be done in a 21-hour job. And the total cost of running that job was under $500 per run. So RL does not have to be a very expensive thing to be able to get non-trivial performance gains. And if you're already working on and working with models that you want to host yourself, if you're already thinking about what you'd like to do to be able to do things we have on-prem solutions, or things where you're doing it with smaller models, and you aren't already thinking that you can improve the models the way that you want, then this is a call to action that you actually can. That it's actually a very tractable thing to get a model that you want to work with actually up to the performance levels that you need using RL. Even if Karpathy doesn't like it. So our FinQA environment is something that we built. It's set up for being able to host the kinds of questions that are being done here. It provides a specific set of tools. It's set up where everything is built into the environment, so there's no external dependencies that might be in some remote data center that you don't have access to. So when you deploy the environment, it's fully self-contained. It's what I'd call that if you've worked with something like Harbor before, or if you've worked with OpenEnv, you're familiar with the same thing about using an environment like this. And this is an environment that we've actually built and published. It's available on Prime Intellect's infrastructure as well. So it's something you can load up right there at Prime Intellect.

6:28

SPEAKER_00

It provides a specific set of tools. It's set up where everything is built into the environment, so there's no external dependencies that might be in some remote data center that you don't have access to. So when you deploy the environment, it's fully self-contained. It's a rollout that if you've worked with something like Harbor before, or if you've worked with OpenEnv, you're familiar with the same thing about using an environment like this. And this is an environment that we've actually built and published. It's available on Prime Intellect's infrastructure as well. So it's something you can load up right there at Prime Intellect.

6:59

SPEAKER_00

Also on OpenEnv and actually saved into the OpenEnv repo on GitHub. And then the OpenEnv, the PyTorch folks and Hugging Face folks team up and host these in Hugging Face spaces. So these kinds of things are accessible and easy to find if you want to take a look at them and see how you might take them and apply them to your needs. And then again, getting started with RL is actually easier these days. We have the FinQA set up where we have 290 samples that way, and we have our more advanced 79 samples called FinQA reasoning that requires multi-table queries.

7:24

SPEAKER_00

And so there's enough of the reasoning that has to be done across that to make that we've identified that these are harder tasks. And so we have essentially two benchmarks that are built inside of this environment. So that's the setup of how we get this done. We're going to talk about the evals and the results that we got working with this now, given the RL that we just did on this 4 billion parameter model. So we did it. It performs better than the 235 billion parameter model with that RL training loop. And the performance in terms of pass at one was essentially double what it had been percentage-wise in terms of solving problems.

7:49

SPEAKER_00

So it's a very significant uplift that was done with this $500 loop. And again, the right data set is really a key. You want to get the questions and answers to be actually things that are really going to help the model learn. But what is also interesting is what was important about what the model needed to learn. So just to give you a little flavor of what that 4 billion parameter model looks like in terms of how it behaves. And if you recall what we talked about earlier, the 235 billion parameter model tried some queries without knowing what the tables were, didn't find anything, and then hallucinated an answer.

8:11

SPEAKER_00

This 4 billion parameter model, having been fine-tuned on this data set, tries a table and actually first discovers the tables by using the tool getTableName. So the tool existed for the other model as well and it just didn't choose to try it. So the first thing it did was actually query to find out what tables it had available to it. So that's already good. All right. The second thing is then from there it went on to actually inspect the schema. Let me find out what's in that table so I know how to make the right SQL query. And so it does getTableInfo to get the information back to know what to query it. Following that, it runs a query. It actually ran into an error.

8:45

SPEAKER_00

It actually asked for the revenue column, but that column was not actually a part of the data in the table. Given that error, it actually corrected.

9:01

SPEAKER_00

It self-corrected.

9:11

SPEAKER_00

It observed the error, responded to that error by actually correcting to find the actual column that it needed. And so you're seeing both the error correction that it had learned how to do as well as the use of the tools to discover the right information in the first place. So between those two, those behaviors are the real keys to succeeding at these questions. And this is actually something not quite intuitive about where it is that the model was failing. The reality is that what it needed to do, and here it is getting the correct answer. The reality is that what it needed to do was to learn how to use tools.

9:34

SPEAKER_00

A couple of interesting things that go along with it that are more fun and also really useful and good for our situation here. The training data that we talked about at the beginning, there were single table questions, multi-table questions included in the overall data set. And as part of the ablation study, one of the things that they said was, let's take a look and see if we train with single table only. Train with multi-table mixed in so the full data set across both types. Or try to do some curriculum learning and actually start with single table, let the model climb a bit, and then progressively add multi-table.

10:03

SPEAKER_00

And it turned out the single table only training was actually the one that yielded the greatest uplift for these questions. So that was a pleasant surprise. And the other surprising thing was that even though the single table only training regime was the best training regime, the uplift that we see in terms of the model's performance on that harder benchmark that has multi-table questions was a similar doubling in percentage improvement. So the harder multi-table Q&A in the FinQA reasoning question set also saw 13.9 to 26.6 percentage jump after this training.

10:28

SPEAKER_00

So interestingly enough, again, the tool discipline, just knowing how to use the tools that are in the environment, turned out to be a bigger deal than anything else in terms of how to make these models actually get better at what they need to do in this space. So it wasn't the reasoning that was the issue, it was the tool use. We focused only on single step for the best performance and were able to fix that core failure mode. And given that core failure mode being fixed, it turned out that that then made the model better in terms of the improvement generalizing to other question sets.

10:44

SPEAKER_00

And so that means that's the key takeaway from this is that sometimes the idea is to find the specific behavior that's really the problem. And one of the things that, going back to what we do at Snorkel, one of the things that our research team has been talking about a lot lately is building rubrics as part of our evals. And then those rubrics, by breaking down the rightness or wrongness of a model's response into a full list of different questions that can be answered, and looking at each of those individual questions, you can then start to use the rubric as a way to find and intuit and find where the actual problem is among all the multiple possible arenas.

10:57

SPEAKER_00

So instead of simply knowing yes or no at the final, which is good for the RL part, you can use the rubric to help you do an analysis of what are the behaviors that you want to actually generate data sets to help you with. So you make decisions about which data sets you need or which data you want to work with based on what you see coming out of the richer feedback that the rubric gives you. And then the RL still gets your single value, as GRPO usually works with a single value. That's part of how it works. So you use that for the actual RL cycle. So that's the summary of what we did with that. We think it's a really interesting result to know.

11:17

SPEAKER_00

So instead of simply knowing yes or no at the final, which is good for the RL part, you can use the rubric to help you do an analysis of what are the behaviors that you want to actually generate data sets to help you with. So you make decisions about which data sets you need or which data you want to work with based on what you see coming out of the richer feedback that the rubric gives you. And then the RL still gets your single value, as a GRPO just usually works with a single value. That's part of how it works. So you use that for the actual RL cycle. So that's the summary of what we did with that. We think it's a really interesting result to know.

11:45

SPEAKER_00

And the opportunity of what you can do with solving the right questions or the right problems really helps. This link here is a blog post that we have about this. So if you have questions about the details of this particular study and you want to see more about it, you can drill down within that. It also links to a partner post from the Agenteca team over at UC Berkeley. So their post also has additional information you can see from them. And this is the significant thing we wanted to talk about.

12:10

SPEAKER_00

So thank you for your time. And I don't know how much time we have left for questions or not. How are we doing? Are we already at time? Looks like it. Yeah, sorry. Okay. So I'm sorry we don't have questions. I'll hang out right outside if anybody has any follow-up questions they want to ask. And thank you very much. Appreciate it. Thank you. Thank you. It's kind of roll out that if you've worked with something like Harbor before, or if you've worked with like OpenEnv, you're familiar with the same thing about using an environment like this. And this is an environment that we've actually built and published. It's available on Prime Intellect's infrastructure as well.

12:47

SPEAKER_00

So it's something you can load up right there at Prime Intellect. Also on OpenEnv and actually saved into the OpenEnv repo on GitHub. And then the OpenEnv, the PyTorch folks and Hugging Face folks team up and host these in Hugging Face spaces. So these kinds of things are accessible and easy to find if you want to take a look at them and see how you might take them and apply them to your needs. And then again, getting started with RL is actually easier and easier these days. We have the FinQA set up where we have 290 samples that way, and we have our more advanced 79 samples called FinQA reasoning that requires multi-table queries.

13:27

SPEAKER_00

And so there's enough of the reasoning that has to be done across that to make that we've identified that these are harder tasks. And so we have essentially two benchmarks that are built inside of this environment.

13:41

SPEAKER_00

So that's the setup of how we get this done. We're gonna go about talking about the evals and the results that we got working with this now, given the RL that we just did on this 4 billion parameter model.

13:54

SPEAKER_00

So we did it. It performs better than the 235 billion parameter model with that RL training loop. And the performance in terms of pass it one was essentially double what it had been percentage-wise in terms of solving problems. So it's a very significant uplift that was done with this $500 loop. And again, the right data set is really a key. You want to get the questions and answers to be actually things that are really gonna help the model learn. But what is also interesting is what was important about what the model needed to learn. So just to give you a little flavor of what that 4 billion parameter model looks like in terms of how it behaves.

14:37

SPEAKER_00

And if you recall what we talked about earlier, the 235 billion parameter model, tried some queries without knowing what the tables were, didn't find anything, and then hallucinated an answer. This 4 billion parameter model, having been fine-tuned on this data set, tries a table and actually first discover the tables by using the tool getTableName. So the tool existed for the other model as well and it just didn't choose to try it. So the first thing it did was actually query to find out what tables it had available to it. So that's already like when. All right. The second thing is then from there it went on to actually inspect the schema.

15:13

SPEAKER_00

Let me find out what's in that table so I know how to make the right SQL query. And so it's like does getTableInfo to get the information back to know what to query it. Following that, it runs a query. Actually ran into an error. It actually asked for the revenue column, but that column was not actually a part of the data in the table. Given that error, it actually corrected. It self-corrected. It observed the error, responded to that error by actually correcting to find the actual column that it needed. And so you're seeing both the error correction that it had learned how to do as well as the use of the tools to discover the right information in the first place.

16:00

SPEAKER_00

So between those two, those behaviors are the real keys to succeeding at these questions. And this is actually something like maybe not quite intuitive about like where it is that the model was failing. The reality is that what it needed to do, and here it is getting the correct answer. The reality is that what it needed to do was to learn how to use tools.

16:23

SPEAKER_00

A couple of interesting things that go along with it that are more fun and also really useful and good for our situation here. The training data that we talked about at the beginning, there were single table questions, multi-table questions included in the overall data set. And for, as part of the ablation study, one of the things that they said was like, let's take a look and see if we train with single table only. Train with multi-table mixed in so the full data set across both types. Or try to do some curriculum learning and actually start with single table, let the model climb a bit, and then progressively add multi-table.

16:57

SPEAKER_00

And it turned out the single table only training was actually the one that yielded the greatest uplift for these kinds of questions. So that was a nice pleasant surprise. And the other surprising thing was that even though the single table only training regime was the best training regime, the uplift that we see in terms of the model's performance on that harder benchmark that has multi-table questions was a similar doubling in percentage improvement. So the harder multi-table Q&A in the FinQA reasoning question set also saw 13.9 to 26.6 percentage jump after this training.

17:50

SPEAKER_00

So interestingly enough, again, the tool discipline, just knowing how to use the tools that are in the environment, turned out to be a bigger deal than anything else in terms of how to make these models actually get better at what they need to do in this space. So, turned out it wasn't the reasoning that was the issue, it was the tool use. We focused only on single step for the best performance and were able to fix that core failure mode. And given that core failure mode being fixed, it turned out that that then made the model better in terms of the improvement generalizing to other question sets.

18:30

SPEAKER_00

And so that means like that's the key to talk to take away from this is that sometimes the idea is to find the specific behavior that's really the problem. And one of the things that, go back to what we do at Snorkel, one of the things that our research team has been talking about a lot lately is building rubrics as part of our evals. And then those rubrics, by breaking down the rightness or wrongness of a model's response into a full list of different questions that can be answered, and looking at each of those individual questions, you can then start to use the rubric as a way to find and intuit,

19:06

SPEAKER_00

like, and find where the actual problem is among all the multiple possible arenas. So instead of simply knowing yes or no at the final, which is good for the RL part, you can use the rubric to help you do an analysis of what are the behaviors that you want to actually generate data sets to help you with. So you make decisions about which data sets you need or which data you want to work with based on what you see coming out of the richer feedback that the rubric gives you. And then the RL still gets your single value, as a GRPO just usually works with a single value. That's part of how it works. So you use that for the actual RL cycle.

19:45

SPEAKER_00

So that's the summary of what we did with that. We think it's a really interesting result to know. And, you know, again, the opportunity of what you can do with solving the right questions or the right problems really helps. This link here is a blog post that we have about this. So if you have questions about the details of this particular study and you want to see more about it, you can drill down within that. It also links to a partner post from the Agenteca team over at UC Berkeley. So their post also has additional information you can see from them. And this is, you know, the significant thing we wanted to talk about.

20:18

SPEAKER_00

So thank you for your time. And I don't know how much time we have left for questions or not. How are we doing? Are we already at time? Looks like it. Yeah, sorry. Okay. So I'm sorry we don't have questions. I'll hang out right outside if anybody has any follow-up questions they want to ask. And thank you very much. Appreciate it. Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note