Open Reader

Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs

completed 19:08 Jul 08, 2026 Watch on YouTube

Current Status

completed

Video ID

HEFSExa0xl0

RAG / Chat

Enabled
Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs
Description

https://github.com/witanlabs/research-log

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Replacing 15 discrete tools with a single JavaScript REPL (persistent state code mode) plus high-fidelity formula/render engines improved coding-agent spreadsheet accuracy from 50% to 92% and eliminated timeouts.
  • Why it matters: This demonstrates a generalizable architecture pattern for building high-performance domain-specific agents: REPL > tool calling, feedback loops via domain engines, and deterministic eval beats LLM-as-judge.
  • Best use: Blueprint for Ken's agent system design; directly applicable to financial modeling, data ops, or any domain requiring sequential reasoning over structured artifacts.

Executive Summary

Nuno Campos from Witan Labs spent four months improving coding agents on spreadsheet tasks, achieving 92% accuracy (up from 50%) on a financial analysis benchmark while eliminating timeouts. The breakthrough was replacing ~15 discrete tools with a single Node.js REPL that maintains persistent state across calls. This let agents write shorter, more iterative scripts instead of monolithic 50-line blocks, interleaving reasoning between actions. The REPL is backed by C# engines for formula calculation and rendering (image output), creating a verification loop analogous to compile-test-fix cycles in software development.

Spreadsheets are deceptively hard for LLMs: humans instantly parse visual layout (revenue table, assumptions, charts), but agents must disambiguate 'revenue' (net vs. gross, quarter, formula vs. input) without spatial context. Early attempts included multi-agent architectures (discovery, edit, verify) and alternative representations (SQL, XML, CSV, HTML)—all failed or proved too rigid. CSV/TSV views and HTML rendering survived as methods inside the REPL but not as standalone representations.

The REPL approach consolidated tool calls (previously 10–15 sequential calls per task) into single composite invocations, dramatically reducing latency. Persistent state allowed agents to define variables once and reuse them, building incrementally. Adding new capabilities became trivial: expose methods in JavaScript and document via TypeScript definitions in the prompt, rather than creating new tools and rebalancing the tool schema.

Evaluation evolved from LLM-as-judge to deterministic black-box testing (golden input/output spreadsheets) wherever possible, because LLM judge score changes were ambiguous (agent improvement vs. evaluator drift). Infrastructure bugs often masqueraded as reasoning failures; trace inspection revealed tool bugs, wrong prompt examples, or plumbing issues rather than model confusion. Domain knowledge prompts (e.g., financial concepts) survived all architecture iterations and consistently improved results by focusing the model's vast knowledge on the task context.

Key Takeaways

  • Claim: The single biggest breakthrough was replacing 15 tools with one Node.js REPL, jumping accuracy from 50% to 74% immediately. | Evidence: Before: 10–15 sequential tool calls per task, frequent timeouts. After: one composite REPL call combining operations, zero timeouts, agents writing shorter iterative scripts (not monolithic 50-line blocks). | Caveat: REPL is the best interface today because LLMs excel at coding, but this may change if computer-use capabilities (mouse/keyboard) match coding proficiency in future model releases. | Implication: If Ken's agents make many sequential or parallel tool calls, he's 'invented a bad scripting language'—switching to a REPL or code-mode interface will likely yield immediate gains. | Timestamp: 05:30
  • Claim: Persistent state in the REPL (vs. ephemeral code mode) enables incremental work: agents define variables once and build on them across calls. | Evidence: Agents wrote shorter scripts with REPL semantics, interleaving reasoning between tool invocations, versus long 50-line scripts in pure code mode. | Caveat: None stated; speaker positions this as strictly superior to code mode for multi-step reasoning tasks. | Implication: Ken should prioritize stateful execution environments for agents performing multi-step analysis or iterative workflows (financial modeling, data pipelines). | Timestamp: 06:45
  • Claim: High-fidelity formula and rendering engines are essential for the verification loop; incomplete engines (e.g., 50% formula coverage) produce worse results. | Evidence: Agents write correct formulas, engine returns errors or wrong results due to missing implementations, agent retries incorrectly. Full-fidelity engines close the compile-test-fix loop. | Caveat: Building domain-specific engines (formula calculation, image rendering) is substantial engineering work; speaker uses C# for spreadsheet logic, JavaScript only for LLM interface. | Implication: Ken should invest in high-fidelity domain engines (not just LLM wrappers) if building agents for specialized domains lacking existing tooling—short-term cost, compounding long-term gains. | Timestamp: 08:15
  • Claim: Adding new capabilities to the REPL is trivial: expose JavaScript methods and update TypeScript definitions in the prompt, no tool schema rebalancing required. | Evidence: Previously, adding a new tool (e.g., formula dependency tracing) meant creating multiple tools and testing interactions; now it's just new methods plus a TypeScript definition file. | Caveat: This assumes the REPL architecture is already in place; initial build is non-trivial (C# engines, sandboxed Node.js runtime). | Implication: For Ken's agent systems, designing for extensibility via in-REPL methods rather than tool proliferation will reduce iteration friction as capabilities expand. | Timestamp: 07:20
  • Claim: Deterministic evaluation (black-box input/output testing) beats LLM-as-judge when possible; LLM judge scores are ambiguous (agent changed vs. evaluator changed). | Evidence: Golden spreadsheets with fixed inputs/outputs used to test agent-produced spreadsheets; same inputs should yield same outputs, removing evaluator variance. | Caveat: Deterministic eval isn't always feasible (some tasks lack ground truth); LLM-as-judge is acceptable when it's the only option. | Implication: Ken should prioritize deterministic benchmarks for agent evaluation wherever domain artifacts allow (code correctness, data outputs, financial calculations) and use LLM judge only as a fallback. | Timestamp: 12:30
  • Claim: Domain knowledge prompts (e.g., financial concepts like ARR, revenue types) survived all architectural iterations and consistently improved results. | Evidence: Same domain prompts worked across REPL, individual tools, CSV, SQL approaches; not teaching the model but focusing its attention on task-relevant knowledge. | Caveat: None stated; speaker frames this as 'pigeonholing' the model's vast knowledge rather than adding new information. | Implication: Ken should invest in domain-specific prompt engineering (financial terms, GTM concepts, operator frameworks) as a portable, high-ROI improvement across agent architectures. | Timestamp: 10:45
  • Claim: Infrastructure bugs often look like reasoning failures; trace inspection revealed tool bugs, wrong prompt examples, or plumbing issues rather than model confusion. | Evidence: Agents retrying failed tools appeared as reasoning loops; fixing bugs (code, examples, tool reliability) resolved apparent model confusion. | Caveat: Requires disciplined trace analysis and willingness to debug infrastructure rather than blame the model. | Implication: Ken should establish systematic trace review processes and assume infrastructure bugs before attributing failures to model limitations—low-hanging fruit for agent performance. | Timestamp: 13:10

Detailed Brief

Why spreadsheets are hard for LLMs and early architecture attempts

  • Claims: Humans parse spreadsheet structure visually (revenue table, assumptions, charts) instantly; LLMs must disambiguate 'revenue' (net/gross, quarter/year, formula/input) without spatial context.; Early multi-agent architecture (discovery, edit with five-step process, verify) changed error types—planning mistakes instead of execution mistakes—but was too rigid (discovery ran once, context didn't flow between agents).
  • Evidence: Opening Excel: humans see structure instantly; LLMs face disambiguation (which revenue? which quarter? formula vs. input?).; Five-step edit process (define end state, plan, execute, verify, iterate) shifted mistakes from execution to planning, easier to debug but ultimately a dead end due to rigidity.
  • Caveats: Multi-agent approach failed because discovery was one-shot and context didn't propagate.; Alternative representations (SQL, XML, CSV, HTML) all had theoretical merit (SQL is popular in training data, XML is Excel's on-disk format) but none worked standalone.
  • Implications: Domain-specific visual parsing is a core challenge for agents working with structured artifacts; spatial/layout information matters.; Rigid agent orchestration (fixed phases, no backtracking) fails for exploratory tasks; agents need dynamic context flow and iterative discovery.

The REPL breakthrough and persistent state advantage

  • Claims: Node.js REPL replaced ~15 tools, consolidating operations into single composite calls; accuracy jumped from 50% to 74% immediately.; Persistent state (variables survive across REPL calls) enabled incremental work; agents wrote shorter scripts with interleaved reasoning, not monolithic 50-line blocks.; REPL is code mode with state: agents define variables once, reason, then extend work in subsequent calls.
  • Evidence: Before REPL: 10–15 sequential tool calls, frequent five-minute timeouts, even parallel tool calling didn't help (couldn't combine results).; After REPL: zero timeouts, agents wrote shorter iterative scripts, all 15 prior tools became JavaScript functions callable in one invocation.; Comparison: pure code mode produced 50-line scripts (do everything at once); REPL produced shorter scripts (reason between steps).
  • Caveats: REPL is best today because current models excel at coding; if computer-use (mouse/keyboard) capabilities catch up, REPL may not remain optimal.; Implementation uses JavaScript for LLM interface but C# for actual spreadsheet logic (sandbox JavaScript, use the right language for the domain).
  • Implications: If agents make many sequential/parallel tool calls, switch to REPL or code mode—it's a qualitatively better interface for current models.; Persistent state accelerates multi-step reasoning tasks; Ken should prioritize stateful execution for financial modeling, data pipelines, iterative workflows.; Separation of concerns (scripting language for LLM, domain language for logic) is a durable design pattern for performance-critical agent systems.

High-fidelity engines and the verification loop

  • Claims: Formula calculation engine and render engine (image output) close the verification loop, analogous to compile-test-fix in software development.; More capable models extract more value from the verification loop (four or five model releases during project, each improved loop utilization).; Incomplete engines (e.g., 50% formula coverage) produce worse results: agents write correct formulas, engine errors/wrong results cause retry loops.
  • Evidence: Agent writes formula it thinks will work (and would in practice), tries to compute, gets error or wrong result due to missing implementation, enters retry loop.; Analogy: coding agents work better when they can compile, lint, test; spreadsheet agents need analogous domain feedback.; C# engines for formula calculation and rendering; JavaScript REPL is interface layer only.
  • Caveats: Building high-fidelity domain engines is substantial work; no shortcuts if domain lacks existing tooling.; Verification loop only as good as the engines; low-fidelity engines mislead the agent.
  • Implications: Ken should invest in domain-specific verification engines (not just LLM wrappers) for specialized workflows—short-term cost, compounding accuracy gains.; The verification loop is the durable part; REPL is today's best interface, but feedback mechanisms will remain valuable as model capabilities evolve.; For domains with existing tools (code compilers, data validators), leverage them; for novel domains, building engines is worth it.

Evaluation evolution and trace-driven debugging

  • Claims: LLM-as-judge alone is ambiguous: score changes could mean agent improved or evaluator drifted; replaced with deterministic black-box testing wherever possible.; Golden spreadsheets (fixed inputs/outputs) used to test agent-produced sheets; same inputs should yield same outputs.; Infrastructure bugs often masquerade as reasoning failures; trace inspection revealed tool bugs, wrong prompt examples, plumbing issues.
  • Evidence: Started LLM-as-judge only, couldn't isolate agent vs. evaluator changes.; Black-box testing: put numbers into golden inputs, check outputs; put same numbers into agent sheet, compare outputs.; Example bugs: tools failing, agents retrying (looks like reasoning loop); wrong prompt examples, agents following faithfully (looks like confusion).
  • Caveats: Deterministic eval not always possible; LLM-as-judge acceptable when it's the only option.; Requires discipline to check traces and assume infrastructure bugs before blaming the model.
  • Implications: Ken should prioritize deterministic benchmarks for agent evaluation (code correctness, data outputs, financial calculations) and use LLM judge as fallback.; Systematic trace review should be a core discipline—low-hanging fruit for performance gains by fixing infrastructure rather than prompts.; Agent confusion is often a symptom of bugs, not model limitations; fix plumbing first.

Generalizable lessons and incremental gains

  • Claims: Domain knowledge prompts (financial concepts, ARR, revenue types) survived all architecture iterations and consistently improved results.; Adding new REPL capabilities is trivial: expose methods, update TypeScript definitions in prompt (no tool schema rebalancing).; Incremental improvements (fuzzy search, formula tracing, prompt refinement, bug fixes) added up to final 92% accuracy (74% post-REPL, then iterative gains to 92%).
  • Evidence: Same domain prompts worked across REPL, tools, CSV, SQL approaches; not teaching but focusing model attention.; Previously: new tool meant multiple new entries, test interactions; now: new methods in REPL, TypeScript definitions.; Path: 50% → 74% (REPL) → 92% (fuzzy search, tracing, prompts, bugs).
  • Caveats: Domain prompts are about focusing vast model knowledge, not adding new information.; Incremental gains require systematic iteration and evaluation infrastructure to measure.; Planning and 'think before you act' still matter; simple improvements shouldn't be neglected.
  • Implications: Ken should invest in domain-specific prompt engineering as a portable, high-ROI improvement (financial, GTM, operator concepts).; REPL architecture reduces iteration friction for capability expansion—design agents for extensibility via in-REPL methods.; Systematic iteration on prompts, tools, and debugging (not just architecture rewrites) compounds to significant gains.

Notable Concepts & Terms

  • REPL (Read-Eval-Print Loop) with persistent state: Code mode that retains variables across tool calls, enabling incremental work and shorter iterative scripts instead of monolithic code blocks; the breakthrough interface for the spreadsheet agent.
  • Verification loop (compile-test-fix analogy): Feedback mechanism where agents execute actions (formulas, rendering), observe results via domain engines, and iterate to correct—analogous to software development workflow; only effective with high-fidelity engines.
  • High-fidelity domain engines: Formula calculation and rendering engines (in C#) that accurately implement spreadsheet semantics; low-fidelity (e.g., 50% formula coverage) misleads agents into retry loops.
  • Deterministic black-box evaluation: Testing agent outputs by comparing results for fixed inputs/outputs (golden spreadsheets) rather than LLM-as-judge; removes evaluator variance and isolates agent performance.
  • Domain knowledge prompts (pigeonholing): Prompts that focus the model's vast knowledge on task-relevant concepts (e.g., financial terms) rather than teaching new information; portable across architectures and consistently improve results.
  • Tool call consolidation: Pattern where many discrete tools are replaced by a single composable interface (REPL/code mode), eliminating sequential call overhead and enabling complex operations in one invocation.
  • Infrastructure bugs masquerading as reasoning failures: Agent confusion or retry loops that appear to be model limitations but are actually tool bugs, wrong prompt examples, or plumbing issues; fixable via trace inspection.

Operator Notes / Why Ken Should Care

  • Direct blueprint for building high-performance domain agents: REPL > tool calling, feedback loops via high-fidelity engines, deterministic eval, systematic trace review.
  • Financial modeling agent architecture applicable to Ken's content/business workflows (e.g., GTM analysis, ARR forecasting, data pipeline agents).
  • Extensibility pattern: design for in-REPL methods rather than tool proliferation to reduce iteration friction as agent capabilities expand.
  • Evaluation strategy: prioritize deterministic benchmarks for measurable workflows (code, data, financial outputs); use LLM-as-judge only as fallback.
  • Debugging discipline: assume infrastructure bugs before attributing failures to model limitations; systematic trace review is low-hanging fruit for performance.
  • Domain prompts as a portable, high-ROI investment: focus model attention on task-relevant concepts (financial, GTM, operator frameworks) across architectures.
  • Persistent state (REPL semantics) accelerates multi-step reasoning; critical for iterative workflows Ken's agents will encounter (analysis, content generation, investment research).
  • Separation of concerns: use scripting language for LLM interface (JavaScript), domain-optimized language for logic (C#)—Ken can apply this to agent system architecture.
  • Verification loops become more valuable as models improve (four releases during project, each extracted more value); invest in feedback mechanisms now for compounding returns.
  • Timeouts as a product constraint: five-minute limit for spreadsheet questions; Ken should establish similar latency targets for agent workflows and optimize ruthlessly.

Watch Map

  • 00:00: Intro: 50% → 92% accuracy on financial analysis benchmark over four months
  • 01:15: Why spreadsheets are hard: visual parsing vs. LLM disambiguation
  • 02:30: Dead end: multi-agent architecture (discovery, edit, verify) too rigid
  • 03:45: Dead ends: SQL, XML, CSV, HTML representations; CSV/TSV and HTML survived as methods
  • 05:30: Breakthrough: Node.js REPL replaces 15 tools, 50% → 74% accuracy, zero timeouts
  • 06:45: Persistent state advantage: shorter iterative scripts vs. 50-line monoliths
  • 07:20: Adding REPL capabilities: expose methods, TypeScript definitions, no tool schema rebalancing
  • 08:15: High-fidelity engines: formula calculation and rendering close verification loop
  • 09:30: Verification loop analogy: compile-test-fix for spreadsheets
  • 10:45: Domain knowledge prompts: portable across architectures, focus model attention
  • 12:30: Evaluation: LLM-as-judge → deterministic black-box testing (golden spreadsheets)
  • 13:10: Infrastructure bugs masquerade as reasoning failures; trace inspection critical
  • 14:30: Generalizable lessons: REPL > tool calling, feedback loops, interfaces will evolve, deterministic eval
  • 16:45: Incremental gains: fuzzy search, tracing, prompts, bugs → 74% to 92%
  • 18:30: Summary and Q&A

Source/Metadata

  • Title: Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs
  • Transcript words: 4012
  • Duration seconds: 1148
  • Timestamp note: Approximate timestamps provided based on transcript structure and typical talk pacing; chapter markers not explicitly present in transcript

Transcript

2746 words en Processed in 197.6s

[SPEAKER_00] Hi everyone, my name is Nunu and I want to talk to you about how we spent the last four months teaching coding agents to master spreadsheets. So our goal was to get coding agents to be as good at spreadsheets as they are at Python, JavaScript, or whatever your favourite language is. We started at around 50% accuracy on a financial analysis benchmark and got to 92%. So I'll chat about what actually moved the needle and what didn't and the dead ends. Spreadsheets are a little bit harder for AI than you might think at first. If you think about how you would open Excel, how you'd find your way in an Excel file that you don't know, it's actually a very visual thing and you just instantly see the structure. There's a revenue table here, assumptions in there, a chart in there, and it just feels intuitive and you don't even think about it. And an LLM doesn't really see any of this. If you ask it what's the revenue, then it has to figure out which revenue do you mean? The net revenue, gross revenue, revenue for this, revenue for that, which quarter, which year, and then is the number it found an actual input? Is it a formula? It's actually a deceptively hard task. One thing we tried close to the beginning was to split the work into three agents. The central one was the edit agent that had a five-step process where you'd define the end state, you'd do a plan, you'd execute, you'd verify all the things you're supposed to do, and this changed the kind of errors we got. Without it, the agent would just make mistakes while actually building a financial model or something, and with this, it would maybe make those mistakes while planning, which was a lot easier to rectify. But this architecture in the end was too rigid because discovery ran once upfront and then you couldn't revisit it and the context wouldn't flow between the different agents. So it just turned out to be one dead end. Then some more dead ends. We ended up probably trying every conceivable way of representing a spreadsheet to an LLM. None really worked as a standalone representation, but two turned out to be useful as methods inside the REPL that we ended up creating. But they all had something going for them in theory, and that's why we tried it. SQL has obviously been around for decades, so it's super popular in LLM training data, so agents are really good at it. It's supposed to be a great way to deal with structured data, but it turns out that it doesn't quite work for this. XML is how Excel files are represented on disk, so maybe that was a good idea. It wasn't. And many others. In the end, we did get two useful things out of this. One was the concept of having these CSV or TSV views of part of a spreadsheet. This turned out to not be that great as the only way to interact with a spreadsheet, but as one piece of the larger solution it turns out to be used very, very often. And HTML was also a step in the right direction as it introduced the idea of layout and formatting. So that ended up resulting in us building a rendering engine to let the agent see what the rendered spreadsheet looked like as an image. And then eventually we hit on what was probably the biggest breakthrough, which was to replace the many tools that we accumulated over time. I think at that time we had around 15 tools with a single tool, which was Node.js REPL. All the 15 tools that we had to start just became different JavaScript functions that the agent could combine in this one REPL call. Why JavaScript? We needed a scripting language that's easy to sandbox and easy for LLMs. But the actual implementation of the code that deals with the spreadsheet is actually in a completely different language in C sharp. And that's the advantage of this architecture. You use the scripting language for what it's good at, which is letting the agents interact with it, and use the right language to then deal with the actual files. Before it would be you'd have 10 or 15 tool calls usually for an agent to explore a spreadsheet and get to an answer. And this would actually very often end up timing out and taking a long time because it was doing things sequentially. Even parallel tool calling didn't really help because you couldn't combine the results in any way. After, the agent would just combine the different things it wanted to do in a single tool call and get all the results at the same time. Some of you will be familiar with the idea of code mode. It showed up in the Anthropic API, and Cloudflare has talked about it. A REPL and code mode are already super useful because that's the basic idea of combining multiple tools into a single tool call. But the REPL actually goes further. The difference is it's basically code mode with persistent state. So the agent calls the REPL tool once, defines a few variables, and then sees the results, spends a few more reasoning tokens. And then the next time it calls the tool, those variables are still there. So it actually can build on its work. What we observed with this is that code mode without the REPL semantics, agents would very often write quite long scripts. Like 50 lines of JavaScript would be pretty common, which is great because it means they're doing many things at the same time. But with a REPL, they would actually write shorter scripts, which meant it could basically do more interleaving of putting reasoning in between each of the things that the agent was doing, which many times resulted in the agent getting to a better answer faster. And then the next time it calls the tool, those variables are still there. So that means it actually can build on its work. And what we observed with this is that pure code mode without the REPL semantics, agents would very often write quite long scripts. Like 50 lines of JavaScript would be pretty common. Which is great, means they're doing many things at the same time. But with a REPL, they would actually write shorter scripts, which meant it could basically do more interleaving of putting reasoning in between each of the things that the agent was doing, which many times resulted in the agent getting to a better answer faster, because it was less static. And another nice thing about this design is in the previous way where we had separate tools, if we figured out, oh, there's a new method we need to give the agent access to, to, I don't know, explore the dependencies between formulas or something. So that would mean creating several more tools that are going to go into the tool schema, and we need to see how they play with each other. Whereas with this approach, all it means is making a few more methods available in the JavaScript REPL, and making the agents aware of that is as simple as creating a TypeScript type definitions file and putting it into the prompt. And that works really well. So the results out of all of this was, we went from 50% before we have the REPL, then 74%, and then over time we made more changes, none as dramatic as the REPL, that eventually got us to 92% on this internal benchmark we have. And these were changes like giving the agent better fuzzy search or formula tracing functions for dependencies or improving the system prompt or just fixing bugs. But it all adds up to a nice result. And another thing I want to call out is the timeouts. So this approach ended up really fixing the tasks that would timeout. We usually ran tasks with a five-minute timeout, because if it takes longer than five minutes to answer a question about a spreadsheet, that's not particularly useful. And this approach essentially resulted in zero timeouts, because it just was a lot more efficient for the agent to do its thing. There's a lot of parallels between spreadsheets and coding. I'm sure you all use ClotCode or Codex or whatever coding agent every day. And it does a much better job when it can run the compiler for your language or the linter or your tests and then iterate based on those results. And when we write code manually, that's true as well, right? If we're not allowed to compile or lint or test the code, then it's not going to produce a great result. And the same is true of spreadsheets. But to enable that for spreadsheet work, we had to build a couple of engines that can close that feedback loop. The two most important ones is one is a formula engine to calculate the formulas. And another one is a render engine to render the contents of a range into an image with all the formatting and layout and so on. And that's the source of truth. It's the verification loop that makes the agent confirm that it did the right thing. And when it didn't do the right thing, go and fix the formula or go and fix the formatting in order to make it correct. But it only really works if the engine is actually high fidelity. So if you use an incomplete engine that implements, say, 50% of the formulas in Excel, then what you end up with is actually worse results. Because the agent is going to write a formula that it thinks would work and in practice would work. And then it's going to try and compute it. And it's going to get the wrong result or going to get an error because it's not implemented in the engine. So that verification loop is really only as good as the engines that power it. So this ends up with two different things. One is the REPL and that's an interface. It's how we present our tools to the agent. And a REPL is the best interface that we could come up with today because coding is what the current state of the art models are the best at. But that's not necessarily going to be true forever, right? The labs are working on computer use a lot. So eventually maybe the models will be as good at computer use with a mouse and keyboard as they are at coding. And at that point, maybe a REPL is not going to be the best interface. But it's the best one today. What won't change is the need for that verification loop. And what's behind it is actually, I think, the more durable part. Because the more capable the models are, like there have been four or five model releases while we've been doing this work. And every time we've seen the more capable the model is, the more they can get out of that verification loop. And another thing we ended up doing is adding domain knowledge to the prompts. And that actually ended up surviving all of the different iterations of the tools. And it always produced improved results. And this is not so much because the LLMs out of the box don't know what revenue or ARR means. It's more because they know many, many things. And you need to pigeonhole them a little bit into what you want them to focus on for the specific task that you have. And it's actually super portable. Almost the exact same prompt would work for the REPL or the individual tools or any of the other approaches. I also want to touch a little bit on evaluation. It ended up being a lot of work to evaluate this stuff. And it was actually a really important part of what actually enabled us to be sure whether the CSV or the SQL representation were good is if we can actually evaluate it. And evaluating it correctly turned out to be a bit of a journey as well. We started with LLM as a judge only. And that works to some extent and sometimes it's the only option you really have. But the annoying part is sometimes you can't really tell if when a score changes is it because the agent changed something or the evaluator changed what it outputs. So we ended up doing a bunch of work to replace it with deterministic comparisons wherever that was possible, which is not always possible. But where we could, for instance, take a golden spreadsheet that had a set of inputs And evaluating it correctly turned out to be a bit of a journey as well. We started with LLM as a judge only. And that works to some extent and sometimes it's the only option you really have. But the annoying part is sometimes you can't really tell if when a score changes is it because the agent changed something or the evaluator changed what it outputs. So we ended up doing a bunch of work to replace it with deterministic comparisons wherever that was possible, which is not always possible. But where we could, for instance, take a golden spreadsheet that had a set of inputs and a set of outputs and then use that as a black box to test a spreadsheet that the model produced saying, hey, if you put some numbers into these inputs, you get something out of these outputs. And then you put the same numbers into the spreadsheet the model produced and you see if you get the same outputs. And that ends up being sometimes more trustworthy than just using an LLM to grade that work. And as with everything, there's bugs and infrastructure bugs when you're building agents many times end up looking like reasoning failures. And it may seem like the model is doing something wrong. But actually many times it turns out it's a bug where you just have a bug in the code or the skill or the prompt has the wrong example and the model is following that very faithfully. Or there's actually a bug in the tools and they fail. And then the model keeps retrying and it seems like the model is done, but it's just trying to work around the issue. So there's a lot of juice to get out of just really looking at those traces and seeing what is going wrong and trying to figure out is this the model not getting it quite right or is this something we can actually fix? So I wanted to end with a summary of what I think generalizes to other tasks. And I think the first thing is if your agent is making many sequential tool calls or even parallel tool calls, then you've invented a bad scripting language. So you might as well just give the agent a real one and that can be code mode or REPL or whatever you want. The second one is I think feedback loops really matter. And if you happen to be working in a domain where you can build those feedback loops with existing tools, then great. Less work for you. But if you're in a domain where those feedback loops don't actually exist, I think it's actually really worth spending the time to build that rendering engine or calculation engine or whatever applies to your particular domain. Okay. The third is I think interfaces are super important. And as I explained before, the REPL really changed the results we got. So you should really spend the time figuring out what the best interface is. But you should expect to have to revisit that because the capability of the models is going to keep changing. And as they get better at other things, you may find that the best interface is something else and you need to find what that next one is. Next, I think we shouldn't really underestimate the power of planning and think before you act. And yes, sometimes the simple things really do make a difference. So spend the time on those as well. Domain knowledge, I think, is really important. And you really need to spend a bunch of time thinking about what's the things that you need to remind the model about. It's not so much teaching the model. It's more reminding it to pay more attention to that than other things. And lastly, evaluation. I think the more you can do deterministic evaluation, the better, which doesn't mean that you should avoid LLM as a judge. It just means if that's the only option you have, then that's exactly what you should do. But if you can evaluate in some other way, do that. And always check your traces and your plumbing because sometimes agent confusion is just bugs and you should fix that. Thank you. One is the REPL and that's an interface. It's how we present our tools to the agent. And, you know, a REPL is the best interface that we could come up with today because coding is what the current state of the art models are the best at. But that's not necessarily going to be true forever, right? The labs are working on computer use a lot. So, you know, eventually maybe the models will be as good at computer use with a mouse and keyboard as they are at coding. And at that point, maybe a REPL is not going to be the best interface. But it's the best one today. What won't change is the need for, you know, for that verification loop. And what's behind it is actually, I think, the more durable part. Because the more capable the models are, like there have been, I don't know, four or five model releases while we've been doing this work. And every time we've seen the more capable the model is, the more they can get out of that verification loop. And another thing we ended up doing is adding domain knowledge to the prompts. And that actually ended up, you know, surviving all of the different iterations of the tools. And it always, you know, produced improved results. And this is not so much because the LLMs out of the box don't know what, I don't know, revenue or ARR means. It's more because they, you know, know many, many things. And you kind of need to pigeonhole them a little bit into what you want them to focus on for the specific task that you have. And it's actually super portable. Like this, almost the exact same prompt would work for the REPL or the individual tools or any of the other approaches. I also want to touch a little bit on evaluation. It ended up being a lot of work to evaluate this stuff. And it was actually, you know, a really important part of what actually enabled us to be sure whether, you know, the CSV or the SQL representation were good is if we can actually evaluate it. And evaluating it correctly turned out to be a bit of a journey as well. We started with LLM as a judge only. And, you know, that works to some extent and sometimes it's the only option you really have. But the annoying part is sometimes you can't really tell if when a score changes is it because the, you know, the agent changed something or the evaluator changed what it outputs. So we ended up doing a bunch of work to replace it with deterministic comparisons wherever that was possible, which is not always possible. But where we could, for instance, you know, take a golden spreadsheet that had a set of inputs and a set of outputs and then use that as kind of a black box to test a spreadsheet that the model produced saying, hey, if you put some numbers into these inputs, you get something out of these outputs. And then you put the same numbers into the spreadsheet the model produced and you see if you get the same outputs. And that ends up being, you know, sometimes more trustworthy than just using an LLM to grade that work. And as with everything, there's bugs and infrastructure bugs when you're building agents many times end up looking like, you know, reasoning failures. And it may seem like, oh, the model is doing something wrong. And, but actually many times it turns out, you know, it's a, it's a bug where you're, you just have a bug in the code or the skill or the prompt has the wrong example and the model is following that very faithfully. Or there's actually, you know, a bug in the tools and they fail. And then the model keeps retrying and it seems like the model is being done, but, you know, it's just trying to work around the issue. So there's, you know, a lot of juice to get out of just really looking at those traces and seeing what is going wrong and trying to figure out is this, you know, the model not getting it quite right or is this something we can actually fix? So I wanted to kind of end with a summary of what I think the model is going to be done. I think generalizes to other tasks. And I think the first thing is if your agent is making many sequential tool calls or even parallel tool calls, then you've kind of invented a bad stripping language. So you might as well just give the agent a real one and that can be code mode or REPL or whatever you want. The second one is I think feedback loops really matter. And if you happen to be working in a domain where you can build those feedback loops with existing tools, then great. Less work for you. But if you're in a domain where those feedback loops don't actually exist, I think it's actually really worth spending the time to build that rendering engine or calculation engine or whatever applies to your particular domain. Okay. The third is I think interfaces are super important. And as, you know, as I explained before, the REPL really changed the results we got. So you should really spend the time figuring out what the best interface is. But you should expect to have to revisit that because the models, the capability of the models is going to keep changing. And as they get better at other things, you may find that the best interface is something else and you need to find what that next one is. Next, I think, you know, we shouldn't really underestimate the power of planning and think before you act. And, yeah, sometimes the simple things really do make a difference. So don't, you know, actually spend the time on those as well. Domain knowledge, I think, is really important. And you really need to spend a bunch of time thinking about what's the things that you need to remind the model about. It's not so much teaching the model. It's more reminding it to pay more attention to that than other things. And lastly, yeah, evaluation. I think the more you can do deterministic evaluation, the better, which doesn't mean that you should, you know, avoid LLM as a judge. It just means if that's the only option you get, then that's exactly what you should do. But if you can evaluate in some other way, do that. And always check your traces and your plumbing because sometimes agent confusion is just bugs and you should fix that. Thank you. . . . . . . . . . . .