[SPEAKER_00] Hi everyone, my name is Nunu and I want to talk to you about how we spent the last four months teaching coding agents to master spreadsheets.
SPEAKER_00
So our goal was to get coding agents to be as good at spreadsheets as they are at Python, JavaScript, or whatever your favourite language is. We started at around 50% accuracy on a financial analysis benchmark and got to 92%. So I'll chat about what actually moved the needle and what didn't and the dead ends.
SPEAKER_00
Spreadsheets are a little bit harder for AI than you might think at first. If you think about how you would open Excel, how you'd find your way in an Excel file that you don't know, it's actually a very visual thing and you just instantly see the structure. There's a revenue table here, assumptions in there, a chart in there, and it just feels intuitive and you don't even think about it. And an LLM doesn't really see any of this.
SPEAKER_00
If you ask it what's the revenue, then it has to figure out which revenue do you mean? The net revenue, gross revenue, revenue for this, revenue for that, which quarter, which year, and then is the number it found an actual input? Is it a formula? It's actually a deceptively hard task. One thing we tried close to the beginning was to split the work into three agents. The central one was the edit agent that had a five-step process where you'd define the end state, you'd do a plan, you'd execute, you'd verify all the things you're supposed to do, and this changed the kind of errors we got.
SPEAKER_00
Without it, the agent would just make mistakes while actually building a financial model or something, and with this, it would maybe make those mistakes while planning, which was a lot easier to rectify. But this architecture in the end was too rigid because discovery ran once upfront and then you couldn't revisit it and the context wouldn't flow between the different agents. So it just turned out to be one dead end. Then some more dead ends. We ended up probably trying every conceivable way of representing a spreadsheet to an LLM. None really worked as a standalone representation, but two turned out to be useful as methods inside the REPL that we ended up creating.
SPEAKER_00
But they all had something going for them in theory, and that's why we tried it. SQL has obviously been around for decades, so it's super popular in LLM training data, so agents are really good at it. It's supposed to be a great way to deal with structured data, but it turns out that it doesn't quite work for this. XML is how Excel files are represented on disk, so maybe that was a good idea. It wasn't. And many others. In the end, we did get two useful things out of this. One was the concept of having these CSV or TSV views of part of a spreadsheet.
SPEAKER_00
This turned out to not be that great as the only way to interact with a spreadsheet, but as one piece of the larger solution it turns out to be used very, very often. And HTML was also a step in the right direction as it introduced the idea of layout and formatting. So that ended up resulting in us building a rendering engine to let the agent see what the rendered spreadsheet looked like as an image.
SPEAKER_00
And then eventually we hit on what was probably the biggest breakthrough, which was to replace the many tools that we accumulated over time. I think at that time we had around 15 tools with a single tool, which was Node.js REPL. All the 15 tools that we had to start just became different JavaScript functions that the agent could combine in this one REPL call. Why JavaScript? We needed a scripting language that's easy to sandbox and easy for LLMs. But the actual implementation of the code that deals with the spreadsheet is actually in a completely different language in C sharp. And that's the advantage of this architecture.
SPEAKER_00
You use the scripting language for what it's good at, which is letting the agents interact with it, and use the right language to then deal with the actual files. Before it would be you'd have 10 or 15 tool calls usually for an agent to explore a spreadsheet and get to an answer. And this would actually very often end up timing out and taking a long time because it was doing things sequentially. Even parallel tool calling didn't really help because you couldn't combine the results in any way. After, the agent would just combine the different things it wanted to do in a single tool call and get all the results at the same time.
SPEAKER_00
Some of you will be familiar with the idea of code mode. It showed up in the Anthropic API, and Cloudflare has talked about it. A REPL and code mode are already super useful because that's the basic idea of combining multiple tools into a single tool call. But the REPL actually goes further. The difference is it's basically code mode with persistent state. So the agent calls the REPL tool once, defines a few variables, and then sees the results, spends a few more reasoning tokens. And then the next time it calls the tool, those variables are still there. So it actually can build on its work.
SPEAKER_00
What we observed with this is that code mode without the REPL semantics, agents would very often write quite long scripts. Like 50 lines of JavaScript would be pretty common, which is great because it means they're doing many things at the same time. But with a REPL, they would actually write shorter scripts, which meant it could basically do more interleaving of putting reasoning in between each of the things that the agent was doing, which many times resulted in the agent getting to a better answer faster. And then the next time it calls the tool, those variables are still there. So that means it actually can build on its work.
SPEAKER_00
And what we observed with this is that pure code mode without the REPL semantics, agents would very often write quite long scripts. Like 50 lines of JavaScript would be pretty common. Which is great, means they're doing many things at the same time. But with a REPL, they would actually write shorter scripts, which meant it could basically do more interleaving of putting reasoning in between each of the things that the agent was doing, which many times resulted in the agent getting to a better answer faster, because it was less static.
SPEAKER_00
And another nice thing about this design is in the previous way where we had separate tools, if we figured out, oh, there's a new method we need to give the agent access to, to, I don't know, explore the dependencies between formulas or something. So that would mean creating several more tools that are going to go into the tool schema, and we need to see how they play with each other. Whereas with this approach, all it means is making a few more methods available in the JavaScript REPL, and making the agents aware of that is as simple as creating a TypeScript type definitions file and putting it into the prompt. And that works really well.
SPEAKER_00
So the results out of all of this was, we went from 50% before we have the REPL, then 74%, and then over time we made more changes, none as dramatic as the REPL, that eventually got us to 92% on this internal benchmark we have. And these were changes like giving the agent better fuzzy search or formula tracing functions for dependencies or improving the system prompt or just fixing bugs. But it all adds up to a nice result. And another thing I want to call out is the timeouts.
SPEAKER_00
So this approach ended up really fixing the tasks that would timeout. We usually ran tasks with a five-minute timeout, because if it takes longer than five minutes to answer a question about a spreadsheet, that's not particularly useful. And this approach essentially resulted in zero timeouts, because it just was a lot more efficient for the agent to do its thing. There's a lot of parallels between spreadsheets and coding. I'm sure you all use ClotCode or Codex or whatever coding agent every day. And it does a much better job when it can run the compiler for your language or the linter or your tests and then iterate based on those results.
SPEAKER_00
And when we write code manually, that's true as well, right? If we're not allowed to compile or lint or test the code, then it's not going to produce a great result. And the same is true of spreadsheets. But to enable that for spreadsheet work, we had to build a couple of engines that can close that feedback loop. The two most important ones is one is a formula engine to calculate the formulas. And another one is a render engine to render the contents of a range into an image with all the formatting and layout and so on. And that's the source of truth. It's the verification loop that makes the agent confirm that it did the right thing.
SPEAKER_00
And when it didn't do the right thing, go and fix the formula or go and fix the formatting in order to make it correct. But it only really works if the engine is actually high fidelity. So if you use an incomplete engine that implements, say, 50% of the formulas in Excel, then what you end up with is actually worse results. Because the agent is going to write a formula that it thinks would work and in practice would work. And then it's going to try and compute it. And it's going to get the wrong result or going to get an error because it's not implemented in the engine. So that verification loop is really only as good as the engines that power it.
SPEAKER_00
So this ends up with two different things. One is the REPL and that's an interface. It's how we present our tools to the agent.
SPEAKER_00
And a REPL is the best interface that we could come up with today because coding is what the current state of the art models are the best at. But that's not necessarily going to be true forever, right? The labs are working on computer use a lot. So eventually maybe the models will be as good at computer use with a mouse and keyboard as they are at coding. And at that point, maybe a REPL is not going to be the best interface. But it's the best one today. What won't change is the need for that verification loop. And what's behind it is actually, I think, the more durable part.
SPEAKER_00
Because the more capable the models are, like there have been four or five model releases while we've been doing this work. And every time we've seen the more capable the model is, the more they can get out of that verification loop. And another thing we ended up doing is adding domain knowledge to the prompts. And that actually ended up surviving all of the different iterations of the tools. And it always produced improved results. And this is not so much because the LLMs out of the box don't know what revenue or ARR means. It's more because they know many, many things.
SPEAKER_00
And you need to pigeonhole them a little bit into what you want them to focus on for the specific task that you have. And it's actually super portable. Almost the exact same prompt would work for the REPL or the individual tools or any of the other approaches. I also want to touch a little bit on evaluation. It ended up being a lot of work to evaluate this stuff. And it was actually a really important part of what actually enabled us to be sure whether the CSV or the SQL representation were good is if we can actually evaluate it. And evaluating it correctly turned out to be a bit of a journey as well. We started with LLM as a judge only.
SPEAKER_00
And that works to some extent and sometimes it's the only option you really have. But the annoying part is sometimes you can't really tell if when a score changes is it because the agent changed something or the evaluator changed what it outputs. So we ended up doing a bunch of work to replace it with deterministic comparisons wherever that was possible, which is not always possible. But where we could, for instance, take a golden spreadsheet that had a set of inputs And evaluating it correctly turned out to be a bit of a journey as well.
SPEAKER_00
We started with LLM as a judge only. And that works to some extent and sometimes it's the only option you really have. But the annoying part is sometimes you can't really tell if when a score changes is it because the agent changed something or the evaluator changed what it outputs. So we ended up doing a bunch of work to replace it with deterministic comparisons wherever that was possible, which is not always possible.
SPEAKER_00
But where we could, for instance, take a golden spreadsheet that had a set of inputs and a set of outputs and then use that as a black box to test a spreadsheet that the model produced saying, hey, if you put some numbers into these inputs, you get something out of these outputs.
SPEAKER_00
And then you put the same numbers into the spreadsheet the model produced and you see if you get the same outputs. And that ends up being sometimes more trustworthy than just using an LLM to grade that work. And as with everything, there's bugs and infrastructure bugs when you're building agents many times end up looking like reasoning failures. And it may seem like the model is doing something wrong. But actually many times it turns out it's a bug where you just have a bug in the code or the skill or the prompt has the wrong example and the model is following that very faithfully. Or there's actually a bug in the tools and they fail.
SPEAKER_00
And then the model keeps retrying and it seems like the model is done, but it's just trying to work around the issue.
SPEAKER_00
So there's a lot of juice to get out of just really looking at those traces and seeing what is going wrong and trying to figure out is this the model not getting it quite right or is this something we can actually fix? So I wanted to end with a summary of what I think generalizes to other tasks. And I think the first thing is if your agent is making many sequential tool calls or even parallel tool calls, then you've invented a bad scripting language. So you might as well just give the agent a real one and that can be code mode or REPL or whatever you want. The second one is I think feedback loops really matter.
SPEAKER_00
And if you happen to be working in a domain where you can build those feedback loops with existing tools, then great. Less work for you. But if you're in a domain where those feedback loops don't actually exist, I think it's actually really worth spending the time to build that rendering engine or calculation engine or whatever applies to your particular domain. Okay. The third is I think interfaces are super important. And as I explained before, the REPL really changed the results we got. So you should really spend the time figuring out what the best interface is.
SPEAKER_00
But you should expect to have to revisit that because the capability of the models is going to keep changing. And as they get better at other things, you may find that the best interface is something else and you need to find what that next one is. Next, I think we shouldn't really underestimate the power of planning and think before you act. And yes, sometimes the simple things really do make a difference. So spend the time on those as well. Domain knowledge, I think, is really important. And you really need to spend a bunch of time thinking about what's the things that you need to remind the model about. It's not so much teaching the model.
SPEAKER_00
It's more reminding it to pay more attention to that than other things. And lastly, evaluation. I think the more you can do deterministic evaluation, the better, which doesn't mean that you should avoid LLM as a judge. It just means if that's the only option you have, then that's exactly what you should do. But if you can evaluate in some other way, do that. And always check your traces and your plumbing because sometimes agent confusion is just bugs and you should fix that.
SPEAKER_00
Thank you. One is the REPL and that's an interface. It's how we present our tools to the agent. And, you know, a REPL is the best interface that we could come up with today because coding is what the current state of the art models are the best at. But that's not necessarily going to be true forever, right? The labs are working on computer use a lot. So, you know, eventually maybe the models will be as good at computer use with a mouse and keyboard as they are at coding. And at that point, maybe a REPL is not going to be the best interface.
SPEAKER_00
But it's the best one today. What won't change is the need for, you know, for that verification loop. And what's behind it is actually, I think, the more durable part. Because the more capable the models are, like there have been, I don't know, four or five model releases while we've been doing this work. And every time we've seen the more capable the model is, the more they can get out of that verification loop.
SPEAKER_00
And another thing we ended up doing is adding domain knowledge to the prompts. And that actually ended up, you know, surviving all of the different iterations of the tools. And it always, you know, produced improved results. And this is not so much because the LLMs out of the box don't know what, I don't know, revenue or ARR means. It's more because they, you know, know many, many things. And you kind of need to pigeonhole them a little bit into what you want them to focus on for the specific task that you have. And it's actually super portable. Like this, almost the exact same prompt would work for the REPL or the individual tools or any of the other approaches.
SPEAKER_00
I also want to touch a little bit on evaluation. It ended up being a lot of work to evaluate this stuff. And it was actually, you know, a really important part of what actually enabled us to be sure whether, you know, the CSV or the SQL representation were good is if we can actually evaluate it. And evaluating it correctly turned out to be a bit of a journey as well. We started with LLM as a judge only. And, you know, that works to some extent and sometimes it's the only option you really have.
SPEAKER_00
But the annoying part is sometimes you can't really tell if when a score changes is it because the, you know, the agent changed something or the evaluator changed what it outputs. So we ended up doing a bunch of work to replace it with deterministic comparisons wherever that was possible, which is not always possible. But where we could, for instance, you know, take a golden spreadsheet that had a set of inputs and a set of outputs and then use that as kind of a black box to test a spreadsheet that the model produced saying, hey, if you put some numbers into these inputs, you get something out of these outputs.
SPEAKER_00
And then you put the same numbers into the spreadsheet the model produced and you see if you get the same outputs. And that ends up being, you know, sometimes more trustworthy than just using an LLM to grade that work.
SPEAKER_00
And as with everything, there's bugs and infrastructure bugs when you're building agents many times end up looking like, you know, reasoning failures. And it may seem like, oh, the model is doing something wrong. And, but actually many times it turns out, you know, it's a, it's a bug where you're, you just have a bug in the code or the skill or the prompt has the wrong example and the model is following that very faithfully. Or there's actually, you know, a bug in the tools and they fail. And then the model keeps retrying and it seems like the model is being done, but, you know, it's just trying to work around the issue.
SPEAKER_00
So there's, you know, a lot of juice to get out of just really looking at those traces and seeing what is going wrong and trying to figure out is this, you know, the model not getting it quite right or is this something we can actually fix?
SPEAKER_00
So I wanted to kind of end with a summary of what I think the model is going to be done. I think generalizes to other tasks. And I think the first thing is if your agent is making many sequential tool calls or even parallel tool calls, then you've kind of invented a bad stripping language. So you might as well just give the agent a real one and that can be code mode or REPL or whatever you want. The second one is I think feedback loops really matter. And if you happen to be working in a domain where you can build those feedback loops with existing tools, then great. Less work for you.
SPEAKER_00
But if you're in a domain where those feedback loops don't actually exist, I think it's actually really worth spending the time to build that rendering engine or calculation engine or whatever applies to your particular domain. Okay. The third is I think interfaces are super important. And as, you know, as I explained before, the REPL really changed the results we got. So you should really spend the time figuring out what the best interface is. But you should expect to have to revisit that because the models, the capability of the models is going to keep changing.
SPEAKER_00
And as they get better at other things, you may find that the best interface is something else and you need to find what that next one is. Next, I think, you know, we shouldn't really underestimate the power of planning and think before you act. And, yeah, sometimes the simple things really do make a difference. So don't, you know, actually spend the time on those as well. Domain knowledge, I think, is really important. And you really need to spend a bunch of time thinking about what's the things that you need to remind the model about. It's not so much teaching the model. It's more reminding it to pay more attention to that than other things.
SPEAKER_00
And lastly, yeah, evaluation. I think the more you can do deterministic evaluation, the better, which doesn't mean that you should, you know, avoid LLM as a judge. It just means if that's the only option you get, then that's exactly what you should do. But if you can evaluate in some other way, do that. And always check your traces and your plumbing because sometimes agent confusion is just bugs and you should fix that. Thank you. . . . . . .
. .
. . .