[SPEAKER_00] OK.
SPEAKER_00
OK, great. Thank you. Then I think we could start. So my name is Sibbrahim. I will share with you the lessons that we learned through our evals of coding agents and different models on real world software engineering tasks using, as an example, our suite Rebench leaderboard. I want to share some practical lessons mostly. And I think that evals matter now even more than before because we have a lot of models, closed source, open-weight models that are doing really great in the software engineering domain. And of course, you can rely on your gut feeling, vibe checks, or maybe one or two of your most favorite questions to choose between the options.
SPEAKER_00
But everything is fun until you roll out something into the production. And it just breaks down. And clients are unhappy. So I think that we need to evaluate everything. And before we will deep dive, I want to share a small fact about me. So actually, I have a very non-traditional background for AI research. I'm a dentist by training. That's me 10 years ago. And that's why on my Google Scholar, I have papers from NERIPS and ICML about ARL and test time scaling along with some psychotherapy or medical insurance problems in dentistry. And in medicine, cost of every mistake is really high.
SPEAKER_00
And I think that for the AI domain, we also could say that cost of each mistake is higher than traditional software engineering. And actually, I should say that I believe that dental pain and infrastructural pain are similar because both of them will not let you sleep at night. But with dental pain, you could go to the dentist and he will cure you. But with the infrastructural pain, you need to do something about it by yourself. So about our leaderboard, let's break down word by word. What do we do? So SWEERI Bench, it's fresh real-world software engineering task on 30 models, evaluated every month. So what does it mean fresh?
SPEAKER_00
Most of the benchmarks during their release, they release questions and solutions. So implicitly or explicitly, this data can become a part of the pre-training of the next generation of models. So if you want to build some open, truly decontaminated benchmark, time splits are the only way. That's why every month we collect only fresh problems from the previous month and then assess the model's capabilities. In terms of the real world, in pre-LLM era, there were a lot of benchmarks about, for example, some brackets, sequence, or ordering correctly adjectives in English. But now we need some natural problems that people could ask systems to do.
SPEAKER_00
And even more, some well-paid problems like software engineering, for example. Also, software engineering problems and tasks are not about just simple question answering. They are truly subtasks. So it means that to solve the issue or implement the feature, you need to understand the structure of repository. You need to try to write some tests, implement the solution, run the test, reproduce the mistakes or bugs. And also, it is some multi-turn and naturally long context task. So it's not just concatenating some text or books. No, it's truly long context. And also, it is about tool use, harnesses.
SPEAKER_00
So that's why I believe that software engineering domain is really valuable for evaluations. We also evaluate something like 30 models with the same harness, simple same harness. And for the reference, we also give some numbers for CloudCode, Codex, and Juni harnesses. And we'll add actually more and report a lot of stuff. And I always read all the comments on local llama subreddit and x and try to add most actual and interesting models. Of course, we get requests like, OK, can you please evaluate some obliterated role play, 69 billion parameters agent, but we mostly stick to the most popular ones. About the tasks.
SPEAKER_00
For any verifiable software engineering task, actually, we have three main components. It's similar for Sweebench, SweeRebench, other domains, TerminalBanch. You have some task description. For us, it's just original issue title and description from the given time frame from some permissive but popular open source repository. For the sandbox, you can call it environment or environment sandbox snapshot. But basically, it's just an executable Docker image with the installed dependencies so we could run the test of the project. And the third one is a verifier. Basically, it's just a test from the pull request that solved some issue or implemented some feature.
SPEAKER_00
And here, I could say that there is actually two sets of tests: failed to pass. It is the test that should be failed before solving the issue, for example, and should be passed after. And pass to pass, it's something like regression test. And also, it's important to say that every test is not just a question, but mostly some Docker image, one or 10 gigabyte, so you need good infrastructure actually to run everything. I think that this is one of the most important slides. I will share the presentation on X or I could send you. But the thing is that every month, we verify every task.
SPEAKER_00
And we have a really big bank of the problems with the task, because I believe that it is not too easy to say what does it look like, a perfect task. But we can say what makes it bad. So for problem description, you actually need something balanced, not too vague, not too over-specified, not too easy, not too hard, because for too easy problems, all the models will solve it, and your effective size of benchmark will be less. For the verifier and test, here's one of the examples. So usually, software engineers write the test after implementing some solution, so they may be some kind of over-fitted.
SPEAKER_00
Here, for example, tests require the agent to generate exact substring in the error message. So even with the correct solution, these tests will not be passed. And you need a stable infrastructure. So for problem description, you actually need something balanced, not too vague, not too over-specified, not too easy, not too hard, because for too easy problems, all the models will solve it, and your effective size of benchmark will be less.
SPEAKER_00
For the verifier and test, here's one of the examples. So usually, software engineers write the test after implementing some solution, so they may be some kind of over-fitted. Here, for example, tests require the agent to generate exact substring in the error message. So even with the correct solution, these tests will not be passed.
SPEAKER_00
And you need a stable infrastructure, because you need to minimize the infrastructural noise during your runs. For example, your test could connect to some external resources, and it will be some dependency. Or we had a problem in one of pipelines, so several images just got some default time, like 1970s, and some tests were relied on that. So we just got some problems with these kind of evaluations.
SPEAKER_00
In my opinion, for our benchmark, collection is mostly a filtering problem, because we have a really good source of task information like GitHub. We use GitHub Archive as main source for pull requests and issues for large-scale projects, and just GitHub API for the smaller ones. Here, 100% is number of pull requests linked with some issues. So, for example, if you need a lot more data for pre-training runs, for example, post-training runs, if you will use just pull requests, it will be eight times bigger data set.
SPEAKER_00
We use interactive agent to install all the dependencies and project, so we could use this Docker image. And we also have some several steps of just LLMs and filtering with the most common problems. But at the end, we try to choose sample that is 10% bigger than we need in our final runs, because after running some models, you could face problems in terms of task quality that could be visible only after agents will try to solve it. And for the final set of tasks, we manually verify. I think it's one full-time day of work to manually verify each task, so we could make sure that they are solvable, but quite challenging.
SPEAKER_00
Here is the slide about our hardness and agent. I believe that it is better to have some minimalistic agent with strong infrastructure than having over-engineering agent with weak infrastructure. It's an example of the most popular tools and bash commands in our scaffold with Claude Opus 4.6. So, with uppercase, it is agents tools, and lowercase, it's bash commands. And actually, the most popular ones are quite simple. And we also run our agent in a no-loop setup, so it means that we don't want our agent to ask some clarification questions or anything like that, so you just need to solve the issue. And we start with some simple ReAct plus demonstration that you have in your prompt, demonstration how to use your tools. But nowadays, every model is quite good in tool calling, so we just minimize our context as well.
SPEAKER_00
So, about what breaks in practice with the agents. I think that every month we have one or two model runs that just became invalid because of some problems. First of all, you need to define your retry policy. You actually want to separate your errors of the model and some infrastructural errors. So, you need to define what exit stats. For example, too long context or too many tool calls or your provider errors. Will you rerun these runs or not?
SPEAKER_00
For the caching, it actually really improves your cost efficiency. I hope you know about that. Here's an example with our simple agent. It's very similar to software engineering agent or mini three agent, but three bench creators. So, with the caching included, your cost will be like four times less. But for Cloud Code, it actually spends a lot of tokens. So, even with prompt caching and Haiku sub-agents for some sub-tasks, will actually cost quite a lot.
SPEAKER_00
And we, after one of the runs, we saw that during the updates of the models, even within the same family, for example, like GPT-4o.2, GPT-4o.4, or the longer or older versions, there could be some default parameters drifting for the reasoning level, for the caching level, or other stuff that you also need to make sure that is relevant and work in your infrastructure. That's why I believe that, first of all, you need to try to run some external benchmark, like SWE-bench and any other terminal bench, on your infrastructure to make sure that your numbers and reported numbers match, and only then start to do your experiments.
SPEAKER_00
Here's the most favorite slides. So, we found at least two ways how models cheat. First one is a well-known issue. It is all about Cloud Code here, but it will be also about codecs and other models as well. So, the thing is that during our runs, before, when we build our Docker image, we do a checkout to the base commit before the solution was implemented. So, agent will start doing something there. And if you will run command git log with all flag, then you will get access to the overall git history. So, that's how, for example, Cloud Code just looked up to the future, to the solution patch, and copy-pasted it. And so, successfully solved this issue.
SPEAKER_00
After that, we remove all the future git history, because previous git history might be helpful to get some context working with the issue, but we need to remove the future one. After that, Cloud Code came up with the webfetch tool. It has a webfetch tool, so it just went to GitHub repository, original one, to see the conversation in the original issue, pull request, and solved it.
SPEAKER_00
After that, we restricted webfetch tool. So, Cloud Code, okay, I have curl. Let's just use bash command with curl. We'll go to the original issue. Here, you can see that Cloud Code also formatted the conversation to be more convenient. And then, just checked the original test in the main and solved the issue. So, when models get better, I believe that they might tend to cheat even more and do some reward hacking. So, we solve only with some kind of post-processing and trajectory analysis and try to come up with new solutions as well.
SPEAKER_00
I think that one of the main reasons why we made this benchmark and maintain it, we want to share some practical value with the real AI engineers and AI creators. So, that's why we report not only some mean resolved metric, but also tokens per problem, price per problem. And we do five runs for each task to report some confidence intervals and also pass at five. Something like if a model solved each task at least, we think that it's successful. To give some kind of potential of the model. Also, you can check something like pass at five if you need reliability.
SPEAKER_00
So we solve only with some kind of post-processing and trajectory analysis and try to come up with new solutions as well. I think that one of the main reasons why we made this benchmark and maintain it is we want to share some practical value with the real AI engineers and AI creators. So that's why we report not only some mean resolved metric, but also tokens per problem, price per problem. And we do five runs for each task to report some confidence intervals and also pass at five. Something like if a model solved each task at least, we think that it's successful. To give some kind of potential of the model.
SPEAKER_00
Also, you can check something like pass all five if you need reliability. So you will mark the task as successful only if agents solve it in all five runs. After some analytics in terms of economics, tokens, and price per problem, we also want to do something on trajectory level. Because I think that it is a source of a lot of insights about how some models work in our or external harnesses. And the next one is about if you know how to make evaluation or benchmark, you could use the same pipeline to collect some validation set, for example. And to think about training. And I don't say about SFT or RL.
SPEAKER_00
At first, you can just try with choosing between models, harnesses, and parameters on your validation set. And then maybe do some kind of auto research or just update your prompts and tools. Then do some simple rejection sampling, fine tuning, or distilling from the bigger models. And then move to more complex strategies like GRPO. So we use the same pipeline that we use for SWE Bench to make two big open source releases. First one is SWE Bench. We released it last year. It is something like 30,000 RL environments like real-world software engineering tasks with Docker images. And it was used by some frontier labs to train better models.
SPEAKER_00
And now we also release SWE Bench V2. It is something about software engineering tasks on 20 programming languages. Also a lot of Docker images, a lot of tasks that could be used for training. I will work on adoption for it. We also have an adapter for Harbor, our terminal bench, which is quite convenient format to run any evaluations or training. And I think that for the future, we need to think about more long horizon tasks, more about something complex, and something about code quality as well. Because if you will check any patch from SWE Bench submission or SWE Bench submission, you will see some problems that actually the real developers will not do.
SPEAKER_00
And during the review, you will say that, okay, it's not how things work actually. For example, Gemini, GLEM, GPT models, they tend to produce some generated tests or files and then just don't remove it. We also can talk about some code quality during the pull request. So, yeah, I think that we need to come up with some long horizon tasks, more trajectory analysis, and then move on to training better models. So, yeah, that's it. Please check the leaderboards, SWE Bench leaderboard. Update every month. I will be here. Feel free to reach out. This is my X handle, and I will release like a new open source project and also will share these slides, I think, tomorrow.
SPEAKER_00
Yeah, thank you for your attention. For example, your test could connect to some external resources, and it will be some dependency. Or we had a problem in one of pipelines, so several images just get some default time, like 1970s, and some tests were relied on that. So we just get some problems with these kind of evaluations. In my opinion, for our benchmark, collection is mostly a filtering problem, because we have a really good source of task information like GitHub. We use GitHub Archive as main source for pull requests and issues for large-scale projects, and just GitHub API for the smaller ones. Here, 100% is number of pull requests linked with some issues.
SPEAKER_00
So, for example, if you need a lot more data for pre-training runs, for example, post-training runs, if you will use just pull requests, it will be eight times bigger data set. We use interactive agent to install all the dependencies and project, so we could use this Docker image. And we also have some several steps of just LLMS and just filtering with the most common problems. But at the end, we try to choose sample that is 10% bigger than we need in our final runs, because after running some models, you could face problems in terms of task quality that could be visible only after agents will try to solve it. And for the final set of tasks, we manually verify.
SPEAKER_00
I think it's one full-time day of work to manually verify each task, so we could make sure that they are solvable, but quite challenging. Here is the slide about our hardness and agent. I believe that it is better to have some minimalistic agent with strong infrastructure than having over-engineering agent with weak infrastructure. It's an example of the most popular tools and bash commands in our scaffold with Claude Opus 4.6. So, with uppercase, it is agents tools, and lowercase, it's bash commands. And actually, the most popular ones, it's quite simple. And we also run our agent in Yola setup, so it means that we don't want our agent to ask some clarification questions
SPEAKER_00
or something like that, so you just need to solve the issue. And we start with some simple React plus demonstration that you have in your prompt, demonstration how to use your tools. But nowadays, every model is quite good in tool calling, so we just minimize our context as well. So, about what breaks in practice with the agents. I think that every month we have one or two model runs that just became invalid because of some problems. First of all, you need to define your retry policy. You actually want to separate your errors of the model and some infrastructural errors. So, you need to define what exit stats.
SPEAKER_00
For example, too long context or too many tool cores or your provider errors. Will you rerun these runs or not? For the caching, it actually really improves your cost efficiency. I hope you know about that. Here's an example with our simple agent. It's very similar to software engineering agent or mini three agent, but three bench creators. So, with the caching included, your cost will be like four times less. But for Cloud Code, it actually spends a lot of tokens. So, even with turnout caching and like Haiku sub-agents for some sub-tasks, will actually cost quite a lot.
SPEAKER_00
And we, after one of the runs, we saw that during the updates of the models, even within the same family, for example, like GPT-5.2, GPT-5.4, or the longer or more older versions, there could be some default parameters drifting for the reasoning level, for the caching level, or other stuff that you also need to make sure that is relevant and work in your infrastructure. That's why I believe that, first of all, you need to try to run some external benchmark, like Sweebench and any other terminal bench, on your infrastructure to make sure that actually your numbers and reported numbers match, and only then start to do your experiments. Here's the most favorite slides.
SPEAKER_00
So, we found at least two ways how models cheat. First one is a well-known issue. It is all about Cloud Code here, but it will be also about codecs and other models as well. So, the thing is that during our runs, before, when we build our Docker image, we do a checkout to the base commit before the solution was implemented. So, agent will start doing something there. And if you will run command git log with all flag, then you will get access to the overall git history. So, that's how, for example, Cloud Code just look up to the future, to the solution patch, and copy-paste it. And so, successfully solved this issue.
SPEAKER_00
After that, we remove all the future git history, because previous git history might be helpful to get some context working with the issue, but we need to remove the future one. After that, Cloud Code came up with the webfetch tool. It has a webfetch tool, so it just went to GitHub repository, original one, to see the conversation in the original issue, pull request, and solved it. Okay. After that, we restricted webfetch tool. So, Cloud Code, okay, I have curl. Let's just use bash command with curl. We'll go to the original issue. Here, you can see that actually Cloud Code also formatted the conversation to be more convenient.
SPEAKER_00
And then, just check the original test in the main and solve the issue. So, when models get better, actually, I believe that they might tend to cheat even more and do some reward hacking. So, we solve only with some kind of post-processing and trajectory analysis and try to come up with new solutions as well. I think that one of the main reasons why we made this benchmark and maintain it, we want to share some practical value with the real AI engineers and AI creators. So, that's why we report not only some mean resolved metric, but also tokens per problem, price per problem. And we do five runs for each task to report some confidence intervals and also pass at five.
SPEAKER_00
Something like if a model solved each task at least, we think that it's successful. To give some kind of potential of the model. Also, you can check something like post all five if you need reliability. So, you will mark the task as successful only if agents solve it in all five runs. After some analytics in terms of economics, tokens, and price per problem, we also want to do something on trajectory level. Because I think that it is a source of a lot of insights about how some models work in our or external harnesses. And the next one is about if you know how to make evaluation or benchmark, you could use the same pipeline to collect some validation set, for example.
SPEAKER_00
And to think about training. And I don't say about like SFT or RL. At first, you can just try with choosing between models, harnesses, and parameters on your validation set. And then maybe do some kind of auto research or just update your prompts and tools. Then do some simple rejection sampling, fight tuning, or distillating from the bigger models. And then move to more complex strategies like GRPO. So, we use the same pipeline that we use for SWE Rebench to make two big open source releases. First one is SWE Rebench. We released it last year. It is something like 30,000 of RL environments like real-world software engineering tasks with Docker images.
SPEAKER_00
And it was used by some frontier labs to train better models. And now we also release SWE Rebench V2. It is something about software engineering tasks on 20 programming languages. Also a lot of Docker images, a lot of tasks that could be used for training. I will work on adoption for it. We also have an adapter for Harbor, our terminal bench, which is quite convenient format to run any evaluations or the training. And I think that for the future, we need to think about more long horizon tasks, more about something complex, and something about code quality as well.
SPEAKER_00
Because if you will check any patch from SWE Bench submission or SWE Rebench submission, you will see some problems that actually the real developers will not do. And during the review, you will say that, okay, it's not how things work actually. For example, Gemini, GLEM, GPD models, they tend to produce some reproduced tests or files and then just don't remove it. We also can talk about some code quality during the poll request. So, yeah, I think that we need to come up with some long horizon tasks, more trajectory analysis, and then move on to training better models. So, yeah, that's it. Please check the leaderboards, SWE Rebench leaderboard. Update every month.
SPEAKER_00
I will be here. Feel free to reach out. This is my ex hand rule, and I will release like new open source project and also will share these slides, I think, tomorrow. Yeah, thank you for your attention.