[SPEAKER_00] Hi, everyone.
SPEAKER_00
My name is Philip. I'm based out of Germany. I'm part of the Google DeepMind team, mostly working on Gemini API and agents, and we are going to talk about why you should not ship skills without evals. And maybe before we start, I need a little bit of your help. So if you could raise your hands, if you use coding agents to write code. So, yeah, hopefully every hand goes up, right? And do you use skills with it? Okay. Do you have evals for those skills? Okay, yeah, that's not a lot of hands. Everyone uses skills. No one has evals. Hopefully we can fix that today. And very important is that code checks fail in production.
SPEAKER_00
And Skillbench is a very popular and nice eval or benchmark, which indexed over 50,000 skills from GitHub and tried to look into them. And almost none of those skills had evals. Most of them were AI-written, not really tested, and it's very hard to know if your skill is good or bad because agents are really nondeterministic. So you might not know if your task fails because your skill is bad or if your task fails because it's way too challenging for the model. So very important before we go into it, I want to really make sure that we know the difference between the agents we use and the agents we build.
SPEAKER_00
Most of us use agents for writing code, doing productivity work. That's the agents we use. It's, like, anti-gravity, cursor, cloud code, and there you are the engineer, and you have context about skills, right? If you write some prompt to help me build a new Gemini API feature, and if your agent does not invoke the skill on the first time, you will notice it very quickly. You stop your task and reprompt it or use slash commands for triggering those skills. When you build an agent inside your application for consumer or customers, they have no idea about what a skill is.
SPEAKER_00
They don't start their prompt with, use customer support skill to help me refund or use refund skill to help me solve my problems. So there's a big difference between the agents we use and how we use skills and the agents we build and how our customers might want to use skills in the context there. And what is a skill? I mean, every one of us knows, hopefully, by now what a skill is. It's really a folder with a skills MD file in it and then some additional assets to make that skill really work. And the big difference with skills is that they work on progressive disclosure. So most of the skills start very small, so you have the title and the description.
SPEAKER_00
The description is normally part of the model's context, so the model knows when to use the skill. Second layer is we have a skills body with more instructions, more details, and hopefully more references to external files, and then you can really go deep in those reference files where there's all of the context the model needs to discover to solve the task. And I like to differentiate between two kinds of skills. So there are capability skills and preference skills. Capability skills teach models something they cannot do consistently at the moment. Maybe it's tracing some logs, creating a new React app, and those capability skills are temporary.
SPEAKER_00
So the better our model gets, the more likely it is that we can remove those skills, and evals will tell us when we can retire a skill and when not. And then we have preference skills. Those are more durable, mostly encode some preferences. So if you have a specific workflow in your team or a specific style language or other preferences which are very specific to your company, you will have or create preference skills, and those preference skills are then protected with evals because most of the foundation models might not integrate the knowledge which is very specific to your use case or your domain.
SPEAKER_00
And preference skills are very valuable, so we really want to make sure that those are working and we don't update our agents to degrade performance. So do skills work? Yes, they do work. And I'm going back to the skills bench, which has an update of 1.1, which has evaluated all kinds of open and closed models in different harnesses showing that skills, on average, improve the performance by roughly 15%. Skills bench covers around 100 different tasks based on coding and also productivity across different languages. It's openly available. They have a very nice website, a very nice leaderboard, and are also very open for community contributions.
SPEAKER_00
And then they did a second analysis on self-generated or AI-generated skills, right? It's very easy if you are in a coding agent and you work on something, tell the model, create a skill, and then it writes a skill MD file. We maybe look at it very closely. It roughly covers what we want to do, and then we just accept it and start using it. And what I found out is that human-written skills are the best we can provide. AI-generated skills can impact performance negatively, and skills or skills.md files should be below 500 lines of words.
SPEAKER_00
So if you have your laptop open and have a skill available, if you open that and if it's above 500 lines, you should definitely look at the skill after our session. And the last topic about what is a skill and how a skill works,
SPEAKER_00
It roughly covers what we want to do, and then we just accept it and start using it. And what I found out is that human-written skills are the best we can provide. AI-generated skills can impact performance negatively, and that skills or skills.md files should be below 500 lines of words. So if you have your laptop open and have a skill available, if you open that and if it's above 500 lines, you should definitely look at the skill after our session. And the last topic about what is a skill and how a skill works, we have different ways of triggering our skill, right? We can have a model-triggered skill, meaning based on the context and the description the model decides to use or read a skill to get more context to solve a task, and then there are user-invoked skills. And I think people underestimate how powerful user-invoked skills are, and they most of the time just accept the overhead by adding it into the context. I have many user-invoked skills for more workflow type of tasks, like creating a pull request, staging documentation, and all of the very normal dev work which could be run in a script should most likely be a user-invoked skill. And when you build agents for a customer, you don't have those user-invoked skills. We are only working in the model-invoked skills, and that's where we are focusing on for the small eval section we are going to look in a second. So writing skills is an important topic. We're going to look at eight examples on how or tips on how you can write good skills. And most importantly, if you work with model-invoked skills, is the description. Because the description are most of the time two sentences we provide to the system instruction to help the model know when it should use a skill or not. And it's bad if your description is too vague because then it might trigger too often or it might not be triggered if you need it. So very important is the why and the how for the model, so why it should use that skill and then how it should use that skill. Very common is use that skill if you are working on a React application, for example. And then, of course, the when. And we should write directives instead of essays. So we should not say something like, hey, the interactions API is recommended for multichat because it handles session state, and it's you should be way more directive. Like use the interactions API if you are working on a chat application. So you need to give the model clear instructions and directives on when it should use the skill and how it should use the skill. And similar to what we have seen in the skills bench results, we should keep the skill lean and layer information. So the description is the cost you always pay on every model invocation. So on every model call, the description is part of the model context. So you always pay that 100, 200 tokens cost, and you don't want to have a super long description because then you always have to pay that. When you have a very long skill MD file, it will be always read into context when the model decides to read the skill or to use the skill, which can be expensive as well. That's why we want to keep it as concise as possible, but still include all of the reference and details for the model to solve that task. And then, of course, the layer three is we can have those reference files where the model needs to go really deep into a very specific task. And a good example for this is if you are working in maybe a multi-cloud environment, and you have a skill to deploy your application, you might need instruction to deploying to AWS and deploying to Google Cloud. Those should not be part of your skill MD file. Those should be references, that you have a reference for AWS, a reference for Google Cloud, maybe a reference for Azure, so that the model can basically explore based on the context where it should go to get all of that information. Then we should set the right level of freedom. So here I see many people very clearly describing the exact workflow in a skill. Step one, go there. Step two, do this. Step three, do this. If you have those type of use cases, you should not use skills. Maybe you should write a script, because if the process or the workflow is always the same, you don't need to waste models and tokens for that exercise. You can create a script. You can tell the model, use that script to run a specific workflow. So rather define goals and constraints. So if you need to deploy to your update or stage your documentation, describe how the model can do that, or for your database, updating a config, you should not say read the config, update the port, and then deploy again. The model knows what to do. Just, hey, if we need to change the config, here's the file, make the change. Then don't skip negative cases. So we always look at when we want to use the skill, but most of the time we don't look at when we don't want to use the skill. So if we have a description for our skill, which says use it for web development tasks, it might over trigger. Maybe you'll work with React, maybe you will also work with Angular, and the model always loads the skill if you're working in a web development environment. But if you are very specific for, hey, only use that skill for React components or for Tailwind CSS, then the model knows, hey, that's very specific for one to use. And with evals, we can also identify those. And then test early.
SPEAKER_00
So, if we have a description for our skill, which says use it for web development tasks, it might over trigger. Maybe you'll work with React, maybe you will also work with Angular, and the model always loads the skill if you're working in a web development environment. But if you are very specific for, hey, only use that skill for React components or for Tailwind CSS, then the model knows, hey, that's very specific for one use. And with evals, we can also identify those. And then test early. So, that's what we are going to look at. We should really try to test when you create a new skill, always try to create 10 or 20 prompts. I like to create five for the happy path. So, when do I want to use that skill? Five when I don't want to use that skill, just to make sure the model is not over triggering the skill and confusing itself. And then if you have already some customer or production traces, try to include those as well because nothing is better than real-world data. And then tip seven, which is quite new, and I have to give all credits to Matt. So, if you don't know Matt, he's a great AI educator, and you should definitely follow him. He published a tweet and also a skill on eliminating all of the no-ops. And what he found is that AI-generated skills tend to include a lot of no-ops. And no-ops basically is an instruction which does nothing to change the agent's behavior. It's like "before" or "make an implementation easy to read." Like the model knows how or when it should make something easy to read or write clear high-quality code. That's what we expect from the model to do without telling it really. So, definitely look at those no-ops. He has published a very good skill in his skills repository. And then last but not least, know when you should retire a skill. Skills are not there to live forever. Models get better, behaviors change, expectations change, the environment changes. So, always try to run evals with and without the skill enabled. And if the model achieves the performance without even triggering the skill, you know you can retire that skill, save the cost for your tokens. And then also, don't keep redundant skills. Save cost at the end and maintenance as well. And to look at a little bit of a practical example and also how you can create your own small eval or eval harness for skills. Earlier this year, we wanted to create a new skill for the Gemini Interactions API. So, the Gemini Interactions API is our new interface for working with Gemini models and with agents. And the Interactions API was released after the last training of Gemini. So, the model, Gemini 3 and 3.1 or even 3.5, has no context about what is the Gemini Interactions API. So, we decided, okay, let's look at creating a skill to help the model create good code for the Interactions API to use the latest models. And to do that, we created 117 test cases. Those are based on data we see from real users trying to generate Gemini code, from synthetically generated test cases, and also from feedback we see people saying, hey, the model is using Gemini 2.0 even if we are already on 3.0. And the end result was that we improved the performance up to almost 90% for generating valid Interactions API code with the latest Gemini models. And to do this, we basically only needed two very simple assets. So, one of those was a JSON file with all of our test cases. It's very clear structure. It's, hey, we have a prompt. That's basically what we expect the user to provide. We have a language because we wanted to test the skill against TypeScript and Python. We have a should_trigger that's basically there to tell us if the agent should read the skill or not read the skill. And then we have different expected checks. We look at them in a little bit. Those are basically very simple asserts for that prompt if it should trigger or not. And then we have a very basic Python script which runs a coding agent. In this case, it was the Gemini CLI which passes the output and returns it. So, we can take a look at the outcome, whether we have valid code for the Interactions API or not. And most of the tests or evals for skills can be regex. It's very amazing how good regex you can write using coding agents. And for us, it was all about, okay, do we use the correct SDK? Do we use the correct model? Do we use the correct methods? Do we use any old patterns? To look at the whole traces or the whole steps taken. And a very easy case is you just create an LLM as a judge with a rubric on what you want to look at, and then take the output, put it through the LLM as a judge, try to get a pass or a fail. And then if it fails, look at the data, and then try to improve your skill based on that. And that's also how we now evolve skills at Google DeepMind. So we don't use YAML, but as an example. We have tests or evals alongside every skill we have internally at Google DeepMind. Every test has multiple cases with a prompt. We all run them in clear workspaces. So you can define your workspace or environment.
SPEAKER_00
and then take the output, put it through the LLM as a judge, try to get a pass or a fail. And then if it fails, look at the data, and then try to improve your skill based on that. And that's also how we now evolve skills at Google DeepMind. So we don't use YAML, but it's just as an example. We have tests or evals alongside every skill we have internally at Google DeepMind. Every test has multiple cases with a prompt. We all run them in clear workspaces. So you can define your workspace or environment if it should include additional files like your application environments. You have startup commands, which preloads or installs libraries into the environment. And then you have script validators. Those are those red checks where we look at all of the traces to see what the skill triggered, was a certain command run, was a certain CLI run. And then we also have LLM as a judge, where we have some expectations which are matched against, like, hey, did it trigger the skill? Did it run a certain bash command to also evaluate it? And we run them on every change to the skill. So if a change happens or a diff to the skill file, the eval will be run. And there will also be a result. And the change will not be merged if it is not improving the test cases. So we always have those regression tests for every change to the skill. And you can only change the skill if it improves the eval or add new evals. And yeah, that's how we manage it. And then last but not least, 10 examples for best practices for skills. You don't need to take photos. They are in the blog post. I can share later. So we had it many, many times. The skill description is very important. We have seen 50% of the failures because the skill was not triggered correctly because the prompt of the user was not detailed enough for the model to understand, hey, I need to use that skill to solve the task. And especially if you build agents for others, they are not aware of the skill descriptions you have for your model and for your skills. So they might write something very shallow. And then the model needs to know, OK, I need to trigger that skill. We should write directives over passive information. So we should always think about it. You should tell the agent what to do or not what to do and not just, hey, if you feel happy today, please use the skill. Include negative tests. We always forget negative tests. Start small. Even 10 to 20 skill eval samples are better than nothing. You will be surprised at how much you will find even from five to 10 examples. And then definitely create outcomes, not paths. We don't want to test if the model loads the skill on the first turn. We really want to test if it can achieve the task based on the prompt. And if it loads the skill, it loads the skill. If not, then not. If it loads the skill after five turns, that's also OK. Then we want to have isolated runs because coding agents are very good at finding or cheating. So if you run inside your existing environment, it might look up previous chats or it might look up some other executions and then try to cheat it and get the context from the skill without even using the skill. Then definitely run more than one trial when running eval. It's that agents, our models are non-deterministic. Maybe the first one works. The second one doesn't. So always run three to six trials per case and to measure reliability. Test across different harnesses. If you work with or if you have employees or people working with different harnesses, not only just evaluate against code or anti-gravity. If you have people working with cursor, try to include them as well because agent harnesses behave differently and of course model behaves differently. So maybe your skill is very good with Gemini, but very bad with codex. And then you have customers, consumers using your harness with codex and then it fails. And then graduate your eval. So if your model is good enough, it doesn't need the skill anymore, keep that eval. You don't need to throw that eval away because you throw the skill away. You can keep that eval to make sure that the model or the agent keeps the performance. And as soon as you start seeing some degradation, you can reintroduce a skill. You can maybe tweak some other tools or pieces to keep the performance up and then really detect when you can retire skill. And you will be very surprised with all of the model updates, how fast you can retire skill, which you might need it six months ago, but not today anymore. And I have some homework for you. So if you are back from holiday on Monday, pick the most used skill and write five test prompts. You can also use your coding agent and ask it to see, look at your trajectories, which are my most used skills, and then try to create some skills. As you have seen, it's very easy to write your eval harness. It's a JSON or YAML file and then some Python script which runs your coding agent or your agent harness and then look at the outcome. Definitely try to look at removing no ops. Maybe it does not change the eval performance, but it helps you save costs because all of the tokens which are not helpful or not changing the agent behavior are money you will spend. So look at writing create skills from Matt. You can find it on GitHub. And then also run ablation tests. So run always evals with your skill loaded and without your skill loaded. Only that way you will know when you can retire a skill.
SPEAKER_00
Some Python script which runs your coding agent or your agent harness and then look at the outcome. Definitely try to look at removing no ops. Maybe it does not change the eval performance, but it helps you save costs because all of the tokens which are not helpful or not changing the agent behavior are money you will spend. So look at writing create skills from Matt. You can find it on GitHub. And then also run ablation tests. So run always evals with your skill loaded and without your skill loaded. Only that way you will know when you can retire a skill or if a skill is really helpful for your performance. So don't ship skills without evals. Thank you. Thank you. Thank you.
SPEAKER_00
in, like, the context there. And what is a skill? I mean, every one of us knows, hopefully, by now what a skill is. It's, like, basically really a folder with a skills MD file in it and then some additional assets to make that skill really work. And the big difference with skills is that they work on progressive disclosure. So most of the skills start very small, so you have the title and the description. The description is normally part of the model's context, so the model knows when to use the skill. Second layer is we have a skills body with more instructions, more details, and hopefully more references to external files,
SPEAKER_00
and then you can really go deep in those reference files where there's all of the context the model needs to discover to solve the task. And I like to differentiate between two kinds of skills. So there are capability skills and preference skills. Capability skills teach models something they cannot do consistently at the moment. Maybe it's, like, I don't know, like tracing some logs, creating a new React app, and those capability skills are temporary. So the better our model gets, the more likely it is that we can remove those skills, and evals will tell us when we can retire a skill and when not. And then we have preference skills. Those are more durable,
SPEAKER_00
mostly encode some preferences. So if you have a specific workflow in your team or a specific style language or other preferences which are very specific to your company, you will have or create preference skills, and those preference skills are then protected with evals because most of, like, the foundation models might not integrate the knowledge which is very specific to your use case or your domain. And preference skills are very valuable, so we really want to make sure that those are working and we don't, like, update our agents to degradate performance. So do skills work? Yes, they do work. And I'm going back to the skills bench, which has an update of 1.1,
SPEAKER_00
which has evaluated all kinds of open and closed models in different harnesses showing that skills, on average, improve the performance by roughly 15%. Skills bench covers around 100 different tasks based on, like, coding and also productivity across different languages. It's openly available. They have a very nice website, a very nice leaderboard, are also very open for community contributions. And then they did a second analysis on self-generated or AI-generated skills, right? It's very easy if you are in a coding agent and you work on something, tell the model, create a skill, and then it writes a skill MD file. We maybe look at it very closely.
SPEAKER_00
It roughly covers what we want to do, and then we just accept it and start using it. And what I found out is that human-written skills are the best we can provide. AI-generated skills can impact performance negatively, and that skills or skills.md files should be below 500 lines of words. So if you have your laptop open and have a skill available, if you open that and if it's above 500 lines, you should definitely look at the skill after our session. And the last topic about what is a skill and how a skill works, we have different ways of triggering our skill, right? We can have a model-triggered skill, meaning based on the context and the description
SPEAKER_00
the model decides to use or read a skill to get more context to solve a task, and then there are user-invoked skills. And I think people underestimate how powerful user-invoked skills are, and they most of the time just accept the overhead by adding it into the context. I have many user-invoked skills for more workflow type of tasks, like creating a pull request, staging documentation, and all of the very normal dev work which could be run in a script should most likely be a user-invoked skill. And when you build agents for a customer, you don't have those user-invoked skills. We are only working in the model-invoked skills, and that's where we are focusing on
SPEAKER_00
for the small eval section we are going to look in a second. So writing skills is an important topic. We're going to look at eight examples on how or tips on how you can write good skills. And most importantly, if you work with model-invoked skills, is the description. Because the description are most of the time two sentences we provide to the system instruction to help the model know when it should use a skill or not. And it's bad if your description is too vague because then it might trigger too often or it might not be triggered if you need it. So very important is the why and the how for the model, so why it should use that skill
SPEAKER_00
and then how it should use that skill. Very common is like use that skill if you are working on a React application, for example. And then, of course, the when. And we should write directives instead of essays. So we should not say something like, hey, the interactions API is recommended for multichat because it handles like session state, and it's like you should be way more directive. Like use the interactions API if you are working on like a chat application. So you need to give the model like clear instructions and directives on when it should use the skill and how it should use the skill. And similar to what we have seen in the skills bench results,
SPEAKER_00
we should keep the skill lean and layer information. So the description is the cost you always pay on every model invocation. So on every model call, the description is part of the model context. So you always pay that 100, 200 tokens cost, and you don't want to have a super long description because then you always have to pay that. When you have a very long skill MD file, it will be always read into context when the model decides to read the skill or to use the skill, which can be expensive as well. That's why we want to keep it as concise as possible, but still include all of the reference and details for the model to solve that task.
SPEAKER_00
And then, of course, the layer three is like, we can have those reference files where the model needs to like go really deep into a very specific task. And a good example for this is like, if you are working in like maybe a multi-cloud environment, and you have a skill to deploy your application, you might need instruction to deploying to AWS and deploying to Google Cloud. Those should not be part of your skill MD file. Those should be references, that you have a reference for AWS, a reference for Google Cloud, maybe a reference for Azure, so that the model can basically explore based on the context where it should go to get all of that information.
SPEAKER_00
Then we should set the right level of freedom. So, here I see many people very clearly describing the exact workflow in a skill. Step one, go there. Step two, do this. Step three, do this. If you have those type of use cases, you should not use skills. Maybe you should write a script, because if the process or the workflow is always the same, you don't need to waste models and tokens for that exercise. You can create a script. You can tell the model, use that script to run a specific workflow. So, rather define goals and constraints. So, if you need to like deploy to your update or stage your documentation, describe how the model can do that,
SPEAKER_00
or like for your database, updating a config, you should not say like read the config, update the port, and then like deploy again. The model knows what to do. Just like, hey, if we need to change the config, here's the file, make the change. Then don't skip negative cases. So, we always look at the when we want to use the skill, but most of the time we don't look at when we don't want to use the skill. So, if we have a description for our skill, which says use it for web development tasks, it might over trigger. Maybe you'll work with React, maybe you will also work with Angular, and the model always loads the skill if you're working like a web development environment.
SPEAKER_00
But if you are very specific for like, hey, only use that skill for React components or for tail in CSS, then the model knows, hey, that's very specific for one to use. And with evals, we can also identify those. And then test early. So, that's what we are going to look at. We should really try to test when you create a new skill, always to re-eye to create 10 of 20 prompts. I like to create five for like the happy path. So, when do I want to use that skill? Five when I don't want to use that skill, just to make sure the model is not over triggering the skill and confusing itself. And then if you have already some customer
SPEAKER_00
or production traces, try to include those as well because nothing is better than real-world data. And then tip seven, which is quite new, and I have to give all credits to Matt. So, if you don't know Matt, it's a great AI educator, and you should definitely follow him. He published a tweet and also a skill on like killing all of the non-ops. And what he found is that AI-generated skills tend to include a lot of no-ops. And no-ops basically is an instruction which does nothing to change the agent's behavior. It's like before or make an implementation easy to read. Like the model knows how or when it should make something easy to read or write clear high-quality code.
SPEAKER_00
I mean, like that's what we expect from the model to do without telling it really. So, definitely look at those no-ops. He has published a very good skill in his like skills repository. And then last but not least, know when you should retire a skill. Skills are not there to live forever. Models get better, behaviors change, expectation change, the environment changes. So, always try to run evals with and without the skill enabled. And if the model achieves the performance without even like triggering the skill, you know you can retire that skill, save the cost for your tokens. And then also, don't keep like redundant.
SPEAKER_00
So, save cost at the end and maintenance also as well. And to look at a little bit of a practical example and also how you can create your own small eval or eval harness for skills. Earlier this year, we wanted to create a new skill for the Gemini Interactions API. So, the Gemini Interactions API is our new interface for working with Gemini models and with agents. And the Interactions API was released after the last training of Gemini. So, the model or Gemini 3 and like 3.1 or even 3.5 has no context about what is the Gemini Interactions API. So, we decided, okay, let's look at creating a skill to help the model create good code for the Interactions API
SPEAKER_00
to use the latest models. And to do that, we created 117 test cases. Those are based on like data we see from real users trying to generate Gemini code, from synthetic generated test cases, and also from like feedback we see people like, hey, the model is like using Gemini 2.0 even if we are already on 3.0. And the end result was that we improved the performance up to like almost 90% for generating valid Interactions API code with the latest Gemini models. And to do this, we basically only needed like two very simple assets. So, one of that was a JSON file with all of our test cases. And it's very like no clear structure. It's like, hey, we have a prompt.
SPEAKER_00
That's basically what we expect the user to provide. We have a language because we wanted to test the skill against TypeScript and Python. We have a should trigger that's basically there to tell us if the agent should read the skill or not read the skill. And then we have different expected checks. We look at them in a little bit. Those are basically very simple asserts for that prompt if it should trigger or not. And then we have a very basic Python script which runs a coding agent. In this case, it was the Gemini CLI which passes the output and returns it. So, we can like take a look at the outcome, whether we have valid code for the Interactions API or not.
SPEAKER_00
And most of the tests or evals for skills can be regex. It's like very amazing how good of regex you can write using coding agents. And it's really for us, it was all about, okay, do we use the correct SDK? Do we use the correct model? Do we use the correct methods? Do we use any old patterns?
SPEAKER_00
Do we use evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil
SPEAKER_00
to look at the whole traces or the whole steps taken. And a very easy case is you just create an LLM as a judge with a rubric on what you want to look at, and then take the output, put it through the LLM as a judge, try to get a pass or a fail. And then if it fails, look at the data, and then try to improve your skill based on that. And that's also how we now evolve skills at Google DeepMind. So we don't use YAML, but it's just as an example. We have tests or evals alongside every skill we have internally at Google DeepMind. Every test has multiple cases with a prompt. We all run them in clear workspaces. So you can define your workspace or environment
SPEAKER_00
if it should include additional files like your application environments. You have startup commands, which basically preloads or installs libraries into the environment. And then you have script validators. Those are those red checks where we look at all of the traces to see what the skill triggered, was a certain command run, was a certain CLI run. And then we also have LLM as a judge, where we have some expectations which are basically matched against, like, hey, did it trigger the skill? Did it run a certain bash command to also evaluate it? And we run them on every change to the skill. So if a change happens to or like a diff to the skill file, the eval will be run.
SPEAKER_00
And there will also be a result. And the change will not be merged if it is not improving the test cases. So we always have those regression tests for every change to the skill. And you can only change the skill if it improves the eval or add new evals. And yeah, that's how we basically manage it. And then last but not least, 10 examples for best practices for skills. You don't need to take photos. They are in the blog post. I can share later. So I mean, we had it many, many times. The skill description is very important. We have seen 50% of the failures because the skill was not triggered correctly because the prompt of the user was not
SPEAKER_00
detailed enough for the model to understand, hey, I need to use that skill to solve the task. And especially if you build agents for others, they are not aware of the skill descriptions you have for your model and for your skills. So they might write something very shallow. And then the model needs to know, OK, I need to trigger that skill. We should write directives over passive information. So we should always think about it. You should tell the agent what to do or not what to do and not just like, hey, if you feel happy today, please use the skill. Include negative tests. We always forget negative tests. Start small. Even like 10 to 20 skill eval samples
SPEAKER_00
are better than nothing. You will be surprised on how much you will find even from like five to 10 examples. And then like definitely create outcomes, not paths. We don't want to test if the model loads the skill on like the first turn. We really want to test if it can achieve the task based on the prompt. And if it loads the skill, it loads the skill. If not, then not. If it loads the skill after five turns, that's also OK. Then we want to have isolated runs because coding agents are very good at finding or cheating. So if you run inside your existing environment, it might look up previous chats or it might look up some other
SPEAKER_00
executions and then like try to cheat it and get the context from the skill without even using the skill. Then definitely run more than one trial when running eval. It's like agents, our models are non-deterministic. Maybe the first one works. The second one doesn't. So always run three to six trials per case and to measure reliability. Test across different harnesses. If you work with like or if you have employees or people working with like different harnesses, not only just evaluate against code or anti-gravity. If you have people working with cursor, try to include them as well because agent harnesses behave differently and of course model behaves differently.
SPEAKER_00
So maybe your skill is very good with a Gemini, but very bad with codex. And then you have customers, consumers using your harness with codex and then it fails. And then gradiate your eval. So if your model is good enough, it doesn't need the skill anymore, keep that eval. You don't need to throw that eval away because you throw the skill away. You can keep that eval to make sure that the model or the agent keeps the performance. And as soon as you start seeing some degradation, you can reintroduce a skill. You can maybe tweak some other tools or pieces to keep like the performance up and then really detect when you can retire skill.
SPEAKER_00
And you will be very surprised with all of the model updates, how fast you can retire skill, which you might need it like six months ago, but not today anymore. And I have some homework for you. So if you are back from holiday on Monday, pick the most used skill and write five test prompts. You can also use your coding agent and ask it to see, look at your trajectories, which are my most used skills, and then try to create some skills. As you have seen, it's like very easy to write your eval harness. It's like a JSON or YAML file and then like some Python script which runs your coding agent or your agent harness and then like look at the outcome.
SPEAKER_00
Definitely try to look at the removing no ops. Maybe it does not change the eval performance, but it helps you save costs because all of the tokens which are not helpful or not changing the agent behavior are money you will like spend. So look at writing create skills from Matt. It's you can find it on GitHub. And then also run ablation tests. So run always evals with your skill loaded and without your skill loaded. Only that way you will know when you can retire a skill or if a skill is really helpful for your performance. So don't ship skills without evals. Thank you.
Thank you. Thank you.