Open Reader

Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind

completed 21:45 Jul 14, 2026 Watch on YouTube

Current Status

completed

Video ID

0vphxNt4wyk

RAG / Chat

Enabled
Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind
Description

There are thousands of agent skills. Almost none of them are tested. They get vibe-checked with two manual runs, maybe a thumbs-up from a colleague, then shipped. You wouldn't merge code without tests — so why are we shipping skills without evals? This talk covers the full lifecycle of building reliable agent skills: what a skill actually is (and isn't), how to write one that triggers correctly, and how to build a lightweight eval harness that catches failures before your users do. ### Philipp Schmid Staff Engineer · Google DeepMind [X/Twitter](https://x.com/_philschmid) · [LinkedIn](https://www.linkedin.com/in/philipp-schmid-a6a2bb196/) · [Website](https://www.philschmid.de/) · [Blog](https://www.philschmid.de) Philipp Schmid is a Staff Engineer at Google DeepMind working on Gemini and Gemma. His work focuses on helping developers build and benefit from AI responsibly. ## About This Session — [View on the schedule](https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_07_01_breakout_track_05_1545_2026_06_26t08_57_00_975z)

Summary

Generated by gpt-5.6-sol

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Treat every agent skill as production code: pair it with outcome-based regression evals, test both triggering and non-triggering cases, and use ablation tests to prove the skill still adds value.
  • Why it matters: Skills can improve agent performance by roughly 15% on average, but poorly written or incorrectly triggered skills can reduce quality, waste context, and create regressions that nondeterministic agent behavior makes difficult to diagnose.
  • Best use: Use the video as a practical blueprint for establishing a skill-evaluation standard, CI merge gates, and a lifecycle process for improving or retiring skills across agent systems.

Executive Summary

Schmid argues that skills should be managed like tested software rather than accepted as plausible-looking Markdown. Skillbench indexed more than 50,000 GitHub skills and found that almost none had evals, despite many being AI-generated. Without controlled tests, teams cannot tell whether a failure comes from a bad skill, a triggering error, the model, or the difficulty of the task.

The highest-risk area is model-invoked skills in customer-facing agents. Internal developers can notice a missed skill and explicitly invoke or reprompt it, but customers do not know the skill catalog or its expected wording. DeepMind observed that about 50% of failures were caused by incorrect triggering, making precise descriptions, directives, and negative boundaries central parts of skill design.

The recommended eval stack is deliberately lightweight: start with 10–20 prompts containing positive and negative cases, add production traces, execute each case in an isolated workspace for three to six trials, and score outcomes with deterministic checks or an LLM judge. DeepMind's Gemini Interactions API skill used 117 cases across Python and TypeScript and raised valid-code performance to almost 90%.

Skills also need lifecycle management. Run every eval with and without the skill, gate skill changes on regression results, and retain the eval even after retiring the skill. Capability skills should disappear as models improve; preference skills encoding company-specific workflows or standards are more durable but still require protection against model and harness changes.

Key Takeaways

  • Claim: A skill without evals creates an attribution problem: agent nondeterminism makes it unclear whether failures come from the skill, its triggering behavior, the model, or an inherently difficult task. | Evidence: Skillbench indexed more than 50,000 skills from GitHub; Schmid says almost none had evals and most were AI-written rather than systematically tested. | Implication: Ken should require a measurable baseline and regression suite before treating any reusable skill as production-ready. | Caveat: The transcript cites Skillbench's findings but does not provide the study's sampling methodology or exact proportion of unevaluated skills.
  • Claim: Skill triggering is a first-class reliability problem for customer-facing agents because users do not know which skills exist or how to request them explicitly. | Evidence: DeepMind saw approximately 50% of failures arise because the skill did not trigger correctly, often when the user's prompt was too shallow for the model to infer the needed skill. | Implication: OpenClaw and other customer-facing agent systems should evaluate routing recall and false-positive activation independently from the quality of the skill's instructions. | Caveat: The 50% figure comes from the speaker's internal experience and is not presented as a universal rate across all agent architectures.
  • Claim: Skills are useful on average, but automatically generating and accepting them can make agents worse. | Evidence: Skillbench 1.1 reportedly found an average performance improvement of roughly 15% across about 100 coding and productivity tasks, while its separate analysis found human-written skills performed best and AI-generated skills could negatively affect performance. | Implication: AI should help draft skills, but acceptance should depend on comparative eval results rather than readability or apparent completeness. | Caveat: The average spans different models, harnesses, languages, and tasks; it does not establish that every skill or deployment will improve by 15%.
  • Claim: Effective skills use progressive disclosure and precise directives instead of placing all knowledge into a long, passive instruction file. | Evidence: Schmid proposes three layers: a short title and description always exposed to the model, a concise skill body loaded when selected, and deeper reference files such as separate AWS, Google Cloud, and Azure deployment guidance. He also recommends reviewing skills.md files over roughly 500 lines. | Implication: Ken should optimize skill architecture for routing precision and context cost, keeping provider- or domain-specific material behind references that agents load only when needed. | Caveat: The 500-line threshold is presented as a practical warning sign rather than a demonstrated hard limit.
  • Claim: A minimally credible skill eval should cover both desired and undesired activation, measure outcomes rather than a prescribed reasoning path, and account for nondeterminism. | Evidence: The suggested starting point is 10–20 prompts, including about five happy-path cases and five cases where the skill should not trigger; mature runs should use isolated environments, three to six trials per case, production traces, and tests across models and harnesses such as Gemini, Codex, Cursor, and other coding agents. | Implication: Ken's eval framework should report task success, activation precision and recall, and reliability distributions by model-harness combination rather than a single pass/fail run. | Caveat: Three to six trials may expose obvious instability but may not be statistically sufficient for high-risk workflows or small performance differences.
  • Claim: Useful skill eval infrastructure can be built with simple assets rather than a large evaluation platform. | Evidence: For the Gemini Interactions API skill, DeepMind created 117 cases from real user behavior, synthetic cases, and feedback; a JSON file stored prompts, language, should-trigger labels, and expected checks, while a Python script ran Gemini CLI. Regex checks validated the SDK, model, methods, and absence of obsolete patterns, helping performance reach almost 90% valid code across Python and TypeScript. | Implication: Ken can bootstrap skill evals immediately with structured fixtures, a harness runner, and deterministic validators, adding LLM judges only where outcomes cannot be asserted directly. | Caveat: The transcript does not state the baseline percentage, exact model configuration, trial count, or whether the nearly 90% result generalized beyond the tested setup.
  • Claim: Skills should have a tested retirement path because model improvements can make capability instructions redundant or harmful. | Evidence: DeepMind runs evals on every internal skill change and blocks merges that do not improve the cases or add new eval coverage. Schmid recommends running every suite with and without the skill, retiring the skill when the unassisted model matches its performance, and preserving the eval to detect future degradation. | Implication: Ken should treat the eval as the durable asset and the skill as a replaceable intervention, using ablation results to control token cost, maintenance burden, and obsolete behavior. | Caveat: Preference skills that encode proprietary workflows, style, or policy are less likely to become unnecessary solely through foundation-model improvements.

Detailed Brief

Choose scripts, user-invoked skills, and model-invoked skills deliberately

  • Claims: A fixed sequence of deterministic steps is usually better implemented as a script than as natural-language skill instructions.; User-invoked skills are well suited to explicit internal workflows, while model-invoked skills are necessary when external users cannot know the agent's skill catalog.; Skill instructions should define goals, constraints, and relevant resources rather than micromanaging every step the model already knows how to perform.
  • Evidence: Creating pull requests and staging documentation are given as examples of workflows suitable for explicit invocation or scripts.; For a configuration change, the skill can identify the relevant file and constraint rather than spelling out read, edit, and redeploy steps.; A vague description such as 'use for web development tasks' can activate for both React and Angular, whereas a boundary such as 'only for React components or Tailwind CSS' narrows routing.
  • Caveats: A script can still be exposed as a tool that an agent selects; the key distinction is whether execution itself requires model judgment.; Overly narrow activation language may improve precision while reducing recall, so description changes still require balanced positive and negative tests.
  • Implications: The skill layer should not become a substitute for deterministic automation or typed tools.; Agent architecture should separate routing policy from execution mechanism so each can be tested and optimized independently.

Build layered validators and contamination-resistant test environments

  • Claims: Deterministic checks should be preferred when correctness can be inspected directly; LLM judges are a fallback for semantic outcomes that resist simple assertions.; Eval runs need clean workspaces because coding agents may recover context from previous chats, files, or executions and appear to pass without using the intended skill.; Trace validation can supplement output validation by detecting commands, CLIs, or other actions relevant to safety and process compliance.
  • Evidence: DeepMind's internal setup supports workspace files, startup commands that install dependencies, script validators over traces, and LLM-judge expectations.; The speaker notes that many coding-skill checks reduce to regex assertions and that coding agents can help generate those expressions.; An LLM judge can apply a rubric to the output and return pass or fail when deterministic validation is inadequate.
  • Caveats: A judge checking whether a skill or command was invoked can conflict with the broader recommendation to score outcomes rather than paths; process assertions should therefore be limited to genuine safety or compliance requirements.; LLM judges introduce their own nondeterminism and bias, so their rubrics and calibration also need testing.
  • Implications: A robust control plane should support both black-box outcome scoring and optional white-box trace assertions.; Workspace isolation is part of eval validity, not merely an infrastructure convenience.

Notable Concepts & Terms

  • Progressive disclosure: A three-layer skill structure that exposes a short routing description by default, loads the core skill only when selected, and retrieves deeper references only when needed.
  • Capability skill: A temporary skill that compensates for something the current model cannot do consistently; it should be retired when ablation tests show the base model has caught up.
  • Preference skill: A more durable skill encoding company-specific style, policy, workflow, or domain conventions unlikely to be absorbed reliably by a general foundation model.
  • Model-invoked skill: A skill the model must select from prompt context and its description, making routing descriptions and negative tests especially important.
  • No-op instruction: Generic language such as 'write clear, high-quality code' that does not measurably alter behavior but consumes tokens and increases maintenance noise.
  • Ablation test: Running the same eval with and without a skill to quantify its incremental value and determine whether it can be removed.
  • Outcome-over-path evaluation: Scoring whether the agent completed the task correctly rather than requiring the skill to be loaded at a particular turn or enforcing an unnecessary reasoning sequence.
  • Graduated eval: An eval retained after its associated skill is retired so later model or harness regressions can trigger a replacement intervention.

Operator Notes / Why Ken Should Care

  • Create a skill registry containing owner, capability-versus-preference classification, supported models and harnesses, token footprint, eval-suite link, and last ablation date.
  • Set an initial CI policy that rejects new or modified production skills without positive cases, negative routing cases, an unassisted baseline, and recorded multi-trial results.
  • Add telemetry fields for skill eligibility, activation, completion result, model, harness, latency, and token consumption so production traces can be converted into regression fixtures.
  • Audit the highest-volume skills first for files over roughly 500 lines, generic no-op language, duplicated references, and deterministic procedures that should become scripts or typed tools.
  • Maintain a model-by-harness compatibility matrix instead of certifying a skill globally; rerun the relevant slice whenever the model, system prompt, tool interface, or agent harness changes.
  • Require human review of LLM-generated skill text and judge rubrics, with deterministic validators used wherever an SDK, command, model identifier, schema, or artifact can be checked directly.
  • Define retirement thresholds that include quality, activation error, cost, and latency so a skill is removed only when the no-skill configuration is operationally equivalent, not merely close on average.

Source/Metadata

  • Title: Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind
  • Transcript words: 7660
  • Duration seconds: 1305
  • Timestamp note: No timestamps or chapters were present. The supplied transcript contains substantial duplicated passages and one corrupted repeated-word segment, but the main arguments and examples remain recoverable.

Transcript

3898 words en Processed in 171.5s

[SPEAKER_00] Hi, everyone. My name is Philip. I'm based out of Germany. I'm part of the Google DeepMind team, mostly working on Gemini API and agents, and we are going to talk about why you should not ship skills without evals. And maybe before we start, I need a little bit of your help. So if you could raise your hands, if you use coding agents to write code. So, yeah, hopefully every hand goes up, right? And do you use skills with it? Okay. Do you have evals for those skills? Okay, yeah, that's not a lot of hands. Everyone uses skills. No one has evals. Hopefully we can fix that today. And very important is that code checks fail in production. And Skillbench is a very popular and nice eval or benchmark, which indexed over 50,000 skills from GitHub and tried to look into them. And almost none of those skills had evals. Most of them were AI-written, not really tested, and it's very hard to know if your skill is good or bad because agents are really nondeterministic. So you might not know if your task fails because your skill is bad or if your task fails because it's way too challenging for the model. So very important before we go into it, I want to really make sure that we know the difference between the agents we use and the agents we build. Most of us use agents for writing code, doing productivity work. That's the agents we use. It's, like, anti-gravity, cursor, cloud code, and there you are the engineer, and you have context about skills, right? If you write some prompt to help me build a new Gemini API feature, and if your agent does not invoke the skill on the first time, you will notice it very quickly. You stop your task and reprompt it or use slash commands for triggering those skills. When you build an agent inside your application for consumer or customers, they have no idea about what a skill is. They don't start their prompt with, use customer support skill to help me refund or use refund skill to help me solve my problems. So there's a big difference between the agents we use and how we use skills and the agents we build and how our customers might want to use skills in the context there. And what is a skill? I mean, every one of us knows, hopefully, by now what a skill is. It's really a folder with a skills MD file in it and then some additional assets to make that skill really work. And the big difference with skills is that they work on progressive disclosure. So most of the skills start very small, so you have the title and the description. The description is normally part of the model's context, so the model knows when to use the skill. Second layer is we have a skills body with more instructions, more details, and hopefully more references to external files, and then you can really go deep in those reference files where there's all of the context the model needs to discover to solve the task. And I like to differentiate between two kinds of skills. So there are capability skills and preference skills. Capability skills teach models something they cannot do consistently at the moment. Maybe it's tracing some logs, creating a new React app, and those capability skills are temporary. So the better our model gets, the more likely it is that we can remove those skills, and evals will tell us when we can retire a skill and when not. And then we have preference skills. Those are more durable, mostly encode some preferences. So if you have a specific workflow in your team or a specific style language or other preferences which are very specific to your company, you will have or create preference skills, and those preference skills are then protected with evals because most of the foundation models might not integrate the knowledge which is very specific to your use case or your domain. And preference skills are very valuable, so we really want to make sure that those are working and we don't update our agents to degrade performance. So do skills work? Yes, they do work. And I'm going back to the skills bench, which has an update of 1.1, which has evaluated all kinds of open and closed models in different harnesses showing that skills, on average, improve the performance by roughly 15%. Skills bench covers around 100 different tasks based on coding and also productivity across different languages. It's openly available. They have a very nice website, a very nice leaderboard, and are also very open for community contributions. And then they did a second analysis on self-generated or AI-generated skills, right? It's very easy if you are in a coding agent and you work on something, tell the model, create a skill, and then it writes a skill MD file. We maybe look at it very closely. It roughly covers what we want to do, and then we just accept it and start using it. And what I found out is that human-written skills are the best we can provide. AI-generated skills can impact performance negatively, and skills or skills.md files should be below 500 lines of words. So if you have your laptop open and have a skill available, if you open that and if it's above 500 lines, you should definitely look at the skill after our session. And the last topic about what is a skill and how a skill works, It roughly covers what we want to do, and then we just accept it and start using it. And what I found out is that human-written skills are the best we can provide. AI-generated skills can impact performance negatively, and that skills or skills.md files should be below 500 lines of words. So if you have your laptop open and have a skill available, if you open that and if it's above 500 lines, you should definitely look at the skill after our session. And the last topic about what is a skill and how a skill works, we have different ways of triggering our skill, right? We can have a model-triggered skill, meaning based on the context and the description the model decides to use or read a skill to get more context to solve a task, and then there are user-invoked skills. And I think people underestimate how powerful user-invoked skills are, and they most of the time just accept the overhead by adding it into the context. I have many user-invoked skills for more workflow type of tasks, like creating a pull request, staging documentation, and all of the very normal dev work which could be run in a script should most likely be a user-invoked skill. And when you build agents for a customer, you don't have those user-invoked skills. We are only working in the model-invoked skills, and that's where we are focusing on for the small eval section we are going to look in a second. So writing skills is an important topic. We're going to look at eight examples on how or tips on how you can write good skills. And most importantly, if you work with model-invoked skills, is the description. Because the description are most of the time two sentences we provide to the system instruction to help the model know when it should use a skill or not. And it's bad if your description is too vague because then it might trigger too often or it might not be triggered if you need it. So very important is the why and the how for the model, so why it should use that skill and then how it should use that skill. Very common is use that skill if you are working on a React application, for example. And then, of course, the when. And we should write directives instead of essays. So we should not say something like, hey, the interactions API is recommended for multichat because it handles session state, and it's you should be way more directive. Like use the interactions API if you are working on a chat application. So you need to give the model clear instructions and directives on when it should use the skill and how it should use the skill. And similar to what we have seen in the skills bench results, we should keep the skill lean and layer information. So the description is the cost you always pay on every model invocation. So on every model call, the description is part of the model context. So you always pay that 100, 200 tokens cost, and you don't want to have a super long description because then you always have to pay that. When you have a very long skill MD file, it will be always read into context when the model decides to read the skill or to use the skill, which can be expensive as well. That's why we want to keep it as concise as possible, but still include all of the reference and details for the model to solve that task. And then, of course, the layer three is we can have those reference files where the model needs to go really deep into a very specific task. And a good example for this is if you are working in maybe a multi-cloud environment, and you have a skill to deploy your application, you might need instruction to deploying to AWS and deploying to Google Cloud. Those should not be part of your skill MD file. Those should be references, that you have a reference for AWS, a reference for Google Cloud, maybe a reference for Azure, so that the model can basically explore based on the context where it should go to get all of that information. Then we should set the right level of freedom. So here I see many people very clearly describing the exact workflow in a skill. Step one, go there. Step two, do this. Step three, do this. If you have those type of use cases, you should not use skills. Maybe you should write a script, because if the process or the workflow is always the same, you don't need to waste models and tokens for that exercise. You can create a script. You can tell the model, use that script to run a specific workflow. So rather define goals and constraints. So if you need to deploy to your update or stage your documentation, describe how the model can do that, or for your database, updating a config, you should not say read the config, update the port, and then deploy again. The model knows what to do. Just, hey, if we need to change the config, here's the file, make the change. Then don't skip negative cases. So we always look at when we want to use the skill, but most of the time we don't look at when we don't want to use the skill. So if we have a description for our skill, which says use it for web development tasks, it might over trigger. Maybe you'll work with React, maybe you will also work with Angular, and the model always loads the skill if you're working in a web development environment. But if you are very specific for, hey, only use that skill for React components or for Tailwind CSS, then the model knows, hey, that's very specific for one to use. And with evals, we can also identify those. And then test early. So, if we have a description for our skill, which says use it for web development tasks, it might over trigger. Maybe you'll work with React, maybe you will also work with Angular, and the model always loads the skill if you're working in a web development environment. But if you are very specific for, hey, only use that skill for React components or for Tailwind CSS, then the model knows, hey, that's very specific for one use. And with evals, we can also identify those. And then test early. So, that's what we are going to look at. We should really try to test when you create a new skill, always try to create 10 or 20 prompts. I like to create five for the happy path. So, when do I want to use that skill? Five when I don't want to use that skill, just to make sure the model is not over triggering the skill and confusing itself. And then if you have already some customer or production traces, try to include those as well because nothing is better than real-world data. And then tip seven, which is quite new, and I have to give all credits to Matt. So, if you don't know Matt, he's a great AI educator, and you should definitely follow him. He published a tweet and also a skill on eliminating all of the no-ops. And what he found is that AI-generated skills tend to include a lot of no-ops. And no-ops basically is an instruction which does nothing to change the agent's behavior. It's like "before" or "make an implementation easy to read." Like the model knows how or when it should make something easy to read or write clear high-quality code. That's what we expect from the model to do without telling it really. So, definitely look at those no-ops. He has published a very good skill in his skills repository. And then last but not least, know when you should retire a skill. Skills are not there to live forever. Models get better, behaviors change, expectations change, the environment changes. So, always try to run evals with and without the skill enabled. And if the model achieves the performance without even triggering the skill, you know you can retire that skill, save the cost for your tokens. And then also, don't keep redundant skills. Save cost at the end and maintenance as well. And to look at a little bit of a practical example and also how you can create your own small eval or eval harness for skills. Earlier this year, we wanted to create a new skill for the Gemini Interactions API. So, the Gemini Interactions API is our new interface for working with Gemini models and with agents. And the Interactions API was released after the last training of Gemini. So, the model, Gemini 3 and 3.1 or even 3.5, has no context about what is the Gemini Interactions API. So, we decided, okay, let's look at creating a skill to help the model create good code for the Interactions API to use the latest models. And to do that, we created 117 test cases. Those are based on data we see from real users trying to generate Gemini code, from synthetically generated test cases, and also from feedback we see people saying, hey, the model is using Gemini 2.0 even if we are already on 3.0. And the end result was that we improved the performance up to almost 90% for generating valid Interactions API code with the latest Gemini models. And to do this, we basically only needed two very simple assets. So, one of those was a JSON file with all of our test cases. It's very clear structure. It's, hey, we have a prompt. That's basically what we expect the user to provide. We have a language because we wanted to test the skill against TypeScript and Python. We have a should_trigger that's basically there to tell us if the agent should read the skill or not read the skill. And then we have different expected checks. We look at them in a little bit. Those are basically very simple asserts for that prompt if it should trigger or not. And then we have a very basic Python script which runs a coding agent. In this case, it was the Gemini CLI which passes the output and returns it. So, we can take a look at the outcome, whether we have valid code for the Interactions API or not. And most of the tests or evals for skills can be regex. It's very amazing how good regex you can write using coding agents. And for us, it was all about, okay, do we use the correct SDK? Do we use the correct model? Do we use the correct methods? Do we use any old patterns? To look at the whole traces or the whole steps taken. And a very easy case is you just create an LLM as a judge with a rubric on what you want to look at, and then take the output, put it through the LLM as a judge, try to get a pass or a fail. And then if it fails, look at the data, and then try to improve your skill based on that. And that's also how we now evolve skills at Google DeepMind. So we don't use YAML, but as an example. We have tests or evals alongside every skill we have internally at Google DeepMind. Every test has multiple cases with a prompt. We all run them in clear workspaces. So you can define your workspace or environment. and then take the output, put it through the LLM as a judge, try to get a pass or a fail. And then if it fails, look at the data, and then try to improve your skill based on that. And that's also how we now evolve skills at Google DeepMind. So we don't use YAML, but it's just as an example. We have tests or evals alongside every skill we have internally at Google DeepMind. Every test has multiple cases with a prompt. We all run them in clear workspaces. So you can define your workspace or environment if it should include additional files like your application environments. You have startup commands, which preloads or installs libraries into the environment. And then you have script validators. Those are those red checks where we look at all of the traces to see what the skill triggered, was a certain command run, was a certain CLI run. And then we also have LLM as a judge, where we have some expectations which are matched against, like, hey, did it trigger the skill? Did it run a certain bash command to also evaluate it? And we run them on every change to the skill. So if a change happens or a diff to the skill file, the eval will be run. And there will also be a result. And the change will not be merged if it is not improving the test cases. So we always have those regression tests for every change to the skill. And you can only change the skill if it improves the eval or add new evals. And yeah, that's how we manage it. And then last but not least, 10 examples for best practices for skills. You don't need to take photos. They are in the blog post. I can share later. So we had it many, many times. The skill description is very important. We have seen 50% of the failures because the skill was not triggered correctly because the prompt of the user was not detailed enough for the model to understand, hey, I need to use that skill to solve the task. And especially if you build agents for others, they are not aware of the skill descriptions you have for your model and for your skills. So they might write something very shallow. And then the model needs to know, OK, I need to trigger that skill. We should write directives over passive information. So we should always think about it. You should tell the agent what to do or not what to do and not just, hey, if you feel happy today, please use the skill. Include negative tests. We always forget negative tests. Start small. Even 10 to 20 skill eval samples are better than nothing. You will be surprised at how much you will find even from five to 10 examples. And then definitely create outcomes, not paths. We don't want to test if the model loads the skill on the first turn. We really want to test if it can achieve the task based on the prompt. And if it loads the skill, it loads the skill. If not, then not. If it loads the skill after five turns, that's also OK. Then we want to have isolated runs because coding agents are very good at finding or cheating. So if you run inside your existing environment, it might look up previous chats or it might look up some other executions and then try to cheat it and get the context from the skill without even using the skill. Then definitely run more than one trial when running eval. It's that agents, our models are non-deterministic. Maybe the first one works. The second one doesn't. So always run three to six trials per case and to measure reliability. Test across different harnesses. If you work with or if you have employees or people working with different harnesses, not only just evaluate against code or anti-gravity. If you have people working with cursor, try to include them as well because agent harnesses behave differently and of course model behaves differently. So maybe your skill is very good with Gemini, but very bad with codex. And then you have customers, consumers using your harness with codex and then it fails. And then graduate your eval. So if your model is good enough, it doesn't need the skill anymore, keep that eval. You don't need to throw that eval away because you throw the skill away. You can keep that eval to make sure that the model or the agent keeps the performance. And as soon as you start seeing some degradation, you can reintroduce a skill. You can maybe tweak some other tools or pieces to keep the performance up and then really detect when you can retire skill. And you will be very surprised with all of the model updates, how fast you can retire skill, which you might need it six months ago, but not today anymore. And I have some homework for you. So if you are back from holiday on Monday, pick the most used skill and write five test prompts. You can also use your coding agent and ask it to see, look at your trajectories, which are my most used skills, and then try to create some skills. As you have seen, it's very easy to write your eval harness. It's a JSON or YAML file and then some Python script which runs your coding agent or your agent harness and then look at the outcome. Definitely try to look at removing no ops. Maybe it does not change the eval performance, but it helps you save costs because all of the tokens which are not helpful or not changing the agent behavior are money you will spend. So look at writing create skills from Matt. You can find it on GitHub. And then also run ablation tests. So run always evals with your skill loaded and without your skill loaded. Only that way you will know when you can retire a skill. Some Python script which runs your coding agent or your agent harness and then look at the outcome. Definitely try to look at removing no ops. Maybe it does not change the eval performance, but it helps you save costs because all of the tokens which are not helpful or not changing the agent behavior are money you will spend. So look at writing create skills from Matt. You can find it on GitHub. And then also run ablation tests. So run always evals with your skill loaded and without your skill loaded. Only that way you will know when you can retire a skill or if a skill is really helpful for your performance. So don't ship skills without evals. Thank you. Thank you. Thank you. in, like, the context there. And what is a skill? I mean, every one of us knows, hopefully, by now what a skill is. It's, like, basically really a folder with a skills MD file in it and then some additional assets to make that skill really work. And the big difference with skills is that they work on progressive disclosure. So most of the skills start very small, so you have the title and the description. The description is normally part of the model's context, so the model knows when to use the skill. Second layer is we have a skills body with more instructions, more details, and hopefully more references to external files, and then you can really go deep in those reference files where there's all of the context the model needs to discover to solve the task. And I like to differentiate between two kinds of skills. So there are capability skills and preference skills. Capability skills teach models something they cannot do consistently at the moment. Maybe it's, like, I don't know, like tracing some logs, creating a new React app, and those capability skills are temporary. So the better our model gets, the more likely it is that we can remove those skills, and evals will tell us when we can retire a skill and when not. And then we have preference skills. Those are more durable, mostly encode some preferences. So if you have a specific workflow in your team or a specific style language or other preferences which are very specific to your company, you will have or create preference skills, and those preference skills are then protected with evals because most of, like, the foundation models might not integrate the knowledge which is very specific to your use case or your domain. And preference skills are very valuable, so we really want to make sure that those are working and we don't, like, update our agents to degradate performance. So do skills work? Yes, they do work. And I'm going back to the skills bench, which has an update of 1.1, which has evaluated all kinds of open and closed models in different harnesses showing that skills, on average, improve the performance by roughly 15%. Skills bench covers around 100 different tasks based on, like, coding and also productivity across different languages. It's openly available. They have a very nice website, a very nice leaderboard, are also very open for community contributions. And then they did a second analysis on self-generated or AI-generated skills, right? It's very easy if you are in a coding agent and you work on something, tell the model, create a skill, and then it writes a skill MD file. We maybe look at it very closely. It roughly covers what we want to do, and then we just accept it and start using it. And what I found out is that human-written skills are the best we can provide. AI-generated skills can impact performance negatively, and that skills or skills.md files should be below 500 lines of words. So if you have your laptop open and have a skill available, if you open that and if it's above 500 lines, you should definitely look at the skill after our session. And the last topic about what is a skill and how a skill works, we have different ways of triggering our skill, right? We can have a model-triggered skill, meaning based on the context and the description the model decides to use or read a skill to get more context to solve a task, and then there are user-invoked skills. And I think people underestimate how powerful user-invoked skills are, and they most of the time just accept the overhead by adding it into the context. I have many user-invoked skills for more workflow type of tasks, like creating a pull request, staging documentation, and all of the very normal dev work which could be run in a script should most likely be a user-invoked skill. And when you build agents for a customer, you don't have those user-invoked skills. We are only working in the model-invoked skills, and that's where we are focusing on for the small eval section we are going to look in a second. So writing skills is an important topic. We're going to look at eight examples on how or tips on how you can write good skills. And most importantly, if you work with model-invoked skills, is the description. Because the description are most of the time two sentences we provide to the system instruction to help the model know when it should use a skill or not. And it's bad if your description is too vague because then it might trigger too often or it might not be triggered if you need it. So very important is the why and the how for the model, so why it should use that skill and then how it should use that skill. Very common is like use that skill if you are working on a React application, for example. And then, of course, the when. And we should write directives instead of essays. So we should not say something like, hey, the interactions API is recommended for multichat because it handles like session state, and it's like you should be way more directive. Like use the interactions API if you are working on like a chat application. So you need to give the model like clear instructions and directives on when it should use the skill and how it should use the skill. And similar to what we have seen in the skills bench results, we should keep the skill lean and layer information. So the description is the cost you always pay on every model invocation. So on every model call, the description is part of the model context. So you always pay that 100, 200 tokens cost, and you don't want to have a super long description because then you always have to pay that. When you have a very long skill MD file, it will be always read into context when the model decides to read the skill or to use the skill, which can be expensive as well. That's why we want to keep it as concise as possible, but still include all of the reference and details for the model to solve that task. And then, of course, the layer three is like, we can have those reference files where the model needs to like go really deep into a very specific task. And a good example for this is like, if you are working in like maybe a multi-cloud environment, and you have a skill to deploy your application, you might need instruction to deploying to AWS and deploying to Google Cloud. Those should not be part of your skill MD file. Those should be references, that you have a reference for AWS, a reference for Google Cloud, maybe a reference for Azure, so that the model can basically explore based on the context where it should go to get all of that information. Then we should set the right level of freedom. So, here I see many people very clearly describing the exact workflow in a skill. Step one, go there. Step two, do this. Step three, do this. If you have those type of use cases, you should not use skills. Maybe you should write a script, because if the process or the workflow is always the same, you don't need to waste models and tokens for that exercise. You can create a script. You can tell the model, use that script to run a specific workflow. So, rather define goals and constraints. So, if you need to like deploy to your update or stage your documentation, describe how the model can do that, or like for your database, updating a config, you should not say like read the config, update the port, and then like deploy again. The model knows what to do. Just like, hey, if we need to change the config, here's the file, make the change. Then don't skip negative cases. So, we always look at the when we want to use the skill, but most of the time we don't look at when we don't want to use the skill. So, if we have a description for our skill, which says use it for web development tasks, it might over trigger. Maybe you'll work with React, maybe you will also work with Angular, and the model always loads the skill if you're working like a web development environment. But if you are very specific for like, hey, only use that skill for React components or for tail in CSS, then the model knows, hey, that's very specific for one to use. And with evals, we can also identify those. And then test early. So, that's what we are going to look at. We should really try to test when you create a new skill, always to re-eye to create 10 of 20 prompts. I like to create five for like the happy path. So, when do I want to use that skill? Five when I don't want to use that skill, just to make sure the model is not over triggering the skill and confusing itself. And then if you have already some customer or production traces, try to include those as well because nothing is better than real-world data. And then tip seven, which is quite new, and I have to give all credits to Matt. So, if you don't know Matt, it's a great AI educator, and you should definitely follow him. He published a tweet and also a skill on like killing all of the non-ops. And what he found is that AI-generated skills tend to include a lot of no-ops. And no-ops basically is an instruction which does nothing to change the agent's behavior. It's like before or make an implementation easy to read. Like the model knows how or when it should make something easy to read or write clear high-quality code. I mean, like that's what we expect from the model to do without telling it really. So, definitely look at those no-ops. He has published a very good skill in his like skills repository. And then last but not least, know when you should retire a skill. Skills are not there to live forever. Models get better, behaviors change, expectation change, the environment changes. So, always try to run evals with and without the skill enabled. And if the model achieves the performance without even like triggering the skill, you know you can retire that skill, save the cost for your tokens. And then also, don't keep like redundant. So, save cost at the end and maintenance also as well. And to look at a little bit of a practical example and also how you can create your own small eval or eval harness for skills. Earlier this year, we wanted to create a new skill for the Gemini Interactions API. So, the Gemini Interactions API is our new interface for working with Gemini models and with agents. And the Interactions API was released after the last training of Gemini. So, the model or Gemini 3 and like 3.1 or even 3.5 has no context about what is the Gemini Interactions API. So, we decided, okay, let's look at creating a skill to help the model create good code for the Interactions API to use the latest models. And to do that, we created 117 test cases. Those are based on like data we see from real users trying to generate Gemini code, from synthetic generated test cases, and also from like feedback we see people like, hey, the model is like using Gemini 2.0 even if we are already on 3.0. And the end result was that we improved the performance up to like almost 90% for generating valid Interactions API code with the latest Gemini models. And to do this, we basically only needed like two very simple assets. So, one of that was a JSON file with all of our test cases. And it's very like no clear structure. It's like, hey, we have a prompt. That's basically what we expect the user to provide. We have a language because we wanted to test the skill against TypeScript and Python. We have a should trigger that's basically there to tell us if the agent should read the skill or not read the skill. And then we have different expected checks. We look at them in a little bit. Those are basically very simple asserts for that prompt if it should trigger or not. And then we have a very basic Python script which runs a coding agent. In this case, it was the Gemini CLI which passes the output and returns it. So, we can like take a look at the outcome, whether we have valid code for the Interactions API or not. And most of the tests or evals for skills can be regex. It's like very amazing how good of regex you can write using coding agents. And it's really for us, it was all about, okay, do we use the correct SDK? Do we use the correct model? Do we use the correct methods? Do we use any old patterns? Do we use evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil evil to look at the whole traces or the whole steps taken. And a very easy case is you just create an LLM as a judge with a rubric on what you want to look at, and then take the output, put it through the LLM as a judge, try to get a pass or a fail. And then if it fails, look at the data, and then try to improve your skill based on that. And that's also how we now evolve skills at Google DeepMind. So we don't use YAML, but it's just as an example. We have tests or evals alongside every skill we have internally at Google DeepMind. Every test has multiple cases with a prompt. We all run them in clear workspaces. So you can define your workspace or environment if it should include additional files like your application environments. You have startup commands, which basically preloads or installs libraries into the environment. And then you have script validators. Those are those red checks where we look at all of the traces to see what the skill triggered, was a certain command run, was a certain CLI run. And then we also have LLM as a judge, where we have some expectations which are basically matched against, like, hey, did it trigger the skill? Did it run a certain bash command to also evaluate it? And we run them on every change to the skill. So if a change happens to or like a diff to the skill file, the eval will be run. And there will also be a result. And the change will not be merged if it is not improving the test cases. So we always have those regression tests for every change to the skill. And you can only change the skill if it improves the eval or add new evals. And yeah, that's how we basically manage it. And then last but not least, 10 examples for best practices for skills. You don't need to take photos. They are in the blog post. I can share later. So I mean, we had it many, many times. The skill description is very important. We have seen 50% of the failures because the skill was not triggered correctly because the prompt of the user was not detailed enough for the model to understand, hey, I need to use that skill to solve the task. And especially if you build agents for others, they are not aware of the skill descriptions you have for your model and for your skills. So they might write something very shallow. And then the model needs to know, OK, I need to trigger that skill. We should write directives over passive information. So we should always think about it. You should tell the agent what to do or not what to do and not just like, hey, if you feel happy today, please use the skill. Include negative tests. We always forget negative tests. Start small. Even like 10 to 20 skill eval samples are better than nothing. You will be surprised on how much you will find even from like five to 10 examples. And then like definitely create outcomes, not paths. We don't want to test if the model loads the skill on like the first turn. We really want to test if it can achieve the task based on the prompt. And if it loads the skill, it loads the skill. If not, then not. If it loads the skill after five turns, that's also OK. Then we want to have isolated runs because coding agents are very good at finding or cheating. So if you run inside your existing environment, it might look up previous chats or it might look up some other executions and then like try to cheat it and get the context from the skill without even using the skill. Then definitely run more than one trial when running eval. It's like agents, our models are non-deterministic. Maybe the first one works. The second one doesn't. So always run three to six trials per case and to measure reliability. Test across different harnesses. If you work with like or if you have employees or people working with like different harnesses, not only just evaluate against code or anti-gravity. If you have people working with cursor, try to include them as well because agent harnesses behave differently and of course model behaves differently. So maybe your skill is very good with a Gemini, but very bad with codex. And then you have customers, consumers using your harness with codex and then it fails. And then gradiate your eval. So if your model is good enough, it doesn't need the skill anymore, keep that eval. You don't need to throw that eval away because you throw the skill away. You can keep that eval to make sure that the model or the agent keeps the performance. And as soon as you start seeing some degradation, you can reintroduce a skill. You can maybe tweak some other tools or pieces to keep like the performance up and then really detect when you can retire skill. And you will be very surprised with all of the model updates, how fast you can retire skill, which you might need it like six months ago, but not today anymore. And I have some homework for you. So if you are back from holiday on Monday, pick the most used skill and write five test prompts. You can also use your coding agent and ask it to see, look at your trajectories, which are my most used skills, and then try to create some skills. As you have seen, it's like very easy to write your eval harness. It's like a JSON or YAML file and then like some Python script which runs your coding agent or your agent harness and then like look at the outcome. Definitely try to look at the removing no ops. Maybe it does not change the eval performance, but it helps you save costs because all of the tokens which are not helpful or not changing the agent behavior are money you will like spend. So look at writing create skills from Matt. It's you can find it on GitHub. And then also run ablation tests. So run always evals with your skill loaded and without your skill loaded. Only that way you will know when you can retire a skill or if a skill is really helpful for your performance. So don't ship skills without evals. Thank you. Thank you. Thank you.