SPEAKER_01
Hello everyone, is everyone excited for the conference? Awesome, we got a full house here. I'm very glad to be here, very honored to be giving the opening or one of the opening workshops. Today, if you've noticed already, the title is slightly different from what we have in the schedule. I've done a rebrand, but the theme for the workshop will remain the same. We went from skill issue to level up your skills. I've decided to move the skill issue title to the keynote that I'm giving tomorrow. You'll have time to know more about what the keynote is going to be about.
SPEAKER_01
But mainly, this workshop is what I've been doing in the last two months at Superbase, writing our own skills. And tomorrow, I'm going to present how we put this actually into production and the lessons we've learned. So, for everyone who's been paying closer attention, you probably noticed that I'm running this slide deck on localhost. Some of you have already noticed. This is no coincidence at all. I essentially vibe coded the presentation. So, if you see something off, it was not my fault. It was Claude. But for you to, if you don't believe me, you can see that you cannot do or you'll have to be a very Google Slides guru to have dark mode enabled.
SPEAKER_01
So, honestly, I like this layout better. I think we're going with dark mode here. If there's any light mode fans out there or the majority of the room, it's light mode, I'm happy to switch back. But for now, let's go with this one. So, to do a little presentation of myself before starting the workshop, my name is Pedro. I'm from Portugal, Lisbon. And I work at Superbase as an AI tooling engineer. Essentially, my day-to-day is to think of how we can make Superbase the most agentic friendly as possible and improve the agent experience. So, you've probably heard about development experience, the DX. We are more focused on DAX, which is the same thing, but for agents.
SPEAKER_01
In this workshop, we're going to talk about skills, because essentially, that's how we've been improving the performance of agents around a product like Superbase or a company like Superbase has multiple products. The secret sauce has been skills. So, we're going to dive into how to write one, how to test it first manually, and then how to automate the testing with evaluations. So, to start with, how many of you have heard about skills? All right. So, almost everyone. So, what I'm going to say, it's probably no news to you.
SPEAKER_01
Skills are folders with instructions and files for you to run workflows, repeated workflows, or give custom information to your agents, or provide a new set of tools, in form of scripts. So, there's a bit of a misconception about skills. Usually, the skill.md, the main file takes the spotlight, but skills can actually be more than just the main file. So, the main file is a markdown file named skill.md, where essentially the main information about the skill lives.
SPEAKER_01
It is composed by this front matter at the top, which can have multiple fields, but the two required ones are the name, which identifies the skill, and then the description, which tells the agent what the skill does. The main thing that skills bring that tools like MCP didn't was this concept of progressive disclosure. Progressive disclosure is when the agent, or all the information about a subject is not loaded straight to context. Instead, you just load the exact amount of information that allows the agent to choose to load the rest of the information once it actually needs it. So, in this case, the skill.md file is designed like this.
SPEAKER_01
So, the front matter will be loaded at first to the context of the agent, not the content of the file. This works as an envelope, so the agent knows from the description what the skill does, and when it should load the rest of the information. So, when should it look for the information inside of the file. Inside this file, you can also reference another file. Usually, these other files are either markdown files or scripts, bash, python, whatever you would like to reference. Starting on the reference files, you usually put them inside a reference folder, and they provide more information. You can think about a skill in this format as a book.
SPEAKER_01
The skill.md, you can think of it as the index on steroids, because besides having these links to the other files, you can think of them as the pages of the book or the other chapters. You can have custom information and then also reference it to the other files. The reference files have nothing special about them. They're basically like a normal, regular markdown file. You can think of them similar to skill.md file, but instead of being the main one, it's the one that got referenced. You can also, funny enough, reference files inside of reference files. So you can make basically a graph out of a skill.
SPEAKER_01
And for scripts, I've actually talked about how MCP and skills differ from each other. And we're comparing apples to oranges when it comes to MCP and skills. One of the misconceptions currently probably was already debunked. The debate now is more about MCP versus CLI. But when the skills were released back in, I think it was November or October last year, they basically started this debate about, well, should we use them instead of MCP?
SPEAKER_01
Because if I can run, if I can provide more information, more context to the agent without actually loading every tool to the context like the MCP, and I can also have scripts so I can have actions just like I have on MCP tools, should we use them? And the answer is you should use both. If you're building anything that is an integration, you should use MCP. Anything that if your agent doesn't have access to bash, you should use MCP to integrate to your service. Skills actually just provide more context to your agent.
SPEAKER_01
And you can define workflows, everything that you would not, that you don't have space to define on the MCP tools, descriptions, you can define them on skills. Also, regarding the comparison, the debate between skills, scripts and the MCP tools, the main difference is that tools don't need an environment to run. The agent can just call a tool, knows how to call a tool, especially if the MCP server is remote. And the tool will run on server side. While the scripts, well, they're basically loaded into your machine, they run on your local environment, and they're tied to whatever environment that you have. So if you're running on Linux, they have to be Linux compatible.
SPEAKER_01
If you're running on Mac OS, the same. Windows, I'm not going to even study about it. But essentially, those are the main differences between the MCP tools and the scripts. I hope everything is clear. If you have any doubts, feel free to, I'm going to have a little demonstration. This workshop is going to be more of a walkthrough than actually code along. But feel free to tag in. I have a GitHub repo prepared, so you'll be able to visit it and explore it. But if you have any doubts at any moment of the workshop, feel free to interrupt me or raise your hands. So moving to this. If you're running on Mac OS, the same. Windows, I'm not going to even study about it.
SPEAKER_01
But essentially, those are the main differences between the MCP tools and the scripts. I hope everything is clear. If you have any doubts, feel free to. I'm going to have a little demonstration. This workshop is going to be more of a walkthrough than actually code along. But feel free to tag in. I have a GitHub repo prepared, so you'll be able to visit it and explore it. But if you have any doubts at any moment of the workshop, feel free to interrupt me or raise your hands. So moving to this. Exactly. So I tested this on a smaller screen. It was working. You can see it was Vipe coded. I tested it.
SPEAKER_01
So how do you test your skills, right? So if this is just a markdown, how do you test your markdown files? So to test a piece of code, it's straightforward, right? We already know you have all sorts of test types. You have unit tests, integration tests, or you can test the whole flow, or what we call end-to-end testing. Essentially, when you're testing a markdown file, you can do exactly the same. You can be as granular as you want. But usually, since we have an LLM in the loop, you'll have something called evaluation.
SPEAKER_01
So for those of you who haven't heard about evaluations, or evals for short, they essentially are a non-deterministic way of testing the output or the behavior of an agent or a model. You can test both an LLM or an agent with evals. Essentially, the most common structure—I'm going to present to you at the end a framework for you to test your evals, a very simple one where you can start. And I'm going to dive deeper on evaluations there. But essentially, they usually are made of an input and an expected output, just like a regular test.
SPEAKER_01
And in between, you can evaluate the steps that the agent took, the reasoning, the tools that it called, which is normally more interesting and easy to evaluate than just a test on the exact output, since this is non-deterministic. So there's essentially a framework that you can follow to test your skills. This one was proposed by OpenAI on their blog post called "Systematically Evaluate Agent Skills." I think they released this back in January or February, so not that long ago. But all this is fairly new. So this is basically pre-history. So you start by defining your metrics—what you want to evaluate on your skills.
SPEAKER_01
If you're building a skill for your product, for example, what exactly do you want the skill to highlight to your agent? Is it going to forward it to the documentation? Are you putting some specific instructions, specific workflow? So depending on what you want to evaluate, you start this eval-driven development, this test-driven development. You start by defining the metrics—what exactly good means when it comes to the skill. Then you create the skill itself, right? So you write the skill.md file, any scripts alongside it, the reference files if you want to. They're all optional. The only required is the skill.md file. And then you move to the testing part.
SPEAKER_01
So you run the evaluations or you run it manually. I recently heard the CEO of Braintrust during a podcast. How many of you know Braintrust? Okay. Not as many as the popular skills. But for those of you who don't know, Braintrust is a platform that allows you to systematically run evals and provides you the full picture of the agent behavior during the evaluation scenario, right? I'm just trying to think about another platform to compare it with, but this is fairly new, to be honest. So you can think of it as an observability tool to check the behavior of your agent during a specific controlled scenario, which are the evaluations. So you move to the testing part.
SPEAKER_01
Basically, you run a set of evaluation scenarios. These are defined by the input and expected output tools that should be called. So basically, how do you expect your agent to behave? And then you move to the grading part. So how did the agents do? Well, essentially, this is very similar to a testing cycle, right? But now, instead of having a deterministic output, you can have a non-deterministic one. It's an LLM in between. But you can still have deterministic parts to evaluate on. And then you iterate and repeat. This is why it's a cycle. It's pretty similar to any of the test development cycles that we have at the moment. All right.
SPEAKER_01
So jumping straight to what we're going to do during this workshop. So we're going to write a skill. I've prepared a little demonstration app, a demo app. It's going to be a performance review application with four employees, I believe. One employee, two managers, and one HR representative. And essentially, we're going to find and fix some errors on the database side. We're going to build a skill to help guide the agents to fix them. All right. And then at the end, as I said, I have a framework to test automatically the same scenario that we're going to test manually using evals. Before moving to the demonstration, how many of you have heard about or used Supabase?
SPEAKER_01
All right. So almost anyone knows or used Supabase. I've seen some hands down. Still, I'm going to give a little brief. So Supabase is essentially a backend as a service. You can think of it as the open source version of Firebase. I only Firebase was coming to my mind. Sorry. To Firebase. And if you don't know Firebase, you're probably living under a rock. No, but essentially, it's a backend as a service. You can use it to build any backend as you would like, coming straight out of the box. We provide a database for you to just plug into your application and run on Postgres. One of the most, if not the most popular open source solution out there for databases.
SPEAKER_01
You can easily integrate authentication into your application. Running storage to save files. Many other things. Sorry. To Firebase. And if you don't know Firebase, you're probably living under a rock. No, but essentially, it's a back-end as a service. You can use it to build any back-end as you would like. Coming straight out of the box. We provide a database for you to just plug into your application and run on Postgres. One of the most, if not the most popular open source solution out there for databases. You can easily integrate, including authentication on your application. Running storage to save files. Many, many other things.
SPEAKER_01
Edge functions, which are lambda functions. For those of you who come from the AWS environment.
SPEAKER_01
And so forth. So, the demo application that I've built was built on top of Superbase, of course. And so you can follow along.
SPEAKER_01
Here is the QR code.
SPEAKER_01
At the back, can anyone, everyone scan the QR code? So, should I make this bigger? Bigger? Right. So, everyone can see. I'm basically editing the presentation at the moment as we speak. Let's see what Claude has to offer us. Bigger? This is the cool thing about coding our presentations. I really recommend it.
SPEAKER_01
I probably spend the same or more time than if I just use something like Google Slides, but at least it's more fun. And Tropic should be thrilled about it. [SPEAKER_01] For sure. All right. Let me know once everyone is in the repo. If you cannot see or scan the QR codes, I should probably make the link bigger as well. So, you asked for a demo. Here's the demo on my slides. So, mainly everyone here in this room uses skills. So, I probably won't have to sell you the power of skills. But if you're still a bit skeptical about skills, this whole presentation without skills would be a lot not pleasant. Let's say, much uglier in a sense. Okay. You should probably see it now.
SPEAKER_01
So basically it navigates to GitHub, who who did it, which is my nickname and Improve Skills Workshop.
SPEAKER_01
AI Europe. It's a very long name. All right. So is everyone at the repo at the moment? Okay. Everyone had no trouble. All right. So not this one. Right. So this is the repo that you should be looking at.
SPEAKER_01
Essentially, it's a big repo, but we'll break it down.
SPEAKER_01
I'm actually going to move to VS Code. All right. So here we have two Next.js apps. Actually, the slides are also embedded here. The Next.js app that matters is inside demo, right? And to give you an insight of what that looks like, it's basically this. So it's a very simple application. You can see that it's a vibe coded application, to be honest. The layout, there's nothing special about it. You have, as I've described earlier, several employees of this fictional company. And you can think of it as an internship or a performance review application where you have all the information as an HR employee. You have all the information about the other employees of the company.
SPEAKER_01
And you can change for the sake of the presentation. You can change between users, right? So what we're going to do first, without a skill, we're going to try to implement a new view here. We're going to implement the reports view. [SPEAKER_02] Essentially, this reports part of the application is going to show both the salary and the average rating for the performance review of each department. So HR can know what's going on, can have an overview of the whole company.
SPEAKER_02
[SPEAKER_01] So before we start to vibe code, because during these workshops, no one actually writes code anymore.
SPEAKER_01
So of course I'm going to vibe code it. We're going to break this application down. So if we navigate to the dashboard, nothing special to see. You have the page, the first page, the main page that you've seen. And then you have here the reports where we should have this set view exists. That is going to, so I've prepared the backend. We're just going to create on the backend the view as a SQL view on the database. And then we should be able to see it on the application. So I've prepared, where is it? Yeah.
SPEAKER_01
So I've prepared the prompt and we're going to live test it. Fingers crossed that this works. Right. First let me navigate to the app. Okay. Let's see, here we have more control.
SPEAKER_01
All right. So essentially, for those in the back, I'm just going to ask Claude to create a department stats view. That shows the ad counts and the average salary broken down by department. All right. So for HR to have a full overview of what's going on in the company. So we're going to hit the prompt. And wait to see what it comes up with. Right. First let me navigate to it's app. Okay. Let's see here we have more control. All right. So essentially, the ones in the back, I'm going to ask Claude to create a department stats view. That shows the ad counts and the average salary broken down by department. All right.
SPEAKER_01
So for HR to have a full overview of what's going on in the company. So we're going to hit the prompt.
[SPEAKER_01] And wait to see what it should come up with. [SPEAKER_01] Right. I forgot about this part. I have this MCP server configured. If you have, you actually have, I totally jumped the readme. If you follow along, sorry about this. If you follow along, if you're following along, you can follow the setup guide to get your application started locally. It's essentially going to install the dependencies, clone the repo, install the dependencies, start locally your super based project. You don't have to have the CLI installed. We were using NPX to start to run it as a binary.
SPEAKER_01
[SPEAKER_02] Just reset the database state so you start from scratch with the seeded data, and then just run the app as running NPM run dev. Should be available on localhost 3000 slash dashboard. You also have this MCP dot JSON file prepared. This essentially is pointing to the MCP server that we super based enable for local projects. No authentication required. So your agent should be able to just load it on demand. This MCP server exposes a set of tools. I don't know how many of you have used the super based MCP server, but I think we currently have something along 29 tools.
SPEAKER_01
I believe for the production one, this one is a smaller version, has 20 tools, but you can basically perform essentially almost anything that the one to connect to your remote project does. Basically list the tables that you have, execute SQL straight on your database, apply migration, and run the database advisors and so forth. So essentially what it started to do was to list my tables. So I've asked for a view and it's going to review the schema that I already have implemented. So I'll let you. And now it's going to run the apply migration tool to create the view. So it's basically doing a schema change on my database and it's going to create the view.
SPEAKER_01
If we inspect the view, we're basically creating, create or replace a view, a department stats, the name that we gave, and we're fetching all the information from.
SPEAKER_01
I think department exactly. From. No, from profiles exactly. And group by department. Okay. Made a mistake. It's going to try again. Okay.
SPEAKER_01
It's going to test it. It's actually something that I really like about.
SPEAKER_01
All right. And here's our view on the database. So we currently have it on the database. Let's see if that's also enabled on the app. Hmm. Okay. It's not. Let's quickly, I've created a SQL. What's the name I gave? This is essentially the problem with live views. It usually doesn't go well at first try. Wait. I want to repeat that. And let's see if it implements, if not, we can just run the SQL query for you to see as a different user to see if everything is working accordingly. So for now, he's going to implement on the Next JS application so we can have a nice interface to check the results.
SPEAKER_01
Let me just put on auto mode so I can continue to talk. [SPEAKER_01] So essentially the agent created the view, tested, said everything is working accordingly. We should not. The app was the feature was implemented.
SPEAKER_02
All goods.
SPEAKER_01
But we're actually going to see if everything is good or not. [SPEAKER_02] Let's give it some space, not to pressure it to create the feature.
SPEAKER_01
Let's just wait a bit more time. In the meantime, if you're following along, you can also play with it, change the layout.
SPEAKER_02
Actually using the cloud code. I don't know.
SPEAKER_01
Just doing a brief survey here during the workshop. How many of you are using cloud code as well? Fairly almost. Okay. How many of you are using cursor with cloud codes or with the plugin or, okay. Yeah. Yeah. At least one person. We're going to have some cursor folks here. I think from Anthropic as well. Open AI is going to be here as Gemini. Of course, Google DeepMind is sponsoring the event. So we're basically going to have the whole gang here. Okay. So we should be. I'm trusting his word. All right. So it says that we should now have correctly displayed the department stats view. So let's see if that's actually true. It looks like it. Yeah.
SPEAKER_01
So we now have these cards with the whole view of the company. So I'm logging in as Julia from HR. We can see that we have five people on the engineering team with an average salary of 107 K. HR as well. Well, only one person, which would be Julia. And product has four people and that average salary. All right. So it says that we should now have correctly displayed the department stats video. So let's see if that's actually true. It looks like it. Yeah. So we now have this cards with the whole view of the company. So I'm logging in as Julia from HR. We can see that we have five people on the engineering team with an average salary of 107 K. HR as well.
SPEAKER_01
Well, only one person, which would be Julia and product has four people and that average salary. So far so good. Looks okay. Let's see. So this is sensible information. The reports, right? We're expecting that the other employees will not have access to it. And even the managers only have access for their departments. Let's see if that's the case. So let's navigate to Bob. Bob is the head of engineering. Oh, okay. So Bob also can see the performance reviews of both the different measures of both HR and product. But well, it's not that bad, right? It's not ideal, but at least it's a manager, right? So it should have access to privileged information anyway.
SPEAKER_01
And oh, it looks like a transparent company. Let's see how bad this is. Okay. This is problematic. So we basically created a view. Claude said everything is working because as you can see, the information is here. It was created. But he missed something that is training data. He missed something, which was for Postgres specifically. When you create a new view and your table has a row level security enabled. So for those of you that know, row level security allows for you to define who can see the information on a specific role on a database level. So without trusting the application, you can filter it directly on the database.
SPEAKER_01
So in this case, we should be limiting the view of the roles by user ID, right? And the user role. So if a user has an employee role and an employee role, you should not have access to the rows that don't belong to them. Right? We have row level security enabled. So if you navigate to our Supabase migrations, you can see that we have row level security enabled both on profiles and on performance reviews, right? And on a performance review it should be all right. [SPEAKER_01] Right. So we have reviewer ID equal current setting. So it should work.
SPEAKER_01
Why is that working? Well, when you create a view on Postgres, by default, the permission, it creates with the permissions or the credentials of the user that created the view. And not with the credentials of the table, let's say with the row level security. So basically by default, it bypasses the row level security that you have in place already on your table. So for this to not happen, we have to add a security invoker. We have to use a security invoker flag to transfer the row level security policies or to enable the RLS policies on the view itself.
SPEAKER_01
So this is why currently everyone can see everyone's because the row level security policies were basically bypassed on the view. So for the sake of the demonstration of this workshop, I've already created a skill. I prepared the skill for the presentation. And essentially the skill has three main security points about Postgres that the agent should be aware of during the presentation. For this one specifically, I actually overfed it to the exact view that we're creating, but models right now are smart enough to generalize this. And if I wanted to create a new view, you will be able to essentially create it with this flag. Since Postgres version 15, this flag was enabled.
SPEAKER_01
And every time it's enabled the row level, the RLS policies are also enabled on the view. As you can see, it's actually quite human readable documents. Most of you have already written skills, so I'm not going to dive deep into this. But as you can see, we have both the title. Let me just move this. We have the title. I called it Supabase security and the description uses the verb use. This is an insight that I got from some experiments that I did. Using verbs, mainly the verb use, increases the chances of the skill being loaded. At least on Claude. I don't know if this is default behavior for Claude.
SPEAKER_01
I don't know if it was trained to recognize verbs more easily, especially use. But I found it more efficient to write use and then the whole purpose of the skill in front of it. And then a regular markdown list. So we have the view case there, but also another checklist point for security on RLS. So public schemas should have RLS enabled by default. Public schemas or exposed schemas are the database schemas that are going to provide information for the application that the user can see. So for example, the user's table, the profiles, the performance reviews, all this information is going to be fetched by the front end.
SPEAKER_01
It's completely secured because Supabase makes it secure by allowing you to fetch information from the front end. But the key part here is that if you don't enable row level security, you will not have this filter on the table. So public schemas should have RLS enabled by default. Public schemas, or exposed schemas are the database schemas that are going to provide information for the application that the user can see. So for example, the user's table, the profiles, the performance reviews, all this information is going to be fetched by the front end. It's completely secured. Supabase makes it secure by allowing you to fetch information from the front end.
SPEAKER_01
But the key part here is that if you don't enable role level security, you will not have this filter on the table.
SPEAKER_01
And you will have to rely on the application logic to make the filter. So enabling role level security at least makes it safer for you as the backend engineer, so that you only expose the information that you actually want from the start. And then a couple more things that I'm not going into. So if we can install this skill on this project by running, where do I have the command? Yeah. So I'll be using Vercel's NPM package called skills. Curious to know how you guys have been packaging your skills. Have you ever used this package? Are you using plugins? Just this one. Yeah. This one mainly. Yeah. It became very popular a few months ago.
SPEAKER_01
I think the only problem is it doesn't really adhere to the project. So you get load. Yes. And then you want it only for your local project. Yeah. You can install it both globally and on your project. And also support for multiple agents. While plugins for now are still tied to the agent that is going to allow them. So cursor has plugins, cloud code has plugins. I think other vendors have as well, but they're specifically distributed and made for those specific models. So we're using this one to install. You can install any skill from a repo online that has a skill.md file or you can use it to install the one locally.
It will auto detect the location that you're trying to fetch from based on the format. [SPEAKER_01] So in this case, we don't have any github. We don't have any HTTP protocol there. So we have a dot slash. [SPEAKER_04] So it'll recognize that it's a local one. And for this, I'm going to move to the main. Yeah. [SPEAKER_01] Okay.
SPEAKER_01
And on a good old fashioned way, going to run on the bash. So it's going to pop this up. Ask me which agent do I want to install on? I'm using cloud code, so I'm going to install it on cloud codes. If you're using any other agent harness, you can also install it as long as it supports it. I'm going to install it on a project level. So it's going to create a dot agent folder with the skill and link it to my dot cloud slash skills folder as well. So cloud knows where to find them. I'm going to symlink and we're ready to install. So if we, oh, let's not expose my key. I'll delete it. This is just for the workshop, so I'll delete it afterwards.
SPEAKER_01
Feel free to use my free credits for the time being. But essentially created the dot agents folder. Yeah. So I also have some more things that we're going to see afterwards. But the essential part, it has the skill that I've showed previously. Yeah, there it is. It's the skill. And then also created a symlink to the dot cloud folder. This is how the package works. And this way allows to either search on dot agents, which is becoming the standard, or on the dot cloud folder that it has. So let's see, let's run the same prompt again on a new session. Let me go back to the apps demo. Yeah. And start a new session. We should have this one enabled. Yeah, there it is.
SPEAKER_01
So Claude is aware of the Supabase security skill now. To run skills you can either just run your performance. You can pray that Claude imports your skill based on the description that you gave. You can include the keywords use and then the name of the skill that you have on the prompts. And this will almost one hundred percent of the times load your skill. Or if you're using Claude code, you can just slash and write the name of your skill. And this one hundred percent guarantees that Claude is going to import the skill. So for our use case, for the presentation, I'm going to, because I cannot afford that it doesn't load the skill.
SPEAKER_01
So let's wait, I need to reset the database to create the view again. And to be ready for the workshop. It's NPX. I'm just resetting the database applying the migrations from the start. I didn't create any migration file. He applied the migration directly to the database. So we now should be good to go. It's going to bring down the database and create a new one based on the schema that we define on the migration files and the seeded data. Yes.
SPEAKER_01
Yes. Yes. Yes. Yes. Yes. Yes.
SPEAKER_01
I didn't, it didn't create any a migration file. He applied the migration directly to the database. So we now should be good to go. It's going to bring down the database and create a new one based on the schema that we define on the migration files and the seeded data. Yes. Yes. Yes. Yes. Yes. Yes. Yes.
SPEAKER_02
[SPEAKER_01] Yes.
SPEAKER_01
[SPEAKER_04] I don't understand what I'm saying. They all how skills get loaded and they get still pretty just the one that they said they run and they understand that you are working with security or control of security just when it's working. Yeah, that's a fair point. So your whole question or observation is that the initial promise of skills when they were presented by Anthropic were. Yeah.
So since this is on the agent side, right? The agent decides when to load this. The best thing that you can do without explicitly either with the slash command or the use and then the name of the skill on your prompt is for you to play and play around with the description and run a bunch of tests either manually or automatically to check what actually works and not for the ones that [SPEAKER_04] that you're expecting. [SPEAKER_04] The agent to behave, right? So you define a bunch of scenarios where you think that the skill should be loaded and when the [SPEAKER_01] skill shouldn't be loaded. [SPEAKER_01] You test it out.
SPEAKER_04
[SPEAKER_01] You can test it on your machine like on this scenario. I don't want the skill to be loaded. Prompt the prompt on cloud code, let's say and check if the skill was loaded or not through the CLI. And then play around with the description to see what actually works or not. This without actually explicitly calling the skill. This is the best thing that you can do to test if the skill is being loaded correctly or not. We're still at the very beginning, very early stage of skills. Even for MCP, all this agent stuff, it's fairly new. So we're still standardizing things. We're still figuring out what works and what doesn't. Progressive disclosure was something that no one was talking about six months ago and now it's fairly it's fair to say that it's one of the north stars of agents development. So in six months from now, probably could be another thing. Or skills could be the standard or maybe Anthropic or OpenAI or someone else found a more efficient way to manage the context or provide more context to the agent. So we'll see.
All right, so the database was reset. Okay, so at least now we have the view but we don't have the information on your database. So now we should be able to run the same prompt again, but with the skill. So if we hit the prompt. You're saying it was quite fast. I don't think.
Yeah, but it didn't create one. Okay, let me try another thing. Instead of instead of this. Let's use to create. Let's see if it works now. Yeah, okay, so it loaded the skill. So now at least we should have the context to create that the RLS or the security invoker flag should be included when creating the view. And the steps should. The rest of the workflow should remain the same. So it will list my tables, right?
Exactly. Identified the tables and now if we look closely, we can see that we are now have. We now have these. The flag here is going to be on the migration. So let's see if with the flag this is the expected result. This is what happens when you try code a CLI. I have the UI duplicated, right? So it created the view. We should be able to see it. But Alice shouldn't. So what's happening? Wait, okay, so [SPEAKER_02] Do I have to reset now? [SPEAKER_02] Hmm. Interesting. [SPEAKER_01] Should have another.
[SPEAKER_01] Probably. Let me just see if I have it. Here. Here. Where did I put it? Let's do so. As a head count. I'm going to cheat here. Let's say. Okay. The both. And the employee should be able to see the image. Okay, basically we left troubleshooting. What is not happening? Probably from a different [SPEAKER_02] Different [SPEAKER_01] Policy that I've defined here. But now it's going to troubleshoot. Let's see if the skill actually improves the efforts here. If not, I have something on my sleeve. Because if you're not aware, Supabase basically has database advisors that you can use.
[SPEAKER_02] To try to identify early on, identify some potential vulnerabilities or schemas that are exposed, information that might be exposed before you're running into production. [SPEAKER_02] So if you can't figure it out by itself, I'm going to include on the skill to also run the advisors to check. So this is the main part of skills: you can. [SPEAKER_01] Oh, that's. You can see, well, it's. It's a very poorly written application, let me say. All right.
[SPEAKER_01] It's essentially the main part of skills. It's not if this specific demo works or not, it's that the behavior changed once it loaded the skill, right? It created with the security invoker part. And with that, that just shows how powerful it is, that you can create, you can change the behavior or guide the agent on demand based on information that you can put. You can think of the skill that uses as a prompt template that you can give to your agent. So let's just quickly troubleshoot. [SPEAKER_01] You can see well it's [SPEAKER_01] It's a very poorly written application let me say [SPEAKER_01] All right
[SPEAKER_01] It's essentially the main part of skills. It's not if this specific demo works or not. It's that the behavior changed once it loaded the skill, right? It created with the security invoker part, and with that, that just shows how powerful it is that you can create, you can change the behavior or guide the agent on demand based on information that you put. You can think of the skill that mds as a prompt template that you can give to your agent. So let's just quickly troubleshoot. [SPEAKER_01] Oh, is it even offering to apply a migration? [SPEAKER_01] Let's see if it doesn't break my app.
[SPEAKER_01] All right, so it seems too complicated, too complex. Anyway, going to move. [SPEAKER_01] What's that, the separate table? [SPEAKER_01] Let's see.
[SPEAKER_01] HR.
[SPEAKER_01] Anyway. [SPEAKER_01] Could you look at the complex, like a slash complex in your, just to see how your complex looks? [SPEAKER_01] Good point. [SPEAKER_01] I have a fairly amount of skills as you can see. I've been playing around with them. I also have some of the pre-installed mcp servers for that. But essentially, it would be more interesting if I've compared the context from before and after loading the skill.
[SPEAKER_01] So right now skills take 1.3 thousand tokens on my context, right? As you saw, I have more than just this one skill, but the skill was loaded so the whole information size skills that md was loaded to context. If we clear and run the context again, the skill amount, so this skill is not, it's not enough for you to see, but as you can see, the skills take quite less space than the mcp would.
[SPEAKER_01] All right. Oh, okay, I have a newer version of the cloud code. So for those of you who are not aware of this, Anthropic recently released the tool, the tool search tool, which is a mechanism for cloud code to load tools on demand. So it doesn't load basically progressive disclosure but for mcp.
[SPEAKER_02] Right. The main difference between this progressive disclosure or the tool search tool on cloud code and skills is that the progressive disclosure is built by design for skills, so it's already baked into the structure of the instance of the skill. While on mcp, it's still not a standard for all tools. So it works for cloud code, but for many other clients, it will just load all the tools straight to your context. So this is a thing for now just for cloud code. If you're interested about it, we are going to have the founder or one of the co-founders of the mcp speaking on the 10th, so on Friday. He's going to give a brief overview of the mcp roadmap. If anything, if nothing changed since last week when he presented this in New York on the mcp dev summit, you should bring this progressive disclosure part to the tools to bring it to the protocol itself.
[SPEAKER_01] Yes. [SPEAKER_01] Let's say that we have a very large database and we have to load in the context the schema of this database because we have to query the database using agents, okay? In your opinion, is it better to use a skill or an mcp to load this schema but progressively? Okay, is it possible to use the schema to progressively disclose, load the schema of this big database? In your experience, yeah. Oh, yes. Okay, so is your question more about how should we access it or the whole architecture of this pipeline to import the data? I just want to ask an agent to query the database and obviously the agent must know the schema of the database before or not?
SPEAKER_04
[SPEAKER_01] Oh, how can you teach the agent to query the database using the skills, using an mcp server or something like that? Okay, and if you use the skills, if you decide to use the skills to load the context of the agents with the schema of the database, is it possible to progressively load the schema within the context? Okay, gotcha. So let me break down the situation for you here. You'll have essentially two parts: one is what's going to be on the context, so what's going to be loaded and the specific information that you want to have on your scenario, and the second part is the actual mechanism, the extraction mechanism that you're going to use to load the information from the database. So for the second part, to load the information from the database, you can either use a script, so a skill that invokes a script, or an mcp tool. I would advise to use an mcp tool because you can use it if you're using in production or on remote project. You don't rely on your local environment. You don't have to manage the keys. And the tool, it's already standardized, and you already have the authentication baked into the protocol. So the agent never managed the application note token. It just runs the tool and it works. For it to progressive disclosure the information on the database, you'll have to.
SPEAKER_04
[SPEAKER_03] You can include it on a skill. Yeah, you'll be using the mcp tool. So on the skill, you'll probably state that use this tool to load, and in the tool implementation, you have to enable it to progressively load it right, so to load into chunks. It might be just enough from the tool parameters. The agent should figure it out by itself that if you put a parameter called buffer, for example, it should be able to load it in chunks, right? Instead of the whole table.
SPEAKER_04
[SPEAKER_03] But if you want to be 100 sure that it's going to load into chunks and use it properly, I would also package it with a skill and describe how to use this tool. So this is actually how both skills and mcp play along together. It's the tool to enable this connection, this integration, and the skill to describe how to use it.
[SPEAKER_03] Yeah, this is how I would implement this type of system. Thank you for the question. And it gave me the opportunity to basically talk about how to use both skills and mcp and not put one against the other. So now as I promised, we should be moving on. I'll have to give it more time to figure out because basically during the workshop, when I was preparing the workshop, I gave it a bunch of vulnerabilities. So if I just kept it simple and that one, the demo would probably work. Since I have more vulnerabilities exposed, that if I had time, I would try to solve it. It didn't for the moment, but you saw on both scenarios that the first one didn't have the security flag, security invoker flag, and the second one had. So at least we can imply that the skill was doing something. It did. The agent saw the information on the skill. It merged it with the system prompt or stored near the system prompt and changed the behavior accordingly to test this. So if you want to move this part, the skill into production, right, so it works on your.
[SPEAKER_01] So if I just kept it simple and that one, the demo would probably work. Since I have more vulnerabilities exposed, if I had time I would try to solve it. It didn't work for the moment, but you saw on both scenarios that the first one didn't have the security flag, security invoker flag, and the second one had. So at least we can imply that the skill was doing something. It did. The agent saw the information on the skill. It merged it with the system prompt or stored near the system prompt and changed the behavior accordingly to test this. So if you want to move this part, the skill into production, right? So it works on your machine. It's a tale older than time that it's working on my machine, but I don't know if it's going to work on your agents, on your machine, on your environment. So to test this or to automate this testing, and with this we can unlock having a pipeline. For example, if you change one thing on your skill, how can you reliably tell that it keeps doing what you're expecting? Didn't break the previous flow? So if I change one of the checklists, how can I ensure that the other ones were still working, right? So for this, evals could step in. So evaluations—it's a very broad term. You can basically evaluate anything. Since this is a markdown file, it's a free text file, you can evaluate basically anything. So it's fairly difficult for you. The most difficult part to create evals, I would say, is actually coming up with the scenarios because you would first have to know what's the expected behavior of your agents. So coming up with representative, actually good scenarios that represent a fairly good amount that cover a fairly good amount of use cases that you want to build are the most difficult. And there's still not a standardized structure to create evaluations. You can use or test it by importing a bunch of prompts and expected outputs from a CSV file, from a JSON file. You can use tools like Brain Trust or LangSmith to test it and to have an analytics and an observability layer on top of it. For this presentation, I followed what Agent Skills Open Standard defines to design the test cases. So if you're not aware of this website, this is the landing page of the Agent Skills Open Standards to try to standardize what a skill is and how it should behave. And they basically propose a very simple structure, a local way to test the skills organized by. You'll have an eval.json that essentially has a set of evals, so an array of eval scenarios. You'll put the prompt that you're going to give the agents, the expected output from the agent. This is only if you have an LLM as a judge. This is a technique used for non-deterministic evaluation. You would have, instead of a human, you can give the outputs of an evaluation run to another LLM, say it, define a success criteria and let the LLM whose role is to judge—in this case, that's why it's called LLM as a judge—to give it a grade basically. So this is one part that you can automate on your evaluations for non-deterministic workflows. You can either assert if a tool was called or you can give the results to an LLM and non-deterministically try to get the agents to grade the performance of the other agents. So basically have agents evaluating agents. So I followed this structure. The answer I gave the same input here, right? So the agent that is going to run this evaluation is going to get the same input that we had. The expected output is that the security invoker is true. So it's present on the app, sorry, on the view. And now I have a bunch of assertions that in this case I'm going to check deterministically, right? I prepared a Python script that essentially just resets the state of the database so we ensure that since we're running this locally and not on isolated containers like a Docker container, for example, we have to make sure that the systems always start from the same ground. So I'm going to reset the app. If you want to run the evaluations as well, you have to pick your own Anthropic key. Create, copy this. You can follow the readme inside the Superbase security here. You'll have how to set this up. But then I will run the cloud code CLI on it. I think it's on print modes, or you can remember what they called, but essentially, we I will run it as a binary headless. So the agent will receive the prompts that are on the evaluation as the task to perform, and I'm also going to give the condition. We're going to test two conditions: one with the skill and another without it. And essentially, so for you to see the condition, this is where the cloud code will run. And if the condition is with skill, we're going to load the skill.md into the system prompt, right? If you would actually like to mimic the behavior, you would run this on the Docker container. You will put the Agent Skills on the dot cloud slash skills directory inside of a Docker container and let organically let the cloud code find them and use them. For this presentation, this is a very simple setup. I've just basically appended it to the system prompt. So we're going to run the evaluations. Do I have the other? Yes, I do. Okay, I think we run it on the database.
SPEAKER_01
Okay. How is it not finding the skill for Superverse? No, okay. Oh, I have. I know what's going on. I have the wrong name. Change it. All right. So we started by running with the skill. So the first result that we should get is with the skill. It stopped. Now it's running without it, and then we're going to compare it. This will output a workspace iteration one folder, and we can compare both the output of with the skill and without it. While the without skill is loading, let's just quickly inspect what the with skill output gave.
SPEAKER_01
And essentially, you can see that it created the view with the security invoker, and then we have this grading.json file with a bunch of information like the assertions that we set on the eval.json. We have them here, and we can see that for this one, it graded as failing even though that created where is it not found. The view as security setting. Okay, I'm actually evaluating something wrong. So the problem here now is with the, yeah, is it the skill view? So since I was expecting this to create a PG class RL options instead of just inspecting the view, it's giving me that it failed. But the key part is it finished. It's not finished. Still running. Take a long time.
SPEAKER_01
Could be okay. Okay, and now we can inspect. Okay. So this is actually a good insight. So with these results, this is the tricky part of writing evals. So as in normal tests, the results will depend on how you implement them. [SPEAKER_02] Right. It's just code. So if you're evaluating something wrong or some or not the expected.
SPEAKER_01
Skill view. So since I was expecting this to create a PG class, RL options instead of just inspecting the view, it's giving me that it failed, but the key part is it finished. It's not finished, still running. Take a long time, could be okay. Okay, and now we can inspect. Okay, so this is actually a good insight. So with these results, this is the tricky part of writing evals. So as normal tests, the results will depend on how you implement them. [SPEAKER_02] Right, it's just code. So if you're evaluating something wrong or not the expected behavior, you're going to have wrong results. It might not be because the system is not working.
SPEAKER_01
We've tested manually and see that with the skill it created with the security, the security flag we can actually just inspect it here with the skill created. Let's see if on this one, surprisingly this time it did. It's non-deterministic behavior of Claude. But since I was evaluating something wrong, I was expecting it to create or inspecting the meta schema to check if the view, the security invoker was there or not instead of just inspecting the view directly. The results came a bit off, so it said that with the skill it failed and without the skill it passed. So if we inspect both outputs, they're basically the same.
SPEAKER_01
So with this, just to show you how tricky it is to write evals because although this can happen with regular tests, it's easier to catch because the output is deterministic. [SPEAKER_02] Right, it's just code. Here, if you're handing to an LLM to evaluate, you can sometimes hallucinate.
SPEAKER_01
So to finish, because we're also almost running out of time, to sum up the structure, this is the one that they recommend. I find it very easy to implement, to get started with. Later on, you can move on to more complex evaluation scenarios like running on Docker or in the sandbox to guarantee that you get a fresh environment with just one skill that you're testing on your set. But essentially, you would just put two conditions with and without the skill, compare the results, and see run them on the harness, the agent harness that you would like, and compare the results out there. This is basically your very first evaluation pipeline to test the skill automatically.
SPEAKER_01
From my end, that's all. I hope you found this workshop useful to get your skills leveled up and ready to production. I'm going, as I said in the beginning, to give a keynote tomorrow, a talk tomorrow about how we've implemented and created the Supabase skill for the product itself, how we're keeping it maintainable while ensuring that it provides value, and we're testing it into production. Thank you. Anyone has any doubts or questions? I'll also be, yeah?
SPEAKER_01
So I have a question about the number of skills that you typically install on your environment. Because with this progressive disclosure, it seems like we can basically keep adding different skills and the agent will automatically find them. Do you have any recommendation on how many skills to have, or is there any limit, or should we just basically keep adding?
SPEAKER_01
Uh, and it will magically work? Yeah, I'm probably not the best person to talk about this because it's easy for you to get into this rabbit hole, or just, especially when you experiment and get a bunch of skills. As you saw, I had plenty of them installed globally and I think it's fair to say and use them all on a daily basis. But it depends. If you're using them on your local machine, I think it's going to be pretty easy for you to get this messy environment where you'll have all of them installed or most of them installed. For your local environment, I wouldn't, for now, since it's very experimental, in my personal opinion, I would not constrain myself on space management or context management. The progressive disclosure is a very powerful thing that you can explore. You're sure if you have skills that you don't use, you're going to have them fill your context window, but the descriptions are so small that you can afford to not delete them if you don't want to. Into production, treat them as any artifact that you would have on your CI, so keep it clean. Into production, into your CI, I would keep only the exact skills that you're using in that specific case. Yeah, another piece of information that I could give you on the production part is that it's now more and more common for you to also export skills or make skills available on your repos as another piece of documentation. So treat skills that you put into production as actual documents, as you would read documentation. So it's important for you to keep them updated, include it on your include the updates workflow on your Cloud or MD or on your agents or MD.
SPEAKER_01
[SPEAKER_00] So you make sure that if anything changes, you will change the skill as well, like you would do on the documentation if a feature or workflow changes. From time to time, you can also create a job to check if the skill is still running a fair workflow. If somehow you could check if the skill has been loaded by your users in, if it hasn't been loaded by your users for a long time, does it still make sense to have it there? So yeah, this is basically the piece of advice that I could give you for skills into production based on my experience. For the rest of it, you'll have to come to the talk tomorrow to learn. We're putting it into production on Supabase. Any more questions? I'm going to be around throughout the whole event, so if you catch me, if you cross paths, feel free to ask me anything. Tell me about what you're building. Love to see if it's with Supabase, even more thrilled to hear about it. From my end, once again, thank you very much. You've been lovely today for 9:00 a.m., pretty cool, good energy. So just from my end, enjoy the rest of the conference and we'll see you around. Thank you.
SPEAKER_01
So in six months for now probably could be another thing So or skills could be the standard or maybe Anthropic or open AI or someone else found a more efficient way to manage the context or provide more context to the to the agent So we'll see basically All right, so The database was was reset Okay, so at least now we have the view but we don't have the information on your database So now we should be able to run the same prompt again, but we but with the With the skill so if if we hit the prompt You're saying it was quite fast I don't think Uh Yeah, but it didn't create one Okay, let me try another thing instead of Instead of This let's Use To Create
SPEAKER_01
Let's see if it works now Yeah, okay, so it loaded the skill so now at least should have the context Uh to Create that the the rls or the security invoker flag should be included when creating the view Uh and The steps should The the the rest of the workflow should remain the same so it will list my tables right? Exactly identified the tables and now if we look if we look closely We can see that we are we now have We now have These the the flag here is going to be on the on the migration so let's see if with the flag Uh, this is the expected result This is what what happens when you vibe code a CLI you know I have The the UI duplicated right so it created the view
SPEAKER_01
We should be able to see it But Alice shouldn't So what's happening? Wait, okay, so
SPEAKER_02
Uh do I have to reset now? Hmm Interesting
SPEAKER_01
Should have another Uh probably Let me just see if I have it Uh here Uh here Uh where did I put it?
SPEAKER_01
Let's Do so As a head count I'm going to cheat here let's say Okay The both And the employee should be able to see the image Okay, basically We left troubleshooting What is not happening? Probably from A different
SPEAKER_02
Uh different
SPEAKER_01
Uh policy that I've defined here Uh But now it's going to troubleshoot let's see if the the skill actually improves the the efforts here if not I have something on my sleeve Uh because if you're not aware of a super base basically has Um Database advisors that you can use
SPEAKER_02
Uh to try to identify early early on identify Um some potential vulnerabilities or schemas that are exposed information that might be exposed Uh before you're running into production Um
SPEAKER_01
So if If you can't figure out by itself I'm going to include on the skill to also run the advisors
SPEAKER_02
Uh to check so this is the the main part of skills is that you can
SPEAKER_01
Oh that's You can uh uh see well it's uh It's a very poorly written application let me say All right Uh It's essentially the the main part of skills uh it's not if if this specific um demo works or not it's that you the the behavior changed uh once the it loaded the skill right it created with the security invoker part uh and with with that that just shows how powerful it is that you can create um you can change the the behavior or or guide the the agent on demand based based on the on information that you if you put you can think of the skill that mds as a prompt template that you can give to to your agent so let's just quickly troubleshoot
SPEAKER_01
Oh is it even offering to apply a migration? Let's see if it doesn't break my my app All right, so it seems too complicated too complex anyway going to going to move What's that the separate table?
SPEAKER_01
Let's see HR Anyway Could you look at the complex like a slash complex in your uh just to see like how your complex looks Good point I have a fairly amount of skills as you as you can see I've been playing around with them I also have the some of the pre-installed mcp servers for that um that super base enables uh but essentially uh it would be more interesting if you if I've just um um if I've compared the context uh from before and after loading the skill so right now skills take 1.3 uh thousand tokens on my context right uh as it as you saw I have more than than just this one
SPEAKER_01
uh skill but the skill was loaded so the whole information size skills that md was loaded to to context if we clear and run the context again the the skill amount so this skill is not it's not enough to for you to see but as you can see the the skills take quite um um less space uh that that the mcp uh would uh from it's all right oh okay I have a newer um version of the cloud code so for those of you who are not aware of this uh entropic recently released the tool the the tool search tool uh which is a mechanism for for cloud code to load tools on demand so it doesn't load basically progressive disclosure but for mcp
SPEAKER_02
right um the key the main difference between names this progressive disclosure or the the tool search tool um on on cloud code and skills is that the progressive disclosure is built by design uh for skills so it's like already baked into the structure of the instance of the skill while on mcp is still not a standard for all tools so it works for cloud code but for many other clients it won't it will just load all the
tools straight to your context so um this is a for now a thing uh for just um for just cloud code if you're interested about it uh we are going to have the the founder or one of the co-founders of the mcp speaking on the 10th so on friday uh is going to give a brief overview of the the mcp roadmap um which if it's something if anything if nothing changed since last week uh when he presented this in new york uh on the mcp dev summit you should bring this uh this progressive disclosure part uh to the tools to bring it to the protocol itself so excuse me yes
SPEAKER_01
um let's say that we have a very large database and uh we have to to load in the context the schema of this database because we we have to query we have to query the database using agents okay in your opinion is it better to use a skill or um an mcp or some for um to to load this schema but progressively okay uh possible to use the schema to progressively disclose um load the schema of this big database in your experience yeah oh yes okay so is your question more about uh how should we access it or the whole architecture of this uh pipeline to import the uh the data i just want to to ask to a uh an agent uh to to query the database and uh obviously uh uh
SPEAKER_01
the agent uh uh must know the the schema of the database before or not he oh how can you teach the the agent to to query the database using the skills using the uh an mcp server or something like that okay and that if you use the skills if you decide to use the skills to uh to load the context of the agents with the schema of the database is it possible to progressively load the schema within the context okay gotcha um so let me break let me break the the situation uh uh break down the situation for you here you you'll have um essentially two parts one is uh what's going
SPEAKER_01
to be on the the context so what's going to be loaded and the specific information that you want to to have um on your on your scenario and the second part is the actual mechanism the extraction mechanism that you're going to use to load the information from the database so for the second part to to load the information from the the database you can either use a script so a skill that invokes a script or an mcp tool i would advise to use an mcp tool because you can use it if you're using on production or on remote project you don't rely on your local environment you don't have to manage the keys
SPEAKER_01
um and the tool it's already standardized and uh you already have the the authentication baked into the protocol so the agent never managed the application note uh token it's on the it just runs the tool uh and and it works for the for it to progressive disclosure the information on the
SPEAKER_03
database it will you'll have to um you can include it on on a skill yeah you'll be using the mcp tool so on the skill you'll probably state that use this tool to load and in the tool implementation you have to enable it to not load to progressively load it right so to load into chunks um it might be just enough from the the the tool parameters um the agent should figure it out by itself that if you put a parameter
SPEAKER_01
called buffer for example should be able to load it in chunks right uh instead of the whole table
SPEAKER_03
but if you want to have 100 sure that it's going to load into chunks and use it properly i would also package with with a skill and describe it how intense to to use this this tool so this is actually how both skills and mcp play along together it's the tool to enable this connection this integration and the skill to describe how to use it yeah this is how i would implement this this type of system uh thank you for for the for the question uh and it got me the opportunity to to basically talk
SPEAKER_01
about how to use both skills and mcp and not put it uh uh one against each other um so now as i promised we should be moving on um i'll have to give it more time uh to to figure out because i basically during the the workshop uh when i was preparing the workshop i've i gave it a bunch of vulnerabilities uh so if i just kept it simple and that one the demo would probably work um since i have more vulnerabilities exposed that i if i had time i would um try to solve it uh it didn't for for the moment but uh but you you saw on both uh scenarios that the first one didn't have the security flag security
SPEAKER_01
invoker flag and the second one had so at least we can um we we can imply that the the the skill was doing something it did uh the agent saw the information on the skill it it merged it with the system prompt or stored near near the system prompt and change the behavior accordingly uh to test this so if you want to move this uh this part the the skill into production right so it works on your machine it's a it's a tale older than time that it's working on my machine but i don't know if it's going to work on your agents on your machine uh on your um environment so to have this uh to test this
SPEAKER_01
or to automate this testing and with this we can unlock having a pipeline for example if you change one thing on your skill uh how can you reliably tell that it's uh it keeps doing what you're expecting didn't break the previous flow so if i uh change one of the checklists how can i ensure that the the other ones were still working right so for the this is where evolves um could step in so uh evaluations it's a very broad term you can basically evaluate anything since this is a markdown file it's a free text file you can evaluate basically anything uh so it's um fairly difficult for you to the most difficult
SPEAKER_01
part to create evolves i would say is actually coming up with the scenarios because you would first have to to know what's the expected behavior uh of your of your agents um so coming up with representative actually good scenarios that represent a fairly amount that cover a fairly amount of uh use cases that you want to to build are the most difficult um and there's still not a standardized structure to create evaluations you can use um you can test it um or by importing a bunch of prompts and expected outputs uh from a csv file from a json file you can use tools uh like brain trust or length use to test it
SPEAKER_01
uh um and to to have a an analytics and an observability layer on top of it uh for this presentation i followed the um i followed the um i followed the what agent skills uh open standard defines as to to design the test cases so if you're not aware of this website this is the landing page of the agent skills open standards uh to try to standardize what a skill is and how should behave and they basically propose a very simple structure local way to test the the skills organized by you'll have an eval.json that essentially has a set of evals so an array of of eval scenarios you'll put the prompt that you're going to give the
SPEAKER_01
agents the expected output from the agent this is only if you have an llm as a judge this is a technique used for non-deterministic evaluation you you would have instead of a human you can give the outputs of of a of an evaluation run uh to another llm say it define a success criteria and let the the llm was who's doing the uh whose role is is to judge in this case that's why it's called llm as a judge uh to give it a grade basically so this is one part that you can automate on your evaluations for non-deterministic workflows you can either assert if a tool was called or you can give the the results to
SPEAKER_01
an llm and non-deterministically try to uh get the the agents to to grade the the performance of the other agents so basically have agents evaluating agents um so i followed this this structure the answer i gave the same um the same input here right uh so the the agent that is going to run this evaluation is going to get the same input that we that we had the expected output it's that the security invoker uh it's true so it's it's present on the um on the app uh sorry on the view and now and then i have a bunch of uh assertions that in this case um i'm going to check it uh deterministically
SPEAKER_01
right i prepared a python script that essentially just resets the the state of the database so we ensure that uh since we're running this locally and not on isolated container like a docker container for example uh we have to make sure that the systems always start from the same ground so i'm going to reset the the app uh if you want to to run the evaluations as well you have to pick your own anthropic key uh create copy this you can follow the the readme inside the the superbase security uh here you'll have how to set this up um but then i i will run the the cloud code cli on it uh i think it's on
SPEAKER_01
print modes or you can remember what they called but essentially like we i will run it as a binary headless um so the agent will receive the um the the prompts that are that's on the evaluation uh as the task to perform and i'm also going to to give the condition uh we're going to test two conditions one with the skill and another way without it and essentially so for you to to see the condition this is where the cloud code will run and if the condition is with skill we're going to load the skill.md into the the the system prompt right um if you if you would actually uh would like to mimic the
SPEAKER_01
behavior you would run this on the docker container you will put the agent skills uh on the dot cloud slash skills um directory inside of a docker container and let organically let the uh cloud code find them and use them uh for for this presentation this is a very simple setup i've just basically append it to to the system prompt so so we're going to run the evaluations uh do i have the other yes i do okay i think we run it on the database okay how is it not finding the the skill for superverse no okay
SPEAKER_01
oh i have i have i know what's going on i have the the wrong name change it all right so we started by running with the skill so the first result that we should get is the with the skill it stopped now it's running without it and then we're going to compare it this will output a workspace iteration one um folder and we can compare it both the output of with the skill and without it uh while the without skill is loading let's just quickly inspect what what the uh with skill output gave um and essentially you can see that it created this the the view with the security invoker and then we have this grading.json file with a bunch of information like the assertions that we
SPEAKER_01
we've put on the eval we've set on the eval.json we have them here and we can see that for this one graded as the as failing even though that created where is it not found the view as security setting okay i'm actually evaluating something wrong so the problem here now it's uh with the yeah is it the skill uh view so since i i was expecting this to create an pg class uh rl options instead of just inspecting the the view it's giving me that uh it failed but the key part is it finished it's not finished still running take a long time could be okay okay and now we can inspect
SPEAKER_01
okay so uh this is actually a good good insight so um with these results this is the the tricky part of of writing evals um so as the uh as like normal tests uh the results will depend on how you implement them
SPEAKER_02
right it's just code uh so if you're evaluating something wrong or some or not the expected behavior you're going to have wrong results it might not be because the the system is is not working so
SPEAKER_01
we've tested manually and see that with the skill it created with the security um that the security flag we can actually just inspect it here with the skill created let's see if if on this one surprisingly this time it did it's a non-deterministic um the non-deterministic behavior of of claude um but since i was evaluating something wrong right i was expecting it to create the or or inspecting uh the um a meta schema to check if the the the view the the security invoker was there or not instead of just inspecting the view directly um the results came a bit off so it said that with the skill it failed
SPEAKER_01
and with the um without the skill it passed so and if if we inspect the both outputs they're basically the same so with this just to show you how tricky it is to to write evals because it's although this can happen on uh with the with regular uh tests um it's easier to catch because the the output is deterministic
SPEAKER_02
right it's just code um here if you're handing to to an llm to evaluate you can sometimes hallucinate
SPEAKER_01
so to finish uh because we're also almost running out of time to sum up the the structure this is the one that they recommend uh i find it very easy to implement to do to getting start with uh later on you can move on to more um complex uh evaluation scenarios like running on the docker or in the sandbox uh to guarantee that uh you get a fresh environment uh with just one skill that you're testing on your set um but essentially you would just put two conditions with and without the skill compare the results and see uh run run them on the harness the agent harness that you would like and compare the results
SPEAKER_01
uh out there this is basically your very first evaluation pipeline to to test the skill automatically from my end that's all i hope you find you found this workshop useful to to get your skills leveled up and ready to productions i'm going as i said in the beginning i'm going to give a keynote tomorrow a keynote no a talk tomorrow about how we've implemented uh uh and created the the super base skill for the product itself how we're keeping its main maintainable while ensuring that provides value and now we're uh testing it into production thank you anyone has uh any doubts questions i'll also be yeah
SPEAKER_01
so i have a question about like the number of skills that you typically install on your environment because with this progressive disclosure it seems like we can basically keep adding different skills and the agent will automatically basically the agent will automatically find them do you have any recommendation on how many skills to have or is there any limit or we should just basically keep adding uh and it will magically work yeah uh i'm probably not the best person to talk about this because uh it's easy for you to um get into this rabbit hole or just like especially when you experiment in getting a bunch
SPEAKER_01
of skills as you saw i had plenty of them um installed globally and um i think it's fair to say that and use them all uh on a daily basis um but it depends if if you're using them on your local machine i think it's pretty um it's going to be pretty easy for you to um get this messy environment where you'll have all of them installed or most of them installed um for look for your local environment i wouldn't for now uh since it's a very experiment in my personal opinion i would not uh constrain myself on like uh space management or context management about this the progressive disclosure it's a very powerful
SPEAKER_01
thing that you can explore in this case you sure if you have skills that you don't use uh you're going to have them uh fill your context window but the descriptions are so small that you can afford to not delete them if you don't want to into production treat them as any artifact that you would have on your ci so keep it clean uh into production into your ci um i would keep them only the exact skills that you that you're using in that specific case yeah another piece of information that i could give you on the production part is that it's now more and more common for you to uh also export skills or make skills
SPEAKER_01
available on your repos as like a another piece of documentation so treat it treat skills that you put into production as actual document as you would read documentation so it's important for you to keep them updated uh include it on your include the the updates workflow on your cloud or md or on your agents or md
SPEAKER_00
so you make sure that if anything changes um you will change this the the skill as well like you would do on on the documentation if a feature or workflow changes um he from time to time you can also create a um a job to to check if the skill is still running a fair workflow uh if somehow you could check if the the skill have been loaded by your users um in uh if it haven't been loaded by us by your users or a long
SPEAKER_01
time does it still make sense to have it there um so yeah this is basically the the piece of advice that i could give you for skills into productions based on my experience uh for the rest of it you'll have to come to the talk tomorrow to learn uh we're putting it into production on superbase any more questions i'm going to be around throughout the whole event so if you catch me um if you if you cross paths feel free to to ask me anything tell me about what you're building love to see if it's with superbase even more thrilled to hear about it um and from my hands once again thank you very much
SPEAKER_01
you've been lovely today for uh 9 00 a.m pretty cool good energy uh so just for my end enjoy the the rest of the the conference and we'll see you around thank you