Combine Skills and MCP to Close the Context Gap — Pedro Rodrigues, Supabase
Description
Agents working with Postgres will confidently create a view over a table with row-level security enabled and silently bypass that security in the process. Not because they can't reason. Because they don't know about the security_invoker flag, and nobody told them. Pedro Rodrigues from Supabase ran this exact test: same agent, same task, MCP alone versus MCP plus a skill. The one without the skill shipped a query that exposed data it shouldn't have. The talk covers what Supabase learned building their agent skill from scratch: critical security rules go directly in skill.md because agents will reliably skip reference files, skills should point to living documentation rather than duplicate it, and opinionated workflow guidance matters more than comprehensive coverage. Their evals ran across Claude and GPT models in three conditions and the result was unanimous. Skills without MCP underperform. MCP without skills misses environment-specific constraints. Together they close the gap that makes agents unreliable on real production systems. Speaker info: - https://x.com/rodriguespn23 - https://www.linkedin.com/in/pedro-neves-rodrigues/ - https://github.com/Rodriguespn
Summary
Generated by claude-haiku-4-5-20251001Combine Skills and MCP to Close the Context Gap — Pedro Rodrigues, Supabase
Main Topics
- Skills vs. MCP Debate: Evolution from MCP vs. Skills debate to MCP vs. CLI discussion
- Supabase Agent Skill Development: Creating comprehensive AI agent guidance for complex products
- Agent Limitations and Guidance: How agents need proper documentation and workflows, not just tools
- Testing and Validation: Using evaluations (evals) to measure skill effectiveness
- Best Practices for Skill Development: Principles for building product-specific agent skills
- Vector Database Use Cases: Integration with semantic search and embeddings
- Skill Distribution Challenges: Current gaps in skill distribution infrastructure
Key Points
The Problem with Agents Alone
- Agents are capable but struggle with:
- New or updated product knowledge beyond their training data
- Security pitfalls (e.g., Supabase's row-level security requirements)
- Admitting knowledge gaps and seeking fresh information
- Following optimized workflows for specific platforms
Real-World Example: SQL View Creation
- Without skill: Agent created insecure SQL view that bypassed row-level security
- With skill: Agent correctly set
security_invoker = trueflag, maintaining security - Lesson: Skills provide essential guidance that MCP tools alone cannot offer
Skills Architecture
Skills are folders containing:
- Front matter: Name and description (helps agent decide when to load)
- skill.md: Main instructions file
- Reference files: Optional bundled resources and scripts
Three Core Principles for Skill Development
- Don't Duplicate Information
- Treat skills as documentation guidance
- Point agents to single source of truth
- Be persistent in directing agents to documentation
- Explore SSH for documentation access (file system-like navigation)
- Understand Agent Laziness
- Agents default to training data to save costs
- Reference files are rarely accessed
- Multiple file lookups are nearly impossible
- Critical information must go in skill.md, not reference files
- Include information that's unlikely to change (like security checklists)
- Be Opinionated
- Know your product's best practices
- Guide agents toward optimal workflows
- Supabase example workflow:
- Run direct DDL operations on development/staging
- Use advisor tool for security/performance checks
- Fix issues first
- Create migration files last
Testing Results
Experimental Setup:
- 6 Supabase scenarios with 4 different agents
- 2 vendors tested (Anthropic, OpenAI)
- 3 conditions: baseline (no tools), MCP only, MCP + skill
Models Tested:
- Claude 3.5 Opus
- Claude 3.5 Sonnet
- GPT-4o
- GPT-4o mini
Results:
- Skills + MCP unanimously outperformed all other conditions
- Test completeness scores improved significantly across all models
- Conclusion: Guidance matters more than just context availability
Notable Quotes
> "The bottleneck is not the context. It's the guidance."
> "If something can get skipped, it will be skipped."
> "You know your product the best. Don't be afraid of guiding the agents on workflows that you think are the most effective."
> "Start minimal, start slow, and then iterate, expand. Don't be afraid to create new versions for it."
> "We haven't spent more time writing a single document since I wrote my master thesis."
Takeaways
For Product Teams Building Skills
- Point to single source of truth – Link to up-to-date documentation
- Be opinionated – Share your product's best practices and workflows
- Start minimal – Begin with critical information, iterate gradually
- Prioritize skill.md – Put essential information there, not in reference files
- Test with evals – Measure skill effectiveness using evaluation frameworks
Key Insights
- Guidance > Context: Agents don't need more context; they need better guidance
- Agent-Agnostic: Skills work across multiple AI vendors (adopting open standards)
- Workflows Matter: Specific workflows for your platform significantly improve outcomes
- Distribution is Unsolved: Skill distribution and package management remains a challenge in the ecosystem
Current State
- Supabase publicly launched their agent skill
- Open standards adoption increasing across vendors
- Vector databases emerging for enhanced semantic search in documentation
- Internal distribution strategies (repo-based, plugin formats) still being refined
Transcript
All right, I have the green light so we can get it started. Hello everyone, you might have noticed that the titles changed a bit from the one that we have on the schedule. That's because when I submitted the talk, the MCP versus skill debate was still going on, was a hot topic. I think now we settled on both different, they have their own roles and now I think the debate is more on the MCP versus CLI. So I thought it would be more useful to come and explain how we wrote our Supabase skill and the lessons that we got from writing this document. Because I've never spent more time writing a single document since I wrote my master thesis. Okay, so I know that writing a skill sounds simple, but it can be very complex, especially when you have a complex product like Supabase. For starters, I'm Pedro, I'm an AI tooling engineer at Supabase. I'm an MCP enthusiast, AI in general. Feel free to connect me on LinkedIn and I'm also a co-founder of the Lisbon AI Week in Lisbon. We'll be on late October this year. And if we're talking about me at the moment, I usually prefer doing this on dark mode. I don't know. How many of you prefer dark mode over light mode? The majority. I thought so. So let's do this presentation on dark mode instead. So I think we can all agree that agents are already smart enough, right? They are very capable of doing mundane tasks by themselves. But when you present a task about something that they haven't seen yet, or you've updated since they were trained, like your product, for example, they need the right guidance. For example, at Supabase we noticed that they would usually either miss some security pitfalls that we have, like role level security instructions that they have to set to not expose your app. They could just usually operate on stale knowledge on their training data. And they are very lazy and very stubborn to admit that they don't know and they do need to find fresh information. And also, we would like to guide them on specific workflows that we think are the most optimized for agents on our products. So, for starters, how many of you know or have written a skill before? Okay, so what I'm going to say is probably not new for most of you, but just to get an introduction on skills. So skills are folders containing instructions, scripts, and resources that agents discover, right? Progressively they discover. This is the main selling point of skills. And they have this envelope called front matter where they have the name and the description. This is how the agent is going to decide when to load the skill. Then they have the actual instructions inside the file, the main file called skill.md. And then optional bundled resources like scripts to perform actions or reference files that the main file can reference for more information that doesn't have to be loaded immediately to context. So we tested out at Supabase. We experienced giving the same agent, in this case with Claude Sonnet 4.6, the same prompt for a simple task. Like we had an app, a collaborative app, and we wanted to create this new SQL view, right? On top of a table that already had role-level security enabled. So the users could only see the information that belonged to them. We gave them just the MCP on one condition and the MCP plus the agent skills. And the result was what we expected. If you don't know in Postgres, if you create a skill over a table, on top of a table that has RLS enabled, if you don't explicitly pass that flag over there, the security invoker equals true, it will bypass the RLS. So basically the view will expose data that is not exposed by default on the table. So the agent with the skill, with the knowledge, was able to get this information implemented correctly and safely, while the one that only had access to the integration, to the MCP tool, did not. So for this we decided, just like this, we wanted to enable agents to know how to work correctly with Supabase. So we decided to announce, actually announcing today, this Supabase agent skill that I've been working on for this couple of months. And to make things official, I'm actually going to try something. I'm going to live tweet it on stage. So it's live. All right. So what exactly is this skill about and what lessons can I share with you? So if you're building a skill for your products, you can build one like we did. And happy to discuss all the details. Like this is free text and we haven't achieved a standard for it. So happy to chat about it later. So to break down some principles that we converged to. The first one is that don't duplicate information. Treat skills as documentation for yourself. And you will not duplicate your documentation, right? You already have documentation on your products. Just point the agents to it, to the most up-to-date. You have to be very stubborn with the model to ask them to search the web or your documentation. Provide the guidance. Tell where and how to find the documentation. But be very persistent with it to go for it. We also did run a little experiments. And this is still, well, as I said, an experiment and yet to be. We're still figuring out the details. But we also announced quite recently and you can see it, you can read it on our blog. We're basically exposing our documentation through SSH. The main reason behind it is that the agents can now look for the documentation like it was a file system. So they're very familiar with file systems in general. Or navigating them, finding files and information using Linux-based tools. If we expose, if we give them this interface to also do this but remotely, our premise is that they will have ease to navigate the agent. So we also would love after the talk or during the conference to see and to hear your opinions on this idea. The second principle that I have for you is that if something can get skipped, it will be skipped. What I mean by this is that besides new information on searching online, agents fetching information online or tool calling, it's expensive for agents, so they mostly default to their training data. The same is true for reference files. We've noticed that even if the agent loaded the skill, even if it had reference files there, it will be very lazy to load them. And even if it loads one reference file, if your problem requires more than one, the information that's in more than one file, it will most likely not load two files. It's almost impossible. Not even starting on three or four. So you have to be very critical about what you put on your skill.md file from the beginning. What I mean by this is that besides new information on searching online, so agents like fetching information online or tool calling, it's expensive for agents, so they mostly default to their training data. The same is true for reference files. We've noticed that even if the agent loaded the skill, even if it had reference files there, it will be very lazy to load them. And even if it loads one reference file, if your problem requires more than one, the information that it's in more than one file, it will most likely not load two files. It's almost impossible. Not even starting on three or four. So you have to be very critical about what you put on your skill.md file from the beginning. You put information that it's likely not to change, like in our case, a security checklist about Supabase that we didn't really want the agent to miss at all. So we decided this cannot be on a reference file. We actually started by putting it on a reference file, and it usually missed it. So we put it on the skill itself. So if you have any information that the agent can just not miss and defines your product, it goes to the skill.md file. Do not afford to put it on a reference file. And lastly, the third principle that I have for you when writing a product skill is to be opinionated. You know your product the best. You know how to work with it. You know how your users, or you should know how your users are using it. Don't be afraid of guiding the agents on workflows that you think are the most effective when working with your product. In our case, for example, managing a database schema, right? I actually haven't asked this and should have in the beginning. How many of you know Supabase and what Supabase does? Okay. So for the ones who don't know, we're basically a back-end as a service. We provide a storage database, authentication, and other features that you would need to create a back-end out of the box. So as I said, we provide a database, and the agent can interact and manipulate your schema, right? We found that this, for our platform, was the best workflow for the agents to efficiently manage the schema. So in this case, run direct DDL operations, like change the schema freely on your development or staging database. Once you're happy about it, we provide an advisor, so basically a tool to give any security or performance issues that the database could have, fix them, and only then create the migration file. So this will prevent the agent from creating a migration file every time it changes the schema. We found this to be the best workflow to manage the schema. So for us, it should be in the skill when working with Supabase. How we tested this skill. So we're living in very interesting times where we now can test free tests. We can now test documents. We can now test documentation, right? This would be completely bonkers to think years ago. Now I've basically been testing a markdown file. And how did we do this? Through evals. So for those of you who don't know evals, evaluations, short for evaluations, are basically tests that you can run, mostly like you would run on your CI. But now instead of evaluating code, you're evaluating an agent, an LLM, and its behavior, what tools it's calling, what's the reasoning, right? And so we ran on an early stage, we ran a set of six specific scenarios for Supabase. So ongoing Supabase projects in different scenarios, right? And we ran it against four different agents from two vendors in three test conditions. So we wanted to test on a baseline. So no MCP, no skills, with just the MCP server. And with the MCP server plus the skill. All this was done based on a test completeness score, a graded score, run on Braintrust. If you don't know them, they're around the responses of this conference. Go to their booth and talk about it. It's a very cool product. So the results are out here. The skills plus the MCP outperform any other conditions on every model that we tested. So we tested on Claude 3.5 Opus, Claude 3.5 Sonnet, and also on OpenAI's GPT-4o and GPT-4o mini. So I think we can conclude that it's pretty unanimous that the skills were actually improving the performance, the test completeness score, because they were providing the right guidance to the agent. So we had already the tools. We already had an MCP server. We just needed the right guidance on how to operate with Supabase. It's agent agnostic. More and more agents are adopting this open standard on skills. And currently, as I said in the beginning, the bottleneck is not the context. It's the guidance. So if you can take something from this talk when building a skill for your product, is that point to your single source of truth. Point to your documentation, basically. Be opinionated. You know your product. Don't be afraid to show it. And start minimal. Any model vendor, whether that's Anthropic, OpenAI, on any blog post about skills, you will read start minimal, start slow, and then iterate, expand. Don't be afraid to create new versions for it. So if you want to know more about it, we actually wrote a blog post. It's live today. You can check it on my Twitter account or on the Supabase blog. Or you can just run this command and install it on your project and start using it now. So that's all. I'll be around. And once again, thank you very much. I do have one more thing to show you. So we're running a giveaway. If you want to have a chance to win a Mac Mini, just scan the QR codes, sign up for Supabase, and good luck. Do I have time for any questions? If there's any? Yeah? Okay. I have one. Since we're all moving into RAGs right now, I'm just wondering how much demand do you guys see for vector database use cases in Supabase right now? It's an interesting question. It depends. You're asking about how many customers are having demand for it? I've been seeing more and more as the days pass. The use cases for vectors are mainly for embeddings, right? And semantic search. There are many use cases for embeddings. The one that I'm most interested in is Semantic Search, right? That you could use to provide even more context. For example, through the SSH, supplementing the docs through the SSH, you can now, instead of naively letting them navigate with the bash tools, you can augment the tools, the already known bash tools, into providing some sort of semantic search. So I do see a very big potential in vectors. On our data, definitely customers have been exploring more of this solution. Thank you for the question. Thank you. Yeah? Thank you for the talk. Very interesting. How do you distribute the skills inside your organization? Do you have a repository, or do you have a package manager? That's an amazing question, because currently one of the downsides, or I would say the constraints of using You can augment the tools, the already known bash tools, into providing some sort of semantic search. So, I do see a very big potential on vectors. On our data, definitely customers have been exploring more this solution. Thank you for the question. Thank you. Yeah? Thank you for the talk. Very interesting. How do you distribute the skills inside your organization? Do you have a repository, or do you have a package manager? That's an amazing question, because currently one of the downsides, or I would say the constraints of using skills is their distribution system, right? We're still finding that there are still some players trying to reclaim either the registry or the way to distribute it. So, Vercel came up with the skills package. We're seeing plugins, right, that you can bundle with MCP servers and other things, but they are model-specific. So, this is to distribute in general, skills in general. It's already a problem for itself that we haven't solved it yet. Internally, we are packaging the skills on the repos themselves. So, if you want to create a plugin, you just create a .cloth plugin or a .cursor plugin or whatever in that repo, right? And then it's available. Or it's discoverable if the repo is open sourced or you have access to it, you can use the skills package to fetch it. So, yeah, and this is how I've seen the other companies distributing these skills, like a repo, a skill, trying to package the skills into the knowledge. Thank you. Thank you. Are there any more? I think we have time for one more. Yeah. Yes. I recently failed self-improving skills. Should we do a call-up? Sure. Let's talk after. All right. So, once again, thank you very much. It was a pleasure to be here. I'll be around. Thank you. How are you? for Superbase. So ongoing Superbase projects in different scenarios, right? And we ran it against four different agents from two vendors in three tests conditions. So we wanted to test on a baseline. So no MCP, no skills with just the MCP server. And with the MCP server plus the skill. All this was done based on a test completeness score, four graded score, run on Braindrust. If you don't know them, they're around the responses of this conference. Go to their booth and talk about it. It's a very cool product. So the results are out here. The skills plus the MCP outperform any other conditions on every model that we tested. So we tested on cloud code for Opus 4.6 and Sonnet 4.6 and also on codecs for GPT 5.4 and GPT 5.4 mini. So I think we can conclude that it's pretty unanimous that the skills were actually improving the performance, the test completeness score, because they were providing the right guidance to the agent. So we had already the tools. We already had an MCP server. We just needed the right guidance on how to operate with Superbase. It's agent agnostic. More and more agents are adopting this open standard on skills. And currently, as I said in the beginning, the bottleneck is not the context. It's the guidance. So if you can take something from this talk when building a skill for your product, is that point for your single source of truth. Point to your documentation, basically. Be opinionated. You know your product. Don't be afraid to show it. And start minimal. Any model vendor, whether that's Entropic, OpenAI, on any blog post about skills, you will read start minimal, start slow, and then iterate, expand. Don't be afraid to create new versions for it. So if you want to know more about it, we actually wrote a blog post. It's live today. You can check it on my Twitter account on Superbase blog. Or you can just run this command and install it on your project and start to use them now. So that's all. I'll be around. And once again, thank you very much. I do have one more thing to show you. So we're running a giveaway. If you want to have a chance to win a Mac Mini, just scan the QR codes, sign up for Superbase, and good luck. Do I have time for any questions? If there's any? Yeah? Okay. I have one. Since we're all moving into rags right now, I'm just wondering how much of a demand do you guys see on vector-executive basis? Superbase right now? It's an interesting question. So, I mean, it depends. You're asking about how many customers are... How much are having demand for ? I have been more and more as the days passed. The use cases for vectors are mainly for embeddings, right? Yeah. And semantic... There are many use cases for embeddings. The one that I'm most interested about is Semantic Search, right? That you could use to provide even more context. For example, through the SSH, supposing the docs through the SSH, you can now, instead of naively let them navigate with the bash tools, you can augment the tools, the already known bash tools, into providing some sort of semantic search. So, I do see a very big potential on vectors. On our data, definitely customers have been exploring more this solution. Thank you for the question. Thank you. Yeah? Thank you for the talk. Very interesting. How do you distribute the skills inside your organization? Do you... Just past 10, you have a repository, you have a repository, or do you have a package manager? That's an amazing question, because currently one of the downsides, or I would say the constraints of using skills is their distribution system, right? We're still finding that there's still some players trying to reclaim either the registry or the way to distribute it. So, Vercel came up with the skills package. We're seeing plugins, right, that you can bundle with MCP servers and other things, but they are model-specific. So, this is to distribute in general, skills in general. It's already a problem for itself that we haven't solved it yet. Internally, we are packaging the skills on the repos themselves. So, if you want to create a plugin, you just create a .cloth plugin or a .cursor plugin or whatever in that repo, right? And then it's available. Or it's discoverable if the repo is open sourced or you have access to it, you can use the skills package to fetch it. So, yeah, and this is how I've seen the other companies distributing these skills, like a repo, a skill, trying to package the skills into the knowledge. Thank you. Thank you. Are there any more? I think we have time for one more. Yeah. Yes. I recently failed self-improving skills. Should we do a call-up? Sure. Let's talk after. All right. So, once again, thank you very much. It was a pleasure to be here. I'll be around. Thank you. How are you.