AI Engineer

AI-Native Organisations Run on Skills: How to Structure and Scale Them — Imad Touil, QuantumBlack

1872 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI-native organizations should treat skills as governed, reusable, composable units of executable organizational know-how, supported by a centralized registry and workflow platform rather than left as ad hoc prompts or team-local artifacts.
  • Why it matters: This is directly applicable to agent-system architecture: unmanaged skills create duplicated logic, variable quality, security exposure, rising token costs, and inconsistent behavior across teams, while governed skills can turn institutional practices into portable, auditable runtime controls.
  • Best use: Use the talk as an operating-model and control-plane blueprint for designing an enterprise skill registry, governance process, and eventually a centralized workflow marketplace; do not treat it as a source of validated performance benchmarks.

Executive Summary

Touil frames skills as the missing organizational layer in the agentic software stack. Agent harnesses already have context managers, tools and MCPs, memory/state, and sub-agents, but those components alone do not encode an organization’s repeatable operating knowledge. Skills do: they specify deterministic, task-focused instructions and scripts that workflows load at runtime.

The argument is that real enterprise delivery is far broader than a simple specify-plan-task-implement loop. Product strategy, research, discovery, data preparation, SDLC variants, infrastructure operations, incident response, and continuous optimization each require different workflows. Since no universal workflow fits every business unit or product type, organizations need a composable catalog of specialized skills that can be reused across workflows and, ideally, across agent harnesses.

The speaker’s central recommendation is a federated creation but centrally governed distribution model. Individuals create and test skills; teams collaborate and improve them; a central platform supplies searchable metadata, discovery via MCP, CLI or IDE delivery, versioning, dependency management, access control, evaluation/observability, ownership, and security review. Architects, engineering leaders, infrastructure leaders, and cyber leaders jointly govern the relevant domains.

The strongest practical example is regulatory implementation: a disclosure-review workflow can compose retention-policy, GDPR, disclosure-standard, and filing-template skills at runtime, producing an auditable report and a remediation loop into the codebase. Touil cautions that auto-evolving skills are premature without this control plane: automation can amplify unmanaged duplication, unsafe code or prompt injections, and policy drift just as readily as it improves productivity.

Key Takeaways

  • Claim: Skills are the layer where organizational know-how becomes deterministic and executable; hooks, MCP services, and sub-agents are enabling components but do not by themselves provide structured workflow behavior. | Evidence: Touil defines workflows as runtime harness blueprints and distinguishes hooks (event-triggered actions), MCP services (tools), and sub-agents (context/task delegation) from skills, which encode the task-specific operating logic. | Implication: Ken should model skills as first-class, versioned operational assets—not merely reusable prompt snippets—and separate them conceptually from tools, agents, and orchestration. | Caveat: Skills are only one workflow component; solving skill management alone does not make an organization’s end-to-end workflows reliable.
  • Claim: Enterprise agent workflows must reflect the full business-to-operations lifecycle rather than a narrow coding-agent SDLC. | Evidence: The talk expands delivery from specification, planning, task decomposition, and implementation into product strategy, market and competitive research, customer interviews, discovery, data-product preparation, differing SDLCs, platform engineering, launch, incident handling, and optimization. | Implication: Build a composable workflow architecture with domain-specific skills and workflow templates instead of forcing all agent work through one generic coding harness. | Caveat: Touil does not prescribe a single canonical enterprise workflow; he explicitly says workflows vary by organization, product, department, platform, and customer-facing versus internal use case.
  • Claim: A well-designed skill should follow microservice-like design principles: reusable, modular, discoverable, portable, specialized, composable, consistent, deterministic, and cost-efficient. | Evidence: Touil analogizes the problem to the microservices movement and argues against a monolithic skill; each skill should target one specific task, compose without conflicts or duplication, and work across workflows and harnesses when standards permit. | Implication: Establish a skill contract: narrow responsibility, clear invocation criteria, declared dependencies and permissions, tests, ownership, and semantic versioning. | Caveat: Cross-harness portability depends on shared standards and compatible implementations; the example of moving a skill from Claude Code to Cursor is conditional on broad adoption of the same standard.
  • Claim: Skills can improve quality and reduce cost by progressively disclosing only the relevant organizational guidance at runtime. | Evidence: Touil cites a recent 'SkillsBench' comparison in auto-engineering and cybersecurity, saying deterministic skills produced clearly better outcomes than the same tasks without skills; he also argues that loading the right skill at the right time reduces context-window/token use. | Implication: Measure skill value locally through task success, human rework, token consumption, latency, policy compliance, and regression performance rather than relying on generic benchmark claims. | Caveat: No benchmark figures, task definitions, or independent validation are provided, so the performance claim is directional rather than a basis for forecasting ROI.
  • Claim: Without governance, proliferating skills become a new class of technical debt and security risk. | Evidence: The speaker identifies duplicated skills across teams, quality degradation as models change, poor discoverability and absent ownership, non-composable designs, prompt injection or unsafe scripts in public skills, and overly broad access to sensitive business logic. | Implication: Treat third-party and internally authored skills as executable supply-chain artifacts: require provenance, code and prompt review, security scanning, explicit owners, lifecycle policy, and least-privilege access. | Caveat: Centralization should govern distribution and policy, not prohibit individual experimentation and team-level creation.
  • Claim: The target operating model is individual creation, team collaboration, and organization-wide reuse through a centralized skills platform. | Evidence: Touil recommends a searchable registry with metadata; MCP-based search; CLI delivery to local IDEs or sandbox environments; dependency visibility; version and lifecycle management; access control; evaluation and observability; and governance shared by architecture, engineering, infrastructure, and cyber leaders. | Implication: Prioritize the registry/control-plane design and accountable governance model before scaling skill creation or enabling autonomous skill improvement. | Caveat: Technology does not resolve ownership or policy decisions by itself; the speaker explicitly identifies governance roles and organizational structure as the point where technology stops being sufficient.
  • Claim: The same registry and governance approach should extend from skills to reusable end-to-end workflows, while auto-evolving skills should wait for adequate guardrails. | Evidence: Touil proposes a centralized workflow platform from which an engineer could pull, run, test, improve, and republish a workflow such as infrastructure provisioning. He calls skills registries and evaluation immediate priorities, but warns that autonomous closed-loop skill evolution magnifies impacts without governance. | Implication: Sequence investment as registry and standards first, evaluation and observability second, controlled workflow reuse third, and autonomous evolution only after robust release gates and rollback mechanisms exist. | Caveat: The talk offers no mature evaluation standard; Touil’s current practical suggestion is static tests and checks against Anthropic best practices.

Detailed Brief

Agentic stack and control-plane framing

  • Claims: The inner loop is the coding-agent harness, composed of context management, tools/MCPs, memory/state, and a skills loader.; The outer loop is workflow orchestration, combining skills, sub-agents, MCP servers, and optional hooks.; The enabling plane beneath both loops includes sandboxed environments, an MCP gateway, a model gateway, a graph or knowledge graph over IT systems and code, a skills registry, and eventually a workflow marketplace.; Context must be assembled from project instructions, tool and MCP schemas, user-conversation memory, and retrieved files or codebase material.
  • Evidence: The speaker compares project instructions to files such as a Claude Code MD or AGENTS.md-style instruction file.; The MCP gateway is positioned as a way to manage and simplify organization-wide MCP tools, while the model gateway manages and optimizes local/open and frontier LLM usage.; The graph is intended to abstract core IT systems, codebases, skills registries, and workflow marketplace assets.
  • Caveats: The stack is a conceptual architecture, not an implementation reference with product choices, APIs, or deployment details.; A knowledge graph is proposed as an abstraction layer, but the transcript does not establish that it is necessary for every organization or use case.
  • Implications: Architectural ownership should distinguish the tool plane, model-routing plane, context/retrieval plane, skill registry, and workflow orchestration plane so that skill governance is not buried inside a single IDE or agent vendor.

Regulatory compliance as the exemplar use case

  • Claims: Compliance requirements are a high-value skill domain because they need to be applied consistently whenever agents manipulate customer data or build data-related features.; A workflow can dynamically compose multiple regulatory skills and return a deterministic, auditable outcome rather than depend on each engineer manually steering an agent.
  • Evidence: Touil’s example combines data-retention policy, disclosure standards, GDPR rules, and filing templates in a regulatory disclosure-review workflow.; The proposed output includes a stored audit report, specific findings, and a feedback loop to improve the codebase.
  • Caveats: Deterministic execution does not establish legal correctness; policy content still requires ongoing expert ownership and validation as regulations change.
  • Implications: Regulatory, security, and platform standards are strong initial candidates for a governed skill library because reuse and auditability matter more than one-off creative flexibility.

Notable Concepts & Terms

  • Skills: Reusable, task-specialized units of instructions and potentially scripts that encode organizational know-how for agent workflows.
  • Agentic software stack: Touil’s two-loop architecture: an inner agent harness and an outer workflow layer, backed by shared enabling infrastructure.
  • Harness blueprint: A workflow definition that shapes how an agent or coding harness behaves at runtime.
  • Progressive disclosure: Loading only the relevant skills and guidance at the moment they are needed, intended to preserve context capacity and reduce token cost.
  • Skills registry: A centralized, searchable catalog with metadata, ownership, versions, dependencies, permissions, evaluation, and observability for skills.
  • MCP gateway: A centralized layer for managing and simplifying access to MCP tools across the organization; it can also support skill discovery from a registry.
  • Internal developer portal (IDP): The existing service-catalog model Touil uses as an analogy for discoverability and ownership; he expects IDP providers to add skill-registry capabilities.
  • Auto-evolving skills: Closed-loop systems that automatically improve skills; presented as potentially powerful but dangerous if deployed before governance and guardrails.

Operator Notes / Why Ken Should Care

  • Define a minimal enterprise skill manifest now: purpose and invocation criteria, inputs/outputs, compatible harnesses, owner, domain, dependencies, permissions, evaluation suite, version, provenance, and deprecation status.
  • Stand up a gated registry pilot around a policy-heavy, repeatedly executed domain such as data handling, security review, deployment, or infrastructure provisioning—not a broad open marketplace.
  • Require CI checks for every publish or update: static structure validation, test fixtures, model-regression evaluation, secret and dependency scanning, prompt-injection review, approval policy, and rollback compatibility.
  • Instrument registry and runtime metrics before claiming productivity gains: discovery-to-reuse rate, duplicate-skill rate, successful invocation rate, task quality, rework, token/latency cost, security findings, and ownership coverage.
  • Separate skill ownership by domain from platform stewardship: domain experts own policy logic, while the central platform team owns registry reliability, distribution, access controls, telemetry, and release gates.
  • Do not enable autonomous skill mutation or automatic republishing until evaluation thresholds, human approval rules for sensitive domains, immutable audit logs, and rapid rollback are operational.

Source/Metadata

  • Title: AI-Native Organisations Run on Skills: How to Structure and Scale Them — Imad Touil, QuantumBlack
  • Transcript words: 6477
  • Duration seconds: 1230
  • Timestamp note: No timestamps or chapter markers were provided. The transcript contains a substantial duplicated passage in its latter portion.
Full transcript 3656 words · 29 min read
0:13

Thank you for joining me today. My name is Imat Will. I'm a distinguished engineer at Quantum Black. And in today's talk, I want to really cover how AI-native organizations run on skills.

0:18

But before I really get started, I just want to do a quick exercise, a show of hands. Can you raise your hand if you have already created and are using skills? Amazing. Now, can you keep your hand up if you are using and sharing them within your teams? Great. Now, keep your hands up if you have governed and maintained skills across your organization. Amazing. I see a few hands. But that is what this talk is about. Today, I'm really looking to break down why this is really critical, why it's important, and how you can actually adopt it across your organization.

0:25

But before I get started, what I really want to cover is the agentics software stack. So the agentics software stack has two loops. The first loop is all you know, right? It's the coding agents or the coding agents harness, right? And then you have some core components. You will have your context manager, the tool and MCPs, memories and states, and skills loader, right? But then there's an outer loop, which is your workflows, right? Those have skills, sub-agents, MCP servers that you use, and some hooks. Sometimes you need them.

0:37

To have this running properly, you will need some enablement components at the bottom. So what you would have is an environment sandbox. You would have your MCP gateway to manage and simplify all of the MCP tools across your organization. A model gateway to, again, manage and optimize for all of your LLMs, either open servers running locally or front-term models, and also a graph, knowledge graph, that abstracts your IT core systems, your code base, your skills registry, and, at the end, the workflow marketplace.

0:44

And then you have your context layer, right? The context layer will bring all of what is needed to get the task done. So this is the project instruction. Think about it as the cloud code MD file, the agency MD file. Your tool and MCP schema are actually to understand which tool to use and when. Your memory, right? Conversation history with the end user, the human is in the loop. And finally, the retrieved contents that you can pull from files, your code base, etc.

0:51

Now, what I want to really focus on today is the workflow, right? And I think this is where we think it's a pretty simple workflow in day-to-day life when you're trying to actually create an end-to-end product software delivery lifecycle in your organization. In reality, we all have seen the four steps. Specify to define what you want to build, the design or plan to plan what you're looking to build. Then you go to the tasks, you break it down into tasks, and then, finally, you start implementing, right? I think this looks familiar. I think this is how most of our coordinations actually are shaped today.

0:56

In reality, that is not how it is composed at scale when you look at the organizational complexity. This is just one step in the journey. This is like building a product increment. When you really look at the overall end-to-end lifecycle of something that's a business one that builds to capture the value out of it all the way to shape it to the client, you will start, first of all, by defining your product strategy, what to build and how to build it, right? You set the different success metrics, you identify and break down your plan, roadmap for your products. And to do this, you may need a lot of insights, right? So then you do your market research, you do a competitive analysis, you bring this as an input with some customer interviews.

1:00

Then you go to the discovery side, right? So then you start discovering, okay, now I understand what to build. I'm going to break this down into some problem statements, find the solution, validate the solution, and then probably experiment and then create user stories.

1:06

And before we start building, in reality, you actually need to prepare your data, right? In some cases, you will actually need to clean up your data catalog that will support the build of your products, or maybe adjust some of the endpoints, connection, and integration to your core systems that will help you actually build your products. And this is where the data product delivery comes. So you build your data pipeline, you validate your data quality, and you put your catalog data assets ready for development.

1:10

Then we go back to the product increment. That is where we start. But then, building up products in every organization that I have been serving for the past 18 years in my career, I can see that in one organization, you will find different SDLCs scattered across the organization. Some of it is actually for a mobile application, others are different departments or different platforms. Some of it is an internal platform that is for your employees, others are actually customer-facing. So it is not one workflow that can actually build anything you want for your organization.

1:16

When you figure out what to build and how to build it, you need to run it. Then you come to the platform engineering ops, right? That is, again, when you have your provisioned infrastructure, thinking about how you build your infrastructure as code modules, et cetera. And then you launch your products. And the moment you launch it, then you start the journey of optimizing the performance of your products, trying to look for any incidents to resolve. And then you start the loop again, right?

1:21

So at scale, when you look at really building a digital platform, not a very simple product that you can build and deploy, the landscape is way more complex than expected. And what you're looking at here is literally probably 10, 20% of what it is. And it's really different from organization to organization. Now, going back to the stack, right? When you look at the workflow, there are four core components. The first one is hooks, MCP services, and sub-agents. Those are given, but they don't really bring the right structured value to your workflows, right? That's why skills is one of the critical components in your workflows.

1:32

Hooks, basically, what it does, it just pre-grounds events to do something along your workflow. The MCP service, we all know that you may need an MCP tool, but tell me, who actually built a lot of MCPs? We just use MCP tools that are actually provided by the tool that we used before, right? So we don't really own those. The sub-agents are just to minimize the context window. We just delegate to sub-agents to execute a specific task that is needed. So at the end of the day, you will find all of your know-how is actually at the skills level. And if you don't have the right structure of your skills, then you're not really having a deterministic workflow.

1:36

And one thing to mention is workflows, think about them as harness blueprints that actually shape the behavior of your coding harness, for example, in the runtime. Now, looking at the rise of skills adoption, right? So just eight months ago, Entropic published the first article about skills, right? Two months later, we had the standard, an open standard that is adopted, and a lot of agent harnesses started adopting this new standard. Around February this year, we have seen most of the agents actually adopt this. Even though you don't see them, right? If you pay attention when the agent is thinking, you can see that he's pulling skills as he's going and doing the task.

1:48

The other indication here that you need to pay attention to is actually the number of skills that are created, right? This is literally just a snapshot that I did across some public GitHub repos and public skills registries, right? There's way more than this publicly and within your organizations, right? So the creation of skills and the demand is rising, but we need to understand why.

1:56

In the latest skills bench, comparing the latest models against running the same tasks against auto engineering and cybersecurity without skills, it did well, right? Because that's what it is expected to do. And it's going to continue improving day after day, right? But then, when we applied skills that are more deterministic, the outcome was clearly higher than without it.

2:03

Now, if you think about the anatomy of skills and how you actually design and implement skills in your organization, this is not a new problem that we are trying to solve here, right? We have solved this with the microservice movements, right? So your microservices need to have all of these design principles, right? This is a software problem that we used to solve. It's similar. So skills need to be reusable, right? They need to be modular, need to be discoverable, so you can discover your skills. So if you are sitting in one team and you need the skills, you can actually automatically discover and capture the skills. It's portable so that you can actually use skills across workflows, but also you can use skills

2:08

improving day after day, right? But then when we applied skills that are more deterministic, the outcome was clearly higher than without it. Now, if you think about the anatomy of skills and how you actually design and implement skills in your organization, this is not a new problem that we are trying to solve here, right? We have solved this with the microservice movements, right? So your microservices need to have all of these design principles, right? This is a software problem that we used to solve. It's similar. So skills need to be reusable, right? They need to be modular, need to be discovered, you can discover your skills. So if you are sitting in one team and you need the skills, you actually can automatically discover and capture the skills. It's portable, that you can actually use skills across workflows, but also you can use skills across harnesses. Again, if everyone adopted the same standard. So if I'm having a skill on cloud code and I want to move it to cursor, it's going to just work. Specialized skills, that is where the value, you should not build one skill, a monolith again. It should be specialized to define one task specifically. Composable skills need to be designed in a way that can actually compose, so you don't have duplication across your skills when you're trying to run them that conflicts with each other. Consistence, that is actually one of the key items of skills, is actually consistency and deterministic. And finally, cost-efficient. And for cost, I can go for another hour talk, but it's the skills come to solve a key problem around the context window, right? It's actually put in with the progressive disclosure pattern, the right skills, the right amount of skills, at the right time to solve the right problem. And that reduces the token usage. This defines a new unit, right, that makes your know-how in your organization executable, portable, and cheap.

2:13

On the right-hand side, you just have a very simple example of data retention policy, right? When it came to regulation, you need to understand and make sure to instruct your agents why you're manipulating, for example, your customer data. You should make sure that this data is manipulated according to the regulation, right? And that brings me to the next example. On the left-hand side, you can see that there's a, think about a catalog of skills. And on the right-hand side is your harness, and the output is on the right-hand side, right? So the composable skills at the regulation level, you have the skill that I just showed earlier, which is the retention policy, but you will need disclosure standards. You will need the GDPR rules that need to be respected. You will need the filing templates, right? So all of this, think about it, is defining how any data, any feature that is built across your web, mobile, different applications across your organization is really respecting these rules. And this gets pulled automatically at runtime by the regulatory disclosure review workflow. And the outcome is expected. It's deterministic. You have an audit report that you can actually store. You have specific identification of if there's anything to improve, and there's a loopback to improve your codebase.

2:19

However, if we don't govern skills, we will start creating a new class of technical debt, right? First of all, you will find out that you are having a lot of duplication in your organization. So if teams are not collaborating and everyone is using the same technology stack, the same infrastructure, you're for sure building the same skills over and over again without sharing them. Quality. If you don't test and make sure that you're maintaining and validating your skills, not against your task, but also against the latest models that come, right, then the quality starts degrading over time. You should be able to discover your skills. In reality, without governance, you cannot really discover it. Think about it as the backstage, the IDP, right? It comes to solve a problem where, okay, I need to understand who owns this microservice back in the days, right? I don't need to talk to anyone. I just need to tap into the service catalog and immediately find who actually owns it. That brings the ownership part. If you don't have an owner, then no one will be able to maintain, scale those skills. Composability is not something that comes by default. You need to have a governance way to align what to build and how to design it. Think about the domain-driven approach that we have been taking also for many years, right? It's similar to how you shape your skills catalog. Security. Again, all of us experiment with the public skills, right? But when you think about it, some skills may have some prompt injection, and skills actually do have scripts because that's the deterministic part of it, because it's going to run a specific script for a specific task. So if you don't have, again, a pipeline that checks your security, you may be pulling something that is insecure. And permissions. Not every skill is actually something that anyone in the organization should access. Some skills may have some business logic that is very sensitive, right? So the access control is also crucial at this stage.

2:25

Now, how to bring this to your organization? First of all, you need to allow, at the individual level, to create, test, improve, and use those skills. Again, you shouldn't be around them. You should be structured way. There are different tools out there. You just need to decide which tool actually you want to agree on, and you use that mechanism at the individual level. The moment you create a skill, you need to be sharing this with your team, right? Your team starts collaborating to improve the skills. And think about it, you build the same technology stack, building the same products. So it's going to evolve very quickly. But then you move on to a very critical point, which is the centralized platform. That is where all of what I've been covering so far comes to play. You need a centralized platform that has a catalog with metadata in it that actually can discover skills and could be searchable. You can have an MCP that actually plugs into this catalog, search for the skill, and a CLI to pull the skills back to either your IDE, if you're locally, or to your sandbox in your factory. Then you have the dependencies. You need to understand the dependencies between the skills as well. You have the versioning and lifecycle, so you understand which version of the skills is actually the latest. And a very good example, when I'm using, for example, building a functionality, I can, the agents automatically capture that there is a latest version of the skill and pull it, right? So this versioning helps also to pull the right latest changes from the skills registry. Access control. Again, as I said, if you don't know who is accessing what, that is a huge gap. And finally, evaluation observability. And then all of this is actually played around a governance.

2:31

And this is where technology stops solving the problem, right? So you figure out all of this, all good. Now, who's going to govern this? And that is where it really depends how your organization is structured today. That is where you should have your architects, your engineering leads, infra leads, et cetera, and cyber leads actually sitting down, owning parts of those domains, and making sure that these skills, when they get updated, are actually according to the policies you want to adhere to within your organization and drive this change. And finally, when you get this right, what you will have, you will have at the organization level all of your teams pulling from one centralized place high-quality skills, executing them, and pulling them back to the centralized platform if they are improved.

2:37

Now, what I want to bring this, because this is a little bit of an unclear view. So what I created, I created a simulation, right? So think about this. This is your organization today, right? And what I have here, I have just a random team. So I have 15 teams created, 5 to 12 per team. You have skills per engineer's contribution. You have the average skills utilization, on average, how much time skills are being pulled per day, the duplication across the team as a ratio, and the skills quality and security ratio. Now, if I run this across six months, what's really happening, and think about it, this is already happening within your organization. If teams are creating and using the skills, right, but we don't have visibility. And skills, again, they are tightly coupled to your productivity uplifts. If, for example, the example that I shared earlier on the regulation, if we don't have a skill about the regulation, that is someone is vibe coding back and forth and trying to figure out exactly how to steer the agent to implement it properly, right? That is burning more tokens from one side cost-wise, but also the productivity is spending more time rather than getting in one shot the right answer. And the quality and security is similar. If you don't have

2:44

how much time skills are being pulled per day, the duplication across the team, as a ratio, and the skills quality and security ratio. Now, if I run this across six months, what's really happening? Think about it, this is already happening within your organization. If teams are creating and using the skills, right, but we don't have visibility. And skills, again, they are tightly coupled to your productivity uplifts. If, for example, the example that I shared earlier on the regulation, if we don't have a skill about the regulation, that is, someone is vibe coding back and

3:19

forth and trying to figure out exactly how to steer the agent to implement it properly, right? That is burning more tokens from one side, cost-wise, but also the productivity is spending more time rather than getting, in one shot, the right answer. And the quality and security is similar. If you don't have clear skills defined and maintained, you will have a low quality in your implementation because then it's up to the human to decide this, and different in the maturity from team to team, you can see the difference. And that is why, for example, if I look just randomly at this, this is, you can see the productivity of this team is a medium, right? If I look at this

3:54

one, it's a bit of, I don't know, it's low, medium, productivity, quality, and security, also a medium, but when it came to the cost, it's really high. Okay. Now let's actually say, okay, how this looks if I governed all of my skills in my organizations. What's going to happen is, of course, some of them were split, right? And this is reality. It's going to be perfect as we expect. But at least what you will see, you will see actually some common ground across all of your teams. The moment you govern, you publish one skill. The next engineer trying to build a new skill, the coding agent harness will identify the skill that is already available and pull it, right? So you

4:34

almost solve all of the issues that I covered about the governance. And one last point is, when it came to skills, it's just one component of your workflows, as I said, right? So that doesn't mean if you figure out skills, that says you're good. No, you need to apply the same approach and solution for your whole workflows. And you may think to apply this again, if you think about it, if you have a centralized platform that has all of your workflows, right? From one side, you're centralized in the workflows, which is also having the skills, but also if the next engineer came and went on, I don't know, provision infrastructure,

5:10

they can tap into a workflow and build that workflow with the required skills and run it and test it. Again, if it's something that needs to be improved in the workflow, you can easily push it back to the centralized platform for your organization to use. Now, before I wrap up, what I want to leave you with is this is just the start at the beginning. You see, this just we're talking about six to eight months. What's coming next, and I would invite you to already explore, is skills registry, right? You should have one, if not already. And the good news is all of the players that's been solving the IDP problem, like internal developer portal,

5:47

they already start centralizing this capability, right? So if you don't have it today, maybe in a couple of months, you will see it coming. But also, there's a lot of tools that's actually solving this specific problem. Second is skills evaluation. There's still a discussion on what is the right approach to evaluate skills. The easy thing that I found so far very valuable is actually tested, like you link static tests or evaluate your skills against the entropic best practices, right? If the skill is not invoked properly, if the skill is not structured properly, there's a high chance there's not going to be high quality. And finally, it's auto evolving. And again,

6:27

this is what everyone is the next hype right now. Like, yeah, I can create a closed loop that can evolve automatically my skills. So what, right? If you automatically start this machine, the impacts will be way more than it is today. Because what I shared earlier is going to be just maintaining auto evolving skills without that governance in place that actually put the guardrails for your organization. And at this point, I would leave you here. Thank you so much for your listening, and looking forward. If you have any question, I will be in the leadership lounge. Feel free to grab me. Thank you so much. Thank you.

7:14

an open standard that is adopted and starts a lot of agents harnesses starts adopting this new standard. Around February this year, we have seen most of the actually agents adopted this. Even though you don't see them, right? If you pay attention when the agent is thinking, you can see that he's pulling skills as he's going and doing the task. The other indication here that you need to pay attention to is actually the number of skills that are created, right? This is just like literally a snapshot that I did across some public gets to have reposts and public skills registries, right? There's way more than

7:47

this publicly and within your organizations, right? So the creation of skills and the demand is rising, but we need to understand why. In the latest skills bench, comparing the latest models against like, you know, running the same tasks against auto engineering and cybersecurity without skills, it did well, right? Because that's what it expects it. And it's going to continue to be improving day after day, right? But then when we applied skills that are more deterministic, the outcome was clearly higher than without it. Now, if you think about now like how like the anatomy of skills and how you actually design and implement skills in your organization, this is not

8:31

like a new problem that we are trying to solve here, right? We have solved this with the microservice kind of movements, right? So your microservices need to have all of these design principles, right? This is like a software kind of like, you know, problem that we use to solve. It's similar. So skills need to be reusable, right? Need to be modular, need to be discovered, like you can discover your skills. So if you are sitting in one team and you need the skills, you actually can automatically discover and capture the skills. It's portable that you can actually use skills across workflows, but also you can use skills

8:59

across harnesses. Again, everyone adopted the same standard. So if I'm having a skill on cloud code and I want to move it to cursor, it's going to just work. Specialized skills, that is where the value, you should not build like a one skill, like a monolith again. It should be specialized to define one tasks specifically. Composable skills need to be designed in a way that's actually can compose. So you don't have duplication across your skills when you're trying to run them. That conflicts each other. Consistence, that is actually one of the key items of skills is actually consistency and deterministic.

9:30

And finally, cost-efficient. And for cost, I can go for another hour talk, but it's basically the skills comes to solve a key problem around the context window, right? It's actually put in with the disclosure, progressive disclosure pattern, the right skills, the right amount of skills in the right time to solve the right problem. And that's reduced the token usage. This defines a new unit, right? That makes your know-how in your organization executable, portable, and cheap. On the right-hand side, you just have a very simple example of data retention policy, right? When it

10:02

came to regulation, you need to understand and make sure to instruct your agents, why you're manipulating, for example, your customer data, you should make sure that this data manipulates according to the regulation, right? And that brings me to the next example. On the left-hand side, you can see that there's like a...think about a catalog of skills. And in the right-hand side is your harness, and the output is on the right-hand side, right? So the composable skills at the regulation level, you have the skill that I just showed earlier, which is the retention policy, but you will need disclosure

10:37

standards. You will need the GDPR rules that need to be respected. You will need the filling templates, right? So all of this, kind of think about it, is like defining how any data, any feature that is built across your web, mobile, like, you know, different applications across your organizations is really respecting these rules. And this gets pulled automatically by on the runtime by the regulatory disclosure review workflow. And the outcome is expected. It's deterministic. You have an audit, uh, an audit report that you can actually store. You have, um, specific, uh, identification of if

11:13

there's anything to improve, and that there's kind of like loopback to improve your, your codebase.

11:20

However, if we don't govern skills, we will start creating a new class of technical lips, right? First of all, you will find out that you are having a lot of duplication in your organization. So if teams are not collaborating and everyone is, think about it, using the same technology stack, the same infrastructure, you're for sure building the same skills over and over again without sharing them. Quality. If you don't test and make sure that you're maintaining and you're validating your skills, not against your task, but also against the latest models that comes, right? Then the quality starts degrading over time. You should be able to discover your skills. In reality,

11:59

without a governance, you cannot really discover it. Think about it as the, the backstage, the IDP, right? It's, it's come to solve a problem where, okay, I need to understand who owns this microservice back in the days, right? I don't need to talk to anyone. I just need to tap into the service catalog and immediately find who actually owns it. That is brings the ownership part. If you don't have an owner, then no one will be able to maintain, scale those skills. Composibility is not something that comes by default. You need to have a governance way to, to align what to build and how to design it.

12:29

It's think about the domain driven approach that we have been taking also for, for many years, right? It's similar to how you shape your skills catalog. Security. Again, some of, all of us, like we experiment with the public skills, right? But when you think about it, some skills may have some prompt injection and skills actually does have scripts because that's the deterministic part of it because it's going to run a specific script for a specific task. So if you don't have, again, a pipeline that checks your security, you may be pulling something that is insecure. And permissions. Not every skills is actually something that's anyone in the organization

13:03

should access. Some skills may have some business logic that is very sensitive, right? So the access control is, is also crucial at this stage. Now, how to bring this to your organization? First of all, you need to allow at the individual level to create, test, improve, and use those skills. Again, you shouldn't be around them. You should be structured way. These are different tools out there. You just need to decide which tool actually you want to agree on. And you use that mechanism at the individual level. The moment you create a skills, you need to be sharing this with your team,

13:34

right? Your team starts collaborating to improve the skills. And think about it, you build the same technology stack, building the same products. So it's going to evolve very quickly. But then you move on to a very critical point, which is the centralized platform. That is where all of what I've been covering so far come to play. You need a centralized platform that has a catalog with metadata in it that actually can discover skills and could be searchable. You can have an MCP that actually plugs to this catalog. Search for the skill and a CLI to pull the skills back to either your IDE, if you're locally,

14:07

or to your sandbox in your factory. Then you have the dependencies. You need to understand the dependencies between the skills as well. You have the versioning and lifecycle. So you understand which version of the skills is actually the latest. And a very good example, when I'm using, for example, building a functionality, I can... The agents automatically capture that there is a latest version of the skill and pull it, right? So this versioning helps also to pull the right latest changes from the skills registry. Access control. Again, as I said, if you don't know who is accessing what, that is a huge

14:42

gap. And finally, evaluation observability. And then all of this is actually played around a governance. And this is where technology stops solving the problem, right? So you figure out all of this, all good. Now, who's going to govern this? And that is where the... You know, it really depends how your organization is structured today. That is where you should have your architects, your engineering leads, infra leads, et cetera, and cyber leads actually sitting down, owning parts of those domains and making sure that this skills when it gets updated, is actually according to the policies you want to adhere within

15:13

your organization and drive this change. And finally, when you get this right, what you will have, you will have at the organization level, all of your teams pulling from one centralized place, high-quality skills, executing them and pulling them back to the centralized platform if it is improved. Now, what I want to bring this, because this is a little bit of an in-clear view. So what I created, I created a simulation, right? So think about this. This is your organization today, right? And what I have here, I have just a random team. So I have 15 teams created, 5 to 12, like, per team. You have

15:51

skills per engineer's contribution. You have the average skills utilization, kind of like on average, like how much time skills are being pulled per day, the duplication across the team, kind of as a ratio, and the skills quality and security ratio. Now, if I run this across six months, what's really happening and think about it, this is already happening within your organization. If teams are creating and using the skills, right, but we don't have visibility. And skills, again, they are tightly coupled to your productivity uplifts. If, for example, the example that I shared earlier on the

16:25

regulation, if we don't have a skill about the regulation, that is someone is vibe coding back and forth and trying to figure out exactly how to steer the agent to implement it properly, right? That is burning more tokens from one side cost-wise, but also the productivity is spending more time rather than getting in one shot the right answer. And the quality and security is similar. If you don't have clear, you know, skills defined and maintained, you will have a low quality in your implementation because then it's up to the human to decide this and different in the maturity from team to team,

16:55

you can see the difference. And that is why, for example, if I look just randomly at this, this is like you can see the productivity of this team is kind of a medium, right? If I look at this one, it's a bit of like, I don't know, it's low, medium, productivity, quality, and security, also a medium, but when it came to the cost, it's really high. Okay. Now let's actually say, okay, how this looks like if I governed all of my skills in my organizations. What's going to happen is, of course, some of them were split, right? And this is reality. It's going to be perfect as we expect. But at least what you will see, you will see actually some common ground across all of your

17:35

teams. The moment you govern, you publish one skill. The next engineer trying to build a new skill, the coding agent harness will identify the skill that is already available and pull it, right? So you almost solve all of the issues that I covered about the governance.

17:53

And one last point is when it came to skills, it's just one component of your workflows, as I said, right? So that doesn't mean if you figure out skills, that says you're good. No, you need to apply the same kind of like approach and solution for your whole workflows. And you may think to apply this again, if you think about it, like if you have a centralized platform that has all of your workflows, right? From one side, you're centralized in the workflows, which is also having the skills, but also if the next engineer came and went on, I don't know, like provision infrastructure,

18:24

they can tap into a workflow and build that workflow with the required skills and run it and test it. Again, if it's something that needs to be improved in the workflow, you can easily push it back to the centralized platform for your organization to use.

18:40

Now, before I wrap up, what I want to leave you with is this is just the start at the beginning. You see, like, this just we're talking about six to eight months. What's coming next and I would invite you to already explore is skills registry, right? You should have one, if not already. And the good news is all of the players that's been solving the IDP problem, like internal developer portal, they already start centralizing this capability, right? So if you don't have it today, maybe in a couple of months, you will see it coming. But also, there's a lot of tools that's actually solving this

19:10

specific problem. Second is skills evaluation. There's still kind of like a discussion on what is the right approach to, you know, to evaluate skills. The easy thing that I found so far very valuable is actually tested, like you link static tests or evaluate your skills against the entropic best practices, right? If the skill is not invoked properly, if the skill is not structured properly, there's a high chance there's not going to be high quality. And finally, it's auto evolving. And again, this is what everyone kind of like is the next hype right now. Like, yeah, I can create like a closed

19:44

loop that can evolve automatically my skills. So what, right? If you automatically start this machine, the impacts will be way more than it is today. Because what I shared earlier is going to be just maintaining auto evolving skills without that governance in place that actually put the guardrails for your organization. And at this point, I would leave you here. Thank you so much for your listening and looking forward. If you have any question, I will be in the leadership lounge. Feel free to grab me. Thank you so much. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note