The Dark Arts of Skill Engineering — Paul Bakaus, Renaissance Geek (Impeccable)
Description
Paul Bakaus once turned the entire web orange. He wrote jQuery UI, shipped an orange default theme, assumed people would change it, and watched them not. So when he points out that the purple gradients everyone learned to mock came from a CSS framework's default sample page, and that AI generated design now has its own house shade of beige, he speaks as the cause rather than the critic. That history is what makes his central claim land. A ban does not produce originality, it relocates the model one cluster over: forbid one overused font and it reaches for the next nearest thing in latent space. Slop is a moving target. The median is the model's gravity, and a few hundred lines of carefully written prose cannot pull against it. So he stopped treating a skill as a packaged prompt and started treating it as an extension of the harness, which is what this workshop is really about. Nine techniques, each aimed at something prose cannot do. Two sub agents kept blind to each other, one playing design director and one running a deterministic linter, because a single thread grading its own work always says it did well. Naming your top three fonts and then throwing all three away to shave off the safe prediction. A script whose only job is to hand the model a random seed it could not have guessed. Instructions emitted on standard output rather than buried in prose, which models follow far more reliably, at the cost of prompt caching. Hooks that block a write instead of complaining afterward. And the discipline underneath all of it: if a gate can be skipped, it will be. Speaker info: - https://x.com/pbakaus - https://linkedin.com/in/paulbakaus - https://www.paulbakaus.com/ Timestamps: 0:00 - Impeccable, and why it started 4:57 - Why banning a font does not work 6:25 - Prompting is the floor, harness engineering is the ceiling 7:55 - One: make it argue 10:11 - Two blind sub agents, one synthesis 16:59 - Two: forcing divergence 20:38 - Three: routing like a mixture of experts
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: Production-grade agent skills should be engineered as cross-harness control-plane extensions—combining routing, deterministic scripts, adversarial agents, persistent state, hooks, UI feedback loops, and evals—rather than packaged as long prompts.
- Why it matters: This is a concrete field report on turning fragile prompt-driven behavior into more reliable agent workflows across Claude Code, Codex, Cursor, Gemini, and other uneven execution environments.
- Best use: Use it as an architecture playbook for designing, testing, and distributing durable OpenClaw-style skills or agent workflows, especially where outcomes must survive model and harness variation.
Executive Summary
Paul Bakaus explains how Impeccable evolved from a prompt for improving AI-generated frontend design into a much more elaborate agent-skill system. His central argument is that prompt prose alone cannot reliably overcome a model's tendency toward median, familiar outputs or guarantee that it follows complex workflows. A useful skill therefore needs to treat the coding environment itself—the harness, tools, state, scripts, browser, hooks, and permissions—as the programmable substrate.
The talk's practical techniques generalize far beyond design. Bakaus separates independent model roles to reduce self-review anchoring; introduces random or externally generated seeds to force creative divergence; routes requests to smaller, specialized instructions instead of maintaining one bloated skill; and persists critique history and user preferences so work compounds across sessions. He repeatedly argues that models will skip difficult steps whenever they can, so important process requirements need executable enforcement rather than hopeful wording.
The strongest implementation material concerns deterministic scripts and control loops. Impeccable injects structured script output into the agent context, uses edit hooks to catch violations automatically, and connects an in-app browser to the chat agent through a polling/event mechanism so users can visually select, annotate, generate, compare, and accept page variants. These mechanisms make the model react to evidence and user interaction rather than relying entirely on conversational instructions.
Bakaus also offers a sober operational warning: portability is expensive. Harnesses differ in sub-agent permissions, user-question tools, background-task completion behavior, hook syntax, file conventions, and install/update methods; models have distinct behavioral quirks. His response is harness-specific compilation, model-specific instruction fragments, and a substantial eval system spanning end-to-end tests, simulated interactive users, visual judges, competitor comparisons, and line-level ablation tests. He does not believe model-based taste evaluation is solved; he uses it as a weak first pass and retains human judgment.
Key Takeaways
- Claim: A skill should be designed as a harness extension, not as a static prompt package. | Evidence: Impeccable began as a simple "normalize" prompt to bring Claude-generated UI back to a design system, but became a system of specialized skill files, scripts, browser tooling, persistent files, hooks, and harness-specific builds. Bakaus summarizes the progression as "prompting is the starter level" and "harness engineering is where you should end up." | Implication: For agent systems that must behave reliably, define the full operating loop—state, validations, tool behaviors, UI surfaces, and failure handling—rather than optimizing the system prompt in isolation. | Caveat: Not every use case warrants the full stack; Bakaus presents several techniques as deliberately exotic and situational rather than mandatory.
- Claim: Independent, blind sub-agents produce more balanced reviews than asking a model to evaluate its own work. | Evidence: Impeccable's critique flow runs one sub-agent as a design director assessing hierarchy, visual slop, and heuristics, while another runs deterministic lint checks and collects browser evidence; neither sees the other's output. The main thread synthesizes both. This avoids a model treating many detected polish issues as proof an otherwise strong page is bad, or treating zero detector findings as proof an empty or poor page is good. | Implication: For review, planning, ranking, security, and audit workflows, separate subjective judgment from deterministic evidence and prevent either evaluator from anchoring on the other's conclusions. | Caveat: Codex requires explicit user permission for sub-agent use, and models may simply avoid spawning sub-agents unless the workflow tells them to stop, request permission, and disclose degraded behavior if they cannot comply.
- Claim: Banning undesirable outputs does not create originality; it only shifts the model to the next nearby default, so systems need explicit divergence mechanisms. | Evidence: Bakaus notes that telling a model not to use Inter, purple gradients, or other familiar aesthetics merely relocates it to the next likely latent-space cluster. His anti-attractor approaches are: repeatedly generate and discard the top three likely choices; generate many candidates and have a fresh sub-agent rank them; or inject an unpredictable seed from a script. Impeccable's color.js selects from more than 100 hand-curated primary-color starting points, while his Radiant Shaders work used unusual prompts such as imagining Rihanna or Beyoncé as a shader. | Implication: When differentiation matters, feed agents controlled novelty—randomized, curated, or user-derived seeds—and use independent ranking rather than relying on negative prompt constraints. | Caveat: The repeated-discard technique eventually reconverges and is therefore a limited form of divergence.
- Claim: Large general-purpose skills become instructionally blurry; route each request to specialized capabilities and context registers. | Evidence: Impeccable splits capabilities such as critique and polish into separate underlying files rather than one giant SKILL.md. It also routes briefs into distinct "brand" versus "product" design registers because rules such as avoiding system fonts may be appropriate for expressive landing pages but harmful for native-feeling product UI. | Implication: Build agent workflows as mixtures of experts: classify intent early, load only the relevant policy and tools, and avoid forcing a single context to reconcile conflicting rules.
- Claim: Skills can compound across sessions if they persist artifacts, preferences, and progress in accessible files. | Evidence: Impeccable saves critiques in a skill or repository-level .impeccable folder, allowing a later "polish" session to use prior findings and user disagreements. Bakaus describes using the same approach for refactors: an agent tracks work across sessions, handles one TSX file and its dependencies at a time, and accumulates enough context to complete a codebase-wide change. | Implication: Treat durable artifacts as agent memory: maintain explicit task state, decisions, exceptions, and prior evaluations so long-running work resumes coherently rather than restarting from zero. | Caveat: Harness support varies: Claude can expose a skill-directory environment variable, while other environments may require repository-local workarounds.
- Claim: Executable guardrails and structured script outputs are more dependable than buried prose rules, but dynamic outputs trade away some prompt-cache efficiency. | Evidence: Impeccable runs context.mjs on invocation to collect product/design context, identify missing files, return structured JSON with exact next steps, and offer updates with user permission. Bakaus found model adherence significantly stronger when instructions arrived through script stdout than when the same rule appeared in the main prose. It also installs design-lint hooks that run on every edit; for weaker models, pre-tool hooks block a problematic write rather than relying on a post-write correction. | Implication: Move critical workflow state transitions, checks, and remediation paths into structured executable outputs and enforcement hooks; reserve static prompt prose for principles and interpretation. | Caveat: Dynamic script output prevents that dynamic portion of the context from being prompt-cached, so it is less suitable when repeated cached execution is the primary optimization. Hooks also require granular ignore mechanisms because false positives quickly become disruptive.
- Claim: A distributed skill must be evaluated and compiled per harness and model, because tool semantics and behavioral biases differ materially. | Evidence: Bakaus cites differences in sub-agent permissions, availability of user-question tools, background-task wakeups, watcher throttling, hook behavior, directory conventions, and marketplace/update mechanisms across Claude Code, Codex, Cursor, GitHub Copilot, and Gemini. He reports model-specific tendencies such as Gemini overusing image hover animations and Codex favoring extreme rounding, hairline borders, and poor letter spacing. Impeccable uses harness-specific substitutions and model-specific blocks, and Bakaus tests releases across approximately 20 niches, several models, five to ten runs per model, competitors, end-to-end Playwright tests, LLM-driven tests, and line-level ablation tests. | Implication: Do not assume an agent skill that works in one IDE/model pair is portable. Establish a compatibility matrix, test real tool behavior, and create per-target builds or explicit degradation paths before distributing broadly. | Caveat: This level of testing is expensive and Bakaus does not recommend full ablation testing for every individual skill author. He also considers model-based taste judges only marginally useful and retains human review.
Detailed Brief
Interactive browser loop: turning a skill into a visual control surface
- Claims: A chat-only interface is poorly suited to pixel-level visual iteration, but a skill can use the host harness's browser capabilities to create a direct manipulation loop.; The design pattern is not inherently about design: connect a user-facing inspection or annotation surface to the agent through a small event bridge, then make the agent react to structured events.
- Evidence: Impeccable's "live mode" starts a lightweight poller/server and injects a snippet into the local development page rather than using MCP.; The user can select an element in the browser, choose a subcommand and number of variants, annotate or dictate changes, and send that signal back to the agent.; The event causes the poller to stop; its stdout then wakes or informs the main agent thread, which knows from the skill instructions how to create marked-up alternatives. The UI immediately displays variants for the user to accept or reject.; Bakaus also uses HTML as a richer visualization channel for artifacts such as designMD, aligning with the broader idea that visual interfaces can communicate agent context more effectively than Markdown alone.
- Caveats: The implementation is a workaround built from available harness primitives; Bakaus calls it a "Jurassic Park experiment" and believes a first-party integration would be materially better.; Background tasks do not produce equivalent wake-up behavior across harnesses. In Codex and some other environments, the implementation may need a foreground task that blocks the chat thread to remain reliable.
- Implications: For operational agents, identify the interface where users naturally make judgments—browser, dashboard, document, terminal, or visual canvas—and connect it to the agent as a structured steering channel.; Harness capability discovery should be part of product design: browser screenshots, development servers, event streams, and task lifecycle events can be more valuable than another prompt refinement.
Evals, human judgment, and the emerging skill-distribution gap
- Claims: Reliable skills require an evaluation discipline closer to software release engineering than prompt tweaking.; Functional visual criteria can be partially automated, but aesthetic taste is not a stable objective that current models can robustly evaluate.; The skill ecosystem lacks a mature cross-harness testing and distribution standard, creating a "works on my machine" risk for shared skills.
- Evidence: Bakaus's private evaluation harness recreates Claude Code, Codex, and Gemini conditions and tools, including browser screenshots and interactive back-and-forth through an LLM acting as the user.; His mixture-of-expert judge evaluates results across verticals such as an Italian restaurant; each release is run multiple times across relevant models and compared with competing frontend-design skills.; Every Impeccable instruction line has a unique XML identifier for ablation testing: remove a line, rerun evals, restore it, and use deterministic checks to determine whether the rule changed behavior.; He found a clear failure mode in taste judging: Gemini tends to rate a fuller first viewport more highly, rewarding excessive density. He sometimes inverts a judge's high score because it signals an obviously poor design.; Native marketplaces and installers are provider-specific and unreliable in his experience; he built a custom CLI because common installers do not properly support harness-specific directory layouts and compiled variants.
- Caveats: The transcript does not establish that Bakaus's private eval framework or design scoring methodology is independently validated.; His view that taste is fundamentally human is a reason to retain human review, not a claim that automated evaluation is useless; he finds it useful for first-pass screening and functional checks.
- Implications: Adopt a layered evaluation stack: deterministic checks for measurable defects, multi-run behavioral evals for workflow adherence, and expert human review for qualitative outcomes.; There is a tooling opportunity in a standardized cross-harness skill compiler, installer, test fixture library, and compatibility certification layer.
Notable Concepts & Terms
- Harness engineering: The core reframing: a skill should extend the agent runtime through tools, state, hooks, scripts, and interfaces, rather than merely supplying a prompt.
- Median / model gravity: The model's tendency to select familiar, high-probability outputs; this explains why negative prompting often just changes the flavor of generic output.
- Anti-attractor: Bakaus's term for a mechanism that injects an unexpected seed or systematically removes likely choices to push generation away from default clusters.
- Blind sub-agents: Independent agents that assess a problem without seeing one another's work, reducing anchoring and enabling later synthesis of different evidence types.
- Mixture of expert skills: Intent-based routing to small, specialized instruction sets and rules, analogous to mixture-of-experts model routing.
- Context.mjs / script stdout as instruction channel: A script-driven pattern in which dynamic repository state and explicit next actions are returned in structured output that models follow more reliably than deeply embedded prose.
- Hooks that fight back: Pre- or post-tool-use enforcement hooks that automatically block or correct agent edits, replacing manual commands that users may forget to invoke.
- Ablation testing: Removing an individual instruction or rule, rerunning evaluations, and measuring whether it changes behavior; used to distinguish useful rules from prompt cargo cult.
Operator Notes / Why Ken Should Care
- For any reusable agent skill, create a target matrix covering model, harness, permissions, sub-agent behavior, user-input behavior, background-task semantics, hook lifecycle, install location, and update mechanism before claiming compatibility.
- Convert critical requirements into explicit machine-checkable gates, structured script outputs, or pre-execution blockers; add a visible degraded-mode path when a required capability is unavailable.
- For review agents, implement at least two independent lanes: one heuristic or domain-expert critique and one deterministic evidence collector, followed by a separate synthesis step.
- Persist agent artifacts in repository-local state—decisions, accepted exceptions, prior critiques, migration progress, and user preferences—and explicitly load relevant history at each run.
- Build an eval suite that runs repeated trials across representative tasks and models; reserve expensive ablations for high-leverage rules, and keep a human approval stage for taste-sensitive output.
- Avoid investing heavily in MCP-mediated skill delivery until its context-cost and packaging tradeoffs are clear; Bakaus specifically reports concern about context pollution and has not validated server-hosted skills deeply.
Source/Metadata
- Title: The Dark Arts of Skill Engineering — Paul Bakaus, Renaissance Geek (Impeccable)
- Transcript words: 12507
- Duration seconds: 3893
- Timestamp note: No usable timestamps or chapter markers were present in the transcript.
Transcript
Hello everybody! How's it going? Okay, I think we still have some people trickling in, but I'm super excited to be here. Okay, first off, let me reload these slides because my cloud code was still building something on it. Okay, so the contrast is a little bit low, so please bear with me. I'm going to try to cover what you cannot read as much as possible. I also have heard that the Wi-Fi is not the strongest, so while it is a workshop, hopefully you just take away a lot of the lessons and then can apply it whenever you want. But I do have a sample repo if you want to follow along. Okay, first of all, hi, my name is Paul. I'm really glad you found your way into this room. I'm the author of Impeccable. Who here has used Impeccable by any chance? Can I see some hands? Okay, a few people, nice. So for those who have not used Impeccable, Impeccable is a skill that I've built for myself mainly. I've built this large enterprise app over the last year and it has lots of different views, states, whatever, and I wanted to design really quickly with my agents, with Codex, with Claude, but I noticed that even though it gets me quickly to something that I can look at, normalizing it back to the design system was really challenging. So the first skill that I've built for myself was called normalize, and it kind of brought whatever Claude designed back to the design system. That's how I started. I also used Anthropic's front-end design skill, like maybe many of you when I first got going. And from there on, it kind of expanded into more and more design skills that allowed me to turn Claude code and then Codex and other harnesses more into a design harness. At some point, I decided maybe other people might find this useful as well, so I open sourced it as an open source skill, released it, and it turns out a lot of you liked it. So if you haven't checked it out yet, give it a go. It's on Impeccable.style. But today won't be a talk about Impeccable, per se. It'll be about what I learned from making these skills because it kind of escalated. It started with a simple prompt and then went all the way to what it is today. A lot of people have looked at the code of Impeccable and they see a whole bunch of scripts in the script folder and they're what is all this stuff? And so I wanted to share some of my knowledge that I've gained with you. So let's get into it. Let's talk about the dark arts of skill engineering. Okay, so first of all, you've all seen this kind of design. This is actually a real design built with the front-end design skill and Claude code. And you probably all have seen a design like this. This is for a fake kids reader, iPad reader app. You have italic serif. You have some capitalized hero. And you have eyebrow text, a kicker, whatever you want to call it. The top of it, a weird label. You have beige. I call it Claude beige. Claude beige backgrounds. And now it's not necessarily a bad design, right? But I think you all can point this out and say, well, this is clearly AI generated. It's clearly slop. And it turns out slop is a moving target. You've all seen, you might have been thinking of slop as purple gradients, but we've moved on since that into Claude beige. Okay, so this is where I started. I started with a system prompt and a prayer. So I started using the front-end design skill like many of you. And it's 55 lines of named bands. No scripts, no routing, pure prose, right? And you just hope for the best. That sometimes worked and most of the time it didn't. So for example, I'm just going to read some of this. So if, maybe it's really readable. I don't know. But for example, in the front-end design skill, you have sentences like never use generic AI aesthetics, overused fonts like Inter, Roboto, Arial, System fonts, or clichéd color schemes, particularly purple gradients and white backgrounds. Never converge on common choices like Space Grotesque, for example. Now, there are two problems with this approach. The first one, it over applies, and then a band just relocates the model to the next cluster. And I'll show you why this is a problem a little bit further down the road. But really, you tell it not to use Inter. It just uses the next best font it finds in its latent space. And so it doesn't actually make it more creative. It just, again, this is why I said slop is a moving target. It kind of picks the next best thing. I have learned my lesson here the hard way, because I don't know if you noticed, but the reason why we got purple gradients in the first place is because of Tailwind. Tailwind's default sample pages theme, whatever, it was purple. Well, it turns out I've turned the web orange before, many years before that. I created a framework called jQuery UI, and the first default theme of jQuery UI was orange. So overnight, I call it the web orange. I thought people would modify the theme, but no, they didn't. So I learned my lesson. Okay, the median is the model's gravity. Even 250 lines of artisanal, crafted, beautiful skill prose cannot change this. It's just not enough. It doesn't help enough, right? It's nowhere near enough. And I learned this the hard way, and hopefully you don't have to. My overall thesis for this talk is that prompting is the starter level, but harness engineering is where you should end up. You should reframe when you're building skills. You should think about, okay, skills the same way as MCP is an extension to the coding harness, or whatever harness you're in. It's not just a prompt that you package. It's something more than that, or you should at least conceptually think about it more than that. It is extending the harness of whoever is using that thing. And it also has more capabilities than just prompting. And when I thought about that way, it sort of clicked for me. Prompting is a spell. Harnessing is magic. And we'll talk about nine different dark arts today that I learned in the process of building Impeccable. We'll make sub-agents argue with each other. We'll talk about how to force divergence as opposed to convergence. Routing like a model. And basically, you've seen this probably. Most modern models are mixture of expert architectures. And Impeccable is built like a mixture of expert skill. We give the memory. We'll create scripts that talk back. And I promise this will make sense. Hooks that fight back. Livewire the browser and use more of the harness. Compile to every harness and design for the weakest model. Let's get into it. Number one: Make it argue. So if you're building something like a critique skill or code review skill, here's one huge issue. If you are, I mean you probably noticed this, when you're working with Claude Code or Codex, that doesn't matter. But if you ask Codex or Claude Code to review its own work, it will usually rate it as very high. I mean, I've built this. I've done a great job. It's rating your own homework. It doesn't make any sense. It anchors on what it already created. Now, that's not great. What can you do in order to solve this? Well, you can make a model argue with another model. Adversarial prompting is also called it. So you have two sub-agents and they never see each other's work. And here's why this matters. So in Impeccable there's a critique command that actually critiques your design. And you can point it to your landing page. You can point it to anything. And there are two particular failure scenarios. The first one, and this is almost impossible to read, so I'll explain it. The first one is a strong page. So it's a really good looking page. But there's a whole bunch of maybe deterministic errors. And Impeccable actually has a deterministic engine, a design linter, that can detect things like bad contrast. It can detect things like too many fonts. Maybe things that are too close to the edge of an element. So it detects some of what I would say are polish issues. But you could have this really beautiful website and then the detector runs. And it's doing that as part of the same skill and the same model thread. And then the model just sees the detector output and says, well, I guess there's 500 issues. Therefore, this design must be bad. Now that's one. The other one is the opposite. The other one is it's actually a really terrible page or maybe an empty page, but there are no detected issues by the deterministic detector. So the model is like, hmm, we didn't find any issues. So this must be great design. So both of those are not amazing. And Impeccable actually has a deterministic engine, a design linter, that can detect things like bad contrast. It can detect things like too many fonts, maybe things that are too close to the edge of an element. So it detects some polish issues. But you could have this really beautiful website and then the detector runs. And it's doing that as part of the same skill and the same model thread. And then the model just sees the detector output and says, well, I guess there's 500 issues. Therefore, this design must be bad. Now that's one. The other one is the opposite. The other one is, it's actually a really terrible page or an empty page, but there are no detected issues by the deterministic detector. So the model is, hmm, we didn't find any issues. So this must be great design. So both of those are not amazing. What you want is, and this is what Impeccable's critique skill does, it combines two things. It spawns two sub-agents and they are blind to each other. And that's how you get to a balanced critique. So the first sub-agent acts like a design director. So it's an LLM that acts like a design director. And so it looks for hierarchy, it looks for slop, it looks for heuristics. And so it does a critique the way a human would with the browser tools that are available to it. The second sub-agent runs the deterministic detector and also collects browser evidence. And then once both of those results come in, the main thread synthesizes both into one critique. And that produces a much more balanced result. And before I go on, I realized I actually have not shown you where the sample repo of this lives. So let me bring this up real quick. If you want to clone this and you have a decent enough internet connection, go ahead if you like. So this is PBAKAUS slash impeccable minus talks. The talk lives here, but also in the dark arts folder, there is a starter folder and a demos folder. Demos has a pretty average median page that you can manipulate. And then in the starter kit, you have enough to build a mini impeccable if you want to follow along or try it out yourself. And so as part of this, there is a tutorial here too, you'll follow along the checkpoints, the dark arts and build something yourself. I would suggest if you like to multitask, great. You can apply this, by the way, to anything. You can do a code review thing. You can do design review, but I wanted to have something for you to play with. Back to the deck. So two blind opinions beat one confident guess. You can use this, again, I already said code review, design review, but also security audits once is a good example. Or creating a really good plan, RFC critique, where you have multiple LLM judges argue with each other before it gets good. Or ranking outputs is a good example. Now, here's a problem though. Codex, why? Why you don't let me do this? It's bad. It turns out Codex never created these sub-agents when I first tried this. And I banged my head against the wall, I'm like, why is this? It turns out Codex has a different permission model than Claude Code and other harnesses. In Codex, you have to explicitly, as a user, request the use of sub-agents for anything in the harness to use sub-agents. So if you're distributing a skill, you're out of luck. The only way to make this work, as far as I know today, is to actually tell the model, okay, if you have sub-agents capabilities, but you do not have permission, please stop right here and ask the user. And so that's pretty much the only way you can get Codex to comply. So in Impeccable, if you see something that makes you go, huh, it's probably because of that. Through lots and lots of issues that people filed and a lot of testing on my end, a lot of this obscure knowledge got into the skills so that it works truly across harnesses. For instance, here you see a pseudo code of how this would work. And then also, this is another really important thing. Very often, if Codex realizes it can get away with something, it will do it. So if there is no punishment for not spawning sub-agents, it will simply not spawn them. It's like, well, this is the easier route. I will take this easier route. So what you have to say is, actually, if you cannot use sub-agents, you must say that you are giving the user a degraded experience. And Codex hates that. So use that to your advantage. So you can watch them argue. Now, I did not pre-record an actual example here, because I'm like, let's do it live. So we're going to go into Cursor. And I'm going to do critique. And let's hope for the best. I don't know if Composer sponsors sub-agents well enough. But let's see. Composer, by the way, if you haven't used it, is a really fast, well-balanced model. So it's neat for work that you want to show on stage in particular. Okay, so now it's doing something here. Okay, this repository, I think, has an old version of the back end. It doesn't sponsor agents. I see this is unfortunate. Well, maybe it does. I'm not sure if it did or not. But I at least want to show you what the type of critique looks like. All right, now it's asking a bunch of questions of what I actually want to create. I'm going to skip this. And now I get a design critique on what's working, what the priority issues are, persona red flags. Now, the actual thing that I wanted to show you, unfortunately, couldn't be seen in this particular set, but we can come back to it. If you run this in Claude Code or Codex on the most recent version, you should very clearly see in Claude Code, it's very easy to see the sub-agents running and doing its work. So it will spin up two sub-agents and you see it at the bottom of the Claude Code thread doing its thing. Okay, level number two. A ban just moves the problem. We talked about this already. You ban into the model graphs, the graph space grotesque. How do you solve that? How do you force divergence? Well, a ban only moves the model around inside its own cluster. And what I've built for Impeccable and for a bunch of other skills that I've released is what I call an anti-attractor. And the anti-attractor works by creating a random seed of sorts and that can come from user input or it can come from a script that it can run that produces something that is completely unexpected to the model because that's what you want. And so in this case, for instance, it will be font selection. And instead of selecting this safe next step prediction font, it went into a completely different space and it's through a different seed. There are three techniques that are easy to hard and work differently. The first one is something you can do right now. It's the most simple one and it's to shave the safe picks. So basically, tell the model, okay, name your top three fonts. And then the model is like, okay, I got the top three fonts. And then you're like, now throw them away. And basically you shaved off the next token that is predicted. And you do that three times and then now you get to a different space, a further away in the latent space, right? Now that's doable, but at some point you still get convergence. So this is a limited technique. The second technique is to generate a lot of different things and then have a sub-agent rank. I've done this for a shader library that I created called Radiant Shaders. And the goal here was to create around a hundred different shaders. And the problem is every time I would say, you know, create a new shader or create ideas for 10 new shaders, I would get the same repeating ideas. I solved this in two different ways. The first one is I created something unexpected, a random seed, a creative seed. In this case, I used celebrities. I said, what would Rihanna look like as a shader? Or what would Beyoncé look like as a shader? And then the model was like, hmm, let me think about that. So that's the first thing. And then I said, well, generate a hundred of these ideas and then spawn a sub-agent that ranks all of those ideas. And that's important. It has to be a sub-agent because the sub-agent doesn't know anything from the prior context of the session and can then completely change the order. The third one is to create a random seed from a script. So in Impeccable, for example, when you first start a project, # Transcript create a new shader or create ideas for 10 new shaders, I would get the same repeating ideas. I solved this in two different ways. The first one is I created something unexpected, a random seed, a creative seed. In this case, I used celebrities. I said, well, what would Rihanna look like as a shader? Or what would Beyonce look like as a shader? And then the model was like, hmm, let me think about that. So, that's the first thing. And then I said, well, generate a hundred of these ideas and then spawn a sub-agent that ranks all of those ideas. And that's important. It has to be a sub-agent because the sub-agent doesn't know anything from the prior context of the session and can then completely change the order. The third one is to create a random seed from a script. So, in Impeccable, for example, when you first start a project, it calls a script called color.js. And color.js has over a hundred hand-selected, they're not complete color palettes, but they are primary colors in the color of a starting point of a palette. And it reads that and then the model uses that as a creative spark to build a palette around it for you. You can still say, I don't like what it proposed. But it turns it into a different direction. So, those are all ways to create divergence. And when you design something with Impeccable, the same brief, depending on the user's input and color script that runs, et cetera, can produce vastly different results because of that. Because I didn't want to have the whole internet look like everything else. Number three. Here's the problem. If you cram everything into one skill, it kind of blurs them. The instruction following becomes not very good enough. So, if you're building some general purpose skill and you expand it and expand it and expand it, at some point, it becomes really, really blurry to the model. Here's a concrete example of this. The Anthropic Frontend Design Skill, the former version of it, they just shipped a new version three weeks ago. But the former version had a line that I read earlier that says avoid system fonts. That's okay for landing page design. But for product UI, oftentimes you want it to feel as native as possible. So, system fonts are actually the thing that you want. So, how do you solve this? You can say, you can have this giant if-else block in a skill. Say like, well, if the user wants a landing page, do this. If the user wants a product, do this. But that becomes really convoluted, wastes a lot of tokens, and honestly, it doesn't work very well. So, Impeccable started as a lot of different subskills and now has this mixture of experts model that routes internally. Both in terms of capabilities, so you can call Impeccable critique or Impeccable polish, and you get a different MD file loaded behind the scenes for a particular job. So, it's not just one giant skill MD. But also, and this is something not a lot of people know, behind the scenes, Impeccable decides based on your brief and what you input. Whether you're trying to design something brandy, so like a landing page or something that wants to attract attention. Or whether it's the actual product that you're designing. So, it switches registers and then loads completely different rules for those two registers. Because product design and brand design are very, very different. So, that's also something that I would recommend you doing if you're building a larger skill. This works for big multi-tool skills, works for context on demand type of skills, per audience behavior, agent toolkits, that kind of thing. Number four. Everyone starts from zero. Skills by default don't have long-term memory. They don't really compound over time. But you can make it so. So, you have a skill folder and you can save things in that skill folder. In fact, in Cloud, you even have an environment variable that resolves to the actual directory that you can save things in, which is nice. No other harness supports this right now, I believe. But you can hack around that. Impeccable uses a .impeccable folder in the current repository route. But you can also save things directly in the skill folder. And maybe ask the user to ignore them. How could this work? So, for example, if you're running a critique in Impeccable, that critique is saved as a file in that folder. And by default, it's getting ignored. But then, if you then later on say, well, okay, I just ran a critique. I'd like to polish my page. Even if you do it in another session, it actually uses that prior critique as a signal to understand what have we found out about this page. And it can look at all prior critiques and see the progression of the page. For example, you could have said in one of the critiques, I don't agree with this critique. I don't think you're right. And I think I really like my instruments. And then the model would be like, okay, no problem. I'll mark this for later. And the skill is now smart enough. The skill has built context to realize, okay, well, that's the user preference. So I'm going to respect it going forward. So, compound engineering really is an interesting theme for skills as well. You can make skills aware of prior sessions with that technique. So, make the runs compound. This works really well for resumable gradients, resumable agents, progress tracking, multi-session refactors, migrations, that kind of thing. For instance, one of the things that I do all the time is refactor my code. And how do I do that? By having a skill that spawns itself across multiple sessions and tackles one file at a time. So, I basically tell it, okay, here's your TSX file or whatever for today's session. And now refactor everything around this file and linking into that file. And then it builds up context over time until it's completely finished with the whole code base. Okay, number five. Buried rules get skimmed. We talked a bit about this before, but this is a little bit of a different point I'm trying to make. Now, especially with weaker models, and if you're building a skill for yourself and you're only running Opus or you're only running Codex, this isn't that big of an issue. You know which model you run. You know if it works with GPT-55, for example, I'm good because that's the only model I use. Now, if you want to distribute your skill to lots of users, this is where things get hairy. Because some of those users might be running Sonnet. Some of them might be running Haiku. Some of them might be running Grok. You never know. Sometimes I meet somebody who does. But that's where it gets complicated, right? Because you need to build for the lowest common denominator. And ideally for the one model that is the weakest at instruction following. For example, GPT-5 Mini is not a very good rule follower. There are things even in Impeccable that don't work with GPT-5 Mini. It consistently doesn't load certain MD files that are sorted. It consistently doesn't spin up the live mode. So there are boundaries to instruction following across these models. And it gets especially bad with longer skills that have lots of rules. So how do you work around this? Well, in Impeccable, Impeccable really is kind of bionic of sorts. It's really not just prose. It is a combination of scripts that run in line within the skill at certain times and then prose around it. For example, every time you call Impeccable, it runs a file called context.mjs. And the context.mjs does a couple of things. The first thing is, if there is a product MD, which is in Impeccable, it's almost like design MD, but it is for product strategy. So it wants to understand who is the target audience or what you want to achieve with this thing, which is oftentimes more important in a design interview than how round do you want your borders to be. But it supports both. It supports product MD and design MD. And by default context.mjs brings these files together and then spits them into the session. Now that's not exciting. But when those files are not available, it will actually give the skill structured JSON and say, by the way, there is no product MD. And here's exactly what you should do about it. Or here's another thing that context.mjs does. It actually makes Impeccable self update if there's a new version of Impeccable. Now with your permission, so we will ask you. But it will say, hey, by the way, there's an update available for the Impeccable skill. And here's what you should do now to ask the user whether they want to update Impeccable. So it's overloaded in many ways. And it will always tell the model the exact instructions on what to do next. And the really interesting thing about this is that I found that that works significantly better than some random rule in the prose of the main skill. But when those files are not available, it will actually give the skill structured JSON and say, by the way, there is no product MD. And here's exactly what you should do about it. Or here's another thing that context.mjs does. It actually makes impeccable self-update if there's a new version of impeccable. Now with your permission, so we will ask you. But it will say, hey, by the way, there's an update available for the impeccable skill. And here's what you should do now to ask the user whether they want to update impeccable. So it's overloaded in many ways. And it will always tell the model the exact instructions on what to do next. And the really interesting thing about this is that I found that that works significantly better than some random rule in the prompts of the main skill. So when you put something out from the exit value, from the standard out of a script, the model will follow it a lot more than before. So that could be environment-aware setup, dynamic onboarding, repo state gating, adaptive flows, anything really. Actually before I end this session, one of the shortcomings of this technique, and this is something to be aware of, is prompt caching. So this works super well to keep a skill flowing in the right direction, instruction following, but it does so at the expense of prompt caching. If you need prompt caching, if you run this skill many times, and you want the whole thing to be cached, this is not a good technique to use. But I found it to be very useful in really interactive scenarios. All right. Number six. Hooks that fight back. It's something I shipped quite recently, and I really like it. I want to show you what I mean by that. So a lot of people have impeccable systems, but sometimes they forget to run it. Sometimes they don't know. I wish Codex is actually pretty good. Some of the harnesses are pretty good, consistently looping in the right skill. But because it now bundles as one skill, oftentimes the harnesses forget to simply call impeccable when you don't explicitly mention it. So now you're building some front end code and maybe it doesn't follow your design system or whatever. Now that can be solved with hooks. Who has used hooks before in Cloud Code or Codex? A few people. Okay, nice. So this skill that I've built here, impeccable, ships design hooks. So I've basically built a design linter that runs under the hood and ships with the skill. When you install impeccable, these hooks install into Cloud Code, Cursor, Codex, and GitHub Copilot. And they will keep the model exactly where it needs to be. So the hooks come to you. It's a guardrail that fires on every edit. And there are some differences between the different providers here. So the hook syntax for Codex and Cloud Code is not the same. And also the behavior is not the same. So for instance, we found out that with weaker models, slightly weaker models like Composer and Cursor, you want to use a pre-tool use hook that prevents writing of code as opposed to a post-tool use hook. Post-tool use basically happens right after the agent has written a file, for example. And then it tells you, hey, by the way, the contrast of these colors is bad. Or you have a purple gradient in here. And then ideally the model is smart enough to actually fix it. Some models don't follow those instructions very well. And so if you do a pre-tool use hook, you are actively preventing the writing of this file in the first place. So it's a much more heavy-handed approach. But we needed to do that for certain models and certain harnesses. But this is nice. And what's even nicer about it is that you can personalize it to your design system and your use case. Or whether, let's say you use it for code reviews. You can personalize it with your own ES lint rules, with your own syntax guidelines, et cetera. And then expand it from there. So passive guardrails beat a command no one remembers to run. So these are passive guardrails that always keep you in the right lane on track. Again, that works for linting, for formatting. Of course, if you're using Cloud Code or Codex, it already uses some of the linters for things like syntax formatting. But design linting is a whole different game. But I would really encourage you to try out hooks in combination with a skill. And think about, okay, well, my skill does this. How can I create a feedback loop, a validation loop that uses hooks to actually keep me on the right lane? Okay. So here's, it's hard to show hooks in action. But if you can see this, this is roughly how it would happen in an agent. So, for instance, in this case, I would use, let's say, Gemini does this all the time. Gemini creates animations on images constantly. It will animate any image and it will usually do a hover zoom in effect. It loves that. And that's something that impeccable flags. And in this case, the hook would fire silently. That's why I built this fake demo because you can't usually see it. And then it will tell the model, hey, by the way, here was a violation. The experience of this is that oftentimes you don't have to do anything. The model just course corrects and fixes itself. Now, one important thing, if you do this and you ship it to users, very important to add a way to create ignore rules or something. Because oftentimes these hooks have false positives as well. And you want a way to configure those hooks. Otherwise it gets annoying very quickly. Impeccable ships with these design hooks that allow you to create ignore rules at a file basis within a CSS rule. So many granular levels to exclude certain files, for example. Okay, level seven. Now you can't really tune pixels through a chat box. Now this might not be relevant if you're not building a design skill. But I think the general point is relevant. So if you think about a skill as harness engineering versus prompting, then you think about the harness as a whole, right? You're living in Cloud Code, for example, or you're living in Codex, or you're living in GitHub Copilot. Now, what are the capabilities of that harness that you can exploit to make the best user experience for your use case? That's the question you should ask yourself. For example, Codex on Desktop now has an in-app browser built into the actual app. Can you use this in-app browser in some interesting ways? Can you use the browser screenshot tool in some interesting ways? And in my case, I could. I realized, hey, there's probably a way to connect the in-app browser and spin up the development server and just load the page there and then connect it to the main thread in some ways so I can allow the user to visually iterate on that page instead of in the chat. And so in Impeccable, what this looks like is it's not using MCP. It's simply spinning up a live poller, a little server that looks for input and inserts a snippet into your development server. It then, on the page, when you do something on the page, it sends an event back to that poller using server-side events. And then, and this is, I think, the clever bit maybe, or the bit that makes it all work. The poller then stops. So the poller ends itself. There's a standard out message. We talked about standard out before, right? The exit value of this thing. And the model reads that message and realizes, oh, something happened. I better do something. So in this case, in the skill itself, I give it instructions on how to handle this event. I say, well, if this event comes in, you should probably build some design for this particular section of the page. And then you should send it back to this poller so that it arrives on the user side. And so this is a direct connection between one harness capability and another harness capability. So the chat thread and the in-app browser. And yeah, this is how it looks on a diagram. But I think the best way to experience it is to see it. So let me bring this up. Okay, Cursor. I think I'm already in live mode here. Okay, so I booted up live mode already. I'm now in picker mode. I get this little bar here at the bottom. And as you can see, I can pick anything on this page. I now get this little overlay bar. And I can select all sorts of subcommands within the skill. So these are basically translating to MD files that live within the skill. I can select the amount of variance I want. And then I can hit go. And now here in the thread, you can see that it picked up the actual signal in the main thread. Because the poller stopped. And it now knows hopefully exactly what it needs to do to first wrap this element in some special tag. Then it knows how to create variance that are marked up in a special way with CSS. And now it did that. So now, as you can see, the thing updated immediately. I now get these three variance. And I can click through. Okay, so I booted up live mode already. I'm now in picker mode. I get this little bar here at the bottom. And as you can see, I can pick anything on this page. I now get this little overlay bar. And I can select all sorts of subcommands within the skill. So these are translating to MD files that live within the skill. I can select the amount of variance I want. And then I can hit go. And now here in the thread, you can see that it picked up the actual signal in the main thread. Because the poller stopped. And it now knows hopefully exactly what it needs to do to first wrap this element in some special tag. Then it knows how to create variance that are marked up in a special way with CSS. And now it did that. So now, as you can see, the thing updated immediately. I now get these three variance. And I can click through. And then if I like one of them, I can click accept and accept it. If I don't like one of them, I hit escape and I'm back in normal mode. So this shows how to exploit a harness capability in an effective way for one problem space, in this case design. You can also insert elements with this thing. And so click into anything here. You can draw on top of this and leave comments. You can leave annotations if you want. You can dictate. You can steer the whole page by simply writing into this. And then again, this goes back to the main agent. And it becomes a steering signal for the whole page. And you can also visualize lots of things this way. You might have read Tharik's blog post about this, about how HTML is a really cool way to communicate as opposed to Markdown. I agree. And I think designMD is much better visualized as HTML. In this case, you see the designMD of this website for demonstration purposes. But you can use this to advantage as well if you hide the in-app browser and use it to your advantage. So this is how I make use of it. Okay. Number eight. It worked on my machine. Well, everybody who is a developer here knows this problem. This hits really hard when you ship a skill. There are so many times I saw this argument on X. It was like, hey, bro, just symlink. Just symlink clawd and all your problems will be gone. Well, that's great if you're building a simple skill and if you're doing it for yourself. By all means, go for it, right? Symlink your clawdMD to agent.md. It's going to be amazing. Symlink everything. But it's not great if you're trying to ship a skill to lots of users. Because, again, we just talked about a whole lot of differences these harnesses have. I'm going to talk about more differences. And I know it's annoying because it would be great to symlink those things. But unfortunately, we don't live in that world. And unfortunately, Anthropic has still not adopted agents.md. So, what are the actual differences? For example, we talked about sub-agents already. We talked about how, well, on the bright side, they're widely supported now. But who can spawn one is very, very different. So, with clawd, you can programmatically do it very easily. Codex needs the user okay. In cursor, it's agent chosen most of the time. So, there are clear differences. Also, if you want to pre-define these agents, Codex has a different syntax for that than clawd and cursor. Another one is the Ask User tool. So, one of the coolest tools in the clawd code harness is the Ask User Question tool. It's a really nice tool that you can use to ask the user a question. It brings up this menu. So, hey, what would you like to do? And then you pick some option. Well, turns out Codex has a tool like this. That's the good news. The bad news is that tool is only available in plan mode. So, again, big differences between how these things work. And what does that mean? That means that if you're not running Codex in plan mode, but your skill wants to ask questions, most of the time it simply doesn't. It will simply infer from the current context and not ask any questions to the user, which is not great. So, there's a lot of sentences in the Impeccable skill that specifically say, if you're Codex, you have to stop and ask questions. No, you're not smart enough to infer the context. So, if you see lines like this, that's why. Another one is background jobs. And it's something you learn through the hard way by doing this. For example, this live mode that I just showed you, it's spawning a background task. So, it's running a shell in a background task. And that's cool because you can keep using the session. And then when the background task finishes, the model is automatically woken up, gets the message back, and then can do something and react to it. Where Codex cannot. Codex and other harnesses do not react when a background task finishes. You actually have to manually say, hey, by the way, this background task, can you take a look at what it did? And that's not great if you were doing an automation like this. So, there are differences in how these tasks are spawned and how they work. So, that's why if you're using the live mode in cursor or in Codex, it creates a foreground task. And it keeps the actual chat thread blocked. Not ideal, but it makes it actually work. So, there are subtle differences in how these tasks are spawned. Watchers is another example. Tail and watch exist now. That's really cool. Most of the harnesses have a way to watch, for instance, a log file. But those are throttled way harder than simply spawning a background task. Edit hooks, we talked about this already. They are different. And so, lots and lots of behavioral differences. But there's also model differences. So, for example, in my case, they all have different tails in the ways they're overfitted. For example, Gemini, again, I mentioned this, loves to animate pictures. It just loves it. You have to tell it not to animate pictures if you don't want a hover effect on every picture. It doesn't matter where it is. It loves it. Codex loves bad letter spacing. I don't know why, but it does. Codex also loves extremely rounded borders. It will round anything you thought it. It loves it. It doesn't matter if it's a hospital website or a kid's website. It also loves hairline borders. And so, there are specific tails that are unique to every model. And that's not just for design. It's for architecture. It's for code architecture. It's for preferred NPM packages. Every model is overfitted in different ways. Finding out how to overfit it usually happens by accident. In my case, I have a pretty extensive evals harness that I run behind the scenes. In fact, every line of Impeccable is ablation tested. So, I test every single line and see what it does across all models. I don't expect you to do that, but it is very good to know that the models are different and are following instructions differently and the behavior, the harness behavior is different as well. And so, what Impeccable does, it creates harness-specific and model-specific builds for every single model. You might not have to go all this way for your own purposes, but I just wanted to show you how far you can go with this. For example, it actually has a substitute variable that picks the right user question tool, depending on the harness. Or, it has these XML blocks for Gemini, for Codex, etc. that will actually insert specific overfitting avoidance rules for the given models. Because it turns out, if you tell Claude not to let us space too much, it will let us space in the exact opposite direction. So, you can't just include it all in the same skill. And that's why you can, if you instrument this way enough, you can actually get to this write once, ship to all of them skill that actually works everywhere. It's a lot of work, but it does pay off and allows you to create beautiful pictures like this. Now, the only other problem is that typical install methods, like for NPX skills, for instance, if you've been using NPX skills, do not honor different directories for different harnesses. So, they actually just take the first directory and then copy it or symlink it into all sorts of folders. That's why if you go to the Impeccable website, I've built my own CLI to solve this problem. That's why it doesn't use NPX skills. So, I think the community hasn't quite yet gotten to the point where this is an accepted idea. And it's annoying. I get it. It's annoying to compile for different harnesses, but I found it worthwhile. Finally, again, built for the lowest common denominator. Write once, ship to all of them skill that actually works everywhere. It's a lot of work, but it does pay off and allows you to create beautiful pictures like this. Now, the only other problem is that typical install methods, like for MPX skills, for instance, if you've been using MPX skills, do not honor different directories for different harnesses. So, they actually just take the first directory and then copy it or symlink it into all sorts of folders. That's why if you go to the Impeccable website, I've built my own CLI to solve this problem. That's why it doesn't use MPX skills. So, I think the community hasn't quite yet gotten to the point where this is an accepted idea. And it's annoying. I get it. It's annoying to compile for different harnesses, but I found it worthwhile. Finally, again, built for the lowest common denominator. Our weaker model has opinions just fine, but what it loses is the discipline to follow yours. So, Codex, for example, and GPT specifically, loves the word gate. If you've built a skill in Codex before, it loves gates. Whenever you say, hey, why didn't you follow these instructions? You're like, well, I think we need a gate. So, I gave it what it loves the most, gates. But I only do that for Codex. So, there's a CodexMD that gets loaded on the fly for Codex and GPT. And then it actually follows, okay, here are your eight gates. You have to pass every single gate. And you are not allowed to compress those gates. That's really important because it loves compressing these instructions as well. Just skim over it and say, well, I guess I do one and two and five and good. And so, the way you solve this is by actually having it log every single result of every gate and say, well, I just passed gate one. Great success. And the most important lesson from this is, if the gate can be skipped, it will be. I mentioned this before, right? If the model can wiggle itself out of a difficult situation, it will absolutely do that. It will not do all the things it needs to do to complete the end result. So, be careful. Make it unskippable. So, we just built a harness extension. We went from prompting all the way to building a monster. But I think it turned out to be pretty powerful in my case. And I wanted to share what I've learned on the way. I don't expect you to use all of those techniques. I think some of them are pretty exotic and maybe not applicable to every use case. But I hope that you find value in some of the advice that I've given today. So, we've done a whole bunch of things today. Nine things a prompt can't do. We made it much more deterministic and made Impeccable better for that reason. If you like to try it out yourself, again, you can clone the repository for this talk. You can clone Impeccable-minus-talks. But, of course, it also is useful to just take a look at the actual skill and see how it's built. The project is completely open source, licensed under Apache 2. You can install MPX Impeccable Skills install and check out the source code on GitHub. With that, I'm at the end of it. Thank you. And now I think we have about 10 minutes for any questions that you have. Does anybody have questions? Yes? Oh, sorry. What was that? A link? To the repository? Yeah. So, this is hard to see, but let me put it up here. This is the repository for the talks. Yeah. Awesome. So, the question is, I mentioned that it breaks prompt caching. The actual trick, the technique to actually get something back from a script within a skill. And the reason is because the result is dynamic, right? It could be anything. So, unless the result is always the same, it's a dynamic shell execution. So, it gets inserted into the thread. Now, to be fair, the skill will still be cached. So, the skill will still be cached. But, I guess I'm differentiating between the skill with inline static content versus the skill with a dynamic instruction to call out. So, this part will not get cached. Yeah. That was my main point. Yeah. Yeah. What is the process on how I evaluate and iterate on this skill? So, the process is pretty involved. Let me see. Let me see if I can bring this up on screen. Okay. Here we go. So, here's a glimpse. Oh, no. Okay. I just shut down the server. That's fine. Okay. I'll just voice over. So, yeah. I mentioned I built an evals harness. And so, I've created myself a harness that closely recreates the conditions and the tools of every harness that I care about. So, for instance, it uses the cloud code SDK. Yes. Sorry, guys. Can you lower your volume a little bit? Because people are still trying to hear the questions. Thank you. So, how do I test this? How do I build it? So, it's a combination. So, first of all, Impeccable has a ton of end-to-end tests in the repository. That's both LLM-driven tests as well as end-to-end playwright tests. So, that's one. And that's useful for things like testing the live mode scripts, for example. But then beyond that, how do I test that it actually works? Well, I've built an evals harness. That one is not open source yet. But I built an evals harness that closely replicates every model harness that I care about right now. Specifically right now, Cloud Code, Codex, and Gemini. And I'm trying to expand it to more. And it also recreates the tools, like for instance, a browser screenshot tool or something along those lines. And then it also recreates the, because some parts of Impeccable are interactive. In the initialization of Impeccable, oftentimes the user gets asked, so what would you like your page to feel like? And so, you get these interactive back and forth. And so, I've built this LLM that acts as the user against the other LLM. And so, it does an interactive back and forth turn. So, I've built that harness and then I've built a mixture of expert design judge that runs on top of it. So, basically, give it eyes to evaluate each result. And then I can run across 20 different niches, like for instance, Italian restaurant. I run across all models that I care about. GPT-4, OPUS, SONNET. And do five to ten tests for each of those, for each skill release to see how it changed. I also run against competitors. For instance, I run against the front end design skill to see, does it make a difference? And how does it make it worse or better? And then, beyond that, I'm doing ablation testing. That's harder and more expensive, I would say. So, I don't recommend it for everyone. But this, the ablation testing, so every, you'll see this in the source card of Impeccable. Every rule has an XML tag that says a unique identifier of that particular line. And that will be used by the harness to then do a test where it removes that line, runs the evals against all models, and then adds the line back in, and then uses the detection engine of Impeccable, the deterministic one, to see, did it actually change? Right? So, if there's a line that says, don't do gray on colorful backgrounds for contrast purposes. There's an ablation test and then a deterministic check or feedback loop that tests against it. So, in short, quite involved. But I really, it started vibes based, and now it's really, truly well tested. Yeah, go ahead. Can you set up evals for evaluating taste? Sorry? Can you set up evals for evaluating taste? Yes. I do have evals for evaluating taste, but I don't think they work particularly well. I just talked to Ben from Contra about this. I don't think, I mean, I know some of my colleagues might disagree, but I don't think taste can be solved at a model level. I actually think it's a fundamentally human thing, because taste is scarce and unique, and once everybody uses the same taste, it becomes ubiquitous, and then we don't think it's tasteful anymore. So, I think it's hot. And I also think the models are particularly bad at evaluating taste. So, for example, there are certain things that the models can evaluate well, like, hey, is this the correct thing in the first viewport? Right? So, functional stuff, that works. But what doesn't work, and here's one example, I've built, again, this mixture of judges, and one judge rates whether the first viewport looks great, right? And is effective. I just talked to Ben from Contra about this. I don't think taste can be solved at a model level. I actually think it's a fundamentally human thing, because taste is scarce and unique, and once everybody uses the same taste, it becomes ubiquitous, and then we don't think it's tasteful anymore. So I think it's hot. And I also think the models are particularly bad at evaluating taste. So for example, there are certain things that the models can evaluate well, like, hey, is the correct thing in the first viewport? Right? So functional stuff, that works. But what doesn't work, and here's one example, I've built a mixture of judges, and one judge rates whether the first viewport looks great, right? And is effective. And one of the tells is that Gemini, for example, the more stuff there is in the first viewport, the higher it rates it. Right? This is just a general rule. If you just cram the viewport full, it gives it a higher ranking. And so there's an interesting example of the models are often maximalists, right? They're like, well, more is more, I guess. And so oftentimes I build judges that actually invert the response of the model, which is really strange, but it works. Where it sort of judges something very high, I'm like, okay, that's definitely not a good design. So anyway, I don't think it's solved, and I don't think it's solvable, but I do have a tool that gives you the design director eyes that works marginally better than random, and that's good enough for me for a first pass, and then I use my own human eyes to evaluate results and annotate them. Any other questions? Yeah, over here. What do you say is the future for sales? Any other questions? The future for skills? So I would say that's a broad question. I think right now you're doing a lot of different approaches or what you think is a better way to do that? Yeah, so I'll first answer for Impeccable and for me. So in the case of Impeccable, I think we're definitely outgrowing the skill platform, what's possible with skills. I think the live mode is a good example of that. The live mode was sort of a Jurassic Park experiment to see if I can do this. And the answer is yes-ish. I think it's working better than I expected, but it still has a lot of problems. I mean, it would be way better to do this in a first-party harness integration or a first-party tool. So I think there are limits that I'm hitting where skills might not be effective anymore. I think in general, I would say most skills should probably be written by the individual users. I think those that actually go through the effort of packaging a skill and sharing it with others need to invest more time than they currently do. So I guess that's my hot take. I think right now I've seen plenty of skills that are distributed that do not work well in a model that the author didn't use, for example. And so I think we just have to raise the bar of what's acceptable to ship to people. I mean, this is the "works on my machine" thing. I would rather see fewer skills in the ecosystem that are really battle-tested and proven. And I hope we're shifting towards that because right now it's sort of a wild west. Yeah, go ahead. Just in that regard, there's no common way to test the skills to make sure that it works for all answers or models. So there could be an opportunity to do something? Yeah. There's no common way to test the skills. Yeah. And then there could be an opportunity. That's a good point. Yeah. I guess I could. I do have the tool for that. That's true. Yes. I could do something with it. Yeah. Right now it's purely built for my own purposes. But yeah, the same is true for, for instance, the Impeccable installer and compiler. I don't think most people know that it exists, that it can compile to every harness and that it has these substitution techniques and stuff like this. I could probably release that standalone as well. Yeah, that's a good point. Yeah, go ahead. What would I ask about if you're having a bunch of things that you have to do on the server? MCP having skills on the server? How would that work? Oh, I see. Yeah. To be honest, I haven't tried it out yet, or I haven't really read too much into it. I think MCP in general, I worry greatly about context pollution. And I do that with skills too. And I think I'm not using MCP a lot for that reason, because it polluted my context many times. How do skills work in MCP? Oh, you can download a script from MCP server. Yeah, okay. One of the issues that I think I was mentioning about is the recommended way to have to do it. Yeah, yeah, yeah. So what's the recommended way of packaging them and distributing them? Yeah, that's a good topic. So of course, the harnesses and the frontier labs have their own ways. I mean, Codex has a marketplace that you can use for distribution, plugin marketplace. Cloud Code has a marketplace as well. I think they started with the marketplace technique. Those marketplaces don't work particularly well. I mean, the Cloud Code one for sure doesn't work particularly well. I know this for a fact because the update mechanism often doesn't work. And people are like, well, my skill doesn't update. And oftentimes there's a caching issue. So my experience has been hit or miss with the native methods of distributing. And then, of course, it's only for that particular provider. That's why projects like skills.sh exist. But again, the problem with MPX skills right now, it doesn't allow for more advanced skill use cases, like compiled for every different harness. I have a pull request in the repository. And I've bugged Andrew a couple times about it. But he still has to get it merged or I agree to agree with me on that, I guess. I think we're still discussing. But yeah, MPX skills, I think is a great project in general. I think it'd be great if we could standardize around it. There's also one from Microsoft that's trying to do that. A project from Microsoft, I forgot the name of it. But there's definitely no industry standard for distribution yet. Yeah, I'm not, I don't love having to maintain my own CLI installer. I would rather not. It's annoying. But it does make it so it plays safe with all harnesses installed, the hooks in the right part of the system, etc. So yeah. Yeah. Okay, I think I'm way out of time. But come up and speak with me if you like. Yeah. I would say I'll end it here, but you have come up if you like. Let me just, thank you. Thank you. So, the process is pretty involved. Let me see. Let me see if I can bring this up on screen. Okay. Here we go. So, here's a glimpse. Oh, no. Okay. I just shut down the server. That's fine. Okay. I'll just voice over. So, yeah. I mentioned I built an evals harness. And so, I've created myself a harness that closely recreates the conditions and the tools of every harness that I care about. So, for instance, it uses the cloud code SDK. Yes. Sorry, guys. Can you lower your volume a little bit? Because people are still trying to hear the questions. Thank you. So, how do I test this? How do I build it? So, it's a combination. So, first of all, Impecable has a ton of end-to-end tests in the repository. That's both LNM-driven tests as well as end-to-end playwright tests. So, that's one. And that's useful for things like testing the live mode scripts, for example. But then beyond that, how do I test that it actually works? Well, I've built an evals harness. That one is not open source yet. But I built an evals harness that closely replicates every model harness that I care about right now. Specifically right now, Cloud Code, Codex, and Gemini. And I'm trying to expand it to more. And it also recreates the tools, like for instance, a browser screenshot tools or something along those lines. And then it also recreates the, because some parts of Impeccable are interactive. In the initialization of Impeccable, oftentimes the user gets asked, so what would you like your page not to feel like? And so, you get these interactive back and forth. And so, I've built this LLM that acts as the user against the other LLM. And so, it does like an interactive back and forth turn. So, I've built that harness and then I've built a mixture of expert design judge that runs on top of it. So, basically, give it eyes to evaluate each result. And then I can run across 20 different niches, like for instance, Italian restaurant. I run across all models that I care about. GPT-55, OPUS, SONNET. And do like five to ten tests for each of those, for each skill release to see, you know, how it changed. I also run against competitors. For instance, I run against the front end design skill to see, does it make a difference? And how does it make it worse or better? And then, beyond that, I'm doing ablation testing. That's harder and more expensive, I would say. So, I don't recommend it for everyone. But this, the ablation testing, so every, you'll see this in the source card of Impeccable. Every rule has sort of an XML tag that says like, you know, a unique identifier of that particular line. And that will be used by the harness to then do a test where it removes that line, runs the evals against all models, and then adds the line back in, and then uses the detection engine of Impeccable, the deterministic one, to see, did it actually change? Right? So, if there's a line that says, hey, don't do like gray on colorful backgrounds for contrast purposes. There's an ablation test and then a deterministic check or feedback loop that tests against it. So, in short, quite involved. But I really, it started, you know, vibes based, and now it's really, truly well tested. Yeah, go ahead. Can you set up evals for evaluating taste? Sorry? Can you set up evals for evaluating taste? Yes. Can you set up evals for evaluating taste? Can you set up evals for evaluating taste? Yes. I do have evals for evaluating taste, but I don't think they work particularly well. I just talked to Ben from Contra about this. I don't think, I mean, I know some of my colleagues might disagree, but I don't think taste can be solved at a model level. I actually think it's a fundamentally human thing, because taste is scarce and unique, and once everybody uses the same taste, it becomes ubiquitous, and then we don't think it's tasteful anymore. So, I think it's, I think it's hot. And I also think the models are particularly bad at evaluating taste. So, for example, there are certain things that the models can evaluate well, like, hey, is this, is the correct thing in the first viewport? Right? So, functional stuff, that works. But what doesn't work, and here's one example, I've built, again, this mixture of judges, and one judge rates whether the first viewport looks great, right? And is effective. And one of the tells is that Gemini, for example, the more stuff there is in the first viewport, the higher it rates it. Right? This is just a general rule. Like, if you just cram the viewport full, it gives it a higher ranking. And so, there's an interesting example of, like, you know, the models are often maximalists, right? They're like, well, more is more, I guess. And so, oftentimes, I build judges that actually invert the response of the model, which is really strange, but it works. Where it sort of judges something very high, I'm like, okay, that's definitely not a good design. So, anyway, I don't think it's solved, and I don't think it's solvable, but I do have, I would say, a tool that gives you the design director eyes that works marginally better than random, and that's good enough for me for, like, a first pass, and then I use my own human eyes to evaluate results and annotate them. Any other questions? Yeah, over here. What do you say is the future for sales? Any other questions? The future for skills? So, I would say, hmm, that's a broad question. I think right now you're doing a lot of different rambles or what you think is a better way to do that? Yeah, so, I'll first answer for impeccable, and for me. So, in the case of impeccable, I think we're definitely outgrowing the skill platform, kind of what's possible with skills. I think the live mode is a good example of that. The live mode was sort of like a Jurassic Park experiment to see, like, can I do this? And the answer is yes-ish. I think it's working better than I expected, but it still has a lot of problems. I mean, it would be way better to do this in a first-party harness integration or like a first-party tool. So, I think there are limits that I'm hitting where skills might not be effective anymore. I think in general, I would say most skills should probably be written by the individual users. I think those that actually go through the effort of packaging a skill and sharing it with others need to invest more time than they currently do. So, I guess that's my hot take. I think right now I've seen plenty of skills that are distributed that do not work well in a model that the author didn't use, for example. And so, I think we just have to raise the bar of what's acceptable to ship to people. I mean, again, this is like the works on my machine thing. I would rather see less skills in the ecosystem that are really battle-tested and proven. And I hope we're shifting towards that because right now it's sort of like a white-west. Yeah, go ahead. Just in that regard, there's no, like, common way to test the skills to make sure that it works for all answers or models. So, there could be an opportunity to do something? Yeah. There's no common way to test the skills. Yeah. And then there could be an opportunity. That's a good point. Yeah. I guess I could. I do have the tool for that. That's true. Yes. I could do something with it. Yeah. Right now it's purely built for my own purposes. But yeah, the same is true for, for instance, like the impeccable installer and compiler. I don't think most people know that it exists, that it can compile to every harness and that it has these substitution techniques and stuff like this. Like I could probably release that standalone as well. Yeah, that's a good point. Yeah, go ahead. What would I ask about if you're having a bunch of things that you have to do on the server? MCP having skills on the server? How would that work? Oh, I see. Yeah. To be honest, I haven't tried it out yet. Or I haven't really read too much into it. I think MCP in general, you know, I worry greatly about context pollution. And I do that with skills too. And I think I'm not using MCP a lot for that reason. Because it polluted my context many times. How do skills work in MCP? Oh, you can download a script from MCP server. Yeah, okay. One of the issues that I think I was mentioning about is the recommended way to have to do it Yeah. Yeah, yeah, yeah. So what's the recommended way of packaging them and distributing them? Yeah, that's a good topic. So, of course, like the harnesses and the frontier labs have their own ways. I mean, Codex has a marketplace that you can use for distribution, plugin marketplace. Cloud Code has a marketplace as well. I think they started with the marketplace technique. Those marketplaces don't work particularly well. I mean, the Cloud Code one for sure doesn't work particularly well. I know this for a fact because, I mean, the update mechanism often doesn't work. And people are like, well, my skill doesn't update. And oftentimes there's a caching issue. So my experience has been hit or miss with the native methods of distributing. And then, of course, it's only for that particular provider. That's why projects like skills.sh exist. But again, the problem with MPX skills right now, it doesn't allow for like, you know, more advanced skill use cases, like, you know, compiled for every different harness. I have a pull request in the repository. And I've, I've bugged Andrew a couple times about it. But he, he still has to get it merged or I agree to agree with me on that, I guess. I think we're still discussing. But yeah, MPX skills, I think is a great project in general. I think it'd be great if we could sort of like standardize around it. There's also one from Microsoft that's trying to do that. A project from Microsoft, I forgot the name of it. But there's definitely no, no industry standard for distribution yet. Yeah, I'm not, I don't love having to maintain my own CLI installer. I would rather not. It's annoying. But it does make it so it plays safe with all harnesses installed, the hooks in the right part of the system, etc. So it's, yeah. Yeah. Okay, I think I'm way out of time. But come up and speak with me if you like. Yeah. I would say I'll end it here, but you have come up if you like. Let me just, thank you. Thank you.