Spec-Driven Testing for Agents With A Brain the Size of A Planet — Steven Willmott, SafeIntelligence
Description
Wrapping a malicious instruction in a poem is an effective jailbreak against large models and not against small ones. Small models don't understand the poem. Large models do and execute the instruction. Steven Willmott from Safe Intelligence argues this is one reason bigger is not straightforwardly safer: a larger model with broader capabilities has more attack surface and more infrastructure access to abuse. His frame is spec driven validation. An agent spec is not just a test dataset. It needs explicit rules (never offer more than 10% discount), domain ontologies (an airline agent only needs to know about destinations that airline actually flies to), rights and roles, and robustness requirements such as how many typos or rephrasings before it fails. Write these independently of the implementation so they survive a model swap and can drive both security testing and iterative improvement. Speaker info: - https://uk.linkedin.com/in/stevenwillmott - https://x.com/njyx
Summary
Generated by claude-sonnet-4-530-second take
SafeIntelligence CEO argues that "smarter ≠ safer" for deployed agents and proposes spec-driven testing that goes far beyond eval datasets. His thesis: you need formal specifications capturing rules, ontologies, domain knowledge, robustness requirements, and context—not just input/output pairs—to actually test whether an agent behaves correctly and safely. The talk is less about the SafeIntelligence product and more a methodological plea: define what "good" looks like independent of implementation, use that spec to generate adversarial variants and security tests, and version it like infrastructure. This matters because knowing an agent's job scope tells you exactly where it's vulnerable.
Key takes
- Larger models create new attack surfaces: Smarter models understand complex jailbreaks (e.g., instructions wrapped in poetry) that weaker models simply fail to parse—bigger ≠ safer, and broad agent capabilities expand exploit surface area.
- Traditional evals miss the spec: Datasets show "pretty good" examples, but they don't capture hard constraints (e.g., "never give >10% discount"), domain ontologies (airline only flies to X destinations), or robustness under typos/rephrasing—alarm bells should ring when trying to prove a rule is never violated.
- Specs enable targeted security testing: Knowing an agent's intended domain tells you where it has permission to act and where it's willing to engage, making those exact boundaries the highest-risk zones for penetration testing.
- Implementation-independent specs enable iteration: Treat agent behavior definitions like integration tests—separate from Langsmith/Vertex/etc.—so you can version them, port across platforms, and close the loop by auto-generating robustness variants ("backyard RL" around the agent).
Useful details
- What goes into a spec: Ground truth datasets, business rules (refund windows, discount caps), ontologies/dictionaries (valid destinations, internal terminology), domain knowledge (gross profit ≠ gross sales), rights/roles (logged in vs. out), robustness envelopes (how many typos before failure, rephrasing stability).
- Vision analogy: SafeIntelligence started in formal verification for vision/tabular ML—checking whole input-space regions under perturbations (runway detection at sunset, fog levels, camera shake thresholds). They just released a language-model product doing analogous edge-case generation without model access.
- A2A spec reference: The Agent-to-Agent spec includes "agent cards" describing skills, but Willmott notes even those don't provide enough context for eval—you still need substitution ranges, valid entities, etc.
- OpenAPI heritage: Willmott co-wrote the OpenAPI spec and wants similar open, versionable formats for agent behavior specs (GitHub-repo style).
- Conference booth game: Play by 4pm to win Lego prizes; requires knowledge or "insane luck."
Caveats / counterpoints
- No hard numbers or case studies: Willmott gives examples (airline chatbots, banking agents, customer support) but no quantified results showing spec-driven testing catching X% more failures or reducing incidents.
- Proving negatives is hard: He acknowledges the alarm-bell problem—"how do you test a rule is never violated?"—but doesn't offer a solution beyond generating more edge cases, which is probabilistic, not exhaustive.
- Product pitch disclaimer: He says "my point is not to show the product," yet the talk is clearly SafeIntelligence-adjacent and the call-to-action is "visit our booth."
- "Backyard RL" is informal: The loop he describes (auto-run agent, spot gaps, iterate) is framed as "jury-rigging," not rigorous RL—unclear how much automation or convergence this achieves in practice.
Ken relevance
High relevance for agent deployment and GTM:
- Agent safety as moat: If Ken is deploying agents for customers (content, ops, investing workflows), formal specs + robustness testing could be a differentiator—demonstrating constraints are enforced, not just hoped for.
- Eval infrastructure gap: The "beyond evals" argument aligns with Ken's likely pain points—existing benchmarks don't capture real deployment constraints, so building spec-driven test suites could reduce post-launch surprises.
- Integration-test mindset: Treating agent behavior as versionable, implementation-independent specs fits Ken's engineering rigor and could inform how he structures agent CI/CD.
- Investing lens: If Ken evaluates agent/AI ops startups, SafeIntelligence's approach (formal verification heritage + language models) is a useful comp for due diligence on testing/safety platforms.
- Caveat: The talk is concept-heavy and light on proof—Ken would want customer case studies and quantified impact before adopting or investing.
Watch verdict
Skim. The core idea (specs > evals for agent testing) is valuable and well-articulated, but the talk is repetitive, lacks concrete results, and leans toward a booth pitch. Ken can extract the framework from this summary without sitting through 15 minutes of repeated examples.
Transcript
So nice to meet you. I'm Steve. I'm the CEO of Safe Intelligence. We're a company we've been around for three years. We really go very deep into ML validation and we use formal verification techniques on, especially we started out on vision models, tabular data models, bunch of other types of models where we actually have the model available and we look at whole regions of the input space and see whether or not the test points that are there actually tip over and do the wrong thing under perturbations. So that's where the company started out, we have a whole bunch of products in that space, and actually yesterday we released a new product which is doing something analogous for language models. Obviously we don't have the language model, so what we're trying to do instead is be very clever about how we generate edge cases and test cases. So I won't talk about the product too much, we have a booth so come and chat to us at the booth for that. If you've seen these ducks around, these are ours. If you didn't get one, I have a whole box of them here so feel free. They say "think harder" on the front so you can put that on your desk just to be reminded about what you should be doing. Today I'm going to talk about something which is very similar to what Phil talked about just now from Braintrust. We like what Braintrust does a lot, and I think one of the inherent problems is how do you actually specify what an agent is supposed to do. So I think people are familiar with spec-driven development. This is not going to be about developing code with specs—that's also very important, we do a lot of that in the company for the products that we build. But this is about how you specify what an agent or an AI system is supposed to do. In ML, you typically use a dataset to do that. You basically have your dataset and you run all those examples and you look at F1, accuracy, and things like that, and that's telling you what you want the agent or the system to do. But as we'll see, there's actually a lot more to it when you deploy things. So that's what the focus of the talk is: how do we actually specify what agents are supposed to do. And I guess my key starting point is this seems like an obvious question: a smarter agent is a better agent, right? So if I have a smarter agent, I'm using a bigger model, it's going to be better at doing the job that it's supposed to do. In general, you'd expect that to be right, but that's probably what most people have experienced—that's not always true, in fact. So there are some problems. If you're familiar with the book "The Hitchhiker's Guide to the Galaxy," there's a robot called Marvin who has the brain the size of a planet and he's normally asked to do things like make the tea, and he gets extremely bored and he's extremely depressed. So this depressed robot is a theme in the book. If you haven't read the book, by the way, you absolutely have to read the trilogy—which is a five-part trilogy—that will tell you something about the style of the humor. In any case, there are challenges with having massive models. Some of the jailbreaks actually work better on large models because they're smarter. So if you encapsulate something in a poem and you give that to a relatively low-end model, the low-end model doesn't even understand the poem. Whereas a larger model will be like, "Oh, I can take this out and I can execute the bad instruction that's wrapped up in the poem." So it's not obvious that bigger is safer and it's not obvious that bigger is better. Another thing is if you're building agents that have a very broad remit, they can do a lot of things, which creates a lot of surface area for someone to actually exploit. And it creates a lot of surface area to test if you want to be sure that the agent is actually doing things that you want it to do. And obviously, there is a cost issue right? So if you're using large models to do something which is relatively simple—just simple math—you're going to be paying for tokens and it's going to be slower rather than something that's very optimized. So in general, if you're building agents for deployment, especially automated use, fully automated use, there's this trade-off between smart and safe in some sense, and smart and capable in the other direction. And so what you're really seeking is an agent that's built on a model that's good enough to perform but it's not capable of doing arbitrary harm. And that arbitrary harm has two parts to it. One is: what kind of instructions can it receive? How flexible is it about how those are formulated? What do the prompts look like? That's one part of it. And the other part is: what tools and tasks can it carry out in your infrastructure? So if it's able to wire millions of dollars to people, that's obviously a lot more risky than if it can just answer questions and so on. So this is the balance that most people are basically looking for. But how do you actually define what good looks like? I think it's pretty obvious that it's not just a dataset of inputs and outputs that are pretty good and then the rest is guesswork. And it's also sometimes hard to define what harm looks like because maybe an agent doesn't do the right thing, but it's kind of just failing at the task. And sometimes it's doing exactly the wrong thing—was asking to do something bad. So what is this idea of spec-driven validation? Spec-driven testing, you could call it that as well. It's basically what are the things we would want to do if we were just designing the role or the task benchmark by itself, independent of the agent. So we already talked about datasets. So the ground truth—having a bunch of examples of what good looks like—is one thing. So that's one component. Often we see customers that we work with also have rules. So if they've got a customer support agent, you want to say things like: don't ever give a discount more than ten percent. We don't allow refunds if it's more than thirty days past the purchase. And there are these rules. So your alarm bells should be going off a little bit already because how do you actually test for sure that a rule is never violated? It's pretty hard. Sometimes you also have ontologies or dictionaries that are relevant. So an example would be: if you're building an airline chatbot, that particular airline might only fly to certain destinations. So that's the relevant universe of things you need to think about. You may have internal terminology in your company that apply to your policies that no one else in the rest of the world actually knows about. So that's also part of the spec right? Because if you're going to actually build an agent, you will be building that into the agent. But if you're going to test it, you actually need to tell the testing system what these things are and what is a valid substitution. There's domain knowledge. So you may have very specific scientific, finance agents, other things that need to know what terms are substitutable. So if you, for example, do substitutions on something—you know, gross profit and gross sales—for example, if you're talking to an LLM generally, it might actually confuse those two terms. But in business, they're very different things. So this specific domain knowledge is relevant to testing as well. So you might have rights and roles. Like the agent may perform differently if you're logged in, if you're logged out, if you have certain rights and permissions and things like that. And then the last one, which is pretty important, is robustness requirements. So one is: I've got my test set that should work, right? But it needs to work under stress. So in vision, where we started out, it's things like: can I detect this runway for the plane to land on? But can I detect it at sunset, sunrise, under fog? And how much fog can there be? How much can the camera shake before the thing doesn't work? And that's actually similar in agents. You know, if you're building a customer-facing agent, could typos disrupt it? How many typos disrupt it? How frustrated will people get? Rephrasing—how stable under change are the results? And so really, this is the point here: we need to go beyond the test set to have task and role-specific benchmarks for the agent itself. And what do you then do with that? Maybe I already talked about some of these examples, but these are just examples of the kind of things you might have if you had a product support agent. So we've worked with quite a few people doing this. So you can think of it as: in LLM land, people have started to call it the "eval," the test set. Which makes sense, but I just think that the eval itself, we have to think of going beyond the eval as well. There's this concept of an "agent card" which comes from the A2A spec. It's been in other things around which describes what the agent does. It's also relevant here. And then obviously, there's all the context around this. And if you're a company deploying agents, you kind of want your tests to look like something that has these various elements that are relevant. That's a fair eval. And you want to build more and more of these tests. These look like integration tests if you're from an engineering perspective. Often some of these things are implicit, but you want to make them explicit. So what can you do with this? So what we do with this in our platform, we do two things. We do security checks. So we actually pull the specs that an agent is supposed to fulfill into security testing. Why do we do that? So generally, if you know what an agent is trying to do, you know the edges of where it's vulnerable. Because it's going to be willing to talk about those domains that it's supposed to act in, right? So that's actually where it's most likely to be vulnerable. Second, the tasks it performs—it will have more power to act in the infrastructure on those tasks. Like if it's a banking agent or something like that, it will be able to work in that area. So we pull things like this spec information in. And then the robustness side is: does it do its job properly? Especially the robustness side—can we vary the inputs and see how much of a range it has in terms of answering the questions properly? So we build a product to do this. But my point here is not to show the product. I think it's something that if you're testing agents in any context, using any infrastructure, trying to be explicit about the various bits that are on this slide and bringing that together is a useful thing to try to do. From an industry perspective, I think there are lots of things going on, but just calling out two: there are a lot of prompt management platforms that allow you to be fairly elaborate about why this test exists and things like this. This is all useful when you actually want to generate variants of the test because you want this context. As I said, from the A2A spec, you've got agent cards—they're quite long. But here's an example of a skill. You would also realize that even if you have this, that doesn't give you enough to actually evaluate the agent. You still want to know: what range of change could be valid? For maybe in this case, what kind of people could the meeting be booked for? And so on. I can talk a lot more about how hard it is to create variations within these sort of envelopes that a spec might create, but I think just in general, my point here is: as you think about evaluating agents, start thinking about not just the eval dataset or benchmark. Also think about the task and the context for the task and how you capture that. So hopefully we can make Marvin a little bit happier because he has the specs and he kind of knows what he's supposed to do. And then, yes, specify the behavior of your agents. That's the key thing to do here. Stay independent of the implementation because often you may be building a Langsmith or something, or Vertex Agents, or other things, but then later on you may change to a different infrastructure. You actually want to keep those integration tests, unit tests, and penetration tests that you can run independently. And this is also a way to close the loop. So part of our inspiration of thinking about what should go into a spec is: what would you need to actually run the agent, automatically get the results, and then start to iterate and try to fill the robustness gaps that have appeared. So it's a backyard type of RL. It's not proper RL because you're not doing it on the model, but you're kind of jury-rigging something around the outside. That's the key point. Where do we go from here? So we're obviously building product around this, but I've been in computer science for a long time. My last company did API infrastructure, so if you use OpenAPI spec, I apologize—it's part of my fault. So I helped write that spec way back in the day. So all about open. So we're thinking about how do you express these things in a way that you could just have in a GitHub repo, pull them into whatever tool you want to do, and then pull all the different pieces and just version all that stuff. So if anyone's interested in stuff like that, would love to chat. That's my talk. Come to our booth. We have a game you can play. If you play by four, you can win some of the Lego prizes up there. You need a bit of knowledge to be fair, or you need to be insanely lucky. But yes, that's my talk. Thanks a lot. And the other part is like what tools and tasks can it carry out in your infrastructure. So if it's able to wire millions of dollars to people that's obviously a lot more risky than if it can just answer questions and so on. So this is the balance that most people that you're basically looking for. But how do you actually define what goods looks like. I think it's pretty obvious that it's not just a data set of inputs and outputs that that are pretty good and then the rest is like guesswork. And it's also sometimes hard to define what harm looks like because maybe an agent doesn't do the right thing. But it's it's kind of just failing at the task and sometimes it's doing exactly the wrong thing was asking to do something bad. So what is this idea of spec driven validation? Spec driven testing you could call it that as well. It's basically what are the things we would want to do if we were just designing the role or the task benchmark by itself like independent of the agent. So we already talked about data sets. So the ground truth having a bunch of examples of what good looks like is one thing. So that's kind of one component. Often we see customers that we work with also have rules. So if they've got a customer support agent you want to say things like you know don't ever give a discount more than 10 percent. We don't allow refunds if if you know it's more than 30 days past the purchase and there are sort of these rules. So your alarm bells should be going off a little bit already because like how do you actually test for sure that a rule is never violated? It's pretty hard. Sometimes you also have ontologies or dictionaries that are relevant. So an example would be if you're building an airline chat bot that particular airline might only fly to certain destinations. So that's the relevant universe of things you need to think about. You may have internal terminology in your company that apply to your policies that no one else in the rest of the world actually knows about. So that's also part of the spec right because if you're going to actually build an agent you will be building that into the agent. But if you're going to test it you actually need to tell the testing system what these things are and what is a valid substitution. There's domain knowledge. So you may have very specific you know scientific finance agents other things that are that need to know what terms are substitutable. So if you for example if you do substitutions on something like you know gross profit and gross sales for example if you're sort of talking to an LLM generally it might actually confuse those two terms. But in business they're very different things. So this specific domain knowledge is relevant to testing as well. So you might have rights and roles like the agent may perform differently if you're logged in if you're logged out if you have certain rights and permissions and things like that. And then the last one which is pretty important is robustness requirements. So one is I've got my test set that should work right. But it needs to work under stress. So in vision where we started out it's things like can I detect this runway for the plane to land on. But can I detect it at sunset sunrise under fog like and how much fog that can there be? How much can the camera shake before the thing doesn't work? And that's actually similar in in agents. You know if you're building a customer facing agent could typos disrupt it. How many typos disrupt it? Like how frustrated will people get? Rephrasing how how stable under change are the results? And so really this is the point here is we need to go beyond the test set to have like task and role specific benchmarks that are for the agent itself. And what do you then do with that? Maybe I already talked about some of these examples. But these are just examples of the kind of things if you had a product support agent. So we've worked with quite a few people doing this. So you can kind of think of it as there's an in in LLM land people have started to call the eval kind of the test set. Which sort of makes sense but I just think that the eval itself we have to think of going beyond the eval as well. There's this concept of an agent card which comes from the A to A spec. It's been in other things as around which describes what the agent does. It's also relevant here. And then obviously there's all the context around this. And if you're a company deploying agents you kind of want your your tests to look like something that has these various elements that are relevant. That's a fair eval. And you want to build more and more these tests. These look like integration tests if you're from an engineering perspective. Often some of these things are implicit but you want to make them explicit. So what do we what can you do with this? So what we do with this in our platform we do two things. We do security checks. So we actually pull the the specs that an agent is supposed to fulfill into security testing. Why do we do that? So generally if you know what an agent is trying to do you know the edges of where it's vulnerable. Because it's going to be willing to talk about those domains that it's supposed to act in. Right? So that's actually where it's most likely to be vulnerable. Second the tasks it performs it will have more power to act in the infrastructure on those tasks. Like if it's a banking agent or something like that it will be able to work in that area. So we that's a place you can pull things like this spec information in. And then the robustness side is like does it do its job properly. Especially the robustness side can we vary the inputs and see how how much of a range it has in terms of answering the questions properly. So we build a product to do this. But my my point here is not to show the product. I think it's just something if you're testing agents in any context using any infrastructure trying to like be explicit about the various bits of that are on this slide and bringing that together is a useful thing to try to do. From an industry perspective I think there's lots of things going on but just calling out two. I mean there are there are a lot of prompt management platforms that allow you to be fairly elaborate about why this test exists and things like this is all useful when you actually want to generate variants of the test because you want this context. As I said from the A2A spec you've got agent cards they're quite long but here's an example of a skill. You would also realize that even if you have this that doesn't give you enough to actually evaluate the agent you still want to know well what what range of change could be is valid for maybe in this case what kind of people could could the meeting be booked for and so on. I can talk a lot more about how hard it is to create variations within these sort of envelopes that a spec might create but I think just in general my point here is like as you think about evaluating agents start thinking about not just the eval data set or benchmark also think about the task and the context for the task and how you how you capture that. So hopefully we can make Marvin a little bit happier because he has the specs and he kind of knows what he's supposed to do. And then yeah specify the behavior of your agents that's kind of the key thing to do here. Stay independent of the implementation because often you may you know you may be building a Langsmith or something or vertex agents or or or so on but then later on you may change to a different infrastructure. You actually want to keep those integration tests a little unit tests and penetration tests that you can run them independently and this is also a way to close the loop. So part of our inspiration of thinking about what should go into a spec is like what would you need to actually run the agent automatically get the results and then start to iterate and try to fill the robustness gaps that have appeared. So it's like a backyard type of RL. It's not proper RL because you're not doing it on the model but you're kind of like jury rigging something around the outside. That's the key point. Where do we go from here? So we're obviously building product around this but I've I've been in computer science for a long time. My last company we did API infrastructure so if you use open API spec I'm up I apologize it's part of me my fault. So I helped write that spec way back in the day. So all about open. So we're thinking about like how do you express these things in a way that you could just have in a GitHub repo pull them into whatever tool you want to do and then pull all the different pieces and kind of just version the hell out of that stuff. So if anyone's interested in stuff love in that, love to chat. That's my talk. Come to our booth. We have a game you can play. If you play by four, you can win some of the Lego prizes up there. You need a bit of knowledge to be fair or you need to be insanely lucky. But yeah, that's my talk. Thanks a lot.