Dark Factory: OpenClaw Ships Faster Than You Can Read the Diff — Vincent Koc, Comet ML
Description
Static benchmarks made sense for static software. Agents that adapt to users, rewrite their own harnesses, and shift behavior over time break that assumption. This talk is about what evaluation looks like when the system you're measuring keeps changing underneath you. Vincent Koc traces the arc from prompt engineering to context engineering to intent engineering, where agents self-optimize toward what users actually want. The eval problem compounds at each step: production traces reveal behavioral drift, test suites go stale, and the 20% of edge cases that break your product rarely show up in handcrafted datasets. The alternative he proposes: define the end state, let agents curate their own suites from traces, and treat evals as a living system rather than a point-in-time snapshot. Speaker info: - https://x.com/vincent_koc
Summary
Generated by claude-haiku-4-5-20251001Dark Factory: OpenClaw Ships Faster Than You Can Read the Diff
Main Topics
- Evolution of AI Evaluation Methodologies: From static benchmarks to adaptive, intent-based systems
- The Agentic AI Paradigm Shift: How AI applications are becoming self-optimizing and malleable
- OpenClaw and Harness Technology: Self-adapting systems that evolve with application changes
- Intent Engineering: Machines optimizing based on user intent rather than static rules
- The Gap in Current Evaluation Practices: Missing chaos engineering and observability in AI/ML
Key Points
The Problem with Static Evals
- Current AI evaluation relies heavily on static benchmarks and hand-crafted test cases
- Traditional approach: create examples → offline evaluation → deploy → hope nothing breaks
- Missing the "chaos engineering" phase present in software engineering practices
- AI applications are dynamic, but we treat them as static software
Evolution of AI Development Approaches
| Phase | Characteristic | Limitation |
|-------|---|---|
| Prompt Engineering | Random word-smithing; "bash words and hope" | Unscientific, luck-based |
| Context Engineering | RAG, tool calling, more structure | Better but still incomplete |
| Intent Engineering | Self-optimizing based on user intent | Requires new evaluation frameworks |
The Case for Adaptive Evaluations
- Benchmarks should change with applications, not remain static
- Agent behavior varies by user, context, and customer base
- Traditional metrics (one plus one equals two) don't capture modern AI complexity
- Need to measure ambiguity and personality in agents
New Evaluation Approaches
- Self-Curated Test Suites: Agents themselves generate test cases from traces
- Always-On Evaluation: Continuous monitoring and optimization rather than point-in-time testing
- Telemetry in the Loop: Harnesses aware of costs, errors, and performance can self-correct
- Trace-Based Learning: Monitor what agents actually do, adapt when patterns change
Notable Quotes
> "Anything we do in technology, anything on the edge is going to be janky, it's going to be weird. And that's fun in my opinion."
> "Our AI applications are not static, but we're treating them as static software."
> "Why are benchmarks static? Why don't we test in a more adaptive manner?"
> "Code is cheap. Tokens are there. The velocity of creating software and applications increases."
> "It's that 20% that's going to mess up your business... someone who's going to come and ask a weird question."
> "Our evals don't become the data set or the starting point, our evals become what is the end state that we want to get to?"
> "People need to start looking at evals not as this static data set thing but actually as code as software or as a living agent."
Takeaways
- Paradigm Shift Required: Move from static benchmarks to dynamic, self-adapting evaluation systems
- Evals as Living Code: Treat evaluations as software that evolves with your applications, not fixed data sets
- Intent Over Examples: Define desired end states and outcomes rather than hardcoding test cases
- Embrace Chaos Engineering: Actively test edge cases and unexpected usage patterns (the problematic 20%)
- Always-On Monitoring: Implement continuous evaluation and self-correction mechanisms
- Agent-Driven Testing: Let agents themselves learn from traces and generate relevant test cases
- Observability Critical: Understanding what happens inside agentic applications is more important than ever, especially as they become more self-optimizing
Key Research Concepts
- Eval Calcification Problem: The growing difficulty of maintaining evaluations unless we adopt smarter approaches
- Karpathy's Auto-Research: Auto-optimization framework that can be applied to agent tuning
- MCP Tools: Modular components that can be tested independently within larger agentic systems
- OpenClaw: Reference system demonstrating harnesses that self-adapt and shift
Transcript
[SPEAKER_00] Cool. Hey everyone. Thanks for joining this session. Sorry if my sound is a little croaky. I've done three talks back to back. So one on Wednesday, one yesterday keynote and then workshop style session today. So I'm Vincent. I'm going to be talking about malleable evals from static AI measuring to adaptive systems. Now, let's jump into who I am, what I do. I call myself the friendly canker. I use AI, I use technology. I'm always on the edge. For those of you that haven't seen my keynote, I do, yeah, I just live on the edge and do some fun stuff. So this is me using VR goggles back in 2013 when people hadn't even heard of VR. It came with a warning label. It said I only use it for five minutes, I use it for three hours, then I vomited for three hours after that. So measurement, anything we do in technology, anything on the edge is going to be janky, it's going to be weird. And that's fun in my opinion. Now, whenever we talk about evals to people and a little bit of pretext, my role at Comet, I work in evals. I do eval research. I work with universities. We benchmark and run evals for large sets of companies and organizations, everything from Uber to Netflix to banks even in the UK. But the thing that's been going on right now is that this joke that evals is a little bit dead. And it's a little bit of a joke, but there's a little bit of truth to it as well. And I'm going to hopefully walk you through the mindset shift and hopefully explain a little bit less about evals, but what's actually happening in the agentic AI space and then how do we translate that back to evals. So when we think about software engineering as a practice, when we're thinking about how do we measure things, we kind of look at it from the sense that we're going to start with this thing is meant to do something. So we would start with a set of examples or some unit tests. We might do a manual regression suite, which is like, hey, when we do A and B, sometimes C happens and C is unfavorable. We don't do that. We don't make people vomit when they put their VR goggles on. We could do things with CI/CD pipelines, make sure that the thing ships out and works the way it's meant to and intended to. But mostly in engineering we do things known as chaos engineering and observability. For those of you that are unfamiliar with the term chaos engineering, it's basically where you're doing all kinds of random stuff and breaking it and having fun with the technology and seeing where you can stretch it and where you can go. Now, when we apply this to AI and data science space, as we traditionally know it in the last little while, 2025 included, we do things with static benchmarks, evaluations. I'll give an example. How compliant is my AI in risk? I'm going to ask it a bunch of questions and make sure it doesn't talk about selling me some financial services because that's a big no-no. We would then handcraft a set of questions and examples and sit there and tune this thing up, make sure it's absolutely perfect. Before we deploy the AI system or the models, we'll do some offline evaluation where we're cycling through those tests. But we're missing that chaos engineering space. We're missing what comes next and how do we mess up with it and how do we know where we can stretch this thing? And I think that's an honest gap that we see in this space. And that's why we're just hyper fixated on benchmarks and evaluations. If you go to any conference in the academic space, all people talk about is benchmarks. I created a benchmark for adding numbers and what LLMs think about it. It's great. But how is this actually helping me? So then you end up with a huge set of datasets to try and explain what is happening with your agent until something goes wrong. And it's a matter of time before something goes wrong and it will. And you're back to the drawing board trying to figure out what's going on. And the reason for that is that our AI applications are not static, but we're treating them as static software. Yes, when we ship software, we might change unit tests. They're a little quicker to do. But realistically speaking, even software is becoming malleable. So flipping to my keynote I gave yesterday, where I'm one of the core contributors of something called OpenClaw. The harness changes itself, the harness will shift, you want to create skills, you want to do other things, it will adapt. So that adaptation that we're seeing inside the things where software is being shipped at lightning speed. How does your benchmarks keep up with that? How does your benchmarks adapt to that space? This is one of many papers that are out there. I don't remember when this one was published, but this concept of adaptive testing for LLM evals. This concept is somewhat revolutionary, but what happens if our benchmarks would change with our applications? I didn't write this, but great that someone did. But it poses the question of why are benchmarks static? Why don't we test in a more adaptive manner? So this could be great. This is more selectively testing and being a bit more smart about how we test. But it's still taking us on that journey. I think it's a mindset shift. Now, rewind to what we're seeing in the AI space for a minute. We had prompt engineering, if we focus purely on LLM space. We had this prompt engineering world where I'm going to doom scroll and wordsmith instructions. I'm going to bash random words into an AI and hope it improves. So if I'm building this banking app or creative app, I'm going to stick all kinds of random words and see what comes out the other end and makes it creative. It's a little akin and I'm not trying to downplay medicine in any way. It's hey, I'm going to make medication for, I don't know, liver disease and turns out it cures pain. Okay, these are painkillers now. That's great. And the same thing we're doing with prompt engineering. We're just bashing words into it and hoping it changes. And for some reason this died in 2023, but people still do it. It's intense. And then we went into this world of context engineering. I think this started making evals a little bit more relevant because it was a little bit more complicated. There were steps involved. There was data coming in and search was a thing. And we're starting to steer the agents in a direction with things like RAG and tool calling. And the beautiful part is there. It's like, okay, I'm an organization with this big agentic system. Maybe I can break this agent up into its parts. Maybe I can say, oh, I have this MCP tool that does some sales agent thing. I can test that thing is doing what it's meant to. And then we went into this world of context engineering. I think this started making evals a little bit more relevant because it was more complicated. There were steps involved. There was data coming in and search was a thing. And we're starting to steer the agents in a direction with things like RAG and tool calling. And the beautiful part is there. It's like, well, okay, I'm an organization of this big agentic system. Maybe I can break this agent up into its parts. Maybe I can go, oh, I have this MCP tool that does some sales agent thing. I can test that thing is doing what it's meant to. Right. I can just go off and be sure that that thing is happening. So this process of tool calling and breaking this larger agentic piece up into its sum of its parts made evaluation somewhat steerable and a little bit more understandable, but it still didn't quite hit the head on its head. But then in 2025, where are we going next? I mean, if you look around, we could see that code is cheap. That's not a changing thing. Tokens available, debatable if you think tokens are cheap or not, but they're there, which makes tokens cheap. And then tokens become fast food. Essentially, we can consume more tokens. Therefore, we generate more. The velocity of creating software and applications increases. And models become really good. I think this is the thing that I did. I think a lot of people just have not yet comprehended that a lot of the AI applications that are now running, the models can do absolutely amazing things. I've been working on a lot of optimization problems and we can take these models that are somewhat seen as generic systems and be able to do amazing things like solve RKGI 2, which is puzzles. And if I recommend anyone who's interested in evals, look at the RKGI 2 puzzles. I've tried the RKGI 3. Some of these puzzles are really hard for humans to solve, but machines can pattern recognize and then an LLM can actually pattern recognize and start to solve those. So what that brings us to is intent engineering. This concept that machines can self-optimize based on intent. Right. And we're seeing this with the harnesses that we're seeing coming out where we've got this with Open Claw, but we're also getting this with other types of harnesses inside Claude and Codex where it's trying to understand you and it's trying to adapt to you and give you a better experience. Now the problem with this is that when we have intent for machines, the evaluations become even more complicated because it's like, how do I know my experience is different from your experience and different from someone else's experience? How do we start to build testing around this sort of methodology and understanding? And I think the complicated part of this is that it exacerbates the need for evaluation even more. There was this joke, like I was saying earlier, people saying, oh, evaluations are dead. They're going to go away. Observability is dead. They're going to go away. But realistically, now more than ever, people want to know what's happening inside of these agentic applications within their different layers because then it gives them some understanding of what's going on. We use words like, oh, these agents are insecure or we're not sure what's happening. So how do we actually turn that into something meaningful? So going back to my earlier slides on this concept, let me just recap where I was for everyone. We said there was this static benchmarks. We hand-created evaluations. We would do these offline evaluations and we had this big gap. So we're actually moving towards this intent-based outcome. So if we think about this concept of intent engineering that I'm mentioning, how do we actually map that to something? So instead of saying, one plus one equals two or users asking this very specific question, this is the answer and this is what we're going to compare towards. It's how do we define ambiguity in an agent? How do we define personality inside of an agent? And how does that look like for an organization? And some of the research is showing things like we can build rubrics. We can do it how we evaluate our pictures and things like that in schools. We can self-curate sweets from traces as in not me, but the agent can. Once we start tracing these applications, let's say 80% of the time it's the same stuff that's happened with my agent. But now suddenly my customer base has changed and because my customers have changed, they're going to start looking at things. They're going to start asking questions differently. Things are going to start changing inside of my agent. But why are we not measuring this? Why are we not taking these traces and feeding them into agents and going something has changed and then telling that to the user or telling that to the owners of these agents and changing the suites, the tests. We can do online always-on evaluation optimization. So to that point, once we start looking at the traces, once we have agents doing the evals, not static benchmarks, we can have this as an always-on sort of service. And then lastly, we can do this sort of telemetry in the loop. I have a paper that's been written on this, which is essentially when we're writing software applications or MCPs or anything like that, agentic systems, if the harness is aware of the telemetry, it's aware of what's breaking. It's aware of how much it's costing and you can set some conditions around it. It can kind of self-correct itself. So we're starting to see this with harnesses where it's had an error, it's had an issue, it's going to fix itself, it's going to continue on. So I think this is a case of instead of trying to predict what's gone wrong, how can we be more smart about using that data back into the agent to be able to make it heal itself to some degree. So I'm calling this an eval calcification problem. I'm still stirring it on my head. Sounds like a really nice paper title. But this idea that it's just going to become harder and harder unless we can get smart about it. And I think one of the concepts I want you to chew on and think about this is one of the auto-research outputs that you can do if you haven't tried it, Karpathy's auto-research. This really basic sort of auto-optimization using Python, you set a goal, you set a target, and it kind of tunes itself, it tweaks itself. You could do this with absolutely anything, right? You could do this with, I don't know, what's the best mix for what's the tastiest barbecue or the cheapest barbecue mix that you want to make? It could be anything, right? You just set a reward signal. But the key here is that your users are going to have a point of intent that you want to optimize towards. And then how do we get the machine to correct itself and kind of look towards that as an eval? So then our evals don't become the data set or the starting point. Our evals become what is the end state that we want to get to? And then we just let the machines do the work. We have evaluations where it's just the agent and we're just defining the end state. So just one last thing to bake into your minds before I'm going to finish a little early so we can do questions or anything else is fine. It could be anything, right? You just set a reward signal. But the key here is that your users are going to have a point of intent that you want to optimize towards. And then how do we get the machine to correct itself and look towards that as an eval? So then our evals don't become the data set or the starting point, our evals become what is the end state that we want to get to? And then we just let the machines do the work. We have evaluations where it's just the agent and we're just defining the end state. So just one last thing to bake into your minds before I'm going to finish a little early so we can do questions or anything else is fine. Is that you can imagine this space where you go 80% is the static stuff like it's been defined in an intentful manner. But that 20% is always going to keep changing and it's that 20% that's going to mess up your business. It's going to be someone who's going to come and ask a weird question or use your agent in a really strange way and it's going to be absolute hell for you. So how do you create agents to manage and maintain that 20% and keep an eye on it and then adapt and change your evals? So I think people need to start looking at the evals not as this static data set thing but actually as code as software or as a living agent. Not as a point in time but as a self optimizing growing solution. That's more or less it. I was going to present more of an in-depth demonstration of this where we've applied it at Comment but the end state is not quite finished yet and it will be over the next coming weeks.