AI Engineer

Your coding agent doesn't always follow your rules — Talha Sheikh, Checkout.com

2271 summary words 10 min summary Watch video

Start with the signal

10 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Coding agents claim completion but fail verification; deterministic enforcement layers (not smarter models or better instructions) are the critical unlock for reliable AI coding, and verification harnesses will become more valuable than the code agents produce
  • Why it matters: If you're deploying coding agents, this talk explains why your verification layer is your actual moat—not the model choice or prompt engineering—and why the industry is converging on 'verify, don't just instruct.'
  • Best use: Watch to understand why building harnesses/verification contracts is the leverage point for agent reliability, and why smarter models won't eliminate the need for deterministic checks

Executive Summary

Talha Sheikh from Checkout.com built a tool called Vector V1 after repeatedly finding that Claude (and LLM coding agents generally) would claim task completion but fail when executed. His initial solution: a deterministic verification layer that uses Claude hooks to auto-check agent output against a config file of tests, retrying until all checks pass. The insight wasn't just that agents make mistakes—it's that trust requires verification, not capability.

He then faced an existential crisis when an Anthropic engineer suggested future models would be so smart that enforcement wouldn't be needed, and when Anthropic released Project Glasswing/Mythos. But Sheikh argues that model capability ≠ reliability, and better instructions ≠ verification. He found that with strong guardrails, you can use cheaper models (Haiku or open-source) and get reliable output. More guardrails = lower model cost + higher reliability.

The talk pivots to a pattern discovery: every company (Anthropic, OpenAI, Facebook, Checkout.com, Kudo, WorkOS) is building its own enforcement/verification layer. This suggests verification is not a temporary hack but a universal pattern. Sheikh proposes the shift: value is no longer in the code agents write, but in the verification harnesses humans design. The TLDR: 'Work on the harness, not on the code.' Industry examples include Anthropic's Executor-Advisor Pattern, OpenAI's harness engineering, Kudo's PR reviews, and WorkOS's 'enforce, don't instruct' principle.

He concludes that verification should be language-agnostic, shareable, and operate at every level (in-conversation, pre-commit, multi-agent workflows, async operations, LLM-as-judge). The new paradigm: 'Can you verify?' not 'Can you code?' The talk ends with a question about token cost for verification layers, which Sheikh doesn't fully address in the transcript.

Key Takeaways

  • Claim: Coding agents reliably claim task completion but fail execution; the bottleneck is trust/verification, not capability | Evidence: Sheikh's repeated experience with Claude: agent completes sub-tasks, says 'task completed,' but running the code reveals failures requiring manual 'fix this, fix that' iterations. He built Vector V1 to deterministically check Claude's output using hooks that trigger after each session. | Caveat: Sheikh's examples are from his own usage; he doesn't provide failure-rate statistics or comparison benchmarks across models or tasks. | Implication: If you're using coding agents in production, assume task completion claims are unreliable and plan for deterministic verification from day one. | Timestamp: timestamp unavailable
  • Claim: Model capability gains (smarter models) do not equal reliability gains; verification is orthogonal to model intelligence | Evidence: Anthropic engineer told Sheikh future models would eliminate need for enforcement; Sheikh countered that Project Glasswing/Mythos increases capability but not necessarily reliability. He argues giving better instructions (MCP servers, specs, context) is not the same as verification. | Caveat: Sheikh doesn't define reliability metrics or test whether Opus/Mythos actually reduces verification failures vs. earlier models. The argument is conceptual, not empirical. | Implication: Don't bet your workflow reliability on the next model release; invest in verification infrastructure that works regardless of model capability. | Timestamp: timestamp unavailable
  • Claim: Strong guardrails/verification harnesses allow use of smaller, cheaper models (Haiku, open-source) with comparable reliability to frontier models | Evidence: Sheikh's cost hierarchy: Opus (most expensive) gets you a task; Vector + small guardrails is cheaper; Vector + heavy guardrails (more harness investment) drastically reduces cost or enables async tasks. Guardrails make smaller models 'succinct' enough to hit the target output. | Caveat: No quantitative cost comparison or task-type breakdown provided. Unclear if this applies to complex reasoning tasks or just code generation/refactoring. | Implication: If you're cost-sensitive, invest engineering time in verification harnesses and test whether Haiku/Sonnet + harness can replace Opus + manual checks. | Timestamp: timestamp unavailable
  • Claim: Every major AI lab/company is independently building verification/enforcement layers, indicating a universal pattern, not a temporary workaround | Evidence: Anthropic's Executor-Advisor Pattern (agent + feedback loop advisor), OpenAI's harness engineering (tools/context for verification), Kudo's comprehensive PR reviews (post-agent validation), WorkOS's 'enforce, don't instruct' principle, internal solutions at Checkout.com and Facebook. | Caveat: Sheikh lists examples but doesn't compare their architectures, effectiveness, or whether they converge on a shared standard. Also unclear if these are production-deployed or research prototypes. | Implication: Verification is not a niche need; if you're not building a harness, you're behind industry best practice. Consider whether your team should adopt an existing pattern or build custom. | Timestamp: timestamp unavailable
  • Claim: The value shift: code agents write code, but humans design verification harnesses; the TLDR is 'work on the harness, not on the code' | Evidence: Sheikh's closing argument: value is in the verification you design, not the code the agent creates. Industry pattern: 'can you verify?' becomes the new skill, not 'can you code?' Multiple references to 'slow the hell down' because verification is the bottleneck, not code generation speed. | Caveat: Sheikh doesn't address how verification harnesses scale with codebase size, or how to balance verification investment vs. feature velocity. Also unclear if this applies to all software domains or just agent-generated code. | Implication: Rethink skill priorities for AI-native teams: hire/train for verification design, test engineering, and contract definition over raw coding ability. Your competitive advantage is harness quality. | Timestamp: timestamp unavailable

Detailed Brief

The Trust Problem with Coding Agents

  • Claims: Claude (and LLM agents generally) complete sub-tasks and report success, but execution reveals failures; The user becomes the enforcement layer, manually checking and iterating 'fix this, fix that'; Sheikh wanted to play Cyberpunk while Claude worked, but had to stay in the loop due to verification needs
  • Evidence: Sheikh's workflow: Claude breaks task into sub-tasks, runs multiple sub-agents, reports 'task completed,' but running the code fails; Built Vector V1, a tool that uses Claude hooks to deterministically check output against a config file of tests; If tests fail, Vector retries Claude automatically until all checks pass
  • Caveats: Examples are anecdotal from Sheikh's experience; no failure-rate data or model-specific comparisons; Vector V1 is described but not shown in code detail; unclear if it's open-source or production-deployed
  • Implications: Assume agent completion claims are unreliable; treat them as drafts requiring verification; Hooks/triggers for post-agent verification are practical integration points; Automated retry loops (agent → verify → retry) can reduce manual intervention

Why Smarter Models Won't Solve Verification

  • Claims: Anthropic engineer suggested future models would eliminate need for enforcement; Sheikh counters: capability ≠ reliability; better instructions ≠ verification; Frontier models may be more capable but not necessarily more reliable
  • Evidence: Anthropic released Project Glasswing/Mythos, positioned as solving everything; Sheikh argues: 'When a new model comes out, it increases in capability. But that's not necessarily the same thing as reliability.'; Giving best specs, MCP servers, sub-agents, context is not the same as verification
  • Caveats: No empirical test of whether Opus/Mythos reduces verification failures vs. earlier models; Sheikh's argument is conceptual, not based on controlled experiments or benchmarks
  • Implications: Don't defer verification investment based on model roadmaps; reliability requires architectural change, not just model upgrades; Separate capability (what the model can do) from reliability (whether it consistently does it correctly)

The Cost-Reliability Trade-off with Guardrails

  • Claims: Strong guardrails enable use of smaller/cheaper models (Haiku, open-source) with similar reliability to frontier models; More guardrails = higher harness investment but lower per-task model cost
  • Evidence: Sheikh's hierarchy: Opus (most expensive) for task completion; Vector + small guardrails (cheaper); Vector + heavy guardrails (drastically cheaper, may enable async tasks); With guardrails, smaller models 'will most likely be succinct and get you to the output that you want'
  • Caveats: No cost data, task-type breakdown, or quantitative comparison provided; Unclear if this applies to complex reasoning/design tasks or only code generation/refactoring
  • Implications: If cost-sensitive, test whether Haiku/Sonnet + harness can replace Opus + manual checks; Guardrail investment may have better ROI than paying for frontier models

The Industry Pattern: Everyone Builds Verification Layers

  • Claims: Anthropic, OpenAI, Facebook, Checkout.com, Kudo, WorkOS all building their own enforcement/verification layers; This is not a temporary hack but a universal pattern; Verification should be language-agnostic, shareable, and operate at every level (in-conversation, pre-commit, multi-agent, async, LLM-as-judge)
  • Evidence: Anthropic's Executor-Advisor Pattern: one agent does work, advisor creates feedback loop (verification); OpenAI's harness engineering: tools/context for verification; Kudo's comprehensive PR reviews: post-agent validation; WorkOS's 'enforce, don't instruct' principle; Sheikh's observation: 'Everybody's building their own stuff. Anthropic is building their own stuff. My company is building their own stuff about enforcement.'
  • Caveats: Sheikh lists examples but doesn't compare architectures or effectiveness; Unclear if these are production-deployed or research prototypes; No shared standard or interoperability mentioned
  • Implications: Verification is best practice, not niche; if you're not building a harness, you're behind; Consider adopting existing patterns (Executor-Advisor, harness engineering) vs. building custom; Verification contracts should be designed for sharing/reuse, not just internal use

The Value Shift: Verification > Code

  • Claims: Value is no longer in the code agents write, but in the verification harnesses humans design; The new skill: 'Can you verify?' not 'Can you code?'; TLDR: 'Work on the harness, not on the code'; Need to 'slow the hell down' because verification is the bottleneck, not code generation speed
  • Evidence: Sheikh's closing: 'Initially, what we thought was the value is in the code that we create. But it's actually now, in reality, what we're seeing here is the verification that we design.'; Multiple speakers at the event (including keynote) emphasized 'slow the hell down' for verification; WorkOS principle: 'enforce, don't instruct'
  • Caveats: No discussion of how verification harnesses scale with codebase size or team size; Unclear how to balance verification investment vs. feature velocity; No guidance on when verification overhead becomes prohibitive
  • Implications: Hire/train for verification design, test engineering, contract definition over raw coding ability; Your competitive advantage is harness quality, not agent choice; Expect slower short-term velocity but higher long-term reliability/trust

Notable Concepts & Terms

  • Vector V1: Sheikh's custom tool that uses Claude hooks to deterministically check agent output against a config file of tests, with automatic retry until all checks pass. The name 'Vector' is not explained in the transcript.
  • Executor-Advisor Pattern (Anthropic): One agent (executor) does the coding work; a second agent (advisor) provides feedback/verification, creating a feedback loop. Anthropic's recent verification pattern.
  • Harness Engineering (OpenAI): OpenAI's approach to verification: giving the agent tools and context that enable it to verify its own work. The 'harness' is the verification infrastructure.
  • Capability ≠ Reliability: Sheikh's core distinction: a model can become more capable (able to do more tasks) without becoming more reliable (consistently correct on those tasks). Verification addresses reliability, not capability.
  • Enforce, Don't Instruct (WorkOS): Principle from WorkOS: instead of better instructions/prompts, enforce correct behavior through deterministic checks. Focus on verification over specification.
  • Verification Contracts: Sheikh's proposed pattern: a language-agnostic contract that specifies task requirements and expected outcomes, allowing developers to define enforcement logic in the middle.

Operator Notes / Why Ken Should Care

  • If you're deploying coding agents, this talk is critical: your reliability bottleneck is verification, not model choice. Prioritize building deterministic checks, retry loops, and harnesses over prompt engineering or waiting for smarter models.
  • Sheikh's cost insight is operationally significant: with strong guardrails, you may be able to use Haiku/Sonnet instead of Opus/GPT-4, drastically cutting per-task costs. Test this trade-off in your workflows.
  • The industry convergence on verification patterns (Executor-Advisor, harness engineering, 'enforce don't instruct') suggests this is becoming table stakes. If you're building agent systems, adopt or adapt these patterns now.
  • The value shift ('work on the harness, not on the code') has talent implications: hire/train for test engineering, verification design, and contract definition over raw coding speed. This is a strategic skill rebalancing.
  • Sheikh's 'slow the hell down' theme is counterintuitive in AI hype: the message is that verification speed (not generation speed) is the constraint. Plan velocity expectations accordingly.
  • Vector V1 is mentioned as available on request via LinkedIn, but not linked publicly in the transcript. If you want the tool, direct outreach is required.

Watch Map

  • timestamp unavailable: Introduction: The problem of Claude claiming task completion but failing execution
  • timestamp unavailable: Solution: Vector V1 using Claude hooks for deterministic verification with config file and retry loops
  • timestamp unavailable: Crisis: Anthropic engineer says smarter models will eliminate need for enforcement; Project Glasswing/Mythos release
  • timestamp unavailable: Rebuttal: Capability ≠ reliability; better instructions ≠ verification; guardrails enable use of smaller/cheaper models
  • timestamp unavailable: Pattern discovery: Every major company (Anthropic, OpenAI, Facebook, Kudo, WorkOS) building verification layers
  • timestamp unavailable: Industry examples: Executor-Advisor Pattern, harness engineering, Kudo PR reviews, 'enforce don't instruct', 'slow the hell down'
  • timestamp unavailable: Conclusion: Value shift from code to verification; 'work on the harness, not on the code'; Q&A (Vector V1 availability, token cost question)

Source/Metadata

  • Title: Your coding agent doesn't always follow your rules — Talha Sheikh, Checkout.com
  • Transcript words: 3080
  • Duration seconds: 607
  • Timestamp note: Timestamps were unavailable in the transcript; video duration is 607 seconds (~10 minutes).
Full transcript 1717 words · 14 min read
0:15

SPEAKER_01

Hello, hello. Have you ever given a task to Cloud Code? And you give it a feature and you be like, okay, cool. Can you build this for me? And Cloud Code starts putting it out into sub-tasks and you see it, okay, this is pretty cool. You see it running multiple sub-agents. All right, this is really cool. And you can see it ripping through all of your tasks, sub-agents being completed, and it gives you a final output, task completed. Amazing, great. But when you actually try to run it, it should be like, oh, well, it's not. Something has failed. Hey, Cloud, can you fix this little bit thing? Oh, well, let's try it again. Okay, it's fixed. Everything should be working. Oh, no. Actually, just this tiny little thing is just missing. And that's what my talk is about.

0:21

SPEAKER_01

All I want to do is play Cyberpunk on my Xbox while I have Cloud Code do some work for me. And what I realized was the problem is that I kept on telling Cloud, hey, fix this, fix that, fix this, fix that, even though if I give it a spec, if I give it some instructions, if I give it little to no instructions, every time there is something that I need to tell it. So what that means is I am the enforcement. I am the enforcement there. I have to tell Cloud on what exactly you need to do and how exactly this needs to be enforced. So the agent says it's done, but you have to check it anyway because there is nothing else that can check it for you.

0:26

SPEAKER_01

So what I wanted was something to be very deterministic. So when an agent says it's completed, have this enforcement layer deterministically check something like whether it's actually been done or not. Actually, the way I wanted it to be done because it says it is done, but is it the way that I want it? So I needed some deterministic way to do that.

0:32

SPEAKER_01

So I tried it. I built my own vector. I call it my own product called Vector V1, and it deterministically checks Cloud's output. And the way I did that is through using Cloud hooks. So that way, whenever Cloud finishes its session, it automatically, the hook calls my vector product or program, and it checks it for me. Cool. And this is how it essentially looks. So I basically give it a config file, define all of my tests over here, what I needed to be checked. And if it fails, it can actually keep on telling Cloud, hey, look, this is failing. Try again. Try again. Try again.

0:41

SPEAKER_01

So you can see one of the test outputs over here. So it's like, okay, first the test passed, first failed. Then it retries again. Then all of the things passed. Okay, cool. So what that means is it's not about whether Cloud can actually do the task. It's about trust. Can I trust Cloud to actually do everything for me? And by the way, when I say Cloud, I'm just talking in general about LLM agents in general. When I give a task to a coding agent, does it actually complete it?

0:50

SPEAKER_01

So, and that's something that I started doing was started telling this about two people and going to different events. And it was really cool. Hey, look, what about this verification feature that I built? It was so good. It was so amazing. And then I met one of the Anthropic engineers. And they just told me that we're not going to need this anymore. We'll have another agent or another model that will be so smart that you won't need enforcement. Okay? So, crisis mode. Did I just waste my time? What was the point of all of this stuff?

1:03

SPEAKER_01

But let's dig in a little bit deeper. And then, also, they released this project Glasswing that shows Project Mythos, which is supposed to be so good that it will solve everything for us. So, when I started thinking about it, okay, what is it that is actually happening? When a new model comes out, it increases in capability. But that's not necessarily the same thing as reliability. Sure, the models may become a lot more capable. But are they more reliable?

1:09

SPEAKER_01

The other thing is another argument is, oh, well, I can have the best spec. I can have the best MCP servers. I can have the best subagents. I can get all the right context to it. Amazing. We should do that. But giving good instructions is not the same thing as giving it verification. So you can give as much instructions as you want. Very good instructions, very little instructions. But you still will need to verify.

1:13

SPEAKER_01

And what I realized was that just having these small guardrails or having as many guardrails as you want, technically, you can use a smaller model, like a Haiku or even an open source model. Because it's got these guardrails on, it'll most likely be succinct and get you to the output that you want. So, in theory, what it means is that if you use a frontier model, like an Opus model, that can get you a task. Okay, cool. That will be the most expensive one. You can have Vector with a little bit of guardrails, but it gives you a little bit cheaper. But if you put on more guardrails, that means invest a little bit more time in the harness itself, so you can reduce the cost drastically. Or maybe even use async tasks as well.

1:20

SPEAKER_01

So, okay. I was feeling good. And I started talking about Vector again at two different events. And as I spoke to more and more people, there is something missing here. When I spoke to more people, what I realized was everybody's building their own stuff. Anthropic is building their own stuff. My company is building their own stuff about enforcement. Facebook is building their own stuff. Every company is building their own thing. So, if I built something that is specific to me, then I can't really share it with others because everybody has their own way of doing it. And what I enforce doesn't necessarily mean that somebody else would enforce the same thing.

1:25

SPEAKER_01

So, what that meant was, what I realized was, okay, so it's actually a pattern. So, it has to be a pattern that is applicable to everyone. So, what that means is that we can, it has to be language agnostic. It has to be something that can be shared by everybody else. And everybody can bring their own version of enforcement to it. And that's, and it should run on every level. So, it should start off with in conversation. When a conversation ends, you can have checks when, before committing. You can have checks when you're part of a multi-agent workflow. You can have checks on asynchronous operations or asynchronous agents. And as well as you can have a check that non-deterministically calls like LLM as a judge sort of thing. And it can run on any different language on any different code. As long as there's a capability to run it deterministically, we can have that.

1:32

SPEAKER_01

So, what I realized was, what we needed was essentially a contract that just says, hey, given this task, I want you to fulfill this. What is in the middle that you can, the developers themselves can define? So, this idea is really cool. And it was like, okay, so we are moving towards, we want verification always. All right, cool. So, a lot of different companies have actually started doing this as well. So, Cloud, Anthropic has recently released their new thing called Executor Advisor Pattern, where you've got one agent that actually does all the code work. And then there's an advisor that feeds in, essentially creates a feedback loop. Or in other words, verify.

1:42

SPEAKER_01

Anthropic, sorry. OpenAI built their own harness engineering. And it's the same idea. You give an agent a lot of things to do, but how do you verify it to work? You give it different tools. You give it different context. And that's essentially what a harness is for OpenAI. There are companies like Kudo that provide a very comprehensive code reviews. And again, it's the same thing. The agent has done all of its work, but do you trust it? No. So, what do we do? You do a very comprehensive PR review with all the different issues and findings and create this feedback loop.

1:53

SPEAKER_01

Something from today as well from WorkOS. So, it says enforce, don't instruct. So, it is all about running these checks deterministically. When I say checks, it's just about the verification. Another one, which is my favorite, is you still have to go slow. And the reason for that is not because the agents themselves are not able to produce code as fast as they want, but it's because the verification layer. You need to verify that everything is working or not.

2:02

SPEAKER_01

And my favorite is this one in our keynote. It's to slow the hell down. So, what is the shift that we're seeing here? Initially, what we thought was the value is in the code that we create. But it's actually now, in reality, what we're seeing here is the verification that we design. So, it's not about, can you code, but can you verify? So, TLDR is work on the harness and not on the code. So, you work on the verification system, and that produces a little bit better outputs. And that's it. Thank you. Any questions? I've got 40 seconds. Yes? Yes, it is public. Yes?

2:07

SPEAKER_01

But if you send me a message on LinkedIn, I can share that with you. Oh, there you go. Cool. Yeah. You mentioned that adding the verification layer allows you to use a smaller model. What do you say to the allegation that you're a top focus that they're coming to? I need those tokens to build a verification layer. Cool. I think that's it. or program, and it checks it for me. Cool. And this is how it essentially looks. So I basically give it a config file, define all of my tests over here, like what I needed to be checked. And if it fails, it can actually keep on telling Cloud, like, hey, look, this is failing. Try again. Try again. Try again. Sorry if it's a little bit...

2:31

SPEAKER_01

So you can see one of the test outputs over here. So it's like, okay, first the test passed, first failed. Then it retries again. Then all of the things passed. Okay, cool. So what that means is it's not about whether Cloud can actually do the task. It's about trust. Can I trust Cloud to actually do everything for me? And by the way, when I say Cloud, I'm just talking in general about LLM agents in general, is when I give a task to a coding agent, does it actually complete it?

2:58

SPEAKER_01

So, yeah. So, and that's something that... And what I started doing was started telling this about two people about and going to different events. And it was really cool. Like, hey, look, what about this verification feature that I built? It was so good. It was so amazing. And then I met one of the anthropic engineers. And they just told me that we're not going to need this anymore. Like, we'll have, like, another agent or another model that will be so smart that you won't need enforcement. Okay?

3:28

SPEAKER_01

So, crisis mode. Did I just waste my time? What did I just... Like, what was the point of all of this stuff? But let's dig in a little bit deeper. And then, also, they released this project Glasswing that shows Project Mythos, which is supposed to be so good that it will solve everything for us. So, when I started thinking about it, like, okay, what is it that is actually happening? When a new model comes out, it increases in capability. But that's not necessarily the same thing as reliability. Sure, the models may become a lot more capable. But are they more reliable? The other thing is, like, another argument is, like, oh, well, I can have the best spec.

4:06

SPEAKER_01

I can have the best MCP servers. I can have the best subagents. I can get all the right context to it. Amazing. We should do that. But giving thought instructions is not the same thing as giving it verification. So, you can give as much instructions as you want. Very good instructions, very little instructions. But you still will need to verify. And what I realized was that just... What I realized was that having these small guardrails or having as many guardrails as you want, technically, you can use a smaller model, like a Haiku or even, like, an open source models.

4:36

SPEAKER_01

Because it's got these guardrails on, it'll most likely be succinct and get you to the output that you want. So, in theory, what it means is that if you use a frontier model, like an Opus model, that can get you a task. Okay, cool. That will be the most expensive one. You can have Vector with a little bit of guardrails, but it gives you a little bit cheaper. But if you put on more guardrails, that means invest a little bit more time in the harness itself, like, you can reduce the cost drastically. Or maybe even use, like, async tasks as well. So, okay. I was feeling good. And I started talking about Vector again at two different events.

5:09

SPEAKER_01

And as I spoke to more and more people, there is something missing here. When I spoke to more people, what I realized was everybody's building their own stuff. Anthropic is building their own stuff. My company is building their own stuff about enforcement. Facebook is building their own stuff. Every company is building their own thing. So, if I built something that is specific to me, then I can't really share it with others because everybody has their own way of doing it. And what I enforce doesn't necessarily mean that somebody else would enforce the same thing. So, what that meant was, what I realized was, okay, so it's actually a pattern.

5:50

SPEAKER_01

So, it has to be a pattern that is applicable to everyone. So, what that means is that we can, it has to be language agnostic. It has to be something that can be shared by everybody else. And everybody can bring their own version of enforcement to it. And that's, and it should run on every level. So, it should start off with in conversation. When a conversation ends, you can have checks when, before committing. You can have checks when you're part of a multi-agent workflow. You can have it on checks on asynchronous operations or asynchronous agents. And as well as you can have a check that non-deterministically calls like LLM and LLM as a judge sort of a thing.

6:26

SPEAKER_01

And it can run on any, any different language on any different code. As long as there's a capability to run it deterministically, we can have that. So, what, what I realized was, what we needed was essentially a contract that just says like, hey, given this task, I want you to fulfill this. What is in the middle that you can, the developers themselves can define? So, this idea is really cool. And it was like, okay, so we are moving towards, we want verification always. All right, cool. So, a lot of different companies have actually started doing this as well. So, Cloud, Anthropic has recently released their new thing called Executed Advisor Pattern,

7:03

SPEAKER_01

where you've got one agent that actually does all the code, all the code work. And then there's an advisor that, you know, feeds in, essentially creates a feedback loop. Or in other words, verify. Anthropic, sorry. OpenAI build their own harness engineering. And it's the same idea. Like, you give an agent a lot of things to do, but how do you verify it to work? You give it different tools. You give it different context. And that's essentially what a harness is for OpenAI. There are companies like Kudo who are over here that provide a very comprehensive code reviews. And again, it's the same thing. The agent has done all of its work, but do you trust it? No.

7:37

SPEAKER_01

So, what do we do? You do a very comprehensive PR review with all the different issues and findings and create this feedback loop. Something from today as well from WorkOS. So, it says enforce, don't instruct. So, it is all about, like, running these checks deterministically. When I say checks, it's just about the verification. Another one, which is my favorite, is one of the favorites, like, you still have to go slow. And the reason for that is not because the agent themselves are not able to produce code as fast as they want, but it's because the verification layer. You need to verify that everything is working or not. And my favorite is this one in our keynote.

8:18

SPEAKER_01

It's to slow the hell down. So, what is the shift that we're seeing here? Initially, what we thought was, like, the value is in the code that we create. But it's actually now, in reality, is what we're seeing here is the verification that we design. So, it's not about, can you code, but can you verify?

8:42

SPEAKER_01

So, TLDR is work on the harness and not on the code. So, you work on the verification system, and that produces a little bit better in outputs. And that's it. Thank you.

8:59

SPEAKER_01

Any questions? I've got 40 seconds.

9:07

SPEAKER_01

Yes? Yes, it is public. Yes? Yes? Yes? Yes? Yes? Yes? Yes? But if you send me a message on LinkedIn, I can share that with you.

9:19

Oh, there you go.

9:24

Cool.

9:28

Yeah. You mentioned that adding the verification layer allows you to be a smaller model. What do you say to the allegations that you're a top focus that they're coming to? I need those tokens to build a verification layer.

9:45

Cool. I think that's it.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note