Open Reader

The AI bugpocalypse is here. Now what? - Jack Cable, Corridor

completed 19:43 Jul 12, 2026 Watch on YouTube

Current Status

completed

Video ID

7JgIS42mz7U

RAG / Chat

Enabled
The AI bugpocalypse is here. Now what? - Jack Cable, Corridor
Description

Something shifted in the past year that most security teams haven't fully reckoned with yet: AI models can now find serious vulnerabilities in production code, at scale, with minimal human skill required. Not in toy examples. In libraries that have been reviewed hundreds of times by the best researchers in the world. Jack Cable, Co-Founder and CEO of Corridor, will walk through what this means for the 80% of organizations that have never had to defend against adversaries doing in-house vuln discovery: where the real exposure is, what the available playbooks actually get right, and what concrete steps security teams can take right now to reduce their blast radius before open-weight models make this everybody's problem. Speakers: - Jack Cable (Corridor): Jack Cable is a hacker who serves as the Co-Founder and CEO at Corridor, the security platform for AI coding. X/Twitter: https://x.com/jackhcable LinkedIn: https://www.linkedin.com/in/jackcable/

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI coding is simultaneously expanding the software attack surface and automating vulnerability discovery and exploitation, so defenders must pair AI-enabled security controls with secure-by-design architectural changes rather than rely on one-off patching.
  • Why it matters: As coding agents gain autonomy and produce larger changes, conventional human review becomes the bottleneck while AI-generated code continues to introduce contextual flaws such as authorization bugs.
  • Best use: Use this as a security operating thesis for agentic software development: define pre-merge guardrails, increase visibility over coding-agent activity, and prioritize systemic reductions in exploitable vulnerability classes.

Executive Summary

Jack Cable argues that the "bugpocalypse" is not chiefly about unprecedented vulnerability types. It is about a sharp increase in the speed and scale at which AI can write code, discover familiar weakness classes, and execute larger portions of an attack chain autonomously. The defensive response therefore cannot be to restrict coding agents outright; adoption pressure and development velocity will win. The practical question is how to let teams use agents with enforceable security guardrails.

His central strategic distinction is between reactive vulnerability hunting and structural risk reduction. Many heavily exploited weaknesses, including buffer overflows and other memory-safety failures, have been known for decades and can be prevented through design choices. He presents migration toward memory-safe languages, particularly for critical new code and libraries, as a durable control: it removes entire categories of vulnerability even as attacker models improve.

Cable cautions that model intelligence does not translate automatically into secure code. Citing Baxbench, he says leading models introduce vulnerabilities in roughly 20% to 40% of coding tasks, with the residual risk increasingly concentrated in contextual business-logic failures such as authorization errors. Because models lack an organization's proprietary context and threat model, autonomous development needs security tooling before pull requests, agent-use observability, and confidence mechanisms before AI-reviewed changes can be merged with less human oversight.

At the policy level, he argues that powerful cyber-capable models should be widely usable by defenders because adversaries can already access increasingly capable open-weight alternatives. His priorities are preventing vulnerabilities in newly produced code, hardening open-source dependencies through systemic remediation rather than patch whack-a-mole, and maintaining a competitive ecosystem of U.S.-made open-weight models that organizations can fine-tune.

Key Takeaways

  • Claim: The security problem is compounding from both directions: AI agents are expanding the volume and autonomy of code production while frontier models are improving at vulnerability discovery and exploitation. | Evidence: Cable cites prior Stack Overflow figures that 84% of developers used AI coding tools and 30% to 40% of companies encouraged them, then describes current workflows in which teams launch multiple agents from Slack or other work surfaces to operate in the background for hours and generate large changes. | Implication: Security architecture should assume agentic development is becoming standard operational infrastructure, not treat it as an exceptional tool that can be governed only through manual approval. | Caveat: The adoption figures are explicitly from the prior year, and Cable says he has not seen the latest data.
  • Claim: Most AI-discovered software flaws are not novel classes of vulnerabilities; defenders can get leverage by eliminating known classes structurally. | Evidence: Cable points to MITRE/CISA Known Exploited Vulnerabilities categories and uses buffer overflows, documented more than 30 years ago, as an example of a recurrent weakness that current models still find. | Implication: Prioritize engineering investments that establish enforceable safety properties over programs that merely accelerate discovery and patching of the same recurring defect patterns.
  • Claim: Memory-safe languages provide an especially high-leverage example of secure-by-design because they can eliminate a large share of vulnerabilities inherent to C/C++-style memory-unsafe code. | Evidence: Cable estimates that 60% to 70% of vulnerabilities in products written in memory-unsafe languages could be prevented with memory-safe languages. He cites adoption or use by Google, Microsoft, Amazon, and Linux kernel Rust efforts; Google's Android chart reportedly shows memory-safety vulnerabilities falling from about 75% in 2019 to roughly 30% in 2022 as new code shifted toward memory-safe languages. | Implication: For critical services, libraries, and new components, language and framework selection should be treated as a security control with persistent returns, especially where autonomous agents will generate or modify code. | Caveat: The argument applies to memory-safety classes, not to all application vulnerabilities; language migration does not resolve business-logic or authorization failures.
  • Claim: Frontier coding models still produce insecure code at meaningful rates because security depends on organizational and business context that models often do not possess. | Evidence: Cable cites Baxbench, from ETH Zurich and UC Berkeley researchers, as finding that even the best models introduce vulnerabilities in approximately 20% to 40% of code-writing cases. He also cites an Opus 4.6 smart-contract vulnerability that reportedly enabled theft of a couple million dollars, and identifies authorization issues as a growing contextual failure mode. | Implication: Do not use model capability or generic secure-coding prompts as a substitute for product-specific authorization policies, threat modeling, and adversarial validation. | Caveat: The benchmark rate is not a direct estimate of production incident frequency; real-world exposure also depends on task type, permissions, testing, review, and deployment controls.
  • Claim: AI review is likely to become the dominant review layer soon, but it must be backed by security controls and visibility before organizations reduce human oversight. | Evidence: Cable predicts that within six to 12 months the majority of shipped code will be reviewed by AI rather than humans, reasoning that code review becomes the bottleneck as coding-agent throughput rises. Corridor's stated product focus is vulnerability prevention before the pull request and visibility into AI coding-tool usage. | Implication: Build a controlled AI review pipeline now: instrument agent actions, gate sensitive changes, and establish evidence thresholds for when AI-reviewed changes can be merged or deployed autonomously. | Caveat: The six-to-12-month timeline is Cable's forecast, not an established industry measurement.
  • Claim: Open-source software will be a prime target for AI-enabled vulnerability discovery, so defenders need systemic hardening of foundational dependencies rather than only one-off CVE remediation. | Evidence: Cable argues that open code lets adversaries run capable models directly against widely used libraries and likely uncover novel flaws. His congressional recommendations include hardening the open-source foundation, including rewrites that reduce future exploitability rather than only patching individual findings. | Implication: Dependency-risk programs should identify high-blast-radius open-source components and fund modernization, maintainership, safer interfaces, or targeted rewrites alongside continuous scanning and patch management.
  • Claim: Restricting access to advanced cyber-capable models is unlikely to keep them from adversaries; the defensive priority should be broad defender access with safeguards. | Evidence: Cable and colleagues, in a letter led by Alex Stamos, urged the White House to lift export controls on the named Mythos and Fable models. He argues that open-weight models are rapidly closing the gap through documented distillation approaches and that adversaries already have powerful model access. | Implication: At the organizational level, plan around the assumption that AI-assisted offensive capability will diffuse; invest in defensive model access, evaluation, and operational controls instead of relying on ecosystem-level scarcity. | Caveat: This is a policy position, and the transcript does not quantify the net effect of export controls or model-release safeguards on attacker capability.

Detailed Brief

A practical control model for agentic development

  • Claims: Cable frames the decision for security teams as no longer whether developers should be allowed to use coding agents, but how to enable their use without creating unmanaged code and security risk.; Security must not become the speed blocker, because teams under competitive pressure will route around controls that prevent them from using autonomous development tools.
  • Evidence: He describes the progression from autocomplete to synchronous coding agents to agents that work for one or more hours and produce broad changes.; He identifies pre-pull-request prevention and visibility into coding-agent use as the two core capabilities his company is pursuing.
  • Caveats: The talk does not specify a concrete technical implementation for identity, sandboxing, secrets handling, CI policy enforcement, provenance, or rollback of agent-generated changes.
  • Implications: The control plane for coding agents should be designed as an enabling platform with embedded policy and telemetry, not as a post hoc audit process.; Autonomy levels should be tied to change risk and available evidence rather than granted uniformly across repositories and deployment paths.

Policy and ecosystem posture

  • Claims: Cable's three policy priorities are secure new code, stronger open-source foundations, and a U.S. ecosystem of competitive open-weight models.; He views open-weight models as strategically important because organizations may need to fine-tune models, which he says requires open weights.
  • Evidence: He says his congressional testimony occurred shortly before releases of the named Mythos and Fable models and related export-control actions.; He credits Anthropic with adding safeguards to Fable intended to tilt its use toward defenders.
  • Caveats: The transcript offers no detail on which safeguards work, how they are evaluated, or how fine-tuning and open-weight availability should be balanced against misuse risk.
  • Implications: External security strategy should assume a mixed model ecosystem: closed frontier models for some tasks and controllable, fine-tunable open-weight models for workloads requiring private adaptation or deployment flexibility.

Notable Concepts & Terms

  • Bugpocalypse: Cable's label for the expected surge in discovered and exploitable flaws as AI accelerates code generation and vulnerability research.
  • Secure by Design: The CISA-associated approach of building out common vulnerability classes through engineering choices rather than shifting remediation responsibility to users after release.
  • Memory safety: A language-level guarantee aimed at preventing categories such as buffer overflows; Cable presents it as the clearest example of durable structural security leverage.
  • Known Exploited Vulnerabilities (KEV) catalog: CISA's catalog used here to illustrate that many actively exploited weaknesses are long-known, recurring classes rather than novel AI-created phenomena.
  • Baxbench: A benchmark cited by Cable, attributed to ETH Zurich and UC Berkeley researchers, for the claim that strong coding models still introduce vulnerabilities in roughly 20% to 40% of cases.
  • Pre-pull-request prevention: Cable's preferred point of intervention: detect or prevent insecure agent behavior before code enters the conventional review and merge workflow.
  • Open-weight models: Models whose weights can be accessed and adapted; Cable considers them necessary for fine-tuning and U.S. AI competitiveness, while also acknowledging their dual-use implications.
  • Distillation attacks: The process Cable cites by which open-weight providers can train on closed-model outputs, reducing the time between frontier closed-model release and comparable open alternatives.

Operator Notes / Why Ken Should Care

  • Set a written autonomy policy for coding agents that distinguishes low-risk changes from changes affecting authentication, authorization, payments, secrets, infrastructure, or production data.
  • Require centralized telemetry for every coding agent: initiating identity, model and tool version, repository scope, permissions, prompt or task reference, files changed, test evidence, and merge/deploy outcome.
  • Create a pre-merge security gate for agent-authored changes, with specialized checks for authorization and business-logic invariants rather than relying only on generic static analysis.
  • Inventory critical dependencies and native-code components; rank candidates for Rust or other memory-safe modernization by exploitability, blast radius, and expected lifespan.
  • Run internal evaluations of available models on the organization's own authorization rules, service boundaries, and threat models before relying on them for autonomous generation or review.
  • Avoid treating model-access restrictions as a complete threat control; maintain a roadmap for AI-assisted defensive testing of proprietary systems and high-risk open-source dependencies.

Source/Metadata

  • Title: The AI bugpocalypse is here. Now what? - Jack Cable, Corridor
  • Transcript words: 5629
  • Duration seconds: 1183
  • Timestamp note: No timestamps or chapters were present. The supplied transcript repeats much of the talk after the closing, so its stated word count includes substantial duplicated content.

Transcript

2897 words en Processed in 168.7s

Hey there, I'm Jack Cable, and today I'm going to be talking about the effects of the AI bugpocalypse. As you may have seen, frontier models are getting better than ever before at discovering and exploiting vulnerabilities in our software. This is leading to what many are calling a bugpocalypse, where we're finding more and more vulnerabilities, particularly in the open-source libraries that power all of the software we rely upon. So today I want to break down what exactly is happening and how defenders can get ahead of the exploitation that is occurring. As far as my background, right now I'm the co-founder and CEO at Corridor, a company I started about 18 months ago focused on securing AI coding. Before this, I served as a senior technical advisor in government at CISA, the Cybersecurity and Infrastructure Security Agency, where I worked with top software companies to help them build their products to be more secure by design. I'm also an ethical hacker. I got into the top 100 rank of hackers on HackerOne when I was in high school and studied computer science at Stanford. So I've seen firsthand how these simple repeat classes of vulnerabilities can be introduced and exploited, and have been a close participant in many of the most recent advancements, and seen what this means for both our adversaries as well as defenders. Just to set the stage, as everyone here knows, I imagine, AI coding tools are scaling faster than any software category in history. We've seen Cursor, Cloud Code, grow exponentially. And with that also come these improvements in how frontier models can find and exploit vulnerabilities. So we're seeing both ends of the equation shifting. On one hand, models can do a better job finding vulnerabilities. On the other hand, our attack surfaces are growing immensely as AI becomes the default code writer. So what I want to explore in this talk is how do we balance that? How do we make sure that we're not going to have immensely more vulnerabilities than we've ever had before? And just to give some sense, I'll move myself here, of some of the statistics, pulled some from last year where about 84% of developers were using AI coding tools, 30% to 40% of companies encouraging use of AI coding assistance. That was from Stack Overflow. I haven't seen the latest numbers this year, but what I would expect once those come out is that it is the vast, vast majority of developers and companies who are using coding agents. And part of this is the increasing level of autonomy by which these coding agents are being used. It's no longer autocomplete. Often it's not even a developer synchronously within Cursor. When we do our own development right now, it's spinning up agents from within Slack or wherever folks are working and having many agents run at once in the background. This is a tremendous shift in how software is being built. And at the same time, like I mentioned, the frontier models are getting significantly better. And you can look at it from pretty much any part of the cyber attack chain, ranging from finding vulnerabilities, where models can now do significantly better than even I could. And I've reported hundreds of vulnerabilities to various companies. So everything from finding vulnerabilities to exploiting them. This is a chart here that comes from Anthropic, showing Mythos compared to a number of other models that they and others have put out. And we can see that we're seeing quite rapid advancements in model capabilities, and particularly to execute more autonomous attack chains. So as we think about adversaries who are using these models, they're not just going to be discovering vulnerabilities, but they're going to be automating every part of the attack process. So it's our job as defenders to understand, okay, what are the points where we can make software systems more resilient to all of these attacks? And to me, this brings back a lot of the work that I was doing in government around the Secure by Design initiative. And so the overall question that I'm worried about is how can we make sure that frontier AI models aren't introducing exponentially more vulnerabilities over time? Even pre-AI, we've had this heavy increase in common, relatively simple classes of vulnerabilities that are being exploited by adversaries. AI is making this significantly easier. So I think the only way that we're going to win as defenders is if we use the same techniques to harden our systems. And I would say that there is good news here, that a lot of the vulnerabilities, pretty much all of the vulnerabilities, that even frontier AI models are finding aren't anything new. Yes, it's new that a given vulnerability was found in a specific file within a piece of software, but that vulnerability class isn't necessarily novel. And we can actually use that to our advantage, and I'll get into that. So overall, the thesis here is that we are seeing both attackers get more tools in their toolkit. At the same time, the way in which software is being built is fundamentally changing. So really, the question then becomes, how can we apply AI to shore up these software systems? I want to take a quick detour to some of the Secure by Design work that I kicked off with others in government. This is a paper that we put out in March of 2023, just as LLMs were starting to become more readily available, but their application in coding at that time wasn't much more than autocomplete. And while that's useful, it wasn't necessarily the step change that we have now. And what we focused on laying out with this vision was this idea that it isn't really rocket science when it comes to preventing vulnerabilities in software. While it's true that it's hard to build a perfectly secure system, we do know how to build systems that are fundamentally more resilient to common classes of vulnerabilities. And just to make this concrete, this is a set of vulnerability classes coming from MITRE. It's the top classes that are exploited in CISA's Known Exploited Vulnerabilities catalog. And if you go down this list, you might notice that pretty much all of these are basic types of vulnerabilities that not only have we known about for decades, but we've known how to prevent at scale for decades. Take buffer overflows, number two on that list. And by the way, these are the same vulnerabilities that models like Mythos are finding in software. Buffer overflows were first documented about, I believe, 30-plus years ago. So we've had documented instances of how to find and exploit these vulnerabilities. And we also now have languages that are memory safe. Languages like Rust, Go, pretty much any language that's not C or C++, are built in a way such that it's impossible to introduce memory safety vulnerabilities. They have guarantees that prevent those from being introduced. So we have techniques by which we can prevent them, and yet they continue getting introduced over and over again. So let's look at memory safety, for instance. The statistics range, but approximately 60% to 70% of vulnerabilities in products written in memory-unsafe languages can be completely prevented using memory-safe languages. This is based on CVE data out there. And not only that, we've seen a lot of companies, Google, Microsoft, Amazon, even open-source software. The Linux kernel is being rewritten in parts in Rust. We've seen real evidence that by shifting to memory-safe languages, you can reduce overall vulnerabilities. On the right here is a chart from Google showing the rate of memory safety vulnerabilities over time in the Android operating system. And what's interesting is that they aren't even necessarily rewriting code in a memory-safe language. They're just writing new code in a memory-safe language, and even then the percent of memory safety vulnerabilities has dropped quite dramatically from about 75% in 2019 to maybe 30% in 2022. So to me, that's personally quite exciting, because it means that it's not a given that we're going to continue having these basic vulnerabilities over and over again. And part of the high-level policy conversation, I think, as a result has to be not just how can we deploy these frontier models to find one-off vulnerabilities in software. That is something that we should be doing. But at the same time, I don't want to miss out on opportunities to make our software fundamentally more secure. We could pour millions of dollars into essentially playing whack-a-mole with vulnerabilities and patching them one-off in some of the open-source libraries that we all rely on. Or we could do a one-time rewrite, for instance, to move some of these critical libraries into a language like Rust, and then that will pay dividends for years to come. So this is really at the core of how I'm thinking about this, is what are some of the fundamental changes that companies, that open-source developers, can be making that can reduce exploitation both by models today and models to come. Because the advantage of doing a rewrite, for instance, is that if you have some of these fundamental guarantees, then even if the models get smarter, Rust has programmatic guarantees such that we know that memory safety vulnerabilities in those circumstances won't be possible to be introduced or discovered. And all of this comes in the context, too, of the fact that AI is increasingly capable at, of course, both writing code and then introducing vulnerabilities. You might have seen a couple months ago one example where Opus 4.6, by all accounts a very smart model, introduced a vulnerability in a smart contract that led to a couple million dollars being stolen. So while the models are very smart and capable, oftentimes security is very contextual, and the model just might not have the context in order to know that it's introducing a vulnerability. And this is reflected in academic benchmarks. One, for instance, here, Baxbench, you can find that at Baxbench.com, by researchers at ETH Zurich, UC Berkeley, finds that even the best models introduce vulnerabilities about 20% to 40% of the time when writing code. And this shouldn't necessarily come as a surprise. For one, models are trained on all of the world's existing code, and humans haven't been great at not introducing vulnerabilities in code in the past. But two, increasingly, and this lines up with some of what we're seeing among our customers, the vulnerabilities being introduced are often less the basic one-liner vulnerabilities and more contextual issues. Things like authorization bugs that require an in-depth understanding of a company's business logic. And that's something, even if the model is very smart, it's not being trained on your company's proprietary information or how your own threat model works. And that is why I believe we're still seeing quite a high rate of vulnerability introduction, even by, by all accounts, very intelligent models. So let's now think about, okay, given that the vast majority of software development is being done with AI, how can we make sure that AI is capable of writing secure-by-design software? And part of this is a shift we're seeing in the level of autonomy that AI is now given when it comes to software development tasks. We're moving up this ladder that started with autocomplete to agents within Cursor, Cloud Code, that can synchronously produce code. Now, increasingly, these autonomous agents can work for an hour, hours at a time, and produce quite large code changes. Of course, the next step then becomes agents that are reviewing code. And we at Corridor believe that within the next six to 12 months, the majority of code that is being shipped will be reviewed not by humans but by AI. I think that's just a natural consequence of the rate at which companies need to move, given that code review is now the bottleneck, and I don't think we're going to accept that for very long. So really, our perspective at Corridor is around preventing vulnerabilities before the pull request, as well as giving visibility into how AI coding tools are being used. And I think that's really essential, that security cannot be the blocker when it comes to companies accelerating their development. Acceleration is always going to win out. So when we talk to security teams, the conversation is less around should you allow your development teams access to coding agents? The answer is obviously yes. It's more around how can you do that with guardrails in place? Because what we're seeing is that without guardrails, yes, the coding agents can introduce vulnerabilities. And in order to get to a point where development can be more autonomous, that code can start to be reviewed by AI and merged in without as much human oversight, we really need to have tooling in place that allows security teams to have that assurance and to give the blessing to their engineering team to accelerate. I want to close with some of the policy perspective. And this is in part tied to the recent export controls on Mythos and Fable models. This is part of a letter led by my colleague Alex Stamos, where we urged the White House to lift the export controls on these models. And the perspective there is that the benefit to defenders far outweighs the risk. These are very powerful and, let's face it, dual-use models that can both be used to secure systems and also to exploit them. To Anthropic's credit, they have done a lot of work with the Fable release to have some safeguards in place such that they're more skewed towards defenders than adversaries. But this is also coming in the face of increasingly powerful open-weight models. You've probably seen documented distillation attacks where open-weight model providers can train on the output of closed-weight models. And that, as a result, is quickly shrinking the timeline between when a frontier closed-weight model comes out and when open-weight models catch up to that. So whether we like it or not, adversaries already have access to incredibly powerful models. They're already using them today to exploit systems. So to me, it becomes more a question of how can we rapidly get the capabilities into the hands of defenders. And I think that requires having these models be more widely available. One cool thing I had the opportunity of doing a couple weeks ago was testifying to the U.S. Congress on the risks of both frontier models as well as AI coding. My recommendations had a couple elements. One was to prevent vulnerabilities in new code going forward. I think this is something that both every company as well as the U.S. government should be focusing on, and making sure that as development accelerates, security isn't being left behind there. Second is to harden the open-source foundation. This is incredibly important, especially since open-source software is going to be the proving ground for a lot of adversaries who want to test out these models and exploit vulnerabilities they find due to the exact nature that it's open source. So you can just go and run a very smart model on them and, in all likelihood, find many novel vulnerabilities. So I think both the U.S. government, but then private companies as well, have a responsibility to help shore this up. And like I mentioned, it's not just about one-off vulnerability discoveries or patches. It really has to be more systemic and start to get into rewrites that can fundamentally reduce the risk of vulnerabilities that can be found whether by models today or in the future. And then the last area of my recommendations was to foster an ecosystem of American-made open-weight models. I think in order for us to stay competitive here, it can't just be closed-weight models alone. There's a number of reasons here. One of those is that for many companies, while there's a place for closed-weight models, you also might want to do things like fine-tuning models. And that is only possible with an open-weight model. So I think it is really essential that we have frontier open-weight models coming out of the United States. We haven't seen as much of that to date, but I think that's a critical element of American competitiveness when it comes to AI. So those were my overall recommendations. Of course, all of this is in the context of these increasingly powerful models. The hearing was, I believe, a few days before Mythos and Fable came out, and then, of course, all of the export control actions that were taken. So this is an incredibly rapidly evolving space. But I think that's why it's all the more important to go back to the basics. What are the fundamental controls that can protect against any vulnerabilities that can be discovered by models today or in the future? And I think that's where we really ought to be spending our time, using these models to make our systems more resilient. So that's my talk. I'm happy to be reached. My email is here. And thanks, everyone, for tuning in. I haven't seen the latest numbers this year, but what I would expect once those come out, right, is that is the vast, vast majority of developers and companies who are using coding agents, right? And part of this is the increasing level of autonomy by which these coding agents are being used. It's no longer, you know, autocomplete. Optin's not even a developer synchronously within cursor, right? Like when we do our own development right now, it's spinning up agents from within Slack or wherever folks are working and having many agents run at once in the background. This is a tremendous shift in how software is being built. And at the same time, right, like I mentioned, the frontier models are getting significantly better. And you can look at it from, you know, pretty much any part of the cyber attack chain, ranging from finding vulnerabilities where models can, you know, now do, you know, significantly better than even I could. And, you know, I've reported hundreds of vulnerabilities to various companies. So everything from finding vulnerabilities to exploiting them. This is a chart here that comes from Anthropic, right, showing Mythos compared to a number of other models that they and others have put out. And we can see that we're seeing quite rapid advancements in models capabilities, right, and particularly to execute more kind of autonomous attack chains. So as we think about adversaries who are using these models, right, they're not just going to be discovering vulnerabilities, but they're going to be automating every part of the attack process. So it's our job, right, as defenders to understand, okay, what are the points where we can make software systems more resilient to all of these attacks, right? And to me, this brings back a lot of the work that I was doing in government around the Secure by Design initiative, right? And so the overall question, right, that I'm worried about is how can we make sure that frontier AI models aren't introducing exponentially more vulnerabilities over time, right? Even pre-AI, we've had this, you know, heavy increase in common, relatively simple classes of vulnerabilities that are being exploited by adversaries. AI is making this significantly easier, right? So I think the only way that we're going to win as defenders is if we use the same techniques, right, to harden our systems. And I would say that there is good news here, right, that a lot of the vulnerabilities, pretty much all of the vulnerabilities, that even frontier AI models are finding aren't anything new. Yes, it's new that a given vulnerability was found in a, you know, specific file within a piece of software, but that vulnerability class isn't necessarily novel. And we can actually use that to our advantage, and I'll get into that. Right, so overall, the thesis here is that, right, we are seeing both attackers get more tools in their toolkit. At the same time, the way in which software is being built is fundamentally changing. So really, the question then becomes, how can we apply AI to shore up these software systems? Right, and I want to take a quick detour to some of the secure by design work that I kicked off with others in government. Right, this is a paper that we put out in March of 2023. So just as, you know, LLMs were starting to become more readily available, but, right, their application in coding at that time wasn't much more than, you know, autocomplete. And while that's useful, it wasn't necessarily the step change that we have now. And what we focused on kind of laying out with this vision, right, was this idea that it isn't really rocket science when it comes to preventing vulnerabilities in software, right? While it's true that it's hard to build a perfectly secure system, we do know how to build systems that are fundamentally more resilient to common classes of vulnerabilities. And just to, you know, make this concrete, this is a set of vulnerability classes coming from MITRE. It's the top classes that are exploited in CISA's known exploit vulnerabilities catalog, right? And if you go down this list, you might notice, right, that pretty much all of these are basic types of vulnerabilities that not only have we known about for decades, but we've known how to prevent at scale for decades, right? Take buffer overflows, right, number two on that list. And by the way, these are the same vulnerabilities that models like Mythos are finding in software. And buffer overflows were first documented about, I believe, 30-plus years ago, right? So we've had documented instances of how to find and exploit these vulnerabilities. And we also now have languages that are memory safe, right? Languages like Rust, Go, pretty much any language that's not C or C++ is built in a way such that it's impossible to introduce memory safety vulnerabilities, right? They have guarantees that prevent those from being introduced. So we have techniques by which we can prevent them, and yet they continue getting introduced over and over again, right? So let's look at memory safety, for instance. There's, you know, the statistics range, but approximately 60-70% of vulnerabilities in products written in memory unsafe languages can be completely prevented using memory safe languages, right? And this is, you know, based on CVE data out there. And not only that, right, we've seen a lot of companies, Google, Microsoft, Amazon, even, you know, open source software. The Linux kernel is being rewritten in parts in Rust. We've seen real evidence that by shifting to memory safe languages, you can reduce overall vulnerabilities, right? On the right here is a chart from Google showing the rate of memory safety vulnerabilities over time in the Android operating system. And what's interesting, right, is that this isn't even, you know, they're not even necessarily rewriting code in a memory safe language. They're just writing new code in a memory safe language, and even then, right, the percent of memory safety vulnerabilities has dropped quite dramatically from, you know, about 75% in 2019 to maybe 30% in 2022. So to me, that's personally quite exciting, right, because it means that it's not a given that we're going to continue having these basic vulnerabilities over and over again, right? And part of the, you know, high-level policy conversation, I think, as a result has to be not just how can we deploy these frontier models to find one-off vulnerabilities in software. That is something that we should be doing. But at the same time, right, I don't want to miss out on opportunities to make our software fundamentally more secure, right? We could pour millions of dollars into essentially playing whack-a-mole with vulnerabilities and patching them one-off in some of the open-source libraries that we all rely on. Or we could do a one-time rewrite, for instance, to move some of these critical libraries into a language like Rust, right, and then that will pay dividends for years to come. So this is really at the core of how I'm thinking about this, right, is what are some of the fundamental changes that companies, that open-source developers can be making that can reduce exploitation both by models today, right, and models to come. Because the advantage, right, of doing a rewrite, for instance, is that if you have some of these fundamental guarantees, then even if the models get smarter, right, Rust has programmatic guarantees such that we know that memory-safe de-vulnerabilities in those circumstances won't be possible to be introduced or discovered. Right, and all of this comes in the context, too, of the fact that AI is increasingly capable at, of course, both writing code and then introducing vulnerabilities. You might have seen a couple months ago one example where Opus 4.6, right, by all accounts a very smart model, introduced a vulnerability in a smart contract that led to a couple million dollars being stolen, right? So while the models are very smart and capable, oftentimes security is very contextual, and the model just might not have the context in order to know that it's introducing a vulnerability, right? And this is reflected in academic benchmarks. One, for instance, here, Baxbench, you can find that at Baxbench.com by researchers at ETH Zurich, UC Berkeley, finds that even the best models introduce vulnerabilities about 20 to 40% of the time when writing code, right? And this shouldn't necessarily come as a surprise. For one, right, models are trained on all of the world's existing code, and humans haven't been great at not introducing vulnerabilities in code in the past. But two, right, increasingly, and this kind of lines up with some of what we're seeing among our customers, is that the vulnerabilities being introduced are often less so the basic one-liner vulnerabilities and more so contextual issues, right? Things like authorization bugs that require an in-depth understanding of a company's business logic. And that's something, right, even if the model is very smart, it's not being trained on your company's proprietary information or how your own kind of, you know, threat model works. And that is why, you know, I believe we're still seeing quite a high rate of vulnerability introduction, even by, you know, by all accounts very intelligent models. So, let's now think about, okay, given that the best majority of software development is being done with AI, how can we make sure that AI is capable of writing secure-by-design software, right? And part of this is a shift we're seeing in the level of autonomy that AI is now given when it comes to software development tasks, right? We're kind of moving up this ladder that started with autocomplete to, you know, agents within cursor cloud code that can synchronously produce code. Now, increasingly, these autonomous agents that can work for an hour, hours at a time, and produce quite large code changes. Of course, the next step then becomes agents that are reviewing code. And we at Corder believe that, you know, within the next six to 12 months, the majority of code that is being shipped will be reviewed not by human but by AI. I think that's just a kind of natural consequence of the rate at which companies need to move, given that code review is now the bottleneck, and I don't think we're going to accept that for very long. So, really, right, our perspective at Corder is around preventing vulnerabilities before the pull request, as well as giving visibility into how AI coding tools are being used. And I think that's really essential, right, is that security cannot be the blocker when it comes to companies accelerating their development, right? Acceleration is always going to win out. So, when we talk to security teams, the conversation is less around, should you allow your, you know, development teams access to coding agents? The answer is obviously yes, right? It's more around how can you do that with guardrails in place, right? Because what we're seeing is that without guardrails, yes, the coding agents can introduce vulnerabilities. And in order to get to a point where, you know, development can be more autonomous, that code can start to be reviewed by AI and merged in without as much human oversight, we really need to have tooling in place that allows security teams to have that assurance. And to give the blessing to their engineering team to accelerate. I want to close with some of the policy perspective, right? And this is in part tied to the recent export controls on Mythos and Fable models, right? And this is part of a letter led by my colleague Alex Stamos, where we urged the White House to lift the export controls on these models. And the perspective there is that the benefit to defenders far outweighs the risk, right? These are very powerful and, let's face it, right, dual-use models that can both be used to secure systems and also to exploit them. To Anthropik's credit, they have done a lot of work with the Fable release, right, to have some safeguards in place such that they're more skewed towards defenders than adversaries. But this is also coming in the face of increasingly powerful open-weight models. You've probably seen, you know, documented distillation attacks where open-weight model providers can train on the output of closed-weight models. And that, as a result, is quickly shrinking the, you know, timeline between when a frontier closed-weight model comes out and when open-weight models catch up to that. So whether we like it or not, right, adversaries already have access to incredibly powerful models. They're already using them today to exploit systems, right? So to me, it becomes more a question of how can we rapidly get the kind of capabilities in the hands of defenders. And I think that requires having these models be more widely available. One cool thing I had the opportunity of doing a couple weeks ago was testifying to the U.S. Congress on the, you know, risks of both frontier models as well as AI coding. My recommendations had a couple elements, right? One was to prevent vulnerabilities in new code going forward. I think this is something that both every company as well as the U.S. government should be focusing on and making sure that as development accelerates, right, security isn't being left behind there. Second is to harden the open-source foundation. This is incredibly important, right, especially since open-source software is going to be the kind of proving ground for a lot of adversaries, right, who want to test out these models and exploit vulnerabilities they find due to, right, the exact nature that's open-source. So you can just go and run a very smart model on them and in all likelihood find many novel vulnerabilities. So I think both the U.S. government but then private companies as well have a responsibility to help shore this up. And like I mentioned, right, it's not just about one-off, you know, vulnerability discoveries or patches. It really has to be more systemic and start to get into rewrites that can fundamentally reduce the risk of vulnerabilities that can be found whether by models today or in the future. And then the last area of my recommendations were to foster an ecosystem of American-made open-weight models, right? I think in order for us to stay competitive here, it can't just be closed-weight models alone, right? There's a number of reasons here. You know, one of those is that for many companies, while there's a place for closed-weight models, you also might want to do things like fine-tuning models. And that is only possible with an open-weight model. So I think it is really essential, right, that we have frontier open-weight models coming out of the United States. We haven't seen as much of that to date, but I think that's a critical element of American competitiveness when it comes to AI. So those were my overall recommendations, right? Of course, all of this is in the context of these increasingly powerful models. The hearing was, I believe, a few days before, you know, Mythos and Fable came out. And then, of course, all of the export control actions that were taken. So this is an incredibly rapidly evolving space. But I think that's why it's all the more important to go back to the basics, right? What are the fundamental controls that can protect against any vulnerabilities that can be discovered by models today or in the future? And I think that's where we really ought to be spending our time using these models to make our systems more resilient. So that's my talk. I'm happy to be reached. My email is here. And thanks, everyone, for tuning in.