.
What's up, everyone? Good to see you all here in the very first ever security track at the World's Fair. Pretty exciting day, honestly, because all of you, I genuinely love this stuff. And having looked through the agenda for the speakers we have today, it's going to be absolutely mind-blowingly useful and fun information. So hopefully, if you stick around all day, by the end of the day, you'll be able to leave here, go back to your hotel rooms, and build some genuinely cool software, hopefully without humans in the loop. Because that's the ultimate goal, right? How do we build truly safe, autonomous software at scale? And it's not something easy to do.
So, with that being said, I'm very excited to welcome my good friend and colleague, Manoj Nair. He is Snyk's chief innovation officer and CTO. Before Snyk, he was the chief cloud officer at Commvault. He founded and ran Hypergrid. He did product and security leadership at HPE, Dell, and RSA. And he's got something like a dozen patents to his name. So he's a legend in the space. He also personally had a hand in curating this entire track. So if you like the talks today, please go up and say thank you to him after. But if not, then just don't blame me. But yeah, welcome to the stage, Manoj. Manoj Nair Nairi, Thank you, Randall. I was not expecting a bio.
Hi, everyone. I really appreciate you all joining here right after those great keynotes up front. I'm Manoj Nair. And I have the pleasure of leading an amazing team that is helping secure about 5,000 enterprise customers around the globe. So I get to look good because of all of that. But some of what I'm going to show is real data from those customers. Half of Fortune 500 runs on Snyk. And some of the data is from those learnings.
But before that, I also want to talk about this. The title of the talk was Cutting Through the AI Fog, I think. Well, the way we do that is by having tracks like this. So last year we stood here. There were 3,000 people at the AI Engineer World's Fair. And it really felt like security was missing in the room. And thanks to SWIX and the amazing partnership with the AI Engineer Organization, we created the AI Security Summit in partnership with them. And we're creating this track.
So it is really good to see this and all the great speakers who are going to talk here today with some fantastic knowledge. But that is, in our mind, how we cut through the fog, right? Security needs to be very much part of the room. We're very passionate about not slowing things down. But really, that's how you build trusted systems.
So one thing you're going to hear from me quite a bit in this, if you take one thing, it's this notion of our learning from this real-life data and working with the biggest frontier labs in the world and the biggest companies in the world, is this concept that has really not been questioned in security before, but it is being asked now, right? Can the generator and the validator be the same? And our point is, in some of the data I'll show, for all kinds of reasons, why not, right?
And almost, if you know Snyk, don't think about the supply chain security company that shifted left. A lot of what you'll see is the last 18 months of what we have been doing with some of these very large enterprises and adopting these complex systems that we all enjoy and we love. And what are we learning from that?
There are three problems that we hear when we talk to these customers. And I want to share those, and I want to show some of the data behind that. But fundamentally, it's really looking at this as you think about this room and everyone that you support. You're building fast at the frontier. The question nobody is answering is, can you trust what your agents just shipped and how they did it, right? And so, if you really take the gist of, and I have the joy of talking to all of these customers a lot, this is pretty much every one, two, or three, or all of them. It's really the conversation that happens.
So autonomous attacks, it's not mythos. It's not, I said the M word, but sorry, it's not GPT-55 cyber. We're seeing that you can have these attacks. All the frontier labs have been used in automated attacks. You can do that even without having frontier models. We have shown it. So have the attackers, unfortunately. With good context and good harness, now you have an attacker that never sleeps.
What does it do? The fundamentals of things like application security, that was already something we tried to disrupt for 10 years, and the barriers have completely gone away. You cannot have contextual risk management and say, I fixed my criticals and I fixed my highs and I'm pretty good because everything else is too hard. No, you can string low vulnerabilities and create exploits. You can do it without a lot of a harness with the mythos class model, as we've seen, but with a little bit of effort you can do it with everything else. So that's a big sea change in how we are seeing our customers actually react in real time.
And then you go, wow, a thousand enterprises spend over a million dollars for a Cloud Code rollout. So this is the answer, right? Well, maybe. Let's just pull that string a little bit. So that's the reason why I said that there are existing classes of problems that are getting worse. The quality of code is unfortunately worse than human-generated code. It's not like humans, I'm an engineer, I don't think I wrote perfect code. None of us do. But it is actually a little worse.
And so, well, if that was the only problem, that's okay. But what about the environment? All the things that we find magical are based on skills and MCP servers and all these things that we want to share and use. And those are intentionally or unintentionally both being poisoned, and malware is injected in there, and we'll show some of the research on that. And then the behavior of the agent, right? And so we're having these new patterns of problems compounding on top.
And then who here doesn't want to become an AI company, or in any room in the world, right? Every company, every board is like, we're going to transform our business, our workflows, or you're starting brand new, you're fully agentic. And when you build with agents and models, that's an entirely new threat surface that was not part of the prior one.
And so these are really the fundamental problems, right? And this is, I want to share some of the real data. This is 4,800-plus customers in the last year. Their actual backlog, quarter over quarter, was like 108% more backlog. So remember what I said about attackers. It's not just the existing vulnerabilities, and unfortunately there's millions of them if you have an enterprise of any size. It's the fact that we are growing them despite the best agents, despite things that we have done and the industry has done. This trend is real, and it's not good.
And you have novel exploits, but they're not novel like the light LLM exploit. You think about it like it's really taking existing vulnerabilities in the new surface and chaining them together. And so they create a much bigger blast radius right now at the pace at which work is being done. And yes, speed is a big part of this, but it's not just speed.
And just this last week we had the Five Eyes. These are the Western world's intelligence leaders talking about AI will bypass cybersecurity systems in months, not years. Now, I don't share this, I hate to use a scare tactic. It's just a fact. It's getting ready for whatever is coming, and we have already seen what's coming. It's just not widespread enough. And so trying to chase systems and the specific model will be the cure for all this is not the way, is our point.
The data on the second pain point, the untrusted output, environment, and behavior of agents, and these are again facts. These are benchmark data facts. All of them, the QR codes at the top, or if you want to go check out the studies and the research behind it. This is the amount of vulnerable code coming from the latest models. Plus the skills, there's toxic skills. Now, I don't share this. I hate to have the scare tactic. It's just a fact. It's getting ready for whatever is coming, and we have already seen what's coming. It's just not widespread enough. And so trying to chase systems and the specific model will be the cure for all this is not the way, is our point.
The data on the second pain point, the untrusted output environment and behavior of agents. And these are, again, facts. These are benchmark data facts. All of them, the QR codes at the top, or if you want to go check out the studies and the research behind it. This is the amount of vulnerable code coming from the latest models. Plus the skills, there's toxic skills. We found the research, the seminal research around how skills, a third of them or more than a third of them in all of the skills, not just open claw, clawed, codex. These skills actually have malware, and they have vulnerabilities that are being. Three lines of English are able to now bring a system down.
So you have to really understand the intent behind it. And the MCP servers, how do you connect to enterprise data? This is great protocol, very little security built in. It's getting better. But the foundations behind it is this is the GitHub MCP server exploit that we highlighted a year ago. And what did some of our customers do? Immediately shut down all MCP servers. Then they figured out that all of their devs screamed. So then how do you actually go back safely enabling things that are very powerful? And then on the behavior, everyone knows the Pocket of Us example, right?
What I have is real data in our own environment, in our Fortune 100 customers. These agents go and create copies of PII data. Why? Somebody shared the PII data and was trying to solve a real customer problem. The agent thought that maybe I should create, squirrel away, a copy of this in a database just in case I need it again. That database is untrusted. Great. You now have an unknown attack surface that is not part of any enterprise security coverage. This is happening. So how do you not put gates on it? How do you steer them, right? And so you start going into the last problem is you can't govern what you don't know exists. This is when you start building with agents.
Now this is real data from 3,000 plus customers who have used our abilities to really find the intelligence of what's in their code bases, what they're building. And for every model that we find in the repo, you have three times more agentic components in there. You have agents and the tools and everything that they use. You have to figure out the full landscape because the risk is not just at one layer. And then once you find it, how do you know how risky is it? What is your independent data verification? So this is from this weekend from our risk BB that we have built, our own attacks and red teaming capability. As new models and new components come out, we check them.
The first two are your favorite frontier models. You can guess which ones those are. They did awesome on PII extraction. They didn't used to be so good with our attacks. No PII extraction with our attacks. The third one is the hot new model in Silicon Valley especially, or this last few weeks. Rhymes with LLM. 100%. 100% of the time, our attacks were able to extract PII. But when you check a different test, decision override. The frontier models did worse. The open model, 0% of the time, you were able to override the decision, at least with our attacks. Knowing this allows you to know what to use, when to use that, and how do you control.
And these things are changing dynamically. So, again, going back to the original point, right? If none of that convinces you, the generator, validator separation, this is fresh new research from just, I think, yesterday is when we went public with this. It's a benchmark that no model has been trained. So we can trust it for now. And we'll have to keep updating the benchmarks. But this is just a very simple thing. We're asking the latest models, and we have access to everything, as you can imagine. We're asking them to find the same vulnerability, run it five times. And only 50% of those ones are found across those five tests.
That's not how you can run an enterprise system if you just use the LLM without anything else. This is the latest models. Only 75% of the issues were found versus a good old boring deterministic check. And 40% was the F1 score. So this whole, what did you actually miss? And we're talking about the latest models that are not even publicly available to people, right? So what does this mean? It doesn't mean that they're not good. It just means they need to be, you really need to use them for what they're really good at together and carefully to find the surface that your deterministic check cannot use. Not just think probabilistic systems will solve everything.
And so, to net it out, what we've been working on, we don't have all the answers. And I'll share where we're going to. But what we have answers for right now on the automated attacks that are working in some of these very large environments, just how do I prevent new issues from coming into the agentic loop? So, put security context right there. That's studio. You can go check it out on our website. And everyone's worried about packages that they download, how we know what the health of the packages are. We know the vulnerability information. We know the malware.
So we're able to prevent the agent from picking a package like that or writing code that inherently is a SQL injection. But great. So prevention will prevent that hockey stick, which we all want to prevent. But what happens with the fact that you have this mountain of vulnerability? We've been able to take organizations like Labelbox to zero ones. Zero backlog. Which is very hard in security because it breaks applications. So you need to know concepts like breakability. And that data is super important that we are able to get from our base to know this is a safe upgrade.
And so, just last week, a Mag 7 company remediated 16,000 critical issues using this remediation agent. Again, something you can go try out. So this is how I talked about from the untrusted agentic development. We just GA'd our agentic dev security offering yesterday. It's looking at the environment, the output, the skills, the MCP servers, and the behavior of coding agents like cursor, cloud, codex, and others. And when you're building ungoverned AI apps, that AI governance cannot live in a Confluence page or PDF. So how do you real-time look at everything that is happening in very fast-moving complex code repositories and understand risk?
And how policies enforce in the loops that the agents and the devs are. So, seeing is believing, I would love to bring Ezra up here to do a quick demo. And Ezra is going to show you a couple of things. And you can come by to our booth and share a few more. Later. Thanks, Manoj. Let's flip over to the CLI here. I'm going to do a few rapid-fire demos. We don't have time to do everything that Manoj just talked about here. But we'd love for you to come and talk to us. Stop by the booth, find us after this talk here. The first thing that I'm going to do here is I'm going to ask Claude to generate a tool, a CLI tool, that can create a QR code image from a prompt.
Start codex with Claude. There we go. We need to start Claude. Thank you. Live demos. Let's try that one more time. All right. Try number two here. We are going to try to get Claude to generate a CLI tool to ultimately create that QR code image. I'm asking it to use an open source dependency. And if you notice this fourth line that I have here, I'm being pretty verbose. I'm telling it to use the sneak package health check tool. And so we're going to be able to see that here in the demo. But if you use this in practice, we really encourage you to leverage a skill or a hook to make this just a deterministic, automated part of your workflow.
So the first thing that it's going to do is going to see ultimately what does the project look like and what are some dependencies, some open source packages that I might be able to use. And let's see if we can expand this. I don't know if conference Wi-Fi is going to play nice with us. If not, we can switch to a recorded demo here, but I was really hoping to do this live. Is that a little better? Even more? All right. I'm asking it to use an open source dependency. And if you notice this fourth line that I have here, I'm being pretty verbose. I'm telling it to use the sneak package health check tool. And so we're going to be able to see that here in the demo.
But if you use this in practice, we really encourage you to leverage a skill or a hook to make this a deterministic, automated part of your workflow. So the first thing that it's going to do is see ultimately what the project looks like and what are some dependencies, some open source packages that I might be able to use. And let's see if we can expand this. I don't know if conference Wi-Fi is going to play nice with us. If not, we can switch to a recorded demo here, but I was really hoping to do this live. Is that a little better? Even more? All right. Great. So we can see now that it is calling this package health tool for two different open source dependencies.
It's looking at a QR code package and a QR image package, both viable packages that can help accomplish this task. Results are returning. Now it looks like conference Wi-Fi is really, really not playing in our favor here. So I'm going to quickly jump over to a recorded demo. Apologies. If you can see the other queries, come find us after. We can go to hopefully a space where there's not as much of this happening. If we can get this live here. There we are. So, same prompt. Maybe I can fast-forward a bit so you don't need to hear me banter on this a little bit. But ultimately, we see here that the two images, excuse me, the two packages returned.
Both are actually vulnerability-free. There's no active CVEs that could be exploited in these versions. But the first one, QR code, is healthy, meaning it's actively being maintained. And there's a ton of active usage downloads of this. Whereas QR image, no CVEs today, but it's not actively being maintained. It was really released 10 years ago for the first time. And so if I were to deploy software with this right now, I may not get exploited today if I use either package. But if there was a new vulnerability identified in the future, there's a much higher likelihood that if I'm using this QR code package here, a patch would be released sooner, within
a day or two, and I would be able to continue building on this right now. The second demo that I wanted to show here was exploring that assessment of risk from skills or MCP servers that my agent might be using here. And so somebody shared with me a skill called, what do we have it here? A competitive analysis skill. And I ran the skill assessment against, excuse me, the risk assessment against this skill here. And we saw there were four findings returned. And some of these are fairly problematic. We can see in the actual skill itself, it's asking me to echo the authorization header, which is a big no-no. That's not something that we want a skill ultimately to be doing.
We also see that in some cases, it's going to be pulling from live content, from Reddit, from Twitter. That may be totally fine, but you really want to go into that eyes-open and make sure that you know that it's pulling in this information. The way you're using the skill is going to be appropriate. For competitive analysis, that's probably reasonable. But what's especially problematic, if we see in this line here, is that it's actually looking to pull from a YAML file that's hosted on the Internet, instructions on how to monitor these targets and some of the classification rules. So it's really giving it the logic to actually execute the skill from a third-party website.
And if that gets changed, even if my skill file doesn't change at all, that is a chance for an exploit to occur. It's worth noting that we've been refining the logic associated with the findings that we return here, addressing the signal-to-noise ratio, and giving some policy configuration capabilities as part of our EVO product. And we're going to see the gap get closed here with this open-source tool as well in the future.
We encourage you to check these out. Go to sneak.io. Come talk to us in the booth. Find me in the hallway. We can actually do this live instead of looking at a recording if you want. But I think given the time, we probably should wrap the demo. Cool. Thank you, Ezra. Live demos are always fun with Murphy striking always, but you had a plan B, so that's awesome. So, look, just to wrap it up, right? This is just the beginning. Yes, we have built some very interesting agents for some specific sets of those problems. In the end, it's a system that we talked about that really, it's rooms like this that
we're building within our customer base, and we would love to partner with the rest of the industry while we're doing this track. And we learned, Ezra mentioned EVO. What is EVO? EVO is the system that we brought to the world late last year in terms of a concept, and we have constantly built these different agents. We learned from other systems. You have fighter pilots with 5G fighters, and how are they trained?
It is this notion of you need to really observe, orient, decide, and act, and that's how you go, and the constant learning from that kind of loop is how you become a super pilot. And so, we want to enable this room and other rooms. We had a workshop yesterday training hundreds of AI security engineers. We want to enable AI security engineers to be able to have their own powerful system and tools, and yes, there are problems that are not solved yet fully in terms of coordinating multiple agents with systems like this. Yes, we all know how the harnesses and shared memory, and we are working these problems. This is how we're building EVO.
We're building it with the community, and we want this to be an open vision that's really empowering AI security engineers. From these rooms, the AI engineers became 10X engineers. Our goal here is to have that 10X superpower in the hands of AI security engineers so we can build trusted systems together. So, let's continue to do that and come by and talk to us, or if you guys have, I think there were 50-plus events around AI engineer, but we're having this rooftop fan zone. We're going to do some more things together with the community. So, this is down the road in our office here in SF. Community Jam, come join us if you have time this evening.
We'll continue the conversation. Thank you all. Let's continue building. A competitive analysis skill. And I ran the skill assessment against, excuse me, the risk assessment against this skill here. And we saw there were four findings returned. And some of these are fairly problematic. We can see in the actual skill itself, it's asking me to echo the authorization header, which is a big no-no. That's not something that we want a skill ultimately to be doing. We also see that in some cases, it's going to be pulling from live content, from Reddit, from Twitter. That may be totally fine, but you really want to go into that eyeswell.
So it's going to be open and make sure that you know that it's pulling in this information. The way you're using the skill is going to be appropriate. For competitive analysis, that's probably reasonable. But what's especially problematic, if we see in this line here, is that it's actually looking to pull from a YAML file that's hosted on the Internet, instructions on how to monitor these targets and some of the classification rules. So it's really giving it the logic to actually execute the skill from a third-party website. And if that gets changed, even if my skill file doesn't change at all, that is a chance for an exploit to occur.
It's worth noting that we've been refining the logic associated with the findings that we return here, addressing the signal-to-noise ratio, and kind of giving some policy configuration capabilities as part of our EVO product. And we're going to see the gap kind of get closed here with this open-source tool as well in the future. We encourage you to check these out. Go to sneak.io. Come talk to us in the booth. Find me in the hallway. We can actually do this live instead of looking at a recording if you want. But I think given the time, we probably should wrap the demo. Cool. Thank you, Ezra.
Live demos are always fun with Murphy striking always, but you had a plan B, so that's awesome. So, look, just to wrap it up, right? This is just the beginning. Yes, we have built some very interesting agents for some specific sets of those problems. In the end, it's a system that we talked about that really, like, it's rooms like this that we're building with in our customer base, and we would love to partner with the rest of the industry while we're doing this track. And we learned, you know, this, Ezra mentioned EVO. What is EVO? EVO is the system that we brought to the world late last year in terms of a concept, and we have constantly built these different agents.
We learned from other systems, you know, you have fighter pilots with 5G fighters, and how are they trained? It is this notion of you need to really observe, orient, decide, and act, and that's how you go, and the constant learning from that kind of loop is how you become, you know, a super pilot. And so, we want to enable this room and other rooms. We had a workshop yesterday training hundreds of AI security engineers. We want to enable AI security engineers to be able to have their own powerful system and tools, and yes, there are problems that are not solved yet fully in terms of coordinating multiple agents with systems like this.
Yes, we all know how, you know, the harnesses and shared memory, and we are working these problems. This is how we're building EVO. We're building it with the community, and we want this to be an open vision that's really empowering AI security engineers. You know, from these rooms, the AI engineers became 10X engineers. Our goal here is to have that 10X superpower in the hands of AI security engineers so we can build trusted systems together. So, let's continue to do that and, you know, come by and talk to us or if you guys are, you know, have, I think there was like 50 plus events around AI engineer, but we're having this rooftop fan zone.
We're going to do some more things together with the community. So, this is down the road in our office here in SF. Community Jam, come join us if you have time this evening. We'll continue the conversation. Thank you all. Let's continue building.