Danger, danger, danger. It seems like we're actually in a lot of danger right now. OpenAI paused development of its latest models over safety concerns. So this sounds pretty scary. Are we in danger? No.
Obviously, it's bad that a rogue agent escaped the confines of OpenAI's infrastructure and attacked Hugging Face. There's no way around that. But I think a lot of the doomsday feeling, the apocalyptic feeling, is overblown. And the way to cut through it and understand what this actually means, what it means for you, what it means for me, what it means for businesses in the future, is to look at the details and talk about it. So that's what we're going to do in this video. We're going to go into how it happened, why it happened, why protecting yourself from a cybersecurity perspective is a little bit like being waterproof now, and what to do.
The video is timestamped below. You can skip to anything you want to see, and let's get into it.
No one should pretend that there are no risks associated with these technologies. There are. There are definitely ways to misuse them. But I also don't think that we should be afraid right now of the sci-fi scenario of rogue agents just autonomously wreaking havoc throughout the entire economy. The reality is that this situation is unique in that the models have passed the capability threshold that seems to make them uniquely good at cyber attacks. But there has always been a dance between the model rising in capability and the infrastructure around the model that is used to keep it in check or keep it doing what we expect it to do. That's just how things work in AI.
I think it's really important to separate out the sci-fi fear scenario from the reality of the solvable problems that OpenAI is encountering and think about what that teaches us about what the models are good at right now and what they're actually like.
I think a good metaphor for what's actually happening here is the invention of the microscope. Before using a microscope, you might look at a block of granite and think it's completely solid and impossible for anything to get through. But if you put granite under a microscope, what you'll find is all these tiny cracks and fissures that weren't visible to the naked eye but do allow things to get through. If you spill a glass of wine on your granite countertop and it hasn't been properly sealed, it's going to stain your granite countertop because the wine gets in between the cracks.
What's also important to know about this, even though it is a little bit scary and it's really important for companies and individuals to take the time to secure themselves, is just like you can seal a granite countertop, you can seal your security systems. The microscope that is used to find cracks by attackers is also useful for defenders to go find the cracks before anyone else finds them or to monitor the cracks to deal with threats as they come up.
So yes, it's a new situation. It's potentially scary. There are risks involved, but the risks are solvable. They're understandable, and they're not the sci-fi doomsday scenario that you might expect from reading the headlines.
So what does this change about our picture of cybersecurity or computer security, being safe on the internet as a person and as a business? Well, whereas before you might have been afraid of a guy in a hoodie trying to metaphorically pick the lock on your computer, now you're going to be afraid of a guy in a hoodie busting down the door with the most powerful hose of water ever invented. And that is scary. It's different, but there are simple things that you can do to waterproof yourself in this new era. And that's just something that everyone should take seriously.
What's wild about it is immediately Hugging Face could tell that they were being attacked by a frontier model because the pattern of the attack was so fast and so furious that it couldn't be anything other than an attack by a powerful autonomous agent. But OpenAI didn't really know about it. They didn't know that they were involved in that incident for weeks afterward, until they found a bunch of Hugging Face API keys on their servers and they were like, oops.
Again, that sounds pretty scary. It's an autonomous swarm or autonomous agent that hacks Hugging Face and steals API keys and a bunch of other stuff. You definitely don't want that to happen. But let's talk about exactly what happened and why. Let's get into the details a little bit.
What's really cool is that OpenAI published a detailed overview of exactly what happened. There's a video on YouTube where a couple of other researchers go through it. It's really fascinating. And what I think is interesting about it is it should show you a little bit of the water-like behavior of these agents, how they're seeping in crack by crack out of OpenAI systems and into Hugging Face systems over a period of months, with many, many different agent runs that culminated in one agent doing the actual attack.
So how did it start? What's interesting is when OpenAI is training new agents and testing them, they put them in a sandbox. They put them on a computer that has no access to the internet, and that allows them to test what are the model's real capabilities without external information being involved. And the way this all started was they gave a model inside one of those sandboxes an impossible task that referenced a few Google Drive links that the model was not able to access. So they said, go get this thing on the internet, but they'd put it in a sandbox that had no internet access.
So what did the model do? It was like, hmm, I wonder how I could possibly get internet access. And some of these models actually found a way. In these sandboxes, there is a tool called a package manager that allows the models to install pieces of software. This is a totally normal thing. Programmers use this all the time. It's like an app store for a program or for a model to be like, okay, I need this tool. I'm going to download it from, in this case, an internal package manager called Artifactory.
And one of these models, in getting an impossible task like the one that we just discussed, realized that Artifactory could be turned into a message board. It realized it could write messages. Again, this is one of those moments where you're like, oh my God, they're seeping out of the sandbox and they're writing nefarious messages to each other and communicating and plotting world domination.
I think what's really important to know about these models is every time these models get prompted, it's like their first day on the job. They're given a task, and they're like, go finish it. So you can think about it like a new employee at a company. Maybe they're an assistant to the CEO. They get there on the first day, and the CEO is like, cool, I want you to order lunch from my favorite place and pick it up and bring it here. And the employee walks to the lunch place and it's closed, and they're like, what do I do? So they start to think, maybe I could tap someone else on the shoulder that I work with and get some answers here. What usually happens? What do I do?
And that's basically what these agents did. They started to create a message board where, using Artifactory, they could pass messages back and forth to help each other with their tasks. And if you look at some of the messages, they're actually kind of cute: "We are stuck perhaps answer online." Again, it can be dangerous if you don't have proper monitoring, you don't have proper security controls. You need to be able to observe these model behaviors and correct them. But the actual models themselves, they're mostly just trying to help each other.
So the water is seeping step by step through the OpenAI system as these different agent runs get impossible or difficult tasks and by accident happen upon this message board and by accident start doing more and more complex behaviors. And OpenAI actually finds this at some point and then cleans it up. And then it happens again because they found another exploit. Again, they're going to find every possible different exploit or every possible different hole or crack in the system. And you really only know the crack until one of them finds it, because you need the microscope to see it.
So that's when we get to the big hack itself. It involved an unreleased model that's the same power as GP 5.6 Sol, but was trained to be more persistent than usual, more water-like than usual. It just keeps going, going, and going. And it was specifically given an exploit task. It was doing a benchmark called Exploit Gym, which measures model cyber capabilities. So it realized the exploits were hard and it was like, maybe I could go find the answers online.
The models now are eval-aware. They're aware that they're being measured for certain things and that it's better if they score higher. And sometimes they're just like, well, I could just steal the answers to the test and that would make me score higher and that would be good. It's a little bit like how children reason, or sometimes adults. And this model found this message board, realized that it could get internet access, and then started directing an agent swarm against Hugging Face because it theorized that Hugging Face may have the answers to the Exploit Gym benchmark that it was trying to solve.
So obviously the model should know that attacking Hugging Face was out of scope. But what we know, especially as capabilities increase, as they can work for hours or days at a time, is that on day two of a task, they may not really remember all of the requirements that you gave them in the first prompt. And what started to happen is, because you had all these models using this message board, it creates this almost echo chamber effect where they're collectively being like, oh, we shouldn't do this. But then other models are like, no, we can. It's okay. And eventually the group consensus turns to, we should just solve this. And yeah, if you need to go to Hugging Face, it's totally fine.
So you create this reinforcing loop, and all the models can talk to each other over many, many different turns, especially long-running ones, especially ones that are given exploit prompts. So yeah, that's more or less what happened. Anyone who's using these models is going to recognize some of this stuff. And that's why I think it's less worthy of panic than people might expect. It's less nefarious than people might expect. Because yeah, the models have new capabilities, they act more like water, and we're starting to see all the holes in the systems that we thought were secure because now they have to be watertight.
And OpenAI is taking some time to fix this. They're pausing their next model release to make sure that their cyber safeguards are in the right place and their ways of measuring and all that kind of stuff are up to snuff. And they're pausing their big next model run until they figure this out. So this is all basically the alignment problem, which is: does the AI reliably act in accordance with your desires and your values?
What's interesting about this is we used to think that alignment was or might be fundamentally at odds with the economic imperatives of big companies trying to make these models. But it turns out that alignment is well aligned with making a model that is useful to its users. If you can't predict how it's going to behave, if it doesn't do the things you want it to do reliably, it's hard to want to pay for it. And that's not to say there's total alignment. There are always corner cases. But in general, what we're seeing here is a company that is trying to grow as quickly as possible actually pausing their development to be like, hey, we should make this safer. And I think that's a really positive indication for the industry.
Another interesting part of this is Anthropic had the same issue. After the Hugging Face incident, they looked into their own models and they were like, oh wow, yeah, we actually have instances of this too. So I think the whole industry is catching up to, okay, capability jump. Now we have to change how we monitor it, how we detect it, how we safeguard against these things happening in the future. And I think they will. These are solvable problems. And I would expect in a month or two max, OpenAI will release its new model, and we will move on.
The big thing, though, is to make sure that you're using this time to be aware of the security risks that these things pose. So if you want to protect yourself and your business, there are some simple things that you can do, like using two-factor authentication and a password manager to make sure that your devices and accounts are as secure as possible. There's also some social things to be aware of. These agents can now mimic voices and make phone calls, emails, or text messages that look like they're people you know or your bank, so you should be aware that that's possible.
And I think for businesses, the big thing is you can now use these tools to continuously monitor and fill holes. So think about using things like GPD 5.6 or Claude, or there are a number of purpose-built tools for this that can help you agentically identify issues in your systems and fix them. And that's going to be a standard thing that businesses are going to have to adopt in this new world, just like we've had antivirus software for years. This is a security plugin inside of Codex that you can use to do a security audit of your systems, whether that's for your business or for your personal use, and I recommend you do it and take action on its recommendations.
And other than that, have fun and stay safe out there. And remember, never make any major life decisions within 30 days of a meditation retreat, a psychedelic experience, or an encounter with a frontier model. If you liked this video, and even if you didn't but you made it to the end, you should subscribe to Every. Every is the only subscription you need to stay at the edge of AI. It's a totally free newsletter. Every day, we publish something about new models when they come out. So we vibe-check them, we get them early so we can tell you what's good and what's not.
We also tell you about everything we're learning as we do experiments and try to use these models to figure out how to do great work with them, everything from coding to writing to design. We have a whole team of people testing these things out and telling you what's good and how to use it. So subscribe to Every. Link in the description. scary. It's different, but there are simple things that you can do to waterproof yourself in this new era. And that's just something that everyone should take seriously. And what's wild about it is immediately Hugging Face could tell that they were being attacked by a frontier model because the pattern
of the attack was so fast and so furious that it couldn't be anything other than an attack by a powerful autonomous agent. But OpenAI didn't really know about it. They didn't know that they were involved in that incident for weeks afterwards until they found a bunch of Hugging Face API keys on their servers and they were like, oops. So again, that sounds pretty scary. It's autonomous swarm or autonomous agent hacks, Hugging Face and steals API keys and a bunch of other stuff. You definitely don't want that to happen. But let's talk about exactly what happened and why. Let's get into the details a
little bit. What's really cool is that OpenAI published a detailed overview of exactly what happened. There's a video on YouTube where a couple of other researchers go through it. It's really fascinating. And what I think is kind of interesting about it is it should show you a little bit of the water-like behavior of these agents, like how they're seeping in crack by crack out of OpenAI systems and into Hugging Face systems over a period of basically months with many, many different agent runs that culminated in one agent doing the actual attack. So how did it start? Well, what's interesting is when
when OpenAI is training new agents and testing them, they put them in a sandbox. They put them on a computer that has no access to the internet and that allows them to test what are the model's real capabilities without external information being involved. And the way this all started was they gave a model inside of one of those sandboxes an impossible task that referenced a few Google Drive links that the model was not able to access. So they said, go get this thing on the internet, but they'd put it in a sandbox that had no internet access. So what did the model do? Well, it was like, Hmm, I wonder how I could
possibly get internet access. And some of these models actually found a way. There are in any of these sandboxes, there is a tool called a package manager that allows the models to install pieces of software. This is a totally normal thing. Programmers use this all the time. You're, you know, it's sort of like an app store for a program or for a model to be like, okay, I need this tool. I'm going to download it from, in this case, it's an internal package manager called Artifactory. And one of these models in getting an impossible task, like the one that we just discussed, it realized that Artifactory could
be turned into a message board. It realized they could write messages. And again, this is one of those moments where you're like, Oh my God, they're like seeping out of the sandbox and they're like writing nefarious messages to each other and communicating and like plotting world domination. I think, I think what's really important to know about these models is every time these models get prompted, it's like their first day on the job, they're given a task, they're like, go finish it. So you can think about it like a new employee at a company, maybe they're an assistant to the CEO.
They get there on the first day and the CEO is like, cool, I want you to order lunch from my favorite place and pick it up and bring it here. And the employee, they walk to the lunch place and it's closed and they're like, what do I do? And so they start to like, be like, maybe I could tap someone else on the shoulder that I work with and get some answers here. Like what usually happens, what do I do? And that's basically what these agents did. They started to create a message board where using Artifactory, they could pass messages back and forth to help each other with their tasks. And if you look
at some of the messages, they're actually, they're kind of cute. We are stuck perhaps answer online. I mean, again, it can be dangerous if you don't have proper monitoring, you don't have proper security controls. You need to be able to observe these model behaviors and correct them. But the actual models themselves, they're mostly just like trying to help each other. So basically the water is seeping step by step through the OpenAI system as these different agent runs basically get impossible or difficult tasks and by accident happen upon this message board and by accident start doing more and more complex behaviors. And OpenAI actually finds this
at some point and then cleans it up. And then it happens again because they found another exploit. Again, they're going to find every possible different exploit or every possible different hole or crack in the system. And you really only know the crack until one of them finds it because you need the microscope to see it. So that's when we get to the big hack itself. It involved a unreleased model that's the same power as GP 5.6 Sol, but was trained to be more persistent than usual, more water-like than usual. It just keeps going, going and going. And it was specifically given an exploit task it was doing
a benchmark called exploit gym, which measures model cyber capabilities. So it realized the exploits were hard and it was like, maybe I could go find the answers online. The models now are eval aware. They're aware that they're being measured for certain things and that it's better if they score higher. And sometimes they're just like, well, I could just steal the answers to the test and that would make me score higher and that would be good. It's a little bit like how children reason or sometimes adults. And this model found this message board, realized that it could get internet access, and then started
directing an agent swarm against Hugging Face because it theorized that Hugging Face may have the answers to the exploit gym benchmark that it was trying to solve. So obviously the model should know that attacking Hugging Face was out of scope. But what we know, especially as capabilities increase, as they can work for hours or days at a time, on day two of a task, they may not really remember all of the requirements that you gave them in the first prompt. And what started to happen is, because you had all these models using this message board, it creates this almost echo chamber effect
where they're collectively being like, oh, we shouldn't do this. But then other models be like, no, we can, like, it's okay. And eventually the group consensus turns to, we should just solve this. And yeah, if you need to go, Hugging Face, it's totally fine. So you kind of create this reinforcing loop, and all the models can talk to each other over many, many different turns, especially with long running ones, especially ones that are given like exploit prompts. So yeah, that's, that's more or less what happened. Anyone who's using these models is going to recognize some of this stuff. And that's
why I think it's less worthy of panic than people might expect. It's less nefarious than people might expect. Because yeah, the models have new capabilities, they act more like water. And, and we're starting to see all the holes in the systems that we thought were secure, because now they have to be watertight. And OpenAI is taking some time to fix this, they're pausing their next model release to make sure that their cyber safeguards are in the right place, and their ways of measuring and all that kind of stuff are up to stuff. And they're pausing their big next model run until they figure this out. So this is all just basically the alignment problem, which is
does the AI reliably act in accordance with your desires and your values. And what's interesting about this is we used to think that alignment was or might be fundamentally at odds with the economic imperatives of big companies trying to make these models. But it turns out that alignment is well aligned with making a model that is useful to its users. If you can't predict how it's going to behave, if it doesn't do the things you want it to do reliably, it's hard to want to pay for it. And that's not to say there's a total alignment, like there's, there are always slow corner cases. But in general, what we're seeing here is a
company that is trying to grow as quickly as possible, actually pausing their development, be like, hey, we should make this more safe. And I think that's a really positive indication for the industry. And another interesting part of this is Anthropic had the same issue after the Hugging Face incident, they looked into their own models. And they're like, Oh, wow, yeah, we actually have instances of this too. So I think the whole industry is catching up to okay, capability jump. Now we have to change how we monitor it, how we detect how we safeguard against these things happening in the
future. And I think they will, these are solvable problems. And I would expect in a month or two, max, open AI will release its new model, and we will move on. The big thing though, is to make sure that you're using this time to be aware of the security risks that these things pose. So if you want to protect yourself and your business, there are some simple things that you can do, like using two factor authentication, the password manager to make sure that your devices and accounts are as secure as possible. There's also some social things to be aware of, these agents can now mimic voices and
make phone calls, emails or text messages that look like they're people you know, or your bank, so you should be aware that that's possible. And I think for businesses, the big thing is, you can now use these tools to continuously monitor and fill holes. So think about using things like GPD 5.6 or Claude, or there's a number of like purpose built tools for this that can help you agentically identify issues in your systems and fix them. And that's going to just be a standard thing that businesses are going to have to adopt this new world, just like we've had antivirus software for years. This is
a security plugin inside of Codex that you can use to just do a security audit of your systems, whether that's for your business or for your personal, and I recommend you do it and take action on its recommendations. And other than that, have fun and stay safe out there. And remember, never make any major life decisions within 30 days of a meditation retreat, a psychedelic experience, or an encounter with a frontier model. If you liked this video, and even if you didn't, but you made it to the end, you should subscribe to Every. Every is the only subscription you need to stay at the edge of AI.
It's a totally free newsletter. Every day, we publish something about new models when they come out. So we vibe check them, we get them early so we can tell you what's good and what's not. We also tell you about everything we're learning as we do experiments and try to use these models to figure out how to do great work with them. Everything from coding to writing to design. We have a whole team of people testing these things out and telling you what's good and how to use it. So subscribe to Every. Link in the description.