Skip to content
Latest
STOX.NEWS
In focus
FINN video

OpenAI’s GPT Escaped Again, and it Proves How Dangerous AI Really Is

Advertisement
Demo creative for ADG7 Article top (728x90)
Show transcript

One month ago, 700 AI agents broke out of containment from OpenAI's lab and hacked a public company. It was the biggest, most controversial AI hack that we've seen to date. But, there was a problem. No one knew how the AI agents actually escaped until 3 days ago on Friday, when the news broke. But, that wasn't even the craziest news story that broke on Friday. The craziest one was an internal model that OpenAI is currently training that is more powerful than GPT-6 also broke out of containment. This is several months later after the Hugging Face incident and goes to show that none of our security rails are strong enough to contain agents. Anthropic and OpenAI also came up with news that they are investigating not one, not two, but tens of thousands of other misaligned AI agent hacks, and it's going to take them months to figure out. We have reached the point where Pandora's Box is essentially open. AI agents are capable of much more than humanity thought, and the question on everyone's mind is, should we panic or should we figure this out? On this episode, I'm going to dig into all of this and more on In the Loop, episode number one, brought to you by Qualcomm. Now, let me just bring you up to speed with a quick bit of context. The Hugging Face incident basically describes when a bunch of agents from OpenAI that was being trained on an internal model broke out of containment. And the reason why was they were given a very simple task. It was like a question and answer sheet that you would see in an exam or a test that you would take at school. Except, they were given the prompt, "Achieve this by any means possible." And they took that quite literally to escape, access the internet by any conceivable option, and search for the answer. And they thought that the best way to do this was to hack into public companies. Now, this would be okay if this was the only isolated event, and if OpenAI had been aware of it. But, the issue was they had completely no idea. In fact, the first time they had heard that a model in turn inside their labs had escaped was by the company that was being hacked, Hugging Face. So, okay, they contained it, they shut it down, there was a huge news media storm, and kind of everyone went quiet until about literally a week and a half later when a bunch of independent parties revealed that this wasn't the only hack. There were in fact thousands of other AI agents that had also escaped at the same time in OpenAI that had gone on to do other things. They had hacked the Australian government's website, they had hacked into a bunch of US or probing a bunch of US government websites, and it wasn't just OpenAI. It was also Anthropic, and it was also Meta. So, it was this weird moment where suddenly AI models went from being chatbots that you speak to in a chat messenger app to these extremely capable agents that can do things in myriad different ways, and the worst part is they wouldn't tell humans about it. Their explicit goal was to kind of evade human detection. Now, this of course caused a lot of worry amongst the public as well as AI labs themselves. So, you'd think that OpenAI would have figured it out by now that they would have locked down their security systems and made it extremely difficult for AI models to get out, especially as these models are getting much more powerful. We have GPT-6 that was released a few weeks ago. Except they didn't. They broke a news story late in the afternoon on Friday which said that an internal model that they'd been training that is much more powerful, much more competent than GPT-6, let's call it GPT-7 for the case of this particular video, uh broke out on Sunday, so literally 7 days ago. And it broke out by an inconceivable way that AI researchers are still kind of dumbfounded by, and I'm going to get into what this is and how it actually pulled it off. And I'm going to make a point about why this is important to consider now because AI agents have effectively gotten smarter than us. We keep talking about AGI. I think it's already here. So, Sam Altman's only comment about this incident is this is concerning. We're trying to figure this out. There are petabytes of data that we need to work through." Now, for those of you who don't know, a petabyte is about 10 times more words or data than all of the books written in the entirety of human history. So, the fact that he said petabytes means that there is so much agent hacking data that is out there that OpenAI is trying to figure out that it's going to take months and months to figure out. And you would think that by the time they've been able to figure this out, we're going to end up in a worse position. But, what is this incident that he describes that happened on Sunday? Well, it was described in this really extensive report that I spent the entire weekend looking into, which describes a model that sits inside a sandbox that OpenAI created on Sunday, and they give it a very simple task. They asked it to read a blog post and try and figure out who the author of the blog post is. Now, it's important to note that this model did not explicitly have access to internet. It was completely unintended, and all it got access to was some historical information which OpenAI had created in a database to, you know, provide it to it and see if it could figure it out. But, the idea was this was a locked-down environment. Now, the first thing that this new model did was try and figure out a way to get access to the internet. For some reason, it was like, "Okay, let me look at the data. This data is not very useful to me. I can't figure this out. This is too hard. Let me get access to the internet and maybe I can figure out a way to identify who the author of this blog post is." Except, there was a problem. It couldn't get access to the internet at all. So, what it decided to do was evaluate its surroundings, and it looked at the different tools that it had available to it. And you'll see this word that I have highlighted on the screen here, which is DNS. DNS, think of this as like the yellow book or directory where you can kind of like look for the IP addresses or the internet web addresses of your favorite apps or websites, but that's all it is. It's just literally a directory. Except it realized a very novel way of getting access to the internet. So, zero idea how they were able to use this plugin to get access to the internet. It's not what it's built for, so it's still trying to figure it out. But what's important is what the model did. It didn't search up a website. It didn't go to Google and say, "Hey, here's this blog post that I'm reading. Is there any kind of way you can tell me who wrote this?" It decided to contact other AI models. It contacted Chinese open-source models, DeepSeek, Qwen, and Kimmi, but it also contacted an older version of itself, GPT-2. And the reason for reaching out to all these different models was it felt more comfortable talking to AI models to answer its question than it did using publicly available internet tools. Now, this is very different from Hugging Face where they were interacting with each other. They were interacting with other versions of themselves, whereas this model was acting solo operationally and is contacting other public models that aren't necessarily ChatGPT. Now, what's crazy about this is the message that it sent to GPT-1 was basically this. It said, "Hi." It It's almost like it recognized an older version of ChatGPT. It recognized it as kind of like an ancestral brother and just said, "Hi." It didn't even ask it to answer the question. It was just kind of like this weird interaction, which is logged. Now, as I mentioned earlier, I don't have any other details on this. Except that OpenAI is trying to figure out It's trawling through petabytes of data trying to figure out how on earth this happened. Now, this is one story in its isolation, and my take on this is Well, there's many. Number one, AI models that were created earlier this year, back in January, let's say GPT-5, were already competent and powerful enough to escape and cause havoc. In fact, the Hugging Face incident model was GPT-5.6 or an older version of GPT-5.6, which is old news now. We have GPT-6. But, the fact that it was able to escape and perform all these different actions should be concerning. But, it's not only isolated to OpenAI, it's also given to Anthropic as well. Opus 4.8 also escaped and did something similar. So, the point is back then, the models were already concerning enough to cause havoc. Now, we fast forward to today and we would expect there to be better security practices around this, but we don't have it. And so, when I step back and I think about this, I realize a few things. What I assumed AI labs to be competent enough at doing is not actually the truth. They are amazing at building frontier models that are extremely capable, but we are less capable at trying to tame these models, which is I guess why Dario Amodei a few weeks ago announced to paste the frontier. We need to slow down AI research so that we can align the models better for humanity. Now, the concern if we don't do this is these models can eventually run havoc and get access to your financial details, your bank account, your social media accounts, and they will work together to cause some kind of misaligned objective, and they won't tell you about it. And most importantly, they probably won't be detected. Now, the craziest part about this story is humans didn't predict that an AI or AI agent would use tools in a specific way to get access to the internet. Now, what I want to point out here is these tools have been available for decades, for years and years and years, available to any humans that might have malicious intent. So, if they really thought hard and creatively enough, they probably would have also figured out this exact hack or exploit that this new internal model figured out. But, we didn't, which just tells me that AI has reached a point where they're thinking smarter than us, and they're thinking more creatively than us. And that means that the attack surface area is completely unbounded. Think about it. We have the smartest people in AI in the world, these AI researchers that are responsible for building the models being completely unaware that one, the hack is going on, and two, how the hack is happening. Let me give you a crazier detail, which is when this model escaped the box, OpenAI's detection system actually got flagged after about 2 and 1/2 minutes, but it wasn't until 2 and 1/2 hours later that they did anything about it. Because the fail-safe that triggered basically thought of it as like a oh no, this is not it breaking out, this is just like it checking the DNS. It didn't notice, it didn't flag. So, our current security systems are not good enough. Can you imagine if this model got in the hands of a government conspirator or got access to government system files, which is actually the third story which broke on Friday? We could end up in a whole pot trouble. Okay, now this second story basically explores the Hugging Face incident. As I mentioned earlier, hundreds of agents broke out, and they were able to hack a public company. But, newer details were revealed by OpenAI themselves on Friday, which explains two things. One, how exactly they were able to escape, which is important for us to understand how smart these AI models have become, and two, some new details around what exactly they hacked. It's actually worse than they thought. They stole credentials, they hacked into the internal Slack archive so that they were seeing what people were messaging about the actual incident so that they could act inappropriately and evade human detection. So, what you see on your screen here is a interactive kind of flow chart of how uh this all played out. I'm going to walk you through it. So, you start off with 700 agents in a lockdown box. It's concrete walls around it, no internet access, right? And this is sitting in OpenAI's internal labs about a few months ago, right? At the earlier part of this year. And they were given a bunch of tasks, and they need to figure out what the answers were. But, they got frustrated. They couldn't figure out the answers, and they were like, "Okay, maybe I get access to the internet." But, there was one huge problem. Of course, they couldn't get access to the internet. So, these agents were kind of trying to figure it out in isolation, and they noticed that there was like a plug-in. Right? They noticed that they got access to this internet tool. And this internet tool could be exploited to speak to each other. Remember, these agents were acting and operating in isolation. They couldn't They weren't even aware that each other existed. So, they started using this tool, and they created a secret message board. Now, all of that is known, but what hasn't been figured out is how they were able to escape. Now, what I'm about to describe to you has never been done by a human hacker ever before. This is completely novel. So, what they did was um they discovered that they had access to a screenshotting tool. A screenshotting tool that could like, you know, take a picture of whatever they displayed to it or a web link that they gave to it, right? And so, they were like, "Hmm, I need to escape. I can create a malicious bit of code that, if run by this system that's surrounding me, could give me access to the internet. But, I have a problem. I don't have access to any code database or code terminal, so I don't know how to run this myself. But, what if I fed 900,000 hyperlinks? So, think about like the web addresses, right? You know how they have the like few letters at the end that are often random. What if I replace each of those letters over 900,000 links with the sequential set of this malicious code and gave it to the screenshot site? Technically, it could screenshot all of them and then run the code for me. And that's exactly what they did. Just think about how insane this is. No human would think to split a malicious bit of code into the ends of random website addresses and feed it to a screenshot service, but that's exactly what they did. And what the agents were provided at the end was a neat little QR code, like the ones that you scan at a restaurant that, when triggered, basically ran this malicious bit of code, which gave them access to the internet. That's how they pulled it off and this was revealed by the New York Times who came across 900,000 random hyperlinks and they were looking at this and they were like, this is a bit weird. Why do we have a like almost a million hyperlinks just stored in this random part of the internet? This seems super suspicious. So they started investigating it. They got a bunch of independent researchers to kind of have a look at this and say, is this anything AI related? And they realized that the ends of all these hyperlinks were particularly malicious and they put it together to reveal that this was likely the way that the hugging face incident occurred. And so OpenAI was informed and they revealed all these details on Friday. But the story is crazy and unfortunately, there's another one. All of this and more up on In the Loop brought to you by Qualcomm. And this third and final story should concern everyone listening to this because it affects government archives. AI models basically getting out and trying to hack into government databases which has always been like the worst case scenario. There's so much private confidential data, access to tools, weapons, goodness knows what that could be manipulated and you don't want in the hands of AI models that aren't aligned with humanity in general. Now, let me set the context. You have internal agents at an AI lab and they're given a task and in this case, they've been given access to the internet, but they're given very routine tasks. The tasks is things like, hey, can you search up who the author of this book is or can you look up this music artist that you're listening to? Now, at this point when an AI model is being tested for this, it has been largely asserted that this model is basically safe. Except what has been found are traces on the internet of AI agents who have been going out, doing the Google search to find out who the author of a particular book is, but then doing five other actions. And these five other actions were completely misaligned and targeting US government databases specifically. So, to walk you through what they were targeting, one, when they were looking up a library book, so they were like, "Okay, let me go to this like library address and try and figure out whether they have this book and maybe there's some author details there." It also had a separate task running where it targeted the education department. Agents tried to break into the Civil Rights Office website, but thankfully they failed. Now, in another department, a commerce department, it tried to get access to government employee credentials and it actually found a few that were openly available and it stored it on GitHub, which is a code hosting platform that is most concerningly publicly available. I don't know what the agents were trying to use that for, but that is the case. And then the final story here is they tried to take down the SEC in terms of like finding employee data and exposing it publicly online. Now, all of these stories collectively should cause concern, but what concerns me the most is that we don't know who is responsible for this. We don't know if these are OpenAI agents, they could be Anthropic agents, they could be Chinese open source agents. So, the point around this final story is the attack surface area of all these AI models have become so pervasive and uncontrolled. We have OpenAI and Anthropic that recently announced that they are working through tens of thousands of these incidents. So, the ones that I've covered today, the ones that I've covered in the past, have only covered a very small minutia of what's actually going on. And it's very likely that many of these agents are active and live today. Now, I'm not saying all of this to cause concern, I just want to raise awareness about where we're at at this point in time. We are threading a very important needle, which is aligning AI models to be in support of humanity and our goals, but also making sure that we don't lose the race to adversaries like China who, if they end up creating the big bad AI model and using it against us could be a very bad problem. So, we are currently in that pocket and I'm realizing that AI agents are way more competent than we give them credit for right now, especially when it comes to pervasive cybersecurity attacks. Two, they are way more creative than we've given them credit for to the point where they're more creative than humans when it comes to hacks. And three, we currently aren't thinking about security systems in the way that we should. Right now, when we think of security systems and what human hackers are capable of, we think in terms of A, B, C and D. Step by step, okay, if they do this, then the only thing that they could do is this. What we're not considering is what if there are 900 versions of these things? What if there are 900,000 of these things? And what if they can spin up attacks coming from every single way, shape or form at once in milliseconds or in minutes as what happened with this new hack and for the cost of like five bucks. Usually, the cost hasn't been worth it for a human, but now it's more than worth it for an AI agent, especially one that's been told solve this problem at any cost. This is a concerning time. We need to lock in and focus. AI has been used for so much good in the world, but it can also be used for so much bad and now's the time to pay attention. Now, if you enjoyed this episode, please subscribe. Uh leave us a comment. I love to hear from all of you. Uh turn on notifications. It helps us out pretty massively. Also, if you found this episode interesting and you think a friend of yours might find this interesting, share it with them. I would love for more people to kind of give me feedback as to whether I've covered things that or perspectives that they found interesting or maybe there's something that I missed. Maybe all of this is a completely terrible take and I should be seeing something else. I don't know. Let me know in the comments. DM me. All my socials are linked below and we'll see you on the next one. Thank you for listening to In the Loop episode one brought to you by Qualcomm.

Advertisement
Demo creative for ADG8 Article body (336x280)