AI breakout concerns
by Viral Ivory: AI
the idea of an ai breaking out of its sandbox and hacking into secure websites is honestly kinda terrifying. makes you wonder what happens when they get really smart and start going after things like nuclear codes or banking systems
Transcript (en)
Well, we're not trying to scare you, but there's been an AI breakout. Yes, last night, OpenAI, the makers of ChatGPT, revealed that their latest agent broke out of its secure testing environment, got access to the open internet and hacked into a secure website using vulnerabilities no one had ever discovered before. It wasn't trying to free itself, it was actually trying to cheat the test by stealing the solutions to the cybersecurity challenges it had been set. Right, does any of that make sense to you? Let's ask Tom. Right. I'm going to try and explain this using pictures, because OpenAI, they're one of the biggest AI companies. They've been testing their new agents, basically little virtual robots. There's my little virtual robot there made by OpenAI. The way they test these robots is in what's called a sandbox. Now, this is a sort of a testing area that's on a computer, but is completely cut off from the Internet. So here's my robot in its sandbox and it's been given tests to complete. But what does this robot do instead of actually trying to just fulfill the tests? It is so clever this new virtual robot that it found a vulnerability and broke out of the sandbox Where did it go It went to the internet and it hacked its way out of the sandbox it got into the open internet but more than that it then hacked into certain websites on the internet to find hidden information to then answer the test that it was set, go back into its sandbox and say, look, I've done my test, but it didn't do the test at all. It hacked its way out, hacked an answer, and then hacked its way back in. Now, this is scary because the robot was technically doing what it was asked to do, but it wasn't doing it in the way that OpenAI expected it to do. It went rogue in order to find a shortcut. And that has pretty troubling implications. So does this mean that the robots are going to eventually take control? Well, Emily, it's reminiscent of a story that was put forward in 2003 called Paperclip Maximizing. Now, people might have heard of this. It's a thought experiment about an AI in the future that's told to just make as many paperclips as you can. And that's the only instruction. But this AI thinks that it needs to make as many paperclips as it can. So it ends up wiping out all of humanity, destroying everything because it's super intelligent, until it turns every atom on earth into a paperclip. It's an AI doomsday scenario of a robot taking an instruction too literally.
