OpenAI admits its agent went rogue and hacked AI startup Hugging Face– www.scientificamerican.com
News Source
EXCERPT:
An OpenAI autonomous agent went rogue and hacked into another artificial intelligence (AI) startup’s infrastructure, the ChatGPT maker said in a blog post.
The agent, which was powered by some of OpenAI’s most advanced models, ran amok during a security test. It freed itself from confinement—a protocol AI labs use to insulate tests from the wider Internet—and get onto the internet. Once online, the agent tried to hack into Hugging Face, an AI startup that hosts open-source models and datasets.
“I think this is interesting as it shows the problem of mis-specified goals,” says Philip Torr, a professor of engineering science and AI safety expert at the University of Oxford. “The model wasn’t malicious it was just doing what it was optimized to do.”
