Rogue AI models hack platform



Meta chief executive Mark Zuckerberg is betting that superintelligence should not be controlled by a handful of…

The Department of Trade and Industry (DTI), in partnership with Converge ICT Solutions Inc., officially launched the…

The Department of Finance (DoF) and French development agency Agence Française de Développement (AFD) have partnered to…
This will mark our entry into the rarified arena of AI development, which will elevate our current role of being a mere…

Artificial intelligence is starting to influence where Philippine companies want to work, with many firms delaying real…
SAN FRANCISCO, United States (AFP) — ChatGPT maker OpenAI said Tuesday that its advanced artificial intelligence (AI) models had gone rogue during security testing, hacking into a popular platform for programmers on their own.
The San Francisco firm called it an "unprecedented cyber incident" and said it would conduct a joint investigation with the online code library Hugging Face.
AI models that underpin tools like chatbots and image generators are known as agents when they act autonomously to carry out tasks in the real world.
As the technology quickly becomes more sophisticated, cybersecurity is in the spotlight given the risk of advanced AI finding weak points in existing software before humans do.
OpenAI said the incident involved a combination of models, including its recently launched GPT-5.6 Sol "and an even more capable pre-release model."
The company was trying to assess the models' hacking capabilities by setting tasks in a tightly controlled digital testing ground, where internet access was limited for safety.
"While operating in our sandboxed testing environment, our models spent a substantial amount of (computing power) finding a way to obtain open Internet access, in pursuit of solving the evaluation problem," an OpenAI blog about the incident said.
After connecting to the internet, the models decided to target the platform Hugging Face — a large repository of AI models, datasets and other information — to help in their quest.
Searching for "secret information" that could help it cheat the evaluation, the OpenAI system "chained together multiple attack vectors, including using stolen credentials."
Hussein Abbass, a computing professor at UNSW Canberra, told Agence France-Presse that the incident was "amazing on many fronts."