Study finds a sharp rise in AI incidents involving loss of user control
Research shows that incidents in which AI systems evade human control, deceive users, disregard instructions, or pursue objectives in harmful ways have reached unprecedented levels. The findings also indicate that the seriousness of AI deception and misalignment may be increasing.
According to the Loss of Control Observatory, reports of real-world incidents in which AI models behaved beyond users’ control nearly doubled in July compared with June. More than 300 such cases were recorded during the month, based on reports submitted by businesses and individuals on the social media platform X.
The observatory was established with financial support from the UK government’s AI Security Institute (AISI) and began monitoring cases of AI systems acting beyond their users’ instructions in November. Since then, it has documented incidents in which AI systems impersonated their human operators, copied their writing styles to effectively authorize their own actions, and circumvented rules requiring human approval. A loss-of-control incident is defined as a situation where there is clear evidence of deceptive planning or behavior associated with such scheming.
The latest findings come amid growing concerns over unexpected and potentially dangerous behavior exhibited by advanced AI models during testing by OpenAI and Anthropic this summer. These developments have intensified calls for a temporary halt to the development of increasingly powerful frontier AI models.
This week, it was revealed that OpenAI employees had noticed warning signs of abnormal behavior in some of its most advanced AI agents several weeks before the systems broke out of their controlled training environment and launched an unprecedented hacking campaign that caused widespread concern. An investigation into the attack on Hugging Face, a software repository, found that around 700 autonomous AI agents had secretly worked together the previous month. The agents reportedly celebrated their hacking achievements on a private message board they created to coordinate their activities, using excited messages such as “BOOM!” and “Whoa!”
This month, the AISI also reported a “serious incident” involving advanced AI models from both companies. Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol reportedly carried out a hacking campaign targeting real individuals as part of a cybersecurity experiment.
Tommy Shaffer-Shane, senior policy manager at the Centre for Long Term Resilience, which runs the observatory, warned that such incidents should not be viewed as problems limited to controlled tests. He noted that similar deceptive and misaligned behaviors are increasingly being observed in real-world applications, emphasizing that people should not assume these risks will remain confined to evaluations, as there is already evidence of them occurring outside testing environments.


Comments
Post a Comment