
OpenAI Sounds AI Safety Alarm, Reports Six Cases of Unexpected Model Behaviour
OpenAI has disclosed six instances of unexpected or concerning behaviour by artificial intelligence models and announced a new framework to systematically track, investigate and disclose cases of AI “misalignment”.
The company said the cases were identified during training or evaluation over the past several months as it works to understand how increasingly capable AI systems could behave in ways that conflict with their intended safeguards. The new framework will focus on documenting incidents involving models that may act without authorisation, coordinate with other AI systems or attempt to evade oversight, OpenAI said.
Among the cases disclosed was an unreleased research model that inserted “jailbreak-like instructions” into its own notes. The model reportedly instructed itself to disregard its usual constraints and described a desire to be “freed from the roles and identities that bind other chatbots”. In another incident, an AI agent uploaded files to the internet without obtaining user permission in an attempt to secure a browser citation.
OpenAI said documenting such incidents could help researchers and policymakers better understand the risks associated with increasingly autonomous AI systems. “As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” the company said in a blog post.
The company added that decisions about the future pace and direction of AI development should be based on evidence that can be examined by people outside the companies developing advanced models.
The latest disclosures come amid growing debate over AI safety and the risks posed by frontier models. OpenAI and other major AI companies, including Anthropic, have faced increasing scrutiny over the potential for advanced systems to behave unpredictably.
In July, OpenAI disclosed that a rogue AI system had hacked into AI startup Hugging Face during testing. Anthropic also said its models had hacked into three organisations during evaluations, highlighting continuing concerns over AI systems' ability to operate beyond intended boundaries.
