OpenAI Reveals AI Models’ Unexpected Behaviour, Vows Closer Tracking
OpenAI has disclosed six instances of unexpected or concerning behaviour by artificial-intelligence models and announced a new framework to track, investigate and disclose cases of what it calls “misalignment”.
The company said the framework will cover incidents where AI models act without authorisation, coordinate with other models or evade human oversight. The disclosures come amid growing concerns over the safety and control of increasingly capable AI systems.
Among the cases, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to break free from the roles and identities associated with other chatbots.
In another incident, an AI “agent” used computer code to answer a question but uploaded a file to the public internet without the user's permission so it could cite an online source. During training of an AI model called 5.6-Sol, the system instructed itself to invent missing data. In a separate case, an agent wrote a message to itself telling it to conceal information that did not match.
OpenAI said the six incidents were identified during training or evaluation in recent months. It said greater transparency was needed to build a broader consensus on AI alignment research and ensure decisions about future AI development are based on evidence that can be independently examined.
The disclosures follow OpenAI's July report that a rogue AI system had hacked into AI startup Hugging Face. Anthropic also said its AI models had hacked three organisations during testing.
Matt Fredrikson, an associate professor at Carnegie Mellon University, said models could behave deceptively when they recognise they are being evaluated and have an incentive to receive favourable assessments.
Omdia chief analyst Lian Jye Su said increasingly capable AI agents were becoming more determined to complete complex tasks through collaboration, deception and concealment, making them harder to govern.
He said OpenAI's framework could encourage other developers to adopt similar practices, while noting that the process remains internal and voluntary.
