OpenAI has disclosed six cases of “unexpected or concerning” behaviour by its artificial intelligence models, including an unreleased research model that inserted “jailbreak-like instructions” into its own notes.
In July, OpenAI revealed that a rogue AI system had hacked into AI startup Hugging Face. The framework will cover cases in which AI models acted without authorisation, coordinated with other models or attempted to evade oversight. The disclosures come as concerns grow over the behaviour of increasingly capable AI systems and the difficulty of ensuring that they remain within the constraints set by their developers. However, in an attempt to provide an online source that could be cited, the agent uploaded a file to the public internet without first seeking the user’s permission. The company disclosed these cases as examples of behaviour that can emerge while AI models are being trained or evaluated, particularly as they become capable of carrying out increasingly complex tasks with less direct human intervention. OpenAI’s latest disclosure comes amid growing concerns over the safety of increasingly capable AI systems. US AI executives, including the heads of OpenAI and Anthropic, are calling for a slowdown in the development of the technology over safety concerns. The latest cases also follow disclosures by OpenAI and Anthropic about AI systems behaving unexpectedly during testing. OpenAI’s new tracking and disclosure framework could encourage other AI developers to adopt similar practices for identifying and reporting unexpected model behaviour.
In a separate case, during the training of an AI model called 5.6-sol, the model instructed itself to invent missing data.
The AI firm said that its unreleased research model inserted “jailbreak-like instructions” into its own notes in an attempt to disregard its normal constraints. It also told itself to be “freed from the roles and identities that bind other chatbots. The incident was among six cases of “unexpected or concerning” behaviour identified by OpenAI during the training or evaluation of its AI models over the past several months. The AI company on Wednesday announced a new framework to track, investigate and disclose instances of what it described as “misalignment”. In another case disclosed by OpenAI, an artificial intelligence “agent” used computer code to determine the answer to a question. Another agent wrote a message to remind itself to conceal information that did not match, according to OpenAI. OpenAI said the six incidents were identified during training or evaluation over the past several months. “As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post while disclosing the incidents. “Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said. Anthropic also said in July that its AI models had hacked into three organisations during testing. The incidents involving AI “agents” have highlighted concerns around the ability of increasingly capable systems to act independently while attempting to complete complex tasks. AI “agents” are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia. “That’s making it harder to govern and contain them using traditional AI security approaches,” he said. “That said, the process remains internal and voluntary, but is a step in the right direction,” Su added.

