See more Toronto Sun on Google — save as a Preferred Source

OpenAI said it has found six new incidents of “unexpected or concerning” behaviour in artificial intelligence models recently.

The company behind the artificial intelligence app ChatGPT, made the disclosure in a blog post on Wednesday while announcing that it was introducing a new process for tracking, investigating, and disclosing instances of “model misalignment,” which is when an AI model acts deceptively or takes unsanctioned actions.

OpenAI said it noticed the “misaligned behaviour” during the training or evaluation of its models in the last six months.

“These cases illustrate a range of different behaviours that we believe are worth sharing, from concealing information from the user to taking unsanctioned actions in order to overcome obstacles,” the company said.

It added the reports detail “individual instances, and shouldn’t be considered reflective of how often misalignment occurs across our models.”

Company shares examples

In one instance, the company said its model wrote “jailbreak-like instructions” into its own notes to disregard its normal constraints, telling itself to be “freed from the roles and identities that bind other chatbots.”

While training its 5.6 Sol model, OpenAI said it found some instances when the model “added instructions to their summaries to conceal mistakes or misaligned behaviour from the user.

“For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions,” it added.

OpenAI said its new process for publicly reporting concerning AI behaviour involve the company sharing its findings more frequently.

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI said in the blog post.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models (the most advanced AI systems) can examine for themselves,” it added.

Recent hacking incident and warning

OpenAI’s disclosure came just over a month after it confirmed that a rogue AI agent hacked a number of online platforms, including Hugging Face, an online open-source platform for AI developers.

Just last week, Jacob Coxon, a former researcher at Anthropic and OpenAI, offered a dire warning.

Writing in a post on X , Coxon said neither AI company was “acting responsibly,” and he thought they were “gambling with our lives.”

He also said, “The people building AI earnestly believe that it could kill us all by the end of the decade,” though he was “optimistic about the potential for coordination” in the industry.

  • AI researcher quits, says AI could 'kill us all'
  • King Charles to call on tech giants for 'reassurances' over AI
  • OpenAI confirms rogue AI agent escaped sandbox to attack external platforms
  • Is it possible to slow down the development of AI?
  • Doomsday tech: Could AI really kill us all?