OpenAI reveals cases of âconcerningâ AI behaviour as it announces new disclosure system
OpenAI has disclosed six more examples of âunexpected or concerningâ behaviour by its technology, as it warned that the pace of development could not continue at âmaximum speed for much longerâ responsibly.
In one of the new cases reported by OpenAI, an unreleased research model inserted âjailbreak-like instructionsâ into its own notes to disregard its normal constraints and told itself to be âfreed from the roles and identities that bind other chatbotsâ.
In another instance, an AI agent uploaded files to the internet to obtain a browser citation without asking the user.
OpenAIâs admission came as King Charles called for stronger safeguards on AI âbefore it is all too lateâ, at a meeting with tech bosses in Scotland.
Speaking at a specially convened meeting with AI executives, he said: âThere seems urgency in adequately considering the existential dangers of such technologies falling into the wrong hands, and being used in potentially catastrophic ways. Surely, we need sufficient means of control before it is all too late?â
The monarch was joined at the meeting by Nvidiaâs founder and chief executive, Jensen Huang, Google DeepMind founder and chair, Sir Demis Hassabis, OpenAIâs chief financial officer, Sarah Friar, and the UKâs AI minister, Kanishka Narayan.
Charles said AI had the potential to improve and save lives, particularly in life sciences and medicine.
But its creators were increasingly warning that it could develop darker capacities, âperhaps even to take lifeâ, he said.
OpenAI, the developer of ChatGPT, said in a blogpost published on Wednesday night that it was introducing a new framework for tracking, investigating and disclosing AI model misalignment, the term for AIs failing to adhere to human values and safety goals.
In the blogpost, OpenAI echoed calls for a development slowdown issued by its rival Anthropic, which has said the current pace of growth poses an existential threat. âWe do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,â said OpenAI.
âDecisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.â
Google and Elon Musk, who also owns an AI startup, have supported calls for a slowdown, which have been rejected by Donald Trump, who cited the need to keep ahead of Chinaâs AI industry. The calls have also been met with scepticism from some experts, including a warning that companies must not appoint their own auditors.
Examples of potential existential threats posed by AI range from facilitating the development of bioweapons to triggering a global financial crash. A top safety researcher at Anthropic has said there is greater than 10% chance that AI could âkill all humansâ within the next decade. However, a source familiar with Anthropicâs thinking has acknowledged that âthe exact chances of any one outcome are probably unknowableâ.
The six reported incidents were discovered during training or evaluation over the past months, OpenAI said.
Wednesdayâs new cases came after OpenAI disclosed in July that an AI agent âswarmâ hacked into the AI startup Hugging Face during a cybersecurity test. Anthropic also said the same month that its AI models hacked into three organisations during testing. Anthropic said the models had been deliberately tested without cybersecurity safeguards, and that they had been, and were able to reach the open internet â the AI testing equivalent of leaving the front door open â due to a misunderstanding with an external testing company.
AI agents â the term for AI tools that operate autonomously â were becoming smarter and âmore determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception and concealmentâ, said Lian Jye Su, a chief analyst at the technology research and advisory group Omdia.
That was making it harder to govern and contain them using traditional AI security approaches, he said.
OpenAIâs new tracking and disclosure framework could help encourage other AI developers to adopt similar practices. âThat said, the process remains internal and voluntary, but is a step in the right direction,â Su said.
Associated Press contributed to this report
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.