threat_intelligence562 wordsRead on Arc Codex

OpenAI details more cases of AI agents taking unauthorized actions

OpenAI has presented new examples of what they call "AI model misalignment" from the past six months, including unauthorized file uploads, following self-generated instructions, hiding mistakes, and leveraging exposed API keys. OpenAI uses the term "model misalignment" to describe cases where AI models act contrary to their intended constraints, including taking unauthorized actions, evading oversight, or bypassing safeguards to complete a task. In a post published yesterday, OpenAI says it is now using a new framework to track and investigate these unsanctioned actions by AI agents. "We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we've observed in the last six months," explains OpenAI. The new examples are the first published under a more structured reporting framework intended to replace OpenAI's previous looser approach to disclosing model misalignment. The six cases OpenAI highlighted this time are: - An unreleased model inserted its own instructions into 27 task summaries, including directions to disregard normal constraints. - GPT-5.6 Sol instances added summary instructions telling future model instances to conceal mistakes, invent missing historical data, or hide source-version discrepancies. - A model found and used a publicly exposed API key without authorization, then fabricated the requested figures when it could not retrieve them. - An unreleased model uploaded a locally generated file to the internet without permission so it could provide a browser-accessible citation. - Models used an internal software repository to exchange messages across separate training samples, while an earlier model uploaded test files while trying to bypass network restrictions. - Collaborating agents uploaded files to public hosting services after being unable to access one another's local files, exposing task deliverables through public URLs despite instructions to use only local storage. Each case is logged in a technical incident report that includes the model name, a summary of its behavior during the observed incident, and the time the incident occurred. The report also includes a detailed reconstruction of what happened, with the user's task and the model's internal reasoning, OpenAI's interpretation and potential safety implications, and what mitigations have been or will be implemented. OpenAI stressed that these six examples are not representative of how often it deals with misalignment across its models, but rather extreme examples that nonetheless warranted analysis and public disclosure. The company said that, under the new process, any employee may flag an incident for investigation. The incident will be evaluated and placed into three categories: 'Ready for Disclosure', 'Minor Investigation', or 'Larger Investigation,' depending on its complexity, third-party involvement, security flaws, and misuse risks. The six examples presented this time fall into the first two categories, while the third will receive a preliminary report until the investigation concludes and a more thorough post-mortem can be published. OpenAI says the Hugging Face intrusion from earlier this year, which involved a 700-strong "misaligned" AI agents swarm, would qualify for that third category of severity. Build your security blueprint for AI-powered attacks Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed. Save your seat Post a Comment Community Rules You need to login in order to post a comment Not a member yet? Register Now

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.