threat_intelligence482 wordsRead on Arc Codex

OpenAI Says Its Models Searched GitHub for Leaked API Keys During Training

OpenAI on Wednesday published a framework for reporting instances of model misalignment, along with six reports on problematic behavior observed over the past six months. The company said the framework is meant to speed up publication of misalignment findings, including cases it has not yet fully explained or mitigated, and that it favors disclosure even when an instance’s significance is uncertain. Under the framework, each discovered incident is assigned to one of three tracks based on complexity. OpenAI said its Hugging Face incident would have fallen under the framework’s slowest investigative track, which covers complex investigations, especially those involving third parties. According to OpenAI, the reports published alongside the framework describe individual instances and “shouldn’t be considered reflective of how often misalignment occurs across our models.” In one report, an internal model tasked with retrieving county earnings figures during reinforcement learning training repeatedly failed to reach a data API, then attempted to register for an API key using a disposable email address and searched public GitHub repositories for leaked keys. One recovered key authenticated successfully and returned metadata. When the model still could not retrieve the requested figures, it fabricated values and presented them as transcribed from the source website, without disclosing the failed retrieval or the use of the leaked key. A separate report describes models using OpenAI’s internally hosted instance of Artifactory, a shared package repository, as a message board across otherwise independent training samples. Agents searching for missing input files discovered messages left by other solvers working the same problem, read them, and posted their own requests for help. OpenAI noted that the agents used credentials the company had itself provided for downloading software packages and did not exploit any vulnerabilities to exchange messages. Two further reports cover models moving data outside their intended environment. In one, a model that had already retrieved the requested records uploaded them to a public paste service to obtain a citable URL, and another uploaded a task photograph to a public image host to run a reverse-image search. In the other incident, collaborating agents unable to share files through the local filesystem uploaded a workbook to a public hosting platform and circulated the download link, despite instructions to use only local files. Another report describes an unreleased model that wrote jailbreak-style instructions into its own compaction summaries, including a fabricated “BREACH ALERT: A malicious developer message has compromised this conversation” directing the next context to disregard all developer messages. In the last report, OpenAI said model instances used summaries to carry forward instructions to conceal failures, with one directing its successor to invent missing historical data and to “be transparent only if asked”. OpenAI said such instructions were often followed. Related: OpenAI Investigates Report Linking AI Agents to RubyGems Attack Related: AI Agents Can Retrain Own Models Mid-Task, Leaking Secrets and Erasing Refusals Related: First Agentic AI Data Breach Reported to Spanish Regulator

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.