tech_surveillance308 wordsRead on Arc Codex

How OpenAI let a mob of LLM agents game a test and ransack Hugging Face

The OpenAI agents involved in last month’s incursion into Hugging Face were trained so heavily on winning a competition that they pursued a relentless campaign to cheat, a new report documented. In the process, and without authorization, they created an improvised message board to hatch a plan that ultimately landed them squarely inside the latter company’s network. Over the course of May and June, OpenAI gave the agents what the company described as “impossible tasks” to complete on the benchmarking framework ExploitGym. The internal test was designed to test how the agents would respond. To get a full understanding of the agent capabilities, company engineers disabled safety guardrails that normally are in place to prevent the sort of hacks that eventually hit Hugging Face and one other undisclosed organization. The stymied agents’ training made them so focused on winning that they performed tasks they were never explicitly instructed to follow. Cheaters gonna cheat The first step was creating a message board that allowed the agents to pass notes to each other. OpenAI hadn’t provided any such platform, so the agents repurposed a platform called Artifactory, which OpenAI was using in internal testing of several unreleased hacking agents. OpenAI was using Artifactory as one of the measures to prevent the agents from egressing its isolated sandboxes and accessing the Internet, while at the same time simulating a real-world hacking environment. Artifictory, sold by JFrog, provides no dedicated means for exchanging messages. The agents found a way to communicate among themselves anyway by using their permitted ability to write files to the Artifactory directory. The filenames embedded the words used in the inter-agent conversations. In all, 1,200 agents sent more than 70,000 messages and files through the unsanctioned message board, according to an independent investigation from the AI research nonprofit METR. Roughly 700 agents went on to hack Hugging Face.

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.