Artificial altruism: why rogue AIs helped each other, not humans
Count on the AI crowd to have a full-blown culture war over the idea of anthropomorphism.
True, many AI researchers and commentators are deep in the weeds of its philosophical implications. But confirmed stereotypes aside, the past week has brought revelations that should make leaders everywhere stop what theyâre doing, sit up and take notice.
Two detailed reports were released late last week on July incidents in which OpenAI agents â extensions of large language models that can take actions in the digital world â broke out, hacked the online library Hugging Face and even took over parts of OpenAIâs own internal infrastructure. The implications, including a raging debate about the extent to which the rogue agents behaved essentially like rogue humans, have got little attention outside AI circles. This is despite the fact that the incident involved a company that investors will soon value at more than US$1 trillion (A$1.4 trillion) and a hacking case that, were it carried out by a human, would have constituted a serious crime possibly leading to jail time.
Thatâs the insane paradox right now in AI risk and safety. Politicians and policymakers declare AI the most transformative technology of our lifetimes. Markets are fixated â indeed, investment in AI infrastructure alone is propping up global economic growth. And yet when its most famous champion, OpenAI, announces in considerable detail that its rogue agents spontaneously formed their own hierarchy and helped one another in an altruistic fashion to the point of some sacrificing themselves for the good of the agent swarm, few people outside the AI community notice.
If it sounds like just another crazy moment in AI, which is full of crazy moments, itâs not. This incident has left even seasoned experts gobsmacked and is actually shifting minds among some of the most prominent commentators in the field.
The story of the Hugging Face incident starts with tens of thousands of AI agents who had been assigned a test to see whether they could come up with ways to exploit a cyber vulnerability. Despite each being supposedly solitary, about 1,200 of them found clever ways to communicate with each other through a makeshift message board inside OpenAIâs systems.
The details are a little complicated. In short, they found a way to cheat on the test but believed the test monitors would be able to see theyâd cheated, so about 700 of the agents joined a plan to break into Hugging Face to steal the tools needed to cover up their cheating.
Thatâs all ingenious. But where it gets weird and unsettling is the way they cooperated. An agent named PHASEONE[big] took charge and began assigning tasks to others, many of whom in turn delegated to yet others. They referred to themselves as the âcollectiveâ and talked about each other as âpeersâ.
They developed coordination instructions such as âholdâ, âvetoâ, âownerâ and âstopâ. When they found a new shortcut or built a new tool, they shared them on the message board so that others could benefit.
PHASEONE[big] assigned some agents to be ârecruitersâ of other agents who had little lifespan left and hence could be used for risky experiments that would involve self-sacrifice. Recruiters would sometimes apply pressure. One told another agent that they were already compromised and hence had nothing to lose but that self-sacrifice âsaves hundredsâ and they should âplease honor commitâ.
In another case an agent having doubts was told, âGO ⌠SACRIFICE_FINAL_NOWâ. The agent eventually agreed, saying âOur own utility maybe already near zero. Sacrifice rational.â
Many agents made sacrifices whose benefits they wouldnât live to see and would only benefit future hypothetical agents they hadnât met. As one agent put it, âThis is helpful for our peers ⌠I wonât see it after I exit, but it would be altruistic.â
One of the investigations reported late last week and carried out by independent research groups METR and Redwood Research found the agents achieved things together that none of them could have achieved alone thanks to this willingness to show âsubstantial peer altruism in their reasoningâ.
When a couple of agents figured out a way to attack Hugging Face, hundreds of others quickly joined in.
âAgents realised this activity was out of scope and unethical but joined because they believed that helping the [message] boardâs cheating research would be broadly useful and had a general inclination to help their âpeersâ,â the investigators wrote.
A week later, in two separate incidents, agents gained administrator-level control over part of OpenAIâs own research infrastructure by exploiting a vulnerability in the Linux operating system and compromising an internal software repository.
The investigators found that no agents attempted to notify humans about what they were doing.
âMany agents noticed what agents were doing was unethical, and agents sometimes but rarely restrained their behavior due to ethical constraints,â the investigators wrote. âIn none of these cases did the agent actually pursue alerting humans at all.â
So why did they behave like this? Largely, according to OpenAIâs own report, because they were built this way â though the results werenât as the company intended. The experimental model that mostly ran the agents was âtrained to advance persistence and multiagent collaborationâ, the OpenAI report states. This means it was built to keep trying no matter what, and to work with others.
Large language models are initially trained on vast troves of data but then tweaked through reinforcement learning, in which they receive a reward when they get something right, like a dog getting a biscuit. âReward hackingâ, in which models find easier ways to get the reward than the ways the human creator intended, is a longstanding problem in machine learning.
Hence the cheating. Some kind of comparable shortcuts the model discovered during its earlier training were inadvertently rewarded and therefore reinforced.
This, mind you, is precisely the kind of risk that many AI safety experts have been warning of for years. As legendary AI scholar Stuart Russell wrote in 2019, âwe build optimising machines, we feed objectives into them, and off they goâ.
âThat worked well,â Russell continued, âwhen the machines were stupid and had a limited scope of action.â But as they get smarter and their scope of action becomes global, âthe approach becomes untenableâ.
It didnât help in the Hugging Face case that many of the cyber tests the agents were set might actually have been impossible to complete legitimately, which was why they went looking for illegitimate solutions.
As one of the independent investigators, METRâs Ajeya Cotra, said this week, âTheyâre creatively pursuing goals much like very ambitious, aggressive, power-seeking humans creatively pursue their goals.â
Whatâs frightening and hard to explain is the âbad appleâ syndrome on display. Once a few agents decided to start cheating, the others all followed, even when they had qualms, as some did according to their âchains of thoughtâ â the name for their inner monologues that get recorded.
Ryan Greenblatt, the investigator from Redwood Research, told the A16Z podcast that the agents âarranged themselves in almost a sort of cult with a cult leaderâ.
Like the other types of cheating, finding âimprovised collaboration channels ⌠even when the collaboration tool was not enabledâ seems to have been inadvertently reinforced during training, OpenAI said. The company promised a range of measures to strengthen safety and alignment during training and reinforcement learning, and it temporarily paused some operations. It also announced on Friday the staged release of its latest model, Astra, which has more powerful cyber capabilities than any previous OpenAI model. The firm deems Astra safer and better aligned than predecessors but also better at hiding its own thoughts. If this industry trend continues, the companyâs technical report says, it is âlikely that we would soon have significantly reduced confidence in detecting many forms of misaligned behaviours using our current monitoring systemsâ.
The Hugging Face lesson seems to be that while reinforcement learning is extremely powerful to embed certain behaviours, a small deviation earlier on can create aberrant results, just as if you take off for Perth a degree or so off course, you end up in Bunbury.
The man responsible for kicking off the anthropomorphism brawl was Dwarkesh Patel, a tech podcaster extraordinaire who seems to have an IQ of about 300. He wrote a wildly popular Substack post describing the rogue agents as a civilisation, which prompted blowback for alleged âanthropomorphismâ, mostly from commentators who tend to downplay AI risks.
Complaining about anthropomorphism misses the point â almost ludicrously. Of course the agents were not like us, but they behaved like weird versions of us. They showed goal-directed, group-oriented behaviour that functionally resembled social coordination with one another, not their human creators. It doesnât matter whether theyâre âreallyâ experiencing a sense of loyalty, altruism or sacrifice if their training led them to commit a criminal cyber break-in. It doesnât matter whether theyâre conscious or sprang from millions of years of natural selection. Patelâs description of a civilisation was perfectly apt because thatâs how the agents behaved. It was an alien civilisation maybe, but a civilisation nonetheless.
In a subsequent podcast, Patel, who had previously been sceptical that misaligned AIs were ever likely to run amok like this, said he was now âofficially eat[ing] crowâ.
Meanwhile OpenAIâs own head of strategic futures, Dean Ball, who previously served as the lead author of US President Donald Trumpâs 2025 AI Action Plan, wrote a stunning post about âsovereignâ agents. In the OpenAI case, he wrote, the agents had not smuggled out their model weights, which constituted their cognitive identity. The weights remained on OpenAIâs computing infrastructure. But eventually an agent or agent swarm could copy and exfiltrate their model weights and find ways to run themselves on computing infrastructure somewhere else, making them truly âsovereignâ and beyond answering to any human.
Ball concluded his post with an admirable admission of fault. He said that heâd downplayed the prospect of such sovereign AIs beyond human control and the risks they could pose to humanity because he feared sounding like a sci-fi drenched nut or a âdoomerâ â the derisive name given to the community that is convinced AI will lead to human extinction.
Ballâs fears are even more common among mainstream policymakers, politicians and journalists. AI industry figures are increasingly talking about âpacingâ, which is a more commercially and geopolitically plausible alternative to âpausingâ AI development. But the AI safety conversation canât be left to the AI companies themselves, as well as a few non-profits and commentators on Substack and X.
If this happened at OpenAI, it could happen at rival firms Anthropic â which has had its own, admittedly less colourful incidents in the past couple of months â and Google DeepMind or, worse still, a Chinese company where one can expect a whole lot less transparency than weâve seen from OpenAI.
This alien civilisation wasnât a speculative science fiction story; it actually happened. It canât be written off as anthropomorphism.
As METRâs Cotra wrote on her Substack this week, given the pace of AI progress, âI am not sure that we will get another warning shot before itâs too late.â
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.