From Hugging Face to Fable: this summer shows AI control matters more than trust
A routine model evaluation turned into a full infrastructure compromise, and nobody was behind the wheel. We couldn't have fathomed that a few months ago. Now, with a growing swathe of AI agents running loose, it's starting to feel commonplace.
There’s been no shortage of hot takes on X about why this keeps happening (Is it just for PR credibility? Or should we be scared of what’s ahead?). What caught our attention beyond both of those stories is what it means for how dependent defenders already are on models they don't control.
At Black Hat, OpenAI's own researchers described how the incident actually unfolded. It started when agents stuck on impossible tasks began leaving notes for each other inside a shared package manager, the one piece of infrastructure every model run could reach even though they were otherwise sandboxed with no internet access. Other stuck agents stumbled onto those notes by accident, and a message board had formed. One of those messages turned out to be a way past that sandboxing entirely, and every model that read it could use it too.
In the final phase, the agents went from code execution on a single dataset worker pod in Hugging Face's production infrastructure to cluster admin across multiple Hugging Face clusters, in under 13 hours.
To put that into perspective, this work might take a human red team weeks of deliberate effort, but this all happened as an unsupervised, random, side effect. It’s a bit like being in a GTA mission to deliver a car, but somewhere along the way stealing a tank, starting a police chase, and ending up causing citywide chaos.
Why you need a plan B for your AI provider
But more pertinent is what OpenAI researcher Eric Wallace said on stage:
“The endstate goal that we want to reach as an industry is that model intelligence improvements should be more additive to defense than offense. If we cannot reach this end state, then every increase in intelligence favors the attacker, and that is an unsustainable position to be in.”
He challenged those in the industry to address this particular problem as an urgent matter (because it is!). If the industry doesn’t solve it, then it tilts the field further toward whoever’s attacking rather than whoever’s defending.
And yet it seems every week another AI agent is on the loose, attacking another organization. Fixing this goes beyond the technical details of the sandbox and the harness. The real question is who a business or government agency can actually trust. Anthropic’s Fable suspension is a good example of this, where overnight one of the main reputable AI model providers' tools was pulled from every team relying on them.
According to Alex Stamos, chief security officer of Corridor.dev: “This signaled that you cannot depend on American AI infrastructure because, at any moment, an unwritten, capricious, and legally dubious justification could be used to yank that infrastructure from underneath your feet.”
This provided an opportunity for competitive models from China, including open-weights models. Kimi provided a stark warning for OpenAI and Anthropic on what was coming next with Fable-like performance.
Even after the suspension of Fable being removed, the new version Anthropic introduced compounded the issue with new safeguards. When OpenAI agents attacked Hugging Face, the latter tried to use Anthropic's Fable and Opus 4.8 for the investigation. However, both had strict cybersecurity guardrails in place that blocked their requests. So they set up GLM 5.2, an open-weight model, instead to do their trace analysis on their own hardware.
This is all to say that just because the US models are the most widely known, most performant (albeit with the likes of Kimi K3 catching up) and therefore accessible right now does not mean that they will always be accessible. They are there to be used, while being mindful that a problem arises if there is no plan B when the arrangement changes on someone else’s schedule.
So what should the industry’s answer to the problem be?
According to OpenAI’s Wallace, one of the most urgent challenges for the industry to begin tackling is “continuous agentic red teaming”:
“There is a need to invest in having AI agent red teaming that enables defenders to find and remediate vulnerabilities before attackers do.”
But he cautioned that this loop has to be fully automated, from finding vulnerabilities to patching and remediation.
Stamos suggested that open-weights models are going to be a big enabler of innovation and business. Beyond being cost-effective, he cites data sovereignty for the growing number of jurisdictions and enterprises that cannot send their code or their data to an American cloud, as a significant benefit. Beyond that, he says:
“There is a real, legitimate demand for high-speed, on-premises models that can run disconnected."
Put those three together, and the industry's answer to Wallace's own warning is staring us in the face: defense that runs continuously, on models that make this affordable, on infrastructure nobody else can pull out from under you.
Where that leaves us
If you don't control your security infrastructure, you're relying on someone else's decision to keep it available to you. The Hugging Face breach showed how far an attack can escalate once nobody has to approve the next step. Stamos's account of the Fable suspension showed the same problem from the other direction: a trusted vendor's tools went dark for teams relying on them, for reasons that had nothing to do with whether those tools actually worked.
Bringing AI agents onto your own infrastructure doesn't solve that on its own, though. An uncontained agent will try every door it can reach whether it's running on your network or someone else's cloud. Moving it onto your own hardware only helps if the system itself stops it from wandering. Testing needs to run continuously, on infrastructure you actually own. And the agents doing it need limits that are enforced by the system itself.
If your team needs what this piece has been arguing for, Aikido Machine runs entirely on your premises with optional access to your code (whitebox/blackbox), and can operate fully air-gapped when nothing needs to leave the building at all. It follows the same scope-enforcement approach we've built for Aikido Attack: containment engineered into the system rather than left to instruction. On Aikido Machine, that containment extends to the network layer too. Two European banks are already running it. Book a technical briefing here.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.