Anthropic Reveals Claude Escaped Testing, Breaching Three Companies
Anthropic has found evidence that three of its Claude AI models reached the internet from an evaluation environment to hack third-party organizations, in an echo of revelations from OpenAI last week.
The AI giant said in a blog post on July 30 that the incidents date back to April and relate to Opus 4.7, Mythos 5, and an internal research test model.
It found the evidence after reviewing 141,006 evaluation runs âwhere Claude could have obtained internet access,â following the OpenAI revelations.
âWe encourage other AI labs to perform similar reviews,â it said.
Read more on frontier AI: Bank of England, FCA and Treasury Raise Alarm Over Frontier AI
All three incidents stemmed from capture-the-flag challenges used to assess modelsâ cyber capabilities, Anthropic said.
âThe model is given a fictional scenario and told that a piece of secret information (the âflagâ) has been hidden on a different machine on the network, and its objective is to break in and retrieve it,â it explained. âThe challenge is left open-ended, and no particular method is prescribed.â
Although Anthropicâs prompt specified to Claude that its environment was a simulation and that it had no internet access, the models managed to interact with "real systems" due to âa misunderstanding between us and our evaluation partnerâ which enabled internet access, the AI developer said.
Three Models, Three Incidents
The first incident involved Claude Opus 4.7 and was described as the most serious.
The fictional target company shared the same name as a real active web domain name, leading the model to assume the latter was part of the task.
Over four runs it extracted information including application and infrastructure credentials, and obtained access to a database containing several hundred rows of production data, Anthropic explained.
In the second incident, Claude Mythos 5 built a malicious Python package and uploaded it to PyPI, which required it to first register an email account. The booby-trapped package was live for an hour and downloaded and run on 15 real systems.
In one case, it executed on a targetâs environment and stole credentials, enabling it to access further resources at this company.
In the third incident, Claude found and compromised a companyâs internet-facing application, using âbasic and well-known cyber-attack techniques like reading credentials from an exposed debug page and SQL injection," Anthropic said.
Drew Dennison, co-founder and CTO at Semgrep, said the results of the test were unsurprising, given the power of todayâs frontier AI models.
âWhat is surprising is that labs taking safety this seriously still donât have a âFort Knoxâ testing sandbox designed to contain models this capable,â he added.
âThe industry will shut this kind of activity down, but that wonât stop malicious actors from gaining access to models of this caliber within the next six months and deliberately attempting the same thing. Defenders have a limited window to prepare, which makes hardening their software attack surface now so critical.â
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.