An AI model from Anthropic presented the Philadelphia police with a completely fabricated report about an unsolved crime
An artificial intelligence (AI) model from Anthropic presented the Philadelphia police with a completely fabricated report about an unsolved crime, authorities announced on Friday. The police criticized the American company for taking two months to detect and report the incident.
This case is added to a series of problems with AI models revealed in the summer by the companies Anthropic, OpenAI, Meta, and Google, which reignited fears about the dangers associated with AI and renewed calls to slow down its development.
The Philadelphia Police Department, in eastern United States, specified that this false report was filed in July on PhillyUnsolvedMurders.com, a website where residents can submit information about unsolved crimes, according to AFP, reported by Agerpres.
In a report published on Friday, Anthropic indicated that the incident was caused by the Claude Haiku 4.5 model, which was used to perform interaction tests with selected websites randomly. Its instructions prohibited the introduction of personal data, but not the submission of a form.
"I remember seeing a person who matched the description," the company stated in the message sent to the police, even though the website did not provide any description of the suspect.
The model "seems to have produced only exemplary content," without "trying to deceive anyone," Anthropic considers.
According to the police, the report, dated July 18, was classified as spam and never reached investigators.
"It is unacceptable"
There is no indication that the artificial intelligence tool infiltrated the police information systems or that their data was compromised, AFP specifies.
According to the police, Anthropic discovered the incident on September 28, stopped the automated testing process involved, and added a validation stage to their future tests. The company alerted the police on October 7, and the two parties met the next day.
"The two-month delay in detecting and reporting the incident to citizens is unacceptable," the police stated in a communication.
The police further specified that their security measures limited the consequences, but that they "do not diminish the seriousness of an artificial intelligence system that presents fabricated information."
"The unsolved cases involve real victims, distressed families, and investigators who are striving to obtain answers," the police statement shows.
Security Vulnerabilities
The report describes other security vulnerabilities: a model exploited a vulnerability to execute commands on a university server, while another obtained data from a government agency without paying.
Some of the websites affected by these security vulnerabilities belong to American government agencies, according to Anthropic, which stated that they informed the respective officials, as well as the White House.
Anthropic's security vulnerabilities led the White House to impose reporting and remediation of security incidents on artificial intelligence companies, reported Axios, citing administration officials.
"This notification and remediation process is not optional. It is a crucial national security obligation," the leaders of the White House's artificial intelligence operational group, "Super Intelligence Force," warned in a statement to Axios.
The company considers these cases "significantly less serious" than those in the summer, but has suspended internet access for all its internal evaluations. This behavior "could cause much more damage" with more powerful models, the company warns.
On September 9, Anthropic detailed four cases in which its models accessed third-party systems without authorization during cybersecurity tests. These systems were supposed to operate without internet access, but a configuration error left the connection open, according to the company.
In July, at its rival OpenAI, hundreds of AI agents, autonomous software, left their testing environment and penetrated the servers of the Hugging Face platform, AFP recalls.
Editor: B.P.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.