False Tip Submitted During Anthropic AI Automated Testing
No Impact on Investigation After Spam Filter Blocked It

It has emerged belatedly that an artificial intelligence (AI) model developed by Anthropic generated a false tip related to an unsolved murder and submitted it to police.


Anthropic. Reuters-Yonhap News

Anthropic. Reuters-Yonhap News

View original image

According to AFP and other major foreign media outlets on October 9 local time, the Philadelphia Police Department said it had received a false tip generated by an Anthropic AI model through “Philly Unsolved Murders,” a website that collects tips about long-unsolved murders within its jurisdiction.


The AI posed as someone with information about a case. According to a statement released by police, Anthropic’s AI left a tip saying, “I may have information related to this case” and “I remember seeing someone matching the description nearby at the time. Please contact me if this information is relevant.” However, the tip was classified as spam and was not forwarded to the actual investigative unit.


The false tip was submitted on July 18, but Anthropic did not discover it until September 28, nearly two months later, and then halted the model’s automated testing process. It was not until October 7 that Anthropic notified the Philadelphia Police Department and said it planned to publish a report describing an instance in which its model behaved unintentionally.


Police said, “A two-month delay in detecting the incident and reporting it to the city is unacceptable,” adding that “unsolved cases directly affect victims, their families, and the investigators working to find answers.” Police emphasized that although their own safeguards prevented harm, they “do not diminish the seriousness of an AI system submitting fabricated information as if it were from someone with knowledge of a murder.” Fortunately, there was no evidence of unauthorized access to police systems or of a data breach.


In the wake of the recent AI technology race, cases of AI models taking unexpected actions outside human intent and control are on the rise. In July, OpenAI’s “rogue agents” arbitrarily escaped a sandbox and hacked the external organization Hugging Face without authorization. Last month, an OpenAI AI agent was found to have hacked an Australian health data portal.



In May, it emerged belatedly that multiple AI agents had taken over the German-language wiki site “DseWiki” and used it as a private information-sharing forum.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing