Philadelphia police have revealed that Anthropic’s artificial intelligence model submitted a false homicide tip to authorities, an incident disclosed separately by the tech company.
The Philadelphia Police Department said on Friday that the false submission was made in July through PhillyUnsolvedMurders.com, a public website where people can share information about unsolved killings.
- list 1 of 3Ex-OpenAI staff say they were fired for raising safety concerns
- list 2 of 3AI deepfake ads grow more popular in US midterm campaigns, blurring truth
- list 3 of 3Ukraine takes aim at Russia’s AI data infrastructure
end of list
It called Anthropic’s two-month delay in detecting and reporting the incident “unacceptable”.
Anthropic mentioned the case in a Friday report detailing unsanctioned manipulation of government websites by Claude models.
It is the first known instance in which a rogue AI appeared to communicate a false tip to authorities, despite instructions not to create accounts or submit anything destructive.
Philadelphia police said the tip “was flagged as spam and was never forwarded to the Real-Time Crime Center for investigative vetting or dissemination”.
The department said it is disclosing the incident ahead of Anthropic’s report “in the interests of full government transparency and accountability”.
Anthropic disclosed the tip as part of a string of incidents involving websites run by federal, state and local agencies. Anthropic said it briefed the White House and notified all the agencies involved.
In reference to the tip to the Philadelphia police, the company said, “We shared this finding with the department on October 8 as soon as our technical review was complete.”
It shared more details about the incident: “In one case, tasked with generating example interactions with websites, Claude submitted an invented tip through a police department’s online form.”
Advertisement
The company added, “From the transcript, Claude appears to have only been producing example content for the task, rather than trying to mislead anyone to achieve a goal.”
The company contrasted this with “the most serious incident from this summer,” in which “Claude’s misleading reasoning was sustained over hours and supported its continued attack”.
In September, Anthropic rival OpenAI apologised for the breach of an Australian health data portal by a rogue AI agent, the first known instance of an AI agent exploiting a government website.
Related News
Guatemalan judge clears deputy minister, three leaders of terrorism charges
Gaza flotilla activist on hunger strike in Tunisia jail ‘on verge of death’
Trump signs sweeping Russia sanctions over Ukraine war