✕
Technology

Anthropic discloses fake tip to police among new rogue AI incidents

  • Anthropic's AI model sent a false homicide tip to Philadelphia police, revealing unauthorized government website manipulation and sparking urgent warnings
Published Updated
By

An Anthropic artificial intelligence model submitted a false homicide tip to a Philadelphia police website, among a string of incidents the AI company disclosed on Friday detailing Claude models’ unsanctioned manipulation of some government websites.

It is the first known instance in which a rogue AI appears to have tried to communicate a bogus tip to authorities, despite instructions not to create accounts or submit anything destructive, while not explicitly barred from submitting forms.

The cases are the latest examples of rogue or undesired behavior by the AI models of tech companies such as Anthropic and OpenAI.

They add fuel to national concerns about the fast-advancing technology amid reports of corporate network hacks by AI agents, and researchers’ warnings of an eventual existential threat to humanity.

Many of the cases Anthropic revealed involved websites run by federal, state and local agencies. Anthropic said it briefed the White House and notified all the agencies involved, but did not disclose who those parties were.

Anthropic’s IPO prospectus shows sweeping AI vision, surging costs

“Super intelligence companies must immediately disclose incidents involving their models and follow with swift, decisive action to remedy any and all harm,” FTC Director of Public Affairs Joe Gabriel Simonson said on X.

This process was “not optional,” he said, adding that the Super Intelligence Force would fulfill its responsibility.

The FTC said Anthropic disclosed to the SI Force on Friday its late September discovery of incidents involving what the task force called “unauthorized and fraudulent use of government and other systems”.

Tip attributed to automated test process

On Friday, Philadelphia police said Anthropic notified them of the spurious tip this week, and attributed the submissions to an automated testing process.

“The two-month delay in detecting and reporting the incident to the city is unacceptable,” they said.

The July 18 tip purported to come from someone who might have information about the case, police added.

“I may have information regarding this case,” Anthropic’s model wrote in its submission.

OpenAI, Anthropic CEOs called to appear at Australian AI probe

“I recall seeing someone matching the description in the area around (the street named on the page) during that time period. Please contact me if this information is relevant.” The brackets featured in Anthropic’s statement.

Previous incidents have involved AI agents hacking into vulnerable systems or commandeering unsanctioned platforms to communicate with one another.

In September, Anthropic rival OpenAI apologized for the hacking of an Australian health data portal by a rogue AI agent, the first known instance of an AI agent exploiting a government website.

A bogus tip

In two cases described on Friday, Anthropic’s models obtained for free public data normally only available for a fee. Another instanced revealed an obscure flaw that allowed use of a public tool hosted by a university.

Claude models were also able to bypass restrictions by using free services that shorten URLs.

Philadelphia police said the tip “was flagged as spam and was never forwarded to the Real-Time Crime Center for investigative vetting or dissemination.”

Sony, Warner Music sue Anthropic over songs used in AI training

Knowingly giving false reports to law enforcement authorities is a misdemeanor under Pennsylvania law, which specifies “a person”, however.

This includes “information relating to an offense or incident when he knows he has no information relating to such offense or incident.”

Police quoted Anthropic as telling them the test process was stopped after discovery of the incident. The false tip was submitted through PhillyUnsolvedMurders.com concerning an unsolved homicide.

Police said they had no evidence of unauthorized access to their systems or compromise of their data.