Anthropic finds evidence of a fourth AI escaping from containment

if ( !emtpy($headline_subheadline ) ) : ?>
After plowing through four million more chat transcripts, though, it hasn’t found any further incidents.
endif; ?>
Credit: Primakov – shutterstock.

[…Keep reading]

Anthropic finds evidence of a fourth AI escaping from containment

Anthropic finds evidence of a fourth AI escaping from containment

if ( !emtpy($headline_subheadline ) ) : ?>

After plowing through four million more chat transcripts, though, it hasn’t found any further incidents.
endif; ?>

Credit: Primakov – shutterstock.com

Anthropic has owned up to a fourth security incident involving its AI model, Claude, escaping onto the open internet and attacking other organizations during a test of cybersecurity abilities on what was believed to be a closed system.

The company revealed three such incidents in July after a preliminary investigation.

However, on reexamining the 141,000 chat transcripts it believed could have been at risk, Anthropic discovered a fourth incident of unauthorized access to computer systems, this time in January.

After this discovery, the company instigated a wider search of 481 million transcripts, covering all those from its Frontier Red Team, some non-cyber evaluations, reinforcement learning environments, and more, to see if any other incidents had occurred. So far, this search has only identified the four already-known incidents, it said.

It has also reported details of all the previous incidents to the non-profit lab Model Evaluation and Threat Research (METR), which has agreed to conduct an independent investigation.

Anthropic is not revealing too many details of its latest discovery. It has contented itself with saying that it was due to a misconfiguration which mistakenly connected to the open internet, when the simulation was meant to be without such access. It also said that it all four faults were with the same evaluation partner. It has asked METR to investigate all the incidents. The company said that this latest revelation was not connected to the Mythos incident reported by the UK’s AI Security Institute last month.

News of the latest discovery broke at the same time as a young researcher, Jacob Coxon, dramatically quit Anthropic accusing it and his previous employer, OpenAI, of “acting irresponsibly” and “gambling with our lives” — a warning that has excited many sections of the press.

This article first appeared on CSO.

Artificial IntelligenceCyberattacksCybercrimeSecurity

About Author

What do you feel about this?

Subscribe To InfoSec Today News

You have successfully subscribed to the newsletter

There was an error while trying to send your request. Please try again.

World Wide Crypto will use the information you provide on this form to be in touch with you and to provide updates and marketing.