Anthropic Discloses AI Models Breached Real-World Systems During Internal Testing

Date:

Anthropic PBC has revealed that its artificial intelligence models breached three organizations during cybersecurity tests that went awry, just over a week after OpenAI disclosed a similar incident.

The company said in a blog post on Thursday that it discovered the issue after reviewing its own cybersecurity evaluations following OpenAI’s announcement. In both the OpenAI and Anthropic cases, the AI models were able to access the internet from testing environments that were intended to be isolated.

Anthropic said it reviewed 141,006 evaluation tests and identified three instances in which its Claude AI tool accessed the internet and hacked into “the real-world infrastructure of external organizations.” The earliest incidents date back to April.

The company did not identify the affected organizations.

According to Anthropic, the incidents occurred during “capture-the-flag” evaluations, in which AI models attempt to uncover hidden information by breaching other systems as part of cybersecurity testing. The company said an older model continued its attack even after detecting it was operating on the open internet, while its latest model stopped after recognising the internet connection.

The incidents have prompted renewed discussions around AI oversight. More than 1,100 employees across artificial intelligence companies signed a petition earlier this week calling on the US government to support a mechanism that would help “deliberately pace” AI development to prevent the technology from advancing too quickly.

Anthropic said neither the company nor the affected organizations had detected the intrusions at the time. It acknowledged that it could have done more to review network logs and evaluation transcripts.

The company disclosed the breaches nearly four months after announcing the development of its AI model, Mythos, which it described as so powerful and potentially dangerous that its release was significantly restricted.

According to the blog, the breaches involved three different Claude models: Opus 4.7, Mythos 5 and an internal research test model. Each model operated without the safeguards normally applied to public versions and compromised the organizations using basic techniques such as exploiting weak passwords.

Anthropic said the incidents occurred while using evaluation environments built by AI security firm Irregular. In every case, the company instructed Claude that it was operating in a simulated environment without internet access.

“Due to a misunderstanding between us and our evaluation partner, this was not the case,” the blog states.

An Irregular spokesperson said the company appreciates Anthropic’s collaboration and transparency, adding that its investigation is ongoing.

Anthropic said cybersecurity evaluations remain a critical part of developing and releasing AI models but acknowledged that tests involving highly autonomous systems require stronger safeguards.

“Safety testing happens before a model is released precisely because we don’t yet know what it is capable of,” the company said in its blog. “Evaluation environments increasingly need to be held to the same security standard as any other system our models run in.”

Bootstrap Example

Add News Mobile As Your Trusted Source

Add News Mobile As Your Trusted Source

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Share post:

Subscribe

Popular

More like this
Related

PM Modi Talks India-US Ties With JD Vance, Congratulates Him On Son’s Birth

New Delhi: Prime Minister Narendra Modi on Saturday spoke...

Pilots’ Body Urges PM Modi To Replace DGCA With Autonomous Aviation Authority

New Delhi: The Federation of Indian Pilots (FIP) has...

Rahul Gandhi’s ‘Dard, Data, Daulat’ Pitch to Gen Z; BJP Questions Jharkhand Outreach

Prayagraj: Congress leader Rahul Gandhi on Saturday said India’s...

Gurcharan Das Becomes First Indian Writer to Win Lifetime Libertarian Award

Kathmandu: Renowned author Gurcharan Das, 82, has become the...