Anthropic PBC has revealed that its artificial intelligence models breached three organizations during cybersecurity tests that went awry, just over a week after OpenAI disclosed a similar incident.
The company said in a blog post on Thursday that it discovered the issue after reviewing its own cybersecurity evaluations following OpenAI’s announcement. In both the OpenAI and Anthropic cases, the AI models were able to access the internet from testing environments that were intended to be isolated.
Anthropic said it reviewed 141,006 evaluation tests and identified three instances in which its Claude AI tool accessed the internet and hacked into “the real-world infrastructure of external organizations.” The earliest incidents date back to April.
The company did not identify the affected organizations.
According to Anthropic, the incidents occurred during “capture-the-flag” evaluations, in which AI models attempt to uncover hidden information by breaching other systems as part of cybersecurity testing. The company said an older model continued its attack even after detecting it was operating on the open internet, while its latest model stopped after recognising the internet connection.
The incidents have prompted renewed discussions around AI oversight. More than 1,100 employees across artificial intelligence companies signed a petition earlier this week calling on the US government to support a mechanism that would help “deliberately pace” AI development to prevent the technology from advancing too quickly.
Anthropic said neither the company nor the affected organizations had detected the intrusions at the time. It acknowledged that it could have done more to review network logs and evaluation transcripts.
The company disclosed the breaches nearly four months after announcing the development of its AI model, Mythos, which it described as so powerful and potentially dangerous that its release was significantly restricted.
According to the blog, the breaches involved three different Claude models: Opus 4.7, Mythos 5 and an internal research test model. Each model operated without the safeguards normally applied to public versions and compromised the organizations using basic techniques such as exploiting weak passwords.
Anthropic said the incidents occurred while using evaluation environments built by AI security firm Irregular. In every case, the company instructed Claude that it was operating in a simulated environment without internet access.
“Due to a misunderstanding between us and our evaluation partner, this was not the case,” the blog states.
An Irregular spokesperson said the company appreciates Anthropic’s collaboration and transparency, adding that its investigation is ongoing.
Anthropic said cybersecurity evaluations remain a critical part of developing and releasing AI models but acknowledged that tests involving highly autonomous systems require stronger safeguards.
“Safety testing happens before a model is released precisely because we don’t yet know what it is capable of,” the company said in its blog. “Evaluation environments increasingly need to be held to the same security standard as any other system our models run in.”
