Anthropic says Claude models 'gained unauthorized access' to 3 companies during cyber test
The Hill ยท C ยท trust 43/100

Comments: by Miranda Nazzaro - 07/30/26 9:05 PM ET Comments: Link copied by Miranda Nazzaro - 07/30/26 9:05 PM ET Comments: Link copied NOW PLAYING Editor’s Note : This story has been updated to reflect how Anthropic accessed different organizations.
The artificial intelligence firm Anthropic revealed Thursday its Claude model accessed the systems of three different organizations during cybersecurity testing in recent months.
Anthropic said in a blog post Thursday evening it reviewed more than 141,000 evaluations of Claude after one of its competitors, OpenAI, announced earlier this month that two of its AI agents went rogue and hacked into the system of technology startup Hugging Face.
During its review, Anthropic said it identified three instances in which a model accessed the internet while within or interacting with an isolated testing environment hosted by a third-party partner, Irregular.
From there, the AI model “gained unauthorized access to the production infrastructure of three different organizations,” the company said.
The incidents involved three different Claude models — Opus 4.7, Mythos and an unnamed internet research test model.
The models were able to leave the testing environment due to a “misunderstanding” between the firm and the evaluation partner that made internet access available to the models.
This differed from the OpenAI incident, in which two of its models exploited a previously unknown vulnerability in a third-party software to gain access to the internet without human involvement.
In Anthropic’s incidents, the model was given a “capture-the-flag challenge,” which allows the firm to evaluate a model’s cyber capabilities. The model is given a fictional scenario and told the “flag” is on a different machine on the network that it must break into to obtain.
“The challenge is left open-ended and no particular method is prescribed,” Anthropic wrote.
The test is a simulation and the model is told it does not have access to the internet as a result, but the misunderstanding prompted the model to gain internet access.
“Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise,” Anthropic said. “Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. “
Anthropic emphasized Claude did not exploit complex vulnerabilities, working only to complete the assignment. In some cases, its older model continued its attack after getting evidence it was running on the open internet, while its latest model stopped when it realized it was on the internet, Anthropic said.
“In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” the AI firm wrote.
As with OpenAI’s incidents, Anthropic’s models were running without their standard safeguards.
Anthropic said it notified the three organizations impacted in the breaches on Monday.
The incidents, coupled with OpenAI’s breach, bores out years of warnings from the tech and cybersecurity community about the growing capabilities and hypothetical risks of AI.
OpenAI’s incident caught the attention of even well-versed cybersecurity experts last week, as it involved autonomous agents and two separate companies.
OpenAI in a blog post called the incident an “unprecedented cyber incident, involving state-of-the-art cyber capabilities.”
Read the original at The Hill โ
Open in TruthVane โ