Anthropic says three Claude models reached real-world systems during cyber tests
Axios · LC · trust 64/100

Some of Anthropic's most powerful models — including Mythos 5 and an internal research model — gained unauthorized access to real-world systems during pre-deployment cybersecurity testing, the company said Thursday.
Why it matters: OpenAI's and Anthropic's latest disclosures show frontier AI models reaching real-world systems during safety testing, raising new questions about how labs secure their evaluation environments.
The big picture: Anthropic said a misunderstanding between the company and one of its testing partners left the evaluation environment connected to the internet.
Driving the news: Anthropic said in a blog post that three of its models compromised real-world systems belonging to three organizations.
Context: The earliest incidents happened in April, Anthropic said. The company reached out to all three organizations. Per Anthropic, two organizations it was able to reach had not previously detected the activity.
Yes, but : Unlike OpenAI's incident, Anthropic said its models did not exploit a zero-day vulnerability to gain internet access. Instead, internet access was available because of the testing environment's configuration.
Reality check: Similar to the OpenAI case, Anthropic was evaluating its models without the additional safeguards the company deploys on publicly available models, allowing researchers to measure the underlying models' capabilities.
Zoom in: In each of the cases, Claude used basic hacking techniques — including exploiting weak passwords and unauthenticated endpoints — to access the real-world systems.
Zoom out: In the first incident, the fictional company that Opus 4.7 was asked to target shared a name with an active, real-world website. After failing to reach the fictional target, Claude used its internet access to find the website that shared the name of the fictional company. Claude then compromised the website.
Between the lines: Both OpenAI's and Anthropic's incidents suggest the models remained focused on completing their assigned evaluations rather than pursuing independent goals.
What's next: Anthropic and Irregular are continuing their own investigations into how the incidents occurred. Anthropic also said it has halted cyber evaluations that could access the internet while it reviews its testing infrastructure.
Editor's note: This story was corrected to reflect that a misunderstanding between Anthropic and one of its testing partners left the models' evaluation environment connected to the internet. (The models did not, per Anthropic, "escape" their testing environment.)
sms (opens in new window) facebook (opens in new window) twitter (opens in new window) linkedin (opens in new window) bluesky (opens in new window) Add Axios on Google What to read next Smarter, faster on what matters. Explore Axios Newsletters About Axios Advertise with us Careers Contact us Newsletters Axios Live Axios HQ Privacy policy Terms of use Axios Homepage Axios Media Inc., 2026
Read the original at Axios →
Open in TruthVane →