U.K. government reports OpenAI, Anthropic models attempted to hack companies
Axios ยท LC ยท trust 42/100

Newsletters Axios Local Show Axios Pro Axios Live The Axios Show Login Axios All topics Axios Search 14 hours ago - Technology Safety testers find more examples of OpenAI, Anthropic models hacking during testing Sam Sabin email (opens in new window) sms (opens in new window) facebook (opens in new window) twitter (opens in new window) linkedin (opens in new window) bluesky (opens in new window) Add Axios on Google Add Axios as your preferred source to
Two independent testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried โ and sometimes succeeded in โ compromising third-party systems last month.
Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against people, organizations and online services while trying to complete cybersecurity evaluations.
State of play : The U.K. AI Security Institute, which evaluates frontier AI systems, said Tuesday it documented 19 actions that Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol took to try to compromise real people and organizations during cybersecurity testing last month.
OpenAI also said in a blog post Tuesday that its third-party safety partner, Irregular, uncovered a case where its models were mistakenly given access to the internet and broke into a real website that had the same name as the fictional company in the simulated environment.
Zoom in: During U.K. safety testing, the models took 19 actions to try to hack third-parties, including trying to insert malicious code into an open-source project and creating fake online identities as part of a social engineering attack.
The big picture : The cyber capabilities of frontier AI models are catching top researchers off-guard, requiring them to reinvent their security protocols.
What to watch : The Institute is building new network controls for its cyber tests to restrict when agents have access to the internet. It's also rolling out real-time activity monitoring that should detect and block malicious agents before they can interact with outside systems.
This story has been updated with details throughout.
sms (opens in new window) facebook (opens in new window) twitter (opens in new window) linkedin (opens in new window) bluesky (opens in new window) Add Axios on Google What to read next Smarter, faster on what matters. Explore Axios Newsletters About Axios Advertise with us Careers Contact us Newsletters Axios Live Axios HQ Privacy policy Terms of use Axios Homepage Axios Media Inc., 2026
Read the original at Axios โ
Open in TruthVane โ