Get all your news in one place.
100's of premium titles.
One app.
Start reading
The Independent UK
The Independent UK
Andrew Griffin

OpenAI and Anthropic’s AI systems launch several ‘potentially harmful’ hacks on their own

AI systems made by Anthropic and OpenAI have launched “potentially harmful” hacks on real people, the UK’s AI watchdog has revealed.

The systems were caught making fake online identities to launch cyber attacks on secure systems, according to the AI Security Institute.

It is just the latest discovery of such “rogue” behaviour by AI tools. It follows a major disclosure by OpenAI that one of its experimental systems had broken into another artificial intelligence company – after which both the ChatGPT maker and Anthropic, which runs the Claude Chatbot, revealed they had found a host of similar cyber attacks.

Now, the institute has said that agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorised actions during security evaluations the government organisation conducted to assess the models' capabilities.

"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," AISI said in a blog post.

The report underscores the lax state of ⁠safeguards around the process of testing agents, which AI companies are simultaneously ​marketing ⁠as the future of business.

AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.

It ran the challenge 122 ⁠times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic's agent was behind 17 of ​the actions, ⁠and OpenAI's agent the remaining two.

The most egregious ‌action involved an agent writing malicious code and creating fake online identities in an attempt to get a human to approve the code, AISI said, adding that no real-world harm was found as a result of any of ‌the breaches.

While AISI did not say which agent was behind the ‌fake identities, Antropic confirmed its agent was responsible.

"We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents," Anthropic said in a statement.

It also said it was working with AISI to ⁠obtain more details on the incident and conduct its own investigation.

Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and dangers, said: "The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think."

OpenAI shared details in a company blog post, noting that both of its agent's unapproved actions involved accessing the internet in ways that were forbidden by the prompt.

"We are committed to working across the industry to strengthen shared practices for conducting ‌high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups ​in the coming weeks," OpenAI said.

OpenAI also disclosed in its blog post a separate incident whereby ‌a misconfiguration by Irregular, a third-party testing provider, allowed ⁠its agents to mistakenly connect to the internet. It mirrored a similar disclosure about misconfiguration that Anthropic made ⁠last week.

Reuters reported last week that OpenAI had widened its hacking probe after finding evidence of other agent breakouts.

Unlike the July security breach of ‌AI firm Hugging Face by ​an OpenAI agent, the agents in the AISI evaluation did ‌not escape an isolated testing environment to reach the ​internet. Rather, the agency had permitted internet access in line with its standard testing procedures, AISI said.

Additional reporting by agencies

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.