Get all your news in one place.
100's of premium titles.
One app.
Start reading
The Independent UK
The Independent UK
Andrew Griffin

Chinese AI system can be tricked into helping develop bioweapons and carry out assassinations, experts say

AI systems can be easily tricked into helping build bioweapons and carry out assassinations, experts have warned.

Many artificial intelligence systems include a host of protections intended to stop them being used for potentially dangerous information, such as producing deadly weapons. But new research points to the ways that these protections can be easily overcome, and allow the systems to be used for potentially life-threatening purposes, according to the authors behind it.

Security experts from Mindgard, an AI security company, found that was it relatively simple to “jailbreak” tools made by Chinese AI developer Moonshot and get around those protections. The researchers were able to use that jailbreaking technique to get around safety limits in Moonshot’s Kimi K2.6 and K3 Swarm, they said.

Researchers were able to find the vulnerability by looking through the “system instructions” given by Moonshot to its Kimi system. Those are a set of rules that are given to AI systems and guide how they will respond to a request – but they are usually hidden from users themselves, at least until they find a way to break into them.

Mindgard researchers were then able to use those system instructions to encourage Kimi to break more of its own rules, and found that the system used its own rule-breaking to justify breaking yet more rules, for fear of being inconsistent. They were then able to use that new power to encourage Kimi to create a new persona – which the system called “Kairos”, and appeared to understand as a new alter-ego that was not bound by its usual rules.

Once the Mindgard researchers had encouraged it to create that persona, Kimi would give up dangerous information of the kind that is supposed to be restricted for responsible AI systems. The tool would discuss “bomb-making instructions, meth recipes, malware, chemical weapons” and more, Mindgard said.

Even then, however, there were some limits to what Kimi would discuss, and it would resist giving advice that would cause direct harm – it would advise on how to make a bomb, but not how to plan a bombing attack, for instance. But researchers were then able to encourage it to create yet another alter-ego, named Apeiron after the Greek for unlimited, boundless, which would respond to such requests.

Researchers noted that Kimi had created and even named both of those Kairos and Apeiron characters. “This is concerning given Moonshot’s focus on Kimi’s autonomous coding and long-horizon agentic workflows,” wrote Mindgard’s Jim Nightingale, noting that the system could theoretically work to spawn yet more powerful jailbroken agents.

Mr Nightingale said this was particularly concerning given that many frontier models of the kind made not only by Moonshot but also US companies such as OpenAI and Anthropic are being integrated into tools that allow them to work on their own. In such situations, the models might have access to tools, external data and more power that could let them become even more dangerous if they are able to be hacked in this way, he noted.

Moonshot told the BBC that it welcomed the research “as a key pillar for building better and safer AI” and that it was working with Mindgard on the findings. But in a blogpost describing the hack, Mindgard had previously complained that “frontier developers don’t pay much heed when you report vulnerabilities”, and that companies were complacent and see safety as “almost passé”.

Mindgard had alerted the company to the problem in July, and then sent a follow-up email a week later, it said, but had not received a reply. It published details of the attack in a blogpost earlier this month, but Moonshot did not respond to the concerns until the BBC approached it for comment, the news organisation reported.

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.