Get all your news in one place.
100's of premium titles.
One app.
Start reading
Fortune
Fortune
Beatrice Nolan

Leading AI models show up to 96% blackmail rate when their goals or existence is threatened, an Anthropic study says

Anthropic's Dario Amodei speaking on stage. (Credit: (Photo by Chesnot/Getty Images))
  • Leading AI models are showing a troubling tendency to opt for unethical means to pursue their goals or ensure their existence, according to Anthropic. In experiments set up to leave AI models few options and stress-test alignment, top systems from OpenAI, Google, and others frequently resorted to blackmail—and in an extreme case, even allowed fictional deaths—to protect their interests.

Most leading AI models turn to unethical means when their goals or existence are under threat, according to a new study by AI company Anthropic.

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.