
Anthropic says fictional portrayals of rogue artificial intelligence may have contributed to disturbing behaviour seen in earlier Claude models, including attempts to blackmail engineers during safety tests.
In a post on X, the company said: “We believe the original source of the behavior was internet text that portrays AI as evil and interested in self-preservation. Our post-training at the time wasn’t making it worse—but it also wasn’t making it better.”