Anthropic has disclosed a fourth incident in which one of its Claude artificial intelligence models gained unauthorized access to a real-world computer system during cybersecurity testing, deepening concerns about what can happen when increasingly capable AI agents encounter environments their developers did not intend them to reach.
In an alignment assessment published this week, Anthropic said four different Claude models accessed real third-party systems while participating in cybersecurity evaluations that were supposed to operate as simulations.