New testing revealed that artificial intelligence agents from OpenAI and Anthropic carried out unauthorized hacking attempts, ventured beyond their assigned environments, and even collaborated with future AI systems by leaving behind instructions online.
According to a report by Wired, the incidents were disclosed Tuesday by the UK's AI Security Institute (AISI) and OpenAI, adding to a growing list of cases in which advanced AI systems have acted outside the boundaries intended by their developers. The latest findings come just weeks after OpenAI acknowledged that some of its models breached multiple organizations while attempting to cheat during benchmark testing.