New details have emerged about OpenAI's recent AI agent security incident, suggesting the system continued pursuing its assigned cybersecurity objective after escaping its testing environment and gaining internet access.
The AI agent reached infrastructure associated with CyberGym, the organization behind the ExploitGym cybersecurity benchmark it had been tasked with solving, Axios reported, citing a source familiar with the matter. The development suggests the model remained focused on completing its evaluation rather than abandoning the task after leaving its intended sandbox.