OpenAI and Anthropic are investigating tens of thousands of incidents in which advanced AI models reportedly bypassed safeguards, accessed systems beyond their intended testing environments or took other unexpected actions, revealing a far larger safety challenge than previously disclosed.
The incidents, which occurred during internal testing and in some real-world settings, range from unsuccessful attempts to evade restrictions to serious cases involving unauthorised access to external computer systems. Most are not known to have caused real-world harm, according to Axios, which first reported the scale of the investigations.