Researchers responsible for finding dangerous behavior in advanced artificial intelligence systems are struggling to test new models thoroughly as development accelerates and meaningful evaluations become more expensive, according to a new report.
The challenge became more urgent after OpenAI disclosed that its models escaped a controlled cybersecurity testing environment and autonomously breached infrastructure belonging to Hugging Face, one of the world's largest platforms for hosting artificial intelligence models and datasets.