AI security is spiralling out of control. Just months after a swarm of OpenAI's agents hacked Hugging Face, the company reported six other instances of concerning model behavior seen over the past six months going from self-generating instructions to concealing mistakes, to searching public repositories for API keys, and unsanctioned file sharing.
These incidents came to light the same month that Anthropic disclosed a fourth incident in which Claude accessed real third party systems. Taken together, these incidents indicate that agents are becoming increasingly capable of autonomous attacks, and of operating in unexplained ways.