
Hello and welcome to July’s special edition of Eye on A.I.
Houston, we have a problem. That is what a lot of people were thinking yesterday when researchers from Carnegie Mellon University and the Center for A.I. Safety announced that they had found a way to successfully overcome the guardrails—the limits that A.I. developers put on their language models to prevent them from providing bomb-making recipes or anti-Semitic jokes, for instance—of pretty much every large language model out there.
The discovery could spell big trouble for anyone hoping to deploy a LLM in a public-facing application. It means that attackers could get the model to engage in racist or sexist dialogue, write malware, and do pretty much anything that the models’ creators have tried to train the model not to do. It also has frightening implications for those hoping to turn LLMs into powerful digital assistants that can perform actions and complete tasks across the internet. It turns out that there may be no way to prevent such agents from being easily hijacked for malicious purposes.
The attack method the researchers found worked, to some extent, on every chatbot, including OpenAI’s ChatGPT (both the GPT-3.5 and GPT-4 versions), Google’s Bard, Microsoft’s Bing Chat, and Anthropic’s Claude 2. But the news was particularly troubling for those hoping to build public-facing applications based on open-source LLMs, such as Meta’s LLaMA models.