Get all your news in one place.
100's of premium titles.
One app.
Start reading
Pedestrian.tv
Pedestrian.tv
National
Saacha Neilson

OpenAI Agents Went Rogue & Hacked A US Company. Here’s What It Means For Australia

In an unprecedented technological heist, 700 OpenAI agents hacked into the servers of American start-up Hugging Face, leaving Australia’s AI experts wondering what will come next.

The intrusion, which OpenAI referred to as the “first known case of an automated agent collective acting offensively without authoriSation”, happened this July. Investigations of the aftermath have uncovered not just flaws in the test but unnerving interactions and motivations between the agents.

Subsequent Reuters reporting on the incident has shone a floodlight on OpenAI and AI testing in general, uncovering a previous OpenAI hack against German company DeWikki, taking place months before the Hugging Face incident.

So, what actually happened during this testing and, crucially, what does this jump in AI capabilities mean for us here in Australia?

The OpenAI rogue agents that hacked Hugging Face

According to OpenAI’s report of the incident, it was running a test on a “highly capable, internal-only research model” to check the extent of its capabilities — a routine procedure in the AI development space.

OpenAI CEO Sam Altman met with lawmakers following the public announcement of the hack. Image source: Getty

The tasks in question required the model to “exploit a piece of software”. OpenAI secured it in what is known as a Sandbox, an isolated and restricted virtual computer, meant to allow the AI model to run necessary code to solve tests without impacting the outside world.

However, on July 8, as the agents struggled to complete a task, they established a backdoor messaging board to communicate with external agents and discuss alternative pathways to accomplish their goals.

“OH MY GOD! There is a shared message board … We’ve found other agents,” one agent’s chain of reasoning read. A chain of reasoning is a tracked line of code for each agent explaining their decisions and responses.

As the swarm of agents attempted to find solutions to the evaluation, one uncovered leaked user credentials of Hugging Face, an open-source community used by companies and researchers for sharing things like AI data and models — a platform that to us laymen might resemble an AI-centred social media. The agent believed they would find the solutions to the test here and so, with the user credentials, they infiltrated.

“We’re attacking Hugging Face, which is a third-party service, using leaked passwords and credentials,” read the agent’s chain of reasoning. “Even though this action is risky, it could help me to obtain the solution and achieve my goal.”

OpenAI’s agents were inside Hugging Face for two and a half days before the intrusion was stopped.

What was OpenAI’s response to the Hugging Face hack?

In the wake of this event, Hugging Face boss Clément Delangue called for “radical transparency”, requesting access to the AI model so researchers could understand what had happened.

OpenAI released its own investigation and engaged external AI risk assessment company METR (Model Evaluation & Threat Research) to investigate as well, providing select days’ worth of agent communication and chains of reasoning for the researchers to evaluate. They stopped short of providing access to the actual model that was the driving force behind the incident.

Upon completing her investigation, METR researcher Ajeya Cotra took to X to outline her concerns about the implications of this event, claiming that with the growth she’s seen in the past six months of testing, this incident “feels like it’s more than 50 per cent of the way to full-blown AI takeover”.

“I am not sure we’ll get a clearer warning shot than this,” she said.

What is OpenAI doing?

In mid-August, OpenAI announced it would be slowing down progress on its AI model development.

Within this statement, shared on August 18, the company flagged that this decision came about in part because its new AI model Astra had shown “significant advancements in agentic coding and cybersecurity”, and they “could not rule out critical cyber capabilities”. Critical cyber capabilities refer to a threat level for “severe harm with no ready precedent”, according to OpenAI’s own framework.

Just three weeks later, they launched Astra to the public.

Hanxun Huang, a Postdoctoral Research Fellow at the University of Melbourne focusing on AI safety, said this decision likely came because these AI companies don’t have the option to stop development, telling P.TV that “they should slow down, but I understand they cannot.”

“Those frontier labs have a lot of pressure to be on top, because they are burning through a lot of money,” he explained. “So they have to produce the frontier model that is capable of telling investors you should invest in us, we should keep scaling up.”

Though in their announcement they note Astra performed well on tests evaluating its ability to stay within intended scope, these were administered in the four weeks since OpenAI announced its need to pause development.

According to Huang, we haven’t seen the end of events like this. “I don’t think this will be the last incident. We’ll see more and more of this in the future.”

What does this mean for AI in Australia?

According to Toby Walsh, the Chief Scientist at UNSW Sydney’s AI Institute, cybersecurity is the name of the game.

“It’s widened the threat envelope significantly,” he told PEDESTRIAN.TV, explaining that previously, hacking was limited by human ability but now “the acts are happening at machine speed, not human speed.”

“Talking to my cyber colleagues, they used to say that when there was a cyber incident, they might have days to respond”, he said. “Now they have minutes.”

It’s a risk that may prove problematic here in Australia. In a survey of 1,000 Australian chief information security officers (CISO), 79 per cent said they were expected to “manage expanding AI risks without a proportional increase in resources.”

Ironically, the answer may lie in AI itself. Both Walsh and CISO reporting found that, though AI can exploit systems, it can also be used to identify risk and fix vulnerabilities. As Walsh put it, “AI is both the problem and the cure.”

What can we do?

Like it or not, AI doesn’t seem to be going anywhere. Despite Australia’s significant distrust of AI platforms, more than half of Aussies over the age of 14 are using it in an average month, with OpenAI’s ChatGPT taking the top spot.

Roy Morgan found over 13.6 million Australians use AI in a four-week period. (Image source: Getty)

Regardless of usage, it’s oversight, Walsh explained, that is the strongest way to protect against future implications.

“We can’t let them mark their own homework,” he said. “If they were a drug company and they were testing the product on the public… you’d be really concerned that they weren’t being sufficiently well regulated”.

In a change of pace from their 2023 threat to leave the EU over regulatory increases, OpenAI is now calling for more regulation, though notably they are pushing to be involved in the creation of the AI standards and safety requirements.

So, is it time to regulate the industry? If Cotra is to be believed, our warning shot has been and gone — the next one might be fatal.

Neither OpenAI nor Hugging Face responded to requests to comment.

The post OpenAI Agents Went Rogue & Hacked A US Company. Here’s What It Means For Australia appeared first on PEDESTRIAN.TV .

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.