Get all your news in one place.
100's of premium titles.
One app.
Start reading
Fortune
Fortune
Sharon Goldman

Inside Anthropic’s ‘Red Team’—ensuring Claude is safe, and that Anthropic is heard in the corridors of power

(Credit: Courtesy of Anthropic)

Last month, at the 33rd annual DEF CON, the world’s largest hacker convention, in Las Vegas, Anthropic researcher Keane Lucas took the stage. A former U.S. Air Force captain with a PhD in electrical and computer engineering from Carnegie Mellon, Lucas wasn’t there to unveil flashy cybersecurity exploits. Instead, he showed how Claude, Anthropic’s family of large language models, has quietly outperformed many human competitors in hacking contests—the kind used to train and test cybersecurity skills in a safe, legal environment. His talk highlighted not only Claude’s surprising wins but also its humorous failures, like drifting into musings on security philosophy when overwhelmed, or inventing fake “flags” (the secret codes competitors need to steal and submit to contest judges to prove they’ve successfully hacked a system).

Lucas wasn’t just trying to get a laugh, though. He wanted to show that AI agents are already more capable at simulated cyberattacks than many in the cybersecurity world realize—they are fast, and make good use of autonomy and tools. That makes them a potential tool for criminal hackers or state actors—and means, he argued, that those same tools need to be deployed for defense. 

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.