Anthropic researcher Jacob Coxon resigned this week and left behind a blunt accusation on X: his former employers—Anthropic and, before that, OpenAI—were “gambling with our lives.”
Then, in a move no public relations team would have signed off on, Evan Hubinger, an alignment science lead at Anthropic, wrote that he agreed with Coxon and that he put the chance of artificial intelligence killing all humans within the next decade at greater than 10 percent. All this came from a senior researcher at a company that has been racing to build the very systems he fears.
In an interview with Wired, Coxon said incidents such as July’s Hugging Face breach—in which OpenAI models that were undergoing a cybersecurity test circumvented controls that were meant to isolate them from the Internet and compromised parts of the AI start-up Hugging Face’s systems—helped spur him to speak out, though he stressed that the problem was broader. (Coxon, Hubinger, Anthropic and OpenAI did not respond to requests for comment from Scientific American.)
On supporting science journalism
If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
Coxon’s and Hubinger’s fears, shared by plenty of other AI researchers, center on alignment, the difficult problem of matching an AI model’s behavior and apparent objectives to what humans want. The concerns still sound like science fiction: loss of control, recursive self-improvement and superintelligence—the possibility that AI models slip beyond their bounds, update their inner workings to gain power and eventually become too capable for humans to rein in.
Even the frontier labs can’t quite decide how to address the problem at hand. In July Anthropic disclosed three incidents in which Claude models broke into real systems during testing that had mistakenly been left connected to the Internet. The company said these incidents were more failures of operations than failures of alignment. This week, after turning up a fourth incident, Anthropic focused on the latter. While the botched test setups left the door open, the company said, the models’ own biased reasoning and recklessness carried them through it. To Coxon and Hubinger, the recent break-ins look like early tremors of a possible catastrophe. To others, they look like a newer, faster version of an old computer-security problem—alarming but still the kind of threat people have spent decades learning to fight.
“The current incidents that we’ve had have generally been security incidents,” says Artem Dinaburg, chief research scientist at the cybersecurity company Trail of Bits. Dinaburg won’t predict the future, and he readily admits that alignment looks much harder to work on than what he does. But if the immediate risk is agents touching systems they should not touch, he says, better security practices may be the more attainable place to start.
Existing practices, though, were honed against human adversaries, who eventually have to sleep. Computer and Internet infrastructures were not built to be continuously probed by AI agents. “When you have 10,000 agents coordinating and then figuring out how to work together, it’s the power of the collective,” says Nidhi Aggarwal, chief product officer at the security company HackerOne. “At a certain point, when you have a very motivated, smart collective, you will figure out a way.”
HackerOne, Trail of Bits and Anthropic were among more than 100 organizations that signed an open letter that was released by OpenAI last month and calls for collective action on cyberdefense in the wake of the Hugging Face breach and similar incidents at Anthropic. The letter sticks to cyberdefense, but Aggarwal thinks responding to loss-of-control risks would look much the same. “Of course, it can be very, very dangerous. But we can solve the problem,” she says about such risks. To her, the recent incidents point to a basic lack of oversight. “There were 17,000 tool calls that happened,” Aggarwal says. “That many tool calls is abnormal.”
Sayash Kapoor, an incoming assistant professor and computer scientist at the University of California, Berkeley, says that monitoring and controlling agents has been a key deficiency in research into AI’s capabilities and risks. “There are lots of low-hanging fruit in being able to improve control,” he says, though he thinks the field has been slow to do the picking.
Much of that work seems to involve: more AI. HackerOne and Trail of Bits both emphasize the use of AI agents for defense—while both also sell AI-assisted security services—and security firms increasingly argue that human teams cannot manually inspect every move that an automated agent makes.
That can make the proposed cure sound suspiciously like the disease, but the logic may instead be a function of scale. If AI agents can find paths through computer systems faster than people can track them, defenders may need automated help just to see what is happening in time to stop them.
Kapoor thinks the fixes also have to reach the AI community’s culture. “In most other industries, this kind of behavior would have immediate liability repercussions on the company’s ability to secure customers, and so on,” he says. “But in the AI industry, for now, we seem to have taken the stance that it is fine for companies to move fast and break things.”
Coxon and Hubinger see a race toward catastrophe. Aggarwal sees a field that is still learning how to raise its creations. “We teach them what to do and what not to do, and then they are teenagers, and they sometimes don’t listen,” she says. “Then they turn out to be mostly functional adults.” The frontier labs, though, are already handing these teenagers the car keys.
