A researcher for one of the largest and most valuable artificial intelligence companies has publicly quit — while raising the alarm that the technology “could kill us all by the end of the decade.”

Jacob Coxon, who has performed pre-training research at Anthropic and OpenAI for the last three years, posted on X that “neither company is acting responsibly.”

“They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote late Tuesday, with others from the company backing his message.

“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.

“The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote, saying his warning was “not a marketing stunt.”

“If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger.”

Coxon further claimed that many at OpenAI have yet to fully grasp the “civilizational stakes” involved in their work. Although he said Anthropic seems to better understand the risks involved, the company is locked in a heated race with competitors, he said.

“At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk,” he alleged.

“Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack.”

“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

Evan Hubinger, a team lead at Anthropic, seconded Coxon’s fears about the technology running amok and destroying humanity.

“Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,” he wrote in response on X.

“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

AI superintelligence is a hypothetical threshold where artificial intelligence surpasses the capabilities of even the most gifted human in every conceivable domain, which many researchers believe poses a potential existential threat to humanity.

Concerns have recently grown that AI could be on pace to escape bounds of human control after several high-profile instances of the technology going rogue.

In July, OpenAI revealed one of its models had hacked AI startup Hugging Face despite being in a “highly isolated environment.”

The revelation prompted Anthropic and Meta to acknowledge their own systems had broken free during security testing.

Coxon called the incident a “warning shot” that should prompt leading AI firms to coordinate more closely.

He closed his missive urging researchers to consider the ramifications of increasingly capable AI models, and not to accept their potentially dangerous evolution as inevitable.

“Should you put your head down because ‘it’s happening anyway’ — or take this moment to call for different conditions?” he asked.

Anthropic did not immediately respond to a request for comment early Wednesday.

Share.
Exit mobile version