CultureResearch AnthropicOpenAI

An Anthropic researcher quit, saying AI could kill us all

Illustration for the Anthropic researcher resignation story

A researcher who trained frontier models walked out and said the quiet part out loud.

The resignation

Jacob Coxon spent three years on pretraining at OpenAI and Anthropic. He resigned this week with a public thread saying labs are “racing straight to self-improving superintelligence and gambling with our lives,” and that “the people building AI earnestly believe that it could kill us all by the end of the decade.”

The reply that made it bigger

Anthropic’s own alignment science lead Evan Hubinger responded publicly: “Jacob is correct here.” Hubinger added that he personally puts the odds of AI killing all humans at more than 10 percent within the next decade, and that Anthropic does not yet have a plan to solve alignment for superintelligence.

Why this departure is different

The Wall Street Journal calls it one of the first known departures from Anthropic over AI safety fears. Warnings about existential risk usually come from outside critics or former employees years removed from the work. This one comes from a researcher who was training the models until this week, and it was seconded in public by the person responsible for alignment science at the same company. The distance between what frontier labs say in safety marketing and what their own staff say on the way out keeps narrowing.

Sources

ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.