Days after an Anthropic researcher resigned, warning of dire: A practical reader guide

Days after an Anthropic researcher resigned, warning of dire consequences, another employee on the company’s safety team also stepped down last week, raising similar concerns.

“A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. Competition pushes every frontier company to underinvest in safety; the cost of falling behind is too high. We’re now starting to see increasingly concerning real-world incidents, including hundreds of agents from OpenAI recently hacking HuggingFace, and Anthropic models socially engineering people on the internet. Benton urged the public to demand greater transparency from AI companies, including information on their progress towards recursive self-improvement, safety incidents and near-misses, and minimum safety standards. He also called for independent guarantees that companies are meeting those standards. “The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. “Humanity may not survive this transition. We need a lot more preparation to make this world safe. He further warned that humanity could be permanently disempowered by the AI systems these companies build in the next few years, if priority is not given to safety. “If the pace of progress continues and the industry does not prioritize safety more heavily, I expect much worse to come: humanity could be permanently disempowered by the AI systems these companies build in the next few years.

Referring to his manager at Anthropic, Benton said that the chances of AI killing humans is greater than 10 per cent.

He also raised concerns over the lack of safety being used by AI companies, saying that a company could lose control of its system without the public ever knowing, adding that is not “acceptable” for a technology that possesses extinction-level risks. In a longer substack article, he said that AI companies are racing to build “superintelligence”, or an AI system that can recursively self-improve, warning that these AI systems may have drives and desires that diverge from those of any human overseer, with capabilities that can’t be effectively constrained. He added that many of his colleagues are “terrified” by the risks of systems being built.

“I left Anthropic’s safety team two weeks ago. Now feels like a good moment to explain why,” he wrote in a post on X. “AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. “Many of the colleagues I had at Anthropic are terrified by the risks of the systems they are building. I spent the last three years doing pretraining research at both OpenAI and Anthropic. “They are racing straight to self-improving superintelligence and gambling with our lives.

Evan Hubinger managed me while I was at Anthropic, and when he says that he thinks the chance that AI kills us all is greater than 10% he means it – he’s been working on these problems for almost a decade, long before there was money to be made from LLMs. Because these AI companies are gambling with lives in their rush to improve superintelligence, jacob Coxon, who stepped down as a researcher from Anthropic, said in a post on X that he is quitting the company — as he had previously quit OpenAI —.

Joe Benton, who describes himself on social media as the former manager of Anthropic’s Scalable Oversight team, has raised concerns that AI companies are building machines “smarter” than humans, warning that humanity might not be able to survive this. Benton, who had worked on AI pretraining research at OpenAI and Anthropic, announced his resignation last week, saying that neither company was acting responsibly and that they were “racing straight to self-improving superintelligence” and “gambling with our lives. Neither company is acting responsibly,” Coxon said. Coxon warned against underestimating the power of AI systems — predicting that they would soon be able to “hack anything, revolutionize any field overnight, and acquire real power and resources”.

“I resigned from Anthropic today.