Eigen RadarAI
Analysis

Jacob Coxon leaves Anthropic, citing the race to superintelligence

Jacob Coxon, who spent three years on pre-training research at Anthropic and OpenAI, resigned and left the field, saying both labs were racing irresponsibly towards self-improving superintelligence. Anthropic Alignment Science lead Evan Hubinger wrote that he personally puts the chance of AI killing all humans within the next decade above 10 per cent. These are the researchers' own assessments, not an independent estimate of the risk.

Artificial Intelligence··Evening
An unmarked access badge sits sharply on a security desk as an anonymous researcher walks away from a bright glass-walled laboratory toward a warm-lit corridor.

Coxon leaves the field after three years

Jacob Coxon resigned on 9 September and said he was leaving the AI field after spending a combined three years on pre-training research at Anthropic and OpenAI. Coxon wrote that neither company was acting responsibly and that both were racing to reach self-improving superintelligence first. ABC News and The Guardian report the resignation and Coxon's stated reason as the same development. His account is a personal objection grounded in his work inside the companies, not an independent review of their research programmes.[1], [2]

Hubinger states his personal risk estimate

Anthropic Alignment Science lead Evan Hubinger responded to Coxon on X. He wrote that people at the company genuinely believe AI could kill all humans and put his own estimate above 10 per cent within the next decade. Hubinger also said Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to one. The figure is Hubinger's personal assessment; it was not presented as a measured probability or Anthropic's corporate forecast.[1], [2]

Anthropic points to research on pacing development

Anthropic pointed to its published work on recursively self-improving AI as a basis for tools intended to pace frontier development. The company did not issue a separate response to Coxon's resignation. Coxon's criticism of the race, Hubinger's risk estimate and Anthropic's research response are therefore different kinds of evidence: the first two describe researchers' views, while the third identifies the technical line of work cited by the company. The reports establish that these statements were made; they do not independently calculate the probability of human extinction.[1]

References

  1. News sourceABC NewsAnthropic researcher Jacob Coxon quits, saying the labs are racing to superintelligence↩1↩2↩3
  2. News sourceThe GuardianJacob Coxon resigns as Anthropic researchers warn of human extinction by 2030↩1↩2