He Quit Anthropic Over AI Safety. His Boss Agreed With Him

An Anthropic researcher resigned over AI safety this week, and the company's own alignment lead publicly agreed. Warning Shots #58 breaks it down.

Written by
on
Sep 13, 2026

An Anthropic Researcher Quit Over AI Safety, and His Own Company Agreed With Him: This Week on Warning Shots

An AI safety researcher resigned from Anthropic this week, walking away from equity that had not yet vested to say publicly that the industry is racing ahead without a real plan for keeping advanced AI safe. Within a day, Anthropic's own alignment stress-testing lead posted that he agreed. Within another day, senators and representatives from both parties were talking about AI extinction risk on the record, and prediction markets were repricing the odds of a federal AI safety bill.

On this week's Warning Shots, John Sherman, Liron Shapira and Michael work through what happened, why this particular resignation broke through when others hadn't, and what a swarm of AI agents solving a famous unsolved math problem the same week says about where the underlying technology actually stands. Here is what stood out.

What Jacob Coxon actually said

Jacob Coxon spent roughly four months at Anthropic after three years at OpenAI. According to his own account, he resigned before his equity was scheduled to vest at the six-month mark, giving up that unvested stock rather than staying quiet. His stated reason, in his own words: "I no longer have anything to gain by juicing up Anthropic's valuation," and "the people building AI earnestly believe that it could kill us all by the end of the decade." He also pointed to competitive pressure inside the industry as the mechanism, arguing that when companies are racing each other, safety steps are the ones most likely to get cut.

What made this one different from earlier departures at other labs, according to the hosts, is that Coxon had no obvious financial motive for speaking up. He gave up money to say it rather than gaining money by staying quiet, and the post reportedly drew more than 100 million views within days, an unusually large number for a single individual's resignation announcement.

Anthropic's own alignment lead agreed, on the record

The part of this story that is harder to wave away is what happened next. Evan Hubinger, Anthropic's alignment stress-testing lead, posted publicly: "Jacob is correct here, we really do earnestly believe AI could kill all humans. I personally think it is more than 10 percent within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to get one." This is not an outside critic or a departing employee. According to the hosts, Hubinger's job is literally to stress-test Anthropic's own safety assumptions, and he chose to make that admission in public rather than only internally.

Congress responded within hours

According to the hosts, dozens of senators and representatives from both parties posted publicly about AI extinction risk and the need for oversight within roughly 24 hours of Coxon's post, a pace neither host said they had seen before on this topic. Sen. Bernie Sanders, a longtime advocate of restricting the development of superintelligent AI, was among the most visible responses. Sen. Josh Hawley separately sent a letter raising concerns about how OpenAI has handled its own AI safety testing. Prediction markets tracking whether a federal AI safety bill will pass before 2027 moved up over the following days, though the hosts cautioned that a moved prediction market is not the same as passed legislation, and that the same week's news cycle could just as easily fade without producing a vote.

A math problem, solved by a swarm

The same week, a separate story showed how fast the underlying capability curve is moving. According to reporting on OpenAI's own announcement, an internal, unreleased model directed a swarm of roughly 10,000 AI agents working for about 88 hours to produce a solution to the Navier-Stokes equations, one of the seven Millennium Prize Problems in mathematics, each carrying a $1 million prize from the Clay Mathematics Institute. The effort reportedly began after a rumor spread that a rival researcher was close to a similar result, and it has since become entangled in a dispute over whether OpenAI's system may have benefited from that researcher's own unpublished work. OpenAI has said it cannot rule out that de-identified user data played a role in improving its models, while maintaining that the two approaches appear different once compared side by side.

On the show, the hosts' point was less about who deserves credit and more about the scale of what a swarm of agents can now do inside a few days on a problem that had resisted human mathematicians for decades.

Is the AI 2027 forecast tracking reality

The hosts also discussed former OpenAI researcher Daniel Kokotajlo's recent appearance on the Joe Rogan podcast, where he discussed the AI 2027 forecast he co-authored, a detailed, published scenario for how AI development could unfold through 2027. Kokotajlo, who like Coxon left equity behind to speak more freely after departing OpenAI, argued that several of the forecast's specific predictions have tracked closely to what has actually happened so far. According to the hosts, that track record is part of why they take the forecast's later, more severe predictions more seriously than they otherwise might.

What this adds up to

According to the hosts, the throughline across a resignation, an internal safety lead's public agreement, a swarm of agents solving a problem that stumped mathematicians, and a forecast tracking closer to reality than they'd like, is the same question the show keeps returning to: the industry's own safety researchers are saying, in public and on the record, that they do not yet have a plan, while the underlying systems keep getting more capable regardless. John Sherman closed the episode with a story of his own: at an advertising industry event this week, he asked a room of professionals who use AI daily whether they would go back to a time before it existed. Every hand in the room went up.

Watch Warning Shots #58 on The AI Risk Network.

Read the full deep dive, with sources and graphics, on Substack.

Warning Shots is a weekly show from The AI Risk Network with John Sherman, Liron Shapira of Doom Debates, and Michael of Lethal Intelligence. Figures and disclosures discussed in this episode are reported by the hosts and by this post from public sources, and are presented as such.