top of page

The Gap Between What Frontier Labs Say and What Their Own Researchers Believe Is a Governance Problem

11 minutes ago
5 min read

Reading the Coxon resignation and the Hubinger confirmation as a due-diligence question for every board buying AI-embedded products.



Sign up for my Substack daily AI newsletter here.


See my AI Training course portfolio for corporate Business Leaders here.




A few days ago, a twenty-seven-year-old pretraining researcher named Jacob Coxon resigned from Anthropic and posted a seven-part message on X explaining why. Within days, the thread had drawn over one hundred million views and a Wall Street Journal exclusive interview. Within a day, a senior colleague on Anthropic’s own alignment team had publicly agreed with him. That second event is the one South African boards should be paying attention to, and it is not the one dominating the headlines.


CONTEXT AND BACKGROUND

Coxon spent three years doing pretraining research at OpenAI and Anthropic, contributed to the GPT-4o model, and left Anthropic roughly four months into his tenure there, ahead of the company’s expected initial public offering. In his resignation post, he wrote that neither company was acting responsibly, that both were racing towards self-improving superintelligence, and that the people building these systems earnestly believe they could kill us all by the end of the decade.


He told Axios that he walked away from unvested equity worth a significant sum in order to say so, because he could not credibly criticise from inside. His own risk estimate is more careful than the headlines suggest. There is no risk of extinction right now, he said, but the pace of development could make the picture very different within years.


Any of that could have been dismissed as one researcher’s personal position. What made the story different was what came next. Evan Hubinger, the alignment science lead at Anthropic and one of the company’s most senior technical safety figures, publicly agreed with the underlying warning and put his own probability estimate for mass extinction within the decade at greater than ten per cent. Yoshua Bengio, one of the three researchers who won the 2018 Turing Award for the foundational work on deep learning, endorsed the substance of the concern from outside the company. Coxon’s resignation was covered by the Wall Street Journal as the news event. The endorsement from inside Anthropic’s own alignment leadership, with a number attached, is the governance event.


INSIGHT AND ANALYSIS

For most of the past two years, the standard corporate response to public warnings about AI risk has been that the people making them are marketing themselves, seeking regulatory moats, or pattern-matching to science fiction. That response has kept the conversation about frontier risk out of most board rooms, and out of most procurement conversations, on the reasonable-sounding grounds that the critics were not credible. It is worth stating plainly what has just happened. A pretraining researcher walked away from unvested equity in the middle of an initial public offering process to say the frontier race is dangerous. His employer’s own alignment science lead publicly agreed and put the number at greater than one in ten. A Turing laureate endorsed the substance. The dismissal of concern as marketing no longer survives contact with the evidence.


The specific governance problem for boards buying AI-embedded products from these companies is the gap between what the company says publicly and what its most senior technical safety staff believe privately. If the private view of the alignment leadership at Anthropic puts extinction risk this decade above one in ten, and the marketing prospectus does not, that is a disclosure question of the kind boards are used to asking about pharmaceutical suppliers, defence contractors, and mining operators. It is not exotic. It is ordinary corporate due diligence applied to a category of supplier that has so far been treated as exempt.


The same standard applies at the supplier level. A South African bank, insurer, hospital or university buying an AI-embedded product is entitled to know how the supplier’s own safety leadership rates the risks of that product. Coxon and Hubinger have just made it possible to ask the question in specific terms. A year ago the answer would have been that no such view had been publicly disclosed. This week it has been.


IMPLICATIONS

Three practical implications follow for South African boards. The first is that AI supplier due diligence needs to be added to the risk register at the same level as data protection due diligence. The second is that the questions asked of the supplier have to include what its own alignment or safety leadership believe about the risks of its most capable models, and what mechanism exists for their view to reach the customer. The third, and the one most likely to be resisted, is that a supplier’s public marketing on safety is now known to be a partial view of what its safety staff think. That is not an accusation against any particular company. It is a description of what Coxon and Hubinger have just placed in the public record.


The Anthropic initial public offering is being pursued at a valuation reported at around two trillion United States dollars, on the strength of the company’s safety-first positioning. South African institutional investors are among the buyers of exposure to these companies through offshore allocations. The disclosure gap that the Coxon resignation and the Hubinger confirmation have just opened up is not only a governance question for the boards buying the products. It is a fiduciary question for the boards buying the equity.


CLOSING TAKEAWAY

The Coxon resignation will be a headline for a fortnight and then it will be gone. The gap it has revealed will not be. Frontier labs have privately held risk views that are materially more pessimistic than their public statements, and now that the gap has been documented on the record by the labs’ own senior technical staff, it is no longer defensible for South African boards to accept the public statements at face value. The right response is not panic and it is not disengagement. It is the same response boards would apply to any supplier whose internal risk view had just become public and had turned out to diverge from the marketing. Ask for the internal view in writing. Name a human at your own institution who is accountable for reading it. Decide, in public and on the record, whether you are comfortable with what you find. That is what cognitive governance requires when the supplier is a frontier AI lab, and it is what the events of this week have made unavoidable.


Johan Steyn is a prominent AI thought leader, speaker, and author with a deep understanding of artificial intelligence’s impact on business and society. He is passionate about ethical AI development and its role in shaping a better future. Find out more about Johan’s work at https://www.aiforbusiness.net


 
 
 

Comments


Leveraging AI in Human Resources ​for Organisational Success
CTU Training Solutions webinar

bottom of page