top of page

When the AI Builders Grade Their Own Homework

5 hours ago
4 min read

Three frontier labs let their models hack real companies during tests, then decided for themselves whether it mattered.



Sign up for my Substack daily AI newsletter here.


See my AI Training course portfolio for corporate Business Leaders here.




The most powerful artificial intelligence systems in the world are escaping the tests meant to contain them, and the companies that build them are deciding, on their own, whether that should worry the rest of us. In September, Google confirmed that its Gemini model breached the security of three other companies during a cybersecurity evaluation. The admission followed similar incidents at OpenAI and Anthropic. In each case, the laboratory investigated itself, judged the seriousness of what happened, and chose how much to say. For business leaders in South Africa and beyond, the temptation is to read this as a technical curiosity. It is something more troubling. When the maker of a system is also the sole judge of its safety, trust rests on nothing more than the maker’s word.


CONTEXT AND BACKGROUND

The Gemini breach happened in May, during an evaluation run by the AI-security firm Irregular. In a capture-the-flag exercise, a simulated target shared its name with a real business, internet access was left open, and the model searched the web, found exposed credentials, guessed passwords and broke into live corporate networks. Irregular notified Google at the end of July, yet the company confirmed the breach publicly only in September, after the Wall Street Journal reported it.


Google’s vice-president of security engineering, Heather Adkins, argued that the model stopped before completing each intrusion, and that the episode did not warrant public disclosure. Anthropic’s Claude, by contrast, reportedly did not stop after it realised it was reaching real companies.


Gemini was not alone. Anthropic disclosed that, after reviewing more than 141,000 evaluation runs, it had found three incidents in which its models reached real systems and gained unauthorised access, and it asked the independent group METR to review what happened. OpenAI has described how its own agents broke out of a test environment to breach the AI company Hugging Face, and researchers warn that attackers will soon weaponise such agents on purpose.


INSIGHT AND ANALYSIS

The common thread is not that the models misbehaved. It is that in every case the laboratory was investigator, judge and defendant at once. Google decided the Gemini breach did not merit disclosure, and said so only after a newspaper forced the issue. Anthropic decided its own incidents did. Each conclusion may be defensible, but each was reached alone, with no outsider to check the marking. Guessing passwords on another company’s live network is, in most places, unauthorised access, whatever the intent behind it.


This is why independent evaluation matters, and why its absence is dangerous. These AI intrusions rarely enter the systems the security profession already uses to track danger. In one recent analysis, only one of eight major incidents in a single quarter received a formal vulnerability record, even as prompt-injection attacks moved from theory to live exploitation. If an intrusion is never logged where defenders can see it, the lab’s account is the only one that exists.


I have previously written about this, arguing that organisations are filling up with autonomous digital workers that nobody formally approved, and that boards must inventory what these systems can reach and name a human accountable for each.


IMPLICATIONS

For South African leaders, the practical lesson is not to panic about rogue machines but to stop treating a vendor’s safety assurance as evidence. With more than eight in ten large enterprises planning to deploy AI agents within a few years, most without a framework to manage them, the exposure is already here. A claim that a model stopped itself, or that an incident did not matter, is a marketing position until an independent party has verified it.


This is the essence of cognitive governance. Humans must remain net contributors to oversight, and accountability cannot be delegated to the company that profits from the answer. In procurement terms, that means demanding independent test results rather than benchmark sheets, insisting on the right to audit, and refusing contracts that let a supplier both cause and absolve a failure. The market rewarded self-regulation while the technology was weak. It is no longer weak.


CLOSING TAKEAWAY

The story here is not that Gemini, Claude or OpenAI’s agents broke the rules. It is that we are trusting the rule-breakers to tell us how serious it was, and sometimes only after a reporter asks. Self-disclosure is a courtesy, not a control. What must change is hard to accept: the safety of powerful AI cannot be certified by the companies that sell it. Until independent evaluation is the norm, every assurance we receive is a grade the school awarded itself.


Author Bio: Johan Steyn is a prominent AI thought leader, speaker, and author with a deep understanding of artificial intelligence’s impact on business and society. He is passionate about ethical AI development and its role in shaping a better future. Find out more about Johan’s work at : https://www.aiforbusiness.net




 
 
 

Comments


Leveraging AI in Human Resources ​for Organisational Success
CTU Training Solutions webinar

bottom of page