top of page

The Benchmark That Travelled Was The One That Flattered The Seller

Z.ai published three cybersecurity scores for its new model, and only the winning one made it into the coverage.



Sign up for my Substack daily AI newsletter here.


See my AI Training course portfolio for corporate Business Leaders here.



See my AI Training course portfolio for corporate Business Leaders here.


One disclosure before anything else. The comparisons in this article are a Chinese laboratory measuring itself against models built by Anthropic and OpenAI, and I use several of those tools in my own work. This is not a piece about whose model is better. It is about what happens to a number between publication and the boardroom.


CONTEXT AND BACKGROUND

In August, the Chinese laboratory Z.ai released GLM-5.3 and reported that it scored eighty-four point five per cent on CyberGym, a benchmark measuring how well a model finds and validates software vulnerabilities. That placed it ahead of Anthropic’s Mythos 5 at eighty-three point eight and OpenAI’s GPT-5.6 Sol at eighty-three point six. An open-weight model had apparently taken the lead from the two most capable closed systems in the world, and that is the sentence that travelled. It is worth noting that the weights were not actually available at the time. Z.ai withheld them at launch because of the cybersecurity capabilities, promising release about two weeks later after safety evaluation, so the model that appeared to have opened the frontier was reachable only through a paid subscription.


The same release contained two further charts. On ExploitBench, which measures later stages of the work, GLM-5.3 scored fifty-four point four per cent against seventy-eight for Mythos 5 and seventy-six point five for Sol. On ExploitGym it completed one hundred and five tasks within a two-hour budget, where Mythos 5 completed one hundred and eighty-one and Sol two hundred and sixteen.


INSIGHT AND ANALYSIS

Read together, the three charts say something quite specific. The model is very good at finding flaws and considerably weaker at turning them into working exploits. That is a meaningful distinction, and it is closer to the opposite of what a single score implied.


Nothing here was fabricated. Z.ai published all three results and stated plainly in its own announcement that closed models remain ahead on exploitation. The company was more candid than the coverage was. The distortion happened afterwards, in the compression from three charts to one number.


The methodology deserves the same attention. Z.ai ran the evaluations itself, using Anthropic’s own coding harness, at maximum reasoning effort, with no web tools and a single run per task with no time limit. On the time-normalised test, the budgets were not wall-clock limits but calculations rescaled using tokens-per-second figures that differ by model, drawn from a third party but selected by Z.ai, which assigned one hundred and fifteen tokens per second to its own system against forty for a competitor. That choice alone shapes how many tasks each model appears to complete.


There was no independent replication at release. The company chose the tests, the harness, the settings and the comparisons, and then scored its rivals on its own equipment.


None of this makes the results false, and it is not a Chinese peculiarity. Self-reported benchmarking is the industry norm, and the same questions apply to every capability claim your suppliers make. Publishing the unflattering charts alongside the flattering one is more disclosure than many vendors offer.


IMPLICATIONS

For anyone evaluating AI capability claims, three questions follow from this episode. How many measures were published, and which one is being quoted to you? Who ran the test, on whose equipment, with which settings? And has anyone outside the company reproduced it?


The first question matters most, because the failure here was not deception. It was selection. A vendor with three results will lead with one, a journalist will report the lead, and by the time the claim reaches a board paper it has become a single sentence with no chart attached.


For South African organisations weighing open-weight models against subscription access to closed ones, the practical point is that the decision cannot rest on a leaderboard. What matters is whether the model performs on your work, and the only way to establish that is to test it on your own tasks, which almost nobody does before signing.


CLOSING TAKEAWAY

The most useful thing about this release is that all three numbers were published. The company did not hide the ones that undercut its headline, which means the record is available to anyone who reads past the first chart.

Very few people did. That is the finding, and it will repeat with every model release from every laboratory in every country until buyers start asking what else was measured.


Author Bio: Johan Steyn is a prominent AI thought leader, speaker, and author with a deep understanding of artificial intelligence’s impact on business and society. He is passionate about ethical AI development and its role in shaping a better future. Find out more about Johan’s work at https://www.aiforbusiness.net

 
 
 

Comments


Leveraging AI in Human Resources ​for Organisational Success
CTU Training Solutions webinar

bottom of page