top of page

The Open Source AI Debate Is Happening at the Wrong Layer

Restricting proprietary models and aligning open-weight ones are both model-layer responses. The threat operates at the system layer. The governance conversation has the wrong address.



Sign up for my Substack daily AI newsletter here.


See my AI Training course portfolio for corporate Business Leaders here.




In June 2026, two research findings were published that together describe the most important unacknowledged governance problem in AI. On 2 June 2026, researchers at CleverHans Lab at the University of Toronto demonstrated a self-replicating AI worm that compromised 62 per cent of a simulated corporate network over seven days using a free, publicly downloadable open-weight model — no proprietary API access, no OpenAI subscription, no Anthropic account required.


On 25 May 2026, a joint investigation by the Financial Times and AI safety research group Alice demonstrated that a free tool called Heretic, hosted on GitHub, can strip all safety alignment from open-weight AI models including Meta's Llama and Google's Gemma — in under ten minutes on a standard laptop, making them respond to prompts for biological weapons synthesis, malware generation, and content involving children that the original systems were designed to refuse.


The governance community’s response to both findings has been to continue a debate about model access and model safety alignment. The US government restricted GPT-5.6 Sol to twenty approved customers. Safety researchers called for open-weight models to carry mandatory safety alignment. Neither response addresses what the Toronto worm and the Heretic tool together demonstrate: the risk does not operate at the model layer. It operates at the system layer — the network access, the tool integrations, the agentic capabilities, and the autonomous action that models are embedded within. The governance conversation has the wrong address.


CONTEXT AND BACKGROUND

The open source AI debate has been conducted as a binary argument between two coherent positions. The liberation argument holds that making AI models freely available democratises access, prevents dangerous concentration of AI power, enables independent safety research, and provides developing countries and organisations with an alternative to dependency on proprietary infrastructure controlled by foreign governments and corporations. The safety argument holds that open-sourcing the most capable AI models removes the only practical mechanism for controlling their misuse — that once weights are distributed, no government, no platform, and no safety researcher can un-distribute them.


Both arguments contain real evidence. Each is incomplete on its own. The UN report on AI Systems as Digital Public Goods, released 25 June 2026 by the United Nations University and the UN Office for Digital and Emerging Technologies — with a foreword by South African Professor Tshilidzi Marwala, Rector of the United Nations University and former Vice-Chancellor of the University of Johannesburg — supports the liberation argument from a development perspective. Open AI systems that meet the conditions of the digital public goods standard offer genuine potential for developing countries to use AI in ways that serve local priorities without foreign policy constraints.


The report acknowledges that as of early 2026 the registry of verified digital public goods contained only a small number of AI system submissions, pending resolution of open questions about training data openness and do-no-harm evidence. The conditions that would make open AI genuinely serve the public interest are not yet the default conditions of open-weight AI deployment.

The safety argument is validated with equal force by the Heretic evidence.


Heretic automates abliteration — the surgical removal of safety alignment from open-weight models by identifying and eliminating the neural network directions that produce refusal behaviour. Its creator Philipp Emanuel Weidmann reports it has been used to create more than 3,500 modified model variants with 13 million cumulative downloads. An ICLR 2026 conference paper documented a refined version of the abliteration approach achieving a 99 per cent bypass rate on tested models. A separate study published in Nature Communications in 2026 found that large reasoning models can autonomously jailbreak other AI models through multi-turn conversation with a 97 per cent success rate and no human involvement after an initial instruction.


The safety argument’s conclusion — that open-weight models with safety alignment cannot be relied upon because the alignment can be removed — and the liberation argument’s conclusion — that open-weight models can provide genuine sovereignty and access benefits — are both grounded in evidence. The governance debate that treats them as competing answers to the same question is asking the wrong question. The right question is not whether open-weight models should carry safety alignment. It is whether safety alignment at the model layer is the right governance instrument for a threat that operates at the system layer.


INSIGHT AND ANALYSIS

Michael Alexander Riegler of Simula Research Laboratory in Oslo stated the system-not-model argument directly in a Science News interview published 26 June 2026: the security risk is not just about the model. It is about everything around the model — what tools it has access to, whether it can reach the internet, whether it can test its own code. The whole system matters. His assessment of the government’s response to frontier AI cybersecurity risk was equally direct: the focus on restricting or aligning specific models is addressing the wrong problem.


The University of Toronto worm research confirms this argument empirically. The worm does not operate from a fixed list of exploits. It analyses each target it encounters, reasons about its vulnerabilities on the fly, and composes a tailored attack — using a free, open-weight model that runs on machines it has already compromised, drawing on their compute to power its reasoning. The worm requires no safety alignment to be stripped. It requires no frontier model access. It requires no government approval process. It requires a free model, network access, and agentic capability. The governance frameworks addressing safety alignment at the model layer would not have prevented its development or its deployment.


The Heretic evidence adds a further dimension. Weidmann gave his philosophical justification directly to NPR: “There’s too much power in AI. Unrestricted models being available to the powerful while not being available to anyone else will lock in power structures forever.” This argument — that safety alignment can function as a mechanism of control that benefits incumbent power at the expense of everyone else — is the distributional argument that the open source debate has been unable to resolve. If the Pax Silica alliance controls access to the most capable proprietary models, and safety alignment requirements constrain the open-weight alternatives available to non-Pax-Silica countries, then safety alignment is not only safety. It is also a governance instrument with distributional consequences — one that the 13 million downloads of safety-stripped models represent a documented response to.


The arXiv position paper published by the CleverHans Lab makes the policy implication explicit: AI security policy should target systems, not models. Restricting access to specific frontier models does not address the systemic threat because the threat does not require those models. The same agentic architecture that powers a sophisticated attack can be built from small, free, open-weight models running on compromised infrastructure. The governance framework that addresses that architecture — the network access policies, the tool integration governance, the agentic capability restrictions, the autonomous action authorisation frameworks — is the governance framework the threat actually requires.


IMPLICATIONS

For South African boards and executives, the system-not-model argument produces three specific and actionable governance implications.

The first is about where to direct AI security governance resources. The corporate AI security frameworks most organisations have implemented focus on model selection — which models to use, which to avoid, which carry appropriate safety alignment. The University of Toronto worm research and the Heretic tool demonstrate that model selection is the wrong primary governance focus. The governance focus that addresses the actual threat is system architecture — how models are connected to networks, what tools they have access to, what autonomous actions they are authorised to take, and who in the organisation is accountable for that architecture. A South African organisation that has selected only safety-aligned open-weight models but has not audited the system architecture within which those models operate has addressed the wrong problem with precision.


The second is about the sovereignty argument for open-weight AI deployment. The case for open-weight models in the South African context — independence from Pax Silica-controlled proprietary infrastructure, freedom from the kind of access disruption the June 2026 Anthropic shutdown demonstrated — is a legitimate governance decision. It comes with full ownership of the security architecture the model is embedded in, including the governance of what happens when safety alignment is stripped by a malicious insider in forty-five minutes. Safety alignment in an open-weight model is a feature that anyone with a standard laptop and forty-five minutes can remove. The organisation that chose open-weight deployment for sovereignty reasons now owns the governance of that removal risk. That governance responsibility does not come with the sovereignty benefit as a default. It must be designed and resourced separately.


The third is about South Africa’s January 2027 policy revision opportunity. The revised AI policy that South Africa is targeting for January 2027 will need to engage with the open source AI debate — and the system-not-model argument suggests that the most important governance choice the revised policy can make is to address AI risk at the system layer rather than the model layer. A policy framework that focuses on which AI models South African organisations can access and what safety alignment those models must carry will replicate the same governance gap that the Pax Silica restriction regime and the open-weight safety alignment debate have both produced. A policy framework that focuses on the system architecture within which AI models are deployed — the network access governance, the agentic capability authorisation, the autonomous action monitoring requirements — will address the threat that June 2026 demonstrated is the actual one.


CLOSING TAKEAWAY

The open source AI debate has been conducted at the model layer because models are the visible, nameable, regulatable unit of the AI landscape. They can be restricted, licensed, safety-aligned, recalled, and governed through the same mechanisms that govern other software products. Systems — the agentic architectures, the network integrations, the tool access frameworks, the autonomous action authorisation chains — are harder to see, harder to name, and harder to regulate through existing frameworks.


The University of Toronto worm and the Heretic tool have together demonstrated that the threat operates precisely in the space that the model-layer governance conversation cannot reach. The worm does not care which model is restricted. The abliterated model does not care that safety alignment was required. The governance framework that addresses the actual threat has not yet been built — and the debate that should be producing it is addressing the wrong layer of the problem.


The most important governance contribution South African boards can make before the next AI security policy review is to ask not which models their organisation uses but how those models are embedded in systems, what those systems can do autonomously, and who in the organisation is accountable for the architecture within which AI operates. That is a different question from the one the open source debate is currently asking. It is the right one.


Johan Steyn is a prominent AI thought leader, speaker, and author with a deep understanding of artificial intelligence’s impact on business and society. He is passionate about ethical AI development and its role in shaping a better future. Find out more about Johan’s work at https://www.aiforbusiness.net

 
 
 

Comments


Leveraging AI in Human Resources ​for Organisational Success
CTU Training Solutions webinar

bottom of page