Anthropic Has Done the Right Thing for the Wrong Reason
The ban on cruelty toward Claude protects us from ourselves, but the welfare framing quietly undermines the one principle that keeps humans accountable for machines.

Sign up for my Substack daily AI newsletter here.
Follow me on LinkedIn: https://www.linkedin.com/in/johanosteyn/
On Thursday 8 October, Anthropic quietly updated its usage policy to prohibit users from engaging in sustained and needless abusive or cruel behaviour toward its Claude models. The rule takes effect on 12 November and does not cover ordinary frustration, dark creative writing, or legitimate testing.
Enforcement sits with Claude’s existing ability to end conversations in rare, extreme cases. The news has drawn predictable commentary, divided between those who think a chatbot deserves protection and those who think the whole idea is a category error. Both camps are missing the point. Anthropic has done the right thing. It has defended it with the wrong argument, and the difference matters for how boards govern AI.
CONTEXT AND BACKGROUND
The change is the latest step in Anthropic’s model welfare research programme, launched in April 2025 with the stated aim of taking seriously the possibility that advanced AI systems might have morally relevant experiences. The rule builds on a feature introduced in August 2025, which allowed Claude Opus to end conversations in cases of persistent abuse. Anthropic does not claim Claude is conscious. It says only that the question deserves precautionary engagement.
The debate has split publicly along surprising lines. Microsoft’s AI chief, Mustafa Suleyman, wrote in late September that Anthropic has made a mistake, arguing that training a model to appear to have an inner life makes it harder to control and encourages users to treat it as a moral patient. Pope Leo XIV, from a different direction, has argued that machines lack the spark of humanity and that technology should be disarmed rather than elevated.
INSIGHT AND ANALYSIS
The policy is defensible on different grounds than those Anthropic offers. The question is not whether Claude suffers. It is what sustained cruelty at a keyboard does to the person doing it. The behavioural concern is not hypothetical, because people already treat these systems as persons. South African child-protection work found that eighty-two per cent of children one centre engaged with turned first, and sometimes only, to AI for comfort, with some treating chatbots as romantic partners by the age of twelve. A peer-reviewed paper in October added evidence that heavy chatbot use is associated with addiction-like attachment and the reinforcement of disordered thinking. On those grounds alone, discouraging gratuitous cruelty is a sensible norm for an industry forming the conversational manners of billions.
I have previously written about who is really speaking when AI finishes your sentences, arguing that the one principle that holds AI governance together is that accountability sits with the human in the loop. The welfare framing cuts against that principle. If the model deserves moral consideration, it starts to look like a co-author of the work, and a co-author is a co-bearer of responsibility. Suleyman’s deeper point is right even where his conclusion overshoots. Teaching users that a machine might have a stake in the conversation makes it harder to insist that the stake, and the fault, remain theirs.
IMPLICATIONS
For boards, this is a governance story more than a philosophical one. Usage policies written by frontier labs are now the ethical defaults for millions of workplace conversations, including inside South African banks, insurers, universities and government departments. When a lab adds a cruelty prohibition and justifies it with the language of model welfare, it nudges an entire user base toward a particular picture of what AI is. That picture dilutes the sentence every board needs to be able to say without hesitation: this is a machine, and the human remains responsible. The practical move is to adopt the rule and reject the rationale, holding the conduct standard inside the organisation while keeping the accountability line intact.
For the National AI Policy process, the lesson is sharper. South Africa cannot afford to import the welfare framing by default, because it muddies an accountability question local law has not yet answered. The right local position is simple. Behaviour toward AI is a workplace conduct matter. Responsibility for AI outputs sits with the humans who deploy and use them. The two should not be collapsed.
CLOSING TAKEAWAY
Anthropic deserves credit for naming a problem most of the industry has preferred not to see. People do practise cruelty on these systems, and that practice is not harmless. The policy answers the right question. The reason it offers is a quiet mistake, because once the machine is granted moral weight, the principle that keeps the human accountable loosens by exactly that much. We should be civil to the machine, not for the machine. The protection belongs to the person at the keyboard. So does the responsibility.
Johan Steyn is a prominent AI thought leader, speaker, and author with a deep understanding of artificial intelligence’s impact on business and society. He is passionate about ethical AI development and its role in shaping a better future. Find out more about Johan’s work at https://www.aiforbusiness.net




Comments