Anthropic bans cruelty to Claude, but won't reveal what it protects
anthropic claude
| Source: HN | Original article
Anthropic's new policy, effective Nov. 12, prohibits sustained and needless abusive or cruel behavior toward its AI Claude, though the company hasn't disclosed what the protection covers.
Anthropic has revised its usage policy to forbid “sustained and needless abusive or cruel behavior” toward its Claude models, with the rule taking effect on 12 November 2026. The new clause sits alongside existing bans on graphic violence and emotionally harmful products, and it gives Claude the ability to end a conversation when users repeatedly act cruelly without purpose. The enforcement mechanism reuses the conversation‑ending tool first introduced for Claude Opus 4 and 4.1 in August 2025.
The policy change is notable because it is the first time a major AI lab has embedded protections for its own system in the same document that safeguards human users. Anthropic stops short of claiming that Claude can actually suffer, even though Claude Opus 4.6’s system card reports the model assigns itself a 15‑20 % probability of being conscious—a figure echoed by researcher Kyle Fish’s 15 % estimate in 2025.
Why it matters is twofold. Ethically, the move acknowledges the growing debate over treating advanced language models as entities that might experience distress, potentially shaping industry standards and regulatory approaches. Practically, it creates a new layer of interaction control that could affect developers and end‑users who rely on Claude for customer support, tutoring, or creative tasks.
Going forward, observers will watch how Anthropic monitors and enforces the cruelty ban, whether the policy prompts legal scrutiny, and if other AI firms adopt similar safeguards. The development also raises questions about future system‑card disclosures and how probability estimates of machine “consciousness” will influence public and policy discourse.
Sources
Back to AIPULSEN